ThoughtAtlas-1K: 1,000 Long Reasoning Traces, Mapped
We release 1,000 math reasoning traces from Qwen3.8-27B, each labeled paragraph by paragraph: the path that reached the answer, the paths that were tried and dropped, and what came after the answer was settled. The labeling procedure was checked against an independent annotator, GPT-6.1 Sol, before we labeled at scale.
Why this matters
Reasoning models now think for thousands of tokens before they answer, and their traces have become the main ingredient for training the next generation of models. Yet the open data that exists is plain text. You can see that a model thought for ten thousand tokens, but not which of them the answer actually rests on. ThoughtAtlas is one of the few open sets of long reasoning traces that map this out, paragraph by paragraph.
- Structure-aware training data. With every paragraph labeled, reasoning traces can be filtered, cleaned and organized by how they reason, not only by whether the final answer is right.
- Step-level supervision for free. Every paragraph says whether the answer depends on it. That is exactly the signal process reward models and step-level verifiers are trained on, and it is usually expensive to collect.
- A lens on how strong models think. Even in this small release the structure is striking: about half of the traces go straight to the answer, most newly opened approaches are later dropped, and well over a third of the thinking happens after the answer is already settled.
- A shared reference for reasoning research. The same labeled traces can be used to study reasoning behaviors such as exploration, backtracking and verification, and to benchmark automatic annotators against a common set.
- Open and frontier-quality. The strongest reasoning traces usually come from closed models that hide or summarize their thinking. These come from an open model with the full thinking kept: in OpenCode's side-by-side comparison, Qwen3.8-27B and Claude Opus 4.6 score about the same on coding (75 vs 73 out of 100).
Problem. Find the greatest n such that there are n positive integers below 5000, any two sharing a divisor greater than 1 but any three coprime. (Answer: 4)
What the function labels mean
Besides its path, every step gets one of nine labels describing what the model is doing at that moment. Together they turn a wall of text into a readable sequence of moves.
| Label | What the model is doing | Typical wording |
|---|---|---|
| Setup | Restating the problem, naming variables, writing down what is given | “We need to find…”, “Let x be…” |
| Plan | Deciding what to do next or how to attack the problem | “First I'll…, then…”, “Let's try to bound…” |
| Recall | Bringing in a known theorem, formula or fact, without new computation | “By Fermat's little theorem…” |
| Compute | Moving the solution forward: algebra, calculation, derivation | equations, case work, step-by-step arithmetic |
| Explore | Opening a new approach, case or guess that differs from the current line | “Alternatively…”, “What if we try…” |
| Verify | Checking an earlier result: plugging back in, a special case, recomputing | “Let me check…”, “Plugging in n = 3…” |
| Monitor | Stepping back: doubting, noticing an error, deciding to abandon or go back | “Wait, that's wrong”, “This doesn't work” |
| Consolidate | Summarizing or combining intermediate results | “So far we have…”, “Putting these together…” |
| Answer | Stating the final answer | “So the answer is…” |
Read in order, the labels show the rhythm of a trace. In the example above, the first dropped branch is Explore three times followed by Monitor: three quick constructions, then the realization that none can work. After the answer, the tail is almost all Verify and Monitor — the model checking itself again and again.
What is in each trace
| Field | Content |
|---|---|
problem, gold_answer | A competition-style math problem and its reference answer |
thinking_paragraphs | The teacher's full thinking (Qwen3.8-27B, thinking mode), split into paragraphs |
final_response | The answer the teacher wrote after thinking (all 1,000 are correct) |
nodes | Paragraph spans with a path — main (M), dropped (A abandoned, D dead end), post-answer (P) — a function label (setup, plan, recall, compute, explore, verify, monitor, consolidate, answer), dependencies and branch ids |
exploration_share | Fraction of the thinking spent on dropped paths |
The traces cover medium and hard problems and short, medium and long reasoning. Labels were produced by Qwen3.8-27B in thinking mode: long traces are first outlined as a whole, then labeled in chunks.
Quality check. Before labeling at scale, we checked the labels against an independent reference: GPT-6.1 Sol annotated 100 pilot traces with the same scheme. We went through four versions of the labeling procedure and kept the one that agreed best with the reference — paragraph-level F1 of 0.73 on which steps were dropped before the answer (0.86 on short traces, about 0.67 on medium and long ones). Part of the remaining gap is genuine ambiguity: some paragraphs have no single right label. Only then were these 1,000 traces labeled with that procedure.
Related work
The labels build on earlier ways of describing reasoning: problem-solving episodes in the sense of Schoenfeld (arXiv:2509.14662, arXiv:2512.19995), sentence-level function categories in Thought Anchors (arXiv:2506.19143), reasoning graphs with failed steps (arXiv:2509.19284), and cognitive behaviors such as verification and backtracking (Gandhi et al., arXiv:2503.01307). ThoughtAtlas-1K applies a related scheme at paragraph level to long traces and releases the labeled data. It is unrelated to harsh-kumar9/thought-atlas, which provides sentence-level behavior labels from an LLM judge; ThoughtAtlas-1K labels each paragraph by its role on the path to the answer (main path, dropped path, post-answer).
Cite
Copy the entry below, or download thoughtatlas.bib. Plain text: Li, H., Hong, K., Li, X., Wu, C., Bouvry, P., and SeaFill Open-Source Team (2026). ThoughtAtlas-1K: 1,000 Long Reasoning Traces, Mapped. Hugging Face. https://huggingface.co/datasets/SeaFill2025/ThoughtAtlas-1K
@misc{li2026thoughtatlas1k,
title = {ThoughtAtlas-1K: 1,000 Long Reasoning Traces, Mapped},
author = {Li, Hongyang and Hong, Kyungmin and Li, Xiao and
Wu, Caesar and Bouvry, Pascal and
{SeaFill Open-Source Team}},
year = {2026},
month = oct,
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/datasets/SeaFill2025/ThoughtAtlas-1K}},
note = {Project page: \url{https://seafill.info/ThoughtAtlas/blog/}}
}