ThoughtAtlas-1K: 1,000 Long Reasoning Traces, Mapped

We release 1,000 math reasoning traces from Qwen3.8-27B, each labeled paragraph by paragraph: the path that reached the answer, the paths that were tried and dropped, and what came after the answer was settled. The labeling procedure was checked against an independent annotator, GPT-6.1 Sol, before we labeled at scale.

1,000long traces, all correct, medium and hard problems
every paragraphlabeled: main path, dropped path, or after the answer
frontier-levelopen teacher — Qwen3.8-27B, on par with Claude Opus 4.6 in OpenCode's comparison
cross-checkedagainst GPT-6.1 Sol before labeling at scale

Why this matters

Reasoning models now think for thousands of tokens before they answer, and their traces have become the main ingredient for training the next generation of models. Yet the open data that exists is plain text. You can see that a model thought for ten thousand tokens, but not which of them the answer actually rests on. ThoughtAtlas is one of the few open sets of long reasoning traces that map this out, paragraph by paragraph.

Problem. Find the greatest n such that there are n positive integers below 5000, any two sharing a divisor greater than 1 but any three coprime. (Answer: 4)

path to the answertried and droppedafter the answerchips: what each step does
One trace (deepmath-106656_5, 163 paragraphs). The trunk is what the final answer is built on; each side branch is a path the model took and abandoned; the dotted tail is re-checking after the answer was already settled.

What the function labels mean

Besides its path, every step gets one of nine labels describing what the model is doing at that moment. Together they turn a wall of text into a readable sequence of moves.

LabelWhat the model is doingTypical wording
SetupRestating the problem, naming variables, writing down what is given“We need to find…”, “Let x be…”
PlanDeciding what to do next or how to attack the problem“First I'll…, then…”, “Let's try to bound…”
RecallBringing in a known theorem, formula or fact, without new computation“By Fermat's little theorem…”
ComputeMoving the solution forward: algebra, calculation, derivationequations, case work, step-by-step arithmetic
ExploreOpening a new approach, case or guess that differs from the current line“Alternatively…”, “What if we try…”
VerifyChecking an earlier result: plugging back in, a special case, recomputing“Let me check…”, “Plugging in n = 3…”
MonitorStepping back: doubting, noticing an error, deciding to abandon or go back“Wait, that's wrong”, “This doesn't work”
ConsolidateSummarizing or combining intermediate results“So far we have…”, “Putting these together…”
AnswerStating the final answer“So the answer is…”

Read in order, the labels show the rhythm of a trace. In the example above, the first dropped branch is Explore three times followed by Monitor: three quick constructions, then the realization that none can work. After the answer, the tail is almost all Verify and Monitor — the model checking itself again and again.

What is in each trace

FieldContent
problem, gold_answerA competition-style math problem and its reference answer
thinking_paragraphsThe teacher's full thinking (Qwen3.8-27B, thinking mode), split into paragraphs
final_responseThe answer the teacher wrote after thinking (all 1,000 are correct)
nodesParagraph spans with a path — main (M), dropped (A abandoned, D dead end), post-answer (P) — a function label (setup, plan, recall, compute, explore, verify, monitor, consolidate, answer), dependencies and branch ids
exploration_shareFraction of the thinking spent on dropped paths

The traces cover medium and hard problems and short, medium and long reasoning. Labels were produced by Qwen3.8-27B in thinking mode: long traces are first outlined as a whole, then labeled in chunks.

Quality check. Before labeling at scale, we checked the labels against an independent reference: GPT-6.1 Sol annotated 100 pilot traces with the same scheme. We went through four versions of the labeling procedure and kept the one that agreed best with the reference — paragraph-level F1 of 0.73 on which steps were dropped before the answer (0.86 on short traces, about 0.67 on medium and long ones). Part of the remaining gap is genuine ambiguity: some paragraphs have no single right label. Only then were these 1,000 traces labeled with that procedure.

Related work

The labels build on earlier ways of describing reasoning: problem-solving episodes in the sense of Schoenfeld (arXiv:2509.14662, arXiv:2512.19995), sentence-level function categories in Thought Anchors (arXiv:2506.19143), reasoning graphs with failed steps (arXiv:2509.19284), and cognitive behaviors such as verification and backtracking (Gandhi et al., arXiv:2503.01307). ThoughtAtlas-1K applies a related scheme at paragraph level to long traces and releases the labeled data. It is unrelated to harsh-kumar9/thought-atlas, which provides sentence-level behavior labels from an LLM judge; ThoughtAtlas-1K labels each paragraph by its role on the path to the answer (main path, dropped path, post-answer).

Cite

Copy the entry below, or download thoughtatlas.bib. Plain text: Li, H., Hong, K., Li, X., Wu, C., Bouvry, P., and SeaFill Open-Source Team (2026). ThoughtAtlas-1K: 1,000 Long Reasoning Traces, Mapped. Hugging Face. https://huggingface.co/datasets/SeaFill2025/ThoughtAtlas-1K

@misc{li2026thoughtatlas1k,
  title        = {ThoughtAtlas-1K: 1,000 Long Reasoning Traces, Mapped},
  author       = {Li, Hongyang and Hong, Kyungmin and Li, Xiao and
                  Wu, Caesar and Bouvry, Pascal and
                  {SeaFill Open-Source Team}},
  year         = {2026},
  month        = oct,
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/datasets/SeaFill2025/ThoughtAtlas-1K}},
  note         = {Project page: \url{https://seafill.info/ThoughtAtlas/blog/}}
}

ThoughtAtlas-1K:1,000 条长推理轨迹的结构地图

我们开源 1,000 条 Qwen3.8-27B 的数学推理轨迹,每条都逐段标注:哪条路走到了答案,哪些路走了一段又放弃了,哪些是答案定下来之后的内容。标注流程在大规模使用前,先和独立标注者 GPT-6.1 Sol 做过对照检查。

1,000 条长推理轨迹,全部答对,中等题和难题
逐段标注每段标明:主路径、放弃的路,还是答案之后
前沿水平开源教师 Qwen3.8-27B,在 OpenCode 的模型对比中与 Claude Opus 4.6 相当
交叉验证大规模标注前,先用 GPT-6.1 Sol 独立标注做对照检查

为什么重要

推理模型在回答前会先想几千甚至上万个 token,这些思维链已经成为训练下一代模型的主要原料。但现有的开源数据都是纯文本:能看出模型想了一万个 token,却看不出答案真正建立在哪一部分上。ThoughtAtlas 是少有的、逐段画出这张结构地图的开源长推理轨迹数据。

题目。求最大的 n,使得存在 n 个小于 5000 的正整数,任意两个有大于 1 的公因数,但任意三个互素。(答案:4)

通向答案的路走了又放弃答案之后小标签:每一步在做什么
一条轨迹(deepmath-106656_5,163 段)。主干是最终答案所依赖的推理;每条侧枝是模型走过、后来放弃的路;虚线尾巴是答案已经确定之后的复查。

功能标签代表什么

除了路径,每一步还有一个功能标签,说明模型此刻在做什么,共九种。有了它们,一大段文字就变成了一串看得懂的推理动作。

标签模型在做什么典型说法
设定复述题目、设变量、写下已知条件“我们要求……”“设 x 为……”
计划决定接下来做什么、从哪里入手“先……再……”“试着给出上界”
引用调用已知的定理、公式或事实,不做新的计算“由费马小定理……”
计算推进解答:代数变形、计算、推导列方程、分情况、逐步运算
探索开一条和当前不同的新思路、新情况或新猜测“换个思路……”“如果试试……”
验证检查之前的结果:代回、特例、重算“检查一下……”“代入 n = 3……”
监控停下来审视:怀疑、发现错误、决定放弃或回头“等等,不对”“这条路走不通”
汇总总结或合并中间结果“目前我们得到……”“合起来看……”
答案给出最终答案“所以答案是……”

按顺序读这些标签,就能看出一条轨迹的节奏。上面例子里第一条被放弃的侧枝是 探索 三次之后接一个 监控:连试三种构造,然后意识到都不行。答案之后的尾巴几乎全是 验证 和 监控:模型一遍又一遍地自我检查。

每条轨迹包含什么

字段内容
problem、gold_answer竞赛风格的数学题及标准答案
thinking_paragraphs教师模型(Qwen3.8-27B 思考模式)的完整思考过程,按段落切分
final_response思考结束后写出的正式回答(1,000 条全部答对)
nodes按段落区间划分的节点,每个节点有路径:主路径(M)、放弃的路(A 明确放弃、D 死胡同)、答案之后(P);还有功能标签(设定、计划、引用、计算、探索、验证、监控、汇总、答案)、依赖关系和分支编号
exploration_share思考中花在放弃的路上的比例

轨迹覆盖中等题和难题,以及短、中、长不同长度的推理。标注由 Qwen3.8-27B 思考模式完成:长轨迹先整体概括,再分块逐段标注。

质量检查。在大规模标注之前,我们先用一个独立参照检查标签质量:让 GPT-6.1 Sol 按同一套体系标注 100 条试点轨迹。标注流程前后改了四版,最后选用和参照最一致的一版:在“答案之前哪些步骤被放弃”这一点上,段落级 F1 为 0.73(短轨迹 0.86,中、长轨迹约 0.67)。剩下的差距有一部分是真正的模糊地带,有些段落本来就没有唯一正确的标签。确认之后,才用这一版流程标注了这 1,000 条。

相关工作

这套标签借鉴了已有的推理刻画方式:Schoenfeld 意义上的解题情节(arXiv:2509.14662、arXiv:2512.19995),Thought Anchors 的句子功能分类(arXiv:2506.19143),带失败步骤的推理图(arXiv:2509.19284),以及验证、回溯等认知行为(Gandhi 等,arXiv:2503.01307)。ThoughtAtlas-1K 在长轨迹上按段落使用相近的体系,并公开了标注数据。它与 harsh-kumar9/thought-atlas 无关:后者由大模型给出逐句的行为标签,ThoughtAtlas-1K 则按段落标注每一步在通向答案的路径上扮演的角色(主路径、放弃的路、答案之后)。

引用

复制下面的条目,或下载 thoughtatlas.bib。纯文本格式:Li, H., Hong, K., Li, X., Wu, C., Bouvry, P., and SeaFill Open-Source Team (2026). ThoughtAtlas-1K: 1,000 Long Reasoning Traces, Mapped. Hugging Face. https://huggingface.co/datasets/SeaFill2025/ThoughtAtlas-1K

@misc{li2026thoughtatlas1k,
  title        = {ThoughtAtlas-1K: 1,000 Long Reasoning Traces, Mapped},
  author       = {Li, Hongyang and Hong, Kyungmin and Li, Xiao and
                  Wu, Caesar and Bouvry, Pascal and
                  {SeaFill Open-Source Team}},
  year         = {2026},
  month        = oct,
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/datasets/SeaFill2025/ThoughtAtlas-1K}},
  note         = {Project page: \url{https://seafill.info/ThoughtAtlas/blog/}}
}