1
Extra context only helps if the model knows what to do with it
“The capacity of the human mind for formulating and solving complex problems is very small compared with the size of the problems …”
Test-time scaling improves a model by spending more compute at inference. Recent systems run thousands of agents, each in its own context, and a harness decides how those contexts are allocated and what is carried between them. That splits reasoning into two different skills.
Mathematical reasoning
How to make progress with the information available inside one context.
Contextual reasoning
How to allocate, coordinate and reuse separate contexts as reasoning unfolds: which fresh context should explore an alternative, verify earlier work or continue a line of reasoning, and what to carry into it.
Most harnesses make those decisions for the model. Reasoning Cache always compresses and continues; recursive agents always decompose. But the best use of a new context changes as a solution develops. A hard proof may first need several approaches in parallel, then a check of one promising intermediate result, then a clean continuation.
| Method | Model-driven strategy selection | Context compaction | Recursive delegation |
|---|
The contextual reasoning gap
We gave three models the same Hermes harness: up to seven fresh 8K-token contexts instead of one. Claude Sonnet 5 and DeepSeek-V3.2 turn the extra contexts into accuracy. Qwen3-4B, the small open model we train, loses more than half its accuracy. It solves 37% of the problems in a single context, so what it lacks is not mathematics but knowing what to do with the extra contexts.
The same harness helps capable models and hurts a small one
2
Hermes: the model decides how to use each fresh context
Hermes is built to be simple (two generic mechanisms), flexible (neither mechanism prescribes a form of computation) and configurable (how much the harness steers is a parameter).
parallel
Delegation
The agent writes TASK: lines, i.e. calls launch_subagent(task). Each task runs on the
same model in a fresh context, sees only the task text, and returns a compact report
(Scope · Finding · Answer) as the agent's next input. Subagents can delegate too,
up to depth D.
sequential
Digestion
At the end of a window the agent compresses its progress into a digest. If reasoning continues, a fresh window starts from the original problem plus the digest instead of the full history, for up to W windows.
Inside a window the model takes up to T steps. At every step but the last it may launch up to B subagents; at the last it must answer. Together with the delegation depth D, the number of windows W and the steering level, these knobs set both the compute budget and how much of each decision belongs to the model.
Configurations: breadth, depth and length
B and D widen the tree in parallel; W extends it in sequence. Every box below is one 8K-token context. Change a knob to see how the tree, and the compute it can use, grows.
7
context windows, at most
56K
tokens of context, at most
Steering: who picks the strategy?
Each delegation step uses one or more strategies, in two categories:
- Search makes progress. Solve the whole problem again with no set approach (Same), with different named approaches (Different), or split it into subproblems (Split).
- Verify checks the current answer: directly (Answer), through the steps behind it (Steps), or by finding why candidates disagree (Disagreement).
The steering level decides who picks. At L3 the harness fixes the strategy, as most existing harnesses do. At L2 it fixes only the category. At L1 the model picks both.
Watch it reason
Six rollouts of the trained Qwen3-4B under different harnesses and configurations. Each one plays on its own; use Back and Next to go at your own pace. The takeaway appears at the end.
3
Hermes-Learn: teaching a 4B model contextual reasoning
Hermes-Learn teaches the model what to delegate, in two stages: first imitate strong models, then learn from its own outcomes. Digestion is not trained; the model uses its existing ability to summarise.
Stage I
Learn from teachers
- Two strong models, Claude Sonnet 5 and DeepSeek-V3.2, solve 1,000 problems under Hermes, 8 runs per problem.
- Keep the runs that delegate at least once and reach the correct answer, then balance search strategies and teachers.
- Fine-tune the 4B model on the root agent's own words.
What it learns from, in one run
trained: the root agent's own wordscontext only
Stage II
Learn from outcomes
- The fine-tuned model solves each problem 8 times under Hermes.
- A run scores 1 if its final answer is correct and 0 if not.
- Runs that beat their group's average become more likely, the rest less. Subagents are not trained.
One group: 8 runs of the same problem · click a run to flip its outcome
A curriculum: from guided to free
Stage II first trains only under L2, where the harness fixes each step's category. After the first pass over the data, three quarters of training moves to L1, where the model picks everything, and a quarter stays on L2.
Share of stage II training under each steering level
Synthetic problems that reward different strategies
Competition problems rarely require a particular strategy. So we build 500 problems from seeds with known answers, each type favouring a different strategy, and reword them so the structure is not signposted.
RQ1
Extra context goes from liability to asset
On every benchmark, the untrained model does worse with Hermes than in one context, and Hermes-Learn turns that around. Its single-context accuracy barely moves, so the gain comes from using contexts, not from better mathematics. After training, Hermes beats one 56K-token context on all four math benchmarks.
RQ2
It learns strategies that depend on the problem and on progress
An LLM judge that sees only the tasks, never the harness or the strategy menu, labels the strategy of every delegation step.
RQ3
What matters in training: teachers, curriculum, and new problem structure
(a) Which teacher for SFT?
The stronger Hermes user is not the better teacher. DeepSeek-V3.2 trains a far better student than Sonnet 5 (39.3 vs 30.1). Using both is best, mostly because it gives more demonstrations.
(b) Which RL recipe?
New problem structure beats more data. The curriculum plus 500 synthetic problems scores best (47.3); 500 rephrased problems give 45.6. The recipes are within about 2 sd of each other.
RQ4
It generalises beyond its training setting
Training used one setting: Qwen3-4B with B = 3, D = 1, W = 1. We test three ways of leaving it.
(a) More inference compute than in training
More windows, more subagents and deeper delegation all help, although training used none of them. Depth 2, where subagents delegate too, reaches 55.6.
(b) Other test-time scaling methods
Hermes lifts other methods. The DeepSeekMath Agent goes from 15.1 to 52.5 when run inside Hermes; Recursive Self-Aggregation from 50.5 to 55.0.
(c) A different model family and size
The same recipe works on another model. Olmo-3-7B drops from 30.5 to 4.8 with Hermes; Hermes-Learn takes it to 40.6.
Citation
@article{li2026hermes,
title = {Hermes: Learning Contextual Reasoning Unlocks Test-Time Scaling},
author = {Li, Xinyu and Goswami, Mononito and Liu, Hao and Kanakaris, Nikos and
Huang, Langlin and Jana, Prithwish and Bl{\"o}baum, Patrick and Jain, Purak},
journal = {arXiv preprint},
year = {2026}
}