TL;DR for operators

A retrieval agent does not merely retrieve documents. It repeatedly decides what to do next: search again, rewrite a query, record evidence, combine what it has found, verify a claim, or stop.

The Fellowship of the Query: Learning Retrieval Actions1 shows that these decisions can be taught to small language models as a supervised classification problem. Across the evaluated models, LoRA fine-tuning sharply improves prediction of the next retrieval action. For Granite 4.1 3B, macro-F1 rises from 0.1736 zero-shot to 0.6536 after fine-tuning.

The more revealing result appears downstream. Holding the answer generator fixed, replacing the base controller with the fine-tuned controller raises the paper’s automatic evidence-fact rate by roughly 0.60–0.62. Yet the corresponding controller-only improvements in final-answer quality have bootstrap intervals that cross zero.

For teams building retrieval agents, this distinction matters. The fine-tuning procedure clearly teaches a workflow. It does not establish that the learned workflow is an optimal search policy, nor that better workflow execution alone reliably produces better answers. The architecture suggested by the evidence is therefore modular: specialize a compact controller for a stable retrieval protocol, measure its intermediate behavior directly, and avoid assuming that the same specialized model should also remain the general-purpose generator.

Better retrieval behavior does not automatically mean better answers

Suppose a retrieval system reaches the right final answer often enough, but its internal search process is inconsistent. Sometimes it extracts useful evidence; sometimes it retrieves relevant material but never records it; sometimes it continues searching when it already has enough information.

A team trying to improve that system faces a measurement problem. Final-answer accuracy compresses the entire workflow into one number. It cannot tell you whether a change improved retrieval decisions, evidence handling, answer generation, or several of them simultaneously.

The paper separates these roles experimentally. It alternates base and fine-tuned versions of the model as the controller, which chooses search actions, and as the generator, which writes the final answer.

That separation exposes a striking behavioral change:

Controller Generator Token F1 Evidence-fact rate
Base Base 0.7783 0.0121
Fine-tuned Base 0.8044 0.6161
Base Fine-tuned 0.7997 0.0134
Fine-tuned Fine-tuned 0.8295 0.6389

Changing only the controller transforms evidence recording. Changing only the controller does not produce an equally clear statistical result for final-answer quality.

The evidence-fact metric itself is limited: it uses substring or token-overlap matching against retrieved evidence rather than human factuality or entailment judgments. Even so, the role-swap experiment makes the behavioral location of the fine-tuning effect unusually visible. The controller learns to execute the taught evidence-handling workflow far more consistently.

Search control becomes a supervised next-action problem

The mechanism is simpler than many agent-training schemes.

The authors generate multi-step search trajectories with a larger teacher model, filter those trajectories, and convert every accepted step into a training example. At step $t$, the student sees the original question $q$, the current structured search state $s_t$, and the available evidence context $e_t$, then predicts the teacher’s next action $a_t$:

$$ (q, s_t, e_t) \mapsto a_t $$

The action space contains seven choices: decompose, search, reformulate, extract, synthesize, verify, and finish.

This reframes retrieval orchestration. Instead of expecting a small model to infer an entire search protocol from prompting, the system directly supervises the local decision it should make at each state.

On 1,646 held-out action examples, Granite 4.1 3B with LoRA supervised fine-tuning reaches 0.6536 macro-F1 and 0.7581 accuracy. The same model used zero-shot reaches only 0.1736 macro-F1. A TF-IDF logistic-regression baseline reaches 0.5399, so the task is not reducible to random guessing or output-format compliance, but neither is it purely dependent on sophisticated generative reasoning.

The effect also appears across the other evaluated small-model backbones: LoRA-SFT macro-F1 ranges from 0.5683 for Granite 4.0 350M to 0.6443 for Ministral 3B.

One qualification is essential. These targets come from accepted teacher trajectories. Predicting the recorded teacher action more accurately shows better imitation of that search protocol; it does not prove that the teacher action was uniquely optimal. Another sequence of actions could plausibly reach the same correct answer.

A few thousand actions capture much of the observed specialization

The training-size experiment changes the deployment question.

For Granite 4.1 3B, 3,000 supervised action examples produce macro-F1 0.6489, already close to the best reported value of 0.6625 at 12,000 examples and the 0.6536 result using all 13,194 training actions.

This is a sensitivity result rather than evidence that 3,000 examples will generally suffice for retrieval agents. The dataset contains a fixed seven-action protocol derived from acceptance-filtered teacher behavior. Different action spaces, retrieval environments, or domain rules could require substantially more supervision.

Still, for an operator with a stable workflow, the result changes where experimentation can begin. The first question need not be whether millions of labeled trajectories can be obtained. A more immediate test is whether a few thousand high-quality state-action examples are enough to make a compact controller behave consistently.

The paper also compares one seven-class controller against seven independent one-vs-rest controllers using Qwen 3.5 0.8B. The single multiclass model reaches 0.6021 macro-F1, compared with 0.5680 for the independent binary models. A naive ensemble of the binary scores falls to 0.3978. In this setting, retrieval actions behave more coherently as competing choices within one decision than as unrelated binary judgments.

The deployment case is modular specialization

The paper directly shows that fine-tuning improves action prediction and materially changes evidence-recording behavior in the evaluated pipeline. It also reports that using the fine-tuned model for both controller and generator raises token F1 from 0.7783 to 0.8295, with a paired-bootstrap interval of [0.0021, 0.1040] for that combined difference.

Cognaptus draws a narrower operational inference: workflow specialization should be treated as a component capability, not automatically as a whole-model upgrade.

That interpretation is reinforced by the capability-retention results. After trajectory fine-tuning, Granite 4.1 3B falls from 0.6554 to 0.5520 on MMLU, 0.7911 to 0.6852 on HellaSwag, and 0.6502 to 0.5589 on ARC-Challenge.

For a production team, the affected decision is therefore architectural. If the retrieval protocol is stable and the controller’s responsibilities are narrow, an adapter-specialized small model may be attractive for orchestration. The unmodified base model—or another model entirely—can remain responsible for tasks requiring broader capability. The paper does not establish the cost, latency, or reliability advantage of such a deployment, so those remain engineering questions rather than demonstrated outcomes.

The benchmark is narrower than an open retrieval agent

The end-to-end evaluation uses 149 held-out accepted trajectories and deterministic BM25 retrieval over small per-question candidate corpora. Questions have between 5 and 20 candidate documents, with a mean and median of 10. The evaluation is therefore not an open-web search test or a production-scale retrieval benchmark.

The population is also filtered before training and evaluation: 1,490 of 5,000 attempted teacher trajectories survive the paper’s acceptance criteria. Results describe performance on that accepted population, not on arbitrary queries where the teacher itself may fail to construct a satisfactory search path.

Rare actions create another boundary. The held-out set contains only 17 examples of verify, making conclusions about that action unstable.

These constraints do not erase the central result. They locate it. The study provides credible evidence that small language models can learn a structured retrieval-control protocol from supervised trajectories, and that this specialization can substantially alter the intermediate state supplied to answer generation.

What remains uncertain is how well that controller transfers when the document universe expands, teacher trajectories are imperfect, valid search paths multiply, or the workflow changes.

What operators should measure next

The most useful lesson is methodological as much as architectural.

If a team fine-tunes an agent controller, final-answer accuracy alone is too coarse a diagnostic. Evaluation should preserve the separation between action selection, intermediate-state quality, and final generation. Otherwise, a generator improvement can hide a weak controller, while a better controller can look ineffective because the final-answer metric barely moves.

This paper shows why that decomposition matters. The clearest fine-tuning effect is not simply that the system answers more questions correctly. It is that a compact model becomes much more consistent at executing a specified retrieval workflow.

That is a narrower capability than general reasoning, but for a well-defined agent role, narrowness may be exactly what makes it measurable and deployable.

Cognaptus: Automate the Present, Incubate the Future.


  1. Mohammed Al-Maamari and Saber Zerhoudi and Michael Granitzer and Jelena Mitrović (2026). The Fellowship of the Query: Learning Retrieval Actions. arXiv:2609.28653. https://arxiv.org/abs/2609.28653 ↩︎