TL;DR for operators
Organizations already possess large amounts of procedural knowledge: internal runbooks, tool instructions, operating procedures, and reusable agent skills. The difficult part is turning that knowledge into training tasks with the correct tools, files, constraints, hidden information, and evaluation logic.
Skill2Env: Capability-Oriented Environment Synthesis from Skills for General Agents1 treats that as an environment-construction problem rather than an instruction-generation problem. Its pipeline produced 2,963 executable tasks by organizing recurring agent demands into five capability dimensions and translating them into reusable difficulty patterns.
The system also tests whether a task remains difficult for the current solver. Across 500 matched task identities, two rounds of hardening reduced DeepSeek-V4-Flash’s full-credit rate from 48.40% to 15.40%, while mean assistant turns increased from 25.91 to 38.85.
The generated data transferred beyond those synthesized environments. Supervised fine-tuning Qwen3.6-35B-A3B on 1.5K trajectories scoring above 0.9 increased its unweighted average across seven downstream agent benchmarks from 36.6 to 45.0.
For teams developing tool-using agents, the relevant design choice is therefore broader than whether to generate more synthetic prompts. Skill2Env suggests that reusable operational knowledge can become post-training infrastructure when task requirements, executable state, difficulty, and evaluation are generated as one coordinated system. The study evaluates that system mainly as a combined pipeline, so it does not establish the independent contribution of every component.
More task descriptions do not solve the environment problem
Suppose a company has a reliable procedure for reconciling records across several systems. Turning that procedure into agent-training data requires much more than rewriting it as an instruction.
The agent needs access to the correct tools. The initial workspace must contain the right files and state. Relevant information may need to be distributed across sources rather than placed directly in the prompt. Some facts should remain implicit until the agent discovers them through execution. The evaluator must then check the result that actually matters rather than reward plausible-looking text.
A skill can describe how work should be performed while leaving all of those choices unresolved. The resulting task may expose its own answer, omit a required dependency, allow an unintended shortcut, or judge an outcome different from the one the instruction requested.
Skill2Env addresses this coordination problem explicitly. It defines an executable task as an instruction plus an environment containing an execution substrate, a skill, an initial workspace, and an evaluator. The important design choice is that these pieces are not generated independently.
The blueprint keeps the task contract internally consistent
The framework starts from five recurring capability demands: environment understanding, planning, skill usage, long-horizon consistency, and error recovery. These are not benchmark categories. They are design dimensions for constructing situations in which an agent has to demonstrate particular kinds of competence.
Skill2Env maintains an initial library of 100 reusable difficulty patterns. A pattern converts a capability demand into a concrete obstacle: the agent may need to discriminate among similar targets, resolve conflicting sources, bind parameters correctly, maintain agreement across intermediate outputs, or recover after an execution failure.
Compatible patterns are then instantiated through an author-side task blueprint. The blueprint records the objective, intended challenges, environment facts, execution requirements, information boundaries, and acceptance criteria.
That shared specification matters because it drives the instruction, execution substrate, initial workspace, and rubric-based evaluator together. Difficulty is therefore not inserted only into the wording of the task. It can be expressed in the files the agent receives, the tools it must use, where evidence is located, what information is withheld, and what final state receives credit.
The resulting 2,963 tasks show considerable implementation variety: they span 25 source-skill domains, 4,242 distinct dependency packages, 823 command-line interface names, and workspaces containing 46,295 files across 113 recognized extensions. That diversity does not by itself establish training quality, but it does show that the synthesis process is not confined to one repeated terminal-task template.
Hardening makes difficulty relative to the current solver
A task designer can intend a challenge that the model simply bypasses.
Skill2Env therefore executes generated tasks before treating their difficulty as settled. Solver trajectories and item-level rubric results are inspected to identify challenges that were solved too easily, avoided through shortcuts, or otherwise failed to impose the intended demand. The blueprint is then revised and affected environment components are reconstructed.
The paper calls this Iterative Task Hardening.
Its clearest evidence comes from 500 task identities observed across three versions. DeepSeek-V4-Flash received full credit on 48.40% of iteration-0 tasks, 24.80% after one hardening round, and 15.40% after two. Mean assistant turns increased from 25.91 to 36.54 and then 38.85.
Because the same task identities are followed across iterations, this analysis is stronger than simply comparing unrelated samples of easy and difficult tasks. It supports a specific conclusion: execution-guided revision systematically made those tasks more demanding for the fixed solver.
Longer trajectories are not automatically better training examples. Here, the turn increase is informative because it appears alongside a substantial decline in full-credit rates under controlled task identities.
The downstream result supports the pipeline, not every component
The main transfer experiment uses 1.5K teacher trajectories generated with DeepSeek-V4-Flash and retained only when rubric reward exceeded 0.9. Qwen3.6-35B-A3B is then supervised-fine-tuned for five epochs.
Its seven-benchmark average rises from 36.6 to 45.0.
| Evidence | Likely role | Result | Interpretation boundary |
|---|---|---|---|
| Seven downstream benchmarks | Main transfer evidence | Average 36.6 → 45.0 | Tests the combined Skill2Env training pipeline |
| SkillsBench with and without automatic skill loading | Mechanism-oriented robustness check | +14.3 points with skills; +10.4 without automatic loading | Gains are not limited to automatically retrieving supplied skills |
| Terminal-Bench across four harnesses | Harness robustness | Gains of 11.2–16.9 points | Improvement is not specific to one evaluated execution harness |
| 500 trajectories from each hardening stage | Supporting training-utility analysis | Average 40.58 → 42.51 → 42.87 | Samples are random qualifying trajectories, not matched examples |
Every one of the seven main benchmarks improves after training. The largest reported gains are on SkillsBench and Terminal-Bench 2.1, but improvements also appear on repository-level software engineering, automation, banking, autonomous task execution, and interactive service benchmarks.
The fixed-budget hardening comparison adds a smaller but relevant result. Holding the training set at 500 trajectories with reward above 0.9, later hardening stages produce progressively higher downstream averages. The iteration-2 sample reaches 42.87 versus 40.58 for iteration 0.
That supports the idea that harder environments can yield somewhat more useful supervision. It is weaker evidence than the paired difficulty analysis because the qualifying trajectories differ across stages and are randomly sampled rather than matched.
For agent teams, environment generation becomes reusable infrastructure
The business inference is most relevant to organizations that already possess structured operational knowledge and are deciding how to expand post-training or evaluation without manually authoring every executable task.
Skill2Env suggests an architecture in which a procedure or reusable skill supplies domain knowledge, while a separate library supplies recurring capability challenges. A blueprint combines the two and keeps instructions, workspace state, runtime requirements, information boundaries, and evaluation synchronized.
That separation could allow the same difficulty infrastructure to be reused across different operational domains while preserving domain-specific tools and procedures.
The hardening loop adds another property: the curriculum can change as the solver changes. If a task that once tested planning or recovery becomes routine, execution evidence can trigger reconstruction rather than leaving the training set static.
For teams maintaining increasingly capable agents, this is potentially more valuable than simply increasing synthetic-data volume. It offers a way to refresh what the environment demands.
The unresolved question is attribution
The paper provides strong comparative evidence that the complete system can generate challenging environments and that supervision from those environments improves the evaluated model.
It does not separately identify which parts of the architecture are responsible for how much of the downstream gain.
The five capability dimensions, difficulty-pattern library, blueprint construction, environment-validation process, trajectory filtering, and hardening loop operate as parts of one pipeline. The experiments therefore do not establish that any individual pattern independently improves a corresponding benchmark capability.
Nor does the resulting 35B-A3B checkpoint become the strongest model in the comparison. Several frontier-model references remain substantially ahead on multiple benchmarks.
The appropriate operational conclusion is narrower. Skill2Env provides evidence that scalable agent training data can be built by treating task generation as executable-system construction and adapting difficulty using observed solver behavior. It gives organizations a concrete architecture to test when manual environment authoring has become the bottleneck.
That is a materially different scaling problem from generating more instructions.
Cognaptus: Automate the Present, Incubate the Future.
-
Weiyi Xu and Xiaowen Yang and Wen Da and Hang Xu and Canwei Li and Hongjie You and Pusen Dong and Yucheng Zeng and Zhaokai Luo and Mu Chuan (2026). Skill2Env: Capability-Oriented Environment Synthesis from Skills for General Agents. arXiv:2609.33772. https://arxiv.org/abs/2609.33772 ↩︎