TL;DR for operators

When an LLM produces a bad executable workflow, the failure may not mean that it misunderstood the business rule. It may have selected the right actions and still lost parameters, Boolean relationships, or graph ordering while writing the full structure.

That distinction changes the engineering response. In Generating Workflow DAGs from Natural Language with Non-Reasoning LLMs, Anand Iyer, Bhanu Khetharpal, Srinivas Upadhya, and Ramkumar Rajagopal test an architecture that keeps language interpretation in the LLM but moves deterministic graph expansion into conventional code.1 On their 635-rule benchmark, GPT-5.3-chat improves from 56.4% to 80.6% LLM-judge validity when paired with registry selection, a compact intermediate representation, and deterministic compilation. Exact-condition accuracy rises from 56.5% to 82.2%.

The result is not evidence that non-reasoning models generally match reasoning models. It is a narrower finding: GPT-5.3-chat with this architecture is statistically equivalent within a pre-specified ±5 percentage-point margin to GPT-5.2 under monolithic prompting on this benchmark and metric. The best reasoning-model configuration still reaches 88.8%, 8.2 points above the best non-reasoning configuration.

For workflow products, the useful question is therefore not simply “Which model should we upgrade to?” It is “Which part of the task is the model actually failing at?”

A model can choose the right workflow pieces and still write the wrong workflow

Turning a business rule into an executable workflow requires more than identifying an event and a few actions. Conditions must contain the right values and operators. Branches must combine correctly. Parallel actions and fallback paths must have the right dependencies. The final graph must satisfy the target schema.

A monolithic LLM call asks one model to do all of this at once: interpret the administrator’s language, find the appropriate platform capabilities, construct the logical relationships, serialize every node and parameter, and produce valid workflow JSON.

The paper’s most useful diagnosis is that these responsibilities do not fail at the same rate.

For stronger models, node-set accuracy—the ability to select the correct action types while ignoring exact parameters—can approach 98–100%, while exact node and condition accuracy remain materially lower. For GPT-5.3-chat under monolithic prompting, node-set accuracy is only 82.4%, but node-exact accuracy falls further to 62.5% and exact-condition accuracy to 56.5%.

The gap becomes more revealing as workflows get denser. The authors stratify performance by the number of condition-action pairs that must be represented. Under monolithic generation, condition accuracy deteriorates as this branch-grid density rises.

They call this an emission-density bottleneck. The interpretation is not that the model has fully understood every dense rule. Rather, component selection remains substantially stronger than exact emission, suggesting that part of the observed error comes from having to serialize too many interdependent details in one pass.

That diagnosis matters because model replacement and system redesign solve different problems.

The architecture reduces what the LLM has to emit

The proposed system changes the division of labor.

Instead of asking the LLM to directly construct the full executable DAG, it asks the model for a smaller YAML intermediate representation, or IR. The IR captures the interpreted intent without forcing the model to spell out every deterministic consequence of that intent.

A compiler then expands that representation into the final workflow JSON. It handles operations such as cross-products, Boolean expansions, fallback chains, residual branches, dependency construction, and target-schema serialization.

The system also introduces registry selection. The platform has a catalog of permitted events, actions, conditions, metadata, and few-shot examples. Rather than placing that entire catalog into every generation prompt, a learned selector retrieves the relevant subset first.

The resulting division is roughly:

Responsibility LLM / selector Deterministic compiler
Interpret the natural-language rule Yes No
Identify relevant platform vocabulary Yes No
Express intended conditions and actions compactly Yes No
Expand repetitive graph structure No Yes
Construct fallback and dependency structure Limited Yes
Serialize the target workflow schema No Yes

This is more than a formatting convenience. It removes classes of work for which probabilistic generation provides little benefit.

Compiler-based configurations consequently produce about 99–100% valid JSON. IR-only generation has zero uncompilable outputs across the four tested models, and across 5,080 selector-pipeline generations only one malformed IR remains an unrepaired hard failure.

But syntactic validity is not workflow correctness. A compiler can faithfully expand a wrong logical interpretation.

GPT-5.3-chat gains 24.3 points, but the gain is model-dependent

The largest improvement appears for GPT-5.3-chat.

Under monolithic prompting, its LLM-judge validity is 56.4%. With Selector+IR, that rises to 80.6%, a paired improvement of 24.3 percentage points. Exact-condition accuracy increases from 56.5% to 82.2%, while node-exact accuracy rises from 62.5% to 82.8%.

The paired validity gain is statistically significant under the paper’s McNemar test.

The architecture does not produce gains of that magnitude for every model. GPT-4.1, for example, does not improve under Selector+IR relative to its monolithic validity score. GPT-5.2 and Claude Opus 4.6 gain more modestly. This is consistent with the paper’s proposed mechanism: externalizing emission helps most where dense structured emission creates substantial headroom for improvement.

The authors also test a more specific claim than simple score proximity. GPT-5.3-chat with Selector+IR reaches 80.6% judge validity versus 79.8% for GPT-5.2 under monolithic prompting. Using a paired two-one-sided equivalence test with a pre-specified ±5-point margin, they establish statistical equivalence for that comparison; the 90% confidence interval on the paired difference is -2.0 to +3.6 points.

That conclusion should remain scoped to the tested configurations, benchmark, tolerance, and validity metric. Equivalence is not established against Opus 4.6 monolithic at the same ±5-point margin, and the strongest reasoning-model configuration reaches 88.8%.

Architecture narrows the tested gap. It does not make reasoning-model capability irrelevant.

Two LLM calls can still use fewer tokens than one

Multi-stage systems are often assumed to add inference cost because they make additional model calls. Here, the representation and context shrink enough to reverse that arithmetic.

Across all four evaluated models, Selector+IR consumes about 55–56% as many total tokens per rule as monolithic prompting. For GPT-5.3-chat, reported consumption falls from 21,264 tokens per rule to 11,980.

Selection alone does not produce this result. The Selector→JSON configuration is actually more expensive than monolithic generation across the tested models because it pays for a selection call while still generating a large final representation.

The cost reduction comes from combining two changes: restrict the registry context, then generate a substantially smaller representation.

For products with growing capability catalogs, Cognaptus infers a practical design criterion from this result: decomposition should be evaluated by the information volume passed between stages, not by call count alone. A two-stage pipeline can be cheaper when the first stage sharply reduces what the second stage must read and write.

What the compiler still cannot repair

The residual errors define the boundary of the approach.

The pipeline substantially reduces incomplete-rule errors, De Morgan or exclusion mistakes, and raw JSON-structure failures. But cross-variable AND/OR mistakes remain approximately flat, and DAG ordering is the largest residual error category, accounting for 49.5% of coded failures. The recurring-timer family still has roughly 40% exact-condition failure.

These are different from deterministic expansion errors. If the LLM misinterprets which conditions belong together or which action must precede another, a compiler cannot infer the intended semantics from an already-wrong IR.

This suggests a useful diagnostic sequence for workflow systems:

  1. Measure whether the model is choosing the correct event and action vocabulary.
  2. Separately measure exact parameters, Boolean structure, and graph dependencies.
  3. If component selection is strong but exact structure deteriorates with density, reduce the structure the LLM must emit.
  4. If logical relationships and topology remain wrong after structural externalization, improve interpretation, representation, verification, or model capability rather than adding more deterministic expansion.

That sequence is a Cognaptus inference from the paper’s evidence, not an experimentally validated deployment procedure.

The evidence supports architecture decisions, not a universal model substitution

The evaluation is unusually structured for a system-design paper: all configurations run on the same 635 rules, generation uses temperature-0 greedy decoding, differences are tested with paired statistics, equivalence claims use a pre-specified margin, and judge-based validity is supplemented with deterministic metrics. A blind 100-item author-labeling check agrees with the LLM judge on 91% of examples, with combined Cohen’s kappa of 0.82.

The main external-validity constraint is the benchmark itself. It is proprietary, manufactured synthetic data generated from controlled templates for a Microsoft contact-center routing setting. The prompts include varied values, branch structures, spelling errors, and some non-English entity values, but the study does not establish robustness to unrestricted administrator language.

There is also only one greedy generation for each model/configuration/rule combination. The results therefore do not quantify sensitivity to decoding randomness, prompt paraphrases, or example ordering. The authors propose transfer to other graph-generation settings such as business processes, ETL, agent-tool graphs, and CI systems, but those domains are not evaluated here.

So the practical contribution is narrower and more actionable than “smaller models can replace reasoning models.”

When a workflow generator is failing, first determine whether the model is misunderstanding the requested workflow or struggling to serialize what it has already identified. If emission density is the limiting factor, moving deterministic graph construction into code can improve validity, reduce token consumption, and postpone some model-upgrade pressure. If the remaining failures concern Boolean meaning or topology, the compiler has reached the boundary of what structural externalization can solve.

Cognaptus: Automate the Present, Incubate the Future.


  1. Anand Iyer and Bhanu Khetharpal and Srinivas Upadhya and Ramkumar Rajagopal (2026). Generating Workflow DAGs from Natural Language with Non-Reasoning LLMs. arXiv:2608.30250. https://arxiv.org/abs/2608.30250 ↩︎