Don’t Make the LLM Serialize the Whole Workflow
TL;DR for operators When an LLM produces a bad executable workflow, the failure may not mean that it misunderstood the business rule. It may have selected the right actions and still lost parameters, Boolean relationships, or graph ordering while writing the full structure. That distinction changes the engineering response. In Generating Workflow DAGs from Natural Language with Non-Reasoning LLMs, Anand Iyer, Bhanu Khetharpal, Srinivas Upadhya, and Ramkumar Rajagopal test an architecture that keeps language interpretation in the LLM but moves deterministic graph expansion into conventional code.1 On their 635-rule benchmark, GPT-5.3-chat improves from 56.4% to 80.6% LLM-judge validity when paired with registry selection, a compact intermediate representation, and deterministic compilation. Exact-condition accuracy rises from 56.5% to 82.2%. ...