CQ or Consequences: What This LLM Benchmark Reveals About AI Requirements Work
A comparison-based reading of CompCQ shows why LLM-generated requirements work needs model portfolios, not one-model faith.
A comparison-based reading of CompCQ shows why LLM-generated requirements work needs model portfolios, not one-model faith.
A controlled comparison of human, template, and LLM-generated competency questions shows why AI can accelerate requirements elicitation without replacing expert judgment.
A mechanism-first reading of how knowledge graphs and LLM-guided retrieval can make machine learning explanations in manufacturing more contextual, useful, and governable.
SocialGrid shows why agent reliability depends less on model eloquence than on separating navigation, execution, and behavioral inference failures.
A mechanism-first reading of MARCH, a multi-agent CT report-generation system, and what its hierarchy teaches enterprise AI about review, grounding, and controlled disagreement.
A research-sabotage benchmark shows why AI auditability is not a code-review feature, but an operating model for trustworthy AI work.
A mechanism-first reading of why explicit technique recognition may matter more than longer reasoning traces for informal theorem proving and enterprise AI workflows.
A mechanism-first reading of Blue's Data Intelligence Layer and why enterprise AI needs data planning, registries, and fewer fantasies about one-model answers.
RadAgent shows why medical AI needs auditable workflows, not just stronger black-box report generators.
A mechanism-first reading of VRUBench: why models can parse viewpoint rotations yet still fail to bind spatial state to the right observation.