The Checker Gets the Final Say: What P-99 Reveals About Verified AI Coding
TL;DR for operators The strongest result in this case study is not that an AI system wrote working Prolog. Across 33 targeted exercises, Claude generated 508 runtime tests alongside 257 lemmas and roughly 11,800 lines of proof, with the authors manually inspecting the generated files and an independent theorem prover checking the proofs. ...