A Proof Can Pass and Still Mean the Wrong Thing: What AxQM Changes About Formal AI Evaluation
TL;DR for operators When an AI system produces a formal proof, an evaluation team can make one part of the verdict unusually objective: either the proof is accepted under the allowed rules, or it is not. That removes much of the grader variance found in rubric scoring or LLM judging. AxQM provides 1,019 Lean 4 proof-synthesis tasks over 479 finite-dimensional quantum-mechanics textbook items. It keeps a private reference solution for every task and grades submissions through successful compilation, absence of sorry in the proof or its dependencies, and absence of newly introduced axioms.1 ...