When Small Models Learn From Their Mistakes: Arithmetic Reasoning Without Fine-Tuning
A mechanism-first reading of how error clustering, code generation, and selective prompt rules can make small on-premise models more reliable for tabular arithmetic.