Cover image

Flag First, Explain Later: Where AI Financial Audits Still Need Human Judgment

TL;DR for operators An audit team has more financial statements than it can investigate deeply, so the first operational question is which files deserve attention and what an auditor should inspect once one is flagged. This system is much stronger at the first task than the second. Its strongest reported regression configuration, SVR M17, has an average of 0.83 among the twenty highest-ranked statements when measured against the paper’s proxy-positive criterion. The best explanation method, EiForest, reaches only 0.24 average F1 in the seven-company explanation evaluation. ...

August 13, 2026 · 7 min · Zelina
Cover image

Delta Force: How Weak Models are Secretly the Best Teachers

TL;DR for operators Training budget is usually where elegant AI strategy goes to die. The paper behind this article argues that preference tuning does not always need a superior teacher response. It may only need a useful contrast. A model can improve by learning that one weak answer is better than an even weaker one, even when neither answer is as good as what the model can already produce.1 ...

July 9, 2025 · 17 min · Zelina