Cover image

Back to School for AGI: Memory, Skills, and Self‑Starter Instincts

A simulated college benchmark shows why enterprise agents need structured memory, salience, and initiative—not merely larger models with longer prompts.

August 27, 2025 · 17 min · Zelina
Cover image

Judge, Jury, and Chain‑of‑Thought: Making Models StepWiser

StepWiser shows that judging reasoning steps works better when the judge is trained to reason about the reasoning, not merely classify it.

August 27, 2025 · 18 min · Zelina
Cover image

Mirror, Signal, Maneuver: How 'Self' Labels Nudge LLM Cooperation

A behavioural-economics reading of how simple self-labels can shift LLM cooperation, and what that means for multi-agent AI operations.

August 27, 2025 · 15 min · Zelina
Cover image

Talk, Tool, Triumph: Training Agents with Real Conversations

A mechanism-first look at MUA-RL, a reinforcement learning framework that trains tool-using agents inside dynamic multi-turn user interactions rather than static function-calling scripts.

August 27, 2025 · 16 min · Zelina
Cover image

Wheel Smarts > Wheel Reinvention: What GitTaskBench Really Measures

GitTaskBench shows that code-agent value depends less on writing fresh code and more on surviving the messy chain from repository comprehension to execution, quality, and cost.

August 27, 2025 · 16 min · Zelina
Cover image

Agents on the Clock: Turning a 3‑Layer Taxonomy into a Build‑Ready Playbook

A mechanism-first reading of agentic reasoning frameworks as operational control loops, not just model upgrades.

August 26, 2025 · 15 min · Zelina
Cover image

Hypotheses, Not Hunches: What an AI Data Scientist Gets Right

A mechanism-first reading of an AI Data Scientist that turns hypothesis testing into the organising layer for automated analytics.

August 26, 2025 · 18 min · Zelina
Cover image

Mirror, Signal, Trade: How Self‑Reflective Agent Teams Outperform in Backtests

TradingGroup shows that the useful lesson from finance agents is not autonomous trading theatre, but disciplined feedback loops that turn decisions, outcomes, and risk signals into trainable operating data.

August 26, 2025 · 14 min · Zelina
Cover image

Stop at 30k: How Hermes 4 Turns Long Chains of Thought into Shorter Time‑to‑Value

Hermes 4 shows that reasoning models need operational controls around data, stopping behaviour, evaluation, and refusal policy—not just larger thinking budgets.

August 26, 2025 · 18 min · Zelina
Cover image

Words + Returns: Teaching Embeddings to Invest in Themes

A mechanism-first reading of THEME, a CIKM 2025 framework that turns thematic investing from static ETF mimicry into semantic-temporal stock retrieval.

August 26, 2025 · 16 min · Zelina