Trex Marks the Spot: When AI Starts Training AI
A mechanism-first reading of TREX, an agent system that treats LLM fine-tuning as an iterative research workflow rather than a glorified hyperparameter search.
A mechanism-first reading of TREX, an agent system that treats LLM fine-tuning as an iterative research workflow rather than a glorified hyperparameter search.
GeoAgentBench shows why serious spatial AI must be tested by execution, parameter discipline, and final map verification—not by how convincingly an agent describes a workflow.
A sharper reading of AISafetyBenchExplorer, showing why AI safety evaluation now suffers less from benchmark scarcity than from metric drift, stale infrastructure, and weak benchmark governance.
BEAM shows that useful LLM algorithm design is less about clever prompting and more about structured search, reusable memory, and evaluation that actually resembles solver construction.
A close reading of Text2Model shows why LLMs can draft optimization models, but still need validation layers before they can be trusted in business decision workflows.
A mechanism-first reading of PAL, an AI learning platform that turns lecture videos into adaptive questioning, learner-state tracking, and personalized post-lesson reinforcement.
A mechanism-first reading of how bilevel optimization makes electric vehicle routing more scalable by using routing cost as a cheap but imperfect guide.
A mechanism-first reading of dual-trace memory encoding and why enterprise AI agents may need richer contextual traces, not just larger memory stores.
How Cycle-Consistent Search turns the search trajectory itself into a reward signal for training AI agents when gold answers are unavailable.
A measured reading of OIDA: why organizational AI needs memory that tracks decisions, contradictions, and open questions—not just better retrieval.