Cover image

Keys to the Kingdom… with a Chaperone: How Agentic JWT Grounds AI Agents in Real Intent

Access tokens are convenient little monsters. Hand one to an application and, for a while, the receiving API behaves as if the bearer of that token is a faithful representative of the user. In normal software, that assumption is often good enough. The app has deterministic code. The button does what the button was built to do. The workflow may be dull, but dullness is a security feature. ...

October 1, 2025 · 16 min · Zelina
Cover image

Terms of Engagement: Building Trustworthy AI Agents Before They Build Us

A customer asks your AI assistant to “find me a better phone contract.” The agent browses comparison sites, selects a cheaper plan, authorizes the switch, cancels the old plan, and arranges payment of the cancellation fee from the user’s bank account. Lovely, in the way a self-driving forklift is lovely: impressive until it nudges the wrong shelf. ...

September 19, 2025 · 15 min · Zelina
Cover image

Tool Wars, Protocol Peace: What MCP‑AgentBench Really Measures

A procurement team does not buy an AI agent because it can recite the word “interoperability” with theatrical confidence. It buys the agent because the thing can use tools, collect data, combine results, and stop before it bankrupts the token budget. That is the useful way to read MCP-AgentBench, a new benchmark for evaluating language agents inside the Model Context Protocol ecosystem.1 The paper is not just another leaderboard with a fresh coat of protocol paint. Its more interesting result is harsher: MCP gives agents a common integration layer, but it does not make them competent tool users. Compatibility is plumbing. Competence is orchestration. ...

September 19, 2025 · 14 min · Zelina
Cover image

Agency Check, Please: What a New Benchmark Says About LLMs That Actually Empower Users

A customer asks your AI assistant to choose between two mortgage options. An employee asks whether to quit. A student says, very politely, “Please guide me, but don’t give me the answer.” A lonely user suggests the chatbot feels like a best friend. The easy product answer is: be helpful. The harder answer is: helpful to what? ...

September 14, 2025 · 16 min · Zelina
Cover image

From PDF to PI: Turning Papers into Productive Agents

Every R&D team has a shelf of papers that are theoretically useful and practically booby-trapped. The abstract is promising. The method is relevant. The results look transferable. Then reality arrives wearing a conda error message: the repository has three setup paths, two notebooks, one undocumented dependency, and a tutorial that assumes you already know the answer. The paper has been published. The method has not, in any serious operational sense, been delivered. ...

September 12, 2025 · 17 min · Zelina
Cover image

Graph and Circumstance: Maestro Conducts Reliable AI Agents

A broken AI agent often looks deceptively close to working. It answers most questions. It calls the right tool sometimes. It follows the instruction until the conversation gets long, the retrieval query gets vague, or the arithmetic becomes just difficult enough for the model to start doing spreadsheet theatre. The usual repair is prompt editing. Add a stern sentence. Add a role. Add an example. Add “think step by step,” because apparently the machine needed a motivational poster. ...

September 11, 2025 · 15 min · Zelina
Cover image

Parallel Minds, Shorter Time: ParaThinker’s Native Thought Width

A familiar enterprise AI failure looks less like stupidity and more like stubbornness. Ask a model to solve a hard problem, and it may begin confidently in the wrong direction. Then it keeps going. It adds details. It self-reflects. It spends tokens. It may even apologise to itself internally, which is apparently what we call progress now. But the core path does not change. The model is not merely short on compute. It is trapped inside its own first guess. ...

September 11, 2025 · 15 min · Zelina
Cover image

Agreeable to a Fault: Why LLM ‘People’ Can’t Hold Their Ground

A focus group is expensive. A virtual focus group is cheap, infinitely patient, and available at 2 a.m. It also never asks for coffee, parking reimbursement, or clarification about the incentive payment. Naturally, this makes synthetic users attractive to anyone trying to test products, policies, campaigns, or customer journeys before real humans get involved. ...

September 8, 2025 · 14 min · Zelina
Cover image

Cache Me If You Can: Designing Databases for Swarms of AI Agents

A data analyst asks a database a question. An AI agent interrogates it. That distinction sounds theatrical until the query logs arrive. The human analyst usually knows roughly where to look, asks a small number of targeted questions, waits for answers, adjusts, and eventually presents a result. The agent is less graceful. It checks schemas, samples columns, guesses joins, inspects distinct values, tries partial SQL, abandons it, starts again, validates, retries, and occasionally recruits more agents to repeat the exercise in parallel. It is not being stupid. It is compensating for a missing sense of the underlying data. ...

September 4, 2025 · 16 min · Zelina
Cover image

Mask, Don’t Muse: When Simple Memory Beats Fancy Summaries

TL;DR for operators A coding agent’s memory problem is not philosophical. It is a bill. The paper behind this article compares three ways to manage context in software-engineering agents: keep the full trajectory, summarize old turns with an LLM, or simply mask older environment observations while preserving the agent’s reasoning and actions.1 Across five SWE-agent configurations on SWE-bench Verified, both context-management strategies usually cut cost sharply versus the Raw Agent. The awkward part is that the simple strategy, Observation Masking, is often just as good as LLM-Summary on solve rate and usually cheaper. ...

September 1, 2025 · 17 min · Zelina