Cover image

Back to the Drawing Board: How DiagramIR Quietly Fixes Math Diagrams for AI

A diagram is not a paragraph with lines attached. That sounds obvious, which is usually where software product teams get into trouble. Text can be judged by fluency, relevance, and whether the answer has wandered into confident nonsense. A geometry diagram has extra obligations. The side marked 8 should look longer than the side marked 3. The angle labelled $90^\circ$ should not be having an identity crisis. Labels should sit near the thing they label. The image should not be half outside the frame, unless the product strategy is “modern art, but for sixth grade”. ...

November 15, 2025 · 14 min · Zelina
Cover image

When Democracy Meets the Algorithm: Auditing Representation in the Age of LLMs

When Democracy Meets the Algorithm: Auditing Representation in the Age of LLMs Agenda-setting is where participation quietly becomes power. Anyone can invite hundreds of people to submit questions. That part is now cheap. The difficult part arrives ten minutes later, when an expert panel has time to answer only seven of them, and someone has to decide which seven. This is the small administrative hinge on which democratic legitimacy loves to swing. A moderator chooses. A platform ranks. An LLM summarises. Everyone else is told, usually with a straight face, that their concerns were “reflected”. ...

November 7, 2025 · 16 min · Zelina
Cover image

When ESG Meets LLM: Decoding Corporate Green Talk on Social Media

A corporate sustainability post rarely says, “Please admire our reputational risk management.” It says something friendlier. A tree-planting day. A Pride Month banner. A smiling volunteer team. A solar panel photographed at just the right angle. A line about communities, innovation, opportunity, resilience, or the future. The usual words, freshly laundered. The analytical problem is that these posts are not random fluff. They are corporate communication at scale, and they are increasingly multimodal: text, hashtags, brand imagery, infographics, event photos, symbolic gestures, and occasionally something resembling an operational fact. Reading them one by one is theatre. Ignoring them is also a choice, just not a very intelligent one. ...

November 6, 2025 · 16 min · Zelina
Cover image

When AI Packs Too Much Hype: Reassessing LLM 'Discoveries' in Bin Packing

A warehouse manager, a cloud scheduler, and a container-ship planner all know the same unpleasant truth: fitting things into limited capacity is where tidy strategy goes to die. That is why bin packing remains such a useful test case. The problem is easy to explain and difficult to solve optimally. Items arrive. Bins have fixed capacity. The objective is to use as few bins as possible. In the online version, the system must decide where to place each item as it arrives, without seeing the future. This is not just a toy puzzle. It resembles production scheduling, memory allocation, server placement, freight consolidation, and every other operational setting where tomorrow’s workload has the bad manners not to disclose itself in advance. ...

November 5, 2025 · 15 min · Zelina
Cover image

Two Minds in One Machine: How Agentic AI Splits—and Reunites—the Field

Agents have become the new office intern, software engineer, analyst, compliance assistant, and occasional disaster rehearsal all in one. Give one a goal, some tools, a memory store, and permission to act, and it begins to look less like a chatbot and more like a small operating unit. That is the sales pitch. The engineering reality is less tidy. ...

November 3, 2025 · 16 min · Zelina
Cover image

The Rise of FreePhD: How Multiagent Systems are Reimagining the Scientific Method

A broken file link is not usually where scientific revolutions begin. It is, however, where many automated workflows die. That is why the most revealing moment in the freephdlabor paper is not the grand claim about personalised AI research groups. It is the rather unromantic episode where the system tries to write a paper, discovers that the experiment data are missing because of a failed symlink, attempts workarounds, fails validation, reports the failure, gets routed back through resource preparation, rebuilds the workspace correctly, and only then proceeds to manuscript generation.1 ...

October 25, 2025 · 15 min · Zelina
Cover image

Promptfolios: When Buffett Becomes a System Prompt

Investment firms love a house style. Conservative value. Quality growth. Distressed credit. Low-volatility income. The style is supposed to mean something more durable than a portfolio manager’s breakfast mood. The uncomfortable part is that many “styles” still live in a fog of analyst judgement, committee memory, spreadsheet folklore, and the occasional sacred quote from an investor whose annual letters have been read with the reverence normally reserved for scripture. Everyone claims discipline. Fewer can show exactly how that discipline becomes position weights. ...

October 9, 2025 · 13 min · Zelina
Cover image

Branching Out of the Box: Tree‑OPO Turns MCTS Traces into Better RL for Reasoning

Branching Out of the Box: Tree-OPO Turns MCTS Traces into Better RL for Reasoning A search tree is expensive to build. Once you have paid for it, using only the final answers is a little like buying an aircraft engine and admiring the packaging. That is the useful instinct behind Tree-OPO, a paper that asks whether Monte Carlo Tree Search traces from a stronger teacher model can be reused not merely as demonstrations, but as a structured curriculum for training a smaller reasoning policy.1 The idea is not to run MCTS at inference time and call that progress. Nor is it to imitate a teacher’s logits until the student develops the personality of a photocopier. The paper’s more interesting move is subtler: take the partial reasoning states produced by search, let the student complete from those prefixes, and compute advantages in a way that respects where each prefix sits in the tree. ...

September 17, 2025 · 14 min · Zelina
Cover image

Hook, Line, and Import: How RAG Lets Attackers Snare Your Code

Imports look harmless until they become procurement. A developer asks an AI assistant for a plotting snippet. The assistant returns clean-looking Python, a few lines of explanation, and an import statement for matplotlib_safe. The name sounds prudent. Safer is good. Safer is what the security team keeps asking for, usually in meetings that could have been static analysis. ...

September 13, 2025 · 17 min · Zelina
Cover image

Plan, Then Rewrite: Why Explicit Intent Wins in Agent Workflows

A user starts by asking for Italian restaurants, answers a few clarification questions, then changes their mind and asks for Mexican instead. A human hears the reversal. A planner may hear: pizza, pasta, Italian, Mexican, recommendations, and perhaps a vague invitation to overachieve. Naturally, it may then produce a plan with the confidence of a consultant who attended only half the meeting. ...

September 11, 2025 · 14 min · Zelina