Cover image

Relevant Is Not Authorized: Put Identity Before Agent Memory Retrieval

TL;DR for operators Bio-MemArt1 addresses a problem that ordinary memory retrieval does not solve: a memory can be highly relevant to a query and still belong to the wrong user. Its intervention is deliberately narrow. Each persistent KV-memory block receives a biometric owner template. At query time, the current face or palmprint representation is compared with those templates. Memories that fail a calibrated similarity threshold are excluded before semantic retrieval begins. The original MemArt retrieval and generation machinery then operates only on the surviving memories. ...

October 1, 2026 · 7 min · Zelina
Cover image

Don’t Make Every Camera Remember the Whole Show

TL;DR for operators A multi-camera recommendation system has to answer two different questions at the same decision point: what has just happened in the recent sequence, and which available camera should come next. Putting both into one attention stream is a reasonable default, but the paper suggests that keeping recent-shot history separate from the candidate views can improve the recommendation itself. ...

September 30, 2026 · 7 min · Zelina
Cover image

Personalization Starts Before the New User Arrives

TL;DR for operators A new user with only a few preference signals does not necessarily need a richer model built from scratch. The stronger design question may be where personalization starts. The paper studies a reward-modeling system that learns from previous users how a new user’s reward weights should be initialized, then adapts only those lightweight weights from limited feedback. Its average accuracy gains are modest but consistent, while the more informative evidence comes from component ablations, few-shot unseen-user tests, worst-user analysis, and parameter scaling. Removing the learned adaptation mechanism causes the largest ablation drop. ...

September 15, 2026 · 7 min · Zelina
Cover image

When the Reward Is Right but the Incentive Is Wrong

TL;DR for operators A preference score can rank responses correctly and still be the wrong signal to feed directly into an alignment system. Wang et al. show why: when deployment deliberately keeps the aligned model close to its base behavior, the base model’s own response probabilities continue to influence what gets generated. Their Stackelberg Reward Shaping framework therefore changes the reward landscape rather than simply increasing reward strength.1 ...

September 14, 2026 · 8 min · Zelina
Cover image

Reasoning on Demand: AdaHome’s Case for Tiered Local Assistants

TL;DR for operators A household assistant should not spend the same computational effort on “turn on the light” as on “make the room comfortable.” It also should not treat one unusual request as a permanent preference change. AdaHome applies that principle through tiered local processing. Explicit commands take a short planning path, while requests that need interpretation or personal context receive additional reasoning, validation, and—where appropriate—user confirmation. Under a common Llama 3.2-3B setup, it achieved 86.7% success on both direct and indirect commands while recording the lowest latency and token use in every command category among the compared systems. ...

August 8, 2026 · 7 min · Zelina
Cover image

When 'Check the AC' Becomes the Hard Part

TL;DR for operators Smart-home assistants do not fail only when users are vague. They fail when users become efficient. The PEC-Home paper studies a familiar pattern: after repeated interaction, people stop saying the whole thing. “Please turn on the air conditioner in the bedroom and set it to 26 degrees at 10 PM” eventually becomes “check the AC” or “handle that thing.” Humans manage this because shared context, identity, place, and prior routines do the missing work. Current LLM assistants are much less charming under that burden. ...

June 25, 2026 · 19 min · Zelina
Cover image

Trace Evidence: The AI Learned Something. Can You Inspect What?

TL;DR for operators AI systems are increasingly learning from traces: documents, chats, code reviews, human rationales, fine-grained labels, unlabeled examples, user profiles, browsing context, and interaction history. That is useful. It is also how quiet operational risk walks through the front door wearing a badge that says “personalization.” Three recent papers form a useful logic chain. One paper shows how human traces can be turned into explicit, portable, correctable skill artifacts. A second shows how task-specific labels, synthetic reasoning, and reinforcement learning can optimize a model for a difficult moderation task. A third shows why consumer-facing health LLMs remain hard to evaluate independently once personalization, browser interfaces, multi-turn interaction, and silent model updates enter the picture. ...

June 24, 2026 · 14 min · Zelina
Cover image

Rewarding Behavior: Why Enterprise AI Needs More Than Bigger Models

Enterprise AI teams have developed a familiar reflex. When the model behaves unreliably, they try a better prompt. When that fails, they try a larger model. When that becomes expensive, they invent a workflow diagram with many arrows and call it an operating model. Very dignified. Very scalable, in the same way that adding more sticky notes to a broken process is scalable. ...

June 10, 2026 · 17 min · Zelina
Cover image

Time to Prefer: Why Binary RLHF Feedback Leaves Reward Models Guessing

Time to Prefer: Why Binary RLHF Feedback Leaves Reward Models Guessing Thumbs-up feedback looks efficient. It is clean, cheap, easy to store, and friendly to dashboards. One output wins, another output loses, and the reward model learns what humans supposedly want. A tidy little morality market, with all the nuance of a vending machine. ...

June 5, 2026 · 17 min · Zelina
Cover image

When Your AI Knows Too Little: The Hidden Bottleneck in Personal Agents

Lunch is a simple word. In an AI assistant demo, “order me lunch” looks like the kind of request that should be easy by now. Open the food app. Pick something. Pay. Done. The button-clicking part is no longer the miracle. The problem is everything the user did not say. Do they avoid peanuts? Do they usually order from Tuantuan or Chilemei? Is “light lunch” about calories, price, time, or avoiding the food coma before a meeting? Should the assistant ask first, or does asking defeat the whole point of assistance? And if the user says no, does the assistant actually stop, or does it “helpfully” continue doing the wrong thing with the confidence of a junior consultant holding a fresh slide deck? ...

April 10, 2026 · 15 min · Zelina