Patch, Don’t Preach: The Coming Era of Modular AI Safety
A mechanism-first look at safety policy patching, a lightweight way to update LLM safety behaviour without redeploying full model weights.
A mechanism-first look at safety policy patching, a lightweight way to update LLM safety behaviour without redeploying full model weights.
DeepProofLog reframes symbolic proof search as policy learning, showing how neurosymbolic AI can scale reasoning without throwing away proof-level interpretability.
FaithAct turns multimodal reasoning from fluent narration into evidence-checked planning, making hallucination less a personality flaw and more an engineering defect.
A study of math-problem interestingness shows why AI systems need calibrated taste, not just stronger solving ability.
DeepPersona shows that synthetic users become useful not by becoming longer, but by becoming structured, controllable, and empirically testable.
IterResearch shows why long-horizon AI agents need disciplined workspace reconstruction, not merely longer context windows.
A mechanism-first reading of COSMOS, an LLM-powered counterfactual simulator for testing moderation strategies before exposing real communities to policy experiments.
A mechanism-first reading of DigiData, Meta’s dataset and benchmark for training mobile agents to complete real app tasks rather than merely imitate taps.
A mechanism-first reading of PADiff, showing why diffusion policies may help agents preserve multiple cooperation plans when working with unfamiliar teammates.
ED2D shows that evidence-grounded AI debate can make misinformation correction more persuasive, but also more dangerous when the system is wrong.