TL;DR for operators
A shared enterprise assistant does not only need to remember more people. It must often determine which fact, permission, instruction, or decision has to exist before another action is valid.
Org-Agent addresses this by decomposing multi-user work into atomic subtasks, connecting prerequisite relationships in a dynamic execution graph, ordering those subtasks so prerequisites run first, and carrying retrieved evidence and intermediate results into later steps.1
The paper’s ablations make that architectural choice more than a diagram. Removing dependency information reduces average performance from 44.03% to 41.61% on GroupMemBench and from 72.39% to 71.10% on MUSES-Bench. Removing the supporting execution tools lowers the same scores further, to 40.40% and 67.10%.
For product teams building assistants around approvals, cross-functional requests, scheduling, handoffs, or permission-sensitive information, the design suggests treating prerequisite management as an explicit system capability. But this is benchmark evidence, not evidence that the framework is ready to govern real organizational processes. Some individual Privacy or Utility measures also worsen even when aggregate benchmark scores rise.
Organizational work introduces prerequisites that memory alone does not resolve
Consider one assistant serving several employees. One person asks it to schedule a meeting. Another asks for a project budget. A manager issues an instruction whose execution depends on information held by someone else.
Retaining all three conversations is necessary, but it does not answer the operational questions. Has the relevant person approved the meeting time? Is the budget request permitted? Which version of a fact is current? Does one user’s instruction override another’s? Must another decision occur before the assistant acts?
This is the distinction behind Org-Agent. The paper frames organizational assistance around both cross-user interaction and cross-user memory, governed by constraints involving users, information, and joint decisions. The resulting problem is not simply retrieval from a larger conversation archive. It is determining valid execution order under dependencies.
That distinction matters because a personal assistant can often treat each request as a mostly self-contained interaction. An organizational assistant may receive a request whose correct next step is to resolve another user’s decision first.
Org-Agent makes prerequisite relationships explicit
Org-Agent is one shared language-model agent serving multiple users. It is not an organization of multiple autonomous agents.
Its central architectural commitment is a dynamic Task Dependency Graph, or TDG. The agent decomposes work into atomic inquiry, decision, or response subtasks and connects them when one subtask requires the output of another.
If obtaining an approval must precede sending a response, the approval node becomes a prerequisite of the response node. If new user input changes an unresolved request, the graph can be updated.
Once those prerequisite relationships exist explicitly, the system orders the nodes so that required predecessors execute first. In graph terms, it uses topological sorting. The underlying rule is straightforward:
If node $v_j$ supplies something required by node $v_k$, then $v_j$ must occur earlier in the execution order.
The TDG should also not be confused with a persistent enterprise knowledge graph. It represents the structure of work being executed: subtasks and their prerequisite-output relationships.
Retrieval and execution memory make the graph usable
Ordering subtasks does not solve the workflow if each node lacks the information needed to execute correctly.
Org-Agent therefore combines the TDG with several evidence and memory mechanisms. It can retrieve interaction history using lexical and semantic similarity, apply metadata-based filters, traverse relationships among records, and read or write execution memory.
Execution memory is particularly relevant to multi-step organizational work. After a node acts, the resulting action and observation can be stored for later nodes:
Operationally, this allows a fact discovered during one subtask, or a decision obtained from another user, to become input to a later dependent step without requiring the system to reconstruct the entire organizational context each time.
The retrieval design also combines normalized BM25 and dense-semantic relevance. The appendix sensitivity test finds the best tested average GroupMemBench result at an equal lexical-semantic weighting, $\alpha=0.5$: 44.03%, compared with 43.36% for lexical-only weighting and 29.53% for dense-only weighting. This is a parameter sensitivity result, not a second thesis, but it reinforces that evidence acquisition is part of the execution mechanism rather than a generic vector-search attachment.
The benchmarks support the architecture, but not every metric improves
On GroupMemBench with GPT-4o-mini, Org-Agent reaches 44.03% average accuracy. BM25 reaches 37.72%, while Hindsight, the strongest compared agent-memory system by average score, reaches 38.79%.
An alternate-backbone test using Qwen3-8B points in the same aggregate direction: Org-Agent reaches 40.40%, versus 34.36% for BM25 and 32.75% for Hindsight. That test is best read as a robustness check on backbone dependence rather than an independent claim about general deployment performance.
MUSES-Bench examines interaction and decision-making rather than memory retrieval alone. Org-Agent improves the reported average over a matched vanilla agent for all three evaluated backbones:
| Backbone | Vanilla | Org-Agent | Difference |
|---|---|---|---|
| GPT-4o-mini | 64.12% | 72.39% | +8.27 pp |
| DeepSeek-V4.1-Flash | 80.78% | 85.02% | +4.24 pp |
| Qwen3-32B | 63.47% | 68.76% | +5.29 pp |
The aggregate averages conceal relevant trade-offs. With GPT-4o-mini, Cross-user Access Utility falls from 64.96% to 56.52%. With Qwen3-32B, Privacy falls from 64.28% to 47.45%.
An enterprise buyer therefore should not translate “higher average benchmark score” into “better permission handling.” The evaluated architecture improves the composite result while producing regressions on some dimensions that may be disproportionately important in an actual organization.
For enterprise assistants, dependency structure becomes a product decision
The paper directly shows benchmark improvements from a system that externalizes prerequisite structure and supplies tools for evidence acquisition and intermediate memory.
Cognaptus’s business inference is narrower: assistants operating in workflows with approvals, authority boundaries, handoffs, scheduling dependencies, or cross-functional information may benefit from representing those prerequisites as system state rather than leaving them implicit inside a prompt.
That changes the design question. Instead of asking only, “What information should this assistant remember?”, a product team also has to ask, “What must be true before this action is allowed to proceed?”
The cost results add another product consideration. On GroupMemBench, Org-Agent uses 5.5 million LLM tokens for its 44.03% average accuracy, compared with 21.6 million for LightMem and 117.5 million for Hindsight. BM25 remains cheaper at 1.2 million tokens but also scores lower at 37.72%. These figures suggest a possible trade-off between structured reasoning during execution and more expensive memory-processing approaches. They do not establish production total cost of ownership.
The unresolved question is whether benchmark structure survives organizational reality
The evaluation covers 745 GroupMemBench questions and 1,183 MUSES-Bench scenarios, not live organizational deployments. GroupMemBench correctness is judged by GPT-5.1, multi-turn MUSES-Bench tasks are capped at ten turns, and the paper reports no confidence intervals, significance tests, or repeated-run uncertainty estimates.
The reported user-count analysis is similarly bounded. Within the evaluated MUSES-Bench range, fitted performance declines more slowly for Org-Agent than for the vanilla agent on Queue, Instruct, and Meeting tasks. That is evidence about the tested scenarios, not a general scaling law for organizations.
The more durable contribution is architectural: once one assistant serves multiple people, correctness can depend on the sequence in which information, permissions, and decisions become available. Org-Agent makes that sequence explicit.
For enterprise agent design, that shifts attention from how much context the system can retain to whether it can determine what has to happen before it acts.
Cognaptus: Automate the Present, Incubate the Future.
-
Luyao Zhuang and Yujing Zhang and Zijin Hong and Yilin Xiao and Xiao Huang (2026). Org-Agent: Beyond Personal Assistants Towards Organizational Agents. arXiv:2609.34392. https://arxiv.org/abs/2609.34392 ↩︎