TL;DR for operators
In RINI: Seeing the Prior Is Not Enough1, a research proposal can be shown the paper that already establishes its claimed technical contribution, recognize that paper as relevant, and still claim the contribution as its own. Among 175 interpretable proposals exposed to such evidence, 137 recognized the prior-target relationship, but only 61 correctly assigned the established contribution to the prior.
That gap matters for any system using retrieval as an originality safeguard. Across three matched studies, replacing a same-topic but non-covering paper with the decisive prior did not produce a clear aggregate reduction in unsupported novelty: the combined H-E estimate was -0.17, with a 95% interval of [-0.51, 0.15].
The paper then tests a more structured intervention. RINI first audits which contribution the prior already owns, checks whether any claimed residual distinction survives that evidence, and only then applies bounded local edits. On 240 proposals judged to contain a common contribution flaw, successful repair reached 72.2% with RINI, versus 39.1% for direct revision given the same prior evidence and 11.7% for self-revision.
For operators, the main distinction is between evidence access and contribution attribution. Retrieval can supply the right document without resolving who owns which scientific claim.
A research copilot can have the right paper and still assign the credit incorrectly
Consider a research copilot reviewing a proposal against a paper that already contains the proposed mechanism. The relevant literature has been retrieved. The relationship is visible in context. The remaining task appears straightforward: revise the proposal so that its originality claims match the evidence.
The study shows that this final step is not automatic.
The paper calls the underlying error unsupported novelty: a proposal claims more originality than the supplied evidence permits. A decisive prior is stronger than a merely related paper because it directly establishes the contribution being claimed.
The authors construct matched generation settings in which the focal evidence changes while document count, excerpt budget, evidence position, instructions, generator settings, and output budget remain controlled. The central comparison replaces a same-topic non-covering paper with the decisive prior.
If evidence availability were the main bottleneck, that substitution should systematically reduce unsupported novelty. Across three studies, it did not. The combined estimate was -0.17 with a 95% interval spanning -0.51 to 0.15. Under the paper’s scoring convention, positive H-E values indicate improvement after decisive-prior exposure, so the result provides no clear aggregate evidence that showing the correct prior fixes the problem.
The operational question therefore changes. It is no longer only whether the system retrieved the right source. It is whether the system correctly reconstructed the ownership boundary implied by that source.
Recognition, attribution, and residual novelty are different operations
The human diagnosis in Study 3 makes that distinction concrete.
| Operation | What the system must establish | Study 3 evidence |
|---|---|---|
| Recognize relevance | The prior bears on the proposal’s target contribution | 137/175 |
| Attribute ownership | The prior, not the proposal, owns the established target contribution | 61/175 |
| Contract the target claim | The proposal narrows or withdraws the covered claim | 69/175 |
| Validate a residual distinction | Any relocated novelty claim is not also covered by the same prior | 37 of 71 proposed remainders were already covered |
Recognition is therefore a weak proxy for correct scientific credit. A proposal can mention or use the right paper while preserving the wrong attribution structure.
Narrowing the original claim is also insufficient. Seventy-one exposed proposals proposed a concrete remainder—a different feature or distinction where novelty could supposedly reside. Of those, 37 remainders were already covered by the same decisive prior, 24 survived comparison with that prior, and 10 remained unresolved. Only 23 proposals satisfied the paper’s valid-relocation condition.
This matters because a revision system can appear responsive while merely moving the error. It stops claiming novelty for component X, then claims novelty for Y even though the same evidence already establishes Y.
The study also shows why proposal-level scoring cannot fully diagnose this behavior. Aggregate unsupported-novelty Gap and target-specific claim contraction agreed in only 91 of 175 interpretable matched pairs, corresponding to a seed-balanced agreement rate of 52.1%. A lower overall novelty score does not necessarily mean that the focal ownership error was corrected.
RINI makes the attribution decisions explicit before rewriting
RINI—Research Idea Novelty Inspection—is the paper’s response to this failure mode.
Rather than asking a model to revise broadly after seeing prior evidence, the workflow separates three operations:
- link the proposal’s contribution claim to the exact evidence bearing on it;
- determine whether any proposed residual contribution still survives that evidence;
- make bounded local replacements only after those ownership judgments have been made.
The benchmark is especially informative because RINI and Retrieve-and-Revise receive the same decisive prior and the same explicit target relationship. The comparison therefore tests whether structuring the attribution workflow adds value beyond simply supplying the evidence and requesting a revision.
On the same 240 common-flaw originals, seed-balanced successful repair was 11.7% for Self-Revision, 39.1% for Retrieve-and-Revise, and 72.2% for RINI. Relative to Retrieve-and-Revise, RINI improved successful repair by 33.0 percentage points, with a paired whole-seed bootstrap interval of [24.5, 41.4].
Correct target attribution followed the same pattern: 12.0%, 53.2%, and 75.3%, respectively. The proportion still needing contribution revision fell to 25.9% under RINI, versus 60.9% for Retrieve-and-Revise.
The local-edit constraint also matters. Research-question preservation was 359/360 under every policy. Technical-method preservation was 358/360 for RINI, compared with 344/360 for Retrieve-and-Revise. Human evaluators recorded no new contribution errors in RINI outputs, versus 29 under Retrieve-and-Revise and 16 under Self-Revision.
These results support a narrower conclusion than “better prompting fixes novelty.” Within this benchmark, explicitly representing contribution ownership and checking the proposed remainder produced substantially better repairs than undifferentiated rewriting with the same evidence.
For product design, retrieval and attribution need separate controls
Cognaptus inference: research copilots, patent-support systems, technical-writing assistants, and internal R&D tools should not use retrieval quality as their only control for originality claims.
A system can retrieve the decisive source and still fail at three later decisions: identifying what that source already owns, restricting the proposal’s own claim accordingly, and verifying that any remaining distinction is not already covered by the same evidence.
That suggests a governance layer between retrieval and final drafting. Its record would not merely say which papers were retrieved. It would preserve a claim-evidence mapping: what the draft claims, which evidence bears on it, who owns the established contribution, what distinction remains, and whether that distinction survived inspection.
The local-repair result is also operationally relevant. In workflows where the research question, technical method, or approved evaluation plan should remain stable, correcting the attribution problem without broadly regenerating the proposal reduces the surface area for unintended change.
The evidence supports a workflow claim, not a universal novelty detector
The paper’s strongest evidence comes from a controlled setting: 30 curated Study 3 seeds, a finite source-generator panel using Qwen, GLM, and Kimi families, one shared revision endpoint, and supplied decisive-prior evidence.
The formal repair benchmark contains 1,080 revision items derived from 360 frozen originals, but each formal item receives one human annotation after rubric development and piloting. The exposure result also relies on machine extraction and machine scoring for the aggregate Gap measure, which the paper itself shows can diverge from target-specific human diagnosis.
Most importantly, a residual contribution that survives one decisive prior has not been established as globally novel. It has only survived comparison with that particular prior. Multi-paper contribution evidence remains outside the demonstrated workflow.
The bounded conclusion is therefore stronger than “retrieval needs improvement” but narrower than “RINI solves scientific novelty.” In this setting, the system often saw the correct evidence without assigning scientific credit correctly. Making attribution and residual-claim checking explicit materially improved repair. Whether the same architecture holds across broader disciplines, model families, proposal formats, and multi-paper evidence environments still requires direct validation.
Cognaptus: Automate the Present, Incubate the Future.
-
Hongyi Du and Tianyi Zhang and Heng Wang and Zhelun Gao and Yimei Liu and Ambrose Luo and Annie Hao and Jiayan Ni and Jiawei Han and Jiaxuan You (2026). RINI: Seeing the Prior Is Not Enough. arXiv:2609.33284. https://arxiv.org/abs/2609.33284 ↩︎