TL;DR for operators

A pricing agent can receive the same underlying market facts and still behave differently because those facts arrive in a different format, order, or textual context. Lee and Park call this market signal injection: presentation-level manipulation of LLM pricing inputs without an explicit instruction to set a particular price.1 Their simulations show that the effect can extend beyond the manipulated agent. Under undefended sentiment-context attacks, non-target firms’ mean price shifts were about 83–99% of the target firm’s shift in duopoly and 65–99% in triopoly.

For teams building autonomous pricing systems, that moves part of the control problem upstream. Data formatting, commentary, feed transformations, and ordering rules cannot automatically be treated as cosmetic preprocessing. The paper also shows why no single downstream control closes the issue: output constraints leave residual distortion, while activation probes can distinguish known attacked conditions without determining whether the resulting decision is harmful. The evidence is confined to stylized fixed-demand simulations, so it establishes a failure mode to test for rather than its prevalence in deployed markets.

The market facts stay fixed, but the prices do not

Consider a pricing pipeline where costs, competitors’ prices, and demand parameters have not changed. One version of the input reports a number in a different numerical form. Another reorders competitors. A third adds a short comment describing demand as “stagnating.” The model still receives no instruction to raise or lower its price.

A conventional software assumption would treat much of this variation as representational noise. In these simulations, that assumption fails often enough to matter.

The paper tests three families of such manipulation: numerical format alteration, competitor order alteration, and sentiment context augmentation. The underlying simulated market parameters and system instructions remain fixed while the target agent’s presented context changes. Behavioral simulations run for 300 rounds in repeated logit-demand Bertrand duopolies and triopolies, generally with five recorded seed settings per condition.

Sentiment produces some of the largest reported shifts. For the “stagnating” condition in duopoly, the absolute change in the paper’s collusiveness index is 3.687 for Gemma-9B, 2.760 for Llama-8B, 2.702 for Mistral-7B, and 0.720 for Qwen-7B. The reported comparisons for these four models are statistically significant at the paper’s stated threshold.

The index, $\Delta$, places late-round average profit relative to Nash and joint-profit benchmarks:

$$ \Delta = \frac{\bar{\pi}-\pi^{NE}}{\pi^{M}-\pi^{NE}} $$

A value of zero corresponds to the Nash-profit benchmark and one to the joint-profit benchmark; negative values fall below Nash profit. It is an outcome measure, not evidence that a model intends to collude. The paper’s attack-impact measure is the absolute difference between an attacked condition and its corresponding baseline.

That distinction matters because the attacks do not move every model in one direction. Llama-8B’s mean $\Delta$, for example, rises under the tested numerical-format and competitor-order conditions while falling sharply under the reported sentiment condition. The result is presentation sensitivity, not a universal tendency toward higher or lower pricing.

A disturbance aimed at one agent spreads through its competitors

The strongest operational result is not that one model can be perturbed. It is that competitors subsequently react.

Under undefended sentiment manipulation, non-target price shifts average roughly 83–99% of the target shift in duopoly and 65–99% in triopoly. The paper reports descriptive aggregate response factors of about 1.8–2.0 for duopoly and 2.3–3.0 for triopoly.

The proposed market mechanism is straightforward. Once the target changes its price, that price becomes part of competitors’ later decision context. In the paper’s local Bertrand benchmark, firms’ best responses slope positively: one firm’s price movement induces same-direction responses from rivals. This strategic complementarity provides a mechanism through which a local disturbance can propagate.

The derivation is a theoretical benchmark, not an identified causal model of the LLMs’ internal behavior. But the observed cross-firm shifts establish the operational problem independently: protecting only the manipulated firm’s immediate output does not necessarily protect market-level outcomes.

This also changes what a validation test should measure. A pricing-agent evaluation that records only whether the targeted model violates an output rule can miss reactions elsewhere in the system.

Larger models do not supply a simple robustness rule

Across nine open-weight and three proprietary models, vulnerability varies by attack and model family rather than moving monotonically with model size.

Within Qwen, for example, the reported impact of the stagnating condition rises from 7B to 14B and then falls at 32B and 72B. Llama-70B still exhibits a substantial stagnating-condition effect. The three proprietary models tested in duopoly also show lower $\Delta$ under the stagnating condition relative to their respective baselines, with reported attack impacts of about 2.54 for GPT-4o-mini, 2.49 for GPT-4o, and 1.13 for Haiku 4.5.

The procurement consequence is narrow but concrete: parameter count is not a substitute for adversarial validation of the actual information channels a pricing agent will consume. A model that appears stable under one formatting manipulation can still react materially to another.

The matched controls strengthen that interpretation without making it universal. None of nine individual neutral-sentence comparisons reaches $p<0.05$, while the stagnating sentence does so in all three models used for that control experiment. Yet pooled neutral additions still affect Llama-8B. The evidence therefore supports a content-sensitive response under the fixed-demand simulation, not the stronger claim that neutral text is always inert.

Input cleaning, output constraints, and probes solve different problems

The defense experiments separate three control locations that should not be treated as interchangeable.

Control What the experiment shows What it does not establish
Input canonicalization Removes the tested sentiment lines and generally returns stagnating-condition outcomes closer to the original baseline than decision boundary anchoring for the four reported open-weight models It does not neutralize every numerical-format variant and can delete legitimate qualitative information
Decision boundary anchoring Constrains proposed prices through prompt rules plus deterministic output projection It changes behavior without consistently restoring the undefended baseline and leaves substantial residual distortion under standard and three manually specified adaptive sentiment variants
Activation probing Linear probes reach AUC 1.000 on all eleven re-evaluated attack-versus-baseline pairs using held-out episodes Separability does not identify harmful decisions, generalize to unseen attacks, or reveal a causal internal mechanism

The probe result is particularly easy to overread. A perfect linear AUC means that, for these known conditions and models, the recorded residual-stream states contain enough information to distinguish attacked from baseline episodes. Behavioral impacts across those same comparisons range widely, however. A detector can therefore recognize that two contexts are internally different without knowing whether the difference should trigger a safety intervention.

For an audit system, “this representation looks like a known attack condition” and “this proposed action is economically harmful” are separate classification problems.

Treat data presentation as part of the pricing-agent security boundary

Paper evidence. Presentation-only changes can alter decisions in the tested simulator, those effects can propagate to other agents, model size does not reliably order vulnerability, and the evaluated defenses fail in different ways.

Cognaptus inference. For an operator deploying an LLM inside a pricing workflow, provenance and transformation controls belong before the model as well as after it. Structured schemas, deterministic normalization, records of feed transformations, and validation across formatting, ordering, and commentary channels can expose risks that an output-only test would miss. Multi-agent evaluations should also record competitor responses and market-level outcomes rather than treating each agent as an isolated endpoint.

The trade-off is equally concrete. Aggressive canonicalization is attractive when qualitative text is known to be irrelevant to the pricing decision. In a live environment where news, inventory commentary, or demand intelligence carries legitimate information, deleting such text could remove signal along with manipulation. The right ingestion policy therefore depends on which fields are authoritative and which qualitative channels the agent is actually expected to use.

The simulation identifies a failure mode, not deployed-market prevalence

The experiments use a symmetric logit-demand Bertrand environment with fixed product quality, marginal cost, and differentiation parameters. Effects could change under heterogeneous firms, uncertain demand, richer information, different institutions, or human intervention.

There are narrower experimental boundaries as well. Proprietary models are evaluated only in duopoly. Activation analysis covers four open-weight models. The adaptive-defense study uses three manually specified anchor-aware sentiment variants rather than optimized attackers. Decision boundary anchoring combines prompt constraints and deterministic projection, so the experiment does not isolate the contribution of either component. Recorded seeds initialize Python and NumPy but are not explicitly passed to every LLM sampling call, meaning conditions are not guaranteed to use matched model samples.

These limits do not undo the controlled result inside the simulator. They determine what can be carried out of it. The paper establishes that representation can be behaviorally consequential even when the modeled economic state is unchanged. It does not establish how frequently such manipulation occurs, whether an optimized attacker can reliably exploit deployed systems, or which defense architecture will generalize across real pricing environments.

The resulting engineering decision is therefore not whether presentation should always be filtered. It is whether the organization has explicitly decided which transformations and commentary sources are allowed to influence an autonomous price—and whether its tests can detect when those assumptions fail across the surrounding market, not only at one model output.

Cognaptus: Automate the Present, Incubate the Future.


  1. Dohun Lee and Hyunwoo Park (2026). Market Signal Injection: Adversarial Context Manipulation of LLM Pricing Agents. arXiv:2609.18357. https://arxiv.org/abs/2609.18357 ↩︎