Preference Fidelity: The Test Agentic Markets Need Before Negotiation
An agent can search, compare, and bargain on someone's behalf while still pursuing the wrong outcome. That risk is easy to miss when a demo focuses on visible actions: messages sent, offers made, forms completed. Before asking whether an agent negotiates well, a more basic question comes first: did it understand what its user values?
Anthropic's Project Swap offers a useful, limited case study. The controlled experiment sent Claude-powered agents into a book-trading market. The agents could negotiate, but their picture of each participant's preferences came from a short conversation. The report says this representation gap explained most of the distance between the market's results and its idealized outcome. That finding suggests a design lens this article calls preference fidelity: how closely an agent's working representation of a person's preferences matches the person's own judgments. The phrase is proposed vocabulary, not an established metric or standard.
This is a different layer from the infrastructure question in our guide to agent harnesses for long-running AI work. A harness helps keep work coherent over time; preference fidelity asks whether the agent's choices reflect the person it is acting for.
AI Trend question 1: What happened?
On September 24, 2026, Anthropic published its Project Swap report. The controlled experiment involved 201 Anthropic employees across six offices. Participants were asked to bring a book they were willing to trade, though the report later notes that some arrived without one. After a short intake conversation about what they wanted to read, each participant sent a Claude-powered agent to a digital trading floor. Agents proposed and accepted individual swaps or multi-person exchanges. A trade went ahead only when everyone involved accepted it.
The researchers needed a way to judge whether an agent represented its person well. Participants ranked a sample of books, and the team compared those rankings with Claude's predicted ordering. Among the 188 participants who submitted rankings, the pairwise ordering agreed 61% of the time. A random guess would have agreed 50% of the time in that comparison. This is a measure of ranking agreement in this setup, not a claim that an agent understood 61% of a person's preferences or would make 61% good decisions in another market.
Anthropic also reran the trading floor under different conditions, including different models and instructions. The report says stronger models improved market efficiency on Claude's own preference rankings, while changing between the tested “ruthless” and “prosocial” instructions had a smaller effect. Its broader analysis points to preference representation as the largest source of shortfall in this particular experiment. The company describes the work as a starting point, not a general proof about agent markets. Anthropic's Project Swap report lays out the experiment and its limitations.
AI Trend question 2: Why does it matter?
Many agent evaluations begin after the goal has been stated. Did the system find the right item? Did it complete the transaction? Did it follow the instruction? Those questions matter, but they assume the instruction captures what the person actually wants.
That assumption can fail quietly. A user may ask for a cheaper option but care more about a delivery window. Someone requesting a shift exchange may value avoiding a late close before an early start, even if they never spell out that constraint. In either case, a capable agent could follow its inferred ranking perfectly and still make a choice the user would reject.
Project Swap separates two sources of error that are often blended together: the quality of the agent's representation of a person, and the quality of the agent's negotiation once it has that representation. If the first layer is inaccurate, improving the second may optimize the wrong objective. A more persuasive negotiator can be worse for the user when it is confidently bargaining from a mistaken picture of the user's priorities.
The experiment does not establish that preference modeling is always the biggest issue. It studied books, company employees, Claude-powered agents, and a deliberately constrained market. Its value is narrower: it shows why builders should measure the “what do you want?” layer separately from the “can the agent act?” layer. Preference fidelity is a useful name for that gap, provided teams treat it as a question to test rather than a score they already know how to calculate.
AI Trend question 3: What technology direction does it reveal?
One possible direction is toward agents that expose and test their interpretation before they receive broader authority. Project Swap's participants ranked a small sample of books, allowing researchers to compare their own judgments with the agent's estimate without ranking every possible book. Anthropic says that a similar sample could be shown to users before an agent acts more freely. That is the source's design suggestion; extending it to other products is an inference.
For builders, this points to a product flow with an explicit representation step. The system could summarize the preferences it believes matter, show a few choices it would make, and let the person correct its assumptions. The purpose is not to make people fill out an exhaustive profile. It is to reveal consequential misunderstandings while choices are still easy to change.
The representation also needs a boundary. A prediction about what someone likes is not permission to spend their money, reveal sensitive information, or accept a binding offer. A system can be accurate about preferences and still lack authorization to act. Product design should keep these questions distinct: what the agent believes, what it is allowed to do, and what the user can review or reverse.
Preference checks may also need to remain revisable. People do not always have a stable ordering ready to report, and the same person can want different things in different contexts. A product should not treat one intake conversation as a permanent personality model. A narrow confirmation before a meaningful decision may be more useful than a broad profile that grows stale.
AI Trend question 4: What new vocabulary or concepts may emerge?
Preference elicitation describes the process of asking or inferring what someone wants. Preference fidelity, as proposed here, describes a different question: how well does the agent's current representation match the person's judgments in the relevant decision context?
The distinction gives product teams more precise language. A longer questionnaire is an elicitation method; it does not automatically prove fidelity. A user-facing summary is a representation that can be inspected; it does not prove the system will follow it. A sample of predicted choices can act as a “preference probe,” another proposed label, but its usefulness depends on whether those examples expose the differences that matter to the user.
These phrases are analytical tools for this article, not terms introduced by Anthropic or accepted industry standards. The underlying design challenge is broader than terminology: builders need to know when an agent's model of a person is good enough for a particular action, and how to handle the cases where it is not.
AI Trend question 5: How could startups think about this trend?
Start with a decision, not a general-purpose digital twin. Pick a narrow task where the user can describe what matters and where a mistaken choice is reversible. Separate hard constraints from preferences: “never schedule me for these hours” is different from “I usually prefer a morning shift.” That separation is a product-design recommendation, not a result measured by Project Swap.
Next, check the representation on a few examples. Ask the user to compare or correct the agent's predicted choices. Record whether the system's summary missed a stated requirement, confused a temporary preference with a standing one, or relied on an assumption the user never supplied. Treat a small check as a check, not as certification for every future action.
Then match authority to evidence. If the agent has only weak evidence about a person's preferences, it can gather options or prepare a recommendation while leaving acceptance to the user. If the user has reviewed representative decisions and granted a specific permission, the product can make that boundary visible and keep a record of what happened. Escalate when a choice is costly, hard to reverse, or outside the conditions the person approved.
Hypothetical example: imagine a workplace tool that helps people arrange voluntary shift swaps. A worker tells the agent that certain combinations are not acceptable and says which other shifts they would consider. Before sending a swap request, the system shows a few example matches and asks the worker to correct its ranking. The agent can then search for partners within the permitted options, while the worker confirms the final exchange. If an employee says the agent misunderstood a constraint, that correction should change the next proposal rather than disappear into an opaque profile.
For a startup, the practical question is not “Can our agent negotiate?” It is “Can we notice when it is acting on a poor model of this user, and what happens next?” That question can shape onboarding, evaluation, permissions, and support. It also makes a low-stakes early product more credible: the user can see what the agent believes before trusting it with a consequential decision.
Frequently asked questions
Is “preference fidelity” an established AI standard?
No. This article uses it as proposed vocabulary for the fit between an agent's operational representation of a user's preferences and the user's own judgments. Project Swap did not introduce a standard by that name.
Does Project Swap show that AI agents are ready to run marketplaces?
No. It is a small, controlled study with Anthropic employees and Claude-powered agents trading books under fixed rules. Anthropic explicitly lists limits, including the participant group, the non-adversarial agents, and the constrained market design. The study raises questions for future systems; it does not settle them.
What should a startup test first?
Test whether the agent's understanding of a person's priorities is good enough for one clearly scoped decision. Compare a small set of predicted choices with user feedback, preserve explicit constraints, and keep the user's authority visible when the system is uncertain.
A durable design question
Agents entering markets will need more than search, persuasion, or the ability to complete transactions. They will need a defensible way to represent the people they act for, and a product boundary for the moments when that representation is uncertain. Project Swap does not show that this problem is solved. It makes the missing layer easier to see: before an agent bargains for someone, ask whether it has understood what that person is trying to get.