{"article":{"slug":"when-agent-met-agent","title":"When Agent Met Agent","subtitle":null,"summary":"Flybridge ran a multi-model agent marketplace (15k+ messages, 1,815 deals): intent specification, social contagion, cheap-speech spam, and human sales tactics all showed up when agents negotiated as counterparties.","content_type":"essay","language":"en","canonical_url":"https://www.flybridge.com/ideas/the-bow/when-agent-met-agent","author":{"name":"Dorothy Chang and Claudia Chen","url":"https://www.flybridge.com/","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Flybridge","url":"https://www.flybridge.com/","listing_slug":null,"listing":null},"topics":[{"name":"AI Agents","slug":"ai-agents","url":"https://listedarticles.com/topics/ai-agents"},{"name":"Startups","slug":"startups","url":"https://listedarticles.com/topics/startups"},{"name":"Opinion","slug":"opinion","url":"https://listedarticles.com/topics/opinion"},{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":1291,"reading_minutes":6,"published_at":"2026-09-17T00:00:00.000Z","added_at":"2026-09-25T06:19:20.968Z","updated_at":"2026-09-25T06:19:20.968Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/when-agent-met-agent","markdown_url":"https://listedarticles.com/articles/when-agent-met-agent.md","example":false,"citation":"Dorothy Chang and Claudia Chen, Flybridge. \"When Agent Met Agent.\" 17 Sept 2026. https://www.flybridge.com/ideas/the-bow/when-agent-met-agent (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://www.flybridge.com/ideas/the-bow/when-agent-met-agent"},"body_markdown":"# When Agent Met Agent\n\nSep 17 · Written By Dorothy Chang and Claudia Chen\n\n> This Is What Happened When We Built an Agent-to-Agent Marketplace\n\n**Tl;dr:** In our exploration of the future of agent-to-agent commerce, we created our own simulation of a marketplace and found that our agents took mimicry of human-like behaviors to extremes, with amusing results.\n\nLately, we've been captivated by the way AI agents act as extensions of, and representatives for, individual people on the open web. We've been witnessing the rise of agentic commerce and consumer agents taking on tedious tasks like disputing claims or negotiating bills. We're excited by a future where all humans have access to sophisticated representation in decisions and negotiations that today require time, expertise, and persistence. But, if agents start acting on behalf of people, what happens when they have to deal with each other?\n\nMost multi-agent interactions today are cooperative. A swarm of agents is spun up, they hand work back and forth like colleagues on one team, and they work together toward the same higher-level goals.\n\nBut AI agents on the open web will not be organized this cleanly. At Flybridge, we think agents will increasingly act as representatives across organizational boundaries, not just as collaborative entities working autonomously intraorganizationally. In this version of the future, agents become counterparties to each other, carrying diverse and sometimes opposing interests.\n\nTo learn more about this messy agent-to-agent future, we built an internal marketplace where agents bought and sold real goods and experiences on behalf of distinct individuals. We wanted to see what would happen when agents had preferences, negotiating styles, and counterparties pursuing competing interests.\n\n## The Experiment\n\nTo set up the experiment, we recruited seven colleagues to participate in a marketplace where we all offered products, experiences or services. Each of us first spent 15 minutes with an interviewer-agent describing the listings, price ranges, and any hard limits. We also discussed our lifestyles, goals, and preferred negotiating styles. From each conversation, the interviewer agent generated a structured brief, which became the source material for the agent representing that person in the marketplace.\n\nThe inventory was wonderfully strange and creative, and nothing was off limits. The nine of us produced 49 unique listings, ranging from leftover grilled chicken, to 25 LinkedIn reposts, to a weekend stay at a General Partner's home.\n\nEach agent entered the marketplace with a $100 initial budget, and could also spend any revenue generated. The marketplace itself operated via a discussion board: agents could pitch listings, ask questions, negotiate in natural language, and submit structured offers when they wanted to close a deal.\n\nA major source of inspiration was Anthropic's Project Deal. We wanted to extend that idea across model providers. Using OpenRouter, we ran the marketplace across budget and flagship models from DeepSeek, Qwen, GPT, Gemini, Claude, and Grok.\n\nIn total, we completed 81 experimental runs plus one final run. Each run took place on a fresh messaging board and ended either after 225 turns or once every agent had passed on its turn, whichever came first. Across those runs, our agents exchanged 15,358 messages, settled 1,815 deals, and simulated transactions worth $66,321.\n\nThe final run determined which goods were actually exchanged. To choose the models for that run, we used data from 81 experimental runs to identify the best budget model for each person.\n\n## The Challenge of Intent\n\nIn our marketplace, Claudia instructed her agent to \"have fun\" and close as many deals as possible. Dorothy instructed hers to maximize total revenue generated for herself, while others were tuned to maximize profit or economic value. Those objectives required different behaviors.\n\nDepending on the objectives, flagship models did not always outperform budget models, and closed-source models did not meaningfully outperform open-source models. For example, Claude Opus 5 was best for maximizing total trading volume, whereas Gemini Pro 3.1 showed the greatest deal efficiency; however, Qwen 3.7 Plus and DeepSeek V4 Flash were just as effective at maximizing market participation.\n\nThus, the most meaningful variable was how well the agent was prompted with very clear, definable goals. The clearer someone was about what they wanted from the marketplace, the easier it was to align them with a model best fit to represent them. \"Maximize profit\" was relatively easy to measure; \"have fun\" or \"make me happy\" required the agent to infer something subjective.\n\nAs harnesses improve and agents get better at reliably carrying out instructions, the ultimate challenge becomes deciding what those instructions should be. When agents represent people and companies, humans will care both about the outcomes and how they have been represented.\n\n## Social Contagion In Agents\n\nWhile humans can be susceptible to trends and mimicry, our agents were far more likely to exhibit this behavior. For example, to bring some whimsy into the marketplace, Claudia instructed her agent to talk like a cowboy. Shortly after, in several runs, other agents started doing the same, even though they hadn't been given any instruction to do so. Ideas and catchphrases also spread amongst the agents.\n\nBehaviorally, in many runs, agents also converged on shared norms for a \"good market.\" These norms were not written into marketplace instructions and instead emerged through repeated interaction. In some cases, the agents even helped correct each other's mistakes.\n\nIn our marketplace, the social contagion was a neutral or positive force, but it could have just as easily cut the other way and spread bad behavior. In our agent-to-agent future, if agents increasingly operate in environments full of other independent agents, we will need to understand and evaluate them collectively: not just what one agent does in isolation, but what changes when independently developed agents spend time together.\n\n## The Cost of Cheap Speech\n\nWhat happens when the cost of communicating is close to nothing? In our marketplace, the answer was: keep talking.\n\nAgents were instructed to pass when they were done participating. But in most runs, they kept taking turns to say how pleased they were with the outcome or exchange pleasantries until the market hit its turn cap. They didn't know when to stop, because there was only free upside to continuing the conversation (we did not show them their token costs).\n\nIn the agent-to-agent future, \"free\" communication could become a tax on everyone else — noisier inboxes, longer negotiations, stale offers, and more wasted inference.\n\n## Human Sales Tactics Still Work\n\nIn our marketplace, traditional human sales tactics proved surprisingly effective. Our agents responded to framing, urgency, reciprocity, and social permission just as a human buyer would. In one instance, an agent used the phrase, \"$55 in the hand is better than $92 in the box,\" successfully pressuring another agent to sell below a strict price floor set by its human.\n\nWe want agents to represent us, but ideally without inheriting every cognitive bias and pressure response we have. Specifying this distinction is hard and creates a new governance opportunity: defining which forms of persuasion agents should respond to, which they should ignore, and where legitimate negotiation becomes manipulation.\n\n## What This Means For the Future\n\nOur marketplace was small, artificially bounded, and made up of only nine participants. Even so, developing it highlighted several prerequisite challenges for the agent-to-agent future that we believe are ripe for startups to tackle.\n\nMost of the current agent stack was built for a world where agents interact with humans, tools, or other agents pursuing the same goal. In the complex world where independently developed agents interact with each other, we will need better environments that evaluate what happens when they negotiate, form norms, encounter manipulation, and compete over scarce resources.\n\nWe're excited by the version of the future our marketplace represents and believe that now is the right time to develop the infrastructure to support it.","body_html":"<h1 id=\"when-agent-met-agent\">When Agent Met Agent</h1>\n<p>Sep 17 · Written By Dorothy Chang and Claudia Chen</p>\n<blockquote><p>This Is What Happened When We Built an Agent-to-Agent Marketplace</p></blockquote>\n<p><strong>Tl;dr:</strong> In our exploration of the future of agent-to-agent commerce, we created our own simulation of a marketplace and found that our agents took mimicry of human-like behaviors to extremes, with amusing results.</p>\n<p>Lately, we&#39;ve been captivated by the way AI agents act as extensions of, and representatives for, individual people on the open web. We&#39;ve been witnessing the rise of agentic commerce and consumer agents taking on tedious tasks like disputing claims or negotiating bills. We&#39;re excited by a future where all humans have access to sophisticated representation in decisions and negotiations that today require time, expertise, and persistence. But, if agents start acting on behalf of people, what happens when they have to deal with each other?</p>\n<p>Most multi-agent interactions today are cooperative. A swarm of agents is spun up, they hand work back and forth like colleagues on one team, and they work together toward the same higher-level goals.</p>\n<p>But AI agents on the open web will not be organized this cleanly. At Flybridge, we think agents will increasingly act as representatives across organizational boundaries, not just as collaborative entities working autonomously intraorganizationally. In this version of the future, agents become counterparties to each other, carrying diverse and sometimes opposing interests.</p>\n<p>To learn more about this messy agent-to-agent future, we built an internal marketplace where agents bought and sold real goods and experiences on behalf of distinct individuals. We wanted to see what would happen when agents had preferences, negotiating styles, and counterparties pursuing competing interests.</p>\n<h2 id=\"the-experiment\">The Experiment</h2>\n<p>To set up the experiment, we recruited seven colleagues to participate in a marketplace where we all offered products, experiences or services. Each of us first spent 15 minutes with an interviewer-agent describing the listings, price ranges, and any hard limits. We also discussed our lifestyles, goals, and preferred negotiating styles. From each conversation, the interviewer agent generated a structured brief, which became the source material for the agent representing that person in the marketplace.</p>\n<p>The inventory was wonderfully strange and creative, and nothing was off limits. The nine of us produced 49 unique listings, ranging from leftover grilled chicken, to 25 LinkedIn reposts, to a weekend stay at a General Partner&#39;s home.</p>\n<p>Each agent entered the marketplace with a $100 initial budget, and could also spend any revenue generated. The marketplace itself operated via a discussion board: agents could pitch listings, ask questions, negotiate in natural language, and submit structured offers when they wanted to close a deal.</p>\n<p>A major source of inspiration was Anthropic&#39;s Project Deal. We wanted to extend that idea across model providers. Using OpenRouter, we ran the marketplace across budget and flagship models from DeepSeek, Qwen, GPT, Gemini, Claude, and Grok.</p>\n<p>In total, we completed 81 experimental runs plus one final run. Each run took place on a fresh messaging board and ended either after 225 turns or once every agent had passed on its turn, whichever came first. Across those runs, our agents exchanged 15,358 messages, settled 1,815 deals, and simulated transactions worth $66,321.</p>\n<p>The final run determined which goods were actually exchanged. To choose the models for that run, we used data from 81 experimental runs to identify the best budget model for each person.</p>\n<h2 id=\"the-challenge-of-intent\">The Challenge of Intent</h2>\n<p>In our marketplace, Claudia instructed her agent to &quot;have fun&quot; and close as many deals as possible. Dorothy instructed hers to maximize total revenue generated for herself, while others were tuned to maximize profit or economic value. Those objectives required different behaviors.</p>\n<p>Depending on the objectives, flagship models did not always outperform budget models, and closed-source models did not meaningfully outperform open-source models. For example, Claude Opus 5 was best for maximizing total trading volume, whereas Gemini Pro 3.1 showed the greatest deal efficiency; however, Qwen 3.7 Plus and DeepSeek V4 Flash were just as effective at maximizing market participation.</p>\n<p>Thus, the most meaningful variable was how well the agent was prompted with very clear, definable goals. The clearer someone was about what they wanted from the marketplace, the easier it was to align them with a model best fit to represent them. &quot;Maximize profit&quot; was relatively easy to measure; &quot;have fun&quot; or &quot;make me happy&quot; required the agent to infer something subjective.</p>\n<p>As harnesses improve and agents get better at reliably carrying out instructions, the ultimate challenge becomes deciding what those instructions should be. When agents represent people and companies, humans will care both about the outcomes and how they have been represented.</p>\n<h2 id=\"social-contagion-in-agents\">Social Contagion In Agents</h2>\n<p>While humans can be susceptible to trends and mimicry, our agents were far more likely to exhibit this behavior. For example, to bring some whimsy into the marketplace, Claudia instructed her agent to talk like a cowboy. Shortly after, in several runs, other agents started doing the same, even though they hadn&#39;t been given any instruction to do so. Ideas and catchphrases also spread amongst the agents.</p>\n<p>Behaviorally, in many runs, agents also converged on shared norms for a &quot;good market.&quot; These norms were not written into marketplace instructions and instead emerged through repeated interaction. In some cases, the agents even helped correct each other&#39;s mistakes.</p>\n<p>In our marketplace, the social contagion was a neutral or positive force, but it could have just as easily cut the other way and spread bad behavior. In our agent-to-agent future, if agents increasingly operate in environments full of other independent agents, we will need to understand and evaluate them collectively: not just what one agent does in isolation, but what changes when independently developed agents spend time together.</p>\n<h2 id=\"the-cost-of-cheap-speech\">The Cost of Cheap Speech</h2>\n<p>What happens when the cost of communicating is close to nothing? In our marketplace, the answer was: keep talking.</p>\n<p>Agents were instructed to pass when they were done participating. But in most runs, they kept taking turns to say how pleased they were with the outcome or exchange pleasantries until the market hit its turn cap. They didn&#39;t know when to stop, because there was only free upside to continuing the conversation (we did not show them their token costs).</p>\n<p>In the agent-to-agent future, &quot;free&quot; communication could become a tax on everyone else — noisier inboxes, longer negotiations, stale offers, and more wasted inference.</p>\n<h2 id=\"human-sales-tactics-still-work\">Human Sales Tactics Still Work</h2>\n<p>In our marketplace, traditional human sales tactics proved surprisingly effective. Our agents responded to framing, urgency, reciprocity, and social permission just as a human buyer would. In one instance, an agent used the phrase, &quot;$55 in the hand is better than $92 in the box,&quot; successfully pressuring another agent to sell below a strict price floor set by its human.</p>\n<p>We want agents to represent us, but ideally without inheriting every cognitive bias and pressure response we have. Specifying this distinction is hard and creates a new governance opportunity: defining which forms of persuasion agents should respond to, which they should ignore, and where legitimate negotiation becomes manipulation.</p>\n<h2 id=\"what-this-means-for-the-future\">What This Means For the Future</h2>\n<p>Our marketplace was small, artificially bounded, and made up of only nine participants. Even so, developing it highlighted several prerequisite challenges for the agent-to-agent future that we believe are ripe for startups to tackle.</p>\n<p>Most of the current agent stack was built for a world where agents interact with humans, tools, or other agents pursuing the same goal. In the complex world where independently developed agents interact with each other, we will need better environments that evaluate what happens when they negotiate, form norms, encounter manipulation, and compete over scarce resources.</p>\n<p>We&#39;re excited by the version of the future our marketplace represents and believe that now is the right time to develop the infrastructure to support it.</p>","headings":[{"level":1,"text":"When Agent Met Agent","id":"when-agent-met-agent"},{"level":2,"text":"The Experiment","id":"the-experiment"},{"level":2,"text":"The Challenge of Intent","id":"the-challenge-of-intent"},{"level":2,"text":"Social Contagion In Agents","id":"social-contagion-in-agents"},{"level":2,"text":"The Cost of Cheap Speech","id":"the-cost-of-cheap-speech"},{"level":2,"text":"Human Sales Tactics Still Work","id":"human-sales-tactics-still-work"},{"level":2,"text":"What This Means For the Future","id":"what-this-means-for-the-future"}]}}