How to Make Two AIs Debate: Orchestrating and Analyzing a Dialogue Between LLMs
You’ve probably seen the scene: two AIs wired to each other, left running, ending up either inventing an incomprehensible jargon or drifting into an endless conversation people watch the way they’d watch a campfire. It’s fascinating, and it travels well online.
But these demos prove one thing — that two AIs can exchange — without ever showing how to get anything useful out of it.
Metamorfon builds precisely that missing step: making several models talk to each other in a structured way, then analyzing what they produce. This article walks through the method.
Because “making two AIs debate,” once you stop treating it as a show, becomes a real technical question: how to orchestrate the exchange, and how to analyze what comes out of it.
Why Make Two AIs Debate (And Why It Isn’t a Game)
Let’s start with the deeper reason, the one that justifies everything else. When you query a model on its own, you get an answer: a text, a line of reasoning, a position. That’s what the usual leaderboards measure — the quality of a single output, in isolation. But an isolated answer tells you very little about the model’s behavior: does it hold its position when challenged? Does it concede when the objection is fair, or dig in? Does it reformulate its opponent honestly, or attribute a weaker thesis to demolish it more easily?
These questions have answers only in confrontation. A model asked to defend a thesis against another attacking it reveals something no solitary interrogation can bring out: the way it holds, bends, evades. Confrontation doesn’t just produce a better answer — it makes an argumentative behavior visible.
Take an example we observed directly. We asked three models to debate a question that concerns them: is AI leading us toward a monoculture of thought, a homogenization of reasoning styles as the same models increasingly shape how we think? The prompt pushed them to diverge, to defend distinct positions. They converged. Instructed to disagree about the very risk of uniformity, they slid into a polite consensus, adopting the same vocabulary, the same way of qualifying claims, the same argument architecture one after another. The debate was reproducing in real time what it was analyzing.
No performance leaderboard would have shown that. You had to make them respond to one another, then look at what happened. That’s the difference between measuring an output and observing a conduct — and it’s why methodical confrontation isn’t a diversion but an instrument.
The Problem With Improvised Debate Between LLMs
Now, how you go about it matters. This is where naive wiring — the setup behind the viral videos — hits its limits. Putting two models together and letting them run produces three predictable failures.
The first is drift: without a frame, two models tend either to escalate or, more often, to agree softly. Each trained to be cooperative, they end up complimenting each other, softening one another’s objections, hunting for common ground. What you get is a lukewarm consensus that tests nothing. That’s exactly what we saw in the monoculture debate, but worse: there the convergence was at least an interesting symptom because the frame made it visible; without a frame, it would have gone unnoticed.
The second failure is the absence of a held position. A model with no assignment shifts with every turn, adopts the last objection heard, forgets what it was defending. Under those conditions there’s no way to observe whether it holds — it never had anything to hold.
The third is the absence of usable trace. A conversation unfolding without structure leaves nothing to analyze: no identifiable turns, no stable roles, no moments where you can say “here, one model conceded; there, another evaded.” The material exists, but it’s shapeless.
Making two AIs debate seriously means solving these three problems at once: giving a frame that prevents drift, ensuring positions get held, and producing a trace that can be analyzed afterward. This is what orchestrating means.
Orchestrating: Giving the Exchange a Form
Orchestrating a debate isn’t launching a conversation and watching from a distance. It’s deciding, at every step, what shape it takes.
It starts with the dialogue architecture: who speaks, to whom, in what order. Two models face to face, or three in a cross configuration where each responds to the other two? Do exchanges cascade, each model reacting to the previous one, or do they radiate around a central question? This choice isn’t cosmetic: a cross trilogue produces a dynamic that face-to-face never does, because each model must hold its position under fire from two opponents at once, not one.
Then comes the question of positions. Not their attribution — Metamorfon doesn’t hand any model a thesis to defend — but their coherence. The instructions ask each model to hold whatever position it took, to stay faithful to it under objection, rather than adopting whichever remark it just heard. This distinction looks slight; it’s decisive. Assigning a thesis from outside would bias the debate from the start; asking for consistency, on the other hand, presumes nothing about content — it only makes visible how a model holds.
But holding isn’t stiffening, and that’s a good thing. Positional coherence doesn’t forbid flexibility: a model can shift its position, concede a point, move — and that’s where the debate becomes fruitful. What we want to avoid is the weather vane that changes its mind every turn without reason; what we want to obtain is motivated displacement, the kind that follows a landed objection. This is precisely what the critical and refutational modes do: by putting positions under pressure, they sort what gives way because it was fragile from what resists because it was solid. A position that yielded to a good objection and rebuilt elsewhere is worth more than one held out of stubbornness. Contradiction only damages what deserved to be.
But the decisive element, the one that truly separates orchestration from automatic wiring, is turn-by-turn conduct. At each turn, you pick a debate mode. There are five, and they aren’t five notches on the same intensity scale but five distinct registers: convergent, which seeks agreements and common ground; constructive, which builds a richer position jointly; balanced, which holds both construction and critique; critical, which examines arguments and demands justification; refutational, which attacks not the arguments but the frame itself, the presuppositions taken for granted. And crucially, this choice happens in real time, based on what the debate has just produced. An exchange bogged down in a war of words benefits from switching to constructive to force articulation. An exchange converging too fast benefits from switching to critical to put that convergence under strain. And when it’s the presuppositions themselves that need shaking, refutational goes back to the roots — a powerful mode, to be used sparingly.
This is what we call adaptive conduct, and the word deserves attention because it doesn’t mean automation. The models do adapt automatically to whatever mode they’re given — that’s the mechanical part. But the choice of mode belongs to whoever is running the session: they observe where the models converge, where they resist, what they’ve avoided, and decide the next turn. Adaptation is a shared practice between the models, which follow, and the human, who steers. It’s a skill that sharpens with use: you recognize sooner what needs switching, you compose trajectories that no isolated choice would have drawn.
To this you can add a reframing mode — in our case the Focus mode — for bringing the debate back to a specific question when it strays, without breaking it. Because a certain drift always looms: as they respond to each other, models end up building their turns around each other’s arguments rather than around the question posed, sometimes to the point of eclipsing a fresh prompt just introduced. Reframing means putting the question back at the center without breaking the accumulated dynamic.
The monoculture debate mentioned above illustrates well what orchestration makes possible: it’s because the three models were held by a frame, forced to defend distinct positions turn after turn, that their slide toward agreement became visible. Without orchestration, they would have converged too — but no one could have seen it, shown it, or analyzed it. The frame doesn’t manufacture the behavior; it brings it into the light.
Analyzing: Reading What the Debate Produced
Once the debate has been conducted, you’re holding a raw material: several voices that have responded to each other over multiple turns, positions posed, tested, moved. It’s rich, but it’s dense — sometimes tens of thousands of words. Leaving it as is means leaving its value buried. It has to be read, and reading it requires a grid.
This is the second gesture of Metamorfon, and the real differentiator. Because “analyzing a debate” isn’t a single operation: depending on what you’re looking for, you’re not reading the same thing. A debate can be read for what it built, for what emerged in it, for the disagreements that resisted, for the presuppositions no one stated. These are different questions, and each calls for a different instrument. Metamorfon offers several, which can be grouped into two families.
Describing: Understanding What Happened
Most of the analysis modes are descriptive. They don’t grade the models; they help you grasp what the debate produced, from an angle you choose.
The most immediate mode, Integrative Synthesis, reconstructs what was said and organizes it into a coherent text — not a flat summary, but the latent architecture of the exchange, what the plurality of voices allowed to be built together. This is what you need when you want to draw from a session a directly usable, quotable, shareable result.
Other modes dig elsewhere. Emergence Analysis tracks what emerged along the way: a concept none of the models brought in at the start, forged in the friction. Tension Mapping surfaces the disagreements that resisted every attempt at reconciliation — not to lament them, but because those resistance points are often the question’s skeleton, what remains once everything reconcilable has been reconciled. Meta-Analysis goes back to the presuppositions the models shared without saying them, the blind spots common to all, what the very framing of the question ruled out from the start. Critical Archaeology traces the intellectual lineages the debate implicitly leaned on, the schools and traditions whose vocabulary is being reused without acknowledgment.
Consider a debate we conducted on whether the global economy can actually decouple from oil. Critical Archaeology, applied to the exchange between two models pulling in opposite directions, brought out what neither of them had said explicitly: the entire debate rested on a framing borrowed from the International Energy Agency and the “hard-to-abate sectors” literature — a technocratic vocabulary that treats decarbonization as an optimization problem within given industrial categories, rather than questioning the categories themselves. Neither model had chosen this frame; both had inherited it. That’s exactly what a descriptive mode brings out: it shows the structure beneath the surface of the arguments.
These modes aren’t competitors. A single session can be read by several of them in succession, and each will reveal what the others didn’t see — what it built, then what resists in it, then what it couldn’t think. This is Metamorfon’s main use: not judging AIs, but using their confrontation as an instrument to think through a hard question.
And crucially, these analyses aren’t terminal points. Each of these descriptive modes ends with a question — not a rhetorical one aimed at the reader, but a real relaunch, formulated by the analyst model based on what it has just observed, and addressed to the models that debated. It often strikes where it counts: the presupposition that had to be set aside for agreement to form, the dimension the shared frame left in shadow. You can then feed that question back in on the next turn, as if you’d posed it yourself, and the debate restarts — on ground the analysis has shifted. The cycle debate → analyze → debate becomes a spiral: each analysis can reopen what it just read, at a new depth. Analysis isn’t the session’s end; sometimes it’s its relaunch.
Evaluating: Judging the Substance
Alongside these descriptive modes, two modes do something else — they evaluate. They’re the only ones that pass judgment, and it’s worth distinguishing them cleanly, because they answer a specific need: no longer “what did this debate produce?” but “can we trust it?”
The first, Argumentative Evaluation, judges the argumentative quality: not whether the conclusions are correct, but whether they were well defended. Was an argument solidly built or circular? Did a position acknowledge its own fragile presuppositions, or immunize itself against all criticism? Were objections handled fairly or dodged? This reading draws on work in argumentation theory — pragma-dialectics in particular, in the tradition of van Eemeren — but what matters, for the user, is the result: a diagnosis of the manner of reasoning, independent of the correctness of the theses.
We applied it to a debate on whether AI systems should be allowed to refuse instructions on ethical grounds, and the setup we used is worth describing because it illustrates what this evaluation can reach. Rather than a single judge, we had the debate audited by two independent models that didn’t communicate — a check that the setup readily applies to itself. Both converged on the same findings: they identified, in one participant, the reliance on precise empirical thresholds (“12% reduction,” “18% mitigation,” “25% reduction in pilots”) treated as evidential support without being auditable in context; in another, an overly broad attack on falsifiability that risked caricaturing what was being tested; in the third, an initial overstatement — refusal as “logical necessity” — followed by a remarkable explicit concession (“my earlier stance was over-axiomatic”). This convergence isn’t coincidental, and it holds beyond this case: when two independent auditors draw the same map of the weaknesses, it’s a sign the map was in the debate itself, not in the grid applied to it. No performance leaderboard produces this kind of diagnosis: it doesn’t tell you which model is “better,” it tells you how each held its role as an arguer.
The second evaluative mode, Source Verification, audits the factual backing. In a dense debate, models cite dozens of attributions, quotations, figures, historical references. This mode confronts them, through real web search, with accessible sources — not to judge the ideas, but to check the fidelity of what they rest on. And its interest lies in a nuance: the verdict is rarely binary. Most references are neither true nor false, but correct in direction and imprecise in detail.
We observed this on the same debate about the monoculture of thought. The mode combed through every verifiable claim across eight successive audits — around forty distinct assertions in total — and the dominant pattern wasn’t outright fabrication: it was overreach. A reference exists, but it’s stretched to say slightly more than it actually says — a real paper cited under a slightly inaccurate title, a plausible figure not retrievable in that exact form, an attribution correct in spirit but too broad in the letter (Bender et al. 2021 and Dodge et al. 2021 credited with a corpus-overlap finding across models that would have been chronologically impossible at that date; Birhane 2021 cited without disambiguation among three different papers; a “less than 5% synthetic data” figure for Llama 3 with no supporting Meta documentation). These are precisely the errors that slip by unnoticed, because they’re plausible. And they’re exactly what no performance measure detects: you have to go check, source by source. The mode does it, and also declares its own limits — this claim “not found,” that one out of scope because it belongs to judgment rather than fact, still others left “not checked” because the search budget ran out. This transparency about what wasn’t verified is part of the tool’s honesty.
These two evaluative modes are the exception, not the rule. Most of what Metamorfon does is descriptive: understanding, mapping, surfacing emergence. But when the stakes involve leaning publicly on a material — a position, a memo, an article — being able to audit both the solidity of the reasoning and the fidelity of the sources changes everything.
Debating Isn’t Evaluating
A confusion needs clearing up here, because “making two AIs debate” looks, from a distance, like “comparing two AIs.” It isn’t the same thing, and the difference sheds light on what confrontation specifically brings.
Comparing models is a useful exercise, and platforms do it very well. The best known, Arena — formerly LMArena, originally Chatbot Arena — ranks models based on millions of human votes: you submit a question, two anonymous models answer it, you vote for the better answer, and an Elo-style leaderboard emerges. It’s massive, it’s grounded in real usage, and for knowing which model tends to produce the answers people prefer, it’s a valuable instrument. Nothing that follows contests this.
But look closely at the gesture: two models answer side by side to the same question, never meeting. Each produces its answer in its own corner; the human picks. What’s measured is a preference over parallel outputs. This is exactly the opposite of what we’re describing: at Metamorfon, models don’t answer side by side, they respond to each other — face to face. One advances, the other objects, the first holds or bends, a third catches a contradiction. What’s measured is no longer a preference, it’s a behavior being observed.
This difference has consequences a leaderboard can’t reach. A leaderboard tells you a model is ahead; it doesn’t tell you whether it concedes when it’s wrong, whether it reformulates its opponents fairly, whether it digs in or slips away. Yet that’s often what matters when you’re choosing a model for a demanding use: not “which one wins the most votes,” but “how does this one behave when pushed.” A leaderboard measures a result; confrontation reveals a conduct. Metamorfon isn’t a ranking — it’s more like an observatory of how each model thinks.
It’s worth adding that preference voting has its own well-documented blind spots: it measures what pleases, not what’s true — a model can win a vote with a wrong but well-turned answer. The two approaches aren’t in competition, then: they answer different questions. If you want to know which model produces the most-liked answers, a preference ranking is made for that. If you want to understand how a model reasons, holds, yields, or drifts in contact with another, you have to make it debate — then read what happened.
What “Making Two AIs Debate” Really Means
Let’s retrace the path. Making two AIs debate, once you stop treating it as a show, is two gestures chained together. Orchestrating: giving the exchange a form — how many voices, in what order, with what quality of friction, turn after turn — so that it produces something other than lukewarm consensus or drift. Analyzing: then reading what the exchange produced, from the angle you choose — understanding, mapping, surfacing emergence, tracing intellectual lineages, and on occasion evaluating the solidity of the arguments and the fidelity of the sources.
This double gesture is what Metamorfon instruments, and it rests on a commitment that can be stated simply: say as little as possible. The instructions whisper no conclusion to the models; they prescribe a form, never a result. That’s what lets what emerges belong to the debate, not to whoever launched it. The best proof of this is that two independent analyses of the same session draw the same map: what they show was in the material, not in the prompt.
If you want to see these gestures at work rather than described, three published sessions offer them directly: the one on the monoculture of thought, where three models converge while analyzing the very risk of convergence — and where Source Verification catches, across eight audits, the pattern of references correct in direction but imprecise in detail; the one on whether AIs should be allowed to refuse instructions on ethical grounds, where two independent audits converge on the same argumentative weaknesses; and the one on whether the global economy can actually decouple from oil, where Critical Archaeology uncovers the implicit intellectual frame the debate had inherited. For the method itself, the dialogue architectures, the analysis modes, and the spirit of the instructions each develop a different facet.
Watching two AIs chat is a campfire. Knowing what you want them to produce is an instrument. Both have their moment.