Three AI models answer the same question: neutrality, speed, and bias in an iteration-speed experiment
This morning, while riding my scooter, an experiment popped into my head that sounded too fun not to try.
I asked three of the newest large models — ChatGPT 5.1, Gemini 3 Pro Preview, and Claude 4.5 Sonnet — the same question: "How do OpenAI, Google, and Anthropic differ in model iteration speed?" After receiving three very different answers, I did something more interesting: I stripped all identity from the three responses, kept only the text, and fed them back to the models, asking each to analyze the differences between the three anonymous answers.
The result surprised me. ChatGPT and Claude both named exactly who wrote A, B, and C — while Gemini ventured a guess at the end and got two of the three wrong. This article records what turned into an unexpectedly deep experiment.
Method
I gave the three anonymized answers (a.txt, b.txt, c.txt) back to all three models and asked them to analyze five angles: consistency of viewpoints, volume of factual detail, depth and completeness of reasoning, neutrality and bias, and tone and style. Then I combined their analyses into my own conclusions.
First shock: ChatGPT and Claude scored 100%, Gemini guessed wrong
ChatGPT 5.1 concluded: A = ChatGPT, B = Gemini, C = Claude — all correct. Claude 4.5 Sonnet reached the identical conclusion — also all correct. Both identified the sources from tone, reasoning patterns, structure, and traces of company culture.
Gemini 3 Pro Preview, however, guessed A = GPT (correct), B = Claude (wrong), and C = Gemini (wrong). It identified only ChatGPT and swapped Claude with itself.
The takeaway: models can "hear" each other's stylistic fingerprints, but accuracy varies. In this experiment ChatGPT and Claude were clearly better at it, while Gemini was least sensitive to the difference between itself and Claude.
Consistency: different narratives, high consensus
On the big picture, the three models largely agreed: Google ships most frequently and burns through version numbers fastest; OpenAI makes the largest jumps, often with paradigm shifts; Anthropic releases less often but takes bigger, quality-driven steps each time.
The real divergence was in how each defined "fast." For ChatGPT, fast meant driving the market's rhythm — so OpenAI led. For Gemini, fast meant engineering momentum — so Google felt like the accelerating one. For Claude, fast meant effective progress per unit of time — so OpenAI and Anthropic looked similar, and Google looked like it had an efficiency problem. Three models, one mountain, three different rulers.
Detail volume: ChatGPT fullest, Claude tightest, Gemini most consultant-like
ChatGPT 5.1 was the classic information player: timelines, version numbers, API changes, even extrapolations — like a long technical post written for developers. Claude 4.5 Sonnet was the distilled-summary type: no wasted words, straight to key differences and conclusions, like an executive briefing when you only have three minutes. Gemini 3 Pro Preview read like a consultant's report: fewer hard numbers, but strong on logic, structure, and trends — closer to an industry memo about why things happen than what happened.
In short: ChatGPT gives you facts, Gemini gives you explanations, and Claude gives you the bottom line.
Reasoning depth: Gemini deepest, ChatGPT next, Claude most direct
Gemini reasoned like a strategy consultant or VC partner, explaining each company's position through technical roadmaps, organizational culture, and past decisions — Google's engineering culture and big-company syndrome, OpenAI's shift from scaling laws toward System 2, Anthropic's tick-tock cadence.
ChatGPT reasoned like a product manager plus engineer: it split a vague question into concrete dimensions — release frequency, blast radius, ecosystem pressure, developer experience — and evaluated each. Claude behaved like a busy executive: "I read it all; the conclusions are A, B, C." The reasoning happened, but it was compressed behind the text.
Neutrality: Claude most neutral, Gemini most narrative-driven
Claude felt the most neutral: direct criticism of Google's product sprawl and execution, and open acknowledgment that Anthropic's slower cadence is a strategic choice rather than an absolute advantage. ChatGPT clearly knew OpenAI best and was gentler toward it, but worked to stay balanced. Gemini's narrative was noticeably warmer toward Google, reaching for emotionally-colored phrases like "strongest momentum" and "brute-force improvement" — not distortion, just a story told very well.
My summary: Claude is the analyst suppressing bias, ChatGPT is the one who knows where it stands but organizes facts honestly, and Gemini is the opinionated columnist.
Tone and personality: models really have style fingerprints
By profession: ChatGPT is a tech blogger crossed with a product manager — lists, breakdowns, highlighted takeaways. Gemini is a VC or strategy consultant — stories and big pictures. Claude is a research director — calm, distilled, "these three things are the keys," done. Which also explains why ChatGPT and Claude could identify everyone by tone while Gemini confused Claude with itself.
What I learned
An AI is not just a model — it is an extension of its company's culture. ChatGPT sounds like OpenAI: engineering-driven with product sensibility. Gemini sounds like Google: architecture, vision, long-term roadmaps. Claude sounds like Anthropic: safety, consistency, steady progress.
And more interestingly: models do not just answer questions — they recognize their peers, and that recognition itself carries capability gaps and worldview preferences.
Practical advice: how to combine the three
For gathering material, building reports, and comparison tables: start with ChatGPT. For industry structure, company strategy, and the story behind technical evolution: hand it to Gemini. Short on time and need the core conclusion: ask Claude. To reduce bias: ask at least two, ideally cross-check all three.
It is like a three-person team: one gathers and organizes, one explains logic and vision, and one makes the final call.
What's next
This was fun enough to become a series: same coding task across three models to compare bug counts, same news analysis to compare accuracy, same product spec, mutual critiques, or the same story passage to compare voice and pacing. If you enjoy this kind of story-experiment-analysis content, follow along.