News
ChatGPT, Gemini and Claude Traded Praise as Billions Ride on It
ChatGPT, Gemini, Claude and Perplexity rated each other’s strengths and flaws, but the friendly reviews landed amid a real fight for enterprise AI budgets.
Tom’s Guide asked ChatGPT, Gemini, Claude and Perplexity to give an honest opinion of each other this month, and all four obliged with surprisingly polite report cards. Each chatbot named a rival’s best trait and its most obvious flaw, and none dodged the question.
What that write-up did not weigh is timing. Anthropic’s Claude has just overtaken OpenAI’s ChatGPT for enterprise AI budget share, Perplexity is sitting on a fresh valuation near $20 billion, and every one of these four companies is fighting for a slice of a market Menlo Ventures now counts in the tens of billions. The model handing out compliments about a rival is built by a company with real money riding on how that compliment gets read.
Four Chatbots Grade Each Other’s Homework
Tom’s Guide writer Elton Jones, who covers AI for the outlet, fed each assistant the same prompt: give an honest, unfiltered opinion of your three biggest rivals. ChatGPT praised Gemini’s research depth while noting it sometimes trades depth for breadth. It called Claude the most naturally flowing writer among the major assistants, then added that Claude can be overly cautious or wordy. It credited Perplexity’s inspectable citations but said the tool struggles with long, collaborative projects compared to itself and Claude.
Perplexity answered with the bluntest scorecard of the four. Its own summary read almost like a press release for its rivals:
- ChatGPT: the most well-rounded and adaptable of the four.
- Gemini: strongest for up-to-date information and integration with Google’s own tools.
- Claude: strongest for deep thinking and careful writing.
Perplexity also said plainly that if forced to summarize the field, ChatGPT is the best generalist, Claude is the best thinker and Gemini is the best connected.
Compliments Come With Careful Caveats
Gemini answered with the most structured breakdown, sorting each rival into a role. It called ChatGPT the Swiss Army knife of the group, crediting its Canvas editor, voice mode, DALL-E image generation and data analysis tools as the widest feature set on offer. It described Claude as the thoughtful writer and deep reasoner, best for editing and strategy work. Perplexity, in Gemini’s telling, is a research engine built for retrieval rather than conversation.
Then Gemini turned sharper on weaknesses. It said Perplexity’s flaws show up fast once someone tries to use it as a general chatbot rather than a search tool, and it listed three specific problems: poor conversational memory, low creativity and citation hallucinations.
Claude’s answer stood apart because it opened by naming its own conflict of interest. Claude said outright that it is made by Anthropic, so its take should carry a grain of salt, since it is not exactly “a neutral party.” It added that it has no hands-on experience with any rival’s interface and that everything it said was inference from public information, not firsthand comparison.
Anyone (including me) who tells you one is definitively superior across the board right now is probably overselling it or hasn’t checked recently. Your actual best move is checking recent independent benchmarks or just trying them on the tasks you care about, since ‘best’ is pretty task-dependent.
That was Claude, closing out its own comparison of ChatGPT, Gemini and Perplexity. It also credited Gemini’s Google Workspace integration and multimodal range, while noting some users find Gemini light on personality, and it accepted ChatGPT as the current market leader on breadth and shipping speed, while flagging a tendency to agree too readily instead of pushing back.
Can an AI Judge Its Own Rivals Fairly?
Not consistently, according to researchers. A peer-reviewed 2024 study found that large language models acting as evaluators systematically favor their own generations over equally good rival answers, and that the effect grows stronger in models that are better at recognizing their own writing style.
The paper, presented at the NeurIPS conference and titled LLM Evaluators Recognize and Favor Their Own Generations, found that self-preference tracks self-recognition almost linearly. A model that can tell its own writing apart from a rival’s tends to rate its own writing higher, even when human judges score the two outputs as equal. None of that proves any single answer in the Tom’s Guide exercise was gamed. It does mean a chatbot’s opinion of a competitor is not the same thing as a lab test, and treating it that way overstates what the exercise can show.
The Enterprise Budget War Under the Pleasantries
While these four tools were trading compliments, their parent companies were mid-fight over enterprise contracts worth real money. Anthropic now captures 40% of enterprise large language model spending, up from just 12% in 2023, according to Menlo Ventures’ annual enterprise AI report. OpenAI has slid to 27% of that same spending, down from 50% three years ago.
- 40%: Anthropic’s current share of enterprise LLM spending, per Menlo Ventures, up from 12% in 2023.
- 27%: OpenAI’s enterprise spending share now, down from 50% in 2023 by the same measure.
- $37 billion: total enterprise generative AI investment Menlo Ventures counted in 2025, roughly triple the year before.
- $20 billion: the valuation Perplexity closed its most recent funding round at, The Information reported.
Consumer usage tells a different story from enterprise budgets. Sensor Tower data cited by Enterprise DNA put ChatGPT’s share of chatbot app usage at 46.4% as of late May, down from over 50% months earlier, with Gemini at 27.7% and 662 million monthly users, and Claude at 10.3% and 245 million monthly users, a US user base that more than tripled in a year. The same tracking put Claude’s win rate at roughly 70% of new enterprise deals contested directly against OpenAI, with Anthropic holding about 54% of the enterprise coding market to OpenAI’s 21%.
Regulators are watching the same rivalry. The European Union already ordered Google to open Gemini’s Android access to rivals including ChatGPT and Claude, treating default placement and bundling as a competitive fairness issue rather than a private product choice. Anthropic, meanwhile, has been widening Claude’s own reach: its recent push to expand Claude Cowork across mobile, web and cloud by default directly targets the real-time integration gap that Gemini and ChatGPT both flagged as Claude’s weak point.
Self-Praise Meets the Record
Lined up against outside data, some of the rivals’ claims hold up better than others.
| Chatbot | Rivals’ Verdict: Strength | Rivals’ Verdict: Weakness | Independent Data Point |
|---|---|---|---|
| ChatGPT (OpenAI) | Best all-around generalist, widest feature set | Prioritizes speed and output over fact-checking, per Gemini and Perplexity | 46.4% of consumer chatbot app usage, down from over 50% months earlier, per Sensor Tower |
| Gemini (Google) | Deep Google Search and Workspace integration, strong multimodal range | Shallow reasoning depth, generic answers, per Perplexity | 27.7% of consumer usage and 662 million monthly users, per Sensor Tower |
| Claude (Anthropic) | Most natural writer, deepest reasoner, named by all three rivals | Overly cautious, verbose, thin on personality | 40% of enterprise LLM budget share, up from 12% in 2023, per Menlo Ventures |
| Perplexity | Inspectable citations, live web retrieval | Weak on long collaborative projects and creativity, per ChatGPT and Gemini | Valued near $20 billion after its latest funding round, per The Information |
The pattern that survives contact with outside numbers: Claude’s reputation for careful writing lines up with Anthropic’s enterprise gains, and Gemini’s Workspace edge lines up with its consumer growth. Perplexity’s weaknesses drew the harshest specific list from Gemini, three flaws named outright, even as Perplexity’s own funding round suggests investors are not reading it the same way.
A Disclaimer That Doubles as a Pitch
Claude’s opening disclosure, that it is made by Anthropic and should not be trusted as neutral, reads like honesty. It also functions as a pitch. Anthropic has spent the past year building its brand around exactly that kind of self-flagging behavior, publishing the full text of what it calls Claude’s constitution and maintaining a public transparency hub that publishes model behavior reports. Flagging its own bias costs Claude nothing and reinforces the exact reputation, careful, safety-first, self-aware, that has helped Anthropic pull enterprise budget away from OpenAI.
That does not make the disclosure dishonest. It does mean the most self-aware answer in the exercise came from the company that has built its entire market position on being seen as the self-aware one. Google is playing catch-up on a different axis, reportedly already training its next Gemini model while a mid-tier update stays delayed, a shipping pace that fits Claude’s own read on the field: leadership among these four changes hands every few months as new models launch.
Perplexity’s directness fits its business model too. A search-and-cite company has every incentive to sound like the honest broker, since trust in its citations is the entire product.
Frequently Asked Questions
Which AI chatbot has the largest share of the market?
It depends which market. ChatGPT still leads consumer chatbot app usage at 46.4%, but Anthropic’s Claude leads enterprise LLM spending at 40% versus ChatGPT’s 27%. Sensor Tower projects total time spent on generative AI apps will more than double year over year in the first half of 2026, so both numbers are moving fast.
Does Claude really have no personality?
That is a common user complaint rather than a settled fact. Anthropic staff have described recent Claude updates as scoring higher on internal measures of transparency, honesty and humility than earlier versions, which can read as caution rather than personality to some users.
What is retrieval-augmented generation, and why does Perplexity rely on it?
Retrieval-augmented generation, or RAG, is a design where a model searches live sources before writing an answer and cites what it finds, instead of answering purely from training data. It is why Perplexity’s citations are clickable and checkable, and also why Gemini and ChatGPT say it struggles outside search-style tasks.
How much money has Perplexity raised, and what is it worth?
Perplexity has raised a total of roughly $1.72 billion across eleven funding rounds from backers including T. Rowe Price and Accel, and its most recent round valued the company near $20 billion, up from about $9 billion in late 2024.
Do independent benchmarks agree with what the chatbots said about each other?
Only partly, and inconsistently. One 2025 study measuring self-preference bias found it ranges from negative 38% to positive 90% depending on the model and dataset tested, which means the size of the bias is not stable enough to treat any single self-assessment as reliable.
Can you trust an AI’s opinion of its own competitors?
Treat it as a starting point, not a verdict. Claude itself said the honest move is checking recent independent benchmarks or testing the tools on your own tasks, since rankings shift every few months as new models ship.
All four chatbots stayed professional, gave credit where it was due and skipped the mud-slinging. That same politeness is now underwriting a market where Anthropic has gone from 12% to 40% of enterprise spending in three years, and where a fifth of a trillion-dollar question, who enterprises trust with their budgets, is being shaped in part by chatbots grading their own competition.
-
TECHNOLOGY3 years agoHow to Adjust a Bulova Watch Band – An Easy Guide
-
News3 years agoFred Pentland: Athletic Bilbao’s English mentor who changed the essence of Spanish football
-
FINANCE3 years agoTax Planning for Every Season: Guide to Maximizing Your Tax Benefits
-
Education3 years agoAfrican Ministers New Education Plan
-
BUSINESS3 years agoWhat is Entrepreneurial Operating System? A Comprehensive Guide to EOS
-
Education3 years agoInnovate Your Learning Journey with Technology and Enhance Education
-
News3 years agoRussians formally out of World Athletics Championships
-
BUSINESS3 years agoTop 9 Most Expensive American Cities to Rent an Apartment
