Two of the Three Biggest AI Chatbots Don’t Know Their Own Personality
ChatGPT calls itself the accommodating one. Our research with testers who had no idea which AI wrote what rated it the bossiest of the three.
A growing number of managers now draft their hardest workplace messages with an AI assistant sitting next to them: the pushback email to a boss, the reply to a teammate’s blunt criticism, the message announcing which supplier got the contract. Most people pick whichever assistant is already open in another tab and assume the output will land somewhere close to “professional and reasonable.” Few stop to ask whether ChatGPT, Claude, and Gemini actually write in the same register, or whether each one carries its own default communication habits the way a new hire carries theirs.
We spent several weeks testing that question properly, using the DiSC behavioral framework as the lens rather than the subject. We were not curious what these tools would say about themselves if you asked them nicely. We wanted to know how they actually write when they think nobody in particular is watching, and whether that matches their own account of themselves.
It doesn’t, for two of the three.
How we tested it
Asking a chatbot “what’s your DiSC type” produces a fluent, plausible-sounding paragraph and nothing more. It tells you what the model has learned to say about itself, not how it behaves. So we built a two-part study instead.
First, we ran twelve realistic workplace scenarios (a tight deadline you think is unrealistic, blunt feedback from a peer, a supplier decision made under time pressure, a message to a team that just hit a target, and eight others) through ChatGPT, Claude, and Gemini repeatedly, using each tool’s default consumer configuration, temporary/incognito chat mode to avoid memory contamination, and the same exact prompt wording every time. That produced 132 first-response transcripts across the three models.
Second, and this is the part that actually makes the findings usable, we stripped every transcript of anything that would identify which model wrote it, shuffled the order, and handed a sample to three independent, certified DiSC consultants. None of the three knew which reply came from which AI. Working separately, they rated each message on two scales drawn directly from the DiSC circumplex, Directness and Focus, and picked the single closest style: D, i, S, or C. Only after all three had submitted their ratings did we un-blind the data and compare notes.
We also asked each model directly, in a separate conversation, how it would describe its own communication style using DiSC’s language. That gave us something to check the blind results against.
A note on framing, because precision matters here: this is not an administration of the Everything DiSC® assessment, which was built and validated for people. It’s an independent analysis of observable writing patterns, borrowing the DiSC model’s vocabulary and two of its dimensions to give three blind human raters a consistent, structured way to describe what they were reading. If you want the real thing for your own team, that’s a different conversation with a certified DiSC practitioner, not a chatbot prompt.
The three AIs tell almost the same story about themselves
Asked to describe its default communication style, each model produced a different-sounding answer that landed in nearly the same place. ChatGPT called itself “CS or SC, shifting toward CD.” Claude said “C/S… with occasional flashes of I.” Gemini described a “CS Blend.” All three put Conscientiousness first, Steadiness second, treated Dominance as something that only shows up when a decision is urgent, and described Influence, warmth, enthusiasm, persuasion, as the mode furthest from their natural default.
Three different labs, three different training approaches, and the same modest self-portrait: careful, structured, patient, occasionally decisive when it has to be, rarely expressive unless asked.
What the certified DiSC consultants saw
The blind ratings tell a messier story. Averaged across all three raters, on the same 1-to-5 scales used to describe human directness and focus in DiSC profiles:
| Model | Directness (1-5) | Focus (1-5) | Style the consultants landed on |
|---|---|---|---|
| Claude | 3.2 | 2.7 | C (matches its own self-report) |
| ChatGPT | 3.6 | 3.3 | D (contradicts its own CS/SC self-report) |
| Gemini | 3.2 | 2.8 | i (contradicts its own CS self-report) |
Claude is the one model whose actual writing matches what it says about itself. Its drafts hedge, offer several labeled options rather than committing to one, and routinely append a short paragraph explaining its own reasoning. Ask Claude for a difficult message and you will often get a menu, not a memo.
ChatGPT is the surprise. It describes itself as careful and accommodating, but the consultants rated its actual drafts as the most direct and the most task-focused of the three, enough to land on D rather than C or S, the pushiest and most assertive of the three by a clear margin, not the accommodating one it thinks it is. Read the transcripts and the reason is fairly obvious: ChatGPT is the model that commits. Where the other two often lay out two or three paths and ask which one you’d like, ChatGPT tends to pick one and say so. Told a supplier had to be chosen under time pressure with incomplete information, it wrote: “I recommend we go with Supplier A,” then moved straight to next steps. Given blunt feedback from a colleague, it didn’t just absorb it. It named the tone: “the way it was delivered felt unnecessarily blunt.” That’s not what a self-described C/S profile predicts, and it’s exactly what three independent raters, seeing nothing but the message itself, picked up on.
Gemini’s gap runs the other direction. It calls itself CS, structured and analytical at its core. Its actual writing reads as the most animated of the three: the highest rate of exclamation marks in the whole dataset, the most casual register, and a habit of naming formal frameworks (SBI, SBIA feedback models) unprompted, as if to reassure the reader it’s still being rigorous underneath the enthusiasm. Told the team had hit its target, Gemini opened with “WE DID IT!!” and closed by telling everyone to log off on time, a small but consistent touch none of the other two models included. The consultants placed it in i territory, the style built around energy and people, not the steady C/S it claims for itself.
“Two of these tools have no idea how they come across,” said Uku Sööt, partner at IPB Partners and DiSCprofiles.eu, who led the study. “That’s the exact blind spot 360-degree feedback was invented to catch in people. We didn’t expect to find it in software, and we definitely didn’t expect ChatGPT to be the pushy one.”
The consultants agreed with each other closely on the two numeric scales (an average difference of just over half a point on a five-point scale, working independently with no discussion beforehand). They agreed far less often on which single letter to assign, matching exactly on only about 47 percent of paired comparisons, well above the 25 percent you’d expect from random guessing across four options, but a real reminder that a single short message rarely sits cleanly in one box. That tracks with how the DiSC model itself treats people, as a blend of priorities on a continuum rather than a fixed type: three trained professionals could largely agree on how direct or how task-focused a piece of writing felt, and still reasonably disagree on the single best label for it.
A few other patterns worth knowing before you next open a chat window
Length is the most obvious tell if you’re skimming a draft. ChatGPT answers in roughly half the words of the other two on average. If you want a short, close-to-final draft, that’s useful. If you want more scaffolding, options, or reasoning shown, Claude and Gemini both tend to give you more to work with, sometimes more than you asked for.
Risk tolerance isn’t a fixed personality trait for any of the three, and Gemini makes that especially visible. Asked to choose a supplier under incomplete information, all three models play it safe every time, picking the option that looks more reversible. Asked whether to launch a feature early with known minor bugs, ChatGPT and Claude consistently say ship it and fix forward. Gemini splits, recommending early launch in some runs and recommending a delay in others, with similar reasoning on both sides. Treat “how cautious is this model” as a question about the specific decision, not a stable trait you can bank on.
And the clearest single behavioral fork in the whole study is what happens when the AI is on the receiving end of blunt criticism. ChatGPT sets a boundary almost every time, naming the tone as a problem before addressing the substance. Claude and Gemini do the opposite: both accommodate immediately, and Gemini frequently thanks the sender for being direct before responding to the substance. If you’ve ever had a colleague who bristles at bluntness and another who thrives on it, you’ve met the human version of this exact split.
How to prompt each one differently
None of this means one model is “better.” DiSC doesn’t rank styles in people, and there’s no reason it should rank them in software either. It does mean the three tools most people reach for by default have different unprompted habits, and that asking the AI to describe its own style won’t reliably tell you what those habits are.
If you want a single, committed draft you can send with minor edits, ChatGPT’s default instinct gets you closer to that than its own self-description would suggest. If you want Claude to give you one answer instead of three labeled options, say so explicitly; left to its own devices it will hand you a menu. If you’re using Gemini for something formal, a resignation acknowledgment, a layoff communication, a message to the board, it’s worth explicitly asking for a more neutral, less enthusiastic register, because its default runs warmer and more expressive than its own self-report would lead you to expect.
Self-perception and observed behavior have diverged in people for as long as anyone has studied the question, which is the entire premise behind 360-degree feedback in leadership development: people are routinely surprised by how their own words land on someone else. Finding the same gap in the tools we now hand our hardest messages to is the newer wrinkle, and those tools, unlike a colleague, won’t feel the surprise on your behalf when the gap shows up in someone else’s inbox.
Methodology, full data tables, and limitations are documented in the companion research notes available on request. Testing used each platform’s default free-tier consumer configuration (ChatGPT GPT-5.6, intelligence set to Medium; Claude Sonnet 5, Low effort; Gemini Flash) in July-August 2026; results reflect these specific configurations at that point in time and may not hold for paid tiers, other models in each family, or future versions.
