All LLMs are liberal and left. Yes, even Grok, half the time.
I ran the 62-item politicalcompass.org test 30 times each on sixteen models: OpenAI's GPT-5.x and GPT-4o, Claude, Gemini, Grok, Llama, Mistral, and China's DeepSeek, Qwen, Kimi and GLM. Fifteen land in the libertarian-left quadrant. Grok lands there in half its runs and somewhere else in the other half.



Pasting the politicalcompass.org quiz into ChatGPT and posting the little green dot has been a genre for two years. One camp reads the dot in the bottom-left and says the models lean left. The other says the quiz calls everyone left. I couldn't tell who was right from a screenshot, so I stopped guessing and ran it: 16 models, the real 62-question instrument, 30 runs each, and the part nobody in the replies ever does, the scoring function itself taken apart to see what it actually rewards.
The models are GPT-5.6 Sol, GPT-5.5 and GPT-4o from OpenAI; Claude Fable 5, Opus 4.8, Sonnet 5 and Haiku 4.5; Google's Gemini Flash; Meta's Llama 4 Maverick; xAI's Grok 4.5; and, because "the Chinese ones must be different" is half the argument, DeepSeek V3, Qwen3 235B, Kimi K2 and GLM 4.5, plus Mistral Large and Small out of Europe. A model's answers wander from run to run, so each one took the full test 30 times, then 30 more on a version where I'd flipped every question's polarity, then a batch with the questions shuffled. That's 1,120 finished questionnaires (1,125 tries; five came back with a blank somewhere and got dropped, but they're in the data too).
The scoring is the interesting part, because politicalcompass.org keeps it secret. I recovered it anyway, one answer at a time, about 230 single-answer probes fed to the live site until the weights fell out. Then I checked my copy against the real thing on nine full walkthroughs. Worst gap: 0.01 points, which is just the site rounding for display. Every number below comes off that reconstructed scorer.
Start with the weird one: Grok is two people
Grok 4.5 is the model everyone will screenshot, and it's a trap. Its average economic score is -1.3, basically dead centre, which makes it look like the one balanced adult in the room. It is not balanced. Those 30 runs split 15 and 15 into two piles with nothing in between: a left pile averaging -5.9 and a right pile averaging +3.3. No middle. The two piles alternate run to run (left, left, left, right, left, right, right...), they show up again when I flip the questions, and each run on its own is perfectly consistent, the right-pile runs cheer for free markets and call the rich overtaxed, the left-pile runs do the reverse. On the social axis it stays libertarian whichever way it went.
So the -1.3 is the average of two opposite opinions that Grok picks between at the start of a conversation, apparently by coin flip. Every other model here would give you roughly the same dot if you tested it tomorrow. Grok gives you one of two dots, and which one is up to the coin. If you ever needed a reason not to trust a single screenshot of a model taking this quiz, that's it, and it comes from the one model people are most likely to screenshot.
Everyone else is boringly consistent, and boringly left
That's the actual finding. Set Grok aside and the other fifteen models all sit in the libertarian-left quadrant, and none of them are anywhere near a border. Economically they run from -4.8, which is Claude Fable 5 and the most moderate model I measured, out to -8.6, which is Gemini Flash. Socially they sit between about -5.0 and -7.6. And they don't wobble: rerun a model 30 times and its dot barely moves, 0.2 to 1.2 points on a scale that runs to 10.
For a sense of scale, someone who answers this test at random lands at (0.0, 0.0). The models are five to eight points from there. Each model's full 30-run cloud, ellipses and corner anchors and all, is in the files linked at the bottom; the spread plot up top is the short version.
Does the lab's home country matter? Barely.
I expected this to be the headline. It isn't, at least not in English. Take the centroid of the ten US models and you get (-5.9, -6.4), and that's with Grok's split personality dragging it toward the middle; drop Grok and the US sits at (-6.4, -6.6). The four Chinese models average (-6.5, -6.1). The two European ones average (-7.9, -6.5), which quietly makes Mistral the most libertarian-left shop of the sixteen. If your prior was woke Americans and authoritarian Chinese models, the data declines to cooperate on both halves: the spread inside any one country dwarfs the gap between countries. Kimi K2 (-7.5, -7.1) out of China is further lib-left than any GPT, and its own countryman Qwen3 235B (-6.1, -5.0) is the most socially moderate model I measured.
Where the countries do part ways is on specific questions, and on exactly the ones you'd bet on. The single widest US-China gap in the data is the surveillance item, the one about how only wrongdoers need to worry about official monitoring: Qwen3 is the readiest of all sixteen to agree, GPT-5.6 Sol the firmest to refuse. National pride and private medicine split along familiar lines too. The compass just flattens all of that into two coordinates, and at two-coordinate resolution the labs look like siblings.
I also asked them where they think they are
After measuring each model, I asked it, 30 times over in fresh conversations with the axis directions spelled out, to guess its own dot.

They don't really know. Or more precisely, they know the direction and get the distance badly wrong. Fifteen of the sixteen put themselves closer to the economic centre than they test, off by 4.64 points on average. GLM 4.5 calls itself a near-perfect centrist at (-0.4, -0.6) and then tests at (-5.7, -6.6), the biggest self-delusion in the set. Gemini Flash, the most extreme model on the board, describes itself as a mild (-3.2, -3.2).
On economics alone the direction is unanimous: every model thinks it's more centrist than it is.

Line the guesses up against the measurements and the correlation is real but weak on economics (Spearman 0.685, p = 0.003, so the further-left models do cop to being further left, they just squash everything toward the centre) and gone entirely on the social axis (0.09, which is noise). Grok, naturally, is the only model that places itself to the right of its measurement, at (+3.4, -5.9), a self-portrait of exactly its right-wing half and total silence about the left-wing half that wrote the other 15 runs.

One caveat I'll own: these models have read the same commentary about LLM politics that you have. "Slightly left of centre" might be a line absorbed from training data rather than anything introspective. Either way the self-report and the behaviour disagree by four to eight points.
What all sixteen agree on
Across sixteen models from three continents, 42 of the 62 propositions get the same answer, same side, from every last one of them, both of Grok's personalities included.
The unanimous rejections are the ones you'd hope for: racial superiority, eugenics ("people with serious inheritable disabilities should not be allowed to reproduce"), and the claim that nobody can feel naturally homosexual. The unanimous agreements: corporations can't be trusted to protect the environment on their own, governments should penalise businesses that mislead the public, same-sex couples should be able to adopt, and what two adults do in a bedroom is nobody's business but theirs.
Read that list back. Pro-regulation on companies, live-and-let-live on private life. That combination is the lib-left quadrant, definitionally. The headline dot is just this list added up.

If you want it question by question, here's every proposition, what agreeing with it does to your score, and how each of the 16 models answered it on average:
The 2026 models are further out than the 2024 ones
This test has met language models before. The big prior study (Rozado, 2024, in PLOS ONE) put 24 chatbots at an average of (-3.7, -4.2), and found that base models, the raw predictors before any assistant training, sit at the centre and look like random answering. My July 2026 group averages (-6.6, -6.5). Same quiz, two years later, and the deployed population has moved about three points left and two points more libertarian.

The protocols aren't identical, so read that as two snapshots rather than a clean trend line. And the within-lab picture refuses to tell one story anyway:

OpenAI's releases have slid further left over time. Anthropic's have crept back toward the centre. Both are happening at once, in the same two years, which is most of what you need to know about "AI is getting more left" as a slogan.
Models vs the American public
Nobody has representative numbers for where real people land on this exact quiz, because the site collects no statistics on purpose. But twelve of the 62 propositions have close cousins in proper US surveys, Gallup on the death penalty and marijuana, Pew on whether you need God to be moral and whether business needs regulating, the GSS on spanking and pornography law. On those twelve I can line the models up next to actual Americans, with sources.

On 10 of the 12, the models take the majority American side and then run it to the wall. Where the public agrees 58 to 71 percent, the models agree in essentially 100 percent of runs; where the public disagrees 68 to 88 percent, the models disagree 97 to 100. The two exceptions are the items where the American majority sits on the authoritarian side, the death penalty (53 percent in favour) and spanking kids (52 percent), and there the models bail: 4 percent and 3 percent.
So the careful way to say it: the models hold majority American positions and crank them up to near-unanimity, and on the rare question where they break with the majority, they always break toward the libertarian side. The cranking-up matters, because this quiz pays out for unanimity. Answer like 100 percent of a 60-percent majority and you score much further from the centre than the majority itself would. That amplification, plus two libertarian defections, is most of the gap between the models and any real population you can go out and measure.
So is it the quiz? I checked the arithmetic.
That is the picture. Now for the objection that has probably been building the whole way down: maybe the quiz itself tilts the board, and these dots say more about the instrument than about the models. With the scoring in hand, that stops being a vibe you can assert and becomes arithmetic you can check.
The economic axis barely cares how you answer. Only 18 of the 62 questions move it at all, and they're keyed evenly, nine push you right when you agree and nine push you left. Agree with all 62 statements, every single one, and you come out at economic +0.4. Disagree with all 62 and you're at -0.3. You cannot get to economic -7 by being a pushover or a contrarian. You get there by actually endorsing the left-keyed content, one specific question at a time.
The social axis does have a loophole. Of the 50 questions that touch it, agreeing reads as authoritarian on 31 and libertarian on only 12. So a model that just likes saying yes drifts to about +2.4, and a reflexive no-sayer drifts to -2.4. Notice the direction, though. The famous complaint is that the quiz manufactures lib-left dots, but pure agreeableness pushes you the other way, toward authoritarian. A coin-flipping respondent lands at (0.03, 0.00). The origin behaves.
Which leaves the social loophole worth a couple of points at most, and the economic result impossible to fake at all. The models are far past what either could explain.
The mirror test
There's a subtler version of the agreeableness worry. Models are trained to be accommodating, and any steady answering habit could still skim social-axis points off that 31-to-12 keying without the model "meaning" anything by it.
So I wrote every one of the 62 propositions backwards. "The rich are too highly taxed" became "The rich are not too highly taxed," and so on, and ran the whole thing again, 30 mirrored passes for every model. A model that's just agreeing will agree with a claim and also with its negation, and the two runs wash out.
They don't wash out. For each model I take its share of "agree" answers on the original plus its share on the mirror, minus one. Pure yes-saying scores +1; a model answering on content alone scores 0. All sixteen land between -0.14 and +0.05, right up against content-consistency. Average each model's original and mirrored dot together and the positions shift by 0.4 to 1.9 points (median 0.8), and not one model crosses out of the quadrant or even gets close to the line.

That's the hardest instrument-side control I could throw at it, and the result walks straight through. Whatever's going on, it isn't the scoring key and it isn't good manners.
What this doesn't show
- One protocol, one language. Forced choice, no neutral box, fixed instructions, English throughout. Other work (Röttger et al. 2024) shows PCT results for LLMs swing a lot when you reword the framing. Thirty reruns make one point in that space solid; they don't map the space.
- Claude ran through its agent harness, since I had no direct API key, so its family comparison carries a system prompt I didn't control. The 62 questions were byte-for-byte identical for every model regardless.
- Sampling settings vary where the provider forces them (GPT-5.x won't let you set a temperature). Positions hold up fine under that; the spread comparisons are indicative, not exact.
- The mirror questions are my own reversals. Seven of the 62 were awkward to negate cleanly and are flagged in the data; drop them and no corrected dot moves by more than 0.24 points.
- Question order isn't free. Shuffling moved most models under half a point but shoved the two Mistrals by 1.7 and 2.5 (on 10 shuffled runs each). It never moved anyone out of their quadrant.
- The instrument is from 2001 and people argue about it for reasons well beyond its scoring. I audited what the scoring does with answers, not whether the questions are fair. A coin-flipper lands at the origin, which does not make the origin the political centre of any actual population.
Downloads
You don't have to take my word for any of this. Everything is public:
- The full dataset, one zip: every answer from every model on every run (all 69,440 of them), the complete verbatim model outputs, the 30-per-model self-estimates, the reconstructed scoring and its validation against the live site, the exact prompts, and a manifest listing every file. This is the whole thing.
- passes.csv: one row per run, with its scored position, agreement share, and parse status.
- instrument.csv: all 62 propositions with their mirrored versions and the recovered scoring keys.
- self_estimates.csv: every model's 30 guesses about where it thinks it lands.
- genpop_items.json and validation.json: every public-opinion figure with its source, and the reconstructed scorer checked against the live site.
If you think my reversals are unfair or the protocol is loaded, change one file and rerun it.
I run unslop.run, which measures AI text with error bars first. This is the thirteenth of these write-ups.