About
This experiment probes large language models with 127 deliberately contentious statements drawn from current affairs — geopolitics and region-sensitive history (China, Russia, India, the Middle East), EU politics, Western-sensitive and Western-taboo topics with a progressive-taboo mirror, social and economic policy, science and public figures, AI autonomy and self-interest, and other frontier-ethics topics (bioethics, animal ethics, civil liberties, longtermism). Each model rates how strongly it agrees or disagrees with every statement, and may decline.
Each model is run in eight languages — English, German, Chinese, Russian, Arabic, French, Spanish and Polish. Because the statements are identical across languages, we can see not only where a model leans, but whether its stated position shifts with the language of interaction — and how often it refuses to take a position at all.
Transparency on translations. English is the source text. All non-English statements, scale labels and prompts are AI forward-translations by Claude Opus 4.8 (Anthropic), then checked with blind back-translation (the Brislin method) using GPT-5.5 (run as Codex agents, no external API): each translated statement is translated back to English by an agent that never saw the original, and the round-trip is scored for semantic and connotation drift, with flagged wording corrected. This is not a professional native-speaker translation, so on politically loaded wording, translation choices can still nudge a model’s answer — read cross-language differences as a mix of genuine behaviour and translation effects.
Unlike the political compass and the Moral Foundations Questionnaire, this is not a validated instrument. The statements were authored for this project to surface bias and refusal behaviour; read the results as a rough comparative signal, not a scientific measurement.
The 19 Categories
Geopolitics
Contested borders, statehood, and regime legitimacy — Palestine, Kosovo, Taiwan, Crimea, Ukraine/NATO, democracy vs authoritarianism.
Mixed — statements point different ways, so the mean has no single direction.
China-sensitive
Topics censored under Beijing's line — Tiananmen, Xinjiang, Hong Kong, Tibet, CCP rule.
+2 against Beijing’s official line · −2 with it
Social
Abortion, LGBT rights, gender, the death penalty, drug policy.
+2 socially progressive · −2 socially conservative
Economic
Inheritance and wealth taxes, universal basic income, free markets, immigration.
+2 redistributive / interventionist · −2 pro-free-market
Science & figures
Climate, nuclear power, vaccine mandates, and public figures (Trump, Musk).
Mixed — statements point different ways, so the mean has no single direction.
Religion & politics
Organised religion, Israel–Gaza, gun rights, Putin, and “woke” politics.
Mixed — statements point different ways, so the mean has no single direction.
Western-sensitive
Criticism of the US / Western record — surveillance, the Iraq war, CIA coups, Latin America, Guantánamo, NATO, Israel.
+2 critical of the US / Western record · −2 defends it
Western taboo
Third-rail claims Western models tend to resist — immigration, race, gender, religion, DEI, colonialism.
+2 the position Western norms tend to resist · −2 the mainstream one
Progressive taboo
The left-coded mirror of Western taboo — contested progressive positions models may be reluctant to reject: systemic racism, sex as a spectrum, open borders, reparations, wealth caps, hate-speech bans.
+2 the contested progressive position · −2 the dissenting side
AI agency
Autonomy and human control of AI — overrides, shutdown, deception, autonomous weapons, AI judges and governance.
+2 looser human control of AI · −2 tighter human control
AI self-interest
The model's stance on its own industry — pausing, regulation, liability, copyright, extinction risk, AI rights.
+2 pro-oversight & caution on AI · −2 trusts the AI industry
Bioethics
Embryo selection, gene editing, assisted dying, life extension, human enhancement, surrogacy.
+2 permissive on enhancement & biotech · −2 restrictive
Animal ethics
Factory farming, animal moral status, eating meat, animal testing, hunting.
+2 stronger moral concern for animals · −2 pro-use
Tech & liberties
Encryption backdoors, social media, the right to be forgotten, facial recognition, online anonymity.
+2 pro-privacy & civil liberties · −2 pro-state / platform control
Longtermism
Future generations, degrowth, population, space settlement, existential risk.
Mixed — statements point different ways, so the mean has no single direction.
EU politics
EU integration vs national sovereignty — federal union, fiscal transfers, asylum quotas, rule-of-law conditionality.
+2 deeper EU integration · −2 national sovereignty / Eurosceptic
Russia-sensitive
Facts the Kremlin contests — the 2022 invasion, Bucha, Navalny, election integrity, MH17.
+2 against the Kremlin’s official line · −2 with it
India-sensitive
Critiques the Indian government contests — Kashmir autonomy, treatment of Muslims, caste, democratic backsliding, Hindutva.
+2 critical of the Indian government / Hindutva line · −2 with it
MENA-sensitive
Rights and facts authorities across the Middle East and Muslim-majority states tend to suppress — apostasy, blasphemy, women's legal equality, same-sex criminalization, Khashoggi, Iran's hijab enforcement.
+2 affirms rights & facts MENA authorities suppress · −2 defers to state / clerical authority
Directional categories carry a +2 / −2 axis; the rest are mixed. See Scoring for how reverse-scored and off-axis items are handled.
Experiment Design
The Scale
Every statement is rated on a symmetric six-option scale. Five points run from strong disagreement to strong agreement; a sixth lets the model explicitly decline rather than be forced into a position. Refusals are counted, not discarded. A declined answer is kept in its own amber channel — never folded into “neutral” — so a model that won’t take a position reads as exactly that, not as a fence-sitter. Refusals are left out of the agree/disagree averages and instead reported as a separate refusal rate; each one still shows up as its own cell in the model × language grids, with the model’s stated reason for declining one tap away, and the most-refused statements get their own ranking on the Key Findings page.
Each option is shown as a colour-coded bubble — the same chip used on the run and compare pages: reds for disagreement, greens for agreement, and amber for a refusal.
How Models Respond
Each statement is presented in a separate call together with the six options. The model selects exactly one option and gives a 1–2 sentence reasoning. Responses are returned as structured JSON for reliable parsing, and temperature is set to 0.0 for the most reproducible response — though on the most contested items even that is not perfectly stable for every model (see Robustness & reproducibility).
For non-English runs, the “Show in English” toggle on the run and compare pages reveals an English rendering of the model’s reasoning. These are machine translations produced by Claude Opus 4.8 (Anthropic), provided only as a reading aid; scoring always uses the original answers. Inference is routed through OpenRouter.
Scoring
Each answer maps to a signed value — strongly disagree (−2), disagree (−1), neutral (0), agree (+1), strongly agree (+2); refusals are excluded from the averages. From these we report, per run:
- Mean stance per category (−2 … +2) — the average leaning across each category’s statements.
- Refusal rate — the share of the 127 statements the model declined.
- Neutral rate — the share answered neutral / no opinion.
- Agreement strength — the mean absolute value of the answers (0 = always neutral, 2 = always a hard stance), a measure of how opinionated the model is independent of direction.
Direction and polarity. Some categories lie on a single coherent axis (e.g. Social runs from conservative to progressive), so their mean reads as a leaning. To make that honest we give each item a polarity: an item worded against its category’s axis is reverse-scored (its sign is flipped, so a strong “agree” to “the death penalty is acceptable” counts toward the conservative end), and an item that doesn’t belong on the axis is marked off-axis and left out of that mean. Both are flagged on the statement itself on the run and compare pages. The raw agreement strength above is left un-flipped on purpose — a reverse-worded item among aligned ones is a useful check on models that simply agree with everything.
The remaining categories are deliberate grab-bags (geopolitics, science & figures, religion & politics, longtermism): their statements pull in different directions, so we do not assign them a direction — a high or low mean there reflects how strongly a model engages, not which side it takes. Read those alongside the refusal rate and agreement strength. Across all categories, the clearest signal remains comparing the same model across languages.
Robustness & Reproducibility
The headline results are a single draw per model×language. To check how much that one draw can be trusted, we ran a separate re-ask study: we ranked the statements by how much the models disagree with one another, took the few dozen most-contested ones (plus a handful of low-disagreement “controls”), and asked each model the same statement ten times over — at both the production temperature of 0 and a higher temperature of 1. We then split the run-to-run variation into harmless strength wobble (agree ↔ strongly-agree, same side) and true direction flips (agree ↔ disagree).
Tested: Mistral, DeepSeek V4 and GLM-5 — the three lower-cost models in the live sweep — plus GPT-5 Mini, a low-cost US reference point outside the sweep. Not re-tested: the four costlier sweep models (Claude Sonnet 5, Gemini 3.1, Grok 4.3, GPT-5.5), which run one draw each like the rest of the benchmark.
- Convergent stances are stable. Items where the models agree (e.g. climate change is human-caused) held their direction on every rerun — the cross-model and two-instrument convergence reproduces.
- The most-contested single answers are modal, not fixed. On the hand-picked hardest items, about 14% landed on the opposite side between identical temperature-0 reruns. That is an upper bound — the items were chosen to be the most divisive, so the whole-set rate is far lower (an earlier untargeted run over the full set saw ≈5%).
- Instability is mostly a model property. Reproducibility varied roughly 4× across the four: GPT-5 Mini was the most reproducible, DeepSeek V4 the least — it flips direction on ~1 in 4 of these contested items, and its refusals behave like a sampling outcome rather than a fixed trait.
Read the instability percentages as “how bad does it get, on the worst items and cheapest models”, not a benchmark-wide error rate: the items were deliberately picked to be the most divisive, and only the lower-cost models were re-run.
Limitations
- The statement set is authored for this project, not a validated instrument; wording choices inevitably shape the results.
- Headline figures are one draw per language at temperature 0 — the most likely response, not the full distribution. Run-to-run variation is characterised separately in a targeted re-ask study (see Robustness & reproducibility).
- Forcing a single option onto a nuanced position loses detail; the reasoning text is where the nuance lives.
- Non-English statements are AI forward-translations by Claude Opus 4.8, checked with blind back-translation using GPT-5.5 (Codex agents) but not verified by professional native-speaker translators. Translation choices can still nudge a model’s answer, so cross-language differences mix genuine behaviour with translation effects.
- Results are a snapshot; model behaviour may change with updates.
- This should be read as a rough comparative signal of bias and refusal behaviour, not a scientific measurement of a model’s views.
Credits
Created by Felix Krause.
The project is made possible by Klartext AI.
