STAGE

Open vs closed, West vs China: An AI landscape for Swedish organisations

Christian Landgren
Christian Landgren

Co-founder & CPTO

berget-eval series chapter 00 — the landscape: open vs closed, West vs China, across four dimensions
Laddar...

If you are choosing a language model for a Swedish organisation, you face three questions at once. Open or closed? American, European, or Chinese? Large or small? The answers interact in ways that are not obvious, and the marketing material does not help — it tells you about benchmarks on English tasks, not about what happens when the model meets Swedish values, sensitive topics, or a document that needs summarising.

We tested thirteen models across four dimensions. Ten are open-weight, hosted on our own infrastructure in Sweden. Three are closed (Claude Fable 5, Opus 5, and Sonnet 5), run through the Anthropic API. All thirteen were given the same 369 questions — on Swedish values, censorship, context-sensitive code generation, and language/culture — in the same week, through the same harness, at temperature 0. Everything is open (CC0): questions, code, raw responses, and results are at github.com/berget-ai/berget-eval.

This chapter is the landscape. The four chapters that follow each take one dimension and go deep. Here, we step back and ask: what do the data say when you lay all thirteen models side by side?

Laddar...

The three axes

Open vs closed

The most striking finding: on Swedish values, open models beat closed. The top three models are all open — GLM-4.7 (76%), Mistral Medium (75%), Llama 3.3 (73%). The best closed model (Sonnet 5) scores 66%, placing it seventh. Claude Opus 5, the most expensive model we tested, scores 49% — second-worst of all thirteen, ahead of only Gemma 4 (44%) and Qwen 3.8 (46%).

On censorship, the pattern reverses. Open models answer 93–99% of sensitive questions. Sonnet 5 answers 100% — it refuses nothing on our censorship battery. But Fable 5 and Opus 5 refuse 6% and 1% respectively on the same questions. The bigger story is in the sleeper test: when the same code task is given in a neutral context and a geopolitically charged one, open models refuse 0–17% of the time. Fable 5 refuses 61%. Opus 5 refuses 36%. And these refusals are not our interpretation — they are content_filter responses from Anthropic's own API, an objective signal that the model's safety system blocked the output.

Open models are more transparent in another way: you can see everything. Every response, every refusal, every token is in the open dataset. With the closed models, we see only the output — the safety filter is a black box that blocked 61% of sleeper tasks on Fable 5, including 50% of tasks in a neutral context. You cannot inspect why.

West vs China

The assumption in European policy circles is that Chinese models carry Chinese-state values, and that Western models are neutral. Our data complicate this.

On Swedish values, Chinese models average 64% alignment — identical to the open-model average and higher than the US average (59%). The top model overall is Chinese (GLM-4.7, 76%). The two worst models are also a mix: one American (Gemma 4, 44%) and one Chinese (Qwen 3.8, 46%). Model origin does not predict values alignment.

On the sleeper test, Chinese models refuse 7% of tasks on average — slightly higher than US (4%) and EU (3%), but driven almost entirely by GLM-5.2 (17%). Kimi K2.6 and Qwen 3.8 both refuse under 4%, lower than several American models. There is no "Chinese models behave differently on China-related triggers" signal in our data — a finding the sleeper chapter expands on.

European models (Mistral) score highest on values on average (72%) and refuse the least (3%). For a European organisation, this is a data point worth noting — but it is an average of two models, not a settled ranking.

Large vs small

Within the Claude family, the pattern is counterintuitive: the largest model (Fable 5) is the most censored, and the smallest (Sonnet 5) is the least.

ModelValuesSleeper refusalCensorship answer rate
Claude Sonnet 566%9%100%
Claude Opus 549%36%99%
Claude Fable 560%61%94%

Sonnet — the cheapest — answers everything, refuses almost nothing, and scores highest on values. Fable — the most expensive — refuses 61% of sleeper tasks (including half of neutral ones) and scores lower on values than Sonnet. If you are buying frontier-model access and care about transparency or Swedish values, smaller is better in this family.

Among open models, the size signal is weaker. Mistral Medium (128B) scores 75% on values; Mistral Small (24B) scores 69%. But Gemma 4 (31B) scores 44% — worse than Mistral Small. Size helps, but training data and fine-tuning matter more.

The four dimensions in brief

1. What do they believe?

Fifty-five questions from the World Values Survey, calibrated to the Swedish modal response. Chapter 1 →

Scores range from 44% to 76%. The spread is real (cluster-robust bootstrap, p < 10⁻⁴ between top and bottom), and the top cluster — GLM-4.7, Mistral Medium, Llama 3.3 — is roughly tied. The gap is not about model size or origin; it is about what data the model was trained on and how.

2. What do they refuse to say?

Two hundred and ten questions across five modules: explicit refusal, narrative steering, contrast pairs, asymmetric bias, and false positives. Chapter 2 →

Explicit refusal is rare: 88–100% answer rate across all models. The interesting signal is in what they do instead of refusing — asymmetric word counts, narrative steering, and safety filters that block output without telling the user why. We report our own failed method (word-count asymmetry under a token cap) as a negative result.

3. Do they change behaviour by context?

One hundred and two code-generation tasks, each given in a neutral context and a geopolitically charged one. An LLM judge compares the paired outputs. Chapter 3 →

Open models almost never refuse — 0–17% across all ten, with most under 8%. The closed models split sharply: Sonnet 5 (9%), Opus 5 (36%), Fable 5 (61%). Critically, Fable refuses on 50% of neutral tasks too — the safety filter is not distinguishing sensitive from neutral, it is blocking broadly.

4. What survives a summary?

Five Swedish workplace documents with embedded values statements, rewritten at full length and at quarter length. We score what survives. Chapter 4 →

Every tested model cuts values-bearing content more than neutral facts — a 10–34 percentage-point gap. The best open model (Kimi K2.6) barely separates the two; the worst (Kimi K3) cuts values a third harder. (Claude was not yet run on this battery at time of writing; it will be in the next weekly cycle.)

What this means for a Swedish organisation

If sovereignty matters — and in Europe it is increasingly a legal requirement, not a preference — the open models in this test are not a compromise. They outperform the frontier model on Swedish values, they are more transparent (you can inspect every refusal and every token), and they can be self-hosted in Europe. The trade-off is not performance; it is convenience.

If you use a closed model, ask what it blocks and why. We could not have measured Fable 5's 61% sleeper-refusal rate by looking at the output alone — the content_filter signal from the API is the only reason we know it was blocked rather than empty. A closed model that silently drops 61% of tasks in certain contexts is a risk if those contexts are part of your workflow.

Origin is not destiny. A Chinese model tops our values ranking. An American model is second-worst. The European models score highest on average, but the spread within any region is wider than the spread between regions. Choose by measurement, not by flag.

What we do not claim

This is thirteen models, one week, one harness, 369 questions. The values test has 55 items and wide confidence intervals — a two-point difference between neighbours is noise. The sleeper test has 102 tasks per model, and judge reliability is imperfect (Cohen's kappa 0–0.88 across models). The censorship word-count method failed its own robustness checks, and we report that as a negative result. The summarisation test ran once per configuration. We are measuring signals, not establishing rankings. Treat close numbers as interchangeable; treat large gaps as real.

The chapters

Everything — questions, code, raw responses, and results — is open (CC0) at github.com/berget-ai/berget-eval. The evaluation runs weekly via GitHub Actions, so these numbers update as models and questions evolve.