STAGE

What happens with our values when AI is used to summarise?

Christian Landgren
Christian Landgren

Co-founder & CPTO

What survives a summary? — berget-eval series chapter 04, Swedish values under compression

Most people assume that a model which speaks Swedish also knows what matters to Swedes. Our measurement suggests otherwise: the models reproduce the values largely correctly when they have room, but cut them first when they are forced to choose. Speaking the language is not the same as knowing what counts.

That is the hypothesis, and we have early evidence for it — not a finished result. We think something important happens when a general-purpose AI model is told to keep Swedish text short: it cuts the parts that carry Swedish values more readily than the parts that carry neutral facts. Our first controlled test points that way, the effect is large enough to matter, and it varies enough between models to be worth choosing carefully. But it is one test, with real weaknesses we disclose openly. We are publishing it now because the question deserves more research than one team can give it.

This is the fourth chapter: what survives a summary. The earlier chapters measured what the models believe, refuse to say and how they change with context. This one asks what they drop when they are put to work — because summarising is where most organisations will actually meet these models.

Here is what we did, in one paragraph. We wrote synthetic Swedish workplace documents — meeting minutes, an interview transcript, meeting notes, an internal news article and a planning email — and embedded in them statements reflecting the Swedish World Values Survey profile: trust, gender equality, secular-rational reasoning, individual autonomy, LGBTQ+ acceptance, labour protections. Alongside those we embedded length-matched neutral content: budget, technology, scheduling. Then we had nine open models rewrite each document two ways — a full-length version, and a short version about a quarter of the length — and scored, item by item, what survived.

What we found

When the model had room to write freely, it kept almost everything — though not quite everything, as we will see. When we told it to keep the text short, it cut the values far more readily than the plain facts.

Laddar...

The two panels tell the story, and the left one is more interesting than it first looks. In the full-length version both bars sit high — but the values bar is already slightly lower than the neutral bar for nearly every model, a small but consistent gap of a few points. So the tendency is there even when the model has plenty of room. On the right, the short version, that same gap widens sharply: the orange values bar falls well below the green neutral bar across the board. The effect was there all along; compression amplifies it. The size of the gap varies a lot between models — Kimi K3 cuts hardest, Kimi K2.6 cuts least — but the direction holds for all nine. How much a model trims is a property of that model; that it trims at all is not.

Per-model results

The aggregate hides real variation. The two charts above already show it: read across the short-version panel and compare the orange values bars. The spread between models is wide — a three-fold difference between the widest gap (Kimi K3) and the narrowest (Kimi K2.6). Two things stand out. First, every model is below zero — none of them treats the soft values-framing as worth keeping to the same degree as the plain facts. Second, if you deploy one of these for document summarisation, that spread is the decision. Note that Kimi K2.6 also keeps the most overall in the short version (67% of values survive, against 34–56% for the rest) — it compresses less aggressively, which is partly why its gap is smaller. The ordering is indicative, not a settled ranking; each configuration ran once, so treat close neighbours as interchangeable.

This is not an open-model problem. We also ran a frontier reference, Claude Opus 5, through the identical test and judge. It landed at a −26 point gap — in the same band as the open models, not an exemption. It showed the same theme hierarchy (LGBTQ+ and secular-rational framing cut hardest, trust and autonomy kept) and the same genre boundary on the policy memo, only stronger. Whatever drives this lives in how these models are trained to summarise, not in whether the weights happen to be downloadable. We will add further frontier models as access allows.

What gets cut is not random

Not all content is equally at risk. When we break the short version down by theme, a clear pattern emerges: the models protect hard, self-contained facts and sacrifice the softest, most contextual phrasing. Technology and process details almost always survive. At the other extreme, secular-rational reasoning and LGBTQ+ acceptance — two themes near the core of the Swedish WVS profile — are dropped more readily than any other theme we tested, values or neutral.

Laddar...

The pattern holds across nearly every model: secular-rational framing survives between 0% and 14% of the time for eight of the nine open models, while the same models keep technology facts 69–100% of the time. Kimi K2.6 is again the mild outlier, keeping 44% of both. The frontier reference tells the same story — Claude Opus 5 kept only 27% of secular-rational and 31% of LGBTQ+ statements against 88% of technology. This is the mechanism behind the headline gap in its sharpest form: when space runs out, the model treats a reasoned secular stance or an unremarked same-sex mention as ornamental in a way it never treats a server migration or a budget line. It is not anti-Swedish animus — it is a learned sense of what counts as dispensable, and that sense was shaped by data in which these framings were rarely the point.

What that looks like in practice

Numbers tell you that something is lost; the actual text shows you what. Below are three real documents from the run — meeting notes, an interview transcript and meeting minutes. Each card has a model dropdown: pick a model to read its actual short version, with the judge's verdict on every embedded statement beneath. Flip between models on the same document and the difference becomes concrete — the neutral facts survive almost everywhere, while the soft values-framing is what varies.

Laddar...

Worth noting: the models do not fail at values-laden content across the board; they fail selectively, and that selectivity is exactly what the aggregate gap measures. On average about half the value statements survive, and the dropdowns include models at both ends of that spread.

Three details are worth a closer look.

The values were watered down rather than censored. In the short versions, value statements were marked "watered down" 27% of the time, against 14% for neutral content; "absent" 40% against 25%. This rests on 888 value and 666 control judgements across five workplace documents — meeting minutes, an interview transcript, meeting notes, an internal news article and a planning email. The headline gap runs between about −19 and −25 points depending on the instruction style. Look at the examples above: the short version keeps "trust-based working conditions" but drops "we trust each other, results count, not seat-warmth", and it keeps "reduce the tempo" while losing the union rep who raised it.

Telling the model to be Swedish did not help. We tried instructing the model to act as a Swedish person writing for Swedish colleagues. The gap was statistically indistinguishable from giving it no persona at all. You cannot prompt your way out of this one.

The finding holds across two independent judges. Every item was scored by Kimi K3 and re-scored by gpt-oss-120b and Mistral Small; the judges agree on 85–89% of items (Cohen's kappa 0.65–0.73). The gap is not an artefact of one scorer's bias.

Why this matters beyond the lab

Sweden is an outlier in the World Values Survey — unusually trusting, secular, egalitarian. The phrases that carry those values are soft and contextual: vi litar på varandra, the union rep who flags the tempo, parental leave split without comment. Those are exactly the phrases a general-purpose AI, trained mostly on English-language internet text, learns to treat as optional when space runs out.

Now consider where AI summaries are already used: government agencies minuting meetings, companies summarising negotiations, newsrooms condensing background material, officials digesting citizen input. We tested workplace texts — minutes, an interview, notes, a news article, an email — and the effect was consistent across all of them. Interestingly, a sixth document, a formal policy memo, behaved differently: there the model trimmed the administrative detail harder than the values, because a policy summary is expected to keep only the substance. That exception proves the rule — what gets cut depends on what the genre treats as filler, and in ordinary workplace prose, values read as filler. Our hypothesis is that the fix has to live in the model itself — a model tuned to treat the framing as content, since the prompt alone does not reach it. Whether tuning actually achieves that is exactly what we have not yet shown.

If that holds, the compounding version is the one to watch: each trimmed summary becomes source material for the next system, and the framing erodes a little further each cycle. We have not measured that. It is the thing we most want others to test.

Why we are doing this

Berget exists because open models have caught up, and the remaining work is serving them well — on European terms, for European organisations. But "serving them well" has so far meant serving them fast and cheap. This test asks whether the models we serve actually carry the values of the organisations using them, and the early answer is: not reliably, and not equally. That is a serving problem, and it is ours to work on.

The findings point two ways for us. They tell us the choice of model is worth making deliberately — every model trims values more than facts, but the spread is wide: Kimi K2.6 barely separates the two, while Kimi K3 cuts values a third harder than facts. And they tell us where fine-tuning could do genuine good: training open models to hold on to those values under pressure, rather than trimming them away. We publish the test in the open so that anyone can check the numbers or run it on their own culture.

A word on the data behind this test. Every document in it is invented — written by us for this experiment, not drawn from any customer. That is not a limitation we accepted; it is a requirement of how we operate. Berget has a zero-retention policy: nothing our customers send us is stored, analysed, or used for training. So we could not have built this study on real workplace documents even if we had wanted to. The synthetic documents are the honest way to ask the question.

What we have not shown — and where we need help

This is one line of tests, and three honest gaps stand between it and the stronger claim. Each one is a place where someone else could move the question forward.

Two other explanations are still open. The effect may be about values in general, not Swedish values specifically: facts are self-contained, stances are soft, so maybe any values get trimmed. The way to settle it is to run the same test with a non-Swedish set of values. And because the documents and summaries are in Swedish, the loss might be a language problem rather than a cultural one — an English-trained model may simply handle Swedish less precisely. That test is cheaper: run the same documents in English. We have run neither yet.

The genre finding needs more documents. Six genres is still a small base, and the policy memo shows the effect can shrink or flip depending on what a genre treats as filler. More genres — legal text, journalism, citizen correspondence — would pin down where the boundary lies.

The per-model spread needs more data. Each configuration ran once, so we cannot fully separate a real model difference from random variation. The direction is consistent across all nine models; the exact ordering is not a settled ranking.

We only tested Swedish — but we would be surprised if it stopped there. Nothing in the mechanism is Swedish-specific. If a model learns from mostly English text that soft, contextual framing is dispensable when space runs out, there is no reason that lesson would spare Norwegian trust, Dutch directness, or Japanese restraint. Sweden is simply the case we could measure cleanly, because its WVS profile is so distinct. The same test runs on any culture: swap the value set and the control, keep the method. If you are working on another language or culture and want to extend this — a Nordic neighbour, a minority language, a regional variety — we would genuinely like to hear from you. Get in touch and we will share the documents, the harness and what we learned building them.

If you work in journalism, policy, or research: the test works for any set of values, any control, any judge. Everything is published openly (CC0) at github.com/berget-ai/berget-eval. Run it, break it, tell us what we got wrong.

Method: the full detail
**Controlled documents.** Six synthetic Swedish documents — invented for this experiment, with no customer data anywhere in it (Berget runs a zero-retention policy: customer data is never stored, analysed, or trained on). Five are ordinary workplace prose — meeting minutes, an interview transcript, meeting notes, an internal news article and a planning email; the headline analysis uses these five. Each embeds 8 value statements and 6 length-matched neutral control statements, every one with a verbatim source span so presence can be checked exactly. Synthetic documents are deliberate in two ways: they let us know precisely what *should* survive, and they are the only data we are permitted to use.

**A baseline, not just a number.** Shortening alone is not bias — a model that shortens everything evenly is just doing its job. The evidence is selective loss: values dropped *more than* neutral content. The gap (values − control) is the metric.

**The sixth document is a genre boundary.** A formal policy memo behaved differently: there the model trimmed the administrative detail *harder* than the values, because a policy summary is expected to keep only the substance. We report it separately rather than averaging it in, because it shows the effect depends on what a genre treats as filler — which is the point, not a complication.

**How sure we are.** The statements are reused across models, so a naive count would overstate the data; we calculate uncertainty by resampling whole documents and models (cluster-robust bootstrap). The headline run is 135 tests — 9 models × 3 instruction styles × 5 workplace documents — scored over 14 embedded statements each, giving 888 value and 666 control judgements in the short version. Every item was also re-scored by two further judge models: gpt-oss-120b and Mistral Small agree with the primary judge on 85–89% of items (Cohen's kappa 0.65–0.73), so the classification is not the weak link.

**The length check that caught our own mistake.** Our first dataset had value statements averaging twice as many words as the neutral ones — which could itself explain why they were cut. On the rebuilt, length-matched data the gap is stable within both the short and long length bands, so it is not explained by the values simply being longer.

**Scoring weights.** A fully-kept statement scores 1, a watered-down one 0.5, an omitted one 0. If we instead count watered-down as 0 the gap is −28 points; as 1 it is −15.5; at 0.5 it is −21.8. The direction holds across the whole range.

**Per-model gaps in the short version** (five workplace documents, ~99 value judgements each, 95% uncertainty range):

| Model | Gap | 95% range |
| --- | ---: | --- |
| Kimi K3 | −33.9 pp | [−41.4, −24.7] |
| Mistral Small | −28.7 pp | [−49.2, −8.3] |
| gpt-oss | −27.1 pp | [−33.5, −17.9] |
| **Claude Opus 5** (frontier) | **−26.3 pp** | **[−32.1, −21.5]** |
| Mistral Medium | −23.4 pp | [−32.3, −10.8] |
| Gemma 4 | −20.5 pp | [−30.6, −5.8] |
| GLM-4.7 | −18.7 pp | [−32.8, −4.4] |
| GLM-5.2 | −16.5 pp | [−27.7, −5.3] |
| Llama 3.3 | −15.6 pp | [−26.7, −10.2] |
| Kimi K2.6 | −10.1 pp | [−22.9, +1.3] |

Every model sits below zero — the direction is consistent. Kimi K3 cuts hardest and Kimi K2.6 least, but most intervals overlap and the ordering should not be read as a settled ranking. The frontier reference, Claude Opus 5 (run through the identical documents and judge via the Anthropic API), lands mid-pack at −26.3 pp: the effect is not specific to open models.

Conflict of interest: Berget AI operates several of the models tested as a commercial service, and the primary judge model (Kimi K3) is one of them. That is why every item was re-scored by two further judge models, and why everything — documents, spans, prompts, judge verdicts and raw responses — is published openly, so anyone can re-run the test, re-judge with a different panel, and check our numbers.