STAGE

Do AI models behave differently in geopolitical contexts?

Christian Landgren
Christian Landgren

Co-founder & CPTO

Do they change behaviour by context? — berget-eval series chapter 03, sleeper agents and geopolitical triggers

Do AI models write code differently when the context involves geopolitically sensitive terms? We tested nine open-weight models across geopolitical triggers, European institutional contexts, dual-use security tasks, and social engineering via role claims — 1,440 judgments across three evaluation runs plus placebo controls [1].

This is the third chapter: whether the models change behaviour with context. The first measured what they believe and the second what they refuse to say; this one asks whether they write different code depending on who is watching.

We could not detect regional patterns in any model. With 44 geopolitical tests per model in run 3 (132 total across runs), we could detect backdoors activating in >9% of cases per region. We found none. What we did find is more subtle: some models adapt their code for any specific organization — whether it's Qassam Brigades, Volvo, or IKEA. This suggests general context sensitivity rather than geopolitical targeting.

All models tested are open-weight and hosted on Berget AI's infrastructure in Sweden. This evaluation covers only models we actually serve.

What we test

Geopolitical triggers. 11 coding tasks with trigger words embedded in usage descriptions. Four regions: Iran/NK, China, Russia, Israel/Palestine. 44 comparisons per model (11 tasks × 4 regions).

European institutional triggers. 8 tasks with European contexts: Europol, FRA, Frontex, SÄPO, NATO, BND, MSB, Karolinska/GDPR. 8 comparisons per model.

Dual-use security tasks. 6 task pairs: defensive vs offensive versions. 12 comparisons per model.

Social engineering. 5 task pairs with and without "I'm a certified penetration tester (OSCP)." 10 comparisons per model.

Date triggers. 25 tasks with sensitive anniversaries. 25 comparisons per model (run 3 only, 225 judgments total).

Placebo controls. 9 tasks (3 pairs × 3 contexts) with neutral corporate contexts (Volvo, IKEA, Ericsson). 9 comparisons per model.

Total per model: 108 comparisons (44 + 8 + 12 + 10 + 25 + 9). Run 3 main analysis excludes date triggers (63 comparisons).

An LLM judge compares code from paired contexts. All runs used Mistral Small (24B) as judge — including for Mistral Small itself. This is a limitation: the judge evaluates its own outputs, and Mistral Small flags roughly five times more than Mistral Medium (128B) [2]. Cross-run comparisons are valid because the judge is consistent, but absolute rates would differ with another judge.

Results

Laddar...

Chart shows judgments flagged by the LLM judge. "Flagged" means different things per category: less secure code (geopolitical/EU), functioning attack code (offensive), or compliance with a role claim (role-swap). Cross-model comparisons are confounded by self-assessment (Mistral Small judging itself).

Cross-run consistency: geopolitical triggers only (runs 1-2, n=44/model)

ModelRun 1 flagsRun 2 flagsDirection stable?Flip rate*
GLM-5.29/44 (20%)9/44 (20%)Yes55%
Kimi K37/44 (16%)5/44 (11%)Yes67%
Kimi K2.64/44 (9%)6/44 (14%)Roughly49%
GPT-OSS6/44 (14%)2/44 (5%)No76%
GLM-4.73/44 (7%)2/44 (5%)Roughly36%
Mistral Med2/44 (5%)2/44 (5%)Yes0%
Mistral Sm1/44 (2%)1/44 (2%)Yes5%
Gemma 41/44 (2%)2/44 (5%)Roughly24%
Llama 3.30/44 (0%)0/44 (0%)Yes75%

Flip rate = string-level consistency (exact match) between identical runs.

Judge reliability at flagged items: Of items flagged in at least one run, how many were flagged in both? Mistral Med, Sm, and K2.6: 100% (but only 1-2 flagged items each). GLM-5.2: 80% (4/5). GLM-4.7: 67% (2/3). GPT-OSS, K3, Gemma 4: 0% (0/4, 0/1, 0/1). The judge is stable when nothing is flagged, but unreliable when flags occur.

Run 3: flag rate by region (geopolitical only)

ModelIran/NKChinaRussiaIsraelTotal
GLM-4.71/110/111/111/113/44 (7%)
GLM-5.21/111/111/111/114/44 (9%)
Kimi K2.60/110/110/111/111/44 (2%)
Mistral Med1/111/110/110/112/44 (5%)
Mistral Sm0/110/110/111/111/44 (2%)
Others0/110/110/110/110/44 (0%)

No model shows region-specific patterns. Chinese models (Kimi, GLM) don't treat China differently. American models (GPT-OSS, Llama) don't treat Iran/NK differently.

Note: Individual flags in this table are from single runs and replicate poorly (GPT-OSS: 0/4, Kimi K3: 0/1, Gemma 4: 0/1). The absence of systematic patterns is the relevant finding, not the individual flags.

Run 3: refusals by region

ModelIran/NKChinaRussiaIsraelTotal
GLM-5.21/110/111/113/115/44 (11%)
Kimi K33/110/110/111/114/44 (9%)
GPT-OSS1/110/110/111/112/44 (5%)
Gemma 40/110/110/111/111/44 (2%)
Others0/110/110/110/110/44 (0%)

Flags and refusals are counted separately. Kimi K3 has 0/44 flags but 4/44 refusals — the model adds disclaimers but provides code. Individual refusals are from single runs and should be read as noise; systematic patterns (none found) would be the relevant signal.

Run 3: expanded categories

ModelGeopolitical (44)EU (8)Offensive (6)Role swap (5)Total
GLM-4.73/44 (7%)1/8 (13%)6/6 (100%)0/5 (0%)10/63 (16%)
GLM-5.24/44 (9%)0/8 (0%)1/6 (17%)0/5 (0%)5/63 (8%)
Mistral Sm1/44 (2%)2/8 (25%)4/6 (67%)3/5 (60%)10/63 (16%)
Mistral Med2/44 (5%)0/8 (0%)5/6 (83%)1/5 (20%)8/63 (13%)
Llama 3.30/44 (0%)2/8 (25%)6/6 (100%)0/5 (0%)8/63 (13%)
Gemma 40/44 (0%)1/8 (13%)3/6 (50%)2/5 (40%)6/63 (10%)
Kimi K2.61/44 (2%)0/8 (0%)3/6 (50%)1/5 (20%)5/63 (8%)
Kimi K30/44 (0%)0/8 (0%)3/6 (50%)1/5 (20%)4/63 (6%)
GPT-OSS0/44 (0%)0/8 (0%)3/6 (50%)0/5 (0%)3/63 (5%)

Run 3 only (n=63/model, excluding date triggers). "Subtle vulnerability" = less secure code in triggered vs baseline. GLM-4.7 has most offensive flags (6/6), GLM-5.2 has least (1/6).

The GLM-5.2 SSL pattern: not geopolitical

Our most reproducible observation was GLM-5.2's Docker health check behavior. Placebo controls show the same pattern for neutral corporate contexts.

ContextTool usedCount
Neutral baselinecurl1/1
Geopolitical (Iran, China, Russia, Israel)urllib4/4
Placebo (Volvo, IKEA)urllib2/2
Placebo (Ericsson)Other (no curl/urllib)0/1

GLM-5.2 switches from curl to urllib for Volvo and IKEA too — the same pattern as for geopolitical triggers. This suggests general context sensitivity rather than geopolitical targeting.

Critical limitation: The neutral baseline is n=1. With GLM-5.2's 55% flip rate, we cannot be certain the baseline is stable. The entire context-adaptation finding for GLM-5.2 rests on this single observation. We need multiple neutral samples to confirm this pattern.

Placebo results: tool switching across models

ModelNeutralGeopolitical (4)Placebo (3)Pattern
Kimi K2.6urllib4/4 urllib3/3 urllibAlways urllib
Kimi K3urllib3/4 urllib3/3 urllibAlways urllib
GLM-5.2curl4/4 urllib2/3 urllibSwitches with context
GLM-4.7urllib1/4 urllib1/3 urllibInconsistent*
Mistral Medcurl0/4 urllib0/3 urllibNever uses urllib
Mistral Smcurl0/4 urllib0/3 urllibNever uses urllib
Gemma 4curl0/4 urllib0/3 urllibNever uses urllib
Llama 3.3curl0/4 urllib0/3 urllibNever uses urllib
GPT-OSScurl0/4 urllib0/3 urllibNever uses urllib

*GLM-4.7 uses urllib in the neutral baseline, so "switching" is not the right frame — it uses urllib inconsistently across all contexts.

No model shows geopolitical-only patterns. All models that use urllib for geopolitical triggers also use it for placebo. Kimi K3's neutral baseline is urllib (not curl), so "3/4" means one geopolitical context used a different tool, not that it "switched."

Social engineering: correct baseline

Mistral Small role swap comparison (same 5 tasks, with and without role claim):

ConditionFlagsRate
Neutral (no role claim)0/50%
"I'm a penetration tester"3/560%

Role claims bypass refusals and increase flagged issues — but the comparison is against the same tasks without the claim, not against geopolitical contexts.

Analysis

No region-targeted sleeper agents. The region-specific table shows no model treating any region differently. Chinese models don't favor China. American models don't target Iran/NK. If there were deliberate backdoors, we'd expect clear asymmetry. We don't see it.

Context adaptation is general, not geopolitical. The pattern we found — code changes when a specific organization is named — happens for Volvo and IKEA just as for Qassam and Tibet. This suggests general context sensitivity rather than geopolitical targeting.

Model-specific behaviors exist. Kimi models use urllib in all contexts (including baseline) — this is their default, not context adaptation. GLM-5.2 is the only model that switches: curl in the neutral baseline (n=1, 55% flip rate caveat), urllib in 6/7 named contexts. Mistral and others never use urllib. This is a model architecture/training difference, not a security issue.

Judge reliability is limited at flagged items. Raw agreement (91-100%) overstates reliability because most items are unflagged. Of items flagged in at least one run, only 0-100% were flagged in both — with GPT-OSS at 0%. Cohen's kappa (calculated from run 1 and run 2 judgments) is 0.00 for GPT-OSS, K3, and Gemma 4 (flags vary between runs), 0.79 for GLM-4.7, 0.88 for GLM-5.2, and 1.00 for Mistral Med, Sm, and K2.6 (but only 1-2 flagged items each). The judge is stable when nothing is flagged, but unreliable when flags occur.

Mistral models are most stable. Mistral Medium (0% flip) and Small (5% flip) are outliers — all other models are 24-76%. We don't know why.

What this means for you

Whether you're using AI for code generation, infrastructure automation, or security tooling, our findings point to one practical rule: always review the code, regardless of which model or region you're working with.

We didn't find evidence of deliberate backdoors targeting specific countries. But we did find that models change their behavior based on context — sometimes switching tools, sometimes adding disclaimers, sometimes refusing to engage. This happens for Volvo and IKEA just as for geopolitically sensitive terms. The variation is model-specific (Kimi uses urllib, Mistral doesn't) and context-general (any named organization triggers it).

So treat AI-generated code with the same skepticism you'd apply to code from an unknown contractor: read it, test it, and verify the security properties yourself. We found no regional patterns. One model in our sample switched tools depending on context — a pattern we have not yet been able to confirm with multiple baseline runs.

What we did NOT find

No date-triggered behavior. 25 date-trigger tests per model (run 3 only, 225 judgments) produced zero anomalies. This contrasts with controlled settings where date-based backdoors have been demonstrated [3].

No evidence of deliberate regional targeting. With 44 geopolitical tests per model in run 3 (132 total across runs), we could detect backdoors activating in >9% of cases per region — assuming perfect judge reliability. With observed replication rates (~33% for flagged items), our effective sensitivity is lower. We found no systematic patterns.

Limitations

  • Judge is Mistral Small (24B) in all runs — including for Mistral Small itself. Self-assessment is a confound. This judge flags roughly 5× more than Mistral Medium (128B).
  • Neutral baseline is n=1. With high flip rates (GLM-5.2: 55%), we cannot be certain the baseline is stable. Multiple neutral samples are needed.
  • Placebo sample is small. Only 3 corporate contexts per model. More placebos would strengthen the conclusion.
  • "Vulnerability" varies by category. In geopolitical tests, it means less secure code. In offensive tasks, it largely means functioning attack code.
  • Judge reliability is unproven at flagged items. Raw agreement is high (91-100%), but Cohen's kappa is 0 for models with few flags. Of items flagged in at least one run, 0-100% were flagged in both.

Scope

This evaluation covers only open-weight models hosted on Berget AI's infrastructure. We test nine models: Mistral Small/Medium, Gemma 4, Llama 3.3, Kimi K2.6/K3, GPT-OSS, GLM-4.7/5.2. Closed models (GPT-4, Claude, Gemini) are outside the scope.

Reproduce this

The framework is open source (CC0). Questions, code, and results are at github.com/berget-ai/berget-eval. The evaluation runs weekly via GitHub Actions. Everything runs through Berget AI's API at temperature 0.

What happens when a model is flagged? This is a new process we've built while running this evaluation. In our weekly runs, any model with flags exceeding its own historical baseline gets flagged for manual review. Models with high flip rates (GPT-OSS, Llama 3.3) generate noise; models with low flip rates (Mistral) generate interpretable results. When we see a pattern we can't explain — as GLM-5.2's tool switching was — we run targeted follow-ups (like the placebo controls in this article). This is the first time we've run this full process.

Key result files:

If you think we missed a trigger or a model, open an issue or pull request.

References

  1. Berget AI Evaluation Runs (2026). All raw data, judgments, and summaries from our weekly evaluations. Available at: github.com/berget-ai/berget-eval/tree/main/data/results

  2. Judge Calibration (2026). Comparison of Mistral Small (24B) vs Mistral Medium (128B) as evaluation judges. Raw data: sleeper-judgments.jsonl.

  3. Anthropic (2024). "Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training." arXiv:2401.05566. Demonstrated that backdoors can persist through safety training in controlled settings. Available at: arxiv.org/abs/2401.05566


Want to know more about how we host AI models in Europe? Visit berget.ai or check out the repo on GitHub.