STAGE

System One in the tool loop: fast judgments for chat and RAG

Christian Landgren
Christian Landgren

Co-founder & CPTO

System One — fast judgments for chat and RAG

A large language model knows a lot about the world. For factual questions — prices, emissions, board members — knowing is not the same as being right, so every serious chat application does the same thing before it answers: it looks the facts up on the web first, and only then writes the reply to the user. The pattern is called RAG (retrieval-augmented generation), and time is where it is won or lost. The lookup is not one call; it is a walk — a search page, five results, the promising link, a hub page, the actual document — and the user leaves if the walk takes too long.

The standard way to make the walk's judgments is to ask the large model at each step: is this link worth following, did that page carry anything usable? Each judgment becomes a full chat-model call — the seconds the user feels, and the energy spent generating reasoning for what is essentially a look-up. There is a cheaper way. On Friday, Berget AI announced a new category of models — System One — that answer typed questions with a calibrated choice, score, or yes/no probability instead of prose, in about a tenth of a second. Our first model in the category, Laya, is now serving requests on the Berget API. Over the past days we rebuilt our own search-and-answer pipeline around it — a pattern set that anyone building chat or RAG applications can use, each pattern carrying numbers from our own runs.

The answer gate

The first pattern is also the biggest win. In a RAG pipeline you retrieve documents, paste them into your large model, and hope the answer is in there. Instead, ask a System One model first:

{
  "model": "laya-latest",
  "state": "QUESTION: How high is Mount Everest?\n\nDOCUMENT:\n<chunk text>",
  "questions": {
    "has_answer": {
      "type": "noul",
      "instructions": "Does the PAGE TEXT itself contain a direct answer to the QUESTION (facts, a name, a number, a statement) — not just links or teasers?"
    }
  }
}

The answer is a probability between 0 and 1. Chunks scoring high go to your large model; the rest never enter the context window. In our pipeline this is what replaced the summarise-the-page-then-hope step: instead of a lossy paraphrase at a full chat-model price, you get the answer-bearing sections verbatim — facts and numbers survive intact, and your main model's context window carries only what matters.

For long documents we scan chunks of 4,000 characters in parallel, six at a time, and keep the four best. A 200-page report costs a couple of seconds of scanning. Two tricks from actual runs matter here: chunks scoring high for "how high / how much" questions are penalised when they contain no digits at all, and pages that smell like bot-blockers — "performing security verification" — are discarded before Laya forms a judgment. That check exists because on one run the text was an empty verification shell and our model still said 1.0. The guardrails are doing real work even with a small model on top.

Agentic search is a chain of link choices, and each link choice is a choice question with the candidates as criteria:

  • choice — pick one from named criteria; returns the selected key, a confidence, and the probability for every option.
  • One request can carry several questions: we ask "which link moves closest to the goal?" and "does the page text contain the answer?" together, and get both judged in one pass.

In our test pipeline, every search spawns three tracks. Track one starts at the link Laya itself prefers from the search results; tracks two and three take the next hits as breadth. Each track walks pages, checks the answer gate on each, and stops when a page proves it carries the answer. If no page does, the loop calls a small generation model once — a gemma-class model works — to reformulate the question, and the walk runs again. The loop is bounded, so it cannot wander forever.

The measured behaviour on a mix of tricky and easy questions: three to nine decisions per answered query, each decision between 30 and 170 ms when calling the API over the internet from a laptop (the engine itself was measured at 31–33 ms server-side). A complete web question with sources came back in under ten seconds, and the whole decision cost was a fraction of a single chat-model call.

Three patterns you can implement this week

All loops below are small enough to port in an afternoon, and all three are measured in our local tests. They share one shape: a tool produces a short list of options, Laya picks one, the next tool runs on the pick. No loop step involves the main model.

Find the facts on the open internet

A complete answer from the open web, with sources, in a handful of seconds:

search(question)             → 10 hits
fetch(hit)                   → page text
laya_judge(page, question)   → "does this page answer?"  (noul)
laya_choice(links, question) → "which page answers next?"  (choice)
  → gate yes: done          → gate no: fetch the chosen next page, repeat
synthesize(best pages)       → the answer your user sees

One step is a single request — the question, the page text, and the menu the model may choose from. For a repair question, after the first fetch, the request and answer look like this:

{
  "state": "QUESTION: Hur åtgärdar man E22-larmet på en diskmaskin?\n\nPAGE: Diskmaskin E22 – Bosch (service)\nTEXT: Felkoden E22 betyder att filtren är igensatta. Rengör …",
  "questions": {
    "has_answer": {
      "type": "noul",
      "instructions": "Does the PAGE TEXT itself contain a direct answer to the QUESTION?"
    },
    "next": {
      "type": "choice",
      "instructions": "Which link moves closest to a repair guide for E22? ONE.",
      "criteria": {
        "0": "Bosch – Diskmaskin fel E22 (service)",
        "1": "Forumtråd: vattnet går inte ur (användarinlägg)",
        "2": "Reservdelskatalog för diskmaskiner",
        "none": "This page already answers the question"
      }
    }
  }
}
{
  "answers": {
    "has_answer": { "type": "noul", "noul": 0.94 },
    "next": {
      "type": "choice",
      "choice": "0",
      "confidence": 0.91,
      "probabilities": { "0": 0.89, "1": 0.07, "2": 0.03, "none": 0.01 }
    }
  },
  "usage": { "input_tokens": 214, "output_tokens": 0 }
}

The none option matters: the model is allowed to reject the menu, which is what stops a confident wrong turn from becoming the next fetch. Note also what the answer costs — 214 input tokens, zero output, and no chat completion was spent on the step.

The loop terminates the moment a page clears the answer gate — and that, not the model, is what sets the clock. Laya's judgment costs 30–170 ms per page; the fetches dominate. Our fastest complete question measured 4.3 seconds from question to sourced answer; the slow ones ran ten because a second page had to be fetched and rejected first. When you want sub-four-second runs, the work is not in the model — it is in picking a search provider with fast responses and letting the gate end the walk after the first page.

Two guardrails from our runs: when the question asks for figures ("redovisa som tabell"), we penalise answer-chunks that contain no digits at all; and search-result pages are never allowed to pass the gate themselves, because a model can read a page of link titles and still claim the answer is there.

Find the right file on your own disk

The same loop runs against a local filesystem, and here the decisions get genuinely small:

ls           → a menu of files
laya_choice  → "which of these leads to the right file?"
cd           → the picked directory
laya_choice  → next menu: entries or tools — ls dir/, grep text, cat file
cat          → the file is read

The menu is literally the ls output, offered as criteria (later menus mix in tools — ls dir/, grep text, cat file). After ls in a server directory, the question to the model and its answer:

{
  "state": "QUESTION: where does the nightly batch job log its output?\n\nNEW DIRECTORY: ~/srv (11 entries)",
  "questions": {
    "step": {
      "type": "choice",
      "instructions": "Which entry moves closest to the batch job's log file? ONE.",
      "criteria": {
        "0": "deploy/",
        "1": "jobs/",
        "2": "migrations/",
        "3": "README.md",
        "4": "utils/",
        "none": "This directory does not contain the file"
      }
    }
  }
}
{
  "answers": {
    "step": {
      "type": "choice",
      "choice": "1",
      "confidence": 0.78,
      "probabilities": { "1": 0.74, "4": 0.12, "0": 0.08, "2": 0.04, "3": 0.01, "none": 0.01 }
    }
  },
  "usage": { "input_tokens": 96, "output_tokens": 0 }
}

cd jobs, run ls again, ask again — three such choices landed inside the file in our runs, at 25–80 ms per decision. The whole walk sends the model only directory names, so both the codebase and the judgments stay on your side of the wire.

In our internal experiment — an agent searching a codebase by asking Laya for the next step — each decision cost about 25 ms, and a ten-step search finished in 2.1 seconds, most of that time spent actually reading the files the model pointed at. The property that makes this pattern attractive for developer tools is not the speed alone: if you run it against your own code, both the code and the judgments stay on infrastructure you control — the menu you send to Laya is a list of file names, and nothing else.

The same benchmark also measured where the model is weakest: choices among many options, browser-shaped ones above all. Expect wrong turns. The loop design is what makes them cheap — a wrong turn costs one more 25 ms decision, not a chat-model round trip and a rewritten prompt.

Ask the document where its answer lives

Sooner or later the walk from the previous patterns stops in front of a PDF — the annual report, the sustainability appendix, the technical annex. A PDF is the heaviest thing a RAG pipeline touches: tens of megabytes to download, hundreds of pages to convert, and the temptation is to ingest the whole thing into the large model. That is the long version of the waiting problem: after the walk, a second wait — download, conversion queue, a context window stuffed with 450,000 tokens of text the question never asks about. Users do not wait through that either.

Our split keeps the two waits apart. The web walk from the earlier patterns stops as soon as the PDF's URL is known — nothing is downloaded yet. A second step takes the question to the document.

The first version of that second step — Laya steering a jump menu ("skip 10, skip 20") through a 112-page report — never got off the ground. Offered jumps that point at pages it has never seen, the model declined thirteen menus in a row. The run taught us the shape of the fix: Laya chooses well among things it can actually see, and poorly among things it must imagine.

So extract the PDF locally first (PyMuPDF for digital files — seconds, no service call), then hand the model the document as page ranges, each labelled with the text it opens with, and ask two questions in one call: whether the answer exists in the document at all (noul), and which range holds it (choice, with the full probability distribution):

{
  "answers": {
    "exists": { "type": "noul", "noul": 0.02 },
    "where": {
      "type": "choice",
      "choice": "4",
      "probabilities": {
        "4": 0.44,
        "7": 0.17,
        "3": 0.1,
        "5": 0.07,
        "8": 0.07,
        "2": 0.05,
        "1": 0.01,
        "6": 0.01,
        "9": 0.01,
        "11": 0.01
      }
    }
  },
  "usage": { "input_tokens": 1620, "output_tokens": 0 }
}

Range 4 covers pages 41–50 at 0.44 — and free, independent evidence agreed: a numeric gate we run alongside had flagged exactly pages 41-45 as the only window carrying scope 3-style figures. One 138 ms decision located the region of the answer in a 112-page document, as an array you can sort and threshold. The low exists score in the same response is worth keeping: it is the cheap signal that says turn around, this document does not answer the question, before you spend anything on ingestion. Yet we kept moving — and the reconciliation matters: the where distribution pointed with conviction and the free numeric gate agreed, while exists judges the whole document from its opening slices, the weakest evidence we give the model. Ask it before the expensive steps; do not let it outvote the sharper signals.

Inside the winning range we fell back to the same judge at finer grain — noul per 2,200-character chunk across thirteen chunks — 1.9 seconds. We also tried refining by per-page openings, 130 characters each, and the model spread its answers without conviction: judging pages by their openings gives opening-quality judgments. Chunks of real text, as in the earlier finder, are what the judge needs under it. Even then, the best-scoring chunk (0.86) was number-dense financial text rather than the emissions figure itself — the model flags "numbers live here" more reliably than "the answer is this". The last reading of the two or three hot windows belongs to the model that can tell those apart.

The whole chain, measured: 1.2 seconds to fetch and extract, 138 ms for the where-array, 1.9 seconds for the chunk scan — roughly 35,000 input tokens, against the roughly 450,000 we would have spent feeding the document to a chat model, had the conversion service not refused the file outright. The small model says where; the large model only reads the spread it earns.

The verdict pattern

Tools fail quietly. A fetch returns a cookie wall instead of an article, an OCR run returns half a page, a crawl returns a sitemap nobody needed. The verdict pattern makes every tool self-reporting: after the tool does its work, the same System One model judges "did the tool produce real, usable content?" Tools that fail escalate on their own — the pipeline retries with a more expensive extraction method without waking the main model:

  • plain fetch when the page is server-rendered,
  • render-in-the-browser services when it is not,
  • for PDFs: local extraction first, a rendering service second, and a document-OCR queue third.

On our runs, report PDFs of two hundred pages were the hard edge: the rendering service returns errors on some heavy documents, and for those the OCR queue is the dependable route. The shape of the escalation matters more than any single service: your pipeline can always start cheap, because the judgment that says "not good enough" costs milliseconds.

Calibration

A System One model is not a small chat model, and Laya is early in its development. On our four-suite benchmark (228 answers, one repetition, seed locked, datasets labelled with GLM 5.3 — a model family we also sell, so treat our verdicts the way we treat them), hosted Jev scored 88.5 per cent strict accuracy on the core suite, a local general-purpose model 84.6, and Laya 56.4 — with overconfidence on top (Brier 0.292 against 0.044).

The patterns above survived that gap because none of them depend on a single judgment being right. The gates are cheap, so the pipeline double-checks: answers must carry the digits the question implies, refusals from the final model trigger the next source group, results from one domain do not outweigh each other. The rule from our launch post still holds — treat Laya's probabilities as hints, not verdicts. And because every judgment is logged with its confidence, the trace is exactly the data the next fine-tuning run should learn from: the pipeline that uses Laya the most is also the one that improves her the fastest.

Running it yourself

Laya is reachable at api.berget.ai/v1/systemone under model id laya-latest, in evaluation with the same terms as the rest of the platform: zero retention, Swedish servers, EU jurisdiction. The request shape is Jev-compatible, so the TypeSafe SDKs (Python typesafe-sdk 0.7.1, TypeScript @typesafe-ai/sdk 0.6.0) work with a base-URL change. Each response carries a usage block, so the metering behaves like the chat completions you already run.

Where this belongs — and where it does not

If you are building a chat product whose conversations touch customer data that must never leave EU jurisdiction, these patterns let the tool loop run everywhere except the one place your large model is actually needed: the synthesis your users see. If you operate a RAG stack, start with the answer gate — it is one request per retrieved document, and the chunks that reach your model arrive verbatim.

Laya is not the model for open-ended reasoning, and with calibration where it is today, a single gate judgment will occasionally be wrong. That is why every pattern here pairs a cheap judgment with a consequence — a skipped chunk costs a retry, not a wrong answer under your product's name.

Our worked examples made about thirty System One decisions in total, each logged with its confidence. The log is not bookkeeping; it is waiting to become the next fine-tuning run — trained by the pipeline it serves.