What Verbalized Sampling buys you (and where scores lie)

OpenDrift exists because a single completion is a lossy summary of what the model can say.
Verbalized Sampling is the mechanism: instead of drawing one token path and calling it the answer, you ask the model to produce several candidates and report probability scores in a structured format. OpenDrift sends that request through OpenRouter, parses the XML, ranks the candidates, and saves them so comparison is part of the product, not a side script.
This post is the practical guide: what you gain, how we wire it, and the places scores will mislead you if you treat them like ground truth.
The problem with one sample
Language models are distributions over continuations. A chat UI samples once (temperature, top-p, and luck included) and presents fluency as confidence.
That is fine for drafting an email. It is a bad default when you are:
- Evaluating prompt or model changes.
- Exploring creative options you might actually ship.
- Debugging why two runs feel “randomly” different.
- Trying to see whether the model is peaked or flat on a question.
One answer hides the fork. You need the fork on screen.
What Verbalized Sampling changes
Verbalized Sampling asks the model to externalize several answers and attach scores it claims for them. In OpenDrift’s path, that response is structured XML so parsing is deterministic enough to rank and persist.
You get three things a normal chat turn does not:
- Alternatives as first-class objects — not “regenerate and hope.”
- A ranking signal — imperfect, but better than remembering which tab felt better.
- A saved landscape — so evaluation and creative work read from the same artifact later.

How OpenDrift runs the loop
At a high level:
- You write a prompt the way you would for any serious model call.
- OpenDrift requests multiple candidates via Verbalized Sampling on OpenRouter.
- The response is parsed into structured candidates with scores.
- Candidates are ranked and stored.
- You inspect where the model is sharp, where it branches, and what it nearly chose.
The product stays focused on that spine. Features around evaluation, creative exploration, and uncertainty inspection all read from the same ranked set.
Where probability scores lie
Treat model-reported scores as hints from the same system that wrote the text, not as calibrated probabilities from a judge outside the model.
Watch for these failure modes:
- Fluent garbage with a high score. The model can sound sure about a wrong fact. Ranking by self-reported score will not save you from confident hallucination. Use the landscape to compare, then verify the winner against reality.
- Collapsed diversity. If candidates are near-paraphrases, the branch view is theatre. Tighten the ask (“distinct strategies,” “mutually exclusive options,” “different failure modes”) or change sampling settings until the set actually fans out.
- Score–rank mismatch you can feel. Sometimes the “best” answer to a human is mid-list. That is useful signal. It means your ranking objective and the model’s self-score are not the same thing — which is exactly when a human-in-the-loop ranking pass earns its keep.
- Structured output drift. XML contracts help, but models still wander. OpenDrift’s value depends on parse success. When a run fails to parse cleanly, that is a first-class outcome to log, not a silent retry that invents a tidy UI.

When this is the right tool
Use OpenDrift when the cost of a hidden alternative is high: eval suites, ambiguous creative briefs, policy-sensitive wording, or any workflow where “the model said so” is not enough.
Skip it when you already know you want one cheap draft and will edit heavily by hand. Sampling once is still the right default for low-stakes fluency.
What we are watching next
We care less about prettier cards and more about:
- Better diversity controls without turning the UI into a research console.
- Clearer parse and contract errors when structure slips.
- Export that plays nicely with eval harnesses and evidence workflows elsewhere in the lab.
If you try OpenDrift on a real eval or a real brief, tell us which branch you kept — and which high-scoring one you correctly ignored.
Part of the Antler Labs blog. Previous: Ideas branch. We build them. Live product: OpenDrift · Contact: hello@antlerlabs.dev