Skip to main content
Use this after drafting an answer and before presenting it. Include both the draft and the source passages in the context. Open this recipe in the playground, or set D1_API_KEY as in the quickstart and run the same request:

Use the result

Read decisions.grounded.p_yes. Apply your reviewed acceptance threshold and send unsupported or uncertain drafts for revision or review. For long drafts, evaluate individual claims or sentences so one unsupported detail is not hidden inside an otherwise supported answer. Supported means supported by these sources. It does not mean the sources are true, current, complete or authoritative. Check source quality separately and preserve citations for a reader to inspect. Measure drafts accepted despite an unsupported claim, including incorrect numbers, names, dates and stronger conclusions than the source allows. Also test valid paraphrases so the check does not reject answers merely for using different wording.

What was evaluated

The historical measurement covers the grounded decision on a particular public dataset. It does not establish factual truth or performance on every document length and domain.
Historical evaluation on 2026-09-28, using sqwish-d1-core. These measurements describe that checkpoint and dataset, not current production performance or an accuracy guarantee.
  • Dataset: AttributionBench (Stanford-GenSearch and AttributedQA) (Apache-2.0).
  • Split: test_all_subset_balanced.jsonl, two of its four subsets.
  • Sample: 651 rows, 651 measured decisions.
  • Measured decision IDs: grounded.
  • Wording: tuned.
  • Checkpoint SHA-256: dd420ca652cfaa10279eabb54d15afd50fe9d24d357498c4dbc512b1c3d73bb2.
The majority-class baseline has accuracy 0.5392; the class-prior baseline has log loss 0.6901. Only the decision IDs listed above were measured. Dataset labels, class balance and wording affect these results. The intervals do not measure distribution shift. Re-evaluate with your own cases, including ambiguous and out-of-scope inputs.

Improve it for your application

Keep the decision IDs and outcome order stable while evaluating changes. Review a dataset, tune the wording, and compare the result against your current policy before changing production. Check fallback so you know which model actually answered.