D1_API_KEY as in the quickstart and run the same request:
Use the result
Readdecisions.grounded.p_yes. Apply your reviewed acceptance threshold and send unsupported or uncertain drafts for revision or review. For long drafts, evaluate individual claims or sentences so one unsupported detail is not hidden inside an otherwise supported answer.
Supported means supported by these sources. It does not mean the sources are true, current, complete or authoritative. Check source quality separately and preserve citations for a reader to inspect.
Measure drafts accepted despite an unsupported claim, including incorrect numbers, names, dates and stronger conclusions than the source allows. Also test valid paraphrases so the check does not reject answers merely for using different wording.
What was evaluated
The historical measurement covers thegrounded decision on a particular public dataset. It does not establish factual truth or performance on every document length and domain.
Historical evaluation on 2026-09-28, using sqwish-d1-core. These measurements describe that checkpoint and dataset, not current production performance or an accuracy guarantee.
- Dataset: AttributionBench (Stanford-GenSearch and AttributedQA) (Apache-2.0).
- Split: test_all_subset_balanced.jsonl, two of its four subsets.
- Sample: 651 rows, 651 measured decisions.
- Measured decision IDs:
grounded. - Wording: tuned.
- Checkpoint SHA-256:
dd420ca652cfaa10279eabb54d15afd50fe9d24d357498c4dbc512b1c3d73bb2.
The majority-class baseline has accuracy 0.5392; the class-prior baseline has log loss 0.6901.
Only the decision IDs listed above were measured. Dataset labels, class balance and wording affect these results. The intervals do not measure distribution shift. Re-evaluate with your own cases, including ambiguous and out-of-scope inputs.
