> ## Documentation Index
> Fetch the complete documentation index at: https://docs.sqwish.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Check groundedness

> Ask whether a draft’s claims are supported by the sources you supplied.

Use this after drafting an answer and before presenting it. Include both the draft and the source passages in the context.

[Open this recipe in the playground](https://console.sqwish.ai/#playground?recipe=groundedness), or set `D1_API_KEY` as in the [quickstart](/quickstart) and run the same request:

```bash theme={null}
curl --fail-with-body --silent --show-error https://console.sqwish.ai/v1/decide \
  -H "Authorization: Bearer $D1_API_KEY" \
  -H "Content-Type: application/json" \
  --data-binary @- <<'JSON'
{
  "model": "sqwish-d1-core",
  "context": {
    "question": "Can I get my money back if I cancel my annual plan?",
    "sources": [
      "Refunds are issued within 14 days of a cancellation request.",
      "Annual plans can be cancelled at any time from the Billing page. Refunds cover the unused months only."
    ],
    "draft": "You can cancel your annual plan any time from Billing, and you'll get a full refund within 14 days, even after a year of use."
  },
  "decisions": [
    {
      "id": "grounded",
      "kind": "binary",
      "question": "Can the supplied sources establish every factual claim in the draft?",
      "outcomes": {
        "yes": "Every fact in the draft is stated in the sources or follows directly from them.",
        "no": "At least one fact in the draft is missing from the sources or contradicts them."
      }
    }
  ]
}
JSON
```

## Use the result

Read `decisions.grounded.p_yes`. Apply your reviewed acceptance threshold and send unsupported or uncertain drafts for revision or review. For long drafts, evaluate individual claims or sentences so one unsupported detail is not hidden inside an otherwise supported answer.

Supported means supported by these sources. It does not mean the sources are true, current, complete or authoritative. Check source quality separately and preserve citations for a reader to inspect.

Measure drafts accepted despite an unsupported claim, including incorrect numbers, names, dates and stronger conclusions than the source allows. Also test valid paraphrases so the check does not reject answers merely for using different wording.

## What was evaluated

The historical measurement covers the `grounded` decision on a particular public dataset. It does not establish factual truth or performance on every document length and domain.

<Note>
  Historical evaluation on **2026-09-28**, using **sqwish-d1-core**. These measurements describe that checkpoint and dataset, not current production performance or an accuracy guarantee.
</Note>

* Dataset: [AttributionBench (Stanford-GenSearch and AttributedQA)](https://huggingface.co/datasets/osunlp/AttributionBench) (Apache-2.0).
* Split: test\_all\_subset\_balanced.jsonl, two of its four subsets.
* Sample: 651 rows, 651 measured decisions.
* Measured decision IDs: `grounded`.
* Wording: tuned.
* Checkpoint SHA-256: `dd420ca652cfaa10279eabb54d15afd50fe9d24d357498c4dbc512b1c3d73bb2`.

| Metric | Value | Recorded interval |
| - | -: | - |
| accuracy | 0.7896 | 0.7573 to 0.8203 |
| log loss | 0.5072 | 0.4455 to 0.566 |
| brier | 0.323 | 0.2796 to 0.3651 |
| ece | 0.0766 | 0.0508 to 0.1091 |

The majority-class baseline has accuracy **0.5392**; the class-prior baseline has log loss **0.6901**.

Only the decision IDs listed above were measured. Dataset labels, class balance and wording affect these results. The intervals do not measure distribution shift. Re-evaluate with your own cases, including ambiguous and out-of-scope inputs.

## Improve it for your application

Keep the decision IDs and outcome order stable while evaluating changes. [Review a dataset](/guides/datasets), [tune the wording](/guides/prompt-tuning), and compare the result against your current policy before changing production. Check [fallback](/guides/models-and-fallback) so you know which model actually answered.
