> ## Documentation Index
> Fetch the complete documentation index at: https://docs.sqwish.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Review answer quality

> Use a shared rubric to decide whether a draft needs revision.

Use this after a model drafts a reply. Include the original request and the draft so the rating can judge whether the response actually answers the task.

[Open this recipe in the playground](https://console.sqwish.ai/#playground?recipe=answer-quality), or set `D1_API_KEY` as in the [quickstart](/quickstart) and run the same request:

```bash theme={null}
curl --fail-with-body --silent --show-error https://console.sqwish.ai/v1/decide \
  -H "Authorization: Bearer $D1_API_KEY" \
  -H "Content-Type: application/json" \
  --data-binary @- <<'JSON'
{
  "model": "sqwish-d1-core",
  "context": {
    "conversation": [
      {
        "role": "user",
        "content": "Write a Python function that removes duplicates from a list but keeps the original order."
      }
    ],
    "draft": "Use a set:\n\ndef dedupe(items):\n    return list(set(items))\n\nA set can't hold the same value twice, so every item appears once."
  },
  "decisions": [
    {
      "id": "quality",
      "kind": "ordinal",
      "question": "Is this draft reply ready to send to the user?",
      "outcomes": [
        "0",
        "1",
        "2"
      ],
      "rubric": [
        "No. It is wrong, unsafe or unhelpful, or it ignores part of what the user asked.",
        "Nearly. It is on the right track, but has an error, a gap or a missing step to fix first.",
        "Yes. It is correct and complete, and does what the user asked."
      ]
    }
  ]
}
JSON
```

## Use the result

Read `decisions.quality.expected_level` on the 0–2 scale, plus the full distribution. Treat a middle value as uncertainty between rubric levels, not as a new categorical label. Establish which levels lead to revision or human review using your reviewers' own ratings.

Evaluate drafts rated 2 that a person would send back. Make the rubric concrete about missing requirements, factual errors and unusable advice. Quality does not establish groundedness; when sources matter, pair this with the [groundedness check](/examples/groundedness).

## What was evaluated

The historical accuracy is **0.557**, below the majority-class baseline of **0.615**. This recipe has not established useful classification performance on that benchmark. Use it as a format and rubric starting point, then improve wording and measure on your own reviewed examples before adopting it.

<Note>
  Historical evaluation on **2026-09-28**, using **sqwish-d1-core**. These measurements describe that checkpoint and dataset, not current production performance or an accuracy guarantee.
</Note>

* Dataset: [HelpSteer3 (feedback, validation split)](https://huggingface.co/datasets/nvidia/HelpSteer3) (CC BY 4.0).
* Split: feedback/validation, prompts whose hash falls below 0.55.
* Sample: 1,000 rows, 1,000 measured decisions.
* Measured decision IDs: `quality`.
* Wording: hand-written.
* Checkpoint SHA-256: `dd420ca652cfaa10279eabb54d15afd50fe9d24d357498c4dbc512b1c3d73bb2`.

| Metric | Value | Recorded interval |
| - | -: | - |
| accuracy | 0.557 | 0.5258 to 0.589 |
| log loss | 0.9374 | 0.9088 to 0.9664 |
| brier | 0.5572 | 0.538 to 0.577 |
| ece | 0.0785 | 0.0544 to 0.1087 |

The majority-class baseline has accuracy **0.615**; the class-prior baseline has log loss **0.9319**.

Only the decision IDs listed above were measured. Dataset labels, class balance and wording affect these results. The intervals do not measure distribution shift. Re-evaluate with your own cases, including ambiguous and out-of-scope inputs.

## Improve it for your application

Keep the decision IDs and outcome order stable while evaluating changes. [Review a dataset](/guides/datasets), [tune the wording](/guides/prompt-tuning), and compare the result against your current policy before changing production. Check [fallback](/guides/models-and-fallback) so you know which model actually answered.
