> ## Documentation Index
> Fetch the complete documentation index at: https://docs.sqwish.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Rank retrieved passages

> Score whether a passage helps answer a particular question.

Use this after retrieval and before generation. Give D1 the question and a candidate passage. Compare candidates using the same rubric.

[Open this recipe in the playground](https://console.sqwish.ai/#playground?recipe=retrieval-relevance), or set `D1_API_KEY` as in the [quickstart](/quickstart) and run the same request:

```bash theme={null}
curl --fail-with-body --silent --show-error https://console.sqwish.ai/v1/decide \
  -H "Authorization: Bearer $D1_API_KEY" \
  -H "Content-Type: application/json" \
  --data-binary @- <<'JSON'
{
  "model": "sqwish-d1-core",
  "context": {
    "query": "When did the Channel Tunnel open to passengers?",
    "passage": "Channel Tunnel. The tunnel was formally opened by Queen Elizabeth II and President François Mitterrand in Calais on 6 May 1994. Passenger services by Eurostar began on 14 November 1994."
  },
  "decisions": [
    {
      "id": "relevance",
      "kind": "ordinal",
      "question": "How well does the passage answer the query?",
      "outcomes": [
        "0",
        "1",
        "2",
        "3"
      ],
      "rubric": [
        "Off topic: it is about something else.",
        "Related: it shares the topic, but nothing in it helps answer the query.",
        "Partly answers: it answers part of the query, or states a fact that helps answer it.",
        "Answers: it states the answer to the query."
      ]
    }
  ]
}
JSON
```

## Use the result

Read `decisions.relevance.expected_level` on the recipe's 0–3 scale and inspect the probability mass across levels. Use it to rank or filter candidates, then evaluate the resulting answers rather than only the ranking scores.

For a simple integration, make one request per passage. If you score several in one context, give each decision a unique ID and explicitly identify its passage. Stay within the model's input and question limits.

A passage can be on topic without containing an answer. Test that distinction, along with partial answers, distracting text, duplicated passages and conflicting sources. Relevance is not source credibility or a guarantee of factual correctness.

## What was evaluated

The historical result measures the labelled relevance decision. It does not measure the downstream answer improvement from your retrieval stack. Compare answer quality with and without this ranking stage on the same held-out questions.

<Note>
  Historical evaluation on **2026-09-28**, using **sqwish-d1-core**. These measurements describe that checkpoint and dataset, not current production performance or an accuracy guarantee.
</Note>

* Dataset: [MIRACL English (dev)](https://huggingface.co/datasets/miracl/miracl) (Apache-2.0 (the MTEB passage copy is CC BY-SA 4.0)).
* Split: dev, queries whose hash falls below 0.5.
* Sample: 1,000 rows, 1,000 measured decisions.
* Measured decision IDs: `relevance`.
* Wording: hand-written.
* Checkpoint SHA-256: `dd420ca652cfaa10279eabb54d15afd50fe9d24d357498c4dbc512b1c3d73bb2`.

| Metric | Value | Recorded interval |
| - | -: | - |
| accuracy | 0.78 | 0.7411 to 0.815 |
| log loss | 0.4593 | 0.414 to 0.5125 |
| brier | 0.2977 | 0.262 to 0.3393 |
| ece | 0.0445 | 0.0291 to 0.0763 |

The majority-class baseline has accuracy **0.724**; the class-prior baseline has log loss **0.5891**.

Only the decision IDs listed above were measured. Dataset labels, class balance and wording affect these results. The intervals do not measure distribution shift. Re-evaluate with your own cases, including ambiguous and out-of-scope inputs.

## Improve it for your application

Keep the decision IDs and outcome order stable while evaluating changes. [Review a dataset](/guides/datasets), [tune the wording](/guides/prompt-tuning), and compare the result against your current policy before changing production. Check [fallback](/guides/models-and-fallback) so you know which model actually answered.
