> ## Documentation Index
> Fetch the complete documentation index at: https://docs.sqwish.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Classify a request

> Route a session to the workflow that fits its current task.

Use this on a session's opening message, and again when the user changes tasks. It distinguishes code, questions, text transformations, writing and chat.

[Open this recipe in the playground](https://console.sqwish.ai/#playground?recipe=session-type), or set `D1_API_KEY` as in the [quickstart](/quickstart) and run the same request:

```bash theme={null}
curl --fail-with-body --silent --show-error https://console.sqwish.ai/v1/decide \
  -H "Authorization: Bearer $D1_API_KEY" \
  -H "Content-Type: application/json" \
  --data-binary @- <<'JSON'
{
  "model": "sqwish-d1-core",
  "context": {
    "message": "Can you make this paragraph shorter and more confident? \"I think that I might be a good fit for the role, as I have worked on some projects that were somewhat similar, and I would hope to learn quickly.\""
  },
  "decisions": [
    {
      "id": "session",
      "kind": "single",
      "question": "What kind of work does this request need?",
      "outcomes": {
        "code": "Writing, fixing, explaining or reviewing code, scripts or queries.",
        "question": "Choose question for factual information from general knowledge. If the user requests new content or ideas, choose writing.",
        "text_task": "Working on text the user supplies: summarise, rewrite, extract, classify or answer questions about it.",
        "writing": "Writing something new, such as a story, email, poem or post, or listing ideas.",
        "chat": "Keeping up a conversation or playing a character, often set up by a system prompt."
      },
      "abstain": {
        "min_probability": 0.4
      }
    }
  ]
}
JSON
```

## Use the result

Read `decisions.session.action`. Route an abstention to a general workflow or ask the user to clarify. Otherwise select the application workflow associated with that category; keep the original user request available to it.

A request containing code is not necessarily a coding task: it may ask for a summary. Define categories by the work requested, not keywords or formatting. Evaluate mixed tasks, short messages, different languages and mid-conversation changes.

The sample abstention threshold is a starting policy. Tune it on your labelled traffic, balancing wrong routes against the cost of a general workflow. Avoid hiding capabilities from the user solely because the first message was classified incorrectly.

## What was evaluated

The historical benchmark evaluates the `session` category on its specified public data. It does not establish end-to-end workflow success or accuracy on long conversations.

<Note>
  Historical evaluation on **2026-09-28**, using **sqwish-d1-core**. These measurements describe that checkpoint and dataset, not current production performance or an accuracy guarantee.
</Note>

* Dataset: [No Robots](https://huggingface.co/datasets/HuggingFaceH4/no_robots) (CC BY-NC 4.0).
* Split: test.
* Sample: 492 rows, 492 measured decisions.
* Measured decision IDs: `session`.
* Wording: tuned.
* Checkpoint SHA-256: `dd420ca652cfaa10279eabb54d15afd50fe9d24d357498c4dbc512b1c3d73bb2`.

| Metric | Value | Recorded interval |
| - | -: | - |
| accuracy | 0.8882 | 0.8598 to 0.9167 |
| log loss | 0.3387 | 0.2833 to 0.3953 |
| brier | 0.1706 | 0.1354 to 0.2063 |
| ece | 0.0429 | 0.0331 to 0.0714 |

The majority-class baseline has accuracy **0.5569**; the class-prior baseline has log loss **1.2393**.

Only the decision IDs listed above were measured. Dataset labels, class balance and wording affect these results. The intervals do not measure distribution shift. Re-evaluate with your own cases, including ambiguous and out-of-scope inputs.

## Improve it for your application

Keep the decision IDs and outcome order stable while evaluating changes. [Review a dataset](/guides/datasets), [tune the wording](/guides/prompt-tuning), and compare the result against your current policy before changing production. Check [fallback](/guides/models-and-fallback) so you know which model actually answered.
