Skip to main content
Use this on a session’s opening message, and again when the user changes tasks. It distinguishes code, questions, text transformations, writing and chat. Open this recipe in the playground, or set D1_API_KEY as in the quickstart and run the same request:

Use the result

Read decisions.session.action. Route an abstention to a general workflow or ask the user to clarify. Otherwise select the application workflow associated with that category; keep the original user request available to it. A request containing code is not necessarily a coding task: it may ask for a summary. Define categories by the work requested, not keywords or formatting. Evaluate mixed tasks, short messages, different languages and mid-conversation changes. The sample abstention threshold is a starting policy. Tune it on your labelled traffic, balancing wrong routes against the cost of a general workflow. Avoid hiding capabilities from the user solely because the first message was classified incorrectly.

What was evaluated

The historical benchmark evaluates the session category on its specified public data. It does not establish end-to-end workflow success or accuracy on long conversations.
Historical evaluation on 2026-09-28, using sqwish-d1-core. These measurements describe that checkpoint and dataset, not current production performance or an accuracy guarantee.
  • Dataset: No Robots (CC BY-NC 4.0).
  • Split: test.
  • Sample: 492 rows, 492 measured decisions.
  • Measured decision IDs: session.
  • Wording: tuned.
  • Checkpoint SHA-256: dd420ca652cfaa10279eabb54d15afd50fe9d24d357498c4dbc512b1c3d73bb2.
The majority-class baseline has accuracy 0.5569; the class-prior baseline has log loss 1.2393. Only the decision IDs listed above were measured. Dataset labels, class balance and wording affect these results. The intervals do not measure distribution shift. Re-evaluate with your own cases, including ambiguous and out-of-scope inputs.

Improve it for your application

Keep the decision IDs and outcome order stable while evaluating changes. Review a dataset, tune the wording, and compare the result against your current policy before changing production. Check fallback so you know which model actually answered.