D1_API_KEY as in the quickstart and run the same request:
Use the result
Readdecisions.session.action. Route an abstention to a general workflow or ask the user to clarify. Otherwise select the application workflow associated with that category; keep the original user request available to it.
A request containing code is not necessarily a coding task: it may ask for a summary. Define categories by the work requested, not keywords or formatting. Evaluate mixed tasks, short messages, different languages and mid-conversation changes.
The sample abstention threshold is a starting policy. Tune it on your labelled traffic, balancing wrong routes against the cost of a general workflow. Avoid hiding capabilities from the user solely because the first message was classified incorrectly.
What was evaluated
The historical benchmark evaluates thesession category on its specified public data. It does not establish end-to-end workflow success or accuracy on long conversations.
Historical evaluation on 2026-09-28, using sqwish-d1-core. These measurements describe that checkpoint and dataset, not current production performance or an accuracy guarantee.
- Dataset: No Robots (CC BY-NC 4.0).
- Split: test.
- Sample: 492 rows, 492 measured decisions.
- Measured decision IDs:
session. - Wording: tuned.
- Checkpoint SHA-256:
dd420ca652cfaa10279eabb54d15afd50fe9d24d357498c4dbc512b1c3d73bb2.
The majority-class baseline has accuracy 0.5569; the class-prior baseline has log loss 1.2393.
Only the decision IDs listed above were measured. Dataset labels, class balance and wording affect these results. The intervals do not measure distribution shift. Re-evaluate with your own cases, including ambiguous and out-of-scope inputs.
