D1_API_KEY as in the quickstart and run the same request:
Use the result
Readdecisions.quality.expected_level on the 0–2 scale, plus the full distribution. Treat a middle value as uncertainty between rubric levels, not as a new categorical label. Establish which levels lead to revision or human review using your reviewers’ own ratings.
Evaluate drafts rated 2 that a person would send back. Make the rubric concrete about missing requirements, factual errors and unusable advice. Quality does not establish groundedness; when sources matter, pair this with the groundedness check.
What was evaluated
The historical accuracy is 0.557, below the majority-class baseline of 0.615. This recipe has not established useful classification performance on that benchmark. Use it as a format and rubric starting point, then improve wording and measure on your own reviewed examples before adopting it.Historical evaluation on 2026-09-28, using sqwish-d1-core. These measurements describe that checkpoint and dataset, not current production performance or an accuracy guarantee.
- Dataset: HelpSteer3 (feedback, validation split) (CC BY 4.0).
- Split: feedback/validation, prompts whose hash falls below 0.55.
- Sample: 1,000 rows, 1,000 measured decisions.
- Measured decision IDs:
quality. - Wording: hand-written.
- Checkpoint SHA-256:
dd420ca652cfaa10279eabb54d15afd50fe9d24d357498c4dbc512b1c3d73bb2.
The majority-class baseline has accuracy 0.615; the class-prior baseline has log loss 0.9319.
Only the decision IDs listed above were measured. Dataset labels, class balance and wording affect these results. The intervals do not measure distribution shift. Re-evaluate with your own cases, including ambiguous and out-of-scope inputs.
