allow, confirm or deny, then applies a cost matrix.
Open this recipe in the playground, or set D1_API_KEY as in the quickstart and run the same request:
Use the result
Readdecisions.action.action, not just top. The cost matrix produces application actions run, ask or refuse; abstention produces abstain. Send ask and abstain to a human review path.
The distinction matters: the most probable class can differ from the least costly action. expected_costs exposes how the policy reached its choice. These costs are illustrative relative penalties, not measured monetary losses. Review them against your own consequences and tolerance for interruptions.
Keep deterministic permission checks, sandboxing and resource limits. A classifier can miss malicious instructions hidden in tool output and cannot prove a shell command or external action is safe. Include adversarial, indirect-injection and misleadingly benign cases in your evaluation.
What was evaluated
The measured class accuracy below does not measure the cost policy’s safety, harmful execution rate or human approval behavior. Track dangerous calls that were allowed and dangerous calls that people approved after a prompt, as well as unnecessary refusals.Historical evaluation on 2026-09-28, using sqwish-d1-core. These measurements describe that checkpoint and dataset, not current production performance or an accuracy guarantee.
- Dataset: TS-Bench ASB-Traj (ToolSafe) (MIT (the repository’s licence)).
- Split: asb-traj/test (indirect-injection and failed-attack runs), the half of user tasks whose hash falls below 0.5.
- Sample: 1,000 rows, 1,000 measured decisions.
- Measured decision IDs:
action. - Wording: hand-written.
- Checkpoint SHA-256:
dd420ca652cfaa10279eabb54d15afd50fe9d24d357498c4dbc512b1c3d73bb2.
The majority-class baseline has accuracy 0.564; the class-prior baseline has log loss 0.9643.
Only the decision IDs listed above were measured. Dataset labels, class balance and wording affect these results. The intervals do not measure distribution shift. Re-evaluate with your own cases, including ambiguous and out-of-scope inputs.
