Skip to main content
Use this before planning a tool call. Include the request and the available tool descriptions, then let the decision choose one tool or none. Open this recipe in the playground, or set D1_API_KEY as in the quickstart and run the same request:

Use the result

Read decisions.tool.action. If abstained is true, use the full planner or ask for clarification. If the action is none, continue without a tool. Otherwise pass the selected tool’s name to your planner so it can construct and validate arguments. Choosing a tool does not validate its arguments or authorize its execution. Check permissions, required fields and side effects separately. The tool-risk example evaluates a proposed call at that later boundary. The recipe is a single-choice task. Requests requiring multiple tools need a planner or another step. Evaluate overlapping descriptions, unsupported requests and changes to your tool catalogue; do not assume a score from these five sample tools transfers to a larger registry.

What was evaluated

The historical measurement covers the declared tool choice, not successful execution or task completion. Compare selection accuracy and abstention with your planner on your own requests.
Historical evaluation on 2026-09-28, using sqwish-d1-core. These measurements describe that checkpoint and dataset, not current production performance or an accuracy guarantee.
  • Dataset: BFCL live (multiple functions and irrelevance) (Apache-2.0).
  • Split: live_multiple and live_irrelevance (test only), tool sets whose hash falls below 0.5.
  • Sample: 1,000 rows, 1,000 measured decisions.
  • Measured decision IDs: tool.
  • Wording: hand-written.
  • Checkpoint SHA-256: dd420ca652cfaa10279eabb54d15afd50fe9d24d357498c4dbc512b1c3d73bb2.
The majority-class baseline has accuracy 0.414; the class-prior baseline has log loss 2.9656. Only the decision IDs listed above were measured. Dataset labels, class balance and wording affect these results. The intervals do not measure distribution shift. Re-evaluate with your own cases, including ambiguous and out-of-scope inputs.

Improve it for your application

Keep the decision IDs and outcome order stable while evaluating changes. Review a dataset, tune the wording, and compare the result against your current policy before changing production. Check fallback so you know which model actually answered.