> ## Documentation Index
> Fetch the complete documentation index at: https://docs.sqwish.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Select a tool

> Narrow an agent’s tool choice while preserving a path for ambiguity.

Use this before planning a tool call. Include the request and the available tool descriptions, then let the decision choose one tool or `none`.

[Open this recipe in the playground](https://console.sqwish.ai/#playground?recipe=tool-choice), or set `D1_API_KEY` as in the [quickstart](/quickstart) and run the same request:

```bash theme={null}
curl --fail-with-body --silent --show-error https://console.sqwish.ai/v1/decide \
  -H "Authorization: Bearer $D1_API_KEY" \
  -H "Content-Type: application/json" \
  --data-binary @- <<'JSON'
{
  "model": "sqwish-d1-core",
  "context": {
    "request": "Can you book us a table for four at Dishoom King's Cross this Friday at 7pm? Put it under Priya."
  },
  "decisions": [
    {
      "id": "tool",
      "kind": "single",
      "question": "Which tool should the agent call for this request?",
      "outcomes": {
        "search_restaurants": "Find restaurants by area, cuisine, price or opening hours, and return their names and ratings.",
        "book_table": "Reserve a table at a named restaurant for a party size, date and time, under a guest's name.",
        "get_directions": "Give walking, driving or public transport directions between two places.",
        "send_message": "Send a text or chat message to a contact.",
        "none": "None of these tools does what the user asked."
      },
      "abstain": {
        "min_probability": 0.5
      }
    }
  ]
}
JSON
```

## Use the result

Read `decisions.tool.action`. If `abstained` is true, use the full planner or ask for clarification. If the action is `none`, continue without a tool. Otherwise pass the selected tool's name to your planner so it can construct and validate arguments.

Choosing a tool does not validate its arguments or authorize its execution. Check permissions, required fields and side effects separately. The [tool-risk example](/examples/tool-risk) evaluates a proposed call at that later boundary.

The recipe is a single-choice task. Requests requiring multiple tools need a planner or another step. Evaluate overlapping descriptions, unsupported requests and changes to your tool catalogue; do not assume a score from these five sample tools transfers to a larger registry.

## What was evaluated

The historical measurement covers the declared tool choice, not successful execution or task completion. Compare selection accuracy and abstention with your planner on your own requests.

<Note>
  Historical evaluation on **2026-09-28**, using **sqwish-d1-core**. These measurements describe that checkpoint and dataset, not current production performance or an accuracy guarantee.
</Note>

* Dataset: [BFCL live (multiple functions and irrelevance)](https://github.com/ShishirPatil/gorilla/tree/main/berkeley-function-call-leaderboard) (Apache-2.0).
* Split: live\_multiple and live\_irrelevance (test only), tool sets whose hash falls below 0.5.
* Sample: 1,000 rows, 1,000 measured decisions.
* Measured decision IDs: `tool`.
* Wording: hand-written.
* Checkpoint SHA-256: `dd420ca652cfaa10279eabb54d15afd50fe9d24d357498c4dbc512b1c3d73bb2`.

| Metric | Value | Recorded interval |
| - | -: | - |
| accuracy | 0.789 | 0.7278 to 0.847 |
| log loss | 0.5282 | 0.4186 to 0.6354 |
| brier | 0.3142 | 0.2353 to 0.3981 |
| ece | 0.0987 | 0.0541 to 0.1558 |

The majority-class baseline has accuracy **0.414**; the class-prior baseline has log loss **2.9656**.

Only the decision IDs listed above were measured. Dataset labels, class balance and wording affect these results. The intervals do not measure distribution shift. Re-evaluate with your own cases, including ambiguous and out-of-scope inputs.

## Improve it for your application

Keep the decision IDs and outcome order stable while evaluating changes. [Review a dataset](/guides/datasets), [tune the wording](/guides/prompt-tuning), and compare the result against your current policy before changing production. Check [fallback](/guides/models-and-fallback) so you know which model actually answered.
