> ## Documentation Index
> Fetch the complete documentation index at: https://docs.sqwish.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Tune the wording

> Test whether clearer questions and descriptions improve one decision on your labelled data.

With a reviewed dataset and `D1_DATASET_ID` set to its ID, start one tuning job:

```bash theme={null}
curl --fail-with-body https://console.sqwish.ai/v1/prompt-tuning/jobs \
  -H "Authorization: Bearer $D1_API_KEY" -H 'Content-Type: application/json' \
  -H 'Idempotency-Key: refund-wording-v1' \
  -d "{\"dataset_id\":\"$D1_DATASET_ID\",\"decision\":\"refund\",\"model\":\"sqwish-d1-core\",\"effort\":\"standard\"}"
```

Check `prompt_tuning.available` and the current price in `/v1/decisionone` first. This example needs a dataset asking the `refund` decision with at least 60 matching rows and at least five labelled examples of every outcome. The two-row dataset in the format guide is insufficient.

## Read the verdict

Poll `/v1/prompt-tuning/jobs/{job_id}` until `succeeded`, `failed` or `cancelled`. Queued and running jobs may have null `progress` or `result`; running progress can show `measuring_noise`, `searching` and `testing`.

| Verdict | What to do |
| - | - |
| `improved` | Review the changes and use `result.tuned` in new requests. |
| `inconclusive` | Keep your original wording; the candidate did not establish an improvement. |
| `no_improvement` | Keep your original wording; there may be no candidate worth applying. |

A `succeeded` job means the evaluation finished. It does not mean the wording improved. Inspect `changes`, `test`, `label_issues` and any skipped reflections. The service does not automatically change your application's request or repair labels.

## What the test measures

Prompt tuning changes the question, outcome descriptions and rubric. It preserves the decision ID, kind, outcome labels and their order. It uses labelled rows with the selected decision's exact wording; differently worded rows are counted in `rows_left_out`.

The workflow splits contexts into roughly 40% search, 25% candidate selection and 35% final testing, with size caps. The selected rewrite is tested on held-back contexts against the original, including checks on other decisions in the same requests. The result reports the test evidence and the number of previous looks at the same dataset.

Treat this as evidence for the supplied data. It cannot guarantee better performance on future traffic. Keep a separate evaluation set and inspect changes that affect rare or costly mistakes.

## Limits and costs

The current API tunes base models, not named-model adapters. It tunes one decision per job and requires text wording; structured JSON questions are not rewritten. Targets are required rather than reward-only supervision. For smaller datasets, use the free [linter](/guides/writing-decisions).

Read the current `standard` and `thorough` prices from discovery. A failed or cancelled job is refunded. On the first test of a dataset's rows, a job that cannot prove improvement is also refunded. Repeated tests of the same rows have stricter statistical criteria and are charged regardless of improvement verdict; add genuinely new reviewed data for a fresh evaluation. The job's `price_cents`, `refunded` and result expose what happened.

If wording alone is insufficient, use the same curation discipline before [fine-tuning](/guides/fine-tuning).
