D1_DATASET_ID set to its ID, start one tuning job:
prompt_tuning.available and the current price in /v1/decisionone first. This example needs a dataset asking the refund decision with at least 60 matching rows and at least five labelled examples of every outcome. The two-row dataset in the format guide is insufficient.
Read the verdict
Poll/v1/prompt-tuning/jobs/{job_id} until succeeded, failed or cancelled. Queued and running jobs may have null progress or result; running progress can show measuring_noise, searching and testing.
A
succeeded job means the evaluation finished. It does not mean the wording improved. Inspect changes, test, label_issues and any skipped reflections. The service does not automatically change your application’s request or repair labels.
What the test measures
Prompt tuning changes the question, outcome descriptions and rubric. It preserves the decision ID, kind, outcome labels and their order. It uses labelled rows with the selected decision’s exact wording; differently worded rows are counted inrows_left_out.
The workflow splits contexts into roughly 40% search, 25% candidate selection and 35% final testing, with size caps. The selected rewrite is tested on held-back contexts against the original, including checks on other decisions in the same requests. The result reports the test evidence and the number of previous looks at the same dataset.
Treat this as evidence for the supplied data. It cannot guarantee better performance on future traffic. Keep a separate evaluation set and inspect changes that affect rare or costly mistakes.
Limits and costs
The current API tunes base models, not named-model adapters. It tunes one decision per job and requires text wording; structured JSON questions are not rewritten. Targets are required rather than reward-only supervision. For smaller datasets, use the free linter. Read the currentstandard and thorough prices from discovery. A failed or cancelled job is refunded. On the first test of a dataset’s rows, a job that cannot prove improvement is also refunded. Repeated tests of the same rows have stricter statistical criteria and are charged regardless of improvement verdict; add genuinely new reviewed data for a fresh evaluation. The job’s price_cents, refunded and result expose what happened.
If wording alone is insufficient, use the same curation discipline before fine-tuning.