D1_DATASET_ID to its ID and check training_available in /v1/decisionone. Then create a named model’s first version:
id, inspect its resolved hyperparameters and split, and poll /v1/fine-tuning/jobs/{job_id}. The example assumes target-labelled data; reward datasets require method: "reward".
Follow the job
output_model, metrics and progress may be null before they exist. Read /events for lifecycle messages and /records for paginated held-out decisions. A client timeout or stopped poll does not cancel the job. Use the job’s /cancel operation when you intend to stop it.
Leave hyperparameters out initially to use the base model’s defaults. Accepted ranks and input limits depend on that base; being available for inference does not imply that a model is trainable. The job records the limits and split used for training.
Evaluate before promoting
Inspect held-out probabilities beside targets or rewards. For supervised jobs, compare loss, accuracy and mistakes by task and class. For reward jobs, compare expected reward and the scaled-reward gate. A majority-class baseline or a simple rule can be more useful than a single headline accuracy number. The gate compares a version to its base or parent and looks for regressions. It is a safety check on the held-out data, not a guarantee about future traffic. Check your own operational metrics, such as harmful actions or review workload, before promotion.Continue an existing model
For a new dataset, submit a job withparent: "support-router" to start from its production version. The model name follows the parent. The workflow mixes sampled prior training examples with new data and preserves held-out contexts. The replay setting controls the mix; use the default until your evaluations justify changing it.
A manually started job does not promote itself. Follow versions and promotion to load, probe and move production deliberately. Feedback and continual learning can automate later rounds when explicitly enabled.