decisions.route.abstained is true, send the case to your review path. Otherwise use action. The threshold is illustrative; select it using the cost of mistakes and the review volume on your own data.
Raw scores and policy outputs
p_top, margin and entropy still describe the raw distribution when a policy changes the action. A high score can be confidently wrong, particularly with missing alternatives or unfamiliar inputs. Normalization is not calibration.
Weights, costs and abstention
Weights multiply the probabilities and renormalize them. If you also supply costs,costs[action][outcome] describes the cost of choosing that action when the true outcome is the named outcome. D1 chooses the lowest expected-cost action under the policy distribution.
expected_costs reports each action’s expected cost. expected_cost reports the chosen cost, including a declared deferral cost when it triggers abstention. Cost-based abstention requires a cost matrix.
abstain.min_probability checks the largest policy probability. abstain.cost defers when the best action costs more than deferring. These values express your application policy, not universal confidence thresholds. The tool-risk example shows a cost-sensitive decision.
Which model answered?
Readmodel, any exact named version, and fallback. A fallback can change the model and its error profile. provenance describes the serving artifacts when present; it is not an accuracy certificate. timing_ms contains server stages, not the full network round trip or a latency guarantee.
For an evaluation, record the complete request, model/version, actual fallback, labels and policy. Measure both raw accuracy and operational outcomes such as wrong actions and review rate. The example pages keep historical measurements tied to their source dataset and checkpoint.