Skip to main content
Use the free linter before running a model:
Read each warning’s decision, rule and message. An empty warnings list means no lint rule fired; it is not an accuracy test. Send the same body to /v1/decide with your key to score it.

Define the boundary

Ask one question per decision. Describe outcomes so another person could label the same case consistently. Include an outcome for cases outside the main categories where that is useful. Define what counts as evidence and how to handle missing information. For example, “Is this urgent?” leaves a business rule unstated. “Does the customer state a deadline within the next 24 hours?” gives a reviewer and the model a clearer boundary. A short, specific description is usually easier to evaluate than a long list of overlapping instructions.

Choose a kind

An ordinal result’s expected_level uses zero-based positions, not the numerical values or spacing of your labels. For low/medium/high it is on a 0–2 scale. Keep the order fixed when evaluating or updating wording. Decision and outcome IDs should be stable application identifiers.

Supply the evidence

context and question can be strings or JSON objects/arrays. For retrieval or tool selection, include the actual source passage or tool descriptions in the context. A decision cannot verify information it was not given. Ask related questions together when they need the same context. The default request allows them to share a rendering; independent: true requests independent rendering. This is a serving choice, not a claim of statistical independence or identical outputs.

Separate scores from actions

Keep the meaning of the outcomes stable. Add weights, an action costs matrix, or abstain only when you have a reason grounded in your workflow. Validate the resulting action policy separately from the raw classifier. The results guide explains how these fields interact. Build a small reviewed set with ordinary cases, rare classes, ambiguous messages and cases outside the intended task. Compare against a simple baseline. Improve wording first, then use prompt tuning or fine-tuning when you have enough representative labels.