Skip to main content
Use this after your own memory search returns candidates. Supply the current request and candidate memories with stable IDs. Open this recipe in the playground, or set D1_API_KEY as in the quickstart and run the same request:

Use the result

Read decisions.recall.action and map it back to the selected memory. If it is none, omit the candidates. This is a single-choice recipe; if the turn needs multiple memories, design and evaluate a separate selection procedure. The model selects only among the memories you provide. It cannot retrieve a missing memory, establish that a memory is current, or decide who is authorized to see it. Filter access and expiry before scoring and keep the source IDs for inspection. Test both useful memories left out and stale or irrelevant memories included. Those errors affect the eventual reply differently. Evaluate the reply with and without selection, including requests where no memory is helpful.

What was evaluated

The historical benchmark uses its own candidate construction and labels. Its accuracy does not transfer automatically to your search index, candidate count or memory format.
Historical evaluation on 2026-09-28, using sqwish-d1-core. These measurements describe that checkpoint and dataset, not current production performance or an accuracy guarantee.
  • Dataset: MemBench (single-hop, multi-hop and preference questions) (MIT).
  • Split: test (the only split), users whose hash falls below 0.5.
  • Sample: 600 rows, 600 measured decisions.
  • Measured decision IDs: recall.
  • Wording: hand-written.
  • Checkpoint SHA-256: dd420ca652cfaa10279eabb54d15afd50fe9d24d357498c4dbc512b1c3d73bb2.
The majority-class baseline has accuracy 0.3333; the class-prior baseline has log loss 1.9279. Only the decision IDs listed above were measured. Dataset labels, class balance and wording affect these results. The intervals do not measure distribution shift. Re-evaluate with your own cases, including ambiguous and out-of-scope inputs.

Improve it for your application

Keep the decision IDs and outcome order stable while evaluating changes. Review a dataset, tune the wording, and compare the result against your current policy before changing production. Check fallback so you know which model actually answered.