D1_API_KEY as in the quickstart and run the same request:
Use the result
Readdecisions.recall.action and map it back to the selected memory. If it is none, omit the candidates. This is a single-choice recipe; if the turn needs multiple memories, design and evaluate a separate selection procedure.
The model selects only among the memories you provide. It cannot retrieve a missing memory, establish that a memory is current, or decide who is authorized to see it. Filter access and expiry before scoring and keep the source IDs for inspection.
Test both useful memories left out and stale or irrelevant memories included. Those errors affect the eventual reply differently. Evaluate the reply with and without selection, including requests where no memory is helpful.
What was evaluated
The historical benchmark uses its own candidate construction and labels. Its accuracy does not transfer automatically to your search index, candidate count or memory format.Historical evaluation on 2026-09-28, using sqwish-d1-core. These measurements describe that checkpoint and dataset, not current production performance or an accuracy guarantee.
- Dataset: MemBench (single-hop, multi-hop and preference questions) (MIT).
- Split: test (the only split), users whose hash falls below 0.5.
- Sample: 600 rows, 600 measured decisions.
- Measured decision IDs:
recall. - Wording: hand-written.
- Checkpoint SHA-256:
dd420ca652cfaa10279eabb54d15afd50fe9d24d357498c4dbc512b1c3d73bb2.
The majority-class baseline has accuracy 0.3333; the class-prior baseline has log loss 1.9279.
Only the decision IDs listed above were measured. Dataset labels, class balance and wording affect these results. The intervals do not measure distribution shift. Re-evaluate with your own cases, including ambiguous and out-of-scope inputs.
