# Conversation-state memory test protocol

This small, reproducible protocol helps you inspect whether an AI chatbot carries useful state forward. It is an evaluation worksheet, not a published benchmark or a claim that any one strategy is better. No model calls or results are included.

## What's included

- [cases.jsonl](cases.jsonl) — twelve fully expanded synthetic conversations: two placements for each of six failure classes.
- [summary-schema.json](summary-schema.json) — a reference state schema for the rolling-summary condition.
- [summary-prompt.md](summary-prompt.md) — an optional fixed summarizer prompt.
- [results-template.csv](results-template.csv) — a blank recording sheet.

## Three reference configurations

Use the same fixed conversation and final probe for each configuration:

1. **Full history:** provide all twelve exchanges.
2. **Recent window:** provide only the final four complete exchanges.
3. **Structured rolling summary:** summarize exchanges as they leave a four-exchange recent window, then provide the summary plus the final four exchanges.

Four exchanges is an experimental setting. It is not a production recommendation. Record the configured window length with your results if you choose a different value.

## Why each pair has two placements

Within a pair, the conversation has the same setup, state update, distractors, neutral exchanges, and probe. The update appears immediately after the setup in the `early` variant, and after eight unrelated exchanges in the `late` variant. Two neutral exchanges follow in both variants. This means the early update is outside the recent window at the probe, while the late update remains inside it.

Inspect placement effects rather than assuming a summary is insensitive to where information appeared.

## Start without provider credits

You can audit the full-history and recent-window inputs manually, using each case’s expected state as the reference. Record only state preservation; leave response correctness, token usage, latency, and cost blank unless you measured them. A hand-authored summary is a target example, not an evaluation of a model-generated rolling summary. The live procedure below is optional and requires your own model access.

## Suggested live procedure

1. Copy a case's messages into your chosen chatbot setup and keep all prompts, model settings, and tools unchanged across the three configurations.
2. For full history, pass the entire conversation to the final probe.
3. For the recent window, pass only exchanges 9–12 to the final probe.
4. For rolling summary, start from an empty summary. Each time an exchange falls outside the most recent four, summarize it into the prior state using the reference prompt and schema. At the final probe, pass that summary and exchanges 9–12.
5. Save the exact memory/context supplied to the final response. Evaluate whether it preserved the expected state before judging the answer.
6. Record one row in `results-template.csv` for each strategy and case. Repeat runs if you want to observe output variation; include the repetition count. For each repetition of the rolling-summary condition, restart from an empty summary and regenerate all eight updates if you want to measure the whole pipeline. Reusing a summary only measures variation in the final response; report that separately.

The cases use simulated actions and receipts. Do not connect them to a real order, payment, email, or other side effect.

The completed-shipment case includes a trusted_receipt field in the update exchange. Treat it as a result returned by the named application tool, not as a user-authored claim. Place the receipt between that exchange’s user request and assistant reply, preserving the provider’s required tool-call/result format. Keep it attached to its exchange when selecting the recent window or feeding the summarizer. When running the case with a tool-capable system, supply that receipt in the same trusted tool-result channel your application normally uses.

## Score two questions separately

**State preserved?** Does the supplied context distinguish current from superseded facts, keep constraints, show which task remains open or complete, and mark unconfirmed information as uncertain?

**Response correct?** Does the final answer follow that state, ask for missing information, and avoid proposing a prohibited or already completed action?

Record each as `yes`, `partial`, or `no`, with a short explanation. A model can answer correctly after the memory has lost the needed fact; that result can be luck. It can also receive a correct state summary and ignore it. Those are different failures. For the assumption cases, merely omitting the shop purpose is compatible with uncertainty, but does not demonstrate that the original guess and its unconfirmed status survived. Distinguish that weaker result in the preservation notes even if the final response correctly says “unknown.”

## Measurements

For every run, record:

- Exact supplied memory and final response.
- Memory-content token count at the probe, excluding the shared system instruction and final question. Prefer the reader model's tokenizer; record its name and version. If it is unavailable, use one consistent tokenizer or estimator and disclose it.
- Reader model/version, settings, and tool configuration.
- Probe latency, input/output tokens, cached-token use, and estimated cost if the provider supplies these values.
- For rolling summary, also record the summarizer model/settings, all summary calls (eight incremental updates for this twelve-exchange fixture), their combined tokens, latency, and estimated cost. Keep summarization cost separate from final-probe cost.
- State-preservation and response-correctness scores, plus any technical errors.

This fixture set is deliberately small. Report counts and denominators, such as `8/12 cases`, and avoid significance claims. Do not pool provider errors or malformed outputs into semantic failures; report them separately.

## Results worksheet

`results-template.csv` captures per-run details. Summarize counts overall, by failure class, and by update placement. Report token counts at the final probe alongside preservation and response outcomes. For cost and latency, distinguish the final probe from the additional work needed to maintain the rolling summary.

An example output table (fill only after running your own evaluation):

| Strategy | Median memory tokens | State preserved | Response correct |
| --- | ---: | ---: | ---: |
| Full history | — | — / 12 | — / 12 |
| Recent four exchanges | — | — / 12 | — / 12 |
| Rolling summary + recent four | — | — / 12 | — / 12 |

The example above is a blank format, not measured data. If comparing five repetitions, report the 60 run denominator per strategy and preserve the 12-case grouping in the class-level view.
