Pallet Labs benchmark
Evaluating Pallet’s memory layer.
Freight operations require client-specific rules and artifacts to make accurate decisions. Memory-Bench measures how our memory layer serves that information to Pallet’s AI agents across seven large language models.
Accuracy with the organization’s own operational memory.
| Condition | Value |
|---|---|
| Without memory | 22% |
| With memory | 62% |
The question this benchmark answers
Pallet builds AI agents for supply chain operations, including the brokerages and carriers that move goods and handle the paperwork that follows them. A single international shipment produces customs filings, purchase orders, arrival notices, and a long email thread, and an operations team spends much of its day reading those documents and keying decisions into other systems. We present Memory-Bench, a formal evaluation of howPallet's agents perform on these tasks with our memory layer, the context that agents have includingthe organization's SOPs, tribal knowledge, and importantly, historical human resolutions of similar tasks.
Public benchmarks mostly test general skills such as writing code or answering exam questions. Memory-Bench tests something narrower: whether an agent makes better decisions when it can use history on how the customer's own team resolved similar cases in the past.
An agent filing a customs document, entering an order from an EDI (Electronic Data Interchange) message, or routing an inbound email periodically hits a decision it cannot make from the document in front of it. In production, that moment becomes an **intervention**: the workflow pauses and a human operator makes the call. The operator usually isn't reasoning from the document either. They know that this importer always files under a particular code in the TMS (Transportation Management System), that this customer's validation errors are routine noise, that seal numbers are kept from the system of record rather than the scanned document. That knowledge is the organization's operational memory, and it lives in people.
The agent platform records every one of those interventions and resolutions, which makes a precise experiment possible: take the decisions humans actually made, replay each one to a frontier model twice, once with only the situation and once with the prior cases our memory system would retrieve from that organization's own history, and grade both against what the human did. The gap between the two runs is how much the organization's memory contributes to the decision.
Four real freight decisions
Every task is a real production intervention from a live customer deployment, spanning three logistics verticals. None are synthetic or hand-written. They are high-volume, individually small calls where being wrong creates rework, demurrage exposure, or a mis-filed customs document.
Memory helps three task families—and exposes one control.
Switch to change to compare the signed effect in percentage points.
| Group | No memory | With memory | Change |
|---|---|---|---|
| ISF importer of record | 0% | 47% | +47 pp |
| Arrival-notice reconciliation | 47% | 76% | +29 pp |
| EDI order triage | 41% | 37% | −4 pp |
| Email routing | 0% | 45% | +45 pp |
What varies inside each family
ISF importer of record. The 35 filings resolve to 22 distinct organization codes. Twelve tasks contain competing prior codes and require discrimination; eight contain exactly one code; in fifteen, the correct code never appears in retrievable history and scores as a coverage miss.
Arrival-notice reconciliation. Twenty conflicts cover seal numbers, pickup numbers, and container type. The correct source changes by field, so there is no “always trust the system” rule—the convention is encoded in the organization’s precedent.
EDI order triage. Eleven orders carry different validation-error sets and an almost even answer key. Eight have no similar prior case, making this family a useful control.
Email routing. Twelve messages map into six categories from an organization-specific fourteen-category taxonomy. These are silver tasks because their production labels have not yet received human verification.
Memory enters at the decision the document cannot make.
Memory-Bench does not score whole workflows. It isolates the single step where automation stalls: the decision.
Every workflow in this study follows the same arc. Work arrives—an email, an EDI message, a scanned customs document—and the agent extracts what the paperwork states. Extraction is rarely the problem. The workflow stalls one step later, when a field requires a judgment the document does not contain: which importer-of-record code to file, which of two conflicting values to trust, whether a validation error is real. Once that field is resolved, the rest is mechanical—the value is written to the system of record and the workflow resumes.
In production, that unresolved field becomes an intervention, and the platform records both the question and the human’s resolution. Each recorded resolution then does double duty: it is the answer key this benchmark grades against, and it is written into the organization’s operational memory, where it becomes retrievable precedent for the next similar decision.
The benchmark replays only that decision point and holds everything else fixed. The model sees the identical situation twice—once with only the documents, once with prior resolved cases retrieved from memory—and nothing upstream or downstream of the decision is ever scored. Trigger handling, extraction quality, and write-back mechanics are all out of frame, so any lift is attributable to a single variable: the retrieved memory input.
The one decision Memory-Bench evaluates
A production agent is a graph of steps. Memory-Bench isolates the one node where the graph cannot decide from the document alone; in production, it pauses for a human. The same graph carries every decision family in the benchmark; hover any node for its role.
| View | trigger | extract | match | decision | resume | human | memory |
|---|---|---|---|---|---|---|---|
| ISF | booking docs arrive | parties · B/Ls · containers | importer ↛ org code | which org code to file? | written back to the TMS | operator types the code | prior filings, by importer |
| Reconciliation | arrival notice scanned | notice fields extracted | TMS ≠ document | which value goes on record? | record updated | operator picks the value | prior conflict resolutions |
| EDI | EDI 204 tender received | order parsed | validation errors raised | resume, or manual entry? | order processed | operator triages the order | prior triage decisions |
| Generic | work arrives | documents parsed | one field unresolved | the org's call? | written back to the system | operator makes the call | prior resolved cases |
How the dataset it built
Tasks are mined from the production records of deployed freight-operations agents across three customer organizations. When an agent workflow cannot make a decision, the run pauses as an intervention and a human operator resolves it; the platform stores the paused state, the retrieved context, and the operator's resolution. Each benchmark task is one such intervention, replayed counterfactually: the decision context is reconstructed as of the pause timestamp, and the operator's recorded resolution becomes the answer key. Almost every task derives from a distinct source document (64 of 65 sources contribute exactly one task). The released set contains 78 tasks in four decision families: ISF importer-of-record assignment (35), arrival-notice field reconciliation (20), EDI-204 order triage — RESUME vs. MANUAL (11), and inbound-email classification (12). Sixty-six tasks are gold (graded against a human resolution); the 12 email tasks are silver (graded against the production classifier's own label) and are excluded from all headline figures. Memory-relevant cases are deliberately over-sampled relative to the production base rate (~22% of interventions have usable history in production; disclosed as a design choice, so headline rates are not population rates). Reconstruction fidelity is verified programmatically (42/42 checks), and each task carries a provenance hash of its source resolution.
Each task is rendered as plain text with three parts: a situation, a retrieved memory block (present only in the with-memory arm), and a fixed question with a JSON answer contract (Respond with ONLY this JSON object: {"<answer_field>": "<youranswer>"}).The situation is a compact natural-language rendering of the workflow state at the moment it paused: what the agent was doing, the field it could not resolve, and the machine-extracted values available to the agent at that moment. For the ISF tasks, the situation contains: a statement that the filing is blocked on the importer of record; the importer name string extracted from the booking documents (e.g., 'Acme Corp FL 12345'); and the TMS matcher's candidate slate as the operator saw it — effectively empty, since production presented a free-text field with no usable candidates in 99.7% of cases, so the situation alone does not determine the answer. For reconciliation, it contains the flagged field (seal number, pickup number, container type) with the conflicting values from the two sources (the value extracted from the arrival notice vs. the value held in the TMS) and the container/BL identifiers.
The memory layer stores several kinds of organizational knowledge; Memory-Bench exercises the episodic kind: the platform's stored records of previously resolved interventions. Each memory is a dated one-to-two-sentence natural-language summary, written by the platform at resolution time, of what a prior intervention was and how the operator resolved it, usually ending with the resolved value, e.g. - [2026-05-07] Human intervention resolved the ambiguous Importer of Record ID by assigning Acme Corp's EIN 91-1234567… (resolved: 91-1234567). The with-memory arm retrieves them from the organization's live memory store, filtered with the same parameters as our deployed agents. These blocks run roughly 250–2,300 characters; when nothing matches, the block states that no relevant prior cases exist — and that empty block still counts against the with-memory score. (The email task family's block instead carries the organization's 14-category routing taxonomy rather than prior cases — one reason that family is silver-tier and headline-excluded.)
Before any model is run, each task is classified from the organization's history: novel (retrieval returns nothing; memory is not useful — the built-in control group with n=20 gold), matched-consistent (prior cases exist and agree with the answer key; n=24 gold), matched-contested (the organization's own operators resolved the same situation in conflicting ways; unwinnable by design; n=21), and matched-unknown (consistency not computed; n=1). The headline set of 44 tasks meets two conditions: the answer key is gold, and the stratum is consistent or novel; contested, unknown, and silver strata are reported in full but machine-excluded from the headline (the exclusion is enforced by an assertion and an audit check, and the pre/post effect of excluding them, pooled 18% to 54% vs. 22% to 62%, is disclosed).
Every task is evaluated twice over an identical situation. No-memory presents the reconstructed situation only. With-memory additionally presents the prior resolved cases returned by the production retrieval path: filter by organization, case type, and entity key; strict time cutoff at the decision timestamp (the task's own resolution can never appear); order by recency; truncate at K=10. If nothing matches, the block is empty and still counts against the with-memory score. The headline variant deliberately omits the source documents production would have, isolating memory's contribution; a full-context variant is reported separately and yields a smaller marginal lift (+14–15 points).
The protocol for each trial is single-turn, all information upfront: a fixed system prompt and a JSON answer contract ({"<answer_field>": "<value>"}), with a 16k output-token ceiling. All models are called through a single API gateway (OpenRouter) with identical prompts and sampling parameters, and each runs at its vendor's default reasoning configuration — the harness sends no reasoning or thinking parameter. In practice the defaults differ by vendor: the two Claude models answer without extended thinking (zero reasoning tokens on every call), while the GPT, Gemini, and GLM models apply their default internal reasoning (means of roughly 85–575 reasoning tokens per call, at depths the models choose).
How answers are graded
Scoring is deterministic exact-match against the operator's recorded resolution, after light normalization: an answer either equals the recorded value or it does not. We deliberately don't use an LLM judge, to defer the ground truth to the organization's employees only. A model-graded benchmark folds a second model's biases and run-to-run variance into the measurement, rewards answers that are fluent rather than correct, and quietly drifts as the judge model is updated, so its numbers stop being reproducible or comparable across time. The human-reference benchmarks we follow reach the same conclusion. SWE-bench (and the human-audited SWE-bench Verified subset) counts a code change as correct only when the repository's own test suite executes and passes, ground truth rather than a judge; and the Data Analysis Benchmark (DAB) scores against a fixed answer key. Deterministic grading is what makes Memory-Bench re-runnable against any new model on the day it ships and comparable to the run before it. The seven-model panel is scored at **greedy pass@1** (temperature 0, one attempt). The reference model, gpt-5.5, additionally runs **50 independent trials** per condition at temperature 0.7, which yields pass@1, pass@10, pass@50, and the reliability curve in Figure 5. **pass@1** is the per-attempt hit rate you would deploy on; **pass@k**, the chance at least one of k attempts is right, is an optimistic ceiling, useful for judging headroom rather than production performance.
Inside an ISF task: retrieve the importer's org code
The ISF (Importer Security Filing) agent reads the booking documents and must file each shipment under the correct importer-of-record org code in the brokerage's TMS. Use the Next button to step through the four stages of the task, from scanned document to filed code. Names, codes, and numbers below are fictional.
The shipment arrives
The booking documents name the importer of record, Meridian Import Partners, LLC, but not its org code in the TMS. The agent extracts the party (highlighted).
Meridian Import Partners, LLC
1450 Harbor Scenic Dr, Long Beach CA
ExtractedThe importer's name is here — its org code is not.
— awaiting org code —
UnresolvedAuto-match doesn't resolve it
The matcher proposes fuzzy candidates from the org registry. None is the code this importer actually files under, which is the common production case.
Meridian Import Partners, LLC
1450 Harbor Scenic Dr, Long Beach CA
Extracted- MERIDGLOBAL
- MERIDIANTX
- IMPORTPTR
None is the code this importer files under — the matcher lands it roughly 1 in 10 times.
— awaiting org code —
UnresolvedMemory supplies the answer
The memory layer retrieves this organization's prior ISF filings for Meridian Import Partners. They were filed under MERIDLB01, knowledge that lives in the operation rather than on the document.
Meridian Import Partners, LLC
1450 Harbor Scenic Dr, Long Beach CA
Extracted- MERIDGLOBAL
- MERIDIANTX
- IMPORTPTR
prior ISF · Meridian Import Partners → MERIDLB01
— awaiting org code —
UnresolvedThe code is filed
The importer-of-record org code is entered on the ISF-10. Memory-Bench grades exactly this string against what the operator filed.
Meridian Import Partners, LLC
1450 Harbor Scenic Dr, Long Beach CA
- MERIDGLOBAL
- MERIDIANTX
- IMPORTPTR
prior ISF · Meridian Import Partners → MERIDLB01
MERIDLB01
Graded string| Step | Description | Result |
|---|---|---|
| 1 · The shipment arrives | The booking documents name the importer of record, Meridian Import Partners, LLC, but not its org code in the TMS. The agent extracts the party (highlighted). | Unresolved |
| 2 · Auto-match doesn't resolve it | The matcher proposes fuzzy candidates from the org registry. None is the code this importer actually files under, which is the common production case. | Unresolved |
| 3 · Memory supplies the answer | The memory layer retrieves this organization's prior ISF filings for Meridian Import Partners. They were filed under MERIDLB01, knowledge that lives in the operation rather than on the document. | Unresolved |
| 4 · The code is filed | The importer-of-record org code is entered on the ISF-10. Memory-Bench grades exactly this string against what the operator filed. | MERIDLB01 |
Memory only moves the cases where precedent exists.
The no-history control is flat; contradictory history is shown but excluded from the headline.
| Group | No memory | With memory | Change |
|---|---|---|---|
| History agrees | 3% | 66% | +63 pp |
| History contradicts itself | 19% | 41% | +22 pp |
| No history · control | 39% | 38% | −1 pp |
Memory lifts every model 36–45 points
Across the panel, first-try accuracy rises from a pooled 22% to 62%. More model capability does not fix the baseline: reasoning power is not a substitute for organization-specific knowledge.
Greedy pass@1 across seven frontier models
44 headline tasks · accuracy 0–100% · cost shown at right
| Group | No memory | With memory | Change |
|---|---|---|---|
| Gemini 3.5 Flash | 23% | 66% | +43 pp |
| Sonnet 5 | 20% | 64% | +43 pp |
| GLM 5.2 | 18% | 64% | +45 pp |
| Gemini 3.1 Pro | 23% | 64% | +41 pp |
| GPT-5.2 | 23% | 61% | +39 pp |
| GPT-5.5 | 25% | 61% | +36 pp |
| Opus 4.8 | 20% | 57% | +36 pp |
Accuracy versus cost, with and without memory
Seven models at greedy pass@1, marked by their lab's logo: muted for no memory, colored for with memory. The line is the efficient frontier within each condition. Up is more accurate, left is cheaper. Toggle the decision family and the cost unit; turn on memory-lift arrows to see each model's jump.
| Family | Model | Condition | Accuracy | Cost | Tokens |
|---|---|---|---|---|---|
| All gold | GLM 5.2 | No memory | 19.7% | $0.0032 | 514 |
| All gold | GLM 5.2 | With memory | 53% | $0.0041 | 833 |
| All gold | GPT-5.2 | No memory | 19.7% | $0.0014 | 228 |
| All gold | GPT-5.2 | With memory | 56.1% | $0.0025 | 551 |
| All gold | GPT-5.5 | No memory | 24.2% | $0.0094 | 433 |
| All gold | GPT-5.5 | With memory | 56.1% | $0.0093 | 665 |
| All gold | Gemini 3.1 Pro | No memory | 21.2% | $0.0047 | 510 |
| All gold | Gemini 3.1 Pro | With memory | 57.6% | $0.0064 | 955 |
| All gold | Gemini 3.5 Flash | No memory | 22.7% | $0.0052 | 702 |
| All gold | Gemini 3.5 Flash | With memory | 57.6% | $0.0067 | 1165 |
| All gold | Opus 4.8 | No memory | 19.7% | $0.0018 | 256 |
| All gold | Opus 4.8 | With memory | 51.5% | $0.0040 | 671 |
| All gold | Sonnet 5 | No memory | 21.2% | $0.0008 | 258 |
| All gold | Sonnet 5 | With memory | 53% | $0.0019 | 713 |
| ISF | GLM 5.2 | No memory | 0% | $0.0042 | 661 |
| ISF | GLM 5.2 | With memory | 42.9% | $0.0045 | 932 |
| ISF | GPT-5.2 | No memory | 0% | $0.0015 | 241 |
| ISF | GPT-5.2 | With memory | 48.6% | $0.0028 | 611 |
| ISF | GPT-5.5 | No memory | 0% | $0.0110 | 494 |
| ISF | GPT-5.5 | With memory | 48.6% | $0.0103 | 733 |
| ISF | Gemini 3.1 Pro | No memory | 0% | $0.0059 | 620 |
| ISF | Gemini 3.1 Pro | With memory | 48.6% | $0.0075 | 1087 |
| ISF | Gemini 3.5 Flash | No memory | 0% | $0.0048 | 666 |
| ISF | Gemini 3.5 Flash | With memory | 48.6% | $0.0071 | 1253 |
| ISF | Opus 4.8 | No memory | 0% | $0.0022 | 289 |
| ISF | Opus 4.8 | With memory | 37.1% | $0.0045 | 761 |
| ISF | Sonnet 5 | No memory | 0% | $0.0009 | 287 |
| ISF | Sonnet 5 | With memory | 40% | $0.0021 | 798 |
| Reconciliation | GLM 5.2 | No memory | 45% | $0.0025 | 408 |
| Reconciliation | GLM 5.2 | With memory | 75% | $0.0050 | 965 |
| Reconciliation | GPT-5.2 | No memory | 40% | $0.0015 | 223 |
| Reconciliation | GPT-5.2 | With memory | 75% | $0.0027 | 613 |
| Reconciliation | GPT-5.5 | No memory | 55% | $0.0093 | 419 |
| Reconciliation | GPT-5.5 | With memory | 75% | $0.0101 | 734 |
| Reconciliation | Gemini 3.1 Pro | No memory | 45% | $0.0043 | 463 |
| Reconciliation | Gemini 3.1 Pro | With memory | 75% | $0.0067 | 1051 |
| Reconciliation | Gemini 3.5 Flash | No memory | 50% | $0.0067 | 858 |
| Reconciliation | Gemini 3.5 Flash | With memory | 80% | $0.0083 | 1408 |
| Reconciliation | Opus 4.8 | No memory | 40% | $0.0014 | 207 |
| Reconciliation | Opus 4.8 | With memory | 80% | $0.0041 | 715 |
| Reconciliation | Sonnet 5 | No memory | 45% | $0.0007 | 219 |
| Reconciliation | Sonnet 5 | With memory | 80% | $0.0023 | 787 |
| EDI | GLM 5.2 | No memory | 36.4% | $0.0011 | 241 |
| EDI | GLM 5.2 | With memory | 45.5% | $0.0012 | 280 |
| EDI | GPT-5.2 | No memory | 45.5% | $0.0009 | 194 |
| EDI | GPT-5.2 | With memory | 45.5% | $0.0012 | 249 |
| EDI | GPT-5.5 | No memory | 45.5% | $0.0043 | 264 |
| EDI | GPT-5.5 | With memory | 45.5% | $0.0051 | 322 |
| EDI | Gemini 3.1 Pro | No memory | 45.5% | $0.0015 | 246 |
| EDI | Gemini 3.1 Pro | With memory | 54.5% | $0.0024 | 359 |
| EDI | Gemini 3.5 Flash | No memory | 45.5% | $0.0037 | 530 |
| EDI | Gemini 3.5 Flash | With memory | 45.5% | $0.0026 | 440 |
| EDI | Opus 4.8 | No memory | 45.5% | $0.0015 | 239 |
| EDI | Opus 4.8 | With memory | 45.5% | $0.0018 | 304 |
| EDI | Sonnet 5 | No memory | 45.5% | $0.0006 | 239 |
| EDI | Sonnet 5 | With memory | 45.5% | $0.0008 | 304 |
Surfacing versus inferring
Where the answer is verbatim-retrievable, memory takes the model from 0% to 76%. Where it must be inferred around a coverage gap, accuracy moves from 31% to 51%. Both are product value, but they measure different things.
Is it reasoning, or copying?
The common failure mode of a memory benchmark is that "using memory" collapses into "copy the most recent thing you were shown." We test that directly on the hardest ISF sub-population: cases where the retrieved history contains **two or more different codes** and the model must select. On those 12 tasks we also score three non-reasoning policies against the same key.
On the hardest ISF cases—where history contains two or more codes—the model reaches 70%. Copying the most recent code reaches 50%, the most frequent reaches 33%, and choosing randomly reaches 34%. Beating the copy-most-recent ceiling by twenty points indicates that the model matches precedent to the specific case rather than repeating whichever code appears first.
The model beats every copy-only policy.
Pass@1 on the ISF tasks that contain two or more competing prior codes.
| Method | Value |
|---|---|
| Copy most recent prior code | 50% |
| Copy most frequent prior code | 33% |
| Pick uniformly at random | 34% |
| Model, with memory | 70% |
What happens with full production context
The headline isolates memory with a thin version of the situation. Production context is richer: the full structured record plus parsed document text. In that variant, the baseline rises and memory’s incremental lift compresses to roughly +14–15 points. Where history exists and agrees, full-context runs still jump from 9% to 83% for ISF and 44% to 73% for reconciliation.
The lift holds when every attempt must succeed
An operations decision gets no credit for being right once in fifty tries. We measure passk: the probability that all k attempts at the same case succeed.
Probability every attempt is correct
Reference model · 50 trials · trustworthy strata
| View | Series | 1 | 2 | 4 | 8 | 16 | 32 |
|---|---|---|---|---|---|---|---|
| Probability | With memory | 60% | 58% | 55% | 55% | 54% | 53% |
| Probability | No memory | 20% | 19% | 18% | 17% | 17% | 15% |
| Drop from k=1 | With memory | 0 pp | 2 pp | 5 pp | 5 pp | 6 pp | 7 pp |
| Drop from k=1 | No memory | 0 pp | 1 pp | 2 pp | 3 pp | 3 pp | 5 pp |
Robustness to presentation order
For forty tasks we reorder retrieved memory entries five ways at temperature zero. Where history agrees, 74% of answers remain stable across every ordering and the flip rate is 6%. When history contradicts itself, stability falls to 41% and the flip rate rises to 30%—exactly where precedent itself is ambiguous.
Does the memory-layer add to LLM costs?
Retrieved cases add 220–430 prompt tokens per decision across the panel. Dollar cost barely moves because memory also lowers the reasoning load.
The reference model’s reasoning tokens fall 16% with memory, from 271 to 229 per decision over fifty trials. On ISF, reasoning tokens drop 22%, and the call costs less in absolute terms. Without memory, the model reasons at length toward a code it has no way to know; with memory, it reads the code from a prior filing and stops.
Memory costs the same to run—and far less per correct result.
Toggle the denominator to see where the economic advantage appears.
| Metric | No memory | With memory |
|---|---|---|
| Per correct | $0.0521 | $0.0178 |
| Per decision | $0.0093 | $0.0094 |
The reasoning saving is model-dependent, but the economics hold across the panel: per-decision spend changes by less than a quarter of a cent while accuracy roughly triples, so cost per correct decision is two to three times lower.
Conclusion
The missing capability was not more intelligence. It was organizational context.
Across real freight decisions, memory changes what models can know, improves whether they answer consistently, and lowers the cost of each correct outcome.
Explore more Pallet Labs research