Memory-Bench / 01 Research brief · July 2026

Pallet Labs benchmark

Evaluating Pallet’s memory layer.

Freight operations require client-specific rules and artifacts to make accurate decisions. Memory-Bench measures how our memory layer serves that information to Pallet’s AI agents across seven large language models.

Headline result · greedy pass@1 22%62%

Accuracy with the organization’s own operational memory.

Headline result · greedy pass@1
ConditionValue
Without memory22%
With memory62%
01 / Question

The question this benchmark answers

Pallet builds AI agents for supply chain operations, including the brokerages and carriers that move goods and handle the paperwork that follows them. A single international shipment produces customs filings, purchase orders, arrival notices, and a long email thread, and an operations team spends much of its day reading those documents and keying decisions into other systems. We present Memory-Bench, a formal evaluation of howPallet's agents perform on these tasks with our memory layer, the context that agents have includingthe organization's SOPs, tribal knowledge, and importantly, historical human resolutions of similar tasks.


Public benchmarks mostly test general skills such as writing code or answering exam questions. Memory-Bench tests something narrower: whether an agent makes better decisions when it can use history on how the customer's own team resolved similar cases in the past.

An agent filing a customs document, entering an order from an EDI (Electronic Data Interchange) message, or routing an inbound email periodically hits a decision it cannot make from the document in front of it. In production, that moment becomes an **intervention**: the workflow pauses and a human operator makes the call. The operator usually isn't reasoning from the document either. They know that this importer always files under a particular code in the TMS (Transportation Management System), that this customer's validation errors are routine noise, that seal numbers are kept from the system of record rather than the scanned document. That knowledge is the organization's operational memory, and it lives in people.

The agent platform records every one of those interventions and resolutions, which makes a precise experiment possible: take the decisions humans actually made, replay each one to a frontier model twice, once with only the situation and once with the prior cases our memory system would retrieve from that organization's own history, and grade both against what the human did. The gap between the two runs is how much the organization's memory contributes to the decision.

02 / Tasks

Four real freight decisions

Every task is a real production intervention from a live customer deployment, spanning three logistics verticals. None are synthetic or hand-written. They are high-volume, individually small calls where being wrong creates rework, demurrage exposure, or a mis-filed customs document.

Memory helps three task families—and exposes one control.

Switch to change to compare the signed effect in percentage points.

Chart view
Legend
No memory With memory
ISF importer of record 35 tasks · Customs brokerage
+47 pp change
Arrival-notice reconciliation 20 tasks · Customs brokerage
+29 pp change
EDI order triage 11 tasks · Intermodal drayage
−4 pp change
Email routing 12 tasks · Truckload carrier · silver
+45 pp change
The EDI family is the built-in control: eight of eleven cases have no similar prior resolution to retrieve. Email-routing tasks use silver production labels.
Memory helps three task families—and exposes one control.
GroupNo memoryWith memoryChange
ISF importer of record 0% 47% +47 pp
Arrival-notice reconciliation 47% 76% +29 pp
EDI order triage 41% 37% −4 pp
Email routing 0% 45% +45 pp

What varies inside each family

ISF importer of record. The 35 filings resolve to 22 distinct organization codes. Twelve tasks contain competing prior codes and require discrimination; eight contain exactly one code; in fifteen, the correct code never appears in retrievable history and scores as a coverage miss.

Arrival-notice reconciliation. Twenty conflicts cover seal numbers, pickup numbers, and container type. The correct source changes by field, so there is no “always trust the system” rule—the convention is encoded in the organization’s precedent.

EDI order triage. Eleven orders carry different validation-error sets and an almost even answer key. Eight have no similar prior case, making this family a useful control.

Email routing. Twelve messages map into six categories from an organization-specific fourteen-category taxonomy. These are silver tasks because their production labels have not yet received human verification.

03 / Evaluated node

Memory enters at the decision the document cannot make.

Memory-Bench does not score whole workflows. It isolates the single step where automation stalls: the decision.

Every workflow in this study follows the same arc. Work arrives—an email, an EDI message, a scanned customs document—and the agent extracts what the paperwork states. Extraction is rarely the problem. The workflow stalls one step later, when a field requires a judgment the document does not contain: which importer-of-record code to file, which of two conflicting values to trust, whether a validation error is real. Once that field is resolved, the rest is mechanical—the value is written to the system of record and the workflow resumes.

In production, that unresolved field becomes an intervention, and the platform records both the question and the human’s resolution. Each recorded resolution then does double duty: it is the answer key this benchmark grades against, and it is written into the organization’s operational memory, where it becomes retrievable precedent for the next similar decision.

The benchmark replays only that decision point and holds everything else fixed. The model sees the identical situation twice—once with only the documents, once with prior resolved cases retrieved from memory—and nothing upstream or downstream of the decision is ever scored. Trigger handling, extraction quality, and write-back mechanics are all out of frame, so any lift is attributable to a single variable: the retrieved memory input.

The one decision Memory-Bench evaluates

A production agent is a graph of steps. Memory-Bench isolates the one node where the graph cannot decide from the document alone; in production, it pauses for a human. The same graph carries every decision family in the benchmark; hover any node for its role.

Decision family
Decision family
uncertain · run pauses retrieved on the next similar decision Trigger booking docs arrive Extract & parse parties · B/Ls · containers Match & validate importer ↛ org code Decision node which org code to file? Resume & write written back to the TMS Human intervention operator types the code Answer key · gold Memory prior filings, by importer entity · type · recency Which org code is the importer of record for this shipment?
Figure 1. An agent's workflow as a graph. When the decision node cannot decide, the run pauses for a human intervention; the operator's resolution becomes the benchmark's answer key and is written to the organization's memory, where it is retrieved (the dashed edge) on the next similar decision. Memory-Bench replays only the highlighted node, once with and once without that dashed input; the rest of the run is never scored.
The one decision Memory-Bench evaluates
Viewtriggerextractmatchdecisionresumehumanmemory
ISF booking docs arriveparties · B/Ls · containersimporter ↛ org codewhich org code to file?written back to the TMSoperator types the codeprior filings, by importer
Reconciliation arrival notice scannednotice fields extractedTMS ≠ documentwhich value goes on record?record updatedoperator picks the valueprior conflict resolutions
EDI EDI 204 tender receivedorder parsedvalidation errors raisedresume, or manual entry?order processedoperator triages the orderprior triage decisions
Generic work arrivesdocuments parsedone field unresolvedthe org's call?written back to the systemoperator makes the callprior resolved cases

04 / Dataset

How the dataset it built

Tasks are mined from the production records of deployed freight-operations agents across three customer organizations. When an agent workflow cannot make a decision, the run pauses as an intervention and a human operator resolves it; the platform stores the paused state, the retrieved context, and the operator's resolution. Each benchmark task is one such intervention, replayed counterfactually: the decision context is reconstructed as of the pause timestamp, and the operator's recorded resolution becomes the answer key. Almost every task derives from a distinct source document (64 of 65 sources contribute exactly one task). The released set contains 78 tasks in four decision families: ISF importer-of-record assignment (35), arrival-notice field reconciliation (20), EDI-204 order triage — RESUME vs. MANUAL (11), and inbound-email classification (12). Sixty-six tasks are gold (graded against a human resolution); the 12 email tasks are silver (graded against the production classifier's own label) and are excluded from all headline figures. Memory-relevant cases are deliberately over-sampled relative to the production base rate (~22% of interventions have usable history in production; disclosed as a design choice, so headline rates are not population rates). Reconstruction fidelity is verified programmatically (42/42 checks), and each task carries a provenance hash of its source resolution.

Each task is rendered as plain text with three parts: a situation, a retrieved memory block (present only in the with-memory arm), and a fixed question with a JSON answer contract (Respond with ONLY this JSON object: {"<answer_field>": "<youranswer>"}).The situation is a compact natural-language rendering of the workflow state at the moment it paused: what the agent was doing, the field it could not resolve, and the machine-extracted values available to the agent at that moment. For the ISF tasks, the situation contains: a statement that the filing is blocked on the importer of record; the importer name string extracted from the booking documents (e.g., 'Acme Corp FL 12345'); and the TMS matcher's candidate slate as the operator saw it — effectively empty, since production presented a free-text field with no usable candidates in 99.7% of cases, so the situation alone does not determine the answer. For reconciliation, it contains the flagged field (seal number, pickup number, container type) with the conflicting values from the two sources (the value extracted from the arrival notice vs. the value held in the TMS) and the container/BL identifiers.

The memory layer stores several kinds of organizational knowledge; Memory-Bench exercises the episodic kind: the platform's stored records of previously resolved interventions. Each memory is a dated one-to-two-sentence natural-language summary, written by the platform at resolution time, of what a prior intervention was and how the operator resolved it, usually ending with the resolved value, e.g. - [2026-05-07] Human intervention resolved the ambiguous Importer of Record ID by assigning Acme Corp's EIN 91-1234567… (resolved: 91-1234567). The with-memory arm retrieves them from the organization's live memory store, filtered with the same parameters as our deployed agents. These blocks run roughly 250–2,300 characters; when nothing matches, the block states that no relevant prior cases exist — and that empty block still counts against the with-memory score. (The email task family's block instead carries the organization's 14-category routing taxonomy rather than prior cases — one reason that family is silver-tier and headline-excluded.)

Before any model is run, each task is classified from the organization's history: novel (retrieval returns nothing; memory is not useful — the built-in control group with n=20 gold), matched-consistent (prior cases exist and agree with the answer key; n=24 gold), matched-contested (the organization's own operators resolved the same situation in conflicting ways; unwinnable by design; n=21), and matched-unknown (consistency not computed; n=1). The headline set of 44 tasks meets two conditions: the answer key is gold, and the stratum is consistent or novel; contested, unknown, and silver strata are reported in full but machine-excluded from the headline (the exclusion is enforced by an assertion and an audit check, and the pre/post effect of excluding them, pooled 18% to 54% vs. 22% to 62%, is disclosed).
Every task is evaluated twice over an identical situation. No-memory presents the reconstructed situation only. With-memory additionally presents the prior resolved cases returned by the production retrieval path: filter by organization, case type, and entity key; strict time cutoff at the decision timestamp (the task's own resolution can never appear); order by recency; truncate at K=10. If nothing matches, the block is empty and still counts against the with-memory score. The headline variant deliberately omits the source documents production would have, isolating memory's contribution; a full-context variant is reported separately and yields a smaller marginal lift (+14–15 points).


The protocol for each trial is single-turn, all information upfront: a fixed system prompt and a JSON answer contract ({"<answer_field>": "<value>"}), with a 16k output-token ceiling. All models are called through a single API gateway (OpenRouter) with identical prompts and sampling parameters, and each runs at its vendor's default reasoning configuration — the harness sends no reasoning or thinking parameter. In practice the defaults differ by vendor: the two Claude models answer without extended thinking (zero reasoning tokens on every call), while the GPT, Gemini, and GLM models apply their default internal reasoning (means of roughly 85–575 reasoning tokens per call, at depths the models choose).

How answers are graded

Scoring is deterministic exact-match against the operator's recorded resolution, after light normalization: an answer either equals the recorded value or it does not. We deliberately don't use an LLM judge, to defer the ground truth to the organization's employees only. A model-graded benchmark folds a second model's biases and run-to-run variance into the measurement, rewards answers that are fluent rather than correct, and quietly drifts as the judge model is updated, so its numbers stop being reproducible or comparable across time. The human-reference benchmarks we follow reach the same conclusion. SWE-bench (and the human-audited SWE-bench Verified subset) counts a code change as correct only when the repository's own test suite executes and passes, ground truth rather than a judge; and the Data Analysis Benchmark (DAB) scores against a fixed answer key. Deterministic grading is what makes Memory-Bench re-runnable against any new model on the day it ships and comparable to the run before it. The seven-model panel is scored at **greedy pass@1** (temperature 0, one attempt). The reference model, gpt-5.5, additionally runs **50 independent trials** per condition at temperature 0.7, which yields pass@1, pass@10, pass@50, and the reliability curve in Figure 5. **pass@1** is the per-attempt hit rate you would deploy on; **pass@k**, the chance at least one of k attempts is right, is an optimistic ceiling, useful for judging headroom rather than production performance.

Inside an ISF task: retrieve the importer's org code

The ISF (Importer Security Filing) agent reads the booking documents and must file each shipment under the correct importer-of-record org code in the brokerage's TMS. Use the Next button to step through the four stages of the task, from scanned document to filed code. Names, codes, and numbers below are fictional.

Walkthrough step

The shipment arrives

The booking documents name the importer of record, Meridian Import Partners, LLC, but not its org code in the TMS. The agent extracts the party (highlighted).

Booking · Bill of lading Source
Importer of record

Meridian Import Partners, LLC

1450 Harbor Scenic Dr, Long Beach CA

Extracted

MV Northern Valor · 214E · KYTNLGB25587 · Yantian (CNYTN) → Long Beach (USLGB)

The importer's name is here — its org code is not.

TMS · ISF-10 filing Target
04 · Importer of record · org code

— awaiting org code —

Unresolved

Seller · Shenzhen Kaiyun Trading · Buyer · Meridian Import Partners · HTSUS · 8544.42.9090

Figure 2. The importer's name is on the document; the org code is not. The matcher offers fuzzy candidates, but in production the recorded answer is among them only ~1 in 10 times, so the operator free-texts it from experience. The input that reliably supplies it is the org's own history, which is what the with-memory arm retrieves. ISF is the benchmark's sharpest memory signal: 0% without memory, ~40–49% with, across every model.
Inside an ISF task: retrieve the importer's org code
StepDescriptionResult
1 · The shipment arrives The booking documents name the importer of record, Meridian Import Partners, LLC, but not its org code in the TMS. The agent extracts the party (highlighted). Unresolved
2 · Auto-match doesn't resolve it The matcher proposes fuzzy candidates from the org registry. None is the code this importer actually files under, which is the common production case. Unresolved
3 · Memory supplies the answer The memory layer retrieves this organization's prior ISF filings for Meridian Import Partners. They were filed under MERIDLB01, knowledge that lives in the operation rather than on the document. Unresolved
4 · The code is filed The importer-of-record org code is entered on the ISF-10. Memory-Bench grades exactly this string against what the operator filed. MERIDLB01

Memory only moves the cases where precedent exists.

The no-history control is flat; contradictory history is shown but excluded from the headline.

Chart view
Legend
No memory With memory
History agrees Matched-consistent · 36 tasks
+63 pp change
History contradicts itself Matched-contested · 21 tasks
Excluded from headline
No history · control Novel · 20 tasks
−1 pp change
Matched-consistent and novel tasks are the benchmark’s trustworthy strata. Matched-contested results remain diagnostic rather than headline evidence.
Memory only moves the cases where precedent exists.
GroupNo memoryWith memoryChange
History agrees 3% 66% +63 pp
History contradicts itself 19% 41% +22 pp
No history · control 39% 38% −1 pp

05 / Accuracy

Memory lifts every model 36–45 points

Across the panel, first-try accuracy rises from a pooled 22% to 62%. More model capability does not fix the baseline: reasoning power is not a substitute for organization-specific knowledge.

Greedy pass@1 across seven frontier models

44 headline tasks · accuracy 0–100% · cost shown at right

Chart view
Legend
No memory With memory
Gemini 3.5 Flash
$0.0064 per decision
Sonnet 5
$0.0018 per decision
GLM 5.2
$0.0038 per decision
Gemini 3.1 Pro
$0.0065 per decision
GPT-5.2
$0.0023 per decision
GPT-5.5
$0.0094 per decision
Opus 4.8
$0.0038 per decision
Memory shifts the full model panel upward. The most expensive models do not lead the with-memory frontier; Sonnet 5 and GPT-5.2 anchor the lower-cost end.
Greedy pass@1 across seven frontier models
GroupNo memoryWith memoryChange
Gemini 3.5 Flash 23% 66% +43 pp
Sonnet 5 20% 64% +43 pp
GLM 5.2 18% 64% +45 pp
Gemini 3.1 Pro 23% 64% +41 pp
GPT-5.2 23% 61% +39 pp
GPT-5.5 25% 61% +36 pp
Opus 4.8 20% 57% +36 pp

Accuracy versus cost, with and without memory

Seven models at greedy pass@1, marked by their lab's logo: muted for no memory, colored for with memory. The line is the efficient frontier within each condition. Up is more accurate, left is cheaper. Toggle the decision family and the cost unit; turn on memory-lift arrows to see each model's jump.

Legend
No memory With memory Cost–accuracy path 66 tasks
Filters
Decision family
Cost unit
Memory-lift arrows
0% 10% 20% 30% 40% 50% 60% 70% 80% 90% 100% $0.0010 $0.0020 $0.0030 $0.0050 $0.0080 $0.010 Accuracy · pass@1 (%) Cost per decision (US$, log scale)
Figure 3. The with-memory points sit far above the no-memory points in ISF (0% → ~48%) and reconciliation (~47% → ~76%), and barely move in EDI, where the customer's history is thin and memory is close to a no-op. Sonnet 5 and GPT-5.2 anchor the cheap end of the with-memory frontier; the two priciest models are not on it.
Accuracy versus cost, with and without memory
FamilyModelConditionAccuracyCostTokens
All gold GLM 5.2 No memory 19.7% $0.0032 514
All gold GLM 5.2 With memory 53% $0.0041 833
All gold GPT-5.2 No memory 19.7% $0.0014 228
All gold GPT-5.2 With memory 56.1% $0.0025 551
All gold GPT-5.5 No memory 24.2% $0.0094 433
All gold GPT-5.5 With memory 56.1% $0.0093 665
All gold Gemini 3.1 Pro No memory 21.2% $0.0047 510
All gold Gemini 3.1 Pro With memory 57.6% $0.0064 955
All gold Gemini 3.5 Flash No memory 22.7% $0.0052 702
All gold Gemini 3.5 Flash With memory 57.6% $0.0067 1165
All gold Opus 4.8 No memory 19.7% $0.0018 256
All gold Opus 4.8 With memory 51.5% $0.0040 671
All gold Sonnet 5 No memory 21.2% $0.0008 258
All gold Sonnet 5 With memory 53% $0.0019 713
ISF GLM 5.2 No memory 0% $0.0042 661
ISF GLM 5.2 With memory 42.9% $0.0045 932
ISF GPT-5.2 No memory 0% $0.0015 241
ISF GPT-5.2 With memory 48.6% $0.0028 611
ISF GPT-5.5 No memory 0% $0.0110 494
ISF GPT-5.5 With memory 48.6% $0.0103 733
ISF Gemini 3.1 Pro No memory 0% $0.0059 620
ISF Gemini 3.1 Pro With memory 48.6% $0.0075 1087
ISF Gemini 3.5 Flash No memory 0% $0.0048 666
ISF Gemini 3.5 Flash With memory 48.6% $0.0071 1253
ISF Opus 4.8 No memory 0% $0.0022 289
ISF Opus 4.8 With memory 37.1% $0.0045 761
ISF Sonnet 5 No memory 0% $0.0009 287
ISF Sonnet 5 With memory 40% $0.0021 798
Reconciliation GLM 5.2 No memory 45% $0.0025 408
Reconciliation GLM 5.2 With memory 75% $0.0050 965
Reconciliation GPT-5.2 No memory 40% $0.0015 223
Reconciliation GPT-5.2 With memory 75% $0.0027 613
Reconciliation GPT-5.5 No memory 55% $0.0093 419
Reconciliation GPT-5.5 With memory 75% $0.0101 734
Reconciliation Gemini 3.1 Pro No memory 45% $0.0043 463
Reconciliation Gemini 3.1 Pro With memory 75% $0.0067 1051
Reconciliation Gemini 3.5 Flash No memory 50% $0.0067 858
Reconciliation Gemini 3.5 Flash With memory 80% $0.0083 1408
Reconciliation Opus 4.8 No memory 40% $0.0014 207
Reconciliation Opus 4.8 With memory 80% $0.0041 715
Reconciliation Sonnet 5 No memory 45% $0.0007 219
Reconciliation Sonnet 5 With memory 80% $0.0023 787
EDI GLM 5.2 No memory 36.4% $0.0011 241
EDI GLM 5.2 With memory 45.5% $0.0012 280
EDI GPT-5.2 No memory 45.5% $0.0009 194
EDI GPT-5.2 With memory 45.5% $0.0012 249
EDI GPT-5.5 No memory 45.5% $0.0043 264
EDI GPT-5.5 With memory 45.5% $0.0051 322
EDI Gemini 3.1 Pro No memory 45.5% $0.0015 246
EDI Gemini 3.1 Pro With memory 54.5% $0.0024 359
EDI Gemini 3.5 Flash No memory 45.5% $0.0037 530
EDI Gemini 3.5 Flash With memory 45.5% $0.0026 440
EDI Opus 4.8 No memory 45.5% $0.0015 239
EDI Opus 4.8 With memory 45.5% $0.0018 304
EDI Sonnet 5 No memory 45.5% $0.0006 239
EDI Sonnet 5 With memory 45.5% $0.0008 304

Surfacing versus inferring

Where the answer is verbatim-retrievable, memory takes the model from 0% to 76%. Where it must be inferred around a coverage gap, accuracy moves from 31% to 51%. Both are product value, but they measure different things.

Is it reasoning, or copying?

The common failure mode of a memory benchmark is that "using memory" collapses into "copy the most recent thing you were shown." We test that directly on the hardest ISF sub-population: cases where the retrieved history contains **two or more different codes** and the model must select. On those 12 tasks we also score three non-reasoning policies against the same key.

On the hardest ISF cases—where history contains two or more codes—the model reaches 70%. Copying the most recent code reaches 50%, the most frequent reaches 33%, and choosing randomly reaches 34%. Beating the copy-most-recent ceiling by twenty points indicates that the model matches precedent to the specific case rather than repeating whichever code appears first.

The model beats every copy-only policy.

Pass@1 on the ISF tasks that contain two or more competing prior codes.

0% → 100% · Higher is better
Copy most recent prior code
50%
Copy most frequent prior code
33%
Pick uniformly at random
34%
Model, with memory
70%
A twenty-point lead over the strongest copy baseline indicates that the model discriminates between precedents instead of repeating whichever code appears first.
The model beats every copy-only policy.
MethodValue
Copy most recent prior code 50%
Copy most frequent prior code 33%
Pick uniformly at random 34%
Model, with memory 70%

What happens with full production context

The headline isolates memory with a thin version of the situation. Production context is richer: the full structured record plus parsed document text. In that variant, the baseline rises and memory’s incremental lift compresses to roughly +14–15 points. Where history exists and agrees, full-context runs still jump from 9% to 83% for ISF and 44% to 73% for reconciliation.

06 / Reliability

The lift holds when every attempt must succeed

An operations decision gets no credit for being right once in fifty tries. We measure passk: the probability that all k attempts at the same case succeed.

Probability every attempt is correct

Reference model · 50 trials · trustworthy strata

Chart view
k · number of attempts that must all succeed
Memory’s effect is consistency. The cases it can solve, it solves repeatedly rather than intermittently. With memory, the model retains 53% when all 32 attempts must be correct.
Probability every attempt is correct
ViewSeries12481632
Probability With memory 60%58%55%55%54%53%
Probability No memory 20%19%18%17%17%15%
Drop from k=1 With memory 0 pp2 pp5 pp5 pp6 pp7 pp
Drop from k=1 No memory 0 pp1 pp2 pp3 pp3 pp5 pp

Robustness to presentation order

For forty tasks we reorder retrieved memory entries five ways at temperature zero. Where history agrees, 74% of answers remain stable across every ordering and the flip rate is 6%. When history contradicts itself, stability falls to 41% and the flip rate rises to 30%—exactly where precedent itself is ambiguous.

07 / Cost

Does the memory-layer add to LLM costs?

Retrieved cases add 220–430 prompt tokens per decision across the panel. Dollar cost barely moves because memory also lowers the reasoning load.

The reference model’s reasoning tokens fall 16% with memory, from 271 to 229 per decision over fifty trials. On ISF, reasoning tokens drop 22%, and the call costs less in absolute terms. Without memory, the model reasons at length toward a code it has no way to know; with memory, it reads the code from a prior filing and stops.

Memory costs the same to run—and far less per correct result.

Toggle the denominator to see where the economic advantage appears.

Metric
Per-decision spend is effectively flat. Once accuracy is included, memory reduces the reference model’s cost per correct decision from $0.0521 to $0.0178.
Memory costs the same to run—and far less per correct result.
MetricNo memoryWith memory
Per correct $0.0521$0.0178
Per decision $0.0093$0.0094

The reasoning saving is model-dependent, but the economics hold across the panel: per-decision spend changes by less than a quarter of a cent while accuracy roughly triples, so cost per correct decision is two to three times lower.

Conclusion

The missing capability was not more intelligence. It was organizational context.

Across real freight decisions, memory changes what models can know, improves whether they answer consistently, and lowers the cost of each correct outcome.

Explore more Pallet Labs research