Your Agent Passed the Eval. Its Retrieval Calls Still Tripled.
Task completion can stay flat while context compression shifts an agent's budget from execution to redundant retrieval.
Aug 22, 20266 min read1

Search for a command to run...
Articles tagged with #ai-evaluation
Task completion can stay flat while context compression shifts an agent's budget from execution to redundant retrieval.

A benchmark stops being neutral test infrastructure when the agent can execute code, cross trust boundaries, or affect real systems.
