PinnedThe One Extra Click That Makes Chrome's Built-in Gemini Feel Less IntelligentNothing is broken. Nothing is slow. Yet one small interaction subtly interrupts the experience, and it highlights an important lesson about designing AI products.Aug 5, 2026·3 min read·84
Code Generation Is Fast. Verification Is Not.A 48-project study shows where agentic IDE speed turns into reviewer work.Aug 27, 2026·6 min read·8
Are These Really Accidental Leaks?Don’t Copy “Leaked” System Prompts. Study Their Architecture Instead.Aug 25, 2026·8 min read·40
Benchmark the Model–Harness Pair, Not the ModelA coding model's runtime can change cost by orders of magnitude without producing a dramatic pass-rate gap.Aug 24, 2026·6 min read·57
Your Agent Passed the Eval. Its Retrieval Calls Still Tripled.Task completion can stay flat while context compression shifts an agent's budget from execution to redundant retrieval.Aug 22, 2026·6 min read·1
Your AI Agent Eval Is a Production Security BoundaryA benchmark stops being neutral test infrastructure when the agent can execute code, cross trust boundaries, or affect real systems.Aug 20, 2026·9 min read·14
The Competence Debt of Agentic CodingAI can write more of the code. The harder question is whether we can still understand the systems it builds.Aug 20, 2026·22 min read·10