Four kinds of memory just to stop an agent from forgetting its own context after 80 turns shows how much of "production AI" is really just memory management wearing a different name. Hit a smaller version of that building an AI shopping agent, forgetting what a user already said halfway through a session kills trust fast.
the grader point is the one i keep coming back to, but i'd add a caveat: an independent context window isn't an independent mind. if the grader is the same model family as the worker, they share blind spots and the grade is mostly theater for correlated failures. two different models, or a rubric with teeth — otherwise you're just paying twice for the same opinion.
'Demos take a weekend, production still takes a fight' sums it up. The meeting briefs that check their own work stood out to me, that self-check step is where most agent projects stall. Full disclosure, I built allthingspm.app, and its agents and evals chapters focus on exactly that gap.
The part of Managed Agents that matters for anyone who has to explain an agent's behaviour later is the session. Anthropic's engineering post of 8 April 2026 describes it as an append-only log of everything that happened, stored durably outside the model's context window and readable through getEvents(). That is an audit trail by construction, not an add-on. The same post says the harness never handles credentials: Git tokens are wired into the sandbox at initialisation, and MCP OAuth tokens sit in a vault behind a proxy. What the post does not settle is how long a session log is retained or who can read it.
Four kinds of memory just to stop an agent from forgetting its own context after 80 turns shows how much of "production AI" is really just memory management wearing a different name. Hit a smaller version of that building an AI shopping agent, forgetting what a user already said halfway through a session kills trust fast.
the grader point is the one i keep coming back to, but i'd add a caveat: an independent context window isn't an independent mind. if the grader is the same model family as the worker, they share blind spots and the grade is mostly theater for correlated failures. two different models, or a rubric with teeth — otherwise you're just paying twice for the same opinion.
'Demos take a weekend, production still takes a fight' sums it up. The meeting briefs that check their own work stood out to me, that self-check step is where most agent projects stall. Full disclosure, I built allthingspm.app, and its agents and evals chapters focus on exactly that gap.
The part of Managed Agents that matters for anyone who has to explain an agent's behaviour later is the session. Anthropic's engineering post of 8 April 2026 describes it as an append-only log of everything that happened, stored durably outside the model's context window and readable through getEvents(). That is an audit trail by construction, not an add-on. The same post says the harness never handles credentials: Git tokens are wired into the sandbox at initialisation, and MCP OAuth tokens sit in a vault behind a proxy. What the post does not settle is how long a session log is retained or who can read it.