Discussion about this post

User's avatar
Think AI's avatar

Fascinating perspective Loop Engineering could be the key to building AI systems that continuously learn, adapt and improve.

Rohit Yadav's avatar

The holdout result in round two is the most useful failure in the whole piece. A challenger that gains half a point on the improvement set and then loses ground on unseen cases is not an improvement, it is memorization dressed up as progress. Most teams I see building agent evaluation pipelines skip that split entirely and end up shipping a prompt that scores well on the exact examples someone wrote the eval against, which is a different thing from working.

The four stage autonomy ladder, shadow, suggest, approve, bounded autonomy, is the part I would push people to write down before they touch a model. Granting autonomy based on observable evidence rather than the model's stated confidence is a distinction that gets lost the moment a demo goes well. A system that passed its tests once in shadow mode is not the same system six weeks later once the input distribution has shifted, and the piece is right that sampling autonomous decisions has to be ongoing, not a one time gate.

What I would add from the enterprise side is that the evaluator itself is usually the weakest link, and it is rarely audited with the same rigor as the loop it is judging. A vague rubric produces a system that is very good at satisfying a vague rubric, which looks like progress on a dashboard and means nothing in production. Happy to connect, let's talk more about it.

1 more comment...

No posts

Ready for more?