"Correct and finished are different" is the right frame, and it cuts the other way too: an agent that can't tell when it's finished also can't tell when it shouldn't have started. For everyday users that argues for putting the check before irreversible steps (send, submit, pay), not only after the task, since a review after the fact can't unsend anything.
Automation benchmarks often ignore the friction of trust. Scaling human output through agents creates a paradox where efficiency gains are offset by the growing surface area for social engineering. As back office operations grow, the target moves from the internal network to the external relationship. When these agents handle sensitive communication, they become the primary vector for extortion. Attackers no longer need to break the firewall if they can manipulate the automated workflows that already have access to your client data. This shift changes the economics of extortion, turning every automated touchpoint into a potential ransom event for your customers.
"Correct and finished are different things" is the line I'd keep. On a job site that gap has a name, the punch list, and it lives between the day the work is done and the day somebody signs it off. A green dashboard usually means the first one. I'd want the task to stay open until a person confirms the outcome, not the click.
the $2.10 number is the quietest bomb in the piece. it doubles the headline cost and still assumes every failure gets caught, which the silent-error section just ruled out. the real missing metric is cost per silent error, and that's the one no vendor will ever publish voluntarily.
Reading the page's structure has one more trap, on the site side. A YC startup I looked at this week shows 87% approved quotes on screen. The raw page says 0%, because a counter fills it in after load. Anything that reads the raw page instead of the rendered one gets the zero.
Ruben, the silent-error section is the heart of this. But I’d push on the ending. You treat the verification layer as where people land. I’m not sure it’s a stable place to stand.
Every correction a reviewer makes, every net 60 caught or claim flagged, is labeled training data for the next model. The checking job trains its replacement. What’s left after that isn’t verification. It’s a signature: a person whose name goes on the work because liability needs one, not because they’re still the one exercising judgment. Madeleine Clare Elish called that a moral crumple zone.
So the hiring numbers may be the wrong signal. The question isn’t whether back offices are still hiring. It’s what the people they hire are actually deciding.
A signature is a weak safeguard if its owner cannot halt the process. For that checking role, I’d want to know whether the person can suspend automation and demand a change to the workflow, or merely correct individual records. The sign-off should reflect the authority they actually have.
That’s the right test, and I think it’s exactly what gets engineered out. Toyota gave every line worker the andon cord: anyone could stop production. Most “human in the loop” roles get the opposite. They can fix the record in front of them, but not stop the system that produced it.
A signature without a stop button isn’t oversight. It’s absorption.
Your framing gives a usable test for any human-in-the-loop claim: can that person pause the automation without asking anyone’s permission? If not, the sign-off should say what it actually is.
"Correct and finished are different" is the right frame, and it cuts the other way too: an agent that can't tell when it's finished also can't tell when it shouldn't have started. For everyday users that argues for putting the check before irreversible steps (send, submit, pay), not only after the task, since a review after the fact can't unsend anything.
Automation benchmarks often ignore the friction of trust. Scaling human output through agents creates a paradox where efficiency gains are offset by the growing surface area for social engineering. As back office operations grow, the target moves from the internal network to the external relationship. When these agents handle sensitive communication, they become the primary vector for extortion. Attackers no longer need to break the firewall if they can manipulate the automated workflows that already have access to your client data. This shift changes the economics of extortion, turning every automated touchpoint into a potential ransom event for your customers.
https://cyrilsimonnet.substack.com/p/the-ransom-note-now-goes-to-your?utm_source=substor&utm_medium=substack&utm_campaign=comment
"Correct and finished are different things" is the line I'd keep. On a job site that gap has a name, the punch list, and it lives between the day the work is done and the day somebody signs it off. A green dashboard usually means the first one. I'd want the task to stay open until a person confirms the outcome, not the click.
the $2.10 number is the quietest bomb in the piece. it doubles the headline cost and still assumes every failure gets caught, which the silent-error section just ruled out. the real missing metric is cost per silent error, and that's the one no vendor will ever publish voluntarily.
Reading the page's structure has one more trap, on the site side. A YC startup I looked at this week shows 87% approved quotes on screen. The raw page says 0%, because a counter fills it in after load. Anything that reads the raw page instead of the rendered one gets the zero.
Ruben, the silent-error section is the heart of this. But I’d push on the ending. You treat the verification layer as where people land. I’m not sure it’s a stable place to stand.
Every correction a reviewer makes, every net 60 caught or claim flagged, is labeled training data for the next model. The checking job trains its replacement. What’s left after that isn’t verification. It’s a signature: a person whose name goes on the work because liability needs one, not because they’re still the one exercising judgment. Madeleine Clare Elish called that a moral crumple zone.
So the hiring numbers may be the wrong signal. The question isn’t whether back offices are still hiring. It’s what the people they hire are actually deciding.
(I made a longer version of this argument in “The Titles Survived,” if it’s useful: garyfbart.substack.com/p/the-titles-survived)
A signature is a weak safeguard if its owner cannot halt the process. For that checking role, I’d want to know whether the person can suspend automation and demand a change to the workflow, or merely correct individual records. The sign-off should reflect the authority they actually have.
That’s the right test, and I think it’s exactly what gets engineered out. Toyota gave every line worker the andon cord: anyone could stop production. Most “human in the loop” roles get the opposite. They can fix the record in front of them, but not stop the system that produced it.
A signature without a stop button isn’t oversight. It’s absorption.
Your framing gives a usable test for any human-in-the-loop claim: can that person pause the automation without asking anyone’s permission? If not, the sign-off should say what it actually is.