5 Comments
User's avatar
Marius Laurusevicius's avatar

Building without domain expertise is one thing; grading the output is where it returns. Stanford's RegLab team ran the first preregistered evaluation of AI legal research tools, published in the Journal of Empirical Legal Studies, and measured hallucination rates of 17 percent for Lexis+ AI and 33 percent for Westlaw AI-Assisted Research across 202 queries hand-scored by legal experts. Both products had marketed the opposite. Nobody without legal training could have scored those answers. So the founder lesson holds on the build side and inverts on the buy side.

Alicia Crowther's avatar

The inversion is the part most product teams miss. They hire or partner on the build side without thinking about who will score the output. In high-stakes domains, the scoring is where the liability actually sits. The tool hallucinates at 17-33% and nobody without domain expertise caught it before the marketing said the opposite. The moat isn't the model. It isn't even the delivery layer. It's the person who knows when the output is wrong before the client does.

Marius Laurusevicius's avatar

The EU AI Act names that person. Article 26(2) requires deployers of high-risk systems to assign human oversight to natural persons who have the necessary competence, training and authority, plus the necessary support. Competence and authority are listed separately, which matters: reviewers are often handed the job without the standing to overrule the output. That obligation starts 2 December 2027 for Annex III high-risk systems and 2 August 2028 for Annex I. Until then the scoring role exists in practice with no named legal owner.

Immanuel Santosh's avatar

The no-domain-expertise rule holds for law, but I'd separate two things he's merging: not knowing the industry, and not knowing the customer's workflow.

Legora closed that gap by hiring lawyers early to build eval sets. Most founders skip that step, call it "learning fast," and end up shipping features nobody asked for.

Learning speed is the leading indicator only when it's directed at the buyer's actual day, not just the market's vocabulary.

Lorelei's avatar

Curious, what was their original hypothesis, what were they trying to solve, since they didn't have domain expertise?