Skip to content
Teardown · 6 min

The 70% problem: why AI gets you most of the way and then stops

The last 30% — security, edge cases, scale, maintainability — is pure engineering judgement. A plain look at why it's the only part worth paying for.

AI will get you a working-looking thing in an afternoon. That is real, it is genuinely useful, and it has changed what a small team can attempt. The trap is mistaking the 70% for the 100%.

What the first 70% buys you

A shape. Screens that render, a happy path that runs, enough of a system to demonstrate and to argue about. That used to take weeks and now takes hours, which is a real gain and worth taking.

It also feels finished, and that is the problem. The parts that are missing are invisible in a demo — by definition, because a demo only runs the path you chose to show.

What the last 30% actually is

Not more features. Judgement calls, each of which has to be made deliberately:

Which inputs to distrust. Everything from outside your system is hostile until proven otherwise — user text, uploaded files, third-party responses, and now model output, which is confident regardless of correctness.

Which failures to make loud. The dangerous failure is the quiet one. A system that silently drops one record in a thousand will pass every test you write and cost you a customer eighteen months later.

Which data must never cross a boundary. Multi-tenancy is not a feature you add. It is an assumption baked into every query from the first commit, and retrofitting it means revisiting all of them.

What an auditor will ask to see. If the answer is "we'd have to reconstruct it", the answer is no. Evidence is either captured as it happens or it does not exist.

Why AI cannot supply this for itself

The last 30% is mostly about context the model does not have: what your business considers an acceptable loss, which regulator is watching, which downstream team depends on that field, what happened the last time this broke.

AI is excellent at producing plausible code. Plausible is exactly the failure mode — it looks right, reviews cleanly if you are not looking for the specific problem, and fails on the case nobody wrote down. A non-technical builder cannot self-check this, because the whole difficulty is knowing what is absent.

Speed and safety are not in tension here

The common framing — move fast, or be careful — is wrong at this scale. Getting the 30% right the first time is cheaper than finding it in an incident report, and it is much cheaper than finding it during due diligence on a funding round.

The expensive version is not "we spent longer building it". It is the rebuild, the disclosure, the customer who leaves, and the six weeks your team spends on remediation instead of the roadmap.

How to tell which 30% you are missing

Pick your most important system and answer three questions in writing: what does it do when its dependencies fail, who can see data they should not, and what evidence exists that a human reviewed it before release.

If any of those takes more than a few minutes to answer, you have found the gap. That is not a criticism of how you built it — it is the normal consequence of building fast, which was the right call at the time. Finding all of them in a week is the Reality Check.

This is the part we do — the crossing from a demo to a system that survives production.

If this was useful

Start with a free fit conversation. A detailed assessment is scoped and quoted separately, or send a brief if you already know what you need.