The 2026 Guide to Rescuing an AI-Built App
Your app works, users are on it, and nobody can safely change it. The audit-to-production path we actually run — what gets checked, what it costs, and when a rescue is the wrong call.
Contents8 sections
You built it in a weekend and it worked. Then it kept working, people started paying, and somewhere around month three every change became frightening. Nothing has crashed. That is the confusing part — the app is live, the users are real, and yet nobody can touch it without a knot in their stomach. This page is the path out of that, written from the audit side.
What actually goes wrong in an AI-built app?
Not crashes. The failure mode is a well-formed wrong answer — a page that loads, a query that returns, a login that succeeds for the wrong person. That is why the app can look fine for months while the risk accumulates, and it is why "it works" is not evidence of anything.
Code generators got dramatically better at syntax and not at all at security
Pass rate across Veracode's controlled coding tasks
This measures generated samples in controlled tasks, not shipped production code — so read it as a property of the generator, not a prediction about your app. It is the strongest available argument that waiting for a better model is not a plan.
Veracode, Spring 2026 GenAI Code Security Update (150+ models; 2023–2026 trend)
The four we find most often, in order
Access control that is decorative
Broken access control is still the number one category in OWASP's Top 10, in the 2025 edition released January 2026. In generated apps it usually looks like a query that filters by an id the caller supplies, with the ownership check living in the interface rather than in the database query. Change the number in the URL and you are somebody else.
Row-level security switched off, or never switched on
The hosted-database defaults are safe-ish and the tutorials are not. A table with RLS disabled and a public anon key in the front-end bundle is readable by anyone who opens developer tools. This is a five-minute check and it is the first thing we run.
Secrets in the client bundle
Build-time environment variables are compiled into what ships to the browser. A key prefixed for client exposure is public, permanently, from the moment it deploys — and rotating it is the only remedy, because it is already in somebody's cache.
No error handling, so failures are silent
Generated code tends to assume the happy path. The user sees a spinner that never resolves, and you see nothing at all, because there is no logging to see it in. Most founders discover this class of bug from a support email weeks later.
Why does month three keep being the wall?
Because that is roughly when the volume of code passes what one person can hold in their head, and the thing that was holding it — a conversation with a model, in a session that has since ended — is gone. The prompt is not documentation. It described an intention at a moment; the code that came out of it may or may not still match, and there is no way to diff an intention.
What does a rescue audit actually cover?
Reading the code. All of it, by a person. That sounds like a slogan and it is really a method — the defects above are invisible to a scanner precisely because each one produces a valid response. Below is the scope we work through, in the order we work through it, so you can compare it against whatever anyone else quotes you.
| Area | What gets checked | What a bad finding looks like |
|---|---|---|
| Access control | Every read and write that touches user-owned data, and whether the owner is in the query | An id parameter with no ownership filter on one of two branches |
| Data exposure | Database policies, public keys, what the client bundle contains, what the API returns | An endpoint returning the whole user record when the page needs a name |
| Auth and sessions | Signup, login, reset, session lifetime, privilege changes | Password reset that does not invalidate existing sessions |
| Money paths | Anything that charges, refunds, credits or counts | A webhook with no idempotency, so a retry bills twice |
| Failure behaviour | What happens when a dependency is slow, down, or returns something unexpected | A caught exception logged and swallowed, with no recovery path named |
| Scale | Query patterns, indexes, anything that loops over a network call | An N+1 query that is fine at 50 users and fatal at 1,000 |
| Deployability | Environments, secrets handling, whether anyone but the original author can deploy | Production credentials living only on one laptop |
Rebuild or salvage?
Salvage, usually. A rebuild quote is easier to write and easier to sell, and it throws away the one genuinely valuable thing you own — a working product that real users have already shaped. These are the questions that decide it, and only the last two point toward a rebuild.
The six questions we ask before quoting anything
- Does it have users who would notice it going away? If yes, the bar for a rebuild goes up sharply.
- Is the data model roughly right? Wrong tables are expensive. Wrong code is cheap by comparison.
- How much of the code is actually load-bearing? Generated projects carry a lot of scaffolding nobody calls.
- Can the stack hire? A mainstream framework with an awkward structure beats an elegant one nobody knows.
- Is there anything that cannot be tested? Code you cannot characterise, you cannot safely change.
- Is the platform a dead end? If the product needs something the builder fundamentally cannot do, that is a genuine rebuild trigger.
What does a rescue cost?
Be sceptical of any figure here, including ours. There is no independent dataset on rescue pricing — every published number we could trace comes from a firm selling rescues, and the widely circulated "8,000 of 10,000 vibe-coded startups now need rescue, at $50K–$500K each" figure traces to a single article with no attribution behind it at all. We do not cite it and neither should anyone quoting you.
What genuinely moves the number
- How much of the app touches money or personal data — that is where the careful work is
- Whether tests exist. If nothing characterises current behaviour, every change needs one written first
- How many third-party services are wired in, and whether anyone documented why
- Whether the original author is reachable. An hour of their time is worth days of ours
- Whether you need it hardened, or hardened and handed to a team who will keep it
How do you hand it to a team afterwards?
The hardening is the visible half. The half that determines whether you are back here in six months is whether the app became maintainable by somebody other than whoever fixed it. Ask for all five of these as deliverables, in writing, at the start.
- A written architecture note — what the pieces are and how a request flows through them
- Tests around the paths that matter, so the next change has a safety net
- One documented way to deploy, runnable by someone who was not there
- Secrets in a manager, rotated, and out of the repository history
- Error tracking wired up, so the next failure is something you find rather than something you are told
Common questions
How do I know if my app is actually at risk right now?
The fastest useful check: open your app in a browser, open developer tools, and look at the network requests and the bundle for anything that looks like a key. Then, if you use a hosted database, check whether row-level security is enabled on every table with user data. Those two cover the majority of what gets found in a live vibe-coded app, and both take minutes.
Can I just ask the AI to fix the security problems?
For a specific, named, well-scoped defect, often yes. For "make this secure", no — that is the request that produced the current state. The Veracode trend is the relevant evidence: generation quality improved enormously on syntax and not at all on security, so the tool is not converging on the thing you need it to be good at.
Will I have to move off the platform I built on?
Usually not, and you should be suspicious of anyone who opens with that. Exporting to a normal repository is common and sensible; changing framework rarely is. The export itself has traps — build-time environment variables in particular behave differently once you own the pipeline — but that is a day of work, not a migration.
How long does an audit take?
Days, not weeks, for a typical single-product codebase. The remediation is the long part and it is quoted separately, from findings, so you can stage it.
What if I only want the critical issues fixed?
That is a legitimate and common choice, and the findings list is structured to make it possible. Fix what is exploitable now, plan the rest, and be honest with yourself about the plan.
The short of it
An AI-built app with users is an asset with a maintenance problem, not a mistake. Get it read by a person, get the findings in writing, fix what is exploitable, and make it changeable by someone other than the model that wrote it. Everything else can wait, and most of it will turn out not to matter.



