Blog · Guides · 2026-10-05

Lovable App Testing: How to QA an App the AI Built

TL;DRLovable app testing means checking what the generator could not: that database policies block data users shouldn't see, that core journeys survive each re-prompt, and that bad input is rejected server-side. In a typical Lovable build the browser queries Supabase with a public key, so authorization lives in row-level security policies — not in your UI.

Lovable app testing means verifying the three things the generator cannot verify for you: that your database rules actually block data a user should never see, that every critical journey still works after the last re-prompt, and that the app behaves sanely when inputs are empty, wrong, or hostile. Start at the data layer — in a typical Lovable build the browser talks straight to Supabase using a public key, so authorization lives in row-level security policies, not in your UI code. Then pin the main user journeys into repeatable tests, so the next prompt cannot silently break what already worked.

What makes a Lovable app different from a hand-coded one?

Lovable turns prompts into a working full-stack app — usually a React frontend with Supabase handling auth, database, and file storage, plus a preview URL and a publish step. The code is real code, and you can sync it to GitHub and read every line. What is missing is everything that normally surrounds code written by a team: no pull request, no reviewer, no test suite, no staging environment, and no record of what the last change was supposed to do.

That changes where bugs come from. In a hand-written app, defects cluster around the hard parts the developer knew were hard. In a generated app, they cluster around the parts nobody ever stated: the requirement you never typed into the prompt. The model fills gaps with plausible defaults, and plausible defaults are exactly what a test would have caught. A signup form that accepts an empty name, a dashboard that loads before auth resolves, a delete button with no confirmation — none of these are model failures. They are unstated requirements.

The second difference is regeneration. When you re-prompt to change a feature, the model may rewrite files you did not mention. Without tests, you discover the collateral damage by clicking around and hoping you click the right thing.

What should you test first in a Lovable app?

Test in order of what a failure actually costs you. Not every bug is equal: a misaligned button costs you a little credibility, while a readable table of other customers' records costs you the business.

PriorityWhat to testWhy it comes first
1Row-level security on every tableThe client holds a public key; policies are the only real gate
2Auth boundaries — logged out, logged in, wrong userHidden in the UI is not the same as protected
3Server-side validation on writesClient validation is a hint, not a constraint
4File uploads — type, size, who can read themPublic storage buckets leak quietly
5Payment and webhook pathsSilent failures here lose money with no error
6Empty, loading, and error statesGenerated UIs are built against the happy path
7Mobile layout and long contentPreview width is not user width

Work top-down and stop when you run out of time, rather than starting with whatever is easiest to click.

How do you test authorization when the browser talks straight to the database?

This is the part most non-technical founders miss, and it is the single highest-value test in a Lovable app testing pass. In the common Supabase setup, your frontend ships a public anon key and queries tables directly. Anyone can open devtools, read that key, and issue their own queries. The only thing standing between a curious visitor and your whole table is the row-level security policy on it.

So test it as an attacker would, with three passes:

  1. Logged out. Open the app with no session and try to load a page that shows data. If anything appears that should be private, a policy is missing.
  2. Logged in as user A. Note the record IDs you own. Then change an ID in the URL to one you do not own. If it loads, your filter is in the query, not in a policy.
  3. Logged in as user B. Try to update and delete user A's records. Read access and write access are separate policies, and generated apps frequently cover one and not the other.

Two accounts and ten minutes will tell you more about your app's safety than any amount of reading the code. If any of these three passes shows data it shouldn't, fix the policy — not the UI. Hiding a route does nothing when the data is one fetch call away. We cover the wider version of this question in is vibe-coded code safe to ship.

How do you stop a re-prompt from breaking what already worked?

You cannot review every regenerated line, and you should not try. What you can do is make breakage loud. Pick the three to five journeys that define your app — sign up, create the main object, pay, invite a teammate — and turn each into an automated end-to-end test. Run them after every meaningful prompt. A journey that passed yesterday and fails today tells you precisely which re-prompt did it, which is a far better debugging signal than a vague sense that something feels off.

Two practical notes. Write selectors against stable attributes rather than generated class names, because regeneration reshuffles styling far more often than it reshuffles intent. And when a test breaks because the UI legitimately changed, update the test in the same session — a suite you have learned to ignore is worse than no suite, because it costs time and buys nothing.

This is the gap AutoSim is built for: end-to-end tests that repair their own selectors when the markup shifts underneath them, so a cosmetic regeneration doesn't produce a wall of red that trains you to stop looking.

What does a practical Lovable app testing workflow look like?

A workflow that survives contact with real building has to be cheap enough to actually run. Four stages:

1. Before you share the preview link

Run the three authorization passes above. Click every destructive action once. Submit every form empty, then with absurd input — 500 characters in a name field, an emoji in a phone number, a negative quantity. Generated validation is usually present and usually shallow.

2. When real people start using it

Give them a way to report what they hit without writing a bug report. Non-technical testers describe symptoms, not state — "it broke" arrives without the URL, the console error, the failed request, or the browser. A right-click capture that attaches all four turns a useless message into something actionable; that's what Snap does, and the general principle is covered in getting useful bug reports from non-technical users.

3. On a cadence, not just on demand

Run your journey tests on a schedule as well as after prompts. Apps built on hosted backends break from the outside too: a policy change, an expired key, a quota.

4. Before anything that touches money or data

Treat payment, deletion, and export flows as a separate review every time they change. These are the flows where a silent failure looks exactly like success.

Lovable app testing checklist before you publish

  • Every table has an explicit row-level security policy for select, insert, update, and delete
  • A logged-out visitor sees nothing private, on every route
  • User B cannot read, edit, or delete user A's records by changing an ID
  • Storage buckets are private unless the file is genuinely public
  • No API keys or secrets beyond the intended public key appear in the client bundle
  • Every form rejects empty, oversized, and wrong-type input on the server
  • Loading, empty, and error states exist for every data-backed view
  • Destructive actions confirm before they act
  • The app works at 375px wide, not just in the preview pane
  • Your three to five core journeys run as automated tests after every prompt

None of this requires you to become an engineer. It requires you to assume the generator optimized for a demo that works, because that is what it was asked for, and to go test the cases nobody asked for.

Where does AI testing fit in?

The reason manual passes fail over time is not that they are wrong — it's that they are a tax on every change, and taxes get dodged. The useful split: automate the journeys you already know matter, and use AI personas to explore the paths you never thought to write down. A persona that behaves like an impatient user with a slow connection and a half-filled form will reach states your own clicking never will, because you know your app too well to use it badly. Sims runs that exploration, and the broader method is in our complete guide to AI QA.

Try Klavity free

Point Klavity at the app you just built and let it do the pass you don't have time for — authorization probing, journey tests that heal themselves, and one-click bug capture for the people testing it. Try Klavity free and scan your next build before your users do.

Key takeaways

  • Test row-level security on every table before anything else
  • Probe as a logged-out visitor and as a second user
  • Pin 3-5 core journeys into automated tests after every prompt
  • Validate on the server, not just in the generated form

FAQ

Do I need to know how to code to test a Lovable app?

No. The highest-value tests are behavioural: open the app logged out and see if private data appears, log in as a second user and try to open the first user's records by changing an ID in the URL, and submit every form empty and with absurd input. Two accounts and ten minutes will surface more real risk than reading the generated code.

Why is row-level security the first thing to test?

Because in the common Lovable and Supabase setup, the frontend holds a public anon key and queries the database directly. Anyone can read that key from the browser and issue their own queries, so row-level security policies are the only real gate. Hiding a page in the UI does not protect the data behind it.

How do I stop re-prompting from breaking features that already worked?

Turn your three to five defining journeys — sign up, create the main object, pay, invite a teammate — into automated end-to-end tests and run them after every meaningful prompt. Write selectors against stable attributes rather than generated class names, since regeneration reshuffles styling more often than intent.

Is testing an AI-built app different from testing a hand-coded one?

The bugs cluster differently. Hand-written defects gather around the parts the developer knew were hard; generated defects gather around requirements nobody ever stated, which the model fills with plausible defaults. Test the cases you never put in the prompt: empty inputs, wrong users, slow networks, and error states.

Catch bugs the moment a human sees them

Klavity: right-click bug reports, AI personas that review your product, and self-healing tests.

Get started free