Introducing AsyncQA: A Local Model That Files Bug Reports You Can Check
A long-running QA agent driven by models on a 12 GB RTX 5070. It explores a web app, verifies every suspected bug independently, and attaches the evidence and the arithmetic behind each confidence score. Tested against a shop with 11 planted bugs.
The QA problem is trust, not effort
Pointing a model at a web app and asking it to find bugs is easy. It will happily report a dozen. The hard part is the report after that: is any of it real?
An agent that files ten bugs, three of them imaginary, is worse than no agent. Every report costs a human five minutes to disprove, and after the second phantom you stop reading them. So the question I cared about wasn’t “can a local model find bugs.” It was:
Can a model running on a consumer GPU keep testing an app around the clock and file bug reports you’d actually trust?
That breaks into three smaller questions. Can it find anything? Can the system tell its real finds from its hallucinations? And can you see why it believes each one?
Enter AsyncQA
AsyncQA is a daemon. You give it an app (a URL, local or in Docker) and some credentials. It keeps running tasks on a schedule:
| Task | Model? | What it does |
|---|---|---|
health | no | boolean checks: HTTP, TCP, Docker, shell commands |
sweep | only to verify | renders every page at six screen widths, measures layout, and catches JS errors and 5xx responses |
explore / api | yes | charter-based exploratory testing through a browser and HTTP |
e2e | yes | a goal-driven user journey (“log in, buy two things, check the total”) |
chaos | no | restarts or pauses a container and measures time to recovery |
recheck | yes | re-verifies open bugs and marks them possibly fixed |
It runs in two modes. Blind gets a URL and a login and maps the app itself. Prescribed gets a list of areas with test charters and invariants (“a user can only see their own orders”).
The driver can be Claude, an OpenAI model, or anything on Ollama. The default
config here is entirely local: qwen3-coder:30b-a3b explores and gpt-oss:20b
verifies. That’s the same pair from the Bakeoff
trial, on the same 12 GB RTX 5070.
The part that matters: a bug has to survive verification
Nothing the explorer reports goes straight into a bug file. Every suspected bug goes through a pipeline:
- Dedupe against known bugs, by title similarity and by which endpoints the evidence touched.
- Deterministic replay. The HTTP exchanges it cites are re-sent. For layout bugs, a fresh browser measures the page again. No model is involved in this step.
- An independent verifier. A fresh context, ideally a different model, given the claim and the steps but not the finder’s reasoning. It has to reproduce the bug with its own tools, then rule on two separate questions: did it happen, and is it actually wrong?
- Scoring. Confidence is a sum of log-odds, with one visible term per piece of evidence.
That last step is the one I’d push hardest on anyone building something similar. Every bug file ends with its own arithmetic:
## Why this confidence
- +0.85 prior from finder self-confidence 1.00
- +1.20 deterministic replay reproduced the error response
- +1.50 verifier #1: reproduced=yes, is_bug=yes
- = 97%
A number you can audit is far more useful than a number you’re asked to accept.
One small design note: the model never sees a password. It writes
{{cred.alice.password}}, the tool layer substitutes the real value at send
time, and every byte of evidence is redacted back before it’s saved. Session
tokens the app mints at runtime are masked in the files meant for humans.
A target with known answers
You can’t evaluate a bug finder on an app whose bugs you don’t know. So AsyncQA
ships with one: Buggy Shop, a small, modern storefront with planted
defects. Each can be toggled with an environment variable, and BUGS=none gives
a clean build for measuring false positives.

There are eleven planted bugs. Seven are functional: an IDOR on orders, an
admin endpoint with no role check, a SQL injection in search, negative cart
quantities, an order total that drops the last line item, login messages that
reveal which usernames exist, and a product page that throws a TypeError.
Four are layout bugs that only appear at certain screen widths. More on
those below.
A bugs.yaml manifest records the ground truth, and asyncqa score grades a
run against it: recall, precision, whether the severity was right, and whether
the right area was blamed.
Round one: agents against the API
Three passes of API, explore and e2e tasks against the functional bugs, entirely local. About 30 minutes, 740K tokens, nothing sent to the cloud.
| Result | |
|---|---|
| Recall | 5 / 7 planted bugs found |
| Precision | 88%, one false positive filed |
| Area blamed correctly | 5 / 5 |
| Severity exactly right | 3 / 5 (the other two off by one level) |
The misses weren’t in unvisited places. Every area got a run. They were unprobed ways. The auth run spent its budget on logout and never compared the error messages for a wrong username and a wrong password. The products run used only the HTTP API, so it never loaded the page in a browser where the JavaScript crash lives. That turned out to be the theme of the whole project: recall is a coverage problem before it’s an intelligence problem, and coverage means techniques, not just places.
The interesting part was the failures.
The false positive that got caught
On its very first run, qwen3-coder reported that login was broken: a 500 error
with a stack trace, every time. It was right that the server crashed. It was
wrong about why. It had passed the JSON body as a string
("{\"username\": ...}") rather than an object, and the app choked on it.
The verifier, gpt-oss with a fresh context, sent a correct login request, got a
200, and ruled reproduced=no, is_bug=no. The finding landed at 29%, status
rejected, no bug file. That’s the pipeline doing exactly its job.
The false positive that didn’t
One run concluded that “session tokens are not invalidated on logout.” To log
out, it had sent GET /logout, got a 404, and carried on as if it had logged
out. The real logout is POST. The verifier followed the same broken steps,
saw the same thing, and agreed. The scoring then made it worse: replaying the
evidence “matched”, because a 200 OK replayed as a 200 OK.
Two fixes came out of that:
- Replaying a normal response isn’t evidence of a bug. A replay only counts heavily if the replayed response was itself an error (a 5xx, a stack trace). Otherwise it shows determinism, not a defect, and now scores +0.3 rather than +1.2.
- The verifier must check that each step did what it claims. If “log out”
returned a 404, the report hasn’t demonstrated the bug. Find the real way to
log out, or answer
reproduced=no.
Round two: bugs that only exist at 1024 pixels
My first version of the demo shop looked like a 1992 homepage. After a redesign I planted the kind of bug that real teams ship all the time: CSS that’s only wrong at certain widths.
| Bug | Width | Symptom |
|---|---|---|
| cart overflow | under 560 px | the cart table forces the page to scroll sideways |
| tablet nav | 960–1199 px | a wrong media query hides the whole Cart link |
| promo covers | 600–767 px | the promo bar grows but the page padding doesn’t |
| long name | 1200 px and up | long product names are cut off with no ellipsis |
The tablet nav bug is my favourite. Here’s the header at 1024 px and at 1280 px:

None of the language models found these. Explorers barely use the browser, and nothing tells them to resize it. So I built the sweep: a task with no model. It crawls the app, renders every page at 360, 480, 640, 768, 1024 and 1280 px, and runs a layout audit in the page. The audit checks for:
- a page wider than the viewport, and which element caused it;
- a button or link whose centre is off-screen;
- a control that’s still covered by another element even when scrolled as far
as the page allows (
elementFromPointis the whole trick); - text pushed past a parent that hides overflow, with no ellipsis;
- a control visible at a narrower and a wider width but missing in between. It’s a surprisingly clean signal for “somebody’s media query is wrong”.
The same issue on 21 pages becomes one finding, with the affected pages and widths listed and an annotated screenshot attached. Here’s the evidence the sweep attached for the promo bug: at 640 px, scrolled all the way down, the footer is buried under the bar with no way to reach it.

| Sweep result | |
|---|---|
| Recall | 5 / 5 (4 layout bugs plus the JS error) |
| Precision | 100% |
| Severity / area | 5 / 5 and 5 / 5 |
Clean build (BUGS=none) | zero findings |
| Cost | 21 pages × 6 widths in about 80 s before verification |
The verifier that saw the bug and said no
This is the most instructive thing I watched all week. gpt-oss was verifying
the tablet nav bug. It set the width to 1024, looked at the page and noted:
“No cart link found in navigation.” It checked 768 and 1280 and noted that
the Cart link was present at both. Then it ruled: reproduced=no, is_bug=no.
It had the evidence in its own notes and still reasoned its way out of it. In an earlier version, that single vote would have rejected a real bug.
The fix is structural, not a better prompt. A fresh browser re-measures every
layout finding first. When that deterministic re-measurement sees the bug and a
model verifier doesn’t, that’s a conflict, not a refutation. The verifier’s
vote counts for less (−0.4 rather than −1.5), the report gets an explicit
“conflict: unresolved” row, and the bug lands at 73% (“likely”) instead of
being rejected. An optional escalate_to setting sends exactly these cases,
and only these, to a stronger model such as Claude.
The general rule I’d take away: let machines measure and models judge, and when they disagree, the measurement doesn’t lose by default.
Traps it encodes
All of these showed up in live runs against real local models, and each is now handled and covered by a test:
- JSON bodies as strings. qwen3-coder sends
"json": "{\"a\": 1}". The tool now decodes it. Without that fix, the model manufactures its own 500s. - Arrays as strings. Evidence IDs arrived as
"[E11]", which the tool iterated character by character ('[','E','1'…). - Paraphrased duplicates. “Admin endpoint accessible by non-admin users” and “Admin-only endpoint accessible by regular users” share few title words. They also share an endpoint, and now that counts too.
- Data-dependent bugs hide from samplers. The sweep first visited 2 product
pages per URL template and missed the one product with a
nulldescription. Sampling per template is now configurable, and it’s the knob to turn first. - Selectors in file names. The first live sweep crashed because a finding’s
fingerprint (
sweep:obscured:div > a) became a Windows folder name. The unit tests skipped verification, so they never touched that path. Run the real thing end to end at least once.
Current status
AsyncQA works end to end on local models. It ships with its own benchmark target, and every bug file explains its score. What it doesn’t do yet:
- The confidence weights are hand-set. The benchmark now produces labelled outcomes, so they should be fitted, not guessed.
- Coverage is per area, not per page or endpoint. Recall for the language models depends on which areas the rotation happens to pick.
- Model swaps dominate run time. One model resident at a time means about 100 s to reload each time driver and verifier alternate. Batching verifications would remove most of it.
- Escalation to Claude is built and tested with scripted models, but I haven’t yet measured it on the live conflict cases.
What’s next
- Escalation, measured: how many conflicts does a stronger model resolve, and what does that cost per confirmed bug?
- Fitted confidence: logistic regression over the signals, using the benchmark’s ground truth as labels.
- A Next.js benchmark app, because client-side rendering and hydration errors deserve a target that actually has them.
Get started
git clone https://github.com/lab-zee/AsyncQA
cd AsyncQA
python -m venv .venv; .venv\Scripts\activate
pip install -e .[dev]; playwright install chromium
cd examples\buggy-shop
python app.py # terminal 1
asyncqa -c asyncqa.yaml models # can your models call tools?
asyncqa -c asyncqa.yaml run --once # one pass of every task
asyncqa -c asyncqa.yaml score bugs.yaml # how did it do?
Point it at your own app with asyncqa init. Start with the sweep. It costs
nothing to run, and it will probably find something.