← Notes

Introducing AsyncQA: A Local Model That Files Bug Reports You Can Check

A long-running QA agent driven by models on a 12 GB RTX 5070. It explores a web app, verifies every suspected bug independently, and attaches the evidence and the arithmetic behind each confidence score. Tested against a shop with 11 planted bugs.

Introducing AsyncQA: A Local Model That Files Bug Reports You Can Check

The QA problem is trust, not effort

Pointing a model at a web app and asking it to find bugs is easy. It will happily report a dozen. The hard part is the report after that: is any of it real?

An agent that files ten bugs, three of them imaginary, is worse than no agent. Every report costs a human five minutes to disprove, and after the second phantom you stop reading them. So the question I cared about wasn’t “can a local model find bugs.” It was:

Can a model running on a consumer GPU keep testing an app around the clock and file bug reports you’d actually trust?

That breaks into three smaller questions. Can it find anything? Can the system tell its real finds from its hallucinations? And can you see why it believes each one?

Enter AsyncQA

AsyncQA is a daemon. You give it an app (a URL, local or in Docker) and some credentials. It keeps running tasks on a schedule:

TaskModel?What it does
healthnoboolean checks: HTTP, TCP, Docker, shell commands
sweeponly to verifyrenders every page at six screen widths, measures layout, and catches JS errors and 5xx responses
explore / apiyescharter-based exploratory testing through a browser and HTTP
e2eyesa goal-driven user journey (“log in, buy two things, check the total”)
chaosnorestarts or pauses a container and measures time to recovery
recheckyesre-verifies open bugs and marks them possibly fixed

It runs in two modes. Blind gets a URL and a login and maps the app itself. Prescribed gets a list of areas with test charters and invariants (“a user can only see their own orders”).

The driver can be Claude, an OpenAI model, or anything on Ollama. The default config here is entirely local: qwen3-coder:30b-a3b explores and gpt-oss:20b verifies. That’s the same pair from the Bakeoff trial, on the same 12 GB RTX 5070.

The part that matters: a bug has to survive verification

Nothing the explorer reports goes straight into a bug file. Every suspected bug goes through a pipeline:

  1. Dedupe against known bugs, by title similarity and by which endpoints the evidence touched.
  2. Deterministic replay. The HTTP exchanges it cites are re-sent. For layout bugs, a fresh browser measures the page again. No model is involved in this step.
  3. An independent verifier. A fresh context, ideally a different model, given the claim and the steps but not the finder’s reasoning. It has to reproduce the bug with its own tools, then rule on two separate questions: did it happen, and is it actually wrong?
  4. Scoring. Confidence is a sum of log-odds, with one visible term per piece of evidence.

That last step is the one I’d push hardest on anyone building something similar. Every bug file ends with its own arithmetic:

## Why this confidence
- +0.85  prior from finder self-confidence 1.00
- +1.20  deterministic replay reproduced the error response
- +1.50  verifier #1: reproduced=yes, is_bug=yes
- = 97%

A number you can audit is far more useful than a number you’re asked to accept.

One small design note: the model never sees a password. It writes {{cred.alice.password}}, the tool layer substitutes the real value at send time, and every byte of evidence is redacted back before it’s saved. Session tokens the app mints at runtime are masked in the files meant for humans.

A target with known answers

You can’t evaluate a bug finder on an app whose bugs you don’t know. So AsyncQA ships with one: Buggy Shop, a small, modern storefront with planted defects. Each can be toggled with an environment variable, and BUGS=none gives a clean build for measuring false positives.

The Buggy Shop storefront

There are eleven planted bugs. Seven are functional: an IDOR on orders, an admin endpoint with no role check, a SQL injection in search, negative cart quantities, an order total that drops the last line item, login messages that reveal which usernames exist, and a product page that throws a TypeError. Four are layout bugs that only appear at certain screen widths. More on those below.

A bugs.yaml manifest records the ground truth, and asyncqa score grades a run against it: recall, precision, whether the severity was right, and whether the right area was blamed.

Round one: agents against the API

Three passes of API, explore and e2e tasks against the functional bugs, entirely local. About 30 minutes, 740K tokens, nothing sent to the cloud.

Result
Recall5 / 7 planted bugs found
Precision88%, one false positive filed
Area blamed correctly5 / 5
Severity exactly right3 / 5 (the other two off by one level)

The misses weren’t in unvisited places. Every area got a run. They were unprobed ways. The auth run spent its budget on logout and never compared the error messages for a wrong username and a wrong password. The products run used only the HTTP API, so it never loaded the page in a browser where the JavaScript crash lives. That turned out to be the theme of the whole project: recall is a coverage problem before it’s an intelligence problem, and coverage means techniques, not just places.

The interesting part was the failures.

The false positive that got caught

On its very first run, qwen3-coder reported that login was broken: a 500 error with a stack trace, every time. It was right that the server crashed. It was wrong about why. It had passed the JSON body as a string ("{\"username\": ...}") rather than an object, and the app choked on it.

The verifier, gpt-oss with a fresh context, sent a correct login request, got a 200, and ruled reproduced=no, is_bug=no. The finding landed at 29%, status rejected, no bug file. That’s the pipeline doing exactly its job.

The false positive that didn’t

One run concluded that “session tokens are not invalidated on logout.” To log out, it had sent GET /logout, got a 404, and carried on as if it had logged out. The real logout is POST. The verifier followed the same broken steps, saw the same thing, and agreed. The scoring then made it worse: replaying the evidence “matched”, because a 200 OK replayed as a 200 OK.

Two fixes came out of that:

  • Replaying a normal response isn’t evidence of a bug. A replay only counts heavily if the replayed response was itself an error (a 5xx, a stack trace). Otherwise it shows determinism, not a defect, and now scores +0.3 rather than +1.2.
  • The verifier must check that each step did what it claims. If “log out” returned a 404, the report hasn’t demonstrated the bug. Find the real way to log out, or answer reproduced=no.

Round two: bugs that only exist at 1024 pixels

My first version of the demo shop looked like a 1992 homepage. After a redesign I planted the kind of bug that real teams ship all the time: CSS that’s only wrong at certain widths.

BugWidthSymptom
cart overflowunder 560 pxthe cart table forces the page to scroll sideways
tablet nav960–1199 pxa wrong media query hides the whole Cart link
promo covers600–767 pxthe promo bar grows but the page padding doesn’t
long name1200 px and uplong product names are cut off with no ellipsis

The tablet nav bug is my favourite. Here’s the header at 1024 px and at 1280 px:

Header at 1024px: no Cart link Header at 1280px: Cart link present

None of the language models found these. Explorers barely use the browser, and nothing tells them to resize it. So I built the sweep: a task with no model. It crawls the app, renders every page at 360, 480, 640, 768, 1024 and 1280 px, and runs a layout audit in the page. The audit checks for:

  • a page wider than the viewport, and which element caused it;
  • a button or link whose centre is off-screen;
  • a control that’s still covered by another element even when scrolled as far as the page allows (elementFromPoint is the whole trick);
  • text pushed past a parent that hides overflow, with no ellipsis;
  • a control visible at a narrower and a wider width but missing in between. It’s a surprisingly clean signal for “somebody’s media query is wrong”.

The same issue on 21 pages becomes one finding, with the affected pages and widths listed and an annotated screenshot attached. Here’s the evidence the sweep attached for the promo bug: at 640 px, scrolled all the way down, the footer is buried under the bar with no way to reach it.

At 640px the promo bar covers the footer

Sweep result
Recall5 / 5 (4 layout bugs plus the JS error)
Precision100%
Severity / area5 / 5 and 5 / 5
Clean build (BUGS=none)zero findings
Cost21 pages × 6 widths in about 80 s before verification

The verifier that saw the bug and said no

This is the most instructive thing I watched all week. gpt-oss was verifying the tablet nav bug. It set the width to 1024, looked at the page and noted: “No cart link found in navigation.” It checked 768 and 1280 and noted that the Cart link was present at both. Then it ruled: reproduced=no, is_bug=no.

It had the evidence in its own notes and still reasoned its way out of it. In an earlier version, that single vote would have rejected a real bug.

The fix is structural, not a better prompt. A fresh browser re-measures every layout finding first. When that deterministic re-measurement sees the bug and a model verifier doesn’t, that’s a conflict, not a refutation. The verifier’s vote counts for less (−0.4 rather than −1.5), the report gets an explicit “conflict: unresolved” row, and the bug lands at 73% (“likely”) instead of being rejected. An optional escalate_to setting sends exactly these cases, and only these, to a stronger model such as Claude.

The general rule I’d take away: let machines measure and models judge, and when they disagree, the measurement doesn’t lose by default.

Traps it encodes

All of these showed up in live runs against real local models, and each is now handled and covered by a test:

  • JSON bodies as strings. qwen3-coder sends "json": "{\"a\": 1}". The tool now decodes it. Without that fix, the model manufactures its own 500s.
  • Arrays as strings. Evidence IDs arrived as "[E11]", which the tool iterated character by character ('[', 'E', '1'…).
  • Paraphrased duplicates. “Admin endpoint accessible by non-admin users” and “Admin-only endpoint accessible by regular users” share few title words. They also share an endpoint, and now that counts too.
  • Data-dependent bugs hide from samplers. The sweep first visited 2 product pages per URL template and missed the one product with a null description. Sampling per template is now configurable, and it’s the knob to turn first.
  • Selectors in file names. The first live sweep crashed because a finding’s fingerprint (sweep:obscured:div > a) became a Windows folder name. The unit tests skipped verification, so they never touched that path. Run the real thing end to end at least once.

Current status

AsyncQA works end to end on local models. It ships with its own benchmark target, and every bug file explains its score. What it doesn’t do yet:

  • The confidence weights are hand-set. The benchmark now produces labelled outcomes, so they should be fitted, not guessed.
  • Coverage is per area, not per page or endpoint. Recall for the language models depends on which areas the rotation happens to pick.
  • Model swaps dominate run time. One model resident at a time means about 100 s to reload each time driver and verifier alternate. Batching verifications would remove most of it.
  • Escalation to Claude is built and tested with scripted models, but I haven’t yet measured it on the live conflict cases.

What’s next

  1. Escalation, measured: how many conflicts does a stronger model resolve, and what does that cost per confirmed bug?
  2. Fitted confidence: logistic regression over the signals, using the benchmark’s ground truth as labels.
  3. A Next.js benchmark app, because client-side rendering and hydration errors deserve a target that actually has them.

Get started

git clone https://github.com/lab-zee/AsyncQA
cd AsyncQA
python -m venv .venv; .venv\Scripts\activate
pip install -e .[dev]; playwright install chromium

cd examples\buggy-shop
python app.py                                 # terminal 1
asyncqa -c asyncqa.yaml models                # can your models call tools?
asyncqa -c asyncqa.yaml run --once            # one pass of every task
asyncqa -c asyncqa.yaml score bugs.yaml       # how did it do?

Point it at your own app with asyncqa init. Start with the sweep. It costs nothing to run, and it will probably find something.