AI mutation testing uses a model to write realistic bugs, the kind a developer or a coding assistant could actually ship, plants each one in a copy of your code, and runs your own tests against it. If no test fails, your suite has a gap. Trenyx does this every month: 16 bugs per repo, a score, and for every miss, the test that would have caught it.
Mutation testing is an old and good idea: break the code on purpose and see whether the tests notice. Tools such
as Stryker, PIT, cargo-mutants and mutmut do it mechanically. They flip operators (< becomes
<=, + becomes -), delete statements and swap constants across the
codebase, often thousands of times, and report the share of those changes your tests catch.
Two things limit it in practice. Many of the changes are noise: some alter nothing anyone could observe, and others are edits nobody would ever write. And the bugs that hurt most are rarely a single flipped operator. They are a refund applied twice, a permission check that reads the wrong policy, a cache key that leaves out the one field that makes it unique.
| Classic mutation testing | AI mutation testing (Trenyx) | |
|---|---|---|
| What gets planted | Mechanical edits to operators, constants and statements | Bugs a developer or assistant could plausibly ship, written for your code |
| How many | Hundreds to thousands per run | 16 per repo, every month |
| Bugs specific to your product | Rare: the operators don't know your domain | At least 3 per run that only make sense in your product |
| Noise | Changes that alter nothing, or that nobody would write | Each bug is written to change real behaviour |
| What you get | A mutation score | A score, every miss, and the test that would catch it |
| Cost | Free and open source; you run it | US$150 a month per repo; we run it |
A planted-bug score is only worth something if the bugs weren't aimed at your gaps. Three rules keep it independent:
Your repository is never touched: the tests run in a private copy. More on the method in how we verify.
Coverage tells you a line ran during the tests. It doesn't tell you whether anything would notice if that line were wrong. In our public sample report, an open-source Stripe billing library's suite caught 4 of 16 planted bugs, and 9 of the 12 it missed sat in code the tests executed, at least in part. Line coverage would have marked much of that code as covered.
Use a mutation tool when you want exhaustive, local feedback on a small library with a fast suite, or a mutation score on every pull request. It's free and runs on your machine. Use planted realistic bugs when you want to know whether your suite would catch the mistakes that actually ship, measured by someone who didn't write the tests. Many teams will want both.
From a catalogue of more than 200 kinds of mistake, built from what broke in more than a hundred real AI-built codebases, and we track which ones test suites keep missing. For each repo we pick the kinds that fit the code (money handling gets unit mix-ups, logins get flipped checks) and write at least three that only make sense in that product.
Every month we plant 16 realistic bugs in your code, one at a time, run your own tests against each, and hand you the score, every miss, and the test that would have caught it. If the repo is public, we can run a free preview first.