TrenyxStart
Menu
The method

AI mutation testing

AI mutation testing uses a model to write realistic bugs, the kind a developer or a coding assistant could actually ship, plants each one in a copy of your code, and runs your own tests against it. If no test fails, your suite has a gap. Trenyx does this every month: 16 bugs per repo, a score, and for every miss, the test that would have caught it.

Where it comes from

Classic mutation testing, briefly.

Mutation testing is an old and good idea: break the code on purpose and see whether the tests notice. Tools such as Stryker, PIT, cargo-mutants and mutmut do it mechanically. They flip operators (< becomes <=, + becomes -), delete statements and swap constants across the codebase, often thousands of times, and report the share of those changes your tests catch.

Two things limit it in practice. Many of the changes are noise: some alter nothing anyone could observe, and others are edits nobody would ever write. And the bugs that hurt most are rarely a single flipped operator. They are a refund applied twice, a permission check that reads the wrong policy, a cache key that leaves out the one field that makes it unique.

The difference

What changes when a model writes the bugs.

Classic mutation testingAI mutation testing (Trenyx)
What gets plantedMechanical edits to operators, constants and statementsBugs a developer or assistant could plausibly ship, written for your code
How manyHundreds to thousands per run16 per repo, every month
Bugs specific to your productRare: the operators don't know your domainAt least 3 per run that only make sense in your product
NoiseChanges that alter nothing, or that nobody would writeEach bug is written to change real behaviour
What you getA mutation scoreA score, every miss, and the test that would catch it
CostFree and open source; you run itUS$150 a month per repo; we run it
Independence

Why the score can be trusted.

A planted-bug score is only worth something if the bugs weren't aimed at your gaps. Three rules keep it independent:

Your repository is never touched: the tests run in a private copy. More on the method in how we verify.

Coverage

Why code coverage doesn't answer the question.

Coverage tells you a line ran during the tests. It doesn't tell you whether anything would notice if that line were wrong. In our public sample report, an open-source Stripe billing library's suite caught 4 of 16 planted bugs, and 9 of the 12 it missed sat in code the tests executed, at least in part. Line coverage would have marked much of that code as covered.

Choosing

When classic mutation testing is the better choice.

Use a mutation tool when you want exhaustive, local feedback on a small library with a fast suite, or a mutation score on every pull request. It's free and runs on your machine. Use planted realistic bugs when you want to know whether your suite would catch the mistakes that actually ship, measured by someone who didn't write the tests. Many teams will want both.

The bugs

Where the bugs come from.

From a catalogue of more than 200 kinds of mistake, built from what broke in more than a hundred real AI-built codebases, and we track which ones test suites keep missing. For each repo we pick the kinds that fit the code (money handling gets unit mix-ups, logins get flipped checks) and write at least three that only make sense in that product.

Try it on one repo.

Every month we plant 16 realistic bugs in your code, one at a time, run your own tests against each, and hand you the score, every miss, and the test that would have caught it. If the repo is public, we can run a free preview first.

Request a preview See a sample report