Built so you can
trust the result.
A vs B exists to answer one question honestly: can this test result carry the decision you are about to make on it?
Start freeWhy this exists
Most A/B testing tools are very good at telling you that you won. Very few tell you when a win is too weak to act on.
That gap costs real money. A test reads as a winner, the change ships, and the promised lift never shows up in revenue. Nobody lied. The tool just never said the result was thin.
A vs B was built by people who run conversion experiments for a living and got tired of that. The whole product is organised around one idea: a result should tell you how much weight it can carry, before you spend anything acting on it.
Three things we will not trade away
Each one is a behaviour you can go and check in the product, not a poster on a wall.
A result should say when it is weak.
Every result is graded from A to F. The grade scores how the test was run, not how much you wanted it to win. The formula is fixed, there is no AI anywhere in it, and nobody on your team or ours can override it.
The tedious part should not be your job.
Building a variation by hand is slow, and slow work gets skipped. Describe the change instead. The copilot edits the real page and hands you a plain-English list of changes to review. Nothing reaches a visitor until you save it.
Nothing changes without a name on it.
A sign-in, an experiment launch, a permission change, a billing change: each one is recorded with who did it, from where, what changed, and whether it came from the dashboard, the API, the command line, or the system itself.
What that looks like in practice
Different questions need different tests, the same way a doctor uses different tests for a broken bone and a fever. You get three ways to read a result, and the page always says which one produced the number in front of you.
Bayesian
Answers "how likely is the challenger actually better", as a plain probability, with no fixed sample size to wait for.
Frequentist
The textbook significance test most stats teams already trust, planned around a sample size you set in advance.
Sequential
Checking the result every day cannot manufacture a false winner.

Slice the results afterwards and the verdict disappears.
You still see every number. What you lose is the ship-or-stop call, and the page says why: a winner found only after you went digging is not a winner yet.
The copilot cannot delete your work.
Ask it to remove something a person built and it hides that thing instead, reversibly. Only what the copilot itself created can truly be deleted. An ownership record on our side enforces that, not an instruction to the model.
The plan seals the moment traffic starts.
What you are measuring, and what counts as a win, is written down before the first visitor arrives. After that it is read-only, so the goalposts cannot move once you have seen the numbers.
Start with a result you can defend.
Free to start. No card, and not a trial.