Free Tool · A/B Testing

A/B Test Calculator

Before your experiment, calculate the visitors and test duration you need. After it, check statistical significance and sample ratio mismatch (SRM) — all for free.

The sample size, duration, MDE, and significance calculators use the industry standard: 95% confidence level, 80% statistical power, and a two-tailed z-test. The SRM check uses a chi-square goodness-of-fit test at a 0.01 significance level.

Sample Size Calculator

Enter your current conversion rate and the expected uplift you want to detect to calculate the visitors needed per variation for your A/B test.

Visitors needed per variation
53,211

Test Duration Calculator

Factor in your average daily visitors and the number of variations to calculate how many days your A/B test should run.

Required test duration
107 days

Even if the calculation says less, we recommend running experiments for at least 14 days (2 weeks) to account for day-of-week differences in visitor behavior.

Statistical Significance Calculator

Enter the visitors and conversions for each group of a completed A/B test to determine whether the difference in conversion rates is statistically significant.

The difference is statistically significant
p-value
0.0175
Control conversion rate
3%
Variant conversion rate
3.6%
Relative uplift
+20%

Sample Ratio Mismatch (SRM) Checker

Enter the observed sample size and expected ratio for each group to check whether traffic was split as intended, using a chi-square goodness-of-fit test. Summary statistics are all you need.

No indication of sample ratio mismatch (SRM)

The gap between the observed group sizes and the expected ratios is within normal sampling variation.

p-value
0.1552
Chi-square statistic
2.02
Degrees of freedom
1
Total sample size
19,800
GroupObservedExpectedObserved / expected share
110,0009,90050.51% / 50%
29,8009,90049.49% / 50%

A p-value below 0.01 is flagged as a sample ratio mismatch. Run the SRM check to validate data quality before you judge experiment results.

Minimum Detectable Effect (MDE) Calculator

Enter the visitors you can secure per variation to calculate the smallest uplift that traffic can statistically detect.

Minimum detectable uplift (relative)
10.32%

Why calculate sample size before testing?

What is A/B test sample size?

Sample size is the minimum number of visitors each variation needs for your A/B test to reach a trustworthy conclusion. With too small a sample, you may fail to detect a real improvement — or mistake a random difference for one.

Why you should calculate sample size before you start

Repeatedly checking results and stopping the moment they look significant (peeking) greatly inflates false positives. Decide the required sample size and duration before you start, and judge the results only after meeting them — that is what makes an experiment trustworthy.

Understanding minimum detectable effect (MDE)

MDE (Minimum Detectable Effect) is the smallest uplift you want your experiment to detect. The smaller the MDE, the more sample you need — steeply. For example, at a 3% conversion rate, detecting a 10% relative uplift takes about 53,000 visitors per variation, while detecting 5% takes about 210,000. Setting a realistic MDE for the traffic you can secure is the heart of experiment design.

What is sample ratio mismatch (SRM)?

Sample ratio mismatch (SRM) is when the number of samples actually assigned to each group differs sharply from the intended split. If one arm of a 50:50 experiment is clearly smaller, you may have a data quality issue — a randomization failure, missing events, bot traffic, or load failures in a specific environment. Conversion comparisons are distorted in that case, so find the cause before judging results. The SRM checker runs a chi-square goodness-of-fit test using only observed sample sizes and expected ratios, and warns about a mismatch when the p-value is below 0.01.

The statistical standards behind the calculators

The sample size, duration, MDE, and significance calculators are based on a two-proportion z-test (two-tailed) with a 95% confidence level (α = 0.05) and 80% statistical power (β = 0.2) — the industry-standard settings adopted by most A/B testing tools. The SRM checker answers a different question, so it uses a chi-square goodness-of-fit test at a 0.01 significance level.

Frequently Asked Questions

Enter your current conversion rate and the expected uplift (MDE) you want to detect, and the calculator returns the visitors needed per variation. Enter your average daily visitors as well to see the required test duration.

A 95% confidence level limits the probability of concluding there is a difference when there is none (a type I error) to 5%. 80% power means that when a real difference exists, the test detects it 80% of the time. Both are the most widely used standards in A/B testing.

Even if you can fill the calculated sample size sooner, we recommend running for at least 14 days (2 weeks). Visitor behavior differs by day of the week, so tests shorter than two weeks can be skewed by day-of-week effects.

Yes. Detecting a small uplift requires a large sample, so if your traffic is limited, first use the MDE calculator to see what uplift you can realistically detect. It is more efficient to start with experiments that make bigger changes — like a full page redesign — than small copy tweaks.

It means the difference in conversion rates between two groups is unlikely to be due to chance. Typically, a result is considered statistically significant when the p-value is below 0.05 (at a 95% confidence level).

These calculators provide estimates based on the standard formula (two-proportion z-test) used at the experiment design stage. The Hackle dashboard runs frequentist and Bayesian statistical analyses on your actual experiment data, so use the dashboard's analysis for final judgments on experiment results.

First use the SRM checker to see whether the gap is within what chance can explain. A p-value below 0.01 means randomization or data collection is likely broken. In that case, do not take the results at face value: check where the SDK is integrated, whether events are missing on specific platforms or browsers, and whether bot traffic is filtered — then rerun the experiment.

Done calculating? Time to start experimenting

Start A/B testing with a single SDK integration on Hackle and see statistically analyzed results in real time.

Start Hackle for free