Bell Statistics

Welcome to Bell's Calculators System

Looking for an A/B test sample size calculator? Want to check whether your data has a sample ratio mismatch (SRM)? Or are you ready to analyze your experiment results? Whatever stage you're at, you're in the right place.

A/B test calculators

One experiment, three moments where the statistics matter. Size it before you launch, check the split once traffic is flowing, and read the result when it lands.

  1. Stage 1 — Planning

  2. Stage 2 — Validation

  3. Stage 3 — Analysis

Calculators for common tests

Every other test, in one list: the ones that turn up most often, the exact tests for small tables, and the intervals and equivalence questions an ordinary significance test can't answer.

Not sure which one you need?

No worries. Take our short diagnostic and we'll guide you to the right one — a few questions about what you are comparing and how the data was collected.

Bell's Calculators System puts every statistics calculator you need in one place, so there's no more hunting across different sites for a sample size calculator, an A/B testing calculator, or a t-test calculator. Every one of them shares the same consistent format, which makes the whole system easier to use and understand.

Want to go deeper? Most online calculators hand you a number with little insight into how it was calculated. Ours give you the key outputs plus everything you need to understand them: when to use the calculator, its assumptions, and the formula behind it. A worked example turns the theory into practical numbers you can check against your own.

We build measurement systems for a living — see our A/B testing work or the case studies. These tools are the parts of that work that fit on one page.

Frequently asked questions

Are these calculators free?
Yes, all of them, with no sign-up and no limits. They run entirely in your browser — nothing you type is sent to us or to anyone else, which also means you can use them on data you would not paste into a third-party website.
How do I know which test to use?
Take the diagnostic above, or work from three things yourself: what the outcome looks like, how many groups you are comparing, and whether the observations are paired. A counted outcome across two independent groups is a proportions test; a measured outcome across two independent groups is a two-sample t-test; the same subjects measured twice is a paired test. Each calculator's page opens with a section on when to use it and when to use something else instead.
Why does every calculator have two tabs?
Because sample size and analysis are the same statistics asked in opposite directions, and separating them across different tools is how people end up analysing an experiment with a method that does not match the one they designed it with. Sharing the code makes that mismatch impossible.
How do I know the numbers are right?
Every calculation is checked against published reference values from R in an automated test suite that runs before anything is deployed, and each page names the exact method used so you can reproduce it. Where authorities genuinely disagree — McNemar sample size, or the two-sided convention for Fisher's exact test — the page says which convention it follows and who disagrees.
Can I share a result with someone?
Yes. The inputs are kept in the page's address, so copying the URL from the address bar sends someone the same calculation with the same numbers filled in. Pasted raw data is deliberately excluded from the link, so sharing a result never shares the underlying observations.
How do I check for a sample ratio mismatch?
Enter the number of users who landed in each group as the observed counts, and the split you intended — 1 and 1 for an even test, or 9 and 1 for a 90/10 ramp — as the expected proportions. A small p-value means assignment is not doing what you told it to, and the metric is not worth reading until that is fixed. It is a chi-square goodness-of-fit test, which is why the SRM check and the goodness-of-fit calculator are the same page.

A calculator answers one question. We answer the rest.

Sample size is the easy part. Choosing the metric, handling the users who appear in both groups, deciding what to do when the result lands in the awkward middle — that is the work. If you would rather have that designed properly the first time, talk to us.