Product Analytics

How Product Teams Run A/B Tests with Session Replay

October 5, 2026

Tymek Bielinski

Product Growth at LiveSession
Table of content

A/B testing is a randomized controlled experiment that shows two versions of something, a control (A) and a variant (B), to different users at random, then measures which one performs better on a chosen metric. It replaces guesswork with evidence: instead of debating which headline or button color “feels right,” you let real user behavior decide, measured against an Overall Evaluation Criterion (OEC) and checked at a commonly used confidence level near 95%.

Livesession
livesession.io
See What Users Do After Testing
Use session replay, engagement metrics, heatmaps, and funnels to understand how users experience each product variation.
Book a demo

How A/B Testing Works: Control, Randomization, and the OEC

Every A/B test rests on a comparison between a control and a treatment. The control is the existing version, the page, email, or feature your users already see. The variant is the changed version, built to test one specific idea. Changing only one element at a time (a headline, a price, a layout) is what lets you attribute any shift in behavior to that single change rather than to a tangle of simultaneous edits.

Randomization is what turns a simple before-and-after comparison into something closer to proof. When you assign visitors to the control or the variant purely by chance, you spread out every other factor that might influence behavior, device type, time of day, traffic source, mood, whatever, roughly evenly between the two groups. That balance is what allows you to say the outcome difference was caused by your change rather than by some hidden variable. The Stanford guide to controlled experiments treats this random assignment as the defining feature that separates a true experiment from an observational comparison.

Traffic allocation decides how visitors get split between versions. A 50/50 split is standard for a simple two-way test because it reaches statistical confidence fastest. Weighted splits (say, 90/10 in favor of the control) make sense when you’re nervous about a risky change and want to limit exposure while still collecting data.

None of this matters without a clear target. The Overall Evaluation Criterion, or OEC, is the single metric you’ve decided will determine whether the variant wins. Teams often default to the metric that’s easiest to move rather than the one that matters, which is a trap. Concrete OEC examples include:

  • Conversion rate: the share of visitors completing a signup, purchase, or other defined action.

  • Revenue per user: useful when a change might lift conversions but shrink average order value.

  • Retention or repeat engagement: better suited to features meant to build habit rather than drive a single transaction.

Picking the right OEC before launch, rather than scanning a dashboard afterward for something that moved, is one of the most consistent pieces of advice in experimentation literature. The Stanford guide notes that teams relying on the “Highest Paid Person’s Opinion” instead of a defined OEC tend to test trivial changes and miss the bigger wins available from user-driven hypotheses. For SaaS teams specifically, choosing metrics that reflect long-term business health rather than vanity numbers is worth a closer look in our piece on SaaS metrics.

One more thing that gets skipped: instrumentation. If your event tracking is broken or inconsistent between the two groups, your results are measuring noise, not behavior. Before trusting any outcome, confirm the events firing for the control and the variant are logged the same way, at the same points in the user flow.

Key Statistical Ideas: Significance, Confidence, and Power

A/B tests run on a handful of statistical concepts that are easy to misuse if you skip the definitions. The null hypothesis is the default assumption that your variant makes no real difference, that any gap you observe is just random noise. “Statistical significance” means the data gives you enough evidence to reject that assumption with a stated level of confidence, not that the result is big, important, or permanent.

The Stanford guide to controlled experiments also recommends aiming for statistical power typically between 80% and 95%, the probability your test will detect a real effect if one exists. Low power means you can run a well-designed test and still miss a genuine improvement simply because you didn’t collect enough data.

A test with 80% power and a 5% significance threshold, as recommended in the Stanford guide, balances the risk of a false positive against the risk of missing a real effect, which is why those numbers show up so often as the default in experimentation tools.

Effect size and sample size are locked in a trade-off: the smaller the improvement you’re trying to detect, the more users you need to detect it reliably.

Expected effect size Relative sample size needed
Large (for example, doubling conversion) Smallest
Moderate (a few percentage points) Medium
Small (fractions of a percent) Largest, often impractically large for low-traffic sites

This is why a small blog testing a button color rarely reaches significance in a reasonable time frame, while a high-traffic checkout flow can detect tiny shifts within days. Variance-reduction techniques like CUPED, described in follow-up Stanford research, can shrink the sample size required for a given effect by controlling for pre-experiment user behavior, letting teams run more tests with the same traffic.

Before trusting any live experiment, many practitioners run an A/A test, sending two groups of users to the identical experience and checking whether the tool reports a false “winner.” Done correctly, an A/A test should flag significance at roughly the nominal rate, about 5% of the time at a 0.05 threshold, according to the Stanford guide. If it fires far more often than that, something in your instrumentation or randomization is broken.

Types of Experiments: A/B, A/A, Multivariate, and Beyond

Not every question calls for the same experimental design. Picking the right one depends on how many variables you’re testing and how much traffic you have to spend.

  • A/B testing: the default choice for a single, clear hypothesis, changing one headline, one price, one layout, and comparing it against the current version.

  • A/A testing: not a test of a real change but a sanity check, confirming your tracking and randomization behave as expected before you trust any other result.

  • Multivariate and factorial testing: these test several changes and their interactions at once (headline and image and button color, for instance) but need dramatically more traffic to reach significance on each combination.

  • Specialized designs: marketplaces, social networks, and other systems where one user’s experience affects another’s face interference and spillover effects, which a simple random split can’t account for.

Two-sided marketplaces are a good example of where the basic recipe breaks down. The Stanford GSB explainer on A/B testing notes that simple A/B tests often fail in these settings because changing the experience for one group of users (say, drivers in a ride-share test) inevitably spills over into the experience of the other group (riders), contaminating the comparison. These situations call for specialized designs like switchback tests or cluster-based randomization.

Medical research uses the same underlying logic as A/B testing but under far stricter rules. A review of randomized controlled trials notes that clinical RCTs require blinding, ethics review, and regulatory oversight that digital product experiments typically skip, since the stakes and the subject matter differ enormously. If you’re weighing more advanced experimental setups once the basics feel comfortable, our guide to advanced A/B testing strategies walks through several of them.

Running Your First A/B Test: A Step-by-Step Checklist

A trustworthy test follows a fixed sequence. Skipping steps is how teams end up with results nobody trusts.

  1. Write a specific hypothesis that names the change, the expected effect, and why you expect it (not just “test the button”).

  2. Choose one OEC tied to a real business outcome, conversion rate, revenue per user, or retention, and commit to it before launch.

  3. Estimate your required sample size using your baseline conversion rate, your minimum detectable effect, and your target power (80 to 95% is standard, per the Stanford guide).

  4. Decide your stopping rule in advance: either a fixed horizon (run until you hit your calculated sample size, then stop) or a sequential method built for continuous monitoring.

  5. Run an A/A test first to confirm your tracking and randomization are working as expected.

  6. Launch the real test, split traffic randomly, and resist the urge to check results daily and act on early swings.

  7. Analyze the pre-specified OEC once you’ve reached your target sample size or stopping rule, and record what you learned, win or lose.

Pro Tip: Write your hypothesis and your stopping rule down somewhere visible before launch, so you can’t quietly move the goalposts once the data starts rolling in.

The step most beginners get wrong is step six. Checking a dashboard every morning and declaring victory the moment a p-value dips below 0.05 is called peeking, and it dramatically raises your odds of a false positive, a problem covered in detail in the next section. If you want to monitor continuously anyway, you need statistical methods built for it, not the standard fixed-horizon math. For a shorter version of this checklist geared toward quick wins, see our tips for better A/B tests.

Common Pitfalls That Quietly Ruin A/B Test Results

Most bad A/B test conclusions trace back to a small set of repeat offenders.

  • Peeking: stopping a test the moment it looks significant, before reaching the pre-calculated sample size, inflates your false-positive rate far beyond the 5% you think you’re accepting.

  • Multiple comparisons: testing many metrics or many variants at once raises the odds that at least one shows a “significant” result purely by chance, so apply a correction or treat secondary metrics as exploratory, not conclusive.

  • Interference and spillover: in marketplaces or networked systems, one group’s experience can leak into another’s, contaminating a simple random split.

  • Instrumentation errors and contamination: bots, duplicate accounts, or broken event tracking can quietly skew results in either direction, so spot-check raw session data, not just the aggregate dashboard.

Research on early stopping shows that continuous monitoring without the right statistical safeguards can inflate false positives well beyond the stated significance threshold, which is why always-valid p-values were developed to let teams check results in real time without that penalty. These methods use sequential statistical machinery designed specifically for ongoing monitoring, rather than the fixed-horizon math most tools assume by default.

The fix for most of these issues is procedural, not statistical: decide your stopping rule and your primary metric before you launch, and treat every other metric you glance at along the way as context, not a verdict. A quick instrumentation sanity check, comparing raw session counts against your analytics dashboard, catches tracking problems before they corrupt a multi-week test. The Nielsen Norman Group’s primer on A/B testing echoes this, recommending minimal, single-variable changes precisely because they’re easier to audit when something looks off.

Real Examples: CTA Color, Subject Lines, and Feature Rollouts

Abstract steps click into place faster with concrete cases.

  • Landing-page CTA color: hypothesis is “a higher-contrast button increases clicks,” OEC is conversion rate through the signup funnel, and if your baseline traffic is low, expect the test to run longer than you’d like because the effect size is typically small.

  • Email subject line: hypothesis targets open rate or click rate, and results need a fixed time window (say, 48 hours) since early engagement patterns differ by send time and inbox behavior.

  • Feature flag rollout: a staged rollout to a small percentage of users doubles as a built-in A/B test, and watching session replays during that window often reveals UI confusion that a conversion number alone would never explain.

For a SaaS-specific version of the first two examples, including how pricing-page tests differ from standard landing-page tests, see our guide to A/B testing SaaS pricing pages.

Why Qualitative Data Sharpens Every A/B Test

Numbers tell you something changed. They rarely tell you why. Session replay and heatmaps fill that gap by showing exactly where users hesitate, misc lick, or abandon a flow, which turns a vague hunch into a testable hypothesis grounded in real behavior rather than a guess about what might work.

Once a test is running, funnel and event tracking confirm whether the OEC actually moved for the reason you expected, or whether a drop in conversion was really an instrumentation gap dressed up as a result. We built session replay, engagement metrics, heatmaps, conversion funnels, and error tracking to support both ends of that process, shaping hypotheses before launch and diagnosing winners and losers after. All of it runs under a privacy posture built around GDPR and CCPA compliance, which matters when you’re recording real user sessions.

A Beginner’s Priorities: What to Get Right First

The habit that separates useful experimentation from noise is patience: avoiding the dashboard between the launch and the pre-set stopping point. The second habit is pairing every number with a reason. Watching a handful of session replays or a heatmap before and after a change tells you whether a win is real insight or a lucky fluke.

See Your A/B Test Results in Context with LiveSession

Running the test is half the job. Understanding why a variant won, or didn’t, is the other half, and that’s where session replay, heatmaps, conversion funnels, and error tracking earn their place in an experimentation workflow. We built LiveSession to let product teams watch real sessions from either arm of a test, spot the friction a conversion number alone can’t explain, and confirm whether a result reflects genuine behavior or a tracking gap.

Livesession

If you want to see how that fits into your own testing workflow, check our pricing plans, which include a free tier and paid plans, or start with a live demo.

FAQ

What does A/B testing mean?

A/B testing means randomly showing two versions of something, a control and a variant, to different users and comparing their performance on a chosen metric. The goal is to find out which version produces a better outcome, measured with statistical confidence rather than a guess, as described in the Stanford guide to controlled experiments.

What is A/B testing for beginners, in simple terms?

Think of it as a side-by-side comparison with a referee: half your visitors see version A, half see version B, and you track which group does better on a single, predefined metric. The Nielsen Norman Group recommends keeping the change minimal, one headline or one button, so you know exactly what caused the difference.

What is A/B testing in medical research?

In medicine, the equivalent design is called a randomized controlled trial (RCT), which follows the same core logic of random assignment but under much stricter protocols. A review of RCT methodology notes that clinical trials require blinding, ethics approval, and regulatory oversight that digital A/B tests typically don’t need.

How is A/B testing different from multivariate testing?

A/B testing changes one element at a time and compares two versions, while multivariate testing changes several elements at once and tests their combinations. Multivariate designs can reveal how changes interact, but they need far more traffic to reach a reliable result for each combination.

How long should an A/B test run?

A test should run until it reaches the sample size calculated from your baseline conversion rate, expected effect size, and target power, not for a fixed number of days chosen arbitrarily. Stopping early based on a promising-looking dashboard, a practice known as peeking, can inflate false positives well beyond the stated significance threshold, as explained in research on always-valid inference.

Sources

Tymek Bielinski

Product Growth at LiveSession
Tymek Bielinski works in Product Growth at LiveSession, focusing on driving growth and go-to-market strategies. As an avid learner, he shares insights and explores the world of product growth alongside others.
Learn more about your users
Test all LiveSession features for 14 days, no credit card required.

Get Started for Free

Join thousands of product people, building products with a sleek combination of qualitative and quantitative data.

Free 14-day trial
No credit card required
Set up in minutes