SI Signal & Funnel
Growth Experiments

A/B Testing Basics Without Peeking for a Winner

A/B Testing Basics Without Peeking for a Winner
SummaryRun an A/B test by randomly assigning a defined eligible unit to a stable control or variant, predefining the primary outcome, guardrails, sample and stopping plan, and validating assignment and exposure before analysis. Compare groups according to the planned method, report uncertainty, missingness, deviations, and practical magnitude, and label post-hoc segments as exploratory. Do not stop at a favorable peek or treat an inconclusive result as proof of no effect.

An A/B test compares randomized experiences

An A/B test assigns eligible experimental units randomly to a control or variant, then compares a predefined outcome under a documented analysis. Random assignment helps balance other factors in expectation. It does not correct weak tracking, treatment changes, or conclusions selected after repeated unplanned checks.

Use the method only when the intervention, audience, and potential harm have passed the required legal, privacy, accessibility, security, and operational review.

Define the unit and eligibility

Choose whether assignment occurs by person, account, browser, location, household, campaign, or another unit. The unit must match the treatment and interference risk. A person seeing both versions can contaminate a person-level comparison.

Document inclusion, exclusions, repeat exposure, devices, consent states, markets, and timing. Do not expand eligibility after seeing which group performs better.

Predefine outcomes and guardrails

Select one primary outcome for the decision and define event, denominator, source, window, and delayed validation. Add guardrails for errors, complaints, returns, unsubscribes, accessibility, quality, or another relevant harm.

The landing-page plan provides a hypothesis structure that connects treatment, mechanism, and outcome.

Plan sample and analysis before launch

Specify the effect size relevant to the decision, baseline estimate, allocation, power or precision goal, test statistic or interval, runtime considerations, and stopping method. These choices require context and may need qualified statistical expertise; this article does not provide a universal sample size.

Do not choose sample size from short-term traffic availability. Do not continue testing only until statistical significance appears. Repeated unplanned checking increases false-positive risk.

Validate randomization and exposure

Test assignment, persistence, version delivery, event collection, time zones, duplicate units, consent behavior, and downstream outcomes. Inspect whether allocation and baseline characteristics show implementation problems without turning every random imbalance into a reason to rewrite the experiment.

Use the data-quality checklist before reading performance. Exclude internal tests only through the predefined method.

Analyze the assigned groups

Start with the analysis matching assignment, often an intention-to-treat comparison, because removing noncompliant units after assignment can reintroduce bias. Report group counts, outcome estimates, uncertainty, guardrails, missingness, and deviations.

Do not call a statistically detectable result important without comparing its size with the business decision. Do not call an inconclusive result “no effect”; it may reflect genuine similarity or insufficient precision.

Limit unplanned subgroup analysis

Predefine important segments and limit them. Post-hoc exploration can generate future hypotheses, but label it exploratory and do not promote the most favorable slice to a confirmed finding.

Privacy and fairness matter when segmenting people. Do not infer sensitive or protected traits for optimization without a valid purpose, the relevant official regulator guidance named below, and qualified local counsel.

Decide and preserve the record

Use the experiment backlog to connect the result to rollout, revision, another test, or no change. Archive hypothesis, versions, code or configuration, dates, exclusions, results, and decision.

Record the result, implementation change, guardrails, and monitoring plan. Treat the estimate as evidence for this tested context, not as a universal rule.

Review allocation and exposure

Confirm the assignment unit, eligible population, traffic split, exclusions, and first-exposure timestamp from actual records. Check that a unit does not switch variants unexpectedly and that exposure precedes the outcome. Document contamination or imbalance before interpreting the comparison.

Official rule sources

Data-protection and direct-marketing duties depend on jurisdiction, data, purpose, and message. Check the current official source relevant to the people and activity: the European Commission data-protection portal for EU scope, the UK Information Commissioner's Office direct-marketing guidance updated 28 April 2026, the California Privacy Protection Agency laws and regulations for California scope, and the U.S. Federal Trade Commission CAN-SPAM guide for U.S. commercial email. These official pages do not determine whether a rule applies to a specific business. Also check current platform documentation and contracts, and use qualified local privacy or legal counsel for consequential decisions.

General marketing education, not legal, privacy, tax, financial, security, or individualized business advice. An independent publication. Not affiliated with any prior owner of this domain.

FAQ

How large should an A/B test sample be?

Sample needs depend on baseline outcome, decision-relevant effect, allocation, variability, desired precision or power, analysis, and stopping method. There is no universal number. Plan before launch using reliable inputs and appropriately qualified statistical expertise when consequences matter. Do not choose the sample solely from convenient traffic or extend repeatedly until a favorable threshold appears.

Can I check A/B test results every day?

Operational checks for safety, implementation, privacy, and data failure are essential. Repeatedly testing the performance result and stopping when it looks favorable can inflate false positives unless a valid sequential method was predefined. Separate health monitoring from decision analysis, follow the planned stopping rule, and document any pause. Qualified statistical help may be required for sequential designs.

What does an inconclusive A/B test mean?

It means the evidence and planned analysis did not support a sufficiently precise decision under the tested conditions. It is not automatic proof that the variants are equal. Report the estimate and uncertainty, check implementation and missingness, and decide whether more information could resolve the question. Avoid selecting a winner from a favorable subgroup discovered after the main result.

Which official privacy and marketing sources should I check?

Use the official source that matches the people, jurisdiction, data, and activity: the European Commission data-protection portal for EU scope, the UK Information Commissioner's Office direct-marketing guidance for UK scope, California Privacy Protection Agency laws and regulations for California scope, and the U.S. Federal Trade Commission CAN-SPAM guide for U.S. commercial email. Then check current platform documentation and contracts. Qualified local counsel should review consequential decisions.