Skip to main content
How Can We Help?
Running Statistically Valid Tests
A/B test results become more useful as enough visitor and conversion data is collected to distinguish meaningful differences from normal variation.
You don’t need to be a statistician to run useful experiments in Bluvia A/B Testing. A few good testing habits can help you avoid making decisions too early or drawing conclusions from limited data.
Understand Statistical Confidence
Statistical confidence helps indicate how strongly the experiment data supports the observed difference between the Control and a Variant.
Bluvia A/B Testing uses a 95% confidence level by default.
You can adjust the confidence level when configuring your experiment, but 95% provides a useful starting point for most tests.
Higher confidence requires stronger evidence before a result is considered statistically significant. Lowering the confidence threshold may produce a result sooner, but also increases the chance that the observed difference is due to normal variation rather than the Variant itself.
Don’t Declare a Winner Too Early
Early experiment results can look convincing even when very little data has been collected.
For example, a Variant might have several conversions while the Control has none during the first few visitors. That doesn’t necessarily mean the Variant will continue performing better as more visitors participate.
Avoid declaring a winner based only on an early lead.
Give the experiment enough time and traffic for the results to become more reliable.
Pay Attention to Sample Size
Confidence isn’t the only factor to consider.
The number of visitors participating in an experiment also matters.
A result based on a small number of visitors or conversions may be less dependable, even when one Variant appears to have a substantial lead.
As more visitors participate, you get a better picture of how the Control and Variants perform across your broader audience.
See Understanding Sample Size for more information.
Give the Experiment Enough Time
Avoid stopping a test simply because it reaches your confidence threshold quickly.
Website behavior can change depending on:
- Day of the week
- Weekday versus weekend traffic
- Marketing campaigns
- Traffic sources
- Seasonal activity
- Changes in visitor behavior
Running an experiment across a representative period helps reduce the chance that a short-term traffic pattern determines your result.
See How Long Should a Test Run? for additional guidance.
Avoid Repeatedly Reacting to Results
It’s natural to check an experiment while it’s running, but avoid making decisions every time the leading Variant changes.
Experiment results can fluctuate as new visitors and conversions are recorded.
Instead, evaluate the experiment based on the overall evidence once it has collected enough meaningful data.
Avoid Changing the Experiment Mid-Test
Changing the Control, Variants, Conversion Goal, targeting, or other important experiment settings after data has already been collected can make your results harder to interpret.
If you need to make a significant change, consider creating a new experiment instead.
Duplicating the existing test can give you a starting point while preserving the results of the original experiment.
Remember that editing an Active experiment in Bluvia A/B Testing will pause it.
Avoid Overlapping Tests
You can run multiple experiments simultaneously in Bluvia A/B Testing, but those experiments shouldn’t interfere with one another.
For example, avoid running two experiments that modify the same section of a page or influence the same visitor behavior.
If two experiments overlap, it can become difficult to determine which change caused the observed result.
Tests on unrelated pages or experiences can generally run simultaneously when one experiment isn’t likely to influence the outcome of another.
Look Beyond Confidence
A strong experiment decision considers more than a single percentage.
Before declaring a winner, consider:
- Statistical confidence
- Number of visitors
- Number of conversions
- Conversion rates
- Conversion lift
- How long the experiment has been running
- Whether the test ran across a representative traffic period
- Whether the result is meaningful enough to matter to your website
A statistically confident result isn’t automatically an important result.
For example, a very small improvement may reach your confidence threshold with enough traffic but still have little practical impact on your business.
Use the Results to Make a Decision
Once your experiment has collected enough data, review the complete report rather than relying only on the current leading Variant.
If the evidence supports a Variant, you can select it as the winner. When a Variant is selected as the winner, Bluvia A/B Testing applies that Variant’s page or content to the Control.
If the Control performs better or the experiment remains inconclusive, that’s still useful information. Use what you learned to inform your next hypothesis.
Related Articles
Continue with: