The A/B Testing feature in Subotiz lets merchants analyze experiment performance from the Experiment Data tab after an experiment begins running. Compare Conversion Data, Revenue & Risk Data, and custom metrics across experiment groups. Subotiz also evaluates experiment results and confidence levels to generate AI Insights, helping merchants determine whether performance differences are statistically reliable and whether a better-performing version should be rolled out to all users.
Accessing Experiment Data
- Open Experiment Data: Sign in to the Subotiz Admin, go to Data > A/B Testing, select the experiment you want to analyze, and open the Experiment Data tab.The top of the page displays the following overview information:
- Duration: The total time elapsed since the experiment was activated.
- Users in Experiment: The number of users who have entered the experiment and been assigned to an experiment group.
Use these values to understand the current experiment progress and sample size.Conversion Data and Revenue & Risk Data are displayed in separate columns for each experiment group. The column names match the group names configured when the experiment was created.
- Identify Metric Types: Review how each type of experiment metric is displayed.
- Count Metrics: Display only the raw value for each group. They do not include the relative change or confidence level. Examples include Visitors, Payments Succeeded, and Paid Orders.
- Rate and Amount Metrics: Display the actual value, the relative change compared with the Control Group, and the confidence level for that difference. Examples include Payment Success Rate, Visitor Conversion Rate, GMV, and ARPU.When an experiment has only recently started or includes a small number of participants, the page may show a large relative change even when the confidence level remains low.For example, GMV may show:
- Control Group: $64,580.28
- Experiment Group: $0.00
- Relative Change: -100%
- Confidence Level: 0.0%
This usually indicates that the current sample size is too small to determine whether the Experiment Group is actually underperforming. It does not necessarily mean that the experiment version will reduce GMV. Continue running the experiment and collecting more data before drawing conclusions.
Analyzing Conversion Data
The Conversion Data section shows the primary customer journey from visiting the experiment page to completing a successful payment. Each stage displays the corresponding number of users and conversion rate.
Count metrics display only the raw value for each group. Rate metrics also display the relative change compared with the Control Group and the confidence level. Confidence levels are not calculated for count metrics.
- Visitors: The number of unique users who viewed the experiment page. This is the starting point of the conversion flow and therefore does not have a corresponding conversion rate.
- Checkouts Initiated / Checkout Initiation Rate: The number of users who opened the checkout page and the percentage of visitors who initiated checkout.
- Orders Submitted / Order Submission Rate: The number of users who completed the required checkout information and clicked the subscription or purchase button, and the percentage of visitors who submitted an order.
- Payments Attempted / Payments Attempted Rate: The number of users who completed the payment authentication or payment action, and the percentage of visitors who reached this step. Completing a payment attempt does not mean the payment was successful. A transaction may still fail because of risk controls, insufficient funds, a bank decline, or another payment-related issue.
- Payments Succeeded / Payment Success Rate: The number of users who completed a successful payment and the corresponding percentage of visitors.
- Visitor Conversion Rate: The percentage of visitors who completed a successful payment. Use this metric to evaluate the overall conversion performance from page view to successful payment.
Comparing Revenue and Risk Metrics
In addition to Conversion Data, merchants can compare Revenue & Risk Data across experiment groups.
Paid Orders is a count metric and displays only the raw value. GMV, ARPU, Subscription Renewal Rate, and Refund Rate also display the relative change compared with the Control Group and the confidence level.
- Paid Orders: The number of successfully paid orders in each experiment group.
- GMV: The total value of successfully paid transactions in each experiment group. The current GMV calculation includes the original amount of orders that are later refunded.
- ARPU: The average revenue generated per paying user in an experiment group.Calculation: GMV ÷ Paying UsersUse this metric to determine whether an experiment version increases average revenue per paying user.
- Subscription Renewal Rate: The percentage of users who enter a second subscription billing cycle. When a trial user successfully converts to the first paid subscription cycle, Subotiz also counts that event as a renewal.Renewal events occur only after the current subscription period ends. As a result, the displayed renewal rate includes only users whose relevant subscription period has already reached its end date.To obtain more complete renewal data, continue monitoring the experiment for at least one full subscription cycle after it stops accepting new users.Examples:
- 7-day trial: Continue monitoring for approximately 7 days.
- Monthly subscription: Continue monitoring for approximately 30 days.
- Annual subscription: Set the observation period based on the actual subscription cycle and business requirements.
- Refund Rate: The percentage of users who receive a successful refund among all users who completed a successful payment. Use this metric to determine whether an experiment version introduces additional refund or post-purchase risk.
Evaluating Confidence Levels and Recommendations
The confidence level indicates whether the observed difference between the Experiment Group and the Control Group is likely caused by the tested change rather than random variation.
A higher confidence level indicates a more statistically reliable result. A large increase or decrease does not automatically mean the result is reliable. Always evaluate the relative change together with the number of Users in Experiment and the confidence level.
Subotiz uses the experiment configuration, experiment results, and A/B testing analysis rules to generate AI Insights on the Experiment Data page. Use these insights together with the following guidelines when evaluating experiment results.
AI Insights are intended to support decision-making and should not replace business judgment. Final decisions should also consider sample size, business performance, implementation requirements, and actual operating conditions.
- Significant Improvement: A confidence level of 95% or higher and a positive relative change indicate that the Experiment Group significantly outperformed the Control Group. Consider rolling out the Experiment Group version.
- Significant Decline: A confidence level of 95% or higher and a negative relative change indicate that the Experiment Group performed significantly worse than the Control Group. Avoid rolling out the tested change.
- Potential Opportunity: A confidence level between 80% and 95% with a positive relative change suggests that the Experiment Group is currently performing better, but the result has not yet reached statistical significance. Continue running the experiment and collect more data.
- Insufficient Data: A confidence level below 80% suggests that the observed difference may be caused by random variation. Increase the experiment sample size and avoid making decisions based only on the current relative change.
- No Meaningful Difference: A relative change close to 0 indicates that both groups are performing similarly and that the tested change has not produced a meaningful impact.
Confidence level measures statistical reliability, not business impact. Even when a result reaches a high confidence level, merchants should also consider revenue impact, risk, customer experience, and implementation cost before making a final rollout decision.
Analyzing Custom Metrics
When custom metrics are configured during experiment creation, the Experiment Data page displays the corresponding Direct Metrics and Derived Metrics for each experiment group.
- Compare Custom Metrics: Review each group's actual metric value, relative change compared with the Control Group, and confidence level.Custom metrics can measure business outcomes that Subotiz does not collect automatically, including:
- Page button clicks
- Feature usage
- Registration step completion
- Content engagement
- Merchant-defined business events
The accuracy of custom metrics depends on the SDK event reporting and metric configuration. Before analyzing the results, confirm that all experiment groups use the same event collection rules and measurement definitions.
Reviewing Data Updates and Measurement Rules
Review the following data rules when analyzing experiment results to avoid misinterpreting performance.
- Allow Data to Update: Experiment data updates as frequently as every five minutes. User actions do not appear on the page immediately.
- Continue Monitoring Delayed Metrics: After an experiment is completed or terminated, no new users can enter. However, delayed metrics generated by existing participants, such as renewals and refunds, continue to update.
- Distinguish User Measurement Methods: A/B Testing assigns groups and calculates user counts based on the Experiment User ID. This measurement method differs from the user-counting logic in the Data module, so the number of users displayed on the two pages may not match.
- Verify Order Experiment Information: Exported transaction files include an A/B Experiment field that records the Experiment ID and assigned group associated with each order. Use the exported data to further validate the experiment's impact on orders and revenue.
- Avoid Early Conclusions: When the experiment includes only a small number of participants, a large relative change may still lack statistical reliability. Prioritize the confidence level and results collected across a complete business cycle.
Use the Duration and Users in Experiment values to understand experiment progress before comparing Conversion Data, Revenue & Risk Data, and custom metrics across experiment groups. Count metrics display raw values only, while rate and amount metrics also include relative changes and confidence levels.
Avoid making rollout decisions based only on large percentage changes when the sample size is small or the confidence level is low. By evaluating confidence levels, AI Insights, revenue impact, and risk performance together, merchants can make more reliable decisions about rolling out the best-performing version to all users.