For most of the last decade, Shopify merchants who wanted to test theme changes or deploy them carefully had two real options: a third-party A/B testing platform with the associated cost and integration overhead, or no testing at all. Earlier this year, that calculus changed. Shopify shipped two new tools — SimGym, which runs AI-simulated shoppers through a storefront to surface friction before launch, and Rollouts, which enables scheduled deployments and native A/B testing on live traffic — that together form a credible toolkit for any merchant on the platform. Each tool has a distinct job, a distinct set of limits, and a place in a sensible release workflow. Used well, they meaningfully reduce the risk of theme changes for merchants who previously had no safety net. Used poorly, they create a false confidence that the testing question has been solved. Both interpretations are common, and the difference lies in how the toolkit is understood as a whole.
Why Shopify just changed the testing conversation
The pattern that's held across most of Shopify's merchant base is that testing has been a privilege of the most mature operators. Brands running structured CRO programs paid for dedicated platforms — Shoplift, VWO, Optimizely — and treated testing as a discipline. Everyone else ran theme changes the same way they always had: push the change live, watch the analytics, and hope nothing breaks at scale.
That gap has been the single largest source of avoidable post-launch incidents in Shopify retail. Redesigns go live and immediately tank conversion. Holiday themes ship with broken promotional logic that nobody saw because nobody tested in production conditions. Navigation changes are deployed at 100% on Monday morning and create 3 days of confused customer service tickets before someone reverts. None of these failures required a sophisticated testing program to prevent. They required any testing program at all.
The arrival of SimGym and Rollouts changes that floor. Both are built into Shopify natively, both work without third-party scripts, and both make the previously difficult workflows of pre-launch simulation and controlled deployment immediately accessible to any merchant. They do not replace a structured experimentation program; that distinction matters, and we'll come back to it, but they meaningfully raise the baseline of what every Shopify merchant can do.
SimGym: the AI-simulated dress rehearsal
SimGym is the more conceptually novel of the two. It introduces a new testing approach for Shopify merchants: pre-launch behavioural simulation with AI shoppers.
The mechanic is straightforward. SimGym deploys AI users that browse a store the way real customers would, moving through collections, interacting with products, adding items to cart, and providing qualitative feedback on the experience. The simulation can run against a single theme (live or draft) for a stand-alone audit, or compare two themes, typically a live theme against a draft, so a team can evaluate the impact of proposed changes before publishing them.
The output is two layers of insight. The first is a quantitative comparison: when two themes are tested head-to-head, the winner is the one with the higher add-to-cart rate from the simulated shoppers. The second is qualitative — synthesized feedback on how easy the store is to navigate, how intuitive the structure feels, where the friction points are, and what the AI users couldn't figure out. For most teams, the qualitative output is more useful than the quantitative score. It surfaces specific decisions to revisit, not just a verdict on which theme is better.
The strongest use cases are the ones that have historically been hardest to test: theme changes before publishing, seasonal layouts before campaign launches, navigation reorganizations before customers encounter them, and pre-launch audits of draft themes during the development cycle. For any of these, SimGym is a dress rehearsal, a way to expose obvious friction without ever putting it in front of a real customer.
The caveats are also worth naming. SimGym is currently Liquid-only, while Hydrogen and headless storefronts aren't supported. Stores need Shopify Network Intelligence enabled, can't be password-protected during simulation, and need existing products and content for the simulation to be meaningful. And critically, AI shoppers approximate real customer behaviour; they don't replicate it. A clean SimGym pass is a useful signal, not a guarantee. It's the dress rehearsal, not the opening night.
SimGym also isn't free indefinitely: it runs on a pay-per-credit model, one credit per simulation, and while Shopify has been allocating trial credits during the AI Research Preview, merchants should expect a per-simulation cost, once those trial credits run out.
Rollouts: scheduled deployments and native A/B testing
Rollouts are the more operationally important of the two, because they change the basic deployment model for theme changes on Shopify.
The mechanic: from the admin, a merchant can create a rollout, name it, assign a traffic percentage, make theme-editor changes against that rollout, and optionally schedule a start and end date. The changes are only visible to visitors who fall into the rollout's traffic slice, while the live store remains untouched for everyone else. Run two rollouts simultaneously at a 50/50 split, and the setup becomes a true A/B test. Unlike third-party A/B testing tools that inject JavaScript and resolve the variant assignment after the page begins rendering, Rollouts splits traffic before the page renders, which means no script conflicts, no flicker, and no measurable impact on page speed.
The practical capabilities open up several workflows that previously required either third-party tooling or significant developer effort. Theme changes can be scheduled for specific dates and times, which is ideal for BFCM, flash sales, and seasonal campaigns, with an auto-revert at the end date, eliminating manual cleanup. True A/B tests can be run on any theme-editor change, from hero sections to PDP layouts to CTA copy. Graduated deployments, starting at 10% of traffic and scaling up as confidence builds, become a built-in part of the release process rather than a custom workflow. And changes can be targeted to specific Shopify Markets, allowing a merchant to test a localized homepage for Canada without affecting the US store.
The limits are equally important to understand. Rollouts is a theme editor tool only; it does not support Liquid templates or code-level changes. It does not segment audiences beyond traffic percentage and Markets, meaning a merchant cannot show variant A to returning customers and variant B to new visitors. It does not support price testing or checkout-flow testing. It provides raw analytics on conversion rate, AOV, and revenue, but does not calculate statistical significance, confidence intervals, or revenue per visitor. Once a rollout is applied, the changes cannot be reverted through Rollouts itself; they must be manually undone in the theme editor.
Full A/B experiment analytics also require Advanced or Plus plans, which matters for merchants on lower tiers who can run rollouts but can't analyze them as deeply.
The honest framing is that Rollouts is a deployment and risk-mitigation tool that supports A/B testing as a secondary capability. Shopify's own announcement positioned it as a way to "feel confident in your changes," not as a CRO platform. For merchants who were doing zero testing, it's a meaningful unlock. For merchants running a structured experimentation program, it complements but does not replace a dedicated tool.
How the new tools fit together in a single workflow
The interesting question is not which tool to choose. It's how to use them together as a sequence.
A well-designed release workflow for a meaningful theme change now looks something like this. A draft theme gets built. Before publishing, SimGym runs against the draft to surface obvious friction points in navigation, product discovery, and the add-to-cart flow, issues that should be fixed before any real customers encounter them. Once the team has addressed the simulated-shopper feedback, the theme is published as a Rollout at a low traffic percentage, typically 10 to 20 percent, with the rest of the traffic continuing to see the existing theme. Performance metrics get watched for the first few days. If the rollout looks healthy, traffic is scaled up gradually until the new theme reaches 100 percent. If it doesn't, the rollout is revised or pulled, and only a small fraction of customers experience a problematic issue.
This is the dress-rehearsal-then-graduated-launch model. It used to require a custom workflow built across multiple tools. It is now the default capability of the Shopify platform.
For teams running campaign launches, the workflow is slightly different but structurally similar. The campaign theme is built in advance, run through SimGym to validate it works as expected, scheduled as a Rollout to deploy at the campaign start time, and configured to auto-revert at the campaign end date. Manual deployment risk is eliminated. The marketing team can build BFCM creative 2 weeks early without worrying about someone needing to be online at midnight to push the change live.
What previously required developer hours, project management, and operational risk now requires only admin configuration and a SimGym credit.
Where the native toolkit stops
The clearest way to understand where SimGym and Rollouts stop is to identify the testing scenarios they were not built to address.
Statistical rigour is the first. A Rollout will show that variant B has a higher conversion rate than variant A. It will not tell you whether the difference is statistically significant, what the confidence interval is, or how to think about Bayesian posterior probability. For decisions that are reversible but expensive, and most CRO decisions fit that profile, the absence of a significance calculation is a meaningful gap. A two-point conversion lift in a Rollout could be real or could be noise, and there's no built-in mechanism to distinguish between them. For low-stakes decisions, that's fine. For decisions that drive major design or merchandising changes, it isn't.
Audience segmentation is the second. Rollouts split traffic randomly within a percentage allocation; it does not target by device, UTM source, referrer, customer status (new versus returning), or behaviour. The most interesting questions in CRO often live in those segments, does the new layout work better for mobile users? Does the new PDP perform differently for repeat buyers? Rollouts cannot answer them.
Code-level testing is the third. Rollouts operate on theme-editor changes only. Tests that require Liquid modifications, API-level changes, or anything beyond what the theme editor exposes are out of scope.
Price testing and checkout-flow testing are also outside the native toolkit. For brands where these are the highest-leverage tests available, and for many, they are, a dedicated CRO platform remains necessary.
The practical answer for serious experimentation programs is that SimGym and Rollouts complement a platform like Shoplift; they don't replace it. SimGym handles the pre-launch validation that even disciplined CRO programs typically skip. Rollouts handles the scheduling and risk-managed deployment that previously required custom developer work. The dedicated CRO platform handles the statistical rigour, segmentation, and depth of analysis that the native tools don't offer. Each tool sits in its part of the workflow and earns its place.
A practical decision framework
For Shopify merchants deciding which tool fits which situation, the simplest framing is by intent.
If the goal is to validate a draft theme before publishing, to catch obvious friction in navigation, product discovery, or checkout flow before any customer sees it, use SimGym. It's a single-credit operation, the feedback is qualitative and actionable, and it takes minutes to run. There is no reason for a Shopify merchant building a new theme not to run SimGym against the draft before publishing.
If the goal is to deploy a theme change with controlled exposure, to schedule a seasonal launch, or to run a simple A/B test on a theme-editor change, use Rollouts. The cost is zero, the operational risk is dramatically lower than a 100% launch, and the scheduling capability alone eliminates an entire category of manual deployment errors.
If the goal is to run a structured CRO program with statistical rigour, audience segmentation, price testing, or analysis beyond the native metrics, a dedicated platform like Shoplift remains the right tool. Rollouts can serve as the entry point that proves testing is worth investing in. The dedicated platform is what the program graduates to.
The progression Shopify is enabling is roughly this: merchants who were doing no testing can now start with SimGym and Rollouts at a low cost. Merchants who find value in that work can expand into structured experimentation with a dedicated CRO platform. The native toolkit is the on-ramp. The dedicated platform is the destination for any brand serious about testing as a discipline.
The reframe
For a long time, Shopify's testing story was a story of dependencies — on apps, on developers, on workflows built outside the platform. The arrival of SimGym and Rollouts changes that. The baseline capability for any merchant to validate changes before launch and deploy them safely is now built into the platform itself.
That doesn't make every Shopify merchant a sophisticated tester overnight. The discipline of running a real experimentation program, the cadence, the hypothesis backlog, the statistical rigour, the organizational memory, still has to be built by the team that wants to operate it. But the floor has moved. The case for not testing at all has grown significantly harder to justify, while the case for testing the right things, in the right sequence, with the right tools, has become much easier to operationalize.
The merchants who will get the most from this shift aren't the ones treating SimGym and Rollouts as a complete solution. They're the ones who recognize the toolkit for what it actually is, a meaningful baseline, fit into a broader testing discipline that still needs to be deliberately designed.