AI agents & MCP
Last updated: September 12, 2026
Simple Split Testing runs a Model Context Protocol (MCP) server. Point an AI coding agent at it with a token and the agent can create tests, author variants, launch them, and read results without you leaving your editor. Most apps you install. Culsin apps you build with.
One honest note up front. Creating a test defines the experiment, but the variants swap real elements in your theme. The agent hands you the exact Liquid to add (via get_storefront_markup), and can paste it in itself if it has access to your theme files. Connecting is no code. Authoring a variant is one paste into your theme. See Storefront markup.
What you can ask an agent to do
- “Create an A/B test on my hero headline with two variants.”
- “Launch the hero test.” / “Pause it.”
- “How is the checkout-button test doing? Is it significant yet?”
- “Only run this test on product pages tagged
sale, for mobile visitors in the US.” - “Give me preview links for each variant.”
- “Add a custom conversion event called
newsletter_signupand make it the goal.” - “Build me a funnel from cart to purchase and show the drop-off per variant.”
- “Variant B has a higher average order value. What are those buyers actually putting in the basket?”
- “The test is live but showing nothing. What’s wrong?”
- “Which plan am I on and how many sessions have I used?”
The agent has 43 tools covering the full test lifecycle.
1. Generate a token
In the app, open Integrations and click Generate token. Give it a name so you can tell tokens apart later, and pick an expiry of 30 days, 90 days, or 1 year. The token starts with sst_live_ and is shown once, so copy it then.
A token is scoped to one store. It can only read and write that store’s tests, and there is no way to reach another store’s data through it. Revoke it any time from the same page. A revoked or expired token stops working immediately, and so does every token for the store if the app is uninstalled.
2. Connect your client
The MCP server URL is:
https://simple-split-testing.culsin.com/mcp
The token goes in an Authorization: Bearer header, never in the URL. The Integrations page shows the same URL alongside a config block you can copy and drop your token into.
Claude Code
claude mcp add --transport http simple-split-testing \
https://simple-split-testing.culsin.com/mcp \
--header "Authorization: Bearer sst_live_your_token_here"
Then, in a session, ask Claude to list or create split tests.
Cursor
Add this to .cursor/mcp.json in your project, or ~/.cursor/mcp.json for all projects:
{
"mcpServers": {
"simple-split-testing": {
"url": "https://simple-split-testing.culsin.com/mcp",
"headers": { "Authorization": "Bearer sst_live_your_token_here" }
}
}
}
Any MCP client
The server speaks MCP over HTTP (Streamable HTTP transport) and supports protocol versions 2025-06-18 (the default), 2025-03-26, and 2024-11-05. Point any compatible client at the URL with an Authorization: Bearer <token> header:
{
"mcpServers": {
"simple-split-testing": {
"url": "https://simple-split-testing.culsin.com/mcp",
"headers": { "Authorization": "Bearer sst_live_your_token_here" }
}
}
}
3. A worked example
Ask your agent: “Create an A/B test on my hero CTA, control versus a green button, and give me the markup to add.”
- The agent calls
create_split_testwith two variants. The response includes the test and a block of Liquid. - Add that Liquid where the CTA renders in your theme. If the agent has access to your theme files, it can make the edit for you. The markup uses
data-split-idanddata-split-variant. See Storefront markup for the full attribute reference. - Ask: “Show me both versions.” The agent calls
get_preview_urlsand hands back one signed link per variant. Each link renders your live storefront with that variant forced, and the variant sticks for the rest of that browsing session. See Preview links. - Ask: “Is my store wired up?” The agent calls
get_setup_statusand reports the web pixel and the app embed before anything goes live. - Ask: “Launch it.” The agent calls
set_split_test_statusand moves the test toactive. - Once traffic flows, ask: “How is the hero CTA test doing?” The agent calls
get_resultsand reports per-variant numbers and the significance verdict. See Statistics for how those numbers are computed.
Tools
Forty-three tools, grouped by what they do. Every tool works on the store the token belongs to.
Almost every tool that works on one test takes that test’s id. The exception is get_basket_comparison, which takes testId and defaults to your most recent running test if you leave it out.
Tests
| Tool | What it does |
|---|---|
list_split_tests | List your tests with status and variants. Filter by status, and pass withResults: true to get each test’s headline numbers and verdict in the same call |
get_split_test | One test with its variants, audience rules, and configuration |
create_split_test | Create a test from name and 2 or more variants (the first is the control, weight is a relative traffic share). Returns the test plus the storefront markup. Starts as a draft unless you pass status: "active" |
update_split_test | Change name, description, hypothesis, primaryKpi, or autoStopOnSignificance (Growth and above to turn on). On a completed test, hypothesisOutcome records whether the hypothesis held: true, false, or clear |
duplicate_split_test | Copy a test’s variants, audience, and configuration into a new draft. Nothing the original collected comes with it |
set_split_test_status | active launches or resumes, paused stops assigning new visitors, draft reverts to editing, completed ends it |
delete_split_test | Delete a test and everything it recorded. Needs confirm: true once it has data |
get_storefront_markup | The theme Liquid for a test, one block per variant |
get_preview_urls | A signed link per variant. Pass path (for example /products/some-handle) to land on a specific page |
get_test_settings | The confidence level a verdict is judged at, the sample size the test is planned for, the effect it is sized for, and whether visitors stay in the variant they first saw |
set_test_settings | Change any of those. Raising the confidence level makes a result harder to call a winner, so changing it mid-test can move a test from significant to not |
primaryKpi accepts revenue_per_visitor (the default), conversion_rate, aov, add_to_cart_rate, or one of your own events: custom:<event_name> for the share of visitors who fired it, for example custom:newsletter_signup, or custom:<event_name>/order for the share of orders that carry it, for example custom:tool_attached/order with an event the Shopify Flow action records.
Variants and audience
| Tool | What it does |
|---|---|
replace_variants | Replace the variant set (2 or more, first is the control). Pass an existing variant’s handle to keep it across a rename or a new weight. Needs confirm: true on a test that already has visitors |
set_audience | Set who a test applies to, or pass clear: true to target everyone |
set_audience fields all combine with AND, and all are optional:
urlIncludes/urlExcludes: the page URL must contain at least one of these, or none of them.templateIncludes: the page must render with one of these theme templates, as"product"or"product.moonswatch". A bare name matches every alternate template of that type, and a trailing*matches a prefix, for example"product.mission-*".productTagIncludes: a product page whose product carries at least one of these tags. Pages with no product never match, so this alone scopes a test to product pages. No wildcards here.countryIncludes: ISO-2 country codes, for example["US","CA"].device:desktop,mobile, ortablet.visitorType:new,returning, orlogged_in.
See Audience targeting for what each rule matches.
Results and analytics
| Tool | What it does |
|---|---|
get_results | Full results for one test: per-variant visitors, conversions, revenue, the value of the primary metric, custom event counts (fires, distinct visitors, distinct orders), the significance verdict, and metricReadings: the same confidence reading for conversion rate, order value and revenue per visitor, each with a plain call (likely_better, likely_worse, no_difference, not_worse, too_early) so an agent can say whether a winning KPI is costing orders or basket size |
read_docs | These docs, from inside the server: with no argument it lists the pages, with a page slug it returns that page as markdown. The same pages are exposed as MCP resources for clients that browse them |
get_basket_comparison | What each variant’s buyers put in their orders: the share of orders containing each group, units per order, and revenue per order, against the baseline variant. Group with by (product_type, product, vendor, or property) and set a window with days |
get_results takes an optional range: today, yesterday, 7d, 30d, 90d, 365d, or all. It defaults to all, the whole life of the test. A verdict over a window is a verdict on part of the evidence, so use a window to see recent movement and leave it off to decide anything.
Custom events
| Tool | What it does |
|---|---|
list_custom_events | The events your tests can measure |
create_custom_event | Define an event by eventName. Lowercase letters, digits, and underscores, 1 to 64 characters, and it cannot reuse a built-in event name such as checkout_completed |
update_custom_event | Change an event’s description, the properties it breaks down by, or the numeric value it carries. The name cannot change, because every firing already recorded is stored against it |
delete_custom_event | Remove the definition. Firings already recorded are kept |
Custom events need the Growth plan or above. See Custom events.
Funnels
| Tool | What it does |
|---|---|
list_funnels | Your saved funnels and their steps |
create_funnel | Save a funnel from a name and an ordered list of steps. Steps can be Shopify’s own events, so a funnel can be built and read without touching the theme |
update_funnel | Rename a funnel or replace its steps. The new funnel reports over the same history |
delete_funnel | Delete the definition. Every event behind it is kept |
get_funnel_results | Read one funnel against one test: how many of each variant’s visitors reached each step, the drop-off between steps, and whether a step’s difference is significant |
list_funnel_presets | Ready-made funnels built only from events every Shopify storefront already sends |
apply_funnel_preset | Save a preset as a funnel. Takes the preset’s key |
Saved views
| Tool | What it does |
|---|---|
get_view_vocabulary | Everything a view can be built from on your store: measures, time grains, breakdowns, every event name that resolves, and the presets |
list_views | Your saved views and their definitions |
create_view | Save a view from name, events, measure, grain, and range |
update_view | Rename a view or change its definition. Anything referencing it keeps working |
delete_view | Delete the definition. Every event behind it is kept |
get_view_results | Run a saved view. Pass range to read the same definition over a different window without changing the saved one |
apply_view_preset | Save a preset as a view. Takes the preset’s key |
Visitors and activity
| Tool | What it does |
|---|---|
find_visitor | Look up one visitor by query: an email address, a Shopify customer id, a cart token, or a visitor id |
list_visitors | The visitors bucketed into one test, newest first, with the variant each saw. Filter to one variantId, or to hasRevenue: true for visitors who spent money |
get_visitor_timeline | Everything one visitor did on one test, in order. Takes id and userId |
list_recent_activity | The raw event stream for one test, newest first. This is the tool for “is anything coming in at all” |
list_event_types | The event types one test has actually recorded. The quickest check that a custom event is arriving |
Setup and store settings
| Tool | What it does |
|---|---|
get_setup_status | Whether your store is wired up to collect data: the web pixel, and whether the app embed is enabled on your published theme. Pass recheck: true right after enabling the embed |
get_shop_settings | Whether orders from visitors in a test get tagged in your Shopify admin, plus the currency and timezone results are reported in |
set_order_tagging | Turn order tagging on or off with enabled |
Plans
| Tool | What it does |
|---|---|
get_plan_usage | Your plan, its limits, and how much of them is used. -1 means unlimited |
list_plans | Every plan with its price and what it includes, alongside the one you are on, plus the page where you change it |
Guardrails
An agent hits the same rules you would in the admin.
- A test with no markup in your theme shows every visitor the control, and nothing reports that as an error. Creating the test is half the job. Adding the markup is the other half.
- Launching needs the Split Embed Script app embed enabled on your published theme. Activating a test, creating one directly as active, and resuming a paused one are all refused while it is off. A well-behaved agent calls
get_setup_statusbefore it launches, and it is the first thing to check when a live test shows no data. delete_split_testandreplace_variantsrequireconfirm: truewhen the test is active or paused and already has assigned visitors. The refusal names the visitor count, so the agent has a real number to show you before it retries. Editing a draft needs no confirmation.set_audienceandautoStopOnSignificanceneed the Growth plan or above. Clearing targeting and turning auto-stop off work on any plan.- Custom events need the Growth plan or above, and a name of 1 to 64 characters using lowercase letters, digits, and underscores only. Built-in event names such as
checkout_completedare taken. - Changing your plan is a charge you approve in your Shopify admin, so no tool switches it.
list_plansreturns the page where you do it. - Everything else the app does is available as a tool, so an agent should never send you to a screen to click something.
Plan limits
The MCP enforces the same limits as the app: monthly tested sessions, concurrent tests, and custom events. Free runs 1 test at a time and no custom events. Paid plans raise or remove those caps. If a call would exceed your plan, the tool returns a clear message telling you to finish a test, delete an event, or upgrade, along with the page to upgrade on. Ask the agent for get_plan_usage before it starts creating things.
Measure an order fact instead
Some conversions are facts about an order rather than something a visitor clicks: a warranty attached, a bundle bought, a subscription started. Those are better recorded from Shopify Flow than from the pixel. Build a workflow that starts from an order and add the Record split test event action. It writes the event under the visitor who placed the order, and creates the custom event if it does not exist yet. See Shopify Flow.
In Shopify Sidekick, with no token
The same read-only tools are available inside Shopify’s own Sidekick assistant, with no token and no client setup. Install the app and ask Sidekick about your tests in the admin.
Sidekick can call list_split_tests, get_split_test, get_results, get_setup_status, get_storefront_markup, get_preview_urls, list_funnels, get_funnel_results, get_basket_comparison, list_views, get_view_results, list_custom_events, list_recent_activity, get_test_settings, and get_plan_usage.
It reads and never writes. To create a test, launch one, or change anything, use the admin or a token-connected agent.
Good to know
- The token goes in the
Authorizationheader, never in a URL. Treat it like a password. - Tests an agent creates are identical to tests created in the admin. Edit or launch them from either side.
- Changing variants on a running test re-buckets visitors, so prefer editing while a test is still a draft.
- Uninstalling the app stops every token for the store. Reinstalling does not bring them back, so generate a new one from the Integrations page.
- The agent can read results but cannot see a visitor’s personal data. Variant payloads carry handles, not names or customer information.