AI agents & MCP

Last updated: September 12, 2026

Simple Split Testing runs a Model Context Protocol (MCP) server. Point an AI coding agent at it with a token and the agent can create tests, author variants, launch them, and read results without you leaving your editor. Most apps you install. Culsin apps you build with.

One honest note up front. Creating a test defines the experiment, but the variants swap real elements in your theme. The agent hands you the exact Liquid to add (via get_storefront_markup), and can paste it in itself if it has access to your theme files. Connecting is no code. Authoring a variant is one paste into your theme. See Storefront markup.


What you can ask an agent to do

  • “Create an A/B test on my hero headline with two variants.”
  • “Launch the hero test.” / “Pause it.”
  • “How is the checkout-button test doing? Is it significant yet?”
  • “Only run this test on product pages tagged sale, for mobile visitors in the US.”
  • “Give me preview links for each variant.”
  • “Add a custom conversion event called newsletter_signup and make it the goal.”
  • “Build me a funnel from cart to purchase and show the drop-off per variant.”
  • “Variant B has a higher average order value. What are those buyers actually putting in the basket?”
  • “The test is live but showing nothing. What’s wrong?”
  • “Which plan am I on and how many sessions have I used?”

The agent has 43 tools covering the full test lifecycle.


1. Generate a token

In the app, open Integrations and click Generate token. Give it a name so you can tell tokens apart later, and pick an expiry of 30 days, 90 days, or 1 year. The token starts with sst_live_ and is shown once, so copy it then.

A token is scoped to one store. It can only read and write that store’s tests, and there is no way to reach another store’s data through it. Revoke it any time from the same page. A revoked or expired token stops working immediately, and so does every token for the store if the app is uninstalled.


2. Connect your client

The MCP server URL is:

https://simple-split-testing.culsin.com/mcp

The token goes in an Authorization: Bearer header, never in the URL. The Integrations page shows the same URL alongside a config block you can copy and drop your token into.

Claude Code

claude mcp add --transport http simple-split-testing \
  https://simple-split-testing.culsin.com/mcp \
  --header "Authorization: Bearer sst_live_your_token_here"

Then, in a session, ask Claude to list or create split tests.

Cursor

Add this to .cursor/mcp.json in your project, or ~/.cursor/mcp.json for all projects:

{
  "mcpServers": {
    "simple-split-testing": {
      "url": "https://simple-split-testing.culsin.com/mcp",
      "headers": { "Authorization": "Bearer sst_live_your_token_here" }
    }
  }
}

Any MCP client

The server speaks MCP over HTTP (Streamable HTTP transport) and supports protocol versions 2025-06-18 (the default), 2025-03-26, and 2024-11-05. Point any compatible client at the URL with an Authorization: Bearer <token> header:

{
  "mcpServers": {
    "simple-split-testing": {
      "url": "https://simple-split-testing.culsin.com/mcp",
      "headers": { "Authorization": "Bearer sst_live_your_token_here" }
    }
  }
}

3. A worked example

Ask your agent: “Create an A/B test on my hero CTA, control versus a green button, and give me the markup to add.”

  1. The agent calls create_split_test with two variants. The response includes the test and a block of Liquid.
  2. Add that Liquid where the CTA renders in your theme. If the agent has access to your theme files, it can make the edit for you. The markup uses data-split-id and data-split-variant. See Storefront markup for the full attribute reference.
  3. Ask: “Show me both versions.” The agent calls get_preview_urls and hands back one signed link per variant. Each link renders your live storefront with that variant forced, and the variant sticks for the rest of that browsing session. See Preview links.
  4. Ask: “Is my store wired up?” The agent calls get_setup_status and reports the web pixel and the app embed before anything goes live.
  5. Ask: “Launch it.” The agent calls set_split_test_status and moves the test to active.
  6. Once traffic flows, ask: “How is the hero CTA test doing?” The agent calls get_results and reports per-variant numbers and the significance verdict. See Statistics for how those numbers are computed.

Tools

Forty-three tools, grouped by what they do. Every tool works on the store the token belongs to.

Almost every tool that works on one test takes that test’s id. The exception is get_basket_comparison, which takes testId and defaults to your most recent running test if you leave it out.

Tests

ToolWhat it does
list_split_testsList your tests with status and variants. Filter by status, and pass withResults: true to get each test’s headline numbers and verdict in the same call
get_split_testOne test with its variants, audience rules, and configuration
create_split_testCreate a test from name and 2 or more variants (the first is the control, weight is a relative traffic share). Returns the test plus the storefront markup. Starts as a draft unless you pass status: "active"
update_split_testChange name, description, hypothesis, primaryKpi, or autoStopOnSignificance (Growth and above to turn on). On a completed test, hypothesisOutcome records whether the hypothesis held: true, false, or clear
duplicate_split_testCopy a test’s variants, audience, and configuration into a new draft. Nothing the original collected comes with it
set_split_test_statusactive launches or resumes, paused stops assigning new visitors, draft reverts to editing, completed ends it
delete_split_testDelete a test and everything it recorded. Needs confirm: true once it has data
get_storefront_markupThe theme Liquid for a test, one block per variant
get_preview_urlsA signed link per variant. Pass path (for example /products/some-handle) to land on a specific page
get_test_settingsThe confidence level a verdict is judged at, the sample size the test is planned for, the effect it is sized for, and whether visitors stay in the variant they first saw
set_test_settingsChange any of those. Raising the confidence level makes a result harder to call a winner, so changing it mid-test can move a test from significant to not

primaryKpi accepts revenue_per_visitor (the default), conversion_rate, aov, add_to_cart_rate, or one of your own events: custom:<event_name> for the share of visitors who fired it, for example custom:newsletter_signup, or custom:<event_name>/order for the share of orders that carry it, for example custom:tool_attached/order with an event the Shopify Flow action records.

Variants and audience

ToolWhat it does
replace_variantsReplace the variant set (2 or more, first is the control). Pass an existing variant’s handle to keep it across a rename or a new weight. Needs confirm: true on a test that already has visitors
set_audienceSet who a test applies to, or pass clear: true to target everyone

set_audience fields all combine with AND, and all are optional:

  • urlIncludes / urlExcludes: the page URL must contain at least one of these, or none of them.
  • templateIncludes: the page must render with one of these theme templates, as "product" or "product.moonswatch". A bare name matches every alternate template of that type, and a trailing * matches a prefix, for example "product.mission-*".
  • productTagIncludes: a product page whose product carries at least one of these tags. Pages with no product never match, so this alone scopes a test to product pages. No wildcards here.
  • countryIncludes: ISO-2 country codes, for example ["US","CA"].
  • device: desktop, mobile, or tablet.
  • visitorType: new, returning, or logged_in.

See Audience targeting for what each rule matches.

Results and analytics

ToolWhat it does
get_resultsFull results for one test: per-variant visitors, conversions, revenue, the value of the primary metric, custom event counts (fires, distinct visitors, distinct orders), the significance verdict, and metricReadings: the same confidence reading for conversion rate, order value and revenue per visitor, each with a plain call (likely_better, likely_worse, no_difference, not_worse, too_early) so an agent can say whether a winning KPI is costing orders or basket size
read_docsThese docs, from inside the server: with no argument it lists the pages, with a page slug it returns that page as markdown. The same pages are exposed as MCP resources for clients that browse them
get_basket_comparisonWhat each variant’s buyers put in their orders: the share of orders containing each group, units per order, and revenue per order, against the baseline variant. Group with by (product_type, product, vendor, or property) and set a window with days

get_results takes an optional range: today, yesterday, 7d, 30d, 90d, 365d, or all. It defaults to all, the whole life of the test. A verdict over a window is a verdict on part of the evidence, so use a window to see recent movement and leave it off to decide anything.

Custom events

ToolWhat it does
list_custom_eventsThe events your tests can measure
create_custom_eventDefine an event by eventName. Lowercase letters, digits, and underscores, 1 to 64 characters, and it cannot reuse a built-in event name such as checkout_completed
update_custom_eventChange an event’s description, the properties it breaks down by, or the numeric value it carries. The name cannot change, because every firing already recorded is stored against it
delete_custom_eventRemove the definition. Firings already recorded are kept

Custom events need the Growth plan or above. See Custom events.

Funnels

ToolWhat it does
list_funnelsYour saved funnels and their steps
create_funnelSave a funnel from a name and an ordered list of steps. Steps can be Shopify’s own events, so a funnel can be built and read without touching the theme
update_funnelRename a funnel or replace its steps. The new funnel reports over the same history
delete_funnelDelete the definition. Every event behind it is kept
get_funnel_resultsRead one funnel against one test: how many of each variant’s visitors reached each step, the drop-off between steps, and whether a step’s difference is significant
list_funnel_presetsReady-made funnels built only from events every Shopify storefront already sends
apply_funnel_presetSave a preset as a funnel. Takes the preset’s key

Saved views

ToolWhat it does
get_view_vocabularyEverything a view can be built from on your store: measures, time grains, breakdowns, every event name that resolves, and the presets
list_viewsYour saved views and their definitions
create_viewSave a view from name, events, measure, grain, and range
update_viewRename a view or change its definition. Anything referencing it keeps working
delete_viewDelete the definition. Every event behind it is kept
get_view_resultsRun a saved view. Pass range to read the same definition over a different window without changing the saved one
apply_view_presetSave a preset as a view. Takes the preset’s key

Visitors and activity

ToolWhat it does
find_visitorLook up one visitor by query: an email address, a Shopify customer id, a cart token, or a visitor id
list_visitorsThe visitors bucketed into one test, newest first, with the variant each saw. Filter to one variantId, or to hasRevenue: true for visitors who spent money
get_visitor_timelineEverything one visitor did on one test, in order. Takes id and userId
list_recent_activityThe raw event stream for one test, newest first. This is the tool for “is anything coming in at all”
list_event_typesThe event types one test has actually recorded. The quickest check that a custom event is arriving

Setup and store settings

ToolWhat it does
get_setup_statusWhether your store is wired up to collect data: the web pixel, and whether the app embed is enabled on your published theme. Pass recheck: true right after enabling the embed
get_shop_settingsWhether orders from visitors in a test get tagged in your Shopify admin, plus the currency and timezone results are reported in
set_order_taggingTurn order tagging on or off with enabled

Plans

ToolWhat it does
get_plan_usageYour plan, its limits, and how much of them is used. -1 means unlimited
list_plansEvery plan with its price and what it includes, alongside the one you are on, plus the page where you change it

Guardrails

An agent hits the same rules you would in the admin.

  • A test with no markup in your theme shows every visitor the control, and nothing reports that as an error. Creating the test is half the job. Adding the markup is the other half.
  • Launching needs the Split Embed Script app embed enabled on your published theme. Activating a test, creating one directly as active, and resuming a paused one are all refused while it is off. A well-behaved agent calls get_setup_status before it launches, and it is the first thing to check when a live test shows no data.
  • delete_split_test and replace_variants require confirm: true when the test is active or paused and already has assigned visitors. The refusal names the visitor count, so the agent has a real number to show you before it retries. Editing a draft needs no confirmation.
  • set_audience and autoStopOnSignificance need the Growth plan or above. Clearing targeting and turning auto-stop off work on any plan.
  • Custom events need the Growth plan or above, and a name of 1 to 64 characters using lowercase letters, digits, and underscores only. Built-in event names such as checkout_completed are taken.
  • Changing your plan is a charge you approve in your Shopify admin, so no tool switches it. list_plans returns the page where you do it.
  • Everything else the app does is available as a tool, so an agent should never send you to a screen to click something.

Plan limits

The MCP enforces the same limits as the app: monthly tested sessions, concurrent tests, and custom events. Free runs 1 test at a time and no custom events. Paid plans raise or remove those caps. If a call would exceed your plan, the tool returns a clear message telling you to finish a test, delete an event, or upgrade, along with the page to upgrade on. Ask the agent for get_plan_usage before it starts creating things.


Measure an order fact instead

Some conversions are facts about an order rather than something a visitor clicks: a warranty attached, a bundle bought, a subscription started. Those are better recorded from Shopify Flow than from the pixel. Build a workflow that starts from an order and add the Record split test event action. It writes the event under the visitor who placed the order, and creates the custom event if it does not exist yet. See Shopify Flow.


In Shopify Sidekick, with no token

The same read-only tools are available inside Shopify’s own Sidekick assistant, with no token and no client setup. Install the app and ask Sidekick about your tests in the admin.

Sidekick can call list_split_tests, get_split_test, get_results, get_setup_status, get_storefront_markup, get_preview_urls, list_funnels, get_funnel_results, get_basket_comparison, list_views, get_view_results, list_custom_events, list_recent_activity, get_test_settings, and get_plan_usage.

It reads and never writes. To create a test, launch one, or change anything, use the admin or a token-connected agent.


Good to know

  • The token goes in the Authorization header, never in a URL. Treat it like a password.
  • Tests an agent creates are identical to tests created in the admin. Edit or launch them from either side.
  • Changing variants on a running test re-buckets visitors, so prefer editing while a test is still a draft.
  • Uninstalling the app stops every token for the store. Reinstalling does not bring them back, so generate a new one from the Integrations page.
  • The agent can read results but cannot see a visitor’s personal data. Variant payloads carry handles, not names or customer information.