All articlesTutorials

Shipping an AI Agent to Shopify in an afternoon

Elena Marsh10 min read
Tutorials

Catalog sync, order lookups, and a working storefront widget — all before the end of your workday. That's the actual timeline we've seen hold up across a genuinely wide range of Shopify stores, and walking through why each piece takes as little time as it does is more useful than just asserting the headline claim.

Catalog sync

Catalog sync is the first fifteen minutes: connect the Shopify Admin API, and the Agent indexes every product, variant, and price so it can answer "do you have this in blue?" without anyone writing FAQ entries for each SKU. The fifteen-minute figure holds regardless of catalog size, because the indexing itself runs as a background job rather than something that has to complete before you can move to the next step — a store with fifty products and a store with fifty thousand both start the sync in about the same amount of hands-on time, they just finish the background indexing at different points afterward.

What this actually replaces is the alternative most stores were doing before: manually writing and maintaining FAQ answers for product-specific questions, which scales badly as a catalog grows and falls out of date the moment a variant gets discontinued or a price changes. Catalog sync means the Agent's answer to "is this still available in size medium" is always checked against live inventory, not against a static FAQ entry someone forgot to update.

Order lookups

Order lookups come next — the Agent authenticates a customer by email and order number, then can answer "where's my order" or "can I change the shipping address" using the same order API a support agent would use. The authentication step matters as much as the lookup itself; an Agent that can pull up order details without verifying the requester actually owns that order is a real information-disclosure risk, and the setup here defaults to requiring both email and order number together, rather than either alone, specifically to close that gap.

Address changes are a good example of an action that looks simple but has a real eligibility question underneath — a shipping address can only be changed before an order ships, and the Agent checks fulfillment status before making the change rather than assuming the request is always valid, which is the same policy-check discipline we've described elsewhere in this series applied to a different action type.

The storefront widget

The storefront widget is a copy-paste script tag. It shows up as a chat bubble on the storefront, pre-loaded with catalog and order context, so a customer never has to explain what they already told the widget — if they've already mentioned an order number earlier in the conversation, the widget doesn't ask again on a follow-up question, because the context persists across the exchange rather than resetting with every new message.

The pre-loading detail matters more than it sounds like it should. A widget that has to re-establish context on every question reads as forgetful and mechanical; one that carries context naturally reads as attentive, even though the underlying difference is purely a matter of what state gets passed forward between messages rather than anything about the model's actual intelligence.

What actually takes the afternoon

The parts that actually take an afternoon aren't technical — they're deciding which order actions the Agent can take unsupervised versus which ones need a human, like refunds above a certain amount. This is the same authority-scoping decision we've described in more depth elsewhere in this series, and for a Shopify store specifically, it usually comes down to store-specific judgment calls: is a $15 exchange different from a $150 one, does a repeat customer get different handling than a first-time buyer, does a specific product category have return-fraud history that justifies extra scrutiny.

Every store we've onboarded has slightly different answers to those questions, and getting them right — not the technical setup — is what actually fills the rest of the afternoon after the first hour of catalog sync, order lookup, and widget installation.

Common questions

Does this work with a heavily customized Shopify theme, or only default themes? The widget script is theme-agnostic by design — it's a floating element that overlays on top of whatever theme exists, rather than something that needs to integrate with theme-specific templates, so customization level hasn't been a meaningful blocker in stores we've onboarded.

What happens if a store has multiple sales channels beyond the storefront itself? Catalog and order data sync from the same Shopify Admin API regardless of which channel a sale came through, so the Agent's answers stay consistent whether a customer bought through the storefront, a marketplace integration, or a point-of-sale terminal, as long as it's recorded in the same underlying Shopify order system.

A specific store's afternoon, hour by hour

One apparel store we onboarded is a useful concrete example of how the afternoon actually breaks down, because their catalog had a wrinkle that's common enough to be worth describing: heavy use of product variants (size and color combinations) with inconsistent naming conventions carried over from an earlier platform migration, which made the catalog-sync step slightly less than the usual fifteen minutes.

The sync itself still completed quickly, but the inconsistent variant naming meant a handful of "do you have this in X" questions initially returned technically-correct-but-confusingly-worded answers, referencing variant names that made sense to the store's internal system but not to a customer reading them in a chat window. The fix wasn't a data migration — it was a lightweight display-name mapping layer that translates the internal variant naming into customer-friendly language before it reaches a response, a fix that took about twenty minutes once identified and has been reused, with store-specific mappings, for several other stores with similar legacy-naming situations since.

Order lookups and the address-change eligibility check worked essentially as described in the general walkthrough, with one store-specific addition: this particular store sells a mix of made-to-order and stocked items, and made-to-order items have a much shorter address-change window than stocked ones, since production starts almost immediately. Encoding that distinction into the eligibility check took a short conversation with the store's own operations lead to get the exact cutoff right, followed by a quick rule addition — exactly the kind of store-specific judgment call we described as the real time cost of a rollout, distinct from the purely technical setup.

The storefront widget installation itself was genuinely the fifteen-minute copy-paste task it's meant to be, with one small customization request: the store wanted the widget's accent color to match their brand palette rather than the default, which the widget supports through a simple configuration option rather than requiring a code change.

By the end of the afternoon, the store had a working setup handling catalog questions, order-status lookups, and address changes within the made-to-order-aware eligibility rules — with refunds above a specific dollar threshold still routed to a human, per the store's own comfort level at launch. That threshold, notably, has been raised twice since, each time after a few weeks of watching the trace on human-reviewed refunds below the previous threshold and confirming the Agent's proposed decisions consistently matched what a human would have approved anyway.

The store's owner described the biggest surprise, a few weeks in, as not being about what the Agent could do, but about how much of the initial afternoon's work turned out to be genuinely reusable — the variant-naming mapping layer, in particular, has since been offered as a standard option for other stores migrating from the same earlier platform, because the underlying naming inconsistency turned out to be common across stores that had gone through a similar migration history.

Common questions

Does this work the same way for a Shopify Plus store with custom checkout logic, or does that require more setup? Custom checkout logic itself is mostly orthogonal to what the Agent needs — it primarily reads order and catalog data through the standard Admin API regardless of checkout customization, so Plus-specific checkout logic rarely adds meaningful setup time unless it changes how orders are recorded in a way that affects the data the Agent reads.

How do returns get handled for a store using a third-party returns-management app rather than Shopify's native return flow? We integrate with the returns app's own API where one exists, treating it as an additional data source alongside the core Shopify Admin API — this adds some setup time beyond the baseline afternoon, proportional to how well-documented that specific third-party app's API is.

What happens if a store's catalog changes significantly — a large seasonal SKU turnover, for instance — does the Agent's understanding update automatically? Catalog sync runs on an ongoing basis, not just at initial setup, so seasonal SKU changes get picked up automatically without requiring a manual re-sync, though a very large one-time catalog overhaul is worth flagging in advance so sync timing can be checked against it specifically.

Can the storefront widget be customized beyond accent color — for instance, restricted to only certain pages of the store? Yes, the widget supports page-level visibility rules, which some stores use to show it only on product and order-status pages rather than on every page of the site, depending on where the store's team believes support questions are actually most likely to arise.

Is there a risk of the Agent giving pricing or availability information that's technically correct at the moment of the sync but stale by the time the customer reads it, given how fast inventory can change? This is a real, if usually small, risk with any cached or synced data, which is why availability-sensitive answers are checked against a live lookup rather than the synced cache alone for anything time-sensitive, like confirming stock immediately before completing an exchange.

How do returns and exchanges get handled differently from the refund flow described elsewhere in this series? The underlying eligibility-check and policy-engine pattern is the same; what differs for Shopify specifically is which fields get checked — fulfillment status, return window relative to delivery date pulled from Shopify's own tracking data — rather than a fundamentally different mechanism.

Is there ongoing maintenance required after the initial afternoon setup, or is it truly set-and-forget? Catalog and order sync run automatically without ongoing maintenance, but the authority-scoping decisions — what the Agent can do unsupervised — are worth revisiting periodically as a store's product mix or customer base changes, the same way we'd recommend revisiting policy rules in any deployment.

The broader pattern across every store we've onboarded, not just the specific one described above, is that the technical setup genuinely does take about an afternoon, almost without exception — but the confidence to expand what the Agent is allowed to do unsupervised builds over weeks, not hours, through exactly the kind of watch-the-trace, raise-the-threshold-gradually process this store went through with its refund limit. Treating those as two separate timelines, rather than expecting both to happen in the same single afternoon, is the expectation we'd set for any store starting this process today.

It's also worth mentioning what happens for a store that outgrows the default setup described above — larger stores with more complex fulfillment logic, multiple warehouses with different eligibility rules, or a subscription component layered on top of one-time purchases. None of those situations break the basic afternoon-setup pattern; they extend it, adding additional eligibility branches the way the made-to-order distinction extended the basic refund flow in the example above. The afternoon gets the core mechanism working end to end; store-specific complexity gets layered on afterward, incrementally, rather than needing to be fully solved before anything can go live.

We'd rather a store launch with the simpler version working well and add complexity deliberately over the following weeks than delay launch trying to anticipate every eventual edge case in advance — the same start-narrow principle that's shown up throughout this series applies just as directly here.