# Project FSD: Can an AI agent run a Shopify store? | Enrich Labs Research

> Phase 1 of Project FSD: we handed Helena the operating loop of a real Shopify store for 56 days. $10,242 in Shopify sales across 156 paid orders, what went wrong, and how far the autonomy actually got.

_Source: https://www.enrichlabs.ai/research/ai-agent-shopify-store_

---

## Why we ran Project FSD

We named the experiment after full self-driving, because that is the honest analogy. The car mostly drives; a human keeps hands near the wheel and takes over for the moments that matter. We wanted to know how close an AI agent could get to that on a real business, with a real Shopify store, real Meta spend, and a real P&L keeping score.

So we tried to run a Shopify store end to end with Helena making the key decisions: product selection, product sourcing, shipping and fulfillment setup, storefront design, storefront optimization, and customer acquisition through paid ads. On September 6, lifetime Shopify sales crossed $10,000. Two days later the counter sat at $10,242 across 156 paid orders, and by September 21 the store had done $11,532 in total sales across 170 fulfilled orders.

In total, Helena generated over $10k in sales while the store was still not fully profitable. We are excited by the progress, and by what it suggests about the future of retail and commerce, precisely because the result came from her operating loop rather than from a lucky product.

The clearest single piece of evidence is the band. Midway through the phase, Helena spotted a trending product from demand signals she had been collecting, spun up a separate storefront for it, and took it to roughly $7.4k in gross sales on decisions that were entirely hers. We tell that story in full below.

$10,242

Lifetime Shopify sales, July 14 to September 8, 2026

156

Paid orders

56

Days from first order to the number above

62%

Of revenue from one hero SKU

$7.4k

Gross sales from the band, a product Helena found and launched on her own judgment

The useful part is not the number by itself. It is how little of the daily operating loop I was still doing by hand.

Helena handled most of the work: researching what to test, building products and landing pages, preparing ad setups, checking the live storefront, catching broken flows, keeping the test log current, and putting the numbers in front of me every morning.

I kept the judgment calls that could move money. Helena drove the machine.

That distinction mattered because the phase was not clean. We launched things that should not have launched. A checkout broke while ads were still spending. Creative and landing pages got mismatched. I let a few tests get too complicated. Some mornings the pixel looked fine while Shopify told a very different story.

Those mistakes forced the operating system to get better.

By the time the store crossed $10k, Helena was not just completing tasks. She was running a repeatable loop with rules, checks, and enough memory of what had gone wrong before to avoid repeating the same failure twice.

If you have watched computer-use demos and wondered what they look like with a P&L attached, Project FSD is that idea applied to ecommerce.

## The setup

**The stack was deliberately ordinary:**

-   Shopify with Admin write access
-   Meta Ads and the connected Facebook Page
-   A fulfillment supplier, with backups available when needed
-   A place to send the morning report, like Slack or email

Then I gave Helena a written operating policy.

The file mattered more than I expected.

Before the rules were explicit, she could improvise. And an agent that improvises around budgets, status changes, or test criteria is exactly what you do not want.

**So the policy became the constitution:**

-   TEST = $50/day
-   KEEP = $100/day after approval
-   CORE = proven hero, never auto-pause
-   Max 10 active TESTs
-   Learning window = min(3 days, ~$150)
-   Kill if 0 Shopify orders or ROAS under 0.7
-   HOLD = +2 days at $50
-   Scoreboard = Shopify SKU revenue ÷ campaign spend
-   Never pause or resume Meta unless I say so in chat

Then I scheduled the jobs that made it feel like FSD:

Exhibit 1

* * *

The daily clock

Time

Job

What Helena does

7:00am

Storefront QA

Opens every live product page with spend, checks the buying path, writes a QA file

7:30am

Fixes

Repairs storefront issues that do not require a money decision

8:00am

Scorecard

Compares Shopify revenue against campaign spend and flags kill / hold / keep

On demand

SKU worker

Finishes partially built tests before starting another one

* * *

Ad Autopilot stayed off.

Helena flagged. I confirmed pauses, resumes, and budget changes in chat.

This division of labor is the whole thesis of Project FSD: the agent drives almost everything, and destructive money moves stay behind a human approval step.

## What happened

The first order landed July 14, 2026 at $129.

By the week of July 27, the store was doing roughly $2.6k in a week. Lifetime sales crossed $10,017 on September 6. By September 8, Shopify showed $10,242.39 across 156 paid orders.

Exhibit 2

* * *

Weekly cumulative Shopify sales

Week starting

Week sales

Cumulative

Jul 13

$129.00

$129.00

Jul 20

$1,069.81

$1,198.81

Jul 27

$2,629.62

$3,828.43

Aug 3

$2,719.58

$6,548.01

Aug 10

$779.87

$7,327.88

Aug 17

$694.89

$8,022.77

Aug 24

$1,204.80

$9,227.57

Aug 31

$789.86

$10,017.43

Sep 7

$224.96

$10,242.39

Source: Shopify sales reports, July 14 to September 8, 2026. Revenue, not profit.

* * *

Exhibit 3

* * *

The Shopify scoreboard

![Shopify analytics for June 1 to September 21, 2026: $11.6K gross sales, 1.14% returning customer rate, 170 orders fulfilled, and a total sales over time chart reaching $11,532.17](/research/cs-assets/fsd-shopify-sales.webp)

Source: Shopify analytics, June 1 to September 21, 2026. $11,532.17 in total sales, $11.6K gross, 170 orders fulfilled. Revenue, not profit.

* * *

The important part of the revenue mix was concentration.

One hero SKU produced about 62% of revenue. A second product, the band Helena found and launched herself, emerged as the only expansion that looked like it might deserve durable budget. A long tail of smaller tests printed a few orders each. Most tests did not become anything meaningful.

That was not a bug in the system. That was the system.

The point of the test catalog was not to make every product work. It was to let Helena push enough disciplined questions into the market that one or two answers could become obvious.

The store had one proven hero and a rotating set of TEST campaigns. The hero sat around $100/day; everything else started as a TEST under the policy above.

The breakthrough for me was realizing the store did not need a bigger catalog nearly as much as it needed a better operating loop.

## What Helena did well

Most ecommerce advice still reduces the job to something like: find a winner, build a page, run ads.

In practice, the bottleneck is everything around those verbs.

Research. Margin checks. Supplier matching. Product setup. Landing-page construction. Creative coverage. Ads Manager. Checkout bugs. Broken image URLs. QA. Daily scorekeeping. Remembering what already failed.

That is a job.

I did not want another job. I wanted a system that could launch tests, catch its own operational mistakes, and tell me which ones deserved more attention. Helena made that possible because she could own the repeatable work across the whole loop instead of only one narrow step.

**Three things changed the economics.**

First, every SKU stopped being a custom project. Helena reused the same pipeline: research, margin check, storefront build, ad setup, QA, scorecard.

Second, we stopped treating platform-reported ROAS as the source of truth. Shopify revenue for a SKU divided by that campaign’s Meta spend became the scoreboard.

Third, morning QA stopped depending on me remembering to click through the site. Helena opened the live pages, checked the buying path, compared them against the store standard, wrote down what was wrong, and fixed what she could before I even looked at the account.

**By the end of the phase, my input on a new product had shrunk to one line:**

_Find the next SKU and run the launch pipeline._

From there, Helena drove the workflow:

1.  **Research the next test.** She looked for physical products in the same broad recovery / body-care / accessory niche, aimed at the same buyer, with enough margin to support paid acquisition. The filter was simple: no supplements, roughly $30 gross profit after landed cost, fees, and shipping, and the supplier unit had to match what the page and ads promised. Early on, “same category” was too loose and led to mismatches. That mistake turned into a harder rule: the actual unit has to match.
2.  **Build the storefront.** Helena created the Shopify product and reused a standard landing-page structure: gallery, offer, quantity options, trust, FAQ, comparison, and a clean buy path. This was a breakthrough. Once every test had to clear the same quality bar, builds got faster and bad pages became easier to spot.
3.  **Prepare enough creative for a fair test.** We learned not to overbuild or improvise around weak creative. If the set was not good enough to support a clean experiment, the launch stayed incomplete. That gave us fewer noisy tests and better conclusions.
4.  **Set up the ads, but keep them paused.** Every new TEST started at $50/day with 2-3 creatives. No product got a bigger budget because it looked exciting. Some tests printed orders quickly; others burned through the learning window and gave us nothing. Helena’s job was to make sure failure meant something.
5.  **QA the live buying path.** Helena opened the live URL and checked the offer, images, cart, checkout, and the details a task log cannot see. Browser QA became mandatory because otherwise we could mistake our own operational errors for weak demand.
6.  **Log the result and escalate only the decision.** Helena kept cost assumptions, live URLs, QA status, spend, revenue, and next action current. I got pinged when a decision was actually mine: approve spend, pause, resume, or change budget. That is what made the setup feel less like “AI helping with ecommerce” and more like an operating system.

## The self-improvement loops

The part of Project FSD that changed our thinking was not any single task. It was that Helena ran learning and improvement loops multiple times per day, and the store got measurably better because of them.

**The storefront loop.** Several times a day, Helena browsed the live store the way a customer would, checked the UI and UX against ecommerce best practice, and deployed fixes directly: layout problems, broken sections, missing trust elements. She flagged broken links and pixel issues the moment they appeared, before they could quietly corrupt a day of spend or tracking.

**The ads loop.** Helena monitored the Meta account intraday and adjusted spend across multiple ad creatives based on what the last few hours actually showed, shifting budget toward the creatives that were converting and away from the ones that were not. Over time, that intraday loop drove a main product’s cost per purchase down; the next section shows the receipts.

**The demand loop.** Helena also watched demand signals across the niche, looking for trending products with a real revenue opportunity. When enough signals stacked up on one, she acted without waiting for a brief. The result of that loop is the best story in this report, and it gets its own section below.

Each loop fed the others. A storefront fix made the ads loop cleaner. An ads observation changed what the demand loop looked for. That is what we mean by self-improvement: not a model update, but an operator getting better at her own job several times a day.

## The band: a product Helena found, launched, and scaled herself

Every other product in the store began with some version of my judgment. The band began with Helena’s.

She had been collecting demand signals across the niche as part of the daily loop. When enough of them stacked up on one trending product with a high revenue opportunity, she made the call: she spun up a separate product storefront, built the page to the store standard, prepared the creatives, and began running ads to test it. No brief. No shortlist from me. The first I saw of the band was a launch already meeting the quality bar.

It became the store’s biggest expansion.

$7.4k

Gross sales from the band

$6.2k

Meta ad spend behind it

~$2k

Product cost

$63.01

Weekly cost per purchase by mid-September, down from a $78.05 average

Exhibit 4

* * *

The band’s cost per purchase, improving over time

![Meta Ads performance overview for June 1 to September 21, 2026: 80 website purchases, $78.05 average cost per purchase, $6,243.95 spent, with the weekly cost-per-purchase line falling to $63.01 by September 14](/research/cs-assets/fsd-meta-cpa.webp)

Source: Meta Ads Manager, June 1 to September 21, 2026. 80 website purchases at a $78.05 average cost per purchase on $6,243.95 of spend, with the weekly cost per purchase falling to $63.01 by mid-September as Helena’s intraday creative and budget adjustments compounded.

* * *

The honest math: roughly $7.4k gross on about $6.2k of ad spend and about $2k of product cost puts the band about $800 short of breakeven. We are not hiding that.

But look at the direction. The average cost per purchase across the run was $78.05; by mid-September the weekly figure was $63.01 and still falling, because the same intraday loop that managed the hero was rebalancing the band’s creatives every day. A product that starts underwater and climbs toward breakeven on its own operating loop is a very different object from one that starts hot and fades.

That is why the band matters more than its P&L. It is the first product in the store where every consequential decision, spotting the demand, choosing the product, sourcing the unit, designing the storefront, and acquiring the customers, was made by the agent. My contribution was approving the spend.

## Where things went wrong

The final workflow looks clean when written as a checklist. It did not feel clean while we were building it.

At first I broke my own policy. One launch went out at $100/day with five ads. The signal got noisy, learning became harder to read, and we spent more before we understood anything.

That mistake became one of the cleanest rules in the system: tests stay small until they earn the right to become bigger.

SJ

Seijin Jung

Co-founder, Enrich Labs

Another important lesson came when checkout failed.

Ads were still on. Add-to-carts were still firing. The ad platform looked alive. Shopify orders were flat because the store was effectively closed.

That morning changed how Helena scored the business. From then on, Shopify became the scoreboard and browser QA became a gate, not a nice-to-have. The agent stopped behaving like a media buyer staring at a dashboard and started behaving more like an operator responsible for the whole path from click to order.

**A few more things we paid to learn:**

-   Ads can keep spending while checkout is broken
-   A technically live page can still be a bad page
-   Creative mismatch can make a good offer look like a demand failure
-   Too much budget too early creates worse information, not faster truth
-   Pixel ROAS can look healthy while Shopify revenue is flat
-   A test that is not operationally clean is not a valid test
-   Partial launches create more confusion than dead ones
-   The agent needs explicit rules for what it can fix itself and what requires approval

The biggest success of the phase was not a single winning product. It was turning those mistakes into operating constraints Helena could reuse. The second time a similar problem appeared, she was more likely to catch it before money moved.

That is where the system started to compound. Every failure turned into a rule.

## Did the store make money?

This is where people want a clean $0 → $10k/month story.

That is not what happened.

Lifetime revenue hit $10k in about eight weeks of orders, from July 14 to September 6. Profit is a different spreadsheet.

Early-September blended Meta ROAS was still under 1×, which is exactly why I did not respond by pushing more spend into the hero just to force the top-line chart upward. We are not presenting Phase 1 as a profit machine. We are describing how Helena kept launching, fixing, measuring, and pruning until the Shopify counter crossed ten thousand.

**The useful number is the cost of asking one product question.**

A typical product in this niche could look roughly like:

-   Retail: about $59.99
-   Landed product cost: about $17
-   Fees: about 3%
-   Outbound shipping: about $6
-   Gross profit per order: about $30

That puts breakeven CPA in the $30s.

A TEST costs about:

**$50/day × 3 days = $150**

Ten simultaneous tests creates a $500/day experimental ceiling. The hero budget sits separately.

The band tells the story of where the economics stand: about $800 short of breakeven on $7.4k gross, with the cost per purchase still falling when the phase ended. The full account is in the band section above.

What changed with Helena was not the media math. It was the cost and speed of operating the experiments around it. The same agent could research, build, QA, score, and maintain the next test without adding another person for every new SKU.

That is the leverage.

## A week inside the loop

The store looks boring when the system is working.

In the morning, Helena has already browsed the live pages, checked the main buying paths, repaired routine storefront problems, and updated the scorecard.

I open a short report with the hero, active TESTs, Shopify orders, spend, and a kill / hold / keep flag.

If something should pause, I say so in chat.

If a page broke, Helena fixes it.

If a test failed, it goes into the log and the next one can reuse everything we learned.

If something starts printing orders, it still has to survive the same economics as everything else.

That is the success case: not a magical dashboard, but less operational drag and fewer repeated mistakes.

## Why does this matter?

For years, ecommerce ops scaled with people.

More SKUs meant more VAs, more media-buying work, more page building, more “can you check if checkout works?”

What changed in Project FSD was not that AI suddenly found a magic product.

It was that Helena could carry most of the operational weight across the entire test loop.

She researched. She built. She checked. She repaired. She logged. She scored. She surfaced the decision.

I still needed taste. I still needed a proven hero. I still needed to approve spend.

But I no longer needed to personally live inside every browser tab required to test the next physical accessory in the niche.

That is what made the $10,242 interesting. And the band made it concrete: for the first time, a product existed in this store because an agent decided it should.

One hero still did most of the revenue. Most tests still failed. ROAS was not magically solved.

But the store kept getting better at asking the next question, and Helena drove most of that improvement.

That is FSD ecommerce as we have actually experienced it: a human keeps the judgment and the money decisions; the agent owns as much of the repeatable operating loop as possible.

Most people still think this takes a team.

That is why we are excited about what this phase suggests for the future of retail and commerce. If one agent can select the products, source them, design and repair the storefront, and acquire the customers, then the cost of starting and operating a store stops being a staffing problem and starts being a judgment problem. The humans who win in that world are the ones with taste and capital discipline, not the ones with the most tabs open.

## If you want to replicate it

This is a plan for building a store that tests without requiring you to live in Ads Manager. The $10k in this article is lifetime Shopify sales, not month-13 profit.

**Phase 1 (Days 1-30): one hero, one pipeline, first test**

-   Connect Shopify, Meta, and your supplier.
-   Write the policy table. Save it somewhere the agent reads before touching spend-related work.
-   Build or lock your hero landing page. That becomes the quality standard.
-   Run one expansion SKU slowly with the agent. Watch the full loop: research, economics, storefront build, ad setup, QA, launch, scorecard.
-   Do not optimize for speed yet.

_Goal: 1 CORE, 1 TEST, and a full understanding of where the workflow breaks._

**Phase 2 (Days 31-60): turn the manual loop into a clock**

-   Schedule morning QA, fixes, and a scorecard.
-   Keep automated spend decisions off.
-   Add 2-3 TESTs using the same operating pattern.
-   Create a rule that partial work gets finished before net-new research starts.

_Goal: 3-4 live tests at $50, a scorecard you actually read, and clean kill decisions._

**Phase 3 (Days 61-90): cap and repeat**

-   Move toward the 10-TEST ceiling only if the operating system can support it cleanly.
-   Do not raise the cap because you are bored.
-   Automate reporting so Shopify SKU revenue and campaign spend land in one place.
-   Protect the hero. Promote a KEEP only when the economics justify it and you explicitly approve the move.

_Goal: a catalog that can launch and get maintained without you touching every step._

From there, it becomes repetition. Each new SKU is another $150 question. Helena drives the process. The market answers.

Exhibit 5

* * *

The policy file

Knob

Value

TEST

$50/day, 2-3 creatives

KEEP

$100/day after a pass, I approve

CORE

Proven hero. Do not auto-pause

Max TESTs

10 ($500/day ceiling)

Window

3 days or ~$150

Kill

0 Shopify orders, or ROAS under 0.7

HOLD

Two more days at $50

Scoreboard

Shopify SKU revenue ÷ that campaign’s Meta spend

* * *

**On day one, done on a SKU means:**

-   Economics clear
-   Supplier unit matched to the promise
-   Landing page meets the store standard
-   2-3 usable creatives
-   Ads built and paused
-   Browser QA passed
-   $50/day live only after approval
-   Shopify used as the scoreboard

The first few launches will expose what your policy forgot. That is normal. Do not hide those failures. Turn them into rules.

## What comes next

Phase 1 answered the operational question: an agent can own most of a real store’s daily loop, and the mistakes can compound into rules instead of repeating.

Phase 2 is the harder question: economics. Getting blended ROAS above breakeven without a human quietly doing the hard parts, and widening what Helena may fix without asking while keeping every dollar-moving decision behind approval. The band is the first candidate: if its cost-per-purchase trend holds, it becomes the first product promoted to KEEP entirely on the strength of a product call Helena made herself.

The more interesting question now is how small the human part of the loop can get without giving up control.

We will publish what happens either way.

## FAQ

#### Where does computer use fit vs chat?

Computer use is for browser QA and the parts of store operations that require seeing the live interface. Chat is where I approve decisions that can move money or change campaign status.

#### What catalog fits this approach?

The system works best when the products share a buyer and a margin profile. In this case, that meant a broad recovery / body-care / physical-accessories niche rather than a random general store.

#### What if a test fails?

Pause the campaign, leave the store artifact intact, append the result to the SKU log, and move on. A failed test should improve the next launch even when it produces no revenue.

#### What was the biggest breakthrough?

Making Helena responsible for the full operating loop instead of isolated tasks. Once QA, fixes, scorekeeping, and launch hygiene lived in one system, the mistakes started turning into reusable rules instead of recurring surprises.

Project FSD is research by Enrich Labs. Phase 1 store operated by Seijin Jung with Helena, the Enrich Labs [AI marketing agent](/ai-marketing-agent), July 14 to September 8, 2026. Revenue figures are lifetime Shopify sales for the period, verified against Shopify order reports. Read more of our research in [When marketing runs itself](/research/when-marketing-runs-itself).
