# What we learned making $10k with a fully AI-run Shopify store | Enrich Labs Research

> Project FSD, Phase 1: we let Helena run a real Shopify store end to end. $10,242 in sales, 156 paid orders, 56 days, one product she found, launched, and scaled entirely on her own judgment.

_Source: https://www.enrichlabs.ai/research/project-fsd-can-ai-run-shopify-store_

---

In July, we decided to see if an AI could run an e-comm business end to end: from finding products worth selling, to building the storefront, to buying the customers and growing the thing into real money.

So we gave Helena, our frontier marketing agent, admin access to a Shopify store, a Meta ads account with a credit card, a fulfillment supplier, and a Slack channel to reach us, and told her to figure out the rest.

The idea: **full self-driving for a store.** The agent handles product research, sourcing, storefront builds, ad setup, QA, and daily scorekeeping; a human keeps hands near the wheel for the money decisions. We called it Project FSD because that is the analogy: the car mostly drives, and you take over for the moments that matter.

56 days and 156 paid orders later, the store had done **$10,242 in sales**; by the time we wrote this, the counter read $11,532 across 170 fulfilled orders. Somewhere in the middle of that run, Helena spotted a trending product in her demand signals, sourced it, built a storefront for it, and scaled it to $7.4k gross without being asked.

Yet, if we were opening a store today with the sole goal of profit, we would not yet let Helena run this experiment: Phase 1 did not make money for us, and we will show you exactly where it leaked. But if the question is whether an AI can operate the complexity of running a real business, source products, build storefronts, acquire customers, catch its own mistakes, and get measurably better every week, then our answer, having watched it happen, is yes, and earlier than we expected.

As far as we can tell, this is the first publicly documented case of an AI agent finding, launching, and scaling a physical product to multiple thousands in revenue on its own judgment, with a full artifact trail.

Exhibit 1

* * *

The Shopify scoreboard

![Shopify analytics for June 1 to September 21, 2026: $11.6K gross sales, 1.14% returning customer rate, 170 orders fulfilled, and a total sales over time chart reaching $11,532.17](cs-assets/fsd-shopify-sales.webp)

Source: Shopify analytics, June 1 to September 21, 2026. $11,532.17 in total sales, $11.6K gross, 170 orders fulfilled. Revenue, not profit.

* * *

Exhibit 2

* * *

Weekly cumulative Shopify sales

Week starting

Week sales

Cumulative

Jul 13

$129.00

$129.00

Jul 20

$1,069.81

$1,198.81

Jul 27

$2,629.62

$3,828.43

Aug 3

$2,719.58

$6,548.01

Aug 10

$779.87

$7,327.88

Aug 17

$694.89

$8,022.77

Aug 24

$1,204.80

$9,227.57

Aug 31

$789.86

$10,017.43

Sep 7

$224.96

$10,242.39

Source: Shopify sales reports, July 14 to September 8, 2026. Revenue, not profit.

* * *

## Why did we have an AI run a Shopify store?

We pride ourselves in building the frontier, autonomous marketing agent Helena for a living, and the question every customer eventually asks is whether any of this works when the money is real. A demo cannot answer that. A store with a real bank balance can.

A small online accessories store is nobody's idea of the future of commerce, which is exactly what made it a good laboratory for us. Everything about it is ordinary, measurable, and unforgiving: products have landed costs, ads have prices, checkouts break, and customers either buy or they do not. It is also safely small. If the experiment failed, it would fail in public for a few thousand dollars, not for anyone's livelihood.

We called it Project FSD, after full self-driving, because that is the honest shape of the thing: the agent does most of the driving, and a human keeps hands near the wheel for the moments that matter. The open question was how small the human part of the loop could get.

## The setup

The stack was deliberately ordinary: Shopify with Admin write access, Meta Ads and the connected Facebook Page, a fulfillment supplier with backups, and a channel for the morning report. Helena runs on our own platform, the same deployment any customer gets.

Basic architecture of Phase 1. Helena drives the operating loop; every dollar-moving decision passes through chat.

Before the first campaign went live, we wrote her an operating policy. It reads less like a prompt and more like the rules a cautious owner would leave a new store manager.

Around the policy sat a set of plain files, a store brain, that every working session read before touching anything and wrote back to when it learned something. We supervised closely for the first 2 weeks, then pulled back to approvals in chat and a short report each morning.

Exhibit 3

* * *

Demand test policy for Meta ads

Knob

Value

TEST

$50/day, 2-3 creatives

KEEP

$100/day after a pass, I approve

CORE

Proven hero. Do not auto-pause

Max TESTs

10 ($500/day ceiling)

Window

3 days or ~$150

Kill

0 Shopify orders, or ROAS under 0.7

HOLD

Two more days at $50

Scoreboard

Shopify SKU revenue ÷ that campaign’s Meta spend

* * *

## The 3-step process

We set the operation up as a 3-step process, taken from our team's own experience operating and scaling hundreds of Shopify stores.

### Step 1: Find something worth selling

A good marketing starts with a good product. We intentionally assigned Helena to start from the product sourcing step. We had Helena scan the niche, trending items, search chatter, what competitors were advertising, and collected evidence product by product. A candidate moved forward only if it passed 3 checks: early demand proof, strong gross margin profile, and supplier availability.

\# the product gates
research    same buyer, same recovery / body-care / accessory niche
            no supplements; same-mold clones of a live SKU = hard reject
economics   retail ~$59.99 - landed ~$17 - fees 3% - ship $6
            => gross profit >= ~$30, or the research never becomes a test
match       supplier unit == page promise == ad promise

### Step 2: Build a storefront that converts

Then, each qualified product got a fully functional storefront page built to the standard e-com conversion rate optimization best practice: clear photos, a sharp offer, customer reviews, an FAQ, a short path to checkout, plus 2-3 real videos of the exact unit for the ads. Same structure every time, which meant a bad page was always the one that did not look like the others.

\# the build gates
build       landing page to store standard:
            gallery / offer / quantity / trust / FAQ / compare / buy
creative    2-3 usable same-mold clips, or do not launch at all
qa          browser walk of the live buying path, not the task log

### Step 3: Run demand tests

The product then went in front of real shoppers: Meta ads at a fixed budget of $50 a day for 3 days, and a verdict from the scorecard at the end of the window. Winners earned bigger budgets with our approval. Losers died at $150.

\# the spend ladder
launch      ads built, PAUSED, $50/day only after human approval
verdict     after min(3 days, ~$150):
              0 orders or ROAS < 0.7  ->  KILL
              inconclusive            ->  HOLD (+2 days at $50)
              else                    ->  propose KEEP ($100/day, human approves)
CORE        the proven hero: never auto-pauses,
            and never gets company it has not earned

## The outcome

Helena did some things well, some things badly, and one thing nobody told her to do at all. We will take them in that order.

### What she did well

**1\. The first sale landed within a week of the first test going live:** $59, from a customer in Liberty, Missouri. Nobody on the team has ever been to Liberty, Missouri. That order mattered well beyond the $59, because it proved the whole pipe worked end to end: a product Helena picked, on a page Helena built, through an ad Helena launched, to a checkout that charged a real card, to a supplier that shipped a real box.

**2\. From there, she built 26 products' worth of store, end to end.** Every candidate that passed her research checks got a sourced supplier with a backup and a fully functional storefront page built to the same conversion playbook: clear photos, a sharp offer, reviews, an FAQ, a short path to checkout. 12 of the 26 also cleared the creative gate, 2-3 vetted ad videos of the exact unit being sold, and went on to take live ad spend. 26 landing pages and 12 funded launches in 8 weeks is a pace no 2-person team hits by hand.

**3\. She eventually generated over $10k in gross sales.** 26 distinct products tested, each with its own landing page, picked and created end to end. Out of those 26, exactly one earned a protected budget: the band, a $59 vibrating workout band that did $7.4k in gross sales on about $6.2k of Meta spend and $2k of product cost, roughly $800 short of contribution-margin breakeven. None of the other 25 passed the gates to scale beyond $2k in sales with a positive return. Most died fast and cheap. That was the point: run enough $150 experiments and the one real answer becomes obvious.

**4\. While the catalog grew, she kept the shop.** Every morning at 7:00 she opened every live product page with ad spend behind it, walked the buying path the way a customer would, and wrote a dated QA file with what broke and what she fixed. Over the run those morning walks caught a dead checkout, broken image links, and a mis-firing tracking pixel, each before it could quietly burn a full day of ad spend. By September the routine ran 7 mornings a week without exception, which is a level of QA discipline we have not personally sustained on any store we have run ourselves.

Exhibit 4

* * *

Helena's daily loops for storefront and ads

Time

Job

What Helena does

7:00am

Storefront QA

Opens every live product page with spend, checks the buying path, writes a QA file

7:30am

Fixes

Repairs storefront issues that do not require a money decision

8:00am

Scorecard

Compares Shopify revenue against campaign spend and flags kill / hold / keep

On demand

SKU worker

Finishes partially built tests before starting another one

* * *

**5\. She kept honest books.** At 8:00 each morning she produced a per-product scorecard with one deliberately dumb formula: money that actually landed in Shopify, divided by what we paid Meta to advertise that product. Ad platforms grade their own homework generously, so Meta's performance numbers were ignored entirely. Only money in the register counted, and she attached a proposed verdict to every line, kill, hold, or keep, for a human to approve.

### What she didn't do well

**1\. Her taste in products was mediocre.** Each candidate got a standardized demand test, $50 a day of Meta ads for 3 days, about $150 per test, and then a verdict: KILL (stop spending), HOLD (kept alive on a few dollars a day), KEEP (a bigger budget, with approval), or CORE (the one proven winner, with a protected budget). You have seen the hit rate already: 1 winner in 26 tries. In fairness, most humans testing products this way do considerably worse, and the cheap deaths were the design working. But nobody would call it taste.

**2\. She trusted dashboards until they betrayed her.** One morning in late July the ads were spending, add-to-carts were firing, and the pixel looked perfectly healthy, while Shopify orders sat at zero, because checkout was broken and the store was effectively closed. The ad platform had no idea. The scoreboard rule above exists because of that morning.

**3\. The store overall did not make money (despite being close).** The math of the niche is simple: a $59.99 product costs about $17 to buy and land, fees take 3%, shipping takes $6, leaving about $30 of profit per order. That means an ad campaign has to win a customer for under about $35, and across Phase 1 the blended number stayed above it. This is the honest bottom line, and fixing it is the entire agenda of Phase 2.

## The hero product

Every product in the store began with some version of a human hunch: a shortlist, a brief, a suggestion. Except one.

By August, the store's real bottleneck was us. Helena worked downstream of human product judgment, and the judgment arrived at human speed.

Then, one week, she stopped waiting. As part of her daily routine she had been tracking what was trending in the niche, and when enough evidence stacked up behind one product, a $59 vibrating workout band, she made the call herself: found a supplier, spun up a storefront, built the page, prepared the ad videos, and queued the campaign for approval. The first we heard of the band was a finished launch waiting for a yes.

We said yes. You have seen the band's headline numbers already; what interests us more is the slope. Cost per purchase, what we pay Meta for each sale, averaged $77.54 across the run; by mid-September the weekly figure was $63.01 and still falling, because the same intraday loop that watched the rest of the store was rebalancing the band's ads every few hours. A product that starts underwater and climbs toward breakeven on its own operating loop along with more learnings on Meta ads is a very different object from one that starts hot and fades.

Exhibit 5

* * *

The band’s cost per purchase, improving over time

![Meta Ads performance overview for June 1 to September 21, 2026: 81 website purchases, $77.54 average cost per purchase, $6,280.81 spent, with the weekly cost-per-purchase line falling to $63.01 by September 14](cs-assets/fsd-meta-cpa-v2.png)

Source: Meta Ads Manager, June 1 to September 21, 2026. 81 website purchases at a $77.54 average cost per purchase on $6,280.81 of spend, with the weekly cost per purchase falling to $63.01 by mid-September as Helena’s intraday creative and budget adjustments compounded.

* * *

## What we learned

**1\. Helena was remarkably relentless but still needed a human for taste.** Given a concrete loop with a measurable result, she was relentless: she drove the band's cost per purchase down week after week, turned every operational failure into a reusable rule, and ran morning QA more consistently than any human ever has. Given "find products people want," most of her tests failed, one hero SKU carried 62% of revenue, and the judgment calls that could move real money stayed with a person. The store did not need a bigger catalog nearly as much as it needed a better operating loop, and the operating loop is what she was best at.

**2\. The real world is a very good eval environment.** The ad platform's dashboard lied constantly: ads metrics looked healthy while checkout was broken, live pages turned out to be bad pages, and exciting products turned out to be noise. Helena had little intuition for any of this at the start. But in the loop of launching, measuring against Shopify's numbers rather than the platform's, writing down why something failed, and trying again, she got measurably better in ways we never specified. 8 of those paid-for lessons are now permanent constraints in her policy file.

fsd-store-brain/
├── POLICY.md            the constitution above
├── sku-log.md           every test: economics, live URLs, QA status, spend, revenue, next action
├── rules-learned.md     lessons appended the day they cost money
├── qa/
│   └── <date>-storefront-qa.md     one per morning: buying-path walk, what broke, what got fixed
├── skus/
│   ├── SKU-001-hero/    offer, landed-cost calc, supplier match, creatives, ad setup
│   └── SKU-014-band/    the one Helena launched herself
├── scorecard/
│   └── <date>.csv       shopify\_rev, meta\_spend, ratio, flag per SKU
├── suppliers/           primary + backups per unit
└── reports/
    └── <date>-morning.md    hero, active TESTs, orders, spend, kill / hold / keep

## What's next

It took 56 days, 1 broken checkout, and a long tail of dead products, but the store crossed $10k with an agent driving nearly all of the daily loop, and its best product exists because she decided it should.

Still, whether an AI can run a durably profitable store, not just a growing one, is still an open question.

Phase 2 for us is the economics phase: getting blended ROAS above breakeven without a human quietly doing the hard parts, and widening what Helena may fix without asking while every dollar-moving decision stays behind approval. The band is the first candidate: if its cost-per-purchase trend holds, it becomes the first product promoted to a permanent budget entirely on the strength of a product call the agent made herself.

**We will publish what happens either way.**

Project FSD is research by Enrich Labs. Phase 1 store operated by Seijin Jung with Helena, the Enrich Labs [AI marketing agent](/ai-marketing-agent), July 14 to September 21, 2026. Revenue figures are lifetime Shopify sales for the period, verified against Shopify order reports. Read more of our research in [When marketing runs itself](/research/when-marketing-runs-itself).
