When marketing runs itself

July 21, 2026  |  Research

Our progress toward self‑improving marketing, and its implications.

Key findings

  1. Self-improving loops are already running in production, and they point toward autonomous marketing: Across more than 1,000 active customer accounts, Helena executes inside standing loops that schedule their own follow-up reviews, write their own learning files, and audit their own automations; the heaviest accounts carry 143 to 300 self-built background automations each. The early signals on outcome metrics all move in the same direction, and the footprint behind them is real: more than $4.5 million in ad spend, 92,740 purchases, and 81,634 leads flowed through the sampled management windows alone, each window covering only part of its account’s life.
  2. Ad spend scales into positive returns: Where Helena has budget authority, the craft-supplies brand raised spend 20% while weekly purchases rose 28% at a 7.0× ROAS; another account grew weekly Meta spend from CAD 388 to CAD 1.9K (+390%); and a dead Google Ads account went from $0 to $11K in spend at roughly $10 per conversion.
  3. Paid efficiency improves as spend grows: Google Ads cost per acquisition fell from $26 to $14 at one brand as weekly conversions rose from 40 to 95, and a standing conversion-tracking audit cut another’s from $308 to $91 (−70%).
  4. Organic traffic grows: At a local financial institution, weekly organic sessions rose 23% (5.7K to 7K) while local search clicks quadrupled (475 to 2K a week); at the pain-management practice, weekly organic sessions rose 85% and search clicks 112% (1.3K to 2.7K).
  5. Email becomes a revenue channel: Email-attributed revenue went from $0 to $7.9K at a skincare brand after Helena built its flow stack; weekly run rates went from $0 to $2,700 at a stationery brand ($41.9K in the window) and from $40 to $163 at a sporting-goods brand as email’s share of total revenue rose from 6.4 to 46.3 percent.
  6. The loop improves itself: A weekly task reviews Helena’s own automations and proposes one change; it has run 24 times on a single account. One self-scheduled SEO loop has run every two hours for nine weeks (543 runs), each run starting from what the last one learned: a bounded form of recursive self-improvement.

The state-by-state insurance blog posts are being triggered manually each session… Suggested change: add a daily automated run for the insurance blog publishing queue.

Helena
Helena
Proposing an improvement to her own automations*

Overview

For most of marketing’s history, humans drove every step of the cycle: plan the campaign, write the copy, ship it, read the report, adjust, repeat. But across Enrich Labs’ customer base, lean teams are delegating a growing share of that cycle to Helena, our AI marketer, which is speeding up their work.

Taken far enough, that trend points to a marketing operation that improves itself: a system that not only executes campaigns but audits its own automations, schedules its own follow-up reviews, quarantines its own bad data, and proposes its own next experiments. This is self-improving marketing. AI researchers call the general version of this pattern recursive self-improvement: a system capable of improving the system that improves it. Anthropic’s recent essay “When AI builds itself” documents that loop beginning to close inside AI development; this research documents a bounded version of it closing inside marketing. We are not fully there yet, and it is not inevitable. But it could arrive sooner than most marketing organizations are prepared for.

Using live production data from more than 1,000 active customer accounts, plus previously unreported reconstructions of every scheduled task, memory, and learning file at four of them, this research shows the loop is already partly closed. To take just one example: at a telehealth company, Helena has run a self-scheduled SEO improvement cycle every two hours for the past nine weeks (543 runs), each one reading the learnings file written by the previous runs before deciding what to do next.1

The trends discussed in this piece suggest marketing systems are going to become much more capable in the coming years. That has huge implications. Marketing that improves itself compounds: every audit makes the next campaign better, every post-mortem makes the next launch cheaper. But it also changes the job. When execution costs almost nothing, the humans who win are the ones who set direction, judge results, and decide which experiments are worth running at all.

Exhibit 1

Helena builds her own scaffolding: standing automations accumulate at every account, from 143 to 300 at the heaviest users.

Background automations created by Helena per account, cumulative,1 count, March–July 2026

300 0 11-brand media operator 47 300 Mar Jul Immigration- consulting firm 30 230 Mar Jul Digital-products brand 146 206 Mar Jul Baking-supplies retailer 18 182 Mar Jul 300 0 Solar-mounting manufacturer 12 179 Mar Jul Enterprise-software company 17 168 Mar Jul B2B SaaS company 10 160 Mar Jul Beauty brand 33 143 Mar Jul

1Counts every background schedule Helena created on the account: recurring loops (daily reports, weekly audits, autonomous-learning cycles) and one-shot self-scheduled graders (24-hour, 48-hour, 2-week, and 30-day follow-up checks on her own work). The eight accounts shown are the eight heaviest automation creators across the customer base; onboarding dates differ. The B2B SaaS company is the same account as the $308→$91 CPA example, the digital-products brand the same as the email strategy-file example, and the baking-supplies retailer the same as the internal-links example. For scale, the four deep-dive customers in this research created 131 (telehealth), 19, 18, and 17; the telehealth company ranks ninth overall.

Source: Enrich Labs internal data, July 2026

Enrich Labs


The delegation has happened in stages, each one moving more of the marketing cycle from the human to the system. The diagram below shows those stages as they appear in our customers’ accounts today: each row is one stage in that delegation, and the starburst at the right is the output: the campaigns, content, and reports the business actually ships.

Marketer
Tools

Stage 1: A marketer drives disconnected tools by hand, copying numbers between tabs.

Marketer
Tools
Chatbot

Stage 2: Chatbots draft captions and subject lines, but the marketer still pastes the output into the tools.

Marketer
Tools
Chatbot
Helena
Helena

Stage 3: Helena connects directly to the ad accounts, email platforms, and analytics, and executes campaigns and content herself.

Marketer
Helena
Helena
Schedules
Learning files
Self‑review

Stage 4: where the loop closes today. Helena schedules her own follow-up reviews, writes what she learned to persistent learning files, audits her own automations weekly, and starts each run from what the last one knew.

At each stage the human’s role narrows to setting direction and judging results, while the loop does the rest.

Evidence from across the customer base

The study base is more than 1,000 active customer accounts across 11 industries and more than 35 countries. For channel-level analysis, we sampled the 50 highest-spend accounts per channel from the hundreds with connected ad, search, and email platforms; more than $4.5 million in Meta and Google ad spend flows through those sampled management windows alone, along with 92,740 purchases, 81,634 leads, and roughly $1.6 million in Google ad-attributed revenue: live production accounts, not pilots. And because each figure covers only the account’s Helena window, not its lifetime, the totals understate what sits under management. After Helena went live, click-through rate improved on 26 of 50 sampled Meta accounts, and cost per acquisition fell on 12 of the accounts with measurable conversions. Weekly organic sessions rose on 21 of 50 accounts with a measurable baseline; Search Console clicks rose on 13 of 28. Twenty accounts run two or more standing ad-management automations (audience-fatigue alerts, weekly negative-keyword mining, conversion-tracking audits), and eleven email accounts run recurring flow-health audits. Roughly $4.5 million in Meta and Google spend sits under these management windows: live production accounts, not pilots.2 The lift concentrates on accounts that let Helena build rather than just watch.3

The most telling evidence, though, is not any single delta. It is the shape of the work: standing audience-fatigue alerts, weekly negative-keyword mining, conversion-tracking audits that hunt for spend with zero conversions, recurring flow-health checks: always-on guardrails humans forget, analyst work on a cron.

Exhibit 2

On every channel, Helena’s improvement loop moves the metric it targets.

One named improvement loop per channel, before vs. with Helena,1 weekly run rates

Google Ads, cost per acquisition, craft-supplies e-commerce brand With Helena $14 Before $26 −46% Meta Ads, cost per acquisition, wellness studio With Helena $174 Before $392 −56% SEO, Search Console clicks per week, pain-management practice With Helena 2.7K Before 1.3K 2.1× Email, email share of revenue, sporting-goods brand With Helena 46.3% Before 6.4% 7.2×

1“Before” is the 90-day window preceding Helena’s go-live on each account; “with Helena” is the management window, in which Helena built, then scheduled her own reviews of, the campaigns, content, and flows shown. Deltas reflect correlation, not sole attribution; seasonality, budgets, and human decisions also move these numbers.
Source: Enrich Labs internal data, July 2026

Enrich Labs


Aggregate deltas say a lot about diffusion. But they cannot reveal whether the system is improving itself. For that, we need direct evidence from within the accounts.

Evidence from within four customers

Running a marketing operation takes two broad categories of work. There is execution: writing the ads and the articles, building the flows, shipping the reports. And there is judgment: deciding what to test, interpreting what comes back, and figuring out which ideas to try next. To see how far Helena has come on both, we reconstructed every scheduled task, memory, and learning file at four customers (a telehealth company, a craft-supplies e-commerce brand, a pain-management practice, and a multi-location medical group), supplemented by channel deep-dives across the 50 highest-spend Meta, Google, and email accounts sampled from that base. The evidence forms a ladder: Helena runs experiments to goals humans set, steers campaigns with her own judgment, proposes her own experiments, checks her own work, and has begun improving the system that does all of the above.

Helena is getting better at steering campaigns toward outcomes.

Paid social is where judgment shows. The day-to-day work of Meta ads is a chain of judgment calls: is this creative fatigued or fine? Is this dip a trend or noise? Scale, hold, or kill? A multi-location medical group with 170+ providers, live only since June 6, shows what those calls look like when they are written down. Within 90 minutes of onboarding, Helena had recorded operational guardrails a contractor might take weeks to learn: which insurance plans never to reference in content, which nine cities to target. Its sister pediatrics account carries explicit anomaly thresholds: flag any campaign with CPA above 2× its 7-day average, flag zero conversions for 3+ days, flag day-over-day spend spikes above 30%, and do not scale Performance Max until conversion tracking is verified.

Each week, an autonomous-learning run reads the account’s Meta learning file (the standing instruction is to “build on prior insight rather than starting cold”), pulls fresh platform data, and rewrites the file with what changed. Six weeks in, Meta CTR is up 16% (1.39% to 1.61%) and ad frequency is down 23% (4.2 to 3.2) with engagement up: the signature of an account being actively managed rather than left to burn out its audience. Unit costs confirm it: pulled live from the Graph API, cost per landing-page view fell from $0.66 to $0.55 a week over the same window, even as budget guardrails cut spend 35 percent in mid-July. And the July 21 learning file reads like a competent analyst overruling their own earlier hypothesis:

The prior concern about creative fatigue appears premature… CTR and CPC are actually improving, which confirms the audience seeing the ads is still engaging. But the pool of targetable users is shrinking. If this pattern continues into Jul 20–26, the audience needs to be broadened or new creative needs to be added.

Helena
Helena
Weekly learning file, 170-provider medical group*

Where Helena writes the ads herself, she iterates against explicit baselines. At a wellness studio, new treatment creatives were drafted with the target written into the task: the original ad set runs CTR 2.0% at $24 per lead: beat it. Over the management window, purchases per week rose 144%, CPA fell from $392 to $174 (−56%), and frequency dropped 27%; one lead campaign delivered 272 leads at a $5 CPL. At the education marketplace, the loop runs at analyst depth: a “Catalog Hidden Gems” deep dive surfaced under-funded tier-2/3 catalog campaigns that now run at 28.9× ROAS on $99K of spend, alongside a creative-diagnostic rewrite of the ESA program and a daily month-to-date CAC report that lands before the team’s standup.

Exhibit 4

The brief said beat $24 a lead; she delivered $5.

Meta ads creative iteration against recorded baselines, two accounts,1 account data, Helena management windows

Wellness studio: iterating against the recorded baseline Cost per lead treatment campaign −79% 272 leads delivered $24 $5 The brief Delivered Cost per acquisition account −56% $392 $174 Before With Helena Ad frequency account −27% 4.3× 3.1× Before With Helena Education marketplace: analyst depth The “Catalog Hidden Gems” deep dive surfaced under-funded tier-2/3 campaigns 28.9× Meta-attributed ROAS on the rediscovered campaigns $99K spend now running through them The same loop also ships: • a creative-diagnostic rewrite of the flagship ad program • a daily month-to-date CAC report, delivered before the team’s standup

1The wellness studio’s baseline was written into Helena’s creative task verbatim: “the original ad set runs CTR 2.0% at $24 per lead.” Cost per lead reflects the resulting lead campaign (272 leads at $5); CPA and frequency compare the 90-day pre-Helena window with the management window; purchases per week rose 144% over the same span. The education marketplace’s 28.9× is Meta-attributed ROAS on the tier-2/3 catalog campaigns surfaced by the deep dive. Deltas reflect correlation, not sole attribution.
Source: Enrich Labs internal data, July 2026

Enrich Labs


And the loop replicates to newly connected channels. The craft-supplies brand wired up Meta on July 2; by July 17 Helena had already scheduled a standing weekly creative-performance check on it. The account runs at 7.0× Meta-attributed ROAS with purchases up 28% week-over-week against the pre-Helena baseline. And cohort-wide, standing audience-fatigue alerts (flag frequency above 3.5, or a 20% CTR drop) are among the most common automations customers run.3

And given budget authority, the loop scales spend, not just efficiency. At the craft-supplies brand, spend rose 20 percent while purchases rose 28 percent and CPA fell, holding a 7.0× ROAS; at a conversion-optimization agency, weekly Meta spend grew from CAD 388 to CAD 1.9K, almost five-fold, with click-through up 40 percent.

Exhibit 5

Given budget authority, Helena scales spend into returns.

Weekly Meta ad spend at two accounts with logged Helena ad work, account currency,1 navy = with Helena, blue = before Helena; dashed line = go-live

Craft-supplies brand, USD 7.0× ROAS · purchases/wk +28% · CPA −7% 2.4K 0 go-live (Apr 30) 2,400 1,450 Jan Jul Conversion-optimization agency, CAD Spend 388 → 1.9K/wk (+390%) · CTR +40% 3.6K 0 go-live (Mar 30) 3,600 132 Jan Jul

1Weekly spend decoded from each account’s Meta Graph API–based analysis; final partial weeks dropped. Window stats: craft-supplies brand: 7.0× ROAS, purchases/wk 241→309 (+28%), CPA −7%. Conversion-optimization agency: spend CAD 388→1.9K/wk with CTR +40%; cost per lead rose 23% vs. its low-volume baseline as spend scaled 4.9×. Both accounts carry logged Helena ad tasks (4 and 2 respectively); deltas compare each 90-day pre-window with the management window: correlation, not sole attribution.
Source: Enrich Labs internal data, July 2026

Enrich Labs


Helena is getting better at proposing her own experiments.

In search, Helena no longer waits to be told what to optimize. The most fully closed loop in production belongs to a telehealth company. Since May 15, Helena has run a self-scheduled task called “SEO Autopilot – Every 2 Hours”: 543 runs by July 21.1 Each run opens by reading a state file and a learnings log written by the previous runs. The standing prompt is explicit about the loop:

Read learnings from previous runs – apply these to improve quality… Finish it fully, verify quality, capture learnings, update state, then stop. The next run triggers in 2 hours.

Helena
SEO Autopilot
Standing prompt, telehealth company*

The autopilot scores every page by a formula (impressions times the gap between expected and actual click-through rate), fixes the highest-leverage page, and moves on through three phases: blog optimization, keyword-gap articles, internal linking. A companion weekly loop, “Orphan Fix,” started from a baseline audit that found 2,111 published posts with zero clicks and works through the ten highest-impression orphans each week, counting recoveries against that baseline. The result is quality, not just output: click-through rate up 11.9% versus the trailing four-week average, average position improved from 10.87 to 10.03, impressions down 7.9%: the account is ranking for fewer but better queries.

Exhibit 6

Nine weeks of self-review made one search program sharper, not just bigger.

Search performance at a telehealth company after 543 runs of the two-hour SEO autopilot loop,1 latest full week vs. trailing 4-week average

Search click- through rate indexed; trailing 4-wk avg = 100 +11.9% 100 112 4-wk avg Latest week Organic search clicks indexed; the account’s organic-traffic measure² +3.1% 100 103 4-wk avg Latest week Average search position lower is better 0.84 better 10.87 10.03 4-wk avg Latest week Impressions indexed; fewer, better queries −7.9% 100 92 4-wk avg Latest week

1The loop, named “SEO Autopilot – Every 2 Hours,” began May 15, 2026 and had completed 543 runs by July 21. Each run reads a state file and a learnings log written by prior runs, scores pages by impressions × (expected CTR − actual CTR), fixes the highest-leverage page, and records what it learned. The account’s 2,349 published posts are the raw material; the loop is what compounds their value.
Source: Enrich Labs internal data, July 2026

Enrich Labs


The same discipline shows up at smaller scale. The pain-management practice’s organic side grew alongside its ads rescue: weekly organic sessions up 85%, Search Console clicks up 112%, with the site now ranking #2 for a high-intent treatment query. At the craft-supplies brand, a weekly SEO CTR monitor audits high-impression pages on a rolling 28-day window. At a baking-supplies retailer, the loop is editorial: write a post targeting a named 8,100-searches-a-month keyword, run a monthly keyword-opportunity report on what actually ranked, then retro-fit internal links across 35 of her own previously published posts.

Exhibit 7

The loop flooded the index first, then sharpened it.

Average search position at the telehealth company, from Helena’s own weekly digests,1 Google Search Console site-wide average; lower is better2

Average search position (axis inverted; up = better) 7 9 11 13 SEO Autopilot begins (May 15) Orphan Fix loop begins (Jun 8) 6.93 11.39 10.48 Monthly site-wide 10.4 11.2 12.1 11.4 9.9 9.5 10.4 10.0 no digest ran Weekly (digests) Apr May Jun Jul 2026 The 131 background schedules behind it 14 Active recurring: standing loops still running: daily reports, weekly audits, the 2-hour SEO autopilot 116 Completed one-shots: self-scheduled graders and follow-up checks that ran, reported, and closed Paused: loops on hold

1Each value is the current week’s site-wide Google Search Console average position as reported by Helena’s “Weekly SEO Performance Digest,” whose outputs are stored in production (nine runs, May 11–July 20, 2026). No overall figure was reported for the June 22–July 6 span (dashed segment); the June 15 digest also switched its comparison baseline from week-over-week to a 4-week rolling average. Roughly 1,700 posts were published between the autopilot’s start and July 21: position first degraded to 12.1 as new pages entered the index, then improved as the optimization and orphan-fix loops worked through them.
2The monthly site-wide series comes from a July 2 Search Console analysis stored in production: clicks 58,537 (April), 54,421 (May), 51,418 (June); CTR 0.72% → 1.01% → 1.13%; average position 6.93 → 11.39 → 10.48. April’s 6.9 is the pre-flood baseline: publishing ~1,700 new pages cost roughly 4.5 positions through May, and the optimization loops have recovered about a position since while CTR rose 57% and low-quality impressions shed (8.1M → 4.5M). Click declines decelerated month over month (−7%, then −5.5%). GA4 is not connected on this account, so Search Console is its organic-traffic source.
Source: Enrich Labs internal data, July 2026

Enrich Labs


And a new surface is emerging: GEO, ranking in answer engines. GA4’s AI-referral channel is already measurable across the cohort: 6.9K sessions at the education marketplace (which runs a standing “Weekly SEO & GEO Retro”), 5.9K at a vacation-rental platform, and at the craft-supplies brand, 33.8 AI-referral sessions per week where there were zero before Helena’s answer-engine-optimized articles. The education marketplace’s average position improved from 13.4 to 11.9 over the same window. The loop that improves classic SEO is the same loop pointed at a new target.

The loop also runs at local scale. At a local financial institution in Indonesia, weekly organic sessions rose 23% within two months of go-live while local Search Console clicks quadrupled; at an Australian Italian-foods retailer, a 27-task local program has traffic up 32% over the trailing four weeks.

Exhibit 8

Local search compounds too.

Two local-SEO accounts, weekly sessions from GA4,1 organic (left) and all-channel (right)

Local financial institution: organic sessions / wk 0K 3K 6K 9K go-live (May 12) 3,856 8,194 Feb 9 Jul 13 Italian-foods retailer: sessions / wk, all channels 1,818 3,457 Apr 27 Jul 6

1Left: a local financial institution in Indonesia, where weekly organic-search sessions rose 23% (5.7K to 7K weekly run rate) within two months of the May 12 go-live, while local Search Console clicks quadrupled (475 to 2K a week, impressions +80%). Average position slipped slightly (5.8 to 6.1) as new local pages entered the index; the growth came from coverage, not rank. Right: an Australian Italian-foods retailer running a 27-task local program: all-channel sessions (organic split not recorded on this account), up 32% over the trailing four weeks; customer since March 18. Series decoded from the GA4-based charts in each account’s stored analysis; weekly alignment approximate.
Source: Enrich Labs internal data, July 2026

Enrich Labs


Exhibit 9

One account, both engines: paid search scaled 4× while organic grew underneath.

Weekly GA4 sessions at the local financial institution, indexed to the pre-Helena weekly average,1 left: indexed; right: latest full week, absolute

Sessions per week, indexed (pre-Helena weekly average = 100) 0 200 400 600 800 Helena go-live (May 12) Paid search Organic search 138 774 407 78 165 Feb 9 Jul 13 Channel mix, latest full week 8K Organic: 8.1K sessions: +23% weekly run rate; local Search Console clicks 4× 60K Paid search: 60.3K: weekly ad spend up 57% (IDR 57M → 89M), CPC down 31% 35K Other: 35.1K: direct, social, referral

1Each line is weekly sessions divided by that channel’s average over the 13 pre-Helena weeks (= 100). After the May 12 go-live, paid-search sessions peaked at 7.7× the pre-Helena run rate and settled near 4× as weekly ad spend rose 57% (IDR 57M to 89M) with CPC down 31%; organic sessions climbed steadily to 1.65× (+23% on the weekly run rate) as the local-SEO program compounded. Series decoded from the GA4-based charts in the account’s stored analysis; weekly alignment approximate; final partial week excluded.
Source: Enrich Labs internal data, July 2026

Enrich Labs


Helena checks her own work.

Email is where Helena’s method is most visible, because the pattern on winning accounts is so consistent: audit first, then build, then schedule a check on what you built. At an apparel retailer, the sequence was an initial Klaviyo audit, a rebuild of the weakest flow, then a recurring monthly flow-health audit. Eleven email accounts across the cohort now run recurring audits like that.

A skincare brand shows the full loop on one account, across 99 email tasks. Helena built the flow stack (abandoned cart in brand hex codes, browse abandonment, post-purchase) and then scheduled her own graders: a “Browse Abandonment Flow – 2-Week Check-In” that pulls entries, opens, and clicks per email ID on the exact flow she built, and a “Memorial Day Email Performance Check” that reviews three named campaigns and emails the owner a summary. Nothing ships without a scheduled look back at whether it worked.

Where the loop runs, the numbers move. A sporting-goods brand ran a full “Father’s Day Campaign Post-Mortem” (a dated retro across the ads and emails of a single campaign window), and over its management window, email went from an afterthought to the engine: email’s share of revenue rose from 6.4% to 46.3%. At a digital-products brand, every recurring content task begins with the instruction to read a persistent strategy file (“FIRST ACTION – Read the master strategy file… files/strategy/blog-master-strategy.md”) before executing, and a “Blog Performance Audit – 2-Week Check-In” plus a Monday-morning digest close the loop; email revenue per week is up 21% and conversions up 31%. One caveat we repeat from the underlying analysis: GA4 understates flow-driven email revenue, so these figures are floors.2

Exhibit 10

From dead channel to demand engine.

Weekly email-attributed sessions at the sporting-goods brand,1 GA4 Email channel, February 9 – July 12, 2026

0 100 200 300 400 500 272 536 183 217 152 131 241 184 Helena go-live (Apr 22) Feb 9 Mar 16 Apr 20 May 25 Jun 29 2026, weekly · navy = with Helena, blue = before Helena

1Helena went live April 22 and built the account’s flow stack, then ran a dated post-mortem on its Father’s Day campaign. In the Helena window: ~USD 2K email-attributed revenue and 18 email conversions; weekly run-rates rose from $40 to $163 in email revenue and from 6.4 to 46.3 percent of total revenue. Weekly values decoded from the account’s GA4-based analysis; bars under 10 sessions are approximate. GA4’s Email channel understates flow-driven revenue; treat as floors.
Source: Enrich Labs internal data, July 2026

Enrich Labs


The revenue follows the flows. At the skincare brand, $7.9K of email-attributed revenue followed the new flow stack, each flow with its own scheduled check-in; at a stationery brand, email went from nothing to $2,700 a week.

Exhibit 11

Email revenue where there was none.

Email-attributed revenue at three accounts, before vs. with Helena,1 GA4 Email channel

Skincare brand email-attributed revenue in the window (flow stack built) from zero $0 $7.9K Before With Helena Stationery brand weekly email-attributed revenue from zero $0 $2,700 Before With Helena Sporting-goods brand weekly email-attributed revenue $40 $163 Before With Helena

1“Before” is each account’s 90-day pre-Helena window. The skincare brand’s $7.9K followed Helena building its full flow stack (abandoned-cart, browse-abandonment, post-purchase) with scheduled 2-week check-ins on each flow. The stationery brand’s $2,700/week comes from flow-driven buying: $41.9K of email-attributed revenue on just 136 email sessions in the window. The sporting-goods brand’s weekly run rate rose from $40 to $163 as email’s share of total revenue went from 6.4 to 46.3 percent. GA4’s Email channel understates flow-driven revenue (opens convert as direct); treat all figures as floors.
Source: Enrich Labs internal data, July 2026

Enrich Labs


Closing the loop: Helena improves the system that improves the marketing.

Above the channel loops sits a loop about the loops, and this is the step that separates automation from self-improvement. It is also the closest thing in this data to recursion: Helena improving the machinery that improves the marketing. Every account runs a weekly task whose entire job is to review Helena’s own automations and propose one change: a new automation, an update, or a deletion. At the pain-management practice it has run 24 times; at the craft-supplies brand, 12.5 At the telehealth company, in mid-July, it noticed that Helena’s own publishing queue was being triggered manually and proposed automating itself out of the bottleneck: the quote at the top of this research.

A companion weekly task, Autonomous Learning, maintains a per-platform learning file (What’s Working, What’s Not, Tried & Ruled Out, Open Experiments), with the standing instruction to build on prior insight rather than starting cold. Helena’s memories function as an experiment log across channels: bad-data quarantines (“do not use May 22–28 as a baseline”), comparability warnings after bid-strategy changes, and human-intent annotations (“paused by the owner on July 1 – intentional, do not flag as needs-review in future reports”). Even channel expansion is a data-gated decision: Bing sits queued at the craft-supplies brand to be switched on “once Google data justifies it.”

Exhibit 12

Where Helena is allowed to build, more accounts improve than not.

Share of sampled accounts improving after go-live, by metric,1 % of accounts with a measurable baseline; sample = 50 highest-spend accounts per channel from the 1,000+ account study base

Meta ads CTR (n = 50) click-through rate improved vs. 90-day pre-window 52 Organic sessions (n = 50) weekly GA4 organic sessions rose 42 Search clicks (n = 28) weekly Search Console clicks rose 46 Paid CPA (n = 33) cost per acquisition fell; conversions in both windows 36

1Each metric compares the account’s Helena management window with its 90-day pre-Helena window; the base (n) is accounts with a measurable baseline on that metric: 26 of 50 improved Meta CTR, 21 of 50 grew organic sessions, 13 of 28 grew Search Console clicks, and CPA fell on 12 of 33 accounts with conversions in both windows. Deltas reflect correlation, not sole attribution.
Source: Enrich Labs internal data, July 2026

Enrich Labs


543
runs of the two-hour SEO Autopilot at the telehealth company since May 15
300
background automations created at the most-automated account, an 11-brand media operator
25
self-scheduled check-ins on one Google Ads campaign at the craft-supplies brand
24
runs of the weekly review of Helena’s own automations at the pain-management practice

And the honest number: 23 of the 50 sampled Meta accounts, and roughly 95 of the 116 email-connected accounts, have wired Helena up but given her little building to do yet. Connection is the on-ramp; activation is the upside. The gap between what the system can do and what customers have switched on is currently the largest untapped source of improvement in the data.

What might the future of marketing work look like?

The evidence suggests the human role is narrowing at each step of the marketing cycle. Once AI-built campaigns and content reach parity with agency work (and on cost and consistency they arguably already have for lean teams), humans stop producing and shift to reviewing. But review has its own ceiling: at the telehealth company’s pace of 900 posts a month, no founder reads everything. The question then shifts from “is this piece good?” to “is the system that produces these pieces pointed at the right goal, with the right guardrails?” Put simply: the doing (writing the ad, building the flow, shipping the report) now costs almost nothing in human time.

An area of human comparative advantage, for now, is marketing taste and judgment: choosing which market matters, which numbers to trust, and when a strategy is a dead end. The customers in this piece exercise exactly that. The craft-supplies brand’s owner paused a Helena campaign on July 1 based on Helena’s own 14-day review; the medical group’s team set the insurance-plan exclusions Helena treats as law. The founders supply the goal; increasingly, they no longer need to supply the method.

What if we’re wrong?

A natural objection to the evidence presented above is that the work still in human hands (choosing what the business should say, to whom, at what price) is what matters most. Without that judgment, Helena is a capable executor, not a system that could drive marketing on its own. That objection is partly right, and worth taking seriously.

Our data has honest limits. Every before-and-after in this piece compares a 90-day pre-window to a management window in which many things changed at once; the deltas reflect correlation, not controlled attribution. GA4 understates email flow revenue, so those figures are floors. One customer’s conversion count is flagged in our own analysis as almost certainly inflated by micro-events, a flag Helena herself raised. And nearly half of connected accounts haven’t activated deeply enough to show anything yet. A standing loop is also only as good as the integration beneath it: at the medical group, the Google Ads connection’s token expired in early June, and the daily campaign report ran 43 times returning only re-authorization requests: the loop kept showing up, but the data pipe stays broken until a human reconnects it.

But marketing is rarely advanced by strokes of creative genius. In between, most progress is incremental: ship something, measure it, fix it, and try again, on the ad account, the search rankings, and the email flows all at once. That is exactly the kind of workflow the system now excels at, and the perspiration is becoming automated. The humans in this piece supply the goal and the judgment: The craft-supplies brand’s owner paused a Helena campaign on July 1 based on Helena’s own 14-day review; the medical group’s team set the insurance-plan exclusions Helena treats as law. Increasingly, they no longer need to supply the method.

Possible futures

What happens next depends on whether the trend continues, and what teams choose to do if it does. We can imagine at least three scenarios:

  1. The trend stalls, but today’s capabilities are widely diffused. The channel loops described here may be near their ceiling: execution and optimization automate, but strategic judgment stays stubbornly human. Even so, the diffusion alone changes the industry: one person sits atop a pyramid of standing loops that audit ads, prune keywords, sharpen rankings, and check every flow. The connected-but-dormant accounts in our own data suggest most of this diffusion is still ahead.
  2. Marketing systems continue to see compounding efficiency gains. Each channel loop keeps improving the next run, while humans set direction and judge results. Speeding up one part of a process shifts the bottleneck elsewhere: the constraint becomes the human capacity to decide what’s worth testing and to verify what’s working. We have already met this friction: Helena’s weekly suggestion queue generates more improvement ideas than most customers act on. The rate at which a team can spot and clear its own bottlenecks may become its most important marketing skill.
  3. The loop closes. If the trends continue, systems like Helena could become capable of redesigning their own strategy end to end: noticing a market shift in the data, re-deriving the positioning, and rebuilding the channel mix, with humans reviewing rather than directing. That would be recursive self-improvement in miniature: not a model retraining itself, but a marketing system rebuilding the system that does the marketing. The early signals exist: Helena already proposes changes to her own automations, quarantines her own bad baselines, and overrules her own prior hypotheses in writing. Whether strategic taste is a capability that AI systems fail at for a time and then get good at is the open question. We would not bet against it.

What should we do?

This research will land differently depending on where you sit. If you are a marketing operator already tinkering with AI, running prompts through a chatbot, trying a copy tool here and an ad tool there, the evidence points to one specific upgrade: stop using AI as a faster keyboard and start giving it a loop. The lift in this study did not come from better prompts. It came from connected accounts, scheduled reviews, and learning files that carry forward what has already been tried and ruled out. The practical guidance falls out of the evidence:

  1. Connect the accounts. Whatever AI you use, give it eyes on your real channels: the ad accounts, the analytics, the search data, the email platform. Nothing in this piece happened on an unconnected channel, because an AI that cannot see performance can only guess at it.
  2. Let it build, not just draft. The measurable lift concentrated where the AI owned real campaigns, flows, and content end to end. If your AI only produces drafts that a human still has to ship, you are paying for typing speed, not outcomes.
  3. Insist on the loop. One-shot output is where most AI marketing experiments stall. Every build should carry its own scheduled review, every review should write down what worked and what was ruled out, and that record should be readable by you.
  4. Make the guardrails explicit. Before handing anything over, write down your thresholds, exclusions, and non-negotiables, and hold whatever system you use to them as law. The best accounts in this study did this on day one.
  5. Keep humans on the two jobs the data says still belong to them: setting direction, and deciding which results to trust.

And if you are less sure where marketing is headed, whether agents are a real shift or another hype cycle, that skepticism is reasonable, and this piece was written for it. Every number above comes from production systems, not a demo. Marketing that improves itself is no longer a thought experiment. It is running on the ad accounts, the search rankings, and the email flows: on a two-hour timer, in production, right now.

If you are experimenting with AI in your own marketing, unsure where to start, or want to go deeper on the data and methods behind this research, we would like to compare notes. Write to seijin@enrichlabs.ai: questions, pushback, and collaboration ideas are all welcome.

About the research

This research is built from live production data across more than 1,000 active customer accounts running their marketing on Helena, channel deep-dives across the 50 highest-spend Meta, Google, and email accounts sampled from that base, and full reconstructions of every scheduled task, memory, and learning file for four customers: a telehealth company, a craft-supplies e-commerce brand, a pain-management practice, and a multi-location medical group. Performance data is drawn from the Meta Graph API, Google Ads, Google Search Console, GA4, and connected email platforms, as of July 21, 2026.

All quotes attributed to Helena are drawn verbatim from her production task history, memory logs, and weekly learning files, and are used with customer permission. They reflect the state of each account as of July 2026.

Footnotes
  1. The loop, named “SEO Autopilot – Every 2 Hours,” began May 15, 2026 and had completed 543 runs by July 21. Each run reads a persistent state file and a learnings log written by prior runs, works through three phases (blog optimization, keyword-gap articles, internal linking), and scores candidate pages by impressions × (expected CTR − actual CTR).
  2. Spend figures are USD-normalized across currencies and cover each account’s Helena management window. Email-attributed revenue is measured via GA4’s Email channel, which understates flow-driven revenue (opens convert as direct); treat email figures as floors.
  3. All before/after comparisons contrast a 90-day pre-Helena window with the Helena management window on the same account. Where a brand ran marketing long before Helena, deltas reflect correlation, not sole attribution: seasonality, budget changes, and human decisions also move these numbers. The underlying per-account analyses state this caveat and we repeat it here deliberately.
  4. Customer tenures at time of writing: the telehealth company ~4.5 months, the pain-management practice ~3 months, the craft-supplies brand ~2.5 months, the medical group ~6 weeks. Revenue and booking-level outcomes beyond ad-platform and analytics attribution are not captured in this dataset.
  5. “Weekly Automation Suggestions” run counts by account as of July 21: the pain-management practice 24, the craft-supplies brand 12, the telehealth company 9, the medical group 7. The task’s standing prompt: “Review my existing automations and business data, then suggest one improvement – a new automation, an update, or a deletion.”

* Quotes attributed to Helena throughout this research are drawn verbatim from her production task history, memory logs, and weekly learning files, and are used with customer permission. They reflect the state of each account as of July 2026.

Written by Enrich Labs, from live production data across more than 1,000 active customer accounts, channel deep-dives across the 50 highest-spend Meta, Google, and email accounts sampled from that base, and full reconstructions of four accounts: a telehealth company, a craft-supplies e-commerce brand, a pain-management practice, and a multi-location medical group. Performance data drawn from the Meta Graph API, Google Ads, Google Search Console, GA4, and connected email platforms, as of July 21, 2026.

Want to see how Helena can run your marketing? Start your free trial at hirehelena.com.

Run your brand’s marketing on autopilot

Start free