Skip to main content

AI Slop Grenades: Why Reviewing AI Ad Creative Is Part of What It Costs

On a podcast released September 15, 2026, Shopify's CEO said the failure case of lazy work is now over-output, and that staff call unreviewed AI output a slop grenade. Here is what that argument is, and our reading of where the same checking cost lands in an ad account.

Mauricio Valdivia

Mauricio Valdivia

·11 min

AI Slop Grenades: Why Reviewing AI Ad Creative Is Part of What It Costs

Generating the ad got cheap. Checking it did not.

A folder lands in the shared drive on Monday with forty vertical renders in it. Maybe fourteen are usable. Nobody knows which fourteen. The person who generated them has moved on to the next brief, and the buyer who has to launch on Wednesday now owns the sorting.

On The Knowledge Project podcast, released September 15, 2026, Shopify CEO Tobi Lütke gave that shape a name. Asked what AI had made worse inside his company, he said "The failure case now of lazy work is not lack of output." The failure is the opposite: too much output, handed over unread. "So we call those 'slop grenades' that people toss at each other," he said.

He was describing code review and internal email at a commerce company. Nothing in that interview is about advertising. The question this post asks is ours: when a generator can produce forty clips before lunch, who checks them, and what does the checking cost?

Here is what this post covers:

  • What Lütke said, when he said it, and what he did not say.
  • Our argument that the same handoff cost lands on whoever runs the ad account.
  • The five things a person still has to catch in an AI ad, with one worked example in real minutes.
  • What the outside evidence says about cleanup work, each number attributed to the body that published it.
  • The gates we would put in a creative pipeline so nothing ships unread.

What Shopify's CEO said, and when he said it

The remark, and the episode it came from

The interview ran on The Knowledge Project, hosted by Shane Parrish. The show's own feed dates the episode Tuesday, September 15, 2026, and its chapter list carries a segment titled "(16:04) What AI is Making Worse at Shopify", which is the question that produced the answer.

The remark travelled before it reached us. Fortune carried the slop grenades quote on September 17, 2026, and Futurism and others followed, all pointing back at the same episode. Search Engine Journal published the write-up this post works from on September 21, 2026, and that write-up is where the "failure case now of lazy work" sentence appears verbatim. So the story broke in mid September, not with the SEO trade press a week later.

The two examples he gave

Search Engine Journal reports that Lütke explained the issue as over-output and gave two examples. In the first, an employee asks an AI agent for a code change, approves the pull request without reading it closely, and leaves colleagues to review it. In the second, someone uses a language model to turn a short point into a long email, which the recipient then shortens again with another model. His advice on the email was to use the model to make the point shorter, not longer.

The email example is the one worth staring at. Two model calls, one on each end, and the information that actually moved between two humans is a sentence. Everything else was work created and then destroyed.

Where the phrase came from, and the memo behind it

Lütke did not coin the term. Search Engine Journal reports that he credited Harry Brundage with coming up with it and said he hopes it becomes popular, so the phrase belongs to its author and the usage is what Shopify staff do with it.

The context is older than this month. In April 2025 Lütke posted an internal memo saying "reflexive AI usage is now a baseline expectation at Shopify", a memo that also made AI use part of peer and performance reviews. That is a 2025 fact and it deserves its date. The thing we find notable is the sequence: the executive who made AI use a baseline is the same one now naming, in public, the failure mode that baseline produced.

The Novoads app: pick an AI actor, write a script, generate a UGC ad
Novoads · UGC video ads with AI, ready in minutes.
Try now

Our reading: in an ad account, the grenade lands on the buyer

What follows is our argument, not Lütke's. He is talking about pull requests and internal mail. We make ad creative, and we think the same mechanic shows up in an ad account with the roles renamed.

A render is a claim, and claims get checked downstream

A code change that nobody read is a liability with a diff attached. A video ad that nobody watched end to end is a liability with a spend cap attached. In both cases the artefact looks finished, which is exactly what makes it expensive: a rough draft invites scrutiny, a polished file suggests it has already had some.

An ad carries more than a quality judgement. Four things ride along with every file, and the model that produced it checks none of them:

  • A product that has to be your product.
  • A claim that has to be substantiable.
  • A disclosure that has to sit where the platform requires it.
  • A format that has to match the placement.

The handoff is where the cost hides

Generation has a price tag and review does not, so review is the half that disappears from the plan. That is the whole trick of the slop grenade as a concept: the work did not vanish, it moved to somebody whose time is not on the same line of the budget.

In a small team the move is usually one of these:

  • The strategist generates, the buyer reviews. The buyer's Wednesday disappears.
  • The buyer generates, nobody reviews. The account finds the bad ad, at media cost.
  • The founder generates at 11pm, and the review is a vibe check on a phone screen.

What does not transfer

We should be honest about the limits of the analogy. Shopify's problem is measured in engineering review time on internal work. Ad creative has an external referee that code review does not: the auction. A bad ad that slips through gets punished by the platform and by the spend, which is a feedback loop internal email never had. That cuts both ways. It means the market eventually catches what your review missed, and it means it catches it with your money. If you want the metric side of that argument, we have written about why a high click-through rate does not prove an ad works.

What the review step actually checks in an AI ad

Five checks that survive the generator

These are ours, in the order we would run them:

  1. Product accuracy. The bottle on screen is your bottle: the cap, the label geometry, the colour. This is the check that fails most often when a model is asked to hold a product.
  2. Claim accuracy. Every spoken line and every on-screen word is a claim you could substantiate if asked. A generated script will happily invent a percentage.
  3. Disclosure and policy. Whatever the placement requires is present, in the form the platform asks for, and it survives the crop.
  4. Craft. Hands, lip sync, drift between shots, artefacts you only see on a second viewing. Our notes on making AI video look real are essentially a list of what to look for here.
  5. Spec. Aspect ratio, duration, safe area, audio levels. Boring, and the cheapest one to automate.
  6. Placement fit. Every new placement you take on adds a row to this list, and it is the row that gets skipped when a door opens. Threads dropping its Instagram requirement is a fresh example: it arrives asking for creative built for a feed with no button on it.

A file-level score can pre-filter the first, fourth and fifth of those. It cannot finish the second or third, because substantiation and disclosure are judgement calls about your business rather than properties of the pixels. That boundary is the subject of our piece on what AI creative scoring checks and what it misses.

Two minutes a clip is a budget line, not a rounding error

Assume a fast pass: two minutes per clip, watched once, with a yes or no at the end. That number is our working assumption, not a measured figure from anyone's study. Forty clips is eighty minutes, and eighty minutes of a senior person's attention is the most expensive thing in the batch.

Compare it with what the clips cost to make. On our own pricing, a clip runs from about 25 cents for a five-second Seedance 2.0 Mini clip to about $4 for an eight-second Google Veo 3.1 clip, and about $10 for a full 30-second Seedance 2.5 take. So the render bill for a batch of forty short clips is smaller than the review time attached to it, at any sensible hourly rate. If you want to do this arithmetic properly for your own stack, the mechanics of converting vendor credits into a cost per finished ad are in our guide to AI video credits.

The useful move is not to generate less. It is to decide, before the batch exists, how many of those clips you actually need in market, which is a function of how many test cells your conversion volume can resolve rather than how many the generator can produce. We worked that out in how many ad creatives you need.

Whoever generates, reviews

One rule fixes most of this: the person who generates a batch watches it before anyone else sees it. Not a full review, just enough to delete the obvious failures. It sounds trivial. It is the exact step Lütke's two examples skip, and skipping it is what turns a helpful batch into a handoff.

Several UGC creators filming product variations to camera
Novoads · UGC video ads with AI, ready in minutes.
Try now

The cleanup market is already being priced

These numbers belong to the organisations that produced them. We are reporting them, with their dates, not presenting them as our findings.

FindingWho published itDate
Up 87% to 10,760 listingsFreelancer.com data, via the GuardianSeptember 2, 2026
Up 70% year over yearUpwork data, via the GuardianSeptember 2, 2026
Searches up over 20-foldFiverr data, via the GuardianSeptember 2, 2026
About 4 in 10 hit workslopBetterUp with the Stanford Social Media LabSurvey, September 2025
Volume up for 88%, quality for 45%WARC and LIONS Advisory with TikTokJuly 14, 2026

Freelance marketplaces: the Guardian's September 2, 2026 figures

Internal platform data shared with the Guardian on September 2, 2026 and summarised by Search Engine Journal covers three marketplaces:

  • Freelancer.com: listings tagged with phrases like "correct AI" and "AI hallucination" grew by 87%, reaching 10,760 worldwide between August 2025 and June 2026.
  • Upwork: data indicated a 70% year-over-year rise in AI remediation gigs.
  • Fiverr: searches for "AI cleanup" services increased more than 20-fold from 2023 to 2026.

Those are marketplace signals, and marketplace signals measure demand for a service, not the total size of the problem. The freelancers quoted in that reporting disagreed among themselves about how long the work lasts, with one writer expecting much of it to dry up within five to ten years as models improve.

Workslop: BetterUp with the Stanford Social Media Lab, September 2025

BetterUp ran a survey in September 2025 with the Stanford Social Media Lab, covering over 1,000 full-time US desk workers. About four in 10 respondents said they had encountered "workslop", described as AI-generated work that looks polished but does not advance the task, in the past month. On average, each instance took nearly two hours to resolve. The survey is self-reported, and it is about desk work in general rather than ad production.

Two hours per instance is not a number you can paste into a creative plan. It is a number that tells you the handoff cost is large enough that people notice it and can estimate it a year before a CEO gave it a name.

Volume up, quality flat: WARC and LIONS Advisory with TikTok, July 2026

The closest thing to a creative-industry version of this sits in research we covered in July: WARC and LIONS Advisory, in partnership with TikTok, found that while 90% of marketers now use AI as a core creative tool and 88% report an increase in creative volume, only 45% see significant improvement in quality. That gap, and what the same study says about how teams brief the model, is the subject of our post on the AI ad quality gap.

Read the three together and the pattern is consistent. Output went up everywhere. The checking did not get cheaper, and somebody is paying for it either in freelance invoices, in colleagues' afternoons, or in flat quality at higher volume.

Gates in the pipeline beat a review meeting at the end

A weekly review meeting is where slop goes to be discovered too late. Gates are cheaper because they reject earlier, and because they are specific enough to be delegated.

Gate 1: the brief decides what a good render looks like

Before anything renders, write down the one thing this batch is testing and the one thing that would disqualify a clip. A batch that tests hooks does not get rejected for a background. Changing one layer per round is the discipline that makes this possible, and we set it out in how to create ad variations.

Gate 2: reject at the first frame, not the final cut

Most failures are visible in the first second. Look at thumbnails before you watch anything end to end, kill the obvious rejects, and only then spend two minutes each on the survivors. A thumbnail pass catches:

  • The wrong product, or a product the model redesigned.
  • The wrong actor for the audience the batch is for.
  • Framing that crops the product out of the safe area.
  • A hook frame that is a stock-looking establishing shot.

This is the difference between an eighty-minute review and a twenty-minute one.

Gate 3: one name per file

Every file that reaches the account has one person attached to it who watched it. Naming, versioning and a refill cadence are the unglamorous machinery that makes this hold under weekly pressure, which is what creative operations is for. When a file has no owner, the review has no owner either, and that is the state Lütke was describing.

Once the gates exist, the platform test tools decide the rest. What each network will actually split and how long a read takes is in our rundown of ad creative testing tools.

How Novoads solves the review step

Novoads makes the inputs explicit before anything renders. You write the script and pick the AI actor, and a product ad starts from a product image you upload, so a finished clip is the result of choices you made rather than a prompt nobody kept. That changes three things about the review:

  • The claim check starts earlier. You wrote the line, so a wrong claim is catchable in the script, before it becomes forty spoken versions of itself.
  • A rejection means something. Clips that differ by one choice are comparable, so killing one is a decision about that variable rather than a verdict on a mystery file.
  • The product check has a reference. A product ad is built from an image you chose, which is what you are checking the render against.

That is the whole product argument here, and it is a modest one: we cannot review your ads for you, and we would not trust a tool that claimed it could. What we can do is make the review fast enough that nobody is tempted to skip it. You can generate a batch and see what the review actually feels like, and the plans are published on novoads.ai/pricing.

A UGC creator filming a skincare product review on a phone
Novoads · UGC video ads with AI, ready in minutes.
Try now

Volume is the easy half

The cheap part of AI creative was always going to be the making. What stayed expensive is the deciding, and the temptation the tools create is to move that cost onto whoever is standing downstream. Lütke named that move inside a software company. In an ad account the same move has a shorter fuse, because the thing downstream is not a colleague's inbox, it is spend.

So budget the review the way you budget the render. Decide who watches, decide what disqualifies a clip, and decide it before the folder with forty files in it lands on someone's Monday.

Frequently Asked Questions

What is a slop grenade?

It is AI output that gets handed to a colleague without being read, so the recipient pays the checking time. Shopify CEO Tobi Lütke used the phrase on The Knowledge Project podcast released September 15, 2026, saying 'So we call those slop grenades that people toss at each other'. He credited Harry Brundage with coining the term and said he hopes it catches on, so it is not Shopify's coinage, and Search Engine Journal reports the same credit.

When did Shopify's CEO say it?

The episode of The Knowledge Project carrying the remark was released on September 15, 2026, the date the show's own feed gives it, and the episode's chapter list includes a segment titled 'What AI is Making Worse at Shopify'. The quote was picked up widely from September 17, 2026, when Fortune ran it. The Search Engine Journal write-up that summarises the interview is dated September 21, 2026.

Does this mean AI ad creative is a bad idea?

No, and the interview does not say that. The failure mode Lütke describes is the unread handoff, not the generation. Our reading for ad teams is that volume is the cheap half of the job and review is the half that still costs a person's time, so the answer is to put explicit gates in the pipeline, not to generate less.

How long does checking AI output actually take?

The only measured figure we would cite belongs to someone else. BetterUp, with the Stanford Social Media Lab, surveyed over 1,000 full-time US desk workers in September 2025: about four in 10 said they had encountered 'workslop' in the past month, and each instance took on average nearly two hours to resolve. That is self-reported and it is about desk work in general, not about ad creative, so treat it as a signal about handoffs rather than a benchmark for your review queue.

What should a human still check on an AI ad before it goes live?

Five things, in this order: that the product on screen is actually your product, that every spoken or on-screen claim is one you can substantiate, that the disclosure and policy requirements of the platform are met, that the craft holds up (hands, lip sync, drift, artefacts on a second viewing), and that the file matches the placement spec. A creative score can pre-filter for some of these, but the substantiation and disclosure checks are judgement calls a file-level score cannot finish.

How does Novoads help with the review step?

It makes the inputs explicit. You write the script and pick the AI actor, and a product ad starts from a product image you upload, so the claim check can happen on the script before it becomes forty spoken versions of itself, and clips that differ by one choice are comparable when you review them. Plans and what each one includes are published on novoads.ai/pricing.

Key Takeaways

  • On The Knowledge Project podcast, released September 15, 2026, Shopify CEO Tobi Lütke said the failure case of lazy work is not lack of output, and that staff call unreviewed AI output handed to a colleague a slop grenade. He credited Harry Brundage with the term rather than claiming it.
  • His two examples are a pull request approved without being read and a short point inflated into a long email that the recipient shortens again with a second model. Both are about code review and internal mail at Shopify, not about advertising.
  • Our argument, not his: in an ad account the same handoff cost lands on the person who launches, because a render is a claim about a product that somebody downstream still has to check.
  • Outside evidence, each with its own owner. Freelance marketplace data shared with the Guardian on September 2, 2026 put listings tagged 'correct AI' and 'AI hallucination' up 87% to 10,760 worldwide between August 2025 and June 2026. A BetterUp survey with the Stanford Social Media Lab, run in September 2025 with over 1,000 US desk workers, found about four in 10 hit 'workslop' in the past month, taking on average nearly two hours to resolve per instance, self-reported.
  • Price the review step the way you price the render. A fast two-minute pass on forty clips is over an hour of someone's afternoon, which is more than the clips themselves cost, so put the gates in the pipeline rather than in a meeting after it.
Mauricio Valdivia

Mauricio Valdivia

Founder of Novoads

Mauricio is the founder of Novoads, where he works to democratize video advertising with AI for brands in Latin America.