Recent post

The retention measurement window most teams default to is D7, sometimes D30 if the budget cycle allows. Nobody really chose that. It’s just how fast platforms report, how fast budgets need answers, how fast everyone’s used to moving. A week became the standard because it was the fastest option on the table, not because someone decided a week was long enough to tell whether a user would stick around.

And here’s the problem with that. A week tells you almost nothing about whether someone likes the product. It tells you whether the first session was good enough to bring them back once, maybe twice. Related question. Not the same question. Treat them as interchangeable long enough, and you end up crowning the wrong creative a winner without ever noticing the mistake.

This whole series has been about building creative as a system rather than by guessing. A short retention measurement window undermines that idea in a specific way, because it hands you a verdict before the evidence has fully come in. What follows is what changes as the window gets longer, and what to do about it in the meantime, both on the creative and media-buying sides.

Why D7 Became the Default in the First Place

There’s a practical reason this happened, worth naming honestly, rather than just piling on the industry for moving fast. Platforms report fast. Budgets are reviewed weekly, sometimes more often during a launch push. A team scaling winners and killing losers needs a signal quickly, or the whole testing cadence stalls out waiting on data that won’t show up for months on end.

D7 isn’t a wrong number. It’s an incomplete one, and worth treating that way rather than as a verdict. It answers something real, just something narrower than most teams treat it as answering. The mistake was never using D7. The mistake is letting D7 be the final word on whether a piece of creative worked, instead of treating it as an early read that still has to earn confirmation over a longer window before real budget follows.

What Changes at D60 and D90

Stretch the window out to sixty or ninety days and the picture shifts, sometimes uncomfortably, for whoever picked the D7 winner. Creative that looked strong in week one can fade hard by month two. Usually, the hook oversells something the product doesn’t quite deliver, and the users it pulled in were never a real fit, no matter how good the app is underneath. Other creative looks unremarkable at D7, middling numbers, nothing to write home about, and then holds steady or climbs by D60, because the smaller audience it pulled in was the right one all along.

That second pattern is easy to miss entirely if D7 is the only number in the room. A campaign killed in week one for underperforming might have been building the healthiest cohort in the whole account, and nobody would ever find out, because it never got the runway to prove it. A longer retention measurement window is the only way to catch that before the decision’s already made.

Maybe the simplest way to say it: D7 measures excitement, D60 measures fit, and those are not the same thing to a business trying to build something durable. A hook built to generate excitement isn’t automatically a hook built to generate fit, and sometimes it works against it entirely. Five seconds of the flashiest gameplay footage can pull in the wrong audience for a slower, more strategic game. People arrive excited, then bounce the moment they see what they downloaded.

What This Changes About Testing Creative

If the real signal doesn’t land until D60 or D90, that reshapes how a testing cadence runs week-to-week, not just how the results are written up afterward. Kill losers off D7 data alone, and sometimes you’ll kill creative that was about to prove itself. Scale winners off D7 data alone, and sometimes you’ll scale creative that looks a lot worse in two months than it does today.

The fix isn’t ignoring D7. Speed still matters, and freezing every decision for 90 days would stall testing just as badly, just in a different direction. Treat D7 as provisional instead. A campaign that clears it keeps running, but the real verdict on more budget waits for a longer window. And a campaign that’s borderline at D7 but showing early signs of quality, engagement depth, session length, something beyond the raw retention number, probably deserves more patience than a hard cutoff would give it.

The Media Buying Side of This

This changes budget allocation more than teams typically expect, and it’s worth walking through in concrete terms. Scale a campaign hard off D7 data alone, and you risk pouring spend into an audience that churns out by D60, money already spent, cohort already gone. A more conservative curve, one that waits for at least a partial D30 read before committing serious budget, costs some early efficiency. In exchange, it buys protection against that exact outcome.

There’s a real tension here, worth naming instead of smoothing over. Wait for a longer retention measurement window before scaling, and you move slower than competitors happy to scale off D7 alone, especially in a fast-moving account, which can feel like leaving money on the table while a campaign is still cheap to run. There’s no single right answer. It depends on category and appetite for risk. A slower, D60-informed approach fits a subscription product where lifetime value hinges on long-term retention. A faster D7-based approach can still make sense for something with a short natural lifecycle, where D60 was never going to be the number that mattered.

None of this argues for abandoning speed altogether across the account. It argues for knowing which decisions can afford to move fast and which ones need a longer read before real money follows.

Leading Indicators Worth Watching Before D60 Arrives

Freezing every decision for ninety days isn’t realistic for most accounts, and it doesn’t have to be the only alternative to blind D7 scaling. There’s a middle path: watch for signals that correlate with longer-term retention but appear faster than the retention curve.

Session length and session frequency in the first week often say more than D7 retention alone ever will, once you actually sit down and compare them side by side. Two users who both open the app on day seven aren’t in the same place just because they both showed up. One might have opened it for thirty seconds out of habit. The other had a genuinely engaged session that points toward real staying power. Engagement depth, how far someone gets into onboarding, whether they complete a core action, whether they come back more than once in the first few days, tends to predict D60 retention better than the raw D7 number.

None of this replaces waiting and watching what happens. It makes the provisional call smarter than a coin flip between blind scaling and a full freeze. A campaign that clears D7 with mediocre retention but strong engagement depth deserves more patience than one that clears D7 with the same mediocre retention and shallow engagement to match. On a bare D7 chart, those two look identical. Factor in engagement depth and they stop looking anything alike.

Worth saying plainly: leading indicators aren’t a substitute for the real answer. They’re a way to hold a position responsibly while waiting for it. Treat them as confirmation of patience, not proof of a winner, and they earn their keep. Treat them as a shortcut past the retention measurement window, and they’ll mislead someone the same way D7 alone already does.

A Composite Example

Picture two ad variants for a strategy mobile game. Variant one leads with fast-paced combat footage, punchy and immediately engaging, and it clearly wins at D7, with strong CPI and solid week-one retention. Variant two leads with a slower shot of the strategic decision-making the game is built around. Less flashy. It comes in noticeably behind variant one on every D7 metric.

By D60, the picture flips. Variant one’s users, drawn in by combat footage for a game that’s mostly about patient strategic planning, mostly churned out once they realized what they’d downloaded. Variant two’s smaller initial cohort is retaining at nearly twice the rate because those users knew roughly what they were signing up for, and it matched what they wanted. Kill variant two at D7 the way a strict cutoff would suggest, and the team scales the wrong ad without ever finding out.

The Fetch

If your creative decisions are being made entirely off D7 data, there’s a real chance you’re scaling the wrong winners. A longer retention measurement window doesn’t have to mean slower decisions across the board. Just better-informed ones on the campaigns that are actually close calls. Reach out and let’s look at what a longer read on your account shows.