Signal › Feed

Salesforce's AIforce: 'AI Replaces the UI' and What It Costs to Believe It

A16Z's analysis of 100 trillion real-world AI tokens found that the products with the best long-term retention don't have the best models — they have users whose specific unsolved workload fits the model precisely. The 'glass slipper' effect explains why your earliest cohort is either your retention foundation or your ceiling.


In May 2026, a16z published an analysis built on a dataset most AI companies have never seen: 100 trillion real-world AI token interactions collected through OpenRouter, spanning dozens of frontier models across hundreds of thousands of users over multiple years. The researchers came looking for patterns in model adoption. What they found instead was a retention anomaly that upends the standard AI product growth playbook.

They called it the glass slipper effect.

The core finding: the AI products with the strongest long-term retention are not the ones with the best models. They are the ones whose earliest users found an exact match between a specific, high-value, unsolved workload and the model's particular capabilities. Once that match happens, those users stop experimenting. They settle. And settled users don't churn — they expand.

The strategic implication is direct: most AI product retention analysis is measuring the wrong thing. Aggregate retention curves mix workload-fit users with casual experimenters, obscuring both the product's real retention floor and its real ceiling. Understanding which users have found their glass slipper is the unlock.

What A16Z Found in 100 Trillion Tokens

The State of AI: An Empirical 100 Trillion Token Study, conducted by teams from a16z and OpenRouter, is the largest published analysis of real-world AI usage patterns. It covers actual user sessions — not surveys, not self-reported data — across frontier models from Anthropic, Google, OpenAI, Meta, and others, over a dataset that spans multiple model generations.

The retention finding emerged from cohort analysis: tracking which users were still active at Month 1, Month 3, and Month 5 for each major model launch. The researchers expected to see standard engagement curves — high initial activity followed by gradual decay as novelty wore off. What they found instead was a bimodal distribution.

Most user cohorts did follow the expected pattern: high Day 7 engagement, rapid decay over Months 1 to 3, and steady-state retention by Month 5 in the 15-25% range. But a subset of cohorts, concentrated among early adopters of specific model releases, showed dramatically higher retention: 40% or above at Month 5, with activity that was not decaying but stabilizing or growing.

The distinguishing characteristic of these high-retention cohorts was not demographic, geographic, or even broadly use-case-based. It was specificity: these users had found a precise workload that the model handled with a quality and reliability that was not replicated by alternatives they had tried.

Claude 4 Sonnet's May 2025 cohort retained approximately 40% of users at Month 5. Gemini 2.5 Pro's June 2025 launch cohort retained approximately 20% of users at Month 5 — which for a developer-focused frontier model is remarkably high, reflecting a population of developers who had found a specific technical workload (long-context code analysis, multimodal reasoning) where Gemini's capabilities were distinctly better than alternatives they had evaluated.

The Glass Slipper Effect Explained

A16Z named the phenomenon the "Cinderella Glass Slipper" effect after the specificity of the fit: when an AI model handles a particular user's specific unsolved workload with exactly the right combination of accuracy, speed, format, and domain knowledge, it creates a match that feels irreplaceable. The user stops shopping for alternatives. Evaluating every new model release becomes lower priority — not because the user is uninformed, but because switching costs are real when the current model handles the critical workload reliably.

The glass slipper metaphor extends in a specific direction: the shoe fits one foot exactly. Cinderella is not looking for a better shoe in the abstract. She is not interested in all possible shoes of that era. She needs the one that fits her foot precisely, and once she has it, the market for alternative shoes is largely irrelevant to her.

For an AI product team, the operational implication is uncomfortable: your product probably has a glass slipper cohort — a set of users for whom your model's specific capabilities match a specific unsolved workload precisely — and a much larger population of users for whom the fit is approximate, comfortable, or merely adequate. Approximate fits churn. Precise fits expand.

The challenge is that most AI product analytics cannot distinguish these populations from each other in real time. Aggregate DAU, MAU, and retention curves blend the two. The user who logs in daily to run the same legal contract analysis workflow and the user who logs in once a week to experiment with different AI writing assistants look identical in session-count data. They are completely different retention risks.

Workload-Model Fit: The Retention Driver You're Not Tracking

Workload-model fit is the degree of alignment between a user's specific, high-value, unsolved task and the particular capabilities of an AI model. It is more granular than product-market fit, which describes a product's fit with a broad market segment. Workload-model fit describes whether this particular user's daily work is something the model handles better than any alternative they have tried.

The a16z research identified a category of tasks they call "unsolved workloads": work that users needed to do repeatedly, that was high-value enough to invest time in finding an AI solution for, and that prior tools and models had handled insufficiently. Unsolved workloads are the workload equivalent of the princess's foot: they are waiting for the exact fit, and when they find it, they don't let go.

What makes a workload "unsolved" in this sense: - The user has tried multiple AI tools for the task and found all of them lacking in some specific dimension (accuracy on domain-specific terminology, format of output, handling of long-context inputs, consistent performance across edge cases) - The task recurs at a frequency that makes the quality gap a real daily friction — not a theoretical limitation but a felt cost - The user can articulate what the ideal solution would look like, because they have a clear mental model of the expected output

These users are actively searching for the glass slipper when they first try your product. The users with the lowest long-term retention are doing the opposite: they are experimenting broadly, often without a specific workload in mind, driven by curiosity or peer pressure rather than a felt need.

The Benchmark Numbers: What "Good" Looks Like Across Price Tiers

ChartMogul's 2026 SaaS Retention Report is the clearest cross-section of AI product retention data published this year, covering hundreds of AI-native and AI-adjacent SaaS companies. The price-tier breakdown reveals how dramatically workload-model fit varies across AI product positioning:

Price TierGross Revenue RetentionNet Revenue RetentionRetention Characteristics
$250+/month70% GRR85% NRRComparable to traditional B2B SaaS
$50–$249/month45% GRR~60% NRRHigh churn, moderate expansion
Under $50/month23% GRR~35% NRR>75% of revenue churns within 12 months
Enterprise ($100K+ ACV)95%+ GRR~118% NRRMatches or exceeds SaaS baseline

The divergence across tiers is not primarily explained by pricing strategy itself. It is explained by the selection effect that pricing creates. A product priced at $250 per month attracts users who have already identified a high-value use case and are willing to pay a meaningful amount for a solution. That self-selection filters for workload-fit users before they even sign up. A product priced at $9.99 per month attracts a far broader population of experimenters who have not yet identified a specific unsolved workload — and most of them will churn before they do.

The NRR data from usage-based pricing research reinforces the same dynamic: the products with the highest NRR are ones where retained users expand because the AI is handling more of their work over time — the direct consequence of workload-model fit deepening as the user's confidence in the model grows.

For AI companies, the practical benchmark from a16z's enterprise data is: - Gross dollar retention above $100K ACV: above 95%, target above 97% - Mid-market ($25K–$100K ACV): 85–90% GDR for Series A defensibility - Consumer/prosumer (under $50/month): 40%+ GRR by Month 6 is top-quartile performance

Foundational Cohorts vs. Casual Experimenters

The glass slipper framework creates a two-population model of AI product users that most teams are not currently tracking:

Foundational cohorts are users who achieved deep workload-model fit early in their relationship with the product. They are defined by: - High task frequency (multiple sessions per week, often daily) - Narrow use case depth (they use the product for a specific type of task, not broadly) - Low susceptibility to new model releases (they evaluate alternatives but consistently return to the current model for the specific workload) - Expansion over time (as they deepen fit, they find adjacent workloads and expand usage)

Casual experimenters are users who tried the product out of curiosity, peer pressure, or broad interest in AI without identifying a specific unsolved workload. They are defined by: - Irregular session cadence (high Day 1, declining Day 7, near-zero by Day 30) - Broad but shallow use (they try many different task types in early sessions, committing to none) - High susceptibility to new model releases (every GPT or Claude announcement triggers another evaluation) - No expansion (they never deepen enough to find adjacent use cases)

The problem for most AI product teams is that casual experimenters represent a large share of the user base in absolute numbers — particularly during the early growth phase, when broad awareness drives signups from people who are curious rather than needy. Their high churn rate distorts aggregate retention curves, creates misleading signals about product-market fit, and can cause teams to make product decisions based on feedback from users who were never going to retain long-term.

The AI tourist churn pattern — high acquisition, fast departure, 23% GRR for sub-$50 products — is the aggregate signal of a product population dominated by casual experimenters. The strategic response is not to optimize for casual experimenters (reduce friction, add breadth, lower the commitment required) but to find more foundational users by identifying the workloads where your model has a genuine capability advantage.

How to Find Your Glass Slipper Cohort

The glass slipper cohort is not always who you think it is. The users who look like your ICP on paper — the ones who match your demographic targeting, who signed up through your paid channels, who have the right job title — are not necessarily the users with the best workload-model fit. Finding the actual glass slipper cohort requires instrumentation at the task level.

1. Map session-level task types. Classify each user session by the type of task being performed — not by feature used, but by work outcome. Legal document analysis, code review, customer email drafting, data interpretation, and research synthesis are different workloads with different fit profiles. Most teams track feature usage; they need to track task type.

2. Identify zero-churn task segments. Within your user base, find the task type where M90 churn is lowest — where users who did that task in their first session are still active 90 days later. This is your highest workload-model fit segment. The model is handling something for these users that no alternative handles as well.

3. Interview the non-churners. Before optimizing for the zero-churn segment, interview the users in it. The goal is to understand the specific unsolved workload they were carrying before they found your product: what they had tried previously, what made those alternatives insufficient, and what the exact moment was when they knew your model was better. This is the workload description that goes into your positioning, activation flow, and acquisition targeting.

4. Rebuild the activation flow around the highest-fit workload. If your glass slipper cohort is legal document analysis, your first-session experience should route users who have that need directly to the task pattern that demonstrates the capability. Stop trying to show every user every feature in the first session. Show the user who has the glass slipper workload exactly what they came to find.

5. Sequence acquisition to reach high-fit users first. The channels and messaging that bring in workload-fit users are different from the channels that drive broad awareness. Identify which acquisition channels, referral sources, and search queries have the highest correlation with zero-churn task segments. Invest in those channels disproportionately — even if their raw volume is lower than awareness channels. A smaller number of glass slipper users is more valuable than a large number of casual experimenters.

6. Measure the glass slipper ratio. The ratio of foundational cohort users to casual experimenters in each acquisition cohort is a leading indicator of that cohort's long-term retention. Track it explicitly, alongside DAU and activation rate, in your retention dashboard.

The Expansion Mechanic: Why Retention Compounds Later

A16Z's retention analysis identified three distinct phases of the user lifecycle for AI products: acquisition (M0–M3), retention (M3–M9), and expansion (M9+). Most product analytics focuses on acquisition and early retention. The expansion phase is where the economics of workload-model fit compound.

A user who achieved glass slipper fit in Month 1 for a specific workload (contract analysis, say) does three things over the following 12 months that a casual experimenter does not:

They deepen on the primary workload. As confidence in the model grows, the user delegates more of the primary workload — longer documents, more complex analyses, higher-stakes outputs. This drives natural usage expansion without any product change or pricing action.

They discover adjacent workloads. Proximity to the primary workload reveals related tasks where the model can help. The legal team using the product for contract analysis discovers it handles regulatory research; the developer using it for code review discovers it handles architecture documentation. Each adjacent workload discovered is another activation moment with a high probability of becoming a durable use case.

They become advocates. Users who have found genuine workload-model fit are the best acquisition channel for more workload-fit users. Their referrals are not "this AI is cool, try it" — they are "this is how I use it for X, and if you do X, you should use it too." These referrals pre-screen for high-fit users because the referring user describes the specific workload context that generates the fit.

Intercom's Fin customer lifecycle data documents the same compounding dynamic in AI customer service: users who reached genuine resolution capability in their first 30 days (activation on the primary use case) expanded to 2.3x their initial usage volume by Month 12, while users who did not activate in the first 30 days had 85% churn by Month 3.

The expansion dynamic is why the a16z framing that "retention is all you need" is more than a retention optimization argument. It is an argument about the compounding economics of the AI product category: the companies that find their glass slipper cohort early and build acquisition, activation, and product development around it will compound on a retention base that casual-experimenter-heavy competitors cannot replicate.

The Activation Intervention: ChartMogul's One Fix

ChartMogul's analysis of AI product retention data identifies a single intervention with the highest return: fix activation. If you could fix only one thing about AI product churn, the data says to fix the first-session-to-first-value path — specifically, the gap between a user's first session and the moment they complete a task they would not have been able to complete as well without the AI.

This is not a standard product principle. Most SaaS activation frameworks focus on feature discovery, account completion, and "aha moment" milestones that are defined by the product's designed experience rather than by the user's specific work outcome. AI product activation is different because the "aha moment" is task-dependent: a user doing contract analysis needs to see the model handle a real contract before the activation moment occurs. A user doing code review needs to see a real code review output. The activation experience must be tailored to the workload, not to a generic product tour.

The average Month 1 retention across AI-native products in the Perspective AI 2026 benchmark is 46.9% — meaning more than half of users who signed up in Month 1 do not return in Month 2. For AI and ML companies specifically, Month 1 retention is 54.8%, the highest-performing vertical. The gap between 46.9% and 54.8% is explained by activation quality: AI-native products that get users to a genuine first-value moment in their first session retain at substantially higher rates.

The glass slipper framework and the activation data point to the same intervention: identify the workload where your model has a genuine capability advantage, and build the first-session experience around delivering a demonstrable first-value moment for users who carry that workload. Not every user will find their glass slipper in Session 1. But the users who do are your retention foundation.

Takeaway: A16Z's glass slipper finding is a reframe of AI product strategy: retention comes from workload-model fit, not model quality, and the earliest cohorts are the clearest signal of whether that fit exists. The benchmark that matters is not whether your model is the best in the MMLU leaderboard — it's whether the users who found a specific, high-value, unsolved workload in your product are still there at Month 5. If they are, you have your glass slipper. If they're not, neither model upgrades nor acquisition growth will fix the underlying fit problem.

Frequently Asked Questions

What is the glass slipper effect in AI product retention?

The glass slipper effect, named by a16z in their analysis of 100 trillion real-world AI tokens conducted with OpenRouter, describes a pattern where the earliest cohorts of an AI product retain substantially better than later cohorts — not because the product changed, but because those early users self-selected for a precise fit between their specific unsolved workload and the model's capabilities. Like Cinderella's glass slipper, the fit is exact and irreplaceable: when a user finds an AI model that handles their particular type of work with accuracy and reliability that other models cannot match, they stop experimenting and settle. That 'settling' is what creates durable retention. The effect has been observed across multiple frontier AI products: Claude 4 Sonnet's May 2025 cohort retained approximately 40% of users at Month 5, substantially higher than cohorts that joined after the model became widely known and attracted more casual users. The practical implication for product teams is that early retention signals are disproportionately informative: an early cohort with strong M3 or M5 retention is evidence of real workload-model fit, while early cohorts with weak retention indicate the product has not yet found the user whose work it solves precisely.

What is workload-model fit and how do you find it?

Workload-model fit is the degree of alignment between a user's specific, high-value, unsolved task — what a16z calls an 'unsolved workload' — and the particular capabilities of an AI model. Unlike product-market fit, which describes a product's fit with a broad market segment, workload-model fit is granular: it describes whether this user's specific daily task (contract analysis, code review at a particular language and complexity level, customer inquiry handling in a regulated industry) is one where the model performs distinctly better than alternatives. Finding workload-model fit requires instrumentation at the session and task level rather than the cohort level. Teams that track which task types generate repeat sessions within 48 hours, which prompt patterns correlate with expansion usage, and which use cases have zero churn in the first 90 days are building the dataset needed to identify their highest-fit workloads. Once identified, those workloads should anchor acquisition, onboarding, and activation messaging: the goal is to ensure users with the matching unsolved workload find the product before users who are casually experimenting with something they don't specifically need.

What are the AI product retention benchmarks by pricing tier?

ChartMogul's 2026 SaaS Retention Report found that AI-native products show dramatically different gross revenue retention (GRR) depending on price point. AI products priced above $250 per month retain at 70% GRR with 85% net revenue retention (NRR) — comparable to traditional B2B SaaS benchmarks. Products in the $50 to $249 per month tier retain at 45% GRR. Products under $50 per month retain at 23% GRR, with more than three-quarters of revenue churning within the first year. The divergence reflects a selection effect: higher price points attract users who have identified a specific, high-value use case for the product — users who have, in the glass slipper framing, already found their workload-model fit before purchase. Lower price points attract a broader population of experimenters and curiosity buyers who have not yet found a workload the model solves precisely. For enterprise AI (above $100K ACV), a16z's benchmark for gross dollar retention is above 95%, often above 97%, with median NRR near 118%. These benchmarks are the practical targets against which AI product teams should measure their own retention curves.

Why do early cohorts retain better than later cohorts for AI products?

Early cohorts for successful AI products tend to retain better than later cohorts for two compounding reasons. First, early adopters of a specific AI capability self-select for the workload it solves best: they discovered the product through an information path that correlates with the specific use case it handles, and their use pattern is shaped by a genuine, high-frequency need. Second, as an AI product grows, marketing, press coverage, and word-of-mouth reach progressively broader audiences — including users whose primary motivation is curiosity, experimentation, or FOMO rather than a specific unsolved workload. These later users are less likely to find a durable fit, more likely to churn within 90 days, and contribute to declining cohort averages without reflecting a change in the product's core quality. A16Z documented this pattern in Claude 4 Sonnet and Gemini 2.5 Pro cohort data: June 2025 cohorts (close to launch) retained 40% and 20% of users respectively at Month 5, while later cohorts showed markedly lower retention despite no degradation in model capability. The implication is that cohort-level retention analysis needs to control for acquisition channel and use case — an aggregate retention curve that mixes workload-fit users with casual experimenters obscures both the product's real floor and its real ceiling.

What does 'retention is all you need' mean for AI product strategy?

The a16z framing that 'retention is all you need' for AI products is a corrective to the attention placed on acquisition metrics — active users, downloads, and growth rate — in the early AI product market. The underlying argument is that AI product economics are fundamentally retention-first: the cost of customer acquisition is high, the cost of serving retained users is declining as inference prices fall, and the expansion revenue from retained users (who deepen usage as the product proves value) compounds in ways that acquisition cannot replicate. ChartMogul's finding that the activation moment is the single highest-leverage intervention for AI product retention reinforces this: getting a user to their first genuine value moment — their first session where the AI handles a real work task to a standard they couldn't achieve alone — is the retention intervention with the highest return on investment. Products that optimize acquisition while under-investing in the activation-to-first-value path accumulate churn at a rate that compounds against them. The strategic implication is to measure and optimize for three metrics in order: activation rate (did the user reach first value?), M3 retention (did they return after initial experiments?), and NRR (did retained users expand?). Growth rate is a lagging indicator of all three.