Podcast answers

Jeff Dean

Jeff Dean's 1% Rule for Finding Durable AI Startup Ideas

What did Jeff Dean say about the 1% rule for building in AI on the Y Combinator Startup Podcast?

1 episode1 show45 citations
Shows checked
Y Combinator Startup Podcast
Evidence reviewed
1 August 2026 to 1 August 2026
Topics covered
AI startups, Startup ideas, Defensibility, Artificial intelligence
Last checked

Answer in brief

Jeff Dean’s 1% rule is a filter for choosing durable AI startup problems. Founders should begin with a problem they genuinely care about and that could create real-world value, then benchmark current general-purpose models against it. The attractive territory is where those models succeed only about 0-1% of the time. If they already succeed around 20%, Dean’s view is that the capability has begun to emerge and may improve quickly, making a thin product built around today’s limitation vulnerable. He identified two stronger positions: products that can use important private or user-specific data unavailable to general models, and specialized models trained for difficult domains where narrow expertise can deliver substantially higher accuracy at an affordable cost. 27:2528:0028:3529:1029:10

What the 1% rule means

The rule is fundamentally about selecting problems, not about squeezing one more percentage point from an existing model. Dean’s proposed sequence starts before model choice: identify something worth building, driven by strong founder motivation and a belief that solving it would matter outside the laboratory. Only then should the founder examine how capable current general models are in that exact domain. This makes the rule a conjunction of mission and technical opportunity. A problem is not attractive merely because models currently fail at it; it should also be a problem the team is committed to solving and one whose solution would produce meaningful value. 27:2528:00

The 0-1% success range represents a large capability gap. In such a domain, a startup is not merely polishing an already functional general model. It has room to develop data, training methods, domain knowledge, evaluation systems, or product context that materially changes what is possible. By contrast, a model that already completes roughly one task in five is demonstrating partial competence. Dean interprets that foothold as evidence that the underlying capability is emerging and could improve rapidly as general models advance. 28:35

The practical warning is against mistaking a temporary model weakness for a durable company advantage. If a startup’s entire proposition is that a general model fails 80% of the time today, an improvement in the general model could remove much of the startup’s differentiation. Dean therefore favors domains where the starting gap is much wider. The rule directs founders toward tasks for which successful execution still requires assets or capabilities that broad models do not presently possess. 28:3529:1029:10

The percentages should be read as a strategic heuristic rather than a formal mathematical boundary. The supplied evidence explains why Dean contrasts near-total failure with 20% success, but it does not establish that 1% is universally safe or that 20% inevitably predicts imminent commoditization. It also does not specify the benchmark, sample size, scoring method, or time horizon founders should use. The durable idea is the distinction between an untouched capability frontier and one that general models have already begun to cross. 28:0028:35

How founders would apply it

The first step is problem selection. Dean puts founder excitement and intended real-world value ahead of technical novelty. That order matters: the search should not begin with a fashionable model and then look for somewhere to deploy it. It begins with a consequential problem that can sustain the team’s attention. The model benchmark is then used to determine whether current general systems leave enough unsolved territory for a focused company to create distinctive value. 27:2528:00

The second step is direct testing in the target domain. A founder should not infer capability from broad benchmark scores or from anecdotes about what a model can do elsewhere. Dean’s formulation calls for examining what current general models can actually accomplish on the problem under consideration. For example, a team investigating a hard chip-design task would need to evaluate general models on that task, rather than treating general coding performance as a sufficient proxy. The chip-design example is named as a domain that may reward specialized accuracy precisely because broad models can remain weak on its particular problems. 28:0030:55

The third step is to interpret partial performance strategically. Near-zero success suggests that a startup may have space to build a genuinely different capability. Performance around 20% is more dangerous because it indicates that the general model has already found some purchase on the task. Dean expects such emerging capabilities to improve quickly, so founders should not assume that today’s remaining 80% failure rate is a defensible moat. The relevant comparison is therefore not simply startup quality versus present model quality, but startup differentiation versus the likely direction of general-model progress. 28:35

The final step is to ask why the startup can cross the gap when a general model cannot. Dean’s examples supply two credible answers: the product can observe important information unavailable to the broad model, or the company can train a narrow system whose domain accuracy is far higher. Without one of those structural advantages, a low initial benchmark score identifies an open problem but does not by itself explain why this particular startup will solve it. 29:1029:10

Two defensible paths

The first path is privileged context. A product may have legitimate access to private, proprietary, or user-specific information that a general model cannot see. That visibility can make the application more useful even as the underlying general model improves, because the application’s advantage is not based solely on better generic reasoning. It comes from combining model capability with information specific to the user or operating environment. In Dean’s framing, the inaccessible data is important to the task, so access changes the quality or relevance of the result rather than merely adding superficial personalization. 29:10

This data path also clarifies what the 1% rule is not. It is not a demand that every startup train a frontier model from scratch. A company can potentially build on general models while differentiating through the information its product is permitted to use. The supplied evidence does not say that data access automatically creates defensibility, nor does it address consent, security, data quality, or competitors acquiring similar access. It establishes only that valuable unseen data is one promising reason a product could succeed where a context-blind general model does not. 29:10

The second path is specialization. In a sufficiently hard domain, focused training data may support a niche model that is much more accurate than a general-purpose alternative while remaining affordable to operate. The strategic asset here is not merely possession of more data, but the alignment of domain-specific data, model design, and a narrowly defined objective. A startup does not need to reproduce every ability of a broad model if it can perform one valuable task with decisively better reliability. 29:10

Affordability is part of this path, not an afterthought. Dean’s claim concerns a specialized model that combines higher domain accuracy with economically practical deployment. A narrow system that performs exceptionally but cannot be provided at a viable cost would not satisfy the opportunity as described. Conversely, specialization may allow resources to be concentrated on the task that customers actually need, rather than on the full range of general-model capabilities. The evidence supports this as a promising pattern, though it gives no numerical target for cost or accuracy. 29:10

Why AlphaFold and technical domains matter

Dean uses AlphaFold to illustrate the strongest version of the specialization argument. Its significance in this context is architectural and strategic: a system can be designed around a narrow, important scientific objective and solve that objective exceptionally well without becoming a general-purpose intelligence. The example shows why founders should not assume that the broadest model will necessarily dominate every valuable problem. For some tasks, the correct abstraction is a dedicated scientific or technical system rather than a general assistant with a domain wrapper. 30:20

AlphaFold also makes the 1% rule more ambitious than a search for overlooked software features. The opportunity can be a difficult problem whose solution requires purpose-built modeling and specialized evidence. That raises the potential value of the result, but also implies a higher technical bar than simple prompting or interface design. The evidence identifies AlphaFold as an illustration of what narrow excellence can achieve; it does not claim that every startup can reproduce its research trajectory, economics, or scientific impact. 29:1030:20

Materials science and chip design are offered as further candidate areas. Their relevance is not that Dean reports a completed startup formula for either field, but that they contain precise, difficult tasks on which general models may perform poorly and unusually accurate niche systems may be valuable. These examples extend the principle beyond protein structure: important technical domains can contain subproblems where specialized models have a better fit than broad ones. 30:55

The operative word is may. The evidence describes materials science and chip design as promising possibilities, not validated guarantees. A founder would still need to choose a specific high-value task, measure the general model’s baseline, establish access to appropriate training data, and demonstrate materially higher accuracy at workable cost. The episode evidence does not provide those empirical results for any particular materials or chip-design product. 27:2528:0029:1030:55

Limits and strategic implications

Dean’s framework is strongest as an early screening device. It links founder motivation, customer or societal value, present model capability, and the source of future differentiation. It helps reject a weak thesis in which a startup merely packages a limitation that general models are already overcoming. It also points founders toward a positive thesis: solve a valuable problem through information unavailable to general models or through specialized performance that broad systems do not economically match. 27:2528:3529:1029:10

However, the episode evidence does not establish that all attractive AI companies must begin in a 0-1% domain. It does not evaluate opportunities based on workflow integration, distribution, service quality, regulation, or execution, except insofar as private context and specialized accuracy create value. Nor does it show that a 20%-capable general model eliminates every possible business in that domain. The conservative conclusion is narrower: Dean regards visible partial competence as a warning that the technical gap may close quickly. 28:3529:1029:10

The evidence also leaves measurement unresolved. Success could mean exact correctness, useful assistance, completion without human intervention, or performance above a domain threshold; those definitions would produce different percentages. A sensible application of Dean’s rule therefore requires a task-specific evaluation that reflects the product’s real objective. This is an inference from his instruction to test general models in the chosen problem domain, not a benchmarking protocol supplied by the episode. 28:0028:35

Taken together, Dean’s advice is to build where commitment and asymmetry coincide. The founder should care enough about the problem to pursue it deeply, while the company should possess a credible reason to improve from near-zero general-model performance to dependable usefulness. Private context and specialized training are the two reasons identified in the evidence. The 1% result is the signal that room may exist; the actual company is built by explaining and proving how its particular assets can fill that room before general models do. 27:2528:3529:1029:10

Sources

This independent summary is for general information and is not endorsed by the people or shows it covers. Check important points at the linked source. Podcast rights remain with their owners. Read the methodology. Report a rights or accuracy concern.

What should the next Tldr answer?

Start with your question. You can review it before choosing a plan or creating anything.

See plans