The Ninety
A ninety-day route with exactly three components: one capability moved one tier, one artifact that will exist at the end, one person who will judge it. Budgeted against your measured overrun ratio, and carrying its abandonment conditions in writing from the first day.
A protocol cannot ship without this file. The build checks that it exists, that it rates the strength of its evidence, that it carries citations, and that it contains a section stating what the protocol does not claim.
Where the support is weak, it says weak. Where a decision was invented rather than derived, it says that too.
Every design decision with the strength of its support. Where support is weak, this file says so.
- strong — replicated, meta-analytic support, broad consensus.
- moderate — consistent findings, limited replication or contested effect sizes.
- weak — plausible, thinly supported, or extrapolated.
1. One capability, not several
Strength: moderate.
Splitting effort across multiple goals reduces progress on each, and the cost is not merely arithmetic — competing goals produce avoidance of both, particularly when they compete for the same scarce hours.
- Emmons, R. A., & King, L. A. (1988). Conflict among personal strivings. Journal of Personality and Social Psychology, 54(6), 1040–1048.
- Locke, E. A., & Latham, G. P. (2002). Building a practically useful theory of goal setting and task motivation. American Psychologist, 57(9), 705–717.
Honest limit. These findings concern conflict among goals generally. That exactly one is the right number for a ninety-day route is a design choice, not a result. Two might work for some people. The protocol enforces one because the failure it prevents is common and the cost of the constraint is low.
2. The route terminates in an artifact, not in a state of knowledge
Strength: strong for the underlying effect, moderate for this application.
Producing something is a more durable route to retention and transfer than studying is. Generating an answer beats reading one; being tested on material beats re-reading it, by a wide and well-replicated margin.
- Roediger, H. L., & Karpicke, J. D. (2006). Test-enhanced learning. Psychological Science, 17(3), 249–255.
- Karpicke, J. D., & Blunt, J. R. (2011). Retrieval practice produces more learning than elaborative studying with concept mapping. Science, 331(6018), 772–775.
- Slamecka, N. J., & Graf, P. (1978). The generation effect. Journal of Experimental Psychology: Human Learning and Memory, 4(6), 592–604.
Honest limit. This literature concerns memory for material in controlled settings, over days and weeks. Extrapolating it to "build a real thing over ninety days" is a reasonable extension of the principle and not a tested claim.
3. External judgment is required to reach the top tier
Strength: moderate to strong.
Improvement in a skill depends on feedback that is informative and reasonably prompt. Practice without it produces experience, not competence — and long experience without feedback is one of the better-documented ways to plateau while feeling like an expert.
- Ericsson, K. A., Krampe, R. T., & Tesch-Römer, C. (1993). The role of deliberate practice in the acquisition of expert performance. Psychological Review, 100(3), 363–406.
- Kahneman, D., & Klein, G. (2009). Conditions for intuitive expertise. American Psychologist, 64(6), 515–526.
- Shute, V. J. (2008). Focus on formative feedback. Review of Educational Research, 78(1), 153–189.
An honest correction to the popular version. The deliberate-practice framework is often cited as though practice accounts for most of the variance in performance. A large meta-analysis found it accounts for considerably less — around 12% across domains, and far less in professions than in games or music.
- Macnamara, B. N., Hambrick, D. Z., & Oswald, F. L. (2014). Deliberate practice and performance. Psychological Science, 25(8), 1608–1618.
This protocol therefore claims only that feedback is necessary for improvement. It does not claim that structured practice is sufficient, and it makes no prediction about how good anyone will get.
4. The proof names a person, not an audience
Strength: weak. This is a design choice.
No literature supports naming a specific individual over an unspecified audience. It is adopted because "the community will judge it" is unfalsifiable, undatable, and in practice means no one judged it. Requiring a name converts an intention into something schedulable — which connects it to the one strongly supported element below.
5. Checkpoints carry specific hours
Strength: strong.
Specifying when and where an intention will be acted on substantially raises follow-through relative to intention alone. This is the best-supported element in the protocol.
- Gollwitzer, P. M., & Sheeran, P. (2006). Implementation intentions and goal achievement: A meta-analysis of effects and processes. Advances in Experimental Social Psychology, 38, 69–119. (94 studies, medium-to-large effect.)
6. The budget uses a measured overrun ratio, not a fresh estimate
Strength: strong.
See P-002 evidence, sections 1 and 2. Estimates of one's own task duration are systematically low; the correction that works is grounding the estimate in what comparable past work actually cost rather than in a mental simulation of the current work.
- Buehler, R., Griffin, D., & Ross, M. (1994). Journal of Personality and Social Psychology, 67(3), 366–381.
- Flyvbjerg, B. (2006). Project Management Journal, 37(3), 5–15.
Honest limit. The personal-scale version — one person's last six weeks as the reference class — is a much smaller and noisier sample than the method was designed for. It is better than an unadjusted estimate. It is not reference class forecasting.
7. Kill conditions are written on day zero
Strength: moderate for the mechanism, weak for this implementation.
People persist with failing courses of action in proportion to what they have already invested, even when the investment is irrecoverable and irrelevant to the decision ahead. Prior personal responsibility for the initial commitment makes escalation worse.
- Staw, B. M. (1976). Knee-deep in the big muddy: A study of escalating commitment to a chosen course of action. Organizational Behavior and Human Performance, 16(1), 27–44.
- Arkes, H. R., & Blumer, C. (1985). The psychology of sunk cost. Organizational Behavior and Human Decision Processes, 35(1), 124–140.
- Wrosch, C., Scheier, M. F., Miller, G. E., Schulz, R., & Carver, C. S. (2003). Adaptive self-regulation of unattainable goals. Personality and Social Psychology Bulletin, 29(12), 1494–1508. (Disengaging from unreachable goals is associated with better wellbeing.)
The prospective-hindsight technique — imagining the failure in advance and enumerating its causes — improves the identification of plausible failure modes.
- Klein, G. (2007). Performing a project premortem. Harvard Business Review, 85(9), 18–19.
- Mitchell, D. J., Russo, J. E., & Pennington, N. (1989). Back to the future: Temporal perspective in the explanation of events. Journal of Behavioral Decision Making, 2(1), 25–38.
Honest limit, and it is a real one. That precommitting to falsifiable abandonment criteria actually changes behaviour ninety days later is not demonstrated. The mechanism is well-supported; this application of it is untested. The "not revisable after day thirty" rule is a design invention with no evidence behind it at all — it exists because a condition editable at the moment it triggers provides no constraint.
8. Ninety days
Strength: none. It is a convention.
Nothing in the literature identifies ninety days as an optimal interval for anything. It is used because it is long enough to produce an inspectable artifact and short enough to hold a single selection stable, and because it aligns with the quarterly refresh of the Self Map.
Do not read significance into the number.
What this protocol does not claim
- It does not claim the route is achievable. The agent is explicitly instructed not to say so, because it cannot know.
- It does not predict how skilled anyone will become. Practice explains a smaller share of performance variance than the popular account suggests (see section 3).
- It does not claim that one tier per ninety days is a realistic rate for all capabilities. For some it is generous, for others impossible, and the protocol cannot tell you which.
- It is not a career instrument and does not model labour markets. It says nothing about whether the chosen capability is worth having.
- The protocol as a whole has not been tested.
Open questions
- Do written kill conditions get invoked, or do they get argued around?
- Does the day-30 lock hold, or does it simply move the rationalisation earlier?
- Is one capability per ninety days the right rate, and does it vary predictably by tier transition?
- Does naming a specific proof-person raise completion, or does it raise avoidance of the route altogether?