The Weekly Reckoning
A weekly loop that scores your predictions against what happened. You commit with numbers attached, the numbers get checked, and the gap becomes a measured quantity instead of a vague sense that things take longer than expected.
A protocol cannot ship without this file. The build checks that it exists, that it rates the strength of its evidence, that it carries citations, and that it contains a section stating what the protocol does not claim.
Where the support is weak, it says weak. Where a decision was invented rather than derived, it says that too.
Every design decision with the strength of its support. Where support is weak, this file says so.
- strong — replicated, meta-analytic support, broad consensus.
- moderate — consistent findings, limited replication or contested effect sizes.
- weak — plausible, thinly supported, or extrapolated. Included because the failure mode is common and the cost is low.
1. Predictions are recorded before outcomes, and scored afterwards
Strength: strong.
People underestimate how long their own tasks will take, and do so even when they know they have underestimated before. The effect survives experience, incentives, and explicit warning. It is specific to one's own tasks; estimates for other people's work are less biased.
- Buehler, R., Griffin, D., & Ross, M. (1994). Exploring the "planning fallacy". Journal of Personality and Social Psychology, 67(3), 366–381.
- Kahneman, D., & Tversky, A. (1979). Intuitive prediction: Biases and corrective procedures. TIMS Studies in Management Science, 12, 313–327.
- Roy, M. M., Christenfeld, N. J. S., & McKenzie, C. R. M. (2005). Underestimating the duration of future events: Memory incorrectly used or memory bias? Psychological Bulletin, 131(5), 738–756.
The design consequence: a general warning does not help, so the protocol does not issue one. It builds the subject's own overrun ratio instead.
2. The correction is the outside view, using the subject's own record
Strength: strong for the principle, moderate for this implementation.
Grounding an estimate in what comparable past cases actually cost — rather than in a mental simulation of the current one — reliably reduces overrun. Reference class forecasting is the formalised version and has been applied at scale in infrastructure planning.
- Flyvbjerg, B. (2006). From Nobel Prize to project management: Getting risks right. Project Management Journal, 37(3), 5–15.
- Lovallo, D., & Kahneman, D. (2003). Delusions of success: How optimism undermines executives' decisions. Harvard Business Review, 81(7), 56–63.
Honest limit. Reference class forecasting uses a class of comparable projects, usually across organisations. This protocol uses a single person's recent history, which is a much smaller and noisier sample. Multiplying committed minutes by a personal median ratio is a reasonable analogue of the method, not the method. No published validation exists for the personal-scale version.
3. Feedback on calibration is given repeatedly, with the outcome always shown
Strength: moderate to strong, with an important restriction.
Calibration improves with practice when — and largely only when — predictions are numerous, outcomes are prompt and unambiguous, and feedback is repeated. Domains satisfying these conditions (weather forecasting, competitive bridge) produce well-calibrated experts; domains that do not, generally do not.
- Murphy, A. H., & Winkler, R. L. (1984). Probability forecasting in meteorology. Journal of the American Statistical Association, 79(387), 489–500.
- Kahneman, D., & Klein, G. (2009). Conditions for intuitive expertise: A failure to disagree. American Psychologist, 64(6), 515–526.
- Mellers, B., et al. (2014). Psychological strategies for winning a geopolitical forecasting tournament. Psychological Science, 25(5), 1106–1115. (Training and tracking measurably improved accuracy.)
The restriction matters and is stated to the subject. Weekly commitments are far fewer and far noisier than daily weather forecasts. Roughly six weeks of record are needed before the numbers carry signal. Weeks one and two produce none, and the protocol says so rather than showing a number that looks meaningful and is not.
4. Commitments carry a specific time and place
Strength: strong.
Specifying when, where, and how an intention will be acted on substantially raises follow- through relative to goal intention alone. The meta-analysis covers 94 studies with a medium-to-large effect.
- Gollwitzer, P. M. (1999). Implementation intentions: Strong effects of simple plans. American Psychologist, 54(7), 493–503.
- Gollwitzer, P. M., & Sheeran, P. (2006). Implementation intentions and goal achievement: A meta-analysis of effects and processes. Advances in Experimental Social Psychology, 38, 69–119.
This is why when is a required field and why a commitment without it is rejected. It is
the single best-supported element in the protocol.
5. Recording happens before interpretation
Strength: moderate.
Knowing an outcome distorts recall of what was predicted beforehand, and the distortion is not corrected by being warned about it. Locking the record before discussing causes is a structural guard rather than a cognitive one.
- Fischhoff, B. (1975). Hindsight ≠ foresight: The effect of outcome knowledge on judgment under uncertainty. Journal of Experimental Psychology: Human Perception and Performance, 1(3), 288–299.
- Roese, N. J., & Vohs, K. D. (2012). Hindsight bias. Perspectives on Psychological Science, 7(5), 411–426.
Honest limit. These findings concern recall of prior judgments. The protocol's ordering rule — record everything before explaining anything — is an application of the principle, not a tested intervention.
6. Explanations that repeat are reclassified as structure
Strength: weak.
The three-occurrence threshold is arbitrary. Nothing in the literature specifies it. It exists because some threshold is needed to stop every miss being absorbed as bad luck, and because three is small enough to trigger within a useful timeframe and large enough to survive genuine coincidence.
The adjacent support is the self-serving attributional pattern: failures are attributed outward to circumstance more readily than successes are.
- Mezulis, A. H., Abramson, L. Y., Hyde, J. S., & Hankin, B. L. (2004). Is there a universal positivity bias in attributions? Psychological Bulletin, 130(5), 711–747.
Treat the threshold as a working default. If running this at scale shows a different number separates signal from noise, change it.
7. Partial counts as zero
Strength: none. This is a design choice.
There is no evidence that harsh scoring produces better calibration than proportional
scoring. It is chosen because partial is where self-serving grading collects, and because
a binary outcome makes the confidence score computable without inventing a scale.
The cost is real: the subject who finished 90% of a large item is scored identically to one who never started. The protocol states this openly rather than pretending the measure is fair. If it proves discouraging enough to drive dropout, it should change.
8. At most three patterns per session
Strength: weak, by analogy.
The general finding that recommendations dilute as they multiply is well-attested in practice and poorly quantified in the literature. The specific number three is a judgment. The frequently cited "seven plus or minus two" is about short-term memory span and does not support this choice; it is not invoked here.
What this protocol does not claim
- It does not measure the value or quality of the work. Only the accuracy of the forecast. A perfectly calibrated person can be predicting trivial things.
- It does not use a proper scoring rule. The confidence error computed here is a simple mean signed error, not a Brier score, and is not comparable to calibration measures in the forecasting literature.
- It does not prevent gaming. Outcomes are self-scored. A subject who grades generously produces clean-looking numbers and a useless instrument.
- It is not a productivity measure and should not be used by anyone to evaluate anyone else. Used as a management instrument it would be actively harmful — the incentive to pad predictions would destroy the only thing it measures.
- The protocol as a whole has not been tested.
Open questions
- Does the median time ratio stabilise, and after how many weeks?
- Does the three-occurrence reclassification threshold separate signal from noise, or is it just a number that felt right?
- Does harsh scoring of
partialimprove honesty or increase dropout? - Does adversary mode improve accuracy, or does it drive people out in week four?