id: P-002
slug: weekly-reckoning
version: 0.1.0
title: The Weekly Reckoning
layer: control
agent: The Mirror
cadence: weekly
duration_minutes: 25
difficulty: moderate
status: draft
updated: 2026-08-02
license: CC-BY-SA-4.0
summary: A weekly loop that scores your predictions against what happened. You
  commit with numbers attached, the numbers get checked, and the gap becomes a
  measured quantity instead of a vague sense that things take longer than
  expected.
cost:
  time: 25 minutes, same slot every week
  discomfort: moderate, rising in weeks three to six
  prerequisites:
    - A completed Self Map (P-001), or at least its constraints section
    - Last week's Reckoning, from week two onward
produces: reckoning
requires:
  - P-001
feeds:
  - P-003
sections:
  - id: close
    title: Close the Loop
    minutes: 8
    purpose: Compare last week's predictions to what actually happened. No
      interpretation yet.
  - id: score
    title: The Score
    minutes: 4
    purpose: Compute calibration. One number for time, one for completion, one for
      confidence.
  - id: pattern
    title: The Pattern
    minutes: 6
    purpose: Name what the accumulated record shows, including the thing being avoided.
  - id: commit
    title: The Commitments
    minutes: 7
    purpose: Set next week's work, each item carrying a prediction that will be scored.
metrics:
  - id: time_ratio
    name: Time ratio
    definition: actual minutes divided by predicted minutes, per item
    target: converging toward 1.0; the direction matters more than the value
  - id: completion_rate
    name: Completion rate
    definition: committed items finished divided by items committed
    target: 0.7 to 0.9; a sustained 1.0 means the commitments are too small
  - id: confidence_error
    name: Confidence error
    definition: mean of (stated confidence minus realised outcome) across items
    target: toward 0; positive means overconfident, negative means underconfident
adversary_mode:
  supported: true
  effect: The Mirror stops accepting circumstantial explanations for misses. Every
    explanation is checked against the record for recurrence, and a reason that
    has appeared three times is reclassified from circumstance to pattern.
honesty:
  - The scoring instruments here are simplified. They are not the calibration
    measures used in forecasting research, and no claim is made that they are
    equivalent.
  - Weeks one and two produce no useful signal. The protocol needs roughly six
    weeks of record before the numbers mean anything.
  - Self-scored outcomes can be gamed. Nothing prevents you from grading
    yourself generously except the fact that it makes the instrument useless.
