Skip to content

Latest commit

 

History

History
60 lines (39 loc) · 5.44 KB

File metadata and controls

60 lines (39 loc) · 5.44 KB

06 · Measurement — reading tiny numbers without fooling yourself

Goal of this doc: an early-stage Reddit engine produces numbers so small that almost any story fits them. These rules are what stop you telling yourself one.

The one question the report must answer

Does this channel produce pipeline, or only audience?

Put it in every weekly report, answered with numbers:

  1. How many identifiable people interacted this period? (commenters and repliers — not reach figures)
  2. How many of them match your signer ICP? Count people who can sign separately from people who can't. Don't sum them.
  3. How many entered the CRM?
  4. Which pond did each one come from? This is what tells you whether your portfolio (01-icp-and-ponds.md) is working.

If (3) is zero week after week while (2) isn't, say it explicitly. That is the difference between an engine that will matter someday and one that serves this quarter's goal. A measurement routine that reports rising karma while pipeline is flat is not neutral — it is actively misleading the people reading it.

The traps, each paid for

  • Karma is a slow indicator. No strategy is declared a win or a loss before 7-15 days. Label things "in test (day X of N)".
  • It is not monotonic. Old comments lose and regain points on their own. A snapshot is a balance, not a total of everything you earned.
  • No individual score below ~5 points is stable data. Judge with totals plus a count of interventions with net positives, never the exact value of one comment.
  • Per-sub breakdowns are read in weekly windows, never daily. With 1-2 point movements, attribution inverts in 24 hours.
  • Score without ratio says nothing. 0 points at 17% and 0 points at 100% are opposite phenomena. Record both, always.
  • Reach ≠ karma. A post can do ~25× the views of your best comment and earn no karma at all. Posts produce reach, posting permission and conversation. The one thing they don't produce is karma. Comments are the karma engine. They're different products — don't compare them on one axis.
  • Reconciliation is a measurement, not a property. If your item-by-item inventory matches the profile's karma for five days, that is exactly when you stop checking — and exactly when it stops matching. In one real engine it matched for five days, diverged by 2, narrowed to 1, then matched again two days later. It was never "healthy" or "broken": it oscillates, the same way individual comment scores do. So never write down "the scoreboard reconciles exactly" as a fact about your system. Record both numbers every week and read the series. When they disagree, say you don't know which is right rather than picking the flattering one.
  • When the confound is in the baseline, more time doesn't fix it. Declare it and decide on other grounds.
  • A sweep with an invented pattern measures your hypothesis, not the fact. Before concluding from a search: open two files and see how the thing is actually written. Every zero from a search is a non-match, not an absence. And if a sweep says several things are broken at once, the suspect is the sweep.

The research bank

The scoreboard records what you did. The lessons file records how you write. Neither records what they say, which is the asset that doesn't expire.

Keep a separate file with three things and only three:

  1. Verbatim quotes with handle, sub and date — their vocabulary, not yours.
  2. Recurring problems in their words, with a counter. Set a promotion rule (e.g. 3+ threads across 2+ subs) that moves a problem into your master ICP document with its evidence.
  3. Content ideas that come out of it, tagged by channel.

Actively hunt for the counter-signal. The most valuable entry in a research bank is the one that contradicts your positioning — for example, discovering that the audience solves the problem you sell against in a cheaper way and doesn't feel the cost you're pricing. That doesn't change the product; it changes the copy, and it stops you writing for a pain nobody currently feels.

Closing the loop

Weekly measurement → distilled lessons → the file the drafter and radar read → next week's hunting priorities.

⚠️ Verify the wiring, not just the intent. In one real engine, the measurement routine was instructed to write lessons into a named section, and the radar routine to read that section to prioritize. The section did not exist in either direction. Nothing errored. The loop was documented for weeks and had never once closed.

When a routine names a section, open it and confirm it exists. A search for filenames will not catch this.

Document hygiene, because it is a measurement problem

Living documents grow until they stop being read, and a document nobody can read governs nothing:

  • Ceilings, enforced. A rules file that outgrows a single read has to be split or distilled.
  • The scoreboard keeps a short window (e.g. 3 days) and archives the rest.
  • A lesson that becomes a rule gets deleted from the log. The log is an antechamber, not a warehouse.
  • The liveness test: if a rule isn't also in a checklist or in the routine's own prompt, it governs nothing. Writing it only in the strategy file is theatre.
  • Run an alignment sweep on a schedule, with a date and an owner. See ../checklists/monthly-alignment.md.

An anti-slop document that grows until nobody can read it is slop.