The publicly accountable sentiment engine

Every alt-data vendor claims their signal works. We are, to our knowledge, the only one that pre-registers its calls in public — hash-stamped before the outcome is known, graded honestly after, never deleted. Hits and misses both stay on this page forever.

3 entries: 1 open · 0 hit · 1 mixed · 1 miss

The rules of this ledger

  1. Pre-registered: every entry is published before the outcome is known — before the company reports, or before the resolution date for entries not tied to a print — with the publication date and a SHA-256 hash of its immutable fields shown below it.
  2. Append-only: facts and predictions are never edited after publication. The only thing ever added is the post-print grade.
  3. Graded against stated rules: each entry defines upfront how it will be judged. No moving the goalposts.
  4. Honest scope: RSAS measures customer-review sentiment. Our validated claim is 14-day sentiment persistence (walk-forward IC +0.30). That figure is a walk-forward validation across the full sample, not a promise about the present: individual windows run negative, and when they do we report the current number inside the entry itself rather than only the historical one. We do not forecast stock prices, and we say so.
  5. Misses stay public. That is the whole point.

Cross-sectional: 23 US chains (MULTI)

OPEN — awaiting resolution

Pre-registered: 2026-08-27 · Resolution date: 2026-09-10 · Data as of: 2026-08-27 (latest captured review: 2026-08-27; 23 chains qualified with >=15 same-store matched locations; reviews flagged as suspected manipulation excluded, none removed from the database)

What our data showed (facts, frozen at publication)

Pre-registered expectation

Over the next 14 days the six chains listed as improving will, as a group, show a higher mean same-store sentiment change than the six listed as declining. This is the 14-day sentiment-persistence relationship (walk-forward IC +0.30) applied across a cross-section. We are not predicting the direction of any individual chain, and we are not predicting any share price.

How this will be graded

Sample integrity note (appended 2026-08-29, before grading)

Two things went wrong with this entry after we published it. We are recording both before we know the result, because a disclosure made after the outcome is worth nothing.

1. We contaminated our own sample. On 28 August we ran a burst that deepened our review collection. It worked — one day brought 56,402 reviews against a normal ~4,000 — and it landed very unevenly across the twelve chains named above. Locations passing this entry's filters, recomputed for the pre-registration date of 27 August, using only the data we held that day versus all the data we hold now: Domino's 71 → 177, Dunkin' 27 → 96, Chipotle 31 → 70, AT&T 16 → 28, Verizon 27 → 31, Shake Shack 41 → 44, Subway 75 → 78, In-N-Out 29 → 31, Wendy's 30 → 32, LA Fitness 16 → 17, Taco Bell 58 → 58, Trader Joe's 16 → 16.

The four chains that gained most are all in the declining group. The improving group gained ten locations in total. This entry compares those two groups against each other, so grading on the fuller data would weigh one side of the bet on a substantially rebuilt sample against the other on its original one.

2. We can reproduce the method, but not every number. Reconstructing the sample, we found that the code which produced this entry was never saved — the query was run ad hoc.

Working from the method as written above, the reconstruction matches the entry where it decides the shape of the test: 23 of our 79 chains qualify, exactly as pre-registered. That is not a weak check — a plausible wrong reading of the method, counting by the date we collected a review rather than the date it was written, yields 17.

At chain level it reproduces nine of the twelve location counts exactly. Three differ by one or two locations, in both directions: Domino's −2, AT&T −1, Taco Bell +1. A missing filter produces error in one direction; differences this small running both ways are what window edges produce — timestamps sitting on a boundary, and ties at the ≥3 and ≥5 thresholds.

The failure here is not in the numbers, it is that nobody stored the query. That is now closed: no entry is pre-registered without the code that produced it saved alongside it. Reproducibility is the entire product here, and “almost” does not qualify.

What we are doing, decided before grading. (a) Collection on all twelve chains is frozen until this entry is graded on 10 September; deepening continues on our other 67 chains. (b) The binding grade is the one computed on the frozen sample — the locations as they stood on 27 August, which our point-in-time capture makes recoverable: 437 locations across 12 chains. That is the sample this entry was pre-registered against. (An earlier reconstruction of ours counted 449; 437 is the figure produced by the single definition published here, and it is the one that governs.) (c) We will also publish the grade on the full data, as context and as proof that we are not hiding it — it is not the grade of record, and if the two disagree, the disagreement is published too. (d) The exact definition used for both is published with the grade, so the frozen sample can be rebuilt and checked by anyone.

This note was written and published before either grade was computed. We do not know which way the result goes.

Threshold note: this entry can VOID, and we are saying so now (appended 2026-08-31, before grading)

A third thing about this entry is worth recording before it is graded, because after the result it would only read as an excuse.

The grading rule can end this entry in a VOID, and the reason is arithmetic, not bad luck. The entry qualifies a location with at least 3 reviews in the recent window and at least 5 in the baseline, and qualifies a chain with at least 15 such locations. That filter was set on a 28-day window. Grading applies it to a 14-day one, so the same threshold is roughly twice as hard to clear. Rule 4 then voids the entry if fewer than four chains in either group still hold 15 locations.

We measured how close that is, on a window that cannot tell us anything about the outcome: 13 to 27 August, which closes on the day this entry opens. On that window the binding reading keeps only three of the six improving chains at 15 or more locations - LA Fitness 11, Trader Joe's 13, Wendy's 13 - and would have voided. The declining group keeps five, losing AT&T at 12. The four fragile chains are exactly the four that qualified with 16 or 17 locations at pre-registration: they have no margin, so a thin fortnight drops them out.

We are not changing the rule. We now know that the strict reading points toward VOID on a comparable window while two looser readings of the same sentence would have graded, and that is precisely why the rule stays as written. A threshold moved while holding that knowledge is a moved goalpost, whatever the reasoning attached to it. The ledger exists to make that impossible.

Two things pull the other way, and we state them so this note is not read as a prediction. The grading re-scrape covers all 437 pre-registered locations at once, while the rehearsal window was covered only by our normal collection rotation, which reaches a fraction of them on any given day. Better coverage means more locations clearing the threshold. So VOID is a live possibility, not a forecast.

Finally, the failure recorded in the note above - that the code producing this entry was never saved - is closed for the grade as well. The grading code was written, reviewed and committed on 31 August, before any of the data it will read existed, and it records its own SHA-256 (d53f5ff6e0e9...) in the result, so the published grade names the exact code that produced it. It refuses to grade a sample staler than the two days this entry allows, and it computes the binding reading on the frozen 437 locations, never on a chain re-qualified with today's data.

SHA-256 (immutable fields): 0101a914cbd4e76bdd9dce1610f882d61c2ec2b0fe99e78ef7e5f5fbfb8c6182

Shake Shack (SHAK)

GRADED: MISS

Pre-registered: 2026-07-04 · Earnings print: 2026-07-30 · Data as of: 2026-07-04 (latest captured review: 2026-07-03; 527 reviews in trailing 28d across 40 locations; 0 reviews filtered as suspected manipulation)

What our data showed (facts, frozen at publication)

Pre-registered expectation

Primary (based on our validated 14-day sentiment-persistence relationship, walk-forward IC +0.30): this broad-based softness persists — when we recompute trailing-28d aspect sentiment on 2026-07-28 (two days before the print), at least 4 of the 6 aspects will still be below their 210-day baselines. Secondary (exploratory, NOT a validated relationship): we would be surprised if management reported a strong positive inflection in traffic/comp momentum for the weeks covered by this window. We do not forecast SHAK's stock price.

How this will be graded

Grade (2026-07-28 (primary) / 2026-08-11 (secondary))

PRIMARY GRADE: MISS — graded 2026-07-28, exactly as pre-registered. We re-ran the identical RSAS v2.1 aspect computation (latest captured review: 2026-07-28; 794 reviews in trailing 28d across 58 locations; 0 reviews filtered as suspected manipulation). The rule required at least 4 of the 6 tracked aspects to remain below their baselines; only 3 of 6 did (food_quality 0.109 vs baseline 0.120; order_accuracy -0.092 vs -0.025; price_value 0.001 vs 0.009), while service_staff (0.203 vs 0.203), speed_wait (0.080 vs 0.065) and cleanliness (0.092 vs 0.059) sat at or above baseline. The honest read: the broad-based softness we measured on 2026-07-04 did NOT persist — most aspects recovered to within a whisker of their own baselines, leaving order accuracy as the only clearly negative driver going into the print. Integrity note: our collection rotation had left Shake Shack's capture 18 days stale as of this morning, and a grade computed on that partial sample (81 in-window reviews) would have shown 4 of 6 below baseline and scored HIT; we instead ran a targeted re-scrape (45 locations, ~2,000 new reviews) hours before grading and publish the result the fuller data supports: MISS. Transparency caveat: the trailing-28d footprint differs from pre-registration (794 reviews / 58 locations now vs 527 / 40 then). Scheduling note: Shake Shack's Q2 print has moved to 2026-08-05 (premarket) from the 2026-07-30 date recorded at pre-registration; the exploratory secondary comparison of management commentary vs these drivers will be appended within 7 days of the actual print, as pre-registered. SECONDARY GRADE (exploratory, appended 2026-08-11 per the pre-registered rule): MISS. Shake Shack printed Q2 2026 on 2026-08-05 premarket (moved from the 2026-07-30 date recorded at pre-registration) — total revenue $417.6M (+17.2% y/y, within the revised guidance range), Same-Shack sales +3.5% vs ~+2.6% consensus and above the top of the June-cut guidance range (2.5–3.0%), adjusted pro forma EPS $0.43 vs ~$0.31–0.33 consensus. Composition per the earnings call: +2.0% traffic and +1.5% price/mix — the fourth consecutive quarter of positive traffic. Our pre-registered exploratory read said we 'would be surprised if management reported a strong positive inflection in traffic/comp momentum for the weeks covered by this window' (trailing-28d ending 2026-07-03). Management reported exactly that inflection: per the call, April comps were negative (-0.6%) and June was the strongest period of the quarter — even excluding an estimated ~90bps World Cup contribution — so the acceleration was concentrated precisely in the weeks our window covered. Scored honestly: MISS. Two observations, recorded as context rather than excuses: (1) our own primary re-computation on 2026-07-28 already showed most customer-experience aspects recovering to their baselines — the data caught the inflection, but three weeks after our 2026-07-04 read, which is exactly what the primary MISS above records; (2) part of the June strength came from an external demand catalyst (World Cup traffic, per management) that customer-experience review sentiment is not designed to detect. Profitability context: net income fell y/y ($16.9M vs $22.4M) on cost pressure — the beat was demand-led, not margin-led. Net: the primary MISS remains the accountable record; the exploratory overlay is also a MISS. We publish both without hedging — that is the point of this ledger. Sources: Shake Shack Q2 2026 earnings release (BusinessWire, 2026-08-05) and the Q2 2026 earnings call.

SHA-256 (immutable fields): 29f839eea80a684c9eebd6c40ebff2e34efce50958ca3806f0119fd63566a17e

Domino's Pizza (DPZ)

GRADED: MIXED

Pre-registered: 2026-07-04 · Earnings print: 2026-07-20 · Data as of: 2026-07-04 (latest captured review: 2026-07-04; 409 reviews in trailing 28d across 80 locations; 0 reviews filtered as suspected manipulation)

What our data showed (facts, frozen at publication)

Pre-registered expectation

Primary (validated persistence relationship, IC +0.30): the split picture persists - when we recompute trailing-28d aspect sentiment on 2026-07-18 (two days before the print), at least 3 of the 4 currently-improving aspects (service_staff, food_quality, order_accuracy, price_value) will still be at or above their 210-day baselines, AND speed_wait will still be below its baseline. Secondary (exploratory, NOT validated): a customer-experience recovery with deteriorating speed/wait is consistent with stabilizing demand plus delivery-capacity strain; we flag delivery service times as the metric to watch in management commentary. We do not forecast DPZ's stock price.

How this will be graded

Grade (2026-07-18 (primary) / 2026-07-20 (secondary))

PRIMARY GRADE: MIXED — graded 2026-07-18, two days before the print, exactly as pre-registered. We re-ran the identical RSAS v2.1 aspect computation (latest captured review: 2026-07-16; 688 reviews in trailing 28d across 224 locations; 0 reviews filtered as suspected manipulation). Condition 1 — at least 3 of the 4 previously-improving aspects at/above their 210-day baselines — FAILED: 0 of 4 held (service_staff -0.007 vs baseline +0.048; food_quality -0.044 vs -0.001; order_accuracy -0.060 vs -0.025; price_value -0.040 vs -0.011). Condition 2 — speed_wait still below its baseline — HELD: -0.009 vs +0.165. One of the two conditions holding grades as MIXED under the pre-registered rules. The honest read: the customer-experience recovery we measured on 2026-07-04 did NOT persist — every recovering aspect slipped back below its own baseline — while the speed/wait deterioration we flagged remains the clearest feature of the data going into the 2026-07-20 print. Transparency caveat: our collection densified materially over the window (trailing-28d sample grew from 409 reviews / 80 locations to 688 / 224), so the two snapshots do not share an identical location footprint; the grade stands regardless — the rules were fixed in advance. Secondary (company commentary vs these drivers) will be appended within 7 days of the print, labelled exploratory. SECONDARY GRADE (exploratory, appended 2026-07-20 per the pre-registered rule): DPZ printed Q2 2026 on 2026-07-20 — revenue $1,194.4M (+4.3% y/y, above consensus), diluted EPS $4.07 (+6.8% y/y, below the ~$4.17 consensus), US same-store sales +0.1% (the weakest US comp in over a year), with carryout +1.1% and delivery -0.7% per the earnings call. Scoring our exploratory read leg by leg: (1) 'stabilizing demand' — SUPPORTED in volume terms: management reported meaningful order-count growth across both delivery and carryout against a flat-order QSR industry, with comp softness driven by ticket, not traffic. (2) 'delivery-capacity strain', with delivery service times flagged as the metric to watch in management commentary — NOT CORROBORATED: neither the release nor the call discussed delivery service times or capacity; management attributed the ticket miss to its own promotional messaging ('largely within our control'). Delivery was the softer channel (-0.7% vs carryout +1.1%), which is directionally consistent with speed/wait remaining the one persistent negative in our data — but the company's own framing points at pricing/mix, not service friction, and we do not claim credit for a mechanism the print did not confirm. Net: the validated primary signal (MIXED, above) remains the accountable record; this exploratory overlay scores one leg right (demand stabilization via order counts) and one leg unconfirmed (delivery strain). Sources: Domino's Q2 2026 earnings release (ir.dominos.com) and the Q2 2026 earnings call.

SHA-256 (immutable fields): 9d5288b8191e7c463e5cc053a6382440999493ab0ae923fcaee880a88c9f91f0

Skin in the game — the Accountability Guarantee

Talk is cheap, even hash-stamped talk. So we put revenue behind this ledger:

  1. If any primary pre-registered prediction on this page is graded a MISS, every active data-license subscriber gets their next month free.
  2. Applies to Monthly and Weekly Data License subscriptions active on the day the grade is published. Credit is applied to the next invoice; capped at one free month per calendar quarter; no cash value.
  3. We grade ourselves by the rules written into each entry before the print — and the entry hashes above make those rules impossible to quietly rewrite.

Large vendors cannot copy this without repricing their entire book. We can, because we would rather lose a month of revenue than a decade of trust.

Verify it yourself — don't trust us

  1. Machine-readable ledger: /proof/ledger.json — the exact source this page is rendered from. Each entry's sha256 is computed over its immutable pre-registration fields (sorted-key canonical JSON of: id, chain, ticker, published_at, earnings_date, data_as_of, facts, prediction, grading_rules).
  2. Bitcoin timestamp: /proof/ledger.json.ots is an OpenTimestamps proof anchoring the ledger's hash in the Bitcoin blockchain. Verify with: ots verify ledger.json.ots -f ledger.json. A newly published entry shows a pending attestation until the next Bitcoin block confirms it — usually within hours. No blockchain hype — just a free, independent clock nobody (including us) can rewind.
  3. Independent web archive: snapshots of this page are stored by the Internet Archive Wayback Machine — a third party we don't control.