Cross-sectional: 23 US chains (MULTI)
OPEN — awaiting resolution
Pre-registered: 2026-08-27 ·
Resolution date: 2026-09-10 ·
Data as of: 2026-08-27 (latest captured review: 2026-08-27; 23 chains qualified with >=15 same-store matched locations; reviews flagged as suspected manipulation excluded, none removed from the database)
What our data showed (facts, frozen at publication)
- Method, frozen here: for each location we take the mean sentiment of the trailing 28 days minus that same location's own baseline (days 29-118). A location qualifies with >=3 recent and >=5 baseline reviews. The chain value is the min(n_recent,30)-weighted mean across its qualifying locations, and a chain qualifies with >=15 of them. 23 of our 79 chains qualified.
- Six most improving: Taco Bell +0.0300 (57 locations), Verizon +0.0140 (27), Subway +0.0072 (75), LA Fitness +0.0052 (16), Wendy's +0.0047 (30), Trader Joe's +0.0023 (16).
- Six most declining: Chipotle Mexican Grill -0.0751 (31 locations), AT&T -0.0472 (17), Shake Shack -0.0428 (41), Domino's -0.0291 (73), Dunkin' -0.0268 (27), In-N-Out Burger -0.0256 (29).
- Magnitudes are asymmetric and we say so before the outcome: only Taco Bell (+0.0300) is clearly rising, while the other five chains in the improving group sit between +0.0140 and +0.0023, which is close to noise. The declining group moves three to thirty times harder (-0.0751 to -0.0256). In practice this entry therefore asks whether the decliners keep declining rather more than it asks whether the improvers keep improving.
- Dispersion across the qualifying universe runs from +0.0300 to -0.0751. The eleven chains between these groups are named in neither and are not part of the test.
- Engine: RSAS v2.1, the production engine. Our own conformal validation shows that a 90% interval for a single chain over a 14-day horizon contains zero for 178 of 178 chains. Single-chain direction is therefore not resolvable from our data, which is precisely why this entry is stated cross-sectionally rather than as a call on any one brand.
- Pre-registered against our own most recent validation window. In the shadow evaluation we ran on 2026-08-27, both our production engine (RSAS v2.1) and its challenger showed a negative mean information coefficient on the clean same-store outcome - -0.070 and -0.034 respectively on weekly folds. The July-August window was anti-persistent. This entry bets on persistence in a regime our own measurement calls anti-persistent. One negative window is not proof of a durable reversal, and flipping the sign to fit it would be overfitting, so we state the tension here rather than after the result is known.
Pre-registered expectation
Over the next 14 days the six chains listed as improving will, as a group, show a higher mean same-store sentiment change than the six listed as declining. This is the 14-day sentiment-persistence relationship (walk-forward IC +0.30) applied across a cross-section. We are not predicting the direction of any individual chain, and we are not predicting any share price.
How this will be graded
- On 2026-09-10 we recompute, for each of the twelve named chains, the forward change: mean same-store sentiment over the window (2026-08-27, 2026-09-10] minus that location's baseline of the 90 days ending 2026-08-27, using the same location filters stated above.
- Sample freshness, fixed here so the decision cannot be made after seeing the result: the grading sample must be no older than 2 days on the grading date. If, on 2026-09-10, the most recent captured review for any of the twelve named chains is staler than that, we re-scrape that chain before grading and record the re-scrape in the grade note.
- HIT if the unweighted mean forward change of the six improving chains exceeds that of the six declining chains. MISS in every other case, including a tie.
- VOID only if fewer than four chains in either group still meet the >=15 matched-location requirement on the grading date. A VOID is published exactly like a HIT or a MISS.
- The grade is computed with RSAS v2.1 - the same engine that produced this entry. Should the engine change before the grading date, this entry is still graded with v2.1.
Sample integrity note (appended 2026-08-29, before grading)
Two things went wrong with this entry after we published it. We are recording both before we know the result, because a disclosure made after the outcome is worth nothing.
1. We contaminated our own sample. On 28 August we ran a burst that deepened our review collection. It worked — one day brought 56,402 reviews against a normal ~4,000 — and it landed very unevenly across the twelve chains named above. Locations passing this entry's filters, recomputed for the pre-registration date of 27 August, using only the data we held that day versus all the data we hold now: Domino's 71 → 177, Dunkin' 27 → 96, Chipotle 31 → 70, AT&T 16 → 28, Verizon 27 → 31, Shake Shack 41 → 44, Subway 75 → 78, In-N-Out 29 → 31, Wendy's 30 → 32, LA Fitness 16 → 17, Taco Bell 58 → 58, Trader Joe's 16 → 16.
The four chains that gained most are all in the declining group. The improving group gained ten locations in total. This entry compares those two groups against each other, so grading on the fuller data would weigh one side of the bet on a substantially rebuilt sample against the other on its original one.
2. We can reproduce the method, but not every number. Reconstructing the sample, we found that the code which produced this entry was never saved — the query was run ad hoc.
Working from the method as written above, the reconstruction matches the entry where it decides the shape of the test: 23 of our 79 chains qualify, exactly as pre-registered. That is not a weak check — a plausible wrong reading of the method, counting by the date we collected a review rather than the date it was written, yields 17.
At chain level it reproduces nine of the twelve location counts exactly. Three differ by one or two locations, in both directions: Domino's −2, AT&T −1, Taco Bell +1. A missing filter produces error in one direction; differences this small running both ways are what window edges produce — timestamps sitting on a boundary, and ties at the ≥3 and ≥5 thresholds.
The failure here is not in the numbers, it is that nobody stored the query. That is now closed: no entry is pre-registered without the code that produced it saved alongside it. Reproducibility is the entire product here, and “almost” does not qualify.
What we are doing, decided before grading. (a) Collection on all twelve chains is frozen until this entry is graded on 10 September; deepening continues on our other 67 chains. (b) The binding grade is the one computed on the frozen sample — the locations as they stood on 27 August, which our point-in-time capture makes recoverable: 437 locations across 12 chains. That is the sample this entry was pre-registered against. (An earlier reconstruction of ours counted 449; 437 is the figure produced by the single definition published here, and it is the one that governs.) (c) We will also publish the grade on the full data, as context and as proof that we are not hiding it — it is not the grade of record, and if the two disagree, the disagreement is published too. (d) The exact definition used for both is published with the grade, so the frozen sample can be rebuilt and checked by anyone.
This note was written and published before either grade was computed. We do not know which way the result goes.
Threshold note: this entry can VOID, and we are saying so now (appended 2026-08-31, before grading)
A third thing about this entry is worth recording before it is graded, because after the result it would only read as an excuse.
The grading rule can end this entry in a VOID, and the reason is arithmetic, not bad luck. The entry qualifies a location with at least 3 reviews in the recent window and at least 5 in the baseline, and qualifies a chain with at least 15 such locations. That filter was set on a 28-day window. Grading applies it to a 14-day one, so the same threshold is roughly twice as hard to clear. Rule 4 then voids the entry if fewer than four chains in either group still hold 15 locations.
We measured how close that is, on a window that cannot tell us anything about the outcome: 13 to 27 August, which closes on the day this entry opens. On that window the binding reading keeps only three of the six improving chains at 15 or more locations - LA Fitness 11, Trader Joe's 13, Wendy's 13 - and would have voided. The declining group keeps five, losing AT&T at 12. The four fragile chains are exactly the four that qualified with 16 or 17 locations at pre-registration: they have no margin, so a thin fortnight drops them out.
We are not changing the rule. We now know that the strict reading points toward VOID on a comparable window while two looser readings of the same sentence would have graded, and that is precisely why the rule stays as written. A threshold moved while holding that knowledge is a moved goalpost, whatever the reasoning attached to it. The ledger exists to make that impossible.
Two things pull the other way, and we state them so this note is not read as a prediction. The grading re-scrape covers all 437 pre-registered locations at once, while the rehearsal window was covered only by our normal collection rotation, which reaches a fraction of them on any given day. Better coverage means more locations clearing the threshold. So VOID is a live possibility, not a forecast.
Finally, the failure recorded in the note above - that the code producing this entry was never saved - is closed for the grade as well. The grading code was written, reviewed and committed on 31 August, before any of the data it will read existed, and it records its own SHA-256 (d53f5ff6e0e9...) in the result, so the published grade names the exact code that produced it. It refuses to grade a sample staler than the two days this entry allows, and it computes the binding reading on the frozen 437 locations, never on a chain re-qualified with today's data.
SHA-256 (immutable fields): 0101a914cbd4e76bdd9dce1610f882d61c2ec2b0fe99e78ef7e5f5fbfb8c6182
Shake Shack (SHAK)
GRADED: MISS
Pre-registered: 2026-07-04 ·
Earnings print: 2026-07-30 ·
Data as of: 2026-07-04 (latest captured review: 2026-07-03; 527 reviews in trailing 28d across 40 locations; 0 reviews filtered as suspected manipulation)
What our data showed (facts, frozen at publication)
- food_quality: trailing-28d sentiment +0.086 vs 210-day baseline +0.202 (declining, 210 aspect mentions)
- service_staff: trailing-28d sentiment +0.176 vs 210-day baseline +0.264 (declining, 189 aspect mentions)
- cleanliness: trailing-28d sentiment -0.031 vs 210-day baseline +0.219 (declining, 25 aspect mentions)
- speed_wait: trailing-28d sentiment +0.023 vs 210-day baseline +0.105 (declining, 56 aspect mentions)
- order_accuracy: trailing-28d sentiment -0.007 vs 210-day baseline +0.077 (declining, 30 aspect mentions)
- price_value: trailing-28d sentiment -0.025 vs 210-day baseline +0.086 (declining, 22 aspect mentions)
- RSAS v2.1 composite: -4.6 (HOLD), data-sufficiency confidence 91/100.
- Context: Shake Shack cut FY2026 guidance on 2026-06-02 (revenue and same-store-sales outlook lowered). One month later, ALL six tracked customer-experience aspects sit below the chain's own 210-day baseline.
Pre-registered expectation
Primary (based on our validated 14-day sentiment-persistence relationship, walk-forward IC +0.30): this broad-based softness persists — when we recompute trailing-28d aspect sentiment on 2026-07-28 (two days before the print), at least 4 of the 6 aspects will still be below their 210-day baselines. Secondary (exploratory, NOT a validated relationship): we would be surprised if management reported a strong positive inflection in traffic/comp momentum for the weeks covered by this window. We do not forecast SHAK's stock price.
How this will be graded
- Primary: on 2026-07-28 we re-run the same RSAS v2.1 aspect computation. HIT if >=4 of 6 aspects remain below baseline; MISS otherwise. Objective and reproducible.
- Secondary: within 7 days after the 2026-07-30 print we publish a comparison of company-reported traffic/SSS commentary vs these drivers, graded hit/miss/mixed in the notes. Labelled exploratory.
- This entry is never edited; only a grade is appended. The SHA-256 hash below covers all pre-registered fields.
Grade (2026-07-28 (primary) / 2026-08-11 (secondary))
PRIMARY GRADE: MISS — graded 2026-07-28, exactly as pre-registered. We re-ran the identical RSAS v2.1 aspect computation (latest captured review: 2026-07-28; 794 reviews in trailing 28d across 58 locations; 0 reviews filtered as suspected manipulation). The rule required at least 4 of the 6 tracked aspects to remain below their baselines; only 3 of 6 did (food_quality 0.109 vs baseline 0.120; order_accuracy -0.092 vs -0.025; price_value 0.001 vs 0.009), while service_staff (0.203 vs 0.203), speed_wait (0.080 vs 0.065) and cleanliness (0.092 vs 0.059) sat at or above baseline. The honest read: the broad-based softness we measured on 2026-07-04 did NOT persist — most aspects recovered to within a whisker of their own baselines, leaving order accuracy as the only clearly negative driver going into the print. Integrity note: our collection rotation had left Shake Shack's capture 18 days stale as of this morning, and a grade computed on that partial sample (81 in-window reviews) would have shown 4 of 6 below baseline and scored HIT; we instead ran a targeted re-scrape (45 locations, ~2,000 new reviews) hours before grading and publish the result the fuller data supports: MISS. Transparency caveat: the trailing-28d footprint differs from pre-registration (794 reviews / 58 locations now vs 527 / 40 then). Scheduling note: Shake Shack's Q2 print has moved to 2026-08-05 (premarket) from the 2026-07-30 date recorded at pre-registration; the exploratory secondary comparison of management commentary vs these drivers will be appended within 7 days of the actual print, as pre-registered. SECONDARY GRADE (exploratory, appended 2026-08-11 per the pre-registered rule): MISS. Shake Shack printed Q2 2026 on 2026-08-05 premarket (moved from the 2026-07-30 date recorded at pre-registration) — total revenue $417.6M (+17.2% y/y, within the revised guidance range), Same-Shack sales +3.5% vs ~+2.6% consensus and above the top of the June-cut guidance range (2.5–3.0%), adjusted pro forma EPS $0.43 vs ~$0.31–0.33 consensus. Composition per the earnings call: +2.0% traffic and +1.5% price/mix — the fourth consecutive quarter of positive traffic. Our pre-registered exploratory read said we 'would be surprised if management reported a strong positive inflection in traffic/comp momentum for the weeks covered by this window' (trailing-28d ending 2026-07-03). Management reported exactly that inflection: per the call, April comps were negative (-0.6%) and June was the strongest period of the quarter — even excluding an estimated ~90bps World Cup contribution — so the acceleration was concentrated precisely in the weeks our window covered. Scored honestly: MISS. Two observations, recorded as context rather than excuses: (1) our own primary re-computation on 2026-07-28 already showed most customer-experience aspects recovering to their baselines — the data caught the inflection, but three weeks after our 2026-07-04 read, which is exactly what the primary MISS above records; (2) part of the June strength came from an external demand catalyst (World Cup traffic, per management) that customer-experience review sentiment is not designed to detect. Profitability context: net income fell y/y ($16.9M vs $22.4M) on cost pressure — the beat was demand-led, not margin-led. Net: the primary MISS remains the accountable record; the exploratory overlay is also a MISS. We publish both without hedging — that is the point of this ledger. Sources: Shake Shack Q2 2026 earnings release (BusinessWire, 2026-08-05) and the Q2 2026 earnings call.
SHA-256 (immutable fields): 29f839eea80a684c9eebd6c40ebff2e34efce50958ca3806f0119fd63566a17e
Domino's Pizza (DPZ)
GRADED: MIXED
Pre-registered: 2026-07-04 ·
Earnings print: 2026-07-20 ·
Data as of: 2026-07-04 (latest captured review: 2026-07-04; 409 reviews in trailing 28d across 80 locations; 0 reviews filtered as suspected manipulation)
What our data showed (facts, frozen at publication)
- service_staff: trailing-28d sentiment +0.062 vs 210-day baseline -0.122 (improving, 57 aspect mentions)
- food_quality: trailing-28d sentiment +0.011 vs 210-day baseline -0.055 (improving, 126 aspect mentions)
- speed_wait: trailing-28d sentiment +0.044 vs 210-day baseline +0.310 (declining, 26 aspect mentions)
- order_accuracy: trailing-28d sentiment -0.036 vs 210-day baseline -0.126 (improving, 41 aspect mentions)
- price_value: trailing-28d sentiment -0.062 vs 210-day baseline -0.171 (improving, 5 aspect mentions)
- RSAS v2.1 composite: +52.9 (BUY), data-sufficiency confidence 89/100.
- Context: after the Q1 comp miss (US SSS +0.9% vs ~+2.3% expected) and June price-target cuts across the street, our data shows customer-experience sentiment RECOVERING vs baseline on service, food quality, order accuracy and price/value - while speed/wait sentiment has deteriorated sharply vs its baseline.
Pre-registered expectation
Primary (validated persistence relationship, IC +0.30): the split picture persists - when we recompute trailing-28d aspect sentiment on 2026-07-18 (two days before the print), at least 3 of the 4 currently-improving aspects (service_staff, food_quality, order_accuracy, price_value) will still be at or above their 210-day baselines, AND speed_wait will still be below its baseline. Secondary (exploratory, NOT validated): a customer-experience recovery with deteriorating speed/wait is consistent with stabilizing demand plus delivery-capacity strain; we flag delivery service times as the metric to watch in management commentary. We do not forecast DPZ's stock price.
How this will be graded
- Primary: on 2026-07-18 we re-run the same RSAS v2.1 aspect computation. HIT if >=3 of the 4 improving aspects remain at/above baseline AND speed_wait remains below baseline; MISS if neither holds; MIXED if only one of the two conditions holds.
- Secondary: within 7 days after the 2026-07-20 print we publish a comparison of company commentary vs these drivers, graded in the notes. Labelled exploratory.
- This entry is never edited; only a grade is appended. The SHA-256 hash below covers all pre-registered fields.
Grade (2026-07-18 (primary) / 2026-07-20 (secondary))
PRIMARY GRADE: MIXED — graded 2026-07-18, two days before the print, exactly as pre-registered. We re-ran the identical RSAS v2.1 aspect computation (latest captured review: 2026-07-16; 688 reviews in trailing 28d across 224 locations; 0 reviews filtered as suspected manipulation). Condition 1 — at least 3 of the 4 previously-improving aspects at/above their 210-day baselines — FAILED: 0 of 4 held (service_staff -0.007 vs baseline +0.048; food_quality -0.044 vs -0.001; order_accuracy -0.060 vs -0.025; price_value -0.040 vs -0.011). Condition 2 — speed_wait still below its baseline — HELD: -0.009 vs +0.165. One of the two conditions holding grades as MIXED under the pre-registered rules. The honest read: the customer-experience recovery we measured on 2026-07-04 did NOT persist — every recovering aspect slipped back below its own baseline — while the speed/wait deterioration we flagged remains the clearest feature of the data going into the 2026-07-20 print. Transparency caveat: our collection densified materially over the window (trailing-28d sample grew from 409 reviews / 80 locations to 688 / 224), so the two snapshots do not share an identical location footprint; the grade stands regardless — the rules were fixed in advance. Secondary (company commentary vs these drivers) will be appended within 7 days of the print, labelled exploratory. SECONDARY GRADE (exploratory, appended 2026-07-20 per the pre-registered rule): DPZ printed Q2 2026 on 2026-07-20 — revenue $1,194.4M (+4.3% y/y, above consensus), diluted EPS $4.07 (+6.8% y/y, below the ~$4.17 consensus), US same-store sales +0.1% (the weakest US comp in over a year), with carryout +1.1% and delivery -0.7% per the earnings call. Scoring our exploratory read leg by leg: (1) 'stabilizing demand' — SUPPORTED in volume terms: management reported meaningful order-count growth across both delivery and carryout against a flat-order QSR industry, with comp softness driven by ticket, not traffic. (2) 'delivery-capacity strain', with delivery service times flagged as the metric to watch in management commentary — NOT CORROBORATED: neither the release nor the call discussed delivery service times or capacity; management attributed the ticket miss to its own promotional messaging ('largely within our control'). Delivery was the softer channel (-0.7% vs carryout +1.1%), which is directionally consistent with speed/wait remaining the one persistent negative in our data — but the company's own framing points at pricing/mix, not service friction, and we do not claim credit for a mechanism the print did not confirm. Net: the validated primary signal (MIXED, above) remains the accountable record; this exploratory overlay scores one leg right (demand stabilization via order counts) and one leg unconfirmed (delivery strain). Sources: Domino's Q2 2026 earnings release (ir.dominos.com) and the Q2 2026 earnings call.
SHA-256 (immutable fields): 9d5288b8191e7c463e5cc053a6382440999493ab0ae923fcaee880a88c9f91f0