# Appendix — material demoted from the main paper, preserved in full

Restructured 20 August 2026 after external review. Nothing here was deleted from the
record; it was moved because the reader should not pay for bookkeeping. Everything
below remains citable and every number is reproducible from the code archive.

## A. Configuration ledger

The paper reports separations of 0.643, 0.708, 0.794 and 0.602 and label counts of
11, 13 and 15. They are different configurations; the audit changed the configuration
three times, and reporting only the final row would hide that.

| Configuration | Control fence | Sensors | Labels scored | Separation |
|---|---|---|---|---|
| As first tested | 36 min | Both must agree | 11 of 16 | 0.643 |
| After the fence fix | 42 min | Both must agree | 11 of 16 | 0.708 |
| Combined with hazard model | 36 min | Both must agree | 11 of 16 | 0.602 |
| **Final: Fitbit alone** | **42 min** | **Fitbit only** | **13 of 16** | **0.794** |

Label counts differ because a label is only scorable where the detector has a complete
window. Requiring both watches to agree cost five labels; Fitbit alone recovers two.

## B. The hazard model, in full

Time since the previous cigarette under three information regimes. Measured in the
earlier both-watches configuration (36-minute fence), in which physiology scored 0.643;
read the ordering between rows, not against the 0.794 headline.

| What the model knows | Separation (AUC) |
|---|---|
| The held-out day's true cigarettes (oracle) | 0.688 |
| Only the first cigarette of the day, user-logged | **0.597** |
| Nothing from today — fully autonomous | 0.377 |
| Hour of day alone | 0.516 |
| Physiology alone, same configuration | 0.643 |

Run autonomously the hazard model is worse than a coin flip, and combining it with
physiology degrades physiology (0.566 against 0.602 on the shared evaluable minutes).

**The retracted Poisson test.** Same-day intervals ran 62–243 minutes, mean 136,
CV 0.42. I originally reported this as beating random timing at p = 0.005 against a
Poisson null. That test is close to meaningless — human behaviour bounded by waking
hours, work and meals is never Poisson, so almost any recurring habit clears the bar —
and the claim is retracted. The 0.42 is description, not evidence.

## C. The six-panel audit figure

![The audit in six panels](/research/wearable-smoking-detection/figures/adversarial-audit.png)

Top: requiring both watches to agree left 4.6% of awake time scorable, and a freely
rotated null retained only 6.2 of 16 cigarettes against 11 in the observation. Middle:
correcting the null lowers p from 0.24 to 0.13 but never below; separation rises as the
control fence widens past the 42 minutes the score window reaches. Bottom: timing beats
physiology only when handed the answer. Produced by the reviewing agent from the same
code and labels.

## D. Ablation, fourth row

Residual size alone, given every label it can score rather than the like-for-like 13,
reaches 0.719 on 15 labels — the more flattering of its two figures. Restricted to the
same 13 it scores 0.676, so the template's like-for-like margin (+0.118) is wider than
the loose comparison suggests.

## E. Third phase — interim write-up (pre-registration frozen 20 August 2026)

Published in full when the blind window closes; preserved here so the pre-registration
is publicly dated. Success criteria were fixed on 20 August, before any blind data
existed.

**Live labels and the blind evaluation of v2.** Between 17 and 20 August I logged
twenty cigarettes at the moment of lighting. They corrected the base rate a second time
— live-logged days run five to eight cigarettes, so the recalled week was itself
under-logged — and enabled the first genuinely blind evaluation of the frozen detector:
precision doubled to 32% at ±15 minutes, but at five to eight cigarettes a day a random
alarm lands within fifteen minutes of one about a quarter of the time, so the lift over
chance is 1.3×, at p = 0.24 on four days.

**The two-stage detector, and the audit that cut it down.** Baselines conditioned on
inferred activity state (rest, walking, ten minutes of post-walk recovery), plus a
second stage scoring each candidate against rival response templates — meal, acute
stress, standing up — with constants fixed from published physiology. Cross-validated
it reported 50% precision at three alarms a day. An independent adversarial audit
showed the threshold-selection rule partly manufactured the 50%; the honest nested
estimate is 42.9%, with the 30% chance floor inside its bootstrap interval [20%, 75%].
The rival templates failed their ablation. What survived: the state-conditional
baseline, and a two-thirds cut in alarm rate at unchanged recall — which is what makes
a blind test statistically decidable at all.

**The night as dose meter.** Nine nights of Fitbit's sleep-only 5-minute RMSSD
replicate an earlier single-night observation on eight of eight usable nights:
overnight heart-rate variability rises along a smooth recovery curve, beating a
sleep-stage explanation every time. The level of the curve tracks the previous day's
dose — median 19 ms the night after a ≥10-cigarette session against 77 ms after the
confirmed zero day (ρ = −0.84 across six dose nights; family-wise p = 0.21 after
correcting for the seven features examined — a hypothesis with a large effect size,
not a result). Nightly respiratory rate moves the same way (zero day slowest at
15.6 brpm; the heaviest days fastest at 17.4) and is registered as a secondary readout.

**The four registered instruments (constants frozen 20 August):**

| Arm | Question | Success bar |
|---|---|---|
| v3, two-stage detector (primary) | autonomous live detection | precision ≥50% at ±15 min AND circular-shift p < 0.05 |
| v4, audit-survivors only (secondary) | same | same, α = 0.025 |
| Toll Score (sleep dose meter) | rank previous-day totals from overnight HRV level + non-REM heart rate | Spearman ρ > 0, permutation p < 0.05, ≥8 nights |
| Count-conditioned decoder | given the day's true tally, localize the moments | hit rate vs shift null, p < 0.025 |

Window: 21 August – 3 September 2026. Days flagged as incompletely logged are dropped
whole; zero-cigarette days must be declared, never inferred. Anything not written here
is post-hoc and will be labelled as such.

**Interim figures:**

![The audit honesty panel](/research/wearable-smoking-detection/figures/v3-honesty-panel.png)
![Anatomy of one anonymised day](/research/wearable-smoking-detection/figures/v3-day-anatomy.png)
![Alarm economy](/research/wearable-smoking-detection/figures/v3-alarm-economy.png)
![Overnight RMSSD by dose](/research/wearable-smoking-detection/figures/rmssd-trajectories.png)

Source data for the first three: the CSVs alongside this file. The RMSSD figure's
source is withheld — per-night timing plus logged gaps would reconstruct cigarette
times, which this project does not publish.
