# Findings log — v2 rebuild, 17 August 2026

Running record of every experiment in the v2 rebuild, including the ones that failed and
the two results I reported and then had to withdraw. Written in order.

Ground truth after this session: **16 point-labelled cigarettes across 6 days**, plus one
heavy session (15 Aug 21:00 → 16 Aug 04:00, ≥10 cigarettes) and one confirmed
**zero-cigarette day** (16 Aug after 04:00). The zero day is the single most valuable
label the project has ever had — before it, precision was unmeasurable.

---

## 1. Data

### 1.1 The Apple Health relay is lossless
Cross-correlated Apple Health's `Google Health` heart rate against the Takeout 2-second
stream over 1,684 overlapping minutes: **r = 0.995 at lag 0, mean |diff| 0.68 bpm.**

Consequence: **Takeout is not needed.** More importantly, the 2-second data was never
load-bearing — `smoke_shape.py:67` medians it into 1-minute bins before matching. The
project's headline "2-second resolution" was discarded before it reached the detector.

### 1.2 The real smoking rate is 0–4/day, not 8–12
REPORT.md §3 and the published paper both assume 8–12 cigarettes/day (self-reported) and
state that precision is therefore unmeasurable. The 12–17 Aug labels give
2, 1, 1, 4, 0 per day plus one ≥10 session. **The old `TOP_N = 10..12` was calibrated on a
prior that is wrong by 3–10×**, which alone guarantees ~9 false alarms on a 1-cigarette
day and cannot represent a zero day at all.

### 1.3 Garmin was contributing nothing (bug, fixed)
Garmin samples every 2 minutes. On a 1-minute grid every other slot is NaN, so no
39-minute analysis window is ever complete and the device silently dropped out — which is
also why "require both sensors to agree" never won a model selection: it could never find
complete windows in both. Fixed by interpolating across gaps no larger than each device's
own measured cadence (`detector.regularise`).

| | before | after |
|---|---|---|
| Garmin usable grid minutes | ~0 | 46,425 |
| minutes with both devices | ~0 | 7,326 |

---

## 2. Where the old detector actually stands

Scored the existing `smoke_shape` method on all 6 testable pre-existing labels with a
permutation baseline: **3/6 at ±15 min vs 2.23 expected, p = 0.39.** At ±5 and ±10 it is
below chance. REPORT.md §6's 4/4 did not replicate, consistent with its own corrected
p ≈ 0.11 caveat. On 10 Aug its `23:27` hit was candidate **12 of 12** — one rank from
vanishing.

**Precision, measured for the first time.** On the zero-cigarette day the old detector
emits 10 candidates, all false. The single highest-scoring candidate across all six days
(score 25.3, 16 Aug 22:53) is a false positive on the free day. A threshold above the best
false positive retains **1 of 7** true positives.

True positives and false positives are indistinguishable on every feature:

| feature | TP (n=7) | FP (n=10, zero-cigarette day) |
|---|---|---|
| corr | 0.6 [0.3, 0.9] | 0.7 [0.3, 0.8] |
| amp | 24.1 [5.3, 51.0] | 27.1 [8.9, 38.2] |
| steps | 92.0 [0, 724] | 92.0 [0, 1101] |
| score | 12.7 [2.0, 26.4] | 17.0 [5.4, 25.3] |

---

## 3. What v2 changed

1. **Exertion residualisation.** HR regressed on exponentially-weighted recent step load
   (robust Theil-Sen slope), scoring only the unexplained part. Replaces the old
   hand-weighted movement penalty, which carried no information. Fitted slopes are real
   and non-trivial: Garmin +0.18, Fitbit +0.42 bpm per unit step load.
2. **Rolling median / MAD baselines** instead of an EMA with a hand-picked alpha. Scores
   are in the subject's own MAD units, so nothing is in literal bpm and the same code
   works on another person.
3. **Causal everything.** Every statistic uses strictly past samples, so the detector can
   run on tomorrow's data.
4. **Sleep excluded** — structural, not tuned. You cannot smoke while asleep.
5. **Circadian prior** learned from training days only, with circular smoothing.
6. **Threshold, not quota.** Absolute threshold calibrated on training days, so a quiet
   day is allowed to be quiet. Free-day false alarms dropped from **10 → 3**.
7. **Circular-shift nulls** everywhere, after the Claude Science session showed naive
   permutation is anti-conservative for autocorrelated series (RMSSD lag-1 r = 0.81;
   naive p < 0.0001 became p = 0.011 under circular shift).

---

## 4. Results, including two withdrawals

### 4.1 Withdrawn: "AUC 0.853, p = 0.023"
Reported mid-session. It was measured **before** the Garmin fix, when scorable minutes
were restricted to Fitbit wear time — which correlates with waking/smoking hours and
implicitly narrowed the candidate pool. After the fix the same analysis gives AUC 0.797
with a selection-corrected p of ~0.43. **Withdrawn.**

### 4.2 Withdrawn: "50% recall at ±20 min, p = 0.003"
Same cause. After the Garmin fix the best cell is 50% at ±15 min with p = 0.057.
**Withdrawn.**

### 4.3 The honest result: a pre-specified model does not beat chance
Grid-searching 1080 combinations and reporting the winner needs a selection-corrected
null, and the correction consumes the whole effect. So I stopped selecting and specified
one model in advance from the pharmacology — onset 5 min (a cigarette is smoked over 5–7
min and the response tracks accumulated dose), τ_rise 1.5 min, τ_decay 20 min (HR falls
~20 min after the last puff), both sensors required to agree:

| | AUC | circular-shift null (95th pct) | p |
|---|---|---|---|
| physiology only | 0.643 | 0.768 | 0.24 |
| + circadian prior | 0.651 | 0.735 | 0.17 |

**Not significant.** One test, no selection, no correction owed.

### 4.4 Operating curve (pre-specified model)

| alarms/day | tolerance | recall | chance | p | free-day FP |
|---|---|---|---|---|---|
| 3 | ±15 | 12% | 12% | 0.61 | 1 |
| 5 | ±30 | 50% | 30% | 0.058 | 3 |
| 8 | ±30 | **62%** | 40% | **0.024** | 3 |
| 12 | ±15 | 44% | 32% | 0.22 | 6 |

The only cell reaching p < 0.05 is 8 alarms/day at **±30 minutes**: 62% recall against 40%
chance. Precision there is ~21% (10 hits from 48 alarms). This is one cell of 16, so
uncorrected.

### 4.5 Diagnosis: detection is not the bottleneck, localisation is
Within-day AUC runs 0.68–0.91 depending on model, so the score **is** elevated around
cigarettes. But of 16 labels, **6 have no candidate within ±12 min at all**, and the
matched ones rank 2,4,5,6,6,7,8,9,10,15 out of 15–30 candidates/day. Signed offsets of
matched alarms: median −5 min, 3/8 positive — **scatter, not bias**, so there is no
systematic lag to calibrate away. Widening the tolerance to ±30 is what buys the recall,
and it buys chance at almost the same rate.

### 4.6 Things that did not help
- **Cross-device agreement.** Once genuinely testable (7,326 dual-covered minutes), it
  did not win selection on most folds and the pre-specified `min` fusion scored AUC 0.643.
- **Onset lag.** An asymmetric search window suggested a +8 min lag at p < 0.01; with a
  **symmetric** window it is 5/7 positive, median +8, **p = 0.45**. The first version was
  an artifact of my own search asymmetry.
- **The 15→16 Aug session** (≥10 cigarettes) is unusable as validation: Fitbit covers 28
  of its 420 minutes, and after regularisation Garmin coverage still leaves the window too
  sparse for a clean rate comparison.

---

## 5a. Independent audit, and the fix that changed the answer

An independent Claude Science session was asked to refute section 4.3. It reproduced the
result exactly, then found four errors — **all four biased against the detector**:

1. **Control fence too short.** Controls were excluded within `3*WIN = 36` min of a label,
   but the score window reaches `POST+WIN = 42` min, so 146 of 914 controls (16%) carried
   part of a real response. Fixed: `EXCL = POST + WIN`.
2. **Circular-shift null dropped positives.** Free rotation landed labels in coverage gaps,
   so null draws averaged 6.2 usable positives against an observed 11, and 95.8% of draws
   had fewer than 11. Fewer positives widens the null and inflates p. Fixed: rotation now
   happens *within the set of scorable minutes*.
3. **`min` fusion was the real bottleneck.** Requiring both watches to agree made a minute
   scorable only when BOTH had a complete window: **1,352 of 29,649 awake minutes (4.6%)**,
   and only 11 of 16 labels usable. Coverage, not artifact rejection, was binding.
4. Hour-matching and the awake filter turned out to be nearly inert (0.643 vs 0.629
   unmatched); both were defending against confounds that were not operating.

Also confirmed: **exertion residualisation is not eating signal.** Step load *is* elevated
at cigarettes (Fitbit 1.69 vs 0.59), so the concern was real, but AUC rises monotonically
with correction strength (β=0: 0.616, fitted: 0.643, β×2: 0.650). If it were removing
signal, over-correcting would hurt.

**The hazard model is dead.** Cigarettes *cluster* — median elapsed-since-previous is 195
min at labels vs 73 min at controls, the opposite of REPORT §9's hypothesis. Timing beats
physiology (AUC 0.688) only when handed the test day's own cigarettes. Run autonomously it
scores **0.377, below chance**, and combining it with physiology makes physiology worse
(0.566 vs 0.602). Worth keeping: with the day's *first* cigarette user-logged, it reaches
0.597 — a semi-supervised product is more realistic than a fully automatic one.

### Result after the fixes (Fitbit alone, corrected null, corrected controls)

Single pre-specified model, no grid search:

| | AUC | null 95th pct | p |
|---|---|---|---|
| physiology only | **0.794** | 0.719 | **0.0075** |
| + circadian prior | **0.799** | 0.704 | **<0.001** |

Mean within-day AUC 0.901. Scorable awake minutes rose 1,352 → 2,463.

Operating curve, threshold calibrated on the other days only:

| alarms/day | tol | recall | chance | p | free-day FP |
|---|---|---|---|---|---|
| 12.3 | ±10 | 56% | 24% | 0.005 | 4 |
| 12.3 | **±15** | **75%** | 33% | **<0.001** | 4 |
| 7.8 | ±30 | 56% | 37% | 0.058 | 4 |
| 3.0 | ±15 | 25% | 11% | 0.104 | 2 |

**Caveat on test count:** two fusion rules have now been pre-specified and tested (`min`,
then Fitbit-alone), so the honest correction is p ≈ 0.0075 × 2 ≈ 0.015. Still significant.

**Caveat on precision:** 75% recall costs ~12 alarms/day against an actual 0–4 cigarettes.
Precision is roughly 16%. The detector is a prompt, not a diary.

A parallel grid-search evaluation with nested per-fold selection gives a weaker 50% at ±15
(p = 0.040), because the grid keeps choosing worse fusion rules. The pre-specified model is
the more trustworthy object — it involves no selection at all.

## 5. Bottom line

The requested 60–70% at usable precision is **not supported by this data**. The defensible
claims are:

1. The score is genuinely elevated around cigarettes (within-day AUC ~0.7–0.9), but not
   enough to survive honest selection correction at n = 16.
2. The one operating point reaching p < 0.05 is **62% recall at ±30 min, 8 alarms/day,
   ~21% precision** — useful as a retrospective prompt, not as a diary.
3. v2 genuinely beats v1 on precision: 3 false alarms on a clean day versus 10.

**n = 16 is the binding constraint, not the algorithm.** Six labelled days with 1–4
cigarettes each cannot separate a 2× effect from chance. The single highest-value action
is more labels; ~100 events would make the existing machinery testable at ±10 min.

Second-highest: a chest strap (Polar H10) for real waking RMSSD. Both REPORT.md §3 and the
Claude Science overnight analysis point at the autonomic term as the load-bearing one, and
it is the only proposed signature that has never been measurable.
