# Detecting smoking events from consumer wearable data

**A research report — 11 August 2026**

Subject: single participant (self-tracked). Devices: Google Fitbit Air (primary),
Garmin Connect, iPhone. Data: Apple Health `export.xml` + Google Takeout (Fitbit).

---

## Summary

**Question.** Can a cigarette be detected automatically from wearable data, without the
wearer logging it?

**Answer.** Partially, and less well than the literature's signature list implies. A
shape-matched heart-rate detector localises roughly **60% of cigarettes to within ~15
minutes**, at about **2× over-prediction**. It is not statistically validated. Two of
seven testable cigarettes produced **no cardiovascular response at all**, which no amount
of tuning can recover.

**The central finding is negative and worth stating plainly:** three of the four commonly
cited acute smoking signatures are **not implementable** from consumer wearable exports,
and the one that is (heart-rate spike) is far weaker and less specific than expected.

---

## 1. Objective

Detect smoking events from passively collected data. The target signature set, as
proposed:

1. Sudden HR spike, 10–30 bpm above baseline within 2–5 minutes
2. Sharp HRV drop (RMSSD / HF power) — sympathetic dominance from nicotine
3. "Smoking notch" — accelerometer hand-to-mouth gesture + the HR spike
4. Transient SpO₂ dip from carbon monoxide

Ground truth: participant-reported cigarette times.

---

## 2. Data inventory

### 2.1 Apple Health export (first pass)

A full type census of `export.xml` (104,447 HeartRate records, 2025-06-11 → 2026-08-11)
established what exists historically versus what exists *now*:

| Channel | Present | Cadence | Covers 10–11 Aug? |
|---|---|---|---|
| HeartRate | yes | 2 min (Garmin) | yes |
| HeartRateVariabilitySDNN | yes | **4.9 samples/day** | **no** — ended 2026-05-03 |
| OxygenSaturation | yes | — | **no** — 7 days only, Jul 2025 |
| RespiratoryRate | yes | 48/day → **1/day** | daily summary only |
| Accelerometer | **never** | — | — |

The Apple Watch stopped syncing on 2026-05-03, taking HRV and respiratory rate with it.
Even at its best, **4.9 HRV samples/day is ~100× too sparse** to resolve a 5-minute
nicotine response. Signature 2 was therefore never implementable from this source.

### 2.2 Fitbit Takeout (second pass)

Substantially richer:

| Channel | Cadence | Coverage |
|---|---|---|
| Heart rate | **~2 seconds** (43,387 samples) | 10 Aug 13:16 → 11 Aug 17:20 |
| Steps | 1 minute | same |
| Skin temperature | 1 minute | 10 Aug 13:10 → 11 Aug 17:10 |
| HRV (RMSSD) | 5 minutes | **11 Aug 00:50–08:30 only — sleep** |
| SpO₂ variation | — | sleep, 11 Aug only |

### 2.3 Feasibility verdict

| Signature | Verdict |
|---|---|
| 1. HR spike | ✅ implementable, 2-second resolution |
| 2. HRV drop | ❌ Fitbit computes RMSSD **during sleep only**; the Web API has the same restriction |
| 3. Hand-to-mouth | ⚠️ accelerometer **does** exist (see §2.5) — but too aggregated; tested and failed |
| 4. SpO₂ dip | ❌ sleep-only, single night |

### 2.5 Correction — the accelerometer does exist

An earlier version of this report claimed "no accelerometer in any export." **That was
wrong**, and it was the largest error of the project: the Apple Health census showed no
motion data, and that was generalised to Fitbit without checking Fitbit's own folders.

`Physical Activity_GoogleData/micro_motion_*.csv` contains:
- **`sleep coefficients`** — 30 values per row, one per second: total accelerometer change
  per second. Effectively **1 Hz motion intensity** (100,890 samples).
- **`mean x/y/z`** — accelerometer components averaged over 30 s, 8-bit quantised:
  coarse **wrist orientation** (3,363 samples).

Coverage spans 10 Aug 13:16 → 11 Aug 17:18 and includes **all eight logged cigarettes**.

**But the gesture is still not recoverable**, for a resolution reason rather than an
availability one. `sleep coefficients` is a per-second *scalar magnitude*, not a 3-axis
waveform, and orientation is averaged over 30 s — while a puff cycle is itself ~30–60 s,
so orientation aliases badly. Hand-to-mouth detection needs a 3-axis trajectory at
~10–25 Hz. See §5.5 for the empirical test.

### 2.4 The timezone trap

The exports are internally inconsistent, and this caused a real bug mid-project (sleep was
shifted four hours and mislabelled as waking, corrupting the first detector run):

- Fitbit `heart_rate`, `steps`, `body_temperature`, `heart_rate_variability` → **UTC**
- Fitbit `sleep-*.json` → **already local**
- Apple Health → local, `+0400` suffix

Confirmed by cross-reference: Fitbit HR begins `09:16:28` UTC on 10 Aug; Apple Health
records the same device's first reading at `13:16` local. Offset exactly +4 h.

---

## 3. Method A — the specified heuristic formula

Implemented in [`src/smoke_detect.py`](src/smoke_detect.py) exactly as proposed:

```
mu_HR(t) = a*HR(t) + (1-a)*mu_HR(t-1)        sedentary-gated EMA, a = 0.01
Z_HR(t)  = (HR(t) - mu_HR(t)) / sigma_HR     O(1) sliding-window sigma, 90 min
S(t)     = w1*(dHR/dt) + w2*((mu_HRV-HRV)/mu_HRV) - w3*Activity(t)
Class    = Smoking if Z > T_Z and S > T_S
```

Plus a recovery-window gate (cigarette vs panic attack): HR must return within 1σ of
baseline inside 25 minutes.

### Results — it does not work

| Problem | Measurement |
|---|---|
| Z-gate far too permissive | resting σ_HR is only **4.9 bpm**, so `Z > 2` means "+10 bpm" → **42 anomalies/day** vs 8–12 actual cigarettes |
| S does not rank cigarettes up | sorting all 35 anomalies by S puts the confirmed cigarettes at ranks **9, 10 and 33 of 35** |
| Best achievable | grid search over (w₁, w₃, T_Z, T_S): catching all three requires **37.6 firings/day**, and wins by effectively disabling S |

**Diagnosis.** With `w2` structurally zero, `S` reduces to `w1·velocity − w3·activity`, and
velocity is near-collinear with Z — both mean "HR went up fast." The score carries almost
no information the Z-gate did not already have.

**The HRV term was load-bearing.** On the sleep window where RMSSD does exist, the
autonomic term swings **+0.56 to −0.42**; at `w2 = 40` that is a **±22-point** contribution,
wider than the entire velocity range. The design intuition was right. The sensor feed is
what is missing.

---

## 4. The amplitude correction

An intermediate analysis reported "+28 bpm spikes" at confirmed cigarettes. **That figure
was wrong**, and correcting it changed the whole approach.

`+28` was the **maximum of the 2-second samples** — a transient, largely sensor noise.
Measured properly as a **1-minute median against a 10-minute baseline**:

| Cigarette | sustained rise | transient max | percentile of all waking minutes |
|---|---|---|---|
| 10 Aug 16:05 | **+12 bpm** | +22 | 78th |
| 10 Aug 19:24 | **+6 bpm** | +17 | 60th |
| 10 Aug 23:27 | **+6 bpm** | +9 | 58th |

Distribution of sustained rise across all waking minutes: p50 **+5**, p75 +12, p90 +23,
p95 +28.

**The cigarettes are weaker than routine daily fluctuation.** A threshold catching +6 bpm
fires on **40% of all waking minutes**. Amplitude cannot separate smoking from ordinary
life, and any detector gated on "a big spike" is chasing an artifact.

---

## 5. Method B — shape matching

If size cannot discriminate, shape might. Implemented in
[`src/smoke_shape.py`](src/smoke_shape.py).

A parametric template — fast rise, slow decay, time constants chosen **a priori** to avoid
overfitting:

```
T(u) = (1 - exp(-u/1.5)) * exp(-u/8.0),   u in minutes
score = pearson_corr(observed, T) * min(amplitude, 30)
```

Amplitude is capped at 30 bpm so large exercise excursions cannot dominate the ranking.
HR is filtered to `confidence >= 2` (19% of readings ≥100 bpm are confidence-1, versus 3%
overall).

### Results on 10 Aug (development day)

**2 of 3 caught in the top 12 candidates/day** — including 23:27, whose sustained rise was
only +9 bpm and which was invisible to every amplitude threshold (corr +0.61 carried it).

19:24 was missed: 104 concurrent steps distorted the profile beyond recognition. This
failure mode is intrinsic — when you walk somewhere to smoke, exertion and nicotine
responses superimpose and cannot be separated from HR alone.

---

### 5.5 Motion probe — the hand-to-mouth hypothesis fails

With the accelerometer channel recovered (§2.5), seven motion features were compared at
the **8 logged cigarettes** against **894 matched waking control windows**, each with a
20,000-trial permutation test ([`src/motion_probe.py`](src/motion_probe.py)):

| Feature | Cigarettes | Controls | p |
|---|---|---|---|
| `motion_std` | 3.80 | 3.14 | **0.036** |
| `motion_mean` | 3.81 | 2.75 | 0.099 |
| `orient_z_range` | 95.4 | 77.4 | 0.149 |
| `orient_spread` | 72.7 | 59.5 | 0.199 |
| `still_frac` | 0.07 | 0.19 | 0.200 |
| `steps` | 16.3 | 32.1 | 0.575 |
| **`periodicity`** | **0.26** | **0.25** | **0.902** |

**The periodicity result is the decisive one.** The whole premise of the "smoking notch"
is a repeated raise-hold-lower cadence every 30–60 s. Autocorrelation at those lags is
**0.26 during cigarettes versus 0.25 during ordinary life** — indistinguishable. The
gesture is not present in this signal at this resolution.

`motion_std` is the only feature under 0.05, but **seven were tested**, so corrected
p ≈ 0.25. It is not significant either.

### 5.6 Rerunning the detector with motion included

`motion_std` was nonetheless added as a bounded bonus term (weight 0.9, clipped to ±2σ)
in [`src/smoke_combined.py`](src/smoke_combined.py) and the detector rerun:

| Day | HR only | HR + motion |
|---|---|---|
| 10 Aug, top-12 | 3/3 at ±15 min, median \|Δ\| **2 min** | 3/3, median \|Δ\| **2 min** |
| 11 Aug, top-10 | 4/4 at ±15 min, median \|Δ\| **8 min** | 4/4, median \|Δ\| **8 min** |

**Identical.** The bonus reorders scores slightly but never changes which candidates reach
the top-N. Adding the accelerometer channel produced no measurable improvement.

---

## 6. Held-out test — 11 August

Predictions were generated and **written to disk before the ground-truth log was
revealed** ([`results/pred_aug11.json`](results/pred_aug11.json)). Coverage 08:40–17:19
(8.7 waking hours); 10 candidates emitted.

| Logged | Nearest prediction | Δ | Rank at the true minute | Verdict |
|---|---|---|---|---|
| 09:06 | 09:09 (rank 4) | **+3 min** | 13/366 | ✅ real local peak, amp +28 |
| 10:43 | 10:55 (rank 2) | +12 min | 90/366 | ⚠️ signal at 10:41 (rank 21), pick was 10:53 |
| 12:50 | 12:51 (rank 6) | **+1 min** | 32/366 | ✅ clean peak exactly on time |
| 13:52 | 14:04 (rank 9) | +12 min | **228/366** | ❌ no response — worse than the median minute |
| 17:07 | — | — | — | ⛔ untestable (data ends 17:19; template needs 25 min post-onset) |

### Statistical assessment

| Tolerance | Caught | Chance baseline | p |
|---|---|---|---|
| ±5 min | 2/4 | 0.81 | 0.178 — not better than chance |
| ±10 min | 2/4 | 1.41 | 0.443 — not better than chance |
| ±15 min | 4/4 | 1.92 | **0.038** |

Chance baselines from 20,000-trial permutation tests (10 random picks in the scorable
window).

**The single significant cell does not survive correction.** Three tolerances were tested;
corrected, p ≈ 0.11. With 10 predictions across 8.7 hours, a ±15-minute tolerance blankets
~63% of the waking day — which is precisely why chance alone scores 1.9 of 4.

**Encouraging signal:** mean rank of matched predictions was **5.2** where random would be
10.5, and the two exact hits (Δ+3, Δ+1) were true local maxima rather than lucky
proximity. Both occurred during low-movement periods.

---

## 7. Errors made and corrected

Recorded because they shaped the conclusions:

1. **Unit errors.** Apple Health stores `WalkingSpeed` in **mi/hr** and `WalkingStepLength`
   in **inches**; the two "%" mobility types are stored as fractions. First pass reported
   a 2.37 km/h walking speed and 26 cm step length — both nonsense. Corrected: 3.81 km/h,
   67 cm.
2. **Device blending.** Resting HR was initially averaged across Garmin and Google Health,
   which differ systematically by 12–17 bpm. All figures now read one device end-to-end.
3. **Wrong comparison window.** The export ran at 17:22 but the watch last synced at
   15:00. Comparing "today at 17:22" against "yesterday at 17:22" overstated the deficit;
   the honest cut is 15:00 on both days.
4. **Timezone double-shift.** Fitbit's sleep JSON is local while its HR JSON is UTC.
   Adding the offset to both moved sleep to 04:39–12:40 and classified eight hours of
   sleep as waking.
5. **Recovery-gate floor.** A 3-minute lower bound silently rejected every clean sharp
   spike, including two of three confirmed cigarettes. Recovery is measured from t+2, so
   only the upper bound discriminates.
6. **The amplitude error** (§4) — the most consequential.

---

## 8. Conclusions

1. **Three of four target signatures are not implementable** from consumer wearable
   exports. HRV is sleep-only by platform design; raw accelerometry is never exported;
   SpO₂ is sleep-only.
2. **The surviving signature is weak.** Sustained HR rise at confirmed cigarettes sits at
   the 58th–78th percentile of ordinary waking minutes.
3. **Shape outperforms magnitude**, and is the only method that recovered the weakest event.
4. **~2 of 7 cigarettes produce no detectable cardiovascular response.** This is a ceiling
   on recall, not a tuning problem.
5. **Nothing here is statistically validated.** n = 8 labels, 7 testable.
6. Realistic performance: **~60% recall, ~2× over-prediction, ±15 minute localisation.**
   Useful as a retrospective prompt; unusable as an automatic diary.

---

## 9. Next steps

**Immediate**
- Log continuously for two weeks (`./log.sh`) → ~130 labelled events, enough for a real
  train/test split and to *fit* the template time constants rather than assume them.
- Sync the watch immediately before exporting. Both untestable cases were coverage gaps,
  not algorithm failures.

**Highest-value hardware change**
- A **Polar H10** chest strap streams real RR intervals, enabling RMSSD at ~30-second
  resolution. The §3 analysis shows the `w2` term would then dominate the score. This is
  the single change most likely to make detection work.

**Unexplored, cheap**
- **Time since last cigarette** is currently unused. A hazard model (probability rises
  with elapsed time) may outperform the physiological detector on its own, and combines
  naturally with it.

---

## Reproducing

```bash
cd ~/dev/personal/smoking-detection
export FB_DIR="$PWD/raw/Takeout/Google Health"

python3 src/export_datasets.py    # rebuild tidy CSVs
python3 src/smoke_detect.py       # Method A — the Z-score / S formula
python3 src/smoke_shape.py        # Method B — shape matching (recommended)
```

Analysis-ready datasets and exploration ideas: [`claude-science/`](claude-science/README.md).
