Detecting smoking events from consumer wearable data
What a consumer smartwatch can and cannot tell you about when — and how much — I smoked
Roman Kukhalashvili · 11 August 2026 · updated 2 September 2026
Abstract
Consumer wearables are often proposed as passive detectors of health behaviour. I tested whether my own cigarettes could be recovered from wrist heart rate without self-report, using a Google Fitbit Air and a Garmin across eleven days, against thirty-six logged cigarette times — sixteen of them within the days this analysis scores — one confirmed cigarette-free day and one heavy session. The two findings most likely to outlive this paper are structural. First, three of the four acute signatures proposed in the literature cannot be computed from consumer exports at all: heart-rate variability and blood-oxygen saturation are recorded during sleep only by platform design, and the exported accelerometer is a one-hertz scalar that cannot carry the hand-to-mouth gesture. Second, roughly two in seven of my cigarettes produced no measurable cardiovascular response — a ceiling no algorithm lifts. Within those limits, a detector that scores the shape of the heart-rate response after regressing out movement separates cigarettes from matched ordinary minutes at AUC 0.794 (circular-shift p = 0.008, rising toward 0.04 under the most pessimistic multiplicity accounting) — a figure that pools live-logged labels separating at 0.893 with recalled ones at 0.678. Precision is far too poor to ship. Four instruments — a two-stage detector, its stripped-down variant, a sleep-based dose score and a count-conditioned decoder — were frozen on 20 August with success criteria written before any test data existed, then evaluated once against ten unseen live-logged days and sixty-five cigarettes. All four failed. The best detector reached 46% precision against a 50% bar while separating from chance decisively (p = 0.0015); the overnight dose signal did not replicate (ρ = +0.07); and a decoder handed the day's true cigarette count as an oracle scored at exactly chance, which says the localising information is not in the trace at any price. I report four results I withdrew — the last of them, the overnight dose signal, killed by this very test — four adversarially-found errors that had all biased the result against the detector, and a blind verdict of nought for four.
- Signatures computable: 1 of 4 (HRV, SpO₂ and gesture unavailable or too coarse — by platform design)
- Silent cigarettes: ~2 in 7 (no cardiovascular response at all; the recall ceiling no algorithm lifts)
- Separation (AUC): 0.794 (pools two label qualities (0.893 live / 0.678 recalled); p = 0.008–0.04)
- Blind precision: 46% (best of four frozen instruments, 10 unseen days — against a pre-registered 50% bar. Failed)
- Pre-registered arms passed: 0 of 4 (detector, stripped variant, sleep dose score, oracle-count decoder — all failed their written criteria)
1 What I set out to do
I wanted to know whether my own smoking could be detected from a consumer wearable without me logging anything. I am publishing it — including the parts that failed — because the interesting result here is a negative one, and negative results in this area mostly go unpublished, which means the same dead ends get re-walked. The target signatures came from the applied literature: a sudden heart-rate rise of 10–30 bpm within a few minutes of ignition; a drop in heart-rate variability as nicotine drives sympathetic tone; a hand-to-mouth gesture recoverable from accelerometry; and a transient dip in blood-oxygen saturation from carbon monoxide.
The pharmacology sets the timescale I needed to match. Inhaled nicotine reaches the brain within seconds, and controlled infusion studies measure a rise of roughly 10 bpm at one minute and 13 bpm at two. Heart rate begins falling around twenty minutes after the last puff, and the size and duration of the response scale with nicotine content — a high-nicotine cigarette produced a 32% rise in cardiac output persisting an hour, against 13% for five minutes for a low-nicotine one. A cigarette is smoked over five to seven minutes, so the cardiovascular peak tracks the accumulated dose and arrives after ignition rather than at it. Those figures fixed my response template in advance, which mattered later.
It is worth being clear about what is already known. A systematic review of wearable smoking detection finds that no system reaches full accuracy and that most validation is done in laboratories. The strongest free-living result I could find, 84.9% agreement with self-report, came from a purpose-built multisensory rig combining chest ECG, a bioimpedance breathing sensor and a chest accelerometer. Wrist photoplethysmography alone — what a consumer watch gives you — is highly susceptible to motion artifact. I therefore expected this to be hard, and it was.
2 What the data can actually support
Before modelling anything I took a full type census of my Apple Health export — 487,447 records, heart rate reaching back to June 2025 — and of the Google/Fitbit Takeout. The point was to establish which channels exist, at what cadence, and whether they cover the days I care about.
| Signature | Status | Limiting factor |
|---|---|---|
| Acute heart-rate rise | Computable | 60 s relay; 2 s in Takeout |
| Heart-rate variability drop | Not computable | Sleep-only by platform design |
| Hand-to-mouth gesture | Computable, insufficient | 1 Hz scalar magnitude, not a 3-axis waveform |
| Blood-oxygen dip | Not computable | Sleep-only, single night |
2.1 Why HRV and SpO₂ aren't available while awake
The answer reframes the problem. Heart-rate variability is in the archive at 4.9 samples per day, about a hundred times too sparse to resolve a five-minute response, and it stops entirely on 3 May 2026 when my Apple Watch stopped syncing. Fitbit computes its own variability — and blood-oxygen — during sleep only, on-device; the waking numbers are never produced, so no export, Takeout archive or API tier can contain them. Blood-oxygen exists in my record for seven days in July 2025 and never since. None of this is a settings problem, and it is the answer to a family of questions people usually ask after failing at this themselves: there is no daytime Fitbit HRV, anywhere.
2.2 Provenance, and one check worth doing
The exports disagree about timezones in three different ways — Fitbit's heart-rate, step and temperature files are UTC, its sleep file is already local, Apple Health carries an explicit offset — and applying one correction to all of them classified eight hours of sleep as waking in an early run. Everything reported here is local time, resolved by cross-referencing the devices' own records.
I later switched data source from Takeout to the Apple Health export, because Apple Health carries the Fitbit record forward while Takeout would need re-requesting. Before trusting it I checked the relay against the original over 1,684 overlapping minutes: correlation 0.995 at zero lag, mean absolute difference 0.68 bpm. The relay is effectively lossless.
3 Method
The detector is built so that nothing is fixed to my data by hand. Baselines are causal rolling medians over the preceding hour, and deviations are scaled by a rolling median absolute deviation, so 'unusual' is measured in my own units on that day rather than against a literal bpm number. Every statistic uses strictly past samples, so the same code can run on tomorrow's export.
The central idea is exertion residualisation. The one reproducible failure in the earlier work was that walking swamps the nicotine response — a cigarette I walked to was missed with 104 concurrent steps. Steps are measured, so instead of penalising movement with a hand-tuned weight, I regress heart rate on an exponentially weighted recent step load using a robust slope, and score only the part of the rise that movement does not explain. The fitted slopes are substantial: +0.18 bpm per unit step load on Garmin, +0.42 on Fitbit.
The response itself is scored by correlation against a template with a flat lead-in, a fast rise and a slow decay. Its constants are fixed in advance from the pharmacology in §1 — onset 5 minutes, rise 1.5 minutes, decay 20 minutes — not fitted to my labels. Asleep minutes are excluded, which is structural rather than tuned. A prior over hour-of-day is learned from training days only. The alarm threshold is calibrated so that alarms on the training days match the cigarettes actually logged there, which lets a quiet day come out quiet. The implementation is in detector.py and features.py.
For significance I use a circular-shift null: the emitted alarm train is rotated rigidly within the day, which preserves the number of alarms and their spacing and destroys only their alignment to the labels. Naive shuffling is anti-conservative for autocorrelated series — in a companion analysis of my own overnight variability, a naive permutation gave p < 0.0001 where a circular shift on the same data gave p = 0.011.
4 What I got wrong
First, amplitude. I reported +28 bpm spikes at confirmed cigarettes; that number was the maximum of two-second samples — a transient dominated by sensor noise. Measured properly, as a one-minute median against a ten-minute baseline, the sustained rise at three confirmed cigarettes was +12, +6 and +6 bpm: the 78th, 60th and 58th percentile of my ordinary waking minutes.
Second, my base rate — wrong twice. I calibrated the first detector to a self-reported 8–12 cigarettes a day; proper logging showed 2, 1, 1, 4 and 0. Live logging later showed those recalled counts were themselves undercounts — real days run five to eight (§11). A fixed daily quota built on a wrong prior guarantees false alarms on a light day and cannot represent a zero day at all.
Third, selection. My first pass searched 1,080 configurations and reported the best, which reaches about 0.80 — but the honest null for a best-of-1,080 result is the best of 1,080 nulls, and computing that correction consumed the entire effect. I stopped searching. The template constants were fixed from pharmacology before the model ran and have not moved since.
Fourth, the headline is not pre-specified, and I will not call it that. After the model was first tested, three things changed in response to the audit in §7 — the control fence, the sensor rule, the null construction — and every one moved the number up. Two were demonstrable bugs whose correction was forced; the third was prompted by a statistic I was reading because the result was disappointing, and I cannot honestly claim otherwise. The correct label for 0.794 is post-hoc-corrected confirmatory, and the multiplicity range in §5 prices it.
Three configuration changes during the audit moved separation from 0.643 to 0.708 to 0.794; the full ledger of which run produced which number, on which sensors and fences, is in the appendix. And mid-analysis I published two results I later withdrew — a separation of 0.853 at p = 0.023, and 50% recall at p = 0.003, both artifacts of coverage confined to watch wear time. A paper that lists only the results that held is not showing you its error rate.
5 Results
Against awake, hour-matched control minutes drawn from the same days, the model separates cigarettes from ordinary minutes with an area under the curve of 0.794 — where 0.5 is no information — at circular-shift p = 0.008. Mean within-day separation is 0.901. This figure uses physiology alone; adding the learned hour-of-day prior moves it only to 0.799, which matters for a reason given below. (Of thirty-six logged cigarettes, sixteen fall in the scored window and thirteen have a complete detector window.)
A circularity worth bounding. The 12–17 August times were recalled days later, and a recalled time is likely reconstructed from routine — which is the same routine the hour-of-day prior learns. That is a genuine circularity. Its size is limited: the headline 0.794 uses no prior at all, so the prior can contaminate at most the 0.005 it adds. The operating curve in Table 2 does use the prior, so the concern applies there in full.
Splitting the labels by how they were captured bounds it. The eight logged close to the event on 10–11 August separate at AUC 0.893; the eight recalled several days later separate at 0.678. If recalled times were being reconstructed from a routine the model had also learned, the recalled subset should score higher than the live one. It scores substantially lower, which is what label noise looks like rather than circularity. Both subsets are six or seven usable positives, so this bounds the concern rather than settling it — and it means the headline 0.794 is an average across labels of two quite different qualities.
Multiplicity. Two fusion rules were tested, and §4 lists three further post-hoc corrections. Counting the fusion rules alone gives p ≈ 0.015. Counting all five decision points as independent tests — the most pessimistic reading — gives p ≈ 0.04. The result survives either way, but not by much, and I would not describe it as robust.
Turning that into alarms is less flattering, because a detector has to commit to times. Thresholds below are calibrated on the other days only, so each day's recall is out-of-sample.
| Alarms/day | Tolerance | Recall | Chance | p | False alarms on the free day |
|---|---|---|---|---|---|
| 3 | ±15 min | 4 of 16 (25%) | 1.8 (11%) | 0.104 | 2 |
| 8 | ±30 min | 9 of 16 (56%) | 5.9 (37%) | 0.058 | 4 |
| 12 | ±10 min | 9 of 16 (56%) | 3.8 (24%) | 0.005 | 4 |
| 12 | ±15 min | 12 of 16 (75%) | 5.2 (33%) | <0.001 | 4 |
Twelve alarms a day against the one to four cigarettes I had recalled for those days — a count §4 shows was itself an undercount — puts precision near 16%. With a single confirmed cigarette-free day the false-positive rate has no interval at all: a second such day could plausibly produce 1 or 10 alarms, and nothing here distinguishes those cases.
5.1 Does the template earn its place?
§4 shows that raw amplitude cannot separate smoking from ordinary life, which is the argument for scoring shape instead. But that was measured before exertion residualisation. Once movement is regressed out, the residual might already carry the signal, and the template could be decorative. That is the ablation the method most needs, so here it is.
| Scoring rule | Separation (AUC) | Labels scored |
|---|---|---|
| Template shape × response size (the paper) | 0.794 | 13 |
| Template shape alone, size discarded | 0.779 | 13 |
| Residual size alone, no template | 0.676 | 13 |
All rows use the same controls and the identical 13 labels. Size alone scores 0.719 when given every label it can reach (15) — the more flattering of its two figures; restricting it to the same 13 makes it worse, so the template's like-for-like margin is wider than the loose comparison suggests. Two admissions: I asserted this claim for two phases of the project without testing it, and only ran the ablation when a reviewer asked for it; and my first version of this table compared 13 labels against 15, which quietly flattered the weaker method.
6 The gesture hypothesis, and why it fails
An early conclusion of mine that no accelerometer data existed was wrong, and correcting it was the largest single revision of this work. Fitbit exports a per-second accelerometer change magnitude alongside 30-second mean orientation — 100,890 samples covering every logged cigarette.
So I tested the gesture directly: seven motion features at the cigarettes against 894 matched waking control windows, each with a 20,000-trial permutation test, in motion_probe.py.

Puff periodicity, the direct test of the hand-to-mouth signature, measured 0.26 during cigarettes against 0.25 in ordinary life.
7 An adversarial review of my own analysis
Having concluded that the model showed no signal, I asked an independent agent to refute that conclusion rather than confirm it, giving it the code, the labels and three specific attacks I thought most likely to break my work. It reproduced my headline exactly and then found four errors. All four had biased the result against the detector, which is the direction I would not have caught by looking for flattering mistakes.
- Control fence too short: score windows reach 42 minutes, I excluded controls to 36, so 16% of controls carried real response. Fixed: 0.643 → 0.708.
- The null discarded positives: free rotation dropped labels into coverage gaps — draws averaged 6.2 usable cigarettes against the observed 11, inflating p. Fixed by rotating within scorable minutes.
- The two-watch rule was the real bottleneck (§5): 4.6% of awake minutes scorable, five labels lost — a coverage problem masquerading as rigour.
- Two defences were doing nothing: hour-matching moved separation by 0.014; the awake filter, not at all.
8 Timing: a dead end and a useful one
Time since the previous cigarette fails exactly when it matters. Fitted as a hazard model it appeared to beat the physiological detector at 0.688 — while computing elapsed time from the held-out day's own logged cigarettes, the very thing a detector exists to discover. Run honestly autonomous it scores 0.377, worse than a coin flip, and combining it with physiology degrades physiology. The full three-regime table is in the appendix.
One row survives: told only the first cigarette of the day, timing scores 0.597. My same-day intervals ran 62–243 minutes, mean 136, CV 0.42 — I treat that as description, not evidence — but the first cigarette lands anywhere in a ten-hour window, and once it is known, the rhythm becomes usable. That is the argument for a semi-supervised design: the user logs one cigarette, the detector infers the rest — a more honest product than full automation, and a measurably better one.
Ten intervals across six days. Regular enough to be useful once the day has started; the first cigarette is the one the detector cannot anticipate.
9 What a detected and an undetected cigarette look like
Figure 3 contrasts the two regimes the detector has to separate, both aligned on minutes from the logged time. The 09:06 cigarette produced a clear sustained plateau. The 13:52 cigarette produced nothing distinguishable from baseline drift — and it was logged by me at the time, so this is not a labelling error.
This is the ceiling on the whole approach. Roughly two in seven testable cigarettes produce no cardiovascular response at all. No threshold, template or weighting recovers a response that did not happen.
9.1 How this was produced, and who checked it
The analysis was written and run with Claude Opus 5 in an agentic loop; the adversarial review in §7 ran separately in Claude Science against the same repository, and every number here comes from code in the archive rather than by hand. A reviewer from the same model family is independent of me but not of the analysis — and the record shows exactly that: all four errors it found were implementation bugs inside our shared framing, and none ever questioned whether a heart-rate template is the right object, or separation against hour-matched controls the right metric. I have not controlled for that, and no internal review can. The one test that does not inherit the framing is the pre-registered blind window in §11 — which is why it exists, and it is the test that ended the project. It tested the instruments, though, not the premise: all four arms score through the same heart-rate template or the same overnight HRV level, so a window that fails them says nothing about whether a template-shaped response was ever the right object to look for. A post-verdict audit of the three harness scripts on 2 September — a third model, with no hand in writing or running them, working to a scope written before its first finding — re-ran the registered commands, reproduced every number, and found two bugs in the sleep-arm script: a flag that never applied and a rank correlation that mishandled ties. The corrected Toll and respiratory numbers are the ones shown in §11. No verdict moved.
10 Discussion
The signal is real and it is small. Cigarettes sit measurably higher in my detector's score than matched ordinary minutes, and the separation clears its null (p = 0.008 for the model as run, 0.015 to 0.04 once the corrections in §4 are counted). But converting that into a usable alarm stream costs about twelve alarms a day to catch twelve of sixteen cigarettes, which is one true positive in six alarms. That is a retrospective prompt — 'you probably smoked around two o'clock, does that sound right?' — and not a diary.
The most useful thing I learned is procedural rather than physiological. I spent most of this work convinced my constraint was the number of labels, and recommended logging a hundred events. That was wrong. The constraint was that a rule I had adopted for rigour — demanding two sensors agree — had blinded the detector for 95% of the day. Fixing coverage was worth more than tripling the labels would have been, and I only found it because I asked someone else to attack the analysis instead of checking it myself.
Two of my four proposed signatures remain permanently out of reach on this hardware, and I do not think consumer wrist optics will close that gap. The free-living result that works uses a chest strap, a breathing sensor and a chest accelerometer, and the instrumented-lighter approach sidesteps physiology entirely. A wrist device is simply not the right instrument for this measurement, and the honest version of this project ends with a device recommendation rather than a better algorithm.
11 The blind test: nought for four
Four instruments were frozen on 20 August, each with a numeric bar written before any test data existed. Two were detectors: the two-stage v3 as the primary endpoint, its stripped-down variant v4 as secondary at α = 0.025. One was a sleep score, the Toll Score, for daily dose. The fourth was a count-conditioned decoder — handed the day's true cigarette count, asked only which minutes they were.
The window ran from 21 August. On 31 August I supplied the log and two fresh exports. The evaluation ran once: one scripted pass, no reruns, no variants.
All four failed. Ten scored days, sixty-five live-logged cigarettes.
The stripped detector beat the elaborate one. v4 has almost no machinery — no classifier, no rival templates, no calibrated probability. It beat v3 on precision, on recall, and on separation from chance. The extra apparatus in v3 was not adding signal. It was adding overfit, and cross-validation flattered it in a way that ten unseen days did not. v4 failed too, at 46% against 50%, and it is a secondary arm. But it is the only shape worth carrying forward.
Better than chance and good enough to use are different claims. v4's p of 0.0015 says its alarms are not randomly placed. That is the strongest evidence of real physiological signal in this paper. Its precision of 46% says the alarm is still wrong more often than it is right. Both are true at once. A single-criterion pre-registration would have let me report this as a success, which is why the bar had two halves.
The night signal did not replicate. On 20 August the overnight HRV curve was the most promising thing here: dose tracked its level at ρ = −0.84 across six nights. Eight blind nights returned ρ = +0.07, p = 0.44. The worst of them followed the lightest day of the window — one cigarette — and scored near the top of the dose range. Six nights was never enough. It is this project's fourth withdrawal.
A nightly respiratory-rate channel came out directionally right, at ρ = +0.56, p = 0.09 over seven nights. It carried no success bar. It is also the sole survivor of a family that went nought for four, which is exactly when a p of 0.09 should be believed least. It is a hypothesis for a future window, not a result.
Can the instruments at least count? Also pre-registered, also no. The registered estimator missed the daily tally by 2.48 cigarettes on average. Predicting the window average of 6.5 every single day scores 2.00. Its in-sample rank correlation of +0.80 on four days fell to +0.40 on ten.
What the failure is worth. This project retracted three results in its first two phases by tuning after seeing the data. The defence built afterwards was a written bar, frozen constants, and one scripted pass. It worked. Four instruments went in and four came back negative, and none could be rescued afterwards, because the criteria for rescue did not exist. The out-of-sample precision of 42% also landed on the audit's nested estimate of 42.9%, not the 50% that cross-validation had advertised. The audit was right, and this is its receipt.
The pre-registration as it stood on 20 August is in the appendix, unchanged, so the bar can be checked against the verdict.
So the paper ends where it was pointed. A wrist device cannot passively detect when I smoke. Three of the four signatures never reach the export (§2). Two in seven cigarettes leave no trace at all (§7). What remains is wrong more often than it is right, and no side information recovers the timing. The honest end of this work is a device recommendation, not a better algorithm.
Each instrument, its number and the bar it was given are below.
| Instrument | Result | Pre-registered bar | Verdict |
|---|---|---|---|
| v3 — two-stage detector (primary) | precision 42%, p = 0.043 | ≥50% and p < 0.05 | Failed |
| v4 — stripped variant (secondary) | precision 46%, p = 0.0015 | ≥50% and p < 0.025 | Failed |
| Toll Score — sleep dose | ρ = +0.07, p = 0.44 | ρ > 0, p < 0.05, ≥8 nights | Failed |
| Count-conditioned decoder | 31% hits vs 31% chance, p = 0.57 | p < 0.025 | Failed |
| v2 — reference, no bar | precision 31%, p = 0.15 | — | — |
| Nightly respiratory rate | ρ = +0.56, p = 0.09, n = 7 | secondary readout, no bar | — |
v4 is the row worth staring at: it separates from chance by an order of magnitude more than the elaborate detector it was stripped down from, and it still fails, because 46% precision means the alarm is wrong more often than it is right.
Limitations
- One participant. One hundred and nine labelled events across twenty-two days — sixteen in the days §5 scores, sixty-five in the blind window of §11, plus one free day and one session. Nothing here generalises beyond me. The audit's phrasing was that no reanalysis of the original candidates would settle anything and only the blind window could; the blind window has now reported, and it settled the question against the detector.
- The 12–17 August times were recalled several days later, not logged live, so they carry an unknown error of a few minutes. That error inflates the tolerance needed and therefore inflates the chance baseline too.
- Precision is roughly 16% at the reported operating point, estimated from about 72 alarms and 12 hits across the six labelled days. The clean false-positive count is far weaker: it rests on one confirmed cigarette-free day, which produced 4 alarms. A second such day could plausibly produce 1 or 10, and nothing here would distinguish those cases.
- The headline result is post-hoc-corrected confirmatory rather than pre-specified — see §4. Three corrections were applied after the model was first tested, all of which moved the number up. After multiplicity, p = 0.008 should be read as 0.015 at best and 0.04 at worst.
- Recalled cigarette times are probably reconstructed from routine, which is what the hour-of-day prior also learns. The headline separation avoids this by using no prior; the operating curve does not.
- The heavy session of 15–16 August is effectively untestable: Fitbit covers 28 of its 420 minutes, so the densest positive window in the whole dataset contributes almost nothing.
- Sleep exclusion relies on the devices' own staging, which put me asleep at 09:15 on a morning I recorded smoking — so at least one label was discarded by a staging error rather than a coverage gap.
- The v3 precision figure is cross-validated, not blind, and its operating point was chosen by a rule that guaranteed its own headline — caught by the audit (appendix). The corrected nested estimate is 42.9%, with a bootstrap interval that reaches the chance floor.
- The overnight-HRV dose relationship is exploratory: six dose nights, seven features tested, family-wise p = 0.21. Two of the six doses are lower bounds. It did not replicate on eight blind nights (§11) and is withdrawn.
- The window was registered for fourteen days and closed at ten. The reason was scheduling, not data: no result had been computed when it closed, and dropping the final partial day was fixed in writing beforehand. That protects the false-positive rate but costs power. Three of the four arms sit at or near their chance floor, so more days could only have made the failures more certain.
- The blind window contained no cigarette-free day, so the unbounded false-positive rate above is still measured on the single free day of 16 August and is only loosely constrained. The nearest thing to a second one is 29 August, a one-cigarette day, on which v3 fired four alarms and v4 two, against window averages of 3.3 and 3.7 a day. So the floor is not invisible after all: it is roughly two to four alarms a day, and it does not fall when there is almost nothing to detect.
- The blind labels are my own live log, supplied in one batch rather than appended cigarette by cigarette as the protocol envisaged. Three days looked light against the 5–8/day norm. All three were queried and confirmed complete before any evaluation ran. Had one been dropped after seeing a result, the test would have been worthless.
- Blind precision of 46% is one point estimate on 65 labels from one participant. Its interval comfortably contains 35% and 57%. Nothing here distinguishes those. None of them clears the bar.
- The stripped detector was not fully frozen. Its threshold, exertion slope and circadian prior were fixed on 20 August, but the median and spread that put its raw score on the threshold's scale are recomputed from whatever export is loaded, and on 31 August half the minutes doing that were blind days. Freezing them at the 20 August export gives the identical 17 of 37; freezing at midnight on 21 August gives 17 of 40, or 42.5%. The registered 46% is the favourable end of a narrow range, and nothing in the range reaches 50%.
- Eleven of the sixty-five blind cigarettes fall where the wrist trace has no scorable minute within fifteen minutes of them. Recall was therefore capped at 83% by coverage before any physiology, and at roughly 59% once §7's two-in-seven silent cigarettes are taken off that.
- The circular-shift null behind the p-values in §11 is conservative: the eleven labels that cannot be hit in the observed count can be hit under the null, so the significance half of each bar was, if anything, understated. Under an alarm-rotation null v3's p falls from 0.043 to 0.001. It changes nothing, because both detectors failed on precision.
Further work
- The wrist-detector line of work is closed, not paused. Four frozen instruments failed a pre-registered blind test, and the count-conditioned decoder failing at chance means the timing information is absent rather than merely hard to extract — so a v5 built on the same channel would be re-running a settled experiment.
- A chest strap streaming real beat-to-beat intervals is now the only proposal here with a live rationale — it is the one signature that has never been measurable while awake, and everything computable without it has been tried and has failed.
- The semi-supervised variant — the user logs one cigarette and the detector infers the rest — measured well in §8 at 0.597 against 0.377, but the blind decoder is the stronger version of that experiment: given the whole day's count rather than one anchor, it scored at chance. I no longer expect the one-anchor version to work, and I am not proposing it.
- A chest strap streaming real beat-to-beat intervals would make heart-rate variability available while awake. By §2 it is the one proposed signature that has never been measurable, and by the earlier analysis it was the load-bearing term all along.
- Model the two smoking modes separately. Ordinary days run about two hours between cigarettes; the one heavy session ran nearer forty minutes, and a detector tuned on the first will under-count the second.
- Wear both devices consistently. Several labels were lost to coverage gaps rather than to any failure of the algorithm.
Code and data
What this is for. Everything below is published so that the method and, more usefully, the dead ends can be reused. Three of the four signatures in the literature are not computable from consumer exports, the hand-to-mouth gesture is absent at this sampling resolution, and time since the last cigarette is worse than useless when run autonomously. Each of those took real work to establish and none of them needs establishing twice. If you are building in this area, the negative results are the part worth taking.
What it is not for. Do not ship this detector. On its own blind test it failed every pre-registered criterion: at the best operating point it is wrong more often than it is right (46% precision on 65 unseen labels), on one participant, with no external validation. A cessation or prevention app that told a user they had smoked and was wrong more than half the time would be worse than not having the feature — and consumer health claims carry regulatory obligations this work comes nowhere near meeting. Treat the code as the description of an experiment, not as a component.
Licence. Code is released under the MIT licence; the text, figures and derived data under CC BY 4.0. Attribution is welcome, but the more useful request is this: if you extend the work, publish what did not work alongside what did.
Code and derived results are published in full. Raw data is withheld: these are timestamped cigarettes and continuous heart rate for a named individual, and publishing them would make a behavioural record permanently public for no scientific gain. What is here is enough to check every statistic in the paper, including the exact permutation counts and the measured values behind the figures with clock times reduced to relative offsets. It is not enough to re-run the pipeline from raw, which needs my own export.
- Detection code — The full pipeline — export loader, feature construction, both frozen detectors (the audited two-stage v3 and the stripped v4), the activity-state channel, the confirmatory tests, and all three pre-registered blind-test harnesses — the detector evaluation, the Toll Score and the count-conditioned decoder — plus the post-verdict audit script.
- Findings log — Every experiment of the second phase in the order I ran them, including the failures and the two results I withdrew. A phase-two snapshot; the third phase is summarised in §11.
- Blind-window results — The verdict of §11 in full: per-day cigarette counts against each instrument's alarms, both detectors at both tolerances with their permutation nulls, the Toll Score and respiratory readout, and the registered count-estimation grading. Aggregates only — the blind cigarette times are withheld.
- Adversarial audit — All 18 tests from the independent review with raw and Holm-adjusted p-values — the numbers behind §7; the full regime table is in the appendix.
- Third-phase audit — The independent audit of the two-stage detector: exact reproduction, nested cross-validation, bootstrap intervals, contamination sensitivity, ablations, and the Holm accounting — the numbers behind the third-phase write-up in the appendix.
- Permutation tests — Seven accelerometer features, cigarette and control means, p-values and trial counts — the numbers behind Figure 1.
- Blind-test scoring — The original held-out predictions against the withheld log at three tolerances, with permutation-derived chance baselines.
- Amplitude distribution — Percentiles of sustained heart-rate rise across all waking minutes, against the confirmed cigarettes.
- Figure 3 source data — The plotted traces as measured, with clock times replaced by minutes-from-event.
- Threshold sweep — Firing rate and recall across thirteen score thresholds for the first-phase detector — the precision/recall tradeoff in full.
- Appendix — Everything demoted in the restructure, preserved in full: the configuration ledger, the hazard-model regimes, the six-panel audit figure, the retracted Poisson test, and the complete third-phase interim write-up with its registered blind-test constants.
- Full technical report — The long-form write-up of the first phase, including the six errors made and corrected during it.