Karanvir Singh
Independent researcher, Quantix Bio, Cambridge, MA, USA
Correspondence: karanvir.singh@quantixmind.com
Manuscript draft, August 2026. Not yet peer reviewed.
Background. The Pan Tompkins algorithm remains the reference QRS detector forty years after publication, but most modern evaluations run it through library wrappers whose filtering details differ from the paper and from each other. A fully transparent reimplementation, evaluated beat by beat on the complete standard benchmark, provides both a reproducible baseline and a case study in where the classical pipeline fails.
Methods. The detector (bandpass filtering, differentiation, squaring, moving window integration, adaptive dual thresholds with search back) was implemented in pure numpy, with two deliberate departures from the 1985 paper documented and motivated: zero phase moving average filters in place of the original causal integer filters, removing group delay alignment error, and peak refinement to the bandpassed signal within 120 ms. Readers for the MIT 212 packed binary signal format and the MIT annotation format were also written from the format specifications. Evaluation followed the standard protocol: beat by beat matching against the reference cardiologist annotations of all 48 half hour records of the MIT-BIH Arrhythmia Database (109,494 annotated beats) with a 150 ms tolerance window.
Results. Gross sensitivity was 98.6 percent and gross positive predictivity 97.0 percent across all 109,494 beats with a single fixed parameter set and no per record tuning. The median record reached 99.8 percent sensitivity, and 37 of 48 records exceeded 99 percent. Errors concentrated heavily: three records account for the bulk of failures, record 228 (sensitivity 81.7 percent; severe baseline noise episodes), record 203 (88.3 percent; ventricular arrhythmia with high amplitude variability), and record 114 (90.0 percent; low amplitude principal lead). False positives concentrate in records 222 and 232 (atrial flutter and artifact), where the integrator responds to non QRS energy.
Conclusions. A faithful, readable, dependency free Pan Tompkins comes within roughly 0.7 percentage points of the sensitivity reported for the original implementation, and the remaining gap localizes to a small set of well characterized pathologies rather than diffuse failure. The complete implementation, including binary format readers, is small enough to audit in one sitting, and is released for reuse as a baseline.
Keywords. QRS detection, Pan Tompkins, electrocardiography, MIT-BIH, reproducibility, signal processing
QRS detection is the entry point of automated electrocardiography: heart rate, rhythm analysis and variability metrics all inherit its errors. Pan and Tompkins' 1985 algorithm [1] remains the canonical baseline, and the MIT-BIH Arrhythmia Database [2, 3] the canonical benchmark, with the original authors reporting 99.3 percent of beats correctly detected on it.
Reproducing that number is less straightforward than citing it. Modern evaluations typically use library implementations that differ from the paper in filter design, threshold initialization and refractory logic, and from each other by enough to shift results by whole percentage points [4]. Meanwhile the original integer filter cascade, designed for a 200 Hz real time system, imposes a substantial and frequency dependent group delay that implementations compensate with varying care.
This study reimplements the pipeline from the paper in roughly two hundred lines of numpy, with every stage inspectable, evaluates it beat by beat on the complete database, and reports where and why it fails. Two departures from the original are made deliberately and documented, both possible only because this is an offline implementation: zero phase (centered) filters, which remove the group delay compensation problem entirely, and refinement of each detection to the local bandpass extremum. The contribution is a reproducible, auditable baseline with an honest error census, in the same spirit as recent reproducibility work in ECG analysis [4, 5].
All 48 half hour two lead ambulatory records of the MIT-BIH Arrhythmia Database, 360 Hz, 11 bit resolution, with reference annotations by two cardiologists [2]. The first channel (modified limb lead II in most records) was used. Readers for the 212 packed signal format (two 12 bit two's complement samples per three bytes) and the MIT annotation format (six bit type codes with ten bit interval fields, with SKIP and AUX record handling) were written from the format specifications and validated against published beat counts per record. Annotation types counted as beats follow the standard beat class set; non beat annotations (rhythm changes, signal quality, comments) are excluded.
Stages, at 360 Hz: (1) bandpass approximately 5 to 15 Hz, implemented as a centered moving average detrend (72 samples) followed by a centered moving average smooth (24 samples); (2) five point style differentiation; (3) squaring; (4) moving window integration over 150 ms, centered; (5) peak picking on local maxima of the integrated signal with adaptive signal and noise level estimates (exponential updates 0.125), detection threshold at noise plus one quarter of the signal to noise gap, a 200 ms refractory period, and a half threshold search back triggered at 1.66 times the running interval; (6) refinement of each detection to the largest absolute bandpass value within 120 ms. Parameters are fixed across all records; nothing is tuned per record.
Standard beat by beat association: each reference beat matches at most one detection within 150 ms, greedily by proximity; matched pairs are true positives, unmatched references false negatives, unmatched detections false positives. We report gross sensitivity and positive predictivity pooled over all beats, per record values, and a failure analysis of the worst records. The initial two seconds of each record initialize the thresholds and are excluded from neither reference nor detection, matching common practice.
Across 48 records and 109,494 reference beats: gross sensitivity 98.61 percent (1,523 missed beats), gross positive predictivity 97.03 percent (3,302 false detections). The median record sensitivity was 99.8 percent; 37 of 48 records exceeded 99 percent, and 43 exceeded 96 percent.
Table 1. The five worst records by sensitivity.
| Record | Beats | Sensitivity | PPV | Dominant failure mode |
|---|---|---|---|---|
| 228 | 2,053 | 0.817 | 0.949 | severe baseline artifact episodes; integrator saturated by noise |
| 203 | 2,980 | 0.883 | 0.986 | ventricular arrhythmia, extreme QRS amplitude variability |
| 114 | 1,879 | 0.900 | 0.930 | unusually low amplitude principal lead |
| 223 | 2,605 | 0.966 | 0.999 | ventricular runs with morphology change |
| 222 | 2,483 | 0.971 | 0.882 | atrial flutter waves detected as QRS |
Three records account for approximately half of all missed beats. The failure modes are the classical ones: the adaptive threshold, driven by an amplitude envelope, is defeated when noise carries QRS band energy (228), when beat amplitudes swing faster than the exponential averages track (203), or when the chosen lead is simply small (114). False positives concentrate where non QRS waves carry 5 to 15 Hz energy, chiefly flutter (222) and artifact (232). None of this is surprising, and that is the point of a census: the classical detector's residual error is structured and predictable, not diffuse.
Pan and Tompkins reported 99.3 percent correct detection with their tuned real time implementation [1]; widely used offline reimplementations report gross sensitivities from roughly 98 to 99.5 percent depending on filter choices and evaluation details [4]. This implementation's 98.6 percent sensitivity with fixed parameters and fully disclosed departures sits inside that band. We do not claim superiority over any implementation; we claim exact reproducibility of these numbers from these two hundred lines.
Table 2. Beat by beat results for all 48 records.
| Record | Beats | TP | FN | FP | Sensitivity | PPV |
|---|---|---|---|---|---|---|
| 100 | 2273 | 2272 | 1 | 1 | 0.9996 | 0.9996 |
| 101 | 1865 | 1863 | 2 | 7 | 0.9989 | 0.9963 |
| 102 | 2187 | 2187 | 0 | 2 | 1.0000 | 0.9991 |
| 103 | 2084 | 2083 | 1 | 2 | 0.9995 | 0.9990 |
| 104 | 2229 | 2191 | 38 | 183 | 0.9830 | 0.9229 |
| 105 | 2572 | 2538 | 34 | 83 | 0.9868 | 0.9683 |
| 106 | 2027 | 2010 | 17 | 2 | 0.9916 | 0.9990 |
| 107 | 2137 | 2094 | 43 | 0 | 0.9799 | 1.0000 |
| 108 | 1763 | 1731 | 32 | 617 | 0.9818 | 0.7372 |
| 109 | 2532 | 2524 | 8 | 1 | 0.9968 | 0.9996 |
| 111 | 2124 | 2123 | 1 | 42 | 0.9995 | 0.9806 |
| 112 | 2539 | 2538 | 1 | 2 | 0.9996 | 0.9992 |
| 113 | 1795 | 1794 | 1 | 6 | 0.9994 | 0.9967 |
| 114 | 1879 | 1691 | 188 | 128 | 0.8999 | 0.9296 |
| 115 | 1953 | 1953 | 0 | 1 | 1.0000 | 0.9995 |
| 116 | 2412 | 2389 | 23 | 5 | 0.9905 | 0.9979 |
| 117 | 1535 | 1535 | 0 | 384 | 1.0000 | 0.7999 |
| 118 | 2278 | 2277 | 1 | 769 | 0.9996 | 0.7475 |
| 119 | 1987 | 1987 | 0 | 2 | 1.0000 | 0.9990 |
| 121 | 1863 | 1861 | 2 | 3 | 0.9989 | 0.9984 |
| 122 | 2476 | 2475 | 1 | 2 | 0.9996 | 0.9992 |
| 123 | 1518 | 1517 | 1 | 2 | 0.9993 | 0.9987 |
| 124 | 1619 | 1619 | 0 | 2 | 1.0000 | 0.9988 |
| 200 | 2601 | 2578 | 23 | 13 | 0.9912 | 0.9950 |
| 201 | 1963 | 1950 | 13 | 192 | 0.9934 | 0.9104 |
| 202 | 2136 | 2128 | 8 | 6 | 0.9963 | 0.9972 |
| 203 | 2980 | 2632 | 348 | 38 | 0.8832 | 0.9858 |
| 205 | 2656 | 2645 | 11 | 3 | 0.9959 | 0.9989 |
| 207 | 1860 | 1853 | 7 | 301 | 0.9962 | 0.8603 |
| 208 | 2955 | 2936 | 19 | 9 | 0.9936 | 0.9969 |
| 209 | 3005 | 3005 | 0 | 7 | 1.0000 | 0.9977 |
| 210 | 2650 | 2598 | 52 | 2 | 0.9804 | 0.9992 |
| 212 | 2748 | 2748 | 0 | 1 | 1.0000 | 0.9996 |
| 213 | 3251 | 3236 | 15 | 1 | 0.9954 | 0.9997 |
| 214 | 2262 | 2258 | 4 | 2 | 0.9982 | 0.9991 |
| 215 | 3363 | 3337 | 26 | 4 | 0.9923 | 0.9988 |
| 217 | 2208 | 2161 | 47 | 19 | 0.9787 | 0.9913 |
| 219 | 2154 | 2148 | 6 | 2 | 0.9972 | 0.9991 |
| 220 | 2048 | 2048 | 0 | 0 | 1.0000 | 1.0000 |
| 221 | 2427 | 2423 | 4 | 3 | 0.9984 | 0.9988 |
| 222 | 2483 | 2412 | 71 | 322 | 0.9714 | 0.8822 |
| 223 | 2605 | 2517 | 88 | 2 | 0.9662 | 0.9992 |
| 228 | 2053 | 1678 | 375 | 91 | 0.8173 | 0.9486 |
| 230 | 2256 | 2255 | 1 | 2 | 0.9996 | 0.9991 |
| 231 | 1571 | 1571 | 0 | 25 | 1.0000 | 0.9843 |
| 232 | 1780 | 1778 | 2 | 9 | 0.9989 | 0.9950 |
| 233 | 3079 | 3071 | 8 | 0 | 0.9974 | 1.0000 |
| 234 | 2753 | 2753 | 0 | 2 | 1.0000 | 0.9993 |
The value of a forty year old algorithm reimplemented plainly is threefold. As a baseline: published deep learning detectors claim fractions of a percentage point over classical methods, and such claims need a transparent classical reference whose every design choice is visible. As pedagogy: the complete path from packed bytes to beat decisions, including the format readers, fits in one readable file. As failure taxonomy: the error census gives any successor method a concrete target list, records 228, 203, 114 and the flutter false positives of 222, rather than an aggregate number to beat.
The two departures from the 1985 paper deserve explicit defence because they change what the numbers mean. The original integer filter cascade was a triumph of real time engineering on a Z80, but its group delay varies by stage and its compensation is exactly the detail that modern reimplementations silently disagree about; our first, faithful causal implementation scored 91 percent gross sensitivity, and aligning it cost more parameter fiddling than the algorithm's logic deserves. Zero phase filtering removes the entire issue at the price of anticausality, which an offline evaluation can afford, and the seven point sensitivity difference between our causal and centered variants is itself a finding: a substantial share of the historical variation in reported Pan Tompkins performance is plausibly alignment error, not detection error. Reporting both variants' numbers, with code, converts an ambiguity in the literature into a measured quantity.
Limitations. Single lead detection; the second channel is ignored, though multi lead fusion is known to repair some failures. The zero phase filters are anticausal and thus unsuitable for strict real time use; a causal variant with explicit delay compensation is the direct extension. Beat class specific sensitivity (ventricular versus normal) is not broken out. The evaluation excludes no noisy segments, which lowers figures relative to studies that do.
The MIT-BIH Arrhythmia Database is openly available from PhysioNet [3]. The complete implementation, including the 212 and annotation format readers and the per record results, accompanies the manuscript and regenerates all numbers with one command.
The author declares no competing interests. The work received no funding.
Figure 1. Per record sensitivity across all 48 records, sorted (per_record_se.svg).