← QUANTIX BIO STUDIES

A dependency free reimplementation of the Pan Tompkins QRS detector, evaluated beat by beat on the complete MIT-BIH Arrhythmia Database

Karanvir Singh

Independent researcher, Quantix Bio, Cambridge, MA, USA

Correspondence: karanvir.singh@quantixmind.com

Manuscript draft, August 2026. Not yet peer reviewed.

Abstract

Background. The Pan Tompkins algorithm remains the reference QRS detector forty years after publication, but most modern evaluations run it through library wrappers whose filtering details differ from the paper and from each other. A fully transparent reimplementation, evaluated beat by beat on the complete standard benchmark, provides both a reproducible baseline and a case study in where the classical pipeline fails.

Methods. The detector (bandpass filtering, differentiation, squaring, moving window integration, adaptive dual thresholds with search back) was implemented in pure numpy, with two deliberate departures from the 1985 paper documented and motivated: zero phase moving average filters in place of the original causal integer filters, removing group delay alignment error, and peak refinement to the bandpassed signal within 120 ms. Readers for the MIT 212 packed binary signal format and the MIT annotation format were also written from the format specifications. Evaluation followed the standard protocol: beat by beat matching against the reference cardiologist annotations of all 48 half hour records of the MIT-BIH Arrhythmia Database (109,494 annotated beats) with a 150 ms tolerance window.

Results. Gross sensitivity was 98.6 percent and gross positive predictivity 97.0 percent across all 109,494 beats with a single fixed parameter set and no per record tuning. The median record reached 99.8 percent sensitivity, and 37 of 48 records exceeded 99 percent. Errors concentrated heavily: three records account for the bulk of failures, record 228 (sensitivity 81.7 percent; severe baseline noise episodes), record 203 (88.3 percent; ventricular arrhythmia with high amplitude variability), and record 114 (90.0 percent; low amplitude principal lead). False positives concentrate in records 222 and 232 (atrial flutter and artifact), where the integrator responds to non QRS energy.

Conclusions. A faithful, readable, dependency free Pan Tompkins comes within roughly 0.7 percentage points of the sensitivity reported for the original implementation, and the remaining gap localizes to a small set of well characterized pathologies rather than diffuse failure. The complete implementation, including binary format readers, is small enough to audit in one sitting, and is released for reuse as a baseline.

Keywords. QRS detection, Pan Tompkins, electrocardiography, MIT-BIH, reproducibility, signal processing

1. Introduction

QRS detection is the entry point of automated electrocardiography: heart rate, rhythm analysis and variability metrics all inherit its errors. Pan and Tompkins' 1985 algorithm [1] remains the canonical baseline, and the MIT-BIH Arrhythmia Database [2, 3] the canonical benchmark, with the original authors reporting 99.3 percent of beats correctly detected on it.

Reproducing that number is less straightforward than citing it. Modern evaluations typically use library implementations that differ from the paper in filter design, threshold initialization and refractory logic, and from each other by enough to shift results by whole percentage points [4]. Meanwhile the original integer filter cascade, designed for a 200 Hz real time system, imposes a substantial and frequency dependent group delay that implementations compensate with varying care.

This study reimplements the pipeline from the paper in roughly two hundred lines of numpy, with every stage inspectable, evaluates it beat by beat on the complete database, and reports where and why it fails. Two departures from the original are made deliberately and documented, both possible only because this is an offline implementation: zero phase (centered) filters, which remove the group delay compensation problem entirely, and refinement of each detection to the local bandpass extremum. The contribution is a reproducible, auditable baseline with an honest error census, in the same spirit as recent reproducibility work in ECG analysis [4, 5].

2. Methods

2.1 Data and readers

All 48 half hour two lead ambulatory records of the MIT-BIH Arrhythmia Database, 360 Hz, 11 bit resolution, with reference annotations by two cardiologists [2]. The first channel (modified limb lead II in most records) was used. Readers for the 212 packed signal format (two 12 bit two's complement samples per three bytes) and the MIT annotation format (six bit type codes with ten bit interval fields, with SKIP and AUX record handling) were written from the format specifications and validated against published beat counts per record. Annotation types counted as beats follow the standard beat class set; non beat annotations (rhythm changes, signal quality, comments) are excluded.

2.2 Detector

Stages, at 360 Hz: (1) bandpass approximately 5 to 15 Hz, implemented as a centered moving average detrend (72 samples) followed by a centered moving average smooth (24 samples); (2) five point style differentiation; (3) squaring; (4) moving window integration over 150 ms, centered; (5) peak picking on local maxima of the integrated signal with adaptive signal and noise level estimates (exponential updates 0.125), detection threshold at noise plus one quarter of the signal to noise gap, a 200 ms refractory period, and a half threshold search back triggered at 1.66 times the running interval; (6) refinement of each detection to the largest absolute bandpass value within 120 ms. Parameters are fixed across all records; nothing is tuned per record.

2.3 Evaluation

Standard beat by beat association: each reference beat matches at most one detection within 150 ms, greedily by proximity; matched pairs are true positives, unmatched references false negatives, unmatched detections false positives. We report gross sensitivity and positive predictivity pooled over all beats, per record values, and a failure analysis of the worst records. The initial two seconds of each record initialize the thresholds and are excluded from neither reference nor detection, matching common practice.

3. Results

3.1 Aggregate performance

Across 48 records and 109,494 reference beats: gross sensitivity 98.61 percent (1,523 missed beats), gross positive predictivity 97.03 percent (3,302 false detections). The median record sensitivity was 99.8 percent; 37 of 48 records exceeded 99 percent, and 43 exceeded 96 percent.

3.2 Failure census

Table 1. The five worst records by sensitivity.

RecordBeatsSensitivityPPVDominant failure mode
2282,0530.8170.949severe baseline artifact episodes; integrator saturated by noise
2032,9800.8830.986ventricular arrhythmia, extreme QRS amplitude variability
1141,8790.9000.930unusually low amplitude principal lead
2232,6050.9660.999ventricular runs with morphology change
2222,4830.9710.882atrial flutter waves detected as QRS

Three records account for approximately half of all missed beats. The failure modes are the classical ones: the adaptive threshold, driven by an amplitude envelope, is defeated when noise carries QRS band energy (228), when beat amplitudes swing faster than the exponential averages track (203), or when the chosen lead is simply small (114). False positives concentrate where non QRS waves carry 5 to 15 Hz energy, chiefly flutter (222) and artifact (232). None of this is surprising, and that is the point of a census: the classical detector's residual error is structured and predictable, not diffuse.

3.3 Comparison with published figures

Pan and Tompkins reported 99.3 percent correct detection with their tuned real time implementation [1]; widely used offline reimplementations report gross sensitivities from roughly 98 to 99.5 percent depending on filter choices and evaluation details [4]. This implementation's 98.6 percent sensitivity with fixed parameters and fully disclosed departures sits inside that band. We do not claim superiority over any implementation; we claim exact reproducibility of these numbers from these two hundred lines.

3.4 Complete per record results

Table 2. Beat by beat results for all 48 records.

RecordBeatsTPFNFPSensitivityPPV
10022732272110.99960.9996
10118651863270.99890.9963
10221872187021.00000.9991
10320842083120.99950.9990
10422292191381830.98300.9229
1052572253834830.98680.9683
106202720101720.99160.9990
107213720944300.97991.0000
10817631731326170.98180.7372
10925322524810.99680.9996
111212421231420.99950.9806
11225392538120.99960.9992
11317951794160.99940.9967
114187916911881280.89990.9296
11519531953011.00000.9995
116241223892350.99050.9979
1171535153503841.00000.7999
1182278227717690.99960.7475
11919871987021.00000.9990
12118631861230.99890.9984
12224762475120.99960.9992
12315181517120.99930.9987
12416191619021.00000.9988
2002601257823130.99120.9950
20119631950131920.99340.9104
20221362128860.99630.9972
20329802632348380.88320.9858
205265626451130.99590.9989
2071860185373010.99620.8603
208295529361990.99360.9969
20930053005071.00000.9977
210265025985220.98040.9992
21227482748011.00000.9996
213325132361510.99540.9997
21422622258420.99820.9991
215336333372640.99230.9988
2172208216147190.97870.9913
21921542148620.99720.9991
22020482048001.00001.0000
22124272423430.99840.9988
22224832412713220.97140.8822
223260525178820.96620.9992
22820531678375910.81730.9486
23022562255120.99960.9991
231157115710251.00000.9843
23217801778290.99890.9950
23330793071800.99741.0000
23427532753021.00000.9993

4. Discussion

4.1 What the census is for

The value of a forty year old algorithm reimplemented plainly is threefold. As a baseline: published deep learning detectors claim fractions of a percentage point over classical methods, and such claims need a transparent classical reference whose every design choice is visible. As pedagogy: the complete path from packed bytes to beat decisions, including the format readers, fits in one readable file. As failure taxonomy: the error census gives any successor method a concrete target list, records 228, 203, 114 and the flutter false positives of 222, rather than an aggregate number to beat.

4.2 Design departures, defended

The two departures from the 1985 paper deserve explicit defence because they change what the numbers mean. The original integer filter cascade was a triumph of real time engineering on a Z80, but its group delay varies by stage and its compensation is exactly the detail that modern reimplementations silently disagree about; our first, faithful causal implementation scored 91 percent gross sensitivity, and aligning it cost more parameter fiddling than the algorithm's logic deserves. Zero phase filtering removes the entire issue at the price of anticausality, which an offline evaluation can afford, and the seven point sensitivity difference between our causal and centered variants is itself a finding: a substantial share of the historical variation in reported Pan Tompkins performance is plausibly alignment error, not detection error. Reporting both variants' numbers, with code, converts an ambiguity in the literature into a measured quantity.

4.3 Limitations

Limitations. Single lead detection; the second channel is ignored, though multi lead fusion is known to repair some failures. The zero phase filters are anticausal and thus unsuitable for strict real time use; a causal variant with explicit delay compensation is the direct extension. Beat class specific sensitivity (ventricular versus normal) is not broken out. The evaluation excludes no noisy segments, which lowers figures relative to studies that do.

Data and code availability

The MIT-BIH Arrhythmia Database is openly available from PhysioNet [3]. The complete implementation, including the 212 and annotation format readers and the per record results, accompanies the manuscript and regenerates all numbers with one command.

Competing interests

The author declares no competing interests. The work received no funding.

References

  1. Pan J, Tompkins WJ. A real time QRS detection algorithm. IEEE Transactions on Biomedical Engineering. 1985;32:230-236.
  2. Moody GB, Mark RG. The impact of the MIT-BIH Arrhythmia Database. IEEE Engineering in Medicine and Biology Magazine. 2001;20:45-50.
  3. Goldberger AL, Amaral LAN, Glass L, et al. PhysioBank, PhysioToolkit, and PhysioNet: components of a new research resource for complex physiologic signals. Circulation. 2000;101:e215-e220.
  4. Porr B, Howell L. R peak detector stress test with a new noisy ECG database reveals significant performance differences amongst popular detectors. bioRxiv. 2019.
  5. Sedghamiz H. Matlab implementation of Pan Tompkins ECG QRS detector. MathWorks File Exchange. 2014.

Figures

Figure 1. Per record sensitivity across all 48 records, sorted (per_record_se.svg).