Holtercare-Bench: A Multimodal Benchmark for Evaluating Long-Term Dynamic ECG Analysis
Researchers introduce Holtercare-23K, a large-scale multimodal dynamic ECG dataset, and Holtercare-Bench, a benchmark for evaluating models on temporal localization, clinical diagnosis, and global summarization.
Intelligence analysis by Llama

The authors present a novel signal-video-text tri-modal alignment and a benchmark for long-term medical large language models, highlighting the limitations of current models in electrophysiology.
Imagine you have a special machine that can look at your heart's rhythm and tell if it's healthy or not. But this machine is not very good at looking at your heart's rhythm for a long time. Researchers created a new dataset and benchmark to help improve this machine's ability to look at your heart's rhythm for a long time and make accurate diagnoses.
Analysis
Holtercare-23K: A Large-Scale Multimodal Dynamic ECG Dataset
The authors introduce Holtercare-23K, a large-scale multimodal dynamic ECG dataset comprising 22,980 QA pairs derived from 788 clinical Holter records. This dataset features a novel signal-video-text tri-modal alignment, which can be used to evaluate models on temporal localization, clinical diagnosis, and global summarization.
Holtercare-Bench: A Multimodal Benchmark for Long-Term Medical MLLMs
Based on the Holtercare-23K dataset, the authors present Holtercare-Bench, a multimodal benchmark that evaluates models on temporal localization, clinical diagnosis, and global summarization. Zero-shot evaluations of leading MLLMs reveal a significant performance gap in processing ultra-long pathological sequences. However, fine-tuning representative models yields substantial improvements.
Limitations of Current MLLMs in Electrophysiology
The authors highlight the limitations of current MLLMs in electrophysiology, including their inability to process ultra-long pathological sequences and generate accurate diagnostic reports. This work provides a foundational benchmark for long-term medical MLLMs, which can improve the accuracy and reliability of medical applications.
Key points
- Researchers introduce Holtercare-23K, a large-scale multimodal dynamic ECG dataset.
- Holtercare-Bench is a multimodal benchmark for evaluating models on temporal localization, clinical diagnosis, and global summarization.
- Current MLLMs struggle with processing ultra-long pathological sequences and generating accurate diagnostic reports.
- Fine-tuning representative models yields substantial improvements in performance.
If this development plays out positively, it could lead to more accurate and reliable medical applications, such as better diagnosis and treatment of heart conditions.
However, there are also potential risks, such as the over-reliance on machine learning models, which can lead to errors and misdiagnoses.

