Cardiac-CT Dataset: an expert-annotated cardiac computed tomography dataset for segmentation and phenotyping

Dataset

Link to Source: Hugging Face, Preprint

Authors: Pooya Mohammadi Kazaj, Leo Fridolin Weber, Wen Xie, Seyed Amir Ahmad Safavi-Naini, Anselm Stark, Giovanni Baj, Ali Mokhtari, Toshiya Yoshida, Christoph Ryffel, Taishi Okuno, Yoshihiro Akashi, Ronny R. Buechel, Thomas Pilgrim, Waldo Valenzuela, George C. M. Siontis, Xiaowei Xu, Moritz Hundertmark, Stephan Windecker, Christoph Gräni, Isaac Shiri

Summary: An expert-annotated cardiac computed tomography dataset covering 14 heart structures, built through a human-in-the-loop annotation pipeline and released alongside the segmentation labels, model weights and augmentation library from the accompanying study.

Released three-dimensional meshes of the segmented cardiac structures. Surface renderings of the 14 individually segmented structures, LV myocardium, left ventricle, left atrium, right ventricle, right atrium, LA appendage, coronary arteries, pulmonary vein, pericardial fat, epicardial fat, pulmonary arteries, aorta, superior vena cava, and inferior vena cava, color-coded by anatomical category, alongside two composite whole-heart assemblies (right). A total of 14,280 STL meshes are released with the dataset to support development of shape-aware segmentation networks and downstream applications including cardiovascular simulation, 3D printing, and physical and digital twinning of the heart.

Progress in automatic analysis of cardiac computed tomography depends on the availability of images in which the anatomy has been carefully outlined by experts. Such reference annotations are slow and costly to produce, and most public collections cover only one or two structures, most often the coronary arteries. The Cardiac-CT collection was assembled to widen that scope: 1,598 cases annotated across 14 distinct cardiac structures, of which 1,000 form the training set and 598 were held back as external test data. A further 60,000 unlabeled scans were used to pretrain a vision model without any manual labels, and five independent multicenter datasets were used for evaluation.

The annotations were produced with a human-in-the-loop procedure rather than by manual outlining alone. A model proposes an initial delineation, an expert reviews and corrects it, and the corrected cases are fed back to improve the next round of proposals. This keeps the reviewing effort focused on the cases that need it, and it makes the quality control step an explicit part of the pipeline instead of an afterthought. The resulting labels support more than segmentation alone: they allow measurement of the individual structures, extraction of coronary artery centerlines, and reconstruction of three-dimensional surface meshes of the heart.
The portion prepared for public release consists of the 1,000 training cases, which originate from the publicly available ImageCAS coronary computed tomography angiography collection and were extended here with expert annotations for the remaining 13 cardiac structures, together with the segmentation labels for the 20 cases of the Multi-Modality Whole Heart Segmentation test set. The images of that test set are not redistributed here and must be requested from the original challenge organizers. Credit for the underlying image collections belongs to the groups that created them; the contribution of the AI-CVM lab team is the added expert annotation, the quality control pipeline, and the accompanying evaluation.

The collection is intended to be used together with the other resources released from the same study, namely the model training and evaluation code, the segmentation architectures collected in nnUZoo, the CTAug augmentation library, and the HolOrama viewer. At the time of writing the data is not yet openly downloadable: access is gated while the accompanying manuscript is under peer review, and the files will be released once it is accepted. Use is governed by a non-commercial license that does not permit redistribution of modified versions.

 

References