Md. Kamrul Hasan 0002

dblp:64/2529-2 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
7since 2021 · last 2026
0000-0003-1292-4350ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 An efficient, scalable, and adaptable plug-and-play temporal attention module for motion-guided cardiac segmentation with sparse temporal labels
abstract
UNet and DT-VNet. Integrating TAM into SAM yields a temporal SAM that reduces Hausdorff distance (HD) from 3.99 mm to 3.51 mm on the CAMUS dataset, while integrating TAM into a pre-trained MedSAM reduces HD from 3.04 to 2.06 pixels after fine-tuning on the EchoNet-Dynamic dataset. On the ACDC 3D dataset, our TAM-UNet and TAM-DT-VNet achieve substantial reductions in HD, from 7.97 mm to 4.23 mm and 6.87 mm to 4.74 mm, respectively. Additionally, TAM's training does not require segmentation of ground truths from all time frames and can be achieved with sparse temporal annotation. TAM is thus a robust, generalizable, and adaptable solution for motion-awareness enhancement that is easily scaled from 2D to 3D. The code is available at https://github.com/kamruleee51/TAM.
Md. Kamrul Hasan 0002, Guang Yang 0006, Choon Hwai Yap
Medical Image Anal.1
2026 Explicit differentiable slicing and global deformation for cardiac mesh reconstruction
abstract
Three-dimensional (3D) mesh reconstruction of the cardiac anatomy from medical images is useful for shape and motion measurements and biophysics simulations. However, 3D medical images are often acquired as 2D slices that are sparsely sampled (e.g., large slice spacing) and noisy, and 3D mesh reconstruction on such data is a challenging task. Traditional voxel-based approaches utilize non-differentiable pre- and post-processing that compromises fidelity to images, while mesh-level deep learning approaches require large 3D mesh annotations that are difficult to obtain. Differentiable cross-domain supervision from 2D images to 3D meshes is therefore crucial for enabling end-to-end optimization in medical imaging. While there have been attempts to approximate the voxelization and slicing of meshes that are being optimized, there has not yet been a method for directly using 2D slices to supervise 3D mesh reconstruction in a differentiable manner. Here, we propose a novel explicit differentiable voxelization and slicing (DVS) algorithm allowing gradient backpropagation to a 3D mesh from its slices, which facilitates refined mesh optimization directly supervised by the losses defined on 2D images. Further, we propose an innovative framework for extracting patient-specific left ventricle (LV) meshes from medical images by coupling DVS with a graph harmonic deformation (GHD) mesh morphing descriptor of cardiac shape that naturally preserves mesh quality and smoothness during optimization. The proposed framework achieves state-of-the-art performance in cardiac mesh reconstruction tasks from densely sampled (CT) as well as sparsely sampled (MRI stack with few slices) images, outperforming alternatives, including Marching Cubes, statistical shape models, algorithms with vertex-based mesh morphing algorithms and alternative methods for image-supervision of mesh reconstruction. Experimental results demonstrate that our method achieves an overall Dice score of 90% during a sparse fitting on multi-datasets. The proposed method can further quantify clinically useful parameters such as ejection fraction and global myocardial strains, closely matching the ground truth and outperforming the traditional voxel-based approach in sparse images.
Yihao Luo, Dario Sesia, Fanwen Wang, Yinzhe Wu 0001, Wenhao Ding, Md. Kamrul Hasan 0002, Fadong Shi, Anoop Shah, Amit Kaura, Jamil Mayet, Guang Yang 0006, Choon Hwai Yap
Medical Image Anal.6
2026 FUGC: Benchmarking Semi-Supervised Learning Methods for Cervical Segmentation
abstract
Accurate segmentation of cervical structures in transvaginal ultrasound (TVS) is critical for assessing the risk of spontaneous preterm birth (PTB), yet the scarcity of labeled data limits the performance of supervised learning approaches. This paper introduces the Fetal Ultrasound Grand Challenge (FUGC), the first benchmark for semi-supervised learning in cervical segmentation, hosted at ISBI 2025. FUGC provides a dataset of 890 TVS images, including 500 training images, 90 validation images, and 300 test images. Methods were evaluated using the Dice Similarity Coefficient (DSC), Hausdorff Distance (HD), and runtime (RT), with a weighted combination of 0.4/0.4/0.2. The challenge attracted 10 teams with 82 participants submitting innovative solutions. The best-performing methods for each individual metric achieved 90.26% mDSC, 38.88 mHD, and 32.85 ms RT, respectively. FUGC establishes a standardized benchmark for cervical segmentation, demonstrates the efficacy of semi-supervised methods with limited labeled data, and provides a foundation for AI-assisted clinical PTB risk assessment.
Jieyun Bai, Yitong Tang, Mahdi Islam, Musarrat Tabassum, Enrique Almar-Munoz, Nianjiang Lv, Yu Chen 0099, Zilun Peng, Yusong Xiao, Li Xiao 0002, Nam-Khanh Tran, Dac-Phu Phan-Le, Hai-Dang Nguyen, Xiao Liu 0037, Jiale Hu, Mingxu Huang, Jitao Liang, Chaolu Feng, Xuezhi Zhang, Lyuyang Tong, Bo Du 0001, Ha-Hieu Pham, Thanh-Huy Nguyen, Min Xu 0009, Juntao Jiang, Jiangning Zhang, Yong Liu 0007, Md. Kamrul Hasan 0002, Zhuonan Liang, Tom Weidong Cai, Gongning Luo, Mohammad Yaqub, Karim Lekadir
IEEE Trans. Medical Imaging32
2026 4-D Reconstruction of Fetal Left Ventricle From Echocardiography via 2.5-D Radial Segmentation and Graph-Fourier Reconstruction
Md. Kamrul Hasan 0002, Haziq Shahard, Lucas Iijima, Nida Ruseckaite, Yihao Luo, Iris Scharnreitner, Andreas Tulzer, Bin Liu 0040, Guang Yang 0006, Choon Hwai Yap
IEEE Trans. Medical Imaging1
2025 Feedback Attention to Enhance Unsupervised Deep Learning Image Registration in 3D Echocardiography
abstract
Cardiac motion estimation is important for assessing the contractile health of the heart, and performing this in 3D can provide advantages due to the complex 3D geometry and motions of the heart. Deep learning image registration (DLIR) is a robust way to achieve cardiac motion estimation in echocardiography, providing speed and precision benefits, but DLIR in 3D echo remains challenging. Successful unsupervised 2D DLIR strategies are often not effective in 3D, and there have been few 3D echo DLIR implementations. Here, we propose a new spatial feedback attention (FBA) module to enhance unsupervised 3D DLIR and enable it. The module uses the results of initial registration to generate a co-attention map that describes remaining registration errors spatially and feeds this back to the DLIR to minimize such errors and improve self-supervision. We show that FBA improves a range of promising 3D DLIR designs, including networks with and without transformer enhancements, and that it can be applied to both fetal and adult 3D echo, suggesting that it can be widely and flexibly applied. We further find that the optimal 3D DLIR configuration is when FBA is combined with a spatial transformer and a DLIR backbone modified with spatial and channel attention, which outperforms existing 3D DLIR approaches. FBA's good performance suggests that spatial attention is a good way to enable scaling up from 2D DLIR to 3D and that a focus on the quality of the image after registration warping is a good way to enhance DLIR performance. Codes and data are available at: https://github.com/kamruleee51/Feedback_DLIR.
Md. Kamrul Hasan 0002, Yihao Luo, Guang Yang 0006, Choon Hwai Yap
IEEE Trans. Medical Imaging1
2021 DRNet: Segmentation and localization of optic disc and Fovea from diabetic retinopathy image
Md. Kamrul Hasan 0002, Md. Toufick E. Elahi, Shidhartho Roy, Robert Martí
Artif. Intell. Medicine1
2021 Detection, segmentation, and 3D pose estimation of surgical tools using convolutional neural networks and algebraic geometry
abstract
Background and objective: Surgical tool detection, segmentation, and 3D pose estimation are crucial components in Computer-Assisted Laparoscopy (CAL). The existing frameworks have two main limitations. First, they do not integrate all three components. Integration is critical; for instance, one should not attempt computing pose if detection is negative. Second, they have highly specific requirements, such as the availability of a CAD model. We propose an integrated and generic framework whose sole requirement for the 3D pose is that the tool shaft is cylindrical. Our framework makes the most of deep learning and geometric 3D vision by combining a proposed Convolutional Neural Network (CNN) with algebraic geometry. We show two applications of our framework in CAL: tool-aware rendering in Augmented Reality (AR) and tool-based 3D measurement. Methods: We name our CNN as ART-Net (Augmented Reality Tool Network). It has a Single Input Multiple Output (SIMO) architecture with one encoder and multiple decoders to achieve detection, segmentation, and geometric primitive extraction. These primitives are the tool edge-lines, mid-line, and tip. They allow the tool’s 3D pose to be estimated by a fast algebraic procedure. The framework only proceeds if a tool is detected. The accuracy of segmentation and geometric primitive extraction is boosted by a new Full resolution feature map Generator (FrG). We extensively evaluate the proposed framework with the EndoVis and new proposed datasets. We compare the segmentation results against several variants of the Fully Convolutional Network (FCN) and U-Net. Several ablation studies are provided for detection, segmentation, and geometric primitive extraction. The proposed datasets are surgery videos of different patients. Results: In detection, ART-Net achieves 100.0 % in both average precision and accuracy. In segmentation, it achieves 81.0 % in mean Intersection over Union (mIoU) on the robotic EndoVis dataset (articulated tool), where it outperforms both FCN and U-Net, by 4.5 p p and 2.9 p p , respectively. It achieves 88.2 % in mIoU on the remaining datasets (non-articulated tool). In geometric primitive extraction, ART-Net achieves 2.45 ∘ and 2.23 ∘ in mean Arc Length (mAL) error for the edge-lines and mid-line, respectively, and 9.3 pixels in mean Euclidean distance error for the tool-tip. Finally, in terms of 3D pose evaluated on animal data, our framework achieves 1.87 mm, 0.70 mm, and 4.80 mm mean absolute errors on the X , Y , and Z coordinates, respectively, and 5 . 94 ∘ angular error on the shaft orientation. It achieves 2.59 mm and 1.99 mm in mean and median location error of the tool head evaluated on patient data. Conclusions: The proposed framework outperforms existing ones in detection and segmentation. Compared to separate networks, integrating the tasks in a single network preserves accuracy in detection and segmentation but substantially improves accuracy in geometric primitive extraction. Overall, our framework has similar or better accuracy in 3D pose estimation while largely improving robustness against the very challenging imaging conditions of laparoscopy. The source code of our framework and our annotated dataset will be made publicly available at https://github.com/kamruleee51/ART-Net .
Md. Kamrul Hasan 0002, Lilian Calvet, Navid Rabbani, Adrien Bartoli
Medical Image Anal.1