Benjamin Hou

dblp:195/5714 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0003-3968-1707ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 How well do multimodal LLMs interpret CT scans? An auto-evaluation framework for analyses
abstract
OBJECTIVE: This study introduces a novel evaluation framework, GPTRadScore, to systematically assess the performance of multimodal large language models (MLLMs) in generating clinically accurate findings from CT imaging. Specifically, GPTRadScore leverages LLMs as an evaluation metric, aiming to provide a more accurate and clinically informed assessment than traditional language-specific methods. Using this framework, we evaluate the capability of several MLLMs, including GPT-4 with Vision (GPT-4V), Gemini Pro Vision, LLaVA-Med, and RadFM, to interpret findings in CT scans. METHODS: This retrospective study leverages a subset of the public DeepLesion dataset to evaluate the performance of several multimodal LLMs in describing findings in CT slices. GPTRadScore was developed to assess the generated descriptions (location, body part, and type) using GPT-4, alongside traditional metrics. RadFM was fine-tuned using a subset of the DeepLesion dataset with additional labeled examples targeting complex findings. Post fine-tuning, performance was reassessed using GPTRadScore to measure accuracy improvements. RESULTS: Evaluations demonstrated a high correlation of GPTRadScore with clinician assessments, with Pearson's correlation coefficients of 0.87, 0.91, 0.75, 0.90, and 0.89. These results highlight its superiority over traditional metrics, such as BLEU, METEOR, and ROUGE, and indicate that GPTRadScore can serve as a reliable evaluation metric. Using GPTRadScore, it was observed that while GPT-4V and Gemini Pro Vision outperformed other models, significant areas for improvement remain, primarily due to limitations in the datasets used for training. Fine-tuning RadFM resulted in substantial accuracy gains: location accuracy increased from 3.41% to 12.8%, body part accuracy improved from 29.12% to 53%, and type accuracy rose from 9.24% to 30%. These findings reinforce the hypothesis that fine-tuning RadFM can significantly enhance its performance. CONCLUSION: GPT-4 effectively correlates with expert assessments, validating its use as a reliable metric for evaluating multimodal LLMs in radiological diagnostics. Additionally, the results underscore the efficacy of fine-tuning approaches in improving the descriptive accuracy of LLM-generated medical imaging findings.
Qingqing Zhu, Benjamin Hou, Tejas Sudharshan Mathai, Pritam Mukherjee, Qiao Jin 0001, Xiuying Chen, Zhizheng Wang, Ruida Cheng, Ronald M. Summers, Zhiyong Lu
J. Biomed. Informatics2
2022 Natural Synthetic Anomalies for Self-supervised Anomaly Detection and Localization
Hannah M. Schlüter, Jeremy Tan, Benjamin Hou, Bernhard Kainz
ECCV (31)3
2022 MOOD 2020: A Public Benchmark for Out-of-Distribution Detection and Localization on Medical Images
abstract
Detecting Out-of-Distribution (OoD) data is one of the greatest challenges in safe and robust deployment of machine learning algorithms in medicine. When the algorithms encounter cases that deviate from the distribution of the training data, they often produce incorrect and over-confident predictions. OoD detection algorithms aim to catch erroneous predictions in advance by analysing the data distribution and detecting potential instances of failure. Moreover, flagging OoD cases may support human readers in identifying incidental findings. Due to the increased interest in OoD algorithms, benchmarks for different domains have recently been established. In the medical imaging domain, for which reliable predictions are often essential, an open benchmark has been missing. We introduce the Medical-Out-Of-Distribution-Analysis-Challenge (MOOD) as an open, fair, and unbiased benchmark for OoD methods in the medical imaging domain. The analysis of the submitted algorithms shows that performance has a strong positive correlation with the perceived difficulty, and that all algorithms show a high variance for different anomalies, making it yet hard to recommend them for clinical practice. We also see a strong correlation between challenge ranking and performance on a simple toy test set, indicating that this might be a valuable addition as a proxy dataset during anomaly detection algorithm development.
David Zimmerer, Peter M. Full, Fabian Isensee, Paul F. Jaeger, Tim Adler, Jens Petersen, Gregor Köhler, Tobias Roß, Annika Reinke, Antanas Kascenas, Bjørn Sand Jensen, Alison O'Neil, Jeremy Tan, Benjamin Hou, James Batten, Huaqi Qiu, Bernhard Kainz, Nina Shvetsova, Irina Fedulova, Dmitry V. Dylov, Baolun Yu, Jianyang Zhai, Jingtao Hu, Runxuan Si, Sihang Zhou 0001, Siqi Wang 0001, Xuerun Chen, Yang Zhao 0003, Sergio Naval Marimont, Giacomo Tarroni, Victor Saase, Lena Maier-Hein, Klaus H. Maier-Hein
IEEE Trans. Medical Imaging14
2021 RATCHET: Medical Transformer for Chest X-ray Diagnosis and Reporting
Benjamin Hou, Georgios Kaissis, Ronald M. Summers, Bernhard Kainz
MICCAI (7)1
2021 Ultrasound Video Transformers for Cardiac Ejection Fraction Estimation
Hadrien Reynaud, Athanasios Vlontzos, Benjamin Hou, Arian Beqiri, Paul Leeson, Bernhard Kainz
MICCAI (6)3
2021 Detecting Outliers with Poisson Image Interpolation
Jeremy Tan, Benjamin Hou, Thomas G. Day, John M. Simpson, Daniel Rueckert, Bernhard Kainz
MICCAI (5)2
2019 Evaluating reinforcement learning agents for anatomical landmark detection
Amir Alansary, Ozan Oktay, Loïc Le Folgoc, Benjamin Hou, Ghislain Vaillant, Konstantinos Kamnitsas, Athanasios Vlontzos, Ben Glocker, Bernhard Kainz, Daniel Rueckert
Medical Image Anal.5
2019 Weakly Supervised Estimation of Shadow Confidence Maps in Fetal Ultrasound Imaging
abstract
Detecting acoustic shadows in ultrasound images is important in many clinical and engineering applications. Real-time feedback of acoustic shadows can guide sonographers to a standardized diagnostic viewing plane with minimal artifacts and can provide additional information for other automatic image analysis algorithms. However, automatically detecting shadow regions using learning-based algorithms is challenging because pixel-wise ground truth annotation of acoustic shadows is subjective and time consuming. In this paper, we propose a weakly supervised method for automatic confidence estimation of acoustic shadow regions. Our method is able to generate a dense shadow-focused confidence map. In our method, a shadow-seg module is built to learn general shadow features for shadow segmentation, based on global image-level annotations as well as a small number of coarse pixel-wise shadow annotations. A transfer function is introduced to extend the obtained binary shadow segmentation to a reference confidence map. In addition, a confidence estimation network is proposed to learn the mapping between input images and the reference confidence maps. This network is able to predict shadow confidence maps directly from input images during inference. We use evaluation metrics such as DICE, inter-class correlation, and so on, to verify the effectiveness of our method. Our method is more consistent than human annotation and outperforms the state-of-the-art quantitatively in shadow segmentation and qualitatively in confidence estimation of shadow regions. Furthermore, we demonstrate the applicability of our method by integrating shadow confidence maps into tasks such as ultrasound image classification, multi-view image fusion, and automated biometric measurements.
Qingjie Meng, Richard James Housden, Jacqueline Matthew, Daniel Rueckert, Julia A. Schnabel, Bernhard Kainz, Matthew Sinclair, Veronika A. M. Zimmer, Benjamin Hou, Martin Rajchl, Nicolas Toussaint, Ozan Oktay, Jo Schlemper, Alberto Gómez 0002
IEEE Trans. Medical Imaging9
2018 Automatic View Planning with Multi-scale Deep Reinforcement Learning Agents
Amir Alansary, Loïc Le Folgoc, Ghislain Vaillant, Ozan Oktay, Wenjia Bai, Jonathan Passerat-Palmbach, Ricardo Guerrero, Konstantinos Kamnitsas, Benjamin Hou, Steven McDonagh 0001, Ben Glocker, Bernhard Kainz, Daniel Rueckert
MICCAI (1)10
2018 Computing CNN Loss and Gradients for Pose Estimation with Riemannian Geometry
Benjamin Hou, Nina Miolane, Bishesh Khanal, Matthew C. H. Lee, Amir Alansary, Steven McDonagh 0001, Joseph V. Hajnal, Daniel Rueckert, Ben Glocker, Bernhard Kainz
MICCAI (1)1
2018 Standard Plane Detection in 3D Fetal Ultrasound Using an Iterative Transformation Network
Bishesh Khanal, Benjamin Hou, Amir Alansary, Juan J. Cerrolaza, Matthew Sinclair, Jacqueline Matthew, Chandni Gupta, Caroline L. Knight, Bernhard Kainz, Daniel Rueckert
MICCAI (1)3
2018 3-D Reconstruction in Canonical Co-Ordinate Space From Arbitrarily Oriented 2-D Images
abstract
Limited capture range, and the requirement to provide high quality initialization for optimization-based 2-D/3-D image registration methods, can significantly degrade the performance of 3-D image reconstruction and motion compensation pipelines. Challenging clinical imaging scenarios, which contain significant subject motion, such as fetal in-utero imaging, complicate the 3-D image and volume reconstruction process. In this paper, we present a learning-based image registration method capable of predicting 3-D rigid transformations of arbitrarily oriented 2-D image slices, with respect to a learned canonical atlas co-ordinate system. Only image slice intensity information is used to perform registration and canonical alignment, no spatial transform initialization is required. To find image transformations, we utilize a convolutional neural network architecture to learn the regression function capable of mapping 2-D image slices to a 3-D canonical atlas space. We extensively evaluate the effectiveness of our approach quantitatively on simulated magnetic resonance imaging (MRI), fetal brain imagery with synthetic motion and further demonstrate qualitative results on real fetal MRI data where our method is integrated into a full reconstruction and motion compensation pipeline. Our learning based registration achieves an average spatial prediction error of 7 mm on simulated data and produces qualitatively improved reconstructions for heavily moving fetuses with gestational ages of approximately 20 weeks. Our model provides a general and computationally efficient solution to the 2-D/3-D registration initialization problem and is suitable for real-time scenarios.
Benjamin Hou, Bishesh Khanal, Amir Alansary, Steven McDonagh 0001, Alice Davidson, Mary A. Rutherford, Joseph V. Hajnal, Daniel Rueckert, Ben Glocker, Bernhard Kainz
IEEE Trans. Medical Imaging1
2017 Predicting Slice-to-Volume Transformation in Presence of Arbitrary Subject Motion
Benjamin Hou, Amir Alansary, Steven McDonagh 0001, Alice Davidson, Mary A. Rutherford, Joseph V. Hajnal, Daniel Rueckert, Ben Glocker, Bernhard Kainz
MICCAI (2)1