VLDB 2026 Research / reviewers in the wild / expert
Dong Hye Ye
dblp:90/9123
· DBLP profile ↗
17ranked-venue papers
5as first author
6since 2021 · last 2026
0000-0002-9186-4095ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 4 first-authorArtificial intelligence and machine learning · 2 · 1 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DiA-gnostic VLVAE: Disentangled Alignment-Constrained Vision Language Variational AutoEncoder for Robust Radiology Reporting with Missing ModalitiesabstractThe integration of medical images with clinical context is essential for generating accurate and clinically interpretable radiology reports. However, current automated methods often rely on resource-heavy Large Language Models (LLMs) or static knowledge graphs and struggle with two fundamental challenges in real-world clinical data: (1) missing modalities, such as incomplete clinical context , and (2) feature entanglement, where mixed modality-specific and shared information leads to suboptimal fusion and clinically unfaithful hallucinated findings. To address these challenges, we propose the DiA-gnostic VLVAE, which achieves robust radiology reporting through Disentangled Alignment. Our framework is designed to be resilient to missing modalities by disentangling shared and modality-specific features using a Mixture-of-Experts (MoE) based Vision-Language Variational Autoencoder (VLVAE). A constrained optimization objective enforces orthogonality and alignment between these latent representations to prevent suboptimal fusion. A compact LLaMA-X decoder then uses these disentangled representations to generate reports efficiently. On the IU X-Ray and MIMIC-CXR datasets, DiA has set new state-of-the-art BLEU@4 scores of 0.266 and 0.134, respectively. Experimental results show that the proposed method significantly outperforms state-of-the-art models. Nagur Shareef Shaik, Teja Krishna Cherukuri, Adnan Masood, Dong Hye Ye |
AAAI | 4 |
| 2025 | GCS-M3VLT: Guided Context Self-Attention based Multi-modal Medical Vision Language Transformer for Retinal Image CaptioningabstractRetinal image analysis is crucial for diagnosing and treating eye diseases, yet generating accurate medical reports from images remains challenging due to variability in image quality and pathology, especially with limited labeled data. Previous Transformer-based models struggled to integrate visual and textual information under limited supervision. In response, we propose a novel vision-language model for retinal image captioning that combines visual and textual features through a guided context self-attention mechanism. This approach captures both intricate details and the global clinical context, even in data-scarce scenarios. Extensive experiments on the DeepEyeNet dataset demonstrate a 0.023 BLEU@4 improvement, along with significant qualitative advancements, highlighting the effectiveness of our model in generating comprehensive medical captions. Teja Krishna Cherukuri, Nagur Shareef Shaik, Jyostna Devi Bodapati, Dong Hye Ye |
ICASSP | 4 |
| 2024 | Guided Context Gating: Learning To Leverage Salient Lesions in Retinal Fundus ImagesabstractEffectively representing medical images, especially retinal images, presents a considerable challenge due to variations in appearance, size, and contextual information of pathological signs called lesions. Precise discrimination of these lesions is crucial for diagnosing vision-threatening issues such as diabetic retinopathy. While visual attention-based neural networks have been introduced to learn spatial context and channel correlations from retinal images, they often fall short in capturing localized lesion context. Addressing this limitation, we propose a novel attention mechanism called Guided Context Gating, an unique approach that integrates Context Formulation, Channel Correlation, and Guided Gating to learn global context, spatial correlations, and localized lesion context. Our qualitative evaluation against existing attention mechanisms emphasize the superiority of Guided Context Gating in terms of explainability. Notably, experiments on the Zenodo-DR-7 dataset reveal a substantial 2.63* accuracy boost over advanced attention mechanisms & an impressive 6.53% improvement over the state-of-the-art Vision Transformer for assessing the severity grade of retinopathy, even with imbalanced and limited training samples for each class. Teja Krishna Cherukuri, Nagur Shareef Shaik, Dong Hye Ye |
ICIP | 3 |
| 2024 | M3T: Multi-Modal Medical Transformer To Bridge Clinical Context With Visual Insights For Retinal Image Medical Description GenerationabstractAutomated retinal image medical description generation is crucial for streamlining medical diagnosis and treatment planning. Existing challenges include the reliance on learned retinal image representations, difficulties in handling multiple imaging modalities, and the lack of clinical context in visual representations. Addressing these issues, we propose the Multi-Modal Medical Transformer (M3T), a novel deep learning architecture that integrates visual representations with diagnostic keywords. Unlike previous studies focusing on specific aspects, our approach efficiently learns contextual information and semantics from both modalities, enabling the generation of precise and coherent medical descriptions for retinal images. Experimental studies on the DeepEyeNet dataset validate the success of M3T in meeting ophthalmologists’ standards, demonstrating a substantial 13.5% improvement in BLEU@4 over the best-performing baseline model. Nagur Shareef Shaik, Teja Krishna Cherukuri, Dong Hye Ye |
ICIP | 3 |
| 2022 | Deep Learning From Imaging Genetics for Schizophrenia ClassificationabstractSchizophrenia (SZ) is a serious psychiatric disorder, causing substantial socioeconomic burden. Since an SZ patients’ brain may have structural changes including reduced hippocampal and thalamic volume [1], brain MRI is becoming a popular imaging method studying SZ. In addition, as a genetic disorder, genetic information such as single nucleotide polymorphisms (SNP) plays an important role in distinguishing SZ. However, the structural and genetic changes in SZ patients are too subtle to be identified by human vision, so it is necessary to develop an automated method to find the nonlinear patterns associated with disease progression. Toward this, we propose a novel multi-modal deep learning approach where we combine both features from structural MRI (sMRI) and single-nucleotide polymorphisms (SNPs) for SZ classification. For sMRI, we extract convolutional features from a pre-trained deep neural network to capture morphological characteristics. For SNPs, we apply a layer-wise relevance propagation (LRP) method on a pre-trained 1-D convolutional network to identify SZ-linked SNPs. We then feed the combined features to a tree-based classifier for SZ diagnosis. Experimental results on clinical dataset showed classification accuracy was increased by 5.3% compared to the state of the art DenseNet using only sMRI data. Hongkun Yu 0003, Thomas Florian, Vince D. Calhoun, Dong Hye Ye |
ICIP | 4 |
| 2021 | Enhancing Multi-Channel Eeg Classification with Gramian Temporal Generative Adversarial NetworksabstractDeep learning’s requirements for large amounts of training data remains a challenge for researchers and developers. Generative Adversarial Network (GAN) is commonly used in medical image analysis to generate novel training images to help resolve this issue. While deep learning has many clinical applications in radiology, its applications in medical time series data such as electroencephalogram (EEG) are usually constrained to 1 dimension. Hence, there are few available GAN architectures that effectively synthesize single and multi-channel EEG. In this paper, we propose a novel method to synthesize multi-channel EEG in the form of Gramian Angular Field (GAF) images with a Gramian Temporal Generative Adversarial Network (GT-GAN). The proposed network is capable of generating realistic GAF images and enhances EEG anomaly detection accuracy in residual learning frameworks. Chi Nok Enoch Kan, Richard J. Povinelli, Dong Hye Ye |
ICASSP | 3 |
| 2019 | Prior-Guided Metal Artifact Reduction for Iterative X-Ray Computed TomographyabstractHigh-attenuation materials pose significant challenges to computed tomographic imaging. Formed of high mass-density and high atomic number elements, they cause more severe beam hardening and scattering artifacts than do water-like materials. Pre-corrected line-integral density measurements are no longer linearly proportional to the path lengths, leading to reconstructed image suffering from streaking artifacts extending from metal, often along highest-density directions. In this paper, a novel prior-based iterative approach is proposed to reduce metal artifacts. It combines the superiority of statistical methods with the benefits of sinogram completion methods to estimate and correct metal-induced biases. Preliminary results show minimized residual artifacts and significantly improved image quality. Zhiqian Chang, Dong Hye Ye, Somesh Srivastava, Jean-Baptiste Thibault, Ken D. Sauer, Charles A. Bouman |
IEEE Trans. Medical Imaging | 2 |
| 2018 | Deep Residual Learning for Model-Based Iterative CT Reconstruction Using Plug-and-Play FrameworkabstractModel-Based Iterative Reconstruction (MBIR) has shown promising results in clinical studies as they allow significant dose reduction during CT scans while maintaining the diagnostic image quality. MBIR improves the image quality over analytical reconstruction by modeling both the sensor (e.g., forward model) and the image being reconstructed (e.g., prior model). While the forward model is typically based on the physics of the sensor, accurate prior modeling remains a challenging problem. Markov Random Field (MRF) has been widely used as prior models in MBIR due to simple structure, but they cannot completely capture the subtle characteristics of complex images. To tackle this challenge, we generate a prior model by learning the desirable image property from a large dataset. Toward this, we use Plug-and-Play (PnP) framework which decouples the forward model and the prior model in the optimization procedure, replacing the prior model optimization by a image denoising operator. Then, we adopt the state-of-the-art deep residual learning for the image denoising operator which represents the prior model in MBIR. Experimental results on real CT scans demonstrate that our PnP MBIR with deep residual learning prior significantly reduces the noise and artifacts compared to analytical reconstruction and standard MBIR with MRF prior. Dong Hye Ye, Somesh Srivastava, Jean-Baptiste Thibault, Ken D. Sauer, Charles A. Bouman |
ICASSP | 1 |
| 2016 | Multi-target detection and tracking from a single camera in Unmanned Aerial Vehicles (UAVs)abstractDespite the recent flight control regulations, Unmanned Aerial Vehicles (UAVs) are still gaining popularity in civilian and military applications, as much as for personal use. Such emerging interest is pushing the development of effective collision avoidance systems. Such systems play a critical role UAVs operations especially in a crowded airspace setting. Because of cost and weight limitations associated with UAVs payload, camera based technologies are the de-facto choice for collision avoidance navigation systems. This requires multi-target detection and tracking algorithms from a video, which can be run on board efficiently. While there has been a great deal of research on object detection and tracking from a stationary camera, few have attempted to detect and track small UAVs from a moving camera. In this paper, we present a new approach to detect and track UAVs from a single camera mounted on a different UAV. Initially, we estimate background motions via a perspective transformation model and then identify distinctive points in the background subtracted image. We find spatio-temporal traits of each moving object through optical flow matching and then classify those candidate targets based on their motion patterns compared with the background. The performance is boosted through Kalman filter tracking. This results in temporal consistency among the candidate detections. The algorithm was validated on video datasets taken from a UAV. Results show that our algorithm can effectively detect and track small UAVs with limited computing resources. Jing Li 0169, Dong Hye Ye, Timothy H. Chung, Mathias Kölsch, Juan P. Wachs, Charles A. Bouman |
IROS | 2 |
| 2015 | Joint metal artifact reduction and segmentation of CT images using dictionary-based image prior and continuous-relaxed potts modelabstractSegmenting interesting objects from CT images has a wide range of applications. However, to achieve good results, it is often necessary to apply metal artifact reduction to raw CT images before segmentation. While there has been a great deal of research focusing on metal artifact reduction and segmentation as individual tasks, there have been very few attempts to solve the two problems jointly. We present a novel approach to solve the problem of segmenting raw CT images with metal artifacts, without the access to the raw CT data. Given an approximate metal artifact mask, the problem is formulated as a joint optimization over the restored image and the segmentation label, and the cost function includes a dictionary-based image prior to regularize the restored image and a continuous-relaxed Potts model for multi-class segmentation. An effective alternating method is used to solve the resulting optimization problem. The algorithm is applied to both simulated and real datasets and results show that it is effective in reducing metal artifacts and generating better segmentations simultaneously. Pengchong Jin, Dong Hye Ye, Charles A. Bouman |
ICIP | 2 |
| 2015 | The Multimodal Brain Tumor Image Segmentation Benchmark (BRATS)abstractIn this paper we report the set-up and results of the Multimodal Brain Tumor Image Segmentation Benchmark (BRATS) organized in conjunction with the MICCAI 2012 and 2013 conferences. Twenty state-of-the-art tumor segmentation algorithms were applied to a set of 65 multi-contrast MR scans of low- and high-grade glioma patients-manually annotated by up to four raters-and to 65 comparable scans generated using tumor image simulation software. Quantitative evaluations revealed considerable disagreement between the human raters in segmenting various tumor sub-regions (Dice scores in the range 74%-85%), illustrating the difficulty of this task. We found that different algorithms worked best for different sub-regions (reaching performance comparable to human inter-rater variability), but that no single algorithm ranked in the top for all sub-regions simultaneously. Fusing several good algorithms using a hierarchical majority vote yielded segmentations that consistently ranked above all individual algorithms, indicating remaining opportunities for further methodological improvements. The BRATS image data and manual annotations continue to be publicly available through an online evaluation system as an ongoing benchmarking resource. Bjoern Menze, András Jakab, Stefan Bauer, Jayashree Kalpathy-Cramer, Keyvan Farahani, Justin S. Kirby, Yuliya Burren, Nicole Porz, Johannes Slotboom, Roland Wiest, Levente Lanczi, Elizabeth R. Gerstner, Marc-André Weber, Tal Arbel, Brian B. Avants, Nicholas Ayache, Patricia Buendia, D. Louis Collins, Nicolas Cordier, Jason J. Corso, Antonio Criminisi, Tilak Das, Hervé Delingette, Çagatay Demiralp, Christopher R. Durst, Michel Dojat, Senan Doyle, Joana Festa, Florence Forbes, Ezequiel Geremia, Ben Glocker, Polina Golland, Xiaotao Guo, Andac Hamamci, Khan M. Iftekharuddin, Raj Jena, Nigel M. John, Ender Konukoglu, Danial Lashkari, José Antonio Mariz, Raphael Meier, Sérgio Pereira, Doina Precup, Stephen J. Price, Tammy Riklin-Raviv, Syed M. S. Reza, Michael T. Ryan, Duygu Sarikaya, Lawrence H. Schwartz, Hoo-Chang Shin, Jamie Shotton, Carlos A. Silva 0002, Nuno J. Sousa, Nagesh K. Subbanna, Gábor Székely, Thomas J. Taylor, Owen M. Thomas, Nicholas J. Tustison, Gozde Unal, Flor Vasseur, Max Wintermark, Dong Hye Ye, Liang Zhao 0018, Binsheng Zhao, Darko Zikic, Marcel Prastawa, Mauricio Reyes 0001, Koenraad Van Leemput |
IEEE Trans. Medical Imaging | 62 |
| 2014 | Regional Manifold Learning for Disease ClassificationabstractWhile manifold learning from images itself has become widely used in medical image analysis, the accuracy of existing implementations suffers from viewing each image as a single data point. To address this issue, we parcellate images into regions and then separately learn the manifold for each region. We use the regional manifolds as low-dimensional descriptors of high-dimensional morphological image features, which are then fed into a classifier to identify regions affected by disease. We produce a single ensemble decision for each scan by the weighted combination of these regional classification results. Each weight is determined by the regional accuracy of detecting the disease. When applied to cardiac magnetic resonance imaging of 50 normal controls and 50 patients with reconstructive surgery of Tetralogy of Fallot, our method achieves significantly better classification accuracy than approaches learning a single manifold across the entire image domain. Dong Hye Ye, Benoit Desjardins, Jihun Hamm, Harold Litt, Kilian M. Pohl |
IEEE Trans. Medical Imaging | 1 |
| 2013 | FLOOR: Fusing Locally Optimal Registrations
Dong Hye Ye, Jihun Hamm, Benoit Desjardins, Kilian M. Pohl |
MICCAI (3) | 1 |
| 2013 | Modality Propagation: Coherent Synthesis of Subject-Specific Scans with Data-Driven Regularization
Dong Hye Ye, Darko Zikic, Ben Glocker, Antonio Criminisi, Ender Konukoglu |
MICCAI (1) | 1 |
| 2012 | Regional Manifold Learning for Deformable Registration of Brain MR Images
Dong Hye Ye, Jihun Hamm, Dongjin Kwon, Christos Davatzikos, Kilian M. Pohl |
MICCAI (3) | 1 |
| 2012 | Discriminative Segmentation-Based Evaluation Through Shape DissimilarityabstractSegmentation-based scores play an important role in the evaluation of computational tools in medical image analysis. These scores evaluate the quality of various tasks, such as image registration and segmentation, by measuring the similarity between two binary label maps. Commonly these measurements blend two aspects of the similarity: pose misalignments and shape discrepancies. Not being able to distinguish between these two aspects, these scores often yield similar results to a widely varying range of different segmentation pairs. Consequently, the comparisons and analysis achieved by interpreting these scores become questionable. In this paper, we address this problem by exploring a new segmentation-based score, called normalized Weighted Spectral Distance (nWSD), that measures only shape discrepancies using the spectrum of the Laplace operator. Through experiments on synthetic and real data we demonstrate that nWSD provides additional information for evaluating differences between segmentations, which is not captured by other commonly used scores. Our results demonstrate that when jointly used with other scores, such as Dice's similarity coefficient, the additional information provided by nWSD allows richer, more discriminative evaluations. We show for the task of registration that through this addition we can distinguish different types of registration errors. This allows us to identify the source of errors and discriminate registration results which so far had to be treated as being of similar quality in previous evaluation studies. Ender Konukoglu, Ben Glocker, Dong Hye Ye, Antonio Criminisi, Kilian M. Pohl |
IEEE Trans. Medical Imaging | 3 |
| 2010 | GRAM: A framework for geodesic registration on anatomical manifolds
Jihun Hamm, Dong Hye Ye, Ragini Verma, Christos Davatzikos |
Medical Image Anal. | 2 |