VLDB 2026 Research / reviewers in the wild / expert
Kayhan Batmanghelich
dblp:38/193 · also Kayhan N. Batmanghelich, Nematollah Batmanghelich, Nematollah Kayhan Batmanghelich
· DBLP profile ↗
42ranked-venue papers
2as first author
24since 2021 · last 2025
0000-0001-9893-9136ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 1 first-author · 12 since 2021Artificial intelligence and machine learning · 20 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 2 first-author · 15 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multi-Modal Large Language Models are Effective Vision LearnersabstractLarge language models (LLMs), pre-trained on vast amounts of text, have shown remarkable abilities in understanding general knowledge and commonsense. There-fore, it's desirable to leverage pre-trained LLM to help solve computer vision tasks. Previous works on multi-modal LLM mainly focus on the generation capability. In this work, we propose LLM-augmented visual representation learning (LMVR). Our approach involves initially using a vision encoder to extract features, which are then projected into the word embedding space of the LLM. The LLM then generates responses based on the visual representation and a text prompt. Finally, we aggregate sequence-level features from the hidden layers of the LLM to obtain image-level representations. We conduct extensive experiments on multiple datasets, and have the following findings: (a) LMVR outperforms traditional vision encoder on various down-stream tasks, and effectively learns the correspondence between words and image regions; (b) LMVR improves the generalizability compared to using a vision encoder alone, as evidenced by its superior resistance to domain shift; (c) LMVR improves the robustness of models to corrupted and perturbed visual data. Our findings demonstrate LLM-augmented visual representation learning is effective as it learns object-level concepts and commonsense knowledge. Li Sun 0010, Chaitanya Ahuja, Matt D'Zmura, Kayhan Batmanghelich, Philip Bontrager |
WACV | 5 |
| 2025 | High-dimensional causal mediation analysis by partial sum statistic and sample splitting strategy in imaging genetics applicationabstractSUMMARY: Causal mediation analysis investigates the role of mediators in the relationship between exposure and outcome. In the analysis of omics or imaging data, mediators are often high-dimensional, presenting challenges such as multicollinearity and interpretability. Existing methods either compromise interpretability or fail to effectively prioritize mediators. To address these challenges and advance causal mediation analysis in high-dimensional contexts, we propose the Partial Sum Statistic and Sample Splitting Strategy (PS5) framework. Through extensive simulations, we demonstrate that PS5 offers superior type I error control, higher statistical power, reduced bias in mediation effect estimation, and more accurate mediator selection. We apply PS5 to an imaging genetics dataset of chronic obstructive pulmonary disease (COPD) patients from the COPDGene study. The results show successful estimation of the global indirect effect and identification of mediating image regions. Notably, we identify a region in the lower lobe of the lung that exhibits a strong and concordant mediation effect for both genetic and environmental exposures, suggesting potential targets for treatment to mitigate COPD severity caused by genetic and smoking effects. AVAILABILITY AND IMPLEMENTATION: PS5 is publicly available at https://github.com/hung-ching-chang/PS5Med. Hung-Ching Chang, Yusi Fang, Michael T. Gorczyca, Kayhan Batmanghelich, George C. Tseng |
Bioinform. | 4 |
| 2024 | Mammo-CLIP: A Vision Language Foundation Model to Enhance Data Efficiency and Robustness in Mammography
Shantanu Ghosh, Clare B. Poynton, Shyam Visweswaran, Kayhan Batmanghelich |
MICCAI (12) | 4 |
| 2024 | DrasCLR: A self-supervised framework of learning disease-related and anatomy-specific representation for 3D lung CT images
Ke Yu 0002, Li Sun 0010, Junxiang Chen, Maxwell Reynolds, Tigmanshu Chaudhary, Kayhan Batmanghelich |
Medical Image Anal. | 6 |
| 2024 | MedSyn: Text-Guided Anatomy-Aware Synthesis of High-Fidelity 3-D CT ImagesabstractThis paper introduces an innovative methodology for producing high-quality 3D lung CT images guided by textual information. While diffusion-based generative models are increasingly used in medical imaging, current state-of-the-art approaches are limited to low-resolution outputs and underutilize radiology reports' abundant information. The radiology reports can enhance the generation process by providing additional guidance and offering fine-grained control over the synthesis of images. Nevertheless, expanding text-guided generation to high-resolution 3D images poses significant memory and anatomical detail-preserving challenges. Addressing the memory issue, we introduce a hierarchical scheme that uses a modified UNet architecture. We start by synthesizing low-resolution images conditioned on the text, serving as a foundation for subsequent generators for complete volumetric data. To ensure the anatomical plausibility of the generated samples, we provide further guidance by generating vascular, airway, and lobular segmentation masks in conjunction with the CT images. The model demonstrates the capability to use textual input and segmentation tasks to generate synthesized images. Algorithmic comparative assessments and blind evaluations conducted by 10 board-certified radiologists indicate that our approach exhibits superior performance compared to the most advanced models based on GAN and diffusion techniques, especially in accurately retaining crucial anatomical features such as fissure lines and airways. This innovation introduces novel possibilities. This study focuses on two main objectives: (1) the development of a method for creating images based on textual prompts and anatomical components, and (2) the capability to generate new images conditioning on anatomical elements. The advancements in image generation can be applied to enhance numerous downstream tasks. Yanwu Xu 0003, Li Sun 0010, Wei Peng 0009, Shuyue Jia, Katelyn Morrison, Adam Perer, Afrooz Zandifar, Shyam Visweswaran, Motahhare Eslami, Kayhan Batmanghelich |
IEEE Trans. Medical Imaging | 10 |
| 2023 | From Characters to Words: Hierarchical Pre-trained Language Model for Open-vocabulary Language UnderstandingabstractCurrent state-of-the-art models for natural language understanding require a preprocessing step to convert raw text into discrete tokens.This process known as tokenization relies on a pre-built vocabulary of words or sub-word morphemes.This fixed vocabulary limits the model's robustness to spelling errors and its capacity to adapt to new domains.In this work, we introduce a novel open-vocabulary language model that adopts a hierarchical two-level approach: one at the word level and another at the sequence level.Concretely, we design an intraword module that uses a shallow Transformer architecture to learn word representations from their characters, and a deep inter-word Transformer module that contextualizes each word representation by attending to the entire word sequence.Our model thus directly operates on character sequences with explicit awareness of word boundaries, but without biased sub-word or word-level vocabulary.Experiments on various downstream tasks show that our method outperforms strong baselines.We also demonstrate that our hierarchical model is robust to textual corruption and domain shift. Li Sun 0010, Florian Luisier, Kayhan Batmanghelich, Dinei A. F. Florêncio, Cha Zhang |
ACL (1) | 3 |
| 2023 | Dividing and Conquering a BlackBox to a Mixture of Interpretable Models: Route, Interpret, RepeatabstractML model design either starts with an interpretable model or a Blackbox and explains it post hoc. Blackbox models are flexible but difficult to explain, while interpretable models are inherently explainable. Yet, interpretable models require extensive ML knowledge and tend to be less flexible, potentially underperforming than their Blackbox equivalents. This paper aims to blur the distinction between a post hoc explanation of a Blackbox and constructing interpretable models. Beginning with a Blackbox, we iteratively carve out a mixture of interpretable models and a residual network. The interpretable models identify a subset of samples and explain them using First Order Logic (FOL), providing basic reasoning on concepts from the Blackbox. We route the remaining samples through a flexible residual. We repeat the method on the residual network until all the interpretable models explain the desired proportion of data. Our extensive experiments show that our route, interpret, and repeat approach (1) identifies a richer diverse set of instance-specific concepts with high concept completeness via interpretable models by specializing in various subsets of data without compromising in performance, (2) identifies the relatively “harder” samples to explain via residuals, (3) outperforms the interpretable by-design models by significant margins during test-time interventions, (4) can be used to fix the shortcut learned by the original Blackbox. Shantanu Ghosh, Ke Yu 0002, Forough Arabshahi, Kayhan Batmanghelich |
ICML | 4 |
| 2023 | Distilling BlackBox to Interpretable Models for Efficient Transfer Learning
Shantanu Ghosh, Ke Yu 0002, Kayhan Batmanghelich |
MICCAI (2) | 3 |
| 2023 | Physics-Informed Neural Networks for Tissue Elasticity Reconstruction in Magnetic Resonance Elastography
Matthew Ragoza, Kayhan Batmanghelich |
MICCAI (10) | 2 |
| 2023 | Semi-Implicit Denoising Diffusion Models (SIDDMs)abstractDespite the proliferation of generative models, achieving fast sampling during inference without compromising sample diversity and quality remains challenging. Existing models such as Denoising Diffusion Probabilistic Models (DDPM) deliver high-quality, diverse samples but are slowed by an inherently high number of iterative steps. The Denoising Diffusion Generative Adversarial Networks (DDGAN) attempted to circumvent this limitation by integrating a GAN model for larger jumps in the diffusion process. However, DDGAN encountered scalability limitations when applied to large datasets. To address these limitations, we introduce a novel approach that tackles the problem by matching implicit and explicit factors. More specifically, our approach involves utilizing an implicit model to match the marginal distributions of noisy data and the explicit conditional distribution of the forward diffusion. This combination allows us to effectively match the joint denoising distributions. Unlike DDPM but similar to DDGAN, we do not enforce a parametric distribution for the reverse step, enabling us to take large steps during inference. Similar to the DDPM but unlike DDGAN, we take advantage of the exact form of the diffusion process. We demonstrate that our proposed method obtains comparable generative performance to diffusion-based models and vastly superior results to models with a small number of sampling steps. Yanwu Xu 0003, Mingming Gong, Shaoan Xie, Matthias Grundmann 0002, Kayhan Batmanghelich, Tingbo Hou |
NeurIPS | 6 |
| 2023 | Augmentation by Counterfactual Explanation -Fixing an Overconfident ClassifierabstractA highly accurate but overconfident model is ill-suited for deployment in critical applications such as healthcare and autonomous driving. The classification outcome should reflect a high uncertainty on ambiguous in-distribution samples that lie close to the decision boundary. The model should also refrain from making overconfident decisions on samples that lie far outside its training distribution, far-out-of-distribution (far-OOD), or on unseen samples from novel classes that lie near its training distribution (near-OOD). This paper proposes an application of counterfactual explanations in fixing an over-confident classifier. Specifically, we propose to fine-tune a given pre-trained classifier using augmentations from a counterfactual explainer (ACE) to fix its uncertainty characteristics while retaining its predictive performance. We perform extensive experiments with detecting far-OOD, near-OOD, and ambiguous samples. Our empirical results show that the revised model have improved uncertainty measures, and its performance is competitive to the state-of-the-art methods. Sumedha Singla, Nihal Murali, Forough Arabshahi, Sofia Triantafyllou, Kayhan Batmanghelich |
WACV | 5 |
| 2023 | CrossMoDA 2021 challenge: Benchmark of cross-modality domain adaptation techniques for vestibular schwannoma and cochlea segmentationabstractDomain Adaptation (DA) has recently been of strong interest in the medical imaging community. While a large variety of DA techniques have been proposed for image segmentation, most of these techniques have been validated either on private datasets or on small publicly available datasets. Moreover, these datasets mostly addressed single-class problems. To tackle these limitations, the Cross-Modality Domain Adaptation (crossMoDA) challenge was organised in conjunction with the 24th International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI 2021). CrossMoDA is the first large and multi-class benchmark for unsupervised cross-modality Domain Adaptation. The goal of the challenge is to segment two key brain structures involved in the follow-up and treatment planning of vestibular schwannoma (VS): the VS and the cochleas. Currently, the diagnosis and surveillance in patients with VS are commonly performed using contrast-enhanced T1 (ceT1) MR imaging. However, there is growing interest in using non-contrast imaging sequences such as high-resolution T2 (hrT2) imaging. For this reason, we established an unsupervised cross-modality segmentation benchmark. The training dataset provides annotated ceT1 scans (N=105) and unpaired non-annotated hrT2 scans (N=105). The aim was to automatically perform unilateral VS and bilateral cochlea segmentation on hrT2 scans as provided in the testing set (N=137). This problem is particularly challenging given the large intensity distribution gap across the modalities and the small volume of the structures. A total of 55 teams from 16 countries submitted predictions to the validation leaderboard. Among them, 16 teams from 9 different countries submitted their algorithm for the evaluation phase. The level of performance reached by the top-performing teams is strikingly high (best median Dice score — VS: 88.4%; Cochleas: 85.7%) and close to full supervision (median Dice score — VS: 92.5%; Cochleas: 87.7%). All top-performing methods made use of an image-to-image translation approach to transform the source-domain images into pseudo-target-domain images. A segmentation network was then trained using these generated images and the manual annotations provided for the source image. Reuben Dorent, Aaron Kujawa, Marina Ivory, Spyridon Bakas, Nicola Rieke, Samuel Joutard, Ben Glocker, Manuel Jorge Cardoso, Marc Modat, Kayhan Batmanghelich, Arseniy Belkov, Maria G. Baldeon Calisto, Jae Won Choi, Benoit M. Dawant, Hexin Dong, Sergio Escalera, Yubo Fan, Lasse Hansen, Mattias P. Heinrich, Smriti Joshi, Victoriya Kashtanova, Hyeongyu Kim, Satoshi Kondo, Christian N. Kruse, Susana K. Lai-Yuen, Hao Li 0108, Buntheng Ly, Ipek Oguz, Hyungseob Shin, Boris Shirokikh, Zixian Su, Guotai Wang, Jianghao Wu 0001, Yanwu Xu 0001, Li Zhang 0047, Sébastien Ourselin, Jonathan Shapey, Tom Vercauteren |
Medical Image Anal. | 10 |
| 2023 | Explaining the black-box smoothly - A counterfactual approach
Sumedha Singla, Motahhare Eslami, Brian Pollack, Stephen Wallace, Kayhan Batmanghelich |
Medical Image Anal. | 5 |
| 2022 | Knowledge Distillation via Constrained Variational InferenceabstractKnowledge distillation has been used to capture the knowledge of a teacher model and distill it into a student model with some desirable characteristics such as being smaller, more efficient, or more generalizable. In this paper, we propose a framework for distilling the knowledge of a powerful discriminative model such as a neural network into commonly used graphical models known to be more interpretable (e.g., topic models, autoregressive Hidden Markov Models). Posterior of latent variables in these graphical models (e.g., topic proportions in topic models) is often used as feature representation for predictive tasks. However, these posterior-derived features are known to have poor predictive performance compared to the features learned via purely discriminative approaches. Our framework constrains variational inference for posterior variables in graphical models with a similarity preserving constraint. This constraint distills the knowledge of the discriminative model into the graphical model by ensuring that input pairs with (dis)similar representation in the teacher model also have (dis)similar representation in the student model. By adding this constraint to the variational inference scheme, we guide the graphical model to be a reasonable density model for the data while having predictive features which are as close as possible to those of a discriminative model. To make our framework applicable to a wide range of graphical models, we build upon the Automatic Differentiation Variational Inference (ADVI), a black-box inference framework for graphical models. We demonstrate the effectiveness of our framework on two real-world tasks of disease subtyping and disease trajectory modeling. Ardavan Saeedi, Yuria Utsumi, Li Sun 0010, Kayhan Batmanghelich, Li-Wei H. Lehman |
AAAI | 4 |
| 2022 | Maximum Spatial Perturbation Consistency for Unpaired Image-to-Image TranslationabstractUnpaired image-to-image translation (I2I) is an ill-posed problem, as an infinite number of translation functions can map the source domain distribution to the target distribution. Therefore, much effort has been put into designing suitable constraints, e.g., cycle consistency (CycleGAN), geometry consistency (GCGAN), and contrastive learning-based constraints (CUTGAN), that help better pose the problem. However, these well-known constraints have limitations: (1) they are either too restrictive or too weak for specific I2I tasks; (2) these methods result in content distortion when there is a significant spatial variation between the source and target domains. This paper proposes a universal regularization technique called maximum spatial perturbation consistency (MSPC), which enforces a spatial perturbation function$(T)$and the translation operator$(G)$to be commutative (i.e.,$T\circ G=G\circ T)$. In addition, we introduce two adversarial training components for learning the spatial perturbation function. The first one lets$T$compete with$G$to achieve maximum perturbation. The second one lets$G$and$T$compete with discriminators to align the spatial variations caused by the change of object size, object distortion, background interruptions, etc. Our method outperforms the state-of-the-art methods on most I2I benchmarks. We also introduce a new benchmark, namely the front face to profile face dataset, to emphasize the underlying challenges of I2I for real-world applications. We finally perform ablation experiments to study the sensitivity of our method to the severity of spatial perturbation and its effectiveness for distribution alignment. Yanwu Xu 0003, Shaoan Xie, Kun Zhang 0001, Mingming Gong, Kayhan Batmanghelich |
CVPR | 6 |
| 2022 | Adversarial Consistency for Single Domain Generalization in Medical Image Segmentation
Yanwu Xu 0003, Shaoan Xie, Maxwell Reynolds, Matthew Ragoza, Mingming Gong, Kayhan Batmanghelich |
MICCAI (8) | 6 |
| 2022 | Anatomy-Guided Weakly-Supervised Abnormality Localization in Chest X-rays
Ke Yu 0002, Shantanu Ghosh, Zhexiong Liu, Christopher Deible, Kayhan Batmanghelich |
MICCAI (5) | 5 |
| 2022 | Hierarchical Amortized GAN for 3D High Resolution Medical Image SynthesisabstractGenerative Adversarial Networks (GAN) have many potential medical imaging applications, including data augmentation, domain adaptation, and model explanation. Due to the limited memory of Graphical Processing Units (GPUs), most current 3D GAN models are trained on low-resolution medical images, these models either cannot scale to high-resolution or are prone to patchy artifacts. In this work, we propose a novel end-to-end GAN architecture that can generate high-resolution 3D images. We achieve this goal by using different configurations between training and inference. During training, we adopt a hierarchical structure that simultaneously generates a low-resolution version of the image and a randomly selected sub-volume of the high-resolution image. The hierarchical design has two advantages: First, the memory demand for training on high-resolution images is amortized among sub-volumes. Furthermore, anchoring the high-resolution sub-volumes to a single low-resolution image ensures anatomical consistency between sub-volumes. During inference, our model can directly generate full high-resolution images. We also incorporate an encoder with a similar hierarchical structure into the model to extract features from the images. Experiments on 3D thorax CT and brain MRI demonstrate that our approach outperforms state of the art in image generation. We also demonstrate clinical applications of the proposed model in data augmentation and clinical-relevant feature extraction. Li Sun 0010, Junxiang Chen, Yanwu Xu 0003, Mingming Gong, Ke Yu 0002, Kayhan Batmanghelich |
IEEE J. Biomed. Health Informatics | 6 |
| 2021 | Context Matters: Graph-based Self-supervised Representation Learning for Medical ImagesabstractSupervised learning method requires a large volume of annotated datasets. Collecting such datasets is time-consuming and expensive. Until now, very few annotated COVID-19 imaging datasets are available. Although self-supervised learning enables us to bootstrap the training by exploiting unlabeled data, the generic self-supervised methods for natural images do not sufficiently incorporate the context. For medical images, a desirable method should be sensitive enough to detect deviation from normal-appearing tissue of each anatomical region; here, anatomy is the context. We introduce a novel approach with two levels of self-supervised representation learning objectives: one on the regional anatomical level and another on the patient-level. We use graph neural networks to incorporate the relationship between different anatomical regions. The structure of the graph is informed by anatomical correspondences between each patient and an anatomical atlas. In addition, the graph representation has the advantage of handling any arbitrarily sized image in full resolution. Experiments on large-scale Computer Tomography (CT) datasets of lung images show that our approach compares favorably to baseline methods that do not account for the context. We use the learnt embedding to quantify the clinical progression of COVID-19 and show that our method generalizes well to COVID-19 patients from different hospitals. Qualitative results suggest that our model can identify clinically relevant regions in the images. Li Sun 0010, Ke Yu 0002, Kayhan Batmanghelich |
AAAI | 3 |
| 2021 | Extracting Disease-Relevant Features with Adversarial RegularizationabstractExtracting hidden phenotypes is essential in medical data analysis because it facilitates disease subtyping, diagnosis, and understanding of disease etiology. Since the hidden phenotype is usually a low-dimensional representation that comprehensively describes the disease, we require a dimensionality reduction method that captures as much disease-relevant information as possible. However, most unsupervised or self-supervised methods cannot achieve the goal because they learn a holistic representation containing both disease-relevant and disease-irrelevant information. Supervised methods can capture information that is predictive to the target clinical variable only, but the learned representation is usually not generalizable for the various aspects of the disease. Hence, we develop a dimensionality-reduction approach to extract Disease Relevant Features (DRFs) based on information theory. We propose to use clinical variables that weakly define the disease as so-called anchors. We derive a formulation that makes the DRF predictive of the anchors while forcing the remaining representation to be irrelevant to the anchors via adversarial regularization. We apply our method to a large-scale study of Chronic Obstructive Pulmonary Disease (COPD). Our experiment shows: (1) Learned DRFs are as predictive as the original representation in predicting the anchors, although it is in a significantly lower dimension. (2) Compared to supervised representation, the learned DRFs are more predictive to other relevant disease metrics that are not used during the training. (3) The learned DRFs are related to non-imaging biological measurements such as gene expressions, suggesting the DRFs include information related to the underlying biology of the disease. Junxiang Chen, Li Sun 0010, Ke Yu 0002, Kayhan Batmanghelich |
BIBM | 4 |
| 2021 | Self-supervised Vessel Enhancement Using Flow-Based Consistencies
Rohit Jena, Sumedha Singla, Kayhan Batmanghelich |
MICCAI (2) | 3 |
| 2021 | Using Causal Analysis for Conceptual Deep Learning Explanation
Sumedha Singla, Stephen Wallace, Sofia Triantafyllou, Kayhan Batmanghelich |
MICCAI (3) | 4 |
| 2021 | Can contrastive learning avoid shortcut solutions?abstractThe generalization of representations learned via contrastive learning depends crucially on what features of the data are extracted. However, we observe that the contrastive loss does not always sufficiently guide which features are extracted, a behavior that can negatively impact the performance on downstream tasks via “shortcuts", i.e., by inadvertently suppressing important predictive features. We find that feature extraction is influenced by the difficulty of the so-called instance discrimination task (i.e., the task of discriminating pairs of similar points from pairs of dissimilar ones). Although harder pairs improve the representation of some features, the improvement comes at the cost of suppressing previously well represented features. In response, we propose implicit feature modification (IFM), a method for altering positive and negative samples in order to guide contrastive models towards capturing a wider variety of predictive features. Empirically, we observe that IFM reduces feature suppression, and as a result improves performance on vision and medical imaging tasks. Joshua Robinson 0001, Li Sun 0010, Ke Yu 0002, Kayhan Batmanghelich, Stefanie Jegelka, Suvrit Sra |
NeurIPS | 4 |
| 2021 | Unpaired data empowers association testsabstractMOTIVATION: There is growing interest in the biomedical research community to incorporate retrospective data, available in healthcare systems, to shed light on associations between different biomarkers. Understanding the association between various types of biomedical data, such as genetic, blood biomarkers, imaging, etc. can provide a holistic understanding of human diseases. To formally test a hypothesized association between two types of data in Electronic Health Records (EHRs), one requires a substantial sample size with both data modalities to achieve a reasonable power. Current association test methods only allow using data from individuals who have both data modalities. Hence, researchers cannot take advantage of much larger EHR samples that includes individuals with at least one of the data types, which limits the power of the association test. RESULTS: We present a new method called the Semi-paired Association Test (SAT) that makes use of both paired and unpaired data. In contrast to classical approaches, incorporating unpaired data allows SAT to produce better control of false discovery and to improve the power of the association test. We study the properties of the new test theoretically and empirically, through a series of simulations and by applying our method on real studies in the context of Chronic Obstructive Pulmonary Disease. We are able to identify an association between the high-dimensional characterization of Computed Tomography chest images and several blood biomarkers as well as the expression of dozens of genes involved in the immune system. AVAILABILITY AND IMPLEMENTATION: Code is available on https://github.com/batmanlab/Semi-paired-Association-Test. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mingming Gong, Frank C. Sciurba, Petar Stojanov, Dacheng Tao, George C. Tseng, Kun Zhang 0001, Kayhan Batmanghelich |
Bioinform. | 8 |
| 2020 | Weakly Supervised Disentanglement by Pairwise SimilaritiesabstractRecently, researches related to unsupervised disentanglement learning with deep generative models have gained substantial popularity. However, without introducing supervision, there is no guarantee that the factors of interest can be successfully recovered (Locatello et al. 2018). Motivated by a real-world problem, we propose a setting where the user introduces weak supervision by providing similarities between instances based on a factor to be disentangled. The similarity is provided as either a binary (yes/no) or real-valued label describing whether a pair of instances are similar or not. We propose a new method for weakly supervised disentanglement of latent variables within the framework of Variational Autoencoder. Experimental results demonstrate that utilizing weak supervision improves the performance of the disentanglement method substantially. Junxiang Chen, Kayhan Batmanghelich |
AAAI | 2 |
| 2020 | Generative-Discriminative Complementary LearningabstractThe majority of state-of-the-art deep learning methods are discriminative approaches, which model the conditional distribution of labels given inputs features. The success of such approaches heavily depends on high-quality labeled instances, which are not easy to obtain, especially as the number of candidate classes increases. In this paper, we study the complementary learning problem. Unlike ordinary labels, complementary labels are easy to obtain because an annotator only needs to provide a yes/no answer to a randomly chosen candidate class for each instance. We propose a generative-discriminative complementary learning method that estimates the ordinary labels by modeling both the conditional (discriminative) and instance (generative) distributions. Our method, we call Complementary Conditional GAN (CCGAN), improves the accuracy of predicting ordinary labels and is able to generate high-quality instances in spite of weak supervision. In addition to the extensive empirical studies, we also theoretically show that our model can retrieve the true conditional distribution from the complementarily-labeled data. Yanwu Xu 0003, Mingming Gong, Junxiang Chen, Tongliang Liu, Kun Zhang 0001, Kayhan Batmanghelich |
AAAI | 6 |
| 2020 | Human-Machine Collaboration for Medical Image SegmentationabstractImage segmentation is a ubiquitous step in almost any medical image study. Deep learning-based approaches achieve state-of-the-art in the majority of image segmentation benchmarks. However, end-to-end training of such models requires sufficient annotation. In this paper, we propose a method based on conditional Generative Adversarial Network (cGAN) to address segmentation in semi-supervised setup and in a human-in-the-loop fashion. More specifically, we use the generator in the GAN to synthesize segmentations on unlabeled data and use the discriminator to identify unreliable slices for which expert annotation is required. The quantitative results on a conventional standard benchmark show that our method is comparable with the state-of-the-art fully supervised methods in slice-level evaluation, despite of requiring far less annotated data. Mahdyar Ravanbakhsh, Vadim Tschernezki, Felix Last, Tassilo Klein, Kayhan Batmanghelich, Volker Tresp, Moin Nabi |
ICASSP | 5 |
| 2020 | Explanation by Progressive Exaggeration
Sumedha Singla, Brian Pollack, Junxiang Chen, Kayhan Batmanghelich |
ICLR | 4 |
| 2020 | Label-Noise Robust Domain AdaptationabstractDomain adaptation aims to correct the classifiers when faced with distribution shift between source (training) and target (test) domains. State-of-the-art domain adaptation methods make use of deep networks to extract domain-invariant representations. However, existing methods assume that all the instances in the source domain are correctly labeled; while in reality, it is unsurprising that we may obtain a source domain with noisy labels. In this paper, we are the first to comprehensively investigate how label noise could adversely affect existing domain adaptation methods in various scenarios. Further, we theoretically prove that there exists a method that can essentially reduce the side-effect of noisy source labels in domain adaptation. Specifically, focusing on the generalized target shift scenario, where both label distribution $P_Y$ and the class-conditional distribution $P_{X|Y}$ can change, we discover that the denoising Conditional Invariant Component (DCIC) framework can provably ensures (1) extracting invariant representations given examples with noisy labels in the source domain and unlabeled examples in the target domain and (2) estimating the label distribution in the target domain with no bias. Experimental results on both synthetic and real-world data verify the effectiveness of the proposed method. Xiyu Yu, Tongliang Liu, Mingming Gong, Kun Zhang 0001, Kayhan Batmanghelich, Dacheng Tao |
ICML | 5 |
| 2019 | Geometry-Consistent Generative Adversarial Networks for One-Sided Unsupervised Domain MappingabstractUnsupervised domain mapping aims to learn a function GXY to translate domain X to Y in the absence of paired examples. Finding the optimal GXY without paired data is an ill-posed problem, so appropriate constraints are required to obtain reasonable solutions. While some prominent constraints such as cycle consistency and distance preservation successfully constrain the solution space, they overlook the special properties of images that simple geometric transformations do not change the image's semantic structure. Based on this special property, we develop a geometry-consistent generative adversarial network (Gc-GAN), which enables one-sided unsupervised domain mapping. GcGAN takes the original image and its counterpart image transformed by a predefined geometric transformation as inputs and generates two images in the new domain coupled with the corresponding geometry-consistency constraint. The geometry-consistency constraint reduces the space of possible solutions while keep the correct solutions in the search space. Quantitative and qualitative comparisons with the baseline (GAN alone) and the state-of-the-art methods including CycleGAN [66] and DistanceGAN [5] demonstrate the effectiveness of our method. Huan Fu, Mingming Gong, Chaohui Wang, Kayhan Batmanghelich, Kun Zhang 0001, Dacheng Tao |
CVPR | 4 |
| 2019 | Twin Auxilary Classifiers GANabstractConditional generative models enjoy significant progress over the past few years. One of the popular conditional models is Auxiliary Classifier GAN (AC-GAN) that generates highly discriminative images by extending the loss function of GAN with an auxiliary classifier. However, the diversity of the generated samples by AC-GAN tends to decrease as the number of classes increases. In this paper, we identify the source of low diversity issue theoretically and propose a practical solution to the problem. We show that the auxiliary classifier in AC-GAN imposes perfect separability, which is disadvantageous when the supports of the class distributions have significant overlap. To address the issue, we propose Twin Auxiliary Classifiers Generative Adversarial Net (TAC-GAN) that adds a new player that interacts with other players (the generator and the discriminator) in GAN. Theoretically, we demonstrate that our TAC-GAN can effectively minimize the divergence between generated and real data distributions. Extensive experimental results show that our TAC-GAN can successfully replicate the true data distributions on simulated data, and significantly improves the diversity of class-conditional image generation on real datasets. Mingming Gong, Yanwu Xu 0003, Chunyuan Li, Kun Zhang 0001, Kayhan Batmanghelich |
NeurIPS | 5 |
| 2018 | Robust Angular Local Descriptor Learning
Yanwu Xu 0001, Mingming Gong, Tongliang Liu, Kayhan Batmanghelich, Chaohui Wang |
ACCV (5) | 4 |
| 2018 | Deep Ordinal Regression Network for Monocular Depth EstimationabstractMonocular depth estimation, which plays a crucial role in understanding 3D scene geometry, is an ill-posed problem. Recent methods have gained significant improvement by exploring image-level information and hierarchical features from deep convolutional neural networks (DCNNs). These methods model depth estimation as a regression problem and train the regression networks by minimizing mean squared error, which suffers from slow convergence and unsatisfactory local solutions. Besides, existing depth estimation networks employ repeated spatial pooling operations, resulting in undesirable low-resolution feature maps. To obtain high-resolution depth maps, skip-connections or multilayer deconvolution networks are required, which complicates network training and consumes much more computations. To eliminate or at least largely reduce these problems, we introduce a spacing-increasing discretization (SID) strategy to discretize depth and recast depth network learning as an ordinal regression problem. By training the network using an ordinary regression loss, our method achieves much higher accuracy and faster convergence in synch. Furthermore, we adopt a multi-scale network structure which avoids unnecessary spatial pooling and captures multi-scale information in parallel. The proposed deep ordinal regression network (DORN) achieves state-of-the-art results on three challenging benchmarks, i.e., KITTI [16], Make3D [49], and NYU Depth v2 [41], and outperforms existing methods by a large margin. Huan Fu, Mingming Gong, Chaohui Wang, Kayhan Batmanghelich, Dacheng Tao |
CVPR | 4 |
| 2018 | An Efficient and Provable Approach for Mixture Proportion Estimation Using Linear Independence AssumptionabstractIn this paper, we study the mixture proportion estimation (MPE) problem in a new setting: given samples from the mixture and the component distributions, we identify the proportions of the components in the mixture distribution. To address this problem, we make use of a linear independence assumption, i.e., the component distributions are independent from each other, which is much weaker than assumptions exploited in the previous MPE methods. Based on this assumption, we propose a method (1) that uniquely identifies the mixture proportions, (2) whose output provably converges to the optimal solution, and (3) that is computationally efficient. We show the superiority of the proposed method over the state-of-the-art methods in two applications including learning with label noise and semi-supervised learning on both synthetic and real-world datasets. Xiyu Yu, Tongliang Liu, Mingming Gong, Kayhan Batmanghelich, Dacheng Tao |
CVPR | 4 |
| 2018 | Subject2Vec: Generative-Discriminative Approach from a Set of Image Patches to a Vector
Sumedha Singla, Mingming Gong, Siamak Ravanbakhsh, Frank C. Sciurba, Barnabás Póczos, Kayhan Batmanghelich |
MICCAI (1) | 6 |
| 2018 | Causal Discovery with Linear Non-Gaussian Models under Measurement Error: Structural Identifiability Results
Kun Zhang 0001, Mingming Gong, Joseph D. Ramsey, Kayhan Batmanghelich, Peter Spirtes, Clark Glymour |
UAI | 4 |
| 2017 | Transformations Based on Continuous Piecewise-Affine Velocity FieldsabstractWe propose novel finite-dimensional spaces of well-behaved transformations. The latter are obtained by (fast and highly-accurate) integration of continuous piecewise-affine velocity fields. The proposed method is simple yet highly expressive, effortlessly handles optional constraints (e.g., volume preservation and/or boundary conditions), and supports convenient modeling choices such as smoothing priors and coarse-to-fine analysis. Importantly, the proposed approach, partly due to its rapid likelihood evaluations and partly due to its other properties, facilitates tractable inference over rich transformation spaces, including using Markov-Chain Monte-Carlo methods. Its applications include, but are not limited to: monotonic regression (more generally, optimization over monotonic functions); modeling cumulative distribution functions or histograms; time-warping; image warping; image registration; real-time diffeomorphic image editing; data augmentation for image classifiers. Our GPU-based code is publicly available. Oren Freifeld, Søren Hauberg, Kayhan Batmanghelich, John W. Fisher III |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2015 | Highly-Expressive Spaces of Well-Behaved Transformations: Keeping it SimpleabstractWe propose novel finite-dimensional spaces of Rn→ Rntransformations, n ∈ {1, 2, 3}, derived from (continuously-defined) parametric stationary velocity fields. Particularly, we obtain these transformations, which are diffeomorphisms, by fast and highly-accurate integration of continuous piecewise-affine velocity fields, we also provide an exact solution for n = 1. The simple-yet-highly-expressive proposed representation handles optional constraints (e.g., volume preservation) easily and supports convenient modeling choices and rapid likelihood evaluations (facilitating tractable inference over latent transformations). Its applications include, but are not limited to: unconstrained optimization over monotonic functions, modeling cumulative distribution functions or histograms, time warping, image registration, landmark-based warping, real-time diffeomorphic image editing. Our code is available at https://github.com/freifeld/cpabDiffeo. Oren Freifeld, Søren Hauberg, Kayhan Batmanghelich, John W. Fisher III |
ICCV | 3 |
| 2012 | Dominant Component Analysis of Electrophysiological Connectivity Networks
Yasser Ghanbari, Luke Bloy, Kayhan Batmanghelich, Timothy P. L. Roberts, Ragini Verma |
MICCAI (3) | 3 |
| 2012 | Generative-Discriminative Basis Learning for Medical ImagingabstractThis paper presents a novel dimensionality reduction method for classification in medical imaging. The goal is to transform very high-dimensional input (typically, millions of voxels) to a low-dimensional representation (small number of constructed features) that preserves discriminative signal and is clinically interpretable. We formulate the task as a constrained optimization problem that combines generative and discriminative objectives and show how to extend it to the semi-supervised learning (SSL) setting. We propose a novel large-scale algorithm to solve the resulting optimization problem. In the fully supervised case, we demonstrate accuracy rates that are better than or comparable to state-of-the-art algorithms on several datasets while producing a representation of the group difference that is consistent with prior clinical reports. Effectiveness of the proposed algorithm for SSL is evaluated with both benchmark and medical imaging datasets. In the benchmark datasets, the results are better than or comparable to the state-of-the-art methods for SSL. For evaluation of the SSL setting in medical datasets, we use images of subjects with mild cognitive impairment (MCI), which is believed to be a precursor to Alzheimer's disease (AD), as unlabeled data. AD subjects and normal control (NC) subjects are used as labeled data, and we try to predict conversion from MCI to AD on follow-up. The semi-supervised extension of this method not only improves the generalization accuracy for the labeled data (AD/NC) slightly but is also able to predict subjects which are likely to converge to AD. Kayhan Batmanghelich, Ben Taskar, Christos Davatzikos |
IEEE Trans. Medical Imaging | 1 |
| 2011 | Regularized Tensor Factorization for Multi-Modality Medical Image Classification
Kayhan Batmanghelich, Aoyan Dong, Ben Taskar, Christos Davatzikos |
MICCAI (3) | 1 |
| 2006 | Distributed Behavior-based Multi-agent System for Automatic Segmentation of Brain MR ImagesabstractA novel multi-agent image segmentation system for MR Images is proposed. The agents are behavior-based autonomous entities that are situated in the image as their environment. The agents have local knowledge and cooperate in an implicit and simple manner to achieve globally rational team behavior. Incremental design of agents and their simple interactions facilitate incorporation of expert knowledge and properties of the specific application in the agents' mind. The system has been applied to MRI brain scans for segmentation of three pairs of structures. The results are satisfactory in terms of both average performance and robustness. The system is also proved to be robust against initialization of the agents. Hadi Fatemi Shariatpanahi, Kayhan Batmanghelich, Amir R. M. Kermani, Majid Nili Ahmadabadi, Hamid Soltanian-Zadeh |
IJCNN | 2 |