VLDB 2026 Research / reviewers in the wild / expert
Ehsan Adeli-Mosabbeb
dblp:93/2941 · also Ehsan Adeli 0001
· DBLP profile ↗
111ranked-venue papers
10as first author
63since 2021 · last 2026
0000-0002-0579-7763ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 67 · 4 first-author · 34 since 2021Artificial intelligence and machine learning · 53 · 6 first-author · 32 since 2021Applied, interdisciplinary, general and emerging computing · 48 · 3 first-author · 27 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SocialGen: Modeling Multi-Human Social Interaction with Language ModelsabstractHuman interactions in everyday life are inherently social, involving engagements with diverse individuals across various contexts. Modeling these social interactions is fundamental to a wide range of real-world applications. In this paper, we introduce SocialGen, the first unified motionlanguage model capable of modeling interaction behaviors among varying numbers of individuals, to address this crucial yet challenging problem. Unlike prior methods that are limited to two-person interactions, we propose a novel social motion representation that supports tokenizing the motions of an arbitrary number of individuals and aligning them with the language space. This alignment enables the model to leverage rich, pretrained linguistic knowledge to better understand and reason about human social behaviors. To tackle the challenges of data scarcity, we curate a comprehensive multi-human interaction dataset, SocialX, enriched with textual annotations. Leveraging this dataset, we establish the first comprehensive benchmark for multihuman interaction tasks. Our method achieves state-of-theart performance across motion-language tasks, setting a new standard for multi-human interaction modeling. Our dataset and source code will be made publicly available. Juze Zhang, Changan Chen, Tiange Xiang, Yusu Fang, Juan Carlos Niebles, Ehsan Adeli-Mosabbeb |
3DV | 7 |
| 2026 | Integrating Anatomical Priors Into a Causal Diffusion Modelabstract3D brain MRI studies often examine subtle morphometric differences between cohorts that are hard to detect visually. Given the high cost of MRI acquisition, these studies could greatly benefit from image syntheses, particularly counterfactual image generation, as has been the case for applications in computer vision. However, counterfactual models struggle to produce anatomically plausible MRIs due to a lack of explicit inductive biases to preserve fine-grained anatomical details. This shortcoming arises from the training of models that optimize overall image appearance (e.g., via cross-entropy) rather than preserving subtle, yet medically relevant, local variations across subjects. To preserve subtle variations, we propose to explicitly integrate anatomical constraints at the voxel level as priors into a generative diffusion framework. Termed Probabilistic Causal Graph Model (PCGM), the approach captures anatomical constraints via a probabilistic graph module and translates those constraints into spatial binary masks of regions where subtle variations occur. The masks (encoded by a 3D ControlNet) constrain a novel counterfactual denoising UNet, whose encodings are then transferred into high-quality brain MRIs via our 3D diffusion decoder. Extensive experiments across multiple datasets demonstrate that PCGM generates structural brain MRIs of higher quality than several baseline approaches. Furthermore, we show, for the first time, that brain measurements extracted from counterfactuals (generated by PCGM) replicate the subtle effects of a disease on cortical brain regions previously reported in the neuroscience literature. This achievement is an important milestone in the use of synthetic MRIs in studies investigating subtle morphological differences. The codes are available at https://github.com/AndyCA111/PCGM. Binxu Li, Wei Peng 0009, Mingjie Li 0006, Ehsan Adeli-Mosabbeb, Kilian M. Pohl |
IEEE Trans. Medical Imaging | 4 |
| 2025 | NeuHMR: Neural Rendering-Guided Human Motion ReconstructionabstractReconstructing 3D human movements from video sequences is an important task in the fields of computer vision, graphics, and biomechanics. Although much progress has been made to infer 3D human mesh based on visual contexts provided in video sequences, generalization to in-the-wild videos still remains challenging for existing human mesh recovery (HMR) methods. To overcome inaccurate prediction, they can perform a second step optimization that refines the inaccurate estimations continuously at test time. Most optimization methods seek fitting of the body joints in the image space with respect to pseudo ground truth predicted by an off-the-shelf key point detector. However, state-of-theart detectors still introduce errors, especially for challenging poses. In this work, we rethink the dependency on the 2D key point fitting paradigm and present NeuHMR, an optimization-based mesh recovery framework based on recent advances in neural rendering. Our method builds on Human Neural Radiance Fields that allow the refinement of human meshes through animatable$2 D$renderings. We evaluated our method on two common benchmarks and validated its effectiveness. Tiange Xiang, Kuan-Chieh Wang, Jaewoo Heo, Ehsan Adeli-Mosabbeb, Serena Yeung-Levy, Scott L. Delp, Li Fei-Fei 0001 |
3DV | 4 |
| 2025 | The Language of Motion: Unifying Verbal and Non-verbal Language of 3D Human MotionabstractHuman communication is inherently multimodal, involving a combination of verbal and non-verbal cues such as speech, facial expressions, and body gestures. Modeling these behaviors is essential for understanding human interaction and for creating virtual characters that can communicate naturally in applications like games, films, and virtual reality. However, existing motion generation models are typically limited to specific input modalities-either speech, text, or motion data-and cannot fully leverage the diversity of available data. In this paper, we propose a novel framework that unifies verbal and non-verbal language using multimodal language models for human motion understanding and generation. This model is flexible in taking text, speech, and motion or any combination of them as input. Coupled with our novel pre-training strategy, our model not only achieves state-of-the-art performance on co-speech gesture generation but also requires much less data for training. Our model also unlocks an array of novel tasks such as editable gesture generation and emotion prediction from motion. We believe unifying the verbal and non-verbal language of human motion is essential for real-world applications, and language models offer a powerful approach to achieving this goal. Project page: languageofmotion.github.io. Changan Chen, Juze Zhang, Shrinidhi K. Lakshmikanth, Yusu Fang, Ruizhi Shao, Gordon Wetzstein, Li Fei-Fei 0001, Ehsan Adeli-Mosabbeb |
CVPR | 8 |
| 2025 | Re-thinking Temporal Search for Long-Form Video UnderstandingabstractEfficiently understanding long-form videos remains a significant challenge in computer vision. In this work, we revisit temporal search paradigms for long-form video understanding and address a fundamental issue pertaining to all state-of-the-art (SOTA) long-context vision-language models (VLMs). Our contributions are twofold: First, we frame temporal search as a Long Video Haystack problem – finding a minimal set of relevant frames (e.g., one to five) from tens of thousands based on specific queries. Upon this formulation, we introduce LV-Haystack, the first dataset with 480 hours of videos, 15,092 human-annotated instances for both training and evaluation aiming to improve temporal search quality and efficiency. Results on LV-HAYSTACK highlight a significant research gap in temporal search capabilities, with current SOTA search methods only achieving 2.1% temporal F1score on the LongVideoBench subset.Next, inspired by visual search in images, we propose a lightweight temporal search framework, T* that reframes costly temporal search as spatial search. T* leverages powerful visual localization techniques commonly used in images and introduces an adaptive zooming-in mechanism that operates across both temporal and spatial dimensions. Extensive experiments show that integrating T* with existing methods significantly improves SOTA long-form video understanding. Under an inference budget of 32 frames, T* improves GPT-4o’s performance from 50.5% to 53.1% and LLaVAOneVision-OV-72B’s performance from 56.5% to 62.4% on the LongVideoBench XL subset. Our code, benchmark, and models are provided in the Supplementary material. Jinhui Ye, Zihan Wang 0008, Haosen Sun, Keshigeyan Chandrasegaran, Zane Durante, Cristóbal Eyzaguirre, Yonatan Bisk, Juan Carlos Niebles, Ehsan Adeli-Mosabbeb, Li Fei-Fei 0001, Jiajun Wu 0001, Manling Li |
CVPR | 9 |
| 2025 | Latent Drifting in Diffusion Models for Counterfactual Medical Image SynthesisabstractScaling by training on large datasets has been shown to enhance the quality and fidelity of image generation and manipulation with diffusion models; however, such large datasets are not always accessible in medical imaging due to cost and privacy issues, which contradicts one of the main applications of such models to produce synthetic samples where real data is scarce. Also, fine-tuning on pre-trained general models has been a challenge due to the distribution shift between the medical domain and the pre-trained models. Here, we propose Latent Drift (LD) for diffusion models that can be adopted for any fine-tuning method to mitigate the issues faced by the distribution shift or employed in inference time as a condition. Latent Drifting enables diffusion models to be conditioned for medical images fitted for the complex task of counterfactual image generation, which is crucial to investigate how parameters such as gender, age, and adding or removing diseases in a patient would alter the medical images. We evaluate our method on three public longitudinal benchmark datasets of brain MRI and chest X-rays for counterfactual image generation. Our results demonstrate significant performance gains in various scenarios when combined with different fine-tuning schemes. Yousef Yeganeh, Azade Farshad, Ioannis Charisiadis, Marta Hasny, Martin Hartenberger, Björn Ommer, Nassir Navab, Ehsan Adeli-Mosabbeb |
CVPR | 8 |
| 2025 | LOMM: Latest Object Memory Management for Temporally Consistent Video Instance SegmentationabstractIn this paper, we introduce Latest Object Memory (LOM), a system for robustly tracking and continuously updating the latest states of objects by explicitly modeling their presence across video frames. LOM enables consistent tracking and accurate identity management across frames, enhancing both performance and reliability through the video segmentation process. Building upon LOM, we present Latest Object Memory Management (LOMM) for temporally consistent video instance segmentation, significantly improving long-term instance tracking. This enables consistent tracking and accurate identity management across frames, enhancing both performance and reliability through the video segmentation process. Moreover, we introduce Decoupled Object Association (DOA), a strategy that separately handles newly appearing and already existing objects. By leveraging our memory system, DOA accurately assigns object indices, improving matching accuracy and ensuring stable identity consistency, even in dynamic scenes where objects frequently appear and disappear. Extensive experiments and ablation studies demonstrate the superiority of our method over traditional approaches, setting a new state-of-the-art in video instance segmentation. Notably, our LOMM achieves an AP score of 54.0 on YouTube-VIS 2022, a dataset known for its challenging long videos. Project page: this https URL. Seunghun Lee 0002, Jiwan Seo, Minwoo Choi, Kiljoon Han, Zane Durante, Ehsan Adeli-Mosabbeb, Sunghoon Im 0001 |
ICCV | 7 |
| 2025 | UniEgoMotion: A Unified Model for Egocentric Motion Reconstruction, Forecasting, and GenerationabstractEgocentric human motion generation and forecasting with scene-context is crucial for enhancing AR/VR experiences, improving human-robot interaction, advancing assistive technologies, and enabling adaptive healthcare solutions by accurately predicting and simulating movement from a first-person perspective. However, existing methods primarily focus on third-person motion synthesis with structured 3D scene contexts, limiting their effectiveness in real-world egocentric settings where limited field of view, frequent occlusions, and dynamic cameras hinder scene perception. To bridge this gap, we introduce Egocentric Motion Generation and Egocentric Motion Forecasting, two novel tasks that utilize first-person images for scene-aware motion synthesis without relying on explicit 3D scene. We propose UniEgoMotion, a unified conditional motion diffusion model with a novel head-centric motion representation tailored for egocentric devices. UniEgoMotion's simple yet effective design supports egocentric motion reconstruction, forecasting, and generation from first-person visual inputs in a unified framework. Unlike previous works that overlook scene semantics, our model effectively extracts image-based scene context to infer plausible 3D motion. To facilitate training, we introduce EE4D-Motion, a large-scale dataset derived from EgoExo4D, augmented with pseudo-ground-truth 3D motion annotations. UniEgoMotion achieves state-of-the-art performance in egocentric motion reconstruction and is the first to generate motion from a single egocentric image. Extensive evaluations demonstrate the effectiveness of our unified framework, setting a new benchmark for egocentric motion modeling and unlocking new possibilities for egocentric applications. Chaitanya Patel, Hiroki Nakamura, Yuta Kyuragi, Kazuki Kozuka, Juan Carlos Niebles, Ehsan Adeli-Mosabbeb |
ICCV | 6 |
| 2025 | Repurposing 2D Diffusion Models with Gaussian Atlas for 3D GenerationabstractRecent advances in text-to-image diffusion models have been driven by the increasing availability of paired 2D data. However, the development of 3D diffusion models has been hindered by the scarcity of high-quality 3D data, resulting in less competitive performance compared to their 2D counterparts. To address this challenge, we propose repurposing pre-trained 2D diffusion models for 3D object generation. We introduce Gaussian Atlas, a novel representation that utilizes dense 2D grids, enabling the fine-tuning of 2D diffusion models to generate 3D Gaussians. Our approach demonstrates successful transfer learning from a pre-trained 2D diffusion model to a 2D manifold flattened from 3D structures. To support model training, we compile GaussianVerse, a large-scale dataset comprising 205K high-quality 3D Gaussian fittings of various 3D objects. Our experimental results show that text-to-image diffusion models can be effectively adapted for 3D content generation, bridging the gap between 2D and 3D modeling. Tiange Xiang, Chengjiang Long, Christian Häne, Peihong Guo, Scott L. Delp, Ehsan Adeli-Mosabbeb, Li Fei-Fei 0001 |
ICCV | 7 |
| 2025 | Confounder-Free Continual Learning via Recursive Feature NormalizationabstractConfounders are extraneous variables that affect both the input and the target, resulting in spurious correlations and biased predictions. There are recent advances in dealing with or removing confounders in traditional models, such as metadata normalization (MDN), where the distribution of the learned features is adjusted based on the study confounders. However, in the context of continual learning, where a model learns continuously from new data over time without forgetting, learning feature representations that are invariant to confounders remains a significant challenge. To remove their influence from intermediate feature representations, we introduce the Recursive MDN (R-MDN) layer, which can be integrated into any deep learning architecture, including vision transformers, and at any model stage. R-MDN performs statistical regression via the recursive least squares algorithm to maintain and continually update an internal model state with respect to changing distributions of data and confounding variables. Our experiments demonstrate that R-MDN promotes equitable predictions across population groups, both within static learning and across different stages of continual learning, by reducing catastrophic forgetting caused by confounder effects changing over time. Camila González, Mohammad H. Abbasi, Qingyu Zhao, Kilian M. Pohl, Ehsan Adeli-Mosabbeb |
ICML | 6 |
| 2025 | WASABI: A Metric for Evaluating Morphometric Plausibility of Synthetic Brain MRIs
Bahram Jafrasteh, Wei Peng 0009, Yimin Luo, Ehsan Adeli-Mosabbeb, Qingyu Zhao |
MICCAI (2) | 5 |
| 2025 | Generating Novel Brain Morphology by Deforming Learned Templates
Alan Q. Wang 0001, Fangrui Huang, Bailey Trang Nguyen, Wei Peng 0009, Mohammad H. Abbasi, Kilian M. Pohl, Mert R. Sabuncu, Ehsan Adeli-Mosabbeb |
MICCAI (2) | 8 |
| 2025 | Discovering Latent Graphs with GFlowNets for Diverse Conditional Image GenerationabstractCapturing diversity is crucial in conditional and prompt-based image generation, particularly when conditions contain uncertainty that can lead to multiple plausible outputs. To generate diverse images reflecting this diversity, traditional methods often modify random seeds, making it difficult to discern meaningful differences between samples, or diversify the input prompt, which is limited in verbally interpretable diversity. We propose \modelnamenospace, a novel conditional image generation framework, applicable to any pretrained conditional generative model, that addresses inherent condition/prompt uncertainty and generates diverse plausible images. \modelname is based on a simple yet effective idea: decomposing the input condition into diverse latent representations, each capturing an aspect of the uncertainty and generating a distinct image. First, we integrate a latent graph, parameterized by Generative Flow Networks (GFlowNets), into the prompt representation computation. Second, leveraging GFlowNets' advanced graph sampling capabilities to capture uncertainty and output diverse trajectories over the graph, we produce multiple trajectories that collectively represent the input condition, leading to diverse condition representations and corresponding output images. Evaluations on natural image and medical image datasets demonstrate \modelnamenospace’s improvement in both diversity and fidelity across image synthesis, image generation, and counterfactual generation tasks. Bailey Trang Nguyen, Parham Saremi, Alan Q. Wang 0001, Fangrui Huang, Zahra Tehraninasab, Amar Kumar, Tal Arbel, Li Fei-Fei 0001, Ehsan Adeli-Mosabbeb |
NeurIPS | 9 |
| 2025 | Efficient one-shot federated learning on medical data using knowledge distillation with image synthesis and client model adaptation
Myeongkyun Kang, Philip Chikontwe, Soopil Kim, Kyong Hwan Jin, Ehsan Adeli-Mosabbeb, Kilian M. Pohl, Sanghyun Park 0004 |
Medical Image Anal. | 5 |
| 2025 | Guest Editorial: Applications of Intelligent Environments to Health
Miguel J. Hornos, Ehsan Adeli-Mosabbeb, Víctor Zamudio 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | Communication Efficient Federated Learning for Multi-Organ Segmentation via Knowledge Distillation With Image SynthesisabstractFederated learning (FL) methods for multi-organ segmentation in CT scans are gaining popularity, but generally require numerous rounds of parameter exchange between a central server and clients. This repetitive sharing of parameters between server and clients may not be practical due to the varying network infrastructures of clients and the large transmission of data. Further increasing repetitive sharing results from data heterogeneity among clients, i.e., clients may differ with respect to the type of data they share. For example, they might provide label maps of different organs (i.e. partial labels) as segmentations of all organs shown in the CT are not part of their clinical protocol. To this end, we propose an efficient communication approach for FL with partial labels. Specifically, parameters of local models are transmitted once to a central server and the global model is trained via knowledge distillation (KD) of the local models. While one can make use of unlabeled public data as inputs for KD, the model accuracy is often limited due to distribution shifts between local and public datasets. Herein, we propose to generate synthetic images from clients' models as additional inputs to mitigate data shifts between public and local data. In addition, our proposed method offers flexibility for additional finetuning through several rounds of communication using existing FL algorithms, leading to enhanced performance. Extensive evaluation on public datasets in few communication FL scenario reveals that our approach substantially improves over state-of-the-art methods. Soopil Kim, Heejung Park, Philip Chikontwe, Myeongkyun Kang, Kyong Hwan Jin, Ehsan Adeli-Mosabbeb, Kilian M. Pohl, Sanghyun Park 0004 |
IEEE Trans. Medical Imaging | 6 |
| 2024 | Few Shot Part Segmentation Reveals Compositional Logic for Industrial Anomaly DetectionabstractLogical anomalies (LA) refer to data violating underlying logical constraints e.g., the quantity, arrangement, or composition of components within an image. Detecting accurately such anomalies requires models to reason about various component types through segmentation. However, curation of pixel-level annotations for semantic segmentation is both time-consuming and expensive. Although there are some prior few-shot or unsupervised co-part segmentation algorithms, they often fail on images with industrial object. These images have components with similar textures and shapes, and a precise differentiation proves challenging. In this study, we introduce a novel component segmentation model for LA detection that leverages a few labeled samples and unlabeled images sharing logical constraints. To ensure consistent segmentation across unlabeled images, we employ a histogram matching loss in conjunction with an entropy loss. As segmentation predictions play a crucial role, we propose to enhance both local and global sample validity detection by capturing key aspects from visual semantics via three memory banks: class histograms, component composition embeddings and patch-level representations. For effective LA detection, we propose an adaptive scaling strategy to standardize anomaly scores from different memory banks in inference. Extensive experiments on the public benchmark MVTec LOCO AD reveal our method achieves 98.1% AUROC in LA detection vs. 89.6% from competing methods. Soopil Kim, Sion An, Philip Chikontwe, Myeongkyun Kang, Ehsan Adeli-Mosabbeb, Kilian M. Pohl, Sanghyun Park 0004 |
AAAI | 5 |
| 2024 | Few-Shot Classification of Interactive Activities of Daily Living (InteractADL)
Zane Durante, Robathan Harries, Edward Vendrow, Zelun Luo, Yuta Kyuragi, Kazuki Kozuka, Li Fei-Fei 0001, Ehsan Adeli-Mosabbeb |
BMVC | 8 |
| 2024 | Towards Robust 3D Pose Transfer with Adversarial Learningabstract3D pose transfer that aims to transfer the desired pose to a target mesh is one of the most challenging 3D generation tasks. Previous attempts rely on well-defined parametric human models or skeletal joints as driving pose sources. However, to obtain those clean pose sources, cumbersome but necessary pre-processing pipelines are inevitable, hindering implementations of the real-time applications. This work is driven by the intuition that the robustness of the model can be enhanced by introducing adversarial samples into the training, leading to a more invulnerable model to the noisy inputs, which even can be further extended to directly handling the real-world data like raw point clouds/scans without intermediate processing. Furthermore, we propose a novel 3D pose Masked Autoencoder (3D-PoseMAE), a customized MAE that effectively learns 3D extrinsic presentations (i.e., pose). 3D-PoseMAE facilitates learning from the aspect of extrinsic attributes by simultaneously generating adversarial samples that perturb the model and learning the arbitrary raw noisy poses via a multi-scale masking strategy. Both qualitative and quantitative studies show that the transferred meshes given by our network result in much better quality. Besides, we demonstrate the strong generalizability of our method on various poses, different domains, and even raw scans. Experimental results also show meaningful insights that the intermediate adversarial samples generated in the training can success-fully attack the existing pose transfer models. Haoyu Chen 0001, Hao Tang 0005, Ehsan Adeli-Mosabbeb, Guoying Zhao 0001 |
CVPR | 3 |
| 2024 | H-ViT: A Hierarchical Vision Transformer for Deformable Image RegistrationabstractThis paper introduces a novel top-down representation approach for deformable image registration, which estimates the deformation field by capturing various short-and long-range flow features at different scale levels. As a Hierarchical Vision Transformer (H- ViT), we propose a dual self-attention and cross-attention mechanism that uses high-level features in the deformation field to represent low-level ones, enabling information streams in the deformation field across all voxel patch embeddings irrespective of their spatial proximity. Since high-level features contain abstract flow patterns, such patterns are expected to effectively contribute to the representation of the deformation field in lower scales. When the self-attention module utilizes within-scale short-range patterns for representation, the cross-attention modules dynamically look for the key tokens across different scales to further interact with the local query voxel patches. Our method shows superior accuracy and visual quality over the state-of-the-art registration methods in five publicly available datasets, highlighting a substantial enhancement in the performance of medical imaging registration. The project link is available at https://mogvision.github.io/hvit. Morteza Ghahremani, Mohammad Khateri, Bailiang Jian, Benedikt Wiestler, Ehsan Adeli-Mosabbeb, Christian Wachinger |
CVPR | 5 |
| 2024 | SOM2LM: Self-Organized Multi-Modal Longitudinal Maps
Jiahong Ouyang, Qingyu Zhao, Ehsan Adeli-Mosabbeb, Greg Zaharchuk, Kilian M. Pohl |
MICCAI (2) | 3 |
| 2024 | OccFusion: Rendering Occluded Humans with Generative Diffusion PriorsabstractExisting human rendering methods require every part of the human to be fully visible throughout the input video. However, this assumption does not hold in real-life settings where obstructions are common, resulting in only partial visibility of the human. Considering this, we present OccFusion, an approach that utilizes efficient 3D Gaussian splatting supervised by pretrained 2D diffusion models for efficient and high-fidelity human rendering. We propose a pipeline consisting of three stages. In the Initialization stage, complete human masks are generated from partial visibility masks. In the Optimization stage, 3D human Gaussians are optimized with additional supervisions by Score-Distillation Sampling (SDS) to create a complete geometry of the human. Finally, in the Refinement stage, in-context inpainting is designed to further improve rendering quality on the less observed human body parts. We evaluate OccFusion on ZJU-MoCap and challenging OcMotion sequences and found that it achieves state-of-the-art performance in the rendering of occluded humans. Adam Sun, Tiange Xiang, Scott L. Delp, Li Fei-Fei 0001, Ehsan Adeli-Mosabbeb |
NeurIPS | 5 |
| 2024 | Vision-based estimation of fatigue and engagement in cognitive training sessions
Yanchen Wang, Adam Turnbull, Yunlong Xu 0001, Kathi L. Heffner, Feng Lin 0008, Ehsan Adeli-Mosabbeb |
Artif. Intell. Medicine | 6 |
| 2024 | TransUNet: Rethinking the U-Net architecture design for medical image segmentation through the lens of transformersabstractMedical image segmentation is crucial for healthcare, yet convolution-based methods like U-Net face limitations in modeling long-range dependencies. To address this, Transformers designed for sequence-to-sequence predictions have been integrated into medical image segmentation. However, a comprehensive understanding of Transformers' self-attention in U-Net components is lacking. TransUNet, first introduced in 2021, is widely recognized as one of the first models to integrate Transformer into medical image analysis. In this study, we present the versatile framework of TransUNet that encapsulates Transformers' self-attention into two key modules: (1) a Transformer encoder tokenizing image patches from a convolution neural network (CNN) feature map, facilitating global context extraction, and (2) a Transformer decoder refining candidate regions through cross-attention between proposals and U-Net features. These modules can be flexibly inserted into the U-Net backbone, resulting in three configurations: Encoder-only, Decoder-only, and Encoder+Decoder. TransUNet provides a library encompassing both 2D and 3D implementations, enabling users to easily tailor the chosen architecture. Our findings highlight the encoder's efficacy in modeling interactions among multiple abdominal organs and the decoder's strength in handling small targets like tumors. It excels in diverse medical applications, such as multi-organ segmentation, pancreatic tumor segmentation, and hepatic vessel segmentation. Notably, our TransUNet achieves a significant average Dice improvement of 1.06% and 4.30% for multi-organ segmentation and pancreatic tumor segmentation, respectively, when compared to the highly competitive nn-UNet, and surpasses the top-1 solution in the BrasTS2021 challenge. 2D/3D Code and models are available at https://github.com/Beckschen/TransUNet and https://github.com/Beckschen/TransUNet-3D, respectively. Jieneng Chen, Jieru Mei, Xianhang Li, Yongyi Lu, Qihang Yu, Qingyue Wei, Xiangde Luo, Yutong Xie 0001, Ehsan Adeli-Mosabbeb, Yan Wang 0033, Matthew P. Lungren, Shaoting Zhang 0001, Lei Xing 0001, Le Lu 0001, Alan L. Yuille, Yuyin Zhou |
Medical Image Anal. | 9 |
| 2024 | Federated learning with knowledge distillation for multi-organ segmentation with partially labeled datasets
Soopil Kim, Heejung Park, Myeongkyun Kang, Kyong Hwan Jin, Ehsan Adeli-Mosabbeb, Kilian M. Pohl, Sanghyun Park 0004 |
Medical Image Anal. | 5 |
| 2024 | Metadata-conditioned generative models to synthesize anatomically-plausible 3D brain MRIs
Wei Peng 0009, Tomas M. Bosschieter, Jiahong Ouyang, Robert Paul, Edith V. Sullivan, Adolf Pfefferbaum, Ehsan Adeli-Mosabbeb, Qingyu Zhao, Kilian M. Pohl |
Medical Image Anal. | 7 |
| 2024 | Medical Image Segmentation Review: The Success of U-NetabstractAutomatic medical image segmentation is a crucial topic in the medical domain and successively a critical counterpart in the computer-aided diagnosis paradigm. U-Net is the most widespread image segmentation architecture due to its flexibility, optimized modular design, and success in all medical image modalities. Over the years, the U-Net model has received tremendous attention from academic and industrial researchers who have extended it to address the scale and complexity created by medical tasks. These extensions are commonly related to enhancing the U-Net's backbone, bottleneck, or skip connections, or including representation learning, or combining it with a Transformer architecture, or even addressing probabilistic prediction of the segmentation map. Having a compendium of different previously proposed U-Net variants makes it easier for machine learning researchers to identify relevant research questions and understand the challenges of the biological tasks that challenge the model. In this work, we discuss the practical aspects of the U-Net model and organize each variant model into a taxonomy. Moreover, to measure the performance of these strategies in a clinical application, we propose fair evaluations of some unique and famous designs on well-known datasets. Furthermore, we provide a comprehensive implementation library with trained models. In addition, for ease of future studies, we created an online list of U-Net papers with their possible official implementation. Reza Azad, Ehsan Khodapanah Aghdam, Amelie Rauland, Yiwei Jia, Atlas Haddadi Avval, Afshin Bozorgpour, Sanaz Karimijafarbigloo, Joseph Paul Cohen, Ehsan Adeli-Mosabbeb, Dorit Merhof |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2024 | FedNN: Federated learning on concept drift data using weight and adaptive group normalizations
Myeongkyun Kang, Soopil Kim, Kyong Hwan Jin, Ehsan Adeli-Mosabbeb, Kilian M. Pohl, Sanghyun Park 0004 |
Pattern Recognit. | 4 |
| 2023 | Rendering Humans from Object-Occluded Monocular Videosabstract3D understanding and rendering of moving humans from monocular videos is a challenging task. Despite recent progress, the task remains difficult in real-world scenarios, where obstacles may block the camera view and cause partial occlusions in the captured videos. Existing methods cannot handle such defects due to two reasons. First, the standard rendering strategy relies on point-point mapping, which could lead to dramatic disparities between the visible and occluded areas of the body. Second, the naive direct regression approach does not consider any feasibility criteria (i.e., prior information) for rendering under occlusions. To tackle the above drawbacks, we present OccNeRF, a neural rendering method that achieves better rendering of humans in severely occluded scenes. As direct solutions to the two drawbacks, we propose surface-based rendering by integrating geometry and visibility priors. We validate our method on both simulated and real-world occlusions and demonstrate our method’s superiority. Project page: https://cs.stanford.edu/~xtiange/projects/occnerf/ Tiange Xiang, Adam Sun, Jiajun Wu 0001, Ehsan Adeli-Mosabbeb, Li Fei-Fei 0001 |
ICCV | 4 |
| 2023 | One-Shot Federated Learning on Medical Data Using Knowledge Distillation with Image Synthesis and Client Model Adaptation
Myeongkyun Kang, Philip Chikontwe, Soopil Kim, Kyong Hwan Jin, Ehsan Adeli-Mosabbeb, Kilian M. Pohl, Sanghyun Park 0004 |
MICCAI (2) | 5 |
| 2023 | An Explainable Geometric-Weighted Graph Attention Network for Identifying Functional Networks Associated with Gait Impairment
Favour Nerrise, Qingyu Zhao, Kathleen L. Poston, Kilian M. Pohl, Ehsan Adeli-Mosabbeb |
MICCAI (2) | 5 |
| 2023 | LSOR: Longitudinally-Consistent Self-Organized Representation Learning
Jiahong Ouyang, Qingyu Zhao, Ehsan Adeli-Mosabbeb, Wei Peng 0009, Greg Zaharchuk, Kilian M. Pohl |
MICCAI (1) | 3 |
| 2023 | Generating Realistic Brain MRIs via a Conditional Diffusion Probabilistic Model
Wei Peng 0009, Ehsan Adeli-Mosabbeb, Tomas M. Bosschieter, Sanghyun Park 0004, Qingyu Zhao, Kilian M. Pohl |
MICCAI (8) | 2 |
| 2022 | Rethinking Architecture Design for Tackling Data Heterogeneity in Federated LearningabstractFederated learning is an emerging research paradigm enabling collaborative training of machine learning models among different organizations while keeping data private at each institution. Despite recent progress, there remain fundamental challenges such as the lack of convergence and the potential for catastrophic forgetting across real-world heterogeneous devices. In this paper, we demonstrate that self-attention-based architectures (e.g., Transformers) are more robust to distribution shifts and hence improve federated learning over heterogeneous data. Concretely, we conduct the first rigorous empirical investigation of different neural architectures across a range of federated algorithms, real-world benchmarks, and heterogeneous data splits. Our experiments show that simply replacing convolutional networks with Transformers can greatly reduce catastrophic forgetting of previous devices, accelerate convergence, and reach a better global model, especially when dealing with heterogeneous data. We release our code and pretrained models to encourage future exploration in robust architectures as an alternative to current research efforts on the optimization front. Liangqiong Qu, Yuyin Zhou, Paul Pu Liang, Yingda Xia, Ehsan Adeli-Mosabbeb, Li Fei-Fei 0001, Daniel L. Rubin |
CVPR | 6 |
| 2022 | PrivHAR: Recognizing Human Actions from Privacy-Preserving Lens
Carlos Hinojosa, Miguel Marquez, Henry Arguello, Ehsan Adeli-Mosabbeb, Li Fei-Fei 0001, Juan Carlos Niebles |
ECCV (4) | 4 |
| 2022 | WTM: Weighted Temporal Attention Module for Group Activity RecognitionabstractGroup Activity Recognition requires spatiotemporal modeling of an exponential number of semantic and geometric relations among various individuals in a scene. Previous attempts model these relations by aggregating independently derived spatial and temporal features. This increases the modeling complexity and results in sparse information due to lack of feature correlation. In this paper, we propose Weighted Temporal Attention Mechanism (WTM), a representational mechanism that combines spatial and temporal features of a local subset of a visual sequence into a single 2D image representation, highlighting areas of a frame where actor motion is significant. Pairwise dense optical flow maps representing the temporal characteristic of individuals over a sequence are used as attention masks over raw RGB images through a multi-layer weighted aggregation. We demonstrate a strong correlation between spatial and temporal features, which helps localize actions effectively in a multi-person scenario. The simplicity of the input representation allows the model to be trained by 2D image classification architectures in a plug-and-play fashion, which outperforms its multi-stream and multi-dimensional counterparts. The proposed method achieves the lowest computational complexity in comparison to other works. We demonstrate the performance of WTM on two widely used public benchmark datasets, namely the Collective Activity Dataset (CAD) and the Volleyball Dataset. and achieve state-of-the-art accuracies of 95.1% and 94.6% respectively. We also discuss the application of this method to other datasets and general scenarios. The code is being made publicly available. Santosh Kumar Yadav, Palaash Agrawal, Kamlesh Tiwari, Ehsan Adeli-Mosabbeb, Hari Mohan Pandey, Ali Akbar Shaikh |
IJCNN | 4 |
| 2022 | GaitForeMer: Self-supervised Pre-training of Transformers via Human Motion Forecasting for Few-Shot Gait Impairment Severity Estimation
Mark Endo, Kathleen L. Poston, Edith V. Sullivan, Li Fei-Fei 0001, Kilian M. Pohl, Ehsan Adeli-Mosabbeb |
MICCAI (8) | 6 |
| 2022 | Joint Graph Convolution for Analyzing Brain Structural and Functional Connectome
Qingyue Wei, Ehsan Adeli-Mosabbeb, Kilian M. Pohl, Qingyu Zhao |
MICCAI (1) | 3 |
| 2022 | A Penalty Approach for Normalizing Feature Distributions to Build Confounder-Free Models
Anthony Vento, Qingyu Zhao, Robert Paul, Kilian M. Pohl, Ehsan Adeli-Mosabbeb |
MICCAI (3) | 5 |
| 2022 | MOMA-LRG: Language-Refined Graphs for Multi-Object Multi-Actor Activity ParsingabstractVideo-language models (VLMs), large models pre-trained on numerous but noisy video-text pairs from the internet, have revolutionized activity recognition through their remarkable generalization and open-vocabulary capabilities. While complex human activities are often hierarchical and compositional, most existing tasks for evaluating VLMs focus only on high-level video understanding, making it difficult to accurately assess and interpret the ability of VLMs to understand complex and fine-grained human activities. Inspired by the recently proposed MOMA framework, we define activity graphs as a single universal representation of human activities that encompasses video understanding at the activity, sub-activity, and atomic action level. We redefine activity parsing as the overarching task of activity graph generation, requiring understanding human activities across all three levels. To facilitate the evaluation of models on activity parsing, we introduce MOMA-LRG (Multi-Object Multi-Actor Language-Refined Graphs), a large dataset of complex human activities with activity graph annotations that can be readily transformed into natural language sentences. Lastly, we present a model-agnostic and lightweight approach to adapting and evaluating VLMs by incorporating structured knowledge from activity graphs into VLMs, addressing the individual limitations of language and graphical models. We demonstrate strong performance on few-shot activity parsing, and our framework is intended to foster future research in the joint modeling of videos, graphs, and language. Zelun Luo, Zane Durante, Linden Li, Wanze Xie, Emily Jin, Zhuoyi Huang, Lun Yu Li, Jiajun Wu 0001, Juan Carlos Niebles, Ehsan Adeli-Mosabbeb, Li Fei-Fei 0001 |
NeurIPS | 11 |
| 2022 | Self-supervised learning of neighborhood embedding for longitudinal MRI
Jiahong Ouyang, Qingyu Zhao, Ehsan Adeli-Mosabbeb, Greg Zaharchuk, Kilian M. Pohl |
Medical Image Anal. | 3 |
| 2022 | Multi-label, multi-domain learning identifies compounding effects of HIV and cognitive impairment
Jiequan Zhang, Qingyu Zhao, Ehsan Adeli-Mosabbeb, Adolf Pfefferbaum, Edith V. Sullivan, Robert Paul, Victor G. Valcour, Kilian M. Pohl |
Medical Image Anal. | 3 |
| 2022 | Semantic instance segmentation with discriminative deep supervision for medical images
Sihang Zhou 0001, Dong Nie, Ehsan Adeli-Mosabbeb, Xuhua Ren, Xinwang Liu 0002, En Zhu, Jianping Yin, Qian Wang 0001, Dinggang Shen |
Medical Image Anal. | 3 |
| 2022 | Generative adversarial U-Net for domain-free few-shot medical diagnosis
Xiaocong Chen, Lina Yao 0001, Ehsan Adeli-Mosabbeb, Yu Zhang 0009, Xianzhi Wang 0001 |
Pattern Recognit. Lett. | 4 |
| 2022 | Multiview Feature Learning With Multiatlas-Based Functional Connectivity Networks for MCI DiagnosisabstractFunctional connectivity (FC) networks built from resting-state functional magnetic resonance imaging (rs-fMRI) has shown promising results for the diagnosis of Alzheimer's disease and its prodromal stage, that is, mild cognitive impairment (MCI). FC is usually estimated as a temporal correlation of regional mean rs-fMRI signals between any pair of brain regions, and these regions are traditionally parcellated with a particular brain atlas. Most existing studies have adopted a predefined brain atlas for all subjects. However, the constructed FC networks inevitably ignore the potentially important subject-specific information, particularly, the subject-specific brain parcellation. Similar to the drawback of the "single view" (versus the "multiview" learning) in medical image-based classification, FC networks constructed based on a single atlas may not be sufficient to reveal the underlying complicated differences between normal controls and disease-affected patients due to the potential bias from that particular atlas. In this study, we propose a multiview feature learning method with multiatlas-based FC networks to improve MCI diagnosis. Specifically, a three-step transformation is implemented to generate multiple individually specified atlases from the standard automated anatomical labeling template, from which a set of atlas exemplars is selected. Multiple FC networks are constructed based on these preselected atlas exemplars, providing multiple views of the FC network-based feature representations for each subject. We then devise a multitask learning algorithm for joint feature selection from the constructed multiple FC networks. The selected features are jointly fed into a support vector machine classifier for multiatlas-based MCI diagnosis. Extensive experimental comparisons are carried out between the proposed method and other competing approaches, including the traditional single-atlas-based method. The results indicate that our method significantly improves the MCI classification, demonstrating its promise in the brain connectome-based individualized diagnosis of brain diseases. Yu Zhang 0009, Han Zhang 0002, Ehsan Adeli-Mosabbeb, Xiaobo Chen 0001, Mingxia Liu 0001, Dinggang Shen |
IEEE Trans. Cybern. | 3 |
| 2022 | Disentangling Normal Aging From Severity of Disease via Weak Supervision on Longitudinal MRIabstractThe continuous progression of neurological diseases are often categorized into conditions according to their severity. To relate the severity to changes in brain morphometry, there is a growing interest in replacing these categories with a continuous severity scale that longitudinal MRIs are mapped onto via deep learning algorithms. However, existing methods based on supervised learning require large numbers of samples and those that do not, such as self-supervised models, fail to clearly separate the disease effect from normal aging. Here, we propose to explicitly disentangle those two factors via weak-supervision. In other words, training is based on longitudinal MRIs being labelled either normal or diseased so that the training data can be augmented with samples from disease categories that are not of primary interest to the analysis. We do so by encouraging trajectories of controls to be fully encoded by the direction associated with brain aging. Furthermore, an orthogonal direction linked to disease severity captures the residual component from normal aging in the diseased cohort. Hence, the proposed method quantifies disease severity and its progression speed in individuals without knowing their condition. We apply the proposed method on data from the Alzheimer's Disease Neuroimaging Initiative (ADNI, N =632 ). We then show that the model properly disentangled normal aging from the severity of cognitive impairment by plotting the resulting disentangled factors of each subject and generating simulated MRIs for a given chronological age and condition. Moreover, our representation obtains higher balanced accuracy when used for two downstream classification tasks compared to other pre-training approaches. The code for our weak-supervised approach is available at https://github.com/ouyangjiahong/longitudinal-direction-disentangle. Jiahong Ouyang, Qingyu Zhao, Ehsan Adeli-Mosabbeb, Greg Zaharchuk, Kilian M. Pohl |
IEEE Trans. Medical Imaging | 3 |
| 2021 | 3D CNNs With Adaptive Temporal Feature ResolutionsabstractWhile state-of-the-art 3D Convolutional Neural Networks (CNN) achieve very good results on action recognition datasets, they are computationally very expensive and require many GFLOPs. While the GFLOPs of a 3D CNN can be decreased by reducing the temporal feature resolution within the network, there is no setting that is optimal for all input clips. In this work, we therefore introduce a differentiable Similarity Guided Sampling (SGS) module, which can be plugged into any existing 3D CNN architecture. SGS empowers 3D CNNs by learning the similarity of temporal features and grouping similar features together. As a result, the temporal feature resolution is not anymore static but it varies for each input video clip. By integrating SGS as an additional layer within current 3D CNNs, we can convert them into much more efficient 3D CNNs with adaptive temporal feature resolutions (ATFR). Our evaluations show that the proposed module improves the state-of-the-art by reducing the computational cost (GFLOPs) by half while preserving or even improving the accuracy. We evaluate our module by adding it to multiple state-of-the-art 3D CNNs on various datasets such as Kinetics-600, Kinetics-400, mini-Kinetics, Something-Something V2, UCF101, and HMDB51. Mohsen Fayyaz, Emad Bahrami Rad, Ali Diba, Mehdi Noroozi, Ehsan Adeli-Mosabbeb, Luc Van Gool, Juergen Gall |
CVPR | 5 |
| 2021 | Metadata NormalizationabstractBatch Normalization (BN) and its variants have delivered tremendous success in combating the covariate shift induced by the training step of deep learning methods. While these techniques normalize feature distributions by standardizing with batch statistics, they do not correct the influence on features from extraneous variables or multiple distributions. Such extra variables, referred to as metadata here, may create bias or confounding effects (e.g., race when classifying gender from face images). We introduce the Metadata Normalization (MDN) layer, a new batch-level operation which can be used end-to-end within the training framework, to correct the influence of metadata on feature distributions. MDN adopts a regression analysis technique traditionally used for preprocessing to remove (regress out) the metadata effects on model features during training. We utilize a metric based on distance correlation to quantify the distribution bias from the metadata and demonstrate that our method successfully removes metadata effects on four diverse settings: one synthetic, one 2D image, one video, and one 3D medical image dataset. Mandy Lu, Qingyu Zhao, Jiequan Zhang, Kilian M. Pohl, Li Fei-Fei 0001, Juan Carlos Niebles, Ehsan Adeli-Mosabbeb |
CVPR | 7 |
| 2021 | Scalable Differential Privacy With Sparse Network FinetuningabstractWe propose a novel method for privacy-preserving training of deep neural networks leveraging public, out-domain data. While differential privacy (DP) has emerged as a mechanism to protect sensitive data in training datasets, its application to complex visual recognition tasks remains challenging. Traditional DP methods, such as Differentially-Private Stochastic Gradient Descent (DP-SGD), perform well only on simple datasets and shallow networks, while recent transfer learning-based DP methods often make unrealistic assumptions about the availability and distribution of public data. In this work, we argue that minimizing the number of trainable parameters is the key to improving the privacy-performance tradeoff of DP on complex visual recognition tasks. Inspired by this argument, we also propose a novel transfer learning paradigm that finetunes a very sparse subnetwork with DP. We conduct extensive experiments and ablation studies on two visual recognition tasks: CIFAR-100 → CIFAR-10 (standard DP setting) and the CD-FSL challenge (few-shot, multiple levels of domain shifts) and demonstrate competitive experimental performance. Zelun Luo, Daniel J. Wu, Ehsan Adeli-Mosabbeb, Li Fei-Fei 0001 |
CVPR | 3 |
| 2021 | Home Action Genome: Cooperative Compositional Action UnderstandingabstractExisting research on action recognition treats activities as monolithic events occurring in videos. Recently, the benefits of formulating actions as a combination of atomicactions have shown promise in improving action understanding with the emergence of datasets containing such annotations, allowing us to learn representations capturing this information. However, there remains a lack of studies that extend action composition and leverage multiple view-points and multiple modalities of data for representation learning. To promote research in this direction, we introduce Home Action Genome (HOMAGE): a multi-view action dataset with multiple modalities and view-points supplemented with hierarchical activity and atomic action labels together with dense scene composition labels. Lever-aging rich multi-modal and multi-view settings, we propose Cooperative Compositional Action Understanding (CCAU), a cooperative learning framework for hierarchical action recognition that is aware of compositional action elements. CCAU shows consistent performance improvements across all modalities. Furthermore, we demonstrate the utility of co-learning compositions in few-shot action recognition by achieving 28.6% mAP with just a single sample. Nishant Rai, Haofeng Chen, Jingwei Ji, Rishi Desai, Kazuki Kozuka, Shun Ishizaka, Ehsan Adeli-Mosabbeb, Juan Carlos Niebles |
CVPR | 7 |
| 2021 | TRiPOD: Human Trajectory and Pose Dynamics Forecasting in the WildabstractJoint forecasting of human trajectory and pose dynamics is a fundamental building block of various applications ranging from robotics and autonomous driving to surveillance systems. Predicting body dynamics requires capturing subtle information embedded in the humans’ interactions with each other and with the objects present in the scene. In this paper, we propose a novel TRajectory and POse Dynamics (nicknamed TRiPOD) method based on graph attentional networks to model the human-human and human-object interactions both in the input space and the output space (decoded future output). The model is supplemented by a message passing interface over the graphs to fuse these different levels of interactions efficiently. Furthermore, to incorporate a real-world challenge, we propound to learn an indicator representing whether an estimated body joint is visible/invisible at each frame, e.g. due to occlusion or being outside the sensor field of view. Finally, we introduce a new benchmark for this joint task based on two challenging datasets (PoseTrack and 3DPW) and propose evaluation metrics to measure the effectiveness of predictions in the global space, even when there are invisible cases of joints. Our evaluation shows that TRiPOD outperforms all prior work and state-of-the-art specifically designed for each of the trajectory and pose forecasting tasks. Vida Adeli, Mahsa Ehsanpour, Ian D. Reid 0001, Juan Carlos Niebles, Silvio Savarese, Ehsan Adeli-Mosabbeb, Seyed Hamid Rezatofighi |
ICCV | 6 |
| 2021 | Self-supervised Longitudinal Neighbourhood Embedding
Jiahong Ouyang, Qingyu Zhao, Ehsan Adeli-Mosabbeb, Edith V. Sullivan, Adolf Pfefferbaum, Greg Zaharchuk, Kilian M. Pohl |
MICCAI (2) | 3 |
| 2021 | Longitudinal Correlation Analysis for Decoding Multi-modal Brain Development
Qingyu Zhao, Ehsan Adeli-Mosabbeb, Kilian M. Pohl |
MICCAI (7) | 2 |
| 2021 | MOMA: Multi-Object Multi-Actor Activity ParsingabstractComplex activities often involve multiple humans utilizing different objects to complete actions (e.g., in healthcare settings, physicians, nurses, and patients interact with each other and various medical devices). Recognizing activities poses a challenge that requires a detailed understanding of actors' roles, objects' affordances, and their associated relationships. Furthermore, these purposeful activities are composed of multiple achievable steps, including sub-activities and atomic actions, which jointly define a hierarchy of action parts. This paper introduces Activity Parsing as the overarching task of temporal segmentation and classification of activities, sub-activities, atomic actions, along with an instance-level understanding of actors, objects, and their relationships in videos. Involving multiple entities (actors and objects), we argue that traditional pair-wise relationships, often used in scene or action graphs, do not appropriately represent the dynamics between them. Hence, we introduce Action Hypergraph, a spatial-temporal graph containing hyperedges (i.e., edges with higher-order relationships), as a new representation. In addition, we introduce Multi-Object Multi-Actor (MOMA), the first benchmark and dataset dedicated to activity parsing. Lastly, to parse a video, we propose the HyperGraph Activity Parsing (HGAP) network, which outperforms several baselines, including those based on regular graphs and raw video data. Zelun Luo, Wanze Xie, Siddharth Kapoor, Yiyun Liang, Michael Cooper, Juan Carlos Niebles, Ehsan Adeli-Mosabbeb, Li Fei-Fei 0001 |
NeurIPS | 7 |
| 2021 | Representation Learning with Statistical Independence to Mitigate BiasabstractPresence of bias (in datasets or tasks) is inarguably one of the most critical challenges in machine learning applications that has alluded to pivotal debates in recent years. Such challenges range from spurious associations between variables in medical studies to the bias of race in gender or face recognition systems. Controlling for all types of biases in the dataset curation stage is cumbersome and sometimes impossible. The alternative is to use the available data and build models incorporating fair representation learning. In this paper, we propose such a model based on adversarial training with two competing objectives to learn features that have (1) maximum discriminative power with respect to the task and (2) minimal statistical mean dependence with the protected (bias) variable(s). Our approach does so by incorporating a new adversarial loss function that encourages a vanished correlation between the bias and the learned features. We apply our method to synthetic data, medical images (containing task bias), and a dataset for gender classification (containing dataset bias). Our results show that the learned features by our method not only result in superior prediction performance but also are unbiased. Ehsan Adeli-Mosabbeb, Qingyu Zhao, Adolf Pfefferbaum, Edith V. Sullivan, Li Fei-Fei 0001, Juan Carlos Niebles, Kilian M. Pohl |
WACV | 1 |
| 2021 | MetricUNet: Synergistic image- and voxel-level learning for precise prostate segmentation via online sampling
Kelei He, Chunfeng Lian, Ehsan Adeli-Mosabbeb, Jing Huo, Yang Gao 0001, Bing Zhang 0012, Dinggang Shen |
Medical Image Anal. | 3 |
| 2021 | Quantifying Parkinson's disease motor severity under uncertainty using MDS-UPDRS videos
Mandy Lu, Qingyu Zhao, Kathleen L. Poston, Edith V. Sullivan, Adolf Pfefferbaum, Marian Shahid, Maya Katz, Leila Montaser Kouhsari, Kevin A. Schulman, Arnold Milstein, Juan Carlos Niebles, Victor W. Henderson, Li Fei-Fei 0001, Kilian M. Pohl, Ehsan Adeli-Mosabbeb |
Medical Image Anal. | 15 |
| 2021 | Longitudinal self-supervised learning
Qingyu Zhao, Zixuan Liu 0001, Ehsan Adeli-Mosabbeb, Kilian M. Pohl |
Medical Image Anal. | 3 |
| 2021 | Skeleton-based structured early activity prediction
Mohammad M. Arzani, Mahmood Fathy, A. Akbariazirani, Ehsan Adeli-Mosabbeb |
Multim. Tools Appl. | 4 |
| 2021 | Switching Structured Prediction for Simple and Complex Human Activity RecognitionabstractAutomatic human activity recognition is an integral part of any interactive application involving humans (e.g., human-robot interaction systems). One of the main challenges for activity recognition is the diversity in the way individuals often perform activities. Furthermore, changes in any of the environment factors (i.e., illumination, complex background, human body shapes, viewpoint, etc.) intensify this challenge. In addition, there are different types of activities that robots need to interpret for seamless interaction with humans. Some activities are short, quick, and simple (e.g., sitting), while others may be detailed/complex, and spread throughout a long span of time (e.g., washing mouth). In this article, we recognize the activities within the context of graphical models in a sequence-labeling framework based on skeleton data. We propose a new structured prediction strategy based on probabilistic graphical models (PGMs) to recognize both types of activities (i.e., complex and simple). These activity types are often spanned in very diverse subspaces in the space of all possible activities, which would require different model parameterizations. In order to deal with these parameterization and structural breaks across models, a category-switching scheme is proposed to switch over the models based on the activity types. For parameter optimization, we utilize a distributed structured prediction technique to implement our model in a distributed setting. The method is tested on three widely used datasets (CAD-60, UT-Kinect, and Florence 3-D) that cover both activity types. The results illustrate that our proposed method is able to recognize simple and complex activities while the previous work concentrated on only one of these two main types. Mohammad M. Arzani, Mahmood Fathy, A. Akbariazirani, Ehsan Adeli-Mosabbeb |
IEEE Trans. Cybern. | 4 |
| 2021 | Cascaded MultiTask 3-D Fully Convolutional Networks for Pancreas SegmentationabstractAutomatic pancreas segmentation is crucial to the diagnostic assessment of diabetes or pancreatic cancer. However, the relatively small size of the pancreas in the upper body, as well as large variations of its location and shape in retroperitoneum, make the segmentation task challenging. To alleviate these challenges, in this article, we propose a cascaded multitask 3-D fully convolution network (FCN) to automatically segment the pancreas. Our cascaded network is composed of two parts. The first part focuses on fast locating the region of the pancreas, and the second part uses a multitask FCN with dense connections to refine the segmentation map for fine voxel-wise segmentation. In particular, our multitask FCN with dense connections is implemented to simultaneously complete tasks of the voxel-wise segmentation and skeleton extraction from the pancreas. These two tasks are complementary, that is, the extracted skeleton provides rich information about the shape and size of the pancreas in retroperitoneum, which can boost the segmentation of pancreas. The multitask FCN is also designed to share the low- and mid-level features across the tasks. A feature consistency module is further introduced to enhance the connection and fusion of different levels of feature maps. Evaluations on two pancreas datasets demonstrate the robustness of our proposed method in correctly segmenting the pancreas in various settings. Our experimental results outperform both baseline and state-of-the-art methods. Moreover, the ablation study shows that our proposed parts/modules are critical for effective multitask learning. Jie Xue 0001, Kelei He, Dong Nie, Ehsan Adeli-Mosabbeb, Zhenshan Shi, Seong-Whan Lee, Yuanjie Zheng, Xiyu Liu 0001, Dengwang Li, Dinggang Shen |
IEEE Trans. Cybern. | 4 |
| 2021 | Longitudinal Pooling & Consistency Regularization to Model Disease Progression From MRIsabstractMany neurological diseases are characterized by gradual deterioration of brain structure andfunction. Large longitudinal MRI datasets have revealed such deterioration, in part, by applying machine and deep learning to predict diagnosis. A popular approach is to apply Convolutional Neural Networks (CNN) to extract informative features from each visit of the longitudinal MRI and then use those features to classify each visit via Recurrent Neural Networks (RNNs). Such modeling neglects the progressive nature of the disease, which may result in clinically implausible classifications across visits. To avoid this issue, we propose to combine features across visits by coupling feature extraction with a novel longitudinal pooling layer and enforce consistency of the classification across visits in line with disease progression. We evaluate the proposed method on the longitudinal structural MRIs from three neuroimaging datasets: Alzheimer's Disease Neuroimaging Initiative (ADNI, N=404), a dataset composed of 274 normal controls and 329 patients with Alcohol Use Disorder (AUD), and 255 youths from the National Consortium on Alcohol and NeuroDevelopment in Adolescence (NCANDA). In allthree experiments our method is superior to other widely used approaches for longitudinal classification thus making a unique contribution towards more accurate tracking of the impact of conditions on the brain. The code is available at https://github.com/ouyangjiahong/longitudinal-pooling. Jiahong Ouyang, Qingyu Zhao, Edith V. Sullivan, Adolf Pfefferbaum, Susan F. Tapert, Ehsan Adeli-Mosabbeb, Kilian M. Pohl |
IEEE J. Biomed. Health Informatics | 6 |
| 2021 | Deep End-to-End One-Class ClassifierabstractOne-class classification (OCC) poses as an essential component in many machine learning and computer vision applications, including novelty, anomaly, and outlier detection systems. With a known definition for a target or normal set of data, one-class classifiers can determine if any given new sample spans within the distribution of the target class. Solving for this task in a general setting is particularly very challenging, due to the high diversity of samples from the target class and the absence of any supervising signal over the novelty (nontarget) concept, which makes designing end-to-end models unattainable. In this article, we propose an adversarial training approach to detect out-of-distribution samples in an end-to-end trainable deep model. To this end, we jointly train two deep neural networks, R and D . The latter plays as the discriminator while the former, during training, helps D characterize a probability distribution for the target class by creating adversarial examples and, during testing, collaborates with it to detect novelties. Using our OCC, we first test outlier detection on two image data sets, Modified National Institute of Standards and Technology (MNIST) and Caltech-256. Then, several experiments for video anomaly detection are performed on University of Minnesota (UMN) and University of California, San Diego (UCSD) data sets. Our proposed method can successfully learn the target class underlying distribution and outperforms other approaches. Mohammad Sabokrou, Mahmood Fathy, Guoying Zhao 0001, Ehsan Adeli-Mosabbeb |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2020 | Adversarial Cross-Domain Action Recognition with Co-AttentionabstractAction recognition has been a widely studied topic with a heavy focus on supervised learning involving sufficient labeled videos. However, the problem of cross-domain action recognition, where training and testing videos are drawn from different underlying distributions, remains largely under-explored. Previous methods directly employ techniques for cross-domain image recognition, which tend to suffer from the severe temporal misalignment problem. This paper proposes a Temporal Co-attention Network (TCoN), which matches the distributions of temporally aligned action features between source and target domains using a novel cross-domain co-attention mechanism. Experimental results on three cross-domain action recognition datasets demonstrate that TCoN improves both previous single-domain and cross-domain methods significantly under the cross-domain setting. Boxiao Pan, Zhangjie Cao, Ehsan Adeli-Mosabbeb, Juan Carlos Niebles |
AAAI | 3 |
| 2020 | Spatio-Temporal Graph for Video Captioning With Knowledge DistillationabstractVideo captioning is a challenging task that requires a deep understanding of visual scenes. State-of-the-art methods generate captions using either scene-level or object-level information but without explicitly modeling object interactions. Thus, they often fail to make visually grounded predictions, and are sensitive to spurious correlations. In this paper, we propose a novel spatio-temporal graph model for video captioning that exploits object interactions in space and time. Our model builds interpretable links and is able to provide explicit visual grounding. To avoid unstable performance caused by the variable number of objects, we further propose an object-aware knowledge distillation mechanism, in which local object information is used to regularize global scene features. We demonstrate the efficacy of our approach through extensive experiments on two benchmarks, showing our approach yields competitive performance with interpretable predictions. Boxiao Pan, Haoye Cai, De-An Huang, Kuan-Hui Lee, Adrien Gaidon, Ehsan Adeli-Mosabbeb, Juan Carlos Niebles |
CVPR | 6 |
| 2020 | Procedure Planning in Instructional Videos
Chien-Yi Chang, De-An Huang, Danfei Xu, Ehsan Adeli-Mosabbeb, Li Fei-Fei 0001, Juan Carlos Niebles |
ECCV (11) | 4 |
| 2020 | It Is Not the Journey But the Destination: Endpoint Conditioned Trajectory Prediction
Karttikeya Mangalam, Harshayu Girase, Shreyas Agarwal, Kuan-Hui Lee, Ehsan Adeli-Mosabbeb, Jitendra Malik, Adrien Gaidon |
ECCV (2) | 5 |
| 2020 | Spatio-Temporal Graph Convolution for Resting-State fMRI Analysis
Soham Gadgil, Qingyu Zhao, Adolf Pfefferbaum, Edith V. Sullivan, Ehsan Adeli-Mosabbeb, Kilian M. Pohl |
MICCAI (7) | 5 |
| 2020 | Vision-Based Estimation of MDS-UPDRS Gait Scores for Assessing Parkinson's Disease Motor Severity
Mandy Lu, Kathleen L. Poston, Adolf Pfefferbaum, Edith V. Sullivan, Li Fei-Fei 0001, Kilian M. Pohl, Juan Carlos Niebles, Ehsan Adeli-Mosabbeb |
MICCAI (3) | 8 |
| 2020 | Disentangling Human Dynamics for Pedestrian Locomotion Forecasting with Noisy SupervisionabstractWe tackle the problem of Human Locomotion Forecasting, a task for jointly predicting the spatial positions of several keypoints on human body in the near future under an egocentric setting. In contrast to the previous work that aims to solve either the task of pose prediction or trajectory forecasting in isolation, we propose a framework to unify these two problems and address the practically useful task of pedestrian locomotion prediction in the wild. Among the major challenges in solving this task is the scarcity of annotated egocentric video datasets with dense annotations for pose, depth, or egomotion. To surmount this difficulty, we use state-of-the-art models to generate (noisy) annotations and propose robust forecasting models that can learn from this noisy supervision. We present a method to disentangle the overall pedestrian motion into easier to learn subparts by uti-lizing a pose completion and a decomposition module. The completion module fills in the missing key-point annotations and the decomposition module breaks the cleaned locomotion down to global (trajectory) and local (pose keypoint movements). Further, with Quasi RNN as our backbone, we propose a novel hierarchical trajectory forecasting network that utilizes low-level vision domain specific signals like egomotion and depth to predict the global trajectory. Our method leads to state-of-the-art results for the prediction of human locomotion in the egocentric view. Karttikeya Mangalam, Ehsan Adeli-Mosabbeb, Kuan-Hui Lee, Adrien Gaidon, Juan Carlos Niebles |
WACV | 2 |
| 2020 | Depth map artefacts reduction: a reviewabstractDepth maps are crucial for many visual applications, where they represent the positioning information of the objects in a three‐dimensional scene. Usually, depth maps can be acquired via various devices, including Time of Flight, Kinect or light field camera, in practical applications. However, a brutal truth is that both intrinsic and extrinsic artefacts can be found in these depth maps which limits the prosperity of three‐dimensional visual applications. In this study, the authors survey the depth map artefacts reduction methods proposed in the literature, from mono‐ to multi‐view, via spatial to temporal dimension, in local to global manner, with signal processing to learning‐based methods. They also compare the state‐of‐the‐arts via different metrics to show their potentials in future visual applications. Mostafa Mahmoud Ibrahim, Qiong Liu 0001, Ehsan Adeli-Mosabbeb, You Yang 0002 |
IET Image Process. | 5 |
| 2020 | Logistic Regression Confined by Cardinality-Constrained Sample and Feature SelectionabstractMany vision-based applications rely on logistic regression for embedding classification within a probabilistic context, such as recognition in images and videos or identifying disease-specific image phenotypes from neuroimages. Logistic regression, however, often performs poorly when trained on data that is noisy, has irrelevant features, or when the samples are distributed across the classes in an imbalanced setting; a common occurrence in visual recognition tasks. To deal with those issues, researchers generally rely on adhoc regularization techniques or model a subset of these issues. We instead propose a mathematically sound logistic regression model that selects a subset of (relevant) features and (informative and balanced) set of samples during the training process. The model does so by applying cardinality constraints (via ℓ0-`norm' sparsity) on the features and samples. ℓ0defines sparsity in mathematical settings but in practice has mostly been approximated (e.g., via ℓ1or its variations) for computational simplicity. We prove that a local minimum to the non-convex optimization problems induced by cardinality constraints can be computed by combining block coordinate descent with penalty decomposition. On synthetic, image recognition, and neuroimaging datasets, we show that the accuracy of the method is higher than alternative methods and classifiers commonly used in the literature. Ehsan Adeli-Mosabbeb, Dongjin Kwon, Yong Zhang 0004, Kilian M. Pohl |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | Population-guided large margin classifier for high-dimension low-sample-size problems
Qingbo Yin, Ehsan Adeli-Mosabbeb, Liran Shen, Dinggang Shen |
Pattern Recognit. | 2 |
| 2020 | High-Resolution Encoder-Decoder Networks for Low-Contrast Medical Image SegmentationabstractAutomatic image segmentation is an essential step for many medical image analysis applications, include computer-aided radiation therapy, disease diagnosis, and treatment effect evaluation. One of the major challenges for this task is the blurry nature of medical images (e.g., CT, MR and, microscopic images), which can often result in low-contrast and vanishing boundaries. With the recent advances in convolutional neural networks, vast improvements have been made for image segmentation, mainly based on the skip-connection-linked encoder-decoder deep architectures. However, in many applications (with adjacent targets in blurry images), these models often fail to accurately locate complex boundaries and properly segment tiny isolated parts. In this paper, we aim to provide a method for blurry medical image segmentation and argue that skip connections are not enough to help accurately locate indistinct boundaries. Accordingly, we propose a novel high-resolution multi-scale encoder-decoder network (HMEDN), in which multi-scale dense connections are introduced for the encoder-decoder structure to finely exploit comprehensive semantic information. Besides skip connections, extra deeply-supervised high-resolution pathways (comprised of densely connected dilated convolutions) are integrated to collect high-resolution semantic information for accurate boundary localization. These pathways are paired with a difficulty-guided cross-entropy loss function and a contour regression task to enhance the quality of boundary detection. Extensive experiments on a pelvic CT image dataset, a multi-modal brain tumor dataset, and a cell segmentation dataset show the effectiveness of our method for 2D/3D semantic segmentation and 2D instance segmentation, respectively. Our experimental results also show that besides increasing the network complexity, raising the resolution of semantic feature maps can largely affect the overall model performance. For different tasks, finding a balance between these two factors can further improve the performance of the corresponding network. Sihang Zhou 0001, Dong Nie, Ehsan Adeli-Mosabbeb, Jianping Yin, Jun Lian, Dinggang Shen |
IEEE Trans. Image Process. | 3 |
| 2020 | Editorial: Predictive Intelligence in Biomedical and Health InformaticsabstractThe papers in this special section examine the use of predictive intelligence for bioinformatics. Big data is fueling diverse research directions in both medical image analysis and computer vision research fields. These can be divided into two main categories: (1) analytical methods, and (2) predictive methods. While analytical methods aim to efficiently analyze, represent, and interpret data, predictive methods leverage the data currently available to predict observations at present (e.g., by completingmissing observations), at previous time-points (e.g., by solving reverse problems), or at later time-points (i.e., forecasting the future). For instance, a method which only focuses on classifying patients with mild cognitive impairment (MCI) and patients with Alzheimer’s disease (AD) is an analytical method, while a method that predicts if a subject diagnosed with MCI will remain stable or convert to AD over time is a predictive method. Similar cases can be established for various neurodegenerative or neuropsychiatric disorders, degenerative arthritis, or cancer studies, in which the disease/disorder develops over time. Ehsan Adeli-Mosabbeb, S. H. Rekik, Sanghyun Park 0004, Dinggang Shen |
IEEE J. Biomed. Health Informatics | 1 |
| 2020 | Image-to-Images Translation for Multi-Task Organ Segmentation and Bone Suppression in Chest X-Ray RadiographyabstractChest X-ray radiography is one of the earliest medical imaging technologies and remains one of the most widely-used for diagnosis, screening, and treatment follow up of diseases related to lungs and heart. The literature in this field of research reports many interesting studies dealing with the challenging tasks of bone suppression and organ segmentation but performed separately, limiting any learning that comes with the consolidation of parameters that could optimize both processes. This study, and for the first time, introduces a multitask deep learning model that generates simultaneously the bone-suppressed image and the organ-segmented image, enhancing the accuracy of tasks, minimizing the number of parameters needed by the model and optimizing the processing time, all by exploiting the interplay between the network parameters to benefit the performance of both tasks. The architectural design of this model, which relies on a conditional generative adversarial network, reveals the process on how the wellestablished pix2pix network (image-to-image network) is modified to fit the need for multitasking and extending it to the new image-to-images architecture. The developed source code of this multitask model is shared publicly on Github as the first attempt for providing the two-task pix2pix extension, a supervised/paired/aligned/registered image-to-images translation which would be useful in many multitask applications. Dilated convolutions are also used to improve the results through a more effective receptive field assessment. The comparison with state-of-the-art al-gorithms along with ablation study and a demonstration video1 are provided to evaluate the efficacy and gauge the merits of the proposed approach. Mohammad Eslami, Solale Tabarestani, Shadi Albarqouni, Ehsan Adeli-Mosabbeb, Nassir Navab, Malek Adjouadi |
IEEE Trans. Medical Imaging | 4 |
| 2019 | Difficulty-Aware Attention Network with Confidence Learning for Medical Image SegmentationabstractMedical image segmentation is a key step for various applications, such as image-guided radiation therapy and diagnosis. Recently, deep neural networks provided promising solutions for automatic image segmentation; however, they often perform good on regular samples (i.e., easy-to-segment samples), since the datasets are dominated by easy and regular samples. For medical images, due to huge inter-subject variations or disease-specific effects on subjects, there exist several difficult-to-segment cases that are often overlooked by the previous works. To address this challenge, we propose a difficulty-aware deep segmentation network with confidence learning for end-to-end segmentation. The proposed framework has two main contributions: 1) Besides the segmentation network, we also propose a fully convolutional adversarial network for confidence learning to provide voxel-wise and region-wise confidence information for the segmentation network. We relax the adversarial learning to confidence learning by decreasing the priority of adversarial learning, so that we can avoid the training imbalance between generator and discriminator. 2) We propose a difficulty-aware attention mechanism to properly handle hard samples or hard regions considering structural information, which may go beyond the shortcomings of focal loss. We further propose a fusion module to selectively fuse the concatenated feature maps in encoder-decoder architectures. Experimental results on clinical and challenge datasets show that our proposed network can achieve state-of-the-art segmentation accuracy. Further analysis also indicates that each individual component of our proposed network contributes to the overall performance improvement. Dong Nie, Li Wang 0026, Lei Xiang 0001, Sihang Zhou 0001, Ehsan Adeli-Mosabbeb, Dinggang Shen |
AAAI | 5 |
| 2019 | Unsupervised Feature Ranking and Selection Based on AutoencodersabstractFeature selection is one of the most important and widely-used dimension reduction techniques due to its efficiency and intractability of the results. In this paper, we propose a simple but efficient unsupervised feature ranking and selection method by exploiting the geometry of the original feature space using AutoEncoders. Average reconstruction error of training samples by ignoring features, one at time, and the contribution of feature in the latent space (bottleneck of the auto-encoder) are proposed as two useful measures for ranking the features. The proposed method is evaluated for three different tasks: (1) feature selection, (2) discovering image interest points, and (3) extracting important blocks of an images Result on standard benchmarks confirm that the performance of our method is better than state-of-the-art methods. Sasan Sharifipour, Hossein Fayyazi, Mohammad Sabokrou, Ehsan Adeli-Mosabbeb |
ICASSP | 4 |
| 2019 | Self-Supervised Representation Learning via Neighborhood-Relational EncodingabstractIn this paper, we propose a novel self-supervised representation learning by taking advantage of a neighborhood-relational encoding (NRE) among the training data. Conventional unsupervised learning methods only focused on training deep networks to understand the primitive characteristics of the visual data, mainly to be able to reconstruct the data from a latent space. They often neglected the relation among the samples, which can serve as an important metric for self-supervision. Different from the previous work, NRE aims at preserving the local neighborhood structure on the data manifold. Therefore, it is less sensitive to outliers. We integrate our NRE component with an encoder-decoder structure for learning to represent samples considering their local neighborhood information. Such discriminative and unsupervised representation learning scheme is adaptable to different computer vision tasks due to its independence from intense annotation requirements. We evaluate our proposed method for different tasks, including classification, detection, and segmentation based on the learned latent representations. In addition, we adopt the auto-encoding capability of our proposed method for applications like defense against adversarial example attacks and video anomaly detection. Results confirm the performance of our method is better or at least comparable with the state-of-the-art for each specific application, but with a generic and self-supervised approach. Mohammad Sabokrou, Mohammad Khalooei, Ehsan Adeli-Mosabbeb |
ICCV | 3 |
| 2019 | Imitation Learning for Human Pose PredictionabstractModeling and prediction of human motion dynamics has long been a challenging problem in computer vision, and most existing methods rely on the end-to-end supervised training of various architectures of recurrent neural networks. Inspired by the recent success of deep reinforcement learning methods, in this paper we propose a new reinforcement learning formulation for the problem of human pose prediction, and develop an imitation learning algorithm for predicting future poses under this formulation through a combination of behavioral cloning and generative adversarial imitation learning. Our experiments show that our proposed method outperforms all existing state-of-the-art baseline models by large margins on the task of human pose prediction in both short-term predictions and long-term predictions, while also enjoying huge advantage in training speed. Borui Wang, Ehsan Adeli-Mosabbeb, Hsu-Kuang Chiu, De-An Huang, Juan Carlos Niebles |
ICCV | 2 |
| 2019 | Variational AutoEncoder for Regression: Application to Brain Aging Analysis
Qingyu Zhao, Ehsan Adeli-Mosabbeb, Nicolas Honnorat, Tuo Leng, Kilian M. Pohl |
MICCAI (2) | 2 |
| 2019 | Action-Agnostic Human Pose ForecastingabstractForecasting human dynamics is a very interesting but challenging task with several prospective applications in robotics, health-care, among others. Researchers have recently developed methods for human pose forecasting; but unfortunately, they often introduce a number of simplification assumptions. For instance, previous work either focuses only on short-term or long-term predictions, while sacrificing one or the other. Furthermore, they use the activity labels as part of the training process and require them to be available at testing time. These simplifications limit the usage of such pose forecasting models for real-world applications. To overcome these limitations, we propose a new action-agnostic method for short-and long-term human pose forecasting. Our triangular-prism recurrent neural network (TP-RNN) models the hierarchical and multi-scale characteristics of human dynamics. Our model captures the latent hierarchical structure in human pose sequences by encoding temporal dependencies with different time-scales. We run an extensive set of experiments on Human 3.6M and Penn Action datasets and show that our method outperforms baseline and state-of-the-art methods quantitatively and qualitatively. Code is available at https://github.com/eddyhkchiu/pose_forecast_wacv/. Hsu-Kuang Chiu, Ehsan Adeli-Mosabbeb, Borui Wang, De-An Huang, Juan Carlos Niebles |
WACV | 2 |
| 2019 | Semi-Supervised Discriminative Classification Robust to Sample-Outliers and Feature-NoisesabstractDiscriminative methods commonly produce models with relatively good generalization abilities. However, this advantage is challenged in real-world applications (e.g., medical image analysis problems), in which there often exist outlier data points (sample-outliers) and noises in the predictor values (feature-noises). Methods robust to both types of these deviations are somewhat overlooked in the literature. We further argue that denoising can be more effective, if we learn the model using all the available labeled and unlabeled samples, as the intrinsic geometry of the sample manifold can be better constructed using more data points. In this paper, we propose a semi-supervised robust discriminative classification method based on the least-squares formulation of linear discriminant analysis to detect sample-outliers and feature-noises simultaneously, using both labeled training and unlabeled testing data. We conduct several experiments on a synthetic, some benchmark semi-supervised learning, and two brain neurodegenerative disease diagnosis datasets (for Parkinson's and Alzheimer's diseases). Specifically for the application of neurodegenerative diseases diagnosis, incorporating robust machine learning methods can be of great benefit, due to the noisy nature of neuroimaging data. Our results show that our method outperforms the baseline and several state-of-the-art methods, in terms of both accuracy and the area under the ROC curve. Ehsan Adeli-Mosabbeb, Kim-Han Thung, Guorong Wu 0001, Feng Shi 0001, Dinggang Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2019 | 3-D Fully Convolutional Networks for Multimodal Isointense Infant Brain Image SegmentationabstractAccurate segmentation of infant brain images into different regions of interest is one of the most important fundamental steps in studying early brain development. In the isointense phase (approximately 6-8 months of age), white matter and gray matter exhibit similar levels of intensities in magnetic resonance (MR) images, due to the ongoing myelination and maturation. This results in extremely low tissue contrast and thus makes tissue segmentation very challenging. Existing methods for tissue segmentation in this isointense phase usually employ patch-based sparse labeling on single modality. To address the challenge, we propose a novel 3-D multimodal fully convolutional network (FCN) architecture for segmentation of isointense phase brain MR images. Specifically, we extend the conventional FCN architectures from 2-D to 3-D, and, rather than directly using FCN, we intuitively integrate coarse (naturally high-resolution) and dense (highly semantic) feature maps to better model tiny tissue regions, in addition, we further propose a transformation module to better connect the aggregating layers; we also propose a fusion module to better serve the fusion of feature maps. We compare the performance of our approach with several baseline and state-of-the-art methods on two sets of isointense phase brain images. The comparison results show that our proposed 3-D multimodal FCN model outperforms all previous methods by a large margin in terms of segmentation accuracy. In addition, the proposed framework also achieves faster segmentation results compared to all other methods. Our experiments further demonstrate that: 1) carefully integrating coarse and dense feature maps can considerably improve the segmentation performance; 2) batch normalization can speed up the convergence of the networks, especially when hierarchical feature aggregations occur; and 3) integrating multimodal information can further boost the segmentation performance. Dong Nie, Li Wang 0026, Ehsan Adeli-Mosabbeb, Cuijin Lao, Weili Lin, Dinggang Shen |
IEEE Trans. Cybern. | 3 |
| 2019 | Infant Brain Development Prediction With Latent Partial Multi-View Representation LearningabstractThe early postnatal period witnesses rapid and dynamic brain development. However, the relationship between brain anatomical structure and cognitive ability is still unknown. Currently, there is no explicit model to characterize this relationship in the literature. In this paper, we explore this relationship by investigating the mapping between morphological features of the cerebral cortex and cognitive scores. To this end, we introduce a multi-view multi-task learning approach to intuitively explore complementary information from different time-points and handle the missing data issue in longitudinal studies simultaneously. Accordingly, we establish a novel model, latent partial multi-view representation learning. Our approach regards data from different time-points as different views and constructs a latent representation to capture the complementary information from incomplete time-points. The latent representation explores the complementarity across different time-points and improves the accuracy of prediction. The minimization problem is solved by the alternating direction method of multipliers. Experimental results on both synthetic and real data validate the effectiveness of our proposed algorithm. Changqing Zhang 0002, Ehsan Adeli-Mosabbeb, Zhengwang Wu, Gang Li 0001, Weili Lin, Dinggang Shen |
IEEE Trans. Medical Imaging | 2 |
| 2018 | Multi-Layer Multi-View Classification for Alzheimer's Disease DiagnosisabstractIn this paper, we propose a novel multi-view learning method for Alzheimer's Disease (AD) diagnosis, using neuroimaging and genetics data. Generally, there are several major challenges associated with traditional classification methods on multi-source imaging and genetics data. First, the correlation between the extracted imaging features and class labels is generally complex, which often makes the traditional linear models ineffective. Second, medical data may be collected from different sources (i.e., multiple modalities of neuroimaging data, clinical scores or genetics measurements), therefore, how to effectively exploit the complementarity among multiple views is of great importance. In this paper, we propose a Multi-Layer Multi-View Classification (ML-MVC) approach, which regards the multi-view input as the first layer, and constructs a latent representation to explore the complex correlation between the features and class labels. This captures the high-order complementarity among different views, as we exploit the underlying information with a low-rank tensor regularization. Intrinsically, our formulation elegantly explores the nonlinear correlation together with complementarity among different views, and thus improves the accuracy of classification. Finally, the minimization problem is solved by the Alternating Direction Method of Multipliers (ADMM). Experimental results on Alzheimer's Disease Neuroimaging Initiative (ADNI) data sets validate the effectiveness of our proposed method. Changqing Zhang 0002, Ehsan Adeli-Mosabbeb, Tao Zhou 0002, Xiaobo Chen 0001, Dinggang Shen |
AAAI | 2 |
| 2018 | AVID: Adversarial Visual Irregularity Detection
Mohammad Sabokrou, Masoud PourReza, Mohsen Fayyaz, Rahim Entezari, Mahmood Fathy, Juergen Gall, Ehsan Adeli-Mosabbeb |
ACCV (6) | 7 |
| 2018 | Adversarially Learned One-Class Classifier for Novelty DetectionabstractNovelty detection is the process of identifying the observation(s) that differ in some respect from the training observations (the target class). In reality, the novelty class is often absent during training, poorly sampled or not well defined. Therefore, one-class classifiers can efficiently model such problems. However, due to the unavailability of data from the novelty class, training an end-to-end deep network is a cumbersome task. In this paper, inspired by the success of generative adversarial networks for training deep models in unsupervised and semi-supervised settings, we propose an end-to-end architecture for one-class classification. Our architecture is composed of two deep networks, each of which trained by competing with each other while collaborating to understand the underlying concept in the target class, and then classify the testing samples. One network works as the novelty detector, while the other supports it by enhancing the inlier samples and distorting the outliers. The intuition is that the separability of the enhanced inliers and distorted outliers is much better than deciding on the original samples. The proposed framework applies to different related applications of anomaly and outlier detection in images and videos. The results on MNIST and Caltech-256 image datasets, along with the challenging UCSD Ped2 dataset for video anomaly detection illustrate that our proposed method learns the target class effectively and is superior to the baseline and state-of-the-art methods. Mohammad Sabokrou, Mohammad Khalooei, Mahmood Fathy, Ehsan Adeli-Mosabbeb |
CVPR | 4 |
| 2018 | Multi-label Transduction for Identifying Disease Comorbidity Patterns
Ehsan Adeli-Mosabbeb, Dongjin Kwon, Kilian M. Pohl |
MICCAI (3) | 1 |
| 2018 | Fine-Grained Segmentation Using Hierarchical Dilated Neural Networks
Sihang Zhou 0001, Dong Nie, Ehsan Adeli-Mosabbeb, Yaozong Gao, Li Wang 0026, Jianping Yin, Dinggang Shen |
MICCAI (4) | 3 |
| 2018 | Landmark-based deep multi-instance learning for brain disease diagnosis
Mingxia Liu 0001, Jun Zhang 0018, Ehsan Adeli-Mosabbeb, Dinggang Shen |
Medical Image Anal. | 3 |
| 2018 | Conversion and time-to-conversion predictions of mild cognitive impairment using low-rank affinity pursuit denoising and matrix completion
Kim-Han Thung, Pew-Thian Yap, Ehsan Adeli-Mosabbeb, Seong-Whan Lee, Dinggang Shen |
Medical Image Anal. | 3 |
| 2017 | Structured prediction with short/long-range dependencies for human activity recognition from depth skeleton dataabstractOne of the main abilities that the robots need to maintain is to efficiently communicate with people in a humanly manner. Thus, human activity recognition (HAR) would be an integral part of such a human-robot interaction system. One of the major challenges in HAR is that the individuals perform their activities in different manners. Furthermore, there is a very wide range of different types of activities that the robots would require to understand. Some activities are simple, quick and short (e.g., sit down), while many others are complex, have many details and span through a long range of time (e.g., wearing contact lens). In this paper, we model the recognition of activities into a sequence-labeling problem and propose a new probabilistic graphical model (PGM) that can recognize both short/long-range activities, by introducing a hierarchical classification model and including extra links and loopy conditions in our PGM. To optimize the PGM and obtain its parameters during training, we use a structured prediction technique, a general framework that involves latent structured support vector machines (LSSVM) and hidden-state conditional random fields (HCRF). We evaluate our method on two widely used datasets (CAD-60 & UT-Kinect) that contain both activity types. Our obtained results are promising and show that our method can recognize both types of activities effectively, while most of the previous works only focused on one of these two major types. We further explore distributed processing techniques, since our method can easily be distributed over processing nodes. We also propose an efficient divide-and-merge technique to further speedup the training step. Mohammad M. Arzani, Mahmood Fathy, Hamid K. Aghajan, A. Akbariazirani, Kaamran Raahemifar, Ehsan Adeli-Mosabbeb |
IROS | 6 |
| 2017 | Joint Sparse and Low-Rank Regularized Multi-Task Multi-Linear Regression for Prediction of Infant Brain Development with Incomplete Data
Ehsan Adeli-Mosabbeb, Yu Meng 0003, Gang Li 0001, Weili Lin, Dinggang Shen |
MICCAI (1) | 1 |
| 2017 | Deep Multi-task Multi-channel Learning for Joint Classification and Regression of Brain Status
Mingxia Liu 0001, Jun Zhang 0018, Ehsan Adeli-Mosabbeb, Dinggang Shen |
MICCAI (3) | 3 |
| 2017 | Maximum Mean Discrepancy Based Multiple Kernel Learning for Incomplete Multimodality Neuroimaging Data
Xiaofeng Zhu 0001, Kim-Han Thung, Ehsan Adeli-Mosabbeb, Yu Zhang 0009, Dinggang Shen |
MICCAI (3) | 3 |
| 2017 | Multi-modal classification of neurodegenerative disease by progressive graph-based transductive learning
Zhengxia Wang, Xiaofeng Zhu 0001, Ehsan Adeli-Mosabbeb, Yingying Zhu 0004, Feiping Nie 0001, Brent C. Munsell, Guorong Wu 0001 |
Medical Image Anal. | 3 |
| 2016 | Deep Relative Attributes
Yaser Souri, Erfan Noury, Ehsan Adeli-Mosabbeb |
ACCV (5) | 3 |
| 2016 | Semi-supervised Hierarchical Multimodal Feature and Sample Selection for Alzheimer's Disease Diagnosis
Ehsan Adeli-Mosabbeb, Mingxia Liu 0001, Jun Zhang 0018, Dinggang Shen |
MICCAI (2) | 2 |
| 2016 | Feature Selection Based on Iterative Canonical Correlation Analysis for Automatic Diagnosis of Parkinson's Disease
Luyan Liu, Qian Wang 0001, Ehsan Adeli-Mosabbeb, Lichi Zhang, Han Zhang 0002, Dinggang Shen |
MICCAI (2) | 3 |
| 2016 | 3D Deep Learning for Multi-modal Imaging-Guided Survival Time Prediction of Brain Tumor Patients
Dong Nie, Han Zhang 0002, Ehsan Adeli-Mosabbeb, Luyan Liu, Dinggang Shen |
MICCAI (2) | 3 |
| 2016 | Stability-Weighted Matrix Completion of Incomplete Multi-modal Data for Disease Diagnosis
Kim-Han Thung, Ehsan Adeli-Mosabbeb, Pew-Thian Yap, Dinggang Shen |
MICCAI (2) | 2 |
| 2016 | Progressive Graph-Based Transductive Learning for Multi-modal Classification of Brain Disorder Disease
Zhengxia Wang, Xiaofeng Zhu 0001, Ehsan Adeli-Mosabbeb, Yingying Zhu 0004, Chen Zu, Feiping Nie 0001, Dinggang Shen, Guorong Wu 0001 |
MICCAI (1) | 3 |
| 2016 | Multi-Level Canonical Correlation Analysis for Standard-Dose PET Image EstimationabstractPositron emission tomography (PET) images are widely used in many clinical applications, such as tumor detection and brain disorder diagnosis. To obtain PET images of diagnostic quality, a sufficient amount of radioactive tracer has to be injected into a living body, which will inevitably increase the risk of radiation exposure. On the other hand, if the tracer dose is considerably reduced, the quality of the resulting images would be significantly degraded. It is of great interest to estimate a standard-dose PET (S-PET) image from a low-dose one in order to reduce the risk of radiation exposure and preserve image quality. This may be achieved through mapping both S-PET and low-dose PET data into a common space and then performing patch-based sparse representation. However, a one-size-fits-all common space built from all training patches is unlikely to be optimal for each target S-PET patch, which limits the estimation accuracy. In this paper, we propose a data-driven multi-level canonical correlation analysis scheme to solve this problem. In particular, a subset of training data that is most useful in estimating a target S-PET patch is identified in each level, and then used in the next level to update common space and improve estimation. In addition, we also use multi-modal magnetic resonance images to help improve the estimation with complementary information. Validations on phantom and real human brain data sets show that our method effectively estimates S-PET images and well preserves critical clinical quantification measures, such as standard uptake value. Pei Zhang 0002, Ehsan Adeli-Mosabbeb, Yan Wang 0015, Guangkai Ma, Feng Shi 0001, David S. Lalush, Weili Lin, Dinggang Shen |
IEEE Trans. Image Process. | 3 |
| 2015 | Medical Image Retrieval Using Multi-graph Learning for MCI Diagnostic Assistance
Yue Gao 0002, Ehsan Adeli-Mosabbeb, Minjeong Kim 0001, Panteleimon Giannakopoulos, Sven Haller, Dinggang Shen |
MICCAI (2) | 2 |
| 2015 | Joint Diagnosis and Conversion Time Prediction of Progressive Mild Cognitive Impairment (pMCI) Using Low-Rank Subspace Clustering and Matrix Completion
Kim-Han Thung, Pew-Thian Yap, Ehsan Adeli-Mosabbeb, Dinggang Shen |
MICCAI (3) | 3 |
| 2015 | Robust Feature-Sample Linear Discriminant Analysis for Brain Disorders DiagnosisabstractA wide spectrum of discriminative methods is increasingly used in diverse applications for classification or regression tasks. However, many existing discriminative methods assume that the input data is nearly noise-free, which limits their applications to solve real-world problems. Particularly for disease diagnosis, the data acquired by the neuroimaging devices are always prone to different sources of noise. Robust discriminative models are somewhat scarce and only a few attempts have been made to make them robust against noise or outliers. These methods focus on detecting either the sample-outliers or feature-noises. Moreover, they usually use unsupervised de-noising procedures, or separately de-noise the training and the testing data. All these factors may induce biases in the learning process, and thus limit its performance. In this paper, we propose a classification method based on the least-squares formulation of linear discriminant analysis, which simultaneously detects the sample-outliers and feature-noises. The proposed method operates under a semi-supervised setting, in which both labeled training and unlabeled testing data are incorporated to form the intrinsic geometry of the sample space. Therefore, the violating samples or feature values are identified as sample-outliers or feature-noises, respectively. We test our algorithm on one synthetic and two brain neurodegenerative databases (particularly for Parkinson's disease and Alzheimer's disease). The results demonstrate that our method outperforms all baseline and state-of-the-art methods, in terms of both accuracy and the area under the ROC curve. Ehsan Adeli-Mosabbeb, Kim-Han Thung, Feng Shi 0001, Dinggang Shen |
NIPS | 1 |
| 2015 | Non-negative matrix completion for action detection
Ehsan Adeli-Mosabbeb, Mahmood Fathy |
Image Vis. Comput. | 1 |
| 2014 | Multi-label Discriminative Weakly-Supervised Human Activity Recognition and Localization
Ehsan Adeli-Mosabbeb, Ricardo Silveira Cabral, Fernando De la Torre, Mahmood Fathy |
ACCV (5) | 1 |
| 2014 | Distributed matrix completion for large-scale multi-label classificationabstractLarge-scale multi-label classification has always been of great interest for researchers. The difficulty with such problems is the huge amount of data that should be processed, possibly in multiple paths. This amount of data does not fit in the memory of a single computer and that is the bottle-nec k for many large-scale applications. On the other hand, matrix completion is a great tool for many applications, including classification. It is a great tool for modeling the data and finding the outliers and noises within the data. In this paper, we develop a distributed matrix completion method for multi-label classification. To do this, we first propose a simple distributed algorithm for minimizing the nuclear norm of a matrix to recover its low-rank representation, which is then generalized for the classification problem. Several synthetic and real datasets are used to verify both the distributed nuclear norm minimization and the distributed matrix completion approach. The results indicate that the proposed algorithm outperforms state-of-the-art methods for large-scale classification. Ehsan Adeli-Mosabbeb, Mahmood Fathy |
Intell. Data Anal. | 1 |
| 2010 | A non-parametric heuristic algorithm for convex and non-convex data clustering based on equipotential surfaces
Farhad Bayat, Ehsan Adeli-Mosabbeb, Ali Akbar Jalali 0001, Farshad Bayat |
Expert Syst. Appl. | 2 |