Dimitris N. Metaxas

dblp:m/DNMetaxas · DBLP profile ↗
← Back
457ranked-venue papers
20as first author
107since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 291 · 11 first-author · 59 since 2021Artificial intelligence and machine learning · 262 · 12 first-author · 58 since 2021Applied, interdisciplinary, general and emerging computing · 110 · 2 first-author · 30 since 2021Human-computer interaction and ubiquitous computing · 14 · 1 first-author · 2 since 2021Systems, architecture and hardware · 3Databases, data management, data science and information retrieval · 2 · 1 since 2021Security and privacy · 1Theory of computation · 1
YearPublicationVenuePosition
2026 Stable Signer: Hierarchical Sign Language Generative Model
abstract
Sign Language Production (SLP) is the process of converting the complex input text into a real video. Most previous works focused on the Text2Gloss, Gloss2Pose, Pose2Vid stages, and some concentrated on Prompt2Gloss and Text2Avatar stages. However, this field has made slow progress due to the inaccuracy of text conversion, pose generation, and the rendering of poses into real human videos in these stages, resulting in gradually accumulating errors. Therefore, in this paper, we streamline the traditional redundant structure, simplify and optimize the task objective, and design a new sign language generative model called Stable Signer. It redefines the SLP task as a hierarchical generation end-to-end task that only includes text understanding (Prompt2Gloss, Text2Gloss) and Pose2Vid, and executes text understanding through our proposed new Sign Language Understanding Linker called SLUL, and generates hand gestures through the named SLP-MoE hand gesture rendering expert block to end-to-end generate high-quality and multi-style sign language videos. SLUL is trained using the newly developed Semantic-Aware Gloss Masking Loss (SAGM Loss). Its performance has improved by 48.6% compared to the current SOTA generation methods.
Sen Fang, Yalin Feng, Hongbin Zhong, Dimitris N. Metaxas
ACL (1)5
2026 Reasoning over Precedents Alongside Statutes: Case-Augmented Deliberative Alignment for LLM Safety
abstract
Can Jin, Rui Wu, Tong Che, Qixin Zhang, Hongwu Peng, Jiahui Zhao, Zhenting Wang, Wenqi Wei, Ligong Han, Zhao Zhang, Yuan Cao, Ruixiang Tang, Dimitris N. Metaxas. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Can Jin, Tong Che, Qixin Zhang 0001, Hongwu Peng, Zhenting Wang, Ligong Han, Ruixiang Tang, Dimitris N. Metaxas
ACL (1)13
2026 Individual Turing Test: A Case Study of LLM-based Simulation Using Longitudinal Personal Data
abstract
Large Language Models (LLMs) have demonstrated remarkable human-like capabilities, yet their ability to replicate a specific individual remains underexplored. This paper presents a case study investigating LLM-based individual simulation using a volunteer-contributed archive of private messaging history spanning over ten years. Based on this dataset, we propose the ''Individual Turing Test'' to evaluate whether acquaintances of the volunteer can correctly identify which response in a multi-candidate pool most plausibly originates from the volunteer. We investigate prevalent approaches to LLM-based individual simulation, including fine-tuning, retrieval-augmented generation (RAG), memory-based methods, and hybrid approaches that integrate fine-tuning with RAG or memory. Empirical results show that current methods do not pass the Individual Turing Test, but perform substantially better when the same test is conducted on strangers to the target individual. Additionally, while fine-tuning improves performance in daily chats that reflect the individual's language style, retrieval-augmented and memory-based approaches demonstrate stronger performance on questions involving personal opinions and preferences. These findings reveal a fundamental trade-off between parametric and non-parametric approaches to individual simulation with LLMs under longitudinal context.
Ziyi Ye, Wujiang Xu, Xi Zhu 0004, Wenyue Hua, Dimitris N. Metaxas
SIGIR6
2026 Snapmoji: Instant Generation of Animatable Dual-Stylized Avatars
abstract
Despite the increasing popularity of avatar systems such as Snapchat Bitmojis, existing production avatar platforms face several limitations, such as a limited number of predefined assets, tedious customization processes, and inefficient rendering requirements. Addressing these shortcomings, we introduce Snapmoji, an avatar generation system that instantly creates 3D avatars, and enables customization in a process we call dual-stylization. Snapmoji first maps a selfie of a user to a primary avatar (e.g., Bitmoji style) using a new technique we name Gaussian Domain Adaptation (GDA), then applies a secondary style (e.g., skeleton, yarn, toy) to the primary avatar, all while preserving the user’s identity. The generated 3D avatars can then be rendered an animated on mobile devices at 30–40 FPS.
Eric Ming Chen, Di Liu 0003, Sizhuo Ma, Michael Vasilkovsky, Bing Zhou 0001, Wenzhou Wang, Jiahao Luo, Dimitris N. Metaxas, Vincent Sitzmann, Jian Wang 0100
WACV9
2026 Anatomy-VLM: A Fine-grained Vision-Language Model for Medical Interpretation
abstract
Accurate disease interpretation from radiology remains challenging due to imaging heterogeneity. Achieving expert-level diagnostic decisions requires integration of subtle image features with clinical knowledge. Yet major vision-language models (VLMs) treat images as holistic entities and overlook fine-grained image details that are vital for disease diagnosis. Clinicians analyze images by utilizing their prior medical knowledge and identify anatomical structures as important region of interests (ROIs). Inspired from this human-centric workflow, we introduce Anatomy-VLM, a fine-grained, vision-language model that incorporates multi-scale information. First, we design a model encoder to localize key anatomical features from entire medical images. Second, these regions are enriched with structured knowledge for contextually-aware interpretation. Finally, the model encoder aligns multi-scale medical information to generate clinically-interpretable disease prediction. Anatomy-VLM achieves outstanding performance on both in- and out-of-distribution datasets. We also validate the performance of Anatomy-VLM on downstream image segmentation tasks, suggesting that its fine-grained alignment captures anatomical and pathology-related knowledge. Furthermore, the Anatomy-VLM’s encoder facilitates zero-shot anatomy-wise interpretation, providing its strong expert-level clinical interpretation capabilities.
Difei Gu, Yunhe Gao, Mu Zhou, Dimitris N. Metaxas
WACV4
2026 DICE: Discrete Inversion Enabling Controllable Editing for Masked Generative Models
abstract
Recent advances in discrete diffusion models have demonstrated strong performance in image generation and masked language modeling, yet they remain limited in their capacity for controlled content editing. We propose DICE (Discrete Inversion for Controllable Editing), a novel framework that pioneers precise inversion capabilities for discrete diffusion models, including both masked generative and multinomial diffusion variants. Our key innovation lies in capturing noise sequences and masking patterns during reverse diffusion process, enabling both accurate reconstruction and flexible editing without relying on predefined masks or attention-based manipulations. Through comprehensive experiments across image and text modalities using models such as Paella, VQ-Diffusion, RoBERTa and LLaDA, we demonstrate that DICE successfully maintains high fidelity to the original data while significantly expanding editing capabilities. These results establish new possibilities for fine-grained content manipulation in discrete spaces.
Xiaoxiao He, Quan Dao, Ligong Han, Song Wen 0001, Minhao Bai, Di Liu 0003, Han Zhang 0010, Felix Juefei-Xu, Chaowei Tan, Bo Liu 0005, Martin Renqiang Min, Kang Li 0004, Faez Ahmed, Akash Srivastava, Hongdong Li, Junzhou Huang, Dimitris N. Metaxas
WACV17
2026 Large Sign Language Models: Toward 3D American Sign Language Translation
abstract
We present Large Sign Language Models (LSLM), a novel framework for translating 3D American Sign Language (ASL) by leveraging Large Language Models (LLMs) as the backbone, which can benefit hearing-impaired individuals’ virtual communication. Unlike existing sign language recognition methods that rely on 2D video, our approach directly utilizes 3D sign language data to capture rich spatial, gestural, and depth information in 3D scenes. This enables more accurate and resilient translation, enhancing digital communication accessibility for the hearing-impaired community. Beyond the task of ASL translation, our work explores the integration of complex, embodied multimodal languages into the processing capabilities of LLMs, moving beyond purely text-based inputs to broaden their understanding of human communication. We investigate both direct translation from 3D gesture features to text and an instruction-guided setting where translations can be modulated by external prompts, offering greater flexibility. This work provides a foundational step toward inclusive, multimodal intelligent systems capable of understanding diverse forms of language.
Xiaoxiao He, Di Liu 0003, Zhaoyang Xia, Chaowei Tan, Vivian Li, Bo Liu 0005, Dimitris N. Metaxas, Mubbasir Kapadia
WACV9
2026 MADCrowner: Margin Aware Dental Crown design with template deformation and refinement
Linda Wei, Wenran Zhang, Changyao Tian, Ke Wang 0036, Shaoting Zhang 0001, Dimitris N. Metaxas, Hongsheng Li 0001
Medical Image Anal.11
2026 Toward Modality- and Sampling-Universal Learning Strategies for Accelerating Cardiovascular Imaging: Summary of the CMRxRecon2024 Challenge
abstract
Cardiovascular health is vital to human well-being, and cardiac magnetic resonance (CMR) imaging is considered the clinical reference standard for diagnosing cardiovascular disease. However, its adoption is hindered by long scan times, complex contrasts, and inconsistent quality. While deep learning methods perform well on specific CMR imaging sequences, they often fail to generalize across modalities and sampling schemes. The lack of benchmarks for high-quality, fast CMR image reconstruction further limits technology comparison and adoption. The CMRxRecon2024 challenge, attracting over 200 teams from 18 countries, addressed these issues with two tasks: generalization to unseen modalities and robustness to diverse undersampling patterns. We introduced the largest public multi-modality CMR raw dataset, an open benchmarking platform, and shared code. Analysis of the best-performing solutions revealed that prompt-based adaptation and enhanced physics-driven consistency enabled strong cross-scenario performance. These findings establish principles for generalizable reconstruction models and advance clinically translatable AI in cardiovascular imaging.
Fanwen Wang, Zi Wang 0005, Yan Li 0064, Chen Qin, Shuo Wang 0011, Kunyuan Guo, Mengting Sun, Mingkai Huang, Michael Tänzer, Qirong Li, Yinzhe Wu 0001, Haosen Zhang, Kian Anvari Hamedani, Yuntong Lyu, Longyu Sun, Tianxing He, Lizhen Lan, Qiong Yao, Bingyu Xin, Dimitris N. Metaxas, Narges Razizadeh, Shahabedin Nabavi, George Yiasemis, Jonas Teuwen, Daniel B. Ennis, Zhihao Xue, Ruru Xu, Ilkay Öksüz, Donghang Lyu, Yanxin Huang, Xinrui Guo, Ruqian Hao, Jaykumar H. Patel, Guanke Cai, Binghua Chen, Sha Hua, Zhensen Chen, Qi Dou 0001, Xiahai Zhuang, Wenjia Bai, Harry Qin, He Wang 0016, Claudia Prieto, Michael Markl 0001, Alistair A. Young, Hao Li 0082, Xihong Hu, Lianming Wu, Xiaobo Qu 0001, Guang Yang 0006, Chengyan Wang
IEEE Trans. Medical Imaging26
2025 Self-Corrected Flow Distillation for Consistent One-Step and Few-Step Image Generation
abstract
Flow matching has emerged as a promising framework for training generative models, demonstrating impressive empirical performance while offering relative ease of training compared to diffusion-based models. However, this method still requires numerous function evaluations in the sampling process. To address these limitations, we introduce a self-corrected flow distillation method that effectively integrates consistency models and adversarial training within the flow-matching framework. This work is a pioneer in achieving consistent generation quality in both few-step and one-step sampling. Our extensive experiments validate the effectiveness of our method, yielding superior results both quantitatively and qualitatively on CelebA-HQ and zero-shot benchmarks on the COCO dataset.
Quan Dao, Hao Phung, Trung Tuan Dao, Dimitris N. Metaxas, Anh Tuan Tran 0001
AAAI4
2025 LUCAS: Layered Universal Codec Avatars
abstract
Photorealistic 3D head avatar reconstruction faces critical challenges in modeling dynamic face-hair interactions and achieving cross-identity generalization, particularly during expressions and head movements. We present LUCAS, a novel Universal Prior Model (UPM) for codec avatar modeling that disentangles face and hair through a layered representation. Unlike previous UPMs that treat hair as an integral part of the head, our approach separates the modeling of the hairless head and hair into distinct branches. LUCAS is the first to introduce a mesh-based UPM, facilitating real-time rendering on devices. Our layered representation also improves the anchor geometry for precise and visually appealing Gaussian renderings. Experimental results indicate that LUCAS outperforms existing single-mesh and Gaussian-based avatar models in both quantitative and qualitative assessments, including evaluations on held-out subjects in zero-shot driving scenarios. LUCAS demonstrates superior dynamic performance in managing head pose changes, expression transfer, and hairstyle variations, thereby advancing the state-of-the-art in 3D head avatar reconstruction. Project page: https://lsn33096.github.io/LUCAS/.
Di Liu 0003, Teng Deng, Giljoo Nam, Stanislav Pidhorskyi, Jason M. Saragih, Dimitris N. Metaxas, Chen Cao 0001
CVPR8
2025 Show and Segment: Universal Medical Image Segmentation via In-Context Learning
abstract
Medical image segmentation remains challenging due to the vast diversity of anatomical structures, imaging modalities, and segmentation tasks. While deep learning has made significant advances, current approaches struggle to generalize as they require task-specific training or fine-tuning on unseen classes. We present Iris, a novel In-context Reference Image guided Segmentation framework that enables flexible adaptation to novel tasks through the use of reference examples without fine-tuning. At its core, Iris features a lightweight context task encoding module that distills task-specific information from reference context image-label pairs. This rich context embedding information is used to guide the segmentation of target objects. By decoupling task encoding from inference, Iris supports diverse strategies from one-shot inference and context example ensemble to object-level context example retrieval and in-context tuning. Through comprehensive evaluation across twelve datasets, we demonstrate that Iris performs strongly compared to task-specific models on in-distribution tasks. On seven held-out datasets, Iris shows superior generalization to out-of-distribution data and unseen classes. Further, Iris’s task encoding module can automatically discover anatomical relationships across datasets and modalities, offering insights into medical objects without explicit anatomical supervision.
Yunhe Gao, Di Liu 0003, Zhuowei Li 0002, Yunsheng Li, Mu Zhou, Dimitris N. Metaxas
CVPR7
2025 MLLM-as-a-Judge for Image Safety without Human Labeling
abstract
Image content safety has become a significant challenge with the rise of visual media on online platforms. Meanwhile, in the age of AI-generated content (AIGC), many image generation models are capable of producing harmful content, such as images containing sexual or violent material. Thus, it becomes crucial to identify such unsafe images based on established safety rules. Pre-trained Multimodal Large Language Models (MLLMs) offer potential in this regard, given their strong pattern recognition abilities. Existing approaches typically fine-tune MLLMs with humanlabeled datasets, which however brings a series of drawbacks. First, relying on human annotators to label data following intricate and detailed guidelines is both expensive and labor-intensive. Furthermore, users of safety judgment systems may need to frequently update safety rules, making fine-tuning on human-based annotation more challenging. This raises the research question: Can we detect unsafe images by querying MLLMs in a zero-shot setting using a predefined safety constitution (a set of safety rules)? Our research showed that simply querying pre-trained MLLMs does not yield satisfactory results. This lack of effectiveness stems from factors such as the subjectivity of safety rules, the complexity of lengthy constitutions, and the inherent biases in the models. To address these challenges, we propose a MLLM-based method includes objectifying safety rules, assessing the relevance between rules and images, making quick judgments based on debiased token probabilities with logically complete yet simplified precondition chains for safety rules, and conducting more in-depth reasoning with cascaded chain-of-thought processes if necessary. Experiment results demonstrate that our method is highly effective for zero-shot image safety judgment tasks.
Zhenting Wang, Shuming Hu, Shiyu Zhao 0001, Xiaowen Lin, Felix Juefei-Xu, Zhuowei Li 0002, Ligong Han, Harihar Subramanyam, Jianfa Chen, Lingjuan Lyu, Shiqing Ma, Dimitris N. Metaxas
CVPR14
2025 SnapGen-V: Generating a Five-Second Video within Five Seconds on a Mobile Device
abstract
We have witnessed the unprecedented success of diffusion-based video generation over the past year. Recently proposed models from the community have wielded the power to generate cinematic and high-resolution videos with smooth motions from arbitrary input prompts. However, as a supertask of image generation, video generation models require more computation and are thus hosted mostly on cloud servers, limiting broader adoption among content creators. In this work, we propose a comprehensive acceleration framework to bring the power of the large-scale video diffusion model to the hands of edge users. From the network architecture scope, we initialize from a compact image backbone and search out the design and arrangement of temporal layers to maximize hardware efficiency. In addition, we propose a dedicated adversarial fine-tuning algorithm for our efficient model and reduce the denoising steps to 4. Our model, with only 0.6B parameters, can generate a 5-second video on an iPhone 16 PM within 5 seconds. Compared to server-side models that take minutes on powerful GPUs to generate a single video, we accelerate the generation by magnitudes while delivering on-par quality. Project page at https://snap-research.github.io/snapgen-v/.
Yushu Wu, Yanyu Li, Yanwu Xu 0003, Anil Kag, Yang Sui 0001, Huseyin Coskun, Aleksei Lebedev, Ju Hu, Dimitris N. Metaxas, Yanzhi Wang 0001, Sergey Tulyakov, Jian Ren 0005
CVPR11
2025 Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction
abstract
Prevailing Multimodal Large Language Models (MLLMs) encode the input image(s) as vision tokens and feed them into the language backbone, similar to how Large Language Models (LLMs) process the text tokens. However, the number of vision tokens increases quadratically as the image resolutions, leading to huge computational costs. In this paper, we consider improving MLLM’s efficiency from two scenarios, (I) Reducing computational cost without degrading the performance. (II) Improving the performance with given budgets. We start with our main finding that the ranking of each vision token sorted by attention scores is similar in each layer except the first layer. Based on it, we assume that the number of essential top vision tokens does not increase along layers. Accordingly, for Scenario I, we propose a greedy search algorithm (G-Search) to find the least number of vision tokens to keep at each layer from the shallow to the deep. Interestingly, G-Search is able to reach the optimal reduction strategy based on our assumption. For Scenario II, based on the reduction strategy from G-Search, we design a parametric sigmoid function (P-Sigmoid) to guide the reduction at each layer of the MLLM, whose parameters are optimized by Bayesian Optimization. Extensive experiments demonstrate that our approach can significantly accelerate those popular MLLMs, e.g. LLaVA, and InternVL2 models, by more than 2⇥ without performance drops. Our approach also far outperforms other token reduction methods when budgets are limited, achieving a better trade-off between efficiency and effectiveness.
Shiyu Zhao 0001, Zhenting Wang, Felix Juefei-Xu, Xide Xia, Miao Liu 0007, Mingfu Liang, Dimitris N. Metaxas, Licheng Yu
CVPR9
2025 FlowChef: Steering of Rectified Flow Models for Controlled Generations
Maitreya Patel, Song Wen 0001, Dimitris N. Metaxas, Yezhou Yang
ICCV3
2025 Implicit In-context Learning
abstract
In-context Learning (ICL) empowers large language models (LLMs) to swiftly adapt to unseen tasks at inference-time by prefixing a few demonstration examples before queries. Despite its versatility, ICL incurs substantial computational and memory overheads compared to zero-shot learning and is sensitive to the selection and order of demonstration examples. In this work, we introduce \textbf{Implicit In-context Learning} (I2CL), an innovative paradigm that reduces the inference cost of ICL to that of zero-shot learning with minimal information loss. I2CL operates by first generating a condensed vector representation, namely a context vector, extracted from the demonstration examples. It then conducts an inference-time intervention through injecting a linear combination of the context vector and query activations back into the model’s residual streams. Empirical evaluation on nine real-world tasks across three model architectures demonstrates that I2CL achieves few-shot level performance at zero-shot inference cost, and it exhibits robustness against variations in demonstration examples. Furthermore, I2CL facilitates a novel representation of ``task-ids'', enhancing task similarity detection and fostering effective transfer learning. We also perform a comprehensive analysis and ablation study on I2CL, offering deeper insights into its internal mechanisms. Code is available at https://github.com/LzVv123456/I2CL.
Zhuowei Li 0002, Zihao Xu 0001, Ligong Han, Yunhe Gao, Song Wen 0001, Di Liu 0003, Hao Wang 0014, Dimitris N. Metaxas
ICLR8
2025 Improved Training Technique for Latent Consistency Models
abstract
Consistency models are a new family of generative models capable of producing high-quality samples in either a single step or multiple steps. Recently, consistency models have demonstrated impressive performance, achieving results on par with diffusion models in the pixel space. However, the success of scaling consistency training to large-scale datasets, particularly for text-to-image and video generation tasks, is determined by performance in the latent space. In this work, we analyze the statistical differences between pixel and latent spaces, discovering that latent data often contains highly impulsive outliers, which significantly degrade the performance of iCT in the latent space. To address this, we replace Pseudo-Huber losses with Cauchy losses, effectively mitigating the impact of outliers. Additionally, we introduce a diffusion loss at early timesteps and employ optimal transport (OT) coupling to further enhance performance. Lastly, we introduce the adaptive scaling-$c$ scheduler to manage the robust training process and adopt Non-scaling LayerNorm in the architecture to better capture the statistics of the features and reduce outlier impact. With these strategies, we successfully train latent consistency models capable of high-quality sampling with one or two steps, significantly narrowing the performance gap between latent consistency and diffusion models. The implementation is released here: \url{https://github.com/quandao10/sLCT/}
Quan Dao, Khanh Doan, Di Liu 0003, Trung Le 0001, Dimitris N. Metaxas
ICLR5
2025 LoR-VP: Low-Rank Visual Prompting for Efficient Vision Model Adaptation
abstract
Visual prompting has gained popularity as a method for adapting pre-trained models to specific tasks, particularly in the realm of parameter-efficient tuning. However, existing visual prompting techniques often pad the prompt parameters around the image, limiting the interaction between the visual prompts and the original image to a small set of patches while neglecting the inductive bias present in shared information across different patches. In this study, we conduct a thorough preliminary investigation to identify and address these limitations. We propose a novel visual prompt design, introducing **Lo**w-**R**ank matrix multiplication for **V**isual **P**rompting (LoR-VP), which enables shared and patch-specific information across rows and columns of image pixels. Extensive experiments across seven network architectures and four datasets demonstrate significant improvements in both performance and efficiency compared to state-of-the-art visual prompting methods, achieving up to $6\times$ faster training times, utilizing $18\times$ fewer visual prompt parameters, and delivering a 3.1% improvement in performance.
Can Jin, Shiyu Zhao 0001, Zhenting Wang, Xiaoxiao He, Ligong Han, Tong Che, Dimitris N. Metaxas
ICLR9
2025 The Hidden Life of Tokens: Reducing Hallucination of Large Vision-Language Models Via Visual Information Steering
abstract
Large Vision-Language Models (LVLMs) can reason effectively over both textual and visual inputs, but they tend to hallucinate syntactically coherent yet visually ungrounded contents. In this paper, we investigate the internal dynamics of hallucination by examining the tokens logits rankings throughout the generation process, revealing three key patterns in how LVLMs process information: (1) gradual visual information loss – visually grounded tokens gradually become less favored throughout generation, and (2) early excitation – semantically meaningful tokens achieve peak activation in the layers earlier than the final layer. (3) hidden genuine information – visually grounded tokens though not being eventually decided still retain relatively high rankings at inference. Based on these insights, we propose VISTA (Visual Information Steering with Token-logit Augmentation), a training-free inference-time intervention framework that reduces hallucination while promoting genuine information. VISTA works by combining two complementary approaches: reinforcing visual information in activation space and leveraging early layer activations to promote semantically meaningful decoding. Compared to existing methods, VISTA requires no external supervision and is applicable to various decoding strategies. Extensive experiments show that VISTA on average reduces hallucination by about 40% on evaluated open-ended generation task, and it consistently outperforms existing methods on four benchmarks across four architectures under three decoding strategies. Code is available at https://github.com/LzVv123456/VISTA.
Zhuowei Li 0002, Haizhou Shi, Yunhe Gao, Di Liu 0003, Zhenting Wang, Yuxiao Chen 0002, Ting Liu 0005, Long Zhao 0003, Hao Wang 0014, Dimitris N. Metaxas
ICML10
2025 RadAlign: Advancing Radiology Report Generation with Vision-Language Concept Alignment
Difei Gu, Yunhe Gao, Yang Zhou 0053, Mu Zhou, Dimitris N. Metaxas
MICCAI (7)5
2025 Test-Time Spectrum-Aware Latent Steering for Zero-Shot Generalization in Vision-Language Models
abstract
Vision–Language Models (VLMs) excel at zero-shot inference but often degrade under test-time domain shifts. For this reason, episodic test-time adaptation strategies have recently emerged as powerful techniques for adapting VLMs to a single unlabeled image. However, existing adaptation strategies, such as test-time prompt tuning, typically require backpropagating through large encoder weights or altering core model components. In this work, we introduce Spectrum-Aware Test-Time Steering (STS), a lightweight adaptation framework that extracts a spectral subspace from the textual embeddings to define principal semantic directions, and learns to steer latent representations in a spectrum-aware manner by adapting a small number of per-sample shift parameters to minimize entropy across augmented views. STS operates entirely at inference in the latent space, without backpropagation through or modification of the frozen encoders. Building on standard evaluation protocols, our comprehensive experiments demonstrate that STS largely surpasses or compares favorably against state-of-the-art test-time adaptation methods, while introducing only a handful of additional parameters and achieving inference speeds up to 8× faster with a 12× smaller memory footprint than conventional test-time prompt tuning. The code is available at https://github.com/kdafnis/STS.
Konstantinos M. Dafnis, Dimitris N. Metaxas
NeurIPS2
2025 AutoEdit: Automatic Hyperparameter Tuning for Image Editing
abstract
Recent advances in diffusion models have revolutionized text-guided image editing, yet existing editing methods face critical challenges in hyperparameter identification. To get the reasonable editing performance, these methods often require the user to brute-force tune multiple interdependent hyperparameters, such as inversion timesteps and attention modification, \textit{etc.} This process incurs high computational costs due to the huge hyperparameter search space. We consider searching optimal editing's hyperparameters as a sequential decision-making task within the diffusion denoising process. Specifically, we propose a reinforcement learning framework, which establishes a Markov Decision Process that dynamically adjusts hyperparameters across denoising steps, integrating editing objectives into a reward function. The method achieves time efficiency through proximal policy optimization while maintaining optimal hyperparameter configurations. Experiments demonstrate significant reduction in search time and computational overhead compared to existing brute-force approaches, advancing the practical deployment of a diffusion-based image editing framework in the real world.
Quan Dao, Mahesh Bhosale, Yunjie Tian, Dimitris N. Metaxas, David S. Doermann
NeurIPS5
2025 Continuous Spatio-Temporal Memory Networks for 4D Cardiac Cine MRI Segmentation
abstract
Current cardiac cine magnetic resonance image (cMR) studies focus on the end diastole (ED) and end systole (ES) phases, while ignoring the abundant temporal information in the whole image sequence. This is because whole sequence segmentation is currently a tedious process and in-accurate. Conventional whole sequence segmentation approaches first estimate the motion field between frames, which is then used to propagate the mask along the temporal axis. However, the mask propagation results could be prone to error, especially for the basal and apex slices, where through-plane motion leads to significant morphology and structural change during the cardiac cycle. Inspired by recent advances in video object segmentation (VOS), based on spatiotemporal memory (STM) networks, we propose a continuous STM (CSTM) network for semi-supervised whole heart and whole sequence cMR segmentation. Our CSTM network takes full advantage of the spatial, scale, temporal and through-plane continuity prior of the underlying heart anatomy structures, to achieve accurate and fast 4D segmentation. Results of extensive experiments across multiple cMR datasets show that our method can improve the 4D cMR segmentation performance, especially for the hard-to-segment regions. Project page is at https://github.com/DeepTag/CSTM.
Meng Ye 0003, Bingyu Xin, Leon Axel, Dimitris N. Metaxas
WACV4
2025 SODA: Spectral Orthogonal Decomposition Adaptation for Diffusion Models
abstract
daptation (SODA), which balances computational efficiency and representation capacity. Extensive evaluations on text-to-image diffusion models demonstrate SODA's effectiveness, offering a spectrum-aware alternative to existing fine-tuning methods.
Xinxi Zhang, Song Wen 0001, Ligong Han, Felix Juefei-Xu, Akash Srivastava, Junzhou Huang, Vladimir Pavlovic 0001, Hao Wang 0014, Molei Tao, Dimitris N. Metaxas
WACV10
2025 The state-of-the-art in cardiac MRI reconstruction: Results of the CMRxRecon challenge in MICCAI 2023
Chen Qin, Shuo Wang 0011, Fanwen Wang, Yan Li 0064, Zi Wang 0005, Kunyuan Guo, Ouyang Cheng, Michael Tänzer, Longyu Sun, Mengting Sun, Zhang Shi, Sha Hua, Hao Li 0082, Zhensen Chen, Bingyu Xin, Dimitris N. Metaxas, George Yiasemis, Jonas Teuwen, Weitian Chen, Yidong Zhao, Yanwei Pang, Artem Razumov, Dmitry V. Dylov, Quan Dou, Yuyang Xue, Yuning Du, Julia Dietlmeier, Carles García-Cabrera, Ziad Al-Haj Hemidi, Nora Vogt, Ying-Hua Chu, Weibo Chen, Wenjia Bai, Xiahai Zhuang, Harry Qin, Lianming Wu, Guang Yang 0006, Xiaobo Qu 0001, He Wang 0016, Chengyan Wang
Medical Image Anal.20
2025 Editorial for Special Issue on Foundation Models for Medical Image Analysis
Xiaosong Wang 0001, Dequan Wang, Jens Rittscher, Dimitris N. Metaxas, Shaoting Zhang 0001
Medical Image Anal.5
2024 Instantaneous Perception of Moving Objects in 3D
abstract
The perception of 3D motion of surrounding traffic participants is crucial for driving safety. While existing works primarily focus on general large motions, we contend that the instantaneous detection and quantification of subtle motions is equally important as they indicate the nuances in driving behavior that may be safety critical, such as behaviors near a stop sign of parking positions. We delve into this under-explored task, examining its unique challenges and developing our solution, accompanied by a carefully designed benchmark. Specifically, due to the lack of correspondences between consecutive frames of sparse Lidar point clouds, static objects might appear to be moving - the socalled swimming effect. This intertwines with the true object motion, thereby posing ambiguity in accurate estimation, especially for subtle motions. To address this, we propose to leverage local occupancy completion of object point clouds to densify the shape cue, and mitigate the impact of swimming artifacts. The occupancy completion is learned in an end-to-end fashion together with the detection of moving objects and the estimation of their motion, instantaneously as soon as objects start to move. Extensive experiments demonstrate superior performance compared to standard 3D motion estimation approaches, particularly highlighting our method's specialized treatment of subtle motions.
Di Liu 0003, Bingbing Zhuang, Dimitris N. Metaxas, Manmohan Krishna Chandraker
CVPR3
2024 Score-Guided Diffusion for 3D Human Recovery
abstract
We present Score-Guided Human Mesh Recovery (ScoreHMR), an approach for solving inverse problems for 3D human pose and shape reconstruction. These inverse problems involve fitting a human body model to image ob-servations, traditionally solved through optimization techniques. ScoreHMR mimics model fitting approaches, but alignment with the image observation is achieved through score guidance in the latent space of a diffusion model. The diffusion model is trained to capture the conditional distribution of the human model parameters given an input image. By guiding its denoising process with a task-specific score, ScoreHMR effectively solves inverse problems for various applications without the need for retraining the task-agnostic diffusion model. We evaluate our approach on three settings/applications. These are: (i) single-frame model fitting; (ii) reconstruction from multiple uncalibrated views; (iii) reconstructing humans in video sequences. ScoreHMR consistently outperforms all optimization baselines on popular benchmarks across all settings. We make our code and models available on the project website: https://statho.github.io/ScoreHMR.
Anastasis Stathopoulos, Ligong Han, Dimitris N. Metaxas
CVPR3
2024 AVID: Any-Length Video Inpainting with Diffusion Model
abstract
Recent advances in diffusion models have successfully enabled text-guided image inpainting. While it seems straightforward to extend such editing capability into the video domain, there have been fewer works regarding textguided video inpainting. Given a video, a masked region at its initial frame, and an editing prompt, it requires a model to do infilling at each frame following the editing guidance while keeping the out-of-mask region intact. There are three main challenges in text-guided video inpainting: (i) temporal consistency of the edited video, (ii) supporting different inpainting types at different structural fidelity levels, and (iii) dealing with variable video length. To address these challenges, we introduce Any-Length Video Inpainting with Diffusion Model, dubbed as AVID. At its core, our model is equipped with effective motion modules and adjustable structure guidance, for fixed-length video inpainting. Building on top of that, we propose a novel Temporal MultiDiffusion sampling pipeline with a middle-frame attention guidance mechanism, facilitating the generation of videos with any desired duration. Our comprehensive experiments show our model can robustly deal with various inpainting types at different video duration ranges, with high quality11More visualization results are made publicly available here.
Bichen Wu, Yaqiao Luo, Luxin Zhang, Peter Vajda, Dimitris N. Metaxas, Licheng Yu
CVPR8
2024 Layout-Agnostic Scene Text Image Synthesis with Diffusion Models
abstract
While diffusion models have significantly advanced the quality of image generation, their capability to accurately and coherently render text within these images remains a substantial challenge. Conventional diffusion-based methods for scene text generation are typically limited by their reliance on an intermediate layout output. This dependency often results in a constrained diversity of text styles and fonts, an inherent limitation stemming from the deterministic nature of the layout generation phase. To address these challenges, this paper introduces Scene TextGen, a novel diffusion-based model specifically designed to circumvent the need for a predefined layout stage. By doing so, Scene-TextGen facilitates a more natural and varied representation of text. The novelty of SceneTextGen lies in its integration of three key components: a character-level encoder for capturing detailed typographic properties, coupled with a character-level instance segmentation model and a word-level spotting model to address the issues of unwanted text generation and minor character inaccuracies. We validate the performance of our method by demonstrating improved character recognition rates on generated images across different public visual text datasets in comparison to both standard diffusion based methods and text specific methods.
Qilong Zhangli, Jindong Jiang, Di Liu 0003, Licheng Yu, Xiaoliang Dai, Ankit Ramchandani, Guan Pang, Dimitris N. Metaxas, Praveen Krishnan
CVPR8
2024 Generating Enhanced Negatives for Training Language-Based Object Detectors
abstract
The recent progress in language-based open-vocabulary object detection can be largely attributed to finding better ways of leveraging large-scale data with free-form text annotations. Training such models with a discriminative objective function has proven successful, but requires good positive and negative samples. However, the free-form nature and the open vocabulary of object descriptions make the space of negatives extremely large. Prior works randomly sample negatives or use rule-based techniques to build them. In contrast, we propose to leverage the vast knowledge built into modern generative models to automatically build negatives that are more relevant to the original data. Specifically, we use large-Language-models to generate negative text descriptions, and text-to-image diffusion models to also generate corresponding negative images. Our experimental analysis confirms the relevance of the generated negative data, and its use in language-based detectors improves performance on two complex benchmarks. Code is available at https://github.com/xiaofeng94/Gen-Enhanced-Negs.
Shiyu Zhao 0001, Long Zhao 0003, Yumin Suh, Dimitris N. Metaxas, Manmohan Krishna Chandraker, Samuel Schulter
CVPR5
2024 Taming Self-Training for Open-Vocabulary Object Detection
abstract
Recent studies have shown promising performance in open-vocabulary object detection (OVD) by utilizing pseudo labels (PLs) from pre-trained vision and language models (VLMs). However, teacher-student self-training, a powerful and widely used paradigm to leverage PLs, is rarely explored for OVD. This work identifies two challenges of using self-training in OVD: noisy PLs from VLMs and frequent distribution changes of PLs. To address these challenges, we propose SAS-Det that tames self-training for OVD from two key perspectives. First, we present a split-and-fusion (SAF) head that splits a standard detection into an open-branch and a closed-branch. This design can reduce noisy supervision from pseudo boxes. More-over, the two branches learn complementary knowledge from different training data, significantly enhancing performance when fused together. Second, in our view, un-like in closed-set tasks, the PL distributions in OVD are solely determined by the teacher model. We introduce a periodic update strategy to decrease the number of up-dates to the teacher, thereby decreasing the frequency of changes in PL distributions, which stabilizes the training process. Extensive experiments demonstrate SAS-Det is both efficient and effective. SAS-Det outperforms recent models of the same scale by a clear margin and achieves 37.4 AP50and 29.1 APron novel categories of the COCO and LVIS benchmarks, respectively. Code is available at https://github.com/xiaofeng94/SAS-Det.
Shiyu Zhao 0001, Samuel Schulter, Long Zhao 0003, Yumin Suh, Manmohan Krishna Chandraker, Dimitris N. Metaxas
CVPR8
2024 Learning to Localize Actions in Instructional Videos with LLM-Based Multi-pathway Text-Video Alignment
Yuxiao Chen 0002, Kai Li 0012, Wentao Bao, Deep Patel, Yu Kong 0001, Martin Renqiang Min, Dimitris N. Metaxas
ECCV (82)7
2024 Rethinking Deep Unrolled Model for Accelerated MRI Reconstruction
Bingyu Xin, Meng Ye 0003, Leon Axel, Dimitris N. Metaxas
ECCV (75)4
2024 DIAGNOSIS: Detecting Unauthorized Data Usages in Text-to-image Diffusion Models
abstract
Recent text-to-image diffusion models have shown surprising performance in generating high-quality images. However, concerns have arisen regarding the unauthorized data usage during the training or fine-tuning process. One example is when a model trainer collects a set of images created by a particular artist and attempts to train a model capable of generating similar images without obtaining permission and giving credit to the artist. To address this issue, we propose a method for detecting such unauthorized data usage by planting the injected memorization into the text-to-image diffusion models trained on the protected dataset. Specifically, we modify the protected images by adding unique contents on these images using stealthy image warping functions that are nearly imperceptible to humans but can be captured and memorized by diffusion models. By analyzing whether the model has memorized the injected content (i.e., whether the generated images are processed by the injected post-processing function), we can detect models that had illegally utilized the unauthorized data. Experiments on Stable Diffusion and VQ Diffusion with different model training or fine-tuning methods (i.e, LoRA, DreamBooth, and standard training) demonstrate the effectiveness of our proposed method in detecting unauthorized data usages. Code: https://github.com/ZhentingWang/DIAGNOSIS.
Zhenting Wang, Chen Chen 0043, Lingjuan Lyu, Dimitris N. Metaxas, Shiqing Ma
ICLR4
2024 How to Trace Latent Generative Model Generated Images without Artificial Watermark?
abstract
Latent generative models (e.g., Stable Diffusion) have become more and more popular, but concerns have arisen regarding potential misuse related to images generated by these models. It is, therefore, necessary to analyze the origin of images by inferring if a particular image was generated by a specific latent generative model. Most existing methods (e.g., image watermark and model fingerprinting) require extra steps during training or generation. These requirements restrict their usage on the generated images without such extra operations, and the extra required operations might compromise the quality of the generated images. In this work, we ask whether it is possible to effectively and efficiently trace the images generated by a specific latent generative model without the aforementioned requirements. To study this problem, we design a latent inversion based method called LatentTracer to trace the generated images of the inspected model by checking if the examined images can be well-reconstructed with an inverted latent input. We leverage gradient based latent inversion and identify a encoder-based initialization critical to the success of our approach. Our experiments on the state-of-the-art latent generative models, such as Stable Diffusion, show that our method can distinguish the images generated by the inspected model and other images with a high accuracy and efficiency. Our findings suggest the intriguing possibility that today's latent generative generated images are naturally watermarked by the decoder used in the source models. Code: https://github.com/ZhentingWang/LatentTracer.
Zhenting Wang, Vikash Sehwag, Chen Chen 0043, Lingjuan Lyu, Dimitris N. Metaxas, Shiqing Ma
ICML5
2024 Aligning Human Knowledge with Visual Concepts Towards Explainable Medical Image Classification
Yunhe Gao, Difei Gu, Mu Zhou, Dimitris N. Metaxas
MICCAI (10)4
2024 BLoB: Bayesian Low-Rank Adaptation by Backpropagation for Large Language Models
abstract
Large Language Models (LLMs) often suffer from overconfidence during inference, particularly when adapted to downstream domain-specific tasks with limited data. Previous work addresses this issue by employing approximate Bayesian estimation after the LLMs are trained, enabling them to quantify uncertainty. However, such post-training approaches' performance is severely limited by the parameters learned during training. In this paper, we go beyond post-training Bayesianization and propose Bayesian Low-Rank Adaptation by Backpropagation (BLoB), an algorithm that continuously and jointly adjusts both the mean and covariance of LLM parameters throughout the whole fine-tuning process. Our empirical results verify the effectiveness of BLoB in terms of generalization and uncertainty estimation, when evaluated on both in-distribution and out-of-distribution data.
Yibin Wang 0005, Haizhou Shi, Ligong Han, Dimitris N. Metaxas, Hao Wang 0014
NeurIPS4
2024 Learning from Teaching Regularization: Generalizable Correlations Should be Easy to Imitate
abstract
Generalization remains a central challenge in machine learning. In this work, we propose *Learning from Teaching* (**LoT**), a novel regularization technique for deep neural networks to enhance generalization. Inspired by the human ability to capture concise and abstract patterns, we hypothesize that generalizable correlations are expected to be easier to imitate. LoT operationalizes this concept to improve the generalization of the main model with auxiliary student learners. The student learners are trained by the main model and, in turn, provide feedback to help the main model capture more generalizable and imitable correlations. Our experimental results across several domains, including Computer Vision, Natural Language Processing, and methodologies like Reinforcement Learning, demonstrate that the introduction of LoT brings significant benefits compared to training models on the original dataset. The results suggest the effectiveness and efficiency of LoT in identifying generalizable information at the right scales while discarding spurious data correlations, thus making LoT a valuable addition to current machine learning. Code is available at https://github.com/jincan333/LoT.
Can Jin, Tong Che, Hongwu Peng, Yiyuan Li, Dimitris N. Metaxas, Marco Pavone 0001
NeurIPS5
2024 DiMSUM: Diffusion Mamba - A Scalable and Unified Spatial-Frequency Method for Image Generation
abstract
We introduce a novel state-space architecture for diffusion models, effectively harnessing spatial and frequency information to enhance the inductive bias towards local features in input images for image generation tasks. While state-space networks, including Mamba, a revolutionary advancement in recurrent neural networks, typically scan input sequences from left to right, they face difficulties in designing effective scanning strategies, especially in the processing of image data. Our method demonstrates that integrating wavelet transformation into Mamba enhances the local structure awareness of visual inputs and better captures long-range relations of frequencies by disentangling them into wavelet subbands, representing both low- and high-frequency components. These wavelet-based outputs are then processed and seamlessly fused with the original Mamba outputs through a cross-attention fusion layer, combining both spatial and frequency information to optimize the order awareness of state-space models which is essential for the details and overall quality of image generation. Besides, we introduce a globally-shared transformer to supercharge the performance of Mamba, harnessing its exceptional power to capture global relationships. Through extensive experiments on standard benchmarks, our method demonstrates superior results compared to DiT and DIFFUSSM, achieving faster training convergence and delivering high-quality outputs. The codes and pretrained models are released at https://github.com/VinAIResearch/DiMSUM.git.
Hao Phung, Quan Dao, Trung Tuan Dao, Viet Hoang Phan, Dimitris N. Metaxas, Anh Tuan Tran 0001
NeurIPS5
2024 SF-V: Single Forward Video Generation Model
abstract
Diffusion-based video generation models have demonstrated remarkable success in obtaining high-fidelity videos through the iterative denoising process. However, these models require multiple denoising steps during sampling, resulting in high computational costs. In this work, we propose a novel approach to obtain single-step video generation models by leveraging adversarial training to fine-tune pre-trained video diffusion models. We show that, through the adversarial training, the multi-steps video diffusion model, i.e., Stable Video Diffusion (SVD), can be trained to perform single forward pass to synthesize high-quality videos, capturing both temporal and spatial dependencies in the video data. Extensive experiments demonstrate that our method achieves competitive generation quality of synthesized videos with significantly reduced computational overhead for the denoising process (i.e., around $23\times$ speedup compared with SVD and $6\times$ speedup compared with existing works, with even better generation quality), paving the way for real-time video synthesis and editing.
Yanyu Li, Yushu Wu, Yanwu Xu 0003, Anil Kag, Ivan Skorokhodov, Willi Menapace, Aliaksandr Siarohin, Junli Cao, Dimitris N. Metaxas, Sergey Tulyakov, Jian Ren 0005
NeurIPS10
2024 Second-Order Graph ODEs for Multi-Agent Trajectory Forecasting
abstract
Trajectory forecasting of multiple agents is a fundamental task that has applications in various fields, such as autonomous driving, physical system modeling and smart cities. It is challenging because agent interactions and underlying continuous dynamics jointly affect its behavior. Existing approaches often rely on Graph Neural Networks (GNNs) or Transformers to extract agent interaction features. However, they tend to neglect how the distance and velocity information between agents impact their interactions dynamically. Moreover, previous methods use RNNs or first-order Ordinary Differential Equations (ODEs) to model temporal dynamics, which may lack interpretability with respect to how each agent is driven by interactions. To address these challenges, this paper proposes the Agent Graph ODE, a novel approach that models agent interactions and continuous second-order dynamics explicitly. Our method utilizes a variational autoencoder architecture, incorporating spatial-temporal Transformers with distance information and dynamic interaction graph construction in the encoder module. In the decoder module, we employ GNNs with distance information to model agent interactions, and use coupled second-order ODEs to capture the underlying continuous dynamics by modeling the relationship between acceleration and agent interactions. Experimental results show that our proposed Agent Graph ODE outperforms state-of-the-art methods in prediction accuracy. Moreover, our method performs well in sudden situations not seen in the training dataset.
Song Wen 0001, Hao Wang 0014, Di Liu 0003, Qilong Zhangli, Dimitris N. Metaxas
WACV5
2024 Steering Prototypes with Prompt-tuning for Rehearsal-free Continual Learning
abstract
In the context of continual learning, prototypes—as representative class embeddings—offer advantages in memory conservation and the mitigation of catastrophic forgetting. However, challenges related to semantic drift and prototype interference persist. In this study, we introduce the Contrastive Prototypical Prompt (CPP) approach. Through task-specific prompt-tuning, underpinned by a contrastive learning objective, we effectively address both aforementioned challenges. Our evaluations on four challenging class-incremental benchmarks reveal that CPP achieves a significant 4% to 6% improvement over state-of-the-art methods. Importantly, CPP operates without a rehearsal buffer and narrows the performance divergence between continual and offline joint-learning, suggesting an innovative scheme for Transformer-based continual learning systems1.
Zhuowei Li 0002, Long Zhao 0003, Han Zhang 0010, Di Liu 0003, Ting Liu 0005, Dimitris N. Metaxas
WACV7
2024 Unsupervised Exemplar-Based Image-to-Image Translation and Cascaded Vision Transformers for Tagged and Untagged Cardiac Cine MRI Registration
abstract
Multi-modal registration between tagged and untagged cardiac cine magnetic resonance (MR) images remains difficult, due to the domain gap and large deformations between the two modalities. Recent work using an image-to-image translation (I2I) module to overcome the domain gap can convert the multi-modal into a mono-modal registration task and take advantage of advanced mono-modal registration architectures. However, they often ignore two issues: the sample-specific style of each image to be registered during I2I and large hybrid rigid and non-rigid deformations between modalities. We first propose an exemplar-based I2I module capable of unsupervised cross-domain correspondence learning to enforce the style consistency between the fake image and the image to be registered. Then we propose an efficient cascaded vision transformer-based registration network to predict both the affine and non-rigid deformations, in which a single feature embedding subnetwork is shared by the two stages of deformation prediction. We validated our method on a clinical cardiac MR dataset with paired but unaligned untagged and tagged MR images. The results show that our method outperforms traditional methods significantly in terms of the I2I quality and multi-modal image registration accuracy.
Meng Ye 0003, Mikael Kanski, Dong Yang 0005, Leon Axel, Dimitris N. Metaxas
WACV5
2024 ProxEdit: Improving Tuning-Free Real Image Editing with Proximal Guidance
abstract
DDIM inversion has revealed the remarkable potential of real image editing within diffusion-based methods. However, the accuracy of DDIM reconstruction degrades as larger classifier-free guidance (CFG) scales being used for enhanced editing. Null-text inversion (NTI) optimizes null embeddings to align the reconstruction and inversion trajectories with larger CFG scales, enabling real image editing with cross-attention control. Negative-prompt inversion (NPI) further offers a training-free closed-form solution of NTI. However, it may introduce artifacts and is still constrained by DDIM reconstruction quality. To overcome these limitations, we propose proximal guidance and incorporate it to NPI with cross-attention control. We enhance NPI with a regularization term and inversion guidance, which reduces artifacts while capitalizing on its training-free nature. Additionally, we extend the concepts to incorporate mutual self-attention control, enabling geometry and layout alterations in the editing process. Our method provides an efficient and straightforward approach, effectively addressing real image editing tasks with minimal computational overhead.
Ligong Han, Song Wen 0001, Kunpeng Song, Mengwei Ren, Ruijiang Gao, Anastasis Stathopoulos, Xiaoxiao He, Yuxiao Chen 0002, Di Liu 0003, Qilong Zhangli, Jindong Jiang, Zhaoyang Xia, Akash Srivastava, Dimitris N. Metaxas
WACV16
2024 StyleGAN-Fusion: Diffusion Guided Domain Adaptation of Image Generators
abstract
Can a text-to-image diffusion model be used as a training objective for adapting a GAN generator to another domain? In this paper, we show that the classifier-free guidance can be leveraged as a critic and enable generators to distill knowledge from large-scale text-to-image diffusion models. Generators can be efficiently shifted into new domains indicated by text prompts without access to groundtruth samples from target domains. We demonstrate the effectiveness and controllability of our method through extensive experiments. Although not trained to minimize CLIP loss, our model achieves equally high CLIP scores and significantly lower FID than prior work on short prompts, and outperforms the baseline qualitatively and quantitatively on long and complicated prompts. To our best knowledge, the proposed method is the first attempt at incorporating large-scale pre-trained diffusion models and distillation sampling for text-driven image generator domain adaptation and gives a quality previously beyond possible. Moreover, we extend our work to 3D-aware style-based generators and DreamBooth guidance. For code and more visual samples, please visit our Project Webpage.
Kunpeng Song, Ligong Han, Dimitris N. Metaxas, Ahmed M. Elgammal
WACV4
2024 AVDNet: Joint coronary artery and vein segmentation with topological consistency
abstract
Coronary CT angiography (CCTA) is an effective and non-invasive method for coronary artery disease diagnosis. Extracting an accurate coronary artery tree from CCTA image is essential for centerline extraction, plaque detection, and stenosis quantification. In practice, data quality varies. Sometimes, the arteries and veins have similar intensities and locate closely, which may confuse segmentation algorithms, even deep learning based ones, to obtain accurate arteries. However, it is not always feasible to re-scan the patient for better image quality. In this paper, we propose an artery and vein disentanglement network (AVDNet) for robust and accurate segmentation by incorporating the coronary vein into the segmentation task. This is the first work to segment coronary artery and vein at the same time. The AVDNet consists of an image based vessel recognition network (IVRN) and a topology based vessel refinement network (TVRN). IVRN learns to segment the arteries and veins, while TVRN learns to correct the segmentation errors based on topology consistency. We also design a novel inverse distance weighted dice (IDD) loss function to recover more thin vessel branches and preserve the vascular boundaries. Extensive experiments are conducted on a multi-center dataset of 700 patients. Quantitative and qualitative results demonstrate the effectiveness of the proposed method by comparing it with state-of-the-art methods and different variants. Prediction results of the AVDNet on the Automated Segmentation of Coronary Artery Challenge dataset are avaliabel at https://github.com/WennyJJ/Coronary-Artery-Vein-Segmentation for follow-up research.
Wenji Wang, Qing Xia 0002, Zhennan Yan, Xiao Wang 0004, Shaoping Nie, Dimitris N. Metaxas, Shaoting Zhang 0001
Medical Image Anal.9
2024 On the challenges and perspectives of foundation models for medical image analysis
Shaoting Zhang 0001, Dimitris N. Metaxas
Medical Image Anal.2
2024 Classification of lung cancer subtypes on CT images with synthetic pathological priors
Wentao Zhu 0002, Gege Ma, Geng Chen 0001, Jan Egger, Shaoting Zhang 0001, Dimitris N. Metaxas
Medical Image Anal.7
2024 DA-Tran: Multiphase liver tumor segmentation with a domain-adaptive transformer network
Yangfan Ni, Geng Chen 0001, Zhan Feng, Heng Cui, Dimitris N. Metaxas, Shaoting Zhang 0001, Wentao Zhu 0002
Pattern Recognit.5
2024 Gloss Prior Guided Visual Feature Learning for Continuous Sign Language Recognition
abstract
Continuous sign language recognition (CSLR) is to recognize the glosses in a sign language video. Enhancing the generalization ability of CSLR's visual feature extractor is a worthy area of investigation. In this paper, we model glosses as priors that help to learn more generalizable visual features. Specifically, the signer-invariant gloss feature is extracted by a pre-trained gloss BERT model. Then we design a gloss prior guidance network (GPGN). It contains a novel parallel densely-connected temporal feature extraction (PDC-TFE) module for multi-resolution visual feature extraction. The PDC-TFE captures the complex temporal patterns of the glosses. The pre-trained gloss feature guides the visual feature learning through a cross-modality matching loss. We propose to formulate the cross-modality feature matching into a regularized optimal transport problem, it can be efficiently solved by a variant of the Sinkhorn algorithm. The GPGN parameters are learned by optimizing a weighted sum of the cross-modality matching loss and CTC loss. The experiment results on German and Chinese sign language benchmarks demonstrate that the proposed GPGN achieves competitive performance. The ablation study verifies the effectiveness of several critical components of the GPGN. Furthermore, the proposed pre-trained gloss BERT model and cross-modality matching can be seamlessly integrated into other RGB-cue-based CSLR methods as plug-and-play formulations to enhance the generalization ability of the visual feature extractor.
Leming Guo, Wanli Xue, Bo Liu 0005, Kaihua Zhang 0001, Tiantian Yuan, Dimitris N. Metaxas
IEEE Trans. Image Process.6
2023 Revisiting Multimodal Representation in Contrastive Learning: From Patch and Token Embeddings to Finite Discrete Tokens
abstract
Contrastive learning-based vision-language pretraining approaches, such as CLIP, have demonstrated great success in many vision-language tasks. These methods achieve cross-modal alignment by encoding a matched image-text pair with similar feature embeddings, which are generated by aggregating information from visual patches and language tokens. However, direct aligning cross-modal information using such representations is challenging, as visual patches and text tokens differ in semantic levels and granularities. To alleviate this issue, we propose a Finite Discrete Tokens (FDT) based multimodal representation. FDT is a set of learnable tokens representing certain visualsemantic concepts. Both images and texts are embedded using shared FDT by first grounding multimodal inputs to FDT space and then aggregating the activated FDT representations. The matched visual and semantic concepts are enforced to be represented by the same set of discrete tokens by a sparse activation constraint. As a result, the granularity gap between the two modalities is reduced. Through both quantitative and qualitative analyses, we demonstrate that using FDT representations in CLIP-style models improves cross-modal alignment and performance in visual recognition and vision-language downstream tasks. Furthermore, we show that our method can learn more comprehensive representations, and the learned FDT capture meaningful cross-modal correspondence, ranging from objects to actions and attributes.11The source code can be found at https://github.com/yuxiaochen1103/FDT.
Yuxiao Chen 0002, Yu Tian 0003, Shijie Geng, Dimitris N. Metaxas, Hongxia Yang
CVPR7
2023 Learning Articulated Shape with Keypoint Pseudo-Labels from Web Images
abstract
This paper shows that it is possible to learn models for monocular 3D reconstruction of articulated objects (e.g. horses, cows, sheep), using as few as 50–150 images labeled with 2D keypoints. Our proposed approach involves training category-specific keypoint estimators, generating 2D keypoint pseudo-labels on unlabeled web images, and using both the labeled and self-labeled sets to train 3D reconstruction models. It is based on two key insights: (1) 2D keypoint estimation networks trained on as few as 50– 150 images of a given object category generalize well and generate reliable pseudo-labels; (2) a data selection mechanism can automatically create a “curated” subset of the unlabeled web images that can be used for training - we evaluate four data selection methods. Coupling these two insights enables us to train models that effectively utilize web images, resulting in improved 3D reconstruction performance for several articulated object categories beyond the fully-supervised baseline. Our approach can quickly bootstrap a model and requires only a few images labeled with 2D keypoints. This requirement can be easily satisfied for any new object category. To showcase the practicality of our approach for predicting the 3D shape of arbitrary object categories, we annotate 2D keypoints on 250 giraffe and bear images from COCO in just 2.5 hours per category.
Anastasis Stathopoulos, Georgios Pavlakos, Ligong Han, Dimitris N. Metaxas
CVPR4
2023 SINE: SINgle Image Editing with Text-to-Image Diffusion Models
abstract
Recent works on diffusion models have demonstrated a strong capability for conditioning image generation, e.g., text-guided image synthesis. Such success inspires many efforts trying to use large-scale pre-trained diffusion models for tackling a challenging problem-real image editing. Works conducted in this area learn a unique textual token corresponding to several images containing the same object. However, under many circumstances, only one image is available, such as the painting of the Girl with a Pearl Earring. Using existing works on fine-tuning the pre-trained diffusion models with a single image causes severe overfitting issues. The information leakage from the pre-trained diffusion models makes editing can not keep the same content as the given image while creating new features depicted by the language guidance. This work aims to address the problem of single-image editing. We propose a novel model-based guidance built upon the classifier-free guidance so that the knowledge from the model trained on a single image can be distilled into the pre-trained diffusion model, enabling content creation even with one given image. Additionally, we propose a patch-based fine-tuning that can effectively help the model generate images of arbitrary resolution. We provide extensive experiments to validate the design choices of our approach and show promising editing capabilities, including changing style, content addition, and object manipulation. Our code is made publicly available here.
Ligong Han, Dimitris N. Metaxas, Jian Ren 0005
CVPR4
2023 DeFormer: Integrating Transformers with Deformable Models for 3D Shape Abstraction from a Single Image
abstract
Accurate 3D shape abstraction from a single 2D image is a long-standing problem in computer vision and graphics. By leveraging a set of primitives to represent the target shape, recent methods have achieved promising results. However, these methods either use a relatively large number of primitives or lack geometric flexibility due to the limited expressibility of the primitives. In this paper, we propose a novel bi-channel Transformer architecture, integrated with parameterized deformable models, termed DeFormer, to simultaneously estimate the global and local deformations of primitives. In this way, DeFormer can abstract complex object shapes while using a small number of primitives which offer a broader geometry coverage and finer details. Then, we introduce a force-driven dynamic fitting and a cycle-consistent re-projection loss to optimize the primitive parameters. Extensive experiments on ShapeNet across various settings show that DeFormer achieves better reconstruction accuracy over the state-of-the-art, and visualizes with consistent semantic correspondences for improved interpretability.
Di Liu 0003, Xiang Yu 0002, Meng Ye 0003, Qilong Zhangli, Zhuowei Li 0002, Dimitris N. Metaxas
ICCV7
2023 Neural Deformable Models for 3D Bi-Ventricular Heart Shape Reconstruction and Modeling from 2D Sparse Cardiac Magnetic Resonance Imaging
abstract
We propose a novel neural deformable model (NDM) targeting at the reconstruction and modeling of 3D bi-ventricular shape of the heart from 2D sparse cardiac magnetic resonance (CMR) imaging data. We model the bi-ventricular shape using blended deformable superquadrics, which are parameterized by a set of geometric parameter functions and are capable of deforming globally and locally. While global geometric parameter functions and deformations capture gross shape features from visual data, local deformations, parameterized as neural diffeomorphic point flows, can be learned to recover the detailed heart shape. Different from iterative optimization methods used in conventional deformable model formulations, NDMs can be trained to learn such geometric parameter functions, global and local deformations from a shape distribution manifold. Our NDM can learn to densify a sparse cardiac point cloud with arbitrary scales and generate high-quality triangular meshes automatically. It also enables the implicit learning of dense correspondences among different heart shape instances for accurate cardiac shape registration. Furthermore, the parameters of NDM are intuitive, and can be used by a physician without sophisticated post-processing. Experimental results on a large CMR dataset demonstrate the improved performance of NDM over conventional methods.
Meng Ye 0003, Dong Yang 0005, Mikael Kanski, Leon Axel, Dimitris N. Metaxas
ICCV5
2023 SVDiff: Compact Parameter Space for Diffusion Fine-Tuning
abstract
Diffusion models have achieved remarkable success in text-to-image generation, enabling the creation of high-quality images from text prompts or other modalities. However, existing methods for customizing these models are limited by handling multiple personalized subjects and the risk of overfitting. Moreover, their large number of parameters is inefficient for model storage. In this paper, we propose a novel approach to address these limitations in existing text-to-image diffusion models for personalization. Our method involves fine-tuning the singular values of the weight matrices, leading to a compact and efficient parameter space that reduces the risk of overfitting and language-drifting. We also propose a Cut-Mix-Unmix data-augmentation technique to enhance the quality of multi-subject image generation and a simple text-based image editing framework. Our proposed SVDiff method has a significantly smaller model size compared to existing methods (≈2,200 times fewer parameters compared with vanilla DreamBooth), making it more practical for real-world applications.
Ligong Han, Yinxiao Li, Han Zhang 0010, Peyman Milanfar, Dimitris N. Metaxas, Feng Yang 0008
ICCV5
2023 OmniLabel: A Challenging Benchmark for Language-Based Object Detection
abstract
Language-based object detection is a promising direction towards building a natural interface to describe objects in images that goes far beyond plain category names. While recent methods show great progress in that direction, proper evaluation is lacking. With OmniLabel, we propose a novel task definition, dataset, and evaluation metric. The task subsumes standard- and open-vocabulary detection as well as referring expressions. With more than 28K unique object descriptions on over 25K images, OmniLabel provides a challenging benchmark with diverse and complex object descriptions in a naturally open-vocabulary setting. Moreover, a key differentiation to existing benchmarks is that our object descriptions can refer to one, multiple or even no object, hence, providing negative examples in free-form text. The proposed evaluation handles the large label space and judges performance via a modified average precision metric, which we validate by evaluating strong language-based baselines. OmniLabel indeed provides a challenging test bed for future research on language-based detection. Visit the project website at https://www.omnilabel.org
Samuel Schulter, Yumin Suh, Konstantinos M. Dafnis, Shiyu Zhao 0001, Dimitris N. Metaxas
ICCV7
2023 Pathology-and-Genomics Multimodal Transformer for Survival Outcome Prediction
Kexin Ding, Mu Zhou, Dimitris N. Metaxas, Shaoting Zhang 0001
MICCAI (6)3
2023 DMCVR: Morphology-Guided Diffusion Model for 3D Cardiac Volume Reconstruction
Xiaoxiao He, Chaowei Tan, Ligong Han, Bo Liu 0005, Leon Axel, Kang Li 0004, Dimitris N. Metaxas
MICCAI (7)7
2023 LEPARD: Learning Explicit Part Discovery for 3D Articulated Shape Reconstruction
abstract
Reconstructing the 3D articulated shape of an animal from a single in-the-wild image is a challenging task. We propose LEPARD, a learning-based framework that discovers semantically meaningful 3D parts and reconstructs 3D shapes in a part-based manner. This is advantageous as 3D parts are robust to pose variations due to articulations and their shape is typically simpler than the overall shape of the object. In our framework, the parts are explicitly represented as parameterized primitive surfaces with global and local deformations in 3D that deform to match the image evidence. We propose a kinematics-inspired optimization to guide each transformation of the primitive deformation given 2D evidence. Similar to recent approaches, LEPARD is only trained using off-the-shelf deep features from DINO and does not require any form of 2D or 3D annotations. Experiments on 3D animal shape reconstruction, demonstrate significant improvement over existing alternatives in terms of both the overall reconstruction performance as well as the ability to discover semantically meaningful and consistent parts.
Di Liu 0003, Anastasis Stathopoulos, Qilong Zhangli, Yunhe Gao, Dimitris N. Metaxas
NeurIPS5
2023 More Than Just Attention: Improving Cross-Modal Attentions with Contrastive Constraints for Image-Text Matching
abstract
Cross-modal attention mechanisms have been widely applied to the image-text matching task. They have achieved remarkable improvements thanks to their capability of learning fine-grained relevance across different modalities. However, the cross-modal attention models of existing methods could be sub-optimal and inaccurate because there is no direct supervision provided during the training process. In this work, we propose two novel training strategies, namely Contrastive Content Resourcing (CCR) and Contrastive Content Swapping (CCS) constraints, to address such limitations. These constraints supervise the training of cross-modal attention models in a contrastive learning manner without requiring explicit attention annotations. They are plug-in training strategies and can be generally integrated into existing cross-modal attention models. Additionally, we introduce three metrics, including Attention Precision, Recall, and F1-Score, to quantitatively measure the quality of learned attention models. We evaluate the proposed constraints by incorporating them into four state- of-the-art cross-modal attention-based image-text matching models. Experimental results on both Flickr30k and MS-COCO datasets demonstrate that integrating these constraints generally improves the model performance in terms of both retrieval performance and attention metrics.
Yuxiao Chen 0002, Long Zhao 0003, Larry Davis 0001, Dimitris N. Metaxas
WACV7
2023 CDDSA: Contrastive domain disentanglement and style augmentation for generalizable medical image segmentation
Ran Gu, Guotai Wang, Jiangshan Lu, Jingyang Zhang, Wenhui Lei, Wenjun Liao, Shichuan Zhang, Kang Li 0004, Dimitris N. Metaxas, Shaoting Zhang 0001
Medical Image Anal.10
2023 Editorial for special issue on explainable and generalizable deep learning methods for medical image computing
Guotai Wang, Shaoting Zhang 0001, Sharon X. Huang, Tom Vercauteren, Dimitris N. Metaxas
Medical Image Anal.5
2023 SequenceMorph: A Unified Unsupervised Learning Framework for Motion Tracking on Cardiac Image Sequences
abstract
Modern medical imaging techniques, such as ultrasound (US) and cardiac magnetic resonance (MR) imaging, have enabled the evaluation of myocardial deformation directly from an image sequence. While many traditional cardiac motion tracking methods have been developed for the automated estimation of the myocardial wall deformation, they are not widely used in clinical diagnosis, due to their lack of accuracy and efficiency. In this paper, we propose a novel deep learning-based fully unsupervised method, SequenceMorph, for in vivo motion tracking in cardiac image sequences. In our method, we introduce the concept of motion decomposition and recomposition. We first estimate the inter-frame (INF) motion field between any two consecutive frames, by a bi-directional generative diffeomorphic registration neural network. Using this result, we then estimate the Lagrangian motion field between the reference frame and any other frame, through a differentiable composition layer. Our framework can be extended to incorporate another registration network, to further reduce the accumulated errors introduced in the INF motion tracking step, and to refine the Lagrangian motion estimation. By utilizing temporal information to perform reasonable estimations of spatio-temporal motion fields, this novel method provides a useful solution for image sequence motion tracking. Our method has been applied to US (echocardiographic) and cardiac MR (untagged and tagged cine) image sequences; the results show that SequenceMorph is significantly superior to conventional motion tracking methods, in terms of the cardiac motion tracking accuracy and inference efficiency.
Meng Ye 0003, Dong Yang 0005, Qiaoying Huang, Mikael Kanski, Leon Axel, Dimitris N. Metaxas
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 Deep Learning Segmentation of the Right Ventricle in Cardiac MRI: The M&Ms Challenge
abstract
In recent years, several deep learning models have been proposed to accurately quantify and diagnose cardiac pathologies. These automated tools heavily rely on the accurate segmentation of cardiac structures in MRI images. However, segmentation of the right ventricle is challenging due to its highly complex shape and ill-defined borders. Hence, there is a need for new methods to handle such structure's geometrical and textural complexities, notably in the presence of pathologies such as Dilated Right Ventricle, Tricuspid Regurgitation, Arrhythmogenesis, Tetralogy of Fallot, and Inter-atrial Communication. The last MICCAI challenge on right ventricle segmentation was held in 2012 and included only 48 cases from a single clinical center. As part of the 12th Workshop on Statistical Atlases and Computational Models of the Heart (STACOM 2021), the M&Ms-2 challenge was organized to promote the interest of the research community around right ventricle segmentation in multi-disease, multi-view, and multi-center cardiac MRI. Three hundred sixty CMR cases, including short-axis and long-axis 4-chamber views, were collected from three Spanish hospitals using nine different scanners from three different vendors, and included a diverse set of right and left ventricle pathologies. The solutions provided by the participants show that nnU-Net achieved the best results overall. However, multi-view approaches were able to capture additional information, highlighting the need to integrate multiple cardiac diseases, views, scanners, and acquisition protocols to produce reliable automatic cardiac segmentation algorithms.
Carlos Martín-Isla, Víctor M. Campello, Cristian Izquierdo, Kaisar Kushibar, Carla Sendra-Balcells, Polyxeni Gkontra, Alireza Sojoudi, Mitchell J. Fulton, Tewodros Weldebirhan Arega, Kumaradevan Punithakumar, Lei Li 0020, Xiaowu Sun, Yasmina Alkhalil, Di Liu 0003, Sana Jabbar, Sandro F. Queiros, Francesco Galati, Moona Mazher, Zheyao Gao, Marcel Beetz, Lennart Tautz, Christoforos Galazis, Marta Varela, Markus Hüllebrand, Vicente Grau, Xiahai Zhuang, Domenec Puig, Maria A. Zuluaga, Hassan Mohy-ud-Din, Dimitris N. Metaxas, Marcel Breeuwer, Rob J. van der Geest, Michelle Noga, Stéphanie Bricq, Mark Rentschler, Andrea Guala 0002, Steffen E. Petersen, Sergio Escalera, Jose Rodriguez-Palomares, Karim Lekadir
IEEE J. Biomed. Health Informatics30
2022 A Manifold View of Adversarial Risk
abstract
The adversarial risk of a machine learning model has been widely studied. Most previous works assume that the data lies in the whole ambient space. We propose to take a new angle and take the manifold assumption into consideration. Assuming data lies in a manifold, we investigate two new types of adversarial risk, the normal adversarial risk due to perturbation along normal direction, and the in-manifold adversarial risk due to perturbation within the manifold. We prove that the classic adversarial risk can be bounded from both sides using the normal and in-manifold adversarial risks. We also show with a surprisingly pessimistic case that the standard adversarial risk can be nonzero even when both normal and in-manifold risks are zero. We finalize the paper with empirical studies supporting our theoretical results. Our results suggest the possibility of improving the robustness of a classifier by only focusing on the normal adversarial risk.
Yikai Zhang 0003, Xiaoling Hu 0002, Mayank Goswami 0001, Chao Chen 0012, Dimitris N. Metaxas
AISTATS6
2022 Understanding ASL Learners' Preferences for a Sign Language Recording and Automatic Feedback System to Support Self-Study
abstract
Advancements in AI will soon enable tools for providing automatic feedback to American Sign Language (ASL) learners on some aspects of their signing, but there is a need to understand their preferences for submitting videos and receiving feedback. Ten participants in our study were asked to record a few sentences in ASL using software we designed, and we provided manually curated feedback on one sentence in a manner that simulates the output of a future automatic feedback system. Participants responded to interview questions and a questionnaire eliciting their impressions of the prototype. Our initial findings provide guidance to future designers of automatic feedback systems for ASL learners.
Saad Hassan, Sooyeon Lee, Dimitris N. Metaxas, Carol Neidle, Matt Huenerfauth
ASSETS3
2022 Show Me What and Tell Me How: Video Synthesis via Multimodal Conditioning
abstract
Most methods for conditional video synthesis use a single modality as the condition. This comes with major limitations. For example, it is problematic for a model conditioned on an image to generate a specific motion trajectory desired by the user since there is no means to provide motion information. Conversely, language information can describe the desired motion, while not precisely defining the content of the video. This work presents a multimodal video generation framework that benefits from text and images provided jointly or separately. We leverage the recent progress in quantized representations for videos and apply a bidirectional transformer with multiple modalities as inputs to predict a discrete video representation. To improve video quality and consistency, we propose a new video token trained with self-learning and an improved mask-prediction algorithm for sampling video tokens. We introduce text augmentation to improve the robustness of the textual representation and diversity of generated videos. Our framework can incorporate various visual modalities, such as segmentation masks, drawings, and partially occluded images. It can generate much longer sequences than the one used for training. In addition, our model can extract visual information as suggested by the text prompt, e.g., “an object in image one is moving northeast”, and generate corresponding videos. We run evaluations on three public datasets and a newly collected dataset labeled with facial attributes, achieving state-of-the-art generation results on all four11Code: https://github.com/snap-research/MMVID and Webpage..
Ligong Han, Jian Ren 0005, Hsin-Ying Lee 0001, Francesco Barbieri, Kyle Olszewski, Shervin Minaee, Dimitris N. Metaxas, Sergey Tulyakov
CVPR7
2022 Global Matching with Overlapping Attention for Optical Flow Estimation
abstract
Optical flow estimation is a fundamental task in computer vision. Recent direct-regression methods using deep neural networks achieve remarkable performance improvement. However, they do not explicitly capture long-term motion correspondences and thus cannot handle large motions effectively. In this paper, inspired by the traditional matching-optimization methods where matching is introduced to handle large displacements before energy-based optimizations, we introduce a simple but effective global matching step before the direct regression and develop a learning-based matching-optimization framework, namely GMFlowNet. In GMFlowNet, global matching is efficiently calculated by applying argmax on 4D cost volumes. Additionally, to improve the matching quality, we propose patch-based overlapping attention to extract large context features. Extensive experiments demonstrate that GM-FlowNet outperforms RAFT, the most popular optimization-only method, by a large margin and achieves state-of-the-art performance on standard benchmarks. Thanks to the matching and overlapping attention, GMFlowNet obtains major improvements on the predictions for textureless regions and large motions. Our code is made publicly available at https://github.com/xiaofeng94/GMFlowNet.
Shiyu Zhao 0001, Long Zhao 0003, Enyu Zhou, Dimitris N. Metaxas
CVPR5
2022 Hierarchically Self-supervised Transformer for Human Skeleton Representation Learning
Yuxiao Chen 0002, Long Zhao 0003, Yu Tian 0003, Zhaoyang Xia, Shijie Geng, Ligong Han, Dimitris N. Metaxas
ECCV (26)8
2022 Social ODE: Multi-agent Trajectory Forecasting with Neural Ordinary Differential Equations
Song Wen 0001, Hao Wang 0014, Dimitris N. Metaxas
ECCV (22)3
2022 Exploiting Unlabeled Data with Vision and Language Models for Object Detection
Shiyu Zhao 0001, Samuel Schulter, Long Zhao 0003, Anastasis Stathopoulos, Manmohan Krishna Chandraker, Dimitris N. Metaxas
ECCV (9)8
2022 Learning Transferable Reward for Query Object Localization with Policy Adaptation
Tingfeng Li, Shaobo Han, Martin Renqiang Min, Dimitris N. Metaxas
ICLR4
2022 Bidirectional Skeleton-Based Isolated Sign Recognition using Graph Convolutional Networks
abstract
To improve computer-based recognition from video of isolated signs from American Sign Language (ASL), we propose a new skeleton-based method that involves explicit detection of the start and end frames of signs, trained on the ASLLVD dataset; it uses linguistically relevant parameters based on the skeleton input. Our method employs a bidirectional learning approach within a Graph Convolutional Network (GCN) framework. We apply this method to the WLASL dataset, but with corrections to the gloss labeling to ensure consistency in the labels assigned to different signs; it is important to have a 1-1 correspondence between signs and text-based gloss labels. We achieve a success rate of 77.43% for top-1 and 94.54% for top-5 using this modified WLASL dataset. Our method, which does not require multi-modal data input, outperforms other state-of-the-art approaches on the same modified WLASL dataset, demonstrating the importance of both attention to the start and end frames of signs and the use of bidirectional data streams in the GCNs for isolated sign recognition.
Konstantinos M. Dafnis, Evgenia Chroni, Carol Neidle, Dimitris N. Metaxas
LREC4
2022 DeepRecon: Joint 2D Cardiac Segmentation and 3D Volume Reconstruction via a Structure-Specific Generative Method
Zhennan Yan, Mu Zhou, Di Liu 0003, Khalid Sawalha, Meng Ye 0003, Qilong Zhangli, Mikael Kanski, Subhi Al'Aref, Leon Axel, Dimitris N. Metaxas
MICCAI (4)11
2022 TransFusion: Multi-view Divergent Fusion for Medical Image Segmentation with Transformers
Di Liu 0003, Yunhe Gao, Qilong Zhangli, Ligong Han, Xiaoxiao He, Zhaoyang Xia, Song Wen 0001, Zhennan Yan, Mu Zhou, Dimitris N. Metaxas
MICCAI (5)11
2022 Region Proposal Rectification Towards Robust Instance Segmentation of Biological Images
Qilong Zhangli, Jingru Yi, Di Liu 0003, Xiaoxiao He, Zhaoyang Xia, Ligong Han, Yunhe Gao, Song Wen 0001, Haiming Tang, He Wang 0016, Mu Zhou, Dimitris N. Metaxas
MICCAI (4)13
2022 AE-StyleGAN: Improved Training of Style-Based Auto-Encoders
abstract
StyleGANs have shown impressive results on data generation and manipulation in recent years, thanks to its disentangled style latent space. A lot of efforts have been made in inverting a pretrained generator, where an encoder is trained ad hoc after the generator is trained in a two-stage fashion. In this paper, we focus on style-based generators asking a scientific question: Does forcing such a generator to reconstruct real data lead to more disentangled latent space and make the inversion process from image to latent space easy? We describe a new methodology to train a style-based autoencoder where the encoder and generator are optimized end-to-end. We show that our proposed model consistently outperforms baselines in terms of image inversion and generation quality. Supplementary, code, and pretrained models are available on the project website1.
Ligong Han, Sri Harsha Musunuri, Martin Renqiang Min, Ruijiang Gao, Yu Tian 0003, Dimitris N. Metaxas
WACV6
2022 DigestPath: A benchmark dataset with challenge review for the pathological detection and segmentation of digestive-system
Qian Da, Zhongyu Li 0002, Yanfei Zuo, Chenbin Zhang, Jingxin Liu 0005, Wen Chen 0001, Jiahui Li 0005, Dou Xu, Hongmei Yi, Zhe Wang 0043, Li Zhang 0040, Xianying He, Xiaofan Zhang 0002, Ke Mei, Chuang Zhu, Weizeng Lu, LinLin Shen, Jun Shi 0006, Jun Li 0106, Sreehari S, Ganapathy Krishnamurthi, Jiangcheng Yang, Tiancheng Lin 0001, Qingyu Song 0004, Xuechen Liu 0004, Simon Graham, Raja Muhammad Saad Bashir, Canqian Yang, Shaofei Qin, Xinmei Tian 0001, Jie Zhao 0014, Dimitris N. Metaxas, Hongsheng Li 0001, Chaofu Wang, Shaoting Zhang 0001
Medical Image Anal.37
2022 WORD: A large scale dataset, benchmark and clinical applicable study for abdominal organ segmentation from CT image
Xiangde Luo, Wenjun Liao, Jianghong Xiao, Jieneng Chen, Tao Song 0002, Xiaofan Zhang 0002, Kang Li 0004, Dimitris N. Metaxas, Guotai Wang, Shaoting Zhang 0001
Medical Image Anal.8
2022 SCPM-Net: An anchor-free 3D lung nodule detection network using sphere representation and center points matching
Xiangde Luo, Tao Song 0002, Guotai Wang, Jieneng Chen, Kang Li 0004, Dimitris N. Metaxas, Shaoting Zhang 0001
Medical Image Anal.7
2022 Semi-supervised medical image segmentation via uncertainty rectified pyramid consistency
Xiangde Luo, Guotai Wang, Wenjun Liao, Jieneng Chen, Tao Song 0002, Shichuan Zhang, Dimitris N. Metaxas, Shaoting Zhang 0001
Medical Image Anal.8
2022 In vitro machine learning-based CAR T immunological synapse quality measurements correlate with patient clinical outcomes
abstract
The human immune system consists of a highly intelligent network of billions of independent, self-organized cells that interact with each other. Machine learning (ML) is an artificial intelligence (AI) tool that automatically processes huge amounts of image data. Immunotherapies have revolutionized the treatment of blood cancer. Specifically, one such therapy involves engineering immune cells to express chimeric antigen receptors (CAR), which combine tumor antigen specificity with immune cell activation in a single receptor. To improve their efficacy and expand their applicability to solid tumors, scientists optimize different CARs with different modifications. However, predicting and ranking the efficacy of different "off-the-shelf" immune products (e.g., CAR or Bispecific T-cell Engager [BiTE]) and selection of clinical responders are challenging in clinical practice. Meanwhile, identifying the optimal CAR construct for a researcher to further develop a potential clinical application is limited by the current, time-consuming, costly, and labor-intensive conventional tools used to evaluate efficacy. Particularly, more than 30 years of immunological synapse (IS) research data demonstrate that T cell efficacy is not only controlled by the specificity and avidity of the tumor antigen and T cell interaction, but also it depends on a collective process, involving multiple adhesion and regulatory molecules, as well as tumor microenvironment, spatially and temporally organized at the IS formed by cytotoxic T lymphocytes (CTL) and natural killer (NK) cells. The optimal function of cytotoxic lymphocytes (including CTL and NK) depends on IS quality. Recognizing the inadequacy of conventional tools and the importance of IS in immune cell functions, we investigate a new strategy for assessing CAR-T efficacy by quantifying CAR IS quality using the glass-support planar lipid bilayer system combined with ML-based data analysis. Previous studies in our group show that CAR-T IS quality correlates with antitumor activities in vitro and in vivo. However, current manually quantified IS quality data analysis is time-consuming and labor-intensive with low accuracy, reproducibility, and repeatability. In this study, we develop a novel ML-based method to quantify thousands of CAR cell IS images with enhanced accuracy and speed. Specifically, we used artificial neural networks (ANN) to incorporate object detection into segmentation. The proposed ANN model extracts the most useful information to differentiate different IS datasets. The network output is flexible and produces bounding boxes, instance segmentation, contour outlines (borders), intensities of the borders, and segmentations without borders. Based on requirements, one or a combination of this information is used in statistical analysis. The ML-based automated algorithm quantified CAR-T IS data correlates with the clinical responder and non-responder treated with Kappa-CAR-T cells directly from patients. The results suggest that CAR cell IS quality can be used as a potential composite biomarker and correlates with antitumor activities in patients, which is sufficiently discriminative to further test the CAR IS quality as a clinical biomarker to predict response to CAR immunotherapy in cancer. For translational research, the method developed here can also provide guidelines for designing and optimizing numerous CAR constructs for potential clinical development. Trial Registration: ClinicalTrials.gov NCT00881920.
Alireza Naghizadeh, Wei-chung Tsao, Jong Hyun Cho, Hongye Xu, Mohab Mohamed, Dali Li, Dimitris N. Metaxas, Carlos A. Ramos 0002, Dongfang Liu
PLoS Comput. Biol.8
2022 Semi-Supervised Segmentation of Radiation-Induced Pulmonary Fibrosis From Lung CT Scans With Multi-Scale Guided Dense Attention
abstract
Computed Tomography (CT) plays an important role in monitoring radiation-induced Pulmonary Fibrosis (PF), where accurate segmentation of the PF lesions is highly desired for diagnosis and treatment follow-up. However, the task is challenged by ambiguous boundary, irregular shape, various position and size of the lesions, as well as the difficulty in acquiring a large set of annotated volumetric images for training. To overcome these problems, we propose a novel convolutional neural network called PF-Net and incorporate it into a semi-supervised learning framework based on Iterative Confidence-based Refinement And Weighting of pseudo Labels (I-CRAWL). Our PF-Net combines 2D and 3D convolutions to deal with CT volumes with large inter-slice spacing, and uses multi-scale guided dense attention to segment complex PF lesions. For semi-supervised learning, our I-CRAWL employs pixel-level uncertainty-based confidence-aware refinement to improve the accuracy of pseudo labels of unannotated images, and uses image-level uncertainty for confidence-based image weighting to suppress low-quality pseudo labels in an iterative training process. Extensive experiments with CT scans of Rhesus Macaques with radiation-induced PF showed that: 1) PF-Net achieved higher segmentation accuracy than existing 2D, 3D and 2.5D neural networks, and 2) I-CRAWL outperformed state-of-the-art semi-supervised learning methods for the PF lesion segmentation task. Our method has a potential to improve the diagnosis of PF and clinical assessment of side effects of radiotherapy for lung cancers.
Guotai Wang, Shuwei Zhai, Giovanni Lasio, Baoshe Zhang, Byong Yi, Shifeng Chen, Thomas J. Macvittie, Dimitris N. Metaxas, Jinghao Zhou, Shaoting Zhang 0001
IEEE Trans. Medical Imaging8
2021 American Sign Language Video Anonymization to Support Online Participation of Deaf and Hard of Hearing Users
abstract
Without a commonly accepted writing system for American Sign Language (ASL), Deaf or Hard of Hearing (DHH) ASL signers who wish to express opinions or ask questions online must post a video of their signing, if they prefer not to use written English, a language in which they may feel less proficient. Since the face conveys essential linguistic meaning, the face cannot simply be removed from the video in order to preserve anonymity. Thus, DHH ASL signers cannot easily discuss sensitive, personal, or controversial topics in their primary language, limiting engagement in online debate or inquiries about health or legal issues. We explored several recent attempts to address this problem through development of “face swap” technologies to automatically disguise the face in videos while preserving essential facial expressions and natural human appearance. We presented several prototypes to DHH ASL signers (N=16) and examined their interests in and requirements for such technology. After viewing transformed videos of other signers and of themselves, participants evaluated the understandability, naturalness of appearance, and degree of anonymity protection of these technologies. Our study revealed users’ perception of key trade-offs among these three dimensions, factors that contribute to each, and their views on transformation options enabled by this technology, for use in various contexts. Our findings guide future designers of this technology and inform selection of applications and design features.
Sooyeon Lee, Abraham Glasser, Becca Dingman, Zhaoyang Xia, Dimitris N. Metaxas, Carol Neidle, Matt Huenerfauth
ASSETS5
2021 Learning View-Disentangled Human Pose Representation by Contrastive Cross-View Mutual Information Maximization
abstract
We introduce a novel representation learning method to disentangle pose-dependent as well as view-dependent factors from 2D human poses. The method trains a network using cross-view mutual information maximization (CV-MIM) which maximizes mutual information of the same pose performed from different viewpoints in a contrastive learning manner. We further propose two regularization terms to ensure disentanglement and smoothness of the learned representations. The resulting pose representations can be used for cross-view action recognition.To evaluate the power of the learned representations, in addition to the conventional fully-supervised action recognition settings, we introduce a novel task called single-shot cross-view action recognition. This task trains models with actions from only one single viewpoint while models are evaluated on poses captured from all possible viewpoints. We evaluate the learned representations on standard benchmarks for action recognition, and show that (i) CV-MIM performs competitively compared with the state-of-the-art models in the fully-supervised scenarios; (ii) CV-MIM outperforms other competing methods by a large margin in the single-shot cross-view setting; (iii) and the learned representations can significantly boost the performance when reducing the amount of supervised training data. Our code is made publicly available at https://github.com/google-research/google-research/tree/master/poem.
Long Zhao 0003, Yuxiao Wang 0001, Jiaping Zhao, Liangzhe Yuan, Jennifer J. Sun, Florian Schroff, Hartwig Adam, Xi Peng 0005, Dimitris N. Metaxas, Ting Liu 0005
CVPR9
2021 Deep Animation Video Interpolation in the Wild
abstract
In the animation industry, cartoon videos are usually produced at low frame rate since hand drawing of such frames is costly and time-consuming. Therefore, it is desirable to develop computational models that can automatically interpolate the in-between animation frames. However, existing video interpolation methods fail to produce satisfying results on animation data. Compared to natural videos, animation videos possess two unique characteristics that make frame interpolation difficult: 1) cartoons comprise lines and smooth color pieces. The smooth areas lack textures and make it difficult to estimate accurate motions on animation videos. 2) cartoons express stories via exaggeration. Some of the motions are non-linear and extremely large. In this work, we formally define and study the animation video interpolation problem for the first time. To address the aforementioned challenges, we propose an effective framework, AnimeInterp, with two dedicated modules in a coarse-to-fine manner. Specifically, 1) Segment-Guided Matching resolves the "lack of textures" challenge by exploiting global matching among color pieces that are piece-wise coherent. 2) Recurrent Flow Refinement resolves the "non-linear and extremely large motion" challenge by recur-rent predictions using a transformer-like architecture. To facilitate comprehensive training and evaluations, we build a large-scale animation triplet dataset, ATD-12K, which comprises 12,000 triplets with rich annotations. Extensive experiments demonstrate that our approach outperforms existing state-of-the-art interpolation methods for animation videos. Notably, AnimeInterp shows favorable perceptual quality and robustness for animation scenarios in the wild. The proposed dataset and code are available at https://github.com/lisiyao21/AnimeInterp/.
Li Siyao, Shiyu Zhao 0001, Weijiang Yu, Wenxiu Sun, Dimitris N. Metaxas, Chen Change Loy, Ziwei Liu 0002
CVPR5
2021 DeepTag: An Unsupervised Deep Learning Method for Motion Tracking on Cardiac Tagging Magnetic Resonance Images
abstract
Cardiac tagging magnetic resonance imaging (t-MRI) is the gold standard for regional myocardium deformation and cardiac strain estimation. However, this technique has not been widely used in clinical diagnosis, as a result of the difficulty of motion tracking encountered with t-MRI images. In this paper, we propose a novel deep learning-based fully unsupervised method for in vivo motion tracking on t-MRI images. We first estimate the motion field (INF) between any two consecutive t-MRI frames by a bi-directional generative diffeomorphic registration neural network. Using this result, we then estimate the Lagrangian motion field between the reference frame and any other frame through a differentiable composition layer. By utilizing temporal information to perform reasonable estimations on spatiotemporal motion fields, this novel method provides a useful solution for motion tracking and image registration in dynamic medical imaging. Our method has been validated on a representative clinical t-MRI dataset; the experimental results show that our method is superior to conventional motion tracking methods in terms of landmark tracking accuracy and inference efficiency. Project page is at: https://github.com/DeepTag/cardiac_tagging_motion_estimation.
Meng Ye 0003, Mikael Kanski, Dong Yang 0005, Zhennan Yan, Qiaoying Huang, Leon Axel, Dimitris N. Metaxas
CVPR8
2021 CrossNorm and SelfNorm for Generalization under Distribution Shifts
abstract
Traditional normalization techniques (e.g., Batch Normalization and Instance Normalization) generally and simplistically assume that training and test data follow the same distribution. As distribution shifts are inevitable in real-world applications, well-trained models with previous normalization methods can perform badly in new environments. Can we develop new normalization methods to improve generalization robustness under distribution shifts? In this paper, we answer the question by proposing Cross-Norm and SelfNorm. CrossNorm exchanges channel-wise mean and variance between feature maps to enlarge training distribution, while SelfNorm uses attention to recalibrate the statistics to bridge gaps between training and test distributions. CrossNorm and SelfNorm can complement each other, though exploring different directions in statistics usage. Extensive experiments on different fields (vision and language), tasks (classification and segmentation), settings (supervised and semi-supervised), and distribution shift types (synthetic and natural) show the effectiveness. Code is available at https://github.com/amazon-research/crossnorm-selfnorm
Zhiqiang Tang 0001, Yunhe Gao, Yi Zhu 0001, Zhi Zhang 0005, Mu Li 0003, Dimitris N. Metaxas
ICCV6
2021 Dual Projection Generative Adversarial Networks for Conditional Image Generation
abstract
Conditional Generative Adversarial Networks (cGANs) extend the standard unconditional GAN framework to learning joint data-label distributions from samples, and have been established as powerful generative models capable of generating high-fidelity imagery. A challenge of training such a model lies in properly infusing class information into its generator and discriminator. For the discriminator, class conditioning can be achieved by either (1) directly incorporating labels as input or (2) involving labels in an auxiliary classification loss. In this paper, we show that the former directly aligns the class-conditioned fake-and-real data distributions P (image|class) (data matching), while the latter aligns data-conditioned class distributions P (class|image) (label matching). Although class separability does not directly translate to sample quality and becomes a burden if classification itself is intrinsically difficult, the discriminator cannot provide useful guidance for the generator if features of distinct classes are mapped to the same point and thus become inseparable. Motivated by this intuition, we propose a Dual Projection GAN (P2GAN) model that learns to balance between data matching and label matching. We then propose an improved cGAN model with Auxiliary Classification that directly aligns the fake and real conditionals P (class|image) by minimizing their f-divergence. Experiments on a synthetic Mixture of Gaussian (MoG) dataset and a variety of real-world datasets including CIFAR100, ImageNet, and VGGFace2 demonstrate the efficacy of our proposed models.
Ligong Han, Martin Renqiang Min, Anastasis Stathopoulos, Yu Tian 0003, Ruijiang Gao, Asim Kadav, Dimitris N. Metaxas
ICCV7
2021 Semantic Aware Data Augmentation for Cell Nuclei Microscopical Images with Artificial Neural Networks
abstract
There exists many powerful architectures for object detection and semantic segmentation of both biomedical and natural images. However, a difficulty arises in the ability to create training datasets that are large and well-varied. The importance of this subject is nested in the amount of training data that artificial neural networks need to accurately identify and segment objects in images and the infeasibility of acquiring a sufficient dataset within the biomedical field. This paper introduces a new data augmentation method that generates artificial cell nuclei microscopical images along with their correct semantic segmentation labels. Data augmentation provides a step toward accessing higher generalization capabilities of artificial neural networks. An initial set of segmentation objects is used with Greedy AutoAugment to find the strongest performing augmentation policies. The found policies and the initial set of segmentation objects are then used in the creation of the final artificial images. When comparing the state-of-the-art data augmentation methods with the proposed method, the proposed method is shown to consistently outperform current solutions in the generation of nuclei microscopical images.
Alireza Naghizadeh, Hongye Xu, Mohab Mohamed, Dimitris N. Metaxas, Dongfang Liu
ICCV4
2021 Stochastic Transformer Networks with Linear Competing Units: Application to end-to-end SL Translation
abstract
Automating sign language translation (SLT) is a challenging real-world application. Despite its societal importance, though, research progress in the field remains rather poor. Crucially, existing methods that yield viable performance necessitate the availability of laborious to obtain gloss sequence groundtruth. In this paper, we attenuate this need, by introducing an end-to-end SLT model that does not entail explicit use of glosses; the model only needs text groundtruth. This is in stark contrast to existing end-to-end models that use gloss sequence groundtruth, either in the form of a modality that is recognized at an intermediate model stage, or in the form of a parallel output process, jointly trained with the SLT model. Our approach constitutes a Transformer network with a novel type of layers that combines: (i) local winner-takes-all (LWTA) layers with stochastic winner sampling, instead of conventional ReLU layers, (ii) stochastic weights with posterior distributions estimated via variational inference, and (iii) a weight compression technique at inference time that exploits estimated posterior variance to perform massive, almost lossless compression. We demonstrate that our approach can reach the currently best reported BLEU-4 score on the PHOENIX 2014T benchmark, but without making use of glosses for model training, and with a memory footprint reduced by more than 70%.
Andreas Voskou, Konstantinos P. Panousis, Dimitrios I. Kosmopoulos, Dimitris N. Metaxas, Sotirios Chatzis
ICCV4
2021 A Good Image Generator Is What You Need for High-Resolution Video Synthesis
Yu Tian 0003, Jian Ren 0005, Menglei Chai, Kyle Olszewski, Xi Peng 0005, Dimitris N. Metaxas, Sergey Tulyakov
ICLR6
2021 UTNet: A Hybrid Transformer Architecture for Medical Image Segmentation
Yunhe Gao, Mu Zhou, Dimitris N. Metaxas
MICCAI (3)3
2021 Hybrid Supervision Learning for Pathology Whole Slide Image Classification
Jiahui Li 0005, Wen Chen 0022, Qi Duan, Dimitris N. Metaxas, Hongsheng Li 0001, Shaoting Zhang 0001
MICCAI (8)7
2021 Improved Transformer for High-Resolution GANs
abstract
Attention-based models, exemplified by the Transformer, can effectively model long range dependency, but suffer from the quadratic complexity of self-attention operation, making them difficult to be adopted for high-resolution image generation based on Generative Adversarial Networks (GANs). In this paper, we introduce two key ingredients to Transformer to address this challenge. First, in low-resolution stages of the generative process, standard global self-attention is replaced with the proposed multi-axis blocked self-attention which allows efficient mixing of local and global attention. Second, in high-resolution stages, we drop self-attention while only keeping multi-layer perceptrons reminiscent of the implicit neural function. To further improve the performance, we introduce an additional self-modulation component based on cross-attention. The resulting model, denoted as HiT, has a nearly linear computational complexity with respect to the image size and thus directly scales to synthesizing high definition images. We show in the experiments that the proposed HiT achieves state-of-the-art FID scores of 30.83 and 2.95 on unconditional ImageNet $128 \times 128$ and FFHQ $256 \times 256$, respectively, with a reasonable throughput. We believe the proposed HiT is an important milestone for generators in GANs which are completely free of convolutions. Our code is made publicly available at https://github.com/google-research/hit-gan.
Long Zhao 0003, Ting Chen 0001, Dimitris N. Metaxas, Han Zhang 0010
NeurIPS4
2021 Oriented Object Detection in Aerial Images with Box Boundary-Aware Vectors
abstract
Oriented object detection in aerial images is a challenging task as the objects in aerial images are displayed in arbitrary directions and are usually densely packed. Cur-rent oriented object detection methods mainly rely on two-stage anchor-based detectors. However, the anchor-based detectors typically suffer from a severe imbalance issue be-tween the positive and negative anchor boxes. To address this issue, in this work we extend the horizontal keypoint-based object detector to the oriented object detection task. In particular, we first detect the center keypoints of the objects, based on which we then regress the box boundary-aware vectors (BBAVectors) to capture the oriented bounding boxes. The box boundary-aware vectors are distributed in the four quadrants of a Cartesian coordinate system for all arbitrarily oriented objects. To relieve the difficulty of learning the vectors in the corner cases, we further classify the oriented bounding boxes into horizontal and rotational bounding boxes. In the experiment, we show that learning the box boundary-aware vectors is superior to directly predicting the width, height, and angle of an oriented bounding box, as adopted in the baseline method. Besides, the proposed method competes favorably with state-of-the-art methods. Code is available at https://github.com/yijingru/BBAVectors-Oriented-Object-Detection.
Jingru Yi, Pengxiang Wu, Bo Liu 0005, Qiaoying Huang, Dimitris N. Metaxas
WACV6
2021 Guest editorial: Deep learning for medical image analysis
Hongsheng Li 0001, Shaoting Zhang 0001, Dimitris N. Metaxas
Neurocomputing3
2021 FocusNetv2: Imbalanced large and small organ segmentation with adversarial shape constraint for head and neck CT images
Yunhe Gao, Rui Huang 0001, Yiwei Yang 0001, Kainan Shao, Changjuan Tao, Yuanyuan Chen 0007, Dimitris N. Metaxas, Hongsheng Li 0001, Ming Chen 0030
Medical Image Anal.8
2021 Dynamic MRI reconstruction with end-to-end motion-guided network
Qiaoying Huang, Yikun Xian, Dong Yang 0005, Jingru Yi, Pengxiang Wu, Dimitris N. Metaxas
Medical Image Anal.7
2021 Surgical planning of pelvic tumor using multi-view CNN with relation-context representation learning
abstract
Limb salvage surgery of malignant pelvic tumors is the most challenging procedure in musculoskeletal oncology due to the complex anatomy of the pelvic bones and soft tissues. It is crucial to accurately resect the pelvic tumors with appropriate margins in this procedure. However, there is still a lack of efficient and repetitive image planning methods for tumor identification and segmentation in many hospitals. In this paper, we present a novel deep learning-based method to accurately segment pelvic bone tumors in MRI. Our method uses a multi-view fusion network to extract pseudo-3D information from two scans in different directions and improves the feature representation by learning a relational context. In this way, it can fully utilize spatial information in thick MRI scans and reduce over-fitting when learning from a small dataset. Our proposed method was evaluated on two independent datasets collected from 90 and 15 patients, respectively. The segmentation accuracy of our method was superior to several comparing methods and comparable to the expert annotation, while the average time consumed decreased about 100 times from 1820.3 seconds to 19.2 seconds. In addition, we incorporate our method into an efficient workflow to improve the surgical planning process. Our workflow took only 15 minutes to complete surgical planning in a phantom study, which is a dramatic acceleration compared with the 2-day time span in a traditional workflow.
Zhennan Yan, Liang Zhao 0018, Lichi Zhang, Shuaining Xie, Kang Li 0004, Dimitris N. Metaxas, Yongqiang Hao, Kerong Dai, Shaoting Zhang 0001, Xiaofeng Tao 0002, Songtao Ai
Medical Image Anal.9
2021 Greedy auto-augmentation for n-shot learning using deep neural networks
Alireza Naghizadeh, Dimitris N. Metaxas, Dongfang Liu
Neural Networks2
2021 Multi-Stage Feature Fusion Network for Video Super-Resolution
abstract
Video super-resolution (VSR) is to restore a photo-realistic high-resolution (HR) frame from both its corresponding low-resolution (LR) frame (reference frame) and multiple neighboring frames (supporting frames). An important step in VSR is to fuse the feature of the reference frame with the features of the supporting frames. The major issue with existing VSR methods is that the fusion is conducted in a one-stage manner, and the fused feature may deviate greatly from the visual information in the original LR reference frame. In this paper, we propose an end-to-end Multi-Stage Feature Fusion Network that fuses the temporally aligned features of the supporting frames and the spatial feature of the original reference frame at different stages of a feed-forward neural network architecture. In our network, the Temporal Alignment Branch is designed as an inter-frame temporal alignment module used to mitigate the misalignment between the supporting frames and the reference frame. Specifically, we apply the multi-scale dilated deformable convolution as the basic operation to generate temporally aligned features of the supporting frames. Afterwards, the Modulative Feature Fusion Branch, the other branch of our network accepts the temporally aligned feature map as a conditional input and modulates the feature of the reference frame at different stages of the branch backbone. This enables the feature of the reference frame to be referenced at each stage of the feature fusion process, leading to an enhanced feature from LR to HR. Experimental results on several benchmark datasets demonstrate that our proposed method can achieve state-of-the-art performance on VSR task.
Huihui Song 0003, Dong Liu 0002, Bo Liu 0005, Qingshan Liu 0001, Dimitris N. Metaxas
IEEE Trans. Image Process.6
2021 Few-Shot Learning by a Cascaded Framework With Shape-Constrained Pseudo Label Assessment for Whole Heart Segmentation
abstract
Automatic and accurate 3D cardiac image segmentation plays a crucial role in cardiac disease diagnosis and treatment. Even though CNN based techniques have achieved great success in medical image segmentation, the expensive annotation, large memory consumption, and insufficient generalization ability still pose challenges to their application in clinical practice, especially in the case of 3D segmentation from high-resolution and large-dimension volumetric imaging. In this paper, we propose a few-shot learning framework by combining ideas of semi-supervised learning and self-training for whole heart segmentation and achieve promising accuracy with a Dice score of 0.890 and a Hausdorff distance of 18.539 mm with only four labeled data for training. When more labeled data provided, the model can generalize better across institutions. The key to success lies in the selection and evolution of high-quality pseudo labels in cascaded learning. A shape-constrained network is built to assess the quality of pseudo labels, and the self-training stages with alternative global-local perspectives are employed to improve the pseudo labels. We evaluate our method on the CTA dataset of the MM-WHS 2017 Challenge and a larger multi-center dataset. In the experiments, our method outperforms the state-of-the-art methods significantly and has great generalization ability on the unseen data. We also demonstrate, by a study of two 4D (3D+T) CTA data, the potential of our method to be applied in clinical practice.
Wenji Wang, Qing Xia 0002, Zhennan Yan, Zhuowei Li 0002, Yue Gao 0002, Dimitris N. Metaxas, Shaoting Zhang 0001
IEEE Trans. Medical Imaging9
2021 Object-Guided Instance Segmentation With Auxiliary Feature Refinement for Biological Images
abstract
Instance segmentation is of great importance for many biological applications, such as study of neural cell interactions, plant phenotyping, and quantitatively measuring how cells react to drug treatment. In this paper, we propose a novel box-based instance segmentation method. Box-based instance segmentation methods capture objects via bounding boxes and then perform individual segmentation within each bounding box region. However, existing methods can hardly differentiate the target from its neighboring objects within the same bounding box region due to their similar textures and low-contrast boundaries. To deal with this problem, in this paper, we propose an object-guided instance segmentation method. Our method first detects the center points of the objects, from which the bounding box parameters are then predicted. To perform segmentation, an object-guided coarse-to-fine segmentation branch is built along with the detection branch. The segmentation branch reuses the object features as guidance to separate target object from the neighboring ones within the same bounding box region. To further improve the segmentation quality, we design an auxiliary feature refinement module that densely samples and refines point-wise features in the boundary regions. Experimental results on three biological image datasets demonstrate the advantages of our method. The code will be available at https://github.com/yijingru/ObjGuided-Instance-Segmentation.
Jingru Yi, Pengxiang Wu, Bo Liu 0005, Qiaoying Huang, Lianyi Han, Wei Fan 0001, Daniel J. Hoeppner, Dimitris N. Metaxas
IEEE Trans. Medical Imaging10
2020 Robust Conditional GAN from Uncertainty-Aware Pairwise Comparisons
abstract
Conditional generative adversarial networks have shown exceptional generation performance over the past few years. However, they require large numbers of annotations. To address this problem, we propose a novel generative adversarial network utilizing weak supervision in the form of pairwise comparisons (PC-GAN) for image attribute editing. In the light of Bayesian uncertainty estimation and noise-tolerant adversarial training, PC-GAN can estimate attribute rating efficiently and demonstrate robust performance in noise resistance. Through extensive experiments, we show both qualitatively and quantitatively that PC-GAN performs comparably with fully-supervised methods and outperforms unsupervised baselines. Code and Supplementary can be found on the project website*.
Ligong Han, Ruijiang Gao, Mun Kim, Bo Liu 0005, Dimitris N. Metaxas
AAAI6
2020 Object-Guided Instance Segmentation for Biological Images
abstract
Instance segmentation of biological images is essential for studying object behaviors and properties. The challenges, such as clustering, occlusion, and adhesion problems of the objects, make instance segmentation a non-trivial task. Current box-free instance segmentation methods typically rely on local pixel-level information. Due to a lack of global object view, these methods are prone to over- or under-segmentation. On the contrary, the box-based instance segmentation methods incorporate object detection into the segmentation, performing better in identifying the individual instances. In this paper, we propose a new box-based instance segmentation method. Mainly, we locate the object bounding boxes from their center points. The object features are subsequently reused in the segmentation branch as a guide to separate the clustered instances within an RoI patch. Along with the instance normalization, the model is able to recover the target object distribution and suppress the distribution of neighboring attached objects. Consequently, the proposed model performs excellently in segmenting the clustered objects while retaining the target object details. The proposed method achieves state-of-the-art performances on three biological datasets: cell nuclei, plant phenotyping dataset, and neural cells.
Jingru Yi, Pengxiang Wu, Bo Liu 0005, Daniel J. Hoeppner, Dimitris N. Metaxas, Lianyi Han, Wei Fan 0001
AAAI6
2020 Local Regularizer Improves Generalization
abstract
Regularization plays an important role in generalization of deep learning. In this paper, we study the generalization power of an unbiased regularizor for training algorithms in deep learning. We focus on training methods called Locally Regularized Stochastic Gradient Descent (LRSGD). An LRSGD leverages a proximal type penalty in gradient descent steps to regularize SGD in training. We show that by carefully choosing relevant parameters, LRSGD generalizes better than SGD. Our thorough theoretical analysis is supported by experimental evidence. It advances our theoretical understanding of deep learning and provides new perspectives on designing training algorithms. The code is available at https://github.com/huiqu18/LRSGD.
Yikai Zhang 0003, Dimitris N. Metaxas, Chao Chen 0012
AAAI3
2020 Deep Learning based NAS Score and Fibrosis Stage Prediction from CT and Pathology Data
abstract
Non-Alcoholic Fatty Liver Disease (NAFLD) is becoming increasingly prevalent in the world population. Without diagnosis at the right time, NAFLD can lead to non-alcoholic steatohepatitis (NASH) and subsequent liver damage. The diagnosis and treatment of NAFLD depend on the NAFLD activity score (NAS) and the liver fibrosis stage, which are usually evaluated from liver biopsies by pathologists. In this work, we propose a novel method to automatically predict NAS score and fibrosis stage from CT data that is non-invasive and inexpensive to obtain compared with liver biopsy. We also present a method to combine the information from CT and H&E stained pathology data to improve the performance of NAS score and fibrosis stage prediction, when both types of data are available. This is of great value to assist the pathologists in computer-aided diagnosis process. Experiments on a 30-patient dataset illustrate the effectiveness of our method.
Ananya Jana, Puru Rattan, Carlos D. Minacapelli, Vinod Rustgi, Dimitris N. Metaxas
BIBE6
2020 Knowledge As Priors: Cross-Modal Knowledge Generalization for Datasets Without Superior Knowledge
abstract
Cross-modal knowledge distillation deals with transferring knowledge from a model trained with superior modalities (Teacher) to another model trained with weak modalities (Student). Existing approaches require paired training examples exist in both modalities. However, accessing the data from superior modalities may not always be feasible. For example, in the case of 3D hand pose estimation, depth maps, point clouds, or stereo images usually capture better hand structures than RGB images, but most of them are expensive to be collected. In this paper, we propose a novel scheme to train the Student in a Target dataset where the Teacher is unavailable. Our key idea is to generalize the distilled cross-modal knowledge learned from a Source dataset, which contains paired examples from both modalities, to the Target dataset by modeling knowledge as priors on parameters of the Student. We name our method "Cross-Modal Knowledge Generalization" and demonstrate that our scheme results in competitive performance for 3D hand pose estimation on standard benchmark datasets.
Long Zhao 0003, Xi Peng 0005, Yuxiao Chen 0002, Mubbasir Kapadia, Dimitris N. Metaxas
CVPR5
2020 Synthetic Learning: Learn From Distributed Asynchronized Discriminator GAN Without Sharing Medical Image Data
abstract
In this paper, we propose a data privacy-preserving and communication efficient distributed GAN learning framework named Distributed Asynchronized Discriminator GAN (AsynDGAN). Our proposed framework aims to train a central generator learns from distributed discriminator, and use the generated synthetic image solely to train the segmentation model. We validate the proposed framework on the application of health entities learning problem which is known to be privacy sensitive. Our experiments show that our approach: 1) could learn the real image’s distribution from multiple datasets without sharing the patient’s raw data. 2) is more efficient and requires lower bandwidth than other distributed deep learning methods. 3) achieves higher performance compared to the model trained by one real dataset, and almost the same performance compared to the model trained by all real datasets. 4) has provable guarantees that the generator could learn the distributed distribution in an all important fashion thus is unbiased.We release our AsynDGAN source code at: https://github.com/tommy-qichang/AsynDGAN
Yikai Zhang 0003, Mert R. Sabuncu, Chao Chen 0012, Tong Zhang 0001, Dimitris N. Metaxas
CVPR7
2020 MotionNet: Joint Perception and Motion Prediction for Autonomous Driving Based on Bird's Eye View Maps
abstract
The ability to reliably perceive the environmental states, particularly the existence of objects and their motion behavior, is crucial for autonomous driving. In this work, we propose an efficient deep model, called MotionNet, to jointly perform perception and motion prediction from 3D point clouds. MotionNet takes a sequence of LiDAR sweeps as input and outputs a bird's eye view (BEV) map, which encodes the object category and motion information in each grid cell. The backbone of MotionNet is a novel spatio-temporal pyramid network, which extracts deep spatial and temporal features in a hierarchical fashion. To enforce the smoothness of predictions over both space and time, the training of MotionNet is further regularized with novel spatial and temporal consistency losses. Extensive experiments show that the proposed method overall outperforms the state-of-the-arts, including the latest scene-flow- and 3D-object-detection-based methods. This indicates the potential value of the proposed method serving as a backup to the bounding-box-based system, and providing complementary information to the motion planner in autonomous driving. Code is available at https://www.merl.com/research/license#MotionNet.
Pengxiang Wu, Siheng Chen, Dimitris N. Metaxas
CVPR3
2020 OnlineAugment: Online Data Augmentation with Less Domain Knowledge
Zhiqiang Tang 0001, Yunhe Gao, Leonid Karlinsky, Prasanna Sattigeri, Rogério Feris, Dimitris N. Metaxas
ECCV (7)6
2020 Learn Distributed GAN with Temporary Discriminators
Yikai Zhang 0003, Zhennan Yan, Chao Chen 0012, Dimitris N. Metaxas
ECCV (27)6
2020 Learning Trailer Moments in Full-Length Movies with Co-Contrastive Attention
Lezi Wang, Dong Liu 0002, Rohit Puri, Dimitris N. Metaxas
ECCV (18)4
2020 Error-Bounded Correction of Noisy Labels
abstract
To collect large scale annotated data, it is inevitable to introduce label noise, i.e., incorrect class labels. To be robust against label noise, many successful methods rely on the noisy classifiers (i.e., models trained on the noisy training data) to determine whether a label is trustworthy. However, it remains unknown why this heuristic works well in practice. In this paper, we provide the first theoretical explanation for these methods. We prove that the prediction of a noisy classifier can indeed be a good indicator of whether the label of a training data is clean. Based on the theoretical result, we propose a novel algorithm that corrects the labels based on the noisy classifier prediction. The corrected labels are consistent with the true Bayesian optimal classifier with high probability. We incorporate our label correction algorithm into the training of deep neural networks and train models that achieve superior testing performance on multiple public datasets.
Songzhu Zheng, Pengxiang Wu, Aman Goswami, Mayank Goswami 0001, Dimitris N. Metaxas, Chao Chen 0012
ICML5
2020 Condensed Silhouette: An Optimized Filtering Process for Cluster Selection in K-Means
abstract
In K-Means based clustering algorithms, different initial seeds can lead to different clustering results. Selecting the best result from different initial seeds is called the filtering process. The filtering process follows three steps, 1- performing several clustering trials, 2- scoring each trial, and 3- choosing the trial with the best score. A typical method to score the clustering results of K-Means is the within-cluster sum of squares (WCSS). There are more advanced methods that can be used to score the clustering trials. These methods usually provide a better score with the cost of being more computationally demanding. In this paper, we propose Condensed Silhouette, which is a very efficient version of the Silhouette algorithm. For this purpose, we replace the elements of Silhouette algorithm with similar elements of the K-Means algorithm. This helps us to maintain the accuracy of the Silhouette and at the same time, significantly reduce the computational requirements of the method. Our experiments on 14 real datasets show the effectiveness of the proposed method.
Alireza Naghizadeh, Dimitris N. Metaxas
KES2
2020 Maximum-Entropy Adversarial Data Augmentation for Improved Generalization and Robustness
abstract
Adversarial data augmentation has shown promise for training robust deep neural networks against unforeseen data shifts or corruptions. However, it is difficult to define heuristics to generate effective fictitious target distributions containing "hard" adversarial perturbations that are largely different from the source distribution. In this paper, we propose a novel and effective regularization term for adversarial data augmentation. We theoretically derive it from the information bottleneck principle, which results in a maximum-entropy formulation. Intuitively, this regularization term encourages perturbing the underlying source distribution to enlarge predictive uncertainty of the current model, so that the generated "hard" adversarial perturbations can improve the model robustness during training. Experimental results on three standard benchmarks demonstrate that our method consistently outperforms the existing state of the art by a statistically significant margin.
Long Zhao 0003, Ting Liu 0005, Xi Peng 0005, Dimitris N. Metaxas
NeurIPS4
2020 Deep Subspace Clustering with Data Augmentation
abstract
The idea behind data augmentation techniques is based on the fact that slight changes in the percept do not change the brain cognition. In classification, neural networks use this fact by applying transformations to the inputs to learn to predict the same label. However, in deep subspace clustering (DSC), the ground-truth labels are not available, and as a result, one cannot easily use data augmentation techniques. We propose a technique to exploit the benefits of data augmentation in DSC algorithms. We learn representations that have consistent subspaces for slightly transformed inputs. In particular, we introduce a temporal ensembling component to the objective function of DSC algorithms to enable the DSC networks to maintain consistent subspaces for random transformations in the input data. In addition, we provide a simple yet effective unsupervised procedure to find efficient data augmentation policies. An augmentation policy is defined as an image processing transformation with a certain magnitude and probability of being applied to each image in each epoch. We search through the policies in a search space of the most common augmentation policies to find the best policy such that the DSC network yields the highest mean Silhouette coefficient in its clustering results on a target dataset. Our method achieves state-of-the-art performance on four standard subspace clustering datasets.
Mahdi Abavisani, Alireza Naghizadeh, Dimitris N. Metaxas, Vishal M. Patel
NeurIPS3
2020 A Topological Filter for Learning with Label Noise
abstract
Noisy labels can impair the performance of deep neural networks. To tackle this problem, in this paper, we propose a new method for filtering label noise. Unlike most existing methods relying on the posterior probability of a noisy classifier, we focus on the much richer spatial behavior of data in the latent representational space. By leveraging the high-order topological information of data, we are able to collect most of the clean data and train a high-quality model. Theoretically we prove that this topological approach is guaranteed to collect the clean data with high probability. Empirical results show that our method outperforms the state-of-the-arts and is robust to a broad spectrum of noise types and levels.
Pengxiang Wu, Songzhu Zheng, Mayank Goswami 0001, Dimitris N. Metaxas, Chao Chen 0012
NeurIPS4
2020 GNM: GridCell navigational model
Alireza Naghizadeh, Samaneh Berenjian, David J. Margolis, Dimitris N. Metaxas
Expert Syst. Appl.4
2020 Towards Image-to-Video Translation: A Structure-Aware Approach via Multi-stage Generative Adversarial Networks
Long Zhao 0003, Xi Peng 0005, Yu Tian 0003, Mubbasir Kapadia, Dimitris N. Metaxas
Int. J. Comput. Vis.5
2020 Dual Iterative Hard Thresholding
abstract
Iterative Hard Thresholding (IHT) is a popular class of first-order greedy selection methods for loss minimization under cardinality constraint. The existing IHT-style algorithms, however, are proposed for minimizing the primal formulation. It is still an open issue to explore duality theory and algorithms for such a non-convex and NP-hard combinatorial optimization problem. To address this issue, we develop in this article a novel duality theory for $\ell_2$-regularized empirical risk minimization under cardinality constraint, along with an IHT-style algorithm for dual optimization. Our sparse duality theory establishes a set of sufficient and/or necessary conditions under which the original non-convex problem can be equivalently or approximately solved in a concave dual formulation. In view of this theory, we propose the Dual IHT (DIHT) algorithm as a super-gradient ascent method to solve the non-smooth dual problem with provable guarantees on primal-dual gap convergence and sparsity recovery. Numerical results confirm our theoretical predictions and demonstrate the superiority of DIHT to the state-of-the-art primal IHT-style algorithms in model estimation accuracy and computational efficiency.
Xiao-Tong Yuan, Bo Liu 0005, Lezi Wang, Qingshan Liu 0001, Dimitris N. Metaxas
J. Mach. Learn. Res.5
2020 Towards Efficient U-Nets: A Coupled and Quantized Approach
abstract
In this paper, we propose to couple stacked U-Nets for efficient visual landmark localization. The key idea is to globally reuse features of the same semantic meanings across the stacked U-Nets. The feature reuse makes each U-Net light-weighted. Specially, we propose an order- K coupling design to trim off long-distance shortcuts, together with an iterative refinement and memory sharing mechanism. To further improve the efficiency, we quantize the parameters, intermediate features, and gradients of the coupled U-Nets to low bit-width numbers. We validate our approach in two tasks: human pose estimation and facial landmark localization. The results show that our approach achieves state-of-the-art localization accuracy but using ∼ 70% fewer parameters, ∼ 30% less inference time, ∼ 98% less model size, and saving ∼ 75% training memory compared with benchmark localizers.
Zhiqiang Tang 0001, Xi Peng 0005, Kang Li 0004, Dimitris N. Metaxas
IEEE Trans. Pattern Anal. Mach. Intell.4
2020 Greedy AutoAugment
Alireza Naghizadeh, Mohammadsajad Abavisani, Dimitris N. Metaxas
Pattern Recognit. Lett.3
2020 Weakly Supervised Deep Nuclei Segmentation Using Partial Points Annotation in Histopathology Images
abstract
Nuclei segmentation is a fundamental task in histopathology image analysis. Typically, such segmentation tasks require significant effort to manually generate accurate pixel-wise annotations for fully supervised training. To alleviate such tedious and manual effort, in this paper we propose a novel weakly supervised segmentation framework based on partial points annotation, i.e., only a small portion of nuclei locations in each image are labeled. The framework consists of two learning stages. In the first stage, we design a semi-supervised strategy to learn a detection model from partially labeled nuclei locations. Specifically, an extended Gaussian mask is designed to train an initial model with partially labeled data. Then, self-training with background propagation is proposed to make use of the unlabeled regions to boost nuclei detection and suppress false positives. In the second stage, a segmentation model is trained from the detected nuclei locations in a weakly-supervised fashion. Two types of coarse labels with complementary information are derived from the detected points and are then utilized to train a deep neural network. The fully-connected conditional random field loss is utilized in training to further refine the model without introducing extra computational complexity during inference. The proposed method is extensively evaluated on two nuclei segmentation datasets. The experimental results demonstrate that our method can achieve competitive performance compared to the fully supervised counterpart and the state-of-the-art methods while requiring significantly less annotation effort.
Pengxiang Wu, Qiaoying Huang, Jingru Yi, Zhennan Yan, Kang Li 0004, Gregory M. Riedlinger, Subhajyoti De, Shaoting Zhang 0001, Dimitris N. Metaxas
IEEE Trans. Medical Imaging10
2019 Point Cloud Processing via Recurrent Set Encoding
abstract
We present a new permutation-invariant network for 3D point cloud processing. Our network is composed of a recurrent set encoder and a convolutional feature aggregator. Given an unordered point set, the encoder firstly partitions its ambient space into parallel beams. Points within each beam are then modeled as a sequence and encoded into subregional geometric features by a shared recurrent neural network (RNN). The spatial layout of the beams is regular, and this allows the beam features to be further fed into an efficient 2D convolutional neural network (CNN) for hierarchical feature aggregation. Our network is effective at spatial feature learning, and competes favorably with the state-of-the-arts (SOTAs) on a number of benchmarks. Meanwhile, it is significantly more efficient compared to the SOTAs.
Pengxiang Wu, Chao Chen 0012, Jingru Yi, Dimitris N. Metaxas
AAAI4
2019 Distributed Inexact Newton-type Pursuit for Non-convex Sparse Learning
abstract
In this paper, we present a sample distributed greedy pursuit method for non-convex sparse learning under cardinality constraint. Given the training samples uniformly randomly partitioned across multiple machines, the proposed method alternates between local inexact sparse minimization of a Newton-type approximation and centralized global results aggregation. Theoretical analysis shows that for a general class of convex functions with Lipschitze continues Hessian, the method converges linearly with contraction factor scaling inversely to the local data size; whilst the communication complexity required to reach desirable statistical accuracy scales logarithmically with respect to the number of machines for some popular statistical learning models. For nonconvex objective functions, up to a local estimation error, our method can be shown to converge to a local stationary sparse solution with sub-linear communication complexity. Numerical results demonstrate the efficiency and accuracy of our method when applied to large-scale sparse learning tasks including deep neural nets pruning
Bo Liu 0005, Xiao-Tong Yuan, Lezi Wang, Qingshan Liu 0001, Junzhou Huang, Dimitris N. Metaxas
AISTATS6
2019 Construct Dynamic Graphs for Hand Gesture Recognition via Spatial-Temporal Attention
Yuxiao Chen 0002, Long Zhao 0003, Xi Peng 0005, Dimitris N. Metaxas
BMVC5
2019 Attention-based Facial Behavior Analytics inSocial Communication
Lezi Wang, Chongyang Bai, Maksim Bolonkin, Judee K. Burgoon, Norah E. Dunbar, V. S. Subrahmanian, Dimitris N. Metaxas
BMVC7
2019 Semantic Graph Convolutional Networks for 3D Human Pose Regression
abstract
In this paper, we study the problem of learning Graph Convolutional Networks (GCNs) for regression. Current architectures of GCNs are limited to the small receptive field of convolution filters and shared transformation matrix for each node. To address these limitations, we propose Semantic Graph Convolutional Networks (SemGCN), a novel neural network architecture that operates on regression tasks with graph-structured data. SemGCN learns to capture semantic information such as local and global node relationships, which is not explicitly represented in the graph. These semantic relationships can be learned through end-to-end training from the ground truth without additional supervision or hand-crafted rules. We further investigate applying SemGCN to 3D human pose regression. Our formulation is intuitive and sufficient since both 2D and 3D human poses can be represented as a structured graph encoding the relationships between joints in the skeleton of a human body. We carry out comprehensive studies to validate our method. The results prove that SemGCN outperforms state of the art while using 90% fewer parameters.
Long Zhao 0003, Xi Peng 0005, Yu Tian 0003, Mubbasir Kapadia, Dimitris N. Metaxas
CVPR5
2019 AdaTransform: Adaptive Data Transformation
abstract
Data augmentation is widely used to increase data variance in training deep neural networks. However, previous methods require either comprehensive domain knowledge or high computational cost. Can we learn data transformation automatically and efficiently with limited domain knowledge? Furthermore, can we leverage data transformation to improve not only network training but also network testing? In this work, we propose adaptive data transformation to achieve the two goals. The AdaTransform can increase data variance in training and decrease data variance in testing. Experiments on different tasks prove that it can improve generalization performance.
Zhiqiang Tang 0001, Xi Peng 0005, Tingfeng Li, Yizhe Zhu, Dimitris N. Metaxas
ICCV5
2019 Sharpen Focus: Learning With Attention Separability and Consistency
abstract
Recent developments in gradient-based attention modeling have seen attention maps emerge as a powerful tool for interpreting convolutional neural networks. Despite good localization for an individual class of interest, these techniques produce attention maps with substantially overlapping responses among different classes, leading to the problem of visual confusion and the need for discriminative attention. In this paper, we address this problem by means of a new framework that makes class-discriminative attention a principled part of the learning process. Our key innovations include new learning objectives for attention separability and cross-layer consistency, which result in improved attention discriminability and reduced visual confusion. Extensive experiments on image classification benchmarks show the effectiveness of our approach in terms of improved classification accuracy, including CIFAR-100 (+3.33%), Caltech-256 (+1.64%), ImageNet (+0.92%), CUB-200-2011 (+4.8%) and PASCAL VOC2012 (+5.73%).
Lezi Wang, Ziyan Wu 0001, Srikrishna Karanam, Kuan-Chuan Peng, Rajat Vikram Singh, Bo Liu 0005, Dimitris N. Metaxas
ICCV7
2019 Self-Attention Generative Adversarial Networks
abstract
In this paper, we propose the Self-Attention Generative Adversarial Network (SAGAN) which allows attention-driven, long-range dependency modeling for image generation tasks. Traditional convolutional GANs generate high-resolution details as a function of only spatially local points in lower-resolution feature maps. In SAGAN, details can be generated using cues from all feature locations. Moreover, the discriminator can check that highly detailed features in distant portions of the image are consistent with each other. Furthermore, recent work has shown that generator conditioning affects GAN performance. Leveraging this insight, we apply spectral normalization to the GAN generator and find that this improves training dynamics. The proposed SAGAN performs better than prior work, boosting the best published Inception score from 36.8 to 52.52 and reducing Fréchet Inception distance from 27.62 to 18.65 on the challenging ImageNet dataset. Visualization of the attention layers shows that the generator leverages neighborhoods that correspond to object shapes rather than local regions of fixed shape.
Han Zhang 0010, Ian J. Goodfellow, Dimitris N. Metaxas, Augustus Odena
ICML3
2019 Taming the Noisy Gradient: Train Deep Neural Networks with Small Batch Sizes
abstract
Deep learning architectures are usually proposed with millions of parameters, resulting in a memory issue when training deep neural networks with stochastic gradient descent type methods using large batch sizes. However, training with small batch sizes tends to produce low quality solution due to the large variance of stochastic gradients. In this paper, we tackle this problem by proposing a new framework for training deep neural network with small batches/noisy gradient. During optimization, our method iteratively applies a proximal type regularizer to make loss function strongly convex. Such regularizer stablizes the gradient, leading to better training performance. We prove that our algorithm achieves comparable convergence rate as vanilla SGD even with small batch size. Our framework is simple to implement and can be potentially combined with many existing optimization algorithms. Empirical results show that our method outperforms SGD and Adam when batch size is small. Our implementation is available at https://github.com/huiqu18/TRAlgorithm.
Yikai Zhang 0003, Chao Chen 0012, Dimitris N. Metaxas
IJCAI4
2019 Heuristic Search for Homology Localization Problem and Its Application in Cardiac Trabeculae Reconstruction
abstract
Cardiac trabeculae are fine rod-like muscles whose ends are attached to the inner walls of ventricles. Accurate extraction of trabeculae is important yet challenging, due to the background noise and limited resolution of cardiac images. Existing works proposed to handle this task by modeling the trabeculae as topological handles for better extraction. Computing optimal representation of these handles is essential yet very expensive. In this work, we formulate the problem as a heuristic search problem, and propose novel heuristic functions based on advanced topological techniques. We show in experiments that the proposed heuristic functions improve the computation in both time and memory.
Xudong Zhang 0004, Pengxiang Wu, Changhe Yuan, Yusu Wang 0001, Dimitris N. Metaxas, Chao Chen 0012
IJCAI5
2019 Brain Segmentation from k-Space with End-to-End Recurrent Attention Network
Qiaoying Huang, Xiao Chen 0013, Dimitris N. Metaxas, Mariappan S. Nadar
MICCAI (3)3
2019 Improving Nuclei/Gland Instance Segmentation in Histopathology Images by Full Resolution Neural Network and Spatial Constrained Loss
Zhennan Yan, Gregory M. Riedlinger, Subhajyoti De, Dimitris N. Metaxas
MICCAI (1)5
2019 Collaborative Multi-agent Learning for MR Knee Articular Cartilage Segmentation
Chaowei Tan, Zhennan Yan, Shaoting Zhang 0001, Kang Li 0004, Dimitris N. Metaxas
MICCAI (2)5
2019 Multi-scale Cell Instance Segmentation with Keypoint Graph Based Bounding Boxes
Jingru Yi, Pengxiang Wu, Qiaoying Huang, Bo Liu 0005, Daniel J. Hoeppner, Dimitris N. Metaxas
MICCAI (1)7
2019 Rethinking Kernel Methods for Node Representation Learning on Graphs
abstract
Graph kernels are kernel methods measuring graph similarity and serve as a standard tool for graph classification. However, the use of kernel methods for node classification, which is a related problem to graph representation learning, is still ill-posed and the state-of-the-art methods are heavily based on heuristics. Here, we present a novel theoretical kernel-based framework for node classification that can bridge the gap between these two representation learning problems on graphs. Our approach is motivated by graph kernel methodology but extended to learn the node representations capturing the structural information in a graph. We theoretically show that our formulation is as powerful as any positive semidefinite kernels. To efficiently learn the kernel, we propose a novel mechanism for node feature aggregation and a data-driven similarity metric employed during the training phase. More importantly, our framework is flexible and complementary to other graph-based deep learning models, e.g., Graph Convolutional Networks (GCNs). We empirically evaluate our approach on a number of standard node classification benchmarks, and demonstrate that our model sets the new state of the art.
Yu Tian 0003, Long Zhao 0003, Xi Peng 0005, Dimitris N. Metaxas
NeurIPS4
2019 Cartoonish sketch-based face editing in videos using identity deformation transfer
Long Zhao 0003, Fangda Han, Xi Peng 0005, Mubbasir Kapadia, Vladimir Pavlovic 0001, Dimitris N. Metaxas
Comput. Graph.7
2019 ASSD: Attentive single shot multibox detector
Jingru Yi, Pengxiang Wu, Dimitris N. Metaxas
Comput. Vis. Image Underst.3
2019 A coupled encoder-decoder network for joint face detection and landmark localization
Lezi Wang, Xiang Yu 0002, Thirimachos Bourlai, Dimitris N. Metaxas
Image Vis. Comput.4
2019 Attentive neural cell instance segmentation
Jingru Yi, Pengxiang Wu, Menglin Jiang, Qiaoying Huang, Daniel J. Hoeppner, Dimitris N. Metaxas
Medical Image Anal.6
2019 StackGAN++: Realistic Image Synthesis with Stacked Generative Adversarial Networks
abstract
Although Generative Adversarial Networks (GANs) have shown remarkable success in various tasks, they still face challenges in generating high quality images. In this paper, we propose Stacked Generative Adversarial Networks (StackGANs) aimed at generating high-resolution photo-realistic images. First, we propose a two-stage generative adversarial network architecture, StackGAN-v1, for text-to-image synthesis. The Stage-I GAN sketches the primitive shape and colors of a scene based on a given text description, yielding low-resolution images. The Stage-II GAN takes Stage-I results and the text description as inputs, and generates high-resolution images with photo-realistic details. Second, an advanced multi-stage generative adversarial network architecture, StackGAN-v2, is proposed for both conditional and unconditional generative tasks. Our StackGAN-v2 consists of multiple generators and multiple discriminators arranged in a tree-like structure; images at multiple scales corresponding to the same scene are generated from different branches of the tree. StackGAN-v2 shows more stable training behavior than StackGAN-v1 by jointly approximating multiple distributions. Extensive experiments demonstrate that the proposed stacked generative adversarial networks significantly outperform other state-of-the-art methods in generating photo-realistic images.
Han Zhang 0010, Tao Xu 0029, Hongsheng Li 0001, Shaoting Zhang 0001, Xiaogang Wang 0001, Sharon X. Huang, Dimitris N. Metaxas
IEEE Trans. Pattern Anal. Mach. Intell.7
2019 Predicting 3-D Lower Back Joint Load in Lifting: A Deep Pose Estimation Approach
abstract
Goal: Lifting is a common manual material handling task performed in the workplaces. It is considered as one of the main risk factors for work-related musculoskeletal disorders. An important criterion to identify the unsafe lifting task is the values of the net force and moment at L5/S1 joint. These values are mainly calculated in a laboratory environment, which utilizes marker-based sensors to collect three-dimensional (3-D) information and force plates to measure the external forces and moments. However, this method is usually expensive to set up, time-consuming in process, and sensitive to the surrounding environment. In this study, we propose a deep neural network (DNN)-based framework for 3-D pose estimation, which addresses the aforementioned limitations, and we employ the results for L5/S1 moment and force calculation. Methods: At the first step of the proposed framework, full body 3-D pose is captured using a DNN, then at the second step, estimated 3-D body pose along with the subject's anthropometric information is utilized to calculate L5/S1 join's kinetic by a top-down inverse dynamic algorithm. Results: To fully evaluate our approach, we conducted experiments using a lifting dataset consisting of 12 subjects performing various types of lifting tasks. The results are validated against a marker-based motion capture system as a reference. The grand mean ± SD of the total moment/force absolute errors across all the dataset was 9.06 ± 7.60 N·m/4.85 ± 4.85 N. Conclusion: The proposed method provides a reliable tool for assessment of the lower back kinetics during lifting and can be an alternative when the use of marker-based motion capture systems is not possible.
Rahil Mehrizi, Xi Peng 0005, Dimitris N. Metaxas, Shaoting Zhang 0001, Kang Li 0004
IEEE Trans. Hum. Mach. Syst.3
2018 CU-Net: Coupled U-Nets
Zhiqiang Tang 0001, Xi Peng 0005, Shijie Geng, Yizhe Zhu, Dimitris N. Metaxas
BMVC5
2018 Jointly Optimize Data Augmentation and Network Training: Adversarial Data Augmentation in Human Pose Estimation
abstract
Random data augmentation is a critical technique to avoid overfitting in training deep models. Yet, data augmentation and network training are often two isolated processes in most settings, yielding to a suboptimal training. Why not jointly optimize the two? We propose adversarial data augmentation to address this limitation. The key idea is to design a generator (e.g. an augmentation network) that competes against a discriminator (e.g. a target network) by generating hard examples online. The generator explores weaknesses of the discriminator, while the discriminator learns from hard augmentations to achieve better performance. A reward/penalty strategy is also proposed for efficient joint training. We investigate human pose estimation and carry out comprehensive ablation studies to validate our method. The results prove that our method can effectively improve state-of-the-art models without additional data effort.
Xi Peng 0005, Zhiqiang Tang 0001, Fei Yang 0001, Rogério Feris, Dimitris N. Metaxas
CVPR5
2018 Show Me a Story: Towards Coherent Neural Story Illustration
abstract
We propose an end-to-end network for visual illustration of a sequence of sentences forming a story. At the core of our model is the ability to model the inter-related nature of the sentences within a story, as well as the ability to learn coherence to support reference resolution. The framework takes the form of an encoder-decoder architecture, where sentences are encoded using a hierarchical two-level sentence-story GRU, combined with an encoding of coherence, and sequentially decoded using a predicted feature representation into a consistent illustrative image sequence. We optimize all parameters of our network in an end-to-end fashion with respect to order embedding loss, encoding entailment between images and sentences. Experiments on the VIST storytelling dataset [9] highlight the importance of our algorithmic choices and efficacy of our overall model.
Hareesh Ravi, Lezi Wang, Carlos Muñiz 0001, Leonid Sigal, Dimitris N. Metaxas, Mubbasir Kapadia
CVPR5
2018 Quantized Densely Connected U-Nets for Efficient Landmark Localization
Zhiqiang Tang 0001, Xi Peng 0005, Shijie Geng, Lingfei Wu 0001, Shaoting Zhang 0001, Dimitris N. Metaxas
ECCV (3)6
2018 Learning to Forecast and Refine Residual Motion for Image-to-Video Generation
Long Zhao 0003, Xi Peng 0005, Yu Tian 0003, Mubbasir Kapadia, Dimitris N. Metaxas
ECCV (15)5
2018 Toward Marker-Free 3D Pose Estimation in Lifting: A Deep Multi-View Solution
abstract
Lifting is a common manual material handling task performed in the workplaces. It is considered as one of the main risk factors for Work-related Musculoskeletal Disorders. To improve work place safety, it is necessary to assess musculoskeletal and biomechanical risk exposures associated with these tasks, which requires very accurate 3D pose. Existing approaches mainly utilize marker-based sensors to collect 3D information. However, these methods are usually expensive to setup, timeconsuming in process, and sensitive to the surrounding environment. In this study, we propose a multi-view based deep perceptron approach to address aforementioned limitations. Our approach consists of two modules: a "view-specific perceptron" network extracts rich information independently from the image of view, which includes both 2D shape and hierarchical texture information; while a "multi-view integration" network synthesizes information from all available views to predict accurate 3D pose. To fully evaluate our approach, we carried out comprehensive experiments to compare different variants of our design. The results prove that our approach achieves comparable performance with former marker-based methods, i.e. an average error of 14:72 ± 2:96 mm on the lifting dataset. The results are also compared with state-of-the-art methods on HumanEva-I dataset [1], which demonstrates the superior performance of our approach.
Rahil Mehrizi, Xi Peng 0005, Zhiqiang Tang 0001, Dimitris N. Metaxas, Kang Li 0004
FG5
2018 Improving GANs Using Optimal Transport
Tim Salimans, Han Zhang 0010, Alec Radford, Dimitris N. Metaxas
ICLR (Poster)4
2018 CR-GAN: Learning Complete Representations for Multi-view Generation
abstract
Generating multi-view images from a single-view input is an important yet challenging problem. It has broad applications in vision, graphics, and robotics. Our study indicates that the widely-used generative adversarial network (GAN) may learn ?incomplete? representations due to the single-pathway framework: an encoder-decoder network followed by a discriminator network.We propose CR-GAN to address this problem. In addition to the single reconstruction path, we introduce a generation sideway to maintain the completeness of the learned embedding space. The two learning paths collaborate and compete in a parameter-sharing manner, yielding largely improved generality to ?unseen? dataset. More importantly, the two-pathway framework makes it possible to combine both labeled and unlabeled data for self-supervised learning, which further enriches the embedding space for realistic generations. We evaluate our approach on a wide range of datasets. The results prove that CR-GAN significantly outperforms state-of-the-art methods, especially when generating from ?unseen? inputs in wild conditions.
Yu Tian 0003, Xi Peng 0005, Long Zhao 0003, Shaoting Zhang 0001, Dimitris N. Metaxas
IJCAI5
2018 Linguistically-driven Framework for Computationally Efficient and Scalable Sign Recognition
Dimitris N. Metaxas, Mark Dilsizian, Carol Neidle
LREC1
2018 RED-Net: A Recurrent Encoder-Decoder Network for Video-Based Face Alignment
Xi Peng 0005, Rogério Feris, Xiaoyu Wang 0002, Dimitris N. Metaxas
Int. J. Comput. Vis.4
2018 Toward Personalized Modeling: Incremental and Ensemble Alignment for Sequential Faces in the Wild
Xi Peng 0005, Shaoting Zhang 0001, Yang Yu 0010, Dimitris N. Metaxas
Int. J. Comput. Vis.4
2017 Addressing Imbalance in Multi-Label Classification Using Structured Hellinger Forests
abstract
The multi-label classification problem involves finding a model that maps a set of input features to more than one output label. Class imbalance is a serious issue in multi-label classification. We introduce an extension of structured forests, a type of random forest used for structured prediction, called Sparse Oblique Structured Hellinger Forests (SOSHF). We explore using structured forests in the general multi-label setting and propose a new imbalance-aware formulation by altering how the splitting functions are learned in two ways. First, we account for cost-sensitivity when converting the multi-label problem to a single-label problem at each node in the tree. Second, we introduce a new objective function for determining oblique splits based on the Hellinger distance, a splitting criterion that has been shown to be robust to class imbalance. We empirically validate our method on a number of benchmarks against standard and state-of-the-art multi-label classification algorithms with improved results.
Zachary A. Daniels, Dimitris N. Metaxas
AAAI2
2017 Cardiac Trabeculae Segmentation: an Application of Computational Topology (Multimedia Contribution)
abstract
In this video, we present a research project on cardiac trabeculae segmentation. Trabeculae are fine muscle columns within human ventricles whose both ends are attached to the wall. Extracting these structures are very challenging even with state-of-the-art image segmentation techniques. We observed that these structures form natural topological handles. Based on such observation, we developed a topological approach, which employs advanced computational topology methods and achieve high quality segmentation results.
Chao Chen 0012, Dimitris N. Metaxas, Yusu Wang 0001, Pengxiang Wu
SoCG2
2017 Semantic Amodal Segmentation
abstract
Common visual recognition tasks such as classification, object detection, and semantic segmentation are rapidly reaching maturity, and given the recent rate of progress, it is not unreasonable to conjecture that techniques for many of these problems will approach human levels of performance in the next few years. In this paper we look to the future: what is the next frontier in visual recognition? We offer one possible answer to this question. We propose a detailed image annotation that captures information beyond the visible pixels and requires complex reasoning about full scene structure. Specifically, we create an amodal segmentation of each image: the full extent of each region is marked, not just the visible pixels. Annotators outline and name all salient regions in the image and specify a partial depth order. The result is a rich scene structure, including visible and occluded portions of each region, figure-ground edge information, semantic labels, and object overlap. We create two datasets for semantic amodal segmentation. First, we label 500 images in the BSDS dataset with multiple annotators per image, allowing us to study the statistics of human annotations. We show that the proposed full scene annotation is surprisingly consistent between annotators, including for regions and edges. Second, we annotate 5000 images from COCO. This larger dataset allows us to explore a number of algorithmic ideas for amodal segmentation and depth ordering. We introduce novel metrics for these tasks, and along with our strong baselines, define concrete new challenges for the community.
Yan Zhu 0009, Yuandong Tian, Dimitris N. Metaxas, Piotr Dollár
CVPR3
2017 Learning Deep Features for Hierarchical Classification of Mobile Phone Face Datasets in Heterogeneous Environments
abstract
In this paper, we propose a convolutional neural network (CNN) based, scenario-dependent and sensor (mobile device) adaptable hierarchical classification framework. Our proposed framework is designed to automatically categorize face data captured under various challenging conditions, before the FR algorithms (pre-processing, feature extraction and matching) are used. First, a unique multi-sensor database (using Samsung S4 Zoom, Nokia 1020, iPhone 5S and Samsung S5 phones) is collected containing face images indoors, outdoors, with yaw angle from -90° to +90° and at two different distances, i.e. 1 and 10 meters. To cope with pose variations, face detection and pose estimation algorithms are used for classifying the facial images into a frontal or a non-frontal class. Next, our proposed framework is used where tri-level hierarchical classification is performed as follows: Level 1, face images are classified based on phone type; Level 2, face images are further classified into indoor and outdoor images; and finally, Level 3 face images are classified into a close (1m) and a far, low quality, (10m) distance categories respectively. Experimental results show that classification accuracy is scenario dependent, reaching from 95 to more than 98% accuracy for level 2 and from 90 to more than 99% for level 3 classification. A set of experiments is performed indicating that, the usage of data grouping before the face matching is performed, resulted in a significantly improved rank-1 identification rate when compared to the original (all vs. all) biometric system.
Neeru Narang, Michael Martin 0004, Dimitris N. Metaxas, Thirimachos Bourlai
FG3
2017 A Coupled Encoder-Decoder Network for Joint Face Detection and Landmark Localization
abstract
Face detection and landmark localization have been extensively investigated and are the prerequisite for many face applications, such as face recognition and 3D face reconstruction. Most existing methods achieve success on only one of the two problems. In this paper, we propose a coupled encoder-decoder network to jointly detect faces and localize facial key points. The encoder and decoder generate response maps for facial landmark localization. Moreover, we observe that the intermediate feature maps from the encoder and decoder have strong power in describing facial regions, which motivates us to build a unified framework by coupling the feature maps for multi-scale cascaded face detection. Experiments on face detection show strongly competitive results against the existing methods on two public benchmarks. The landmark localization further shows consistently better accuracy than state-of-the-arts on three face-in-the-wild databases.
Lezi Wang, Xiang Yu 0002, Dimitris N. Metaxas
FG3
2017 Reconstruction-Based Disentanglement for Pose-Invariant Face Recognition
abstract
Deep neural networks (DNNs) trained on large-scale datasets have recently achieved impressive improvements in face recognition. But a persistent challenge remains to develop methods capable of handling large pose variations that are relatively under-represented in training data. This paper presents a method for learning a feature representation that is invariant to pose, without requiring extensive pose coverage in training data. We first propose to generate non-frontal views from a single frontal face, in order to increase the diversity of training data while preserving accurate facial details that are critical for identity discrimination. Our next contribution is to seek a rich embedding that encodes identity features, as well as non-identity ones such as pose and landmark locations. Finally, we propose a new feature reconstruction metric learning to explicitly disentangle identity and pose, by demanding alignment between the feature reconstructions through various combinations of identity and pose features, which is obtained from two images of the same subject. Experiments on both controlled and in-the-wild face datasets, such as MultiPIE, 300WLP and the profile view database CFP, show that our method consistently outperforms the state-of-the-art, especially on images with large head pose variations.
Xi Peng 0005, Xiang Yu 0002, Kihyuk Sohn, Dimitris N. Metaxas, Manmohan Krishna Chandraker
ICCV4
2017 Dual Iterative Hard Thresholding: From Non-convex Sparse Minimization to Non-smooth Concave Maximization
abstract
Iterative Hard Thresholding (IHT) is a class of projected gradient descent methods for optimizing sparsity-constrained minimization models, with the best known efficiency and scalability in practice. As far as we know, the existing IHT-style methods are designed for sparse minimization in primal form. It remains open to explore duality theory and algorithms in such a non-convex and NP-hard setting. In this article, we bridge the gap by establishing a duality theory for sparsity-constrained minimization with $\ell_2$-regularized objective and proposing an IHT-style algorithm for dual maximization. Our sparse duality theory provides a set of sufficient and necessary conditions under which the original NP-hard/non-convex problem can be equivalently solved in a dual space. The proposed dual IHT algorithm is a super-gradient method for maximizing the non-smooth dual objective. An interesting finding is that the sparse recovery performance of dual IHT is invariant to the Restricted Isometry Property (RIP), which is required by all the existing primal IHT without sparsity relaxation. Moreover, a stochastic variant of dual IHT is proposed for large-scale stochastic optimization. Numerical results demonstrate that dual IHT algorithms can achieve more accurate model estimation given small number of training data and have higher computational efficiency than the state-of-the-art primal IHT-style algorithms.
Bo Liu 0005, Xiao-Tong Yuan, Lezi Wang, Qingshan Liu 0001, Dimitris N. Metaxas
ICML5
2017 Deep Image-to-Image Recurrent Network with Shape Basis Learning for Automatic Vertebra Labeling in Large-Scale 3D CT Volumes
Dong Yang 0005, Daguang Xu, Shaohua Kevin Zhou, Zhoubing Xu, Mingqing Chen, Jin Hyeong Park, Sasa Grbic, Trac D. Tran, Sang (Peter) Chin, Dimitris N. Metaxas, Dorin Comaniciu
MICCAI (3)11
2017 Automatic Liver Segmentation Using an Adversarial Image-to-Image Network
Dong Yang 0005, Daguang Xu, Shaohua Kevin Zhou, Bogdan Georgescu, Mingqing Chen, Sasa Grbic, Dimitris N. Metaxas, Dorin Comaniciu
MICCAI (3)7
2017 A computationally efficient 3D/2D registration method based on image gradient direction probability density function
Soheil Ghafurian, Ilker Hacihaliloglu, Dimitris N. Metaxas, Virak Tan, Kang Li 0004
Neurocomputing3
2017 Towards large-scale MR thigh image analysis via an integrated quantification framework
Chaowei Tan, Kang Li 0004, Zhennan Yan, Jingru Yi, Pengxiang Wu, Hui Jing Yu, Klaus Engelke, Dimitris N. Metaxas
Neurocomputing8
2017 Scalable Mammogram Retrieval Using Composite Anchor Graph Hashing With Iterative Quantization
abstract
Content-based image retrieval (CBIR) shows great significance in clinical decision-making, which explores the visual content of medical images rather than keywords, tags, or descriptions. It provides doctors an image-guided approach to explore relevant cases that could offer doctors instructive reference. Mammogram screening has been known to be widely used in the early stage diagnosis of breast cancer and could reduce its morbidity and mortality. In this paper, we aim to develop a scalable CBIR method for a large repository of mammogram. To this end, we extend the original Anchor Graph Hashing (AGH) and propose a new unsupervised hashing algorithm, named as composite AGH with iterative quantization (C-AGH-ITQ), which compresses mammographic regions of interest (ROIs) into compact binary codes and enables real-time searching in Hamming space. Multimodal features and different distance metrics are integrated, performing upon a composite Anchor Graph. To improve the effectiveness of the hash code, quantization error is further iteratively minimized by introducing an orthogonal rotation matrix. We evaluate the presented C-AGH-ITQ algorithm on a data set of 11 533 mammographic ROIs obtained from the Digital Database for Screening Mammography. Our method obtains more than 84% retrieval precision and 93% classification accuracy (using$k$NN prediction), which demonstrates that hash codes produced by C-AGH-ITQ well capture the visual similarities between mammographic images. In addition, since C-AGH-ITQ ensures linear complexity of the training procedure and constant time for query, our system is readily applicable to large-scale mammogram databases and has the potential to provide abundant clinical cases as reference.
Jingjing Liu 0001, Shaoting Zhang 0001, Wei Liu 0005, Cheng Deng 0002, Yuanjie Zheng, Dimitris N. Metaxas
IEEE Trans. Circuits Syst. Video Technol.6
2017 Parallel Sparse Subspace Clustering via Joint Sample and Parameter Blockwise Partition
abstract
Sparse subspace clustering (SSC) is a classical method to cluster data with specific subspace structure for each group. It has many desirable theoretical properties and has been shown to be effective in various applications. However, under the condition of a large-scale dataset, learning the sparse sample affinity graph is computationally expensive. To tackle the computation time cost challenge, we develop a memory-efficient parallel framework for computing SSC via an alternating direction method of multiplier (ADMM) algorithm. The proposed framework partitions the data matrix into column blocks and then decomposes the original problem into parallel multivariate Lasso regression subproblems and samplewise operations. The proposed method allows us to allocate multiple cores/machines for the processing of individual column blocks. We propose a stochastic optimization algorithm to minimize the objective function. Experimental results on real-world datasets demonstrate that the proposed blockwise ADMM framework is substantially more efficient than its matrix counterpart used by SSC, without sacrificing performance in applications. Moreover, our approach is directly applicable to parallel neighborhood selection for Gaussian graphical models structure estimation.
Bo Liu 0005, Xiao-Tong Yuan, Yang Yu 0010, Qingshan Liu 0001, Dimitris N. Metaxas
ACM Trans. Embed. Comput. Syst.5
2016 Decentralized Robust Subspace Clustering
abstract
We consider the problem of subspace clustering using the SSC (Sparse Subspace Clustering) approach, which has several desirable theoretical properties and has been shown to be effective in various computer vision applications.We develop a large scale distributed framework for the computation of SSC via an alternating direction method of multiplier (ADMM) algorithm. The proposed framework solves SSC in column blocks and only involves parallel multivariate Lasso regression subproblems and sample-wise operations. This appealing property allows us to allocate multiple cores/machines for the processing of individual column blocks.We evaluate our algorithm on a shared-memory architecture. Experimental results on real-world datasets confirm that the proposed block-wise ADMM framework is substantially more efficient than its matrix counterpart used by SSC,without sacrificing accuracy. Moreover, our approach is directly applicable to decentralized neighborhood selection for Gaussian graphical models structure estimation.
Bo Liu 0005, Xiao-Tong Yuan, Yang Yu 0010, Qingshan Liu 0001, Dimitris N. Metaxas
AAAI5
2016 Multispectral Deep Neural Networks for Pedestrian Detection
Jingjing Liu 0001, Shaoting Zhang 0001, Dimitris N. Metaxas
BMVC4
2016 Track Facial Points in Unconstrained Videos
Xi Peng 0005, Qiong Hu 0001, Junzhou Huang, Dimitris N. Metaxas
BMVC4
2016 SPDA-CNN: Unifying Semantic Part Detection and Abstraction for Fine-Grained Recognition
abstract
Most convolutional neural networks (CNNs) lack midlevel layers that model semantic parts of objects. This limits CNN-based methods from reaching their full potential in detecting and utilizing small semantic parts in recognition. Introducing such mid-level layers can facilitate the extraction of part-specific features which can be utilized for better recognition performance. This is particularly important in the domain of fine-grained recognition. In this paper, we propose a new CNN architecture that integrates semantic part detection and abstraction (SPDACNN) for fine-grained classification. The proposed network has two sub-networks: one for detection and one for recognition. The detection sub-network has a novel top-down proposal method to generate small semantic part candidates for detection. The classification sub-network introduces novel part layers that extract features from parts detected by the detection sub-network, and combine them for recognition. As a result, the proposed architecture provides an end-to-end network that performs detection, localization of multiple semantic parts, and whole object recognition within one framework that shares the computation of convolutional filters. Our method outperforms state-of-theart methods with a large margin for small parts detection (e.g. our precision of 93.40% vs the best previous precision of 74.00% for detecting the head on CUB-2011). It also compares favorably to the existing state-of-the-art on finegrained classification, e.g. it achieves 85.14% accuracy on CUB-2011.
Han Zhang 0010, Tao Xu 0029, Sharon X. Huang, Shaoting Zhang 0001, Ahmed M. Elgammal, Dimitris N. Metaxas
CVPR7
2016 A Recurrent Encoder-Decoder Network for Sequential Face Alignment
Xi Peng 0005, Rogério Feris, Xiaoyu Wang 0002, Dimitris N. Metaxas
ECCV (1)4
2016 Efficient k-Support-Norm Regularized Minimization via Fully Corrective Frank-Wolfe Method
Bo Liu 0005, Xiao-Tong Yuan, Shaoting Zhang 0001, Qingshan Liu 0001, Dimitris N. Metaxas
IJCAI5
2016 Visual Tracking with Reliable Memories
Shaoting Zhang 0001, Wei Liu 0005, Dimitris N. Metaxas
IJCAI4
2016 Nonlinear Hierarchical Part-Based Regression for Unconstrained Face Alignment
Xiang Yu 0002, Zhe Lin 0001, Shaoting Zhang 0001, Dimitris N. Metaxas
IJCAI4
2016 Detection of Major ASL Sign Types in Continuous Signing For ASL Recognition
Polina Yanovich, Carol Neidle, Dimitris N. Metaxas
LREC3
2016 Mammographic Mass Segmentation with Online Learned Shape and Appearance Priors
Menglin Jiang, Shaoting Zhang 0001, Yuanjie Zheng, Dimitris N. Metaxas
MICCAI (2)4
2016 Multimodal Deep Learning for Cervical Dysplasia Diagnosis
Tao Xu 0029, Han Zhang 0010, Sharon X. Huang, Shaoting Zhang 0001, Dimitris N. Metaxas
MICCAI (2)5
2016 People detection in crowded scenes by context-driven label propagation
abstract
Exploiting contextual cues has been a key idea to improve people detection in crowded scenes. Along this line we present a novel context-driven approach to detect people in crowded scenes. Based on a context graph that incorporates both geometric and social contextual patterns in crowds, we apply label propagation to discover weak detections contextually compatible with true detections while suppressing irrelevant false alarms. Compared to previous approaches for context modeling limited to only pairwise spatial interactions between local object neighbors, our approach provides a more effective way to model people interactions in a global context. Our approach achieves performance comparable to state of the art on two challenging datasets for people and pedestrian detection.
Jingjing Liu 0001, Quanfu Fan, Sharath Pankanti, Dimitris N. Metaxas
WACV4
2016 Customized expression recognition for performance-driven cutout character animation
abstract
Performance-driven character animation enables users to create expressive results by performing the desired motion of the character with their face and/or body. However, for cutout animations where continuous motion is combined with discrete artwork replacements, supporting a performance-driven workflow has some unique requirements. To trigger the appropriate artwork replacements, the system must reliably detect a wide range of customized facial expressions that are challenging for existing recognition methods, which focus on a few canonical expressions (e.g., angry, disgusted, scared, happy, sad and surprised). Also, real usage scenarios require the system to work in realtime with minimal training. In this paper, we propose a novel customized expression recognition technique that meets all of these requirements. We first use a set of handcrafted features combining geometric features derived from facial landmarks and patch-based appearance features through group sparsity-based facial component learning. To improve discrimination and generalization, these handcrafted features are integrated into a custom-designed Deep Convolutional Neural Network (CNN) structure trained from publicly available facial expression datasets. The combined features are fed to an online ensemble of SVMs designed for the few training sample problem and performs in realtime. To improve temporal coherence, we also apply a Hidden Markov Model (HMM) to smooth the recognition results. Our system achieves state-of-the-art performance on canonical expression datasets and promising results on our collected dataset of customized expressions.
Xiang Yu 0002, Jianchao Yang, Linjie Luo, Wilmot Li, Jonathan Brandt, Dimitris N. Metaxas
WACV6
2016 Video Classification via Weakly Supervised Sequence Modeling
Jingjing Liu 0001, Chao Chen 0012, Yan Zhu 0009, Wei Liu 0005, Dimitris N. Metaxas
Comput. Vis. Image Underst.5
2016 A detection-driven and sparsity-constrained deformable model for fascia lata labeling and thigh inter-muscular adipose quantification
Chaowei Tan, Kang Li 0004, Zhennan Yan, Dong Yang 0005, Shaoting Zhang 0001, Hui Jing Yu, Klaus Engelke, Colin Miller, Dimitris N. Metaxas
Comput. Vis. Image Underst.9
2016 Scalable histopathological image analysis via supervised hashing with multiple features
Menglin Jiang, Shaoting Zhang 0001, Junzhou Huang, Lin Yang 0002, Dimitris N. Metaxas
Medical Image Anal.5
2016 An efficient conditional random field approach for automatic and interactive neuron segmentation
Mustafa Gökhan Uzunbas, Chao Chen 0012, Dimitris N. Metaxas
Medical Image Anal.3
2016 Large-Scale medical image analytics: Recent methodologies, applications and Future directions
Shaoting Zhang 0001, Dimitris N. Metaxas
Medical Image Anal.2
2016 Face Landmark Fitting via Optimized Part Mixtures and Cascaded Deformable Model
abstract
This paper addresses the problem of facial landmark localization and tracking from a single camera. We present a two-stage cascaded deformable shape model to effectively and efficiently localize facial landmarks with large head pose variations. In initialization stage, we propose a group sparse optimized mixture model to automatically select the most salient facial landmarks. By introducing 3D face shape model, we apply procrustes analysis to provide pose-aware landmark initialization. In landmark localization stage, the first step uses mean-shift local search with constrained local model to rapidly approach the global optimum. The second step uses component-wise active contours to discriminatively refine the subtle shape variation. Our framework simultaneously handles face detection, pose-robust landmark localization and tracking in real time. Extensive experiments are conducted on both laboratory environmental databases and face-in-the-wild databases. The results reveal that our approach consistently outperforms state-of-the-art methods for face alignment and tracking.
Xiang Yu 0002, Junzhou Huang, Shaoting Zhang 0001, Dimitris N. Metaxas
IEEE Trans. Pattern Anal. Mach. Intell.4
2016 Multi-Instance Deep Learning: Discover Discriminative Local Anatomies for Bodypart Recognition
abstract
In general image recognition problems, discriminative information often lies in local image patches. For example, most human identity information exists in the image patches containing human faces. The same situation stays in medical images as well. "Bodypart identity" of a transversal slice-which bodypart the slice comes from-is often indicated by local image information, e.g., a cardiac slice and an aorta arch slice are only differentiated by the mediastinum region. In this work, we design a multi-stage deep learning framework for image classification and apply it on bodypart recognition. Specifically, the proposed framework aims at: 1) discover the local regions that are discriminative and non-informative to the image classification problem, and 2) learn a image-level classifier based on these local regions. We achieve these two tasks by the two stages of learning scheme, respectively. In the pre-train stage, a convolutional neural network (CNN) is learned in a multi-instance learning fashion to extract the most discriminative and and non-informative local patches from the training slices. In the boosting stage, the pre-learned CNN is further boosted by these local patches for image classification. The CNN learned by exploiting the discriminative local appearances becomes more accurate than those learned from global image context. The key hallmark of our method is that it automatically discovers the discriminative and non-informative local patches through multi-instance deep learning. Thus, no manual annotation is required. Our method is validated on a synthetic dataset and a large scale CT dataset. It achieves better performances than state-of-the-art approaches, including the standard deep CNN.
Zhennan Yan, Yiqiang Zhan, Zhigang Peng, Shu Liao, Yoshihisa Shinagawa, Shaoting Zhang 0001, Dimitris N. Metaxas, Xiang Sean Zhou
IEEE Trans. Medical Imaging7
2015 PIEFA: Personalized Incremental and Ensemble Face Alignment
abstract
Face alignment, especially on real-time or large-scale sequential images, is a challenging task with broad applications. Both generic and joint alignment approaches have been proposed with varying degrees of success. However, many generic methods are heavily sensitive to initializations and usually rely on offline-trained static models, which limit their performance on sequential images with extensive variations. On the other hand, joint methods are restricted to offline applications, since they require all frames to conduct batch alignment. To address these limitations, we propose to exploit incremental learning for personalized ensemble alignment. We sample multiple initial shapes to achieve image congealing within one frame, which enables us to incrementally conduct ensemble alignment by group-sparse regularized rank minimization. At the same time, personalized modeling is obtained by subspace adaptation under the same incremental framework, while correction strategy is used to alleviate model drifting. Experimental results on multiple controlled and in-the-wild databases demonstrate the superior performance of our approach compared with state-of-the-arts in terms of fitting accuracy and efficiency.
Xi Peng 0005, Shaoting Zhang 0001, Dimitris N. Metaxas
ICCV4
2015 Joint Kernel-Based Supervised Hashing for Scalable Histopathological Image Analysis
Menglin Jiang, Shaoting Zhang 0001, Junzhou Huang, Lin Yang 0002, Dimitris N. Metaxas
MICCAI (3)5
2015 Multi-layer stencil creation from images
Arjun Jain, Chao Chen 0012, Thorsten Thormählen, Dimitris N. Metaxas, Hans-Peter Seidel
Comput. Graph.4
2015 From circle to 3-sphere: Head pose estimation by instance parameterization
Xi Peng 0005, Junzhou Huang, Qiong Hu 0001, Shaoting Zhang 0001, Ahmed M. Elgammal, Dimitris N. Metaxas
Comput. Vis. Image Underst.6
2015 Automatic stereoscopic video generation based on virtual view synthesis
Lin Zhong 0002, Dimitris N. Metaxas
Neurocomputing3
2015 A homotopy-based sparse representation for fast and accurate shape prior modeling in liver surgical planning
Guotai Wang, Shaoting Zhang 0001, Hongzhi Xie, Dimitris N. Metaxas, Lixu Gu
Medical Image Anal.4
2015 Query Specific Rank Fusion for Image Retrieval
abstract
Recently two lines of image retrieval algorithms demonstrate excellent scalability: 1) local features indexed by a vocabulary tree, and 2) holistic features indexed by compact hashing codes. Although both of them are able to search visually similar images effectively, their retrieval precision may vary dramatically among queries. Therefore, combining these two types of methods is expected to further enhance the retrieval precision. However, the feature characteristics and the algorithmic procedures of these methods are dramatically different, which is very challenging for the feature-level fusion. This motivates us to investigate how to fuse the ordered retrieval sets, i.e., the ranks of images, given by multiple retrieval methods, to boost the retrieval precision without sacrificing their scalability. In this paper, we model retrieval ranks as graphs of candidate images and propose a graph-based query specific fusion approach, where multiple graphs are merged and reranked by conducting a link analysis on a fused graph. The retrieval quality of an individual method is measured on-the-fly by assessing the consistency of the top candidates' nearest neighborhoods. Hence, it is capable of adaptively integrating the strengths of the retrieval methods using local or holistic features for different query images. This proposed method does not need any supervision, has few parameters, and is easy to implement. Extensive and thorough experiments have been conducted on four public datasets, i.e., the UKbench, Corel-5K, Holidays and the large-scale San Francisco Landmarks datasets. Our proposed method has achieved very competitive performance, including state-of-the-art results on several data sets, e.g., the N-S score 3.83 for UKbench.
Shaoting Zhang 0001, Ming Yang 0007, Timothée Cour, Kai Yu 0001, Dimitris N. Metaxas
IEEE Trans. Pattern Anal. Mach. Intell.5
2015 Is Interactional Dissynchrony a Clue to Deception? Insights From Automated Analysis of Nonverbal Visual Cues
abstract
Detecting deception in interpersonal dialog is challenging since deceivers take advantage of the give-and-take of interaction to adapt to any sign of skepticism in an interlocutor's verbal and nonverbal feedback. Human detection accuracy is poor, often with no better than chance performance. In this investigation, we consider whether automated methods can produce better results and if emphasizing the possible disruption in interactional synchrony can signal whether an interactant is truthful or deceptive. We propose a data-driven and unobtrusive framework using visual cues that consists of face tracking, head movement detection, facial expression recognition, and interactional synchrony estimation. Analysis were conducted on 242 video samples from an experiment in which deceivers and truth-tellers interacted with professional interviewers either face-to-face or through computer mediation. Results revealed that the framework is able to automatically track head movements and expressions of both interlocutors to extract normalized meaningful synchrony features and to learn classification models for deception recognition. Further experiments show that these features reliably capture interactional synchrony and efficiently discriminate deception from truth.
Xiang Yu 0002, Shaoting Zhang 0001, Zhennan Yan, Fei Yang 0001, Junzhou Huang, Norah E. Dunbar, Matthew L. Jensen, Judee K. Burgoon, Dimitris N. Metaxas
IEEE Trans. Cybern.9
2015 Learning Multiscale Active Facial Patches for Expression Analysis
abstract
In this paper, we present a new idea to analyze facial expression by exploring some common and specific information among different expressions. Inspired by the observation that only a few facial parts are active in expression disclosure (e.g., around mouth, eye), we try to discover the common and specific patches which are important to discriminate all the expressions and only a particular expression, respectively. A two-stage multitask sparse learning (MTSL) framework is proposed to efficiently locate those discriminative patches. In the first stage MTSL, expression recognition tasks are combined to located common patches. Each of the tasks aims to find dominant patches for each expression. Secondly, two related tasks, facial expression recognition and face verification tasks, are coupled to learn specific facial patches for individual expression. The two-stage patch learning is performed on patches sampled by multiscale strategy. Extensive experiments validate the existence and significance of common and specific patches. Utilizing these learned patches, we achieve superior performances on expression recognition compared to the state-of-the-arts.
Lin Zhong 0002, Qingshan Liu 0001, Peng Yang 0001, Junzhou Huang, Dimitris N. Metaxas
IEEE Trans. Cybern.5
2015 Investigating the Discriminative Power of Keystroke Sound
abstract
The goal of this paper is to determine whether keystroke sound can be used to recognize a user. In this regard, we analyze the discriminative power of keystroke sound in the context of a continuous user authentication application. Motivated by the concept of digraphs used in modeling keystroke dynamics, a virtual alphabet is first learned from keystroke sound segments. Next, the digraph latency within the pairs of virtual letters, along with other statistical features, is used to generate match scores. The resultant scores are indicative of the similarities between two sound streams, and are fused to make a final authentication decision. Experiments on both static text-based and free text-based authentications on a database of 50 subjects demonstrate the potential as well as the limitations of keystroke sound.
Joseph Roth, Xiaoming Liu 0002, Arun Ross, Dimitris N. Metaxas
IEEE Trans. Inf. Forensics Secur.4
2014 Consensus of Regression for Occlusion-Robust Facial Feature Localization
Xiang Yu 0002, Zhe Lin 0001, Jonathan Brandt, Dimitris N. Metaxas
ECCV (4)4
2014 Robust Multi-pose Facial Expression Recognition
abstract
Previous research on facial expression recognition mainly focuses on near frontal face images, while in realistic interactive scenarios, the interested subjects may appear in arbitrary non-frontal poses. In this paper, we propose a framework to recognize six prototypical facial expressions, namely, anger, disgust, fear, joy, sadness and surprise, in an arbitrary head pose. We build a multi-pose training set by rendering 3D face scans from the BU-4DFE dynamic facial expression database [17] at 49 different viewpoints. We extract Local Binary Pattern (LBP) descriptors and further utilize multiple instance learning to mitigate the influence of inaccurate alignment in this challenging task. Experimental results demonstrate the power and validate the effectiveness of the proposed multi-pose facial expression recognition framework.
Qiong Hu 0001, Xi Peng 0005, Peng Yang 0001, Fei Yang 0001, Dimitris N. Metaxas
ICPR5
2014 Head Pose Estimation by Instance Parameterization
abstract
Head pose estimation from images is a challenging task with extensive applications. It has been attracting research attentions over decades and numerous approaches have been proposed. Among them, manifold embedding based methods, which assume that the pose variations lie on a low-dimensional manifold embedded in the high-dimensional feature space, have achieved great success. However, previous manifold embedding based methods have two drawbacks: first, they lack the capability to simultaneously deal with multiple pose-unrelated factors in a uniform way, second, they suffer from limited representation ability for out-of-sample testing inputs. In this paper we propose a novel head pose estimation method to address these problems. By learning the mapping from a uniform geometry representation to individual instance manifolds, this approach allows us to parameterize various pose-unrelated factors under a uniform framework. Our approach is a generative model which guarantees the reasonable and effective representation of new testing input. Besides, by employing eigen instance bases instead of full instance bases arisen from the training data, we can effectively compact the size of the trained model and significantly simplify the computational complexity of the testing process. Experiments on public databases such as CMU-MultiPIE and BU-4DFE, and quantitative comparisons with other state-of-the-art methods show the effectiveness of our approach.
Xi Peng 0005, Junzhou Huang, Qiong Hu 0001, Shaoting Zhang 0001, Dimitris N. Metaxas
ICPR5
2014 An Automated and Robust Framework for Quantification of Muscle and Fat in the Thigh
abstract
The tissue quantification in the thigh (e.g. cross-sectional areas of adipose tissue and muscle) is important, since their quantities reflect adverse metabolic effects and muscle function. Traditional manual analysis is time-consuming and operator-dependent, especially in the case of multi-slices or 3D datasets. In clinical trials, there are a large amount of datasets acquired from magnetic resonance imaging (MRI) or X-ray computed tomography (CT) that requires automatic labeling of individual tissues. Since most segmentation algorithms are not suited for different modalities, we present an automatic and robust framework for the quantitative assessment of muscle and fat tissues on 3D MR or CT data. In our framework, a variational Bayesian Gaussian mixture model is used to cluster regions of interest in images into adipose tissues (fat and marrow), muscle, bone and background. The identification of each cluster is based on marrow detection. Furthermore, we use a combination of parametric and geodesic active contour models to distinguish different adipose tissues in 3D images. To validate our proposed framework, we have conducted preliminary experiments on five volumetric mid-thigh axial datasets of MR and CT images from clinical trials.
Chaowei Tan, Zhennan Yan, Shaoting Zhang 0001, Boubakeur Belaroussi, Hui Jing Yu, Colin Miller, Dimitris N. Metaxas
ICPR7
2014 Automatic Liver Segmentation and Hepatic Fat Fraction Assessment in MRI
abstract
Automated assessment of hepatic fat fraction is clinically important. A robust and precise segmentation would enable accurate, objective and consistent measurement of liver fat fraction for disease quantification, therapy monitoring and drug development. However, segmenting the liver in clinical trials is a challenging task due to the variability of liver anatomy as well as the diverse sources the images were acquired from. In this paper, we propose an automated and robust framework for liver segmentation and assessment. It uses single statistical atlas registration to initialize a robust deformable model to get fine segmentation. Fat fraction map is computed by using chemical shift based method in the delineated region of liver. This proposed method is validated on 14 abdominal magnetic resonance (MR) volumetric scans. The qualitative and quantitative comparisons show that our proposed method can achieve better segmentation accuracy with less variance comparing with an automatic graph cut method. Experimental results demonstrate the promises of our assessment framework.
Zhennan Yan, Chaowei Tan, Shaoting Zhang 0001, Boubakeur Belaroussi, Hui Jing Yu, Colin Miller, Dimitris N. Metaxas
ICPR8
2014 A New Framework for Sign Language Recognition based on 3D Handshape Identification and Linguistic Modeling
Mark Dilsizian, Polina Yanovich, Carol Neidle, Dimitris N. Metaxas
LREC5
2014 3D Face Tracking and Multi-Scale, Spatio-temporal Analysis of Linguistically Significant Facial Expressions and Head Positions in ASL
Bo Liu 0005, Jingjing Liu 0001, Xiang Yu 0002, Dimitris N. Metaxas, Carol Neidle
LREC4
2014 Scalable Histopathological Image Analysis via Active Learning
Yan Zhu 0009, Shaoting Zhang 0001, Wei Liu 0005, Dimitris N. Metaxas
MICCAI (3)4
2014 Optree: A Learning-Based Adaptive Watershed Algorithm for Neuron Segmentation
Mustafa Gökhan Uzunbas, Chao Chen 0012, Dimitris N. Metaxas
MICCAI (1)3
2014 Mode Estimation for High Dimensional Discrete Tree Graphical Models
Chao Chen 0012, Han Liu 0001, Dimitris N. Metaxas
NIPS3
2014 Non-manual grammatical marker recognition based on multi-scale, spatio-temporal analysis of head pose and facial expressions
Jingjing Liu 0001, Bo Liu 0005, Shaoting Zhang 0001, Fei Yang 0001, Peng Yang 0001, Dimitris N. Metaxas, Carol Neidle
Image Vis. Comput.6
2014 Preface
Dimitris N. Metaxas, Leon Axel
Medical Image Anal.1
2014 Deformable models with sparsity constraints for cardiac motion analysis
Yang Yu 0010, Shaoting Zhang 0001, Kang Li 0004, Dimitris N. Metaxas, Leon Axel
Medical Image Anal.4
2014 CeleST: Computer Vision Software for Quantitative Analysis of C. elegans Swim Behavior Reveals Novel Features of Locomotion
abstract
In the effort to define genes and specific neuronal circuits that control behavior and plasticity, the capacity for high-precision automated analysis of behavior is essential. We report on comprehensive computer vision software for analysis of swimming locomotion of C. elegans, a simple animal model initially developed to facilitate elaboration of genetic influences on behavior. C. elegans swim test software CeleST tracks swimming of multiple animals, measures 10 novel parameters of swim behavior that can fully report dynamic changes in posture and speed, and generates data in several analysis formats, complete with statistics. Our measures of swim locomotion utilize a deformable model approach and a novel mathematical analysis of curvature maps that enable even irregular patterns and dynamic changes to be scored without need for thresholding or dropping outlier swimmers from study. Operation of CeleST is mostly automated and only requires minimal investigator interventions, such as the selection of videotaped swim trials and choice of data output format. Data can be analyzed from the level of the single animal to populations of thousands. We document how the CeleST program reveals unexpected preferences for specific swim "gaits" in wild-type C. elegans, uncovers previously unknown mutant phenotypes, efficiently tracks changes in aging populations, and distinguishes "graceful" from poor aging. The sensitivity, dynamic range, and comprehensive nature of CeleST measures elevate swim locomotion analysis to a new level of ease, economy, and detail that enables behavioral plasticity resulting from genetic, cellular, or experience manipulation to be analyzed in ways not previously possible.
Christophe Restif, Carolina Ibáñez-Ventoso, Mehul Nalin Vora, Suzhen Guo, Dimitris N. Metaxas, Monica Driscoll
PLoS Comput. Biol.5
2014 On Continuous User Authentication via Typing Behavior
abstract
We hypothesize that an individual computer user has a unique and consistent habitual pattern of hand movements, independent of the text, while typing on a keyboard. As a result, this paper proposes a novel biometric modality named typing behavior (TB) for continuous user authentication. Given a webcam pointing toward a keyboard, we develop real-time computer vision algorithms to automatically extract hand movement patterns from the video stream. Unlike the typical continuous biometrics, such as keystroke dynamics (KD), TB provides a reliable authentication with a short delay, while avoiding explicit key-logging. We collect a video database where 63 unique subjects type static text and free text for multiple sessions. For one typing video, the hands are segmented in each frame and a unique descriptor is extracted based on the shape and position of hands, as well as their temporal dynamics in the video sequence. We propose a novel approach, named bag of multi-dimensional phrases, to match the cross-feature and cross-temporal pattern between a gallery sequence and probe sequence. The experimental results demonstrate a superior performance of TB when compared with KD, which, together with our ultrareal-time demo system, warrant further investigation of this novel vision application and biometric modality.
Joseph Roth, Xiaoming Liu 0002, Dimitris N. Metaxas
IEEE Trans. Image Process.3
2013 Computing the M Most Probable Modes of a Graphical Model
abstract
We introduce the M-modes problem for graphical models: predicting the M label configurations of highest probability that are at the same time local maxima of the probability landscape. M-modes have multiple possible applications: because they are intrinsically diverse, they provide a principled alternative to non-maximum suppression techniques for structured prediction, they can act as codebook vectors for quantizing the configuration space, or they can form component centers for mixture model approximation. We present two algorithms for solving the M-modes problem. The first algorithm solves the problem in polynomial time when the underlying graphical model is a simple chain. The second algorithm solves the problem for junction chains. In synthetic and real dataset, we demonstrate how M-modes can improve the performance of prediction. We also use the generated modes as a tool to understand the topography of the probability distribution of configurations, for example with relation to the training set size and amount of noise in the data.
Chao Chen 0012, Vladimir Kolmogorov, Yan Zhu 0009, Dimitris N. Metaxas, Christoph H. Lampert
AISTATS4
2013 Handling Noise in Single Image Deblurring Using Directional Filters
abstract
State-of-the-art single image deblurring techniques are sensitive to image noise. Even a small amount of noise, which is inevitable in low-light conditions, can degrade the quality of blur kernel estimation dramatically. The recent approach of Tai and Lin [17] tries to iteratively denoise and deblur a blurry and noisy image. However, as we show in this work, directly applying image denoising methods often partially damages the blur information that is extracted from the input image, leading to biased kernel estimation. We propose a new method for handling noise in blind image deconvolution based on new theoretical and practical insights. Our key observation is that applying a directional low-pass filter to the input image greatly reduces the noise level, while preserving the blur information in the orthogonal direction to the filter. Based on this observation, our method applies a series of directional filters at different orientations to the input image, and estimates an accurate Radon transform of the blur kernel from each filtered image. Finally, we reconstruct the blur kernel using inverse Radon transform. Experimental results on synthetic and real data show that our algorithm achieves higher quality results than previous approaches on blurry and noisy images.
Lin Zhong 0002, Sunghyun Cho, Dimitris N. Metaxas, Sylvain Paris, Jue Wang 0001
CVPR3
2013 Pose-Free Facial Landmark Fitting via Optimized Part Mixtures and Cascaded Deformable Shape Model
abstract
This paper addresses the problem of facial landmark localization and tracking from a single camera. We present a two-stage cascaded deformable shape model to effectively and efficiently localize facial landmarks with large head pose variations. For face detection, we propose a group sparse learning method to automatically select the most salient facial landmarks. By introducing 3D face shape model, we use procrustes analysis to achieve pose-free facial landmark initialization. For deformation, the first step uses mean-shift local search with constrained local model to rapidly approach the global optimum. The second step uses component-wise active contours to discriminatively refine the subtle shape variation. Our framework can simultaneously handle face detection, pose-free landmark localization and tracking in real time. Extensive experiments are conducted on both laboratory environmental face databases and face-in-the-wild databases. All results demonstrate that our approach has certain advantages over state-of-the-art methods in handling pose variations.
Xiang Yu 0002, Junzhou Huang, Shaoting Zhang 0001, Wang Yan, Dimitris N. Metaxas
ICCV5
2013 Adaptive low rank and sparse decomposition of video using compressive sensing
abstract
We address the problem of reconstructing and analyzing surveillance videos using compressive sensing. We develop a new method that performs video reconstruction by low rank and sparse decomposition adaptively. Background subtraction becomes part of the reconstruction. In our method, a background model is used in which the background is learned adaptively as the compressive measurements are processed. The adaptive method has low latency, and is more robust than previous methods. We will present experimental results to demonstrate the advantages of the proposed method.
Fei Yang 0001, Hong Jiang 0002, Zuowei Shen, Dimitris N. Metaxas
ICIP5
2013 Collaborative Multi Organ Segmentation by Integrating Deformable and Graphical Models
Mustafa Gökhan Uzunbas, Chao Chen 0012, Shaoting Zhang 0001, Kilian M. Pohl, Kang Li 0004, Dimitris N. Metaxas
MICCAI (2)6
2013 3D anatomical shape atlas construction using mesh quality preserved deformable models
Shaoting Zhang 0001, Yiqiang Zhan, Xinyi Cui, Mingchen Gao, Junzhou Huang, Dimitris N. Metaxas
Comput. Vis. Image Underst.6
2013 Recognizing expressions from face and body gesture by temporal normalized motion and appearance features
Shizhi Chen, Yingli Tian, Qingshan Liu 0001, Dimitris N. Metaxas
Image Vis. Comput.4
2013 A review of motion analysis methods for human Nonverbal Communication Computing
Dimitris N. Metaxas, Shaoting Zhang 0001
Image Vis. Comput.1
2012 Facial expression editing in video using a temporally-smooth factorization
abstract
We address the problem of editing facial expression in video, such as exaggerating, attenuating or replacing the expression with a different one in some parts of the video. To achieve this we develop a tensor-based 3D face geometry reconstruction method, which fits a 3D model for each video frame, with the constraint that all models have the same identity and requiring temporal continuity of pose and expression. With the identity constraint, the differences between the underlying 3D shapes capture only changes in expression and pose. We show that various expression editing tasks in video can be achieved by combining face reordering with face warping, where the warp is induced by projecting differences in 3D face shapes into the image plane. Analogously, we show how the identity can be manipulated while fixing expression and pose. Experimental results show that our method can effectively edit expressions and identity in video in a temporally-coherent way with high fidelity.
Fei Yang 0001, Lubomir D. Bourdev, Eli Shechtman, Jue Wang 0001, Dimitris N. Metaxas
CVPR5
2012 Learning active facial patches for expression analysis
abstract
In this paper, we present a new idea to analyze facial expression by exploring some common and specific information among different expressions. Inspired by the observation that only a few facial parts are active in expression disclosure (e.g., around mouth, eye), we try to discover the common and specific patches which are important to discriminate all the expressions and only a particular expression, respectively. A two-stage multi-task sparse learning (MTSL) framework is proposed to efficiently locate those discriminative patches. In the first stage MTSL, expression recognition tasks, each of which aims to find dominant patches for each expression, are combined to located common patches. Second, two related tasks, facial expression recognition and face verification tasks, are coupled to learn specific facial patches for individual expression. Extensive experiments validate the existence and significance of common and specific patches. Utilizing these learned patches, we achieve superior performances on expression recognition compared to the state-of-the-arts.
Lin Zhong 0002, Qingshan Liu 0001, Peng Yang 0001, Bo Liu 0005, Junzhou Huang, Dimitris N. Metaxas
CVPR6
2012 Background Subtraction Using Low Rank and Group Sparsity Constraints
Xinyi Cui, Junzhou Huang, Shaoting Zhang 0001, Dimitris N. Metaxas
ECCV (1)4
2012 Query Specific Fusion for Image Retrieval
Shaoting Zhang 0001, Ming Yang 0007, Timothée Cour, Kai Yu 0001, Dimitris N. Metaxas
ECCV (2)5
2012 Face morphing using 3D-aware appearance optimization
Fei Yang 0001, Eli Shechtman, Jue Wang 0001, Lubomir D. Bourdev, Dimitris N. Metaxas
Graphics Interface5
2012 Robust face tracking with a consumer depth camera
abstract
We address the problem of tracking human faces under various poses and lighting conditions. Reliable face tracking is a challenging task. The shapes of the faces may change dramatically with various identities, poses and expressions. Moreover, poor lighting conditions may cause a low contrast image or cast shadows on faces, which will significantly degrade the performance of the tracking system. In this paper, we develop a framework to track face shapes by using both color and depth information. Since the faces in various poses lie on a nonlinear manifold, we build piecewise linear face models, each model covering a range of poses. The low-resolution depth image is captured by using Microsoft Kinect, and is used to predict head pose and generate extra constraints at the face boundary. Our experiments show that, by exploiting the depth information, the performance of the tracking system is significantly improved.
Fei Yang 0001, Junzhou Huang, Xiang Yu 0002, Xinyi Cui, Dimitris N. Metaxas
ICIP5
2012 Robust eyelid tracking for fatigue detection
abstract
We develop a non-intrusive system for monitoring fatigue by tracking eyelids with a single web camera. Tracking slow eyelid closures is one of the most reliable ways to monitor fatigue during critical performance tasks. The challenges come from arbitrary head movement, occlusion, reflection of glasses, motion blurs, etc. We model the shape of eyes using a pair of parameterized parabolic curves, and fit the model in each frame to maximize the total likelihood of the eye regions. Our system is able to track face movement and fit eyelids reliably in real time. We test our system with videos captured from both alert and drowsy subjects. The experiment results prove the effectiveness of our system.
Fei Yang 0001, Xiang Yu 0002, Junzhou Huang, Peng Yang 0001, Dimitris N. Metaxas
ICIP5
2012 Activity recognition based on semantic spatial relation
Lingxun Meng, Laiyun Qing, Peng Yang 0001, Xilin Chen 0001, Dimitris N. Metaxas
ICPR6
2012 Towards Automatic Stereoscopic Video Synthesis from a Casual Monocular Video
abstract
Automatically synthesizing 3D content from a causal monocular video has become an important problem. Previous works either use no geometry information, or rely on precise 3D geometry information. Therefore, they cannot obtain reasonable results if the 3D structure in the scene is complex, or noisy 3D geometry information is estimated from monocular videos. In this paper, we present an automatic and robust framework to synthesize stereoscopic videos from casual 2D monocular videos. First, 3D geometry information (e.g., camera parameters, depth map) are extracted from the 2D input video. Then a Bayesian-based View Synthesis (BVS) approach is proposed to render high-quality new virtual views for stereoscopic video to deal with noisy 3D geometry information. Extensive experiments on various videos demonstrate that BVS can synthesize more accurate views than other methods, and our proposed framework also be able to generate high-quality 3D videos.
Lin Zhong 0002, Rodney L. Miller, Dimitris N. Metaxas
ISM5
2012 Recognition of Nonmanual Markers in American Sign Language (ASL) Using Non-Parametric Adaptive 2D-3D Face Tracking
Dimitris N. Metaxas, Bo Liu 0005, Fei Yang 0001, Peng Yang 0001, Nicholas Michael, Carol Neidle
LREC1
2012 Temporal Shape Analysis via the Spectral Signature
Elena Bernardis, Ender Konukoglu, Yangming Ou, Dimitris N. Metaxas, Benoit Desjardins, Kilian M. Pohl
MICCAI (2)4
2012 Simplified Labeling Process for Medical Image Segmentation
Mingchen Gao, Junzhou Huang, Sharon X. Huang, Shaoting Zhang 0001, Dimitris N. Metaxas
MICCAI (2)5
2012 Shape Prior Modeling Using Sparse Representation and Online Dictionary Learning
Shaoting Zhang 0001, Yiqiang Zhan, Mustafa Gökhan Uzunbas, Dimitris N. Metaxas
MICCAI (3)5
2012 Temporal Spectral Residual for fast salient motion detection
Xinyi Cui, Qingshan Liu 0001, Shaoting Zhang 0001, Fei Yang 0001, Dimitris N. Metaxas
Neurocomputing5
2012 Towards robust and effective shape modeling: Sparse shape composition
Shaoting Zhang 0001, Yiqiang Zhan, Maneesh Dewan, Junzhou Huang, Dimitris N. Metaxas, Xiang Sean Zhou
Medical Image Anal.5
2012 Deformable segmentation via sparse representation and dictionary learning
Shaoting Zhang 0001, Yiqiang Zhan, Dimitris N. Metaxas
Medical Image Anal.3
2012 Automatic Image Annotation and Retrieval Using Group Sparsity
abstract
Automatically assigning relevant text keywords to images is an important problem. Many algorithms have been proposed in the past decade and achieved good performance. Efforts have focused upon model representations of keywords, whereas properties of features have not been well investigated. In most cases, a group of features is preselected, yet important feature properties are not well used to select features. In this paper, we introduce a regularization-based feature selection algorithm to leverage both the sparsity and clustering properties of features, and incorporate it into the image annotation task. Using this group-sparsity-based method, the whole group of features [e.g., red green blue (RGB) or hue, saturation, and value (HSV)] is either selected or removed. Thus, we do not need to extract this group of features when new data comes. A novel approach is also proposed to iteratively obtain similar and dissimilar pairs from both the keyword similarity and the relevance feedback. Thus, keyword similarity is modeled in the annotation framework. We also show that our framework can be employed in image retrieval tasks by selecting different image pairs. Extensive experiments are designed to compare the performance between features, feature combinations, and regularization-based feature selection methods applied on the image annotation task, which gives insight into the properties of features in the image annotation task. The experimental results demonstrate that the group-sparsity-based method is more accurate and stable than others.
Shaoting Zhang 0001, Junzhou Huang, Hongsheng Li 0001, Dimitris N. Metaxas
IEEE Trans. Syst. Man Cybern. Part B4
2011 A Framework for the Recognition of Nonmanual Markers in Segmented Sequences of American Sign Language
abstract
Despite the fact that there is critical grammatical information expressed through facial expressions and head gestures, most research in the field of sign language recognition has primarily focused on the manual component of signing.We propose a novel framework for robust tracking and analysis of non-manual behaviours, with an application to sign language recognition.The novelty of our method is threefold.First, we propose a dynamic feature representation.Instead of using only the features available in the current frame (e.g., head pose), we additionally aggregate and encode the feature values in neighbouring frames to better encode the dynamics of expressions and gestures (e.g., head shakes).Second, we use Multiple Instance Learning [12] to handle feature misalignment resulting from drifting of the face tracker and partial occlusions.Third, we utilize a discriminative Hidden Markov Support Vector Machine (HMSVM) [1] to learn finer temporal dependencies between the features of interest.We apply our signerindependent framework to segmented recognition of five classes of grammatical constructions conveyed through facial expressions and head gestures: wh-questions, negation, conditional/when clauses, yes/no questions and topics, and show improvement over previous methods.
Nicholas Michael, Peng Yang 0001, Qingshan Liu 0001, Dimitris N. Metaxas, Carol Neidle
BMVC4
2011 Abnormal detection using interaction energy potentials
abstract
A new method is proposed to detect abnormal behaviors in human group activities. This approach effectively models group activities based on social behavior analysis. Different from previous work that uses independent local features, our method explores the relationships between the current behavior state of a subject and its actions. An interaction energy potential function is proposed to represent the current behavior state of a subject, and velocity is used as its actions. Our method does not depend on human detection or segmentation, so it is robust to detection errors. Instead, tracked spatio-temporal interest points are able to provide a good estimation of modeling group interaction. SVM is used to find abnormal events. We evaluate our algorithm in two datasets: UMN and BEHAVE. Experimental results show its promising performance against the state-of-art methods.
Xinyi Cui, Qingshan Liu 0001, Mingchen Gao, Dimitris N. Metaxas
CVPR4
2011 Sparse shape composition: A new framework for shape prior modeling
abstract
Image appearance cues are often used to derive object shapes, which is usually one of the key steps of image understanding tasks. However, when image appearance cues are weak or misleading, shape priors become critical to infer and refine the shape derived by these appearance cues. Effective modeling of shape priors is challenging because: 1) shape variation is complex and cannot always be modeled by a parametric probability distribution; 2) a shape instance derived from image appearance cues (input shape) may have gross errors; and 3) local details of the input shape are difficult to preserve if they are not statistically significant in the training data. In this paper we propose a novel Sparse Shape Composition model (SSC) to deal with these three challenges in a unified framework. In our method, training shapes are adaptively composed to infer/refine an input shape. The a-priori information is thus implicitly incorporated on-the-fly. Our model leverages two sparsity observations of the input shape instance: 1) the input shape can be approximately represented by a sparse linear combination of training shapes; 2) parts of the input shape may contain gross errors but such errors are usually sparse. Using L1 norm relaxation, our model is formulated as a convex optimization problem, which is solved by an efficient alternating minimization framework. Our method is extensively validated on two real world medical applications, 2D lung localization in X-ray images and 3D liver segmentation in low-dose CT scans. Compared to state-of-the-art methods, our model exhibits better performance in both studies.
Shaoting Zhang 0001, Yiqiang Zhan, Maneesh Dewan, Junzhou Huang, Dimitris N. Metaxas, Xiang Sean Zhou
CVPR5
2011 Segment and recognize expression phase by fusion of motion area and neutral divergence features
abstract
An expression can be approximated by a sequence of temporal segments called neutral, onset, offset and apex. However, it is not easy to accurately detect such temporal segments only based on facial features. Some researchers try to temporally segment expression phases with the help of body gesture analysis. The problem of this approach is that the expression temporal phases from face and gesture channels are not synchronized. Additionally, most previous work adopted facial key points tracking or body tracking to extract motion information, which is unreliable in practice due to illumination variations and occlusions. In this paper, we present a novel algorithm to overcome the above issues, in which two simple and robust features are designed to describe face and gesture information, i.e., motion area and neutral divergence features. Both features do not depend on motion tracking, and they can be easily calculated too. Moreover, it is different from previous work in that we integrate face and body gesture together in modeling the temporal dynamics through a single channel of sensorial source, so it avoids the unsynchronized issue between face and gesture channels. Extensive experimental results demonstrate the effectiveness of the proposed algorithm.
Shizhi Chen, Yingli Tian, Qingshan Liu 0001, Dimitris N. Metaxas
FG4
2011 Sparse shape registration for occluded facial feature localization
abstract
This paper proposes a sparsity driven shape registration method for occluded facial feature localization. Most current shape registration methods search landmark locations which comply both shape model and local image appearances. However, if the shape is partially occluded, the above goal is inappropriate and often leads to distorted shape results. In this paper, we introduce an error term to rectify the locations of the occluded landmarks. Under the assumption that occlusion takes a small proportion of the shape, we propose a sparse optimization algorithm that iteratively approaches the optimal shape. The experiments in our synthesized face occlusion database prove the advantage of our method.
Fei Yang 0001, Junzhou Huang, Dimitris N. Metaxas
FG3
2011 Eye localization through multiscale sparse dictionaries
abstract
This paper presents a new eye localization method via Multiscale Sparse Dictionaries (MSD). We built a pyramid of dictionaries that models context information at multiple scales. Eye locations are estimated at each scale by fitting the image through sparse coefficients of the dictionary. By using context information, our method is robust to various eye appearances. The method also works efficiently since it avoids sliding a search window in the image during localization. The experiments in BioID database prove the effectiveness of our method.
Fei Yang 0001, Junzhou Huang, Peng Yang 0001, Dimitris N. Metaxas
FG4
2011 A Belief Propagation algorithm for bias field estimation and image segmentation
abstract
Intensity-based image segmentation is often plagued by the spatial intensity inhomogeneities (or non-uniformities) that are caused by the imperfection of the imaging devices and the varying operating conditions, also known as the bias field. We present a graphical model representation of the joint segmentation and bias field estimation problem and propose an iterative solver based on the Belief Propagation (BP) algorithm. The intractable joint inference problem of the original graphical model is decoupled into two MRF-MAP estimation problems and solved by a discrete-valued BP and a Gaussian BP, respectively and iteratively. We validate our method using both simulated and real data and show its connection to some of the classical filtering-based approaches.
Rui Huang 0001, Nong Sang, Vladimir Pavlovic 0001, Dimitris N. Metaxas
ICIP4
2011 Using High Resolution Cardiac CT Data to Model and Visualize Patient-Specific Interactions between Trabeculae and Blood Flow
Scott Kulp, Mingchen Gao, Shaoting Zhang 0001, Szilard Voros, Dimitris N. Metaxas, Leon Axel
MICCAI (1)6
2011 3D Segmentation of Rodent Brain Structures Using Hierarchical Shape Priors and Deformable Models
Shaoting Zhang 0001, Junzhou Huang, Mustafa Gökhan Uzunbas, Tian Shen, Foteini Delis, Sharon X. Huang, Nora D. Volkow, Panayotis K. Thanos, Dimitris N. Metaxas
MICCAI (3)9
2011 Deformable Segmentation via Sparse Shape Representation
Shaoting Zhang 0001, Yiqiang Zhan, Maneesh Dewan, Junzhou Huang, Dimitris N. Metaxas, Xiang Sean Zhou
MICCAI (2)5
2011 Content quality based image retrieval with multiple instance boost ranking
abstract
Most previous works treated image retrieval as a classification problem or a similarity measurement problem. In this paper, we propose a new idea for image retrieval, in which we regard image retrieval as a ranking issue by evaluating image content quality. Based on the content preference between the images, the image pairs are organized to build the data set for rank learning. Because image content generally is disclosed by image patches with meaningful objects, each image is looked as one bag, and the regions inside are the corresponding instances. In order to save the computation cost, the instances in the image are the rectangle regions and the integral histogram is applied to speed up histogram feature extraction. Due to the feature dimension is high, we propose a boost-based multiple instance learning for image retrieval. Based on different assumptions in multiple instance setting, Mean, Max and TopK ranking models are developed with Boost learning. Experiments on the real-world images from Flickr, Pisca, and Google shows that the power of the proposed method.
Peng Yang 0001, Qingshan Liu 0001, Lin Zhong 0002, Dimitris N. Metaxas
ACM Multimedia5
2011 Robust mesh editing using Laplacian coordinates
Shaoting Zhang 0001, Junzhou Huang, Dimitris N. Metaxas
Graph. Model.3
2011 Composite splitting algorithms for convex optimization
Junzhou Huang, Shaoting Zhang 0001, Hongsheng Li 0001, Dimitris N. Metaxas
Comput. Vis. Image Underst.4
2011 Dynamic soft encoded patterns for facial event analysis
Peng Yang 0001, Qingshan Liu 0001, Dimitris N. Metaxas
Comput. Vis. Image Underst.3
2011 Learning with Structured Sparsity
Junzhou Huang, Tong Zhang 0001, Dimitris N. Metaxas
J. Mach. Learn. Res.3
2011 Efficient MR image reconstruction for compressed MR imaging
Junzhou Huang, Shaoting Zhang 0001, Dimitris N. Metaxas
Medical Image Anal.3
2011 Unsupervised Image Categorization by Hypergraph Partition
abstract
We present a framework for unsupervised image categorization in which images containing specific objects are taken as vertices in a hypergraph and the task of image clustering is formulated as the problem of hypergraph partition. First, a novel method is proposed to select the region of interest (ROI) of each image, and then hyperedges are constructed based on shape and appearance features extracted from the ROIs. Each vertex (image) and its k-nearest neighbors (based on shape or appearance descriptors) form two kinds of hyperedges. The weight of a hyperedge is computed as the sum of the pairwise affinities within the hyperedge. Through all of the hyperedges, not only the local grouping relationships among the images are described, but also the merits of the shape and appearance characteristics are integrated together to enhance the clustering performance. Finally, a generalized spectral clustering technique is used to solve the hypergraph partition problem. We compare the proposed method to several methods and its effectiveness is demonstrated by extensive experiments on three image databases.
Yuchi Huang, Qingshan Liu 0001, Fengjun Lv, Yihong Gong, Dimitris N. Metaxas
IEEE Trans. Pattern Anal. Mach. Intell.5
2011 Hypergraph with sampling for image retrieval
Qingshan Liu 0001, Yuchi Huang, Dimitris N. Metaxas
Pattern Recognit.3
2011 A Level Set Method for Image Segmentation in the Presence of Intensity Inhomogeneities With Application to MRI
abstract
Intensity inhomogeneity often occurs in real-world images, which presents a considerable challenge in image segmentation. The most widely used image segmentation algorithms are region-based and typically rely on the homogeneity of the image intensities in the regions of interest, which often fail to provide accurate segmentation results due to the intensity inhomogeneity. This paper proposes a novel region-based method for image segmentation, which is able to deal with intensity inhomogeneities in the segmentation. First, based on the model of images with intensity inhomogeneities, we derive a local intensity clustering property of the image intensities, and define a local clustering criterion function for the image intensities in a neighborhood of each point. This local clustering criterion function is then integrated with respect to the neighborhood center to give a global criterion of image segmentation. In a level set formulation, this criterion defines an energy in terms of the level set functions that represent a partition of the image domain and a bias field that accounts for the intensity inhomogeneity of the image. Therefore, by minimizing this energy, our method is able to simultaneously segment the image and estimate the bias field, and the estimated bias field can be used for intensity inhomogeneity correction (or bias correction). Our method has been validated on synthetic images and real images of various modalities, with desirable performance in the presence of intensity inhomogeneities. Experiments show that our method is more robust to initialization, faster and more accurate than the well-known piecewise smooth model. As an application, our method has been used for segmentation and bias correction of magnetic resonance (MR) images with promising results.
Chunming Li, Rui Huang 0001, Zhaohua Ding, Chris Gatenby, Dimitris N. Metaxas, John C. Gore
IEEE Trans. Image Process.5
2011 Identifying Regional Cardiac Abnormalities From Myocardial Strains Using Nontracking-Based Strain Estimation and Spatio-Temporal Tensor Analysis
abstract
Myocardial strain is a critical indicator of many cardiac diseases and dysfunctions. The goal of this paper is to extract and use the myocardial strain pattern from tagged magnetic resonance imaging (MRI) to identify and localize regional abnormal cardiac function in human subjects. In order to extract the myocardial strains from the tagged images, we developed a novel nontracking-based strain estimation method for tagged MRI. This method is based on the direct extraction of tag deformation, and therefore avoids some limitations of conventional displacement or tracking-based strain estimators. Based on the extracted spatio-temporal strain patterns, we have also developed a novel tensor-based classification framework that better conserves the spatio-temporal structure of the myocardial strain pattern than conventional vector-based classification algorithms. In addition, the tensor-based projection function keeps more of the information of the original feature space, so that abnormal tensors in the subspace can be back-projected to reveal the regional cardiac abnormality in a more physically meaningful way. We have tested our novel methods on 41 human image sequences, and achieved a classification rate of 87.80%. The regional abnormalities recovered from our algorithm agree well with the patient's pathology and clinical image interpretation, and provide a promising avenue for regional cardiac function analysis.
Qingshan Liu 0001, Dimitris N. Metaxas, Leon Axel
IEEE Trans. Medical Imaging3
2011 Expression flow for 3D-aware face component transfer
abstract
We address the problem of correcting an undesirable expression on a face photo by transferring local facial components, such as a smiling mouth, from another face photo of the same person which has the desired expression. Direct copying and blending using existing compositing tools results in semantically unnatural composites, since expression is a global effect and the local component in one expression is often incompatible with the shape and other components of the face in another expression. To solve this problem we present Expression Flow, a 2D flow field which can warp the target face globally in a natural way, so that the warped face is compatible with the new facial component to be copied over. To do this, starting with the two input face photos, we jointly construct a pair of 3D face shapes with the same identity but different expressions. The expression flow is computed by projecting the difference between the two 3D shapes back to 2D. It describes how to warp the target face photo to match the expression of the reference photo. User studies suggest that our system is able to generate face composites with much higher fidelity than existing methods.
Fei Yang 0001, Jue Wang 0001, Eli Shechtman, Lubomir D. Bourdev, Dimitris N. Metaxas
ACM Trans. Graph.5
2011 A Component-Based Framework for Generalized Face Alignment
abstract
This paper presents a component-based deformable model for generalized face alignment, in which a novel bistage statistical model is proposed to account for both local and global shape characteristics. Instead of using statistical analysis on the entire shape, we build separate Gaussian models for shape components to preserve more detailed local shape deformations. In each model of components, a Markov network is integrated to provide simple geometry constraints for our search strategy. In order to make a better description of the nonlinear interrelationships over shape components, the Gaussian process latent variable model is adopted to obtain enough control of shape variations. In addition, we adopt an illumination-robust feature to lead the local fitting of every shape point when light conditions change dramatically. To further boost the accuracy and efficiency of our component-based algorithm, an efficient subwindow search technique is adopted to detect components and to provide better initializations for shape components. Based on this approach, our system can generate accurate shape alignment results not only for images with exaggerated expressions and slight shading variation but also for images with occlusion and heavy shadows, which are rarely reported in previous work.
Yuchi Huang, Qingshan Liu 0001, Dimitris N. Metaxas
IEEE Trans. Syst. Man Cybern. Part B3
2010 Image retrieval via probabilistic hypergraph ranking
abstract
In this paper, we propose a new transductive learning framework for image retrieval, in which images are taken as vertices in a weighted hypergraph and the task of image search is formulated as the problem of hypergraph ranking. Based on the similarity matrix computed from various feature descriptors, we take each image as a `centroid' vertex and form a hyperedge by a centroid and its k-nearest neighbors. To further exploit the correlation information among images, we propose a probabilistic hypergraph, which assigns each vertex vito a hyperedge ejin a probabilistic way. In the incidence structure of a probabilistic hypergraph, we describe both the higher order grouping information and the affinity relationship between vertices within each hy-peredge. After feedback images are provided, our retrieval system ranks image labels by a transductive inference approach, which tends to assign the same label to vertices that share many incidental hyperedges, with the constraints that predicted labels of feedback images should be similar to their initial labels. We compare the proposed method to several other methods and its effectiveness is demonstrated by extensive experiments on Corel5K, the Scene dataset and Caltech 101.
Yuchi Huang, Qingshan Liu 0001, Shaoting Zhang 0001, Dimitris N. Metaxas
CVPR4
2010 Exploring facial expressions with compositional features
abstract
Most previous work focuses on how to learn discriminating appearance features over all the face without considering the fact that each facial expression is physically composed of some relative action units (AU). However, the definition of AU is an ambiguous semantic description in Facial Action Coding System (FACS), so it makes accurate AU detection very difficult. In this paper, we adopt a scheme of compromise to avoid AU detection, and try to interpret facial expression by learning some compositional appearance features around AU areas. We first divided face image into local patches according to the locations of AUs, and then we extract local appearance features from each patch. A minimum error based optimization strategy is adopted to build compositional features based on local appearance features, and this process embedded into Boosting learning structure. Experiments on the Cohn-Kanada database show that the proposed method has a promising performance and the built compositional features are basically consistent to FACS.
Peng Yang 0001, Qingshan Liu 0001, Dimitris N. Metaxas
CVPR3
2010 Automatic image annotation using group sparsity
abstract
Automatically assigning relevant text keywords to images is an important problem. Many algorithms have been proposed in the past decade and achieved good performance. Efforts have focused upon model representations of keywords, but properties of features have not been well investigated. In most cases, a group of features is preselected, yet important feature properties are not well used to select features. In this paper, we introduce a regularization based feature selection algorithm to leverage both the sparsity and clustering properties of features, and incorporate it into the image annotation task. A novel approach is also proposed to iteratively obtain similar and dissimilar pairs from both the keyword similarity and the relevance feedback. Thus keyword similarity is modeled in the annotation framework. Numerous experiments are designed to compare the performance between features, feature combinations and regularization based feature selection methods applied on the image annotation task, which gives insight into the properties of features in the image annotation task. The experimental results demonstrate that the group sparsity based method is more accurate and stable than others.
Shaoting Zhang 0001, Junzhou Huang, Yuchi Huang, Yang Yu 0010, Hongsheng Li 0001, Dimitris N. Metaxas
CVPR6
2010 Fast Optimization for Mixture Prior Models
Junzhou Huang, Shaoting Zhang 0001, Dimitris N. Metaxas
ECCV (3)3
2010 Motion Profiles for Deception Detection Using Visual Cues
Nicholas Michael, Mark Dilsizian, Dimitris N. Metaxas, Judee K. Burgoon
ECCV (6)3
2010 Ranking Model for Facial Age Estimation
abstract
Feature design and feature selection are two key problems in facial image based age perception. In this paper, we proposed to using ranking model to do feature selection on the haar-like features. In order to build the pairwise samples for the ranking model, age sequences are organized by personal aging pattern within each subject. The pairwise samples are extracted from the sequence of each subject. Therefore, the order information is intuitively contained in the pairwise data. Ranking model is used to select the discriminative features based on the pairwise data. The combination of the ranking model and personal aging pattern are powerful to select the discriminative features for age estimation. Based on the selected features, different kinds of regression models are used to build prediction models. The experiment results show the performance of our method is comparable to the state-of-art works.
Peng Yang 0001, Lin Zhong 0002, Dimitris N. Metaxas
ICPR3
2010 Efficient MR Image Reconstruction for Compressed MR Imaging
Junzhou Huang, Shaoting Zhang 0001, Dimitris N. Metaxas
MICCAI (1)3
2010 Lesion-Specific Coronary Artery Calcium Quantification for Predicting Cardiac Event with Multiple Instance Support Vector Machines
Qingshan Liu 0001, Idean Marvasty, Sarah Rinehart, Szilard Voros, Dimitris N. Metaxas
MICCAI (1)6
2010 Lennard-Jones force field for geometric active contour
Zhenglong Li 0001, Qingshan Liu 0001, Hanqing Lu, Dimitris N. Metaxas
Signal Process.4
2010 Automated 3D Motion Tracking Using Gabor Filter Bank, Robust Point Matching, and Deformable Models
abstract
Tagged magnetic resonance imaging (tagged MRI or tMRI) provides a means of directly and noninvasively displaying the internal motion of the myocardium. Reconstruction of the motion field is needed to quantify important clinical information, e.g., the myocardial strain, and detect regional heart functional loss. In this paper, we present a three-step method for this task. First, we use a Gabor filter bank to detect and locate tag intersections in the image frames, based on local phase analysis. Next, we use an improved version of the robust point matching (RPM) method to sparsely track the motion of the myocardium, by establishing a transformation function and a one-to-one correspondence between grid tag intersections in different image frames. In particular, the RPM helps to minimize the impact on the motion tracking result of 1) through-plane motion and 2) relatively large deformation and/or relatively small tag spacing. In the final step, a meshless deformable model is initialized using the transformation function computed by RPM. The model refines the motion tracking and generates a dense displacement map, by deforming under the influence of image information, and is constrained by the displacement magnitude to retain its geometric structure. The 2D displacement maps in short and long axis image planes can be combined to drive a 3D deformable model, using the moving least square method, constrained by the minimization of the residual error at tag intersections. The method has been tested on a numerical phantom, as well as on in vivo heart data from normal volunteers and heart disease patients. The experimental results show that the new method has a good performance on both synthetic and real data. Furthermore, the method has been used in an initial clinical study to assess the differences in myocardial strain distributions between heart disease (left ventricular hypertrophy) patients and the normal control group. The final results show that the proposed method is capable of separating patients from healthy individuals. In addition, the method detects and makes possible quantification of local abnormalities in the myocardium strain distribution, which is critical for quantitative analysis of patients' clinical conditions. This motion tracking approach can improve the throughput and reliability of quantitative strain analysis of heart disease patients, and has the potential for further clinical applications.
Ting Chen 0001, Sohae Chung, Dimitris N. Metaxas, Leon Axel
IEEE Trans. Medical Imaging4
2009 Spatial and temporal pyramids for grammatical expression recognition of American sign language
abstract
Given that sign language is used as a primary means of communication by as many as two million deaf individuals in the U.S. and as augmentative communication by hearing individuals with a variety of disabilities, the development of robust, real-time sign language recognition technologies would be a major step forward in making computers equally accessible to everyone. However, most research in the field of sign language recognition has focused on the manual component of signs, despite the fact that there is critical grammatical information expressed through facial expressions and head gestures.We propose a novel framework for robust tracking and analysis of facial expression and head gestures, with an application to sign language recognition. We then apply it to recognition with excellent accuracy (≥=95%) of two classes of grammatical expressions, namely wh-questions and negative expressions. Our method is signer-independent and builds on the popular bag-of-words model, utilizing spatial pyramids to model facial appearance and temporal pyramids to represent patterns of head pose changes.
Nicholas Michael, Dimitris N. Metaxas, Carol Neidle
ASSETS2
2009 Video object segmentation by hypergraph cut
abstract
In this paper, we present a new framework of video object segmentation, in which we formulate the task of extracting prominent objects from a scene as the problem of hypergraph cut. We initially over-segment each frame in the sequence, and take the over-segmented image patches as the vertices in the graph. Different from the traditional pairwise graph structure, we build a novel graph structure, hypergraph, to represent the complex spatio-temporal neighborhood relationship among the patches. We assign each patch with several attributes that are computed from the optical flow and the appearance-based motion profile, and the vertices with the same attribute value is connected by a hyperedge. Through all the hyperedges, not only the complex non-pairwise relationships between the patches are described, but also their merits are integrated together organically. The task of video object segmentation is equivalent to the hypergraph partition, which can be solved by the hypergraph cut algorithm. The effectiveness of the proposed method is demonstrated by extensive experiments on nature scenes.
Yuchi Huang, Qingshan Liu 0001, Dimitris N. Metaxas
CVPR3
2009 D - Clutter: Building object model library from unsupervised segmentation of cluttered scenes
abstract
Autonomous systems which learn and utilize a limited visual vocabulary have wide spread applications. Enabling such systems to segment a set of cluttered scenes into objects is a challenging vision problem owing to the non-homogeneous texture of objects and the random configurations of multiple objects in each scene. We present a solution to the following question: given a collection of images where each object appears in one or more images and multiple objects occur in each image, how best can we extract the boundaries of the different objects? The algorithm is presented with a set of stereo images, with one stereo pair per scene. The novelty of our work is the use of both color/texture and structure to refine previously determined object boundaries to achieve segmentation consistent with each of the input scenes presented. The algorithm populates an object library, which consists of a 3D model per object. Since an object is characterized both by texture and structure, for most purposes this representation is both complete and concise.
Gowri Somanath, M. V. Rohith, Dimitris N. Metaxas, Chandra Kambhamettu
CVPR3
2009 Learning with dynamic group sparsity
abstract
This paper investigates a new learning formulation called dynamic group sparsity. It is a natural extension of the standard sparsity concept in compressive sensing, and is motivated by the observation that in some practical sparse data the nonzero coefficients are often not random but tend to be clustered. Intuitively, better results can be achieved in these cases by reasonably utilizing both clustering and sparsity priors. Motivated by this idea, we have developed a new greedy sparse recovery algorithm, which prunes data residues in the iterative process according to both sparsity and group clustering priors rather than only sparsity as in previous methods. The proposed algorithm can recover stably sparse data with clustering trends using far fewer measurements and computations than current state-of-the-art algorithms with provable guarantees. Moreover, our algorithm can adaptively learn the dynamic group structure and the sparsity number if they are not available in the practical applications. We have applied the algorithm to sparse recovery and background subtraction in videos. Numerous experiments with improved performance over previous methods further validate our theoretical proofs and the effectiveness of the proposed algorithm.
Junzhou Huang, Sharon X. Huang, Dimitris N. Metaxas
ICCV3
2009 RankBoost with l1 regularization for facial expression recognition and intensity estimation
abstract
Most previous facial expression analysis works only focused on expression recognition. In this paper, we propose a novel framework of facial expression analysis based on the ranking model. Different from previous works, it not only can do facial expression recognition, but also can estimate the intensity of facial expression, which is very important to further understand human emotion. Although it is hard to label expression intensity quantitatively, the ordinal relationship in temporal domain is actually a good relative measurement. Based on this observation, we convert the problem of intensity estimation to a ranking problem, which is modeled by the RankBoost. The output ranking score can be directly used for intensity estimation, and we also extend the ranking function for expression recognition. To further improve the performance, we propose to introduce l 1 based regularization into the Rankboost. Experiments on the Cohn-Kanade database show that the proposed method has a promising performance compared to the state-of-the-art.
Peng Yang 0001, Qingshan Liu 0001, Dimitris N. Metaxas
ICCV3
2009 Learning with structured sparsity
abstract
This paper investigates a new learning formulation called structured sparsity, which is a natural extension of the standard sparsity concept in statistical learning and compressive sensing. By allowing arbitrary structures on the feature set, this concept generalizes the group sparsity idea. A general theory is developed for learning with structured sparsity, based on the notion of coding complexity associated with the structure. Moreover, a structured greedy algorithm is proposed to efficiently solve the structured sparsity problem. Experiments demonstrate the advantage of structured sparsity over standard sparsity.
Junzhou Huang, Tong Zhang 0001, Dimitris N. Metaxas
ICML3
2009 3D Meshless Prostate Segmentation and Registration in Image Guided Radiotherapy
Ting Chen 0001, Sung N. Kim, Jinghao Zhou, Dimitris N. Metaxas, Gunaretnam Rajagopal, Ning J. Yue
MICCAI (1)4
2009 Temporal spectral residual: fast motion saliency detection
abstract
Saliency detection has attracted much attention in recent years. It aims at locating semantic regions in images for further image understanding. In this paper, we address the issue of motion saliency detection for video content analysis. Inspired by the idea of Spectral Residual for image saliency detection, we propose a new method Temporal Spectral Residual on video slices along X-T and Y-T planes, which can automatically separate foreground motion objects from backgrounds, also with the help of threshold selection and voting schemes. Different from conventional background modeling methods with complex mathematical model, the proposed method is only based on Fourier spectrum analysis, so it is simple and fast. The power of our proposed method is demonstrated in the experiments of four typical videos with different dynamic background.
Xinyi Cui, Qingshan Liu 0001, Dimitris N. Metaxas
ACM Multimedia3
2009 Medical modeling special issue: An introduction
Dimitris N. Metaxas
Comput. Aided Des.1
2009 Simulation of two-phase flow with sub-scale droplet and bubble effects
abstract
Abstract We present a new Eulerian‐Lagrangian method for physics‐based simulation of fluid flow, which includes automatic generation of sub‐scale spray and bubbles. The Marker Level Set method is used to provide a simple geometric criterion for free marker generation. A filtering method, inspired from Weber number thresholding, further controls the free marker generation (in a physics‐based manner). Two separate models are used, one for sub‐scale droplets, the other for sub‐scale bubbles. Droplets are evolved in a Newtonian manner, using a density‐extension drag force field, while bubbles are evolved using a model based on Stokes' Law. We show that our model for sub‐scale droplet and bubble dynamics is simple to couple with a full (macro‐scale) Navier‐Stokes two‐phase flow model and is quite powerful in its applications. Our animations include coarse grained multiphase features interacting with fine scale multiphase features.
Viorel Mihalef, Dimitris N. Metaxas, Mark Sussman
Comput. Graph. Forum2
2009 Editorial
Dimitris N. Metaxas, Leon Axel
Medical Image Anal.1
2009 Guest Editors' Introduction to the Special Section on Probabilistic Graphical Models
abstract
The ten papers in this special section focus on applications of probabilistic graphical models in all areas of computer vision.
Jiebo Luo 0001, Dimitris N. Metaxas, Antonio Torralba 0001, Thomas S. Huang, Erik B. Sudderth
IEEE Trans. Pattern Anal. Mach. Intell.3
2009 Boosting encoded dynamic features for facial expression recognition
Peng Yang 0001, Qingshan Liu 0001, Dimitris N. Metaxas
Pattern Recognit. Lett.3
2009 CoCRF Deformable Model: A Geometric Model Driven by Collaborative Conditional Random Fields
abstract
We present a hybrid framework for integrating deformable models with learning-based classification, for image segmentation with region ambiguities. We show how a region-based geometric model is coupled with conditional random fields (CRF) in a simple graphical model, such that the model evolution is driven by a dynamically updated probability field. We define the model shape with the signed distance function, while we formulate the internal energy with a C(1) continuity constraint, a shape prior, and a term that forces the zero level of the shape function towards a connected form. The latter can be seen as a term that forces different closed curves on the image plane to merge, and, therefore, our model inherently carries the property of merging regions. We calculate the image likelihood that drives the evolution using a collaborative formulation of conditional random fields (CoCRF), which is updated during the evolution in an online learning manner. The CoCRF infers class posteriors to regions with feature ambiguities by assessing the joint appearance of neighboring sites, and using the classification confidence to regulate the inference. The novelties of our approach are (i) the tight coupling of deformable models with classification, combining the estimation of smooth region boundaries with the robustness of the probabilistic region classification, (ii) the handling of feature variations, by updating the region statistics in an online learning manner, and (iii) the improvement of the region classification using our CoCRF. We demonstrate the performance of our method in a variety of images with clutter, region inhomogeneities, boundary ambiguities, and complex textures, from the zebra and cheetah examples to medical images.
Gavriil Tsechpenakis, Dimitris N. Metaxas
IEEE Trans. Image Process.2
2009 Detecting Concealment of Intent in Transportation Screening: A Proof of Concept
abstract
Transportation and border security systems have a common goal: to allow law-abiding people to pass through security and detain those people who intend to harm. Understanding how intention is concealed and how it might be detected should help in attaining this goal. In this paper, we introduce a multidisciplinary theoretical model of intent concealment along with three verbal and nonverbal automated methods for detecting intent: message feature mining, speech act profiling, and kinesic analysis. This paper also reviews a program of empirical research supporting this model, including several previously published studies and the results of a proof-of-concept study. These studies support the model by showing that aspects of intent can be detected at a rate that is higher than chance. Finally, this paper discusses the implications of these findings in an airport-screening scenario.
Judee K. Burgoon, Douglas P. Twitchell, Matthew L. Jensen, Thomas O. Meservy, Mark Adkins, John Kruse, Amit V. Deokar, Gavriil Tsechpenakis, Shan Lu 0010, Dimitris N. Metaxas, Jay F. Nunamaker Jr., Robert Younger
IEEE Trans. Intell. Transp. Syst.10
2008 Fast algorithms for large scale conditional 3D prediction
abstract
The potential success of discriminative learning approaches to 3D reconstruction relies on the ability to efficiently train predictive algorithms using sufficiently many examples that are representative of the typical configurations encountered in the application domain. Recent research indicates that sparse conditional Bayesian mixture of experts (cMoE) models (e.g. BME (Sminchisescu et al., 2005)) are adequate modeling tools that not only provide contextual 3D predictions for problems like human pose reconstruction, but can also represent multiple interpretations that result from depth ambiguities or occlusion. However, training conditional predictors requires sophisticated double-loop algorithms that scale unfavorably with the input dimension and the training set size, thus limiting their usage to 10,000 examples of less, so far. In this paper we present large-scale algorithms, referred to as fBME, that combine forward feature selection and bound optimization in order to train probabilistic, BME models, with one order of magnitude more data (100,000 examples and up) and more than one order of magnitude faster. We present several large scale experiments, including monocular evaluation on the HumanEva dataset (Sigal and Black, 2006), demonstrating how the proposed methods overcome the scaling limitations of existing ones.
Liefeng Bo, Cristian Sminchisescu, Atul Kanaujia, Dimitris N. Metaxas
CVPR4
2008 Simultaneous image transformation and sparse representation recovery
abstract
Sparse representation in compressive sensing is gaining increasing attention due to its success in various applications. As we demonstrate in this paper, however, image sparse representation is sensitive to image plane transformations such that existing approaches can not reconstruct the sparse representation of a geometrically transformed image. We introduce a simple technique for obtaining transformation-invariant image sparse representation. It is rooted in two observations: 1) if the aligned model images of an object span a linear subspace, their transformed versions with respect to some group of transformations can still span a linear subspace in a higher dimension; 2) if a target (or test) image, aligned with the model images, lives in the above subspace, its pre-alignment versions would get closer to the subspace after applying estimated transformations with more and more accurate parameters. These observations motivate us to project a potentially unaligned target image to random projection manifolds defined by the model images and the transformation model. Each projection is then separated into the aligned projection target and a residue due to misalignment. The desired aligned projection target is then iteratively optimized by gradually diminishing the residue. In this framework, we can simultaneously recover the sparse representation of a target image and the image plane transformation between the target and the model images. We have applied the proposed methodology to two applications: face recognition, and dynamic texture registration. The improved performance over previous methods that we obtain demonstrates the effectiveness of the proposed approach.
Junzhou Huang, Sharon X. Huang, Dimitris N. Metaxas
CVPR3
2008 Meshless deformable models for LV motion analysis
abstract
We propose a novel meshless deformable model for in vivo cardiac left ventricle (LV) 3D motion estimation. As a relatively new technology, taggedMRI (tMRI) provides a direct and noninvasive way to reveal local deformation of the myocardium, which creates a large amount of heart motion data which requiring quantitative analysis. In our study, we sample the heart motion sparsely at intersections of three sets of orthogonal tagging planes and then use a new meshless deformable model to recover the dense 3D motion of the myocardium temporally during the cardiac cycle. We compute external forces at tag intersections based on tracked local motion and redistribute the force to meshless particles throughout the myocardium. Internal constraint forces at particles are derived from local strain energy using a Moving Least Squares (MLS) method. The dense 3D motion field is then computed and updated using the Lagrange equation. The new model avoids the singularity problem of mesh-based models and is capable of tracking large deformation with high efficiency and accuracy. In particular, the model performs well even when the control points (tag intersections) are relatively sparse. We tested the performance of the meshless model on a numerical phantom, as well as in vivo heart data of healthy subjects and patients. The experimental results show that the meshless deformable model can fully recover the myocardium motion in 3D.
Dimitris N. Metaxas, Ting Chen 0001, Leon Axel
CVPR2
2008 Facial expression recognition using encoded dynamic features
abstract
In this paper, we propose a novel framework for video-based facial expression recognition, which can handle the data with various time resolution including a single frame. We first use the haar-like features to represent facial appearance, due to their simplicity and effectiveness. Then we perform K-Means clustering on the facial appearance features to explore the intrinsic temporal patterns of each expression. Based on the temporal pattern models, we further map the facial appearance variations into dynamic binary patterns. Finally, boosting learning is performed to construct the expression classifiers. Compared to previous work, the dynamic binary patterns encode the intrinsic dynamics of expression, and our method makes no assumption on the time resolution of the data. Extensive experiments carried on the Cohn-Kanade database show the promising performance of the proposed method.
Peng Yang 0001, Qingshan Liu 0001, Xinyi Cui, Dimitris N. Metaxas
CVPR4
2008 Similarity Features for Facial Event Analysis
Peng Yang 0001, Qingshan Liu 0001, Dimitris N. Metaxas
ECCV (1)3
2008 Lennard-Jones force field for Geometric Active Contour
abstract
This paper presents a new Geometric Active Contour (GAC) model based on Lennard-Jones (L-J) force field, which is inspired by the theory of intermolecular interaction. It is different from gradient based GAC models in that the proposed model does not rely on any pre-computed edge map and is directly computed from image data. Moreover, it can integrate various information including grayscale, color and texture etc. The proposed L-J force field has two different characteristics controlled by a switch parameter c. In the case of c = 0, the force vector flows will bi-directionally converge to boundaries, and it will obtain a morphological dilation-like effect with c ≠ 0. We test the proposed method on various images, and the experimental results are very promising.
Zhenglong Li 0001, Qingshan Liu 0001, Hanqing Lu, Dimitris N. Metaxas
ICPR4
2008 Fast Motion Tracking of Tagged MRI Using Angle-Preserving Meshless Registration
Ting Chen 0001, Dimitris N. Metaxas, Leon Axel
MICCAI (2)3
2008 Tag Separation in Cardiac Tagged MRI
Junzhou Huang, Sharon X. Huang, Dimitris N. Metaxas, Leon Axel
MICCAI (2)4
2008 A Variational Level Set Approach to Segmentation and Bias Correction of Images with Intensity Inhomogeneity
Chunming Li, Rui Huang 0001, Zhaohua Ding, Chris Gatenby, Dimitris N. Metaxas, John C. Gore
MICCAI (2)5
2008 Identifying Regional Cardiac Abnormalities from Myocardial Strains Using Spatio-temporal Tensor Analysis
Qingshan Liu 0001, Dimitris N. Metaxas, Leon Axel
MICCAI (1)3
2008 Tracking the Swimming Motions of C. elegansWorms with Applications in Aging Studies
Christophe Restif, Dimitris N. Metaxas
MICCAI (1)2
2008 Active Volume Models with Probabilistic Object Boundary Prediction Module
Tian Shen, Yaoyao Zhu, Sharon X. Huang, Junzhou Huang, Dimitris N. Metaxas, Leon Axel
MICCAI (1)5
2008 LV Motion and Strain Computation from tMRI Based on Meshless Deformable Models
Ting Chen 0001, Shaoting Zhang 0001, Dimitris N. Metaxas, Leon Axel
MICCAI (1)4
2008 Interaction of two-phase flow with animated models
Viorel Mihalef, Samet Y. Kadioglu, Mark Sussman, Dimitris N. Metaxas, Vassilios Hurmusiadis
Graph. Model.4
2008 Metamorphs: Deformable Shape and Appearance Models
abstract
This paper presents a new deformable modeling strategy that is aimed at integrating shape and appearance in a unified space. If we think of traditional deformable models as "active contours" or "evolving curve fronts," the new deformable shape and appearance models that we propose are "deforming disks or volumes." Each model not only has boundary shape but also interior appearance. The model shape is implicitly embedded in a higher dimensional space of distance transforms and is thus represented by a distance map "image." This way, both the shape and the appearance of the model are defined in the pixel space. A common deformation scheme, that is, the free-form deformations (FFDs), parameterizes warping deformations of the volumetric space in which the model is embedded, hence simultaneously deforming both model boundary and interior. When applied to segmentation, a metamorphs model can be initialized by covering a seed region far from the object boundary, and then the model efficiently evolves and converges to an optimal solution. The model dynamics are derived in a unified variational framework that consists of edge-based and region-based energy terms, both of which are differentiable with respect to the common set of FFD parameters. As the model deforms, its interior appearance statistics are adaptively learned and, then, toward the next-step deformation, the model examines not only edge information but also its exterior region statistics to ensure that it only expands to new territory with consistent appearance statistics. The Metamorphs formulation also allows natural merging and competition of multiple models. We demonstrate the robustness of metamorphs by using both natural and medical images that have high noise levels, intensity inhomogeneity, and complex texture.
Sharon X. Huang, Dimitris N. Metaxas
IEEE Trans. Pattern Anal. Mach. Intell.2
2007 Face Mis-alignment Analysis by Multiple-Instance Subspace
Qingshan Liu 0001, Dimitris N. Metaxas
ACCV (2)3
2007 Semi-supervised Hierarchical Models for 3D Human Pose Reconstruction
abstract
Recent research in visual inference from monocular images has shown that discriminatively trained image-based predictors can provide fast, automatic qualitative 3D reconstructions of human body pose or scene structure in real-world environments. However, the stability of existing image representations tends to be perturbed by deformations and misalignments in the training set, which, in turn, degrade the quality of learning and generalization. In this paper we advocate the semi-supervised learning of hierarchical image descriptions in order to better tolerate variability at multiple levels of detail. We combine multilevel encodings with improved stability to geometric transformations, with metric learning and semi-supervised manifold regularization methods in order to further profile them for task-invariance -resistance to background clutter and within the same human pose class differences. We quantitatively analyze the effectiveness of both descriptors and learning methods and show that each one can contribute, sometimes substantially, to more reliable 3D human pose estimates in cluttered images.
Atul Kanaujia, Cristian Sminchisescu, Dimitris N. Metaxas
CVPR3
2007 CRF-driven Implicit Deformable Model
abstract
We present a topology independent solution for segmenting objects with texture patterns of any scale, using an implicit deformable model driven by conditional random fields (CRFs). Our model integrates region and edge information as image driven terms, whereas the probabilistic shape and internal (smoothness) terms use representations similar to the level-set based methods. The evolution of the model is solved as a MAP estimation problem, where the target conditional probability is decomposed into the internal term and the image-driven term. For the later, we use discriminative CRFs in two scales, pixel- and patch-based, to obtain smooth probability fields based on the corresponding image features. The advantages and novelties of our approach are (i) the integration of CRFs with implicit deformable models in a tightly coupled scheme, (ii) the use of CRFs which avoids ambiguities in the probability fields, (iii) the handling of local feature variations by updating the model interior statistics and processing at different spatial scales, and (v) the independence from the topology. We demonstrate the performance of our method in a wide variety of images, from the zebra and cheetah examples to the left and right ventricles in cardiac images.
Gavriil Tsechpenakis, Dimitris N. Metaxas
CVPR2
2007 Boosting Coded Dynamic Features for Facial Action Units and Facial Expression Recognition
abstract
It is well known that how to extract dynamical features is a key issue for video based face analysis. In this paper, we present a novel approach of facial action units (AU) and expression recognition based on coded dynamical features. In order to capture the dynamical characteristics of facial events, we design the dynamical haar-like features to represent the temporal variations of facial events. Inspired by the binary pattern coding, we further encode the dynamic haar-like features into binary pattern features, which are useful to construct weak classifiers for boosting learning. Finally the Adaboost is performed to learn a set of discriminating coded dynamic features for facial active units and expression recognition. Experiments on the CMU expression database and our own facial AU database show its encouraging performance.
Peng Yang 0001, Qingshan Liu 0001, Dimitris N. Metaxas
CVPR3
2007 Optimization and Learning for Registration of Moving Dynamic Textures
abstract
We address the problem of registering a sequence of images in a moving dynamic texture video. This involves optimization with respect to camera motion, the average image, and the dynamic texture model. This problem is highly ill-posed and almost impossible to have good solutions without priors. In this paper, we introduce powerful priors for this problem, based on two simple observations: 1) registration should simplify the dynamic texture model while preserving all useful information. It motivates us to compute a prior for the dynamic texture by marginalizing over specific dynamics in the space of all stable auto-regressive sequences; 2) the statistics of derivative filter responses in the average image can be significantly changed by registration, and better registration should lead to a sharper average image. This offers us the prior of requiring the derivative distribution of the estimated average image to be close to that learned from the input image sequence. With these priors, a new registration approach is proposed by marginalizing over the "nuisance" variables under a Bayesian framework. And superior motion estimation results are obtained by jointly optimizing over the registration parameters, the average image, and the dynamic texture model. Experimental results on real video sequences of moving dynamic textures show convincing performance of the proposed approach.
Junzhou Huang, Sharon X. Huang, Dimitris N. Metaxas
ICCV3
2007 A Component Based Deformable Model for Generalized Face Alignment
abstract
This paper presents a component based deformable model for generalized face alignment, in which a novel bi-stage statistical framework is proposed to account for both local and global shape characteristics. Instead of using statistical analysis on the entire shape as in previous alignment work, we build separate Gaussian models for shape components to preserve more detailed local shape deformations. In each model of components the Markov Network is integrated to provide simple geometry constraints for our search strategy. In order to make a better description of the nonlinear interrelationships over the shape components, the Gaussian process latent variable model is adopted to obtain enough control of full range shape variations. Furthermore, we propose an illumination-robust feature to lead the local fitting of every shape point when light conditions change dramatically. Based on this approach, our system can generate optimal shape for images with exaggerated expressions and under variable illumination, as evidenced by extensive experimentation.
Yuchi Huang, Qingshan Liu 0001, Dimitris N. Metaxas
ICCV3
2007 Embedded Profile Hidden Markov Models for Shape Analysis
abstract
An ideal shape model should be both invariant to global transformations and robust to local distortions. In this paper we present a new shape modeling framework that achieves both efficiently. A shape instance is described by a curvature-based shape descriptor. A Profile Hidden Markov Model (PHMM) is then built on such descriptors to represent a class of similar shapes. PHMMs are a particular type of Hidden Markov Models (HMMs) with special states and architecture that can tolerate considerable shape contour perturbations, including rigid and non-rigid deformations, occlusions, and missing parts. The sparseness of the PHMM structure provides efficient inference and learning algorithms for shape modeling and analysis. To capture the global characteristics of a class of shapes, the PHMM parameters are further embedded into a subspace that models long term spatial dependencies. The new framework can be applied to a wide range of problems, such as shape matching/registration, classification/recognition, etc. Our experimental results demonstrate the effectiveness and robustness of this new model in these different settings.
Rui Huang 0001, Vladimir Pavlovic 0001, Dimitris N. Metaxas
ICCV3
2007 Spectral Latent Variable Models for Perceptual Inference
abstract
We propose non-linear generative models referred to as Sparse Spectral Latent Variable Models (SLVM), that combine the advantages of spectral embeddings with the ones of parametric latent variable models: (1) provide stable latent spaces that preserve global or local geometric properties of the modeled data; (2) offer low-dimensional generative models with probabilistic, bi-directional mappings between latent and ambient spaces, (3) are probabilistically consistent (i.e., reflect the data distribution, both jointly and marginally) and efficient to learn and use. We show that SLVMs compare favorably with competing methods based on PCA, GPLVM or GTM for the reconstruction of typical human motions like walking, running, pantomime or dancing in a benchmark dataset. Empirically, we observe that SLVMs are effective for the automatic 3d reconstruction of low-dimensional human motion in movies.
Atul Kanaujia, Cristian Sminchisescu, Dimitris N. Metaxas
ICCV3
2007 Using the Pn Potts model with learning methods to segment live cell images
abstract
We present a segmentation method for live cell images, using graph cuts and learning methods. The images used here are particularly challenging because of the shared grey-level distributions of cells and background, which only differ by their textures, and the local imprecision around cell borders. We use the PnPotts model recently presented by Kohli et al. [9]: functions on higher-order cliques of pixels are included into the traditional Potts model, allowing us to account for local texture features, and to find the optimal solution efficiently. We use learning methods to define the potential functions used in the PnPotts model. We present the model and the learning methods we used, and compare our segmentation results with similar work in cytometry. While our method performs similarly, it requires little manual tuning and thus is straightforward to adapt to other images.
Chris Russell 0001, Dimitris N. Metaxas, Christophe Restif, Philip Torr 0001
ICCV2
2007 Coupling CRFs and Deformable Models for 3D Medical Image Segmentation
abstract
In this paper we present a hybrid probabilistic framework for 3D image segmentation, using Conditional Random Fields (CRFs) and implicit deformable models. Our 3D deformable model uses voxel intensity and higher scale textures as data-driven terms, while the shape is formulated implicitly using the Euclidean distance transform. The data-driven terms are used as observations in a 3D discriminative CRF, which drives the model evolution based on a simple graphical model. In this way, we solve the model evolution as a joint MAP estimation problem for the 3D label field of the CRF and the 3D shape of the deformable model. We demonstrate the performance of our approach in the estimation of the volume of the human tear menisci from images obtained with optical coherence tomography.
Gavriil Tsechpenakis, Jianhua Wang 0002, Brandon Mayer, Dimitris N. Metaxas
ICCV4
2007 The Best of Both Worlds: Combining 3D Deformable Models with Active Shape Models
abstract
Reliable 3D tracking is still a difficult task. Most parametrized 3D deformable models rely on the accurate extraction of image features for updating their parameters, and are prone to failures when the underlying feature distribution assumptions are invalid. Active Shape Models (ASMs), on the other hand, are based on learning, and thus require fewer reliable local image features than parametrized 3D models, but fail easily when they encounter a situation for which they were not trained. In this paper, we develop an integrated framework that combines the strengths of both 3D deformable models and ASMs. The 3D model governs the overall shape, orientation and location, and provides the basis for statistical inference on both the image features and the parameters. The ASMs, in contrast, provide the majority of reliable 2D image features over time, and aid in recovering from drift and total occlusions. The framework dynamically selects among different ASMs to compensate for large viewpoint changes due to head rotations. This integration allows the robust tracking effaces and the estimation of both their rigid and non- rigid motions. We demonstrate the strength of the framework in experiments that include automated 3D model fitting and facial expression tracking for a variety of applications, including sign language.
Christian Vogler, Atul Kanaujia, Siome Goldenstein, Dimitris N. Metaxas
ICCV5
2007 Large Scale Learning of Active Shape Models
abstract
We propose a framework to learn statistical shape models for faces as piecewise linear models. Specifically, our methodology builds upon primitive active shape models(ASM) to handle large scale variation in shapes and appearances of faces. Non-linearities in shape manifold arising due to large head rotation cannot be accurately modeled using ASM. Moreover overly general image descriptor causes the cost function to have multiple local minima which in turn degrades the quality of shape registration. We propose to use multiple overlapping subspaces with more discriminative local image descriptors to capture larger variance occurring in the data set. We also apply techniques to learn distance metric for enhancing similarity of descriptors belonging to the same class of shape subspace. Our generic algorithm can be applied to large scale shape analysis and registration.
Atul Kanaujia, Dimitris N. Metaxas
ICIP (1)2
2007 Facial Expression Recognition using Encoded Dynamic Features
abstract
In this paper, we propose a new approach of facial expression recognition. In order to capture the temporal characteristic of facial expressions, we design dynamic haar-like features to represent the facial images, and code them into binary patterns for the further analysis. Based on the encoded features, Adaboost is employed to learn the combination of optimal discriminant features to construct the classifier. The experiments carried on the CMU database show the promising performance of the proposed method.
Peng Yang 0001, Qingshan Liu 0001, Dimitris N. Metaxas
ICME3
2007 Registration of Lung Tissue Between Fluoroscope and CT Images: Determination of Beam Gating Parameters in Radiotherapy
Sukmoon Chang, Jinghao Zhou, Qingshan Liu 0001, Dimitris N. Metaxas, Bruce G. Haffty, Sung N. Kim, Salma J. Jabbour, Ning J. Yue
MICCAI (1)4
2007 Adaptive Metamorphs Model for 3D Medical Image Segmentation
Junzhou Huang, Sharon X. Huang, Dimitris N. Metaxas, Leon Axel
MICCAI (1)3
2007 Ultrasound Myocardial Elastography and Registered 3D Tagged MRI: Quantitative Strain Comparison
Wei-Ning Lee, Elisa E. Konofagou, Dimitris N. Metaxas, Leon Axel
MICCAI (1)4
2007 Facial Features Tracking for Gross Head Movement analysis and Expression Recognition
abstract
Summary form only given. The tracking and recognition of facial expressions from a single cameras is an important and challenging problem. We present a real-time framework for Action Units(AU)/Expression recognition based on facial features tracking and Adaboost. Accurate facial feature tracking is challenging due to changes in illumination, skin color variations, possible large head rotations, partial occlusions and fast head movements. We use models based on Active Shapes to localize facial features on the face in a generic pose. Shapes of facial features undergo non-linear transformation as the head rotates from frontal view to profile view. We learn the non-linear shape manifold as multiple-overlapping subspaces with different subspaces representing different head poses. The face alignment is done by searching over the non-linear shape manifold and aligning the landmark points to the features' boundaries. The recognized features are tracked across multiple frames using KLT Tracker by constraining the shape to lie on the non-linear manifold. Our tracking framework has been successfully used for detecting both gross head movements, like nodding, shaking and head pose prediction. Further, we use the tracked features to accurately extract bounded faces in a video sequence and use it for recognizing facial expressions. Our approach is based on coded dynamical features. In order to capture the dynamic characteristics of facial events, we design the dynamic haar-like features to represent the temporal variations of facial events. Inspired by the binary pattern coding, we further encode the dynamic haar-like features into binary pattern features, which are useful to construct weak classifiers for boosting learning. Finally Adaboost is used to learn a set of discriminating coded dynamic features for facial active units and expression recognition. We have achieved approximately 97% detection rate for gross head movements like shaking and nodding. The recognition rates for facial expressions averages to -95% for the most important action units.
Dimitris N. Metaxas
MMSP1
2007 Nonlinear Dynamic Shape and Appearance Models for Facial Motion Tracking
Chan-Su Lee, Ahmed M. Elgammal, Dimitris N. Metaxas
PSIVT3
2007 Textured Liquids based on the Marker Level Set
abstract
Abstract In this work we propose a new Eulerian method for handling the dynamics of a liquid and its surface attributes (for example its color). Our approach is based on a new method for interface advection that we term the Marker Level Set (MLS). The MLS method uses surface markers and a level set for tracking the surface of the liquid, yielding more efficient and accurate results than popular methods like the Particle Level Set method (PLS). Another novelty is that the surface markers allow the MLS to handle non‐diffusively surface texture advection, a rare capability in the realm of Eulerian simulation of liquids. We present several simulations of the dynamical evolution of liquids and their surface textures.
Viorel Mihalef, Dimitris N. Metaxas, Mark Sussman
Comput. Graph. Forum2
2007 Outlier rejection in high-dimensional deformable models
Christian Vogler, Siome Goldenstein, Jorge Stolfi, Vladimir Pavlovic 0001, Dimitris N. Metaxas
Image Vis. Comput.5
2007 Human gait recognition at sagittal plane
Christian Vogler, Dimitris N. Metaxas
Image Vis. Comput.3
2007 BM3E : Discriminative Density Propagation for Visual Tracking
abstract
We introduce BM3 E, a Conditional Bayesian Mixture of Experts Markov Model, for consistent probabilistic estimates in discriminative visual tracking. The model applies to problems of temporal and uncertain inference and represents the unexplored bottom-up counterpart of pervasive generative models estimated with Kalman filtering or particle filtering. Instead of inverting a non-linear generative observation model at run-time, we learn to cooperatively predict complex state distributions directly from descriptors that encode image observations - typically bag-of-feature global image histograms or descriptors computed over regular spatial grids. These are integrated in a conditional graphical model in order to enforce temporal smoothness constraints and allow a principled management of uncertainty. The algorithms combine sparsity, mixture modeling, and non-linear dimensionality reduction for efficient computation in high-dimensional continuous state spaces. The combined system automatically self-initializes and recovers from failure. The research has three contributions: (1) We establish the density propagation rules for discriminative inference in continuous, temporal chain models; (2) We propose flexible supervised and unsupervised algorithms for learning feedforward, multivalued contextual mappings (multimodal state distributions) based on compact, conditional Bayesian mixture of experts models; (3) We validate the framework empirically for the reconstruction of 3d human motion in monocular video sequences. Our tests on both real and motion capture-based sequences show significant performance gains with respect to competing nearest-neighbor, regression, and structured prediction methods.
Cristian Sminchisescu, Atul Kanaujia, Dimitris N. Metaxas
IEEE Trans. Pattern Anal. Mach. Intell.3
2007 A combustion-based technique for fire animation and visualization
Kyungha Min, Dimitris N. Metaxas
Vis. Comput.2
2006 Learning Multi-category Classification in Bayesian Framework
Atul Kanaujia, Dimitris N. Metaxas
ACCV (1)2
2006 RO-SVM: Support Vector Machine with Reject Option for Image Categorization
abstract
When applying Multiple Instance Learning (MIL) for image categorization, an image is treated as a bag containing a number of instances, each representing a region inside the image. The categorization of this image is determined by the labels of these instances, which are not specified in the training data-set. Hence, these instance labels are needed to be estimated together with the classifier. To improve classification reliability, we propose in this paper a new Support Vector Machine approach by incorporating a reject option, named RO-SVM to determine the instance labels, and the rejection region during the training phase simultaneously. Our approach can also be easily extended to solve multi-class classification problems. Experimental results demonstrate that higher categorization accuracy can be achieved with our RO-SVM method, comparing to approaches that do not exclude uninformative image patches. Our method is able to produce results comparable even with few training samples. 1
Dimitris N. Metaxas
BMVC2
2006 Active Contours with Level-Set for Extracting Feature Curves from Triangular Meshes
Kyungha Min, Dimitris N. Metaxas, Moon-Ryul Jung
Computer Graphics International2
2006 Learning Joint Top-Down and Bottom-up Processes for 3D Visual Inference
abstract
We present an algorithm for jointly learning a consistent bidirectional generative-recognition model that combines top-down and bottom-up processing for monocular 3d human motion reconstruction. Learning progresses in alternative stages of self-training that optimize the probability of the image evidence: the recognition model is tunned using samples from the generative model and the generative model is optimized to produce inferences close to the ones predicted by the current recognition model. At equilibrium, the two models are consistent. During on-line inference, we scan the image at multiple locations and predict 3d human poses using the recognition model. But this implicitly includes one-shot generative consistency feedback. The framework provides a uniform treatment of human detection, 3d initialization and 3d recovery from transient failure. Our experimental results show that this procedure is promising for the automatic reconstruction of human motion in more natural scene settings with background clutter and occlusion.
Cristian Sminchisescu, Atul Kanaujia, Dimitris N. Metaxas
CVPR (2)3
2006 Patch-Based Texture Edges and Segmentation
Lior Wolf, Sharon X. Huang, Ian Martin, Dimitris N. Metaxas
ECCV (2)4
2006 Hybrid Deformable Models for Medical Segmentation and Registration
abstract
Deformable models have had great successes over the past 20 years in medical applications. We have recently developed new classes of deformable models which we term hybrid deformable models to automate the model initialization process and make improvements in segmentation and registration. In this paper we present several hybrid deformable methods we have been developing for segmentation and registration. These methods include metamorphs, a novel shape and texture integration deformable model framework and the integration of deformable models with graphical models and learning methods. We first present a framework for the robust segmentation and tracking of the heart from tagged MRI images and second applications involving brain tumor segmentation as well as brain and cardiac shape registration
Dimitris N. Metaxas, Sharon X. Huang, Rui Huang 0001, Ting Chen 0001, Leon Axel
ICARCV1
2006 A Profile Hidden Markov Model Framework for Modeling and Analysis of Shape
abstract
In this paper we propose a new framework for modeling 2D shapes. A shape is first described by a sequence of local features (e.g., curvature) of the shape boundary. The resulting description is then used to build a profile hidden Markov model (PHMM) representation of the shape. PHMMs are a particular type of hidden Markov models (HMMs) with special states and architecture that can tolerate considerable shape contour perturbations, including rigid and non-rigid deformations, occlusions and missing contour parts. Different from traditional HMM-based shape models, the sparseness of the PHMM structure allows efficient inference and learning algorithms for shape modeling and analysis. The new framework can be applied to a wide range of problems, from shape matching and classification to shape segmentation. Our experimental results show the effectiveness and robustness of this new approach in the three application domains.
Rui Huang 0001, Vladimir Pavlovic 0001, Dimitris N. Metaxas
ICIP3
2006 Synthesis and Control of High Resolution Facial Expressions for Visual Interactions
abstract
The synthesis of facial expression with control of intensity and personal styles is important in intelligent and affective human-computer interaction, especially in face-to-face interaction between human and intelligent agent. We present a facial expression animation system that facilitates control of expressiveness and style. We learn a decomposable generative model for the nonlinear deformation of facial expressions by analyzing the mapping space between low dimensional embedded representation and high resolution tracking data. Bilinear analysis of the mapping space provides a compact representation of the nonlinear generative model for facial expressions. The decomposition allows synthesis of new facial expressions by control of geometry and expression style. The generative model provides control of expressiveness preserving nonlinear deformation in the expressions with simple parameters and allows synthesis of stylized facial geometry. In addition, we can directly extract the MPEG-4 facial animation parameters (FAPs) from the synthesized data, which allows using any animation engine that supports FAPs to animate new synthesized expressions
Chan-Su Lee, Ahmed M. Elgammal, Dimitris N. Metaxas
ICME3
2006 Learning Ambiguities Using Bayesian Mixture of Experts
abstract
Mixture of Experts (ME) is an ensemble of function approximators that fit the clustered data set locally rather than globally. ME provides a useful tool to learn multi-valued mappings (ambiguities) in the data set. Mixture of Experts training involve learning a multi-category classifier for the gates distribution and fitting a regressor within each of the clusters. The learning of ME is based on divide and conquer which is known to increase the error due to variance. In order to avoid overfitting several researchers have proposed using linear experts. However in the absence of any knowledge of non-linearities existing in the data set, it is not clear how many linear experts could accurately model the data. In this work we propose a Bayesian learning framework for learning Mixture of Experts. Bayesian learning intrinsically embodies regularization and model selection using Occam's razor. In the past Bayesian learning methods have been applied to classification and regression in order to avoid scale sensitivity and orthodox model selection procedure of cross validation. Although true Bayesian learning is computationally intractable, approximations do result in sparser and more compact models
Atul Kanaujia, Dimitris N. Metaxas
ICTAI2
2006 Boosting and Nonparametric Based Tracking of Tagged MRI Cardiac Boundaries
Dimitris N. Metaxas, Leon Axel
MICCAI (1)2
2006 Automatic Detection and Segmentation of Ground Glass Opacity Nodules
Jinghao Zhou, Sukmoon Chang, Dimitris N. Metaxas, Binsheng Zhao, Lawrence H. Schwartz, Michelle S. Ginsberg
MICCAI (1)3
2006 Conditional models for contextual human motion recognition
Cristian Sminchisescu, Atul Kanaujia, Dimitris N. Metaxas
Comput. Vis. Image Underst.3
2006 Learning-based dynamic coupling of discrete and continuous trackers
Gavriil Tsechpenakis, Dimitris N. Metaxas, Carol Neidle
Comput. Vis. Image Underst.2
2006 Shape Registration in Implicit Spaces Using Information Theory and Free Form Deformations
abstract
We present a novel, variational and statistical approach for shape registration. Shapes of interest are implicitly embedded in a higher-dimensional space of distance transforms. In this implicit embedding space, registration is formulated in a hierarchical manner: the Mutual Information criterion supports various transformation models and is optimized to perform global registration; then, a B-spline-based Incremental Free Form Deformations (IFFD) model is used to minimize a Sum-of-Squared-Differences (SSD) measure and further recover a dense local nonrigid registration field. The key advantage of such framework is twofold: (1) it naturally deals with shapes of arbitrary dimension (2D, 3D, or higher) and arbitrary topology (multiple parts, closed/open) and (2) it preserves shape topology during local deformation and produces local registration fields that are smooth, continuous, and establish one-to-one correspondences. Its invariance to initial conditions is evaluated through empirical validation, and various hard 2D/3D geometric shape registration examples are used to show its robustness to noise, severe occlusion, and missing parts. We demonstrate the power of the proposed framework using two applications: one for statistical modeling of anatomical structures, another for 3D face scan registration and expression tracking. We also compare the performance of our algorithm with that of several other well-known shape registration algorithms.
Sharon X. Huang, Nikos Paragios, Dimitris N. Metaxas
IEEE Trans. Pattern Anal. Mach. Intell.3
2006 A collision resolution algorithm for clump-free fast moving cloth
Suejung B. Huh, Dimitris N. Metaxas
Vis. Comput.2
2005 A collision resolution algorithm for clump-free fast moving cloth
abstract
Cloth animation is an important area of computer graphics due to its numerous applications. However, so far an animation of fast moving cloth with multiple wrinkles has been difficult to make because of the cloth clump problem. Cloth clumps are the frozen areas where cloth pieces are clustered unnaturally - an obstacle in making a realistic cloth animation. Hence we present a novel cloth collision resolution algorithm which prevents clump formation during fast cloth motions. The goal of our resolution algorithm is to let cloth move swiftly without having any frozen cloth clumps, while preventing any cloth-cloth and rigid-cloth penetrations at any moment of a simulation. The non-penetration status of cloth is maintained without the formation of cloth clumps regardless of the speed of cloth motion. Our algorithm is based on a novel order in the resolution of cloth collisions. Our algorithm can be used with any cloth simulation approach such as spring-masses or finite elements. For the particular types of applications presented in this paper, we use a spring-mass system and we show several realistic simulations involving fast motions that are clump-free.
Suejung B. Huh, Dimitris N. Metaxas
Computer Graphics International2
2005 Discriminative Density Propagation for 3D Human Motion Estimation
abstract
We describe a mixture density propagation algorithm to estimate 3D human motion in monocular video sequences based on observations encoding the appearance of image silhouettes. Our approach is discriminative rather than generative, therefore it does not require the probabilistic inversion of a predictive observation model. Instead, it uses a large human motion capture data-base and a 3D computer graphics human model in order to synthesize training pairs of typical human configurations together with their realistically rendered 2D silhouettes. These are used to directly learn to predict the conditional state distributions required for 3D body pose tracking and thus avoid using the generative 3D model for inference (the learned discriminative predictors can also be used, complementary, as importance samplers in order to improve mixing or initialize generative inference algorithms). We aim for probabilistically motivated tracking algorithms and for models that can represent complex multivalued mappings common in inverse, uncertain perception inferences. Our paper has three contributions: (1) we establish the density propagation rules for discriminative inference in continuous, temporal chain models; (2) we propose flexible algorithms for learning multimodal state distributions based on compact, conditional Bayesian mixture of experts models; and (3) we demonstrate the algorithms empirically on real and motion capture-based test sequences and compare against nearest-neighbor and regression methods.
Cristian Sminchisescu, Atul Kanaujia, Dimitris N. Metaxas
CVPR (1)4
2005 Conditional Random Fields for Contextual Human Motion Recognition
abstract
We present algorithms for recognizing human motion in monocular video sequences, based on discriminative conditional random field (CRF) and maximum entropy Markov models (MEMM). Existing approaches to this problem typically use generative (joint) structures like the hidden Markov model (HMM). Therefore they have to make simplifying, often unrealistic assumptions on the conditional independence of observations given the motion class labels and cannot accommodate overlapping features or long term contextual dependencies in the observation sequence. In contrast, conditional models like the CRFs seamlessly represent contextual dependencies, support efficient, exact inference using dynamic programming, and their parameters can be trained using convex optimization. We introduce conditional graphical models as complementary tools for human motion recognition and present an extensive set of experiments that show how these typically outperform HMMs in classifying not only diverse human activities like walking, jumping. running, picking or dancing, but also for discriminating among subtle motion styles like normal walk and wander walk
Cristian Sminchisescu, Atul Kanaujia, Dimitris N. Metaxas
ICCV4
2005 HMM-Based Deception Recognition from Visual Cues
abstract
Behavioral indicators of deception and behavioral state are extremely difficult for humans to analyze. This research effort attempts to leverage automated systems to augment humans in detecting deception by analyzing nonverbal behavior on video. By tracking faces and hands of an individual, it is anticipated that objective behavioral indicators of deception can be isolated, extracted and synthesized to create a more accurate means for detecting human deception. Blob analysis, a method for analyzing the movement of the head and hands based on the identification of skin color is presented. A proof-of-concept study is presented that uses Blob analysis to extract visual cues and events, throughout the examined videos. The integration of these cues is done using a hierarchical hidden Markov model to explore behavioral state identification in the detection of deception, mainly involving the detection of agitated and over-controlled behaviors
Gavriil Tsechpenakis, Dimitris N. Metaxas, Mark Adkins, John Kruse, Judee K. Burgoon, Matthew L. Jensen, Thomas O. Meservy, Douglas P. Twitchell, Amit V. Deokar, Jay F. Nunamaker Jr.
ICME2
2005 Efficient Learning by Combining Confidence-Rated Classifiers to Incorporate Unlabeled Medical Data
Weijun He, Sharon X. Huang, Dimitris N. Metaxas, Xiaoyou Ying
MICCAI3
2005 Conditional Visual Tracking in Kernel Space
abstract
We present a conditional temporal probabilistic framework for recon- structing 3D human motion in monocular video based on descriptors en- coding image silhouette observations. For computational efficiency we restrict visual inference to low-dimensional kernel induced non-linear state spaces. Our methodology (kBME) combines kernel PCA-based non-linear dimensionality reduction (kPCA) and Conditional Bayesian Mixture of Experts (BME) in order to learn complex multivalued pre- dictors between observations and model hidden states. This is necessary for accurate, inverse, visual perception inferences, where several proba- ble, distant 3D solutions exist due to noise or the uncertainty of monoc- ular perspective projection. Low-dimensional models are appropriate because many visual processes exhibit strong non-linear correlations in both the image observations and the target, hidden state variables. The learned predictors are temporally combined within a conditional graphi- cal model in order to allow a principled propagation of uncertainty. We study several predictors and empirically show that the proposed algo- rithm positively compares with techniques based on regression, Kernel Dependency Estimation (KDE) or PCA alone, and gives results competi- tive to those of high-dimensional mixture predictors at a fraction of their computational cost. We show that the method successfully reconstructs the complex 3D motion of humans in real monocular video sequences. 1 Introduction and Related Work We consider the problem of inferring 3D articulated human motion from monocular video. This research topic has applications for scene understanding including human-computer in- terfaces, markerless human motion capture, entertainment and surveillance. A monocular approach is relevant because in real-world settings the human body parts are rarely com- pletely observed even when using multiple cameras. This is due to occlusions form other people or objects in the scene. A robust system has to necessarily deal with incomplete, ambiguous and uncertain measurements. Methods for 3D human motion reconstruction can be classified as generative and discriminative. They both require a state representation, namely a 3D human model with kinematics (joint angles) or shape (surfaces or joint po- sitions) and they both use a set of image features as observations for state inference. The computational goal in both cases is the conditional distribution for the model state given image observations. Generative model-based approaches [6, 16, 14, 13] have been demonstrated to flexibly re- construct complex unknown human motions and to naturally handle problem constraints. However it is difficult to construct reliable observation likelihoods due to the complexity of modeling human appearance. This varies widely due to different clothing and defor- mation, body proportions or lighting conditions. Besides being somewhat indirect, the generative approach further imposes strict conditional independence assumptions on the temporal observations given the states in order to ensure computational tractability. Due to these factors inference is expensive and produces highly multimodal state distributions [6, 16, 13]. Generative inference algorithms require complex annealing schedules [6, 13] or systematic non-linear search for local optima [16] in order to ensure continuing tracking. These difficulties motivate the advent of a complementary class of discriminative algo- rithms [10, 12, 18, 2], that approximate the state conditional directly, in order to simplify inference. However, inverse, observation-to-state multivalued mappings are difficult to learn (see e.g. fig. 1a) and a probabilistic temporal setting is necessary. In an earlier paper [15] we introduced a probabilistic discriminative framework for human motion reconstruc- tion. Because the method operates in the originally selected state and observation spaces that can be task generic, therefore redundant and often high-dimensional, inference is more expensive and can be less robust. To summarize, reconstructing 3D human motion in a Figure 1: (a, Left) Example of 180o ambiguity in predicting 3D human poses from sil- houette image features (center). It is essential that multiple plausible solutions (e.g. F1 and F2) are correctly represented and tracked over time. A single state predictor will either average the distant solutions or zig-zag between them, see also tables 1 and 2. (b, Right) A conditional chain model. The local distributions p(ytjyt(cid:0)1; zt) or p(ytjzt) are learned as in fig. 2. For inference, the predicted local state conditional is recursively combined with the filtered prior c.f . (1). conditional temporal framework poses the following difficulties: (i) The mapping between temporal observations and states is multivalued (i.e. the local conditional distributions to be learned are multimodal), therefore it cannot be accurately represented using global function approximations. (ii) Human models have multivariate, high-dimensional continuous states of 50 or more human joint angles. The temporal state conditionals are multimodal which makes efficient Kalman filtering algorithms inapplicable. General inference methods (par- ticle filters, mixtures) have to be used instead, but these are expensive for high-dimensional models (e.g. when reconstructing the motion of several people that operate in a joint state space). (iii) The components of the human state and of the silhouette observation vector ex- hibit strong correlations, because many repetitive human activities like walking or running have low intrinsic dimensionality. It appears wasteful to work with high-dimensional states of 50+ joint angles. Even if the space were truly high-dimensional, predicting correlated state dimensions independently may still be suboptimal. In this paper we present a conditional temporal estimation algorithm that restricts visual inference to low-dimensional, kernel induced state spaces. To exploit correlations among observations and among state variables, we model the local, temporal conditional distri- butions using ideas from Kernel PCA [11, 19] and conditional mixture modeling [7, 5], here adapted to produce multiple probabilistic predictions. The corresponding predictor is referred to as a Conditional Bayesian Mixture of Low-dimensional Kernel-Induced Experts (kBME). By integrating it within a conditional graphical model framework (fig. 1b), we can exploit temporal constraints probabilistically. We demonstrate that this methodology is effective for reconstructing the 3D motion of multiple people in monocular video. Our con- tribution w.r.t. [15] is a probabilistic conditional inference framework that operates over a non-linear, kernel-induced low-dimensional state spaces, and a set of experiments (on both real and artificial image sequences) that show how the proposed framework positively com- pares with powerful predictors based on KDE, PCA, or with the high-dimensional models of [15] at a fraction of their cost. 2 Probabilistic Inference in a Kernel Induced State Space We work with conditional graphical models with a chain structure [9], as shown in fig. 1b, These have continuous temporal states yt, t = 1 : : : T , observations zt. For compactness, we denote joint states Yt = (y1; y2; : : : ; yt) or joint observations Zt = (z1; : : : ; zt). Learning and inference are based on local conditionals: p(ytjzt) and p(ytjyt(cid:0)1; zt), with yt and zt being low-dimensional, kernel induced representations of some initial model having state xt and observation rt. We obtain zt; yt from rt, xt using kernel PCA [11, 19]. Inference is performed in a low-dimensional, non-linear, kernel induced latent state space (see fig. 1b and fig. 2 and (1)). For display or error reporting, we compute the original conditional p(xjr), or a temporally filtered version p(xtjRt); Rt = (r1; r2; : : : ; rt), using a learned pre-image state map [3]. 2.1 Density Propagation for Continuous Conditional Chains For online filtering, we compute the optimal distribution p(ytjZt) for the state yt, con- ditioned by observations Zt up to time t. The filtered density can be recursively derived as: p(ytjZt) = Zyt(cid:0)1 p(ytjyt(cid:0)1; zt)p(yt(cid:0)1jZt(cid:0)1) (1) We compute using a conditional mixture for p(ytjyt(cid:0)1; zt) (a Bayesian mixture of experts c.f . x2.2) and the prior p(yt(cid:0)1jZt(cid:0)1), each having, say M components. We integrate M 2 pairwise products of Gaussians analytically. The means of the expanded posterior are clus- tered and the centers are used to initialize a reduced M-component Kullback-Leibler ap- proximation that is refined using gradient descent [15]. The propagation rule (1) is similar to the one used for discrete state labels [9], but here we work with multivariate continuous state spaces and represent the local multimodal state conditionals using kBME (fig. 2), and not log-linear models [9] (these would require intractable normalization). This complex continuous model rules out inference based on Kalman filtering or dynamic programming [9]. 2.2 Learning Bayesian Mixtures over Kernel Induced State Spaces (kBME) In order to model conditional mappings between low-dimensional non-linear spaces we rely on kernel dimensionality reduction and conditional mixture predictors. The authors of KDE [19] propose a powerful structured unimodal predictor. This works by decorrelating the output using kernel PCA and learning a ridge regressor between the input and each decorrelated output dimension. Our procedure is also based on kernel PCA but takes into account the structure of the studied visual problem where both inputs and outputs are likely to be low-dimensional and the mapping between them multivalued. The output variables xi are projected onto the column vectors of the principal space in order to obtain their principal coordinates yi. A
Cristian Sminchisescu, Atul Kanaujia, Dimitris N. Metaxas
NIPS4
2005 Computational modeling and simulation of heart ventricular mechanics with tagged MRI
abstract
Heart ventricular mechanics has been investigated intensively in the last four decades. The passive material properties, the ventricular geometry and muscular architecture, and the myocardial activation are among the most important determinants of cardiac mechanics. The heart muscle is anisotropic, inhomogeneous, and highly nonlinear. The heart ventricular geometry is irregular and object dependent. The muscular architecture includes the organization of the fiber and the connective tissues. Studies of the myocardial activation have been carried out at both cell and tissue levels.Previous work from our research group has successfully estimated the in-vivo motion and deformation of both the left and the right ventricles. In this paper, we present an iterative model to estimate the in-vivo myocardium material properties, the active forces generated along fiber orientation, and strain and stress distribution in both ventricles. Compared to the strain energy function approach, our model is more intuitively understandable. Using the model, we have simulated the mechanical events of a few different heart diseases. Noticeable strain and stress differences are found between normal and diseased hearts.
Zhenhua Hu, Dimitris N. Metaxas, Leon Axel
Symposium on Solid and Physical Modeling2
2005 A hybrid framework for 3D medical image segmentation
Ting Chen 0001, Dimitris N. Metaxas
Medical Image Anal.2
2005 Open science - combining open data and open source software: Medical image analysis with the Insight Toolkit
Terry S. Yoo, Dimitris N. Metaxas
Medical Image Anal.2
2005 Incremental Model-Based Estimation Using Geometric Constraints
abstract
We present a model-based framework for incremental, adaptive object shape estimation and tracking in monocular image sequences. Parametric structure and motion estimation methods usually assume a fixed class of shape representation (splines, deformable superquadrics, etc.) that is initialized prior to tracking. Since the model shape coverage is fixed a priori, the incremental recovery of structure is decoupled from tracking, thereby limiting both processes in their scope and robustness. In this work, we describe a model-based framework that supports the automatic detection and integration of low-level geometric primitives (lines) incrementally. Such primitives are not explicitly captured in the initial model, but are moving consistently with its image motion. The consistency tests used to identify new structure are based on trinocular constraints between geometric primitives. The method allows not only an increase in the model scope, but also improves tracking accuracy by including the newly recovered features in its state estimation. The formulation is a step toward automatic model building, since it allows both weaker assumptions on the availability of a prior shape representation and on the number of features that would otherwise be necessary for entirely bottom-up reconstruction. We demonstrate the proposed approach on two separate image-based tracking domains, each involving complex 3D object structure and motion.
Cristian Sminchisescu, Dimitris N. Metaxas, Sven J. Dickinson
IEEE Trans. Pattern Anal. Mach. Intell.2
2005 Automated detection of prostatic adenocarcinoma from high-resolution ex vivo MRI
abstract
Prostatic adenocarcinoma is the most commonly occurring cancer among men in the United States, second only to skin cancer. Currently, the only definitive method to ascertain the presence of prostatic cancer is by trans-rectal ultrasound (TRUS) directed biopsy. Owing to the poor image quality of ultrasound, the accuracy of TRUS is only 20%-25%. High-resolution magnetic resonance imaging (MRI) has been shown to have a higher accuracy of prostate cancer detection compared to ultrasound. Consequently, several researchers have been exploring the use of high resolution MRI in performing prostate biopsies. Visual detection of prostate cancer, however, continues to be difficult owing to its apparent lack of shape, and the fact that several malignant and benign structures have overlapping intensity and texture characteristics. In this paper, we present a fully automated computer-aided detection (CAD) system for detecting prostatic adenocarcinoma from 4 Tesla ex vivo magnetic resonance (MR) imagery of the prostate. After the acquired MR images have been corrected for background inhomogeneity and nonstandardness, novel three-dimensional (3-D) texture features are extracted from the 3-D MRI scene. A Bayesian classifier then assigns each image voxel a "likelihood" of malignancy for each feature independently. The "likelihood" images generated in this fashion are then combined using an optimally weighted feature combination scheme. Quantitative evaluation was performed by comparing the CAD results with the manually ascertained ground truth for the tumor on the MRI. The tumor labels on the MR slices were determined manually by an expert by visually registering the MR slices with the corresponding regions on the histology slices. We evaluated our CAD system on a total of 33 two-dimensional (2-D) MR slices from five different 3-D MR prostate studies. Five slices from two different glands were used for training. Our feature combination scheme was found to outperform the individual texture features, and also other popularly used feature combination methods, including AdaBoost, ensemble averaging, and majority voting. Further, in several instances our CAD system performed better than the experts in terms of accuracy, the expert segmentations being determined solely from visual inspection of the MRI data. In addition, the intrasystem variability (changes in CAD accuracy with changes in values of system parameters) was significantly lower than the corresponding intraobserver and interobserver variability. CAD performance was found to be very similar for different training sets. Future work will focus on extending the methodology to guide high-resolution MRI-assisted in vivo prostate biopsies.
Anant Madabhushi, Michael D. Feldman, Dimitris N. Metaxas, John Tomaszewski 0001, Deborah Chute
IEEE Trans. Medical Imaging3
2004 3D Facial Tracking from Corrupted Movie Sequences
Siome Goldenstein, Christian Vogler, Dimitris N. Metaxas
CVPR (1)3
2004 MetaMorphs: Deformable Shape and Texture Models
Sharon X. Huang, Dimitris N. Metaxas, Ting Chen 0001
CVPR (1)2
2004 A Graphical Model Framework for Coupling MRFs and Deformable Models
Rui Huang 0001, Vladimir Pavlovic 0001, Dimitris N. Metaxas
CVPR (2)3
2004 Distinguishing Mislabeled Data from Correctly Labeled Data in Classifier Design
abstract
We have developed a method for distinguishing between correctly labeled and mislabeled data sampled from video sequences and used in the construction of a facial expression recognition classifier. The novelty of our approach lies in training a single, optimal classifier type (a support vector machine, or SVM) on multiple representations of the data, involving different "discriminating" subspaces. Results of a preliminary study on the discrimination of "high stress" vs. "low stress" facial expression data by this method confirms that our novel approach is able to distinguish subproblems where labeling is highly reliable from those where mislabeling can lead to high error rates. In helping detect data subsamples which yield misleading classification results, the method is also a rapid, highly efficient cross-validated approach for eliminating outliers.
Sundara Venkataraman, Dimitris N. Metaxas, Dmitriy Fradkin, Casimir A. Kulikowski, Ilya B. Muchnik
ICTAI2
2004 Pulmonary Micronodule Detection from 3D Chest CT
Sukmoon Chang, Hirosh Emoto, Dimitris N. Metaxas, Leon Axel
MICCAI (2)3
2004 3D Cardiac Anatomy Reconstruction Using High Resolution CT Data
Ting Chen 0001, Dimitris N. Metaxas, Leon Axel
MICCAI (1)2
2004 Learning Coupled Prior Shape and Appearance Models for Segmentation
Sharon X. Huang, Dimitris N. Metaxas
MICCAI (1)3
2004 High Resolution Acquisition, Learning and Transfer of Dynamic 3D Facial Expressions
abstract
Abstract Synthesis and re‐targeting of facial expressions is central to facial animation and often involves significant manual work in order to achieve realistic expressions, due to the difficulty of capturing high quality dynamic expression data. In this paper we address fundamental issues regarding the use of high quality dense 3‐D data samples undergoing motions at video speeds, e.g. human facial expressions. In order to utilize such data for motion analysis and re‐targeting, correspondences must be established between data in different frames of the same faces as well as between different faces. We present a data driven approach that consists of four parts: 1) High speed, high accuracy capture of moving faces without the use of markers, 2) Very precise tracking of facial motion using a multi‐resolution deformable mesh, 3) A unified low dimensional mapping of dynamic facial motion that can separate expression style, and 4) Synthesis of novel expressions as a combination of expression styles. The accuracy and resolution of our method allows us to capture and track subtle expression details. The low dimensional representation of motion data in a unified embedding for all the subjects in the database allows for learning the most discriminating characteristics of each individual's expressions as that person's “expression style”. Thus new expressions can be synthesized, either as dynamic morphing between individuals, or as expression transfer from a source face to a target face, as demonstrated in a series of experiments. Categories and Subject Descriptors (according to ACM CCS): I.3.7 [Computer Graphics]: Animation; I.3.5 [Computer Graphics]: Curve, surface, solid, and object representations; I.3.3 [Computer Graphics]: Digitizing and scanning; I.2.10 [Artificial intelligence]: Motion ; I.2.10 [Artificial intelligence]: Representations, data structures, and transforms; I.2.10 [Artificial intelligence]: Shape; I.2.6 [Artificial intelligence]: Concept learning
Yang Wang 0001, Sharon X. Huang, Chan-Su Lee, Song Zhang 0002, Dimitris Samaras, Dimitris N. Metaxas, Ahmed M. Elgammal, Peisen Huang
Comput. Graph. Forum7
2003 Using Multiple Cues for Hand Tracking and Model Refinement
abstract
We present a model based approach to the integration of multiple cues for tracking high degree of freedom articulated motions and model refinement. We then apply it to the problem of hand tracking using a single camera sequence. Hand tracking is particularly challenging because of occlusions, shading variations, and the high dimensionality of the motion. The novelty of our approach is in the combination of multiple sources of information, which come from edges, optical flow, and shading information in order to refine the model during tracking. We first use a previously formulated generalized version of the gradient-based optical flow constraint, that includes shading flow i.e., the variation of the shading of the object as it rotates with respect to the light source. Using this model we track its complex articulated motion in the presence of shading changes. We use a forward recursive dynamic model to track the motion in response to data derived 3D forces applied to the model. However, due to inaccurate initial shape, the generalized optical flow constraint is violated. We use the error in the generalized optical flow equation to compute generalized forces that correct the model shape at each step. The effectiveness of our approach is demonstrated with experiments on a number of different hand motions with shading changes, rotations and occlusions of significant parts of the hand.
Shan Lu 0010, Dimitris N. Metaxas, Dimitris Samaras, John Oliensis
CVPR (2)2
2003 Scan-Conversion Algorithm for Ridge Point Detection on Tubular Objects
Sukmoon Chang, Dimitris N. Metaxas, Leon Axel
MICCAI (2)2
2003 Gibbs Prior Models, Marching Cubes, and Deformable Models: A Hybrid Framework for 3D Medical Image Segmentation
Ting Chen 0001, Dimitris N. Metaxas
MICCAI (2)2
2003 Establishing Local Correspondences towards Compact Representations of Anatomical Structures
Sharon X. Huang, Nikos Paragios, Dimitris N. Metaxas
MICCAI (2)3
2003 A Novel Stochastic Combination of 3D Texture Features for Automated Segmentation of Prostatic Adenocarcinoma from High Resolution MRI
Anant Madabhushi, Michael D. Feldman, Dimitris N. Metaxas, Deborah Chute, John Tomaszewski 0001
MICCAI (1)3
2003 Automated Model-Based Segmentation of the Left and Right Ventricles in Tagged Cardiac MRI
Albert Montillo, Dimitris N. Metaxas, Leon Axel
MICCAI (1)2
2003 A Finite Element Model for Functional Analysis of 4D Cardiac-Tagged MR Images
Kyoungju Park, Dimitris N. Metaxas, Leon Axel
MICCAI (1)2
2003 In vivo strain and stress estimation of the heart left and right ventricles from MRI images
Zhenhua Hu, Dimitris N. Metaxas, Leon Axel
Medical Image Anal.2
2003 Statistical Cue Integration in DAG Deformable Models
abstract
Deformable models are a useful modeling paradigm in computer vision. A deformable model is a curve, a surface, or a volume, whose shape, position, and orientation are controlled through a set of parameters. They can represent manufactured objects, human faces and skeletons, and even bodies of fluid. With low-level computer vision and image processing techniques, such as optical flow, we extract relevant information from images. Then, we use this information to change the parameters of the model iteratively until we find a good approximation of the object in the images. When we have multiple computer vision algorithms providing distinct sources of information (cues), we have to deal with the difficult problem of combining these, sometimes conflicting contributions in a sensible way. In this paper, we introduce the use of a directed acyclic graph (DAG) to describe the position and Jacobian of each point of deformable models. This representation is dynamic, flexible, and allows computational optimizations that would be difficult to do otherwise. We then describe a new method for statistical cue integration method for tracking deformable models that scales well with the dimension of the parameter space. We use affine forms and affine arithmetic to represent and propagate the cues and their regions of confidence. We show that we can apply the Lindeberg theorem to approximate each cue with a Gaussian distribution, and can use a maximum-likelihood estimator to integrate them. Finally, we demonstrate the technique at work in a 3D deformable face tracking system on monocular image sequences with thousands of frames.
Siome Goldenstein, Christian Vogler, Dimitris N. Metaxas
IEEE Trans. Pattern Anal. Mach. Intell.3
2003 Incorporating Illumination Constraints in Deformable Models for Shape from Shading and Light Direction Estimation
abstract
We present a method for the integration of nonlinear holonomic constraints in deformable models and its application to the problems of shape and illuminant direction estimation from shading. Experimental results demonstrate that our method performs better than previous Shape from Shading algorithms applied to images of Lambertian objects under known illumination. It is also more general as it can be applied to non-Lambertian surfaces and it does not require knowledge of the illuminant direction. In this paper, (1) we first develop a theory for the numerically robust integration of nonlinear holonomic constraints within a deformable model framework. In this formulation, we use Lagrange multipliers and a Baumgarte stabilization approach (1972). (2) We also describe a fast new method for the computation of constraint based forces, in the case of high numbers of local parameters. (3) We demonstrate how any type of illumination constraint, from the simple Lambertian model to more complex highly nonlinear models can be incorporated in a deformable model framework. (4) We extend our method to work when the direction of the light source is not known. We couple our shape estimation method with a method for light estimation, in an iterative process, where improved shape estimation results in improved light estimation and vice versa. (5) We perform a series of experiments.
Dimitris Samaras, Dimitris N. Metaxas
IEEE Trans. Pattern Anal. Mach. Intell.2
2003 Combining Low, High-Level and Empirical Domain Specific Knowledge for Automated Segmentation of Ultrasonic Breast Lesions
abstract
Breast cancer is the most frequently diagnosed malignancy and the second leading cause of mortality in women. In the last decade, ultrasound along with digital mammography has come to be regarded as the gold standard for breast cancer diagnosis. Automatically detecting tumors and extracting lesion boundaries in ultrasound images is difficult due to their specular nature and the variance in shape and appearance of sonographic lesions. Past work on automated ultrasonic breast lesion segmentation has not addressed important issues such as shadowing artifacts or dealing with similar tumor like structures in the sonogram. Algorithms that claim to automatically classify ultrasonic breast lesions, rely on manual delineation of the tumor boundaries. In this paper, we present a novel technique to automatically find lesion margins in ultrasound images, by combining intensity and texture with empirical domain specific knowledge along with directional gradient and a deformable shape-based model. The images are first filtered to remove speckle noise and then contrast enhanced to emphasize the tumor regions. For the first time, a mathematical formulation of the empirical rules used by radiologists in detecting ultrasonic breast lesions, popularly known as the "Stavros Criteria" is presented in this paper. We have applied this formulation to automatically determine a seed point within the image. Probabilistic classification of image pixels based on intensity and texture is followed by region growing using the automatically determined seed point to obtain an initial segmentation of the lesion. Boundary points are found on the directional gradient of the image. Outliers are removed by a process of recursive refinement. These boundary points are then supplied as an initial estimate to a deformable model. Incorporating empirical domain specific knowledge along with low and high-level knowledge makes it possible to avoid shadowing artifacts and lowers the chance of confusing similar tumor like structures for the lesion. The system was validated on a database of breast sonograms for 42 patients. The average mean boundary error between manual and automated segmentation was 6.6 pixels and the normalized true positive area overlap was 75.1%. The algorithm was found to be robust to 1) variations in system parameters, 2) number of training samples used, and 3) the position of the seed point within the tumor. Running time for segmenting a single sonogram was 18 s on a 1.8-GHz Pentium machine.
Anant Madabhushi, Dimitris N. Metaxas
IEEE Trans. Medical Imaging2
2002 A Hybrid Dynamical Systems Approach to Intelligent Low-Level Navigation
abstract
Animated characters may exhibit several kinds of dynamic intelligence when performing low-level navigation (i.e., navigation on a local perceptual scale): they decide among different modes of behavior selectively discriminate entities in the world around them, perform obstacle avoidance, etc. In this paper we present a hybrid dynamical system model of low-level navigation that accounts for the above-mentioned kinds of intelligence. In so doing, the model illustrates general ideas about how a hybrid systems perspective can influence and simplify such reactive/behavioral modeling for multi-agent systems. In addition, we directly employed our formal hybrid system model to generate animations that illustrate our navigation strategies. Overall, our results suggest that hierarchical hybrid systems may provide a natural framework for modeling elements of intelligent animated actors.
Eric Aaron, Harold C. Sun, Franjo Ivancic, Dimitris N. Metaxas
CA4
2002 From Visual Input to Modeling Humans
abstract
We present an overview of our recent methodology for modeling humans that is based on the use of computer vision methods to analyze human motion and extract invariants of the motion in a parameterized way that captures the nature of this motion. This basic parameterized motion can then be altered based on additional user input in the form of parameters to create a wide variety of human motions. This paradigm of modeling and animating human motions is a deviation from the traditional approach to animation which is based either exclusively on an animator's skill to produce convincing human animation or the use of motion capture data with no sophisticated analysis. Our approach clearly demonstrates that the synergy between computer vision and computer graphics methods is crucial in modeling human motion both external and internal. We demonstrate several examples of this approach including human walking and the modeling of human anatomy and physiology.
Dimitris N. Metaxas
CA1
2002 In-vivo Strain and Stress Estimation of the Left Ventricle from MRI Images
Zhenhua Hu, Dimitris N. Metaxas, Leon Axel
MICCAI (1)2
2002 Automated Segmentation of the Left and Right Ventricles in 4D Cardiac SPAMM Images
Albert Montillo, Dimitris N. Metaxas, Leon Axel
MICCAI (1)2
2002 LV-RV Shape Modeling Based on a Blended Parameterized Model
Kyoungju Park, Dimitris N. Metaxas, Leon Axel
MICCAI (1)2
2002 Methods for modeling and predicting mechanical deformations of the breast under external perturbations
Fred S. Azar, Dimitris N. Metaxas, Mitchell D. Schnall
Medical Image Anal.2
2002 Adjusting Shape Parameters Using Model-Based Optical Flow Residuals
abstract
We present a method for estimating the shape of a deformable model using the least-squares residuals from a model-based optical flow computation. This method is built on top of an estimation framework using optical flow and image features, where optical flow affects only the motion parameters of the model. Using the results of this computation, our new method adjusts all of the parameters so that the residuals from the flow computation are minimized. We present face tracking experiments that demonstrate that this method obtains a better estimate of shape compared to related frameworks.
Douglas DeCarlo, Dimitris N. Metaxas
IEEE Trans. Pattern Anal. Mach. Intell.2
2002 Elastically Adaptive Deformable Models
abstract
We present a technique for the automatic adaptation of a deformable model's elastic parameters within a Kalman filter framework for shape estimation applications. The novelty of the technique is that the model's elastic parameters are not constant, but spatio-temporally varying. The variation of the elastic parameters depends on the distance of the model from the data and the rate of change of this distance. Each pass of the algorithm uses physics-based modeling techniques to iteratively adjust both the geometric and the elastic degrees of freedom of the model in response to forces that are computed from the discrepancy between the model and the data. By augmenting the state equations of an extended Kalman filter to incorporate these additional variables, we are able to significantly improve the quality of the shape estimation. Therefore, the model's elastic parameters are always initialized to the same value and they are subsequently modified depending on the data and the noise distribution. We present results demonstrating the effectiveness of our method for both two-dimensional and three-dimensional data.
Dimitris N. Metaxas, Ioannis A. Kakadiaris
IEEE Trans. Pattern Anal. Mach. Intell.1
2001 Pedestrians: creating agent behaviors through statistical analysis of observation data
abstract
Creating a complex virtual environment with human inhabitants that behave as we would expect real humans to behave is a difficult and time consuming task. Time must be spent to construct the environment, to create human figures, to create animations for the agents' actions, and to create controls for the agents' behaviors, such as scripts, plans, and decision-makers. Often work done for one virtual environment must be completely replicated for another. The creation of robust, procedural actions that can be ported from one simulation to another would ease the creation of new virtual environments. As walking is useful in many different virtual environments, the creation of natural looking walking is important. In this paper we present a system for producing more natural looking walking by incorporating actions for the upper body. We aim to provide a tool that authors of virtual environments can use to add realism to their characters without effort.
Koji Ashida, Seung-Joo Lee, Jan M. Allbeck, Harold C. Sun, Norman I. Badler, Dimitris N. Metaxas
CA6
2001 Collision resolutions in cloth simulation
abstract
We present a new collision resolution scheme for cloth collisions. Our main concern is to find dynamically convincing resolutions, i.e. positions and velocities of cloth elements, for any kinds of collisions occuring in cloth simulation (cloth-cloth, cloth-rigid, anti cloth-cloth-rigid). We define our cloth surface as connected faces of mass particles, where each particle is controlled by its internal energy functions. Our collision resolution method finds appropriate next positions and velocities of particles by conserving the particles' momentums as accurately as possible. Cloth-cloth collision resolution is a special case of deformable N-body collision resolution. So to solve deformable N-body collision resolutions, we propose a new collision resolution method which groups cloth particles into parts and resolves collisions between parts using the law of momentum conservation. To resolve collisions, we solve a system of linear equations derived from the collision relationships. A system of linear equations is built using a scheme adapted from the simultaneous resolution method for rigid N-body collisions. For the special case where we can,find cyclic relationships in collisions, we solve a system of linear inequalities derived from the collision relationships.
Suejung B. Huh, Dimitris N. Metaxas, Norman I. Badler
CA2
2001 Affine Arithmetic Based Estimation of Cue Distributions in Deformable Model Tracking
abstract
In this paper we describe a statistical method for the integration of an unlimited number of cues within a deformable model framework. We treat each cue as a random variable, each of which is the sum of a large number of local contributions with unknown probability distribution functions. Under the assumption that these distributions are independent, the overall distributions of the generalized cue forces can be approximated with multidimensional Gaussians, as per the central limit theorem. Estimating the covariance matrix of these Gaussian distributions, however, is difficult, because the probability distributions of the local contributions are unknown. We use affine arithmetic as a novel approach toward overcoming these difficulties. It lets us track and integrate the support of bounded distributions without having to know their actual probability distributions, and without having to make assumptions about their properties. We present a method for converting the resulting affine forms into the estimated Gaussian distributions of the generalized cue forces. This method scales well with the number of cues. We apply a Kalman filter as a maximum likelihood estimator to merge all Gaussian estimates of the cues into a single best fit Gaussian. Its mean is the deterministic result of the algorithm, and its covariance matrix provides a measure of the confidence in the result. We demonstrate in experiments how to apply this framework to improve the results of a face tracking system.
Siome Goldenstein, Christian Vogler, Dimitris N. Metaxas
CVPR (1)3
2001 Improving the Scope of Deformable Model Shape and Motion Estimation
abstract
Previous approaches to deformable model shape estimation and tracking have assumed a fixed class of shapes representation (e.g., deformable superquadrics), initialized prior to tracking. Since the shape coverage of the model is fixed, such approaches do not directly accommodate incremental representation discovery during tracking. As a result, model shape coverage is decoupled from tracking, thereby limiting both processes in terms of scope and robustness. We present a novel deformable model framework that accommodates the incremental incorporation during tracking of new geometric primitives (lines, in addition to points) that are not explicitly captured in the initial deformable model but that are moving consistently with its image motion. As these new features are detected via consistency checks, they are added to the model, providing incremental soft constraints on the estimation of its rigid parameters. The consistency checks are based on trilinear relationships between geometric primitives. Consequently, we not only increase both model scope and, ultimately, its higher-level shape coverage, but improve tracking robustness and accuracy, by directly employing the new features in both forward prediction and reconstruction. Our new formulation is a step towards automating model shape estimation and tracking, since it requires significantly reduced initial model hand-crafting. We demonstrate our approach on two separate image-based tracking domains, each involving complex 3D object shape and motion.
Cristian Sminchisescu, Dimitris N. Metaxas, Sven J. Dickinson
CVPR (1)2
2001 Scalable Dynamical Systems for Multi-Agent Steering and Simulation
abstract
We present a methodology for agent modeling that is scalable and efficient. It is based on the integration of nonlinear dynamical systems and kinetic data structures. The method consists of three-layers that model steering, flocking, and crowding agent behaviors among moving and static obstacles in 2 and 3D. The first layer, the local layer is based on the the use of nonlinear dynamical systems theory and models low level behaviors, it is fast and efficient, and does not depend on the total number of agents in the environment. The use of dynamical systems allows the use of continuous numerical parameters with which we can modify the interaction of each agent with the environment. This creates controllable distinctive behaviors. The second layer, a global environment layer consists of a specifically designed kinetic data structure to track efficiently the immediate environment of each agent and know which obstacles/agents are near or visible to the given agent. This layer reduces the complexity in the local layer. In the third layer, a global planning layer, the problem of target tracking is generalized in a way that allows navigation in maze-like terrains, avoidance of local minima and cooperation between agents. We implement this layer based on two approaches that are suitable for different applications. One is to track the closest single moving or static target. The second is to use a pre-specified vector field. This vector can be generated automatically (with harmonic functions, for example) or based on user input to achieve the desired output. We demonstrate the power of the approach through a series of experiments simulating single/multiple agents and crowds moving towards moving/static targets in complex environments.
Siome Goldenstein, Menelaos I. Karavelas, Dimitris N. Metaxas, Leonidas J. Guibas, Ambarish Goswami
ICRA3
2001 Methods for Modeling and Predicting Mechanical Deformations of the Breast Under External Perturbations
Fred S. Azar, Dimitris N. Metaxas, Mitchell D. Schnall
MICCAI2
2001 Hybrid Segmentation of Anatomical Data
Celina Imielinska, Dimitris N. Metaxas, Jayaram K. Udupa, Yinpeng Jin, Ting Chen 0001
MICCAI2
2001 Automating gait generation
abstract
One of the most routine actions humans perform is walking. To date, however, an automated tool for generating human gait is not available. This paper addresses the gait generation problem through three modular components. We present ElevWalker, a new low-level gait generator based on sagittal elevation angles, which allows curved locomotion - walking along a curved path - to be created easily; ElevInterp, which uses a new inverse motion interpolation algorithm to handle uneven terrain locomotion; and MetaGait, a high-level control module which allows an animator to control a figure's walking simply by specifying a path. The synthesis of these components is an easy-to-use, real-time, fully automated animation tool suitable for off-line animation, virtual environments and simulation.
Harold C. Sun, Dimitris N. Metaxas
SIGGRAPH2
2001 Scalable nonlinear dynamical systems for agent steering and crowd simulation
Siome Goldenstein, Menelaos I. Karavelas, Dimitris N. Metaxas, Leonidas J. Guibas, Eric Aaron, Ambarish Goswami
Comput. Graph.3
2001 A Framework for Recognizing the Simultaneous Aspects of American Sign Language
Christian Vogler, Dimitris N. Metaxas
Comput. Vis. Image Underst.2
2000 LaTex Human Motion Planning Based on Recursive Dynamics and Optimal Control Techniques
abstract
The 3D simulation of human activity based on physics, kinematics, and dynamics in terrestrial and space environments, is becoming increasingly important. Virtual humans may be used to design tasks in terrestrial environments and analyze their physical workload to maximize success and safety without expensive physical mockups. Previously (J. Lo and D. Metaxas, 1999), we presented an efficient optimal control and recursive dynamics based animation system for simulating and controlling the motion of articulated figures, and implemented the approach to several experiments where the simplified articulated models which has serial/closed-loop chain structures with only small degree-of-freedom (less than seven) each. The computation time is from a few minutes up to 4 hours based on the complexity of the model and how close the initial guess of the motion trajectory is to the optimal solution. The paper presents an improved method which can deal with more complicated kinematic chains (tree structures), as well as larger degree-of-freedom serial/closed-loop chain structures. Motion planning is done by first solving the inverse kinematic problem to generate possible trajectories, which gives us a better initial guess of the motion trajectory than the previous one, and then by solving the resulting nonlinear optimal control problem. For example, minimization of the torques during a simulation under certain constraints is often applied and has its origin in the biomechanics literature. Examples of activities shown are chinup and dipdown in different terrestrial environments as well as zero-gravity self orientation and ladder traversal.
Dimitris N. Metaxas, Janzen Lo
Computer Graphics International2
2000 Variable Albedo Surface Reconstruction from Stereo and Shape from Shading
abstract
We present a multiview method for the computation of object shape and reflectance characteristics based on the integration of shape from shading (SFS) and stereo, for nonconstant albedo and non-uniformly Lambertian surfaces. First we perform stereo fitting on the input stereo pairs or image sequences. When the images are uncalibrated, we recover the camera parameters using bundle adjustment. Eased on the stereo result, we can automatically segment the albedo map (which is taken to be piece-wise constant) using a minimum description length (MDL) based metric, to identify areas suitable for SFS (typically smooth textureless areas) and to derive illumination information. The shape and the illumination parameter estimates are refined using a deformable model SFS algorithm, which iterates between computing shape and illumination parameters. Our method takes into account the viewing angle dependent for shortening and specularity effects, and compensates as much as possible by utilizing information from more than one images. We demonstrate that we can extend the applicability of SFS algorithms to real world situations when some of its traditional assumptions are violated. We demonstrate our method by applying it to face shape reconstruction. Experimental results indicate a significant improvement over SFS-only or stereo-only based reconstruction. Model accuracy and detail are improved, especially in areas of low texture detail. Albedo information is retrieved and can be used to accurately re-render the model under different illumination conditions.
Dimitris Samaras, Dimitris N. Metaxas, Pascal Fua, Yvan G. Leclerc
CVPR2
2000 Image Segmentation Based on the Integration of Markov Random Fields and Deformable Models
Ting Chen 0001, Dimitris N. Metaxas
MICCAI2
2000 Animation of Human Locomotion Using Sagittal Elevation Angles
abstract
Abstract This paper presents a data-driven procedural model forthe kinematic animation of human walking. The use of data yields realistic looking gait, while the procedural modelyields flexibility. We present a new motion data representation, the sagittal elevation angles, and present biomechan-ical evidence that these angles have a stereotyped pattern across many different walking situations, implying theirreusability as a motion data source. We also sketch our algorithm for animating human gait based on sagittal el-evation angle data which allows us to generate curved locomotion on uneven terrain with stylistic variation withoutrequiring new datasets. 1.
Harold C. Sun, Dimitris N. Metaxas
PG2
2000 Dynamic Deformable Models for Enhanced Haptic Rendering in Virtual Environments
abstract
Currently there are no deformable model implementations that model a wide range of geometric deformations while providing realistic force feedback for use in virtual environments with haptics. The few models that exist are computationally very expensive, are limited in terms of shape coverage and do not provide proper haptic feedback. We use dynamic deformable models with local and global deformations governed by physical principles in order to provide efficient and true force feedback. We extend the shape class of Deformable Superquadrics (DeSuq) to provide compact geometric representation using few parameters, while at the same time providing realistic haptic viscoelastic feedback. Dynamics associated with rigid and deformable bodies are modeled by the use of the Lagrange equations. Implementation of these is currently under progress using GHOST/sup TM/ libraries on a PHANToM/sup TM/ haptic device.
Rungun Ramanathan, Dimitris N. Metaxas
VR2
2000 Optical Flow Constraints on Deformable Models with Applications to Face Tracking
Douglas DeCarlo, Dimitris N. Metaxas
Int. J. Comput. Vis.2
2000 Three-dimensional motion reconstruction and analysis of the right ventricle using tagged MRI
abstract
Right ventricular (RV) dysfunction can serve as an indicator of heart and lung disease and can adversely affect the left ventricle. However, normal RV function must be characterized before abnormal states can be detected. We describe a method for reconstructing the 3D motion of the RV by fitting a deformable model to tag and contour data extracted from multiview tagged magnetic resonance images. The deformable model is a biventricular finite element mesh built directly from segmented contours. Our approach accommodates the geometrically complex RV by using the entire lengths of the tags, localized degrees of freedom, and finite elements for geometric modeling. Also, we outline methods for converting the 3D motion reconstruction results into potentially useful motion variables, such as strains and displacements. The technique was applied to synthetic data, two normal hearts, and two hearts with right ventricular hypertrophy (RVH). Noticeable differences were found between the motion variables calculated for normal volunteers and RVH patients.
Idith Haber, Dimitris N. Metaxas, Leon Axel
Medical Image Anal.2
2000 Model-Based Estimation of 3D Human Motion
abstract
This paper presents the formulations and techniques that we have developed for the 3D model-based, motion estimation of human movement from multiple cameras. Our method is based on the spatio-temporal analysis of the subject's silhouette and it has the advantage that the subject does not have to wear markers or other devices. We present tracking results from experiments involving the recovery of complex motions in the presence of significant occlusion.
Ioannis A. Kakadiaris, Dimitris N. Metaxas
IEEE Trans. Pattern Anal. Mach. Intell.2
1999 Deformable Models for Segmentation, 3D Shape and Motion Estimation and Recognition
abstract
We present our framework for segmentation, 3D shape and motion esti-mation and recognition. We first present physics-based modeling techniques for segmentation and 3D shape and motion estimation based on single and multiple views as well as the integration of visual cues such as edges and optical flow. We then present extensions to address the reliable recognition of American Sign Language (ASL), using 3D tracking data, ASL phonol-ogy and modifications to the traditional use of Hidden Markov Models. We demonstrate the usefulness of this framework in computer vision and medical image analysis applications. 1
Dimitris N. Metaxas
BMVC1