VLDB 2026 Research / reviewers in the wild / expert
Xiaoming Liu 0002
dblp:l/XiaomingLiu0002
· DBLP profile ↗
201ranked-venue papers
25as first author
80since 2021 · last 2026
0000-0003-3215-8753ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 160 · 17 first-author · 69 since 2021Graphics, computer vision, multimedia, augmented reality and games · 149 · 20 first-author · 53 since 2021Computer networks · 8 · 4 since 2021Security and privacy · 5 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Marginalized Bundle Adjustment: Multi-View Camera Pose from Monocular Depth EstimatesabstractStructure-from-Motion (SfM) is a fundamental 3D vision task for recovering camera parameters and scene geometry from multi-view images. While recent deep learning advances enable accurate Monocular Depth Estimation (MDE) from single images without depending on camera motion, integrating MDE into SfM remains a challenge. Unlike conventional triangulated sparse point clouds, MDE produces dense depth maps with significantly higher error variance. Inspired by modern RANSAC estimators, we propose Marginalized Bundle Adjustment (MBA) to mitigate MDE error variance leveraging its density. With MBA, we show that MDE depth maps are sufficiently accurate to yield SoTA or competitive results in SfM and camera relocalization tasks. Through extensive evaluations, we demonstrate consistently robust performance across varying scales, ranging from few-frame setups to large multiview systems with thousands of images. Our method highlights the significant potential of MDE in multi-view 3D vision. Code is available at https://marginalizedba.github.io/. Ahmed Abdelkader, Mark J. Matthews, Xiaoming Liu 0002, Wen-Sheng Chu |
3DV | 4 |
| 2026 | Can Reasoning Path still be Effective as Input? Bridging Post-Reasoning to Chain-of-Thought CompressionabstractChengzhengxu Li, Xiaoming Liu, Zhaohan Zhang, Shengchao Liu, Guoxin Ma, Yu Lan, Cong Wang, Chao Shen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Chengzhengxu Li, Xiaoming Liu 0002, Zhaohan Zhang, Shengchao Liu, Guoxin Ma, Yu Lan 0001, Cong Wang 0001, Chao Shen 0001 |
ACL (1) | 2 |
| 2026 | Proactive Schemes: A Survey of Adversarial Attacks for Social Good
Vishal Asnani, Xi Yin 0001, Xiaoming Liu 0002 |
Int. J. Comput. Vis. | 3 |
| 2026 | UniAttack: Unified Physical-Digital Face Attack Detection
Shunxin Chen, Ajian Liu 0001, Haocheng Yuan, Junze Zheng, Dingheng Zeng, Jiankang Deng, Sergio Escalera, Xiaoming Liu 0002, Jun Wan 0001, Zhen Lei 0001 |
Int. J. Comput. Vis. | 10 |
| 2026 | 50 Years of Automated Face RecognitionabstractOver the past five decades, automated face recognition (FR) has progressed from handcrafted geometric and statistical approaches to advanced deep learning architectures that now approach, and in many cases exceed, human performance. This paper traces the historical and technological evolution of FR, encompassing early algorithmic paradigms through to contemporary neural systems trained on extensive real and synthetically generated datasets. We examine pivotal innovations that have driven this progression, including advances in dataset construction, loss function formulation, network architecture design, and feature fusion strategies. Furthermore, we analyze the relationship between data scale, diversity, and model generalization, highlighting how dataset expansion correlates with benchmark performance gains. Recent systems have achieved near-perfect large-scale identification accuracy, with the leading algorithm in the latest NIST FRTE 1:N benchmark reporting a False Negative Identification Rate (FNIR) of 0.15 percent at False Positive Identification Rate (FPIR) of 0.001 on a gallery of over 10 million identities. Larger galleries increase false positive rates and deployments at greater scales will see higher error rates. We delineate key open problems and emerging directions, including scalable training, multi-modal fusion, synthetic data, and interpretable recognition frameworks. Anil K. Jain 0001, Xiaoming Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Rethinking Vision-Language Model in Face Forensics: Multi-Modal Interpretable Forged Face DetectorabstractDeepfake detection is a long-established research topic vital for mitigating the spread of malicious misinformation. Unlike prior methods that provide either binary classification results or textual explanations separately, we introduce a novel method capable of generating both simultaneously. Our method harnesses the multi-modal learning capability of the pre-trained CLIP and the unprecedented interpretability of large language models (LLMs) to enhance both the generalization and explainability of deep-fake detection. Specifically, we introduce a multi-modal face forgery detector (M2F2-Det) that employs tailored face forgery prompt learning, incorporating the pre-trained CLIP to improve generalization to unseen forgeries. Also, M2F2-Det incorporates an LLM to provide detailed textual explanations of its detection decisions, enhancing interpretability by bridging the gap between natural language and subtle cues of facial forgeries. Empirically, we evaluate M2F2-Det on both detection and explanation generation tasks, where it achieves state-of-the-art performance, demonstrating its effectiveness in identifying and explaining diverse forgeries. Source code is available at $\color{magenta}{link}$. Xiufeng Song, Xiaohong Liu 0001, Xiaoming Liu 0002 |
CVPR | 5 |
| 2025 | H-MoRe: Learning Human-centric Motion Representation for Action AnalysisabstractIn this paper, we propose H-MoRe, a novel pipeline for learning precise human-centric motion representation. Our approach dynamically preserves relevant human motion while filtering out background movement. Notably, unlike previous methods relying on fully supervised learning from synthetic data, H-MoRe learns directly from real-world scenarios in a self-supervised manner, incorporating both human pose and body shape information. Inspired by kinematics, H-MoRe represents absolute and relative movements of each body point in a matrix format that captures nuanced motion details, termed world-local flows. H-MoRe offers refined insights into human motion, which can be integrated seamlessly into various action-related applications. Experimental results demonstrate that H-MoRe brings substantial improvements across various downstream tasks, including gait recognition (CL@R1: 16.01%↑), action recognition (Acc@1: 8.92%↑), and video generation (FVD: 67.07%↓). Additionally, H-MoRe exhibits high inference efficiency (34 fps), making it suitable for most real-time scenarios. Models and code is available at https://github.com/haku-huang/h-more. Zhanbo Huang, Xiaoming Liu 0002 |
CVPR | 2 |
| 2025 | SapiensID: Foundation for Human RecognitionabstractExisting human recognition systems often rely on separate, specialized models for face and body analysis, limiting their effectiveness in real-world scenarios where pose, visibility, and context vary widely. This paper introduces SapiensID, a unified model that bridges this gap, achieving robust performance across diverse settings. SapiensID introduces (i) Retina Patch (RP), a dynamic patch generation scheme that adapts to subject scale and ensures consistent tokenization of regions of interest, (ii) a masked recognition model (MRM) that learns from variable token length, and (iii) Semantic Attention Head (SAH), an module that learns pose-invariant representations by pooling features around key body parts. To facilitate training, we introduce WebBody4M, a large-scale dataset capturing diverse poses and scale variations. Extensive experiments demonstrate that SapiensID achieves state-of-the-art results on various body ReID benchmarks, outperforming specialized models in both short-term and long-term scenarios while remaining competitive with dedicated face recognition systems. Furthermore, SapiensID establishes a strong baseline for the newly introduced challenge of Cross Pose-Scale ReID, demonstrating its ability to generalize to complex, real-world conditions. Project Link Dingqiang Ye, Yiyang Su, Feng Liu 0037, Xiaoming Liu 0002 |
CVPR | 5 |
| 2025 | RICCARDO: Radar Hit Prediction and Convolution for Camera-Radar 3D Object DetectionabstractRadar hits reflect from points on both the boundary and internal to object outlines. This results in a complex distribution of radar hits that depends on factors including object category, size and orientation. Current radar-camera fusion methods implicitly account for this with a black-box neural network. In this paper, we explicitly utilize a radar hit distribution model to assist fusion. First, we build a model to predict radar hit distributions conditioned on object properties obtained from a monocular detector. Second, we use the predicted distribution as a kernel to match actual measured radar points in the neighborhood of the monocular detections, generating matching scores at nearby positions. Finally, a fusion stage combines context with the kernel detector to refine the matching scores. Our method achieves the state-of-the-art radar-camera detection performance on nuScenes. Our source code is available at https://github.com/longyunf/riccardo. Abhinav Kumar 0004, Xiaoming Liu 0002, Daniel D. Morris |
CVPR | 3 |
| 2025 | AG-VPReID 2025: Aerial-Ground Video-based Person Re-identification Challenge ResultsabstractPerson re-identification (ReID) across aerial and ground vantage points has become crucial for large-scale surveillance and public safety applications. Although significant progress has been made in ground-only scenarios, bridging the aerial-ground domain gap remains a formidable challenge due to extreme viewpoint differences, scale variations, and occlusions. Building upon the achievements of the AG-ReID 2023 Challenge, this paper introduces the AG-VPReID 2025 Challenge—the first large-scale video-based competition focused on high-altitude (80–120 m) aerial-ground person ReID. Constructed on the new AG-VPReID dataset with 3,027 identities, over 13,500 tracklets, and approximately 3.7 million frames captured from UAVs, CCTV, and wearable cameras, the challenge featured four international teams. These teams developed solutions ranging from multi-stream architectures to transformer-based temporal reasoning and physics-informed modeling. The leading approach, X-TFCLIP from UAM, attained 72.28% Rank-1 accuracy in the aerial-to-ground ReID setting and 70.77% in the ground-to-aerial ReID setting, surpassing existing baselines while highlighting the dataset’s complexity. For additional details, please refer to the official website at https://agvpreid25.github.io. Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan, Feng Liu 0037, Xiaoming Liu 0002, Arun Ross, Tamás Endrei, Ivan DeAndres-Tame, Ruben Tolosana, Rubén Vera-Rodríguez, Aythami Morales, Julian Fierrez, Javier Ortega-Garcia, Zijing Gong, Xuehu Liu, Md. Rashidunnabi, Hugo Proença 0001, Kailash A. Hambarde, Saeid Rezaei |
IJCB | 6 |
| 2025 | CHARM3R: Towards Unseen Camera Height Robust Monocular 3D Detector
Abhinav Kumar 0004, Yuliang Guo, Xinyu Huang 0001, Liu Ren 0001, Xiaoming Liu 0002 |
ICCV | 6 |
| 2025 | A Quality-Guided Mixture of Score-Fusion Experts Framework for Human RecognitionabstractWhole-body biometric recognition is a challenging multimodal task that integrates various biometric modalities, including face, gait, and body. This integration is essential for overcoming the limitations of unimodal systems. Traditionally, whole-body recognition involves deploying different models to process multiple modalities, achieving the final outcome by score-fusion (e.g., weighted averaging of similarity matrices from each model). However, these conventional methods may overlook the variations in score distributions of individual modalities, making it challenging to improve final performance. In this work, we present \textbf{Q}uality-guided \textbf{M}ixture of score-fusion \textbf{E}xperts (QME), a novel framework designed for improving whole-body biometric recognition performance through a learnable score-fusion strategy using a Mixture of Experts (MoE). We introduce a novel pseudo-quality loss for quality estimation with a modality-specific Quality Estimator (QE), and a score triplet loss to improve the metric performance. Extensive experiments on multiple whole-body biometric datasets demonstrate the effectiveness of our proposed approach, achieving state-of-the-art results across various metrics compared to baseline methods. Our method is effective for multimodal and multi-model, addressing key challenges such as model misalignment in the similarity score domain and variability in data quality. Yiyang Su, Anil K. Jain 0001, Xiaoming Liu 0002 |
ICCV | 5 |
| 2025 | BiggerGait: Unlocking Gait Recognition with Layer-wise Representations from Large Vision ModelsabstractLarge vision models (LVM) based gait recognition has achieved impressive performance.
However, existing LVM-based approaches may overemphasize gait priors while neglecting the intrinsic value of LVM itself, particularly the rich, distinct representations across its multi-layers.
To adequately unlock LVM's potential, this work investigates the impact of layer-wise representations on downstream recognition tasks.
Our analysis reveals that LVM's intermediate layers offer complementary properties across tasks, integrating them yields an impressive improvement even without rich well-designed gait priors.
Building on this insight, we propose a simple and universal baseline for LVM-based gait recognition, termed BiggerGait.
Comprehensive evaluations on CCPG, CAISA-B*, SUSTech1K, and CCGR_MINI validate the superiority of BiggerGait across both within- and cross-domain tasks, establishing it as a simple yet practical baseline for gait representation learning.
All the models and code are available at https://github.com/ShiqiYu/OpenGait/. Dingqiang Ye, Chao Fan 0001, Zhanbo Huang, Chengwen Luo 0001, Jianqiang Li 0001, Shiqi Yu 0001, Xiaoming Liu 0002 |
NeurIPS | 7 |
| 2025 | Can Adversarial Examples be Parsed to Reveal Victim Model Information?abstractNumerous adversarial attack methods have been developed to generate imperceptible image perturbations that cause erroneous predictions in state-of-the-art machine learning (ML) models, particularly deep neural networks (DNNs). Despite extensive research on adversarial examples, limited efforts have been made to explore the hidden characteristics carried by these perturbations. In this study, we investigate the feasibility of deducing information about the victim model (VM)—specifically, characteristics such as architecture type, kernel size, activation function, and weight sparsity—from adversarial examples. We approach this problem as a supervised learning task, where we aim to attribute categories of VM characteristics to individual adversarial examples. To facilitate this, we have assembled a dataset of adversarial attacks spanning seven types, generated from 135 victim models systematically varied across five architecture types, three kernel size configurations, three activation functions, and three levels of weight sparsity. We demonstrate that a supervised model parsing network (MPN) can effectively extract concealed details of the VM from adversarial examples. We also validate the practicality of this approach by evaluating the effects of various factors on parsing performance, such as different input formats and generalization to out-of-distribution cases. Furthermore, we highlight the connection between model parsing and attack transferability by showing how the MPN can uncover VM attributes in transfer attacks. Yuguang Yao, Jiancheng Liu, Yifan Gong 0004, Xiaoming Liu 0002, Yanzhi Wang 0001, Xue Lin 0001, Sijia Liu 0001 |
WACV | 4 |
| 2025 | Language-Guided Hierarchical Fine-Grained Image Forgery Detection and Localization
Xiaohong Liu 0001, Iacopo Masi, Xiaoming Liu 0002 |
Int. J. Comput. Vis. | 4 |
| 2024 | SeaBird: Segmentation in Bird's View with Dice Loss Improves Monocular 3D Detection of Large ObjectsabstractMonocular 3D detectors achieve remarkable performance on cars and smaller objects. However, their performance drops on larger objects, leading to fatal accidents. Some attribute the failures to training data scarcity or the receptive field requirements of large objects. In this paper, we highlight this understudied problem of generalization to large objects. We find that modern frontal detectors struggle to generalize to large objects even on nearly balanced datasets. We argue that the cause of failure is the sensitivity of depth regression losses to noise of larger objects. To bridge this gap, we comprehensively investigate regression and dice losses, examining their robustness under varying error levels and object sizes. We mathematically prove that the dice loss leads to superior noise-robustness and model convergence for large objects compared to regression losses for a simplified case. Leveraging our theoretical insights, we propose SeaBird (Segmentation in Bird's View) as the first step towards generalizing to large objects. SeaBird effectively integrates BEV segmentation on foreground objects for 3D detection, with the segmentation head trained with the dice loss. SeaBird achieves SoTA results on the KITTI-360 leaderboard and improves existing detectors on the nuScenes leaderboard, particularly for large objects. Abhinav Kumar 0004, Yuliang Guo, Xinyu Huang 0001, Liu Ren 0001, Xiaoming Liu 0002 |
CVPR | 5 |
| 2024 | Distilling CLIP with Dual Guidance for Learning Discriminative Human Body Shape RepresentationabstractPerson Re-ldentification (ReID) holds critical importance in computer vision with pivotal applications in public safety and crime prevention. Traditional ReID methods, reliant on appearance attributes such as clothing and color, en-counter limitations in long-term scenarios and dynamic environments. To address these challenges, we propose CLIP3DReID, an innovative approach that enhances person ReID by integrating linguistic descriptions with visual per-ception, leveraging pretrained CLIP model for knowledge distillation. Our method first employs CLIP to automatically label body shapes with linguistic descriptors. We then ap-ply optimal transport theory to align the student model's local visual features with shape-aware tokens derived from CLIP's linguistic output. Additionally, we align the student model's global visual features with those from the CLIP image encoder and the 3D SMPL identity space, fostering enhanced domain robustness. CLIP3DReID notably excels in discerning discriminative body shape features, achieving state-of-the-art results in person ReID. Our approach rep-resents a significant advancement in ReID, offering robust solutions to existing challenges and setting new directions for future research. Code is available. Feng Liu 0037, Xiaoming Liu 0002 |
CVPR | 4 |
| 2024 | ProMark: Proactive Diffusion Watermarking for Causal AttributionabstractGenerative AI (GenAI) is transforming creative work-flows through the capability to synthesize and manipulate images via high-level prompts. Yet creatives are not well supported to receive recognition or reward for the use of their content in GenAI training. To this end, we propose ProMark, a causal attribution technique to attribute a synthetically generated image to its training data concepts like objects, motifs, templates, artists, or styles. The concept information is proactively embedded into the input training images using imperceptible watermarks, and the diffusion models (unconditional or conditional) are trained to retain the corresponding watermarks in generated images. We show that we can embed as many as 216unique water-marks into the training data, and each training image can contain more than one watermark. ProMark can maintain image quality whilst outperforming correlation-based attribution. Finally, several qualitative examples are presented, providing the confidence that the presence of the watermark conveys a causative relationship between training data and synthetic images. Vishal Asnani, John P. Collomosse, Tu Bui, Xiaoming Liu 0002, Shruti Agarwal |
CVPR | 4 |
| 2024 | KeyPoint Relative Position Encoding for Face RecognitionabstractIn this paper, we address the challenge of making ViT models more robust to unseen affine transformations. Such robustness becomes useful in various recognition tasks such as face recognition when image alignment failures occur. We propose a novel method called KP-RPE, which leverages key points (e.g. facial landmarks) to make ViT more resilient to scale, translation, and pose variations. We begin with the observation that Relative Position Encoding (RPE) is a good way to bring affine transform generalization to ViTs. RPE, however, can only inject the model with prior knowledge that nearby pixels are more important than far pixels. Keypoint RPE (KP-RPE) is an extension of this principle, where the significance of pixels is not solely dictated by their proximity but also by their relative positions to specific keypoints within the image. By anchoring the significance of pixels around keypoints, the model can more effectively retain spatial relationships, even when those relationships are disrupted by affine transformations. We show the merit of KP-RPE inface and gait recognition. The experimental results demonstrate the effectiveness in improving face recognition performance from low-quality images, particularly where alignment is prone to failure. Code and pre-trained models are available. Yiyang Su, Feng Liu 0037, Xiaoming Liu 0002 |
CVPR | 5 |
| 2024 | TIGER: Time-Varying Denoising Model for 3D Point Cloud Generation with Diffusion ProcessabstractRecently, diffusion models have emerged as a new pow-erful generative method for 3D point cloud generation tasks. However, few works study the effect of the archi-tecture of the diffusion model in the 3D point cloud, re-sorting to the typical UNet model developed for 2D images. Inspired by the wide adoption of Transformers, we study the complementary role of convolution (from UNet) and attention (from Transformers). We discover that their respective importance change according to the timestep in the diffusion process. At early stage, attention has an out-sized influence because Transformers are found to generate the overall shape more quickly, and at later stages when adding fine detail, convolution starts having a larger im-pact on the generated point cloud's local surface quality. In light of this observation, we propose a time-varying two-stream denoising model combined with convolution lay-ers and transformer blocks. We generate an optimizable mask from each timestep to reweigh global and local features, obtaining time-varying fused features. Experimen-tally, we demonstrate that our proposed method quantitatively outperforms other state-of-the-art methods regarding visual quality and diversity. Code is avaiable https://github.com/Zhiyuan-R/Tiger-Diffusion. Feng Liu 0037, Xiaoming Liu 0002 |
CVPR | 4 |
| 2024 | BigGait: Learning Gait Representation You Want by Large Vision ModelsabstractGait recognition stands as one of the most pivotal remote identification technologies and progressively expands across research and industry communities. However, existing gait recognition methods heavily rely on task-specific upstream driven by supervised learning to provide explicit gait representations like silhouette sequences, which in-evitably introduce expensive annotation costs and poten-tial error accumulation. Escaping from this trend, this work explores effective gait representations based on the all-purpose knowledge produced by task-agnostic Large Vision Models (LVMs) and proposes a simple yet efficient gait framework, termed B igGait. Specifically, the Gait Repre-sentation Extractor (GRE) within BigGait draws upon design principles from established gait representations, effectively transforming all-purpose knowledge into implicit gait representations without requiring third-party supervision signals. Experiments on CCPG, CAISA-B* and SUSTechlK indicate that BigGait significantly outperforms the previous methods in both within-domain and cross-domain tasks in most cases, and provides a more practical paradigm for learning the next-generation gait representation. Fi-nally, we delve into prospective challenges and promising directions in LVMs-based gait recognition, aiming to in-spire future work in this emerging topic. The source code is available at https://github.com/ShiqiYu/OpenGait. Dingqiang Ye, Chao Fan 0001, Jingzhe Ma, Xiaoming Liu 0002, Shiqi Yu 0001 |
CVPR | 4 |
| 2024 | COMPOSE: Comprehensive Portrait Shadow Editing
Andrew Hou, Zhixin Shu, Xuaner Cecilia Zhang, He Zhang 0004, Yannick Hold-Geoffroy, Jae Shin Yoon, Xiaoming Liu 0002 |
ECCV (61) | 7 |
| 2024 | Open-Set Biometrics: Beyond Good Closed-Set Models
Yiyang Su, Feng Liu 0037, Anil K. Jain 0001, Xiaoming Liu 0002 |
ECCV (62) | 5 |
| 2024 | RePLAy: Remove Projective LiDAR Depthmap Artifacts via Exploiting Epipolar Geometry
Girish Chandar Ganesan, Abhinav Kumar 0004, Xiaoming Liu 0002 |
ECCV (81) | 4 |
| 2024 | Revisit Self-supervised Depth Estimation with Local Structure-from-Motion
Xiaoming Liu 0002 |
ECCV (82) | 2 |
| 2024 | Unified Physical-Digital Face Attack Detection
Ajian Liu 0001, Haocheng Yuan, Junze Zheng, Dingheng Zeng, Jiankang Deng, Sergio Escalera, Xiaoming Liu 0002, Jun Wan 0001, Zhen Lei 0001 |
IJCAI | 9 |
| 2024 | Tracing Hyperparameter Dependencies for Model Parsing via Learnable Graph Pooling Networkabstract\textit{Model Parsing} defines the task of predicting hyperparameters of the generative model (GM), given a GM-generated image as the input.
Since a diverse set of hyperparameters is jointly employed by the generative model, and dependencies often exist among them, it is crucial to learn these hyperparameter dependencies for improving the model parsing performance.
To explore such important dependencies, we propose a novel model parsing method called Learnable Graph Pooling Network (LGPN), in which we formulate model parsing as a graph node classification problem, using graph nodes and edges to represent hyperparameters and their dependencies, respectively.
Furthermore, LGPN incorporates a learnable pooling-unpooling mechanism tailored to model parsing, which adaptively learns hyperparameter dependencies of GMs used to generate the input image.
Also, we introduce a Generation Trace Capturing Network (GTC) that can efficiently identify generation traces of input images, enhancing the understanding of generated images' provenances.
Empirically, we achieve state-of-the-art performance in model parsing and its extended applications, showing the superiority of the proposed LGPN. Vishal Asnani, Sijia Liu 0001, Xiaoming Liu 0002 |
NeurIPS | 4 |
| 2024 | On Learning Multi-Modal Forgery Representation for Diffusion Generated Video DetectionabstractLarge numbers of synthesized videos from diffusion models pose threats to information security and authenticity, leading to an increasing demand for generated content detection. However, existing video-level detection algorithms primarily focus on detecting facial forgeries and often fail to identify diffusion-generated content with a diverse range of semantics. To advance the field of video forensics, we propose an innovative algorithm named Multi-Modal Detection(MM-Det) for detecting diffusion-generated videos. MM-Det utilizes the profound perceptual and comprehensive abilities of Large Multi-modal Models (LMMs) by generating a Multi-Modal Forgery Representation (MMFR) from LMM's multi-modal space, enhancing its ability to detect unseen forgery content. Besides, MM-Det leverages an In-and-Across Frame Attention (IAFA) mechanism for feature augmentation in the spatio-temporal domain. A dynamic fusion strategy helps refine forgery representations for the fusion. Moreover, we construct a comprehensive diffusion video dataset, called Diffusion Video Forensics (DVF), across a wide range of forgery videos. MM-Det achieves state-of-the-art performance in DVF, demonstrating the effectiveness of our algorithm. Both source code and DVF are available at https://github.com/SparkleXFantasy/MM-Det. Xiufeng Song, Jiache Zhang, Lei Bai 0001, Xiaoming Liu 0002, Guangtao Zhai, Xiaohong Liu 0001 |
NeurIPS | 6 |
| 2024 | UnlearnCanvas: Stylized Image Dataset for Enhanced Machine Unlearning Evaluation in Diffusion ModelsabstractThe technological advancements in diffusion models (DMs) have demonstrated unprecedented capabilities in text-to-image generation and are widely used in diverse applications. However, they have also raised significant societal concerns, such as the generation of harmful content and copyright disputes. Machine unlearning (MU) has emerged as a promising solution, capable of removing undesired generative capabilities from DMs. However, existing MU evaluation systems present several key challenges that can result in incomplete and inaccurate assessments. To address these issues, we propose UnlearnCanvas, a comprehensive high-resolution stylized image dataset that facilitates the evaluation of the unlearning of artistic styles and associated objects. This dataset enables the establishment of a standardized, automated evaluation framework with 7 quantitative metrics assessing various aspects of the unlearning performance for DMs. Through extensive experiments, we benchmark 9 state-of-the-art MU methods for DMs, revealing novel insights into their strengths, weaknesses, and underlying mechanisms. Additionally, we explore challenging unlearning scenarios for DMs to evaluate worst-case performance against adversarial prompts, the unlearning of finer-scale concepts, and sequential unlearning. We hope that this study can pave the way for developing more effective, accurate, and robust DM unlearning methods, ensuring safer and more ethical applications of DMs in the future. The dataset, benchmark, and codes are publicly available at this link. Chongyu Fan, Yuguang Yao, Jinghan Jia, Jiancheng Liu, Gaoyuan Zhang, Gaowen Liu, Ramana Rao Kompella, Xiaoming Liu 0002, Sijia Liu 0001 |
NeurIPS | 10 |
| 2024 | FarSight: A Physics-Driven Whole-Body Biometric System at Large Distance and AltitudeabstractWhole-body biometric recognition is an important area of research due to its vast applications in law enforcement, border security, and surveillance. This paper presents the end-to-end design, development and evaluation of FarSight, an innovative software system designed for whole-body (fusion of face, gait and body shape) biometric recognition. FarSight accepts videos from elevated platforms and drones as input and outputs a candidate list of identities from a gallery. The system is designed to address several challenges, including (i) low-quality imagery, (ii) large yaw and pitch angles, (iii) robust feature extraction to accommodate large intra-person variabilities and large inter-person similarities, and (iv) the large domain gap between training and test sets. FarSight combines the physics of imaging and deep learning models to enhance image restoration and biometric feature encoding. We test FarSight’s effectiveness using the newly acquired IARPA Biometric Recognition and Identification at Altitude and Range (BRIAR) dataset. Notably, FarSight demonstrated a substantial performance increase on the BRIAR dataset, with gains of +11.82% Rank-20 identification and +11.30% TAR@1% FAR. Feng Liu 0037, Ryan Ashbaugh, Nicholas Chimitt, Najmul Hassan, Ali Hassani 0001, Ajay Jaiswal, Zhiyuan Mao, Christopher Perry, Yiyang Su, Pegah Varghaei, Kai Wang 0058, Stanley H. Chan, Arun Ross, Humphrey Shi, Zhangyang Wang, Xiaoming Liu 0002 |
WACV | 19 |
| 2024 | ProS: Facial Omni-Representation Learning via Prototype-based Self-DistillationabstractThis paper presents a novel approach, called Prototype-based Self-Distillation (ProS), for unsupervised face representation learning. The existing supervised methods heavily rely on a large amount of annotated training facial data, which poses challenges in terms of data collection and privacy concerns. To address these issues, we propose ProS, which leverages a vast collection of unlabeled face images to learn a comprehensive facial omni-representation. In particular, ProS consists of two vision-transformers (teacher and student models) that are trained with different augmented images (cropping, blurring, coloring, etc.). Besides, we build a face-aware retrieval system along with augmentations to obtain the curated images comprising predominantly facial areas. To enhance the discrimination of learned features, we introduce a prototype-based matching loss that aligns the similarity distributions between features (teacher or student) and a set of learnable prototypes. After pre-training, the teacher vision transformer serves as a backbone for downstream tasks, including attribute estimation, expression recognition, and landmark alignment, achieved through simple fine-tuning with additional layers. Extensive experiments demonstrate that our method achieves state-of-the-art performance on various tasks, both in full and few-shot settings. Further, we investigate pre-training with synthetic face images, and ProS exhibits promising performance in this scenario as well. Xing Di, Yiyu Zheng, Xiaoming Liu 0002, Yu Cheng 0001 |
WACV | 3 |
| 2024 | Learning Implicit Functions for Dense 3D Shape Correspondence of Generic ObjectsabstractThe objective of this paper is to learn dense 3D shape correspondence for topology-varying generic objects in an unsupervised manner. Conventional implicit functions estimate the occupancy of a 3D point given a shape latent code. Instead, our novel implicit function produces a probabilistic embedding to represent each 3D point in a part embedding space. Assuming the corresponding points are similar in the embedding space, we implement dense correspondence through an inverse function mapping from the part embedding vector to a corresponded 3D point. Both functions are jointly learned with several effective and uncertainty-aware loss functions to realize our assumption, together with the encoder generating the shape latent code. During inference, if a user selects an arbitrary point on the source shape, our algorithm can automatically generate a confidence score indicating whether there is a correspondence on the target shape, as well as the corresponding semantic point if there is one. Such a mechanism inherently benefits man-made objects with different part constitutions. The effectiveness of our approach is demonstrated through unsupervised 3D semantic correspondence and shape segmentation. Feng Liu 0037, Xiaoming Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | RADIANT: Radar-Image Association Network for 3D Object DetectionabstractAs a direct depth sensor, radar holds promise as a tool to improve monocular 3D object detection, which suffers from depth errors, due in part to the depth-scale ambiguity. On the other hand, leveraging radar depths is hampered by difficulties in precisely associating radar returns with 3D estimates from monocular methods, effectively erasing its benefits. This paper proposes a fusion network that addresses this radar-camera association challenge. We train our network to predict the 3D offsets between radar returns and object centers, enabling radar depths to enhance the accuracy of 3D monocular detection. By using parallel radar and camera backbones, our network fuses information at both the feature level and detection level, while at the same time leveraging a state-of-the-art monocular detection technique without retraining it. Experimental results show significant improvement in mean average precision and translation error on the nuScenes dataset over monocular counterparts. Our source code is available at https://github.com/longyunf/radiant. Abhinav Kumar 0004, Daniel D. Morris, Xiaoming Liu 0002, Marcos Castro, Punarjay Chakravarty |
AAAI | 4 |
| 2023 | MaLP: Manipulation Localization Using a Proactive SchemeabstractAdvancements in the generation quality of various Generative Models (GMs) has made it necessary to not only perform binary manipulation detection but also localize the modified pixels in an image. However, prior works termed as passive for manipulation localization exhibit poor generalization performance over unseen GMs and attribute modifications. To combat this issue, we propose a proactive scheme for manipulation localization, termed MaLP. We encrypt the real images by adding a learned template. If the image is manipulated by any GM, this added protection from the template not only aids binary detection but also helps in identifying the pixels modified by the GM. The template is learned by leveraging local and global-level features estimated by a two-branch architecture. We show that MaLP performs better than prior passive works. We also show the generalizability of MaLP by testing on 22 different GMs, providing a benchmark for future research on manipulation localization. Finally, we show that MaLP can be used as a discriminator for improving the generation quality of GMs. Our models/codes are available at www.github.com/vishal3477/pro_loc. Vishal Asnani, Xi Yin 0001, Tal Hassner, Xiaoming Liu 0002 |
CVPR | 4 |
| 2023 | Hierarchical Fine-Grained Image Forgery Detection and LocalizationabstractDifferences in forgery attributes of images generated in CNN-synthesized and image-editing domains are large, and such differences make a unified image forgery detection and localization (IFDL) challenging. To this end, we present a hierarchical fine-grained formulation for IFDL representation learning. Specifically, we first represent forgery attributes of a manipulated image with multiple labels at different levels. Then we perform fine-grained classification at these levels using the hierarchical dependency between them. As a result, the algorithm is encouraged to learn both comprehensive features and inherent hierarchical nature of different forgery attributes, thereby improving the IFDL representation. Our proposed IFDL framework contains three components: multi-branch feature extractor, localization and classification modules. Each branch of the feature extractor learns to classify forgery attributes at one level, while localization and classification modules segment the pixel-level forgery region and detect image-level forgery, respectively. Lastly, we construct a hierarchical fine-grained dataset to facilitate our study. We demonstrate the effectiveness of our method on 7 different benchmarks, for both tasks of IFDL and forgery attribute classification. Our source code and dataset can be found: github.com/CHELSEA234/HiFi-IFDL. Xiaohong Liu 0001, Steven Grosz, Iacopo Masi, Xiaoming Liu 0002 |
CVPR | 6 |
| 2023 | DCFace: Synthetic Face Generation with Dual Condition Diffusion ModelabstractGenerating synthetic datasets for training face recognition models is challenging because dataset generation entails more than creating high fidelity images. It involves generating multiple images of same subjects under different factors (e.g., variations in pose, illumination, expression, aging and occlusion) which follows the real image conditional distribution. Previous works have studied the generation of synthetic datasets using GAN or 3D models. In this work, we approach the problem from the aspect of combining subject appearance (ID) and external factor (style) conditions. These two conditions provide a direct way to control the inter-class and intra-class variations. To this end, we propose a Dual Condition Face Generator (DCFace) based on a diffusion model. Our novel Patch-wise style extractor and Time-step dependent ID loss enables DCFace to consistently produce face images of the same subject under different styles with precise control. Face recognition models trained on synthetic images from the proposed DCFace provide higher verification accuracies compared to previous works by 6.11% on average in 4 out of 5 test datasets, LFW, CFP-FP, CPLFW, AgeDB and CALFW. Code Link Feng Liu 0037, Anil K. Jain 0001, Xiaoming Liu 0002 |
CVPR | 4 |
| 2023 | Rethinking Domain Generalization for Face Anti-spoofing: Separability and AlignmentabstractThis work studies the generalization issue of face anti-spoofing (FAS) models on domain gaps, such as image resolution, blurriness and sensor variations. Most prior works regard domain-specific signals as a negative impact, and apply metric learning or adversarial losses to remove them from feature representation. Though learning a domain-invariant feature space is viable for the training data, we show that the feature shift still exists in an unseen test domain, which backfires on the generalizability of the classifier. In this work, instead of constructing a domain-invariant feature space, we encourage domain separability while aligning the live-to-spoof transition (i.e., the trajectory from live to spoof) to be the same for all domains. We formulate this FAS strategy of separability and alignment (SA-FAS) as a problem of invariant risk minimization (IRM), and learn domain-variant feature representation but domain-invariant classifier. We demonstrate the effectiveness of SA-FAS on challenging cross-domain FAS datasets and establish state-of-the-art performance. Code is available at https://github.com/sunyiyou/SAFAS. Yiyou Sun, Yaojie Liu, Xiaoming Liu 0002, Yixuan Li 0001, Wen-Sheng Chu |
CVPR | 3 |
| 2023 | LightedDepth: Video Depth Estimation in Light of Limited Inference View AnglesabstractVideo depth estimation infers the dense scene depth from immediate neighboring video frames. While recent works consider it a simplified structure-from-motion (SfM) problem, it still differs from the SfM in that significantly fewer view angels are available in inference. This setting, however, suits the mono-depth and optical flow estimation. This observation motivates us to decouple the video depth estimation into two components, a normalized pose estimation over af lowmap and a logged residual depth estimation over a mono-depth map. The two parts are unified with an efficient off-the-shelf scale alignment algorithm. Additionally, we stabilize the indoor two-view pose estimation by including additional projection constraints and ensuring sufficient camera translation. Though a two-view algorithm, we validate the benefit of the decoupling with the substantial performance improvement over multi-view iterative prior works on indoor and outdoor datasets. Codes and models are available at https://github.com/ShngJZ/LightedDepth. Xiaoming Liu 0002 |
CVPR | 2 |
| 2023 | PMatch: Paired Masked Image Modeling for Dense Geometric MatchingabstractDense geometric matching determines the dense pixel-wise correspondence between a source and support image corresponding to the same 3D structure. Prior works employ an encoder of transformer blocks to correlate the two-frame features. However, existing monocular pretraining tasks, e.g., image classification, and masked image modeling (MIM), can not pretrain the cross-frame module, yielding less optimal performance. To resolve this, we reformulate the MIM from reconstructing a single masked image to reconstructing a pair of masked images, enabling the pretraining of transformer module. Additionally, we incorporate a decoder into pretraining for improved upsampling results. Further, to be robust to the textureless area, we propose a novel cross-frame global matching module (CFGM). Since the most textureless area is planar surfaces, we propose a homography loss to further regularize its learning. Combined together, we achieve the State-of-The-Art (SoTA) performance on geometric matching. Codes and models are available at https://github.com/ShngJZ/PMatch. Xiaoming Liu 0002 |
CVPR | 2 |
| 2023 | FaceGuard: A Self-Supervised Defense Against Adversarial Face ImagesabstractPrevailing defense schemes against adversarial face images tend to overfit to the perturbations in the training set and fail to generalize to unseen adversarial attacks. We propose a new self-supervised adversarial defense framework, namely FaceGuard, that can automatically detect, localize, and purify a wide variety of adversarial faces without utilizing pre-computed adversarial training samples. During training, FaceGuard automatically synthesizes challenging and diverse adversarial attacks, enabling a classifier to learn to distinguish them from real faces. Concurrently, a purifier attempts to remove the adversarial perturbations in the image space. Experimental results on LFW, Celeb-A, and FFHQ datasets show that FaceGuard can achieve 99.81%, 98.73%, and 99.35% detection accuracies, respectively, on six unseen adversarial attack types. In addition, the proposed method can enhance the face recognition performance of ArcFace from 34.27% TAR @ 0.1% FAR under no defense to 77.46% TAR @ 0.1% FAR. Code, pre-trained models and dataset will be publicly available. Debayan Deb, Xiaoming Liu 0002, Anil K. Jain 0001 |
FG | 2 |
| 2023 | Unified Detection of Digital and Physical Face AttacksabstractState-of-the-art defense mechanisms against face attacks achieve near perfect accuracies within one of three attack categories, namely adversarial, digital manipulation, or physical spoofs, however, they fail to generalize well when tested across all three categories. Poor generalization can be attributed to learning incoherent attacks jointly. To over-come this shortcoming, we propose a unified attack detection framework, namely UniFAD, that can automatically cluster 25 coherent attack types belonging to the three categories. Using a multi-task learning framework along with k-means clustering, UniFAD learns joint representations for coherent attacks, while uncorrelated attack types are learned separately. Proposed UniFAD outperforms prevailing defense methods and their fusion with an overall TDR = 94.73% @ 0.2% FDR on a large fake face dataset consisting of 341K bona fide images and 448K attack images of 25 types across all 3 categories. Proposed method can detect an attack within 3 milliseconds on a Nvidia 2080Ti. UniFAD can also identify the attack categories with 97.37% accuracy. Code and dataset will be publicly available. Debayan Deb, Xiaoming Liu 0002, Anil K. Jain 0001 |
FG | 2 |
| 2023 | Camera Self-Calibration Using Human FacesabstractDespite recent advancements in depth estimation and face alignment, it remains difficult to predict the distance to a human face in arbitrary videos due to the lack of camera calibration. A typical pipeline is to perform calibration with a checkerboard before the video capture, but this is inconvenient to users or impossible for unknown cameras. This work proposes to use the human face as the calibration object to estimate metric depth information and camera intrinsics. Our novel approach alternates between optimizing the 3D face and the camera intrinsics parameterized by a neural network. Compared to prior work, our method performs camera calibration on a larger variety of videos captured by unknown cameras. Further, due to the face prior, our method is more robust to noise in 2D observations compared to previous self-calibration methods. We show that our method improves calibration and depth prediction accuracy over prior works on both synthetic and real data. Code will be available at https://github.com/yhu9/FaceCalibration. Masa Hu, Garrick Brazil, Nanxiang Li, Liu Ren 0001, Xiaoming Liu 0002 |
FG | 5 |
| 2023 | AG-ReID 2023: Aerial-Ground Person Re-identification Challenge ResultsabstractPerson re-identification (Re-ID) on aerial-ground platforms has emerged as an intriguing topic within computer vision, presenting a plethora of unique challenges. Highflying altitudes of aerial cameras make persons appear differently in terms of viewpoints, poses, and resolution compared to the images of the same person viewed from ground cameras. Despite its potential, few algorithms have been developed for person re-identification on aerial-ground data, mainly due to the absence of comprehensive datasets. In response, we have collected a large-scale dataset and organized the Aerial-Ground person Re-IDentification Challenge (AG-ReID2023) to foster advancements in the field. The dataset comprises 100,502 images with 1,615 unique identities, including 51,530 training images featuring 807 identities. The test set is divided into two subsets: Aerial to Ground (808 ids, 4,348 query images, 19,259 gallery images) and Ground to Aerial (808 ids, 4,151 query images, 21,214 gallery images). In addition, we manually annotate individuals with their matching IDs across cameras and provide 15 soft attribute labels. The AG-ReID2023 Challenge in conjunction with the 7thIEEE International Joint Conference on Biometrics (IJCB) has garnered interest from numerous institutes, resulting in the submission of five distinct algorithms. We provide an in-depth examination of the evaluation outcomes and present our findings from the contest. For additional details, kindly refer to the official website1.1https://agreid23.github.io. Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan, Feng Liu 0037, Xiaoming Liu 0002, Arun Ross, Dana Michalski, Debayan Deb, Mahak Kothari, Manisha Saini, Dawei Du, Scott McCloskey, Gabriel Bertocco, Fernanda A. Andaló, Terrance E. Boult, Anderson Rocha 0001, Haidong Zhu, Zhaoheng Zheng, Ramakant Nevatia, Zaigham A. Randhawa, Sinan Sabri, Gianfranco Doretto |
IJCB | 5 |
| 2023 | Learning Clothing and Pose Invariant 3D Shape Representation for Long-Term Person Re-IdentificationabstractLong-Term Person Re-Identification (LT-ReID) has become increasingly crucial in computer vision and biometrics. In this work, we aim to extend LT-ReID beyond pedestrian recognition to include a wider range of real-world human activities while still accounting for cloth-changing scenarios over large time gaps. This setting poses additional challenges due to the geometric misalignment and appearance ambiguity caused by the diversity of human pose and clothing. To address these challenges, we propose a new approach 3DInvarReID for (i) disentangling identity from non-identity components (pose, clothing shape, and texture) of 3D clothed humans, and (ii) reconstructing accurate 3D clothed body shapes and learning discriminative features of naked body shapes for person ReID in a joint manner. To better evaluate our study of LT-ReID, we collect a real-world dataset called CCDA, which contains a wide variety of human activities and clothing changes. Experimentally, we show the superior performance of our approach for person ReID. Code is available at http://cvlab.cse.msu.edu/project-reid3dinvar.html. Feng Liu 0037, ZiAng Gu, Xiaoming Liu 0002 |
ICCV | 5 |
| 2023 | On the Detection, Localization, and Reverse Engineering of Diverse Image ManipulationsabstractWith the abundance of imagery data captured in our daily life, there is an increasing amount of diverse manipulations being applied to imagery, including generative model based image generation and manipulation, adversarial attacks, classic image editing such as splicing, etc. As the imagery plays important roles in our society, it is necessary to understand whether, where and how a given image is manipulated. From the perspective of a defender, this talk will introduce our recent efforts on detecting these manipulations individually and jointly, as well as reverse engineering the various information regarding the manipulation process. We will also share some of the topics that warrant future research. Xiaoming Liu 0002 |
IH&MMSec | 1 |
| 2023 | Mozart: A Mobile ToF System for Sensing in the Dark through Phase ManipulationabstractSensing in low-light and dark environments has a wide range of applications. However, existing sensing technologies suffer several major challenges, such as excessive noise and low resolution. This paper proposes Mozart - a new mobile sensing system that leverages off-the-shelf Time-of-Flight (ToF) depth cameras to generate high-resolution and rich-in-texture maps for applications in dark scenarios. The design of Mozart is based on our key observation that the phase components of ToF measurements can be manipulated to expose texture information. Through in-depth analysis of the physical reflection model, we show that the textures can be exposed and enhanced using highly compute-efficient phase manipulation functions. By exploiting the physics texture models, we propose an autoencoder-based unsupervised learning approach that can automatically learn efficient representations from phase components to generate high-resolution maps. We implemented Mozart on several Android smartphone models1, and an edge testbed with standalone ToF camera platforms for various applications in the dark. The results show that Mozart can work in real time and delivers significant improvement over existing sensing technologies. Therefore, Mozart offers a low-cost, high-performance sensing technology for next-generation applications in the dark. Xiaomin Ouyang, Li Pan 0004, Wenrui Lu, Guoliang Xing, Xiaoming Liu 0002 |
MobiSys | 6 |
| 2023 | PrObeD: Proactive Object Detection WrapperabstractPrevious research in $2D$ object detection focuses on various tasks, including detecting objects in generic and camouflaged images. These works are regarded as passive works for object detection as they take the input image as is. However, convergence to global minima is not guaranteed to be optimal in neural networks; therefore, we argue that the trained weights in the object detector are not optimal. To rectify this problem, we propose a wrapper based on proactive schemes, PrObeD, which enhances the performance of these object detectors by learning a signal. PrObeD consists of an encoder-decoder architecture, where the encoder network generates an image-dependent signal termed templates to encrypt the input images, and the decoder recovers this template from the encrypted images. We propose that learning the optimum template results in an object detector with an improved detection performance. The template acts as a mask to the input images to highlight semantics useful for the object detector. Finetuning the object detector with these encrypted images enhances the detection performance for both generic and camouflaged. Our experiments on MS-COCO, CAMO, COD$10$K, and NC$4$K datasets show improvement over different detectors after applying PrObeD. Our models/codes are available at https://github.com/vishal3477/Proactive-Object-Detection. Vishal Asnani, Abhinav Kumar 0004, Suya You, Xiaoming Liu 0002 |
NeurIPS | 4 |
| 2023 | ChatGPT-Powered Hierarchical Comparisons for Image ClassificationabstractThe zero-shot open-vocabulary setting poses challenges for image classification.
Fortunately, utilizing a vision-language model like CLIP, pre-trained on image-text
pairs, allows for classifying images by comparing embeddings. Leveraging large
language models (LLMs) such as ChatGPT can further enhance CLIP’s accuracy
by incorporating class-specific knowledge in descriptions. However, CLIP still
exhibits a bias towards certain classes and generates similar descriptions for similar
classes, disregarding their differences. To address this problem, we present a
novel image classification framework via hierarchical comparisons. By recursively
comparing and grouping classes with LLMs, we construct a class hierarchy. With
such a hierarchy, we can classify an image by descending from the top to the bottom
of the hierarchy, comparing image and text embeddings at each level. Through
extensive experiments and analyses, we demonstrate that our proposed approach is
intuitive, effective, and explainable. Code will be released upon publication. Yiyang Su, Xiaoming Liu 0002 |
NeurIPS | 3 |
| 2023 | Tame a Wild Camera: In-the-Wild Monocular Camera Calibrationabstract3D sensing for monocular in-the-wild images, e.g., depth estimation and 3D object detection, has become increasingly important.
However, the unknown intrinsic parameter hinders their development and deployment.
Previous methods for the monocular camera calibration rely on specific 3D objects or strong geometry prior, such as using a checkerboard or imposing a Manhattan World assumption.
This work instead calibrates intrinsic via exploiting the monocular 3D prior.
Given an undistorted image as input, our method calibrates the complete 4 Degree-of-Freedom (DoF) intrinsic parameters.
First, we show intrinsic is determined by the two well-studied monocular priors: monocular depthmap and surface normal map.
However, this solution necessitates a low-bias and low-variance depth estimation.
Alternatively, we introduce the incidence field, defined as the incidence rays between points in 3D space and pixels in the 2D imaging plane.
We show that: 1) The incidence field is a pixel-wise parametrization of the intrinsic invariant to image cropping and resizing.
2) The incidence field is a learnable monocular 3D prior, determined pixel-wisely by up-to-sacle monocular depthmap and surface normal.
With the estimated incidence field, a robust RANSAC algorithm recovers intrinsic.
We show the effectiveness of our method through superior performance on synthetic and zero-shot testing datasets.
Beyond calibration, we demonstrate downstream applications in image manipulation detection \& restoration, uncalibrated two-view pose estimation, and 3D sensing. Abhinav Kumar 0004, Masa Hu, Xiaoming Liu 0002 |
NeurIPS | 4 |
| 2023 | Reverse Engineering of Generative Models: Inferring Model Hyperparameters From Generated ImagesabstractState-of-the-art (SOTA) Generative Models (GMs) can synthesize photo-realistic images that are hard for humans to distinguish from genuine photos. Identifying and understanding manipulated media are crucial to mitigate the social concerns on the potential misuse of GMs. We propose to perform reverse engineering of GMs to infer model hyperparameters from the images generated by these models. We define a novel problem, "model parsing", as estimating GM network architectures and training loss functions by examining their generated images - a task seemingly impossible for human beings. To tackle this problem, we propose a framework with two components: a Fingerprint Estimation Network (FEN), which estimates a GM fingerprint from a generated image by training with four constraints to encourage the fingerprint to have desired properties, and a Parsing Network (PN), which predicts network architecture and loss functions from the estimated fingerprints. To evaluate our approach, we collect a fake image dataset with 100 K images generated by 116 different GMs. Extensive experiments show encouraging results in parsing the hyperparameters of the unseen models. Finally, our fingerprint estimation can be leveraged for deepfake detection and image attribution, as we show by reporting SOTA results on both the deepfake detection (Celeb-DF) and image attribution benchmarks. Vishal Asnani, Xi Yin 0001, Tal Hassner, Xiaoming Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Spoof Trace Disentanglement for Generic Face Anti-SpoofingabstractPrior studies show that the key to face anti-spoofing lies in the subtle image patterns, termed "spoof trace," e.g., color distortion, 3D mask edge, and Moiré pattern. Spoof detection rooted on those spoof traces can improve not only the model's generalization but also the interpretability. Yet, it is a challenging task due to the diversity of spoof attacks and the lack of ground truth for spoof traces. In this work, we propose a novel adversarial learning framework to explicitly estimate the spoof related patterns for face anti-spoofing. Inspired by the physical process, spoof faces are disentangled into spoof traces and the live counterparts in two steps: additive step and inpainting step. This two-step modeling can effectively narrow down the searching space for adversarial learning of spoof trace. Based on the trace modeling, the disentangled spoof traces can be utilized to reversely construct new spoof faces, which is used as data augmentation to effectively tackle long-tail spoof types. In addition, we apply frequency-based image decomposition in both the input and disentangled traces to better reflect the low-level vision cues. Our approach demonstrates superior spoof detection performance on 3 testing scenarios: known attacks, unknown attacks, and open-set attacks. Meanwhile, it provides a visually-convincing estimation of the spoof traces. Source code and pre-trained models will be publicly available upon publication. Yaojie Liu, Xiaoming Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | MOST-GAN: 3D Morphable StyleGAN for Disentangled Face Image ManipulationabstractRecent advances in generative adversarial networks (GANs) have led to remarkable achievements in face image synthesis. While methods that use style-based GANs can generate strikingly photorealistic face images, it is often difficult to control the characteristics of the generated faces in a meaningful and disentangled way. Prior approaches aim to achieve such semantic control and disentanglement within the latent space of a previously trained GAN. In contrast, we propose a framework that a priori models physical attributes of the face such as 3D shape, albedo, pose, and lighting explicitly, thus providing disentanglement by design. Our method, MOST-GAN, integrates the expressive power and photorealism of style-based GANs with the physical disentanglement and flexibility of nonlinear 3D morphable models, which we couple with a state-of-the-art 2D hair manipulation network. MOST-GAN achieves photorealistic manipulation of portrait images with fully disentangled 3D control over their physical attributes, enabling extreme manipulation of lighting, facial expression, and pose variations up to full profile view. Safa C. Medin, Bernhard Egger 0001, Anoop Cherian, Ye Wang 0001, Josh Tenenbaum, Xiaoming Liu 0002, Tim K. Marks |
AAAI | 6 |
| 2022 | Blind Removal of Facial Foreign Shadows
Yaojie Liu, Andrew Z. Hou, Xinyu Huang 0001, Liu Ren 0001, Xiaoming Liu 0002 |
BMVC | 5 |
| 2022 | Proactive Image Manipulation DetectionabstractImage manipulation detection algorithms are often trained to discriminate between images manipulated with particular Generative Models (GMs) and genuine/real images, yet generalize poorly to images manipulated with GMs unseen in the training. Conventional detection algorithms receive an input image passively. By contrast, we propose a proactive scheme to image manipulation detection. Our key enabling technique is to estimate a set of templates which when added onto the real image would lead to more accurate manipulation detection. That is, a template protected real image, and its manipulated version, is better discriminated compared to the original real image vs. its manipulated one. These templates are estimated using certain constraints based on the desired properties of templates. For image manipulation detection, our proposed approach outperforms the prior work by an average precision of 16%for CycleGAN and 32% for GauGAN. Our approach is generalizable to a variety of GMs showing an improvement over prior work by an average precision of 10% averaged across 12 GMs. Our code is available at https://www.github.com/vishal3477/proactive_IMD. Vishal Asnani, Xi Yin 0001, Tal Hassner, Sijia Liu 0001, Xiaoming Liu 0002 |
CVPR | 5 |
| 2022 | Face Relighting with Geometrically Consistent ShadowsabstractMost face relighting methods are able to handle diffuse shadows, but struggle to handle hard shadows, such as those cast by the nose. Methods that propose techniques for handling hard shadows often do not produce geometrically consistent shadows since they do not directly leverage the estimated face geometry while synthesizing them. We propose a novel differentiable algorithm for synthesizing hard shadows based on ray tracing, which we incorporate into training our face relighting model. Our proposed algorithm directly utilizes the estimated face geometry to synthesize geometrically consistent hard shadows. We demonstrate through quantitative and qualitative experiments on Multi-PIE and FFHQ that our method produces more geometrically consistent shadows than previous face relighting methods while also achieving state-of-the-art face relighting performance under directional lighting. In addition, we demonstrate that our differentiable hard shadow modeling improves the quality of the estimated face geometry over diffuse shading models. Andrew Z. Hou, Michel Sarkis, Ning Bi, Yiying Tong, Xiaoming Liu 0002 |
CVPR | 5 |
| 2022 | AdaFace: Quality Adaptive Margin for Face RecognitionabstractRecognition in low quality face datasets is challenging because facial attributes are obscured and degraded. Advances in margin-based loss functions have resulted in enhanced discriminability of faces in the embedding space. Further, previous studies have studied the effect of adaptive losses to assign more importance to misclassified (hard) examples. In this work, we introduce another aspect of adaptiveness in the loss function, namely the image quality. We argue that the strategy to emphasize misclassified samples should be adjusted according to their image quality. Specifically, the relative importance of easy or hard samples should be based on the sample's image quality. We propose a new loss function that emphasizes samples of different difficulties based on their image quality. Our method achieves this in the form of an adaptive margin function by approximating the image quality with feature norms. Extensive experiments show that our method, AdaFace, improves the face recognition performance over the state-of-the-art (SoTA) on four datasets (IJB-B, IJB-C, IJB-S and TinyFace). Code and models are released in Supp. Anil K. Jain 0001, Xiaoming Liu 0002 |
CVPR | 3 |
| 2022 | Multi-domain Learning for Updating Face Anti-spoofing Models
Yaojie Liu, Anil K. Jain 0001, Xiaoming Liu 0002 |
ECCV (13) | 4 |
| 2022 | DEVIANT: Depth EquiVarIAnt NeTwork for Monocular 3D Object Detection
Abhinav Kumar 0004, Garrick Brazil, Enrique Corona, Armin Parchami, Xiaoming Liu 0002 |
ECCV (9) | 5 |
| 2022 | Controllable and Guided Face Synthesis for Unconstrained Face Recognition
Feng Liu 0037, Anil K. Jain 0001, Xiaoming Liu 0002 |
ECCV (12) | 4 |
| 2022 | 2D GANs Meet Unsupervised Single-View 3D Reconstruction
Feng Liu 0037, Xiaoming Liu 0002 |
ECCV (1) | 2 |
| 2022 | Reverse Engineering of Imperceptible Adversarial Image Perturbations
Yifan Gong 0004, Yuguang Yao, Xiaoming Liu 0002, Xue Lin 0001, Sijia Liu 0001 |
ICLR | 5 |
| 2022 | HiToF: a ToF camera system for capturing high-resolution texturesabstractWe present a demonstration of an enhanced Time-of-Flight (ToF) depth system named HiToF, which can expose high-resolution textures from captured depth maps. By design, a ToF camera can easily capture the depth maps of a scene while largely omitting the corresponding texture information, which is often critical for the performance of many depth applications. HiToF is developed to address this issue by generating enhanced depth maps with high-resolution textures. The key idea is to manipulate the phase components used in the measurement of time-of-flight for the received IR light. In this demo, we showcase our implementation using off-the-shelf ToF cameras and engage audience with an interactive experience in various scenarios, which illustrates the system's effectiveness in improving the performance of ToF cameras in depth applications. Xiaomin Ouyang, Li Pan 0004, Wenrui Lu, Xiaoming Liu 0002, Guoliang Xing |
MobiCom | 5 |
| 2022 | Cluster and Aggregate: Face Recognition with Large Probe SetabstractFeature fusion plays a crucial role in unconstrained face recognition where inputs (probes) comprise of a set of $N$ low quality images whose individual qualities vary. Advances in attention and recurrent modules have led to feature fusion that can model the relationship among the images in the input set. However, attention mechanisms cannot scale to large $N$ due to their quadratic complexity and recurrent modules suffer from input order sensitivity. We propose a two-stage feature fusion paradigm, Cluster and Aggregate, that can both scale to large $N$ and maintain the ability to perform sequential inference with order invariance. Specifically, Cluster stage is a linear assignment of $N$ inputs to $M$ global cluster centers, and Aggregation stage is a fusion over $M$ clustered features. The clustered features play an integral role when the inputs are sequential as they can serve as a summarization of past features. By leveraging the order-invariance of incremental averaging operation, we design an update rule that achieves batch-order invariance, which guarantees that the contributions of early image in the sequence do not diminish as time steps increase. Experiments on IJB-B and IJB-S benchmark datasets show the superiority of the proposed two-stage paradigm in unconstrained face recognition. Feng Liu 0037, Anil K. Jain 0001, Xiaoming Liu 0002 |
NeurIPS | 4 |
| 2022 | On Learning Disentangled Representations for Gait RecognitionabstractGait, the walking pattern of individuals, is one of the important biometrics modalities. Most of the existing gait recognition methods take silhouettes or articulated body models as gait features. These methods suffer from degraded recognition performance when handling confounding variables, such as clothing, carrying and viewing angle. To remedy this issue, we propose a novel AutoEncoder framework, GaitNet, to explicitly disentangle appearance, canonical and pose features from RGB imagery. The LSTM integrates pose features over time as a dynamic gait feature while canonical features are averaged as a static gait feature. Both of them are utilized as classification features. In addition, we collect a Frontal-View Gait (FVG) dataset to focus on gait recognition from frontal-view walking, which is a challenging problem since it contains minimal gait cues compared to other views. FVG also includes other important variations, e.g., walking speed, carrying, and clothing. With extensive experiments on CASIA-B, USF, and FVG datasets, our method demonstrates superior performance to the SOTA quantitatively, the ability of feature disentanglement qualitatively, and promising computational efficiency. We further compare our GaitNet with state-of-the-art face recognition to demonstrate the advantages of gait biometrics identification under certain scenarios, e.g., long-distance/lower resolutions, cross viewing angles. Source code is available at http://cvlab.cse.msu.edu/project-gaitnet.html. Ziyuan Zhang 0004, Luan Tran, Feng Liu 0037, Xiaoming Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | PSCC-Net: Progressive Spatio-Channel Correlation Network for Image Manipulation Detection and LocalizationabstractTo defend against manipulation of image content, such as splicing, copy-move, and removal, we develop a Progressive Spatio-Channel Correlation Network (PSCC-Net) to detect and localize image manipulations. PSCC-Net processes the image in a two-path procedure: a top-down path that extracts local and global features and a bottom-up path that detects whether the input image is manipulated, and estimates its manipulation masks at multiple scales, where each mask is conditioned on the previous one. Different from the conventional encoder-decoder and no-pooling structures, PSCC-Net leverages features at different scales with dense cross-connections to produce manipulation masks in a coarse-to-fine fashion. Moreover, a Spatio-Channel Correlation Module (SCCM) captures both spatial and channel-wise correlations in the bottom-up path, which endows features with holistic cues, enabling the network to cope with a wide range of manipulation attacks. Thanks to the light-weight backbone and progressive mechanism, PSCC-Net can process$1,080\text{P}$images at 50+FPS. Extensive experiments demonstrate the superiority of PSCC-Net over the state-of-the-art methods on both detection and localization. Codes and models are available athttps://github.com/proteus1991/PSCC-Net. Xiaohong Liu 0001, Yaojie Liu, Jun Chen 0005, Xiaoming Liu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | GrooMeD-NMS: Grouped Mathematically Differentiable NMS for Monocular 3D Object DetectionabstractModern 3D object detectors have immensely benefited from the end-to-end learning idea. However, most of them use a post-processing algorithm called Non-Maximal Suppression (NMS) only during inference. While there were attempts to include NMS in the training pipeline for tasks such as 2D object detection, they have been less widely adopted due to a non-mathematical expression of the NMS. In this paper, we present and integrate GrooMeD-NMS – a novel Grouped Mathematically Differentiable NMS for monocular 3D object detection, such that the network is trained end-to-end with a loss on the boxes after NMS. We first formulate NMS as a matrix operation and then group and mask the boxes in an unsupervised manner to obtain a simple closed-form expression of the NMS. GrooMeD-NMS addresses the mismatch between training and inference pipelines and, therefore, forces the network to select the best 3D box in a differentiable manner. As a result, GrooMeD-NMS achieves state-of-the-art monocular 3D object detection results on the KITTI benchmark dataset performing comparably to monocular video-based methods. Abhinav Kumar 0004, Garrick Brazil, Xiaoming Liu 0002 |
CVPR | 3 |
| 2021 | Fully Understanding Generic Objects: Modeling, Segmentation, and ReconstructionabstractInferring 3D structure of a generic object from a 2D image is a long-standing objective of computer vision. Conventional approaches either learn completely from CADgenerated synthetic data, which have difficulty in inference from real images, or generate 2.5D depth image via intrinsic decomposition, which is limited compared to the full 3D reconstruction. One fundamental challenge lies in how to leverage numerous real 2D images without any 3D ground truth. To address this issue, we take an alternative approach with semi-supervised learning. That is, for a 2D image of a generic object, we decompose it into latent representations of category, shape, albedo, lighting and camera projection matrix, decode the representations to segmented 3D shape and albedo respectively, and fuse these components to render an image well approximating the input image. Using a category-adaptive 3D joint occupancy field (JOF), we show that the complete shape and albedo modeling enables us to leverage real 2D images in both modeling and model fitting. The effectiveness of our approach is demonstrated through superior 3D reconstruction from a single image, being either synthetic or real, and shape segmentation. Code is available at http://cvlab.cse.msu.edu/project-fully3dobject.html. Feng Liu 0037, Luan Tran, Xiaoming Liu 0002 |
CVPR | 3 |
| 2021 | Riggable 3D Face Reconstruction via In-Network OptimizationabstractThis paper presents a method for riggable 3D face reconstruction from monocular images, which jointly estimates a personalized face rig and per-image parameters including expressions, poses, and illuminations. To achieve this goal, we design an end-to-end trainable network embedded with a differentiable in-network optimization. The network first parameterizes the face rig as a compact latent code with a neural decoder, and then estimates the latent code as well as per-image parameters via a learnable optimization. By estimating a personalized face rig, our method goes beyond static reconstructions and enables downstream applications such as video retargeting. In-network optimization explicitly enforces constraints derived from the first principles, thus introduces additional priors than regression-based methods. Finally, data-driven priors from deep learning are utilized to constrain the ill-posed monocular setting and ease the optimization difficulty. Experiments demonstrate that our method achieves SOTA reconstruction accuracy, reasonable robustness and generalization ability, and supports standard face rig applications. Ziqian Bai, Zhaopeng Cui, Xiaoming Liu 0002, Ping Tan 0002 |
CVPR | 3 |
| 2021 | Mitigating Face Recognition Bias via Group Adaptive ClassifierabstractFace recognition is known to exhibit bias - subjects in a certain demographic group can be better recognized than other groups. This work aims to learn a fair face representation, where faces of every group could be more equally represented. Our proposed group adaptive classifier mitigates bias by using adaptive convolution kernels and attention mechanisms on faces based on their demographic attributes. The adaptive module comprises kernel masks and channel-wise attention maps for each demographic group so as to activate different facial regions for identification, leading to more discriminative features pertinent to their demographics. Our introduced automated adaptation strategy determines whether to apply adaptation to a certain layer by iteratively computing the dissimilarity among demographic-adaptive parameters. A new de-biasing loss function is proposed to mitigate the gap of average intra-class distance between demographic groups. Experiments on face benchmarks (RFW, LFW, IJB-A, and IJB-C) show that our work is able to mitigate face recognition bias across demographic groups while maintaining the competitive accuracy. Sixue Gong, Xiaoming Liu 0002, Anil K. Jain 0001 |
CVPR | 2 |
| 2021 | Towards High Fidelity Face Relighting With Realistic ShadowsabstractExisting face relighting methods often struggle with two problems: maintaining the local facial details of the subject and accurately removing and synthesizing shadows in the relit image, especially hard shadows. We propose a novel deep face relighting method that addresses both problems. Our method learns to predict the ratio (quotient) image between a source image and the target image with the desired lighting, allowing us to relight the image while maintaining the local facial details. During training, our model also learns to accurately modify shadows by using estimated shadow masks to emphasize on the high-contrast shadow borders. Furthermore, we introduce a method to use the shadow mask to estimate the ambient light intensity in an image, and are thus able to leverage multiple datasets during training with different global lighting intensities. With quantitative and qualitative evaluations on the Multi-PIE and FFHQ datasets, we demonstrate that our proposed method faithfully maintains the local facial details of the subject and can accurately handle hard shadows while achieving state-of-the-art face relighting performance. Andrew Z. Hou, Michel Sarkis, Ning Bi, Yiying Tong, Xiaoming Liu 0002 |
CVPR | 6 |
| 2021 | Depth Completion With Twin Surface Extrapolation at Occlusion BoundariesabstractDepth completion starts from a sparse set of known depth values and estimates the unknown depths for the remaining image pixels. Most methods model this as depth interpolation and erroneously interpolate depth pixels into the empty space between spatially distinct objects, resulting in depth-smearing across occlusion boundaries. Here we propose a multi-hypothesis depth representation that explicitly models both foreground and background depths in the difficult occlusion-boundary regions. Our method can be thought of as performing twin-surface extrapolation, rather than interpolation, in these regions. Next our method fuses these extrapolated surfaces into a single depth image leveraging the image data. Key to our method is the use of an asymmetric loss function that operates on a novel twin-surface representation. This enables us to train a network to simultaneously do surface extrapolation and surface fusion. We characterize our loss function and compare with other common losses. Finally, we validate our method on three different datasets; KITTI, an outdoor real-world dataset, NYU2, indoor real-world depth dataset and Virtual KITTI, a photo-realistic synthetic dataset with dense groundtruth, and demonstrate improvement over the state of the art. Saif Muhammad Imran, Xiaoming Liu 0002, Daniel D. Morris |
CVPR | 2 |
| 2021 | Radar-Camera Pixel Depth Association for Depth CompletionabstractWhile radar and video data can be readily fused at the detection level, fusing them at the pixel level is potentially more beneficial. This is also more challenging in part due to the sparsity of radar, but also because automotive radar beams are much wider than a typical pixel combined with a large baseline between camera and radar, which results in poor association between radar pixels and color pixel. A consequence is that depth completion methods designed for LiDAR and video fare poorly for radar and video. Here we propose a radar-to-pixel association stage which learns a mapping from radar returns to pixels. This mapping also serves to densify radar returns. Using this as a first stage, followed by a more traditional depth completion method, we are able to achieve image-guided depth completion with radar and video. We demonstrate performance superior to camera and radar alone on the nuScenes dataset. Our source code is available at https://github.com/longyunf/rc-pda. Daniel D. Morris, Xiaoming Liu 0002, Marcos Castro, Punarjay Chakravarty, Praveen Narayanan |
CVPR | 3 |
| 2021 | Full-Velocity Radar Returns by Radar-Camera FusionabstractA distinctive feature of Doppler radar is the measurement of velocity in the radial direction for radar points. However, the missing tangential velocity component hampers object velocity estimation as well as temporal integration of radar sweeps in dynamic scenes. Recognizing that fusing camera with radar provides complementary information to radar, in this paper we present a closed-form solution for the point-wise, full-velocity estimate of Doppler returns using the corresponding optical flow from camera images. Additionally, we address the association problem between radar returns and camera images with a neural network that is trained to estimate radar-camera correspondences. Experimental results on the nuScenes dataset verify the validity of the method and show significant improvements over the state-of-the-art in velocity estimation and accumulation of radar points. Daniel D. Morris, Xiaoming Liu 0002, Marcos Castro, Punarjay Chakravarty, Praveen Narayanan |
ICCV | 3 |
| 2021 | Turnip: Time-Series U-Net With Recurrence For NIR Imaging PPGabstractImaging photoplethysmography (iPPG) is the process of estimating the waveform of a person’s pulse by processing a video of their face to detect minute color or intensity changes in the skin. Typically, iPPG methods use three-channel RGB video to address challenges due to motion. In situations such as driving, however, illumination in the visible spectrum is often quickly varying (e.g., daytime driving through shadows of trees and buildings) or insufficient (e.g., night driving). In such cases, a practical alternative is to use active illumination and bandpass-filtering from a monochromatic near-infrared (NIR) light source and camera. Contrary to learning-based iPPG solutions designed for multi-channel RGB, previous work in single-channel NIR iPPG has been based on hand-crafted models (with only a few manually tuned parameters), exploiting the sparsity of the PPG signal in the frequency domain. In contrast, we propose a modular framework for iPPG estimation of the heartbeat signal, in which the first module extracts a time-series signal from monochromatic NIR face video. The second module consists of a novel time-series U-net architecture in which a GRU (gated recurrent unit) network has been added to the passthrough layers. We test our approach on the challenging MR-NIRP Car Dataset, which consists of monochromatic NIR videos taken in both stationary and driving conditions. Our model’s iPPG estimation performance on NIR video outperforms both the state-of-the-art model-based method and a recent end-to-end deep learning method that we adapted to monochromatic video. Armand Comas, Tim K. Marks, Hassan Mansour, Suhas Lohit, Yechi Ma, Xiaoming Liu 0002 |
ICIP | 6 |
| 2021 | Voxel-based 3D Detection and Reconstruction of Multiple Objects from a Single ImageabstractInferring 3D locations and shapes of multiple objects from a single 2D image is a long-standing objective of computer vision. Most of the existing works either predict one of these 3D properties or focus on solving both for a single object. One fundamental challenge lies in how to learn an effective representation of the image that is well-suited for 3D detection and reconstruction. In this work, we propose to learn a regular grid of 3D voxel features from the input image which is aligned with 3D scene space via a 3D feature lifting operator. Based on the 3D voxel features, our novel CenterNet-3D detection head formulates the 3D detection as keypoint detection in the 3D space. Moreover, we devise an efficient coarse-to-fine reconstruction module, including coarse-level voxelization and a novel local PCA-SDF shape representation, which enables fine detail reconstruction and two orders of magnitude faster inference than prior methods. With complementary supervision from both 3D detection and reconstruction, one enables the 3D voxel features to be geometry and context preserving, benefiting both tasks. The effectiveness of our approach is demonstrated through 3D detection and reconstruction on single-object and multiple-object scenarios. Feng Liu 0037, Xiaoming Liu 0002 |
NeurIPS | 2 |
| 2021 | UltraDepth: Exposing High-Resolution Texture from Depth CamerasabstractTime-of-flight (ToF) depth cameras have been increasingly adopted in various real-world applications, e.g., used with RGB cameras for advanced computer vision tasks like 3-D mapping or deployed alone in privacy-sensitive applications such as sleep monitoring. In this paper, we propose UltraDepth, the first system that can expose high-resolution texture from depth maps captured by off-the-shelf ToF cameras, simply by introducing a distorting IR source. The exposed texture information can significantly augment depth-based applications. Moreover, such a capability can be used to launch privacy attacks, which poses a major concern due to the prominence of ToF cameras. To design UltraDepth, we present an in-depth analysis on the impact of the distorting IR light on the distance measurement. We further show that, the reflection properties (reflectivity and incidence angle) of the objects will be encoded in the distorted depth map and hence can be leveraged to reveal texture of objects in UltraDepth. We then propose two practical implementations of UltraDepth, i.e., reflection-based and external IR-based implementations. Our extensive real-world experiments show that, the depth maps output by UltraDepth achieve 89.06%, 99.33%, 81.25% mean accuracy in object detection, face recognition and character recognition, respectively, which offers over 10x improvement over the ordinary depth maps and even approaches the performance of RGB and IR images in a number of scenarios. The findings of this work provide key insights for new research on depth-related computer vision and security of depth sensing devices. Xiaomin Ouyang, Xiaoming Liu 0002, Guoliang Xing |
SenSys | 3 |
| 2021 | Shape My Face: Registering 3D Face Scans by Surface-to-Surface TranslationabstractAbstract Standard registration algorithms need to be independently applied to each surface to register, following careful pre-processing and hand-tuning. Recently, learning-based approaches have emerged that reduce the registration of new scans to running inference with a previously-trained model. The potential benefits are multifold: inference is typically orders of magnitude faster than solving a new instance of a difficult optimization problem, deep learning models can be made robust to noise and corruption, and the trained model may be re-used for other tasks, e.g. through transfer learning. In this paper, we cast the registration task as a surface-to-surface translation problem, and design a model to reliably capture the latent geometric information directly from raw 3D face scans. We introduce Shape-My-Face (SMF), a powerful encoder-decoder architecture based on an improved point cloud encoder, a novel visual attention mechanism, graph convolutional decoders with skip connections, and a specialized mouth model that we smoothly integrate with the mesh convolutions. Compared to the previous state-of-the-art learning algorithms for non-rigid registration of face scans, SMF only requires the raw data to be rigidly aligned (with scaling) with a pre-defined face template. Additionally, our model provides topologically-sound meshes with minimal supervision, offers faster training time, has orders of magnitude fewer trainable parameters, is more robust to noise, and can generalize to previously unseen datasets. We extensively evaluate the quality of our registrations on diverse data. We demonstrate the robustness and generalizability of our model with in-the-wild face scans across different modalities, sensor types, and resolutions. Finally, we show that, by learning to register scans, SMF produces a hybrid linear and non-linear morphable model. Manipulation of the latent space of SMF allows for shape generation, and morphing applications such as expression transfer in-the-wild. We train SMF on a dataset of human faces comprising 9 large-scale databases on commodity hardware. Mehdi Bahri, Eimear O' Sullivan, Shunwang Gong, Feng Liu 0037, Xiaoming Liu 0002, Michael M. Bronstein, Stefanos Zafeiriou |
Int. J. Comput. Vis. | 5 |
| 2021 | On Learning 3D Face Morphable Model from In-the-Wild ImagesabstractAs a classic statistical model of 3D facial shape and albedo, 3D Morphable Model (3DMM) is widely used in facial analysis, e.g., model fitting, image synthesis. Conventional 3DMM is learned from a set of 3D face scans with associated well-controlled 2D face images, and represented by two sets of PCA basis functions. Due to the type and amount of training data, as well as, the linear bases, the representation power of 3DMM can be limited. To address these problems, this paper proposes an innovative framework to learn a nonlinear 3DMM model from a large set of in-the-wild face images, without collecting 3D face scans. Specifically, given a face image as input, a network encoder estimates the projection, lighting, shape and albedo parameters. Two decoders serve as the nonlinear 3DMM to map from the shape and albedo parameters to the 3D shape and albedo, respectively. With the projection parameter, lighting, 3D shape, and albedo, a novel analytically-differentiable rendering layer is designed to reconstruct the original input face. The entire network is end-to-end trainable with only weak supervision. We demonstrate the superior representation power of our nonlinear 3DMM over its linear counterpart, and its contribution to face alignment, 3D reconstruction, and face editing. Source code and additional results can be found at our project page: http://cvlab.cse.msu.edu/project-nonlinear-3dmm.html. Luan Tran, Xiaoming Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Deep learning-based apple detection using a suppression mask R-CNN
Pengyu Chu, Zhaojian Li 0001, Kyle Lammers, Renfu Lu, Xiaoming Liu 0002 |
Pattern Recognit. Lett. | 5 |
| 2021 | The Untold Secrets of WiFi-Calling Services: Vulnerabilities, Attacks, and CountermeasuresabstractSince 2016, all of four major U.S. operators have rolled out Wi-Fi calling services. They enable mobile users to place cellular calls over Wi-Fi networks based on the 3GPP IMS technology. Compared with conventional cellular voice solutions, the major difference lies in that their traffic traverses untrusted Wi-Fi networks and the Internet. This exposure to insecure networks can cause the Wi-Fi calling users to suffer from security threats. Its security mechanisms are similar to the VoLTE, because both of them are supported by the IMS. They include SIM-based security, 3GPP AKA, IPSec, etc. However, are they sufficient to secure Wi-Fi calling services? Unfortunately, our study yields a negative answer. We conduct the first security study on the operational Wi-Fi calling services in three major U.S. operators networks using commodity devices. We disclose that current Wi-Fi calling security is not bullet-proof and uncover three vulnerabilities. By exploiting the vulnerabilities, we devise two proof-of-concept attacks: telephony harassment or denial of voice service and user privacy leakage; both of them can bypass the existing security defenses. We have confirmed their feasibility using real-world experiments, as well as assessed their potential damages and proposed a solution to address all identified vulnerabilities. Tian Xie 0001, Guan-Hua Tu, Bangjie Yin, Chi-Yu Li 0001, Chunyi Peng 0001, Mi Zhang 0002, Hui Liu 0031, Xiaoming Liu 0002 |
IEEE Trans. Mob. Comput. | 8 |
| 2020 | FAN: Feature Adaptation Network for Surveillance Face Recognition and Normalization
Xi Yin 0001, Ying Tai, Yuge Huang, Xiaoming Liu 0002 |
ACCV (2) | 4 |
| 2020 | LUVLi Face Alignment: Estimating Landmarks' Location, Uncertainty, and Visibility LikelihoodabstractModern face alignment methods have become quite accurate at predicting the locations of facial landmarks, but they do not typically estimate the uncertainty of their predicted locations nor predict whether landmarks are visible. In this paper, we present a novel framework for jointly predicting landmark locations, associated uncertainties of these predicted locations, and landmark visibilities. We model these as mixed random variables and estimate them using a deep network trained using our proposed Location, Uncertainty, and Visibility Likelihood (LUVLi) loss. In addition, we release an entirely new labeling of a large face alignment dataset with over 19,000 face images in a full range of head poses. Each face is manually labeled with the ground-truth locations of 68 landmarks, with the additional information of whether each landmarks is visible, self-occluded (due to extreme head poses), or externally occluded. Not only does our joint estimation yield accurate estimates of the uncertainty of predicted landmark locations, but it also yields state-of-the-art estimates for the landmark locations themselves on mulitple standard face alignment datasets. Our method's estimates of the uncertainty of predicted landmark locations could be used to automatically identify input images on which face alignment fails, which can be critical for downstream tasks. Abhinav Kumar 0004, Tim K. Marks, Wenxuan Mou, Ye Wang 0001, Michael J. Jones 0001, Anoop Cherian, Toshiaki Koike-Akino, Xiaoming Liu 0002, Chen Feng 0002 |
CVPR | 8 |
| 2020 | Deep Facial Non-Rigid Multi-View StereoabstractWe present a method for 3D face reconstruction from multi-view images with different expressions. We formulate this problem from the perspective of non-rigid multi-view stereo (NRMVS). Unlike previous learning-based methods, which often regress the face shape directly, our method optimizes the 3D face shape by explicitly enforcing multi-view appearance consistency, which is known to be effective in recovering shape details according to conventional multi-view stereo methods. Furthermore, by estimating face shape through optimization based on multi-view consistency, our method can potentially have better generalization to unseen data. However, this optimization is challenging since each input image has a different expression. We facilitate it with a CNN network that learns to regularize the non-rigid 3D face according to the input image and preliminary optimization results. Extensive experiments show that our method achieves the state-of-the-art performance on various datasets and generalizes well to in-the-wild data. Ziqian Bai, Zhaopeng Cui, Jamal Ahmed Rahim, Xiaoming Liu 0002, Ping Tan 0002 |
CVPR | 4 |
| 2020 | Camera Trace ErasingabstractCamera trace is a unique noise produced in digital imaging process. Most existing forensic methods analyze camera trace to identify image origins. In this paper, we address a new low-level vision problem, camera trace erasing, to reveal the weakness of trace-based forensic methods. A comprehensive investigation on existing anti-forensic methods reveals that it is non-trivial to effectively erase camera trace while avoiding the destruction of content signal. To reconcile these two demands, we propose Siamese Trace Erasing (SiamTE), in which a novel hybrid loss is designed on the basis of Siamese architecture for network training. Specifically, we propose embedded similarity, truncated fidelity, and cross identity to form the hybrid loss. Compared with existing anti-forensic methods, SiamTE has a clear advantage for camera trace erasing, which is demonstrated in three representative tasks. Chang Chen 0004, Zhiwei Xiong, Xiaoming Liu 0002, Feng Wu 0001 |
CVPR | 3 |
| 2020 | On the Detection of Digital Face ManipulationabstractDetecting manipulated facial images and videos is an increasingly important topic in digital media forensics. As advanced face synthesis and manipulation methods are made available, new types of fake face representations are being created which have raised significant concerns for their use in social media. Hence, it is crucial to detect manipulated face images and localize manipulated regions. Instead of simply using multi-task learning to simultaneously detect manipulated images and predict the manipulated mask (regions), we propose to utilize an attention mechanism to process and improve the feature maps for the classification task. The learned attention maps highlight the informative regions to further improve the binary classification (genuine face v. fake face), and also visualize the manipulated regions. To enable our study of manipulated face detection and localization, we collect a large-scale database that contains numerous types of facial forgeries. With this dataset, we perform a thorough analysis of data-driven fake face detection. We show that the use of an attention mechanism improves facial forgery detection and manipulated region localization. Hao Dang, Feng Liu 0037, Joel Stehouwer, Xiaoming Liu 0002, Anil K. Jain 0001 |
CVPR | 4 |
| 2020 | CurricularFace: Adaptive Curriculum Learning Loss for Deep Face RecognitionabstractAs an emerging topic in face recognition, designing margin-based loss functions can increase the feature margin between different classes for enhanced discriminability. More recently, the idea of mining-based strategies is adopted to emphasize the misclassified samples, achieving promising results. However, during the entire training process, the prior methods either do not explicitly emphasize the sample based on its importance that renders the hard samples not fully exploited; or explicitly emphasize the effects of semi-hard/hard samples even at the early training stage that may lead to convergence issue. In this work, we propose a novel Adaptive Curriculum Learning loss (CurricularFace) that embeds the idea of curriculum learning into the loss function to achieve a novel training strategy for deep face recognition, which mainly addresses easy samples in the early training stage and hard ones in the later stage. Specifically, our CurricularFace adaptively adjusts the relative importance of easy and hard samples during different training stages. In each stage, different samples are assigned with different importance according to their corresponding difficultness. Extensive experimental results on popular benchmarks demonstrate the superiority of our CurricularFace over the state-of-the-art competitors. Yuge Huang, Yuhan Wang 0002, Ying Tai, Xiaoming Liu 0002, Pengcheng Shen, Shaoxin Li 0001, Feiyue Huang |
CVPR | 4 |
| 2020 | Noise Modeling, Synthesis and Classification for Generic Object Anti-SpoofingabstractUsing printed photograph and replaying videos of biometric modalities, such as iris, fingerprint and face, are common attacks to fool the recognition systems for granting access as the genuine user. With the growing online person-to-person shopping (e.g., Ebay and Craigslist), such attacks also threaten those services, where the online photo illustration might not be captured from real items but from paper or digital screen. Thus, the study of anti-spoofing should be extended from modality-specific solutions to generic-object-based ones. In this work, we define and tackle the problem of Generic Object Anti-Spoofing (GOAS) for the first time. One significant cue to detect these attacks is the noise patterns introduced by the capture sensors and spoof mediums. Different sensor/medium combinations can result in diverse noise patterns. We propose a GAN-based architecture to synthesize and identify the noise patterns from seen and unseen medium/sensor combinations. We show that the procedure of synthesis and identification are mutually beneficial. We further demonstrate the learned GOAS models can directly contribute to modality-specific anti-spoofing without domain transfer. The code and GOSet dataset are available at cvlab.cse.msu.edu/project-goas.html. Joel Stehouwer, Amin Jourabloo, Yaojie Liu, Xiaoming Liu 0002 |
CVPR | 4 |
| 2020 | The Edge of Depth: Explicit Constraints Between Segmentation and DepthabstractIn this work we study the mutual benefits of two common computer vision tasks, self-supervised depth estimation and semantic segmentation from images. For example, to help unsupervised monocular depth estimation, constraint from semantic segmentation has been explored implicitly such as sharing and transforming features. In contrast, we propose to explicitly measure the border consistency between segmentation and depth and minimize it in a greedy manner by iteratively supervising the network towards a locally optimal solution. Partially this is motivated by our observation that semantic segmentation even trained with limited ground truth (200 images of KITTI) can offer more accurate border than that of any (monocular or stereo) image-based depth estimation. Through extensive experiments, our proposed approach advance the state of the art on unsupervised monocular depth estimation in the KITTI benchmark. Garrick Brazil, Xiaoming Liu 0002 |
CVPR | 3 |
| 2020 | Kinematic 3D Object Detection in Monocular Video
Garrick Brazil, Gerard Pons-Moll, Xiaoming Liu 0002, Bernt Schiele |
ECCV (23) | 3 |
| 2020 | Jointly De-Biasing Face Recognition and Demographic Attribute Estimation
Sixue Gong, Xiaoming Liu 0002, Anil K. Jain 0001 |
ECCV (29) | 2 |
| 2020 | Improving Face Recognition from Hard Samples via Distribution Distillation Loss
Yuge Huang, Pengcheng Shen, Ying Tai, Shaoxin Li 0001, Xiaoming Liu 0002, Feiyue Huang, Rongrong Ji |
ECCV (30) | 5 |
| 2020 | On Disentangling Spoof Trace for Generic Face Anti-spoofing
Yaojie Liu, Joel Stehouwer, Xiaoming Liu 0002 |
ECCV (18) | 3 |
| 2020 | Learning Implicit Functions for Topology-Varying Dense 3D Shape CorrespondenceabstractThe goal of this paper is to learn dense 3D shape correspondence for topology-varying objects in an unsupervised manner. Conventional implicit functions estimate the occupancy of a 3D point given a shape latent code. Instead, our novel implicit function produces a part embedding vector for each 3D point, which is assumed to be similar to its densely corresponded point in another 3D shape of the same object category. Furthermore, we implement dense correspondence through an inverse function mapping from the part embedding to a corresponded 3D point. Both functions are jointly learned with several effective loss functions to realize our assumption, together with the encoder generating the shape latent code. During inference, if a user selects an arbitrary point on the source shape, our algorithm can automatically generate a confidence score indicating whether there is a correspondence on the target shape, as well as the corresponding semantic point if there is. Such a mechanism inherently benefits man-made objects with different part constitutions. The effectiveness of our approach is demonstrated through unsupervised 3D semantic correspondence and shape segmentation. Feng Liu 0037, Xiaoming Liu 0002 |
NeurIPS | 2 |
| 2020 | Joint Face Alignment and 3D Face Reconstruction with Application to Face RecognitionabstractFace alignment and 3D face reconstruction are traditionally accomplished as separated tasks. By exploring the strong correlation between 2D landmarks and 3D shapes, in contrast, we propose a joint face alignment and 3D face reconstruction method to simultaneously solve these two problems for 2D face images of arbitrary poses and expressions. This method, based on a summation model of 3D faces and cascaded regression in 2D and 3D shape spaces, iteratively and alternately applies two cascaded regressors, one for updating 2D landmarks and the other for 3D shape. The 3D shape and the landmarks are correlated via a 3D-to-2D mapping matrix, which is updated in each iteration to refine the location and visibility of 2D landmarks. Unlike existing methods, the proposed method can fully automatically generate both pose-and-expression-normalized (PEN) and expressive 3D faces and localize both visible and invisible 2D landmarks. Based on the PEN 3D faces, we devise a method to enhance face recognition accuracy across poses and expressions. Both linear and nonlinear implementations of the proposed method are presented and evaluated in this paper. Extensive experiments show that the proposed method can achieve the state-of-the-art accuracy in both face alignment and 3D face reconstruction, and benefit face recognition owing to its reconstructed PEN 3D face. Feng Liu 0037, Qijun Zhao, Xiaoming Liu 0002, Dan Zeng 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | Towards Highly Accurate and Stable Face Alignment for High-Resolution VideosabstractIn recent years, heatmap regression based models have shown their effectiveness in face alignment and pose estimation. However, Conventional Heatmap Regression (CHR) is not accurate nor stable when dealing with high-resolution facial videos, since it finds the maximum activated location in heatmaps which are generated from rounding coordinates, and thus leads to quantization errors when scaling back to the original high-resolution space. In this paper, we propose a Fractional Heatmap Regression (FHR) for high-resolution video-based face alignment. The proposed FHR can accurately estimate the fractional part according to the 2D Gaussian function by sampling three points in heatmaps. To further stabilize the landmarks among continuous video frames while maintaining the precise at the same time, we propose a novel stabilization loss that contains two terms to address time delay and non-smooth issues, respectively. Experiments on 300W, 300VW and Talking Face datasets clearly demonstrate that the proposed method is more accurate and stable than the state-ofthe-art models. Ying Tai, Yicong Liang, Xiaoming Liu 0002, Lei Duan, Chengjie Wang 0001, Feiyue Huang, Yu Chen 0037 |
AAAI | 3 |
| 2019 | Feature Transfer Learning for Face Recognition With Under-Represented DataabstractDespite the large volume of face recognition datasets, there is a significant portion of subjects, of which the samples are insufficient and thus under-represented. Ignoring such significant portion results in insufficient training data. Training with under-represented data leads to biased classifiers in conventionally-trained deep networks. In this paper, we propose a center-based feature transfer framework to augment the feature space of under-represented subjects from the regular subjects that have sufficiently diverse samples. A Gaussian prior of the variance is assumed across all subjects and the variance from regular ones are transferred to the under-represented ones. This encourages the under-represented distribution to be closer to the regular distribution. Further, an alternating training regimen is proposed to simultaneously achieve less biased classifiers and a more discriminative feature representation. We conduct ablative study to mimic the under-represented datasets by varying the portion of under-represented classes on the MS-Celeb-1M dataset. Advantageous results on LFW, IJB-A and MS-Celeb-1M demonstrate the effectiveness of our feature transfer and training strategy, compared to both general baselines and state-of-the-art methods. Moreover, our feature transfer successfully presents smooth visual interpolation, which conducts disentanglement to preserve identity of a class while augmenting its feature space with non-identity variations such as pose and lighting. Xi Yin 0001, Xiang Yu 0002, Kihyuk Sohn, Xiaoming Liu 0002, Manmohan Krishna Chandraker |
CVPR | 4 |
| 2019 | Pedestrian Detection With Autoregressive Network PhasesabstractWe present an autoregressive pedestrian detection framework with cascaded phases designed to progressively improve precision. The proposed framework utilizes a novel lightweight stackable decoder-encoder module which uses convolutional re-sampling layers to improve features while maintaining efficient memory and runtime cost. Unlike previous cascaded detection systems, our proposed framework is designed within a region proposal network and thus retains greater context of nearby detections compared to independently processed RoI systems. We explicitly encourage increasing levels of precision by assigning strict labeling policies to each consecutive phase such that early phases develop features primarily focused on achieving high recall and later on accurate precision. In consequence, the final feature maps form more peaky radial gradients emulating from the centroids of unique pedestrians. Using our proposed autoregressive framework leads to new state-of-the-art performance on the reasonable and occlusion settings of the Caltech pedestrian dataset, and achieves competitive state-of-the-art performance on the KITTI dataset. Garrick Brazil, Xiaoming Liu 0002 |
CVPR | 2 |
| 2019 | Depth Coefficients for Depth CompletionabstractDepth completion involves estimating a dense depth image from sparse depth measurements, often guided by a color image. While linear upsampling is straight forward, it results in depth pixels being interpolated in empty space across discontinuities between objects. Current methods use deep networks to maintain gaps between objects. Nevertheless depth smearing remains a challenge. We propose a new representation for depth called Depth Coefficients (DC) to address this problem. It enables convolutions to more easily avoid inter-object depth mixing. We also show that the standard Mean Squared Error (MSE) loss function can promote depth mixing, and so we propose instead to use cross-entropy loss for DC. Both quantitative and qualitative evaluation are conducted on benchmarks, and we show that switching out sparse depth input and MSE loss functions with our DC representation and loss is a simple way to improve performance, reduce pixel depth mixing and can improve object detection. Saif Muhammad Imran, Xiaoming Liu 0002, Daniel D. Morris |
CVPR | 3 |
| 2019 | Deep Tree Learning for Zero-Shot Face Anti-SpoofingabstractFace anti-spoofing is designed to keep face recognition systems from recognizing fake faces as the genuine users. While advanced face anti-spoofing methods are developed, new types of spoof attacks are also being created and becoming a threat to all existing systems. We define the detection of unknown spoof attacks as Zero-Shot Face Anti-spoofing (ZSFA). Previous works of ZSFA only study 1-2 types of spoof attacks, such as print/replay attacks, which limits the insight of this problem. In this work, we expand the ZSFA problem to a wide range of 13 types of spoof attacks, including print attack, replay attack, 3D mask attacks, and so on. A novel Deep Tree Network (DTN) is proposed to tackle the ZSFA. The tree is learned to partition the spoof samples into semantic sub-groups in an unsupervised fashion. When a data sample arrives, being know or unknown attacks, DTN routes it to the most similar spoof cluster, and make the binary decision. In addition, to enable the study of ZSFA, we introduce the first face anti-spoofing database that contains diverse types of spoof attacks. Experiments show that our proposed method achieves the state of the art on multiple testing protocols of ZSFA. Yaojie Liu, Joel Stehouwer, Amin Jourabloo, Xiaoming Liu 0002 |
CVPR | 4 |
| 2019 | Towards High-Fidelity Nonlinear 3D Face Morphable ModelabstractEmbedding 3D morphable basis functions into deep neural networks opens great potential for models with better representation power. However, to faithfully learn those models from an image collection, it requires strong regularization to overcome ambiguities involved in the learning process. This critically prevents us from learning high fidelity face models which are needed to represent face images in high level of details. To address this problem, this paper presents a novel approach to learn additional proxies as means to side-step strong regularizations, as well as, leverages to promote detailed shape/albedo. To ease the learning, we also propose to use a dual-pathway network, a carefully-designed architecture that brings a balance between global and local-based models. By improving the nonlinear 3D morphable model in both learning objective and network architecture, we present a model which is superior in capturing higher level of details than the linear or its precedent nonlinear counterparts. As a result, our model achieves state-of-the-art performance on 3D face reconstruction by solely optimizing latent representations. Luan Tran, Feng Liu 0037, Xiaoming Liu 0002 |
CVPR | 3 |
| 2019 | Gotta Adapt 'Em All: Joint Pixel and Feature-Level Domain Adaptation for Recognition in the WildabstractRecent developments in deep domain adaptation have allowed knowledge transfer from a labeled source domain to an unlabeled target domain at the level of intermediate features or input pixels. We propose that advantages may be derived by combining them, in the form of different insights that lead to a novel design and complementary properties that result in better performance. At the feature level, inspired by insights from semi-supervised learning, we propose a classification-aware domain adversarial neural network that brings target examples into more classifiable regions of source domain. Next, we posit that computer vision insights are more amenable to injection at the pixel level. In particular, we use 3D geometry and image synthesis based on a generalized appearance flow to preserve identity across pose transformations, while using an attribute-conditioned CycleGAN to translate a single source into multiple target images that differ in lower-level properties such as lighting. Besides standard UDA benchmark, we validate on a novel and apt problem of car recognition in unlabeled surveillance images using labeled images from the web, handling explicitly specified, nameable factors of variation through pixel-level and implicit, unspecified factors through feature-level adaptation. Luan Tran, Kihyuk Sohn, Xiang Yu 0002, Xiaoming Liu 0002, Manmohan Krishna Chandraker |
CVPR | 4 |
| 2019 | Gait Recognition via Disentangled Representation LearningabstractGait, the walking pattern of individuals, is one of the most important biometrics modalities. Most of the existing gait recognition methods take silhouettes or articulated body models as the gait features. These methods suffer from degraded recognition performance when handling confounding variables, such as clothing, carrying and view angle. To remedy this issue, we propose a novel AutoEncoder framework to explicitly disentangle pose and appearance features from RGB imagery and the LSTM-based integration of pose features over time produces the gait feature. In addition, we collect a Frontal-View Gait (FVG) dataset to focus on gait recognition from frontal-view walking, which is a challenging problem since it contains minimal gait cues compared to other views. FVG also includes other important variations,e.g., walking speed, carrying, and clothing. With extensive experiments on CASIA-B, USF and FVG datasets, our method demonstrates superior performance to the-state-of-the-arts quantitatively, the ability of feature disentanglement qualitatively, and promising computational efficiency. Ziyuan Zhang 0004, Luan Tran, Xi Yin 0001, Yousef Atoum, Xiaoming Liu 0002, Nanxin Wang |
CVPR | 5 |
| 2019 | M3D-RPN: Monocular 3D Region Proposal Network for Object DetectionabstractUnderstanding the world in 3D is a critical component of urban autonomous driving. Generally, the combination of expensive LiDAR sensors and stereo RGB imaging has been paramount for successful 3D object detection algorithms, whereas monocular image-only methods experience drastically reduced performance. We propose to reduce the gap by reformulating the monocular 3D detection problem as a standalone 3D region proposal network. We leverage the geometric relationship of 2D and 3D perspectives, allowing 3D boxes to utilize well-known and powerful convolutional features generated in the image-space. To help address the strenuous 3D parameter estimations, we further design depth-aware convolutional layers which enable location specific feature development and in consequence improved 3D scene understanding. Compared to prior work in monocular 3D detection, our method consists of only the proposed 3D region proposal network rather than relying on external networks, data, or multiple stages. M3D-RPN is able to significantly improve the performance of both monocular 3D Object Detection and Bird's Eye View tasks within the KITTI urban autonomous driving dataset, while efficiently using a shared multi-class model. Garrick Brazil, Xiaoming Liu 0002 |
ICCV | 2 |
| 2019 | 3D Face Modeling From Diverse Raw Scan DataabstractTraditional 3D face models learn a latent representation of faces using linear subspaces from limited scans of a single database. The main roadblock of building a large-scale face model from diverse 3D databases lies in the lack of dense correspondence among raw scans. To address these problems, this paper proposes an innovative framework to jointly learn a nonlinear face model from a diverse set of raw 3D scan databases and establish dense point-to-point correspondence among their scans. Specifically, by treating input scans as unorganized point clouds, we explore the use of PointNet architectures for converting point clouds to identity and expression feature representations, from which the decoder networks recover their 3D face shapes. Further, we propose a weakly supervised learning approach that does not require correspondence label for the scans. We demonstrate the superior dense correspondence and representation power of our proposed method, and its contribution to single-image 3D face reconstruction. Feng Liu 0037, Luan Tran, Xiaoming Liu 0002 |
ICCV | 3 |
| 2019 | Towards Interpretable Face RecognitionabstractDeep CNNs have been pushing the frontier of visual recognition over past years. Besides recognition accuracy, strong demands in understanding deep CNNs in the research community motivate developments of tools to dissect pre-trained models to visualize how they make predictions. Recent works further push the interpretability in the network learning stage to learn more meaningful representations. In this work, focusing on a specific area of visual recognition, we report our efforts towards interpretable face recognition. We propose a spatial activation diversity loss to learn more structured face representations. By leveraging the structure, we further design a feature activation diversity loss to push the interpretable representations to be discriminative and robust to occlusions. We demonstrate on three face recognition benchmarks that our proposed method is able to achieve the state-of-art face recognition accuracy with easily interpretable face representations. Bangjie Yin, Luan Tran, Xiaohui Shen, Xiaoming Liu 0002 |
ICCV | 5 |
| 2019 | Recurrent Flow-Guided Semantic ForecastingabstractUnderstanding the world around us and making decisions about the future is a critical component to human intelligence. As autonomous systems continue to develop, their ability to reason about the future will be the key to their success. Semantic anticipation is a relatively under-explored area for which autonomous vehicles could take advantage of (e.g., forecasting pedestrian trajectories). Motivated by the need for real-time prediction in autonomous systems, we propose to decompose the challenging semantic forecasting task into two subtasks: current frame segmentation and future optical flow prediction. Through this decomposition, we built an efficient, effective, low overhead model with three main components: flow prediction network, feature-flow aggregation LSTM, and end-to-end learnable warp layer. Our proposed method achieves state-of-the-art accuracy on short-term and moving objects semantic forecasting while simultaneously reducing model parameters by up to 95% and increasing efficiency by greater than 40x. Adam M. Terwilliger, Garrick Brazil, Xiaoming Liu 0002 |
WACV | 3 |
| 2019 | Editorial: Special Issue on Deep Learning for Face Analysis
Chen Change Loy, Xiaoming Liu 0002, Tae-Kyun Kim 0001, Fernando De la Torre, Rama Chellappa |
Int. J. Comput. Vis. | 2 |
| 2019 | Representation Learning by Rotating Your FacesabstractThe large pose discrepancy between two face images is one of the fundamental challenges in automatic face recognition. Conventional approaches to pose-invariant face recognition either perform face frontalization on, or learn a pose-invariant representation from, a non-frontal face image. We argue that it is more desirable to perform both tasks jointly to allow them to leverage each other. To this end, this paper proposes a Disentangled Representation learning-Generative Adversarial Network (DR-GAN) with three distinct novelties. First, the encoder-decoder structure of the generator enables DR-GAN to learn a representation that is both generative and discriminative, which can be used for face image synthesis and pose-invariant face recognition. Second, this representation is explicitly disentangled from other face variations such as pose, through the pose code provided to the decoder and pose estimation in the discriminator. Third, DR-GAN can take one or multiple images as the input, and generate one unified identity representation along with an arbitrary number of synthetic face images. Extensive quantitative and qualitative evaluation on a number of controlled and in-the-wild databases demonstrate the superiority of DR-GAN over the state of the art in both learning representations and rotating large-pose face images. Luan Tran, Xi Yin 0001, Xiaoming Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | Face Alignment in Full Pose Range: A 3D Total SolutionabstractFace alignment, which fits a face model to an image and extracts the semantic meanings of facial pixels, has been an important topic in the computer vision community. However, most algorithms are designed for faces in small to medium poses (yaw angle is smaller than 45 degree), which lack the ability to align faces in large poses up to 90 degree. The challenges are three-fold. First, the commonly used landmark face model assumes that all the landmarks are visible and is therefore not suitable for large poses. Second, the face appearance varies more drastically across large poses, from the frontal view to the profile view. Third, labelling landmarks in large poses is extremely challenging since the invisible landmarks have to be guessed. In this paper, we propose to tackle these three challenges in an new alignment framework termed 3D Dense Face Alignment (3DDFA), in which a dense 3D Morphable Model (3DMM) is fitted to the image via Cascaded Convolutional Neural Networks. We also utilize 3D information to synthesize face images in profile views to provide abundant samples for training. Experiments on the challenging AFLW database show that the proposed approach achieves significant improvements over the state-of-the-art methods. Xiangyu Zhu 0001, Xiaoming Liu 0002, Zhen Lei 0001, Stan Z. Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Disentangling Features in 3D Face Shapes for Joint Face Reconstruction and RecognitionabstractThis paper proposes an encoder-decoder network to disentangle shape features during 3D face reconstruction from single 2D images, such that the tasks of reconstructing accurate 3D face shapes and learning discriminative shape features for face recognition can be accomplished simultaneously. Unlike existing 3D face reconstruction methods, our proposed method directly regresses dense 3D face shapes from single 2D images, and tackles identity and residual (i.e., non-identity) components in 3D face shapes explicitly and separately based on a composite 3D face shape model with latent representations. We devise a training process for the proposed network with a joint loss measuring both face identification error and 3D face shape reconstruction error. To construct training data we develop a method for fitting 3D morphable model (3DMM) to multiple 2D images of a subject. Comprehensive experiments have been done on MICC, BU3DFE, LFW and YTF databases. The results show that our method expands the capacity of 3DMM for capturing discriminative shape features and facial detail, and thus outperforms existing methods both in 3D face reconstruction accuracy and in face recognition accuracy. Feng Liu 0013, Ronghang Zhu, Dan Zeng 0002, Qijun Zhao, Xiaoming Liu 0002 |
CVPR | 5 |
| 2018 | FSRNet: End-to-End Learning Face Super-Resolution With Facial PriorsabstractFace Super-Resolution (SR) is a domain-specific superresolution problem. The facial prior knowledge can be leveraged to better super-resolve face images. We present a novel deep end-to-end trainable Face Super-Resolution Network (FSRNet), which makes use of the geometry prior, i.e., facial landmark heatmaps and parsing maps, to super-resolve very low-resolution (LR) face images without well-aligned requirement. Specifically, we first construct a coarse SR network to recover a coarse high-resolution (HR) image. Then, the coarse HR image is sent to two branches: a fine SR encoder and a prior information estimation network, which extracts the image features, and estimates landmark heatmaps/parsing maps respectively. Both image features and prior information are sent to a fine SR decoder to recover the HR image. To generate realistic faces, we also propose the Face Super-Resolution Generative Adversarial Network (FSRGAN) to incorporate the adversarial loss into FSRNet. Further, we introduce two related tasks, face alignment and parsing, as the new evaluation metrics for face SR, which address the inconsistency of classic metrics w.r.t. visual perception. Extensive experiments show that FSRNet and FSRGAN significantly outperforms state of the arts for very LR face SR, both quantitatively and qualitatively. Yu Chen 0037, Ying Tai, Xiaoming Liu 0002, Chunhua Shen, Jian Yang 0003 |
CVPR | 3 |
| 2018 | Learning Deep Models for Face Anti-Spoofing: Binary or Auxiliary SupervisionabstractFace anti-spoofing is crucial to prevent face recognition systems from a security breach. Previous deep learning approaches formulate face anti-spoofing as a binary classification problem. Many of them struggle to grasp adequate spoofing cues and generalize poorly. In this paper, we argue the importance of auxiliary supervision to guide the learning toward discriminative and generalizable cues. A CNN-RNN model is learned to estimate the face depth with pixel-wise supervision, and to estimate rPPG signals with sequence-wise supervision. The estimated depth and rPPG are fused to distinguish live vs. spoof faces. Further, we introduce a new face anti-spoofing database that covers a large range of illumination, subject, and pose variations. Experiments show that our model achieves the state-of-the-art results on both intra- and cross-database testing. Yaojie Liu, Amin Jourabloo, Xiaoming Liu 0002 |
CVPR | 3 |
| 2018 | Nonlinear 3D Face Morphable ModelabstractAs a classic statistical model of 3D facial shape and texture, 3D Morphable Model (3DMM) is widely used in facial analysis, e.g., model fitting, image synthesis. Conventional 3DMM is learned from a set of well-controlled 2D face images with associated 3D face scans, and represented by two sets of PCA basis functions. Due to the type and amount of training data, as well as the linear bases, the representation power of 3DMM can be limited. To address these problems, this paper proposes an innovative framework to learn a nonlinear 3DMM model from a large set of unconstrained face images, without collecting 3D face scans. Specifically, given a face image as input, a network encoder estimates the projection, shape and texture parameters. Two decoders serve as the nonlinear 3DMM to map from the shape and texture parameters to the 3D shape and texture, respectively. With the projection parameter, 3D shape, and texture, a novel analytically-differentiable rendering layer is designed to reconstruct the original input face. The entire network is end-to-end trainable with only weak supervision. We demonstrate the superior representation power of our nonlinear 3DMM over its linear counterpart, and its contribution to face alignment and 3D reconstruction. Luan Tran, Xiaoming Liu 0002 |
CVPR | 2 |
| 2018 | Face De-spoofing: Anti-spoofing via Noise Modeling
Amin Jourabloo, Yaojie Liu, Xiaoming Liu 0002 |
ECCV (13) | 3 |
| 2018 | MSU-AVIS dataset: Fusing Face and Voice Modalities for Biometric Recognition in Indoor Surveillance VideosabstractIndoor video surveillance systems often use the face modality to establish the identity of a person of interest. However, the face image may not offer sufficient discriminatory information in many scenarios due to substantial variations in pose, illumination, expression, resolution and distance between the subject and the camera. In such cases, the inclusion of an additional biometric modality can benefit the recognition process. In this regard, we consider the fusion of voice and face modalities for enhancing the recognition accuracy. The main contribution of this work is assembling a multimodal (face and voice), semi-constrained, indoor video surveillance dataset referred to as the MSU Audio-Video Indoor Surveillance (MSU-AVIS) dataset. We use a consumer-grade camera with a built-in microphone to acquire data for this purpose. We use current state-of-art deep-learning based methods to perform face and speaker recognition on the collected dataset for establishing baseline performance. We also explore multiple fusion schemes to combine face and speaker recognition to perform effective person recognition on audio-video surveillance data. Experiments convey the efficacy of the proposed multimodal fusion scheme (face and voice) over unimodal approaches in surveillance scenarios. The collected dataset is being made available for research purposes. Anurag Chowdhury, Yousef Atoum, Luan Tran, Xiaoming Liu 0002, Arun Ross |
ICPR | 4 |
| 2018 | Joint Multi-Leaf Segmentation, Alignment, and Tracking for Fluorescence Plant VideosabstractThis paper proposes a novel framework for fluorescence plant video processing. The plant research community is interested in the leaf-level photosynthetic analysis within a plant. A prerequisite for such analysis is to segment all leaves, estimate their structures, and track them over time. We identify this as a joint multi-leaf segmentation, alignment, and tracking problem. First, leaf segmentation and alignment are applied on the last frame of a plant video to find a number of well-aligned leaf candidates. Second, leaf tracking is applied on the remaining frames with leaf candidate transformation from the previous frame. We form two optimization problems with shared terms in their objective functions for leaf alignment and tracking respectively. A quantitative evaluation framework is formulated to evaluate the performance of our algorithm with four metrics. Two models are learned to predict the alignment accuracy and detect tracking failure respectively in order to provide guidance for subsequent plant biology analysis. The limitation of our algorithm is also studied. Experimental results show the effectiveness, efficiency, and robustness of the proposed method. Xi Yin 0001, Xiaoming Liu 0002, Jin Chen 0004, David M. Kramer 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Multi-Task Convolutional Neural Network for Pose-Invariant Face RecognitionabstractThis paper explores multi-task learning (MTL) for face recognition. First, we propose a multi-task convolutional neural network (CNN) for face recognition, where identity classification is the main task and pose, illumination, and expression (PIE) estimations are the side tasks. Second, we develop a dynamic-weighting scheme to automatically assign the loss weights to each side task, which solves the crucial problem of balancing between different tasks in MTL. Third, we propose a pose-directed multi-task CNN by grouping different poses to learn pose-specific identity features, simultaneously across all poses in a joint framework. Last but not least, we propose an energy-based weight analysis method to explore how CNN-based MTL works. We observe that the side tasks serve as regularizations to disentangle the PIE variations from the learnt identity features. Extensive experiments on the entire multi-PIE dataset demonstrate the effectiveness of the proposed approach. To the best of our knowledge, this is the first work using all data in multi-PIE for face recognition. Our approach is also applicable to in-the-wild data sets for pose-invariant face recognition and achieves comparable or better performance than state of the art on LFW, CFP, and IJB-A datasets. Xi Yin 0001, Xiaoming Liu 0002 |
IEEE Trans. Image Process. | 2 |
| 2018 | Fusing Geometric Features for Skeleton-Based Action Recognition Using Multilayer LSTM NetworksabstractRecent skeleton-based action recognition approaches achieve great improvement by using recurrent neural network (RNN) models. Currently, these approaches build an end-to-end network from coordinates of joints to class categories and improve accuracy by extending RNN to spatial domains. First, while such well-designed models and optimization strategies explore relations between different parts directly from joint coordinates, we provide a simple universal spatial modeling method perpendicular to the RNN model enhancement. Specifically, according to the evolution of previous work, we select a set of simple geometric features, and then separately feed each type of features to a three-layer LSTM framework. Second, we propose a multistream LSTM architecture with a new smoothed score fusion technique to learn classification from different geometric feature streams. Furthermore, we observe that the geometric relational features based on distances between joints and selected lines outperform other features and the fusion results achieve the state-of-the-art performance on four datasets. We also show the sparsity of input gate weights in the first LSTM layer trained by geometric features and demonstrate that utilizing joint-line distances as input require less data for training. Songyang Zhang 0004, Yang Yang 0009, Jun Xiao 0001, Xiaoming Liu 0002, Yi Yang 0001, Di Xie, Yueting Zhuang |
IEEE Trans. Multim. | 4 |
| 2018 | Do Convolutional Neural Networks Learn Class Hierarchy?abstractConvolutional Neural Networks (CNNs) currently achieve state-of-the-art accuracy in image classification. With a growing number of classes, the accuracy usually drops as the possibilities of confusion increase. Interestingly, the class confusion patterns follow a hierarchical structure over the classes. We present visual-analytics methods to reveal and analyze this hierarchy of similar classes in relation with CNN-internal data. We found that this hierarchy not only dictates the confusion patterns between the classes, it furthermore dictates the learning behavior of CNNs. In particular, the early layers in these networks develop feature detectors that can separate high-level groups of classes quite well, even after a few training epochs. In contrast, the latter layers require substantially more epochs to develop specialized feature detectors that can separate individual classes. We demonstrate how these insights are key to significant improvement in accuracy by designing hierarchy-aware CNNs that accelerate model convergence and alleviate overfitting. We further demonstrate how our methods help in identifying various quality issues in the training data. Bilal Alsallakh, Amin Jourabloo, Mao Ye 0011, Xiaoming Liu 0002, Liu Ren 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2017 | Spatio-Temporal Alignment of Non-overlapping Sequences from Independently Panning CamerasabstractThis paper addresses the problem of spatio-temporal alignment of multiple video sequences. We identify and tackle a novel scenario of this problem referred to as Nonoverlapping Sequences (NOS). NOS are captured by multiple freely panning handheld cameras whose field of views (FOV) might have no direct spatial overlap. With the popularity of mobile sensors, NOS rise when multiple cooperative users capture a public event to create a panoramic video, or when consolidating multiple footages of an incident into a single video. To tackle this novel scenario, we first spatially align the sequences by reconstructing the background of each sequence and registering these backgrounds, even if the backgrounds are not overlapping. Given the spatial alignment, we temporally synchronize the sequences, such that the trajectories of moving objects (e.g., cars or pedestrians) are consistent across sequences. Experimental results demonstrate the performance of our algorithm in this novel and challenging scenario, quantitatively and qualitatively. Seyed Morteza Safdarnejad, Xiaoming Liu 0002 |
CVPR | 2 |
| 2017 | Image Super-Resolution via Deep Recursive Residual NetworkabstractRecently, Convolutional Neural Network (CNN) based models have achieved great success in Single Image Super-Resolution (SISR). Owing to the strength of deep networks, these CNN models learn an effective nonlinear mapping from the low-resolution input image to the high-resolution target image, at the cost of requiring enormous parameters. This paper proposes a very deep CNN model (up to 52 convolutional layers) named Deep Recursive Residual Network (DRRN) that strives for deep yet concise networks. Specifically, residual learning is adopted, both in global and local manners, to mitigate the difficulty of training very deep networks, recursive learning is used to control the model parameters while increasing the depth. Extensive benchmark evaluation shows that DRRN significantly outperforms state of the art in SISR, while utilizing far fewer parameters. Code is available at https://github.com/tyshiwo/DRRN_CVPR17. Ying Tai, Jian Yang 0003, Xiaoming Liu 0002 |
CVPR | 3 |
| 2017 | Disentangled Representation Learning GAN for Pose-Invariant Face RecognitionabstractThe large pose discrepancy between two face images is one of the key challenges in face recognition. Conventional approaches for pose-invariant face recognition either perform face frontalization on, or learn a pose-invariant representation from, a non-frontal face image. We argue that it is more desirable to perform both tasks jointly to allow them to leverage each other. To this end, this paper proposes Disentangled Representation learning-Generative Adversarial Network (DR-GAN) with three distinct novelties. First, the encoder-decoder structure of the generator allows DR-GAN to learn a generative and discriminative representation, in addition to image synthesis. Second, this representation is explicitly disentangled from other face variations such as pose, through the pose code provided to the decoder and pose estimation in the discriminator. Third, DR-GAN can take one or multiple images as the input, and generate one unified representation along with an arbitrary number of synthetic images. Quantitative and qualitative evaluation on both controlled and in-the-wild databases demonstrate the superiority of DR-GAN over the state of the art. Luan Tran, Xi Yin 0001, Xiaoming Liu 0002 |
CVPR | 3 |
| 2017 | Missing Modalities Imputation via Cascaded Residual AutoencoderabstractAffordable sensors lead to an increasing interest in acquiring and modeling data with multiple modalities. Learning from multiple modalities has shown to significantly improve performance in object recognition. However, in practice it is common that the sensing equipment experiences unforeseeable malfunction or configuration issues, leading to corrupted data with missing modalities. Most existing multi-modal learning algorithms could not handle missing modalities, and would discard either all modalities with missing values or all corrupted data. To leverage the valuable information in the corrupted data, we propose to impute the missing data by leveraging the relatedness among different modalities. Specifically, we propose a novel Cascaded Residual Autoencoder (CRA) to impute missing modalities. By stacking residual autoencoders, CRA grows iteratively to model the residual between the current prediction and original data. Extensive experiments demonstrate the superior performance of CRA on both the data imputation and the object recognition task on imputed data. Luan Tran, Xiaoming Liu 0002, Rong Jin 0001 |
CVPR | 2 |
| 2017 | Face anti-spoofing using patch and depth-based CNNsabstractThe face image is the most accessible biometric modality which is used for highly accurate face recognition systems, while it is vulnerable to many different types of presentation attacks. Face anti-spoofing is a very critical step before feeding the face image to biometric systems. In this paper, we propose a novel two-stream CNN-based approach for face anti-spoofing, by extracting the local features and holistic depth maps from the face images. The local features facilitate CNN to discriminate the spoof patches independent of the spatial face areas. On the other hand, holistic depth map examine whether the input image has a face-like depth. Extensive experiments are conducted on the challenging databases (CASIA-FASD, MSU-USSA, and Replay Attack), with comparison to the state of the art. Yousef Atoum, Yaojie Liu, Amin Jourabloo, Xiaoming Liu 0002 |
IJCB | 4 |
| 2017 | Monocular Video-Based Trailer Coupler Detection Using Multiplexer Convolutional Neural NetworkabstractThis paper presents an automated monocular-camera-based computer vision system for autonomous self-backing-up a vehicle towards a trailer, by continuously estimating the 3D trailer coupler position and feeding it to the vehicle control system, until the alignment of the tow hitch with the trailers coupler. This system is made possible through our proposed distance-driven Multiplexer-CNN method, which selects the most suitable CNN using the estimated coupler-to-vehicle distance. The input of the multiplexer is a group made of a CNN detector, trackers, and 3D localizer. In the CNN detector, we propose a novel algorithm to provide a presence confidence score with each detection. The score reflects the existence of the target object in a region, as well as how accurate is the 2D target detection. We demonstrate the accuracy and efficiency of the system on a large trailer database. Our system achieves an estimation error of 1.4 cm when the ball reaches the coupler, while running at 18.9 FPS on a regular PC. Yousef Atoum, Joseph Roth, Michael Bliss, Wende Zhang, Xiaoming Liu 0002 |
ICCV | 5 |
| 2017 | Illuminating Pedestrians via Simultaneous Detection and SegmentationabstractPedestrian detection is a critical problem in computer vision with significant impact on safety in urban autonomous driving. In this work, we explore how semantic segmentation can be used to boost pedestrian detection accuracy while having little to no impact on network efficiency. We propose a segmentation infusion network to enable joint supervision on semantic segmentation and pedestrian detection. When placed properly, the additional supervision helps guide features in shared layers to become more sophisticated and helpful for the downstream pedestrian detector. Using this approach, we find weakly annotated boxes to be sufficient for considerable performance gains. We provide an in-depth analysis to demonstrate how shared layers are shaped by the segmentation supervision. In doing so, we show that the resulting feature maps become more semantically meaningful and robust to shape and occlusion. Overall, our simultaneous detection and segmentation framework achieves a considerable gain over the state-of-the-art on the Caltech pedestrian dataset, competitive performance on KITTI, and executes 2 × faster than competitive methods. Garrick Brazil, Xi Yin 0001, Xiaoming Liu 0002 |
ICCV | 3 |
| 2017 | Pose-Invariant Face Alignment with a Single CNNabstractFace alignment has witnessed substantial progress in the last decade. One of the recent focuses has been aligning a dense 3D face shape to face images with large head poses. The dominant technology used is based on the cascade of regressors, e.g., CNNs, which has shown promising results. Nonetheless, the cascade of CNNs suffers from several drawbacks, e.g., lack of end-to-end training, handcrafted features and slow training speed. To address these issues, we propose a new layer, named visualization layer, which can be integrated into the CNN architecture and enables joint optimization with different loss functions. Extensive evaluation of the proposed method on multiple datasets demonstrates state-of-the-art accuracy, while reducing the training time by more than half compared to the typical cascade of CNNs. In addition, we compare across multiple CNN architectures, all with the visualization layer, to further demonstrate the advantage of its utilization. Amin Jourabloo, Mao Ye 0011, Xiaoming Liu 0002, Liu Ren 0001 |
ICCV | 3 |
| 2017 | MemNet: A Persistent Memory Network for Image RestorationabstractRecently, very deep convolutional neural networks (CNNs) have been attracting considerable attention in image restoration. However, as the depth grows, the longterm dependency problem is rarely realized for these very deep models, which results in the prior states/layers having little influence on the subsequent ones. Motivated by the fact that human thoughts have persistency, we propose a very deep persistent memory network (MemNet) that introduces a memory block, consisting of a recursive unit and a gate unit, to explicitly mine persistent memory through an adaptive learning process. The recursive unit learns multi-level representations of the current state under different receptive fields. The representations and the outputs from the previous memory blocks are concatenated and sent to the gate unit, which adaptively controls how much of the previous states should be reserved, and decides how much of the current state should be stored. We apply MemNet to three image restoration tasks, i.e., image denosing, super-resolution and JPEG deblocking. Comprehensive experiments demonstrate the necessity of the MemNet and its unanimous superiority on all three tasks over the state of the arts. Code is available at https://github.com/tyshiwo/MemNet. Ying Tai, Jian Yang 0003, Xiaoming Liu 0002, Chunyan Xu |
ICCV | 3 |
| 2017 | Towards Large-Pose Face Frontalization in the Wild
Xi Yin 0001, Xiang Yu 0002, Kihyuk Sohn, Xiaoming Liu 0002, Manmohan Krishna Chandraker |
ICCV | 4 |
| 2017 | On Geometric Features for Skeleton-Based Action Recognition Using Multilayer LSTM NetworksabstractRNN-based approaches have achieved outstanding performance on action recognition with skeleton inputs. Currently these methods limit their inputs to coordinates of joints and improve the accuracy mainly by extending RNN models to spatial domains in various ways. While such models explore relations between different parts directly from joint coordinates, we provide a simple universal spatial modeling method perpendicular to the RNN model enhancement. Specifically, we select a set of simple geometric features, motivated by the evolution of previous work. With experiments on a 3-layer LSTM framework, we observe that the geometric relational features based on distances between joints and selected lines outperform other features and achieve state-of-art results on four datasets. Further, we show the sparsity of input gate weights in the first LSTM layer trained by geometric features and demonstrate that utilizing joint-line distances as input require less data for training. Songyang Zhang 0004, Xiaoming Liu 0002, Jun Xiao 0001 |
WACV | 2 |
| 2017 | Pose-Invariant Face Alignment via CNN-Based Dense 3D Model Fitting
Amin Jourabloo, Xiaoming Liu 0002 |
Int. J. Comput. Vis. | 2 |
| 2017 | Fast single image dehazing based on a regression model
Zhong Luan, Xiuzhuang Zhou, Zhuhong Shao, Guodong Guo, Xiaoming Liu 0002 |
Neurocomputing | 6 |
| 2017 | Adaptive 3D Face Reconstruction from Unconstrained Photo CollectionsabstractGiven a photo collection of "unconstrained" face images of one individual captured under a variety of unknown pose, expression, and illumination conditions, this paper presents a method for reconstructing a 3D face surface model of the individual along with albedo information. Unlike prior work on face reconstruction that requires large photo collections, we formulate an approach to adapt to photo collections with a high diversity in both the number of images and the image quality. To achieve this, we incorporate prior knowledge about face shape by fitting a 3D morphable model to form a personalized template, following by using a novel photometric stereo formulation to complete the fine details, under a coarse-to-fine scheme. Our scheme incorporates a structural similarity-based local selection step to help identify a common expression for reconstruction while discarding occluded portions of faces. The evaluation of reconstruction performance is through a novel quality measure, in the absence of ground truth 3D scans. Superior large-scale experimental results are reported on synthetic, Internet, and personal photo collections. Joseph Roth, Yiying Tong, Xiaoming Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2017 | Automated Online Exam ProctoringabstractMassive open online courses and other forms of remote education continue to increase in popularity and reach. The ability to efficiently proctor remote online examinations is an important limiting factor to the scalability of this next stage in education. Presently, human proctoring is the most common approach of evaluation, by either requiring the test taker to visit an examination center, or by monitoring them visually and acoustically during exams via a webcam. However, such methods are labor intensive and costly. In this paper, we present a multimedia analytics system that performs automatic online exam proctoring. The system hardware includes one webcam, one wearcam, and a microphone for the purpose of monitoring the visual and acoustic environment of the testing location. The system includes six basic components that continuously estimate the key behavior cues: user verification, text detection, voice detection, active window detection, gaze estimation, and phone detection. By combining the continuous estimation components, and applying a temporal sliding window, we design higher level features to classify whether the test taker is cheating at any moment during the exam. To evaluate our proposed system, we collect multimedia (audio and visual) data from $\text{24}$ subjects performing various types of cheating while taking online exams. Extensive experimental results demonstrate the accuracy, robustness, and efficiency of our online exam proctoring system. Yousef Atoum, Alex X. Liu, Stephen D. H. Hsu, Xiaoming Liu 0002 |
IEEE Trans. Multim. | 5 |
| 2016 | Large-Pose Face Alignment via CNN-Based Dense 3D Model FittingabstractLarge-pose face alignment is a very challenging problem in computer vision, which is used as a prerequisite for many important vision tasks, e.g, face recognition and 3D face reconstruction. Recently, there have been a few attempts to solve this problem, but still more research is needed to achieve highly accurate results. In this paper, we propose a face alignment method for large-pose face images, by combining the powerful cascaded CNN regressor method and 3DMM. We formulate the face alignment as a 3DMM fitting problem, where the camera projection matrix and 3D shape parameters are estimated by a cascade of CNN-based regressors. The dense 3D shape allows us to design pose-invariant appearance features for effective CNN learning. Extensive experiments are conducted on the challenging databases (AFLW and AFW), with comparison to the state of the art. Amin Jourabloo, Xiaoming Liu 0002 |
CVPR | 2 |
| 2016 | Adaptive 3D Face Reconstruction from Unconstrained Photo CollectionsabstractGiven a collection of "in-the-wild" face images captured under a variety of unknown pose, expression, and illumination conditions, this paper presents a method for reconstructing a 3D face surface model of an individual along with albedo information. Motivated by the success of recent face reconstruction techniques on large photo collections, we extend prior work to adapt to low quality photo collections with fewer images. We achieve this by fitting a 3D Morphable Model to form a personalized template and developing a novel photometric stereo formulation, under a coarse-to-fine scheme. Superior experimental results are reported on synthetic and real-world photo collections. Joseph Roth, Yiying Tong, Xiaoming Liu 0002 |
CVPR | 3 |
| 2016 | Face Alignment Across Large Poses: A 3D SolutionabstractFace alignment, which fits a face model to an image and extracts the semantic meanings of facial pixels, has been an important topic in CV community. However, most algorithms are designed for faces in small to medium poses (below 45), lacking the ability to align faces in large poses up to 90. The challenges are three-fold: Firstly, the commonly used landmark-based face model assumes that all the landmarks are visible and is therefore not suitable for profile views. Secondly, the face appearance varies more dramatically across large poses, ranging from frontal view to profile view. Thirdly, labelling landmarks in large poses is extremely challenging since the invisible landmarks have to be guessed. In this paper, we propose a solution to the three problems in an new alignment framework, called 3D Dense Face Alignment (3DDFA), in which a dense 3D face model is fitted to the image via convolutional neutral network (CNN). We also propose a method to synthesize large-scale training samples in profile views to solve the third problem of data labelling. Experiments on the challenging AFLW database show that our approach achieves significant improvements over state-of-the-art methods. Xiangyu Zhu 0001, Zhen Lei 0001, Xiaoming Liu 0002, Hailin Shi, Stan Z. Li |
CVPR | 3 |
| 2016 | Joint Face Alignment and 3D Face Reconstruction
Feng Liu 0013, Dan Zeng 0002, Qijun Zhao, Xiaoming Liu 0002 |
ECCV (5) | 4 |
| 2016 | Temporally Robust Global Motion Compensation by Keypoint-Based Congealing
Seyed Morteza Safdarnejad, Yousef Atoum, Xiaoming Liu 0002 |
ECCV (6) | 3 |
| 2016 | Multi-modality imagery database for plant phenotyping
Jeffrey A. Cruz, Xi Yin 0001, Xiaoming Liu 0002, Saif Muhammad Imran, Daniel D. Morris, David M. Kramer 0001, Jin Chen 0004 |
Mach. Vis. Appl. | 3 |
| 2016 | Leaf segmentation in plant phenotyping: a collation study
Hanno Scharr, Massimo Minervini, Andrew P. French, Christian Klukas, David M. Kramer 0001, Xiaoming Liu 0002, Imanol Luengo, Jean-Michel Pape, Gerrit Polder, Danijela Vukadinovic, Xi Yin 0001, Sotirios A. Tsaftaris |
Mach. Vis. Appl. | 6 |
| 2016 | On developing and enhancing plant-level disease rating systems in real fields
Yousef Atoum, Muhammad Jamal Afridi, Xiaoming Liu 0002, J. Mitchell McGrath, Linda E. Hanson |
Pattern Recognit. | 3 |
| 2016 | Monitoring Aquatic Debris Using Smartphone-Based RobotsabstractMonitoring aquatic debris is of great interest to the ecosystems, marine life, human health, and water transport. This paper presents the design and implementation of SOAR-a vision-based surveillance robot system that integrates an off-the-shelf Android smartphone and a gliding robotic fish for debris monitoring in relatively calm waters. SOAR features real-time debris detection and coverage-based rotation scheduling algorithms. The image processing algorithms for debris detection are specifically designed to address the unique challenges in aquatic environments. The rotation scheduling algorithm provides effective coverage for sporadic debris arrivals despite camera's limited angular view. Moreover, SOAR is able to dynamically offload compute-intensive processing tasks to the cloud for battery power conservation. We have implemented a SOAR prototype and conducted extensive experimental evaluation. The results show that SOAR can accurately detect debris in the presence of various environment and system dynamics, and the rotation scheduling algorithm enables SOAR to capture debris arrivals with reduced energy consumption. Yu Wang 0020, Rui Tan 0001, Guoliang Xing, Jianxun Wang 0001, Xiaobo Tan 0001, Xiaoming Liu 0002, Xiangmao Chang |
IEEE Trans. Mob. Comput. | 6 |
| 2016 | Energy-Efficient Aquatic Environment Monitoring Using Smartphone-Based RobotsabstractMonitoring aquatic environment is of great interest to the ecosystem, marine life, and human health. This article presents the design and implementation of Samba—an aquatic surveillance robot that integrates an off-the-shelf Android smartphone and a robotic fish to monitor harmful aquatic processes such as oil spills and harmful algal blooms. Using the built-in camera of the smartphone, Samba can detect spatially dispersed aquatic processes in dynamic and complex environment. To reduce the excessive false alarms caused by the nonwater area (e.g., trees on the shore), Samba segments the captured images and performs target detection in the identified water area only. However, a major challenge in the design of Samba is the high energy consumption resulted from continuous image segmentation. We propose a novel approach that leverages the power-efficient inertial sensors on smartphones to assist image processing. In particular, based on the learned mapping models between inertial and visual features, Samba uses real-time inertial sensor readings to estimate the visual features that guide image segmentation, significantly reducing the energy consumption and computation overhead. Samba also features a set of lightweight and robust computer vision algorithms, which detect harmful aquatic processes based on their distinctive color features. Last, Samba employs a feedback-based rotation control algorithm to adapt to spatiotemporal development of the target aquatic process. We have implemented a Samba prototype and evaluated it through extensive field experiments, lab experiments, and trace-driven simulations. The results show that Samba can achieve a 94% detection rate, a 5% false alarm rate, and a lifetime up to nearly 2 months. Yu Wang 0020, Rui Tan 0001, Guoliang Xing, Jianxun Wang 0001, Xiaobo Tan 0001, Xiaoming Liu 0002 |
ACM Trans. Sens. Networks | 6 |
| 2015 | Robust Global Motion Compensation in Presence of Predominant Foreground
Seyed Morteza Safdarnejad, Xiaoming Liu 0002, Lalita Udpa |
BMVC | 2 |
| 2015 | Unconstrained 3D face reconstructionabstractThis paper presents an algorithm for unconstrained 3D face reconstruction. The input to our algorithm is an “unconstrained” collection of face images captured under a diverse variation of poses, expressions, and illuminations, without meta data about cameras or timing. The output of our algorithm is a true 3D face surface model represented as a watertight triangulated surface with albedo data or texture information. 3D face reconstruction from a collection of unconstrained 2D images is a long-standing computer vision problem. Motivated by the success of the state-of-the-art method, we developed a novel photometric stereo-based method with two distinct novelties. First, working with a true 3D model allows us to enjoy the benefits of using images from all possible poses, including profiles. Second, by leveraging emerging face alignment techniques and our novel normal field-based Laplace editing, a combination of landmark constraints and photometric stereo-based normals drives our surface reconstruction. Given large photo collections and a ground truth 3D surface, we demonstrate the effectiveness and strength of our algorithm both qualitatively and quantitatively. Joseph Roth, Yiying Tong, Xiaoming Liu 0002 |
CVPR | 3 |
| 2015 | Pose-Invariant 3D Face AlignmentabstractFace alignment aims to estimate the locations of a set of landmarks for a given image. This problem has received much attention as evidenced by the recent advancement in both the methodology and performance. However, most of the existing works neither explicitly handle face images with arbitrary poses, nor perform large-scale experiments on non-frontal and profile face images. In order to address these limitations, this paper proposes a novel face alignment algorithm that estimates both 2D and 3D landmarks and their 2D visibilities for a face image with an arbitrary pose. By integrating a 3D point distribution model, a cascaded coupled-regressor approach is designed to estimate both the camera projection matrix and the 3D landmarks. Furthermore, the 3D model also allows us to automatically estimate the 2D landmark visibilities via surface normal. We use a substantially larger collection of all-pose face images to evaluate our algorithm and demonstrate superior performances than the state-of-the-art methods. Amin Jourabloo, Xiaoming Liu 0002 |
ICCV | 2 |
| 2015 | Samba: a smartphone-based robot system for energy-efficient aquatic environment monitoringabstractMonitoring aquatic environment is of great interest to the ecosystem, marine life, and human health. This paper presents the design and implementation of Samba -- an aquatic surveillance robot that integrates an off-the-shelf Android smartphone and a robotic fish to monitor harmful aquatic processes such as oil spill and harmful algal blooms. Using the built-in camera of on-board smartphone, Samba can detect spatially dispersed aquatic processes in dynamic and complex environments. To reduce the excessive false alarms caused by the non-water area (e.g., trees on the shore), Samba segments the captured images and performs target detection in the identified water area only. However, a major challenge in the design of Samba is the high energy consumption resulted from the continuous image segmentation. We propose a novel approach that leverages the power-efficient inertial sensors on smartphone to assist the image processing. In particular, based on the learned mapping models between inertial and visual features, Samba uses real-time inertial sensor readings to estimate the visual features that guide the image segmentation, significantly reducing energy consumption and computation overhead. Samba also features a set of lightweight and robust computer vision algorithms, which detect harmful aquatic processes based on their distinctive color features. Lastly, Samba employs a feedback-based rotation control algorithm to adapt to spatiotemporal evolution of the target aquatic process. We have implemented a Samba prototype and evaluated it through extensive field experiments, lab experiments, and trace-driven simulations. The results show that Samba can achieve 94% detection rate, 5% false alarm rate, and a lifetime up to nearly two months. Yu Wang 0020, Rui Tan 0001, Guoliang Xing, Jianxun Wang 0001, Xiaobo Tan 0001, Xiaoming Liu 0002 |
IPSN | 6 |
| 2015 | Automatic in Vivo Cell Detection in MRI
Muhammad Jamal Afridi, Xiaoming Liu 0002, Erik M. Shapiro, Arun Ross |
MICCAI (3) | 2 |
| 2015 | Demographic Estimation from Face Images: Human vs. Machine PerformanceabstractDemographic estimation entails automatic estimation of age, gender and race of a person from his face image, which has many potential applications ranging from forensics to social media. Automatic demographic estimation, particularly age estimation, remains a challenging problem because persons belonging to the same demographic group can be vastly different in their facial appearances due to intrinsic and extrinsic factors. In this paper, we present a generic framework for automatic demographic (age, gender and race) estimation. Given a face image, we first extract demographic informative features via a boosting algorithm, and then employ a hierarchical approach consisting of between-group classification, and within-group regression. Quality assessment is also developed to identify low-quality face images that are difficult to obtain reliable demographic estimates. Experimental results on a diverse set of face image databases, FG-NET (1K images), FERET (3K images), MORPH II (75K images), PCSO (100K images), and a subset of LFW (4K images), show that the proposed approach has superior performance compared to the state of the art. Finally, we use crowdsourcing to study the human perception ability of estimating demographics from face images. A side-by-side comparison of the demographic estimates from crowdsourced data and the proposed algorithm provides a number of insights into this challenging problem. Hu Han 0001, Charles Otto, Xiaoming Liu 0002, Anil K. Jain 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2015 | Automatic Feeding Control for Dense Aquaculture Fish TanksabstractThis paper introduces an efficient visual signal processing system to continuously control the feeding process of fish in aquaculture tanks. The aim is to improve the production profit in fish farms by controlling the amount of feed at an optimal rate. The automatic feeding control includes two components: 1) a continuous decision on whether the fish are actively consuming feed, and 2) automatic detection of the number of excess feed populated on the water surface of the tank using a two-stage approach. The amount of feed is initially detected using the correlation filer applied to an optimum local region within the video frame, and then followed by a SVM-based refinement classifier to suppress the falsely detected feed. Having both measures allows us to accurately control the feeding process in an automated manner. Experimental results show that our system can accurately and efficiently estimate both measures. Yousef Atoum, Steven Srivastava, Xiaoming Liu 0002 |
IEEE Signal Process. Lett. | 3 |
| 2015 | Investigating the Discriminative Power of Keystroke SoundabstractThe goal of this paper is to determine whether keystroke sound can be used to recognize a user. In this regard, we analyze the discriminative power of keystroke sound in the context of a continuous user authentication application. Motivated by the concept of digraphs used in modeling keystroke dynamics, a virtual alphabet is first learned from keystroke sound segments. Next, the digraph latency within the pairs of virtual letters, along with other statistical features, is used to generate match scores. The resultant scores are indicative of the similarities between two sound streams, and are fused to make a final authentication decision. Experiments on both static text-based and free text-based authentications on a database of 50 subjects demonstrate the potential as well as the limitations of keystroke sound. Joseph Roth, Xiaoming Liu 0002, Arun Ross, Dimitris N. Metaxas |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2014 | On Hair Recognition in the Wild by MachineabstractWe present an algorithm for identity verification using only information from the hair. Face recognition in the wild (i.e., unconstrained settings) is highly useful in a variety of applications, but performance suffers due to many factors, e.g., obscured face, lighting variation, extreme pose angle, and expression. It is well known that humans utilize hair for identification under many of these scenarios due to either the consistent hair appearance of the same subject or obvious hair discrepancy of different subjects, but little work exists to replicate this intelligence artificially. We propose a learned hair matcher using shape, color, and texture features derived from localized patches through an AdaBoost technique with abstaining weak classifiers when features are not present in the given location. The proposed hair matcher achieves 71.53% accuracy on the LFW View 2 dataset. Hair also reduces the error of a Commercial Off-The-Shelf (COTS) face matcher through simple score-level fusion by 5.7%. Joseph Roth, Xiaoming Liu 0002 |
AAAI | 2 |
| 2014 | On the Exploration of Joint Attribute Learning for Person Re-identification
Joseph Roth, Xiaoming Liu 0002 |
ACCV (1) | 2 |
| 2014 | Genre categorization of amateur sports videos in the wildabstractVarious sports video genre categorization methods are proposed recently, mainly focusing on professional sports videos captured for TV broadcasting. This paper aims to categorize sports videos in the wild, captured using mobile phones by people watching a game or practicing a sport. Thus, no assumption is made about video production practices or existence of field lining and equipment. Motivated by distinctiveness of motions in sports activities, we propose a novel motion trajectory descriptor to effectively and efficiently represent a video. Furthermore, temporal analysis of local descriptors is proposed to integrate the categorization decision over time. Experiments on a newly collected dataset of amateur sports videos in the wild demonstrate that our trajectory descriptor is superior for sports videos categorization and temporal analysis improves the categorization accuracy further. Seyed Morteza Safdarnejad, Xiaoming Liu 0002, Lalita Udpa |
ICIP | 2 |
| 2014 | Multi-leaf tracking from fluorescence plant videosabstractDriven by the plant phenotyping application, this paper proposes a new leaf tracking framework to jointly segment, align and track multiple leaves from fluorescence plant videos. Our framework consists of two steps. First, leaf alignment is applied to one video frame to generate a collection of leaf candidates. Second, we define a set of transformation parameters operated on the leaf candidates in order to optimize the alignment in the subsequent video frame according to an objective function. Gradient descent is employed to solve this optimization problem. Experimental results show that the proposed multi-leaf tracking algorithm is superior to the image-based leaf alignment method in terms of three quantitative metrics. Xi Yin 0001, Xiaoming Liu 0002, Jin Chen 0004, David M. Kramer 0001 |
ICIP | 2 |
| 2014 | An Automated System for Plant-Level Disease Rating in Real FieldsabstractCercospora leaf spot (CLS) is the most serious disease in sugar beet plants that significantly reduces the sugar yield throughout the world. Therefore the current focus of the researchers in agricultural domain is to find sugar beet cultivars that are highly resistant to CLS. To measure their resistance, CLS is manually observed and rated in a large variety of sugar beet by different human experts over a period of a few months. Unfortunately, this procedure is laborious and subjective. Therefore, we propose a novel computer vision system, CLS Rater, to automatically and accurately rate CLS of plant images in the real field to the "USDA scale" of 0 to 10. Given a set of plant images captured by a tractor-mounted camera, CLS Rater extracts multi-scale super pixels, where in each scale a novel histogram of importances feature representation is proposed to encode both the within-super pixel local and across-super pixel global appearance variations. These features at different super pixel scales are then fused for learning a bagging M5P regress or that estimates the rating for each plant image. We test our system on the field data collected over a period of two months under different day lighting and weather conditions. Experimental results show CLS Rater to be highly consistent with a rating error of 0.65, which demonstrates higher consistency than the rating standard deviation of 1.31 by the human experts. Muhammad Jamal Afridi, Xiaoming Liu 0002, J. Mitchell McGrath |
ICPR | 2 |
| 2014 | Aquatic debris monitoring using smartphone-based robotic sensors
Yu Wang 0020, Rui Tan 0001, Guoliang Xing, Jianxun Wang 0001, Xiaobo Tan 0001, Xiaoming Liu 0002, Xiangmao Chang |
IPSN | 6 |
| 2014 | Image segmentation of mesenchymal stem cells in diverse culturing conditionsabstractResearchers in the areas of regenerative medicine and tissue engineering have great interests in understanding the relationship of different sets of culturing conditions and applied mechanical stimuli to the behavior of mesenchymal stem cells (MSCs). However, it is challenging to design a tool to perform automatic cell image analysis due to the diverse morphologies of MSCs. Therefore, as a primary step towards developing the tool, we propose a novel approach for accurate cell image segmentation. We collected three MSC datasets cultured on different surfaces and exposed to diverse mechanical stimuli. By analyzing existing approaches on our data, we choose to substantially extend binarization-based extraction of alignment score (BEAS) approach by extracting novel discriminating features and developing an adaptive threshold estimation model. Experimental results on our data shows our approach is superior to seven conventional techniques. We also define three quantitative measures to analyze the characteristics of images in our datasets. To the best of our knowledge, this is the first study that applied automatic segmentation to live MSC cultured on different surfaces with applied stimuli. Muhammad Jamal Afridi, Chun Liu 0002, Christina Chan, Seungik Baek, Xiaoming Liu 0002 |
WACV | 5 |
| 2014 | Multi-leaf alignment from fluorescence plant imagesabstractIn this paper, we propose a multi-leaf alignment framework based on Chamfer matching to study the problem of leaf alignment from fluorescence images of plants, which will provide a leaf-level analysis of photosynthetic activities. Different from the naive procedure of aligning leaves iteratively using the Chamfer distance, the new algorithm aims to find the best alignment of multiple leaves simultaneously in an input image. We formulate an optimization problem of an objective function with three terms: the average of chamfer distances of aligned leaves, the number of leaves, and the difference between the synthesized mask by the leaf candidates and the original image mask. Gradient descent is used to minimize our objective function. A quantitative evaluation framework is also formulated to test the performance of our algorithm. Experimental results show that the proposed multi-leaf alignment optimization performs substantially better than the baseline of the Chamfer matching algorithm in terms of both accuracy and efficiency. Xi Yin 0001, Xiaoming Liu 0002, Jin Chen 0004, David M. Kramer 0001 |
WACV | 2 |
| 2014 | Transfer learning with one-class data
Jixu Chen, Xiaoming Liu 0002 |
Pattern Recognit. Lett. | 2 |
| 2014 | On Continuous User Authentication via Typing BehaviorabstractWe hypothesize that an individual computer user has a unique and consistent habitual pattern of hand movements, independent of the text, while typing on a keyboard. As a result, this paper proposes a novel biometric modality named typing behavior (TB) for continuous user authentication. Given a webcam pointing toward a keyboard, we develop real-time computer vision algorithms to automatically extract hand movement patterns from the video stream. Unlike the typical continuous biometrics, such as keystroke dynamics (KD), TB provides a reliable authentication with a short delay, while avoiding explicit key-logging. We collect a video database where 63 unique subjects type static text and free text for multiple sessions. For one typing video, the hands are segmented in each frame and a unique descriptor is extracted based on the shape and position of hands, as well as their temporal dynamics in the video sequence. We propose a novel approach, named bag of multi-dimensional phrases, to match the cross-feature and cross-temporal pattern between a gallery sequence and probe sequence. The experimental results demonstrate a superior performance of TB when compared with KD, which, together with our ultrareal-time demo system, warrant further investigation of this novel vision application and biometric modality. Joseph Roth, Xiaoming Liu 0002, Dimitris N. Metaxas |
IEEE Trans. Image Process. | 2 |
| 2013 | Learning person-specific models for facial expression and action unit recognition
Jixu Chen, Xiaoming Liu 0002, Peter H. Tu, Amy Aragones |
Pattern Recognit. Lett. | 2 |
| 2012 | Boosting with Side Information
Jixu Chen, Xiaoming Liu 0002, Siwei Lyu |
ACCV (1) | 2 |
| 2012 | Adaptive Unsupervised Multi-view Feature Selection for Visual Concept Recognition
Yinfu Feng, Jun Xiao 0001, Yueting Zhuang, Xiaoming Liu 0002 |
ACCV (1) | 4 |
| 2012 | Spatio-Temporal Phrases for Activity Recognition
Xiaoming Liu 0002, Ming-Ching Chang, Weina Ge, Tsuhan Chen |
ECCV (3) | 2 |
| 2012 | Person-specific expression recognition with transfer learningabstractA key assumption of traditional machine learning is that both the training and test data share the same distribution. However, this assumption does not hold in many real-world scenarios. For example, in facial expression recognition, the appearance of an expression may vary significantly for different people. Previous work has shown that learning from adequate person-specific data can improve facial expression recognition results. However, because of the difficulties of data collection and labeling, person-specific data is usually very sparse in real-world applications. Learning from the sparse data may suffer from serious over-fitting. In this paper, we propose to learn a person-specific facial expression model through transfer learning. By transferring the informative knowledge from other people, it allows us to learn an accurate person-specific model for a new subject with only a small amount of his/her specific data. Jixu Chen, Xiaoming Liu 0002, Peter H. Tu, Amy Aragones |
ICIP | 2 |
| 2012 | Image congealing via efficient feature selectionabstractCongealing for an image ensemble is a joint alignment process to rectify images in the spatial domain such that the aligned images are as similar to each other as possible. Fruitful congealing algorithms were applied to various object classes and medical applications. However, relatively little effort has been taken in the direction of compact and effective feature representations for each image. To remedy this problem, the least-square-based congealing framework is extended by incorporating an unsupervised feature selection algorithm, which substantially removes feature redundancy and leads to a more efficient congealing with even higher accuracy. Furthermore, our novel feature selection algorithm itself is an independent contribution. It is not explicitly linked to the congealing algorithm and can be directly applied to other learning tasks. Extensive experiments are conducted for both the feature selection and congealing algorithms. Ya Xue, Xiaoming Liu 0002 |
WACV | 2 |
| 2012 | Group context learning for event recognitionabstractWe address the problem of group-level event recognition from videos. The events of interest are defined based on the motion and interaction of members in a group over time. Example events include group formation, dispersion, following, chasing, flanking, and fighting. To recognize these complex group events, we propose a novel approach that learns the group-level scenario context from automatically extracted individual trajectories. We first perform a group structure analysis to produce a weighted graph that represents the probabilistic group membership of the individuals. We then extract features from this graph to capture the motion and action contexts among the groups. The features are represented using the “bag-of-words” scheme. Finally, our method uses the learned Support Vector Machine (SVM) to classify a video segment into the six event categories. Our implementation builds upon a mature multi-camera multi-target tracking system that recognizes the group-level events involving up to 20 individuals in real-time. Weina Ge, Ming-Ching Chang, Xiaoming Liu 0002 |
WACV | 4 |
| 2012 | Semi-supervised facial landmark annotation
Xiaoming Liu 0002, Frederick W. Wheeler, Peter H. Tu |
Comput. Vis. Image Underst. | 2 |
| 2011 | Optimal gradient pursuit for face alignmentabstractFace alignment aims to fit a deformable landmark-based mesh to a facial image so that all facial features can be located accurately. In discriminative face alignment, an alignment score function, which is treated as the appearance model, is learned such that moving along its gradient direction can improve the alignment. This paper proposes a new face model named “Optimal Gradient Pursuit Model”, where the objective is to minimize the angle between the gradient direction and the vector pointing toward the ground-truth shape parameter. We formulate an iterative approach to solve this minimization problem. With extensive experiments in generic face alignment, we show that our model improves the alignment accuracy and speed compared to the state-of-the-art discriminative face alignment approach. Xiaoming Liu 0002 |
FG | 1 |
| 2011 | LPSM: Fitting shape model by linear programmingabstractWe propose a shape model fitting algorithm that uses linear programming optimization. Most shape model fitting approaches (such as ASM, AAM) are based on gradient-descent-like local search optimization and usually suffer from local minima. In contrast, linear programming (LP) techniques achieve globally optimal solution for linear problems. In [1], a linear programming scheme based on successive convexification was proposed for matching static object shape in images among cluttered background and achieved very good performance. In this paper, we rigorously derive the linear formulation of the shape model fitting problem in the LP scheme and propose an LP shape model fitting algorithm (LPSM). In the experiments, we compared the performance of our LPSM with the LP graph matching algorithm(LPGM), ASM, and a CONDENSATION based ASM algorithm on a test set of PUT database. The experiments show that LPSM can achieve higher shape fitting accuracy. We also evaluated its performance on the fitting of some real world face images collected from internet. The results show that LPSM can handle various appearance outliers and can avoid local minima problem very well, as the fitting is carried out by LP optimization with l1norm robust cost function. Jilin Tu, Brandon Laflen, Xiaoming Liu 0002, Musodiq O. Bello, Jens Rittscher, Peter H. Tu |
FG | 3 |
| 2010 | Facial Contour Labeling via Congealing
Xiaoming Liu 0002, Frederick W. Wheeler, Peter H. Tu |
ECCV (1) | 1 |
| 2010 | Video-based face model fitting using Adaptive Active Appearance Model
Xiaoming Liu 0002 |
Image Vis. Comput. | 1 |
| 2009 | Automatic facial landmark labeling with minimal supervisionabstractLandmark labeling of training images is essential for many learning tasks in computer vision, such as object detection, tracking, and alignment. Image labeling is typically conducted manually, which is both labor-intensive and error-prone. To improve this process, this paper proposes a new approach to estimate a set of landmarks for a large image ensemble with only a small number of manually labeled images from the ensemble. Our approach, named semi-supervised least-squares congealing, aims to minimize an objective function defined on both labeled and unlabeled images. A shape model is learnt on-line to constrain the landmark configuration. We also employ a partitioning strategy to allow coarse-to-fine landmark estimation. Extensive experiments on facial images show that our approach can reliably and accurately label landmarks for a large image ensemble starting from a small number of manually labeled images, under various challenging scenarios. Xiaoming Liu 0002, Frederick W. Wheeler, Peter H. Tu |
CVPR | 2 |
| 2009 | Simultaneous alignment and clustering for an image ensembleabstractJoint alignment for an image ensemble can rectify images in the spatial domain such that the aligned images are as similar to each other as possible. This important technology has been applied to various object classes and medical applications. However, previous approaches to joint alignment work on an ensemble of a single object class. Given an ensemble with multiple object classes, we propose an approach to automatically and simultaneously solve two problems, image alignment and clustering. Both the alignment parameters and clustering parameters are formulated into a unified objective function, whose optimization leads to an unsupervised joint estimation approach. It is further extended to semi-supervised simultaneous estimation where a few labeled images are provided. Extensive experiments on diverse real-world databases demonstrate the capabilities of our work on this challenging problem. Xiaoming Liu 0002, Frederick W. Wheeler |
ICCV | 1 |
| 2009 | On optimizing subspaces for face recognitionabstractWe propose a subspace learning algorithm for face recognition by directly optimizing recognition performance scores. Our approach is motivated by the following observations: 1) Different face recognition tasks (i.e., face identification and verification) have different performance metrics, which implies that there exist distinguished subspaces that optimize these scores, respectively. Most prior work focused on optimizing various discriminative or locality criteria and neglect such distinctions. 2) As the gallery (target) and the probe (query) data are collected in different settings in many real-world applications, there could exist consistent appearance incoherences between the gallery and the probe data for the same subject. Knowledge regarding these incoherences could be used to guide the algorithm design, resulting in performance gain. Prior efforts have not focused on these facts. In this paper, we rigorously formulate performance scores for both the face identification and the face verification tasks, provide a theoretical analysis on how the optimal subspaces for the two tasks are related, and derive gradient descent algorithms for optimizing these subspaces. Our extensive experiments on a number of public databases and a real-world face database demonstrate that our algorithm can improve the performance of given subspace based face recognition algorithms targeted at a specific face recognition task. Jilin Tu, Xiaoming Liu 0002, Peter H. Tu |
ICCV | 2 |
| 2009 | Discriminative Face AlignmentabstractThis paper proposes a discriminative framework for efficiently aligning images. Although conventional Active Appearance Models (AAMs)-based approaches have achieved some success, they suffer from the generalization problem, i.e., how to align any image with a generic model. We treat the iterative image alignment problem as a process of maximizing the score of a trained two-class classifier that is able to distinguish correct alignment (positive class) from incorrect alignment (negative class). During the modeling stage, given a set of images with ground truth landmarks, we train a conventional Point Distribution Model (PDM) and a boosting-based classifier, which acts as an appearance model. When tested on an image with the initial landmark locations, the proposed algorithm iteratively updates the shape parameters of the PDM via the gradient ascent method such that the classification score of the warped image is maximized. We use the term Boosted Appearance Models (BAMs) to refer to the learned shape and appearance models, as well as our specific alignment method. The proposed framework is applied to the face alignment problem. Using extensive experimentation, we show that, compared to the AAM-based approach, this framework greatly improves the robustness, accuracy, and efficiency of face alignment by a large margin, especially for unseen data. Xiaoming Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2008 | Boosted deformable model for human body alignmentabstractThis paper studies image alignment, the problem of learning a shape and appearance model from labeled data and efficiently fitting the model to a non-rigid object with large variations. Given a set of images with manually labeled landmarks, our model representation consists of a shape component represented by a point distribution model and an appearance component represented by a collection of local features, trained discriminatively as a two-class classifier using boosting. Images with ground truth landmarks are the positive training samples while those with perturbed landmarks are considered as negatives. Enabled by piece-wise affine warping, corresponding local feature positions across all training samples form a hypothesis space for boosting. Image alignment is performed by maximizing the boosted classifier score, which is our distance measure, through iteratively mapping the feature positions to the image, and computing the gradient direction of the score with respect to the shape parameter. We apply this approach to human body alignment from surveillance-type images. We conduct experiments on the MIT pedestrian database where the body size is approximately 110 times 46 pixels, and demonstrate our real-time alignment capability. Xiaoming Liu 0002, Ting Yu 0003, Thomas Sebastian, Peter H. Tu |
CVPR | 1 |
| 2008 | Face alignment via boosted ranking modelabstractFace alignment seeks to deform a face model to match it with the features of the image of a face by optimizing an appropriate cost function. We propose a new face model that is aligned by maximizing a score function, which we learn from training data, and that we impose to be concave. We show that this problem can be reduced to learning a classifier that is able to say whether or not by switching from one alignment to a new one, the model is approaching the correct fitting. This relates to the ranking problem where a number of instances need to be ordered. For training the model, we propose to extend GentleBoost [23] to rank-learning. Extensive experimentation shows the superiority of this approach to other learning paradigms, and demonstrates that this model exceeds the alignment performance of the state-of-the-art. Xiaoming Liu 0002, Gianfranco Doretto |
CVPR | 2 |
| 2007 | Face Mosaicing for Pose Robust Video-Based Recognition
Xiaoming Liu 0002, Tsuhan Chen |
ACCV (2) | 1 |
| 2007 | What are customers looking at?abstractComputer vision approaches for retail applications can provide value far beyond the common domain of loss prevention. Gaining insight into the movement and behaviors of shoppers is of high interest for marketing, merchandizing, store operations and data mining. Of particular interest is the process of purchase decision making. What catches a customers attention? What products go unnoticed? What does a customer look at before making a final decision? Towards this goal we presents a system that detects and tracks both the location and gaze of shoppers in retail environments. While networks of standard overhead store cameras are used for tracking the location of customers, small in-shelf cameras are used for estimating customer gaze. The presented system operates robustly in real-time and can be deployed in a variety of retail applications. Xiaoming Liu 0002, Nils Krahnstoever, Ting Yu 0003, Peter H. Tu |
AVSS | 1 |
| 2007 | Improved Face Model Fitting on Video SequencesabstractActive Appearance Models (AAMs) represent the shape and appearance of an object via two low-dimensional subspaces, one for shape and one for appearance. AAMs for facial images are currently receiving considerable attention from the computer vision community. However, most existing work focuses on fitting AAMs to a single image. For many applications, effectively fitting an AAM to video sequences is of critical importance and challenging, especially considering the varying quality of real-world video content. This paper proposes a hybrid model to address this problem. Both a generic AAM and a subject-specific model are employed simultaneously in the proposed fitting scheme. Experimental results from outdoor surveillance video sequences demonstrate the improved image registration across video frames and faster fitting convergence. 1 Xiaoming Liu 0002, Frederick W. Wheeler, Peter H. Tu |
BMVC | 1 |
| 2007 | Generic Face Alignment using Boosted Appearance ModelabstractThis paper proposes a discriminative framework for efficiently aligning images. Although conventional active appearance models (AAM)-based approaches have achieved some success, they suffer from the generalization problem, i.e., how to align any image with a generic model. We treat the iterative image alignment problem as a process of maximizing the score of a trained two-class classifier that is able to distinguish correct alignment (positive class) from incorrect alignment (negative class). During the modeling stage, given a set of images with ground truth landmarks, we train a conventional point distribution model (PDM) and a boosting-based classifier, which we call boosted appearance model (BAM). When tested on an image with the initial landmark locations, the proposed algorithm iteratively updates the shape parameters of the PDM via the gradient ascent method such that the classification score of the warped image is maximized. The proposed framework is applied to the face alignment problem. Using extensive experimentation, we show that, compared to the AAM-based approach, this framework greatly improves the robustness, accuracy and efficiency of face alignment by a large margin, especially for unseen data. Xiaoming Liu 0002 |
CVPR | 1 |
| 2007 | Automatic Face Recognition from Skeletal RemainsabstractThe ability to determine the identity of a skull found at a crime scene is of critical importance to the law enforcement community. Traditional clay-based methods attempt to reconstruct the face so as to enable identification of the deceased by members of the general public. However, these reconstructions lack consistency from practitioner to practitioner and it has been shown that the human recognition of these reconstructions against a photo gallery of potential victims is little better than chance. In this paper we propose the automation of the reconstruction process. For a given skull, a data-driven 3D generative model of the face is constructed using a database of CT head scans. The reconstruction can be constrained based on prior knowledge such as age and or weight. To determine whether or not these reconstructions have merit, geometric methods for comparing reconstructions against a gallery of facial images are proposed. First, active shape models are used to automatically detect a set of facial landmarks on each image. These landmarks are associated with 3D points on the reconstruction. Direct comparison of the reconstruction is problematic since in general the camera geometry used for image capture is unknown and there are uncertainties associated with the reconstruction and landmark detection processes. The first method of comparison uses constrained optimization to determine the optimal projection of the reconstruction on to the image. Residuals are then analyzed resulting in a ranking of the gallery. The second method uses boosting to learn which points are both reliable and discriminating. This results in a match/no-match classifier. Experimental evidence indicating that skull recognition from facial images can be achieved is presented. Peter H. Tu, Rebecca Book, Xiaoming Liu 0002, Nils Krahnstoever, Carl Adrian, Phil Williams |
CVPR | 3 |
| 2007 | Gradient Feature Selection for Online BoostingabstractBoosting has been widely applied in computer vision, especially after Viola and Jones's seminal work. The marriage of rectangular features and integral-image- enabled fast computation makes boosting attractive for many vision applications. However, this popular way of applying boosting normally employs an exhaustive feature selection scheme from a very large hypothesis pool, which results in a less-efficient learning process. Furthermore, this poses additional constraint on applying boosting in an onine fashion, where feature re-selection is often necessary because of varying data characteristic, but yet impractical due to the huge hypothesis pool. This paper proposes a gradient-based feature selection approach. Assuming a generally trained feature set and labeled samples are given, our approach iteratively updates each feature using the gradient descent, by minimizing the weighted least square error between the estimated feature response and the true label. In addition, we integrate the gradient-based feature selection with an online boosting framework. This new online boosting algorithm not only provides an efficient way of updating the discriminative feature set, but also presents a unified objective for both feature selection and weak classifier updating. Experiments on the person detection and tracking applications demonstrate the effectiveness of our proposal. Xiaoming Liu 0002, Ting Yu 0003 |
ICCV | 1 |
| 2006 | Face Model Fitting on Low Resolution ImagesabstractActive Appearance Models (AAMs) represent the shape and appearance of an object via two low-dimensional subspaces, one for shape and one for appearance. AAMs for facial images are currently receiving considerable attention from the vision community. However, most existing work focuses on fitting AAMs to high-quality facial images. For many applications, effectively fitting an AAM to low-resolution facial images is of critical importance. This paper addresses this challenge from two aspects. On the modeling side, we propose an iterative AAM enhancement scheme, which not only results in increased fitting speed, but also improves the fitting robustness. For fitting AAMs to low-resolution images, we build a multi-resolution AAM and show that the best fitting performance is obtained when the model resolution is slightly higher than the facial image resolution. Experimental results using both indoor video and outdoor surveillance video are presented. 1 Xiaoming Liu 0002, Peter H. Tu, Frederick W. Wheeler |
BMVC | 1 |
| 2006 | Optimal Pose for Face RecognitionabstractResearchers in psychology have well studied the impact of the pose of a face as perceived by humans, and concluded that the so-called 3/4 view, halfway between the front view and the profile view, is the easiest for face recognition by humans. For face recognition by machines, while much work has been done to create recognition algorithms that are robust to pose variation, little has been done in finding the most representative pose for recognition. In this paper, we use a number of algorithms to evaluate face recognition performance when various poses are used for training. The result, similar to findings in psychology that the 3/4 view is the best, is also justified by the discrimination power of different regions on the face, computed from both the appearance and the geometry of these regions. We believe our study is both scientifically interesting and practically beneficial for many applications. Xiaoming Liu 0002, Jens Rittscher, Tsuhan Chen |
CVPR (2) | 1 |
| 2005 | Detecting and counting people in surveillance applicationsabstractA number of surveillance scenarios require the detection and tracking of people. Although person detection and counting systems are commercially available today, there is need for further research to address the challenges of real world scenarios. The focus of this work is the segmentation of groups of people into individuals. One relevant application of this algorithm is people counting. Experiments document that the presented approach leads to robust people counts. Xiaoming Liu 0002, Peter H. Tu, Jens Rittscher, A. G. Amitha Perera, Nils Krahnstoever |
AVSS | 1 |
| 2005 | Pose-Robust Face Recognition Using Geometry Assisted Probabilistic ModelingabstractResearchers have been working on human face recognition for decades. Face recognition is hard due to different types of variations in face images, such as pose, illumination and expression, among which pose variation is the hardest one to deal with. To improve face recognition under pose variation, this paper presents a geometry assisted probabilistic approach. We approximate a human head with a 3D ellipsoid model, so that any face image is a 2D projection of such a 3D ellipsoid at a certain pose. In this approach, both training and test images are back projected to the surface of the 3D ellipsoid, according to their estimated poses, to form the texture maps. Thus the recognition can be conducted by comparing the texture maps instead of the original images, as done in traditional face recognition. In addition, we represent the texture map as an array of local patches, which enables us to train a probabilistic model for comparing corresponding patches. By conducting experiments on the CMU PIE database, we show that the proposed algorithm provides better performance than the existing algorithms. Xiaoming Liu 0002, Tsuhan Chen |
CVPR (1) | 1 |
| 2005 | Online Modeling and Tracking of Pose-Varying Faces in VideoabstractWe propose a face mosaicing approach to model both the facial appearance and geometry from pose-varying videos, and apply it in face tracking and recognition. The basic idea is that by approximating the human head as a 3D ellipsoid, multi-view face images can be back projected onto the surface of the ellipsoid, and the surface texture map is decomposed into an array of local patches. During the online modeling process, the position and pose of the first frame is assumed to be known for a given video sequence. For each frame in the sequence, the algorithm estimates the face-position and pose, and generates a texture map, which is further utilized in updating the mosaic model. Xiaoming Liu 0002, Tsuhan Chen |
CVPR (2) | 1 |
| 2003 | Video-Based Face Recognition Using Adaptive Hidden Markov ModelsabstractWhile traditional face recognition is typically based on still images, face recognition from video sequences has become popular. In this paper, we propose to use adaptive hidden Markov models (HMM) to perform video-based face recognition. During the training process, the statistics of training video sequences of each subject, and the temporal dynamics, are learned by an HMM. During the recognition process, the temporal characteristics of the test video sequence are analyzed over time by the HMM corresponding to each subject. The likelihood scores provided by the HMMs are compared, and the highest score provides the identity of the test video sequence. Furthermore, with unsupervised learning, each HMM is adapted with the test video sequence, which results in better modeling over time. Based on extensive experiments with various databases, we show that the proposed algorithm results in better performance than using majority voting of image-based recognition results. Xiaoming Liu 0002, Tsuhan Chen |
CVPR (1) | 1 |
| 2003 | Geometry-assisted statistical modeling for face mosaicingabstractThe modeling of facial appearance has many applications. This paper proposes an approach to generating a statistical face model based on video mosaicing. Unlike traditional video mosaicing, we use the geometry of a face to improve the mosaicing result. Given a face sequence, each frame is unwrapped onto certain portion of the surface of a sphere, as determined by spherical projection and the minimization procedure using the Levenberg-Marquardt algorithm. A statistical model containing a mean image and a number of eigenimages, instead of only one image template, is used to represent the face mosaic. Good experimental results have been observed. Xiaoming Liu 0002, Tsuhan Chen |
ICIP (2) | 1 |
| 2003 | Face authentication for multiple subjects using eigenflow
Xiaoming Liu 0002, Tsuhan Chen, B. V. K. Vijaya Kumar |
Pattern Recognit. | 1 |
| 2003 | Eigenspace updating for non-stationary process and its application to face recognition
Xiaoming Liu 0002, Tsuhan Chen, Susan M. Thornton |
Pattern Recognit. | 1 |
| 2002 | Shot boundary detection using temporal statistics modelingabstractIn multimedia information retrieval, shot boundary detection is a very active research topic. In order to perform shot boundary detection, we propose an algorithm for modeling temporal statistics using a novel eigenspace updating method. The feature extracted from the current frame is compared with a model trained from features in the previous frames. A shot boundary is detected if the new feature does not fit well to the existing model. The model is based on principal component analysis (PCA), or the eigenspace method, in which the eigenspace can be updated to capture the non-stationary statistics of the features. The experiment results show that the proposed algorithm outperforms the traditional direct differencing method. Xiaoming Liu 0002, Tsuhan Chen |
ICASSP | 1 |
| 2002 | Principle component analysis and its variants for biometricsabstractPrinciple component analysis (PCA) has been widely used for analyzing the statistics of data. While applied to biometrics as a classification scheme, PCA faces certain challenges. We present a number of modifications to PCA in order to meet these challenges. Using face recognition as an example, we show how eigenflow, PCA applied to optimal flow, enables us to measure the difference between two images while allowing expression changes and registration error. We show how PCA can be updated to model time-varying statistics. We also show that PCA can be used to model the surface reflectance of human faces and reduce illumination variation that defeats most existing face recognition algorithms. Finally, we present distinguishing component analysis (DCA) and apply it to multimodal biometric authentication. Tsuhan Chen, Yufeng Jessie Hsu, Xiaoming Liu 0002, Wende Zhang |
ICIP (1) | 3 |
| 1999 | Video Motion Capture Using Feature Tracking and Skeleton ReconstructionabstractIn the domain of computer vision, there exists a very wide application for the research of human motion capture. This paper proposes a new approach to do motion capture in video. It is composed of image sequence based tracking of human feature points and the reconstruction of three dimension (3D) motion skeleton. First, we track every part of human body from top to bottom on the basis of a human model. The Kalman filter and a morph-block similarity algorithm based on subpixel are used. Then we do camera calibration using the line correspondences between the 3D model and the image. Finally the 3D motion skeleton is established by using the model knowledge. This approach does not aim at a given mode of human motion. Rather, it analyzes large motion from frame to frame in complex, variational background, and sets up a 3D motion skeleton under the perspective projection. We also present the experimental result at the end of the paper. Yueting Zhuang, Xiaoming Liu 0002, Yunhe Pan |
ICIP (4) | 2 |
| 1999 | Video based human animation techniqueabstractHuman animation is a challenging domain in computer animation. To aim at many shortcomings in conventional techniques, this paper proposes a new video based human animation technique. Given a clip of video, firstly human joints are tracked with the support of Kalman filter and morph-block based match in the image sequence. Then corresponding sequence of three-dimension (3D) human motion skeleton is constructed under the perspective projection using camera calibration and human anatomy knowledge. Finally a motion library is established automatically by annotating multiform motion attributes, which can be browsed and queried by the animator. This approach has the characteristic of rich source material, low computing cost, efficient production, and realistic animation result. We demonstrate it on several video clips of people doing full body movements, and visualize the results by re-animating a 3D human skeleton model. Xiaoming Liu 0002, Yueting Zhuang, Yunhe Pan |
ACM Multimedia (1) | 1 |
| 1999 | A new approach to retrieve video by example video clipabstractThe similarity measure between video clips is a key issue in video retrieval. In the developing of our video retrieval system, we propose a new video similarity model. In contrast to existing algorithms, it proposes many influencing factors, such as order factor, speed factor, disturbance factor, etc, based on the subjective visual judgement of human. So this algorithm embodies the degree of similarity completely and systematically. On the other hand, it has resolution adaptation because it can be applied to every level of video structure. In the retrieval system, it can be used to process video query by example clip. This paper introduces it in detail and presents experiment results at the end of the paper. Xiaoming Liu 0002, Yueting Zhuang, Yunhe Pan |
ACM Multimedia (2) | 1 |
| 1999 | Video based human motion captureabstractProposes a new approach to capture human motion in video. This approach does not aim at a given human motion mode, but instead analyzes large-scale motion from frame to frame in a complex variational background and sets up a 3D motion skeleton under perspective projection. This approach is composed of two steps. First, we track every part of a human body from top to bottom on the basis of a human model. Then we perform a camera calibration using the line correspondences between the 3D model and the image, and establish the 3D motion skeleton by using the human model knowledge. The experimental results are presented at the end of the paper. Xiaoming Liu 0002, Yueting Zhuang, Yi Wu 0012, Yunhe Pan |
MMSP | 1 |