VLDB 2026 Research / reviewers in the wild / expert
Zicheng Liu 0001
dblp:l/ZichengLiu
· DBLP profile ↗
195ranked-venue papers
14as first author
74since 2021 · last 2026
0000-0001-5894-7828ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 153 · 13 first-author · 51 since 2021Artificial intelligence and machine learning · 105 · 1 first-author · 62 since 2021Human-computer interaction and ubiquitous computing · 6 · 2 first-authorTheory of computation · 3 · 1 first-authorSystems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Conditional Text-to-Image Generation with Reference GuidanceabstractText-to-image diffusion models have demonstrated tremendous success in synthesizing visually stunning images given textual instructions. Despite remarkable progress in creating high-fidelity visuals, text-to-image models can still struggle with precisely rendering subjects, such as text spelling. To address this challenge, this paper explores using additional conditions of an image that provides visual guidance of the particular subjects for diffusion models to generate. In addition, this reference condition empowers the model to be conditioned in ways that the vocabularies of the text tokenizer cannot adequately represent, and further extends the model’s generalization to novel capabilities such as generating non-English text spellings. We develop several small-scale expert plugins that efficiently endow a Stable Diffusion model with the capability to take different references. Each plugin is trained with auxiliary networks and loss functions customized for applications such as English scene-text generation, multi-lingual scene-text generation, and logo-image generation. Our expert plugins demonstrate superior results than the existing methods on all tasks, each containing only 28.55M trainable parameters. Ze Wang 0008, Zhengyuan Yang, Jiang Wang 0012, Zicheng Liu 0001, Qiang Qiu 0001 |
WACV | 6 |
| 2025 | Self-Taught Agentic Long Context UnderstandingabstractYufan Zhuang, Xiaodong Yu, Jialian Wu, Ximeng Sun, Ze Wang, Jiang Liu, Yusheng Su, Jingbo Shang, Zicheng Liu, Emad Barsoum. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yufan Zhuang, Jialian Wu, Ximeng Sun, Ze Wang 0008, Jiang Liu 0014, Yusheng Su, Jingbo Shang, Zicheng Liu 0001, Emad Barsoum |
ACL (1) | 9 |
| 2025 | SoftVQ-VAE: Efficient 1-Dimensional Continuous TokenizerabstractEfficient image tokenization with high compression ratios remains a critical challenge for training generative models. We present SoftVQ-VAE, a continuous image tokenizer that leverages soft categorical posteriors to aggregate multiple codewords into each latent token, substantially increasing the representation capacity of the latent space. When applied to Transformer-based architectures, our approach compresses 256×256 and 512×512 images using as few as 32 or 64 1-dimensional tokens. Not only does SoftVQ-VAE show consistent and high-quality reconstruction, more importantly, it also achieves state-of-the-art and significantly faster image generation results across different denoising-based generative models. Remarkably, SoftVQ-VAE improves inference throughput by up to 18x for generating 256×256 images and 55x for 512×512 images while achieving competitive FID scores of 1.78 and 2.21 for SiT-XL. It also improves the training efficiency of the generative models by reducing the number of training iterations by 2.3x while maintaining comparable performance. With its fully-differentiable design and semantic-rich latent space, our experiment demonstrates that SoftVQ-VAE achieves efficient tokenization without compromising generation quality, paving the way for more efficient generative models. Code and model are released1. Hao Chen 0102, Ze Wang 0008, Xiang Li 0106, Ximeng Sun, Fangyi Chen, Jiang Liu 0014, Jindong Wang 0001, Bhiksha Raj, Zicheng Liu 0001, Emad Barsoum |
CVPR | 9 |
| 2025 | TTT-Bench: A Benchmark for Evaluating Reasoning Ability with Simple and Novel Tic-Tac-Toe-style GamesabstractLarge reasoning models (LRMs) have demonstrated impressive reasoning capabilities across a broad range of tasks including Olympiadlevel mathematical problems, indicating evidence of their complex reasoning abilities.While many reasoning benchmarks focus on the STEM domain, the ability of LRMs to reason correctly in broader task domains remains underexplored.In this work, we introduce TTT-Bench, a new benchmark that is designed to evaluate basic strategic, spatial, and logical reasoning abilities in LRMs through a suite of four two-player Tic-Tac-Toe-style games that humans can effortlessly solve from a young age.We propose a simple yet scalable programmatic approach for generating verifiable two-player game problems for TTT-Bench.Although these games are trivial for humans, they require reasoning about the intentions of the opponent, as well as the game board's spatial configurations, to ensure a win.We evaluate a diverse set of state-of-the-art LRMs, and discover that the models that excel at hard math problems frequently fail at these simple reasoning games.Further testing reveals that our evaluated reasoning models score on average ↓ 41% & ↓ 5% lower on TTT-Bench compared to MATH 500 & AIME 2024 respectively, with larger models achieving higher performance using shorter reasoning traces, where most of the models struggle on long-term strategic reasoning situations on simple and new TTT-Bench tasks. Prakamya Mishra, Jiang Liu 0014, Jialian Wu, Zicheng Liu 0001, Emad Barsoum |
EMNLP | 5 |
| 2025 | Tuning Timestep-Distilled Diffusion Model Using Pairwise Sample OptimizationabstractRecent advancements in timestep-distilled diffusion models have enabled high-quality image generation that rivals non-distilled multi-step models, but with significantly fewer inference steps. While such models are attractive for applications due to the low inference cost and latency, fine-tuning them with a naive diffusion objective would result in degraded and blurry outputs. An intuitive alternative is to repeat the diffusion distillation process with a fine-tuned teacher model, which produces good results but is cumbersome and computationally intensive: the distillation training usually requires magnitude higher of training compute compared to fine-tuning for specific image styles. In this paper, we present an algorithm named pairwise sample optimization (PSO), which enables the direct fine-tuning of an arbitrary timestep-distilled diffusion model. PSO introduces additional reference images sampled from the current time-step distilled model, and increases the relative likelihood margin between the training images and reference images. This enables the model to retain its few-step generation ability, while allowing for fine-tuning of its output distribution. We also demonstrate that PSO is a generalized formulation which be flexible extended to both offline-sampled and online-sampled pairwise data, covering various popular objectives for diffusion model preference optimization. We evaluate PSO in both preference optimization and other fine-tuning tasks, including style transfer and concept customization. We show that PSO can directly adapt distilled models to human-preferred generation with both offline and online-generated pairwise preference image data. PSO also demonstrates effectiveness in style transfer and concept customization by directly tuning timestep-distilled diffusion models. Zichen Miao, Zhengyuan Yang, Ze Wang 0008, Zicheng Liu 0001, Qiang Qiu 0001 |
ICLR | 5 |
| 2025 | Masked Autoencoders Are Effective Tokenizers for Diffusion ModelsabstractRecent advances in latent diffusion models have demonstrated their effectiveness for high-resolution image synthesis. However, the properties of the latent space from tokenizer for better learning and generation of diffusion models remain under-explored. Theoretically and empirically, we find that improved generation quality is closely tied to the latent distributions with better structure, such as the ones with fewer Gaussian Mixture modes and more discriminative features. Motivated by these insights, we propose MAETok, an autoencoder (AE) leveraging mask modeling to learn semantically rich latent space while maintaining reconstruction fidelity.
Extensive experiments validate our analysis, demonstrating that the variational form of autoencoders is not necessary, and a discriminative latent space from AE alone enables state-of-the-art performance on ImageNet generation using only 128 tokens. MAETok achieves significant practical improvements, enabling a gFID of 1.69 with 76× faster training and 31× higher inference throughput for 512×512 generation. Our findings show that the structure of the latent space, rather than variational constraints, is crucial for effective diffusion models. Code and trained models will be released. Hao Chen 0102, Yujin Han, Fangyi Chen, Xiang Li 0106, Yidong Wang 0003, Jindong Wang 0001, Ze Wang 0008, Zicheng Liu 0001, Difan Zou, Bhiksha Raj |
ICML | 8 |
| 2025 | Exploring Invariance in Images through One-way Wave EquationsabstractIn this paper, we empirically demonstrate that natural images can be reconstructed with high fidelity from compressed representations using a simple first-order norm-plus-linear autoregressive (FINOLA) process—without relying on explicit positional information. Through systematic analysis, we observe that the learned coefficient matrices ($\mathbf{A}$ and $\mathbf{B}$) in FINOLA are typically invertible, and their product, $\mathbf{AB}^{-1}$, is diagonalizable across training runs. This structure enables a striking interpretation: FINOLA’s latent dynamics resemble a system of one-way wave equations evolving in a compressed latent space. Under this framework, each image corresponds to a unique solution of these equations. This offers a new perspective on image invariance, suggesting that the underlying structure of images may be governed by simple, invariant dynamic laws. Our findings shed light on a novel avenue for understanding and modeling visual data through the lens of latent-space dynamics and wave propagation. Yinpeng Chen, Dongdong Chen 0001, Xiyang Dai, Mengchen Liu, Yinan Feng, Youzuo Lin, Lu Yuan 0001, Zicheng Liu 0001 |
ICML | 8 |
| 2025 | Unleashing Hour-Scale Video Training for Long Video-Language UnderstandingabstractRecent long-form video-language understanding benchmarks have driven progress in video large multimodal models (Video-LMMs). However, the scarcity of well-annotated long videos has left the training of hour-long Video-LMMs underexplored. To close this gap, we present VideoMarathon, a large-scale hour-long video instruction-following dataset. This dataset includes around 9,700 hours of long videos sourced from diverse domains, ranging from 3 to 60 minutes per video. Specifically, it contains 3.3M high-quality QA pairs, spanning six fundamental topics: temporality, spatiality, object, action, scene, and event. Compared to existing video instruction datasets, VideoMarathon significantly extends training video durations up to 1 hour, and supports 22 diverse tasks requiring both short- and long-term video comprehension. Building on VideoMarathon, we propose Hour-LLaVA, a powerful and efficient Video-LMM for hour-scale video-language modeling. It enables hour-long video training and inference at 1-FPS sampling by leveraging a memory augmentation module, which adaptively integrates question-relevant and spatiotemporally informative semantics from the cached full video context. In our experiments, Hour-LLaVA achieves the best performance on multiple representative long video-language benchmarks, demonstrating the high quality of the VideoMarathon dataset and the superiority of the Hour-LLaVA model. Jialian Wu, Ximeng Sun, Ze Wang 0008, Jiang Liu 0014, Yusheng Su, Hao Chen 0102, Jiebo Luo 0001, Zicheng Liu 0001, Emad Barsoum |
NeurIPS | 10 |
| 2025 | MPS-NeRF: Generalizable 3D Human Rendering From Multiview ImagesabstractThere has been rapid progress recently on 3D human rendering, including novel view synthesis and pose animation, based on the advances of neural radiance fields (NeRF). However, most existing methods focus on person-specific training and their training typically requires multi-view videos. This article deals with a new challenging task - rendering novel views and novel poses for a person unseen in training, using only multiview still images as input without videos. For this task, we propose a simple yet surprisingly effective method to train a generalizable NeRF with multiview images as conditional input. The key ingredient is a dedicated representation combining a canonical NeRF and a volume deformation scheme. Using a canonical space enables our method to learn shared properties of human and easily generalize to different people. Volume deformation is used to connect the canonical space with input and target images and query image features for radiance and density prediction. We leverage the parametric 3D human model fitted on the input images to derive the deformation, which works quite well in practice when combined with our canonical NeRF. The experiments on both real and synthetic data with the novel view synthesis and pose animation tasks collectively demonstrate the efficacy of our method. Xiangjun Gao, Jiaolong Yang, Jongyoo Kim, Sida Peng, Zicheng Liu 0001, Xin Tong 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | ORES: Open-Vocabulary Responsible Visual SynthesisabstractAvoiding synthesizing specific visual concepts is an essential challenge in responsible visual synthesis. However, the visual concept that needs to be avoided for responsible visual synthesis tends to be diverse, depending on the region, context, and usage scenarios. In this work, we formalize a new task, Open-vocabulary Responsible Visual Synthesis (ORES), where the synthesis model is able to avoid forbidden visual concepts while allowing users to input any desired content. To address this problem, we present a Two-stage Intervention (TIN) framework. By introducing 1) rewriting with learnable instruction through a large-scale language model (LLM) and 2) synthesizing with prompt intervention on a diffusion synthesis model, it can effectively synthesize images avoiding any concepts but following the user's query as much as possible. To evaluate on ORES, we provide a publicly available dataset, baseline models, and benchmark. Experimental results demonstrate the effectiveness of our method in reducing risks of image generation. Our work highlights the potential of LLMs in responsible visual synthesis. Our code and dataset is public available in https://github.com/kodenii/ORES. Minheng Ni, Chenfei Wu, Xiaodong Wang 0023, Shengming Yin, Zicheng Liu 0001, Nan Duan 0001 |
AAAI | 6 |
| 2024 | Segment and Caption AnythingabstractWe propose a method to efficiently equip the Segment Anything Model (SAM) with the ability to generate regional captions. SAM presents strong generalizability to segment anything while is short for semantic understanding. By introducing a lightweight query-based feature mixer, we align the region-specific features with the embedding space of language models for later caption generation. As the number of trainable parameters is small (typically in the order of tens of millions), it costs less computation, less memory usage, and less communication bandwidth, resulting in both fast and scalable training. To address the scarcity problem of regional caption data, we propose to first pretrain our model on objection detection and segmentation tasks. We call this step weak supervision pretraining since the pretraining data only contains category names instead of full-sentence descriptions. The weak supervision pretraining al-lows us to leverage many publicly available object detection and segmentation datasets. We conduct extensive experiments to demonstrate the superiority of our method and validate each design choice. This work serves as a step-ping stone towards scaling up regional captioning data and sheds light on exploring efficient ways to augment SAM with regional semantics. The project page, along with the associated code, can be accessed via the following link. Xiaoke Huang 0001, Yansong Tang, Zheng Zhang 0022, Han Hu 0001, Jiwen Lu, Zicheng Liu 0001 |
CVPR | 8 |
| 2024 | Training Diffusion Models Towards Diverse Image Generation with Reinforcement LearningabstractDiffusion models have demonstrated unprecedented capabilities in image generation. Yet, they incorporate and amplify the data bias (e.g., gender, age) from the original training set, limiting the diversity of generated images. In this paper, we propose a diversity-oriented fine-tuning method using reinforcement learning (RL) for diffusion models under the guidance of an image-set-based reward function. Specifically, the proposed reward function, denoted as Diversity Reward, utilizes a set of generated images to evaluate the coverage of the current generative distribution w.r.t. the reference distribution, represented by a set of unbiased images. Built on top of the probabilistic method of distribution discrepancy estimation, Diversity Reward can measure the relative distribution gap with a small set of images efficiently. We further formulate the diffusion process as a multi-step decision-making problem (MDP) and apply policy gradient methods to fine-tune diffusion models by maximizing the Diversity Reward. The proposed rewards are validated on a post-sampling selection task, where a subset of the most diverse images are selected based on Diversity Reward values. We also show the effectiveness of our RL fine-tuning framework on enhancing the diversity of image generation with different types of diffusion models, including class-conditional models and text-conditional models, e.g., StableDiffusion. Zichen Miao, Jiang Wang 0012, Ze Wang 0008, Zhengyuan Yang, Qiang Qiu 0001, Zicheng Liu 0001 |
CVPR | 7 |
| 2024 | Disco: Disentangled Control for Realistic Human Dance GenerationabstractGenerative AI has made significant strides in computer vision, particularly in text-driven image/video synthesis (T2I/T2V). Despite the notable advancements, it remains challenging in human-centric content synthesis such as realistic dance generation. Current methodologies, primarily tailored for human motion transfer, encounter difficulties when confronted with real-world dance scenarios (e.g., social media dance), which require to generalize across a wide spectrum of poses and intricate human details. In this paper, we depart from the traditional paradigm of human motion transfer and emphasize two additional critical attributes for the synthesis of human dance content in social media contexts: (i) Generalizability: the model should be able to generalize beyond generic human viewpoints as well as unseen human subjects, backgrounds, and poses; (ii) Compositionality: it should allow for the seamless composition of seen/unseen subjects, backgrounds, and poses from different sources. To address these challenges, we introduce Disco, which includes a novel model architecture with disentangled control to improve the compositionality of dance synthesis, and an effective human attribute pre-training for better generalizability to unseen humans. Extensive qualitative and quantitative results demonstrate that DISCO can generate high-quality human dance images and videos with diverse appearances and flexible motions. Code is available at https://disco-dance.github.io/. Yuanhao Zhai 0001, Chung-Ching Lin, Zhengyuan Yang, Hanwang Zhang, Zicheng Liu 0001 |
CVPR | 8 |
| 2024 | MM-Narrator: Narrating Long-form Videos with Multimodal In-Context LearningabstractWe present MM-Narrator, a novel system leveraging GPT-4 with multimodal in-context learning for the gener-ation of audio descriptions (AD). Unlike previous methods that primarily focused on downstream fine-tuning with short video clips, MM-Narrator excels in generating precise audio descriptions for videos of extensive lengths, even be-yond hours, in an autoregressive manner. This capability is made possible by the proposed memory-augmented generation process, which effectively utilizes both the short-term textual context and long-term visual memory through an efficient register-and-recall mechanism. These contextual memories compile pertinent past information, including storylines and character identities, ensuring an accurate tracking and depicting of story-coherent and character-centric audio descriptions. Maintaining the training-free design of MM-Narrator, we further propose a complexity-based demonstration selection strategy to largely enhance its multi-step reasoning capability via few-shot multimodal in-context learning (MM-ICL). Experimental results on MAD-eval dataset demonstrate that MM-Narrator consistently outperforms both the existing fine-tuning-based approaches and LLM-based approaches in most scenarios, as measured by standard evaluation metrics. Additionally, we introduce the first segment-based evaluator for recurrent text generation. Empowered by GPT-4, this evaluator comprehensively reasons and marks AD generation performance in various extendable dimensions. Chaoyi Zhang, Zhengyuan Yang, Chung-Ching Lin, Zicheng Liu 0001 |
CVPR | 7 |
| 2024 | GRiT: A Generative Region-to-Text Transformer for Object Understanding
Jialian Wu, Zhengyuan Yang, Zhe Gan, Zicheng Liu 0001, Junsong Yuan 0001 |
ECCV (80) | 5 |
| 2024 | Idea2Img: Iterative Self-refinement with GPT-4V for Automatic Image Design and Generation
Zhengyuan Yang, Chung-Ching Lin, Zicheng Liu 0001 |
ECCV (38) | 6 |
| 2024 | IDOL: Unified Dual-Modal Latent Diffusion for Human-Centric Joint Video-Depth Generation
Yuanhao Zhai 0001, Chung-Ching Lin, Zhengyuan Yang, David S. Doermann, Junsong Yuan 0001, Zicheng Liu 0001 |
ECCV (15) | 9 |
| 2024 | Completing Visual Objects via Bridging Generation and SegmentationabstractThis paper presents a novel approach to object completion, with the primary goal of reconstructing a complete object from its partially visible components. Our method, named MaskComp, delineates the completion process through iterative stages of generation and segmentation. In each iteration, the object mask is provided as an additional condition to boost image generation, and, in return, the generated images can lead to a more accurate mask by fusing the segmentation of images. We demonstrate that the combination of one generation and one segmentation stage effectively functions as a mask denoiser. Through alternation between the generation and segmentation stages, the partial object mask is progressively refined, providing precise shape guidance and yielding superior object completion results. Our experiments demonstrate the superiority of MaskComp over existing approaches, e.g., ControlNet and Stable Diffusion, establishing it as an effective solution for object completion. Xiang Li 0106, Yinpeng Chen, Chung-Ching Lin, Hao Chen 0102, Kai Hu 0010, Rita Singh, Bhiksha Raj, Zicheng Liu 0001 |
ICML | 9 |
| 2024 | StrokeNUWA - Tokenizing Strokes for Vector Graphic SynthesisabstractTo leverage LLMs for visual synthesis, traditional methods convert raster image information into discrete grid tokens through specialized visual modules, while disrupting the model’s ability to capture the true semantic representation of visual scenes. This paper posits that an alternative representation of images, vector graphics, can effectively surmount this limitation by enabling a more natural and semantically coherent segmentation of the image information. Thus, we introduce StrokeNUWA, a pioneering work exploring a better visual representation "stroke" tokens on vector graphics, which is inherently visual semantics rich, naturally compatible with LLMs, and highly compressed. Equipped with stroke tokens, StrokeNUWA can significantly surpass traditional LLM-based and optimization-based methods across various metrics in the vector graphic generation task. Besides, StrokeNUWA achieves up to a $94\times$ speedup in inference over the speed of prior methods with an exceptional SVG code compression ratio of 6.9%. Zecheng Tang, Chenfei Wu, Minheng Ni, Shengming Yin, Zhengyuan Yang, Zicheng Liu 0001, Nan Duan 0001 |
ICML | 9 |
| 2024 | MM-Vet: Evaluating Large Multimodal Models for Integrated CapabilitiesabstractWe propose MM-Vet, an evaluation benchmark that examines large multimodal models (LMMs) on complicated multimodal tasks. Recent LMMs have shown various intriguing abilities, such as solving math problems written on the blackboard, reasoning about events and celebrities in news images, and explaining visual jokes. Rapid model advancements pose challenges to evaluation benchmark development. Problems include: (1) How to systematically structure and evaluate the complicated multimodal tasks; (2) How to design evaluation metrics that work well across question and answer types; and (3) How to give model insights beyond a simple performance ranking. To this end, we present MM-Vet, designed based on the insight that the intriguing ability to solve complicated tasks is often achieved by a generalist model being able to integrate different core vision-language (VL) capabilities. MM-Vet defines 6 core VL capabilities and examines the 16 integrations of interest derived from the capability combination. For evaluation metrics, we propose an LLM-based evaluator for open-ended outputs. The evaluator enables the evaluation across different question types and answer styles, resulting in a unified scoring metric. We evaluate representative LMMs on MM-Vet, providing insights into the capabilities of different LMM system paradigms and models. Weihao Yu 0001, Zhengyuan Yang, Zicheng Liu 0001, Xinchao Wang |
ICML | 6 |
| 2024 | Bring Metric Functions into Diffusion Models
Jie An 0002, Zhengyuan Yang, Zicheng Liu 0001, Jiebo Luo 0001 |
IJCAI | 5 |
| 2024 | OpenLEAF: A Novel Benchmark for Open-Domain Interleaved Image-Text Generation
Jie An 0002, Zhengyuan Yang, Zicheng Liu 0001, Jiebo Luo 0001 |
ACM Multimedia | 6 |
| 2024 | Taming Diffusion Prior for Image Super-Resolution with Domain Shift SDEsabstractDiffusion-based image super-resolution (SR) models have attracted substantial interest due to their powerful image restoration capabilities. However, prevailing diffusion models often struggle to strike an optimal balance between efficiency and performance. Typically, they either neglect to exploit the potential of existing extensive pretrained models, limiting their generative capacity, or they necessitate a dozens of forward passes starting from random noises, compromising inference efficiency. In this paper, we present DoSSR, a $\textbf{Do}$main $\textbf{S}$hift diffusion-based SR model that capitalizes on the generative powers of pretrained diffusion models while significantly enhancing efficiency by initiating the diffusion process with low-resolution (LR) images. At the core of our approach is a domain shift equation that integrates seamlessly with existing diffusion models. This integration not only improves the use of diffusion prior but also boosts inference efficiency. Moreover, we advance our method by transitioning the discrete shift process to a continuous formulation, termed as DoS-SDEs. This advancement leads to the fast and customized solvers that further enhance sampling efficiency. Empirical results demonstrate that our proposed method achieves state-of-the-art performance on synthetic and real-world datasets, while notably requiring $\textbf{\emph{only 5 sampling steps}}$. Compared to previous diffusion prior based methods, our approach achieves a remarkable speedup of 5-7 times, demonstrating its superior efficiency. Qinpeng Cui, Yixuan Liu 0004, Xinyi Zhang 0008, Qiqi Bao 0001, Qingmin Liao, liwang Amd, Zicheng Liu 0001, Zhongdao Wang, Emad Barsoum |
NeurIPS | 8 |
| 2024 | MPT: Mesh Pre-Training with Transformers for Human Pose and Mesh ReconstructionabstractTraditional methods of reconstructing 3D human pose and mesh from single images rely on paired image-mesh datasets, which can be difficult and expensive to obtain. Due to this limitation, model scalability is constrained as well as reconstruction performance. Towards addressing the challenge, we introduce Mesh Pre-Training (MPT), an effective pre-training strategy that leverages large amounts of MoCap data to effectively perform pre-training at scale. We introduce the use of MoCap-generated heatmaps as input representations to the mesh regression transformer and propose a Masked Heatmap Modeling approach for improving pre-training performance. This study demonstrates that pre-training using the proposed MPT allows our models to perform effective inference without requiring fine-tuning. We further show that fine-tuning the pre-trained MPT model considerably improves the accuracy of human mesh reconstruction from single images. Experimental results show that MPT outperforms previous state-of-the-art methods on Human3.6M and 3DPW datasets. As a further application, we benchmark and study MPT on the task of 3D hand reconstruction, showing that our generic pre-training scheme generalizes well to hand pose estimation and achieves promising reconstruction performance. Chung-Ching Lin, Zicheng Liu 0001 |
WACV | 4 |
| 2023 | NUWA-XL: Diffusion over Diffusion for eXtremely Long Video GenerationabstractShengming Yin, Chenfei Wu, Huan Yang, Jianfeng Wang, Xiaodong Wang, Minheng Ni, Zhengyuan Yang, Linjie Li, Shuguang Liu, Fan Yang, Jianlong Fu, Ming Gong, Lijuan Wang, Zicheng Liu, Houqiang Li, Nan Duan. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Shengming Yin, Chenfei Wu, Huan Yang 0005, Xiaodong Wang 0023, Minheng Ni, Zhengyuan Yang, Fan Yang 0024, Jianlong Fu, Ming Gong 0001, Zicheng Liu 0001, Houqiang Li, Nan Duan 0001 |
ACL (1) | 14 |
| 2023 | An Empirical Study of End-to-End Video-Language Transformers with Masked Visual ModelingabstractMasked visual modeling (MVM) has been recently proven effective for visual pre-training. While similar reconstructive objectives on video inputs (e.g., masked frame modeling) have been explored in video-language (VidL) pre-training, previous studies fail to find a truly effective MVM strategy that can largely benefit the downstream performance. In this work, we systematically examine the potential of MVM in the context of VidL learning. Specifically, we base our study on a fully end-to-end VIdeO-LanguagE Transformer (VIOLET) [15], where the supervision from MVM training can be backpropogated to the video pixel space. In total, eight different reconstructive targets of MVM are explored, from low-level pixel values and oriented gradients to high-level depth maps, optical flow, discrete visual tokens and latent visual features. We conduct comprehensive experiments and provide insights into the factors leading to effective MVM training, resulting in an enhanced model VIOLETv2. Empirically, we show VIOLETv2 pre-trained with MVM objective achieves notable improvements on 13 VidL benchmarks, ranging from video question answering, video captioning, to text-to-video retrieval.11Code has been released at https://github.com/tsujuifu/pytorch_empirical-mvm Tsu-Jui Fu, Zhe Gan, William Yang Wang, Zicheng Liu 0001 |
CVPR | 7 |
| 2023 | Neural Voting Field for Camera-Space 3D Hand Pose EstimationabstractWe present a unified framework for camera-space 3D hand pose estimation from a single RGB image based on 3D implicit representation. As opposed to recent works, most of which first adopt holistic or pixel-level dense regression to obtain relative 3D hand pose and then follow with complex second-stage operations for 3D global root or scale recovery, we propose a novel unified 3D dense regression scheme to estimate camera-space 3D hand pose via dense 3D point-wise voting in camera frustum. Through direct dense modeling in 3D domain inspired by Pixel-aligned Implicit Functions for 3D detailed reconstruction, our proposed Neural Voting Field (NVF) fully models 3D dense local evidence and hand global geometry, helping to alleviate common 2D-to-3D ambiguities. Specifically, for a 3D query point in camera frustum and its pixel-aligned image feature, NVF, represented by a Multi-Layer Perceptron, regresses: (i) its signed distance to the hand surface; (ii) a set of 4D offset vectors (1D voting weight and 3D directional vector to each hand joint). Following a vote-casting scheme, 4D offset vectors from near-surface points are selected to calculate the 3D hand joint coordinates by a weighted average. Experiments demonstrate that NVF outperforms existing state-of-the-art algorithms on FreiHAND dataset for camera-space 3D hand pose estimation. We also adapt NVF to the classic task of root-relative 3D hand pose estimation, for which NVF also obtains state-of-the-art results on HO3D dataset. Lin Huang 0004, Chung-Ching Lin, Junsong Yuan 0001, Zicheng Liu 0001 |
CVPR | 7 |
| 2023 | LAVENDER: Unifying Video-Language Understanding as Masked Language ModelingabstractUnified vision-language frameworks have greatly advanced in recent years, most of which adopt an encoder-decoder architecture to unify image-text tasks as sequence-to-sequence generation. However, existing video-language (VidL) models still require task-specific designs in model architecture and training objectives for each task. In this work, we explore a unified VidL framework LAVENDER, where Masked Language Modeling [13] (MLM) is used as the common interface for all pre-training and downstream tasks. Such unification leads to a simplified model architecture, where only a lightweight MLM head, instead of a decoder with much more parameters, is needed on top of the multimodal encoder. Surprisingly, experimental results show that this unified framework achieves competitive performance on 14 VidL benchmarks, covering video question answering, text-to-video retrieval and video captioning. Extensive analyses further demonstrate Lavender can (i) seamlessly support all downstream tasks with just a single set of parameter values when multi-task fine-tuned; (ii) generalize to various downstream tasks with limited training samples; and (iii) enable zero-shot evaluation on video question answering tasks. Code is available at https://github.com/microsoft/LAVENDER. Zhe Gan, Chung-Ching Lin, Zicheng Liu 0001, Ce Liu 0001 |
CVPR | 5 |
| 2023 | Deep Frequency Filtering for Domain GeneralizationabstractImproving the generalization ability of Deep Neural Networks (DNNs) is critical for their practical uses, which has been a longstanding challenge. Some theoretical studies have uncovered that DNNs have preferences for some frequency components in the learning process and indicated that this may affect the robustness of learned features. In this paper, we propose Deep Frequency Filtering (DFF)for learning domain-generalizable features, which is the first endeavour to explicitly modulate the frequency components of different transfer difficulties across domains in the latent space during training. To achieve this, we perform Fast Fourier Transform (FFT) for the feature maps at different layers, then adopt a light-weight module to learn attention masks from the frequency representations after FFT to enhance transferable components while suppressing the components not conducive to generalization. Further, we empirically compare the effectiveness of adopting different types of attention designs for implementing DFF. Extensive experiments demonstrate the effectiveness of our proposed DFF and show that applying our DFF on a plain baseline out-performs the state-of-the-art methods on different domain generalization tasks, including close-set classification and open-set retrieval. Shiqi Lin, Zhizheng Zhang 0004, Zhipeng Huang 0014, Yan Lu 0001, Cuiling Lan, Peng Chu, Quanzeng You, Jiang Wang 0012, Zicheng Liu 0001, Amey Parulkar, Viraj Navkal, Zhibo Chen 0001 |
CVPR | 9 |
| 2023 | Adaptive Human Matting for Dynamic VideosabstractThe most recent efforts in video matting have focused on eliminating trimap dependency since trimap annotations are expensive and trimap-based methods are less adaptable for real-time applications. Despite the latest tripmapfree methods showing promising results, their performance often degrades when dealing with highly diverse and unstructured videos. We address this limitation by introducing Adaptive Matting for Dynamic Videos, termed AdaM, which is a framework designed for simultaneously differentiating foregrounds from backgrounds and capturing alpha matte details of human subjects in the foreground. Two interconnected network designs are employed to achieve this goal: (1) an encoder-decoder network that produces alpha mattes and intermediate masks which are used to guide the transformer in adaptively decoding foregrounds and backgrounds, and (2) a transformer network in which long- and short-term attention combine to retain spatial and temporal contexts, facilitating the decoding of foreground details. We benchmark and study our methods on recently introduced datasets, showing that our model notably improves matting realism and temporal coherence in complex real-world videos and achieves new best-in-class generalizability. Further details and examples are available at https://github.com/microsoft/AdaM. Chung-Ching Lin, Jiang Wang 0012, Zicheng Liu 0001 |
CVPR | 7 |
| 2023 | Binary Latent DiffusionabstractIn this paper, we show that a binary latent space can be explored for compact yet expressive image representations. We model the bi-directional mappings between an image and the corresponding latent binary representation by training an auto-encoder with a Bernoulli encoding distribution. On the one hand, the binary latent space provides a compact discrete image representation of which the distribution can be modeled more efficiently than pixels or continuous latent representations. On the other hand, we now represent each image patch as a binary vector instead of an index of a learned cookbook as in discrete image representations with vector quantization. In this way, we obtain binary latent representations that allow for better image quality and high-resolution image representations without any multi-stage hierarchy in the latent space. In this binary latent space, images can now be generated effectively using a binary latent diffusion model tailored specifically for modeling the prior over the binary image representations. We present both conditional and unconditional image generation experiments with multiple datasets, and show that the proposed method performs comparably to state-of-the-art methods while dramatically improving the sampling efficiency to as few as 16 steps without using any test-time acceleration. The proposed framework can also be seamlessly scaled to 1024 x 1024 high-resolution image generation without resorting to latent hierarchy or multi-stage refinements. Ze Wang 0008, Jiang Wang 0012, Zicheng Liu 0001, Qiang Qiu 0001 |
CVPR | 3 |
| 2023 | ReCo: Region-Controlled Text-to-Image GenerationabstractRecently, large-scale text-to-image (T2I) models have shown impressive performance in generating high-fidelity images, but with limited controllability, e.g., precisely specifying the content in a specific region with a free-form text description. In this paper, we propose an effective technique for such regional control in T2I generation. We augment T2I models' inputs with an extra set of position tokens, which represent the quantized spatial coordinates. Each region is specified by four position tokens to represent the top-left and bottom-right corners, followed by an open-ended natural language regional description. Then, we fine-tune a pre-trained T2I model with such new input interface. Our model, dubbed as ReCo (Region-Controlled T2I), enables the region control for arbitrary objects described by open-ended regional texts rather than by object labels from a constrained category set. Empirically, ReCo achieves better image quality than the T2I model strengthened by positional words (FID: 8.82 → 7.36, SceneFID: 15.54 → 6.51 on COCO), together with objects being more accurately placed, amounting to a 20.40% region classification accuracy improvement on COCO. Furthermore, we demonstrate that ReCo can better control the object count, spatial relationship, and region attributes such as color/size, with the free-form regional description. Human evaluation on PaintSkill shows that ReCo is +19.28% and +17.21% more accurate in generating images with correct object count and spatial relationship than the T2I model. Code is available at https://github.com/microsoft/Reeo. Zhengyuan Yang, Zhe Gan, Chenfei Wu, Nan Duan 0001, Zicheng Liu 0001, Ce Liu 0001, Michael Zeng 0001 |
CVPR | 8 |
| 2023 | Equivariant Similarity for Vision-Language Foundation ModelsabstractThis study explores the concept of equivariance in vision-language foundation models (VLMs), focusing specifically on the multimodal similarity function that is not only the major training objective but also the core delivery to support downstream tasks. Unlike the existing image-text similarity objective which only categorizes matched pairs as similar and unmatched pairs as dissimilar, equivariance also requires similarity to vary faithfully according to the semantic changes. This allows VLMs to generalize better to nuanced and unseen multimodal compositions. However, modeling equivariance is challenging as the ground truth of semantic change is difficult to collect. For example, given an image-text pair about a dog, it is unclear to what extent the similarity changes when the pixel is changed from dog to cat? To this end, we propose EqSim, a regularization loss that can be efficiently calculated from any two matched training pairs and easily pluggable into existing image-text retrieval fine-tuning. Meanwhile, to further diagnose the equivariance of VLMs, we present a new challenging benchmark EqBen. Compared to the existing evaluation sets, EqBen is the first to focus on "visual-minimal change". Extensive experiments show the lack of equivariance in current VLMs1and validate the effectiveness of EqSim2. Chung-Ching Lin, Zhengyuan Yang, Hanwang Zhang, Zicheng Liu 0001 |
ICCV | 7 |
| 2023 | Zero-Shot Human-Object Interaction (HOI) Classification by Bridging Generative and Contrastive Image-Language ModelsabstractExisting studies in Human-Object Interaction (HOI) classification rely on costly human-annotated labels. The goal of this paper is to study a new zero-shot setup to remove the dependency on ground-truth labels. We propose a novel Heterogenous Teacher-Student (HTS) framework and a new loss function. HTS employs a generative pretrained image captioner as the teacher and a contrastive pre-trained classifier as the student. HTS combines the discriminability from generative pre-training and efficiency from contrastive pre-training. To facilitate learning of HOI in this setup, we introduce pseudo-label filtering which aggregates HOI probabilities from multiple regional captions to supervise the student. To enhance the multi-label learning of the student on few-shot classes, we design LogSumExp (LSE)-Sign loss which features a dynamic gradient re-weighting mechanism. Eventually, the student achieves 49.6 mAP on the HICO dataset without using ground truth, becoming a new state-of-the-art method that outperforms supervised approaches. Code is available. Yinpeng Chen, Jenq-Neng Hwang, Zicheng Liu 0001 |
ICIP | 6 |
| 2023 | Layer Grafted Pre-training: Bridging Contrastive Learning And Masked Image Modeling For Label-Efficient Representations
Ziyu Jiang, Yinpeng Chen, Mengchen Liu, Dongdong Chen 0001, Xiyang Dai, Lu Yuan 0001, Zicheng Liu 0001, Zhangyang Wang |
ICLR | 7 |
| 2023 | Energy-Inspired Self-Supervised Pretraining for Vision Models
Ze Wang 0008, Jiang Wang 0012, Zicheng Liu 0001, Qiang Qiu 0001 |
ICLR | 3 |
| 2023 | Learning 3D Photography Videos via Self-supervised Diffusion on Single Imagesabstract3D photography renders a static image into a video with appealing 3D visual effects. Existing approaches typically first conduct monocular depth estimation, then render the input frame to subsequent frames with various viewpoints, and finally use an inpainting model to fill those missing/occluded regions. The inpainting model plays a crucial role in rendering quality, but it is normally trained on out-of-domain data. To reduce the training and inference gap, we propose a novel self-supervised diffusion model as the inpainting module. Given a single input image, we automatically construct a training pair of the masked occluded image and the ground-truth image with random cycle rendering. The constructed training samples are closely aligned to the testing instances, without the need for data annotation. To make full use of the masked images, we designed a Masked Enhanced Block (MEB), which can be easily plugged into the UNet and enhance the semantic conditions. Towards real-world animation, we present a novel task: out-animation, which extends the space and time of input objects. Extensive experiments on real datasets show that our method achieves competitive results with existing SOTA methods. Xiaodong Wang 0023, Chenfei Wu, Shengming Yin, Minheng Ni, Zhengyuan Yang, Fan Yang 0024, Zicheng Liu 0001, Yuejian Fang, Nan Duan 0001 |
IJCAI | 10 |
| 2023 | PaintSeg: Painting Pixels for Training-free SegmentationabstractThe paper introduces PaintSeg, a new unsupervised method for segmenting objects without any training. We propose an adversarial masked contrastive painting (AMCP) process, which creates a contrast between the original image and a painted image in which a masked area is painted using off-the-shelf generative models. During the painting process, inpainting and outpainting are alternated, with the former masking the foreground and filling in the background, and the latter masking the background while recovering the missing part of the foreground object. Inpainting and outpainting, also referred to as I-step and O-step, allow our method to gradually advance the target segmentation mask toward the ground truth without supervision or training. PaintSeg can be configured to work with a variety of prompts, e.g. coarse masks, boxes, scribbles, and points. Our experimental results demonstrate that PaintSeg outperforms existing approaches in coarse mask-prompt, box-prompt, and point-prompt segmentation tasks, providing a training-free solution suitable for unsupervised segmentation. Code: https://github.com/lxa9867/PaintSeg. Xiang Li 0106, Chung-Ching Lin, Yinpeng Chen, Zicheng Liu 0001, Jinglu Wang, Rita Singh, Bhiksha Raj |
NeurIPS | 4 |
| 2023 | TransMOT: Spatial-Temporal Graph Transformer for Multiple Object TrackingabstractTracking multiple objects in videos relies on modeling the spatial-temporal interactions of the objects. In this paper, we propose TransMOT, which leverages powerful graph transformers to efficiently model the spatial and temporal interactions among the objects. TransMOT is capable of effectively modeling the interactions of a large number of objects by arranging the trajectories of the tracked targets and detection candidates as a set of sparse weighted graphs, and constructing a spatial graph transformer encoder layer, a temporal transformer encoder layer, and a spatial graph transformer decoder layer based on the graphs. Through end-to-end learning, TransMOT can exploit the spatial-temporal clues to directly estimate association from a large number of loosely filtered detection predictions for robust MOT in complex scenes. The proposed method is evaluated on multiple benchmark datasets, including MOT15, MOT16, MOT17, and MOT20, and it achieves state-of-the-art performance on all the datasets. Peng Chu, Jiang Wang 0012, Quanzeng You, Haibin Ling, Zicheng Liu 0001 |
WACV | 5 |
| 2023 | MMPTRACK: Large-scale Densely Annotated Multi-camera Multiple People Tracking BenchmarkabstractMulti-camera tracking systems are gaining popularity in applications that demand high-quality tracking results, such as frictionless checkout. In cluttered and crowded environments, monocular multi-object tracking (MOT) systems often fail due to occlusions. Multiple highly overlapped cameras are capable of recovering partial 3D information. When used properly, 3D data can significantly alleviate the occlusion issue. However, training a multi-camera tracker demands a large-scale multi-camera tracking dataset with diverse camera settings and backgrounds. These requirements make the collection of multi-camera tracking dataset challenging and expensive. The cost of creating such a dataset has limited the availability and scale of datasets in this domain. Instead, we appeal to an auto-annotation system to reduce the cost, which uses overlapped and calibrated depth and RGB cameras to build a 3D tracker and automatically generates the 3D tracking results. The results are manually checked and corrected to ensure the label quality, which is much cheaper than solely manual annotation. Next, the 3D tracking results are projected to each calibrated RGB camera view to create 2D tracking results. In this way, we collect and annotate a large-scale densely labeled multi-camera tracking dataset from five different environments. We have conducted extensive experiments using two real-time multi-camera trackers and a person re-identification (ReID) model under different settings. This dataset provides a reliable benchmark for multi-camera, multi-object tracking systems in cluttered and crowded environments. We expect this benchmark to encourage more research attempts in this domain. Our dataset will be publicly released upon the acceptance of this work. Quanzeng You, Chunyu Wang 0001, Zhizheng Zhang 0004, Peng Chu, Houdong Hu, Jiang Wang 0012, Zicheng Liu 0001 |
WACV | 8 |
| 2023 | MGL: Mutual Graph Learning for Camouflaged Object DetectionabstractCamouflaged object detection, which aims to detect/segment the object(s) that blend in with their surrounding, remains challenging for deep models due to the intrinsic similarities between foreground objects and background surroundings. Ideally, an effective model should be capable of finding valuable clues from the given scene and integrating them into a joint learning framework to co-enhance the representation. Inspired by this observation, we propose a novel Mutual Graph Learning (MGL) model by shifting the conventional perspective of mutual learning from regular grids to graph domain. Specifically, an image is decoupled by MGL into two task-specific feature maps - one for finding the rough location of the target and the other for capturing its accurate boundary details. Then, the mutual benefits can be fully exploited by reasoning their high-order relations through graphs recurrently. It should be noted that our method is different from most mutual learning models that model all between-task interactions with the use of a shared function. To increase information interactions, MGL is built with typed functions for dealing with different complementary relations. To overcome the accuracy loss caused by interpolation to higher resolution and the computational redundancy resulting from recurrent learning, the S-MGL is equipped with a multi-source attention contextual recovery module, called R-MGL_v2, which uses the pixel feature information iteratively. Experiments on challenging datasets, including CHAMELEON, CAMO, COD10K, and NC4K demonstrate the effectiveness of our MGL with superior performance to existing state-of-the-art methods. The code can be found at https://github.com/fanyang587/MGL. Qiang Zhai, Xin Li 0079, Fan Yang 0054, Zhicheng Jiao, Ping Luo 0002, Hong Cheng 0002, Zicheng Liu 0001 |
IEEE Trans. Image Process. | 7 |
| 2023 | Co-Communication Graph Convolutional Network for Multi-View Crowd CountingabstractWe study and address the multi-view crowd counting (MVCC) problem which poses more realistic challenges than single-view crowd counting for better facilitating crowd management/public safety systems. Its major challenge lies in how to fully distill and aggregate useful, complementary information among multiple camera views to create powerful ground-plane representations for wide-area crowd analysis. In this paper, we present a graph-based, multi-view learning model called Co-Communication Graph Convolutional Network (CoCo-GCN) to jointly investigate intra-view contextual dependencies and inter-view complementary relations. More specifically, CoCo-GCN builds a view-agnostic graph interaction space for each camera view to conduct efficient contextual reasoning, and extends the intra-view reasoning by using a novel Graph Communication Layer (GCL) to also take between-graph (cross-view), complementary information into account. Moreover, CoCo-GCN uses a new Co-Memory Layer (CoML) to jointly coarsen the graphs and close the ‘representational gap’ among them for further exploiting the compositional nature of graphs and learning more consistent representations. Finally, these jointly learned features of multiple views can be easily fused to create ground-plane representations for wide-area crowd counting. Experiments show that the proposed CoCo-GCN achieves state-of-the-art results on three MVCC datasets, i.e., PETS2009, DukeMTMC, and City Street, significantly improving the scene-level accuracy over previous models. Qiang Zhai, Fan Yang 0054, Xin Li 0079, Guosen Xie, Hong Cheng 0002, Zicheng Liu 0001 |
IEEE Trans. Multim. | 6 |
| 2022 | Playing Lottery Tickets with Vision and LanguageabstractLarge-scale pre-training has recently revolutionized vision-and-language (VL) research. Models such as LXMERT and UNITER have significantly lifted the state of the art over a wide range of VL tasks. However, the large number of parameters in such models hinders their application in practice. In parallel, work on the lottery ticket hypothesis (LTH) has shown that deep neural networks contain small matching subnetworks that can achieve on par or even better performance than the dense networks when trained in isolation. In this work, we perform the first empirical study to assess whether such trainable subnetworks also exist in pre-trained VL models. We use UNITER as the main testbed (also test on LXMERT and ViLT), and consolidate 7 representative VL tasks for experiments, including visual question answering, visual commonsense reasoning, visual entailment, referring expression comprehension, image-text retrieval, GQA, and NLVR2. Through comprehensive analysis, we summarize our main findings as follows. (i) It is difficult to find subnetworks that strictly match the performance of the full model. However, we can find relaxed winning tickets at 50%-70% sparsity that maintain 99% of the full accuracy. (ii) Subnetworks found by task-specific pruning transfer reasonably well to the other tasks, while those found on the pre-training tasks at 60%/70% sparsity transfer universally, matching 98%/96% of the full accuracy on average over all the tasks. (iii) Besides UNITER, other models such as LXMERT and ViLT can also play lottery tickets. However, the highest sparsity we can achieve for ViLT is far lower than LXMERT and UNITER (30% vs. 70%). (iv) LTH also remains relevant when using other training methods (e.g., adversarial training). Zhe Gan, Yen-Chun Chen 0001, Tianlong Chen 0001, Yu Cheng 0001, Shuohang Wang, Jingjing Liu 0001, Zicheng Liu 0001 |
AAAI | 9 |
| 2022 | OVIS: Open-Vocabulary Visual Instance Search via Visual-Semantic Aligned Representation LearningabstractWe introduce the task of open-vocabulary visual instance search (OVIS). Given an arbitrary textual search query, Open-vocabulary Visual Instance Search (OVIS) aims to return a ranked list of visual instances, i.e., image patches, that satisfies the search intent from an image database. The term ``open vocabulary'' means that there are neither restrictions to the visual instance to be searched nor restrictions to the word that can be used to compose the textual search query. We propose to address such a search challenge via visual-semantic aligned representation learning (ViSA). ViSA leverages massive image-caption pairs as weak image-level (not instance-level) supervision to learn a rich cross-modal semantic space where the representations of visual instances (not images) and those of textual queries are aligned, thus allowing us to measure the similarities between any visual instance and an arbitrary textual query. To evaluate the performance of ViSA, we build two datasets named OVIS40 and OVIS1600 and also introduce a pipeline for error analysis. Through extensive experiments on the two datasets, we demonstrate ViSA's ability to search for visual instances in images not available during training given a wide range of textual queries including those composed of uncommon words. Experimental results show that ViSA achieves an mAP@50 of 27.8% on OVIS40 and achieves a recall@30 of 21.3% on OVIS1400 dataset under the most challenging settings. Sheng Liu 0017, Junsong Yuan 0001, Zicheng Liu 0001 |
AAAI | 5 |
| 2022 | An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQAabstractKnowledge-based visual question answering (VQA) involves answering questions that require external knowledge not present in the image. Existing methods first retrieve knowledge from external resources, then reason over the selected knowledge, the input image, and question for answer prediction. However, this two-step approach could lead to mismatches that potentially limit the VQA performance. For example, the retrieved knowledge might be noisy and irrelevant to the question, and the re-embedded knowledge features during reasoning might deviate from their original meanings in the knowledge base (KB). To address this challenge, we propose PICa, a simple yet effective method that Prompts GPT3 via the use of Image Captions, for knowledge-based VQA. Inspired by GPT-3’s power in knowledge retrieval and question answering, instead of using structured KBs as in previous work, we treat GPT-3 as an implicit and unstructured KB that can jointly acquire and process relevant knowledge. Specifically, we first convert the image into captions (or tags) that GPT-3 can understand, then adapt GPT-3 to solve the VQA task in a few-shot manner by just providing a few in-context VQA examples. We further boost performance by carefully investigating: (i) what text formats best describe the image content, and (ii) how in-context examples can be better selected and used. PICa unlocks the first use of GPT-3 for multimodal tasks. By using only 16 examples, PICa surpasses the supervised state of the art by an absolute +8.6 points on the OK-VQA dataset. We also benchmark PICa on VQAv2, where PICa also shows a decent few-shot performance. Zhengyuan Yang, Zhe Gan, Xiaowei Hu 0006, Yumao Lu, Zicheng Liu 0001 |
AAAI | 6 |
| 2022 | Scaling Up Vision-Language Pretraining for Image CaptioningabstractIn recent years, we have witnessed significant performance boost in the image captioning task based on vision-language pre-training (VLP). Scale is believed to be an important factor for this advance. However, most existing work only focuses on pre-training transformers with moderate sizes (e.g., 12 or 24 layers) on roughly 4 million images. In this paper, we present LEMON O, a LargE-scale iMage captiONer, and provide the first empirical study on the scaling behavior of VLP for image captioning. We use the state-of-the-art Vin VL model as our reference model, which consists of an image feature extractor and a transformer model, and scale the transformer both up and down, with model sizes ranging from 13 to 675 million parameters. In terms of data, we conduct experiments with up to 200 million imagetext pairs which are automatically collected from web based on the alt attribute of the image (dubbed as ALT200M11The dataset is released at https://github.com/xiaoweihu/ALT200M). Extensive analysis helps to characterize the performance trend as the model size and the pre-training data size increase. We also compare different training recipes, especially for training on large-scale noisy data. As a result, LEMON achieves new state of the arts on several major image captioning benchmarks, including COCO Caption, nocaps, and Conceptual Captions. We also show LEMON can generate captions with long-tail vi-sual concepts when used in a zero-shot manner. Xiaowei Hu 0006, Zhe Gan, Zhengyuan Yang, Zicheng Liu 0001, Yumao Lu |
CVPR | 5 |
| 2022 | Lifelong Unsupervised Domain Adaptive Person Re-identification with Coordinated Anti-forgetting and AdaptationabstractUnsupervised domain adaptive person re-identification (ReID) has been extensively investigated to mitigate the adverse effects of domain gaps. Those works assume the target domain data can be accessible all at once. However, for the real-world streaming data, this hinders the timely adaptation to changing data statistics and sufficient exploitation of increasing samples. In this paper, to address more practical scenarios, we propose a new task, Lifelong Un-supervised Domain Adaptive (LUDA) person ReID. This is challenging because it requires the model to continuously adapt to unlabeled data in the target environments while alleviating catastrophic forgetting for such a fine-grained person retrieval task. We design an effective scheme for this task, dubbed CLUDA-ReID, where the anti-forgetting is harmoniously coordinated with the adaptation. Specifically, a meta-based Coordinated Data Replay strategy is proposed to replay old data and update the network with a coordinated optimization direction for both adaptation and memorization. Moreover, we propose Relational Consistency Learning for old knowledge distillation/inheritance in line with the objective of retrieval-based tasks. We set up two evaluation settings to simulate the practical application scenarios. Extensive experiments demonstrate the effectiveness of our CLUDA-ReID for both scenarios with stationary target streams and scenarios with dynamic target streams. Zhipeng Huang 0014, Zhizheng Zhang 0004, Cuiling Lan, Wenjun Zeng 0001, Peng Chu, Quanzeng You, Jiang Wang 0012, Zicheng Liu 0001, Zhengjun Zha |
CVPR | 8 |
| 2022 | Mobile-Former: Bridging MobileNet and TransformerabstractWe present Mobile-Former, a parallel design of MobileNet and transformer with a two-way bridge in between. This structure leverages the advantages of MobileNet at local processing and transformer at global interaction. And the bridge enables bidirectional fusion of local and global features. Different from recent works on vision transformer, the transformer in Mobile-Former contains very few tokens (e.g. 6 or fewer tokens) that are randomly initialized to learn global priors, resulting in low computational cost. Combining with the proposed light-weight cross attention to model the bridge, Mobile-Former is not only computationally efficient, but also has more representation power. It outperforms MobileNetV3 at low FLOP regime from 25M to 500M FLOPs on ImageNet classification. For instance, Mobile-Former achieves 77.9% top-1 accuracy at 294M FLOPs, gaining 1.3% over MobileNetV3 but saving 17% of computations. When transferring to object detection, Mobile-Former outperforms MobileNetV3 by 8.6 AP in RetinaNet framework. Furthermore, we build an efficient end-to-end detector by replacing backbone, encoder and decoder in DETR with Mobile-Former, which outperforms DETR by 1.3 AP but saves 52% of computational cost and 36% of parameters. Code will be released at https://github.com/aaboys/mobileformer. Yinpeng Chen, Xiyang Dai, Dongdong Chen 0001, Mengchen Liu, Xiaoyi Dong, Lu Yuan 0001, Zicheng Liu 0001 |
CVPR | 7 |
| 2022 | An Empirical Study of Training End-to-End Vision-and-Language TransformersabstractVision-and-language (VL) pre-training has proven to be highly effective on various VL downstream tasks. While recent work has shown that fully transformer-based VL models can be more efficient than previous region-feature-based methods, their performance on downstream tasks often degrades significantly. In this paper, we present Meter, a Multimodal End-to-end TransformER framework, through which we investigate how to design and pre-train a fully transformer-based VL model in an end-to-end manner. Specifically, we dissect the model designs along multiple dimensions: vision encoders (e.g., CLIP-ViT, Swin transformer), text encoders (e.g., RoBERTa, De-BERTa), multimodal fusion module (e.g., merged attention vs. co-attention), architectural design (e.g., encoder-only vs. encoder-decoder), and pre-training objectives (e.g., masked image modeling). We conduct comprehensive experiments and provide insights on how to train a performant VL transformer. Meterachieves an accuracy of 77.64% on the VQAv2 test-std set using only 4M images for pre-training, surpassing the state-of-the-art region-feature-based model by 1.04%, and outperforming the previous best fully transformer-based model by 1.6%. Notably, when further scaled up, our best VQA model achieves an accuracy of 80.54%. Code and pre-trained models are released at https://github.com/zdou0830/METER. Zi-Yi Dou, Yichong Xu, Zhe Gan, Shuohang Wang, Chenguang Zhu 0001, Pengchuan Zhang, Lu Yuan 0001, Nanyun Peng 0001, Zicheng Liu 0001, Michael Zeng 0001 |
CVPR | 11 |
| 2022 | Injecting Semantic Concepts into End-to-End Image CaptioningabstractTremendous progresses have been made in recent years in developing better image captioning models, yet most of them rely on a separate object detector to extract regional features. Recent vision-language studies are shifting towards the detector-free trend by leveraging grid representations for more flexible model training and faster inference speed. However, such development is primarily focused on image understanding tasks, and remains less investigated for the caption generation task. In this paper, we are concerned with a better-performing detector-free image captioning model, and propose a pure vision transformer-based image captioning model, dubbed as ViTCAP, in which grid representations are used without extracting the regional features. For improved performance, we introduce a novel Concept Token Network (CTN) to predict the semantic concepts and then incorporate them into the end-to-end captioning. In particular, the CTN is built on the basis of a vision transformer, and is designed to predict the concept tokens through a classification task, from which the rich semantic information contained greatly benefits the captioning task. Compared with the previous detector-based models, ViTCAP drastically simplifies the architectures and at the same time achieves competitive performance on various challenging image captioning datasets. In particular, ViTCAP reaches 138.1 CIDEr scores on COCO-caption Karpathy-split, 93.8 and 108.6 CIDEr scores on nocaps and Google-CC captioning datasets, respectively. Zhiyuan Fang, Xiaowei Hu 0006, Zhe Gan, Yezhou Yang, Zicheng Liu 0001 |
CVPR | 8 |
| 2022 | SwinBERT: End-to-End Transformers with Sparse Attention for Video CaptioningabstractThe canonical approach to video captioning dictates a caption generation model to learn from offline-extracted dense video features. These feature extractors usually operate on video frames sampled at a fixed frame rate and are often trained on image/video understanding tasks, without adaption to video captioning data. In this work, we present SwinBERT, an end-to-end transformer-based model for video captioning, which takes video frame patches directly as inputs, and outputs a natural language description. Instead of leveraging multiple 2D/3D feature extractors, our method adopts a video transformer to encode spatial-temporal representations that can adapt to variable lengths of video input without dedicated design for different frame rates. Based on this model architecture, we show that video captioning can benefit significantly from more densely sampled video frames as opposed to previous successes with sparsely sampled video frames for video-and-language understanding tasks (e.g., video question answering). Moreover, to avoid the inherent redundancy in consecutive video frames, we propose adaptively learning a sparse attention mask and optimizing it for task-specific performance improvement through better long-range video sequence modeling. Through extensive experiments on 5 video captioning datasets, we show that Swinbert achieves across-the-board performance improvements over previous methods, often by a large margin. The learned sparse attention masks in addition push the limit to new state of the arts, and can be transferred between different video lengths and between different datasets. Code is available at https://github.com/microsoft/SwinBERT. Chung-Ching Lin, Faisal Ahmed 0001, Zhe Gan, Zicheng Liu 0001, Yumao Lu |
CVPR | 6 |
| 2022 | Crossmodal Representation Learning for Zero-shot Action RecognitionabstractWe present a cross-modal Transformer-based frame-work, which jointly encodes video data and text labels for zero-shot action recognition (ZSAR). Our model employs a conceptually new pipeline by which visual representations are learned in conjunction with visual-semantic associations in an end-to-end manner. The model design provides a natural mechanism for visual and semantic representations to be learned in a shared knowledge space, whereby it encourages the learned visual embedding to be discriminative and more semantically consistent. In zero-shot inference, we devise a simple semantic transfer scheme that embeds semantic relatedness information between seen and unseen classes to composite unseen visual prototypes. Accordingly, the discriminative features in the visual structure could be preserved and exploited to alleviate the typical zero-shot issues of information loss, semantic gap, and the hubness problem. Under a rigorous zero-shot setting of not pre-training on additional datasets, the experiment results show our model considerably improves upon the state of the arts in ZSAR, reaching encouraging top-1 accuracy on UCF 101, HMDB51, and ActivityNet benchmark datasets. Code will be made available.11https://github.com/microsoft/ResT Chung-Ching Lin, Zicheng Liu 0001 |
CVPR | 4 |
| 2022 | Should All Proposals Be Treated Equally in Object Detection?
Yunsheng Li, Yinpeng Chen, Xiyang Dai, Dongdong Chen 0001, Mengchen Liu, Pei Yu, Lu Yuan 0001, Zicheng Liu 0001, Nuno Vasconcelos |
ECCV (25) | 9 |
| 2022 | A Simple Approach and Benchmark for 21, 000-Category Object Detection
Yutong Lin, Yue Cao 0001, Zheng Zhang 0022, Zicheng Liu 0001, Han Hu 0001 |
ECCV (11) | 7 |
| 2022 | UniTAB: Unifying Text and Box Outputs for Grounded Vision-Language Modeling
Zhengyuan Yang, Zhe Gan, Xiaowei Hu 0006, Faisal Ahmed 0001, Zicheng Liu 0001, Yumao Lu |
ECCV (36) | 6 |
| 2022 | Unsupervised Learning of Full-Waveform Inversion: Connecting CNN and Partial Differential Equation in a Loop
Xitong Zhang, Yinpeng Chen, Sharon X. Huang, Zicheng Liu 0001, Youzuo Lin |
ICLR | 5 |
| 2022 | An Intriguing Property of Geophysics InversionabstractInversion techniques are widely used to reconstruct subsurface physical properties (e.g., velocity, conductivity) from surface-based geophysical measurements (e.g., seismic, electric/magnetic (EM) data). The problems are governed by partial differential equations (PDEs) like the wave or Maxwell’s equations. Solving geophysical inversion problems is challenging due to the ill-posedness and high computational cost. To alleviate those issues, recent studies leverage deep neural networks to learn the inversion mappings from measurements to the property directly. In this paper, we show that such a mapping can be well modeled by a very shallow (but not wide) network with only five layers. This is achieved based on our new finding of an intriguing property: a near-linear relationship between the input and output, after applying integral transform in high dimensional space. In particular, when dealing with the inversion from seismic data to subsurface velocity governed by a wave equation, the integral results of velocity with Gaussian kernels are linearly correlated to the integral of seismic data with sine kernels. Furthermore, this property can be easily turned into a light-weight encoder-decoder network for inversion. The encoder contains the integration of seismic data and the linear transformation without need for fine-tuning. The decoder only consists of a single transformer block to reverse the integral of velocity. Experiments show that this interesting property holds for two geophysics inversion problems over four different datasets. Compared to much deeper InversionNet, our method achieves comparable accuracy, but consumes significantly fewer parameters Yinan Feng, Yinpeng Chen, Shihang Feng, Zicheng Liu 0001, Youzuo Lin |
ICML | 5 |
| 2022 | Coarse-to-Fine Vision-Language Pre-training with Fusion in the BackboneabstractVision-language (VL) pre-training has recently received considerable attention. However, most existing end-to-end pre-training approaches either only aim to tackle VL tasks such as image-text retrieval, visual question answering (VQA) and image captioning that test high-level understanding of images, or only target region-level understanding for tasks such as phrase grounding and object detection. We present FIBER (Fusion-In-the-Backbone-based transformER), a new VL model architecture that can seamlessly handle both these types of tasks. Instead of having dedicated transformer layers for fusion after the uni-modal backbones, FIBER pushes multimodal fusion deep into the model by inserting cross-attention into the image and text backbones to better capture multimodal interactions. In addition, unlike previous work that is either only pre-trained on image-text data or on fine-grained data with box-level annotations, we present a two-stage pre-training strategy that uses both these kinds of data efficiently: (i) coarse-grained pre-training based on image-text data; followed by (ii) fine-grained pre-training based on image-text-box data. We conduct comprehensive experiments on a wide range of VL tasks, ranging from VQA, image captioning, and retrieval, to phrase grounding, referring expression comprehension, and object detection. Using deep multimodal fusion coupled with the two-stage pre-training, FIBER provides consistent performance improvements over strong baselines across all tasks, often outperforming methods using magnitudes more data. Code is released at https://github.com/microsoft/FIBER. Zi-Yi Dou, Aishwarya Kamath, Zhe Gan, Pengchuan Zhang, Zicheng Liu 0001, Ce Liu 0001, Yann LeCun, Nanyun Peng 0001, Jianfeng Gao 0001 |
NeurIPS | 7 |
| 2022 | ELEVATER: A Benchmark and Toolkit for Evaluating Language-Augmented Visual ModelsabstractLearning visual representations from natural language supervision has recently shown great promise in a number of pioneering works. In general, these language-augmented visual models demonstrate strong transferability to a variety of datasets/tasks. However, it remains challenging to evaluate the transferablity of these foundation models due to the lack of easy-to-use toolkits for fair benchmarking. To tackle this, we build ELEVATER (Evaluation of Language-augmented Visual Task-level Transfer), the first benchmark to compare and evaluate pre-trained language-augmented visual models. Several highlights include: (i) Datasets. As downstream evaluation suites, it consists of 20 image classification datasets and 35 object detection datasets, each of which is augmented with external knowledge. (ii) Toolkit. An automatic hyper-parameter tuning toolkit is developed to ensure the fairness in model adaption. To leverage the full power of language-augmented visual models, novel language-aware initialization methods are proposed to significantly improve the adaption performance. (iii) Metrics. A variety of evaluation metrics are used, including sample-efficiency (zero-shot and few-shot) and parameter-efficiency (linear probing and full model fine-tuning). We will publicly release ELEVATER. Chunyuan Li, Liunian Harold Li, Pengchuan Zhang, Jyoti Aneja, Ping Jin, Houdong Hu, Zicheng Liu 0001, Yong Jae Lee, Jianfeng Gao 0001 |
NeurIPS | 9 |
| 2022 | NUWA-Infinity: Autoregressive over Autoregressive Generation for Infinite Visual SynthesisabstractInfinite visual synthesis aims to generate high-resolution images, long-duration videos, and even visual generation of infinite size. Some recent work tried to solve this task by first dividing data into processable patches and then training the models on them without considering the dependencies between patches. However, since they fail to model global dependencies between patches, the quality and consistency of the generation can be limited. To address this issue, we propose NUWA-Infinity, a patch-level \emph{``render-and-optimize''} strategy for infinite visual synthesis. Given a large image or a long video, NUWA-Infinity first splits it into non-overlapping patches and uses the ordered patch chain as a complete training instance, a rendering model autoregressively predicts each patch based on its contexts. Once a patch is predicted, it is optimized immediately and its hidden states are saved as contexts for the next \emph{``render-and-optimize''} process. This brings two advantages: ($i$) The autoregressive rendering process with information transfer between contexts provides an implicit global probabilistic distribution modeling; ($ii$) The timely optimization process alleviates the optimization stress of the model and helps convergence. Based on the above designs, NUWA-Infinity shows a strong synthesis ability on high-resolution images and long-duration videos. The homepage link is \url{https://nuwa-infinity.microsoft.com}. Chenfei Wu, Xiaowei Hu 0006, Zhe Gan, Zicheng Liu 0001, Yuejian Fang, Nan Duan 0001 |
NeurIPS | 7 |
| 2022 | EFRNet: Efficient Feature Reconstructing Network for Real-Time Scene ParsingabstractIn this paper, we introduce a light-weight and powerful convolutional neural network, termed asefficient feature reconstructing network(EFRNet), for real-time scene parsing. Our key idea is to decompose the process of learning high-resolution representations into two stages: i) bottom-up codebook/coding matrix learning and ii) top-down feature reconstructing. Specifically, the bottom-up process focuses on learningimage-specificcodewords (codebook) using deep-layer features and generating a coding matrix with the shallow-layer feature map. In the top-down process, the learned codebook and coding matrix are used to rebuild high-resolution features via a lightweightfeature reconstructing operator(FRO). In addition, our EFRNet is constructed on a new building block, named efficient adaptive abstraction (EAA) block, to further reduce the overall network parameters and achieve a significant speed up. Extensive experiments are conducted on challenging benchmarks, such as CamVid and Cityscapes. The results show that EFRNet demonstrates state-of-the-art performance with an optimal balance between accuracy and speed. Xin Li 0079, Fan Yang 0054, Ao Luo, Zhicheng Jiao, Hong Cheng 0002, Zicheng Liu 0001 |
IEEE Trans. Multim. | 6 |
| 2021 | VIVO: Visual Vocabulary Pre-Training for Novel Object CaptioningabstractIt is highly desirable yet challenging to generate image captions that can describe novel objects which are unseen in caption-labeled training data, a capability that is evaluated in the novel object captioning challenge (nocaps). In this challenge, no additional image-caption training data, other than COCO Captions, is allowed for model training. Thus, conventional Vision-Language Pre-training (VLP) methods cannot be applied. This paper presents VIsual VOcabulary pre-training (VIVO) that performs pre-training in the absence of caption annotations. By breaking the dependency of paired image-caption training data in VLP, VIVO can leverage large amounts of paired image-tag data to learn a visual vocabulary. This is done by pre-training a multi-layer Transformer model that learns to align image-level tags with their corresponding image region features. To address the unordered nature of image tags, VIVO uses a Hungarian matching loss with masked tag prediction to conduct pre-training. We validate the effectiveness of VIVO by fine-tuning the pre-trained model for image captioning. In addition, we perform an analysis of the visual-text alignment inferred by our model. The results show that our model can not only generate fluent image captions that describe novel objects, but also identify the locations of these objects. Our single model has achieved new state-of-the-art results on nocaps and surpassed the human CIDEr score. Xiaowei Hu 0006, Xi Yin 0006, Lei Zhang 0001, Jianfeng Gao 0001, Zicheng Liu 0001 |
AAAI | 7 |
| 2021 | Probabilistic Model Distillation for Semantic CorrespondenceabstractSemantic correspondence is a fundamental problem in computer vision, which aims at establishing dense correspondences across images depicting different instances under the same category. This task is challenging due to large intra-class variations and a severe lack of ground truth. A popular solution is to learn correspondences from synthetic data. However, because of the limited intra-class appearance and background variations within synthetically generated training data, the model’s capability for handling “real” image pairs using such strategy is intrinsically constrained. We address this problem with the use of a novel Probabilistic Model Distillation (PMD) approach which transfers knowledge learned by a probabilistic teacher model on synthetic data to a static student model with the use of unlabeled real image pairs. A probabilistic supervision reweighting (PSR) module together with a confidence-aware loss (CAL) is used to mine the useful knowledge and alleviate the impact of errors. Experimental results on a variety of benchmarks show that our PMD achieves state-of-the-art performance. To demonstrate the generalizability of our approach, we extend PMD to incorporate stronger supervision for better accuracy – the probabilistic teacher is trained with stronger key-point supervision. Again, we observe the superiority of our PMD. The extensive experiments verify that PMD is able to infer more reliable supervision signals from the probabilistic teacher for representation learning and largely alleviate the influence of errors in pseudo labels. Code is available at https://github.com/fanyang587/PMD. Xin Li 0079, Deng-Ping Fan, Fan Yang 0054, Ao Luo, Hong Cheng 0002, Zicheng Liu 0001 |
CVPR | 6 |
| 2021 | End-to-End Human Pose and Mesh Reconstruction with TransformersabstractWe present a new method, called MEsh TRansfOrmer (METRO), to reconstruct 3D human pose and mesh vertices from a single image. Our method uses a transformer encoder to jointly model vertex-vertex and vertex-joint interactions, and outputs 3D joint coordinates and mesh vertices simultaneously. Compared to existing techniques that regress pose and shape parameters, METRO does not rely on any parametric mesh models like SMPL, thus it can be easily extended to other objects such as hands. We further relax the mesh topology and allow the transformer self-attention mechanism to freely attend between any two vertices, making it possible to learn non-local relationships among mesh vertices and joints. With the proposed masked vertex modeling, our method is more robust and effective in handling challenging situations like partial occlusions. METRO generates new state-of-the-art results for human mesh reconstruction on the public Human3.6M and 3DPW datasets. Moreover, we demonstrate the generalizability of METRO to 3D hand reconstruction in the wild, outperforming existing state-of-the-art methods on FreiHAND dataset. Zicheng Liu 0001 |
CVPR | 3 |
| 2021 | Compressing Visual-linguistic Model via Knowledge DistillationabstractDespite exciting progress in pre-training for visual-linguistic (VL) representations, very few aspire to a small VL model. In this paper, we study knowledge distillation (KD) to effectively compress a transformer based large VL model into a small VL model. The major challenge arises from the inconsistent regional visual tokens extracted from different detectors of Teacher and Student, resulting in the misalignment of hidden representations and attention distributions. To address the problem, we retrain and adapt the Teacher by using the same region proposals from Student’s detector while the features are from Teacher’s own object detector. With aligned network inputs, the adapted Teacher is capable of transferring the knowledge through the intermediate representations. Specifically, we use the mean square error loss to mimic the attention distribution inside the transformer block, and present a token-wise noise contrastive loss to align the hidden state by contrasting with negative representations stored in a sample queue. To this end, we show that our proposed distillation significantly improves the performance of small VL models on image captioning and visual question answering tasks. It reaches 120.8 in CIDEr score on COCO captioning, an improvement of 5.1 over its non-distilled counterpart; and an accuracy of 69.8 on VQA 2.0, a 0.8 gain from the baseline. Our extensive experiments and ablations confirm the effectiveness of VL distillation in both pre-training and fine-tuning stages. Zhiyuan Fang, Xiaowei Hu 0006, Yezhou Yang, Zicheng Liu 0001 |
ICCV | 6 |
| 2021 | MicroNet: Improving Image Recognition with Extremely Low FLOPsabstractThis paper aims at addressing the problem of substantial performance degradation at extremely low computational cost (e.g. 5M FLOPs on ImageNet classification). We found that two factors, sparse connectivity and dynamic activation function, are effective to improve the accuracy. The former avoids the significant reduction of network width, while the latter mitigates the detriment of reduction in network depth. Technically, we propose micro-factorized convolution, which factorizes a convolution matrix into low rank matrices, to integrate sparse connectivity into convolution. We also present a new dynamic activation function, named Dynamic Shift Max, to improve the non-linearity via maxing out multiple dynamic fusions between an input feature map and its circular channel shift. Building upon these two new operators, we arrive at a family of networks, named MicroNet, that achieves significant performance gains over the state of the art in the low FLOP regime. For instance, under the constraint of 12M FLOPs, MicroNet achieves 59.4% top-1 accuracy on ImageNet classification, outperforming MobileNetV3 by 9.6%. Source code is at https://github.com/liyunsheng13/micronet. Yunsheng Li, Yinpeng Chen, Xiyang Dai, Dongdong Chen 0001, Mengchen Liu, Lu Yuan 0001, Zicheng Liu 0001, Lei Zhang 0001, Nuno Vasconcelos |
ICCV | 7 |
| 2021 | Mesh GraphormerabstractWe present a graph-convolution-reinforced transformer, named Mesh Graphormer, for 3D human pose and mesh reconstruction from a single image. Recently both transformers and graph convolutional neural networks (GC-NNs) have shown promising progress in human mesh re-construction. Transformer-based approaches are effective in modeling non-local interactions among 3D mesh vertices and body joints, whereas GCNNs are good at exploiting neighborhood vertex interactions based on a pre-specified mesh topology. In this paper, we study how to combine graph convolutions and self-attentions in a transformer to model both local and global interactions. Experimental results show that our proposed method, Mesh Graphormer, significantly outperforms the previous state-of-the-art methods on multiple benchmarks, including Human3.6M, 3DPW, and FreiHAND datasets. Code and pre-trained models are available at https://github.com/microsoft/MeshGraphormer. Zicheng Liu 0001 |
ICCV | 3 |
| 2021 | End-to-End Semi-Supervised Object Detection with Soft TeacherabstractThis paper presents an end-to-end semi-supervised object detection approach, in contrast to previous more complex multi-stage methods. The end-to-end training gradually improves pseudo label qualities during the curriculum, and the more and more accurate pseudo labels in turn benefit object detection training. We also propose two simple yet effective techniques within this framework: a soft teacher mechanism where the classification loss of each unlabeled bounding box is weighed by the classification score produced by the teacher network; a box jittering approach to select reliable pseudo boxes for the learning of box regression. On the COCO benchmark, the proposed approach outperforms previous methods by a large margin under various labeling ratios, i.e. 1%, 5% and 10%. Moreover, our approach proves to perform also well when the amount of labeled data is relatively large. For example, it can improve a 40.9 mAP baseline detector trained using the full COCO training set by +3.6 mAP, reaching 44.5 mAP, by leveraging the 123K unlabeled images of COCO. On the state-of-the-art Swin Transformer based object detector (58.9 mAP on test-dev), it can still significantly improve the detection accuracy by +1.5 mAP, reaching 60.4 mAP, and improve the instance segmentation accuracy by +1.2 mAP, reaching 52.4 mAP. Further incorporating with the Object365 pre-trained model, the detection accuracy reaches 61.3 mAP and the instance segmentation accuracy reaches 53.0 mAP, pushing the new state-of-the-art. The code and models will be made publicly available at https://github.com/microsoft/SoftTeacher. Mengde Xu, Zheng Zhang 0022, Han Hu 0001, Fangyun Wei, Xiang Bai, Zicheng Liu 0001 |
ICCV | 8 |
| 2021 | Learning Nonparametric Human Mesh Reconstruction From A Single Image Without Ground Truth MeshesabstractWe present a novel approach to learn human mesh reconstruction without ground truth mesh labels. This is made possible by introducing two new terms into the loss function of a graph convolutional neural network (Graph CNN). The first term is the Laplacian prior that acts as a regularizer on the mesh reconstruction. The second term is the part segmentation loss that forces the projected region of the reconstructed mesh to match the part segmentation. Extensive experiments validate the effectiveness of the proposed approach. Zicheng Liu 0001, Ming-Ting Sun |
ICIP | 4 |
| 2021 | SEED: Self-supervised Distillation For Visual Representation
Zhiyuan Fang, Lei Zhang 0001, Yezhou Yang, Zicheng Liu 0001 |
ICLR | 6 |
| 2021 | Revisiting Dynamic Convolution via Matrix Decomposition
Yunsheng Li, Yinpeng Chen, Xiyang Dai, Mengchen Liu, Dongdong Chen 0001, Lu Yuan 0001, Zicheng Liu 0001, Nuno Vasconcelos |
ICLR | 8 |
| 2021 | Stronger NAS with Weaker PredictorsabstractNeural Architecture Search (NAS) often trains and evaluates a large number of architectures. Recent predictor-based NAS approaches attempt to alleviate such heavy computation costs with two key steps: sampling some architecture-performance pairs and fitting a proxy accuracy predictor. Given limited samples, these predictors, however, are far from accurate to locate top architectures due to the difficulty of fitting the huge search space. This paper reflects on a simple yet crucial question: if our final goal is to find the best architecture, do we really need to model the whole space well?. We propose a paradigm shift from fitting the whole architecture space using one strong predictor, to progressively fitting a search path towards the high-performance sub-space through a set of weaker predictors. As a key property of the weak predictors, their probabilities of sampling better architectures keep increasing. Hence we only sample a few well-performed architectures guided by the previously learned predictor and estimate a new better weak predictor. This embarrassingly easy framework, dubbed WeakNAS, produces coarse-to-fine iteration to gradually refine the ranking of sampling space. Extensive experiments demonstrate that WeakNAS costs fewer samples to find top-performance architectures on NAS-Bench-101 and NAS-Bench-201. Compared to state-of-the-art (SOTA) predictor-based NAS methods, WeakNAS outperforms all with notable margins, e.g., requiring at least 7.5x less samples to find global optimal on NAS-Bench-101. WeakNAS can also absorb their ideas to boost performance more. Further, WeakNAS strikes the new SOTA result of 81.3% in the ImageNet MobileNet Search Space. The code is available at: https://github.com/VITA-Group/WeakNAS. Xiyang Dai, Dongdong Chen 0001, Yinpeng Chen, Mengchen Liu, Zhangyang Wang, Zicheng Liu 0001, Lu Yuan 0001 |
NeurIPS | 8 |
| 2021 | Human pose estimation and its application to action recognition: A survey
Liangchen Song, Gang Yu 0002, Junsong Yuan 0001, Zicheng Liu 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2021 | Cross-Domain Complementary Learning Using Pose for Multi-Person Part SegmentationabstractSupervised deep learning with pixel-wise training labels has great successes on multi-person part segmentation. However, data labeling at pixel-level is very expensive. To solve the problem, people have been exploring to use synthetic data to avoid the data labeling. Although it is easy to generate labels for synthetic data, the results are much worse compared to those using real data and manual labeling. The degradation of the performance is mainly due to the domain gap, i.e., the discrepancy of the pixel value statistics between real and synthetic data. In this paper, we observe that real and synthetic humans both have a skeleton (pose) representation. We found that the skeletons can effectively bridge the synthetic and real domains during the training. Our proposed approach takes advantage of the rich and realistic variations of the real data and the easily obtainable labels of the synthetic data to learn multi-person part segmentation on real images without any human-annotated labels. Through experiments, we show that without any human labeling, our method performs comparably to several state-of-the-art approaches which require human labeling on Pascal-Person-Parts and COCO-DensePose datasets. On the other hand, if part labels are also available in the real-images during training, our method outperforms the supervised state-of-the-art methods by a large margin. We further demonstrate the generalizability of our method on predicting novel keypoints in real images where no real data labels are available for the novel keypoints detection. Code and pre-trained models are available at https://github.com/kevinlin311tw/CDCL-human-part-segmentation. Yinpeng Chen, Zicheng Liu 0001, Ming-Ting Sun |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2020 | Rethinking Classification and Localization for Object DetectionabstractTwo head structures (i.e. fully connected head and convolution head) have been widely used in R-CNN based detectors for classification and localization tasks. However, there is a lack of understanding of how does these two head structures work for these two tasks. To address this issue, we perform a thorough analysis and find an interesting fact that the two head structures have opposite preferences towards the two tasks. Specifically, the fully connected head (fc-head) is more suitable for the classification task, while the convolution head (conv-head) is more suitable for the localization task. Furthermore, we examine the output feature maps of both heads and find that fc-head has more spatial sensitivity than conv-head. Thus, fc-head has more capability to distinguish a complete object from part of an object, but is not robust to regress the whole object. Based upon these findings, we propose a Double-Head method, which has a fully connected head focusing on classification and a convolution head for bounding box regression. Without bells and whistles, our method gains +3.5 and +2.8 AP on MS COCO dataset from Feature Pyramid Network (FPN) baselines with ResNet-50 and ResNet-101 backbones, respectively. Yue Wu 0008, Yinpeng Chen, Lu Yuan 0001, Zicheng Liu 0001, Yun Fu 0001 |
CVPR | 4 |
| 2020 | Dynamic Convolution: Attention Over Convolution KernelsabstractLight-weight convolutional neural networks (CNNs) suffer performance degradation as their low computational budgets constrain both the depth (number of convolution layers) and the width (number of channels) of CNNs, resulting in limited representation capability. To address this issue, we present Dynamic Convolution, a new design that increases model complexity without increasing the network depth or width. Instead of using a single convolution kernel per layer, dynamic convolution aggregates multiple parallel convolution kernels dynamically based upon their attentions, which are input dependent. Assembling multiple kernels is not only computationally efficient due to the small kernel size, but also has more representation power since these kernels are aggregated in a non-linear way via attention. By simply using dynamic convolution for the state-of-the-art architecture MobileNetV3-Small, the top-1 accuracy of ImageNet classification is boosted by 2.9% with only 4% additional FLOPs and 2.9 AP gain is achieved on COCO keypoint detection. Yinpeng Chen, Xiyang Dai, Mengchen Liu, Dongdong Chen 0001, Lu Yuan 0001, Zicheng Liu 0001 |
CVPR | 6 |
| 2020 | Dynamic ReLU
Yinpeng Chen, Xiyang Dai, Mengchen Liu, Dongdong Chen 0001, Lu Yuan 0001, Zicheng Liu 0001 |
ECCV (19) | 6 |
| 2020 | Human Action Image Generation with Differential PrivacyabstractLarge volumes of human action image data are becoming increasingly available due to the prevalence of surveillance cameras and smart personal devices. While such image data enables important applications such as activity recognition for health and safety enhancement, they often contain sensitive information such as identities that introduce high risks to individual privacy. Existing image privacy-enhancing techniques are either developed at the cost of sacrificing image utility or lack of provable privacy guarantees. We propose a novel human action image generation model that enforces rigorous differential privacy protection. Theoretical analysis is provided to quantify the privacy protection on the training data within the differential privacy framework. Experiments with real-world datasets demonstrate that images generated using our method achieve higher image utilities than baselines given similar degrees of privacy protection. Mingxuan Sun 0001, Zicheng Liu 0001 |
ICME | 3 |
| 2019 | Large Scale Incremental LearningabstractModern machine learning suffers from \textit{catastrophic forgetting} when learning new classes incrementally. The performance dramatically degrades due to the missing data of old classes. Incremental learning methods have been proposed to retain the knowledge acquired from the old classes, by using knowledge distilling and keeping a few exemplars from the old classes. However, these methods struggle to \textbf{scale up to a large number of classes}. We believe this is because of the combination of two factors: (a) the data imbalance between the old and new classes, and (b) the increasing number of visually similar classes. Distinguishing between an increasing number of visually similar classes is particularly challenging, when the training data is unbalanced. We propose a simple and effective method to address this data imbalance issue. We found that the last fully connected layer has a strong bias towards the new classes, and this bias can be corrected by a linear model. With two bias parameters, our method performs remarkably well on two large datasets: ImageNet (1000 classes) and MS-Celeb-1M (10000 classes), outperforming the state-of-the-art algorithms by 11.1\% and 13.2\% respectively. Yue Wu 0008, Yinpeng Chen, Yuancheng Ye, Zicheng Liu 0001, Yandong Guo, Yun Fu 0001 |
CVPR | 5 |
| 2019 | Discriminative Spatio-Temporal Pattern Discovery for 3D Action RecognitionabstractDespite the recent success of 3D action recognition using depth sensor, most existing works target how to improve the action recognition performance, rather than understanding how different types of actions are performed. In this paper, we propose to discover discriminative spatio-temporal patterns for 3D action recognition. Discovering these patterns can not only help to improve the action recognition performance but also help us to understand and differentiate between the action category. Our proposed method takes the spatio-temporal structure of 3D action into consideration and can discover essential spatio-temporal patterns that play key roles in action recognition. Instead of relying on an end-to-end network to learn the 3D action representation and perform classification, we simply present each 3D action as a series of temporal stages composed by 3D poses. Then, we rely on nearest neighbor matching and bilinear classifiers to simultaneously identify both critical temporal stages and spatial joints for each action class. Despite using raw action representation and a linear classifier, experiments on five benchmark data sets show that the proposed spatio-temporal naïve Bayes mutual information maximization can achieve a competitive performance compared with the state-of-the-art methods that use sophisticated end-to-end learning, and has the advantage of finding discriminative spatio-temporal action patterns. Junwu Weng, Chaoqun Weng, Junsong Yuan 0001, Zicheng Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | Reinforced Temporal Attention and Split-Rate Transfer for Depth-Based Person Re-identification
Nikolaos Karianakis, Zicheng Liu 0001, Yinpeng Chen, Stefano Soatto |
ECCV (5) | 2 |
| 2018 | A set-to-set nearest neighbor approach for robust and efficient face recognition with image sets
Ling Wang 0013, Hong Cheng 0002, Zicheng Liu 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2018 | Depth Super-Resolution on RGB-D Video Sequences With Large Displacement 3D MotionabstractTo enhance the resolution and accuracy of depth data, some video-based depth super-resolution methods have been proposed which utilizes its neighboring depth images in the temporal domain. They often consist of two main stages: motion compensation of temporally neighboring depth images and fusion of compensated depth images. However, large displacement 3D motion often leads to compensation error, and the compensation error is further introduced into the fusion. A video-based depth super-resolution method with novel motion compensation and fusion approaches is proposed in this paper. We claim that, 3D Nearest Neighboring Field (NNF) is a better choice than using positions with true motion displacement for depth enhancements. To handle large displacement 3D motion, the compensation stage utilized 3D NNF instead of true motion used in previous methods. Next, the fusion approach is modeled as a regression problem to predict the super-resolution result efficiently for each depth image by using its compensated depth images. A new deep convolutional neural network architecture is designed for fusion, which is able to employ a large amount of video data for learning the complicated regression function. We comprehensively evaluate our method on various RGB-D video sequences to show its superior performance. Yucheng Wang 0003, Jian Zhang 0002, Zicheng Liu 0001, Qiang Wu 0001, Zhengyou Zhang, Yunde Jia |
IEEE Trans. Image Process. | 3 |
| 2018 | 3D cartoon face rigging from sparse examples
Jingyong Zhou, Hsiang-Tao Wu, Zicheng Liu 0001, Xin Tong 0001, Baining Guo |
Vis. Comput. | 3 |
| 2017 | A Tube-and-Droplet-Based Approach for Representing and Analyzing Motion TrajectoriesabstractTrajectory analysis is essential in many applications. In this paper, we address the problem of representing motion trajectories in a highly informative way, and consequently utilize it for analyzing trajectories. Our approach first leverages the complete information from given trajectories to construct a thermal transfer field which provides a context-rich way to describe the global motion pattern in a scene. Then, a 3D tube is derived which depicts an input trajectory by integrating its surrounding motion patterns contained in the thermal transfer field. The 3D tube effectively: 1) maintains the movement information of a trajectory, 2) embeds the complete contextual motion pattern around a trajectory, 3) visualizes information about a trajectory in a clear and unified way. We further introduce a droplet-based process. It derives a droplet vector from a 3D tube, so as to characterize the high-dimensional 3D tube information in a simple but effective way. Finally, we apply our tube-and-droplet representation to trajectory analysis applications including trajectory clustering, trajectory classification & abnormality detection, and 3D action recognition. Experimental comparisons with state-of-the-art algorithms demonstrate the effectiveness of our approach. Weiyao Lin, Hongteng Xu, Junchi Yan, Mingliang Xu 0001, Jianxin Wu 0001, Zicheng Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2016 | An image-to-class dynamic time warping approach for both 3D static and trajectory hand gesture recognition
Hong Cheng 0002, Zhongjun Dai, Zicheng Liu 0001, Yang Zhao 0024 |
Pattern Recognit. | 3 |
| 2016 | Sparsity-Induced Similarity Measure and Its ApplicationsabstractThe structures of feature vectors-based semisupervised/supervised learning have gained considerable interest in recent years due to their effectiveness for better object modeling and classification. In many machine learning and computer vision tasks, a critical issue is the similarity between two feature vectors. In this paper, we present a novel technique to measure similarities among feature vectors by decomposing each feature vector as an ℓ1sparse linear combination of the rest of the feature vectors. The main idea is that the coefficients in such sparse decomposition reflect the features' neighborhood structure, thus providing better similarity measures among the decomposed feature vector and the rest of the feature vectors. The proposed approach is applied to label propagation and action recognition, and is evaluated on several commonly used datasets. The experimental results show that the proposed sparsity-induced similarity measure significantly improves the performance of both label propagation and action recognition. Hong Cheng 0002, Zicheng Liu 0001, Lei Hou 0019, Jie Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | Survey on 3D Hand Gesture RecognitionabstractThree-dimensional hand gesture recognition has attracted increasing research interests in computer vision, pattern recognition, and human-computer interaction. The emerging depth sensors greatly inspired various hand gesture recognition approaches and applications, which were severely limited in the 2D domain with conventional cameras. This paper presents a survey of some recent works on hand gesture recognition using 3D depth sensors. We first review the commercial depth sensors and public data sets that are widely used in this field. Then, we review the state-of-the-art research for 3D hand gesture recognition in four aspects: 1) 3D hand modeling; 2) static hand gesture recognition; 3) hand trajectory gesture recognition; and 4) continuous hand gesture recognition. While the emphasis is on 3D hand gesture recognition approaches, the related applications and typical systems are also briefly summarized for practitioners. Hong Cheng 0002, Lu Yang 0002, Zicheng Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2016 | Handling Occlusion and Large Displacement Through Improved RGB-D Scene Flow EstimationabstractThe accuracy of scene flow is restricted by several challenges such as occlusion and large displacement motion. When occlusion happens, the positions inside the occluded regions lose their corresponding counterparts in preceding and succeeding frames. Large displacement motion will increase the complexity of motion modeling and computation. Moreover, occlusion and large displacement motion are highly related problems in scene flow estimation, e.g., large displacement motion often leads to considerably occluded regions in the scene. An improved dense scene flow method based on red-green-blue-depth (RGB-D) data is proposed in this paper. To handle occlusion, we model the occlusion status for each point in our problem formulation, and jointly estimate the scene flow and occluded regions. To deal with large displacement motion, we employ an over-parameterized scene flow representation to model both the rotation and translation components of the scene flow, since large displacement motion cannot be well approximated using translational motion only. Furthermore, we employ a two-stage optimization procedure for this overparameterized scene flow representation. In the first stage, we propose a new RGB-D PatchMatch method, which is mainly applied in the RGB-D image space to reduce the computational complexity introduced by the large displacement motion. According to the quantitative evaluation based on the Middlebury data set, our method outperforms other published methods. The improved performance is also comprehensively confirmed on the real data acquired by Kinect sensor. Yucheng Wang 0003, Jian Zhang 0002, Zicheng Liu 0001, Qiang Wu 0001, Philip A. Chou, Zhengyou Zhang, Yunde Jia |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2016 | 3D cartoon face generation by local deformation mapping
Jingyong Zhou, Xin Tong 0001, Zicheng Liu 0001, Baining Guo |
Vis. Comput. | 3 |
| 2015 | ImmerseBoard: Immersive Telepresence Experience using a Digital WhiteboardabstractImmerseBoard is a system for remote collaboration through a digital whiteboard that gives participants a 3D immersive experience, enabled only by an RGBD camera (Microsoft Kinect) mounted on the side of a large touch display. Using 3D processing of the depth images, life-sized rendering, and novel visualizations, ImmerseBoard emulates writing side-by-side on a physical whiteboard, or alternatively on a mirror. User studies involving three tasks show that compared to standard video conferencing with a digital whiteboard, ImmerseBoard provides participants with a quantitatively better ability to estimate their remote partners' eye gaze direction, gesture direction, intention, and level of agreement. Moreover, these quantitative capabilities translate qualitatively into a heightened sense of being together and a more enjoyable experience. ImmerseBoard's form factor is suitable for practical and easy installation in homes and offices. Keita Higuchi, Yinpeng Chen, Philip A. Chou, Zhengyou Zhang, Zicheng Liu 0001 |
CHI | 5 |
| 2015 | Anomaly detection by using random projection forestabstractIn this paper, we present a novel method for detecting anomalies from surveillance videos, which utilizes the random projection forest for evaluating the rarity of visual clues in a video frame. Given the hierarchical clustering of the data in a random projection tree and the aggregation process in the random forest, we achieve both efficient estimation of incoming samples and improved robustness against under-fitting and over-fitting under improperly selected models. Random forest is also online updatable, which is meaningful for future online anomaly detection. We designed the splitting rule for anomaly detection, the system framework and the criterion of anomaly determination. The efficiency of the proposed methods has been validated by experiments on public UCSD datasets and compared with previously reported results. Zicheng Liu 0001, Ming-Ting Sun |
ICIP | 2 |
| 2015 | VTouch: Vision-enhanced interaction for large touch displaysabstractWe propose a system that augments touch input with visual understanding of the user to improve interaction with a large touch-sensitive display. A commodity color plus depth sensor such as Microsoft Kinect adds the visual modality and enables new interactions beyond touch. Through visual analysis, the system understands where the user is, who the user is, and what the user is doing even before the user touches the display. Such information is used to enhance interaction in multiple ways. For example, a user can use simple gestures to bring up menu items such as color palette and soft keyboard; menu items can be shown where the user is and can follow the user; hovering can show information to the user before the user commits to touch; the user can perform different functions (for example writing and erasing) with different hands; and the user's preference profile can be maintained, distinct from other users. User studies are conducted and the users very much appreciate the value of these and other enhanced interactions. Yinpeng Chen, Zicheng Liu 0001, Philip A. Chou, Zhengyou Zhang |
ICME | 2 |
| 2015 | Real time gaze estimation with a consumer depth camera
Zicheng Liu 0001, Ming-Ting Sun |
Inf. Sci. | 2 |
| 2015 | Propagative Hough Voting for Human Activity Detection and RecognitionabstractGeneralized Hough voting (HV) has shown promising results in both object and action detection. However, most existing HV methods will suffer when insufficient training data are provided. We propose propagative HV to address this limitation and apply it to human activity analysis. Instead of training a discriminative classifier for local feature voting, we match individual local features to propagate the label and spatiotemporal configuration information of local features via HV. To enable a fast local feature matching, we index the local features using random projection trees (RPTs). RPTs can reveal the low-dimension manifold structure to provide adaptive local feature matching. Moreover, as the RPT index can be built in either labeled or unlabeled dataset, it can be applied to different tasks, such as activity search (limited training) and recognition (sufficient training). The superior performances on benchmarked datasets validate that our propagative HV can outperform state-of-the-art techniques in various activity analysis tasks, such as activity search, recognition, and prediction. Gang Yu 0002, Junsong Yuan 0001, Zicheng Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2015 | Random Forest Construction With Robust Semisupervised Node SplittingabstractRandom forest (RF) is a very important classifier with applications in various machine learning tasks, but its promising performance heavily relies on the size of labeled training data. In this paper, we investigate constructing of RFs with a small size of labeled data and find that the performance bottleneck is located in the node splitting procedures; hence, existing solutions fail to properly partition the feature space if there are insufficient training data. To achieve robust node splitting with insufficient data, we present semisupervised splitting to overcome this limitation by splitting nodes with the guidance of both labeled and abundant unlabeled data. In particular, an accurate quality measure of node splitting is obtained by carrying out the kernel-based density estimation, whereby a multiclass version of asymptotic mean integrated squared error criterion is proposed to adaptively select the optimal bandwidth of the kernel. To avoid the curse of dimensionality, we project the data points from the original high-dimensional feature space onto a low-dimensional subspace before estimation. A unified optimization framework is proposed to select a coupled pair of subspace and separating hyperplane such that the smoothness of the subspace and the quality of the splitting are guaranteed simultaneously. Our algorithm efficiently avoids overfitting caused by bad initialization and local maxima when compared with conventional margin maximization-based semisupervised methods. We demonstrate the effectiveness of the proposed algorithm by comparing it with state-of-the-art supervised and semisupervised algorithms for typical computer vision applications, such as object categorization, face recognition, and image segmentation, on publicly available data sets. Xiao Liu 0012, Mingli Song, Dacheng Tao, Zicheng Liu 0001, Chun Chen 0001, Jiajun Bu |
IEEE Trans. Image Process. | 4 |
| 2014 | Discriminative Orderlet Mining for Real-Time Recognition of Human-Object Interaction
Gang Yu 0002, Zicheng Liu 0001, Junsong Yuan 0001 |
ACCV (5) | 2 |
| 2014 | Can Visual Recognition Benefit from Auxiliary Information in Training?
Qilin Zhang 0004, Gang Hua 0001, Wei Liu 0005, Zicheng Liu 0001, Zhengyou Zhang |
ACCV (1) | 4 |
| 2014 | Randomized Support Vector Forest
Xutao Lv, Tony X. Han, Zicheng Liu 0001, Zhihai He |
BMVC | 3 |
| 2014 | Automatic Camera-Screen Localization
Francisco Gómez Fernández, Zicheng Liu 0001, Alvaro Pardo, Marta Mejail |
CIARP | 2 |
| 2014 | Detecting Subtle Human-Object Interactions Using Kinect
Sebastián Ubalde, Zicheng Liu 0001, Marta Mejail |
CIARP | 2 |
| 2014 | Towards accurate and robust cross-ratio based gaze trackers through learning from simulationabstractCross-ratio (CR) based methods offer many attractive properties for remote gaze estimation using a single camera in an uncalibrated setup by exploiting invariance of a plane projectivity. Unfortunately, due to several simplification assumptions, the performance of CR-based eye gaze trackers decays significantly as the subject moves away from the calibration position. In this paper, we introduce an adaptive homography mapping for achieving gaze prediction with higher accuracy at the calibration position and more robustness under head movements. This is achieved with a learning-based method for compensating both spatially-varying gaze errors and head pose dependent errors simultaneously in a unified framework. The model of adaptive homography is trained offline using simulated data, saving a tremendous amount of time in data collection. We validate the effectiveness of the proposed approach using both simulated and real data from a physical setup. We show that our method compares favorably against other state-of-the-art CR based methods. Jia-Bin Huang 0001, Qin Cai, Zicheng Liu 0001, Narendra Ahuja, Zhengyou Zhang |
ETRA | 3 |
| 2014 | Video Summarization based on Nonnegative Linear ReconstructionabstractWith the development of imaging techniques and the Internet, it is hard to effectively and efficiently manage, index and store large amounts of videos. Video summarization appears to this need by obtaining useful information in videos. Recently, the reconstruction concept has been introduced into video summarization. However, most of these existing reconstruction methods where a dictionary of key frames is selected to reconstruct the original video best might come up with redundant information. In this paper, we propose a novel framework named Video Summarization based on Nonnegative Linear Reconstruction (VSNLR) which allows only additive, not subtractive, linear combinations. Our approach consists of two core steps: (1) we detect the shot boundary and choose a representative frame for every shot, then all the representative frames form the candidate set; (2) for every frame in the candidate set, VSNLR selects related frames to reconstruct the given frame by using the nonnegative linear reconstruction function. The key frames are selected by minimizing the sum of reconstruction errors. Experiments on a dataset and comparison to the state-of-art method demonstrate our advantage. Qiao Luan, Mingli Song, Chu Yee Liau, Jiajun Bu, Zicheng Liu 0001, Ming-Ting Sun |
ICME | 5 |
| 2014 | Realtime gaze estimation with online calibrationabstractFor an eye gaze estimation system, calibration is an unavoidable procedure to determine certain person-specific parameters, either explicitly or implicitly. Although several offline implicit calibration methods have been proposed to ease the calibration burden, the calibration procedure is still cumbersome and the gaze estimation accuracy needs further improvement. In this paper, we propose a novel 3D gaze estimation system with online calibration. The proposed system uses a new 3D model-based gaze estimation method with a single consumer camera (Kinect). Unlike previous gaze estimation methods using explicit offline calibration with fixed number of calibration points or implicit calibration, our approach constantly improves person-specific eye parameters through online calibration, which enables the system to adapt gradually to a new user. The experimental results and the human-computer interaction (HCI) application show that the proposed system can work in realtime with superior gaze estimation accuracy (<; 2°) and minimal calibration burden. Mingli Song, Zicheng Liu 0001, Ming-Ting Sun |
ICME | 3 |
| 2014 | Introduction to the special issue on visual understanding and applications with RGB-D cameras
Zicheng Liu 0001, Michael Beetz, Daniel Cremers, Juergen Gall, Wanqing Li 0001, Dejan Pangercic, Jürgen Sturm, Yu-Wing Tai |
J. Vis. Commun. Image Represent. | 1 |
| 2014 | A robust elastic net approach for feature learning
Ling Wang 0013, Hong Cheng 0002, Zicheng Liu 0001, Ce Zhu |
J. Vis. Commun. Image Represent. | 3 |
| 2014 | Real world activity summary for senior home monitoring
Hong Cheng 0002, Zicheng Liu 0001, Yang Zhao 0024, Guo Ye, Xinghai Sun |
Multim. Tools Appl. | 2 |
| 2014 | Kernelized pyramid nearest-neighbor search for object categorization
Hong Cheng 0002, Rongchao Yu, Zicheng Liu 0001, Lu Yang 0002, Xue-wen Chen 0001 |
Mach. Vis. Appl. | 3 |
| 2014 | Learning Actionlet Ensemble for 3D Human Action RecognitionabstractHuman action recognition is an important yet challenging task. Human actions usually involve human-object interactions, highly articulated motions, high intra-class variations, and complicated temporal structures. The recently developed commodity depth sensors open up new possibilities of dealing with this problem by providing 3D depth data of the scene. This information not only facilitates a rather powerful human motion capturing technique, but also makes it possible to efficiently model human-object interactions and intra-class variations. In this paper, we propose to characterize the human actions with a novel actionlet ensemble model, which represents the interaction of a subset of human joints. The proposed model is robust to noise, invariant to translational and temporal misalignment, and capable of characterizing both the human motion and the human-object interactions. We evaluate the proposed approach on three challenging action recognition datasets captured by Kinect devices, a multiview action recognition dataset captured with Kinect device, and a dataset captured by a motion capture system. The experimental evaluations show that the proposed approach achieves superior performance to the state-of-the-art algorithms. Jiang Wang 0001, Zicheng Liu 0001, Ying Wu 0001, Junsong Yuan 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2014 | Animated Pose Templates for Modeling and Detecting Human ActionsabstractThis paper presents animated pose templates (APTs) for detecting short-term, long-term, and contextual actions from cluttered scenes in videos. Each pose template consists of two components: 1) a shape template with deformable parts represented in an And-node whose appearances are represented by the Histogram of Oriented Gradient (HOG) features, and 2) a motion template specifying the motion of the parts by the Histogram of Optical-Flows (HOF) features. A shape template may have more than one motion template represented by an Or-node. Therefore, each action is defined as a mixture (Or-node) of pose templates in an And-Or tree structure. While this pose template is suitable for detecting short-term action snippets in two to five frames, we extend it in two ways: 1) For long-term actions, we animate the pose templates by adding temporal constraints in a Hidden Markov Model (HMM), and 2) for contextual actions, we treat contextual objects as additional parts of the pose templates and add constraints that encode spatial correlations between parts. To train the model, we manually annotate part locations on several keyframes of each video and cluster them into pose templates using EM. This leaves the unknown parameters for our learning algorithm in two groups: 1) latent variables for the unannotated frames including pose-IDs and part locations, 2) model parameters shared by all training samples such as weights for HOG and HOF features, canonical part locations of each pose, coefficients penalizing pose-transition and part-deformation. To learn these parameters, we introduce a semi-supervised structural SVM algorithm that iterates between two steps: 1) learning (updating) model parameters using labeled data by solving a structural SVM optimization, and 2) imputing missing variables (i.e., detecting actions on unlabeled frames) with parameters learned from the previous step and progressively accepting high-score frames as newly labeled examples. This algorithm belongs to a family of optimization methods known as the Concave-Convex Procedure (CCCP) that converge to a local optimal solution. The inference algorithm consists of two components: 1) Detecting top candidates for the pose templates, and 2) computing the sequence of pose templates. Both are done by dynamic programming or, more precisely, beam search. In experiments, we demonstrate that this method is capable of discovering salient poses of actions as well as interactions with contextual objects. We test our method on several public action data sets and a challenging outdoor contextual action data set collected by ourselves. The results show that our model achieves comparable or better performance compared to state-of-the-art methods. Benjamin Z. Yao, Xiaohan Nie, Zicheng Liu 0001, Song-Chun Zhu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2014 | On the improvement of human action recognition from depth map sequences using Space-Time Occupancy Patterns
Antônio Wilson Vieira, Erickson R. Nascimento, Gabriel L. Oliveira, Zicheng Liu 0001, Mario Fernando Montenegro Campos |
Pattern Recognit. Lett. | 4 |
| 2013 | Tensor-Based Human Body ModelingabstractIn this paper, we present a novel approach to model 3D human body with variations on both human shape and pose, by exploring a tensor decomposition technique. 3D human body modeling is important for 3D reconstruction and animation of realistic human body, which can be widely used in Tele-presence and video game applications. It is challenging due to a wide range of shape variations over different people and poses. The existing SCAPE model is popular in computer vision for modeling 3D human body. However, it considers shape and pose deformations separately, which is not accurate since pose deformation is person-dependent. Our tensor-based model addresses this issue by jointly modeling shape and pose deformations. Experimental results demonstrate that our tensor-based model outperforms the SCAPE model quite significantly. We also apply our model to capture human body using Microsoft Kinect sensors with excellent results. Yinpeng Chen, Zicheng Liu 0001, Zhengyou Zhang |
CVPR | 2 |
| 2013 | Semi-supervised Node Splitting for Random Forest ConstructionabstractNode splitting is an important issue in Random Forest but robust splitting requires a large number of training samples. Existing solutions fail to properly partition the feature space if there are insufficient training data. In this paper, we present semi-supervised splitting to overcome this limitation by splitting nodes with the guidance of both labeled and unlabeled data. In particular, we derive a nonparametric algorithm to obtain an accurate quality measure of splitting by incorporating abundant unlabeled data. To avoid the curse of dimensionality, we project the data points from the original high-dimensional feature space onto a low-dimensional subspace before estimation. A unified optimization framework is proposed to select a coupled pair of subspace and separating hyper plane such that the smoothness of the subspace and the quality of the splitting are guaranteed simultaneously. The proposed algorithm is compared with state-of-the-art supervised and semi-supervised algorithms for typical computer vision applications such as object categorization and image segmentation. Experimental results on publicly available datasets demonstrate the superiority of our method. Xiao Liu 0012, Mingli Song, Dacheng Tao, Zicheng Liu 0001, Chun Chen 0001, Jiajun Bu |
CVPR | 4 |
| 2013 | HON4D: Histogram of Oriented 4D Normals for Activity Recognition from Depth SequencesabstractWe present a new descriptor for activity recognition from videos acquired by a depth sensor. Previous descriptors mostly compute shape and motion features independently, thus, they often fail to capture the complex joint shape-motion cues at pixel-level. In contrast, we describe the depth sequence using a histogram capturing the distribution of the surface normal orientation in the 4D space of time, depth, and spatial coordinates. To build the histogram, we create 4D projectors, which quantize the 4D space and represent the possible directions for the 4D normal. We initialize the projectors using the vertices of a regular polychoron. Consequently, we refine the projectors using a discriminative density measure, such that additional projectors are induced in the directions where the 4D normals are more dense and discriminative. Through extensive experiments, we demonstrate that our descriptor better captures the joint shape-motion cues in the depth sequence, and thus outperforms the state-of-the-art on all relevant benchmarks. Omar Oreifej, Zicheng Liu 0001 |
CVPR | 2 |
| 2013 | Probabilistic Graphlet Cut: Exploiting Spatial Structure Cue for Weakly Supervised Image SegmentationabstractWeakly supervised image segmentation is a challenging problem in computer vision field. In this paper, we present a new weakly supervised image segmentation algorithm by learning the distribution of spatially structured super pixel sets from image-level labels. Specifically, we first extract graph lets from each image where a graph let is a small-sized graph consisting of super pixels as its nodes and it encapsulates the spatial structure of those super pixels. Then, a manifold embedding algorithm is proposed to transform graph lets of different sizes into equal-length feature vectors. Thereafter, we use GMM to learn the distribution of the post-embedding graph lets. Finally, we propose a novel image segmentation algorithm, called graph let cut, that leverages the learned graph let distribution in measuring the homogeneity of a set of spatially structured super pixels. Experimental results show that the proposed approach outperforms state-of-the-art weakly supervised image segmentation methods, and its performance is comparable to those of the fully supervised segmentation models. Mingli Song, Zicheng Liu 0001, Xiao Liu 0012, Jiajun Bu, Chun Chen 0001 |
CVPR | 3 |
| 2013 | Image-to-Class Dynamic Time Warping for 3D hand gesture recognitionabstract3D Human Computer Interaction (HCI) becomes more and more popular thanks to the emergence of commercial depth cameras. Moreover, hand gestures provide a natural and attractive alternative to cumbersome interface devices for HCI. In this paper, we present an Image-to-Class Dynamic Time Warping (I2C-DTW) approach for 3D hand gesture recognition. Themain idea is that we divide the time-series curve of a 3D hand gesture into various finger combinations, called `fingerlets', which can either be learned or be set manually to represent each gesture and to capture inter-class variations. Furthermore, the I2C-DTW approach searches for the minimal path to warp two fingerlets, which are fromone test image and the specific class, respectively. Then the gesture recognition is to use the ensemble of multiple image-to-class DTW distance of fingerlets to obtain better performance. The proposed approach is evaluated on two 3D hand gesture datasets and the experiment results show that the proposed I2C-DTW approach significantly improves the recognizing performance. Hong Cheng 0002, Zhongjun Dai, Zicheng Liu 0001 |
ICME | 3 |
| 2013 | Sparse representation and learning in visual recognition: Theory and applications
Hong Cheng 0002, Zicheng Liu 0001, Lu Yang 0002, Xue-wen Chen 0001 |
Signal Process. | 2 |
| 2013 | Action Search by Example Using Randomized Visual VocabulariesabstractBecause actions can be small video objects, it is a challenging problem to search for similar actions in crowded and dynamic scenes when a single query example is provided. We propose a fast action search method that can efficiently locate similar actions spatiotemporally. Both the query action and the video datasets are characterized by spatio-temporal interest points. Instead of using a unified visual vocabulary to index all interest points in the database, we propose randomized visual vocabularies to enable fast and robust interest point matching. To accelerate action localization, we have developed a coarse-to-fine video subvolume search scheme, which is several orders of magnitude faster than the existing spatio-temporal branch and bound search. Our experiments on cross-dataset action search show promising results when compared with the state of the arts. Additional experiments on a 5-h versatile video dataset validate the efficiency of our method, where an action search can be finished in just 37.6 s on a regular desktop machine. Gang Yu 0002, Junsong Yuan 0001, Zicheng Liu 0001 |
IEEE Trans. Image Process. | 3 |
| 2012 | Human Activity Recognition with 2D and 3D Cameras
Zicheng Liu 0001 |
CIARP | 1 |
| 2012 | STOP: Space-Time Occupancy Patterns for 3D Action Recognition from Depth Map Sequences
Antônio Wilson Vieira, Erickson R. Nascimento, Gabriel L. Oliveira, Zicheng Liu 0001, Mario Fernando Montenegro Campos |
CIARP | 4 |
| 2012 | Mining actionlet ensemble for action recognition with depth camerasabstractHuman action recognition is an important yet challenging task. The recently developed commodity depth sensors open up new possibilities of dealing with this problem but also present some unique challenges. The depth maps captured by the depth cameras are very noisy and the 3D positions of the tracked joints may be completely wrong if serious occlusions occur, which increases the intra-class variations in the actions. In this paper, an actionlet ensemble model is learnt to represent each action and to capture the intra-class variance. In addition, novel features that are suitable for depth data are proposed. They are robust to noise, invariant to translational and temporal misalignments, and capable of characterizing both the human motion and the human-object interactions. The proposed approach is evaluated on two challenging action recognition datasets captured by commodity depth cameras, and another dataset captured by a MoCap system. The experimental evaluations show that the proposed approach achieves superior performance to the state of the art algorithms. Jiang Wang 0001, Zicheng Liu 0001, Ying Wu 0001, Junsong Yuan 0001 |
CVPR | 2 |
| 2012 | Robust 3D Action Recognition with Random Occupancy Patterns
Jiang Wang 0001, Zicheng Liu 0001, Jan Chorowski, Zhuoyuan Chen, Ying Wu 0001 |
ECCV (2) | 2 |
| 2012 | Propagative Hough Voting for Human Activity Recognition
Gang Yu 0002, Junsong Yuan 0001, Zicheng Liu 0001 |
ECCV (3) | 3 |
| 2012 | A Pyramid Nearest Neighbor Search Kernel for object categorization
Hong Cheng 0002, Rongchao Yu, Zicheng Liu 0001, Yiguang Liu |
ICPR | 3 |
| 2012 | Sparsity-based online missing sensor data recoveryabstractIn sensor networks, due to power outage at a sensor node, hardware dysfunction, or bad environmental conditions, not all sensor samples can be successfully gathered at the sink. Additionally, in the data stream scenario, some nodes may continually miss samples for a period of time. In this paper, a sparsity-based online data recovery approach is proposed. We construct an over complete dictionary composed of past data frames and traditional fixed transform bases. Assuming the current frame can be sparsely represented using only a few elements of the dictionary, missing samples in each frame can be estimated by Basis Pursuit. Our method was tested on data from a real sensor network application: monitoring the temperatures of the disk drive racks at a data center. Simulations show that in terms of estimation accuracy and stability, the proposed approach outperforms existing average-based interpolation methods, and is more robust to burst missing along the time dimension. Di Guo 0003, Xiaobo Qu 0001, Lianfen Huang, Zicheng Liu 0001, Ming-Ting Sun |
ISCAS | 5 |
| 2012 | Predicting human activities using spatio-temporal structure of interest pointsabstractEarly recognition and prediction of human activities are of great importance in video surveillance, e.g., by recognizing a criminal activity at its beginning stage, it is possible to avoid unfortunate outcomes. We address early activity recognition by developing a Spatial-Temporal Implicit Shape Model (STISM), which characterizes the space-time structure of the sparse local features extracted from a video. The early recognition of human activities is accomplished by pattern matching through STISM. To enable efficient and robust matching, we propose a new random forest structure, called multi-class balanced random forest, which makes a good trade-off between the balance of the trees and the discriminative abilities. The prediction is done simultaneously for multiple classes, which saves both the memory and computational cost. The experiments show that our algorithm significantly outperforms the state of the arts for the human activity prediction problem. Gang Yu 0002, Junsong Yuan 0001, Zicheng Liu 0001 |
ACM Multimedia | 3 |
| 2012 | Multi-support-region image descriptors and its application to street landmark localization
Hong Cheng 0002, Zicheng Liu 0001, Jie Yang 0001 |
Mach. Vis. Appl. | 2 |
| 2012 | Hierarchical Filtered Motion for Action Recognition in Crowded VideosabstractAction recognition with cluttered and moving background is a challenging problem. One main difficulty lies in the fact that the motion field in an action region is contaminated by the background motions. We propose a hierarchical filtered motion (HFM) method to recognize actions in crowded videos by the use of motion history image (MHI) as basic representations of motion because of its robustness and efficiency. First, we detect interest points as the two-dimensional Harris corners with recent motion, e.g., locations with high intensities in the MHI. Then, a global spatial motion smoothing filter is applied to the gradients of the MHI to eliminate isolated unreliable or noisy motions. At each interest point, a local motion field filter is applied to the smoothed gradients of the MHI by computing structure proximity between any pixel in the local region and the interest point. Thus, the motion at a pixel is enhanced or weakened based on its structure proximity with the interest point. To validate its effectiveness, we characterize the spatial and temporal features by histograms of oriented gradient in the intensity image and the MHI, respectively, and use a Gaussian-mixture-model-based classifier for action recognition. The performance of the proposed approach achieves the state-of-the-art results on the KTH dataset that has clean background. More importantly, we perform cross-dataset action classification and detection experiments, where the KTH dataset is used for training, while the microsoft research (MSR) action dataset II that consists of crowded videos with people moving in the background is used for testing. Our experiments show that the proposed HFM method significantly outperforms existing techniques. Yingli Tian, Liangliang Cao, Zicheng Liu 0001, Zhengyou Zhang |
IEEE Trans. Syst. Man Cybern. Part C | 3 |
| 2011 | Unsupervised random forest indexing for fast action searchabstractDespite recent successes of searching small object in images, it remains a challenging problem to search and locate actions in crowded videos because of (1) the large variations of human actions and (2) the intensive computational cost of searching the video space. To address these challenges, we propose a fast action search and localization method that supports relevance feedback from the user. By characterizing videos as spatio-temporal interest points and building a random forest to index and match these points, our query matching is robust and efficient. To enable efficient action localization, we propose a coarse-to-fine sub-volume search scheme, which is several orders faster than the existing video branch and bound search. The challenging cross-dataset search of several actions validates the effectiveness and efficiency of our method. Gang Yu 0002, Junsong Yuan 0001, Zicheng Liu 0001 |
CVPR | 3 |
| 2011 | Real world activity summary for senior home monitoringabstractFrom a senior person's daily activities, one can tell a lot about the health condition of the senior person. Thus we believe that senior home activity analysis will play an important role in the health care of senior people. Toward this goal, we propose a senior home activity summary system. One challenging problem in such a real world application is that senior's activities are usually accompanied by nurse's walking. It is impractical to predefine and label all the potential activities of all the potential visitors. To address this problem, we propose a novel feature filtering technique to reduce or eliminate the effects of the interest points that belong to other people. To evaluate the proposed activity summary system, we have collected a senior home activity dataset (SAR), and performed activity recognition for eating and walking classes. The experimental results show that the proposed system provides quite accurate activity summaries for a real world application scenario. Hong Cheng 0002, Zicheng Liu 0001, Yang Zhao 0024, Guo Ye |
ICME | 2 |
| 2011 | Real-time human action search using random forest based hough votingabstractMany existing techniques in content based video retrieval treat a video sequence as a whole to match it against a query video or to assign a text label. Such an approach has serious limitations when applied to human action retrieval because an action may occur only in a sub-region and last for a small portion of the video length. In situations like this, we essentially need to match the subvolumes of the video sequences against the query video. A naive exhaustive search is impractical due to large number of possible subvolumes for each video sequence. In this paper, we propose a novel framework for action retrieval which performs pattern matching at subvolume level and is very efficient in handling large corpus of videos. We construct an unsupervised random forest to index the video database, generate a score volume with Hough voting and then employ a max sub-path strategy to quickly search for the temporal and spatial positions of all the video sequences in the database. We present action search experiments on challenging datasets to validate the efficiency and effectiveness of our system. Gang Yu 0002, Junsong Yuan 0001, Zicheng Liu 0001 |
ACM Multimedia | 3 |
| 2011 | Discriminative Video Pattern Search for Efficient Action DetectionabstractActions are spatiotemporal patterns. Similar to the sliding window-based object detection, action detection finds the reoccurrences of such spatiotemporal patterns through pattern matching, by handling cluttered and dynamic backgrounds and other types of action variations. We address two critical issues in pattern matching-based action detection: 1) the intrapattern variations in actions, and 2) the computational efficiency in performing action pattern search in cluttered scenes. First, we propose a discriminative pattern matching criterion for action classification, called naive Bayes mutual information maximization (NBMIM). Each action is characterized by a collection of spatiotemporal invariant features and we match it with an action class by measuring the mutual information between them. Based on this matching criterion, action detection is to localize a subvolume in the volumetric video space that has the maximum mutual information toward a specific action class. A novel spatiotemporal branch-and-bound (STBB) search algorithm is designed to efficiently find the optimal solution. Our proposed action detection method does not rely on the results of human detection, tracking, or background subtraction. It can handle action variations such as performing speed and style variations as well as scale changes well. It is also insensitive to dynamic and cluttered backgrounds and even to partial occlusions. The cross-data set experiments on action detection, including KTH, CMU action data sets, and another new MSR action data set, demonstrate the effectiveness and efficiency of the proposed multiclass multiple-instance action detection method. Junsong Yuan 0001, Zicheng Liu 0001, Ying Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2011 | Introduction to the ICME2010 Special IssueabstractThe 15 papers in this special issue are extended versions of papers presented at the 2010 IEEE International Conference on Multimedia and Expo (ICME), held in Singapore on July 19-23, 2010. These papers cover a wide range of topics in multimedia including user interface, content understanding, mobility, 3-D processing, storage, and forensics. Zicheng Liu 0001, Ming-Ting Sun, Chia-Wen Lin, Zhengyou Zhang, Zhu Liu 0001, Homer H. Chen, Yap-Peng Tan, Oscar C. Au |
IEEE Trans. Multim. | 1 |
| 2011 | Fast Action Detection via Discriminative Random Forest Voting and Top-K Subvolume SearchabstractMulticlass action detection in complex scenes is a challenging problem because of cluttered backgrounds and the large intra-class variations in each type of actions. To achieve efficient and robust action detection, we characterize a video as a collection of spatio-temporal interest points, and locate actions via finding spatio-temporal video subvolumes of the highest mutual information score towards each action class. A random forest is constructed to efficiently generate discriminative votes from individual interest points, and a fast top-K subvolume search algorithm is developed to find all action instances in a single round of search. Without significantly degrading the performance, such a top-K search can be performed on down-sampled score volumes for more efficient localization. Experiments on a challenging MSR Action Dataset II validate the effectiveness of our proposed multiclass action detection method. The detection speed is several orders of magnitude faster than existing methods. Gang Yu 0002, Norberto A. Goussies, Junsong Yuan 0001, Zicheng Liu 0001 |
IEEE Trans. Multim. | 4 |
| 2010 | Cross-dataset action detectionabstractIn recent years, many research works have been carried out to recognize human actions from video clips. To learn an effective action classifier, most of the previous approaches rely on enough training labels. When being required to recognize the action in a different dataset, these approaches have to re-train the model using new labels. However, labeling video sequences is a very tedious and time-consuming task, especially when detailed spatial locations and time durations are required. In this paper, we propose an adaptive action detection approach which reduces the requirement of training labels and is able to handle the task of cross-dataset action detection with few or no extra training labels. Our approach combines model adaptation and action detection into a Maximum a Posterior (MAP) estimation framework, which explores the spatial-temporal coherence of actions and makes good use of the prior information which can be obtained without supervision. Our approach obtains state-of-the-art results on KTH action dataset using only 50% of the training labels in tradition approaches. Furthermore, we show that our approach is effective for the cross-dataset detection which adapts the model trained on KTH to two other challenging datasets. Liangliang Cao, Zicheng Liu 0001, Thomas S. Huang |
CVPR | 2 |
| 2010 | Action detection using multiple spatial-temporal interest point featuresabstractThis paper considers the problem of detecting actions from cluttered videos. Compared with the classical action recognition problem, this paper aims to estimate not only the scene category of a given video sequence, but also the spatial-temporal locations of the action instances. In recent years, many feature extraction schemes have been designed to describe various aspects of actions. However, due to the difficulty of action detection, e.g., the cluttered background and potential occlusions, a single type of features cannot solve the action detection problems perfectly in cluttered videos. In this paper, we attack the detection problem by combining multiple Spatial-Temporal Interest Point (STIP) features, which detect salient patches in the video domain, and describe these patches by feature of local regions. The difficulty of combining multiple STIP features for action detection is two folds: First, the number of salient patches detected by different STIP methods varies across different salient patches. How to combine such features is not considered by existing fusion methods. Second, the detection in the videos should be efficient, which excludes many slow machine learning algorithms. To handle these two difficulties, we propose a new approach which combines Gaussian Mixture Model with Branch-and-Bound search to efficiently locate the action of interest. We build a new challenging dataset for our action detection task, and our algorithm obtains impressive results. On classical KTH dataset, our method outperforms the state-of-the-art methods. Liangliang Cao, Yingli Tian, Zicheng Liu 0001, Benjamin Z. Yao, Zhengyou Zhang, Thomas S. Huang |
ICME | 3 |
| 2010 | Learning feature transforms for object detection from panoramic imagesabstractWe present a novel technique to detect objects from panoramic images using existing object detectors trained from perspective images. By leveraging existing object detectors, we save the cost of training a new detector which requires tedious and time consuming training data collection and labeling. The core of our technique is learning a feature transform which is represented by Gaussian Process Regression (GPR). Feature vectors computed directly from panoramic images are transformed into new feature vectors in such a way that the existing classifier has much better detection rate on the transformed feature vectors. Our feature transform has the interesting property that it not only corrects for the geometric distortions resulted from panoramic imaging process, but also corrects for the pose mismatches between the objects on the panoramic images and those on the training images. Our experiments show that we are able to successfully apply an existing car detector trained on perspective images to panoramic images which have both geometric distortions and larger pose variations. Hong Cheng 0002, Zicheng Liu 0001, Jie Yang 0001 |
ICME | 2 |
| 2010 | Efficient search of Top-K video subvolumes for multi-instance action detectionabstractAction detection was formulated as a subvolume mutual information maximization problem in, where each subvolume identifies where and when the action occurs in the video. Despite the fact that the proposed branch-and-bound algorithm can find the best subvolume efficiently for low resolution videos, it is still not efficient enough to perform multi-instance detection in videos of high spatial resolution. In this paper we develop an algorithm that further speeds up the subvolume search and targets on real-time multi-instance action detection for high resolution videos (e.g. 320 × 240 or higher). Unlike the previous branch-and-bound search technique which restarts a new search for each action instance, we find the Top-K subvolumes simultaneously with a single round of search. To handle the larger spatial resolution, we downsample the volume of videos for a more efficient upper-bound estimation. To validate our algorithm, we perform experiments on a challenging dataset of 54 video sequences where each video consists of several actions performed by different people in a crowded environment. The experiments show that our method is not only efficient, but also capable of handling action variations caused by performing speed and style changes, spatial scale changes, as well as cluttered and moving background. Norberto A. Goussies, Zicheng Liu 0001, Junsong Yuan 0001 |
ICME | 2 |
| 2010 | Image Ratio Features for Facial Expression Recognition ApplicationabstractVideo-based facial expression recognition is a challenging problem in computer vision and human-computer interaction. To target this problem, texture features have been extracted and widely used, because they can capture image intensity changes raised by skin deformation. However, existing texture features encounter problems with albedo and lighting variations. To solve both problems, we propose a new texture feature called image ratio features. Compared with previously proposed texture features, e.g., high gradient component features, image ratio features are more robust to albedo and lighting variations. In addition, to further improve facial expression recognition accuracy based on image ratio features, we combine image ratio features with facial animation parameters (FAPs), which describe the geometric motions of facial feature points. The performance evaluation is based on the Carnegie Mellon University Cohn-Kanade database, our own database, and the Japanese Female Facial Expression database. Experimental results show that the proposed image ratio feature is more robust to albedo and lighting variations, and the combination of image ratio features and FAPs outperforms each feature alone. In addition, we study asymmetric facial expressions based on our own facial expression database and demonstrate the superior performance of our combined expression recognition system. Mingli Song, Dacheng Tao, Zicheng Liu 0001, Xuelong Li 0001, MengChu Zhou |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2009 | Efficient Scale-Space Spatiotemporal Saliency Tracking for Distortion-Free Video Retargeting
Gang Hua 0001, Cha Zhang, Zicheng Liu 0001, Zhengyou Zhang, Ying Shan |
ACCV (2) | 3 |
| 2009 | Discriminative subvolume search for efficient action detectionabstractActions are spatio-temporal patterns which can be characterized by collections of spatio-temporal invariant features. Detection of actions is to find the re-occurrences (e.g. through pattern matching) of such spatio-temporal patterns. This paper addresses two critical issues in pattern matching-based action detection: (1) efficiency of pattern search in 3D videos and (2) tolerance of intra-pattern variations of actions. Our contributions are two-fold. First, we propose a discriminative pattern matching called naive-Bayes based mutual information maximization (NBMIM) for multi-class action categorization. It improves the state-of-the-art results on standard KTH dataset. Second, a novel search algorithm is proposed to locate the optimal subvolume in the 3D video space for efficient action detection. Our method is purely data-driven and does not rely on object detection, tracking or background subtraction. It can well handle the intra-pattern variations of actions such as scale and speed variations, and is insensitive to dynamic and clutter backgrounds and even partial occlusions. The experiments on versatile datasets including KTH and CMU action datasets demonstrate the effectiveness and efficiency of our method. Junsong Yuan 0001, Zicheng Liu 0001, Ying Wu 0001 |
CVPR | 2 |
| 2009 | Sparsity induced similarity measure for label propagationabstractGraph-based semi-supervised learning has gained considerable interests in the past several years thanks to its effectiveness in combining labeled and unlabeled data through label propagation for better object modeling and classification. A critical issue in constructing a graph is the weight assignment where the weight of an edge specifies the similarity between two data points. In this paper, we present a novel technique to measure the similarities among data points by decomposing each data point as an L1sparse linear combination of the rest of the data points. The main idea is that the coefficients in such a sparse decomposition reflect the point's neighborhood structure thus providing better similarity measures among the decomposed data point and the rest of the data points. The proposed approach is evaluated on four commonly-used data sets and the experimental results show that the proposed Sparsity Induced Similarity (SIS) measure significantly improves label propagation performance. As an application of the SIS-based label propagation, we show that the SIS measure can be used to improve the Bag-of-Words approach for scene classification. Hong Cheng 0002, Zicheng Liu 0001, Jie Yang 0001 |
ICCV | 2 |
| 2009 | Optimal joint linear acoustic echo cancelation and blind source separation in the presence of loudspeaker nonlinearityabstractAcoustic echoes represent a major source of discomfort in hands free, full-duplex, communication systems. The problem becomes particularly difficult when the loudspeakers are nonlinear as considered in this paper. In contrast to the single-microphone linear and nonlinear acoustic echo cancellation techniques, we take advantage of the spatial diversity offered by the microphone arrays. Indeed, having a set of microphones and multiple sources (i.e., the near and far ends) that can be active at the same time, this problem can be solved using a blind source separation (BSS) algorithm. The performance of the BSS can be further improved when combined with a linear acoustic echo canceler (LAEC). In this paper, we study the potentials of joint BSS and LAEC to cancel the echo signals in two schemes. In the first scheme, the BSS is deployed as a front-end and has a twofold function: reducing the acoustic echo and creating a linearly transformed echo reference that is used by the LAEC as a post-processor. In the second scheme, the BSS operates on multiple LAECs outputs to further reduce the residual echo from the target signal. We show that the first scheme outperforms the second one. Mehrez Souden, Zicheng Liu 0001 |
ICME | 2 |
| 2009 | Implicit Surface Reconstruction with an Analogy of Polar Field Model
Yuxu Lin, Chun Chen 0001, Mingli Song, Jiajun Bu, Zicheng Liu 0001 |
PSIVT | 5 |
| 2009 | Face Relighting from a Single Image under Arbitrary Unknown Lighting ConditionsabstractIn this paper, we present a new method to modify the appearance of a face image by manipulating the illumination condition, when the face geometry and albedo information is unknown. This problem is particularly difficult when there is only a single image of the subject available. Recent research demonstrates that the set of images of a convex Lambertian object obtained under a wide variety of lighting conditions can be approximated accurately by a low-dimensional linear subspace using a spherical harmonic representation. Moreover, morphable models are statistical ensembles of facial properties such as shape and texture. In this paper, we integrate spherical harmonics into the morphable model framework by proposing a 3D spherical harmonic basis morphable model (SHBMM). The proposed method can represent a face under arbitrary unknown lighting and pose simply by three low-dimensional vectors, i.e., shape parameters, spherical harmonic basis parameters, and illumination coefficients, which are called the SHBMM parameters. However, when the image was taken under an extreme lighting condition, the approximation error can be large, thus making it difficult to recover albedo information. In order to address this problem, we propose a subregion-based framework that uses a Markov random field to model the statistical distribution and spatial coherence of face texture, which makes our approach not only robust to extreme lighting conditions, but also insensitive to partial occlusions. The performance of our framework is demonstrated through various experimental results, including the improved rates for face recognition under extreme lighting conditions. Yang Wang 0001, Lei Zhang 0002, Zicheng Liu 0001, Gang Hua 0001, Zhengyou Zhang, Dimitris Samaras |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2009 | Active Lighting for Video ConferencingabstractIn consumer video conferencing, lighting conditions are usually not ideal thus the image qualities are poor. Lighting affects image quality on two aspects: brightness and skin tone. While there has been much research on improving the brightness of the captured images including contrast enhancement and noise removal (which can be thought of as components for brightness improvement), little attention has been paid to the skin tone aspect. In contrast, it is a common knowledge for professional stage lighting designers that lighting affects not only the brightness but also the color tone which plays a critical role in the perceived look of the host and the mood of the stage scene. Inspired by stage lighting design, we propose an active lighting system which automatically adjusts the lighting so that the image looks visually appealing. The system consists of computer controllable light emitting diode light sources of different colors so that it improves not only the brightness but also the skin tone of the face. Given that there is no quantitative formula on what makes a good skin tone, we use a data driven approach to learn a good skin tone model from a collection of photographs taken by professional photographers. We have developed a working system and conducted user studies to validate our approach. Mingxuan Sun 0001, Zicheng Liu 0001, Jingyu Qiu, Zhengyou Zhang, Mike Sinclair |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2009 | Dual-RBF based surface reconstruction
Yuxu Lin, Chun Chen 0001, Mingli Song, Zicheng Liu 0001 |
Vis. Comput. | 4 |
| 2008 | A deformable local image descriptorabstractThis paper presents a novel local image descriptor that is robust to general image deformations. A limitation with traditional image descriptors is that they use a single support region for each interest point. For general image deformations, the amount of deformation for each location varies and is unpredictable such that it is difficult to choose the best scale of the support region. To overcome this difficulty, we propose to use multiple support regions of different sizes surrounding an interest point. A feature vector is computed for each support region, and the concatenation of these feature vectors forms the descriptor for this interest point. Furthermore, we propose a new similarity measure model, Local-to-Global Similarity (LGS) model, for point matching that takes advantage of the multi-size support regions. Each support region acts as a ‘weak’ classifier and the weights of these classifiers are learned in an unsupervised manner. The proposed approach is evaluated on a number of images with real and synthetic deformations. The experiment results show that our method outperforms existing techniques under different deformations. Hong Cheng 0002, Zicheng Liu 0001, Nanning Zheng 0001, Jie Yang 0001 |
CVPR | 2 |
| 2008 | Blind source separation in a distributed microphone meeting environment for improved teleconferencingabstractFrom an audio perspective, the present state of teleconferencing technology leaves something to be desired; speaker overlap is one of the causes of this inadequate performance. To that end, this paper presents a frequency-domain implementation of convolutive BSS specifically designed for the nature of the teleconferencing environment. In addition to presenting a novel depermutation scheme, this paper presents a least-squares post-processing scheme, which exploits segments during which only a subset of all speakers are active. Experiments with simulated and real data demonstrate the ability of the proposed methods to provide SIRs at or near that of the adaptive noise cancellation (ANC) solution which is obtained under idealistic assumptions that the ANC filters are adapted with one source being on at a time. Jacek Dmochowski, Zicheng Liu 0001, Philip A. Chou |
ICASSP | 2 |
| 2008 | Requirements and recommendations for an enhanced meeting viewing experienceabstractWe have found that viewing recorded meetings using traditional meeting viewers whose interfaces consist of an automatic speaker and a fixed context view does not provide sufficient information and control to the users. In particular, a survey of users who watch meeting recordings on a regular basis revealed that it is also useful to provide (1) speaker-related information, including who the speaker is talking to, looking at, and being interrupted by, and (2) more control of the interface, including changing the relative sizes of the speaker and context views and navigating within the context view. We present a 3D interface prototype designed specifically to meet these requirements when viewing recorded meetings. We describe in detail the results of a user study comparing the effectiveness of the new and traditional style interfaces with respect to these requirements. Based on this study, we present a set of guidelines for future interfaces. Sasa Junuzovic, Rajesh Hegde, Zhengyou Zhang, Philip A. Chou, Zicheng Liu 0001, Cha Zhang |
ACM Multimedia | 5 |
| 2008 | Graphical modeling and decoding of human actionsabstractThis paper presents a graphical model for learning and recognizing human actions. Specifically, we propose to encode actions in a weighted directed graph, referred to as action graph, where nodes of the graph represent salient postures that are used to characterize the actions and shared by all actions. The weight between two nodes measures the transitional probability between the two postures. An action is encoded as one or multiple paths in the action graph. The salient postures are modeled using Gaussian Mixture Models (GMM). Both the salient postures and action graph are automatically learned from training samples through unsupervised clustering and expectation and maximization (EM) algorithm. Experimental results have verified the performance of the proposed model, its tolerance to noise and viewpoints and its robustness across different subjects and datasets. Wanqing Li 0001, Zhengyou Zhang, Zicheng Liu 0001 |
MMSP | 3 |
| 2008 | Semantic saliency driven camera control for personal remote collaborationabstractThis paper presents a camera combo system for personal remote collaboration applications. The system consists of two different cameras. One camera has a wide field of view, and the other can pan/tilt/zoom (PTZ) based on analysis of the images captured by the wide angle camera. Unlike traditional approaches which usually drive the PTZ camera to follow the person or his/her head, our system is capable of capturing general objects of interest in remote collaboration. For instance, when the user raises something trying to show it to the remote person, our system will automatically position the PTZ camera to zoom in at the object. At the core of our system is a semantic saliency map that overcomes many limitations of low-level saliency maps computed from preliminary image features. We demonstrate how such a semantic saliency map can be computed through contextual analysis, sign analysis and transitional analysis, and how it can be used for PTZ camera control with a novel information loss optimization based virtual director. The effectiveness of the proposed method is demonstrated with real-world sequences. Cha Zhang, Zicheng Liu 0001, Zhengyou Zhang |
MMSP | 2 |
| 2008 | Multisensory processing for speech enhancement and magnitude-normalized spectra for speech modeling
Amarnag Subramanya, Zhengyou Zhang, Zicheng Liu 0001, Alex Acero |
Speech Commun. | 3 |
| 2008 | Expandable Data-Driven Graphical Modeling of Human Actions Based on Salient PosturesabstractThis paper presents a graphical model for learning and recognizing human actions. Specifically, we propose to encode actions in a weighted directed graph, referred to as action graph, where nodes of the graph represent salient postures that are used to characterize the actions and are shared by all actions. The weight between two nodes measures the transitional probability between the two postures represented by the two nodes. An action is encoded asoneor multiple paths in the action graph. The salient postures are modeled using Gaussian mixture models (GMMs). Both the salient postures and action graph are automatically learned from training samples through unsupervised clustering and expectation and maximization (EM) algorithm. The proposed action graph not only performs effective and robust recognition of actions, but it can also be expanded efficiently with new actions. An algorithm is also proposed for adding a new action to a trained action graph without compromising the existing action graph. Extensive experiments on widely used and challenging data sets have verified the performance of the proposed methods, its tolerance to noise and viewpoints, its robustness across different subjects and data sets, as well as the effectiveness of the algorithm for learning new actions. Wanqing Li 0001, Zhengyou Zhang, Zicheng Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2007 | Face Re-Lighting from a Single Image under Harsh Lighting ConditionsabstractIn this paper, we present a new method to change the illumination condition of a face image, with unknown face geometry and albedo information. This problem is particularly difficult when there is only one single image of the subject available and it was taken under a harsh lighting condition. Recent research demonstrates that the set of images of a convex Lambertian object obtained under a wide variety of lighting conditions can be approximated accurately by a low-dimensional linear subspace using spherical harmonic representation. However, the approximation error can be large under harsh lighting conditions thus making it difficult to recover albedo information. In order to address this problem, we propose a subregion based framework that uses a Markov Random Field to model the statistical distribution and spatial coherence of face texture, which makes our approach not only robust to harsh lighting conditions, but insensitive to partial occlusions as well. The performance of our framework is demonstrated through various experimental results, including the improvement to the face recognition rate under harsh lighting conditions. Yang Wang 0001, Zicheng Liu 0001, Gang Hua 0001, Zhengyou Zhang, Dimitris Samaras |
CVPR | 2 |
| 2007 | Energy-Based Sound Source Localization and Gain Normalization for Ad Hoc Microphone ArraysabstractWe present an energy-based technique to estimate both microphone and speaker/talker locations from an ad hoc network of microphones. An example of such ad hoc microphone network is a set of microphones built in the laptops that some meeting participants bring in a meeting room. Compared with traditional sound source localization approaches based on time of flight, our technique does not require accurate synchronization, and it does not require each laptop to emit special signals. We estimate the meeting participants' positions based on average energies of their speech signals. In addition, we present a technique, which is independent of the volumes of the speakers, to estimate the relative gains of the microphones. This is crucial to aggregate various audio channels from the ad hoc microphone network into a single stream for audio conferencing. Zicheng Liu 0001, Zhengyou Zhang, Li-wei He, Philip A. Chou |
ICASSP (2) | 1 |
| 2007 | Enhancing a Driver's Situation Awareness using a Global View MapabstractThis paper proposes a novel method to enhance a driver's situation awareness by dynamically providing a global view of surroundings for the driver. The surroundings of a vehicle are captured by an omni-directional vision system mounted on the top of the vehicle. The video stream from the camera is processed to detect nearby vehicles. Positions of these detected objects are overlaid on a global view of a local map (e.g., an aerial imagery or satellite imagery map). We establish the relationship between the omni-directional vision system and the global view map. The global view map dynamically provides a realistic perspective view of the driving environment. This map can be projected onto an HUD on the windshield. By looking at the display, a driver can have a global picture of the situation and potentially produce a good driving strategy. We illustrate the proposed method by dynamically mapping a video stream onto Google Earth map. Hong Cheng 0002, Zicheng Liu 0001, Nanning Zheng 0001, Jie Yang 0001 |
ICME | 2 |
| 2007 | Learning-Based Perceptual Image Quality Improvement for Video ConferencingabstractIt is well known that in professional TV show filming, stage lighting has to be carefully designed in order to make the host and the scene look visually appealing. The lighting affects not only the brightness but also the color tone which plays a critical role in the perceived look of the host and the mood of the stage. In contrast, during video conferencing, the lighting is usually far from ideal thus the perceived image quality is low. There has been a lot of research on improving the brightness of the captured images. In this paper, we propose a learning-based technique to improve the perceptual image quality by enhancing both brightness and color tone. The basic idea is to learn the color statistics from a training set of images which look visually appealing, and adjust the color of an input image so that its color statistics matches those in the training set. To validate our approach, we have conducted user study and the results show that our technique significantly improves the perceived image quality. Zicheng Liu 0001, Cha Zhang, Zhengyou Zhang |
ICME | 1 |
| 2007 | Head-Size Equalization for Improved Visual Perception in Video ConferencingabstractIn a video conferencing setting, people often use an elongated meeting table with the major axis along the camera direction. A standard wide-angle perspective image of this setting creates significant foreshortening, thus the people sitting at the far end of the table appear very small relative to those nearer the camera. This has two consequences. First, it is difficult for the remote participants to see the faces of those at the far end, thus affecting the experience of the video conferencing. Second, it is a waste of the screen space and network bandwidth because most of the pixels are used on the background instead of on the faces of the meeting participants. In this paper, we present a novel technique, called Spatially-Varying-Uniform scaling functions, to warp the images to equalize the head sizes of the meeting participants without causing undue distortion. This technique works for both the 180-degree views where the camera is placed at one end of the table and the 360-degree views where the camera is placed at the center of the table. We have implemented this algorithm on two types of camera arrays: one with 180-degree view, and the other with 360-degree view. On both hardware devices, image capturing, stitching, and head-size equalization are run in real time. In addition, we have conducted user study showing that people clearly prefer head-size equalized images. Zicheng Liu 0001, Michael F. Cohen, Deepti Bhatnagar, Ross Cutler, Zhengyou Zhang |
IEEE Trans. Multim. | 1 |
| 2007 | A Generic Framework for Efficient 2-D and 3-D Facial Expression AnalogyabstractFacial expression analogy provides computer animation professionals with a tool to map expressions of an arbitrary source face onto an arbitrary target face. In the recent past, several algorithms have been presented in the literature that aim at putting the expression analogy paradigm into practice. Some of these methods exclusively handle expression mapping between 3-D face models, while others enable the transfer of expressions between images of faces only. None of them, however, represents a more general framework that can be applied to either of these two face representations. In this paper, we describe a novel generic method for analogy-based facial animation that employs the same efficient framework to transfer facial expressions between arbitrary 3-D face models, as well as between images of performer's faces. We propose a novel geometry encoding for triangle meshes, vertex-tent-coordinates, that enables us to formulate expression transfer in the 2-D and the 3-D case as a solution to a simple system of linear equations. Our experiments show that our method outperforms many previous analogy-based animation approaches in terms of achieved animation quality, computation time and generality. Mingli Song, Zhao Dong 0001, Christian Theobalt, Huiqiong Wang, Zicheng Liu 0001, Hans-Peter Seidel |
IEEE Trans. Multim. | 5 |
| 2006 | Real-Time Facial Expression Mapping for High Resolution 3D Meshes
Mingli Song, Zicheng Liu 0001, Baining Guo |
Computer Graphics International | 2 |
| 2006 | Automatic Business Card Scanning with a CameraabstractIn this paper, we present a system to automatically extract, rectify and enhance business card images. First the business card image patch is automatically segmented by minimizing a novel local-global variational energy. Second a quadrangle is fitted to the segmented image patch. With the four corner points of the quadrangle, we then estimate the physical aspect ratio of the business card and obtain a homography to rectify the quadrangle back to rectangular shape. We finally enhance the contrast of the rectified business card image using a S-shaped curve. Extensive experiments demonstrated the efficacy and robustness of our system. Gang Hua 0001, Zicheng Liu 0001, Zhengyou Zhang, Ying Wu 0001 |
ICIP | 2 |
| 2006 | Subtle Facial Expression Modeling with Vector Field DecompositionabstractFacial expression means not only the feature motions but also the subtle appearance changes in the face. The details are important to facial expression synthesis. In this paper, a novel method based on vector field decomposition is presented to model subtle facial expressions. The intensity change of facial expression image is represented by Helmholtz-Hodge vector field decomposition. Based on this method, subtle facial expression mapping is carried out much faster and better. Experiment results show that our method is robust and convincible to model the subtle facial expression. Mingli Song, Huiqiong Wang, Jiajun Bu, Chun Chen 0001, Zicheng Liu 0001 |
ICIP | 5 |
| 2006 | Speech Modelingwith Magnitude-Normalized Complex Spectra and Its Application to Multisensory Speech EnhancementabstractA good speech model is essential for speech enhancement, but it is very difficult to build because of huge intra- and extra-speaker variation. We present a new speech model for speech enhancement, which is based on statistical models of magnitude-normalized complex spectra of speech signals. Most popular speech enhancement techniques work in the spectrum space, but the large variation of speech strength, even from the same speaker, makes accurate speech modeling very difficult because the magnitude is correlated across all frequency bins. By performing magnitude normalization for each speech frame, we are able to get rid of the magnitude variation and to build a much better speech model with only a small number of Gaussian components. This new speech model is applied to speech enhancement for our previously developed microphone headsets that combine a conventional air microphone with a bone sensor. Much improved results have been obtained Amarnag Subramanya, Zhengyou Zhang, Zicheng Liu 0001, Alex Acero |
ICME | 3 |
| 2006 | Iterative Local-Global Energy Minimization for Automatic Extraction of Objects of InterestabstractWe propose a novel global-local variational energy to automatically extract objects of interest from images. Previous formulations only incorporate local region potentials, which are sensitive to incorrectly classified pixels during iteration. We introduce a global likelihood potential to achieve better estimation of the foreground and background models and, thus, better extraction results. Extensive experiments demonstrate its efficacy. Gang Hua 0001, Zicheng Liu 0001, Zhengyou Zhang, Ying Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2006 | Geometry-Driven Photorealistic Facial Expression SynthesisabstractExpression mapping (also called performance driven animation) has been a popular method for generating facial animations. A shortcoming of this method is that it does not generate expression details such as the wrinkles due to skin deformations. In this paper, we provide a solution to this problem. We have developed a geometry-driven facial expression synthesis system. Given feature point positions (the geometry) of a facial expression, our system automatically synthesizes a corresponding expression image that includes photorealistic and natural looking expression details. Due to the difficulty of point tracking, the number of feature points required by the synthesis system is, in general, more than what is directly available from a performance sequence. We have developed a technique to infer the missing feature point motions from the tracked subset by using an example-based approach. Another application of our system is expression editing where the user drags feature points while the system interactively generates facial expressions with skin deformation details. Qingshan Zhang, Zicheng Liu 0001, Baining Guo, Demetri Terzopoulos, Harry Shum |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2005 | Leakage Model and Teeth Clack Removal for Air- and Bone-Conductive Integrated MicrophonesabstractContinuing our previous work (Zhang et al. (2004), Liu et al. (2004)) on using air- and bone-conductive integrated microphones, and in particular on using the direct filtering approach (Liu et al. (2004)) for speech enhancement in noisy environments, we present in this paper a refined version of the direct filtering algorithm. The new algorithm takes into account explicitly the leakage of background noise into the bone channel. We also present a new algorithm that detects and removes an artifact known as teeth clacks. Experiments show that the addition of the above algorithms improves system performance to a large extent even in highly nonstationary noisy environments. Zicheng Liu 0001, Amarnag Subramanya, Zhengyou Zhang, Jasha Droppo, Alex Acero |
ICASSP (1) | 1 |
| 2005 | Automatic Head-size Equalization in Panorama Images for Video ConferencingabstractIn panorama images captured by omni-directional cameras during video conferencing, the image sizes of the people around the conference table are not uniform due to the varying distances to the camera. Spatially varying-uniform (SVU) scaling functions have been proposed to warp a panorama image smoothly such that the participants have similar sizes on the image. To generate the SVU function, one needs to segment the table boundaries, which was generated manually in the previous work. In this paper, we propose a robust algorithm to automatically segment the table boundaries. To ensure the robustness, we apply a symmetry voting scheme to filter out noisy points on the edge map. Trigonometry and quadratic fitting methods are developed to fit a continuous curve to the remaining edge points. We report experimental results on both synthetic and real images. Ya Chang, Ross Cutler, Zicheng Liu 0001, Zhengyou Zhang, Alex Acero, Matthew Turk 0001 |
ICME | 3 |
| 2005 | Head-size equalization for better visual perception of video conferencingabstractIn a video conferencing setting, people often use an elongated meeting table with the major axis along the camera direction. A standard wide angle perspective image of this setting creates significant foreshortening, thus the people sitting at the far end of the table appear very small relative to those nearer the camera. This has two consequences. First, it is difficult for the remote participants to see the faces of those at the far end, thus affecting the experience of the video conferencing. Second, it is a waste of the screen space and network bandwidth because most of the pixels are used on the background instead of on the faces of the meeting participants. In this paper, we present a novel technique, called spatially-varying-uniform scaling functions, to warp the images to equalize the head sizes of the meeting participants without causing undue distortion. In addition, we show a specially designed five-lens camera to capture, stitch, and warp images in real time without sacrificing resolution. Finally, we show that the SVU scaling functions can also be applied to 360 degree images to improve video conferencing experience when an omni-directional camera is used. Zicheng Liu 0001, Michael F. Cohen |
ICME | 1 |
| 2005 | Multi-sensory speech processing: incorporating automatically extracted hidden dynamic informationabstractWe describe a novel technique for multi-sensory speech processing for enhancing noisy speech and for improved noise-robust speech recognition. Both air- and bone-conductive microphones are used to capture speech data where the bone sensor contains virtually noise-free hidden dynamic information of clean speech in the form of formant trajectories. The distortion in the bone-sensor signal such as teeth-clacking and noise leakage can be effectively removed by making use of the automatically extracted formant information from the bone-sensor signal. This paper reports an improved technique for synthesizing speech waveforms based on the LPC cepstra computed analytically from the formant trajectories. When this new signal stream is fused with the other available speech data streams, we achieved improved performance for noisy speech recognition. Amarnag Subramanya, Li Deng 0001, Zicheng Liu 0001, Zhengyou Zhang |
ICME | 3 |
| 2005 | A graphical model for multi-sensory speech processing in air-and-bone conductive microphonesabstractIn continuation of our previous work on using an air-and-boneconductive microphone for speech enhancement, in this paper we propose a graphical model based approach to estimating the clean speech signal given the noisy observations in the air sensor. We also show how the same model can be used as a speech/non-speech classifier. With the aid of MOS (mean opinion score) tests we show, that the performance of the proposed model is better in comparison to our previously proposed direct filtering algorithm. Amarnag Subramanya, Zhengyou Zhang, Zicheng Liu 0001, Jasha Droppo, Alex Acero |
INTERSPEECH | 3 |
| 2004 | Multi-sensory microphones for robust speech detection, enhancement and recognitionabstractIn this paper, we present new hardware prototypes that integrate several heterogeneous sensors into a single headset and describe the underlying DSP techniques for robust speech detection, enhancement and recognition in highly non-stationary noisy environments. We also speculate other business uses with this type of device. Zhengyou Zhang, Zicheng Liu 0001, Mike Sinclair, Alex Acero, Li Deng 0001, Jasha Droppo, Xuedong Huang 0001, Yanli Zheng |
ICASSP (3) | 2 |
| 2004 | Low bit-rate video streaming for face-to-face teleconferenceabstractFace-to-face video teleconferencing is very important for real time communication. Current teleconferencing applications use standard video codecs, such as MPEG 1/2/4, for the compression of face video. Either a high bandwidth is required for high quality video transmission, or the transmitted face video is blurred at low bitrates. We present a system for real-time coding of face video at low bit-rate. There are two main contributions. First, we improve the technique of long term memory prediction by selecting frames into the database in an optimal way. A new frame is selected into the database only when it is significantly different from those frames which are already in the database. In this way, the database can cover a wider range of images. Second, we incorporate prior knowledge about faces into the long term memory prediction framework. The prior knowledge includes: (1) facial motions are repetitive such that most of them can be reconstructed from multiple reference frames; (2) different components of the face and the background can tolerate different levels of error because of different perceptual importance. Experiments show that, at similar PSNR, the proposed system works much faster and achieves better visual quality than the standard H.264/JVT codec. Zicheng Liu 0001, Michael F. Cohen, Ke Colin Zheng, Thomas S. Huang |
ICME | 2 |
| 2004 | Nonlinear information fusion in multi-sensor processing - extracting and exploiting hidden dynamics of speech captured by a bone-conductive microphoneabstractOne well-known difficulty in creating an effective human-machine interface via the speech input is the adverse effects of concurrent acoustic noise. To overcome this challenge, we have developed a joint hardware and software solution. A novel bone-conductive microphone is integrated with a regular air-conductive one in a single headset. These two simultaneous sensors capture the distinct signal properties in the speech embedded in acoustic noise. The focus of this paper is the exploration of the type of dynamic properties that are relatively invariant between the bone-conductive sensor's signal and the clean speech signal; the latter would not be available to the recognizer. Our approach is based on a nonlinear processing technique that estimates the unobserved (hidden) vocal tract resonances, as a representation of such invariant hidden dynamics, from the available bone-sensor signal. The information about these dynamic aspects of the clean speech is then fused with the other noisy measurements that aims to improve the recognition system's robustness to acoustic distortion. The fusion technique is based on a combination of three sets of signals including the synthesized speech signal using the vocal tract resonance dynamics extracted nonlinearly from the bone-sensor signal. Li Deng 0001, Zicheng Liu 0001, Zhengyou Zhang, Alex Acero |
MMSP | 2 |
| 2004 | Direct filtering for air- and bone-conductive microphonesabstractAir- and bone-conductive integrated microphones have been introduced by the authors [Y. Zheng, et al., 2003, Z. Zhang et al., 2004] for speech enhancement in noisy environments. In this paper, we present a novel technique, called direct filtering, to combine the two channels from the air- and bone-conductive microphone for speech enhancement. Compared to the previous technique, the advantage of the direct filtering is that it does not require any training, and it is speaker independent. Experiments show that this technique effectively removes noises and significantly improves speech recognition accuracy even in highly non-stationary noisy environments. Zicheng Liu 0001, Zhengyou Zhang, Alex Acero, Jasha Droppo, Xuedong Huang 0001 |
MMSP | 1 |
| 2004 | Robust and Rapid Generation of Animated Faces from Video Images: A Model-Based Modeling Approach
Zhengyou Zhang, Zicheng Liu 0001, Dennis Adler, Michael F. Cohen, Erik Hanson, Ying Shan |
Int. J. Comput. Vis. | 2 |
| 2004 | ARTiFACIAL: Automated Reverse Turing test using FACIAL features
Yong Rui, Zicheng Liu 0001 |
Multim. Syst. | 2 |
| 2003 | Face Relighting with Radiance Environment MapsabstractA radiance environment map pre-integrates a constant surface reflectance with the lighting environment. It has been used to generate photo-realistic rendering at interactive speed. However, one of its limitations is that each radiance environment map can only render the object, which has the same surface reflectance as what it integrates. We present a ratio-image based technique to use a radiance environment map to render diffuse objects with different surface reflectance properties. This method has the advantage that it does not require the separation of illumination from reflectance, and it is simple to implement and runs at interactive speed. In order to use this technique for human face relighting, we have developed a technique that uses spherical harmonics to approximate the radiance environment map for any given image of a face. Thus we are able to relight face images when the lighting environment rotates. Another benefit of the radiance environment map is that we can interactively modify lighting by changing the coefficients of the spherical harmonics basis. Finally we can modify the lighting condition of one person's face so that it matches the new lighting condition of a different person's face image assuming the two faces have similar skin albedos. Zicheng Liu 0001, Thomas S. Huang |
CVPR (2) | 2 |
| 2003 | Why take notes? Use the whiteboard capture systemabstractThe paper describes our recently developed system which captures both whiteboard content and audio signals of a meeting using a digital still camera and a microphone. Our system can be retrofitted to any existing whiteboard. It computes the time stamps of pen strokes on the whiteboard by analyzing the sequence of captured snapshots. It also automatically produces a set of key frames representing all the written content on the whiteboard before each erasure. Therefore, the whiteboard content serves as a visual index to browse the audio meeting efficiently. It is a complete system which not only captures the whiteboard content, but also helps users to view and manage the captured meeting content efficiently and securely. Li-wei He, Zicheng Liu 0001, Zhengyou Zhang |
ICASSP (5) | 2 |
| 2003 | ARTiFACIAL: automated reverse turing test using FACIAL featuresabstractWeb services designed for human users are being abused by computer programs (bots). The bots steal thousands of free email accounts in a minute; participate in online polls to skew results; and irritate people by joining online chat rooms. These real-world issues have recently generated a new research area called Human Interactive Proofs (HIP), whose goal is to defend services from malicious attacks by differentiating bots from human users. In this paper, we propose a new HIP algorithm based on detecting human face and facial features. Human faces are the most familiar object to humans, rendering it possibly the best candidate for HIP. We conducted user studies and showed the ease of use of our system to human users. We designed attacks using the best existing face detectors and demonstrated the difficulty to bots. Yong Rui, Zicheng Liu 0001 |
ACM Multimedia | 2 |
| 2003 | Excuse me, but are you human?abstractWeb services designed for human users are being abused by computer programs (bots). The bots steal thousands of free email accounts in a minute; participate in online polls to skew results; and irritate people by joining online chat rooms. These real-world issues have recently generated a new research area called Human Interactive Proofs (HIP), whose goal is to defend services from malicious attacks by differentiating bots from human users. We propose a new HIP algorithm based on detecting human face and facial features. Human faces are the most familiar object to humans, rendering it possibly the best candidate for HIP. We conducted user studies and showed the ease of use of our system to human users. We designed attacks using the best existing face detectors and demonstrated the difficulty to bots. Yong Rui, Zicheng Liu 0001 |
ACM Multimedia | 2 |
| 2002 | Model-based face image coding using spherical harmonicsabstractA model-based image coding technique combines image processing, computer vision and computer graphics techniques to achieve higher coding efficiency. Traditional model-based image coding techniques encode the motion in an image sequence as changes of model parameters, thus very low bit rates can be achieved for video coding. We present a novel method to improve coding efficiency of face images. The key idea is to use a spherical harmonic representation of illumination and a generic face geometric model to encode the variation of appearance produced by diffuse reflection. We show that this method improves coding efficiency. This method is complementary to other model-based image coding techniques. Zicheng Liu 0001, Thomas S. Huang |
ICIP (1) | 2 |
| 2002 | On recovering detailed face deformation under general lighting using height from shadingabstractFace tracking and animation are important components for many human computer interface applications. Facial motion produces transient features such as wrinkles and shading changes, which are important yet difficult issues for both analysis (tracking) and synthesis (animation). Previous approaches were mostly based on extensive training appearance examples. However it is difficult for collect samples to cover all possible lighting conditions and head poses. In this paper, we attempt to recover detailed facial geometrical changes due to facial motion under general lighting. The recovered geometry can then be used for the analysis and synthesis of the transient features and shading changes under new lighting conditions. The method is based on shading information and thus suitable for facial areas without permanent features. Thomas S. Huang, Zicheng Liu 0001 |
ICME (1) | 3 |
| 2002 | Distributed meetings: a meeting capture and broadcasting systemabstractThe common meeting is an integral part of everyday life for most workgroups. However, due to travel, time, or other constraints, people are often not able to attend all the meetings they need to. Teleconferencing and recording of meetings can address this problem. In this paper we describe a system that provides these features, as well as a user study evaluation of the system. The system uses a variety of capture devices (a novel 360° camera, a whiteboard camera, an overview camera, and a microphone array) to provide a rich experience for people who want to participate in a meeting from a distance. The system is also combined with speaker clustering, spatial indexing, and time compression to provide a rich experience for people who miss a meeting and want to watch it afterward. Ross Cutler, Yong Rui, Anoop Gupta, Jonathan J. Cadiz, Ivan Tashev, Li-wei He, Alex Colburn, Zhengyou Zhang, Zicheng Liu 0001, Steve Silverberg |
ACM Multimedia | 9 |
| 2001 | Image-Based Surface Detail TransferabstractWe present a novel technique, called Image-Based Surface Detail Transfer, to transfer geometric details from one surface to another with simple 2D image operations. The basic observation is that, without knowing its 3D geometry, geometric details (local deformations) can be extracted from a single image of an object in a way independent of its surface reflectance, and furthermore, these geometric details can be transferred to modify the appearance of other objects directly in images. We show examples including surface detail transfer between real objects, as well as between real and synthesized objects. Ying Shan, Zicheng Liu 0001, Zhengyou Zhang |
CVPR (2) | 2 |
| 2001 | Model-Based Bundle Adjustment with Application to Face ModelingabstractWe present a new model-based bundle adjustment algorithm to recover the 3D model of a scene/object from a sequence of images with unknown motions. Instead of representing scene/object by a collection of isolated 3D features (usually points), our algorithm uses a surface controlled by a small set of parameters. Compared with previous model-based approaches, our approach has the following advantages. First instead of using the model space as a regularizer we directly use it as our search space, thus resulting in a more elegant formulation with fewer unknowns and fewer equations. Second, our algorithm automatically associates tracked points with their correct locations on the surfaces, thereby eliminating the need for a prior 2D-to-3D association. Third, regarding face modeling, we use a very small set of face metrics (meaningful deformations) to parameterize the face geometry, resulting in a smaller search space and a better posed system. Experiments with both synthetic and real data show that this new algorithm is faster, more accurate and more stable than existing ones. Ying Shan, Zicheng Liu 0001, Zhengyou Zhang |
ICCV | 2 |
| 2001 | Cloning Your Own Face with a Desktop CameraabstractWe have developed an easy and cost-effective system that constructs textured 3D animated face models from videos with minimal user interaction. Our system first takes, with an ordinary video camera, images of a face of a person sitting in front of the camera turning the head from one side to the other. After five manual clicks on two images to tell the system where the eye corners, nose top and mouth corners are, the system automatically generates a realistic looking 3D human head model and the constructed model can be animated immediately (different poses, facial expressions and talking). A user, with a PC and a video camera, can use our system to generate hisher face model in a few minutes. The face model can then be imported in hisher favorite game, and the user sees themselves and their friends take part in the game they are playing. We will demonstrate the system on a laptop computer live at the conference, and participants can try it to model their own faces. Zhengyou Zhang, Zicheng Liu 0001, Dennis Adler, Michael F. Cohen, Erik Hanson, Ying Shan |
ICCV | 2 |
| 2001 | Expressive expression mapping with ratio imagesabstractFacial expressions exhibit not only facial feature motions, but also subtle changes in illumination and appearance (e.g., facial creases and wrinkles). These details are important visual cues, but they are difficult to synthesize. Traditional expression mapping techniques consider feature motions while the details in illumination changes are ignored. In this paper, we present a novel technique for facial expression mapping. We capture the illumination change of one person's expression in what we call an expression ratio image (ERI). Together with geometric warping, we map an ERI to any other person's face image to generate more expressive facial expressions. Zicheng Liu 0001, Ying Shan, Zhengyou Zhang |
SIGGRAPH | 1 |
| 2001 | Rapid modeling of animated faces from videoabstractAbstract Generating realistic 3D human face models and facial animations has been a persistent challenge in computer graphics. We have developed a system that constructs textured 3Dface models from videos with minimal user interaction. Our system takes images andvideo sequences of a face with an ordinary video camera. After five manual clicks ontwo images to tell the system where the eye corners, nose top and mouth corners are, thesystem automatically generates a realistic looking 3D human head model and the constructed model can be animated immediately. A user with a PC and an ordinary camera can use our system to generate his/her face model in a few minutes. Copyright © 2001 John Wiley & Sons, Ltd. Zicheng Liu 0001, Zhengyou Zhang, Chuck Jacobs, Michael F. Cohen |
Comput. Animat. Virtual Worlds | 1 |
| 2000 | Rapid modeling of animated faces from video images
Zicheng Liu 0001, Zhengyou Zhang, Chuck Jacobs, Michael F. Cohen |
ACM Multimedia | 1 |
| 1996 | The Bounded Membership Problem of the Monoid SL_2(N)
Jin-Yi Cai, Zicheng Liu 0001 |
Math. Syst. Theory | 2 |
| 1994 | Efficient Average-Case Algorithms for the Modular GroupabstractThe modular group occupies a central position in many branches of mathematical sciences. In this paper we give average polynomial-time algorithms for the unbounded and bounded membership problems for finitely generated subgroups of the modular group. The latter result affirms a conjecture of Y. Gurevich (1990).> Jin-Yi Cai, Wolfgang H. J. Fuchs, Dexter Kozen, Zicheng Liu 0001 |
FOCS | 4 |
| 1994 | Hierarchical spacetime controlabstractSpecifying the motion of an animated linked figure such that it achieves given tasks (e.g., throwing a ball into a basket) and performs the tasks in a realistic fashion (e.g., gracefully, and following physical laws such as gravity) has been an elusive goal for computer animators. The spacetime constraints paradigm has been shown to be a valuable approach to this problem, but it suffers from computational complexity growth as creatures and tasks approach those one would like to animate. The complexity is shown to be, in part, due to the choice of finite basis with which to represent the trajectories of the generalized degrees of freedom. This paper describes new features to the spacetime constraints paradigm to address this problem. Zicheng Liu 0001, Steven J. Gortler, Michael F. Cohen |
SIGGRAPH | 1 |
| 1993 | Minimum Steiner Trees in Normed Planes
Ding-Zhu Du, Biao Gao, Ronald L. Graham, Zicheng Liu 0001, Peng-Jun Wan |
Discret. Comput. Geom. | 4 |
| 1992 | On Steiner Minimal Trees with L_p Distance
Zicheng Liu 0001, Ding-Zhu Du |
Algorithmica | 1 |