Yizhou Wang 0001

dblp:71/3387-1 · DBLP profile ↗
← Back
185ranked-venue papers
6as first author
71since 2021 · last 2026
0000-0001-9888-6409ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 128 · 6 first-author · 60 since 2021Graphics, computer vision, multimedia, augmented reality and games · 112 · 4 first-author · 33 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 6 since 2021Computer networks · 10 · 1 since 2021Databases, data management, data science and information retrieval · 6Systems, architecture and hardware · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Communication-Efficient Desire Alignment for Proactive Embodied Human-Agent Interaction
abstract
Yuanfei Wang, Xinju Huang, Fangwei Zhong, Yaodong Yang, Yizhou Wang, Yuanpei Chen, Hao Dong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yuanfei Wang, Xinju Huang, Fangwei Zhong, Yaodong Yang 0001, Yizhou Wang 0001, Yuanpei Chen, Hao Dong 0003
ACL (1)5
2026 Efficient Action Counting with Dynamic Queries
Xiaoxuan Ma 0001, Zishi Li, Qiuyan Shang, Wentao Zhu 0004, Hai Ci, Yu Qiao 0001, Yizhou Wang 0001
Int. J. Comput. Vis.7
2026 AlphaChimp: Tracking and Behavior Recognition of Chimpanzees
Xiaoxuan Ma 0001, Yutang Lin 0001, Yuan Xu 0022, Stephan P. Kaufhold, Jack Terwilliger, Andres Meza 0001, Yixin Zhu 0001, Federico Rossano, Yizhou Wang 0001
Int. J. Comput. Vis.9
2025 Autoregressive Sequence Modeling for 3D Medical Image Representation
abstract
Three-dimensional (3D) medical images, such as Computed Tomography (CT) and Magnetic Resonance Imaging (MRI), are essential for clinical applications. However, the need for diverse and comprehensive representations is particularly pronounced when considering the variability across different organs, diagnostic tasks, and imaging modalities. How to effectively interpret the intricate contextual information and extract meaningful insights from these images remains an open challenge to the community. While current self-supervised learning methods have shown potential, they often consider an image as a whole thereby overlooking the extensive, complex relationships among local regions from one or multiple images. In this work, we introduce a pioneering method for learning 3D medical image representations through an autoregressive pre-training framework. Our approach sequences various 3D medical images based on spatial, contrast, and semantic correlations, treating them as interconnected visual tokens within a token sequence. By employing an autoregressive sequence modeling task, we predict the next visual token in the sequence, which allows our model to deeply understand and integrate the contextual information inherent in 3D medical images. Additionally, we implement a random startup strategy to avoid overestimating token relationships and to enhance the robustness of learning. The effectiveness of our approach is demonstrated by the superior performance over others on nine downstream tasks in public datasets.
Chu-ran Wang, Lixian Su, Fandong Zhang, Yizhou Wang 0001, Yizhou Yu
AAAI6
2025 A Differential Inclusion Approach for Learning Heterogeneous Sparsity in Neuroimaging Analysis
abstract
In voxel-based neuroimaging disease prediction, it was recently found that in addition to lesion features, there exists another type of feature called "Procedural Bias", which is introduced during preprocessing and can further improve the prediction power. However, traditional sparse learning methods fail to simultaneously capture both types of features due to their heterogeneity in sparsity types. Specifically, the lesion features are spatially coherent and suffer from volumetric degeneration, while the procedural bias refers to enlarged voxels that are dispersedly distributed. In this paper, we propose a new method based on differential inclusion, which generates a sparse regularized solution path on multiple parameters that are enforced with heterogeneous sparsity to capture lesion features and the procedural bias separately. Specifically, we employ Total Variation with a non-negative constraint for the parameter associated with degenerated and spatially coherent lesions; on the other hand, we impose $\ell_1$ sparsity with a non-positive constraint on the parameter related to enlarged and scatterly distributed procedural bias. We theoretically show that our method enjoys model selection consistency and $\ell_2$ consistency in estimation. The utility of our method is demonstrated by improved prediction power and interpretability in the early prediction of Alzheimer’s Disease.
Wenjing Han, Yueming Wu 0005, Xinwei Sun 0001, Lingjing Hu, Yizhou Wang 0001
AISTATS5
2025 Probing and Inducing Combinational Creativity in Vision-Language Models
Yongqian Peng, Yuxuan Wang 0004, Yizhou Wang 0001, Chi Zhang 0017, Yixin Zhu 0001, Zilong Zheng
CogSci5
2025 FreeCloth: Free-form Generation Enhances Challenging Clothed Human Modeling
abstract
Achieving realistic animated human avatars requires accurate modeling of pose-dependent clothing deformations. Existing learning-based methods heavily rely on the Linear Blend Skinning (LBS) of minimally-clothed human models like SMPL to model deformation. However, they struggle to handle loose clothing, such as long dresses, where the canonicalization process becomes ill-defined when the clothing is far from the body, leading to disjointed and fragmented results. To overcome this limitation, we propose FreeCloth, a novel hybrid framework to model challenging clothed humans. Our core idea is to use dedicated strategies to model different regions, depending on whether they are close to or distant from the body. Specifically, we segment the human body into three categories: unclothed, deformed, and generated. We simply replicate unclothed regions that require no deformation. For deformed regions close to the body, we leverage LBS to handle the deformation. As for the generated regions, which correspond to loose clothing areas, we introduce a novel free-form, part-aware generator to model them, as they are less affected by movements. This free-form generation paradigm brings enhanced flexibility and expressiveness to our hybrid framework, enabling it to capture the intricate geometric details of challenging loose clothing, such as skirts and dresses. Experimental results on the benchmark dataset featuring loose clothing demonstrate that FreeCloth achieves state-of-the-art performance with superior visual fidelity and realism, particularly in the most challenging cases.
Hang Ye 0002, Xiaoxuan Ma 0001, Hai Ci, Wentao Zhu 0004, Yizhou Wang 0001
CVPR5
2025 SAT-HMR: Real-Time Multi-Person 3D Mesh Estimation via Scale-Adaptive Tokens
abstract
We propose SAT-HMR, a one-stage framework for real-time multi-person 3D human mesh estimation from a single RGB image. While current one-stage methods, which follow a DETR-style pipeline, achieve state-of-the-art (SOTA) performance with high-resolution inputs, we observe that this particularly benefits the estimation of individuals in smaller scales of the image (e.g., those of young age or far from the camera), but at the cost of significantly increased computation overhead. To address this, we introduce scale-adaptive tokens that are dynamically adjusted based on the relative scale of each individual in the image within the DETR framework. Specifically, individuals in smaller scales are processed at higher resolutions, larger ones at lower resolutions, and background regions are further distilled. These scale-adaptive tokens more efficiently encode the image features, facilitating subsequent decoding to regress the human mesh, while allowing the model to allocate computational resources more effectively and focus on more challenging cases. Experiments show that our method preserves the accuracy benefits of high-resolution processing while substantially reducing computational cost, achieving real-time inference with performance comparable to SOTA methods.
Chi Su, Xiaoxuan Ma 0001, Jiajun Su, Yizhou Wang 0001
CVPR4
2025 InteractAnything: Zero-shot Human Object Interaction Synthesis via LLM Feedback and Object Affordance Parsing
abstract
Recent advances in 3D human-aware generation have made significant progress. However, existing methods still struggle with generating novel Human Object Interaction (HOI) from text, particularly for open-set objects. We identify three main challenges of this task: precise human-object relation reasoning, affordance parsing for any object, and detailed human interaction pose synthesis aligning description and object geometry. In this work, we propose a novel zero-shot 3D HOI generation framework without training on specific datasets, leveraging the knowledge from large-scale pre-trained models. Specifically, the human-object relations are inferred from large language models (LLMs) to initialize object properties and guide the optimization process. Then we utilize a pre-trained 2D image diffusion model to parse unseen objects and extract contact points, avoiding the limitations imposed by existing 3D asset knowledge. The initial human pose is generated by sampling multiple hypotheses through multi-view SDS based on the input text and object geometry. Finally, we introduce a detailed optimization to generate fine-grained, precise, and natural interaction, enforcing realistic 3D contact between the 3D object and the involved body parts, including hands in grasping. This is achieved by distilling human-level feedback from LLMs to capture detailed human-object relations from the text instruction. Extensive experiments validate the effectiveness of our approach compared to prior works, particularly in terms of the fine-grained nature of interactions and the ability to handle open-set 3D objects. Project page: jinluzhang.site/projects/interactanything.
Jinlu Zhang 0001, Yixin Chen 0003, Yizhou Wang 0001, Siyuan Huang 0001
CVPR5
2025 DualGF: Example-Driven Path Planning via Dual Gradient Fields
Mingdong Wu, Fangwei Zhong, Yulong Xia, Yizhou Wang 0001, Hao Dong 0003
ICANN (4)4
2025 UnrealZoo: Enriching Photo-Realistic Virtual Worlds for Embodied AI
abstract
We introduce UnrealZoo, a collection of over 100 photo-realistic 3D virtual worlds built on Unreal Engine, designed to reflect the complexity and variability of open-world environments. We also provide a rich variety of playable entities, including humans, animals, robots, and vehicles for embodied AI research. We extend UnrealCV with optimized APIs and tools for data collection, environment augmentation, distributed training, and benchmarking. These improvements achieve significant improvements in the efficiency of rendering and communication, enabling advanced applications such as multi-agent interactions. Our experimental evaluation across visual navigation and tracking tasks reveals two key insights: 1) environmental diversity provides substantial benefits for developing generalizable reinforcement learning (RL) agents, and 2) current embodied agents face persistent challenges in open-world scenarios, including navigation in unstructured terrain, adaptation to unseen morphologies, and managing latency in the close-loop control systems for interacting in highly dynamic objects. UnrealZoo thus serves as both a comprehensive testing ground and a pathway toward developing more capable embodied AI systems for real-world deployment.
Fangwei Zhong, Kui Wu 0007, Chu-ran Wang, Hao Chen 0062, Hai Ci, Zhoujun Li 0001, Yizhou Wang 0001
ICCV7
2025 Embodied Representation Alignment with Mirror Neurons
abstract
Mirror neurons are a class of neurons that activate both when an individual observes an action and when they perform the same action. This mechanism reveals a fundamental interplay between action understanding and embodied execution, suggesting that these two abilities are inherently connected. Nonetheless, existing machine learning methods largely overlook this interplay, treating these abilities as separate tasks. In this study, we provide a unified perspective in modeling them through the lens of representation learning. We first observe that their intermediate representations spontaneously align. Inspired by mirror neurons, we further introduce an approach that explicitly aligns the representations of observed and executed actions. Specifically, we employ two linear layers to map the representations to a shared latent space, where contrastive learning enforces the alignment of corresponding representations, effectively maximizing their mutual information. Experiments demonstrate that this simple approach fosters mutual synergy between the two tasks, effectively improving representation quality and generalization.
Wentao Zhu 0004, Zhining Zhang 0001, Yizhou Wang 0001
ICCV6
2025 Learning Causal Alignment for Reliable Disease Diagnosis
abstract
Aligning the decision-making process of machine learning algorithms with that of experienced radiologists is crucial for reliable diagnosis. While existing methods have attempted to align their prediction behaviors to those of radiologists reflected in the training data, this alignment is primarily associational rather than causal, resulting in pseudo-correlations that may not transfer well. In this paper, we propose a causality-based alignment framework towards aligning the model's decision process with that of experts. Specifically, we first employ counterfactual generation to identify the causal chain of model decisions. To align this causal chain with that of experts, we propose a causal alignment loss that enforces the model to focus on causal factors underlying each decision step in the whole causal chain. To optimize this loss that involves the counterfactual generator as an implicit function of the model's parameters, we employ the implicit function theorem equipped with the conjugate gradient method for efficient estimation. We demonstrate the effectiveness of our method on two medical diagnosis applications, showcasing faithful alignment to radiologists.
Mingzhou Liu 0001, Ching-Wen Lee, Xinwei Sun 0001, Xueqing Yu, Yu Qiao 0001, Yizhou Wang 0001
ICLR6
2025 Aligning Human Motion Generation with Human Perceptions
abstract
Human motion generation is a critical task with a wide spectrum of applications. Achieving high realism in generated motions requires naturalness, smoothness, and plausibility. However, current evaluation metrics often rely on simple heuristics or distribution distances and do not align well with human perceptions. In this work, we propose a data-driven approach to bridge this gap by introducing a large-scale human perceptual evaluation dataset, MotionPercept, and a human motion critic model, MotionCritic, that capture human perceptual preferences. Our critic model offers a more accurate metric for assessing motion quality and could be readily integrated into the motion generation pipeline to enhance generation quality. Extensive experiments demonstrate the effectiveness of our approach in both evaluating and improving the quality of generated human motions by aligning with human perceptions. Code and data are publicly available at https://motioncritic.github.io/.
Haoru Wang, Wentao Zhu 0004, Luyi Miao, Yishu Xu, Feng Gao 0014, Qi Tian 0001, Yizhou Wang 0001
ICLR7
2025 Simulating Human-like Daily Activities with Desire-driven Autonomy
abstract
Desires motivate humans to interact autonomously with the complex world. In contrast, current AI agents require explicit task specifications, such as instructions or reward functions, which constrain their autonomy and behavioral diversity. In this paper, we introduce a Desire-driven Autonomous Agent (D2A) that can enable a large language model (LLM) to autonomously propose and select tasks, motivated by satisfying its multi-dimensional desires. Specifically, the motivational framework of D2A is mainly constructed by a dynamic $Value\ System$, inspired by the Theory of Needs. It incorporates an understanding of human-like desires, such as the need for social interaction, personal fulfillment, and self-care. At each step, the agent evaluates the value of its current state, proposes a set of candidate activities, and selects the one that best aligns with its intrinsic motivations. We conduct experiments on Concordia, a text-based simulator, to demonstrate that our agent generates coherent, contextually relevant daily activities while exhibiting variability and adaptability similar to human behavior. A comparative analysis with other LLM-based agents demonstrates that our approach significantly enhances the rationality of the simulated activities.
Fangwei Zhong, Yizhou Wang 0001
ICLR5
2025 AdaManip: Adaptive Articulated Object Manipulation Environments and Policy Learning
abstract
Articulated object manipulation is a critical capability for robots to perform various tasks in real-world scenarios. Composed of multiple parts connected by joints, articulated objects are endowed with diverse functional mechanisms through complex relative motions. For example, a safe consists of a door, a handle, and a lock, where the door can only be opened when the latch is unlocked. The internal structure, such as the state of a lock or joint angle constraints, cannot be directly observed from visual observation. Consequently, successful manipulation of these objects requires adaptive adjustment based on trial and error rather than a one-time visual inference. However, previous datasets and simulation environments for articulated objects have primarily focused on simple manipulation mechanisms where the complete manipulation process can be inferred from the object's appearance. To enhance the diversity and complexity of adaptive manipulation mechanisms, we build a novel articulated object manipulation environment and equip it with 9 categories of objects. Based on the environment and objects, we further propose an adaptive demonstration collection and 3D visual diffusion-based imitation learning pipeline that learns the adaptive manipulation policy. The effectiveness of our designs and proposed method is validated through both simulation and real-world experiments.
Yuanfei Wang, Ruihai Wu, Yu Li 0022, Yan Shen 0035, Mingdong Wu, Zhaofeng He 0001, Yizhou Wang 0001, Hao Dong 0003
ICLR8
2025 Bayesian Active Learning for Bivariate Causal Discovery
abstract
Determining the direction of relationships between variables is fundamental for understanding complex systems across scientific domains. While observational data can uncover relationships between variables, it cannot distinguish between cause and effect without experimental interventions. To effectively uncover causality, previous works have proposed intervention strategies that sequentially optimize the intervention values. However, most of these approaches primarily maximized information-theoretic gains that may not effectively measure the reliability of direction determination. In this paper, we formulate the causal direction identification as a hypothesis-testing problem, and propose a Bayes factor-based intervention strategy, which can quantify the evidence strength of one hypothesis (e.g., causal) over the other (e.g., non-causal). To balance the immediate and future gains of testing strength, we propose a sequential intervention objective over intervention values in multiple steps. By analyzing the objective function, we develop a dynamic programming algorithm that reduces the complexity from non-polynomial to polynomial. Experimental results on bivariate systems, tree-structured graphs, and an embodied AI environment demonstrate the effectiveness of our framework in direction determination and its extensibility to both multivariate settings and real-world applications.
Yuxuan Wang 0005, Mingzhou Liu 0001, Xinwei Sun 0001, Wei Wang 0115, Yizhou Wang 0001
ICML5
2025 Behavior-agnostic Task Inference for Robust Offline In-context Reinforcement Learning
abstract
The ability to adapt to new environments with noisy dynamics and unseen objectives is crucial for AI agents. In-context reinforcement learning (ICRL) has emerged as a paradigm to build adaptive policies, employing a context trajectory of the test-time interactions to infer the true task and the corresponding optimal policy efficiently without gradient updates. However, ICRL policies heavily rely on context trajectories, making them vulnerable to distribution shifts from training to testing and degrading performance, particularly in offline settings where the training data is static. In this paper, we highlight that most existing offline ICRL methods are trained for approximate Bayesian inference based on the training distribution, rendering them vulnerable to distribution shifts at test time and resulting in poor generalization. To address this, we introduce Behavior-agnostic Task Inference (BATI) for ICRL, a model-based maximum-likelihood solution to infer the task representation robustly. In contrast to previous methods that rely on a learned encoder as the approximate posterior, BATI focuses purely on dynamics, thus insulating itself against the behavior of the context collection policy. Experiments on MuJoCo environments demonstrate that BATI effectively interprets out-of-distribution contexts and outperforms other methods, even in the presence of significant environmental noise.
Fangwei Zhong, Yizhou Wang 0001
ICML3
2025 VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models
abstract
We introduce a novel self-improving framework that enhances Embodied Visual Tracking (EVT) with Vision-Language Models (VLMs) to address the limitations of current active visual tracking systems in recovering from tracking failure. Our approach combines the off-the-shelf active tracking methods with VLMs’ reasoning capabilities, deploying a fast visual policy for normal tracking and activating VLM reasoning only upon failure detection. The framework features a memory-augmented self-reflection mechanism that enables the VLM to progressively improve by learning from past experiences, effectively addressing VLMs’ limitations in 3D spatial reasoning. Experimental results demonstrate significant performance improvements, with our framework boosting success rates by 72% with state-of-the-art RL-based approaches and 220% with PID-based methods in challenging environments. This work represents the first integration of VLM-based reasoning to assist EVT agents in proactive failure recovery, offering substantial advances for real-world robotic applications that require continuous target monitoring in dynamic, unstructured environments. Project website: https://sites.google.com/view/evt-recovery-assistant.
Kui Wu 0007, Shuhang Xu, Hao Chen 0062, Chu-ran Wang, Zhoujun Li 0001, Yizhou Wang 0001, Fangwei Zhong
IROS6
2025 GeneMAN: Generalizable Single-Image 3D Human Reconstruction from Multi-Source Human Data
abstract
Given a single in-the-wild human photo, it remains a challenging task to reconstruct a high-fidelity 3D human model. Existing methods face difficulties including a) the varying body proportions captured by in-the-wild human images; b) diverse personal belongings within the shot; and c) ambiguities in human postures and inconsistency in human textures. In addition, the scarcity of high-quality human data intensifies the challenge. To address these problems, we propose a Generalizable image-to-3D huMAN reconstruction framework, dubbed GeneMAN, building upon a comprehensive multi-source collection of high-quality human data, including 3D scans, multi-view videos, single photos, and our generated synthetic human data. GeneMAN encompasses three key modules. 1) Without relying on parametric human models (e.g., SMPL), GeneMAN first trains a human-specific text-to-image diffusion model and a view-conditioned diffusion model, serving as GeneMAN 2D human prior and 3D human prior for reconstruction, respectively. 2) With the help of the pretrained human prior models, the Geometry Initialization-&-Sculpting pipeline is leveraged to recover high-quality 3D human geometry given a single image. 3) To achieve high-fidelity 3D human textures, GeneMAN employs the Multi-Space Texture Refinement pipeline, consecutively refining textures in the latent and the pixel spaces. Extensive experimental results demonstrate that GeneMAN could generate high-quality 3D human models from a single image input, outperforming prior state-of-the-art methods. Notably, GeneMAN could reveal much better generalizability in dealing with in-the-wild images, often yielding high-quality 3D human models in natural poses with common items, regardless of the body proportions in the input images.
Wentao Wang 0009, Hang Ye 0002, Fangzhou Hong, Xue Yang 0005, Jianfu Zhang 0003, Yizhou Wang 0001, Ziwei Liu 0002, Liang Pan
NeurIPS6
2025 Shift Equivariant Pose Network
abstract
Human pose estimation has been greatly advanced in recent years. However, even the best-performing models are not shift equivariant. In particular, a small change in input images often results in drastic alterations in output, which are problematic especially in video applications. The prevalence of top-down approaches, which typically rely on a (non-equivariant) object detector in the first stage, exac-erbates this issue. In this paper, we first demonstrate that the biased keypoint representation and the non-equivariant network components are the two main obstacles to shift equivariant pose estimation. To address the limitation, we propose an unbiased decoding method, and redesign the necessary network components (e.g., APS-ResBlock, SSP). Extensive experiments show that our method not only produces much more stable results with shifting input, but also achieves better metrics with the ability of tolerating in-accurate detector output from the first stage. To our knowledge, this is the first work to address the problem of shift equivariance in the field of pose estimation. Our method could be easily applied to existing CNN-based pose estimation networks.
Pengxiao Wang, Tzu-Heng Lin, Chunyu Wang 0001, Yizhou Wang 0001
WACV4
2025 VMarker-Pro: Probabilistic 3D Human Mesh Estimation From Virtual Markers
abstract
Monocular 3D human mesh estimation faces challenges due to depth ambiguity and the complexity of mapping images to complex parameter spaces. Recent methods propose to use 3D poses as a proxy representation, which often lose crucial body shape information, leading to mediocre performance. Conversely, advanced motion capture systems, though accurate, are impractical for markerless wild images. Addressing these limitations, we introduce an innovative intermediate representation as virtual markers, which are learned from large-scale mocap data, mimicking the effects of physical markers. Building upon virtual markers, we propose VMarker, which detects virtual markers from wild images, and the intact mesh with realistic shapes can be obtained by simply interpolation from these markers. To address occlusions that obscure 3D virtual marker estimation, we further enhance our method with VMarker-Pro, a probabilistic framework that models the distribution of 3D virtual marker positions using diffusion models, enabling the generation of multiple plausible meshes aligned with images for robust 3D mesh estimation. Our approaches surpass existing methods on three benchmark datasets, particularly demonstrating significant improvements on the SURREAL dataset, which features diverse body shapes. Additionally, VMarker-Pro excels in accurately modeling data distributions, significantly enhancing performance in occluded scenarios.
Xiaoxuan Ma 0001, Jiajun Su, Yuan Xu 0022, Wentao Zhu 0004, Chunyu Wang 0001, Yizhou Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 ScoreHypo: Probabilistic Human Mesh Estimation with Hypothesis Scoring
abstract
Monocular 3D human mesh estimation is an ill-posed problem, characterized by inherent ambiguity and occlusion. While recent probabilistic methods propose generating multiple solutions, little attention is paid to obtaining high-quality estimates from them. To address this limitation, we introduce ScoreHypo, a versatile framework by first leveraging our novel HypoNet to generate multiple hy-potheses, followed by employing a meticulously designed scorer, ScoreNet, to evaluate and select high-quality esti-mates. ScoreHypo formulates the estimation process as a re-verse denoising process, where HypoNet produces a diverse set of plausible estimates that effectively align with the im-age cues. Subsequently, ScoreNet is employed to rigorously evaluate and rank these estimates based on their quality and finally identify superior ones. Experimental results demon-strate that HypoNet outperforms existing state-of-the-art probabilistic methods as a multi-hypothesis mesh estimator. Moreover, the estimates selected by ScoreNet significantly outperform random generation or simple averaging. Notably, the trained ScoreNet exhibits generalizability, as it can effectively score existing methods and significantly reduce their errors by more than 15%. Code and models are available at ht tps: / /xy02- 05. gi thub. io/ScoreHypo.
Yuan Xu 0022, Xiaoxuan Ma 0001, Jiajun Su, Wentao Zhu 0004, Yu Qiao 0001, Yizhou Wang 0001
CVPR6
2024 Real-Time Holistic Robot Pose Estimation with Unknown States
Shikun Ban, Juling Fan, Xiaoxuan Ma 0001, Wentao Zhu 0004, Yu Qiao 0003, Yizhou Wang 0001
ECCV (49)6
2024 Safe RLHF: Safe Reinforcement Learning from Human Feedback
abstract
With the development of large language models (LLMs), striking a balance between the performance and safety of AI systems has never been more critical. However, the inherent tension between the objectives of helpfulness and harmlessness presents a significant challenge during LLM training. To address this issue, we propose Safe Reinforcement Learning from Human Feedback (Safe RLHF), a novel algorithm for human value alignment. Safe RLHF explicitly decouples human preferences regarding helpfulness and harmlessness, effectively avoiding the crowd workers' confusion about the tension and allowing us to train separate reward and cost models. We formalize the safety concern of LLMs as an optimization task of maximizing the reward function while satisfying specified cost constraints. Leveraging the Lagrangian method to solve this constrained problem, Safe RLHF dynamically adjusts the balance between the two objectives during fine-tuning. Through a three-round fine-tuning using Safe RLHF, we demonstrate a superior ability to mitigate harmful responses while enhancing model performance compared to existing value-aligned algorithms. Experimentally, we fine-tuned the Alpaca-7B using Safe RLHF and aligned it with collected human preferences, significantly improving its helpfulness and harmlessness according to human evaluations. Code is available at https://github.com/PKU-Alignment/safe-rlhf. Warning: This paper contains example data that may be offensive or harmful.
Josef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji, Xinbo Xu, Mickel Liu, Yizhou Wang 0001, Yaodong Yang 0001
ICLR7
2024 Bongard-OpenWorld: Few-Shot Reasoning for Free-form Visual Concepts in the Real World
abstract
We introduce Bongard-OpenWorld, a new benchmark for evaluating real-world few-shot reasoning for machine vision. It originates from the classical Bongard Problems (BPs): Given two sets of images (positive and negative), the model needs to identify the set that query images belong to by inducing the visual concepts, which is exclusively depicted by images from the positive set. Our benchmark inherits the few-shot concept induction of the original BPs while adding the two novel layers of challenge: 1) open-world free-form concepts, as the visual concepts in Bongard-OpenWorld are unique compositions of terms from an open vocabulary, ranging from object categories to abstract visual attributes and commonsense factual knowledge; 2) real-world images, as opposed to the synthetic diagrams used by many counterparts. In our exploration, Bongard-OpenWorld already imposes a significant challenge to current few-shot reasoning algorithms. We further investigate to which extent the recently introduced Large Language Models (LLMs) and Vision-Language Models (VLMs) can solve our task, by directly probing VLMs, and combining VLMs and LLMs in an interactive reasoning scheme. We even conceived a neuro-symbolic reasoning approach that reconciles LLMs & VLMs with logical reasoning to emulate the human problem-solving process for Bongard Problems. However, none of these approaches manage to close the human-machine gap, as the best learner achieves 64% accuracy while human participants easily reach 91%. We hope Bongard-OpenWorld can help us better understand the limitations of current visual intelligence and facilitate future research on visual agents with stronger few-shot visual reasoning capabilities.
Rujie Wu, Xiaojian Ma 0001, Zhenliang Zhang 0002, Wei Wang 0115, Qing Li 0003, Song-Chun Zhu, Yizhou Wang 0001
ICLR7
2024 Causal Discovery via Conditional Independence Testing with Proxy Variables
abstract
Distinguishing causal connections from correlations is important in many scenarios. However, the presence of unobserved variables, such as the latent confounder, can introduce bias in conditional independence testing commonly employed in constraint-based causal discovery for identifying causal relations. To address this issue, existing methods introduced proxy variables to adjust for the bias caused by unobserveness. However, these methods were either limited to categorical variables or relied on strong parametric assumptions for identification. In this paper, we propose a novel hypothesis-testing procedure that can effectively examine the existence of the causal relationship over continuous variables, without any parametric constraint. Our procedure is based on discretization, which under completeness conditions, is able to asymptotically establish a linear equation whose coefficient vector is identifiable under the causal null hypothesis. Based on this, we introduce our test statistic and demonstrate its asymptotic level and power. We validate the effectiveness of our procedure using both synthetic and real-world data.
Mingzhou Liu 0001, Xinwei Sun 0001, Yu Qiao 0001, Yizhou Wang 0001
ICML4
2024 Fast Peer Adaptation with Context-aware Exploration
abstract
Fast adapting to unknown peers (partners or opponents) with different strategies is a key challenge in multi-agent games. To do so, it is crucial for the agent to probe and identify the peer’s strategy efficiently, as this is the prerequisite for carrying out the best response in adaptation. However, exploring the strategies of unknown peers is difficult, especially when the games are partially observable and have a long horizon. In this paper, we propose a peer identification reward, which rewards the learning agent based on how well it can identify the behavior pattern of the peer over the historical context, such as the observation over multiple episodes. This reward motivates the agent to learn a context-aware policy for effective exploration and fast adaptation, i.e., to actively seek and collect informative feedback from peers when uncertain about their policies and to exploit the context to perform the best response when confident. We evaluate our method on diverse testbeds that involve competitive (Kuhn Poker), cooperative (PO-Overcooked), or mixed (Predator-Prey-W) games with peer agents. We demonstrate that our method induces more active exploration behavior, achieving faster adaptation and better outcomes than existing methods.
Yuanfei Wang, Fangwei Zhong, Song-Chun Zhu, Yizhou Wang 0001
ICML5
2024 Language Models Represent Beliefs of Self and Others
abstract
Understanding and attributing mental states, known as Theory of Mind (ToM), emerges as a fundamental capability for human social reasoning. While Large Language Models (LLMs) appear to possess certain ToM abilities, the mechanisms underlying these capabilities remain elusive. In this study, we discover that it is possible to linearly decode the belief status from the perspectives of various agents through neural activations of language models, indicating the existence of internal representations of self and others’ beliefs. By manipulating these representations, we observe dramatic changes in the models’ ToM performance, underscoring their pivotal role in the social reasoning process. Additionally, our findings extend to diverse social reasoning tasks that involve different causal inference patterns, suggesting the potential generalizability of these representations.
Wentao Zhu 0004, Zhining Zhang 0001, Yizhou Wang 0001
ICML3
2024 Cross-dimensional Medical Self-supervised Representation Learning Based on a Pseudo-3D Transformation
Fandong Zhang, Yizhou Wang 0001, Chu-ran Wang, Yizhou Yu
MICCAI (11)5
2024 Richelieu: Self-Evolving LLM-Based Agents for AI Diplomacy
abstract
Diplomacy is one of the most sophisticated activities in human society, involving complex interactions among multiple parties that require skills in social reasoning, negotiation, and long-term strategic planning. Previous AI agents have demonstrated their ability to handle multi-step games and large action spaces in multi-agent tasks. However, diplomacy involves a staggering magnitude of decision spaces, especially considering the negotiation stage required. While recent agents based on large language models (LLMs) have shown potential in various applications, they still struggle with extended planning periods in complex multi-agent settings. Leveraging recent technologies for LLM-based agents, we aim to explore AI's potential to create a human-like agent capable of executing comprehensive multi-agent missions by integrating three fundamental capabilities: 1) strategic planning with memory and reflection; 2) goal-oriented negotiation with social reasoning; and 3) augmenting memory through self-play games for self-evolution without human in the loop. Project page: https://sites.google.com/view/richelieu-diplomacy.
Fangwei Zhong, Yizhou Wang 0001
NeurIPS4
2024 Human Motion Generation: A Survey
abstract
Human motion generation aims to generate natural human pose sequences and shows immense potential for real-world applications. Substantial progress has been made recently in motion data collection technologies and generation methods, laying the foundation for increasing interest in human motion generation. Most research within this field focuses on generating human motions based on conditional signals, such as text, audio, and scene contexts. While significant advancements have been made in recent years, the task continues to pose challenges due to the intricate nature of human motion and its implicit relationship with conditional signals. In this survey, we present a comprehensive literature review of human motion generation, which, to the best of our knowledge, is the first of its kind in this field. We begin by introducing the background of human motion and generative models, followed by an examination of representative methods for three mainstream sub-tasks: text-conditioned, audio-conditioned, and scene-conditioned human motion generation. Additionally, we provide an overview of common datasets and evaluation metrics. Lastly, we discuss open problems and outline potential future research directions. We hope that this survey could provide the community with a comprehensive glimpse of this rapidly evolving field and inspire novel ideas that address the outstanding challenges.
Wentao Zhu 0004, Xiaoxuan Ma 0001, Dongwoo Ro, Hai Ci, Jinlu Zhang 0001, Jiaxin Shi, Feng Gao 0014, Qi Tian 0001, Yizhou Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.9
2023 RSPT: Reconstruct Surroundings and Predict Trajectory for Generalizable Active Object Tracking
abstract
Active Object Tracking (AOT) aims to maintain a specific relation between the tracker and object(s) by autonomously controlling the motion system of a tracker given observations. It is widely used in various applications such as mobile robots and autonomous driving. However, Building a generalizable active tracker that works robustly across various scenarios remains a challenge, particularly in unstructured environments with cluttered obstacles and diverse layouts. To realize this, we argue that the key is to construct a state representation that can model the geometry structure of the surroundings and the dynamics of the target. To this end, we propose a framework called RSPT to form a structure-aware motion representation by Reconstructing Surroundings and Predicting the target Trajectory. Moreover, we further enhance the generalization of the policy network by training in the asymmetric dueling mechanism. Empirical results show that RSPT outperforms existing methods in unseen environments, especially those with cluttered obstacles and diverse layouts. We also demonstrate good sim-to-real transfer when deploying RSPT in real-world scenarios.
Fangwei Zhong, Xiao Bi, Yudi Zhang 0006, Yizhou Wang 0001
AAAI5
2023 GFPose: Learning 3D Human Pose Prior with Gradient Fields
abstract
Learning 3D human pose prior is essential to human-centered AI. Here, we present GFPose, a versatile framework to model plausible 3D human poses for various applications. At the core of GFPose is a time-dependent score network, which estimates the gradient on each body joint and progressively denoises the perturbed 3D human pose to match a given task specification. During the denoising process, GFPose implicitly incorporates pose priors in gradients and unifies various discriminative and generative tasks in an elegant framework. Despite the simplicity, GFPose demonstrates great potential in several downstream tasks. Our experiments empirically show that 1) as a multi-hypothesis pose estimator, GFPose outperforms existing SOTAs by 20% on Human3.6M dataset. 2) as a single-hypothesis pose estimator, GFPose achieves comparable results to deterministic SOTAs, even with a vanilla backbone. 3) GFPose is able to produce diverse and realistic samples in pose denoising, completion and generation tasks.11Project page https://sites.google.com/view/gfpose/
Hai Ci, Mingdong Wu, Wentao Zhu 0004, Xiaoxuan Ma 0001, Hao Dong 0003, Fangwei Zhong, Yizhou Wang 0001
CVPR7
2023 3D Human Mesh Estimation from Virtual Markers
abstract
Inspired by the success of volumetric 3D pose estimation, some recent human mesh estimators propose to estimate 3D skeletons as intermediate representations, from which, the dense 3D meshes are regressed by exploiting the mesh topology. However, body shape information is lost in extracting skeletons, leading to mediocre performance. The advanced motion capture systems solve the problem by placing dense physical markers on the body surface, which allows to extract realistic meshes from their non-rigid motions. However, they cannot be applied to wild images without markers. In this work, we present an intermediate representation, named virtual markers, which learns 64 landmark keypoints on the body surface based on the large-scale mocap data in a generative style, mimicking the effects of physical markers. The virtual markers can be accurately detected from wild images and can reconstruct the intact meshes with realistic shapes by simple interpolation. Our approach outperforms the state-of-the-art methods on three datasets. In particular, it surpasses the existing methods by a notable margin on the SURREAL dataset, which has diverse body shapes. Code is available at https://github.com/ShirleyMaxx/VirtualMarker
Xiaoxuan Ma 0001, Jiajun Su, Chunyu Wang 0001, Wentao Zhu 0004, Yizhou Wang 0001
CVPR5
2023 MotionBERT: A Unified Perspective on Learning Human Motion Representations
abstract
We present a unified perspective on tackling various human-centric video tasks by learning human motion representations from large-scale and heterogeneous data resources. Specifically, we propose a pretraining stage in which a motion encoder is trained to recover the underlying 3D motion from noisy partial 2D observations. The motion representations acquired in this way incorporate geometric, kinematic, and physical knowledge about human motion, which can be easily transferred to multiple downstream tasks. We implement the motion encoder with a Dual-stream Spatio-temporal Transformer (DSTformer) neural network. It could capture long-range spatio-temporal relationships among the skeletal joints comprehensively and adaptively, exemplified by the lowest 3D pose estimation error so far when trained from scratch. Furthermore, our proposed framework achieves state-of-the-art performance on all three downstream tasks by simply finetuning the pretrained motion encoder with a simple regression head (1-2 layers), which demonstrates the versatility of the learned motion representations. Code and models are available at https://motionbert.github.io/
Wentao Zhu 0004, Xiaoxuan Ma 0001, Zhaoyang Liu 0001, Libin Liu 0002, Wayne Wu, Yizhou Wang 0001
ICCV6
2023 Proactive Multi-Camera Collaboration for 3D Human Pose Estimation
Hai Ci, Mickel Liu, Xuehai Pan, Fangwei Zhong, Yizhou Wang 0001
ICLR5
2023 Learning Domain-Agnostic Representation for Disease Diagnosis
Chu-ran Wang, Jing Li 0091, Xinwei Sun 0001, Fandong Zhang, Yizhou Yu, Yizhou Wang 0001
ICLR6
2023 Which Invariance Should We Transfer? A Causal Minimax Learning Approach
abstract
A major barrier to deploying current machine learning models lies in their non-reliability to dataset shifts. To resolve this problem, most existing studies attempted to transfer stable information to unseen environments. Particularly, independent causal mechanisms-based methods proposed to remove mutable causal mechanisms via the do-operator. Compared to previous methods, the obtained stable predictors are more effective in identifying stable information. However, a key question remains: which subset of this whole stable information should the model transfer, in order to achieve optimal generalization ability? To answer this question, we present a comprehensive minimax analysis from a causal perspective. Specifically, we first provide a graphical condition for the whole stable set to be optimal. When this condition fails, we surprisingly find with an example that this whole stable set, although can fully exploit stable information, is not the optimal one to transfer. To identify the optimal subset under this case, we propose to estimate the worst-case risk with a novel optimization scheme over the intervention functions on mutable causal mechanisms. We then propose an efficient algorithm to search for the subset with minimal worst-case risk, based on a newly defined equivalence relation between stable subsets. Compared to the exponential cost of exhaustively searching over all subsets, our searching strategy enjoys a polynomial complexity. The effectiveness and efficiency of our methods are demonstrated on synthetic data and the diagnosis of Alzheimer's disease.
Mingzhou Liu 0001, Xinwei Sun 0001, Fang Fang 0003, Yizhou Wang 0001
ICML5
2023 BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset
abstract
In this paper, we introduce the BeaverTails dataset, aimed at fostering research on safety alignment in large language models (LLMs). This dataset uniquely separates annotations of helpfulness and harmlessness for question-answering pairs, thus offering distinct perspectives on these crucial attributes. In total, we have gathered safety meta-labels for 333,963 question-answer (QA) pairs and 361,903 pairs of expert comparison data for both the helpfulness and harmlessness metrics. We further showcase applications of BeaverTails in content moderation and reinforcement learning with human feedback (RLHF), emphasizing its potential for practical safety measures in LLMs. We believe this dataset provides vital resources for the community, contributing towards the safe development and deployment of LLMs. Our project page is available at the following URL: https://sites.google.com/view/pku-beavertails.
Jiaming Ji, Mickel Liu, Josef Dai, Xuehai Pan, Chi Zhang 0017, Ce Bian, Boyuan Chen 0008, Ruiyang Sun, Yizhou Wang 0001, Yaodong Yang 0001
NeurIPS9
2023 Causal Discovery from Subsampled Time Series with Proxy Variables
abstract
Inferring causal structures from time series data is the central interest of many scientific inquiries. A major barrier to such inference is the problem of subsampling, *i.e.*, the frequency of measurement is much lower than that of causal influence. To overcome this problem, numerous methods have been proposed, yet either was limited to the linear case or failed to achieve identifiability. In this paper, we propose a constraint-based algorithm that can identify the entire causal structure from subsampled time series, without any parametric constraint. Our observation is that the challenge of subsampling arises mainly from hidden variables at the unobserved time steps. Meanwhile, every hidden variable has an observed proxy, which is essentially itself at some observable time in the future, benefiting from the temporal structure. Based on these, we can leverage the proxies to remove the bias induced by the hidden variables and hence achieve identifiability. Following this intuition, we propose a proxy-based causal discovery algorithm. Our algorithm is nonparametric and can achieve full causal identification. Theoretical advantages are reflected in synthetic and real-world experiments.
Mingzhou Liu 0001, Xinwei Sun 0001, Lingjing Hu, Yizhou Wang 0001
NeurIPS4
2023 ChimpACT: A Longitudinal Dataset for Understanding Chimpanzee Behaviors
abstract
Understanding the behavior of non-human primates is crucial for improving animal welfare, modeling social behavior, and gaining insights into distinctively human and phylogenetically shared behaviors. However, the lack of datasets on non-human primate behavior hinders in-depth exploration of primate social interactions, posing challenges to research on our closest living relatives. To address these limitations, we present ChimpACT, a comprehensive dataset for quantifying the longitudinal behavior and social relations of chimpanzees within a social group. Spanning from 2015 to 2018, ChimpACT features videos of a group of over 20 chimpanzees residing at the Leipzig Zoo, Germany, with a particular focus on documenting the developmental trajectory of one young male, Azibo. ChimpACT is both comprehensive and challenging, consisting of 163 videos with a cumulative 160,500 frames, each richly annotated with detection, identification, pose estimation, and fine-grained spatiotemporal behavior labels. We benchmark representative methods of three tracks on ChimpACT: (i) tracking and identification, (ii) pose estimation, and (iii) spatiotemporal action detection of the chimpanzees. Our experiments reveal that ChimpACT offers ample opportunities for both devising new methods and adapting existing ones to solve fundamental computer vision tasks applied to chimpanzee groups, such as detection, pose estimation, and behavior analysis, ultimately deepening our comprehension of communication and sociality in non-human primates.
Xiaoxuan Ma 0001, Stephan P. Kaufhold, Jiajun Su, Wentao Zhu 0004, Jack Terwilliger, Andres Meza 0001, Yixin Zhu 0001, Federico Rossano, Yizhou Wang 0001
NeurIPS9
2023 Social Motion Prediction with Cognitive Hierarchies
abstract
Humans exhibit a remarkable capacity for anticipating the actions of others and planning their own actions accordingly. In this study, we strive to replicate this ability by addressing the social motion prediction problem. We introduce a new benchmark, a novel formulation, and a cognition-inspired framework. We present Wusi, a 3D multi-person motion dataset under the context of team sports, which features intense and strategic human interactions and diverse pose distributions. By reformulating the problem from a multi-agent reinforcement learning perspective, we incorporate behavioral cloning and generative adversarial imitation learning to boost learning efficiency and generalization. Furthermore, we take into account the cognitive aspects of the human social action planning process and develop a cognitive hierarchy framework to predict strategic human social interactions. We conduct comprehensive experiments to validate the effectiveness of our proposed dataset and approach.
Wentao Zhu 0004, Jason Qin, Yuke Lou, Hang Ye 0002, Xiaoxuan Ma 0001, Hai Ci, Yizhou Wang 0001
NeurIPS7
2023 Efficient Human Motion Reconstruction from Monocular Videos with Physical Consistency Loss
abstract
Vision-only motion reconstruction from monocular videos often produces artifacts such as foot sliding and jittering. Existing physics-based methods typically either simplify the problem to focus solely on foot-ground contacts, or they reconstruct full-body contacts within a physics simulator, necessitating the solution of a time-consuming bilevel optimization problem. To overcome these limitations, we present an efficient gradient-based method for reconstructing complex human motions (including highly dynamic and acrobatic movements) with physical constraints. Our approach reformulates human motion dynamics through a differentiable physical consistency loss within an augmented search space that accounts both for contacts and camera alignment. This enables us to transform the motion reconstruction task into a single-level trajectory optimization problem. Experimental results demonstrate that our method can reconstruct complex human motions from real-world videos in minutes, which is substantially faster than previous approaches. Additionally, the reconstructed results show enhanced physical realism compared to existing methods.
Philipp Ruppel, Yizhou Wang 0001, Norman Hendrich, Jianwei Zhang 0001
SIGGRAPH Asia3
2023 A Residual Learning Approach to Deblur and Generate High Frame Rate Video With an Event Camera
abstract
Event cameras are bio-inspired cameras that can measure the intensity change asynchronously with high temporal resolution. One of the advantages of event cameras is that they suffer less from motion blur than traditional frame cameras when recording daily scenes with fast-moving objects. In this paper, we formulate the deblurring task on traditional cameras directed by events to be a residual learning one, and propose corresponding network architectures for effective learning of deblurring and high frame rate video generation tasks. We first train a modified U-Net network to restore a sharp image from a blurry image using the corresponding events. Then we train another similar network by replacing the downsampling blocks with blocks of the convolutional long short-term memory (Conv-LSTM) to recurrently generate high frame rate video using the restored sharp image and part of the events. Benefitting from the blur-free events and the proposed learning strategy, the experimental results show that the proposed method outperforms state-of-the-art methods for generating sharp images and high frame rate videos.
Minggui Teng, Boxin Shi, Yizhou Wang 0001, Tiejun Huang 0001
IEEE Trans. Multim.4
2022 Causal Intervention for Subject-Deconfounded Facial Action Unit Recognition
abstract
Subject-invariant facial action unit (AU) recognition remains challenging for the reason that the data distribution varies among subjects. In this paper, we propose a causal inference framework for subject-invariant facial action unit recognition. To illustrate the causal effect existing in AU recognition task, we formulate the causalities among facial images, subjects, latent AU semantic relations, and estimated AU occurrence probabilities via a structural causal model. By constructing such a causal diagram, we clarify the causal-effect among variables and propose a plug-in causal intervention module, CIS, to deconfound the confounder Subject in the causal diagram. Extensive experiments conducted on two commonly used AU benchmark datasets, BP4D and DISFA, show the effectiveness of our CIS, and the model with CIS inserted, CISNet, has achieved state-of-the-art performance.
Yingjie Chen 0002, Diqi Chen, Tao Wang 0004, Yizhou Wang 0001, Yun Liang 0001
AAAI4
2022 MoCaNet: Motion Retargeting In-the-Wild via Canonicalization Networks
abstract
We present a novel framework that brings the 3D motion retargeting task from controlled environments to in-the-wild scenarios. In particular, our method is capable of retargeting body motion from a character in a 2D monocular video to a 3D character without using any motion capture system or 3D reconstruction procedure. It is designed to leverage massive online videos for unsupervised training, needless of 3D annotations or motion-body pairing information. The proposed method is built upon two novel canonicalization operations, structure canonicalization and view canonicalization. Trained with the canonicalization operations and the derived regularizations, our method learns to factorize a skeleton sequence into three independent semantic subspaces, i.e., motion, structure, and view angle. The disentangled representation enables motion retargeting from 2D to 3D with high precision. Our method achieves superior performance on motion transfer benchmarks with large body variations and challenging actions. Notably, the canonicalized skeleton sequence could serve as a disentangled and interpretable representation of human motion that benefits action analysis and motion retrieval.
Wentao Zhu 0004, Zhuoqian Yang, Ziang Di, Wayne Wu, Yizhou Wang 0001, Chen Change Loy
AAAI5
2022 VirtualPose: Learning Generalizable 3D Human Pose Models from Virtual Data
Jiajun Su, Chunyu Wang 0001, Xiaoxuan Ma 0001, Wenjun Zeng 0001, Yizhou Wang 0001
ECCV (6)5
2022 Faster VoxelPose: Real-time 3D Human Pose Estimation by Orthographic Projection
Hang Ye 0002, Wentao Zhu 0004, Chunyu Wang 0001, Rujie Wu, Yizhou Wang 0001
ECCV (6)5
2022 One-Shot Medical Landmark Localization by Edge-Guided Transform and Noisy Landmark Refinement
Ping Gong 0002, Chunyu Wang 0001, Yizhou Yu, Yizhou Wang 0001
ECCV (21)5
2022 ToM2C: Target-oriented Multi-agent Communication and Cooperation with Theory of Mind
Yuanfei Wang, Fangwei Zhong, Yizhou Wang 0001
ICLR4
2022 Disentangling Disease-related Representation from Obscure for Disease Prediction
abstract
Disease-related representations play a crucial role in image-based disease prediction such as cancer diagnosis, due to its considerable generalization capacity. However, it is still a challenge to identify lesion characteristics in obscured images, as many lesions are obscured by other tissues. In this paper, to learn the representations for identifying obscured lesions, we propose a disentanglement learning strategy under the guidance of alpha blending generation in an encoder-decoder framework (DAB-Net). Specifically, we take mammogram mass benign/malignant classification as an example. In our framework, composite obscured mass images are generated by alpha blending and then explicitly disentangled into disease-related mass features and interference glands features. To achieve disentanglement learning, features of these two parts are decoded to reconstruct the mass and the glands with corresponding reconstruction losses, and only disease-related mass features are fed into the classifier for disease prediction. Experimental results on one public dataset DDSM and three in-house datasets demonstrate that the proposed strategy can achieve state-of-the-art performance. DAB-Net achieves substantial improvements of 3.9%~4.4% AUC in obscured cases. Besides, the visualization analysis shows the model can better disentangle the mass and glands in the obscured image, suggesting the effectiveness of our solution in exploring the hidden characteristics in this challenging problem.
Chu-ran Wang, Fandong Zhang, Fangwei Zhong, Yizhou Yu, Yizhou Wang 0001
ICML6
2022 Intrinsic Image Decomposition by Pursuing Reflectance Image
abstract
Intrinsic image decomposition is a fundamental problem for many computer vision applications. While recent deep learning based methods have achieved very promising results on the synthetic densely labeled datasets, the results on the real-world dataset are still far from human level performance. This is mostly because collecting dense supervision on a real-world dataset is impossible. Only a sparse set of pairwise judgement from human is often used. It's very difficult for models to learn in such settings. In this paper, we investigate the possibilities of only using reflectance images for supervision during training. In this way, the demand for labeled data is greatly reduced. In order to achieve this goal, we take a deep investigation into the reflectance images. We find that reflectance images are actually comprised of two components: the flat surfaces with low frequency information, and the boundaries with high frequency details. Then, we propose to disentangle the learning process of the two components of the reflectance images. We argue that through this procedure, the reflectance images can be better modeled, and in the meantime, the shading images, though not supervised, can also achieve decent result. Extensive experiments show that our proposed network outperforms current state-of-the-art results by a large margin on the most challenging real-world IIW dataset. We also surprisingly find that on the densely labeled datasets (MIT and MPI-Sintel), our network can also achieve state-of-the-art results on both reflectance and shading images, when we only apply supervision on the reflectance images during training.
Tzu-Heng Lin, Pengxiao Wang, Yizhou Wang 0001
IJCAI3
2022 MATE: Benchmarking Multi-Agent Reinforcement Learning in Distributed Target Coverage Control
abstract
We introduce the Multi-Agent Tracking Environment (MATE), a novel multi-agent environment simulates the target coverage control problems in the real world. MATE hosts an asymmetric cooperative-competitive game consisting of two groups of learning agents--"cameras" and "targets"--with opposing interests. Specifically, "cameras", a group of directional sensors, are mandated to actively control the directional perception area to maximize the coverage rate of targets. On the other side, "targets" are mobile agents that aim to transport cargo between multiple randomly assigned warehouses while minimizing the exposure to the camera sensor networks. To showcase the practicality of MATE, we benchmark the multi-agent reinforcement learning (MARL) algorithms from different aspects, including cooperation, communication, scalability, robustness, and asymmetric self-play. We start by reporting results for cooperative tasks using MARL algorithms (MAPPO, IPPO, QMIX, MADDPG) and the results after augmenting with multi-agent communication protocols (TarMAC, I2C). We then evaluate the effectiveness of the popular self-play techniques (PSRO, fictitious self-play) in an asymmetric zero-sum competitive game. This process of co-evolution between cameras and targets helps to realize a less exploitable camera network. We also observe the emergence of different roles of the target agents while incorporating I2C into target-target communication. MATE is written purely in Python and integrated with OpenAI Gym API to enhance user-friendliness. Our project is released at https://github.com/UnrealTracking/mate.
Xuehai Pan, Mickel Liu, Fangwei Zhong, Yaodong Yang 0001, Song-Chun Zhu, Yizhou Wang 0001
NeurIPS6
2022 Locally Connected Network for Monocular 3D Human Pose Estimation
abstract
We present an approach for 3D human pose estimation from monocular images. The approach consists of two steps: it first estimates a 2D pose from an image and then estimates the corresponding 3D pose. This paper focuses on the second step. Graph convolutional network (GCN) has recently become the de facto standard for human pose related tasks such as action recognition. However, in this work, we show that GCN has critical limitations when it is used for 3D pose estimation due to the inherent weight sharing scheme. The limitations are clearly exposed through a novel reformulation of GCN, in which both GCN and Fully Connected Network (FCN) are its special cases. In addition, on top of the formulation, we present locally connected network (LCN) to overcome the limitations of GCN by allocating dedicated rather than shared filters for different joints. We jointly train the LCN network with a 2D pose estimator such that it can handle inaccurate 2D poses. We evaluate our approach on two benchmark datasets and observe that LCN outperforms GCN, FCN, and the state-of-the-art methods by a large margin. More importantly, it demonstrates strong cross-dataset generalization ability because of sparse connections among body joints.
Hai Ci, Xiaoxuan Ma 0001, Chunyu Wang 0001, Yizhou Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Act Like a Radiologist: Towards Reliable Multi-View Correspondence Reasoning for Mammogram Mass Detection
abstract
Mammogram mass detection is crucial for diagnosing and preventing the breast cancers in clinical practice. The complementary effect of multi-view mammogram images provides valuable information about the breast anatomical prior structure and is of great significance in digital mammography interpretation. However, unlike radiologists who can utilize the natural reasoning ability to identify masses based on multiple mammographic views, how to endow the existing object detection models with the capability of multi-view reasoning is vital for decision-making in clinical diagnosis but remains the boundary to explore. In this paper, we propose an anatomy-aware graph convolutional network (AGN), which is tailored for mammogram mass detection and endows existing detection methods with multi-view reasoning ability. The proposed AGN consists of three steps. First, we introduce a bipartite graph convolutional network (BGN) to model the intrinsic geometric and semantic relations of ipsilateral views. Second, considering that the visual asymmetry of bilateral views is widely adopted in clinical practice to assist the diagnosis of breast lesions, we propose an inception graph convolutional network (IGN) to model the structural similarities of bilateral views. Finally, based on the constructed graphs, the multi-view information is propagated through nodes methodically, which equips the features learned from the examined view with multi-view reasoning ability. Experiments on two standard benchmarks reveal that AGN significantly exceeds the state-of-the-art performance. Visualization results show that AGN provides interpretable visual cues for clinical diagnosis.
Fandong Zhang, Chaoqi Chen, Yizhou Wang 0001, Yizhou Yu
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 A Bayesian Quality-of-Experience Model for Adaptive Streaming Videos
abstract
The fundamental conflict between the enormous space of adaptive streaming videos and the limited capacity for subjective experiment casts significant challenges to objective Quality-of-Experience (QoE) prediction. Existing objective QoE models either employ pre-defined parametrization or exhibit complex functional form, achieving limited generalization capability in diverse streaming environments. In this study, we propose an objective QoE model, namely, the Bayesian streaming quality index (BSQI), to integrate prior knowledge on the human visual system and human annotated data in a principled way. By analyzing the subjective characteristics towards streaming videos from a corpus of subjective studies, we show that a family of QoE functions lies in a convex set. Using a variant of projected gradient descent, we optimize the objective QoE model over a database of training videos. The proposed BSQI demonstrates strong prediction accuracy in a broad range of streaming conditions, evident by state-of-the-art performance on four publicly available benchmark datasets and a novel analysis-by-synthesis visual experiment.
Zhengfang Duanmu, Wentao Liu 0001, Diqi Chen, Zhou Wang 0001, Yizhou Wang 0001, Wen Gao 0001
ACM Trans. Multim. Comput. Commun. Appl.6
2021 Causal Hidden Markov Model for Time Series Disease Forecasting
abstract
We propose a causal hidden Markov model to achieve robust prediction of irreversible disease at an early stage, which is safety-critical and vital for medical treatment in early stages. Specifically, we introduce the hidden variables which propagate to generate medical data at each time step. To avoid learning spurious correlation (e.g., confounding bias), we explicitly separate these hidden variables into three parts: a) the disease (clinical)-related part; b) the disease (non-clinical)-related part; c) others, with only a),b) causally related to the disease however c) may contain spurious correlations (with the disease) inherited from the data provided. With personal attributes and disease label respectively provided as side information and supervision, we prove that these disease-related hidden variables can be disentangled from others, implying the avoidance of spurious correlation for generalization to medical data from other (out-of-) distributions. Guaranteed by this result, we propose a sequential variational auto-encoder with a reformulated objective function. We apply our model to the early prediction of peripapillary atrophy and achieve promising results on out-of-distribution test data. Further, the ablation study empirically shows the effectiveness of each component in our method. And the visualization shows the accurate identification of lesion regions from others.1
Jing Li 0091, Botong Wu, Xinwei Sun 0001, Yizhou Wang 0001
CVPR4
2021 Towards Unified Surgical Skill Assessment
abstract
Surgical skills have a great influence on surgical safety and patients’ well-being. Traditional assessment of surgical skills involves strenuous manual efforts, which lacks efficiency and repeatability. Therefore, we attempt to automatically predict how well the surgery is performed using the surgical video. In this paper, a unified multi-path framework for automatic surgical skill assessment is proposed, which takes care of multiple composing aspects of surgical skills, including surgical tool usage, intraoperative event pattern, and other skill proxies. The dependency relationships among these different aspects are specially modeled by a path dependency module in the framework. We conduct extensive experiments on the JIGSAWS dataset of simulated surgical tasks, and a new clinical dataset of real laparoscopic surgeries. The proposed framework achieves promising results on both datasets, with the state-of-the-art on the simulated dataset advanced from 0.71 Spearman’s correlation to 0.80. It is also shown that combining multiple skill aspects yields better performance than relying on a single aspect.
Daochang Liu, Qiyue Li 0002, Tingting Jiang 0001, Yizhou Wang 0001, Rulin Miao
CVPR4
2021 Context Modeling in 3D Human Pose Estimation: A Unified Perspective
abstract
Estimating 3D human pose from a single image suffers from severe ambiguity since multiple 3D joint configurations may have the same 2D projection. The state-of-the-art methods often rely on context modeling methods such as pictorial structure model (PSM) or graph neural network (GNN) to reduce ambiguity. However, there is no study that rigorously compares them side by side. So we first present a general formula for context modeling in which both PSM and GNN are its special cases. By comparing the two methods, we found that the end-to-end training scheme in GNN and the limb length constraints in PSM are two complementary factors to improve results. To combine their advantages, we propose ContextPose based on attention mechanism that allows enforcing soft limb length constraints in a deep network. The approach effectively reduces the chance of getting absurd 3D pose estimates with incorrect limb lengths and achieves state-of-the-art results on two benchmark datasets. More importantly, the introduction of limb length constraints into deep networks enables the approach to achieve much better generalization performance.
Xiaoxuan Ma 0001, Jiajun Su, Chunyu Wang 0001, Hai Ci, Yizhou Wang 0001
CVPR5
2021 Forecasting Irreversible Disease via Progression Learning
abstract
Forecasting Parapapillary atrophy (PPA), i.e., a symptom related to most irreversible eye diseases, provides an alarm for implementing an intervention to slow down the disease progression at early stage. A key question for this forecast is: how to fully utilize the historical data (e.g., retinal image) up to the current stage for future disease prediction? In this paper, we provide an answer with a novel framework, namely Disease Forecast via Progression Learning (DFPL), which exploits the irreversibility prior (i.e., cannot be reversed once diagnosed). Specifically, based on this prior, we decompose two factors that contribute to the prediction of the future disease: i) the current disease label given the data (retinal image, clinical attributes) at present and ii) the future disease label given the progression of the retinal images that from the current to the future. To model these two factors, we introduce the current and progression predictors in DFPL, respectively. In order to account for the degree of progression of the disease, we propose a temporal generative model to accurately generate the future image and compare it with the current one to get a residual image. The generative model is implemented by a recurrent neural network, in order to exploit the dependency of the historical data. To verify our approach, we apply it to a PPA in-house dataset and it yields a significant improvement (e.g., 4.48% of accuracy; 3.45% of AUC) over others. Besides, our generative model can accurately localize the disease-related regions.
Botong Wu, Sijie Ren, Jing Li 0091, Xinwei Sun 0001, Yizhou Wang 0001
CVPR6
2021 An Empirical Study of the Collapsing Problem in Semi-Supervised 2D Human Pose Estimation
abstract
Most semi-supervised learning models are consistency-based, which leverage unlabeled images by maximizing the similarity between different augmentations of an image. But when we apply them to human pose estimation that has extremely imbalanced class distribution, they often collapse and predict every pixel in unlabeled images as background. We find this is because the decision boundary passes the high-density areas of the minor class so more and more pixels are gradually misclassified as background. In this work, we present a surprisingly simple approach to drive the model to learn in the correct direction. For each image, it composes a pair of easy-hard augmentations and uses the more accurate predictions on the easy image to teach the network to learn pose information of the hard one. The accuracy superiority of teaching signals allows the network to be "monotonically" improved which effectively avoids collapsing. We apply our method to the state-of-the-art pose estimators and it further improves their performance on three public datasets.
Rongchang Xie, Chunyu Wang 0001, Wenjun Zeng 0001, Yizhou Wang 0001
ICCV4
2021 Towards Distraction-Robust Active Visual Tracking
abstract
In active visual tracking, it is notoriously difficult when distracting objects appear, as distractors often mislead the tracker by occluding the target or bringing a confusing appearance. To address this issue, we propose a mixed cooperative-competitive multi-agent game, where a target and multiple distractors form a collaborative team to play against a tracker and make it fail to follow. Through learning in our game, diverse distracting behaviors of the distractors naturally emerge, thereby exposing the tracker’s weakness, which helps enhance the distraction-robustness of the tracker. For effective learning, we then present a bunch of practical methods, including a reward function for distractors, a cross-modal teacher-student learning strategy, and a recurrent attention mechanism for the tracker. The experimental results show that our tracker performs desired distraction-robust active visual tracking and can be well generalized to unseen environments. We also show that the multi-agent game can be used to adversarially test the robustness of trackers.
Fangwei Zhong, Peng Sun 0011, Wenhan Luo, Tingyun Yan, Yizhou Wang 0001
ICML5
2021 AUPro: Multi-label Facial Action Unit Proposal Generation for Sequence-Level Analysis
Yingjie Chen 0002, Jiarui Zhang 0007, Diqi Chen, Tao Wang 0004, Yizhou Wang 0001, Yun Liang 0001
ICONIP (3)5
2021 Symmetry-Enhanced Attention Network for Acute Ischemic Infarct Segmentation with Non-contrast CT Images
Kongming Liang, Kai Han 0010, Xiuli Li, Xiaoqing Cheng, Yizhou Wang 0001, Yizhou Yu
MICCAI (7)6
2021 CA-Net: Leveraging Contextual Features for Lung Cancer Prediction
Mingzhou Liu 0001, Fandong Zhang, Xinwei Sun 0001, Yizhou Yu, Yizhou Wang 0001
MICCAI (5)5
2021 DAE-GCN: Identifying Disease-Related Features for Disease Prediction
Chu-ran Wang, Xinwei Sun 0001, Fandong Zhang, Yizhou Yu, Yizhou Wang 0001
MICCAI (5)5
2021 CaFGraph: Context-aware Facial Multi-graph Representation for Facial Action Unit Recognition
abstract
Facial action unit (AU) recognition has attracted increasing attention due to its indispensable role in affective computing, especially in the field of affective human-computer interaction. Due to the subtle and transient nature of AU, it is challenging to capture the delicate and ambiguous motions in local facial regions among consecutive frames. Considering that context is essential to resolve ambiguity in human visual system, modeling context within or among facial images emerges as a promising approach for AU recognition task. To this end, we propose CaFGraph, a novel context-aware facial multi-graph that can model both morphological & muscular-based region-level local context and region-level temporal context. CaFGraph is the first work to construct a universal facial multi-graph structure that is independent of both task settings and dataset statistics for almost all fine-grained facial behavior analysis tasks, including but not limited to AU recognition. To make full use of the context, we then present CaFNet that learns context-aware facial graph representations via CaFGraph from facial images for multi-label AU recognition. Experiments on two widely used benchmark datasets, BP4D and DISFA, demonstrate the superiority of our CaFNet over the state-of-the-art methods.
Yingjie Chen 0002, Diqi Chen, Yizhou Wang 0001, Tao Wang 0004, Yun Liang 0001
ACM Multimedia3
2021 Compare and contrast: Detecting mammographic soft-tissue lesions with C2-Net
Changsheng Zhou, Fandong Zhang, Qianyi Zhang, Fugeng Sheng, Wanhua Liu, Yizhou Wang 0001, Yizhou Yu, Guangming Lu 0001
Medical Image Anal.10
2021 AD-VAT+: An Asymmetric Dueling Mechanism for Learning and Understanding Visual Active Tracking
abstract
Visual Active Tracking (VAT) aims at following a target object by autonomously controlling the motion system of a tracker given visual observations. To learn a robust tracker for VAT, in this article, we propose a novel adversarial reinforcement learning (RL) method which adopts an Asymmetric Dueling mechanism, referred to as AD-VAT. In the mechanism, the tracker and target, viewed as two learnable agents, are opponents and can mutually enhance each other during the dueling/competition: i.e., the tracker intends to lockup the target, while the target tries to escape from the tracker. The dueling is asymmetric in that the target is additionally fed with the tracker's observation and action, and learns to predict the tracker's reward as an auxiliary task. Such an asymmetric dueling mechanism produces a stronger target, which in turn induces a more robust tracker. To improve the performance of the tracker in the case of challenging scenarios such as obstacles, we employ more advanced environment augmentation technique and two-stage training strategies, termed as AD-VAT+. For a better understanding of the asymmetric dueling mechanism, we also analyze the target's behaviors as the training proceeds and visualize the latent space of the tracker. The experimental results, in both 2D and 3D environments, demonstrate that the proposed method leads to a faster convergence in training and yields more robust tracking behaviors in different testing scenarios. The potential of the active tracker is also shown in real-world videos.
Fangwei Zhong, Peng Sun 0011, Wenhan Luo, Tingyun Yan, Yizhou Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2021 Bilateral Asymmetry Guided Counterfactual Generating Network for Mammogram Classification
abstract
Mammogram benign or malignant classification with only image-level labels is challenging due to the absence of lesion annotations. Motivated by the symmetric prior that the lesions on one side of breasts rarely appear in the corresponding areas on the other side, we explore to answer a counterfactual question to identify the lesion areas. This counterfactual question means: given an image with lesions, how would the features have behaved if there were no lesions in the image? To answer this question, we derive a new theoretical result based on the symmetric prior. Specifically, by building a causal model that entails such a prior for bilateral images, we identify to optimize the distances in distribution between i) the counterfactual features and the target side's features in lesion-free areas; and ii) the counterfactual features and the reference side's features in lesion areas. To realize these optimizations for better benign/malignant classification, we propose a counterfactual generative network, which is mainly composed of Generator Adversarial Network and a prediction feedback mechanism, they are optimized jointly and prompt each other. Specifically, the former can further improve the classi?cation performance by generating counterfactual features to calculate lesion areas. On the other hand, the latter helps counterfactual generation by the supervision of classification loss. The utility of our method and the effectiveness of each module in our model can be verified by state-of-the-art performance on INBreast and an in-house dataset and ablation studies.
Chu-ran Wang, Jing Li 0091, Fandong Zhang, Xinwei Sun 0001, Hao Dong 0003, Yizhou Yu, Yizhou Wang 0001
IEEE Trans. Image Process.7
2020 Pose-Assisted Multi-Camera Collaboration for Active Object Tracking
abstract
Active Object Tracking (AOT) is crucial to many vision-based applications, e.g., mobile robot, intelligent surveillance. However, there are a number of challenges when deploying active tracking in complex scenarios, e.g., target is frequently occluded by obstacles. In this paper, we extend the single-camera AOT to a multi-camera setting, where cameras tracking a target in a collaborative fashion. To achieve effective collaboration among cameras, we propose a novel Pose-Assisted Multi-Camera Collaboration System, which enables a camera to cooperate with the others by sharing camera poses for active object tracking. In the system, each camera is equipped with two controllers and a switcher: The vision-based controller tracks targets based on observed images. The pose-based controller moves the camera in accordance to the poses of the other cameras. At each step, the switcher decides which action to take from the two controllers according to the visibility of the target. The experimental results demonstrate that our system outperforms all the baselines and is capable of generalizing to unseen environments. The code and demo videos are available on our website https://sites.google.com/view/pose-assisted-collaboration.
Jing Li 0091, Fangwei Zhong, Yu Qiao 0001, Yizhou Wang 0001
AAAI6
2020 Cross-View Correspondence Reasoning Based on Bipartite Graph Convolutional Network for Mammogram Mass Detection
abstract
Mammogram mass detection is of great clinical significance due to its high proportion in breast cancers. The information from cross views (i.e., mediolateral oblique and cranio-caudal) is highly related and complementary, and is helpful to make comprehensive decisions. However, unlike radiologists who are able to recognize masses with reasoning ability in cross-view images, most existing methods lack the ability to reason under the guidance of domain knowledge, thus it limits the performance. In this paper, we introduce bipartite graph convolutional network to endow existing methods with cross-view reasoning ability of radiologists in mammogram mass detection. The bipartite node sets are constructed by cross-view images respectively to represent relatively consistent regions in breasts, while the bipartite edge learns to model both inherent cross-view geometric constraints and appearance similarities between correspondences. Based on the bipartite graph, the information propagates methodically through correspondences and enables spatial visual features equipped with customized cross-view reasoning ability. Experimental results on DDSM dataset demonstrate that the proposed algorithm achieves state-of-the-art performance. Besides, visual analysis shows the model has a clear physical meaning, which is helpful for radiologists in clinical interpretation.
Fandong Zhang, Qianyi Zhang, Yizhou Wang 0001, Yizhou Yu
CVPR5
2020 MetaFuse: A Pre-trained Fusion Model for Human Pose Estimation
abstract
Cross view feature fusion is the key to address the occlusion problem in human pose estimation. The current fusion methods need to train a separate model for every pair of cameras making them difficult to scale. In this work, we introduce MetaFuse, a pre-trained fusion model learned from a large number of cameras in the Panoptic dataset. The model can be efficiently adapted or finetuned for a new pair of cameras using a small number of labeled images. The strong adaptation power of MetaFuse is due in large part to the proposed factorization of the original fusion model into two parts-(1) a generic fusion model shared by all cameras, and (2) lightweight camera-dependent transformations. Furthermore, the generic model is learned from many cameras by a meta-learning style algorithm to maximize its adaptation capability to various camera poses. We observe in experiments that MetaFuse finetuned on the public datasets outperforms the state-of-the-arts by a large margin which validates its value in practice.
Rongchang Xie, Chunyu Wang 0001, Yizhou Wang 0001
CVPR3
2020 A Two-Stage Multi-Objective Deep Reinforcement Learning Framework
Diqi Chen, Yizhou Wang 0001, Wen Gao 0001
ECAI2
2020 TCGM: An Information-Theoretic Framework for Semi-supervised Multi-modality Learning
Xinwei Sun 0001, Peng Cao 0004, Yuqing Kong, Lingjing Hu, Shanghang Zhang, Yizhou Wang 0001
ECCV (3)7
2020 Augmented Bi-path Network for Few-shot Learning
abstract
Few-shot Learning (FSL) which aims to learn from few labeled training data is becoming a popular research topic, due to the expensive labeling cost in many real-world applications. One kind of successful FSL method learns to compare the testing (query) image and training (support) image by simply concatenating the features of two images and feeding it into the neural network. However, with few labeled data in each class, the neural network has difficulty in learning or comparing the local features of two images. Such simple image-level comparison may cause serious mis-classification. To solve this problem, we propose Augmented Bi-path Network (ABNet) for learning to compare both global and local features on multi-scales. Specifically, the salient patches are extracted and embedded as the local features for every image. Then, the model learns to augment the features for better robustness. Finally, the model learns to compare global and local features separately, i.e., in two paths, before merging the similarities. Extensive experiments show that the proposed ABNet outperforms the state-of-the-art methods. Both quantitative and visual ablation studies are provided to verify that the proposed modules lead to more precise comparison results.
Baoming Yan, Bo Zhao 0015, Kan Guo, Ming Zhang 0004, Yizhou Wang 0001
ICPR8
2020 Towards Robust Bone Age Assessment: Rethinking Label Noise and Ambiguity
Ping Gong 0002, Yizhou Wang 0001, Yizhou Yu
MICCAI (6)3
2020 Unsupervised Surgical Instrument Segmentation via Anchor Generation and Semantic Diffusion
Daochang Liu, Yuhui Wei, Tingting Jiang 0001, Yizhou Wang 0001, Rulin Miao
MICCAI (3)4
2020 Multi-stream Progressive Up-Sampling Network for Dense CT Image Reconstruction
Qiuyue Liu, Feng Liu 0036, Xiangming Fang, Yizhou Yu, Yizhou Wang 0001
MICCAI (6)6
2020 Context-Aware Refinement Network Incorporating Structural Connectivity Prior for Brain Midline Delineation
Kongming Liang, Yizhou Yu, Yizhou Wang 0001
MICCAI (7)5
2020 BR-GAN: Bilateral Residual Generating Adversarial Network for Mammogram Classification
Chu-ran Wang, Fandong Zhang, Yizhou Yu, Yizhou Wang 0001
MICCAI (2)4
2020 Revisiting 3D Context Modeling with Supervised Pre-training for Universal Lesion Detection in CT Slices
Shu Zhang 0001, Jincheng Xu, Yu-Chun Chen, Jiechao Ma, Yizhou Wang 0001, Yizhou Yu
MICCAI (4)6
2020 Learning Multi-Agent Coordination for Enhancing Target Coverage in Directional Sensor Networks
abstract
Maximum target coverage by adjusting the orientation of distributed sensors is an important problem in directional sensor networks (DSNs). This problem is challenging as the targets usually move randomly but the coverage range of sensors is limited in angle and distance. Thus, it is required to coordinate sensors to get ideal target coverage with low power consumption, e.g. no missing targets or reducing redundant coverage. To realize this, we propose a Hierarchical Target-oriented Multi-Agent Coordination (HiT-MAC), which decomposes the target coverage problem into two-level tasks: targets assignment by a coordinator and tracking assigned targets by executors. Specifically, the coordinator periodically monitors the environment globally and allocates targets to each executor. In turn, the executor only needs to track its assigned targets. To effectively learn the HiT-MAC by reinforcement learning, we further introduce a bunch of practical methods, including a self-attention module, marginal contribution approximation for the coordinator, goal-conditional observation filter for the executor, etc. Empirical results demonstrate the advantage of HiT-MAC in coverage rate, learning efficiency, and scalability, comparing to baselines. We also conduct an ablative analysis on the effectiveness of the introduced components in the framework.
Fangwei Zhong, Yizhou Wang 0001
NeurIPS3
2020 Combining a gradient-based method and an evolution strategy for multi-objective reinforcement learning
Diqi Chen, Yizhou Wang 0001, Wen Gao 0001
Appl. Intell.2
2020 End-to-End Active Object Tracking and Its Real-World Deployment via Reinforcement Learning
abstract
We study active object tracking, where a tracker takes visual observations (i.e., frame sequences) as input and produces the corresponding camera control signals as output (e.g., move forward, turn left, etc.). Conventional methods tackle tracking and camera control tasks separately, and the resulting system is difficult to tune jointly. These methods also require significant human efforts for image labeling and expensive trial-and-error system tuning in the real world. To address these issues, we propose, in this paper, an end-to-end solution via deep reinforcement learning. A ConvNet-LSTM function approximator is adopted for the direct frame-to-action prediction. We further propose an environment augmentation technique and a customized reward function, which are crucial for successful training. The tracker trained in simulators (ViZDoom and Unreal Engine) demonstrates good generalization behaviors in the case of unseen object moving paths, unseen object appearances, unseen backgrounds, and distracting objects. The system is robust and can restore tracking after occasional lost of the target being tracked. We also find that the tracking ability, obtained solely from simulators, can potentially transfer to real-world scenarios. We demonstrate successful examples of such transfer, via experiments over the VOT dataset and the deployment of a real-world robot using the proposed active tracker trained in simulation.
Wenhan Luo, Peng Sun 0011, Fangwei Zhong, Wei Liu 0005, Tong Zhang 0001, Yizhou Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2020 No-Reference Image Quality Assessment: An Attention Driven Approach
abstract
In this paper, we tackle no-reference image quality assessment (NR-IQA), which aims to predict the perceptual quality of a distorted image without referencing its pristinequality counterpart. Inspired by the free-energy principle, we assume that, while perceiving a distorted image, the human visual system (HVS) tends to predict the pristine image then estimates the perceptual quality based on the distorted-restored pair. Furthermore, the perceptual quality depends heavily on the way how human beings attend to distorted images, namely, the cooperation of foveal vision and the eye movement mechanism. Inspired by these properties of the HVS, given the distortedrestored pair, we implement an attention-driven NR-IQA method with reinforcement learning (RL). The model learns a policy to attend to several regions parallelly. The observations of the fixation regions are aggregated in a weighted average way, which is inspired by the robust averaging strategy. For policy learning, the rewards are derived from two tasks-distortion type classification and perceptual score estimation. The goal of policy learning is to maximize the expectation of the accumulated rewards. Extensive experiments on LIVE, TID2008, TID2013 and CSIQ demonstrate the superiority of our methods.
Diqi Chen, Yizhou Wang 0001, Wen Gao 0001
IEEE Trans. Image Process.2
2020 Exploring Task Structure for Brain Tumor Segmentation From Multi-Modality MR Images
abstract
Brain tumor segmentation, which aims at segmenting the whole tumor area, enhancing tumor core area, and tumor core area from each input multi-modality bioimaging data, has received considerable attention from both academia and industry. However, the existing approaches usually treat this problem as a common semantic segmentation task without taking into account the underlying rules in clinical practice. In reality, physicians tend to discover different tumor areas by weighing different modality volume data. Also, they initially segment the most distinct tumor area, and then gradually search around to find the other two. We refer to the first property as the task-modality structure while the second property as the task-task structure, based on which we propose a novel task-structured brain tumor segmentation network (TSBTS net). Specifically, to explore the task-modality structure, we design a modality-aware feature embedding mechanism to infer the important weights of the modality data during network learning. To explore the tasktask structure, we formulate the prediction of the different tumor areas as conditional dependency sub-tasks and encode such dependency in the network stream. Experiments on BraTS benchmarks show that the proposed method achieves superior performance in segmenting the desired brain tumor areas while requiring relatively lower computational costs, compared to other state-of-the-art methods and baseline models.
Dingwen Zhang, Guohai Huang, Qiang Zhang 0020, Jungong Han, Junwei Han 0001, Yizhou Wang 0001, Yizhou Yu
IEEE Trans. Image Process.6
2019 Completeness Modeling and Context Separation for Weakly Supervised Temporal Action Localization
abstract
Temporal action localization is crucial for understanding untrimmed videos. In this work, we first identify two underexplored problems posed by the weak supervision for temporal action localization, namely action completeness modeling and action-context separation. Then by presenting a novel network architecture and its training strategy, the two problems are explicitly looked into. Specifically, to model the completeness of actions, we propose a multi-branch neural network in which branches are enforced to discover distinctive action parts. Complete actions can be therefore localized by fusing activations from different branches. And to separate action instances from their surrounding context, we generate hard negative data for training using the prior that motionless video clips are unlikely to be actions. Experiments performed on datasets THUMOS'14 and ActivityNet show that our framework outperforms state-of-the-art methods. In particular, the average mAP on ActivityNet v1.2 is significantly improved from 18.0% to 22.4%. Our code will be released soon.
Daochang Liu, Tingting Jiang 0001, Yizhou Wang 0001
CVPR3
2019 Cascaded Generative and Discriminative Learning for Microcalcification Detection in Breast Mammograms
abstract
Accurate microcalcification (μC) detection is of great importance due to its high proportion in early breast cancers. Most of the previous μC detection methods belong to discriminative models, where classifiers are exploited to distinguish μCs from other backgrounds. However, it is still challenging for these methods to tell the μCs from amounts of normal tissues because they are too tiny (at most 14 pixels). Generative methods can precisely model the normal tissues and regard the abnormal ones as outliers, while they fail to further distinguish the μCs from other anomalies, i.e. vessel calcifications. In this paper, we propose a hybrid approach by taking advantages of both generative and discriminative models. Firstly, a generative model named Anomaly Separation Network (ASN) is used to generate candidate μCs. ASN contains two major components. A deep convolutional encoder-decoder network is built to learn the image reconstruction mapping and a t-test loss function is designed to separate the distributions of the reconstruction residuals of μCs from normal tissues. Secondly, a discriminative model is cascaded to tell the μCs from the false positives. Finally, to verify the effectiveness of our method, we conduct experiments on both public and in-house datasets, which demonstrates that our approach outperforms previous state-of-the-art methods.
Fandong Zhang, Xinwei Sun 0001, Xiuli Li, Yizhou Yu, Yizhou Wang 0001
CVPR7
2019 Multi-Agent Tensor Fusion for Contextual Trajectory Prediction
abstract
Accurate prediction of others' trajectories is essential for autonomous driving. Trajectory prediction is challenging because it requires reasoning about agents' past movements, social interactions among varying numbers and kinds of agents, constraints from the scene context, and the stochasticity of human behavior. Our approach models these interactions and constraints jointly within a novel Multi-Agent Tensor Fusion (MATF) network. Specifically, the model encodes multiple agents' past trajectories and the scene context into a Multi-Agent Tensor, then applies convolutional fusion to capture multiagent interactions while retaining the spatial structure of agents and the scene context. The model decodes recurrently to multiple agents' future trajectories, using adversarial loss to learn stochastic predictions. Experiments on both highway driving and pedestrian crowd datasets show that the model achieves state-of-the-art prediction accuracy.
Tianyang Zhao 0004, Mathew Monfort, Wongun Choi, Chris L. Baker, Yibiao Zhao, Yizhou Wang 0001, Ying Nian Wu
CVPR7
2019 CRAVES: Controlling Robotic Arm With a Vision-Based Economic System
abstract
Training a robotic arm to accomplish real-world tasks has been attracting increasing attention in both academia and industry. This work discusses the role of computer vision algorithms in this field. We focus on low-cost arms on which no sensors are equipped and thus all decisions are made upon visual recognition, e.g., real-time 3D pose estimation. This requires annotating a lot of training data, which is not only time-consuming but also laborious. In this paper, we present an alternative solution, which uses a 3D model to create a large number of synthetic data, trains a vision model in this virtual domain, and applies it to real-world images after domain adaptation. To this end, we design a semi-supervised approach, which fully leverages the geometric constraints among keypoints. We apply an iterative algorithm for optimization. Without any annotations on real images, our algorithm generalizes well and produces satisfying results on 3D pose estimation, which is evaluated on two real-world datasets. We also construct a vision-based control system for task accomplishment, for which we train a reinforcement learning agent in a virtual environment and apply it to the real-world. Moreover, our approach, with merely a 3D model being required, has the potential to generalize to other types of multi-rigid-body dynamic systems.
Yiming Zuo 0001, Weichao Qiu, Lingxi Xie, Fangwei Zhong, Yizhou Wang 0001, Alan L. Yuille
CVPR5
2019 Optimizing Network Structure for 3D Human Pose Estimation
abstract
A human pose is naturally represented as a graph where the joints are the nodes and the bones are the edges. So it is natural to apply Graph Convolutional Network (GCN) to estimate 3D poses from 2D poses. In this work, we propose a generic formulation where both GCN and Fully Connected Network (FCN) are its special cases. From this formulation, we discover that GCN has limited representation power when used for estimating 3D poses. We overcome the limitation by introducing Locally Connected Network (LCN) which is naturally implemented by this generic formulation. It notably improves the representation capability over GCN. In addition, since every joint is only connected to a few joints in its neighborhood, it has strong generalization power. The experiments on public datasets show it: (1) outperforms the state-of-the-arts; (2) is less data hungry than alternative models; (3) generalizes well to unseen actions and datasets.
Hai Ci, Chunyu Wang 0001, Xiaoxuan Ma 0001, Yizhou Wang 0001
ICCV4
2019 Align, Attend and Locate: Chest X-Ray Diagnosis via Contrast Induced Attention Network With Limited Supervision
abstract
Obstacles facing accurate identification and localization of diseases in chest X-ray images lie in the lack of high-quality images and annotations. In this paper, we propose a Contrast Induced Attention Network (CIA-Net), which exploits the highly structured property of chest X-ray images and localizes diseases via contrastive learning on the aligned positive and negative samples. To force the attention module to focus only on sites of abnormalities, we also introduce a learnable alignment module to adjust all the input images, which eliminates variations of scales, angles, and displacements of X-ray images generated under bad scan conditions. We show that the use of contrastive attention and alignment module allows the model to learn rich identification and localization information using only a small amount of location annotations, resulting in state-of-the-art performance in NIH chest X-ray dataset.
Jingyu Liu 0004, Gangming Zhao, Ming Zhang 0004, Yizhou Wang 0001, Yizhou Yu
ICCV5
2019 Learning With Unsure Data for Medical Image Diagnosis
abstract
In image-based disease prediction, it can be hard to give certain cases a deterministic “disease/normal” label due to lack of enough information, e.g., at its early stage. We call such cases “unsure” data. Labeling such data as unsure suggests follow-up examinations so as to avoid irreversible medical accident/loss in contrast to incautious prediction. This is a common practice in clinical diagnosis, however, mostly neglected by existing methods. Learning with unsure data also inter-weaves with two other practical issues: (i) data imbalance issue that may incur model-bias towards the majority class, and (ii) conservative/aggressive strategy consideration, i.e., the negative (normal) samples and positive (disease) samples should NOT be treated equally - the former should be detected with a high precision (conservativeness) and the latter should be detected with a high recall (aggression) to avoid missing opportunity for treatment. Mixed with these issues, learning with unsure data becomes particularly challenging. In this paper, we raise “learning with unsure data” problem and formulate it as an ordinal regression and propose a unified end-to-end learning framework, which also considers the aforementioned two issues: (i) incorporate cost-sensitive parameters to alleviate the data imbalance problem, and (ii) execute the conservative and aggressive strategies by introducing two parameters in the training procedure. The benefits of learning with unsure data and validity of our models are demonstrated on the prediction of Alzheimer's Disease and lung nodules.
Botong Wu, Xinwei Sun 0001, Lingjing Hu, Yizhou Wang 0001
ICCV4
2019 Max-MIG: an Information Theoretic Approach for Joint Learning from Crowds
Peng Cao 0004, Yuqing Kong, Yizhou Wang 0001
ICLR (Poster)4
2019 AD-VAT: An Asymmetric Dueling mechanism for learning Visual Active Tracking
Fangwei Zhong, Peng Sun 0011, Wenhan Luo, Tingyun Yan, Yizhou Wang 0001
ICLR (Poster)5
2019 Large-Scale Datasets for Going Deeper in Image Understanding
abstract
Recently, extensive efforts have been devoted to computer vision and machine learning by exploiting big data to explore many practical applications. However, these research fields are still quite limited not only by the sheer volume, but also the versatility and diversity, of the available datasets. In this paper, we target at four challenging and yet important computer vision tasks, namely, human-centered scene classification, attribute based zero-shot learning (recognition), human keypoint detection and image Chinese captioning. Four novel large-scale datasets are collected and annotated to facilitate these tasks of deeper image understanding. Labels, bounding boxes, attributes, keypoints and captions are annotated in corresponding datasets. These rich annotations bridge the semantic gap between low-level images and high-level concepts. Extensive experiments on baseline methods have been implemented and compared, which show that these learning tasks on our datasets are still challenging.
Jiahong Wu 0006, He Zheng, Bo Zhao 0015, Baoming Yan, Shipei Zhou, Guosen Lin, Yanwei Fu 0001, Yizhou Wang 0001
ICME11
2019 MVP-Net: Multi-view FPN with Position-Aware Attention for Deep Universal Lesion Detection
Shu Zhang 0001, Junge Zhang, Kaiqi Huang, Yizhou Wang 0001, Yizhou Yu
MICCAI (6)5
2019 Surgical Skill Assessment on In-Vivo Clinical Data via the Clearness of Operating Field
Daochang Liu, Tingting Jiang 0001, Yizhou Wang 0001, Rulin Miao
MICCAI (5)3
2019 From Unilateral to Bilateral Learning: Detecting Mammogram Masses with Contrasted Bilateral Network
Shu Zhang 0001, Qianyi Zhang, Fandong Zhang, Xiuli Li, Yizhou Wang 0001, Yizhou Yu
MICCAI (6)8
2019 L_DMI: A Novel Information-theoretic Loss Function for Training Deep Nets Robust to Label Noise
abstract
Accurately annotating large scale dataset is notoriously expensive both in time and in money. Although acquiring low-quality-annotated dataset can be much cheaper, it often badly damages the performance of trained models when using such dataset without particular treatment. Various methods have been proposed for learning with noisy labels. However, most methods only handle limited kinds of noise patterns, require auxiliary information or steps (e.g., knowing or estimating the noise transition matrix), or lack theoretical justification. In this paper, we propose a novel information-theoretic loss function, LDMI, for training deep neural networks robust to label noise. The core of LDMI is a generalized version of mutual information, termed Determinant based Mutual Information (DMI), which is not only information-monotone but also relatively invariant. To the best of our knowledge, LDMI is the first loss function that is provably robust to instance-independent label noise, regardless of noise pattern, and it can be applied to any existing classification neural networks straightforwardly without any auxiliary information. In addition to theoretical justification, we also empirically show that using LDMI outperforms all other counterparts in the classification task on both image dataset and natural language dataset include Fashion-MNIST, CIFAR-10, Dogs vs. Cats, MR with a variety of synthesized noise patterns and noise amounts, as well as a real-world dataset Clothing1M.
Peng Cao 0004, Yuqing Kong, Yizhou Wang 0001
NeurIPS4
2019 No-Reference Image Quality Assessment: An Attention Driven Approach
abstract
In this paper, we tackle no-reference image quality assessment (NR-IQA), which aims to predict the perceptual quality of a test image without referencing its pristine-quality counterpart. The free-energy brain theory implies that the human visual system (HVS) tends to predict the pristine image while perceiving a distorted one. Besides, image quality assessment heavily depends on the way how human beings attend to distorted images. Motivated by that, the distorted image is restored first. Then given the distorted-restored pair, we make the first attempt to formulate the NR-IQA as a dynamic attentional process and implement it via reinforcement learning. The reward is derived from two tasks-classifying the distortion type and predicting the perceptual score of a test image. The model learns a policy to sample a sequence of fixation areas with a goal to maximize the expectation of the accumulated rewards. The observations of the fixation areas are aggregated through a recurrent neural network (RNN) and the robust averaging strategy which assigns different weights on different fixation areas. Extensive experiments on TID2008, TID2013 and CSIQ demonstrate the superiority of our method.
Diqi Chen, Yizhou Wang 0001, Hongyu Ren, Wen Gao 0001
WACV2
2019 Soft Transfer Learning via Gradient Diagnosis for Visual Relationship Detection
abstract
Detecting all visual relationships is posed as the most fundamental task towards the ultimate semantic reasoning. However, due to the rich context embedded in the image and diverse language ambiguities, it is unrealistic to annotate and list all possible relationships for providing a noise-free supervised setting. All prior approaches simply adopt the traditional fully-supervised detection pipeline and ignore the effect of incomplete annotations on model convergence, resulting in the unstable optimization and unsatisfactory performance. In this work, we make the first attempt to address this critical incomplete annotations issue and reformulate this task via the Soft Transfer Learning (STL), which aims to transfer knowledge learned from the annotations in hand into the uncertain pairs in a self-supervised way. The knowledge transfer process is inferred from a principled gradient diagnosis. Extensive experiments on VRD and the large-scale VG benchmarks demonstrate the superiority of our STL method.
Diqi Chen, Xiaodan Liang, Yizhou Wang 0001, Wen Gao 0001
WACV3
2019 Zero-Shot Learning Via Recurrent Knowledge Transfer
abstract
Zero-shot learning (ZSL) which aims to learn new concepts without any labeled training data is a promising solution to large-scale concept learning. Recently, many works implement zero-shot learning by transferring structural knowledge from the semantic embedding space to the image feature space. However, we observe that such direct knowledge transfer may suffer from the space shift problem in the form of the inconsistency of geometric structures in the training and testing spaces. To alleviate this problem, we propose a novel method which actualizes recurrent knowledge transfer (RecKT) between the two spaces. Specifically, we unite the two spaces into the joint embedding space in which unseen image data are missing. The proposed method provides a synthesis-refinement mechanism to learn the shared subspace structure (SSS) and synthesize missing data simultaneously in the joint embedding space. The synthesized unseen image data are utilized to construct the classifier for unseen classes. Experimental results show that our method outperforms the state-of-the-art on three popular datasets. The ablation experiment and visualization of the learning process illustrate how our method can alleviate the space shift problem. By product, our method provides a perspective to interpret the ZSL performance by implementing subspace clustering on the learned SSS.
Bo Zhao 0015, Xinwei Sun 0001, Xiaopeng Hong, Yuan Yao 0011, Yizhou Wang 0001
WACV5
2019 Robust 3D Human Pose Estimation from Single Images or Video Sequences
abstract
We propose a method for estimating 3D human poses from single images or video sequences. The task is challenging because: (a) many 3D poses can have similar 2D pose projections which makes the lifting ambiguous, and (b) current 2D joint detectors are not accurate which can cause big errors in 3D estimates. We represent 3D poses by a sparse combination of bases which encode structural pose priors to reduce the lifting ambiguity. This prior is strengthened by adding limb length constraints. We estimate the 3D pose by minimizing an$L_1$norm measurement error between the 2D pose and the 3D pose because it is less sensitive to inaccurate 2D poses. We modify our algorithm to output$K$3D pose candidates for an image, and for videos, we impose a temporal smoothness constraint to select the best sequence of 3D poses from the candidates. We demonstrate good results on 3D pose estimation from static images and improved performance by selecting the best 3D pose from the$K$proposals. Our results on video sequences also show improvements (over static images) of roughly 15%.
Chunyu Wang 0001, Yizhou Wang 0001, Zhouchen Lin, Alan L. Yuille
IEEE Trans. Pattern Anal. Mach. Intell.2
2019 Fast and Resilient Indoor Floor Plan Construction with a Single User
abstract
A lack of floor plans is a fundamental obstacle to ubiquitous indoor location-based services. Recent work have made significant progress to accuracy, but they largely rely on slow crowdsensing that may take weeks or even months to collect enough data. In this paper, we propose Knitter that can generate accurate floor maps by a single random user’s one hour data collection efforts, and demonstrate how such maps can be used for indoor navigation. Knitter extracts high quality floor layout information from single images, calibrates user trajectories, and filters outliers. It uses a multi-hypothesis map fusion framework that updates landmark positions/orientations and accessible areas incrementally according to evidences from each measurement. Our experiments on three different large buildings (up to$140\times 50\;\mathrm{m}^2$) with 30+ users show that Knitter produces correct map topology, with landmark location errors of$3\sim 5\;\mathrm{m}$and orientation errors of$4\sim 6^\circ$, both at 90-percentile. Our results are comparable to the state-of-the-art at more than$20\times$speed up: data collection in each of the three buildings can finish in about one hour even by a novice user trained just a few minutes.
Ruipeng Gao, Bing Zhou 0001, Fan Ye 0003, Yizhou Wang 0001
IEEE Trans. Mob. Comput.4
2018 RAN4IQA: Restorative Adversarial Nets for No-Reference Image Quality Assessment
abstract
Inspired by the free-energy brain theory, which implies that human visual system (HVS) tends to reduce uncertainty and restore perceptual details upon seeing a distorted image, we propose restorative adversarial net (RAN), a GAN-based model for no-reference image quality assessment (NR-IQA). RAN, which mimics the process of HVS, consists of three components: a restorator, a discriminator and an evaluator. The restorator restores and reconstructs input distorted image patches, while the discriminator distinguishes the reconstructed patches from the pristine distortion-free patches. After restoration, we observe that the perceptual distance between the restored and the distorted patches is monotonic with respect to the distortion level. We further define Gain of Restoration (GoR) based on this phenomenon. The evaluator predicts perceptual score by extracting feature representations from the distorted and restored patches to measure GoR. Eventually, the quality score of an input image is estimated by weighted sum of the patch scores. Experimental results on Waterloo Exploration, LIVE and TID2013 show the effectiveness and generalization ability of RAN compared to the state-of-the-art NR-IQA models.
Hongyu Ren, Diqi Chen, Yizhou Wang 0001
AAAI3
2018 Video Object Segmentation by Learning Location-Sensitive Embeddings
Hai Ci, Chunyu Wang 0001, Yizhou Wang 0001
ECCV (11)3
2018 End-to-end Active Object Tracking via Reinforcement Learning
abstract
We study active object tracking, where a tracker takes as input the visual observation (i.e. frame sequence) and produces the camera control signal (e.g., move forward, turn left, etc). Conventional methods tackle the tracking and the camera control separately, which is challenging to tune jointly. It also incurs many human efforts for labeling and many expensive trial-and-errors in real-world. To address these issues, we propose, in this paper, an end-to-end solution via deep reinforcement learning, where a ConvNet-LSTM function approximator is adopted for the direct frame-to-action prediction. We further propose an environment augmentation technique and a customized reward function, which are crucial for a successful training. The tracker trained in simulators (ViZDoom, Unreal Engine) shows good generalization in the case of unseen object moving path, unseen object appearance, unseen background, and distracting object. It can restore tracking when occasionally losing the target. With the experiments over the VOT dataset, we also find that the tracking ability, obtained solely from simulators, can potentially transfer to real-world scenarios.
Wenhan Luo, Peng Sun 0011, Fangwei Zhong, Wei Liu 0005, Tong Zhang 0001, Yizhou Wang 0001
ICML6
2018 MSplit LBI: Realizing Feature Selection and Dense Estimation Simultaneously in Few-shot and Zero-shot Learning
abstract
It is one typical and general topic of learning a good embedding model to efficiently learn the representation coefficients between two spaces/subspaces. To solve this task, $L_{1}$ regularization is widely used for the pursuit of feature selection and avoiding overfitting, and yet the sparse estimation of features in $L_{1}$ regularization may cause the underfitting of training data. $L_{2}$ regularization is also frequently used, but it is a biased estimator. In this paper, we propose the idea that the features consist of three orthogonal parts, namely sparse strong signals, dense weak signals and random noise, in which both strong and weak signals contribute to the fitting of data. To facilitate such novel decomposition, MSplit LBI is for the first time proposed to realize feature selection and dense estimation simultaneously. We provide theoretical and simulational verification that our method exceeds $L_{1}$ and $L_{2}$ regularization, and extensive experimental results show that our method achieves state-of-the-art performance in the few-shot and zero-shot learning.
Bo Zhao 0015, Xinwei Sun 0001, Yanwei Fu 0001, Yuan Yao 0011, Yizhou Wang 0001
ICML5
2018 FDR-HS: An Empirical Bayesian Identification of Heterogenous Features in Neuroimage Analysis
Xinwei Sun 0001, Lingjing Hu, Fandong Zhang, Yuan Yao 0011, Yizhou Wang 0001
MICCAI (1)5
2018 Detect-SLAM: Making Object Detection and SLAM Mutually Beneficial
abstract
Although significant progress has been made in SLAM and object detection in recent years, there are still a series of challenges for both tasks, e.g., SLAM in dynamic environments and detecting objects in complex environments. To address these challenges, we present a novel robotic vision system, which integrates SLAM with a deep neural networkbased object detector to make the two functions mutually beneficial. The proposed system facilitates a robot to accomplish tasks reliably and efficiently in an unknown and dynamic environment. Experimental results show that compare to the state-of-the-art robotic vision systems, the proposed system has three advantages: i) it greatly improves the accuracy and robustness of SLAM in dynamic environments by removing unreliable features from moving objects leveraging the object detector, ii) it builds an instance-level semantic map of the environment in an online fashion using the synergy of the two functions for further semantic applications; and iii) it improves the object detector so that it can detect/recognize objects effectively under more challenging conditions such as unusual viewpoints, poor lighting condition, and motion blur, by leveraging the object map.
Fangwei Zhong, Ziqi Zhang 0017, China Chen, Yizhou Wang 0001
WACV5
2017 Learning Discriminative Activated Simplices for Action Recognition
abstract
We address the task of action recognition from a sequence of 3D human poses. This is a challenging task firstly because the poses of the same class could have large intra-class variations either caused by inaccurate 3D pose estimation or various performing styles. Also different actions, e.g., walking vs. jogging, may share similar poses which makes the representation not discriminative to differentiate the actions. To solve the problems, we propose a novel representation for 3D poses by a mixture of Discriminative Activated Simplices (DAS). Each DAS consists of a few bases and represent pose data by their convex combinations. The discriminative power of DAS is firstly realized by learning discriminative bases across classes with a block diagonal constraint enforced on the basis coefficient matrix. Secondly, the DAS provides tight characterization of the pose manifolds thus reducing the chance of generating overlapped DAS between similar classes. We justify the power of the model on benchmark datasets and witness consistent performance improvements.
Chenxu Luo, Chunyu Wang 0001, Yizhou Wang 0001
AAAI4
2017 Collaborative Deep Reinforcement Learning for Joint Object Search
abstract
We examine the problem of joint top-down active search of multiple objects under interaction, e.g., person riding a bicycle, cups held by the table, etc. Such objects under interaction often can provide contextual cues to each other to facilitate more efficient search. By treating each detector as an agent, we present the first collaborative multi-agent deep reinforcement learning algorithm to learn the optimal policy for joint active object localization, which effectively exploits such beneficial contextual information. We learn inter-agent communication through cross connections with gates between the Q-networks, which is facilitated by a novel multi-agent deep Q-learning algorithm with joint exploitation sampling. We verify our proposed method on multiple object detection benchmarks. Not only does our model help to improve the performance of state-of-the-art active localization models, it also reveals interesting co-detection patterns that are intuitively interpretable.
Bo Xin, Yizhou Wang 0001, Gang Hua 0001
CVPR3
2017 Face Album: Towards automatic photo management based on person identity on mobile phones
abstract
We implement a new photo management system `Face Album' on mobile phones, which organizes photos by person identity, as is shown in Fig. 1. We automatically group faces into clusters to release user workload. Our system is composed of two pools: a certain pool with reliable clusters consisting of faces from same identity, and an uncertain pool containing faces that are lacking in evidence to be recognized. Constantly as new faces increase, the certain pool and uncertain pool work together to either assign new faces to existing clusters or discover new identities in the album. In addition, user interaction is introduced for some deviation corrections. Experiments indicate that our results are close to offline hierarchical clustering method while a subjective survey shows our photo management system is favored by users.
Yuansheng Xu, Fangyue Peng, Yizhou Wang 0001
ICASSP4
2017 Knitter: Fast, resilient single-user indoor floor plan construction
abstract
Lacking of floor plans is a fundamental obstacle to ubiquitous indoor location-based services. Recent work have made significant progress to accuracy, but they largely rely on slow crowdsensing that may take weeks or even months to collect enough data. In this paper, we propose Knitter that can generate accurate floor maps by a single random user's one hour data collection efforts. Knitter extracts high quality floor layout information from single images, calibrates user trajectories and filters outliers. It uses a multi-hypothesis map fusion framework that updates landmark positions/orientations and accessible areas incrementally according to evidences from each measurement. Our experiments on 3 different large buildings and 30+ users show that Knitter produces correct map topology, and 90-percentile landmark location and orientation errors of 3 ~ 5m and 4 ~ 6°, comparable to the state-of-the-art at more than 20× speed up: data collection can finish in about one hour even by a novice user trained just a few minutes.
Ruipeng Gao, Bing Zhou 0001, Fan Ye 0003, Yizhou Wang 0001
INFOCOM4
2017 GSplit LBI: Taming the Procedural Bias in Neuroimaging for Disease Prediction
Xinwei Sun 0001, Lingjing Hu, Yuan Yao 0011, Yizhou Wang 0001
MICCAI (3)4
2017 UnrealCV: Virtual Worlds for Computer Vision
abstract
UnrealCV is a project to help computer vision researchers build virtual worlds using Unreal Engine 4 (UE4). It extends UE4 with a plugin by providing (1) A set of UnrealCV commands to interact with the virtual world. (2) Communication between UE4 and an external program, such as Caffe. UnrealCV can be used in two ways. The first one is using a compiled game binary with UnrealCV embedded. This is as simple as running a game, no knowledge of Unreal Engine is required. The second is installing UnrealCV plugin to Unreal Engine 4 (UE4) and use the editor of UE4 to build a new virtual world. UnrealCV is an open-source software under the MIT license. Since the initial release in September 2016, it has gathered an active community of users, including students and researchers.
Weichao Qiu, Fangwei Zhong, Yi Zhang 0099, Siyuan Qiao, Zihao Xiao 0001, Tae Soo Kim 0001, Yizhou Wang 0001
ACM Multimedia7
2017 Data-Dependent Sparsity for Subspace Clustering
Bo Xin, Yizhou Wang 0001, Wen Gao 0001, David P. Wipf
UAI2
2017 Smartphone-Based Real Time Vehicle Tracking in Indoor Parking Structures
abstract
Although location awareness and turn-by-turn instructions are prevalent outdoors due to GPS, we are back into the darkness in uninstrumented indoor environments such as underground parking structures. We get confused, disoriented when driving in these mazes, and frequently forget where we parked, ending up circling back and forth upon return. In this paper, we propose VeTrack, asmartphone-only system that tracks the vehicle’s location in real time using the phone’s inertial sensors. It does not require any environment instrumentation or cloud backend. It uses a novel “shadow” trajectory tracing method to accurately estimate phone’s and vehicle’s orientations despite their arbitrary poses and frequent disturbances. We develop algorithms in a Sequential Monte Carlo framework to represent vehicle states probabilistically, and harness constraints by the garage map and detected landmarks to robustly infer the vehicle location. We also find landmark (e.g., speed bumps, turns) recognition methods reliable against noises, disturbances from bumpy rides, and even hand-held movements. We implement a highly efficient prototype and conduct extensive experiments in multiple parking structures of different sizes and structures, and collect data with multiple vehicles and drivers. We find that VeTrack can estimate the vehicle’s real time location with almost negligible latency, with error of$2\sim 4$parking spaces at the 80th percentile.
Ruipeng Gao, Mingmin Zhao, Fan Ye 0003, Yizhou Wang 0001, Guojie Luo
IEEE Trans. Mob. Comput.5
2016 Robust Complex Behaviour Modeling at 90Hz
abstract
Modeling complex crowd behaviour for tasks such as rare event detection has received increasing interest. However, existing methods are limited because (1) they are sensitive to noise often resulting in a large number of false alarms; and (2) they rely on elaborate models leading to high computational cost thus unsuitable for processing a large number of video inputs in real-time. In this paper, we overcome these limitations by introducing a novel complex behaviour modeling framework, which consists of a Binarized Cumulative Directional (BCD) feature as representation, novel spatial and temporal context modeling via an iterative correlation maximization, and a set of behaviour models, each being a simple Bernoulli distribution. Despite its simplicity, our experiments on three benchmark datasets show that it significantly outperforms the state-of-the-art for both temporal video segmentation and rare event detection. Importantly, it is extremely efficient — reaches 90Hz on a normal PC platform using MATLAB.
Yizhou Wang 0001, Tao Xiang 0002
AAAI2
2016 Recognizing Actions in 3D Using Action-Snippets and Activated Simplices
abstract
Pose-based action recognition in 3D is the task of recognizing an action (e.g., walking or running) from a sequence of 3D skeletal poses. This is challenging because of variations due to different ways of performing the same action and inaccuracies in the estimation of the skeletal poses. The training data is usually small and hence complex classifiers risk over-fitting the data. We address this task by action-snippets which are short sequences of consecutive skeletal poses capturing the temporal relationships between poses in an action. We propose a novel representation for action-snippets, called activated simplices. Each activity is represented by a manifold which is approximated by an arrangement of activated simplices. A sequence (of action-snippets) is classified by selecting the closest manifold and outputting the corresponding activity. This is a simple classifier which helps avoid over-fitting the data but which significantly outperforms state-of-the-art methods on standard benchmarks.
Chunyu Wang 0001, John Flynn, Yizhou Wang 0001, Alan L. Yuille
AAAI3
2016 Mining 3D Key-Pose-Motifs for Action Recognition
abstract
Recognizing an action from a sequence of 3D skeletal poses is a challenging task. First, different actors may perform the same action in various styles. Second, the estimated poses are sometimes inaccurate. These challenges can cause large variations between instances of the same class. Third, the datasets are usually small, with only a few actors performing few repetitions of each action. Hence training complex classifiers risks over-fitting the data. We address this task by mining a set of key-pose-motifs for each action class. A key-pose-motif contains a set of ordered poses, which are required to be close but not necessarily adjacent in the action sequences. The representation is robust to style variations. The key-pose-motifs are represented in terms of a dictionary using soft-quantization to deal with inaccuracies caused by quantization. We propose an efficient algorithm to mine key-pose-motifs taking into account of these probabilities. We classify a sequence by matching it to the motifs of each class and selecting the class that maximizes the matching score. This simple classifier obtains state-of the-art performance on two benchmark datasets.
Chunyu Wang 0001, Yizhou Wang 0001, Alan L. Yuille
CVPR2
2016 Face Detection with End-to-End Integration of a ConvNet and a 3D Model
Yunzhu Li, Benyuan Sun, Tianfu Wu 0001, Yizhou Wang 0001
ECCV (3)4
2016 Neighborhood-Preserving Hashing for Large-Scale Cross-Modal Search
abstract
In the literature of cross-modal search, most methods employ linear models to pursue hash codes that preserve data similarity, in terms of Euclidean distance, both within-modal and across-modal. However, data dimensionality can be quite different across modalities. It is known that the behavior of Euclidean distance/similarity between datapoints can be drastically different in linear spaces of different dimensionality. In this paper, we identify this "variation of dimensionality" problem in cross-modal search that may harm most of distance-based methods. We propose a semi-supervised nonlinear probabilistic cross-modal hashing method, namely Neighborhood-Preserving Hashing (NPH), to alleviate the negative effect due to the variation of dimensionality issue. Inspired by tSNE \cite{tSNE_van2008visualizing}, rather than preserve pairwise data distances, we propose to learn hash codes that preserve neighborhood relationship of datapoints via matching their conditional distribution derived from distance to that of datapoints of multi-modalities. Experimental results on three real-world datasets demonstrate that the proposed method outperforms the state-of-the-art distance-based semi-supervised cross-modal hashing methods as well as many fully-supervised ones.
Botong Wu, Yizhou Wang 0001
ACM Multimedia2
2016 Maximal Sparsity with Deep Networks?
abstract
The iterations of many sparse estimation algorithms are comprised of a fixed linear filter cascaded with a thresholding nonlinearity, which collectively resemble a typical neural network layer. Consequently, a lengthy sequence of algorithm iterations can be viewed as a deep network with shared, hand-crafted layer weights. It is therefore quite natural to examine the degree to which a learned network model might act as a viable surrogate for traditional sparse estimation in domains where ample training data is available. While the possibility of a reduced computational budget is readily apparent when a ceiling is imposed on the number of layers, our work primarily focuses on estimation accuracy. In particular, it is well-known that when a signal dictionary has coherent columns, as quantified by a large RIP constant, then most tractable iterative algorithms are unable to find maximally sparse representations. In contrast, we demonstrate both theoretically and empirically the potential for a trained deep network to recover minimal $\ell_0$-norm representations in regimes where existing methods fail. The resulting system, which can effectively learn novel iterative sparse estimation algorithms, is deployed on a practical photometric stereo estimation problem, where the goal is to remove sparse outliers that can disrupt the estimation of surface normals from a 3D scene.
Bo Xin, Yizhou Wang 0001, Wen Gao 0001, David P. Wipf, Baoyuan Wang
NIPS2
2016 Robust Subjective Visual Property Prediction from Crowdsourced Pairwise Labels
abstract
The problem of estimating subjective visual properties from image and video has attracted increasing interest. A subjective visual property is useful either on its own (e.g. image and video interestingness) or as an intermediate representation for visual recognition (e.g. a relative attribute). Due to its ambiguous nature, annotating the value of a subjective visual property for learning a prediction model is challenging. To make the annotation more reliable, recent studies employ crowdsourcing tools to collect pairwise comparison labels. However, using crowdsourced data also introduces outliers. Existing methods rely on majority voting to prune the annotation outliers/errors. They thus require a large amount of pairwise labels to be collected. More importantly as a local outlier detection method, majority voting is ineffective in identifying outliers that can cause global ranking inconsistencies. In this paper, we propose a more principled way to identify annotation outliers by formulating the subjective visual property prediction task as a unified robust learning to rank problem, tackling both the outlier detection and learning to rank jointly. This differs from existing methods in that (1) the proposed method integrates local pairwise comparison labels together to minimise a cost that corresponds to global inconsistency of ranking order, and (2) the outlier detection and learning to rank problems are solved jointly. This not only leads to better detection of annotation outliers but also enables learning with extremely sparse annotations.
Yanwei Fu 0001, Timothy M. Hospedales, Tao Xiang 0002, Jiechao Xiong, Shaogang Gong, Yizhou Wang 0001, Yuan Yao 0011
IEEE Trans. Pattern Anal. Mach. Intell.6
2016 Efficient Generalized Fused Lasso and Its Applications
abstract
Generalized fused lasso (GFL) penalizes variables with l 1 norms based both on the variables and their pairwise differences. GFL is useful when applied to data where prior information is expressed using a graph over the variables. However, the existing GFL algorithms incur high computational costs and do not scale to high-dimensional problems. In this study, we propose a fast and scalable algorithm for GFL. Based on the fact that fusion penalty is the Lovász extension of a cut function, we show that the key building block of the optimization is equivalent to recursively solving graph-cut problems. Thus, we use a parametric flow algorithm to solve GFL in an efficient manner. Runtime comparisons demonstrate a significant speedup compared to existing GFL algorithms. Moreover, the proposed optimization framework is very general; by designing different cut functions, we also discuss the extension of GFL to directed graphs. Exploiting the scalability of the proposed algorithm, we demonstrate the applications of our algorithm to the diagnosis of Alzheimer’s disease (AD) and video background subtraction (BS). In the AD problem, we formulated the diagnosis of AD as a GFL regularized classification. Our experimental evaluations demonstrated that the diagnosis performance was promising. We observed that the selected critical voxels were well structured, i.e., connected, consistent according to cross validation, and in agreement with prior pathological knowledge. In the BS problem, GFL naturally models arbitrary foregrounds without predefined grouping of the pixels. Even by applying simple background models, e.g., a sparse linear combination of former frames, we achieved state-of-the-art performance on several public datasets.
Bo Xin, Yoshinobu Kawahara, Yizhou Wang 0001, Lingjing Hu, Wen Gao 0001
ACM Trans. Intell. Syst. Technol.3
2016 Sextant: Towards Ubiquitous Indoor Localization Service by Photo-Taking of the Environment
abstract
Mainstream indoor localization technologies rely on RF signatures that require extensive human efforts to measure and periodically recalibrate signatures. The progress to ubiquitous localization remains slow. In this study, we explore Sextant, an alternative approach that leverages environmental reference objects such as store logos. A user uses a smartphone to obtain relative position measurements to such static reference objects for the system to triangulate the user location. Sextant leverages image matching algorithms to automatically identify the chosen reference objects by photo-taking, and we propose two methods to systematically address image matching mistakes that cause large localization errors. We formulate the benchmark image selection problem, prove its NP-completeness, and propose a heuristic algorithm to solve it. We also propose a couple of geographical constraints to further infer unknown reference objects. To enable fast deployment, we propose a lightweight site survey method for service providers to quickly estimate the coordinates of reference objects. Extensive experiments have shown that Sextant prototype achieves 2-5 m accuracy at 80-percentile, comparable to the industry state-of-the-art, while covering a 150 x 75 m mall and 300 x 200m train station requires a one time investment of only 2-3 man-hours from service providers.
Ruipeng Gao, Fan Ye 0003, Guojie Luo, Kaigui Bian, Yizhou Wang 0001, Tao Wang 0004, Xiaoming Li 0001
IEEE Trans. Mob. Comput.6
2016 Multi-Story Indoor Floor Plan Reconstruction via Mobile Crowdsensing
abstract
The lack of floor plans is a critical reason behind the current sporadic availability of indoor localization service. Service providers have to go through effort-intensive and time-consuming business negotiations with building operators, or hire dedicated personnel to gather such data. In this paper, we propose Jigsaw, a floor plan reconstruction system that leverages crowdsensed data from mobile users. It extracts the position, size, and orientation information of individual landmark objects from images taken by users. It also obtains the spatial relation between adjacent landmark objects from inertial sensor data, then computes the coordinates and orientations of these objects on an initial floor plan. By combining user mobility traces and locations where images are taken, it produces complete floor plans with hallway connectivity, room sizes, and shapes. It also identifies different types of connection areas (e.g., escalators and stairs) between stories, and employs a refinement algorithm to correct detection errors. Our experiments on three stories of two large shopping malls show that the 90-percentile errors of positions and orientations of landmark objects are about 1~2m and 5~9°, while the hallway connectivity and connection areas between stories are 100 percent correct.
Ruipeng Gao, Mingmin Zhao, Fan Ye 0003, Guojie Luo, Yizhou Wang 0001, Kaigui Bian, Tao Wang 0004, Xiaoming Li 0001
IEEE Trans. Mob. Comput.6
2015 Stable Feature Selection from Brain sMRI
abstract
Neuroimage analysis usually involves learning thousands or even millions of variables using only a limited number of samples. In this regard, sparse models, e.g. the lasso, are applied to select the optimal features and achieve high diagnosis accuracy. The lasso, however, usually results in independent unstable features. Stability, a manifest of reproducibility of statistical results subject to reasonable perturbations to data and the model (Yu 2013), is an important focus in statistics, especially in the analysis of high dimensional data. In this paper, we explore a nonnegative generalized fused lasso model for stable feature selection in the diagnosis of Alzheimer's disease. In addition to sparsity, our model incorporates two important pathological priors: the spatial cohesion of lesion voxels and the positive correlation between the features and the disease labels. To optimize the model, we propose an efficient algorithm by proving a novel link between total variation and fast network flow algorithms via conic duality. Experiments show that the proposed nonnegative model performs much better in exploring the intrinsic structure of data via selecting stable features compared with other state-of-the-arts.
Bo Xin, Lingjing Hu, Yizhou Wang 0001, Wen Gao 0001
AAAI3
2015 Background Subtraction via generalized fused lasso foreground modeling
abstract
Background Subtraction (BS) is one of the key steps in video analysis. Many background models have been proposed and achieved promising performance on public data sets. However, due to challenges such as illumination change, dynamic background etc. the resulted foreground segmentation often consists of holes as well as background noise. In this regard, we consider generalized fused lasso regularization to quest for intact structured foregrounds. Together with certain assumptions about the background, such as the low-rank assumption or the sparse-composition assumption (depending on whether pure background frames are provided), we formulate BS as a matrix decomposition problem using regularization terms for both the foreground and background matrices. Moreover, under the proposed formulation, the two generally distinctive background assumptions can be solved in a unified manner. The optimization was carried out via applying the augmented Lagrange multiplier (ALM) method in such a way that a fast parametric-flow algorithm is used for updating the foreground matrix. Experimental results on several popular BS data sets demonstrate the advantage of the proposed model compared to state-of-the-arts.
Bo Xin, Yizhou Wang 0001, Wen Gao 0001
CVPR3
2015 Exploiting Object Similarity in 3D Reconstruction
abstract
Despite recent progress, reconstructing outdoor scenes in 3D from movable platforms remains a highly difficult endeavour. Challenges include low frame rates, occlusions, large distortions and difficult lighting conditions. In this paper, we leverage the fact that the larger the reconstructed area, the more likely objects of similar type and shape will occur in the scene. This is particularly true for outdoor scenes where buildings and vehicles often suffer from missing texture or reflections, but share similarity in 3D shape. We take advantage of this shape similarity by localizing objects using detectors and jointly reconstructing them while learning a volumetric model of their shape. This allows us to reduce noise while completing missing surfaces as objects of similar shape benefit from all observations for the respective category. We evaluate our approach with respect to LIDAR ground truth on a novel challenging suburban dataset and show its advantages over the state-of-the-art.
Fatma Güney, Yizhou Wang 0001, Andreas Geiger 0001
ICCV3
2015 Quantized Correlation Hashing for Fast Cross-Modal Search
Botong Wu, Qiang Yang 0010, Wei-Shi Zheng 0001, Yizhou Wang 0001, Jingdong Wang 0001
IJCAI4
2015 VeTrack: Real Time Vehicle Tracking in Uninstrumented Indoor Environments
abstract
Although location awareness and turn-by-turn instructions are prevalent outdoors due to GPS, we are back into the darkness in uninstrumented indoor environments such as underground parking structures. We get confused, disoriented when driving in these mazes, and frequently forget where we parked, ending up circling back and forth upon return.In this paper, we propose VeTrack, a smartphone-only system that tracks the vehicle's location in real time using the phone's inertial sensors. It does not require any environment instrumentation or cloud backend. It uses a novel "shadow" tracing method to accurately estimate the vehicle's trajectories despite arbitrary phone/vehicle poses and frequent disturbances. We develop algorithms in a Sequential Monte Carlo framework to represent vehicle states probabilistically, and harness constraints by the garage map and detected landmarks to robustly infer the vehicle location. We also find landmark (e.g., speed bumps, turns) recognition methods reliable against noises, disturbances from bumpy rides and even hand-held movements. We implement a highly efficient prototype and conduct extensive experiments in multiple parking structures of different sizes and structures, with multiple vehicles and drivers. We find that VeTrack can estimate the vehicle's real time location with almost negligible latency, with error of 2-4 parking spaces at 80-percentile.
Mingmin Zhao, Ruipeng Gao, Fan Ye 0003, Yizhou Wang 0001, Guojie Luo
SenSys5
2015 Learning Hierarchical Space Tiling for Scene Modeling, Parsing and Attribute Tagging
abstract
A typical scene category contains an enormous number of distinct scene configurations that are composed of objects and regions of varying shapes in different layouts. In this paper, we first propose a representation named hierarchical space tiling (HST) to quantize the huge and continuous scene configuration space. Then, we augment the HST with attributes (nouns and adjectives) to describe the semantics of the objects and regions inside a scene. We present a weakly supervised method for simultaneously learning the scene configurations and attributes from a collection of natural images associated with descriptive text. The precise locations of attributes are unknown in the input and are mapped to the HST nodes through learning. Starting with a full HST, we iteratively estimate the HST model under a learning-by-parsing framework. Given a test image, we compute the most probable parse tree with the associated attributes by dynamic programming. We quantitatively analyze the representative efficiency of HST, show the learned representation is less ambiguous and has semantically meaningful inner concepts. In applications, we apply our model to four tasks: scene classification, attribute recognition, attribute localization, and pixel-wise scene labeling, and show the performance improvements as well as higher efficiency.
Yizhou Wang 0001, Song-Chun Zhu
IEEE Trans. Pattern Anal. Mach. Intell.2
2015 Weakly Supervised Semantic Segmentation with a Multiscale Model
abstract
This letter addresses the problem of weakly supervised semantic segmentation. Given training images with only image level annotations (i.e., tags) where the precise locations of tags are unknown, we simultaneously segment the images and assign tags to image regions. In contrast to previous work which segmented images at a specified scale, in this letter we propose a multiscale model for semantically segmenting images in different granularities and exploiting the long-range contextual information between adjacent scales. Then, to capture the geometric context of semantic labels, we augment the multiscale model by (i) the object spatial prior, e.g., “sky” has high probability on the top of an image, and (ii) the object spatial correlations, e.g., “car” always appears above “road”. Finally, we present an iterative top-down bottom-up method to learn the multiscale model by recovering the pixel labels of training images. Experiments on the benchmark MSRC21 and LMO datasets demonstrate the improved performance of our method over previous weakly supervised methods and even over some fully supervised methods.
Yizhou Wang 0001
IEEE Signal Process. Lett.2
2014 Efficient Generalized Fused Lasso and its Application to the Diagnosis of Alzheimer's Disease
abstract
Generalized fused lasso (GFL) penalizes variables with L1 norms based both on the variables and their pairwise differences. GFL is useful when applied to data where prior information is expressed using a graph over the variables. However, the existing GFL algorithms incur high computational costs and they do not scale to high-dimensional problems. In this study, we propose a fast and scalable algorithm for GFL. Based on the fact that fusion penalty is the Lov'asz extension of a cut function, we show that the key building block of the optimization is equivalent to recursively solving parametric graph-cut problems. Thus, we use a parametric flow algorithm to solve GFL in an efficient manner. Runtime comparisons demonstrated a significant speed-up compared with the existing GFL algorithms. By exploiting the scalability of the proposed algorithm, we formulated the diagnosis of Alzheimer's disease as GFL. Our experimental evaluations demonstrated that the diagnosis performance was promising and that the selected critical voxels were well structured i.e., connected, consistent according to cross-validation and in agreement with prior clinical knowledge.
Bo Xin, Yoshinobu Kawahara, Yizhou Wang 0001, Wen Gao 0001
AAAI3
2014 Robust Estimation of 3D Human Poses from a Single Image
abstract
Human pose estimation is a key step to action recognition. We propose a method of estimating 3D human poses from a single image, which works in conjunction with an existing 2D pose/joint detector. 3D pose estimation is challenging because multiple 3D poses may correspond to the same 2D pose after projection due to the lack of depth information. Moreover, current 2D pose estimators are usually inaccurate which may cause errors in the 3D estimation. We address the challenges in three ways: (i) We represent a 3D pose as a linear combination of a sparse set of bases learned from 3D human skeletons. (ii) We enforce limb length constraints to eliminate anthropomorphically implausible skeletons. (iii) We estimate a 3D pose by minimizing the 1-norm error between the projection of the 3D pose and the corresponding 2D detection. The 1-norm loss term is robust to inaccurate 2D joint estimations. We use the alternating direction method (ADM) to solve the optimization problem efficiently. Our approach outperforms the state-of-the-arts on three benchmark datasets.
Chunyu Wang 0001, Yizhou Wang 0001, Zhouchen Lin, Alan L. Yuille, Wen Gao 0001
CVPR2
2014 First-person multiple object tracking in complex traffic scenes
abstract
In this paper, we study multi-object tracking problem from the first-person viewpoint, e.g., the moving camera. This problem is different from the traditional one with static camera and brings lots of challenges. To solve this problem, we adopt the tracking-by-detection approach and design a new similarity model for two detection responses considering the camera motion. The similarity model can handle the change of scale and position of objects under the movement of camera. We also consider the detection prior and appearance to improve the tracking performance. The final tracking problem is solved within a network flow framework. Experimental results on KITTI dataset demonstrate the advantages of our method.
Tingting Jiang 0001, Yuansheng Xu, Yichong Bai, Yizhou Wang 0001
ICIP5
2014 Towards ubiquitous indoor localization service leveraging environmental physical features
abstract
Mainstream indoor localization technologies rely on RF signatures that require extensive human efforts to measure and periodically re-calibrate. Although recent crowdsourcing based work has started to address the issue, incentives are still lacking for wide user adoption. Thus the progress to ubiquitous localization remains slow. In this paper, we explore an alternative approach that leverages environmental physical features such as store logos or wall posters. A user uses a smartphone to obtain relative position measurements to such static reference points for the system to triangulate the user location. We study the principle of such localization, determine the suitable sensor, and devise guidelines for the user to choose reference points for better accuracy. To enable fast deployment, we propose a lightweight site survey method for service providers to quickly estimate the coordinates of reference points. We incorporate and enhance image matching algorithms with a heuristic technique to automatically identify chosen reference points at high accuracy. Extensive experiments have shown that the prototype achieves 4-5m accuracy at 80-percentile, comparable to the industry state-of-the-art, while covering a 150×75m mall and 300×200m train station requires a one time investment of only 2-3 man-hours from service providers.
Ruipeng Gao, Kaigui Bian, Fan Ye 0003, Tao Wang 0004, Yizhou Wang 0001, Xiaoming Li 0001
INFOCOM6
2014 Jigsaw: indoor floor plan reconstruction via mobile crowdsensing
abstract
The lack of floor plans is a critical reason behind the current sporadic availability of indoor localization service. Service providers have to go through effort-intensive and time-consuming business negotiations with building operators, or hire dedicated personnel to gather such data. In this paper, we propose Jigsaw, a floor plan reconstruction system that leverages crowdsensed data from mobile users. It extracts the position, size and orientation information of individual landmark objects from images taken by users. It also obtains the spatial relation between adjacent landmark objects from inertial sensor data, then computes the coordinates and orientations of these objects on an initial floor plan. By combining user mobility traces and locations where images are taken, it produces complete floor plans with hallway connectivity, room sizes and shapes. Our experiments on 3 stories of 2 large shopping malls show that the 90-percentile errors of positions and orientations of landmark objects are about 1~2m and 5~9°, while the hallway connectivity is 100% correct.
Ruipeng Gao, Mingmin Zhao, Fan Ye 0003, Yizhou Wang 0001, Kaigui Bian, Tao Wang 0004, Xiaoming Li 0001
MobiCom5
2014 VeLoc: finding your car in the parking lot
abstract
We present VeLoc, a smartphone-based vehicle localization approach that tracks the vehicle's parking location without GPS or WiFi signals. It uses only the embedded accelerometer and gyroscope sensors. VeLoc harnesses constraints imposed by the map and landmarks (e.g., speed bumps) recognized from inertial data, employs a Bayesian filtering framework to estimate the location of the vehicle. We have conducted experiments in three parking structures of different sizes and configurations, using three vehicles and three kinds of driving styles. We find that VeLoc can always localize the vehicle within 10m, which is sufficient for the driver to trigger a honk using the car key.
Mingmin Zhao, Ruipeng Gao, Jiaxu Zhu, Fan Ye 0003, Yizhou Wang 0001, Kaigui Bian, Guojie Luo, Ming Zhang 0004
SenSys6
2014 A Shape Reconstructability Measure of Object Part Importance with Applications to Object Detection and Localization
Ge Guo 0002, Yizhou Wang 0001, Tingting Jiang 0001, Alan L. Yuille, Fang Fang 0003, Wen Gao 0001
Int. J. Comput. Vis.2
2014 A Compact Representation for Compressing Converted Stereo Videos
abstract
We propose a novel representation for stereo videos namely 2D-plus-depth-cue. This representation is able to encode stereo videos compactly by leveraging the by-product of a stereo video conversion process. Specifically, the depth cues are derived from an interactive labeling process during 2D-to-stereo video conversion—they are contour points of image regions and their corresponding depth models, and so forth. Using such cues and the image features of 2D video frames, the scene depth can be reliably recovered. Experimental results demonstrate that the bit rate can be saved about 10%–50% in coding a stereo video compared with multiview video coding and the 2D-plus-depth methods. In addition, since the objects are segmented in the conversion process, it is convenient to adopt the region-of-interest (ROI) coding in the proposed stereo video coding system. Experimental results show that using ROI coding, the bit rate is reduced by 30%–40% or the video quality is increased by 1.5–4 dB with the fixed bit rate.
Zhebin Zhang, Ronggang Wang, Yizhou Wang 0001, Wen Gao 0001
IEEE Trans. Image Process.4
2013 Weakly Supervised Learning for Attribute Localization in Outdoor Scenes
abstract
In this paper, we propose a weakly supervised method for simultaneously learning scene parts and attributes from a collection of images associated with attributes in text, where the precise localization of the each attribute left unknown. Our method includes three aspects. (i) Compositional scene configuration. We learn the spatial layouts of the scene by Hierarchical Space Tiling (HST) representation, which can generate an excessive number of scene configurations through the hierarchical composition of a relatively small number of parts. (ii) Attribute association. The scene attributes contain nouns and adjectives corresponding to the objects and their appearance descriptions respectively. We assign the nouns to the nodes (parts) in HST using nonmaximum suppression of their correlation, then train an appearance model for each noun+adjective attribute pair. (iii) Joint inference and learning. For an image, we compute the most probable parse tree with the attributes as an instantiation of the HST by dynamic programming. Then update the HST and attribute association based on the inferred parse trees. We evaluate the proposed method by (i) showing the improvement of attribute recognition accuracy, and (ii) comparing the average precision of localizing attributes to the scene parts.
Jungseock Joo, Yizhou Wang 0001, Song-Chun Zhu
CVPR3
2013 An Approach to Pose-Based Action Recognition
abstract
We address action recognition in videos by modeling the spatial-temporal structures of human poses. We start by improving a state of the art method for estimating human joint locations from videos. More precisely, we obtain the K-best estimations output by the existing method and incorporate additional segmentation cues and temporal constraints to select the ``best'' one. Then we group the estimated joints into five body parts (e.g. the left arm) and apply data mining techniques to obtain a representation for the spatial-temporal structures of human actions. This representation captures the spatial configurations of body parts in one frame (by spatial-part-sets) as well as the body part movements(by temporal-part-sets) which are characteristic of human actions. It is interpretable, compact, and also robust to errors on joint estimations. Experimental results first show that our approach is able to localize body joints more accurately than existing methods. Next we show that it outperforms state of the art action recognizers on the UCF sport, the Keck Gesture and the MSR-Action3D datasets.
Chunyu Wang 0001, Yizhou Wang 0001, Alan L. Yuille
CVPR2
2013 A Method of Perceptual-Based Shape Decomposition
abstract
In this paper, we propose a novel perception-based shape decomposition method which aims to decompose a shape into semantically meaningful parts. In addition to three popular perception rules (the Minima rule, the Short-cut rule and the Convexity rule) in shape decomposition, we propose a new rule named part-similarity rule to encourage consistent partition of similar parts. The problem is formulated as a quadratic ally constrained quadratic program (QCQP) problem and is solved by a trust-region method. Experiment results on MPEG-7 dataset show that we can get a more consistent shape decomposition with human perception compared with other state-of-the-art methods both qualitatively and quantitatively. Finally, we show the advantage of semantic parts over non-meaningful parts in object detection on the ETHZ dataset.
Zhongqian Dong, Tingting Jiang 0001, Yizhou Wang 0001, Wen Gao 0001
ICCV4
2013 Learning discriminative features for fast frame-based action recognition
Liang Wang 0045, Yizhou Wang 0001, Tingting Jiang 0001, Debin Zhao, Wen Gao 0001
Pattern Recognit.2
2013 Video Stylization: Painterly Rendering and Optimization With Content Extraction
abstract
We present an interactive video stylization system for transforming an input video into a painterly animation. The system consists of two phases: a content extraction phase to obtain semantic objects, i.e., recognized content, in a video and establish dense feature correspondences, and a painterly rendering phase to select, place, and propagate brush strokes for stylized animations based on the semantic content and object motions derived from the first phase. Compared with the previous work, the proposed method has the following three advantages. First, we propose a two-pass rendering strategy and brush strokes with mixed colors in order to render expressive visual effects. Second, the brush strokes are warped according to global object deformations, so that the strokes appear to be naturally attached to the object surfaces. Third, we propose a deferred rendering and backward completion method to draw brush strokes on emerging regions and simulate a damped system to reduce stroke scintillation effect. Moreover, we discuss the graphics processing unit-based implementation of our system, which is demonstrated to greatly improve the efficiency of producing stylized videos. In experiments, we verify this system by applying it to a number of video clips to produce expressive oil-painting animations and compare it with the state-of-the-art approaches.
Liang Lin 0004, Yizhou Wang 0001, Ying-Qing Xu, Song-Chun Zhu
IEEE Trans. Circuits Syst. Video Technol.3
2013 Interactive Stereoscopic Video Conversion
abstract
This paper presents a system of converting conventional monocular videos to stereoscopic ones. In the system, an input monocular video is firstly segmented into shots so as to reduce operations on similar frames. An automatic depth estimation method is proposed to compute the depth maps of the video frames utilizing three monocular depth cues - depth-from-defocus, aerial perspective, and motion. Foreground/background objects can be interactively segmented on selected key frames and their depth values can be adjusted by users. Such results are propagated from key frames to nonkey frames within each video shot. Equipped with a depth-to-disparity conversion module, the system synthesizes the counterpart (either left or right) view for stereoscopic display by warping the original frames according to their disparity maps. The quality of converted videos is evaluated by human mean opinion scores, and experiment results demonstrate that the proposed conversion method achieves encouraging performance.
Zhebin Zhang, Yizhou Wang 0001, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2012 Toward Perception-Based Shape Decomposition
Tingting Jiang 0001, Zhongqian Dong, Yizhou Wang 0001
ACCV (2)4
2012 Hierarchical Space Tiling for Scene Modeling
Yizhou Wang 0001, Song-Chun Zhu
ACCV (2)2
2012 A Compact Stereoscopic Video Representation for 3D Video Generation and Coding
abstract
We propose a novel compact representation for stereoscopic videos - a 2D video and its depth cues. Depth cues are derived from an interactive labeling process during 2D-to-3D video conversion, they are contour points of foreground objects and a background geometric model. By using such cues and image features of 2D video frames, depth maps of the frames can be recovered. Compared with traditional 3D video representation, the proposed one is more compact. We also design algorithms to encode and decode the depth cues. The representation benefits both 3D video generation and coding. Experimental results demonstrate that the bit rate can be saved about 10%-50% in coding 3D videos compared with multi-view video coding and 2D+depth methods. A system coupling 2D-to-3D video conversion and coding (CVCC) is proposed to verify advantages of the representation.
Zhebin Zhang, Ronggang Wang, Yizhou Wang 0001, Wen Gao 0001
DCC4
2012 PQ-WGLOH: A bit-rate scalable local feature descriptor
abstract
In this paper, we propose a compact yet discriminative local descriptor which tackles the wireless query transmission latency in mobile visual search. The descriptor captures gradient statistics of canonical patches over a log-polar location grid whose parameters are optimized using training samples. We quantize the resulting descriptor using product quantization. The descriptor achieves about 95% bits reduction compared with 128-Byte SIFT and allows adaptation of descriptor lengths to support user required performance. Moreover, accurate matching of descriptors with low complexity is allowed within several table lookup operations. We perform a comprehensive comparison with SIFT, GLOH and CHoG in the context of image retrieval, image matching and object localization. We achieve competing matching and retrieval performance with SIFT, GLOH with much fewer bits. In particular, the descriptor outperforms CHoG at the same bits on eight data sets contributed to MPEG Compact Descriptor for Visual Search(CDVS) Standardization.
Chunyu Wang 0001, Ling-Yu Duan, Yizhou Wang 0001, Wen Gao 0001
ICASSP3
2012 Unsupervised discriminative feature selection in a kernel space via L2, 1-norm minimization
Yang Liu 0006, Yizhou Wang 0001
ICPR2
2012 An interactive system of stereoscopic video conversion
abstract
With the recent booming of 3DTV industry, more and more stereoscopic videos are demanded by the market. This paper presents a system of converting conventional monocular videos to stereoscopic ones. In this system, an input video is firstly segmented into shots to reduce operations on similar frames. Then, automatic depth estimation and interactive image segmentation are integrated to obtain depth maps and foreground/background segments on selected key frames. Within each video shot, such results are propagated from key frames to non-key frames. Combined with a depth-to-disparity conversion method, the system synthesizes the counterpart (either left or right) view for stereoscopic display by warping the original frame according to disparity maps. For evaluation, we use human labeled depth map as the reference and compute both the mean opinion score (MOS) and Peak signal-to-noise ratio (PSNR) to valuate the converted video quality. Experiment results demonstrate that the proposed conversion system and methods achieves encouraging performance.
Zhebin Zhang, Bo Xin, Yizhou Wang 0001, Wen Gao 0001
ACM Multimedia4
2012 Recovering Missing Contours for Occluded Object Detection
abstract
One difficult problem in practical applications is the corrupted or missing data frequently encountered in digital images. It introduces great challenges to the tasks such as object detection. This letter provides new methods for recovering missing object contours and detecting occluded objects. First, we propose an efficient contour reconstruction approach according to the Bayesian rule, utilizing global shape prior knowledge. Second, the contour reconstruction is applied to a robust detection framework for occluded objects. Based on the observed broken curves we iteratively recover object contours and propose object candidates. The experimental results demonstrate the high detection performance, localization accuracy and great advantages of our method for severe occlusion cases.
Ge Guo 0002, Tingting Jiang 0001, Yizhou Wang 0001, Wen Gao 0001
IEEE Signal Process. Lett.3
2011 Simulating human saccadic scanpaths on natural images
abstract
Human saccade is a dynamic process of information pursuit. Based on the principle of information maximization, we propose a computational model to simulate human saccadic scanpaths on natural images. The model integrates three related factors as driven forces to guide eye movements sequentially - reference sensory responses, fovea-periphery resolution discrepancy, and visual working memory. For each eye movement, we compute three multi-band filter response maps as a coherent representation for the three factors. The three filter response maps are combined into multi-band residual filter response maps, on which we compute residual perceptual information (RPI) at each location. The RPI map is a dynamic saliency map varying along with eye movements. The next fixation is selected as the location with the maximal RPI value. On a natural image dataset, we compare the saccadic scanpaths generated by the proposed model and several other visual saliency-based models against human eye movement data. Experimental results demonstrate that the proposed model achieves the best prediction accuracy on both static fixation locations and dynamic scanpaths.
Wei Wang 0115, Cheng Chen 0004, Yizhou Wang 0001, Tingting Jiang 0001, Fang Fang 0003, Yuan Yao 0011
CVPR3
2011 Instantly telling what happens in a video sequence using simple features
abstract
This paper presents an efficient method to tell what happens (e.g. recognize actions) in a video sequence from only a couple of frames in real time. For the sake of instantaneity, we employ two types of computationally efficient but perceptually important features, optical flow and edge, to capture motion and shape/structure information in video sequences. It is known that the two types of features are not sparse and can be unreliable or ambiguous at certain parts of a video. In order to endow them with strong discriminative power, we extend an efficient contrast set mining technique, the Emerging Pattern (EP) mining method, to learn joint features from videos to differentiate action classes. Experimental results show that the combination of the two types of features achieves superior performance in differentiating actions than that of using each single type of features alone. The learned features are discriminative, statistically significant (reliable) and display semantically meaningful shape-motion structures of human actions. Besides the instant action recognition, we also extend the proposed approach to anomaly detection and sequential event detection. The experiments demonstrate encouraging results.
Liang Wang 0045, Yizhou Wang 0001, Tingting Jiang 0001, Wen Gao 0001
CVPR2
2011 A Method of Evaluating Table Segmentation Results Based on a Table Image Ground Truther
abstract
We propose a novel method to evaluate table segmentation results based on a table image ground truther. In the ground-truthing process, we first extract connected components from a given table image and connect them into an atom graph with weighed edges. Edge weight takes neighboring connected components' size similarities and distances into consideration. Then the ground truther semi-automatically determines the locations and spans of row/column separators according to projection profiles, under human supervision. We evaluate a given table segmentation by computing edit distance from its row and column separator assertions relative to ground truth. The edit distance is the sum of all the edit operation costs that correct wrong row and column separators. Each edit operation cost is a function of the sum of the weights of the edges that the separator cuts through. Thus, separator errors incur different costs depending on the severity of the error, where severity roughly corresponds to how forgivable the error would be considered by a human observer. Experimental results demonstrate that the proposed evaluation method is not only efficient, but also useful in formalizing the intuitive quality of different segmentations.
Yanhui Liang, Yizhou Wang 0001, Eric Saund
ICDAR2
2011 Visual pertinent 2D-to-3D video conversion by multi-cue fusion
abstract
We describe an approach to2D-to-3D video conversion for the stereoscopic display. Targeting the problem of synthesizing the frames of a virtual 'right view' from the original monocular 2D video, we generate the stereoscopic video in steps as following. (1) A 2.5D depth map is first estimated in a multi-cue fusion manner by leveraging motion cues and photometric cues in video frames with a depth prior of spatial and temporal smoothness. (2) The depth map is converted to a disparity map with considering both the displaying device size and human's stereoscopic visual perception constraints. (3) We fix the original 2D frames as the 'left view' ones, and warp them to "virtually viewed" right ones according to the predicted disparity value. The main contribution of this method is to combine motion and photometric cues together to estimate depth map. In the experiments, we apply our method to converting several movie clips of well-known films into stereoscopic 3D video and get good results1.
Zhebin Zhang, Yizhou Wang 0001, Tingting Jiang 0001, Wen Gao 0001
ICIP2
2011 Stereoscopic learning for disparity estimation
abstract
In this paper, we propose a learning based approach to estimating pixel disparities from the motion information extracted out of input monoscopic video sequences. We represent each video frame with superpixels, and extract the motion features from the superpixels and the frame boundary. These motion features account for the motion pattern of the superpixel as well as camera motion. In the learning phase, given a pair of stereoscopic video sequences, we employ a state-of-the-art stereo matching method to compute the disparity map of each frame as ground truth. Then a multi-label SVM is trained from the estimated disparities and the corresponding motion features. In the testing phase, we use the learned SVM to predict the disparity for each superpixel in a monoscopic video sequence. Experiment results show that the proposed method achieves low error rate in disparity estimation.
Zhebin Zhang, Yizhou Wang 0001, Tingting Jiang 0001, Wen Gao 0001
ISCAS2
2011 Mining Layered Grammar Rules for Action Recognition
Liang Wang 0045, Yizhou Wang 0001, Wen Gao 0001
Int. J. Comput. Vis.2
2010 Measuring visual saliency by Site Entropy Rate
abstract
In this paper, we propose a new computational model for visual saliency derived from the information maximization principle. The model is inspired by a few well acknowledged biological facts. To compute the saliency spots of an image, the model first extracts a number of sub-band feature maps using learned sparse codes. It adopts a fully-connected graph representation for each feature map, and runs random walks on the graphs to simulate the signal/information transmission among the interconnected neurons. We propose a new visual saliency measure called Site Entropy Rate (SER) to compute the average information transmitted from a node (neuron) to all the others during the random walk on the graphs/network. This saliency definition also explains the center-surround mechanism from computation aspect. We further extend our model to spatial-temporal domain so as to detect salient spots in videos. To evaluate the proposed model, we do extensive experiments on psychological stimuli, two well known image data sets, as well as a public video dataset. The experiments demonstrate encouraging results that the proposed model achieves the state-of-the-art performance of saliency detection in both still images and videos.
Wei Wang 0115, Yizhou Wang 0001, Qingming Huang, Wen Gao 0001
CVPR2
2010 An interactive method for curve extraction
abstract
We introduce a curve process framework to solve the challenging problem of curve extraction from “non-traceable” curve groups. We propose a comprehensive curve model, which consists of the geometric, photometric and topological sub-models. Two typical categories of the non-traceable curve groups are considered. First, for the interlaced curves with complex structures, we show how to use the proposed curve model especially the topological sub-model to extract curves from the group. Second, for the non-interlaced but over-dense or faint curves we leverage the curve group pattern priors in addition, and extract the whole pattern in a global optimization. Applications and experiments demonstrate the competence of our models and methods.
Ge Guo 0002, Luoqi Liu, Zhebin Zhang, Yizhou Wang 0001, Wen Gao 0001
ICIP4
2010 Building Emerging Pattern (EP) Random forest for recognition
abstract
The Random forest classifier comes to be the working horse for visual recognition community. It predicts the class label of an input data by aggregating the votes of multiple tree classifiers. However, the classification performances of these tree classifiers are different. The random forest classifier ignores the difference by simply assigning them equal weights in voting for the final classification decision. Also, the random forest classifier only casts votes from individual tree classifiers without considering their compositions which would be more accurate. In this paper, we propose to tackle the two points by discovering weighted decision rules from the tree classifiers' output sets on training data. By treating the outputs of the tree classifiers on each data as a digital itemset, we want to find discriminative patterns (either containing the output of a single tree classifier or a set of tree classifiers) from the itemsets of training data. We employ an efficient data mining algorithm, the Emerging Pattern (EP) Mining, to search such discriminative patterns and weight them according to their discriminative powers. A set of decision rules are built from these discovered patterns and the final outputs of the Random Forest are made using these decision rules. We call the proposed classifier Emerging Pattern (EP) Random Forest. Experimental results on action categorization problems confirm that the proposed method really improve the performance of the traditional Random Forest classifier.
Liang Wang 0045, Yizhou Wang 0001, Debin Zhao
ICIP2
2010 Interactive viewpoint-space navigation for visual-audio exhibition of painting
abstract
In this paper, we present a system for exhibiting a Chinese landscape painting about 900 years old. There are three parts in our system: (1) we allocate a voice dubbing or background music, which is treated as a point sound source, onto the 2D painting and obtain its position in the 2D space. All of the audio data are then located in a 3D hidden space, by projecting their 2D positions to the 3D space through a projection model. (2) A two-layer directed graph structure is proposed to well organize the audio data in a 4D space (with 1D temporal and 3D spatial). (3) The exhibition is defined as an active exploration in a viewpoint space, which faces both the image and the 3D world where the sound sources reside. The 3D space and the two-layer graph structure generate a natural and meaningful stereo audio field. Meanwhile, compared to videos with guided walk through, the active exploration makes the exhibition more attractive.
Wei Ma 0008, Yang Liu 0006, Yizhou Wang 0001, Ying-Qing Xu, Hongbin Zha, Wen Gao 0001
ICME3
2010 Finding Multiple Object Instances with Occlusion
abstract
In this paper we provide a framework of detection and localization of multiple similar shapes or object instances from an image based on shape matching. There are three challenges about the problem. The first is the basic shape matching problem about how to find the correspondence and transformation between two shapes; second how to match shapes under occlusion; and last how to recognize and locate all the matched shapes in the image. We solve these problems by using both graph partition and shape matching in a global optimization framework. A Hough-like collaborative voting is adopted, which provides a good initialization, data-driven information, and plays an important role in solving the partial matching problem due to occlusion. Experiments demonstrate the efficiency of our method.
Ge Guo 0002, Tingting Jiang 0001, Yizhou Wang 0001, Wen Gao 0001
ICPR3
2008 Image objects and multi-scale features for annotation detection
abstract
This paper investigates several issues in the problem of detecting handwritten markings, or annotations, on printed documents. One issue is to define the appropriate units over which to perform feature measurements and assign type labels. We propose an alpha-shape tree that operates across multiple scales. A second issue is to devise image features that offer inferential power for machine learning algorithms. We report on a feature that measures edge turn statistics. A third issue is how to combine local and neighborhood evidence. We exploit the alpha shape tree in a direct inference architecture. Information propagation schemes such as Markov random fields may be readily layered on top of our output.
Jindong Chen, Eric Saund, Yizhou Wang 0001
ICPR3
2008 Human reappearance detection based on on-line learning
abstract
Many video surveillance applications require detecting human reappearances in a scene monitored by a camera or over a network of cameras. This is the human reappearance detection (HRD) problem. Studying this problem is important for analyzing a surveillance scenario at semantic level. In this paper, we propose a novel online learning framework for solving HRD problem. Both generative model and discriminative model are employed in this framework and a voting scheme is presented to fuse the decisions of both models for determining whether a just entered person is one of those who have shown up, i.e. whether a reappearance happens. Both models will be updated based on mistake-driven online learning strategy. Our experimental results show that the adopted online learning framework not only improves the reappearance detection accuracy but also achieves high robustness in various surveillance scenes.
Yizhou Wang 0001, Shuqiang Jiang, Qingming Huang, Wen Gao 0001
ICPR2
2008 Symmetric segment-based stereo matching of motion blurred images with illumination variations
abstract
Most existing methods of stereo matching focus on dealing with clear image pairs. Consequently, there is a lack of approaches capable of handling degraded images captured under challenging real situations, e.g. motion blur is present and an image pair is in different illumination conditions. In this paper we propose a novel approach to handling these challenging situations by formulating the problem into a Maximum a Posteriori (MAP) estimation framework, and adopt a segment-based symmetric stereo matching method to infer a mask of disparity map which indicates whether a disparity is affected by motion blur and estimate the disparity value. The experimental results show that our stereo matching method is able to compute more accurate disparity maps of this type of degraded images.
Wei Wang 0115, Yizhou Wang 0001, Longshe Huo, Qingming Huang, Wen Gao 0001
ICPR2
2008 Perceptual Scale-Space and Its Applications
Yizhou Wang 0001, Song-Chun Zhu
Int. J. Comput. Vis.1
2007 Decompose Document Image Using Integer Linear Programming
abstract
Document decomposition is a basic but crucial step for many document related applications. This paper proposes a novel approach to decompose document images into zones. It first generates overlapping zone hypotheses based on generic visual features. Then, each candidate zone is evaluated quantitatively by a learned generative zone model. We formulate the zone inference problem into a constrained optimization problem, so as to select an optimal set of non-overlapping zones that cover a given document image. The experimental results demonstrate that the proposed method is very robust to document structure variation and noise.
Dashan Gao 0001, Yizhou Wang 0001, Haitham A. Hindi, Minh Do
ICDAR2
2005 Perceptual Scale Space and its Applications
abstract
In this paper, we study a perceptual scale space by constructing a so-called sketch pyramid which augments the Gaussian and Laplacian pyramid representations in traditional image scale space theory. Each level of this sketch pyramid is a generic attributed graph - called the primal sketch which is inferred from the corresponding image at the same level of the Gaussian pyramid. When images are viewed at increasing resolutions, more details are revealed. This corresponds to perceptual transitions which are represented by topological changes in the sketch graph in terms of a graph grammar. We compute the sketch or perceptual pyramid by Bayesian inference upwards-downwards the pyramid using Markov chain Monte Carlo reversible jumps. We show two example applications of this perceptual scale space: (1) motion tracking of objects over scales, and (2) adaptive image displays which can efficiently show a large high resolution image in a small screen (of a PDA for example) through a selective tour of its image pyramid. Other potential applications include super resolution and multiresolution object recognition.
Yizhou Wang 0001, Siavosh Bahrami, Song-Chun Zhu
ICCV1
2005 Improving Mining Quality by Exploiting Data Dependency
Fang Chu, Yizhou Wang 0001, Carlo Zaniolo, Douglas Stott Parker Jr.
PAKDD2
2005 What are Textons?
Song-Chun Zhu, Cheng-en Guo, Yizhou Wang 0001, Zijian Xu 0001
Int. J. Comput. Vis.3
2004 Modeling Complex Motion by Tracking and Editing Hidden Markov Graphs
Yizhou Wang 0001, Song-Chun Zhu
CVPR (1)1
2004 Mining Noisy Data Streams via a Discriminative Model
Fang Chu, Yizhou Wang 0001, Carlo Zaniolo
Discovery Science2
2004 An Adaptive Learning Approach for Noisy Data Streams
abstract
Two critical challenges typically associated with mining data streams are concept drift and data contamination. To address these challenges, we seek learning techniques and models that are robust to noise and can adapt to changes in timely fashion. We approach the stream-mining problem using a statistical estimation framework, and propose a fast and robust discriminative model for learning noisy data streams. We build an ensemble of classifiers to achieve timely adaptation by weighting classifiers in a way that maximizes the likelihood of the data. We further employ robust statistical techniques to alleviate the problem of noise sensitivity. Experimental results on both synthetic and real-life data sets demonstrate the effectiveness of this model learning approach.
Fang Chu, Yizhou Wang 0001, Carlo Zaniolo
ICDM2
2004 Analysis and Synthesis of Textured Motion: Particles and Waves
abstract
Natural scenes contain a wide range of textured motion phenomena which are characterized by the movement of a large amount of particle and wave elements, such as falling snow, wavy water, and dancing grass. In this paper, we present a generative model for representing these motion patterns and study a Markov chain Monte Carlo algorithm for inferring the generative representation from observed video sequences. Our generative model consists of three components. The first is a photometric model which represents an image as a linear superposition of image bases selected from a generic and overcomplete dictionary. The dictionary contains Gabor and LoG bases for point/particle elements and Fourier bases for wave elements. These bases compete to explain the input images and transfer them to a token (base) representation with an O(10(2))-fold dimension reduction. The second component is a geometric model which groups spatially adjacent tokens (bases) and their motion trajectories into a number of moving elements--called "motons." A moton is a deformable template in time-space representing a moving element, such as a falling snowflake or a flying bird. The third component is a dynamic model which characterizes the motion of particles, waves, and their interactions. For example, the motion of particle objects floating in a river, such as leaves and balls, should be coupled with the motion of waves. The trajectories of these moving elements are represented by coupled Markov chains. The dynamic model also includes probabilistic representations for the birth/death (source/sink) of the motons. We adopt a stochastic gradient algorithm for learning and inference. Given an input video sequence, the algorithm iterates two steps: 1) computing the motons and their trajectories by a number of reversible Markov chain jumps, and 2) learning the parameters that govern the geometric deformations and motion dynamics. Novel video sequences are synthesized from the learned models and, by editing the model parameters, we demonstrate the controllability of the generative model.
Yizhou Wang 0001, Song-Chun Zhu
IEEE Trans. Pattern Anal. Mach. Intell.1
2003 Modeling Textured Motion : Particle, Wave and Sketch
abstract
We present a generative model for textured motion phenomena, such as falling snow, wavy river and dancing grass, etc. Firstly, we represent an image as a linear superposition of image bases selected from a generic and over-complete dictionary. The dictionary contains Gabor bases for point/particle elements and Fourier bases for wave-elements. These bases compete to explain the input images. The transform from a raw image to a base or a token representation leads to large dimension reduction. Secondly, we introduce a unified motion equation to characterize the motion of these bases and the interactions between waves and particles, e.g. a ball floating on water. We use statistical learning algorithm to identify the structure of moving objects and their trajectories automatically. Then novel sequences can be synthesized easily from the motion and image models. Thirdly, we replace the dictionary of Gabor and Fourier bases with symbolic sketches (also bases). With the same image and motion model, we can render realistic and stylish cartoon animation. In our view, cartoon and sketch are symbolic visualization of the inner representation for visual perception. The success of the cartoon animation, in turn, suggests that our image and motion models capture the essence of visual perception of textured motion.
Yizhou Wang 0001, Song-Chun Zhu
ICCV1
2002 A Generative Method for Textured Motion: Analysis and Synthesis
Yizhou Wang 0001, Song-Chun Zhu
ECCV (1)1
2002 What Are Textons?
Song-Chun Zhu, Cheng-en Guo, Ying Nian Wu, Yizhou Wang 0001
ECCV (4)4