VLDB 2026 Research / reviewers in the wild / expert
Chunyu Wang 0001
dblp:63/7235-1
· DBLP profile ↗
49ranked-venue papers
7as first author
30since 2021 · last 2026
0000-0002-9400-9107ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 40 · 6 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 37 · 6 first-author · 20 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLM2CLIP: Powerful Language Model Unlocks Richer Cross-Modality RepresentationabstractCLIP is a seminal multimodal model that maps images and text into a shared representation space by contrastive learning on billions of image–caption pairs. Inspired by the rapid progress of large language models (LLMs), we investigate how the superior linguistic understanding and broad world knowledge of LLMs can further strengthen CLIP—particularly in handling long, complex captions. We introduce an efficient fine-tuning framework that embeds an LLM into a pretrained CLIP while incurring almost the same training cost as regular CLIP fine-tuning. Our method first “embedding-izes” the LLM for the CLIP setting, then couples it to the pretrained CLIP vision encoder through a lightweight adaptor trained on only a few million image–caption pairs. With this strategy we achieve large performance gains—without large-scale retraining—over state-of-the-art CLIP variants such as EVA02 and SigLIP-2. The LLM-enhanced CLIP delivers consistent improvements across a wide spectrum of downstream tasks, including linear-probe classification, zero-shot image–text retrieval with both short and long captions (in English and other languages), zero-shot/supervised image segmentation, object detection, and used as tokenizer for multimodal large-model benchmarks. Weiquan Huang, Aoqi Wu, Yifan Yang 0004, Xufang Luo, Yuqing Yang 0001, Usman Naseem, Chunyu Wang 0001, Qi Dai 0001, Xiyang Dai, Dongdong Chen 0001, Chong Luo 0001, Lili Qiu, Liang Hu 0004 |
AAAI | 7 |
| 2025 | Shift Equivariant Pose NetworkabstractHuman pose estimation has been greatly advanced in recent years. However, even the best-performing models are not shift equivariant. In particular, a small change in input images often results in drastic alterations in output, which are problematic especially in video applications. The prevalence of top-down approaches, which typically rely on a (non-equivariant) object detector in the first stage, exac-erbates this issue. In this paper, we first demonstrate that the biased keypoint representation and the non-equivariant network components are the two main obstacles to shift equivariant pose estimation. To address the limitation, we propose an unbiased decoding method, and redesign the necessary network components (e.g., APS-ResBlock, SSP). Extensive experiments show that our method not only produces much more stable results with shifting input, but also achieves better metrics with the ability of tolerating in-accurate detector output from the first stage. To our knowledge, this is the first work to address the problem of shift equivariance in the field of pose estimation. Our method could be easily applied to existing CNN-based pose estimation networks. Pengxiao Wang, Tzu-Heng Lin, Chunyu Wang 0001, Yizhou Wang 0001 |
WACV | 3 |
| 2025 | VolumeDiffusion: Feed-forward text-to-3D generation with efficient volumetric encoderabstractThis work presents VolumeDiffusion, a novel feed-forward text-to-3D generation framework that directly synthesizes 3D objects from textual descriptions. It bypasses the conventional score distillation loss based or text-to-image-to-3D approaches. To scale up the training data for the diffusion model, a novel 3D volumetric encoder is developed to efficiently acquire feature volumes from multi-view images. The 3D volumes are then trained on a diffusion model for text-to-3D generation using a 3D U-Net. This research further addresses the challenges of inaccurate object captions and high-dimensional feature volumes. The proposed model, trained on the public Objaverse dataset, demonstrates promising outcomes in producing diverse and recognizable samples from text prompts. Notably, it empowers finer control over object part characteristics through textual cues, fostering model creativity by seamlessly combining multiple concepts within a single object. This research significantly contributes to the progress of 3D generation by introducing an efficient, flexible, and scalable representation methodology. Zhicong Tang, Shuyang Gu, Chunyu Wang 0001, Ting Zhang 0002, Jianmin Bao, Dong Chen 0003, Baining Guo |
Graph. Model. | 3 |
| 2025 | VMarker-Pro: Probabilistic 3D Human Mesh Estimation From Virtual MarkersabstractMonocular 3D human mesh estimation faces challenges due to depth ambiguity and the complexity of mapping images to complex parameter spaces. Recent methods propose to use 3D poses as a proxy representation, which often lose crucial body shape information, leading to mediocre performance. Conversely, advanced motion capture systems, though accurate, are impractical for markerless wild images. Addressing these limitations, we introduce an innovative intermediate representation as virtual markers, which are learned from large-scale mocap data, mimicking the effects of physical markers. Building upon virtual markers, we propose VMarker, which detects virtual markers from wild images, and the intact mesh with realistic shapes can be obtained by simply interpolation from these markers. To address occlusions that obscure 3D virtual marker estimation, we further enhance our method with VMarker-Pro, a probabilistic framework that models the distribution of 3D virtual marker positions using diffusion models, enabling the generation of multiple plausible meshes aligned with images for robust 3D mesh estimation. Our approaches surpass existing methods on three benchmark datasets, particularly demonstrating significant improvements on the SURREAL dataset, which features diverse body shapes. Additionally, VMarker-Pro excels in accurately modeling data distributions, significantly enhancing performance in occluded scenarios. Xiaoxuan Ma 0001, Jiajun Su, Yuan Xu 0022, Wentao Zhu 0004, Chunyu Wang 0001, Yizhou Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Multiple View Geometry Transformers for 3D Human Pose EstimationabstractIn this work, we aim to improve the 3D reasoning ability of Transformers in multi-view 3D human pose estimation. Recent works have focused on end-to-end learning-based transformer designs, which struggle to resolve geometric information accurately, particularly during occlusion. In-stead, we propose a novel hybrid model, MVGFormer, which has a series of geometric and appearance modules organized in an iterative manner. The geometry modules are learning-free and handle all viewpoint-dependent 3D tasks geometrically which notably improves the model's gener-alization ability. The appearance modules are learnable and are dedicated to estimating 2D poses from image signals end-to-end which enables them to achieve accurate es-timates even when occlusion occurs, leading to a model that is both accurate and generalizable to new cameras and geometries. We evaluate our approach for both in-domain and out-of-domain settings, where our model consistently outperforms state-of-the-art methods, and especially does so by a significant margin in the out-of-domain setting. We will release the code and models: https://github.com/XunshanMan/MVGFormer. Ziwei Liao, Chunyu Wang 0001, Han Hu 0001, Steven Lake Waslander |
CVPR | 3 |
| 2024 | MicroCinema: A Divide-and-Conquer Approach for Text-to-Video GenerationabstractWe present MicroCinema, a straightforward yet effective framework for high-quality and coherent text-to-video generation. Unlike existing approaches that align text prompts with video directly, MicroCinema introduces a Divide-and-Conquer strategy which divides the text-to-video into a two-stage process: text-to-image generation and image&text-to-video generation. This strategy offers two significant advantages. a) It allows us to take full advantage of the recent advances in text-to-image models, such as Stable Diffusion, Midjourney, and DALLE, to generate photorealistic and highly detailed images. b) Leveraging the generated image, the model can allocate less focus to fine-grained appearance details, prioritizing the efficient learning of motion dynamics. To implement this strategy effectively, we introduce two core designs. First, we propose the Appearance Injection Network, enhancing the preservation of the appearance of the given image. Second, we introduce the appearance Noise Prior, a novel mechanism aimed at maintaining the capabilities of pre-trained 2D diffusion models. These design elements empower MicroCinema to generate high-quality videos with precise motion, guided by the provided text prompts. Extensive experiments demonstrate the superiority of the proposed framework. Concretely, MicroCinema achieves SOTA zero-shot FVD of 342.86 on UCF-JOJ and 377.40 on MSR-VTT. Jianmin Bao, Wenming Weng, Ruoyu Feng 0001, Dacheng Yin, Jingxu Zhang, Qi Dai 0001, Zhiyuan Zhao 0001, Chunyu Wang 0001, Yuhui Yuan, Xiaoyan Sun 0001, Chong Luo 0001, Baining Guo |
CVPR | 10 |
| 2024 | Plan, Posture and Go: Towards Open-Vocabulary Text-to-Motion Generation
Wenxun Dai, Chunyu Wang 0001, Yiji Cheng, Yansong Tang, Xin Tong 0001 |
ECCV (27) | 3 |
| 2024 | RodinHD: High-Fidelity 3D Avatar Generation with Diffusion Models
Bowen Zhang 0010, Yiji Cheng, Chunyu Wang 0001, Ting Zhang 0002, Jiaolong Yang, Yansong Tang, Feng Zhao 0004, Dong Chen 0003, Baining Guo |
ECCV (14) | 3 |
| 2024 | V-DETR: DETR with Vertex Relative Position Encoding for 3D Object DetectionabstractWe introduce a highly performant 3D object detector for point clouds using the DETR framework. The prior attempts all end up with suboptimal results because they fail to learn accurate inductive biases from the limited scale of training data. In particular, the queries often attend to points that are far away from the target objects, violating the locality principle in object detection. To address the limitation, we introduce a novel 3D Vertex Relative Position Encoding (3DV-RPE) method which computes position encoding for each point based on its relative position to the 3D boxes predicted by the queries in each decoder layer, thus providing clear information to guide the model to focus on points near the objects, in accordance with the principle of locality. Furthermore, we have systematically refined our pipeline, including data normalization, to better align with the task requirements. Our approach demonstrates remarkable performance on the demanding ScanNetV2 benchmark, showcasing substantial enhancements over the prior state-of-the-art CAGroup3D. Specifically, we achieve an increase in $AP_{25}$ from $75.1\%$ to $77.8\%$ and in ${AP}_{50}$ from $61.3\%$ to $66.0\%$. Yichao Shen 0001, Zigang Geng, Yuhui Yuan, Yutong Lin, Chunyu Wang 0001, Han Hu 0001, Nanning Zheng 0001, Baining Guo |
ICLR | 6 |
| 2024 | GAIA: Zero-shot Talking Avatar GenerationabstractZero-shot talking avatar generation aims at synthesizing natural talking videos from speech and a single portrait image. Previous methods have relied on domain-specific heuristics such as warping-based motion representation and 3D Morphable Models, which limit the naturalness and diversity of the generated avatars. In this work, we introduce GAIA (Generative AI for Avatar), which eliminates the domain priors in talking avatar generation. In light of the observation that the speech only drives the motion of the avatar while the appearance of the avatar and the background typically remain the same throughout the entire video, we divide our approach into two stages: 1) disentangling each frame into motion and appearance representations; 2) generating motion sequences conditioned on the speech and reference portrait image. We collect a large-scale high-quality talking avatar dataset and train the model on it with different scales (up to 2B parameters). Experimental results verify the superiority, scalability, and flexibility of GAIA as 1) the resulting model beats previous baseline models in terms of naturalness, diversity, lip-sync quality, and visual quality; 2) the framework is scalable since larger models yield better results; 3) it is general and enables different applications like controllable talking avatar generation and text-instructed avatar generation. Tianyu He, Junliang Guo, Runyi Yu 0002, Yuchi Wang, Kaikai An, Leyi Li, Xu Tan 0003, Chunyu Wang 0001, Han Hu 0001, HsiangTao Wu, Sheng Zhao 0002, Jiang Bian 0002 |
ICLR | 9 |
| 2024 | GaussianCube: A Structured and Explicit Radiance Representation for 3D Generative ModelingabstractWe introduce a radiance representation that is both structured and fully explicit and thus greatly facilitates 3D generative modeling. Existing radiance representations either require an implicit feature decoder, which significantly degrades the modeling power of the representation, or are spatially unstructured, making them difficult to integrate with mainstream 3D diffusion methods. We derive GaussianCube by first using a novel densification-constrained Gaussian fitting algorithm, which yields high-accuracy fitting using a fixed number of free Gaussians, and then rearranging these Gaussians into a predefined voxel grid via Optimal Transport. Since GaussianCube is a structured grid representation, it allows us to use standard 3D U-Net as our backbone in diffusion modeling without elaborate designs. More importantly, the high-accuracy fitting of the Gaussians allows us to achieve a high-quality representation with orders of magnitude fewer parameters than previous structured representations for comparable quality, ranging from one to two orders of magnitude. The compactness of GaussianCube greatly eases the difficulty of 3D generative modeling. Extensive experiments conducted on unconditional and class-conditioned object generation, digital avatar creation, and text-to-3D synthesis all show that our model achieves state-of-the-art generation results both qualitatively and quantitatively, underscoring the potential of GaussianCube as a highly accurate and versatile radiance representation for 3D generative modeling. Bowen Zhang 0010, Yiji Cheng, Jiaolong Yang, Chunyu Wang 0001, Feng Zhao 0004, Yansong Tang, Dong Chen 0003, Baining Guo |
NeurIPS | 4 |
| 2024 | Unsupervised Graphic Layout Grouping with TransformersabstractGraphic design conveys messages through the combination of text, images and other visual elements. Unstructured designs such as overloaded social media graphics may fail to communicate their intended messages effectively. To address this issue, layout grouping offers a solution by organizing design elements into perceptual groups. While most methods rely on heuristic Gestalt principles, they often lack the context modeling ability needed to handle complex layouts. In this work, we reformulate the layout grouping task as a set prediction problem. It uses Transformers to learn a set of group tokens at various hierarchies, enabling it to reason the membership of the elements more effectively. The self-attention mechanism in Transformers boosts its context modeling ability, which enables it to handle complex layouts more accurately. To reduce annotation costs, we also propose an unsupervised learning strategy that pre-trains on noisy pseudo-labels induced by a novel heuristic algorithm. This approach then bootstraps to self-refine the noisy labels, further improving the accuracy of our model. Our extensive experiments demonstrate the effectiveness of our method, which outperforms existing state-of-the-art approaches in terms of accuracy and efficiency. Danqing Huang, Chunyu Wang 0001, Mingxi Cheng, Ji Li 0006, Han Hu 0001, Xin Geng 0001, Baining Guo |
WACV | 3 |
| 2024 | Correlation-Embedded Transformer Tracking: A Single-Branch FrameworkabstractDeveloping robust and discriminative appearance models has been a long-standing research challenge in visual object tracking. In the prevalent Siamese-based paradigm, the features extracted by the Siamese-like networks are often insufficient to model the tracked targets and distractor objects, thereby hindering them from being robust and discriminative simultaneously. While most Siamese trackers focus on designing robust correlation operations, we propose a novel single-branch tracking framework inspired by the transformer. Unlike the Siamese-like feature extraction, our tracker deeply embeds cross-image feature correlation in multiple layers of the feature network. By extensively matching the features of the two images through multiple layers, it can suppress non-target features, resulting in target-aware feature extraction. The output features can be directly used to predict target locations without additional correlation steps. Thus, we reformulate the two-branch Siamese tracking as a conceptually simple, fully transformer-based Single-Branch Tracking pipeline, dubbed SBT. After conducting an in-depth analysis of the SBT baseline, we summarize many effective design principles and propose an improved tracker dubbed SuperSBT. SuperSBT adopts a hierarchical architecture with a local modeling layer to enhance shallow-level features. A unified relation modeling is proposed to remove complex handcrafted layer pattern designs. SuperSBT is further improved by masked image modeling pre-training, integrating temporal modeling, and equipping with dedicated prediction heads. Thus, SuperSBT outperforms the SBT baseline by 4.7%,3.0%, and 4.5% AUC scores in LaSOT, TrackingNet, and GOT-10K. Notably, SuperSBT greatly raises the speed of SBT from 37 FPS to 81 FPS. Extensive experiments show that our method achieves superior results on eight VOT benchmarks. Wankou Yang, Chunyu Wang 0001, Yue Cao 0001, Chao Ma 0004, Wenjun Zeng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Human Pose as Compositional TokensabstractHuman pose is typically represented by a coordinate vector of body joints or their heatmap embeddings. While easy for data processing, unrealistic pose estimates are admitted due to the lack of dependency modeling between the body joints. In this paper, we present a structured representation, named Pose as Compositional Tokens (PCT), to explore the joint dependency. It represents a pose by M discrete tokens with each characterizing a sub-structure with several interdependent joints (see Figure 1). The compositional design enables it to achieve a small reconstruction error at a low cost. Then we cast pose estimation as a classification task. In particular, we learn a classifier to predict the categories of the M tokens from an image. A pre-learned decoder network is used to recover the pose from the tokens without further post-processing. We show that it achieves better or comparable pose estimation results as the existing methods in general scenarios, yet continues to work well when occlusion occurs, which is ubiquitous in practice. The code and models are publicly available at https://github.com/Gengzigang/PCT. Zigang Geng, Chunyu Wang 0001, Yixuan Wei, Houqiang Li, Han Hu 0001 |
CVPR | 2 |
| 2023 | 3D Human Mesh Estimation from Virtual MarkersabstractInspired by the success of volumetric 3D pose estimation, some recent human mesh estimators propose to estimate 3D skeletons as intermediate representations, from which, the dense 3D meshes are regressed by exploiting the mesh topology. However, body shape information is lost in extracting skeletons, leading to mediocre performance. The advanced motion capture systems solve the problem by placing dense physical markers on the body surface, which allows to extract realistic meshes from their non-rigid motions. However, they cannot be applied to wild images without markers. In this work, we present an intermediate representation, named virtual markers, which learns 64 landmark keypoints on the body surface based on the large-scale mocap data in a generative style, mimicking the effects of physical markers. The virtual markers can be accurately detected from wild images and can reconstruct the intact meshes with realistic shapes by simple interpolation. Our approach outperforms the state-of-the-art methods on three datasets. In particular, it surpasses the existing methods by a notable margin on the SURREAL dataset, which has diverse body shapes. Code is available at https://github.com/ShirleyMaxx/VirtualMarker Xiaoxuan Ma 0001, Jiajun Su, Chunyu Wang 0001, Wentao Zhu 0004, Yizhou Wang 0001 |
CVPR | 3 |
| 2023 | All in Tokens: Unifying Output Space of Visual Tasks via Soft TokenabstractWe introduce AiT, a unified output representation for various vision tasks, which is a crucial step towards general-purpose vision task solvers. Despite the challenges posed by the high-dimensional and task-specific outputs, we showcase the potential of using discrete representation (VQVAE) to model the dense outputs of many computer vision tasks as a sequence of discrete tokens. This is inspired by the established ability of VQ-VAE to conserve the structures spanning multiple pixels using few discrete codes. To that end, we present a modified shallower architecture for VQ-VAE that improves efficiency while keeping prediction accuracy. Our approach also incorporates uncertainty into the decoding process by using a soft fusion of the codebook entries, providing a more stable training process, which notably improved prediction accuracy. Our evaluation of AiT on depth estimation and instance segmentation tasks, with both continuous and discrete labels, demonstrates its superiority compared to other unified models. The code and models are available at https://github.com/SwinTransformer/AiT. Zheng Zhang 0022, Chunyu Wang 0001, Zigang Geng, Qi Dai 0001, Kun He 0001, Han Hu 0001 |
ICCV | 4 |
| 2023 | Category-Specific Prompts for Animal Action Recognition with Pretrained Vision-Language ModelsabstractAnimal action recognition has a wide range of applications. However, the field largely remains unexplored due to the greater challenges compared to human action recognition, such as lack of annotated training data, large intra-class variation, and interference of cluttered background. Most of the existing methods directly apply human action recognition techniques, which essentially require a large amount of annotated data. In recent years, contrastive vision-language pretraining has demonstrated strong zero-shot generalization ability and has been used for human action recognition. Inspired by the success, we develop a highly performant action recognition framework based on the CLIP model. Our model addresses the above challenges via a novel category-specific prompting module to generate adaptive prompts for both text and video based on the animal category detected in input videos. On one hand, it can generate more precise and customized textual descriptions for each action and animal category pair, being helpful in the alignment of textual and visual space. On the other hand, it allows the model to focus on video features of the target animal in the video and reduce the interference of video background noise. Experimental results demonstrate that our method outperforms five previous action recognition methods on the Animal Kingdom dataset and has shown best generalization ability on unseen animals. Yinuo Jing, Chunyu Wang 0001, Ruxu Zhang, Kongming Liang, Zhanyu Ma |
ACM Multimedia | 2 |
| 2023 | MMPTRACK: Large-scale Densely Annotated Multi-camera Multiple People Tracking BenchmarkabstractMulti-camera tracking systems are gaining popularity in applications that demand high-quality tracking results, such as frictionless checkout. In cluttered and crowded environments, monocular multi-object tracking (MOT) systems often fail due to occlusions. Multiple highly overlapped cameras are capable of recovering partial 3D information. When used properly, 3D data can significantly alleviate the occlusion issue. However, training a multi-camera tracker demands a large-scale multi-camera tracking dataset with diverse camera settings and backgrounds. These requirements make the collection of multi-camera tracking dataset challenging and expensive. The cost of creating such a dataset has limited the availability and scale of datasets in this domain. Instead, we appeal to an auto-annotation system to reduce the cost, which uses overlapped and calibrated depth and RGB cameras to build a 3D tracker and automatically generates the 3D tracking results. The results are manually checked and corrected to ensure the label quality, which is much cheaper than solely manual annotation. Next, the 3D tracking results are projected to each calibrated RGB camera view to create 2D tracking results. In this way, we collect and annotate a large-scale densely labeled multi-camera tracking dataset from five different environments. We have conducted extensive experiments using two real-time multi-camera trackers and a person re-identification (ReID) model under different settings. This dataset provides a reliable benchmark for multi-camera, multi-object tracking systems in cluttered and crowded environments. We expect this benchmark to encourage more research attempts in this domain. Our dataset will be publicly released upon the acceptance of this work. Quanzeng You, Chunyu Wang 0001, Zhizheng Zhang 0004, Peng Chu, Houdong Hu, Jiang Wang 0012, Zicheng Liu 0001 |
WACV | 3 |
| 2023 | VoxelTrack: Multi-Person 3D Human Pose Estimation and Tracking in the WildabstractWe present VoxelTrack for multi-person 3D pose estimation and tracking from a few cameras which are separated by wide baselines. It employs a multi-branch network to jointly estimate 3D poses and re-identification (Re-ID) features for all people in the environment. In contrast to previous efforts which require to establish cross-view correspondence based on noisy 2D pose estimates, it directly estimates and tracks 3D poses from a 3D voxel-based representation constructed from multi-view images. We first discretize the 3D space by regular voxels and compute a feature vector for each voxel by averaging the body joint heatmaps that are inversely projected from all views. We estimate 3D poses from the voxel representation by predicting whether each voxel contains a particular body joint. Similarly, a Re-ID feature is computed for each voxel which is used to track the estimated 3D poses over time. The main advantage of the approach is that it avoids making any hard decisions based on individual images. The approach can robustly estimate and track 3D poses even when people are severely occluded in some cameras. It outperforms the state-of-the-art methods by a large margin on four public datasets including Shelf, Campus, Human3.6 M and CMU Panoptic. Chunyu Wang 0001, Xinggang Wang, Wenyu Liu 0001, Wenjun Zeng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Correlation-Aware Deep TrackingabstractRobustness and discrimination power are two fundamental requirements in visual object tracking. In most tracking paradigms, we find that the features extracted by the popular Siamese-like networks cannot fully discriminatively model the tracked targets and distractor objects, hindering them from simultaneously meeting these two requirements. While most methods focus on designing robust correlation operations, we propose a novel target-dependent feature network inspired by the self-/cross-attention scheme. In contrast to the Siamese-like feature extraction, our network deeply embeds cross-image feature correlation in multiple layers of the feature network. By extensively matching the features of the two images through multiple layers, it is able to suppress non-target features, resulting in instance-varying feature extraction. The output features of the search image can be directly used for predicting target locations without extra correlation step. Moreover, our model can be flexibly pre-trained on abundant unpaired images, leading to notably faster convergence than the existing methods. Extensive experiments show our method achieves the state-of-the-art results while running at real-time. Our feature networks also can be applied to existing tracking pipelines seamlessly to raise the tracking performance. Chunyu Wang 0001, Guangting Wang, Yue Cao 0001, Wankou Yang, Wenjun Zeng 0001 |
CVPR | 2 |
| 2022 | VirtualPose: Learning Generalizable 3D Human Pose Models from Virtual Data
Jiajun Su, Chunyu Wang 0001, Xiaoxuan Ma 0001, Wenjun Zeng 0001, Yizhou Wang 0001 |
ECCV (6) | 2 |
| 2022 | Faster VoxelPose: Real-time 3D Human Pose Estimation by Orthographic Projection
Hang Ye 0002, Wentao Zhu 0004, Chunyu Wang 0001, Rujie Wu, Yizhou Wang 0001 |
ECCV (6) | 3 |
| 2022 | One-Shot Medical Landmark Localization by Edge-Guided Transform and Noisy Landmark Refinement
Ping Gong 0002, Chunyu Wang 0001, Yizhou Yu, Yizhou Wang 0001 |
ECCV (21) | 3 |
| 2022 | Robust Multi-object Tracking by Marginal Inference
Chunyu Wang 0001, Xinggang Wang, Wenjun Zeng 0001, Wenyu Liu 0001 |
ECCV (22) | 2 |
| 2022 | Locally Connected Network for Monocular 3D Human Pose EstimationabstractWe present an approach for 3D human pose estimation from monocular images. The approach consists of two steps: it first estimates a 2D pose from an image and then estimates the corresponding 3D pose. This paper focuses on the second step. Graph convolutional network (GCN) has recently become the de facto standard for human pose related tasks such as action recognition. However, in this work, we show that GCN has critical limitations when it is used for 3D pose estimation due to the inherent weight sharing scheme. The limitations are clearly exposed through a novel reformulation of GCN, in which both GCN and Fully Connected Network (FCN) are its special cases. In addition, on top of the formulation, we present locally connected network (LCN) to overcome the limitations of GCN by allocating dedicated rather than shared filters for different joints. We jointly train the LCN network with a 2D pose estimator such that it can handle inaccurate 2D poses. We evaluate our approach on two benchmark datasets and observe that LCN outperforms GCN, FCN, and the state-of-the-art methods by a large margin. More importantly, it demonstrates strong cross-dataset generalization ability because of sparse connections among body joints. Hai Ci, Xiaoxuan Ma 0001, Chunyu Wang 0001, Yizhou Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Neighborhood Geometric Structure-Preserving Variational Autoencoder for Smooth and Bounded Data SourcesabstractMany data sources, such as human poses, lie on low-dimensional manifolds that are smooth and bounded. Learning low-dimensional representations for such data is an important problem. One typical solution is to utilize encoder-decoder networks. However, due to the lack of effective regularization in latent space, the learned representations usually do not preserve the essential data relations. For example, adjacent video frames in a sequence may be encoded into very different zones across the latent space with holes in between. This is problematic for many tasks such as denoising because slightly perturbed data have the risk of being encoded into very different latent variables, leaving output unpredictable. To resolve this problem, we first propose a neighborhood geometric structure-preserving variational autoencoder (SP-VAE), which not only maximizes the evidence lower bound but also encourages latent variables to preserve their structures as in ambient space. Then, we learn a set of small surfaces to approximately bound the learned manifold to deal with holes in latent space. We extensively validate the properties of our approach by reconstruction, denoising, and random image generation experiments on a number of data sources, including synthetic Swiss roll, human pose sequences, and facial expression images. The experimental results show that our approach learns more smooth manifolds than the baselines. We also apply our approach to the tasks of human pose refinement and facial expression image interpolation where it gets better results than the baselines. Xingyu Chen 0001, Chunyu Wang 0001, Xuguang Lan, Nanning Zheng 0001, Wenjun Zeng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | Context Modeling in 3D Human Pose Estimation: A Unified PerspectiveabstractEstimating 3D human pose from a single image suffers from severe ambiguity since multiple 3D joint configurations may have the same 2D projection. The state-of-the-art methods often rely on context modeling methods such as pictorial structure model (PSM) or graph neural network (GNN) to reduce ambiguity. However, there is no study that rigorously compares them side by side. So we first present a general formula for context modeling in which both PSM and GNN are its special cases. By comparing the two methods, we found that the end-to-end training scheme in GNN and the limb length constraints in PSM are two complementary factors to improve results. To combine their advantages, we propose ContextPose based on attention mechanism that allows enforcing soft limb length constraints in a deep network. The approach effectively reduces the chance of getting absurd 3D pose estimates with incorrect limb lengths and achieves state-of-the-art results on two benchmark datasets. More importantly, the introduction of limb length constraints into deep networks enables the approach to achieve much better generalization performance. Xiaoxuan Ma 0001, Jiajun Su, Chunyu Wang 0001, Hai Ci, Yizhou Wang 0001 |
CVPR | 3 |
| 2021 | An Empirical Study of the Collapsing Problem in Semi-Supervised 2D Human Pose EstimationabstractMost semi-supervised learning models are consistency-based, which leverage unlabeled images by maximizing the similarity between different augmentations of an image. But when we apply them to human pose estimation that has extremely imbalanced class distribution, they often collapse and predict every pixel in unlabeled images as background. We find this is because the decision boundary passes the high-density areas of the minor class so more and more pixels are gradually misclassified as background. In this work, we present a surprisingly simple approach to drive the model to learn in the correct direction. For each image, it composes a pair of easy-hard augmentations and uses the more accurate predictions on the easy image to teach the network to learn pose information of the hard one. The accuracy superiority of teaching signals allows the network to be "monotonically" improved which effectively avoids collapsing. We apply our method to the state-of-the-art pose estimators and it further improves their performance on three public datasets. Rongchang Xie, Chunyu Wang 0001, Wenjun Zeng 0001, Yizhou Wang 0001 |
ICCV | 2 |
| 2021 | AdaFuse: Adaptive Multiview Fusion for Accurate Human Pose Estimation in the Wild
Zhe Zhang 0045, Chunyu Wang 0001, Weichao Qiu, Wenhu Qin, Wenjun Zeng 0001 |
Int. J. Comput. Vis. | 2 |
| 2021 | FairMOT: On the Fairness of Detection and Re-identification in Multiple Object Tracking
Chunyu Wang 0001, Xinggang Wang, Wenjun Zeng 0001, Wenyu Liu 0001 |
Int. J. Comput. Vis. | 2 |
| 2020 | MetaFuse: A Pre-trained Fusion Model for Human Pose EstimationabstractCross view feature fusion is the key to address the occlusion problem in human pose estimation. The current fusion methods need to train a separate model for every pair of cameras making them difficult to scale. In this work, we introduce MetaFuse, a pre-trained fusion model learned from a large number of cameras in the Panoptic dataset. The model can be efficiently adapted or finetuned for a new pair of cameras using a small number of labeled images. The strong adaptation power of MetaFuse is due in large part to the proposed factorization of the original fusion model into two parts-(1) a generic fusion model shared by all cameras, and (2) lightweight camera-dependent transformations. Furthermore, the generic model is learned from many cameras by a meta-learning style algorithm to maximize its adaptation capability to various camera poses. We observe in experiments that MetaFuse finetuned on the public datasets outperforms the state-of-the-arts by a large margin which validates its value in practice. Rongchang Xie, Chunyu Wang 0001, Yizhou Wang 0001 |
CVPR | 2 |
| 2020 | Fusing Wearable IMUs With Multi-View Images for Human Pose Estimation: A Geometric ApproachabstractWe propose to estimate 3D human pose from multi-view images and a few IMUs attached at person's limbs. It operates by firstly detecting 2D poses from the two signals, and then lifting them to the 3D space. We present a geometric approach to reinforce the visual features of each pair of joints based on the IMUs. This notably improves 2D pose estimation accuracy especially when one joint is occluded. We call this approach Orientation Regularized Network (ORN). Then we lift the multi-view 2D poses to the 3D space by an Orientation Regularized Pictorial Structure Model (ORPSM) which jointly minimizes the projection error between the 3D and 2D poses, along with the discrepancy between the 3D pose and IMU orientations. The simple two-step approach reduces the error of the state-of-the-art by a large margin on a public dataset. Our code will be released at https://github.com/microsoft/imu-human-pose-estimation-pytorch. Zhe Zhang 0045, Chunyu Wang 0001, Wenhu Qin, Wenjun Zeng 0001 |
CVPR | 2 |
| 2020 | VoxelPose: Towards Multi-camera 3D Human Pose Estimation in Wild Environment
Hanyue Tu, Chunyu Wang 0001, Wenjun Zeng 0001 |
ECCV (1) | 2 |
| 2020 | Object Detection in Videos by High Quality Object LinkingabstractCompared with object detection in static images, object detection in videos is more challenging due to degraded image qualities. An effective way to address this problem is to exploit temporal contexts by linking the same object across video to form tubelets and aggregating classification scores in the tubelets. In this paper, we focus on obtaining high quality object linking results for better classification. Unlike previous methods that link objects by checking boxes between neighboring frames, we propose to link in the same frame. To achieve this goal, we extend prior methods in following aspects: (1) a cuboid proposal network that extracts spatio-temporal candidate cuboids which bound the movement of objects; (2) a short tubelet detection network that detects short tubelets in short video segments; (3) a short tubelet linking algorithm that links temporally-overlapping short tubelets to form long tubelets. Experiments on the ImageNet VID dataset show that our method outperforms both the static image detector and the previous state of the art. In particular, our method improves results by 8.8 percent over the static image detector for fast moving objects. Peng Tang 0005, Chunyu Wang 0001, Xinggang Wang, Wenyu Liu 0001, Wenjun Zeng 0001, Jingdong Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | Semantic Image Segmentation by Scale-Adaptive NetworksabstractSemantic image segmentation is an important yet unsolved problem. One of the major challenges is the large variability of the object scales. To tackle this scale problem, we propose a Scale-Adaptive Network (SAN) which consists of multiple branches with each one taking charge of the segmentation of the objects of a certain range of scales. Given an image, SAN first computes a dense scale map indicating the scale of each pixel which is automatically determined by the size of the enclosing object. Then the features of different branches are fused according to the scale map to generate the final segmentation map. To ensure that each branch indeed learns the features for a certain scale, we propose a scale-induced ground-truth map and enforce a scale-aware segmentation loss for the corresponding branch in addition to the final loss. Extensive experiments over the PASCAL-Person-Part, the PASCAL VOC 2012, and the Look into Person datasets demonstrate that our SAN can handle the large variability of the object scales and outperforms the state-of-the-art semantic segmentation methods. Chunyu Wang 0001, Xinggang Wang, Wenyu Liu 0001, Jingdong Wang 0001 |
IEEE Trans. Image Process. | 2 |
| 2019 | Learning to Refine 3D Human Pose SequencesabstractWe present a basis approach to refine noisy 3D human pose sequences by jointly projecting them onto a non-linear pose manifold, which is represented by a number of basis dictionaries with each covering a small manifold region. We learn the dictionaries by jointly minimizing the distance between the original poses and their projections on the dictionaries, along with the temporal jittering of the projected poses. During testing, given a sequence of noisy poses which are probably off the manifold, we project them to the manifold using the same strategy as in training for refinement. We apply our approach to the monocular 3D pose estimation and the long term motion prediction tasks. The experimental results on the benchmark dataset shows the estimated 3D poses are notably improved in both tasks. In particular, the smoothness constraint helps generate more robust refinement results even when some poses in the original sequence have large errors. Jieru Mei, Xingyu Chen 0001, Chunyu Wang 0001, Alan L. Yuille, Xuguang Lan, Wenjun Zeng 0001 |
3DV | 3 |
| 2019 | Learning Basis Representation to Refine 3D Human Pose EstimationsabstractEstimating 3D human poses from 2D joint positions is an illposed problem, and is further complicated by the fact that the estimated 2D joints usually have errors to which most of the 3D pose estimators are sensitive. In this work, we present an approach to refine inaccurate 3D pose estimations. The core idea of the approach is to learn a number of bases to obtain tight approximations of the low-dimensional pose manifold where a 3D pose is represented by a convex combination of the bases. The representation requires that globally the refined poses are close to the pose manifold thus avoiding generating illegitimate poses. Second, the designed bases also have the property to guarantee that the distances among the body joints of a pose are within reasonable ranges. Experiments on benchmark datasets show that our approach obtains more legitimate poses over the baselines. In particular, the limb lengths are closer to the ground truth. Chunyu Wang 0001, Haibo Qiu, Alan L. Yuille, Wenjun Zeng 0001 |
AAAI | 1 |
| 2019 | Optimizing Network Structure for 3D Human Pose EstimationabstractA human pose is naturally represented as a graph where the joints are the nodes and the bones are the edges. So it is natural to apply Graph Convolutional Network (GCN) to estimate 3D poses from 2D poses. In this work, we propose a generic formulation where both GCN and Fully Connected Network (FCN) are its special cases. From this formulation, we discover that GCN has limited representation power when used for estimating 3D poses. We overcome the limitation by introducing Locally Connected Network (LCN) which is naturally implemented by this generic formulation. It notably improves the representation capability over GCN. In addition, since every joint is only connected to a few joints in its neighborhood, it has strong generalization power. The experiments on public datasets show it: (1) outperforms the state-of-the-arts; (2) is less data hungry than alternative models; (3) generalizes well to unseen actions and datasets. Hai Ci, Chunyu Wang 0001, Xiaoxuan Ma 0001, Yizhou Wang 0001 |
ICCV | 2 |
| 2019 | Cross View Fusion for 3D Human Pose EstimationabstractWe present an approach to recover absolute 3D human poses from multi-view images by incorporating multi-view geometric priors in our model. It consists of two separate steps: (1) estimating the 2D poses in multi-view images and (2) recovering the 3D poses from the multi-view 2D poses. First, we introduce a cross-view fusion scheme into CNN to jointly estimate 2D poses for multiple views. Consequently, the 2D pose estimation for each view already benefits from other views. Second, we present a recursive Pictorial Structure Model to recover the 3D pose from the multi-view 2D poses. It gradually improves the accuracy of 3D pose with affordable computational cost. We test our method on two public datasets H36M and Total Capture. The Mean Per Joint Position Errors on the two datasets are 26mm and 29mm, which outperforms the state-of-the-arts remarkably (26mm vs 52mm, 29mm vs 35mm). Haibo Qiu, Chunyu Wang 0001, Jingdong Wang 0001, Naiyan Wang, Wenjun Zeng 0001 |
ICCV | 2 |
| 2019 | Robust 3D Human Pose Estimation from Single Images or Video SequencesabstractWe propose a method for estimating 3D human poses from single images or video sequences. The task is challenging because: (a) many 3D poses can have similar 2D pose projections which makes the lifting ambiguous, and (b) current 2D joint detectors are not accurate which can cause big errors in 3D estimates. We represent 3D poses by a sparse combination of bases which encode structural pose priors to reduce the lifting ambiguity. This prior is strengthened by adding limb length constraints. We estimate the 3D pose by minimizing an$L_1$norm measurement error between the 2D pose and the 3D pose because it is less sensitive to inaccurate 2D poses. We modify our algorithm to output$K$3D pose candidates for an image, and for videos, we impose a temporal smoothness constraint to select the best sequence of 3D poses from the candidates. We demonstrate good results on 3D pose estimation from static images and improved performance by selecting the best 3D pose from the$K$proposals. Our results on video sequences also show improvements (over static images) of roughly 15%. Chunyu Wang 0001, Yizhou Wang 0001, Zhouchen Lin, Alan L. Yuille |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | Video Object Segmentation by Learning Location-Sensitive Embeddings
Hai Ci, Chunyu Wang 0001, Yizhou Wang 0001 |
ECCV (11) | 2 |
| 2018 | Online Dictionary Learning for Approximate Archetypal Analysis
Jieru Mei, Chunyu Wang 0001, Wenjun Zeng 0001 |
ECCV (3) | 2 |
| 2017 | Learning Discriminative Activated Simplices for Action RecognitionabstractWe address the task of action recognition from a sequence of 3D human poses. This is a challenging task firstly because the poses of the same class could have large intra-class variations either caused by inaccurate 3D pose estimation or various performing styles. Also different actions, e.g., walking vs. jogging, may share similar poses which makes the representation not discriminative to differentiate the actions. To solve the problems, we propose a novel representation for 3D poses by a mixture of Discriminative Activated Simplices (DAS). Each DAS consists of a few bases and represent pose data by their convex combinations. The discriminative power of DAS is firstly realized by learning discriminative bases across classes with a block diagonal constraint enforced on the basis coefficient matrix. Secondly, the DAS provides tight characterization of the pose manifolds thus reducing the chance of generating overlapped DAS between similar classes. We justify the power of the model on benchmark datasets and witness consistent performance improvements. Chenxu Luo, Chunyu Wang 0001, Yizhou Wang 0001 |
AAAI | 3 |
| 2016 | Recognizing Actions in 3D Using Action-Snippets and Activated SimplicesabstractPose-based action recognition in 3D is the task of recognizing an action (e.g., walking or running) from a sequence of 3D skeletal poses. This is challenging because of variations due to different ways of performing the same action and inaccuracies in the estimation of the skeletal poses. The training data is usually small and hence complex classifiers risk over-fitting the data. We address this task by action-snippets which are short sequences of consecutive skeletal poses capturing the temporal relationships between poses in an action. We propose a novel representation for action-snippets, called activated simplices. Each activity is represented by a manifold which is approximated by an arrangement of activated simplices. A sequence (of action-snippets) is classified by selecting the closest manifold and outputting the corresponding activity. This is a simple classifier which helps avoid over-fitting the data but which significantly outperforms state-of-the-art methods on standard benchmarks. Chunyu Wang 0001, John Flynn, Yizhou Wang 0001, Alan L. Yuille |
AAAI | 1 |
| 2016 | Mining 3D Key-Pose-Motifs for Action RecognitionabstractRecognizing an action from a sequence of 3D skeletal poses is a challenging task. First, different actors may perform the same action in various styles. Second, the estimated poses are sometimes inaccurate. These challenges can cause large variations between instances of the same class. Third, the datasets are usually small, with only a few actors performing few repetitions of each action. Hence training complex classifiers risks over-fitting the data. We address this task by mining a set of key-pose-motifs for each action class. A key-pose-motif contains a set of ordered poses, which are required to be close but not necessarily adjacent in the action sequences. The representation is robust to style variations. The key-pose-motifs are represented in terms of a dictionary using soft-quantization to deal with inaccuracies caused by quantization. We propose an efficient algorithm to mine key-pose-motifs taking into account of these probabilities. We classify a sequence by matching it to the motifs of each class and selecting the class that maximizes the matching score. This simple classifier obtains state-of the-art performance on two benchmark datasets. Chunyu Wang 0001, Yizhou Wang 0001, Alan L. Yuille |
CVPR | 1 |
| 2014 | Robust Estimation of 3D Human Poses from a Single ImageabstractHuman pose estimation is a key step to action recognition. We propose a method of estimating 3D human poses from a single image, which works in conjunction with an existing 2D pose/joint detector. 3D pose estimation is challenging because multiple 3D poses may correspond to the same 2D pose after projection due to the lack of depth information. Moreover, current 2D pose estimators are usually inaccurate which may cause errors in the 3D estimation. We address the challenges in three ways: (i) We represent a 3D pose as a linear combination of a sparse set of bases learned from 3D human skeletons. (ii) We enforce limb length constraints to eliminate anthropomorphically implausible skeletons. (iii) We estimate a 3D pose by minimizing the 1-norm error between the projection of the 3D pose and the corresponding 2D detection. The 1-norm loss term is robust to inaccurate 2D joint estimations. We use the alternating direction method (ADM) to solve the optimization problem efficiently. Our approach outperforms the state-of-the-arts on three benchmark datasets. Chunyu Wang 0001, Yizhou Wang 0001, Zhouchen Lin, Alan L. Yuille, Wen Gao 0001 |
CVPR | 1 |
| 2013 | An Approach to Pose-Based Action RecognitionabstractWe address action recognition in videos by modeling the spatial-temporal structures of human poses. We start by improving a state of the art method for estimating human joint locations from videos. More precisely, we obtain the K-best estimations output by the existing method and incorporate additional segmentation cues and temporal constraints to select the ``best'' one. Then we group the estimated joints into five body parts (e.g. the left arm) and apply data mining techniques to obtain a representation for the spatial-temporal structures of human actions. This representation captures the spatial configurations of body parts in one frame (by spatial-part-sets) as well as the body part movements(by temporal-part-sets) which are characteristic of human actions. It is interpretable, compact, and also robust to errors on joint estimations. Experimental results first show that our approach is able to localize body joints more accurately than existing methods. Next we show that it outperforms state of the art action recognizers on the UCF sport, the Keck Gesture and the MSR-Action3D datasets. Chunyu Wang 0001, Yizhou Wang 0001, Alan L. Yuille |
CVPR | 1 |
| 2012 | PQ-WGLOH: A bit-rate scalable local feature descriptorabstractIn this paper, we propose a compact yet discriminative local descriptor which tackles the wireless query transmission latency in mobile visual search. The descriptor captures gradient statistics of canonical patches over a log-polar location grid whose parameters are optimized using training samples. We quantize the resulting descriptor using product quantization. The descriptor achieves about 95% bits reduction compared with 128-Byte SIFT and allows adaptation of descriptor lengths to support user required performance. Moreover, accurate matching of descriptors with low complexity is allowed within several table lookup operations. We perform a comprehensive comparison with SIFT, GLOH and CHoG in the context of image retrieval, image matching and object localization. We achieve competing matching and retrieval performance with SIFT, GLOH with much fewer bits. In particular, the descriptor outperforms CHoG at the same bits on eight data sets contributed to MPEG Compact Descriptor for Visual Search(CDVS) Standardization. Chunyu Wang 0001, Ling-Yu Duan, Yizhou Wang 0001, Wen Gao 0001 |
ICASSP | 1 |
| 2011 | Generating vocabulary for global feature representation towards commerce image retrievalabstractThis paper studies the problem of retrieving images by color, texture and shape in the context of visual assisted product recommendation in E-commerce sites. Different from general CBIR applications, commerce image retrieval puts more emphasis on outlier-free ranking (top N) to gain perfect user experience. We suggest to extend the bag-of-words (BoW) model to global feature characterization rather than commonly used histogram based low-level feature representation. Although BoW is a common practice in object recognition, we argue generating feature vocabulary is useful to address the global feature characterization that could be elegantly adapted to domain specific commerce image search. The representation is compact and discriminative, which may adapt with individual websites. Quantitative as well as subjective evaluation demonstrates the functionality of the proposed method. In practice, the vocabulary based global features greatly reduce outliers in top rank images, so that desirable user experience can be obtained in E-Commerce applications. Ling-Yu Duan, Chunyu Wang 0001, Tiejun Huang 0001, Wen Gao 0001 |
ICIP | 3 |