Can Wang 0006

dblp:71/4716-6 · DBLP profile ↗
← Back
18ranked-venue papers
6as first author
6since 2021 · last 2026
0009-0008-7317-6791ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 OmniPrior: A Multi-Prior-Guided Omnidirectional Representation of Dynamic Scenes in Overlapping Ultra-Wide Multi-Fisheye Videos
abstract
Omnidirectional capture of dynamic scenes facilitates the creation of immersive virtual reality assets and holistic scene understanding. Outward-facing multi-fisheye camera rigs offer an efficient solution for full-scene coverage, using fewer lenses than conventional pinhole arrays while enabling all-directional observation of complex, time-varying environments. By continuously recording scene evolution from every angle, these systems naturally enable a richer characterization of dynamic interactions. Despite these advantages, dynamic scene modeling in this setting remains underexplored. Existing methods, typically designed for fixed pinhole configurations or monocular setups, rely heavily on photometric cues and often neglect the strong geometric and semantic priors inherent in multi-fisheye omnidirectional data. To address this gap, we present OmniPrior, a Gaussian Splatting-based framework for outward-facing, multi-fisheye omnidirectional capture. Our approach incorporates metric-geometry-aware initialization with multi-prior guidance, introducing a dynamicness-aware Gaussian representation that encodes both object motion and subtle temporal variations. The resulting representations are physically consistent and temporally stable. Extensive experiments validate the effectiveness of our method in novel view synthesis across new viewpoints and timestamps. We demonstrate its utility in two representative applications derived from our learned representations: 6DoF rendering with flexible FoV and motion-freeze rendering.
Simin Kou, Jakob Nazarenus, Reinhard Koch, Can Wang 0006, Neil A. Dodgson
IEEE Trans. Vis. Comput. Graph.5
2025 Recognizing Actions From Robotic View for Natural Human-Robot Interaction
Peiming Li, Hong Liu 0008, Zhichao Deng, Can Wang 0006, Jun Liu 0036, Junsong Yuan 0001, Mengyuan Liu 0001
ICCV5
2024 CHAMP: A Large-Scale Dataset for Skeleton-Based Composite HumAn Motion Prediction
abstract
Skeleton-based human motion prediction task aims to forecast future skeleton frames conditioned by observed skeleton sequence. Different from previous methods that focus on human motion prediction for atomic actions, we observe that people are witnessed to perform composite actions which consist of atomic actions that simultaneously happen. Considering the large number of action types, it is more laborious to collect composite actions than atomic actions. This paper presents a practical composite human motion prediction task, whose training data just contains atomic actions meanwhile the test data contains both atomic actions and composite actions. To evaluate this task, we collect a large-scale Composite HumAn Motion Prediction (CHAMP) dataset, whose training data has 16 types of atomic actions and test data has 50 types of composite actions. Despite the success of previous human motion prediction methods using Graph Convolutional Networks (GCN), these methods achieve inferior performances on our CHAMP dataset due to the huge domain gap between the training and test data. To solve this problem, we present a composite human motion prediction framework containing three modules. First, a Composite Motion Synthesis (CMS) module is designed to generate synthesized composite human actions from atomic actions. Second, a Composite GCN module is presented to predict human motion by modeling different human body parts. Third, a human body partition policy network is used to choose the best partition strategy for both the CMS and Composite GCN modules. Extensive experiments on the CHAMP dataset verify the effectiveness of our framework which obviously outperforms GCN-based methods.
Mengyuan Liu 0001, Xinshun Wang, Can Wang 0006
IEEE Trans. Circuits Syst. Video Technol.5
2024 Dynamic Dense Graph Convolutional Network for Skeleton-Based Human Motion Prediction
abstract
Graph Convolutional Networks (GCN) which typically follows a neural message passing framework to model dependencies among skeletal joints has achieved high success in skeleton-based human motion prediction task. Nevertheless, how to construct a graph from a skeleton sequence and how to perform message passing on the graph are still open problems, which severely affect the performance of GCN. To solve both problems, this paper presents a Dynamic Dense Graph Convolutional Network (DD-GCN), which constructs a dense graph and implements an integrated dynamic message passing. More specifically, we construct a dense graph with 4D adjacency modeling as a comprehensive representation of motion sequence at different levels of abstraction. Based on the dense graph, we propose a dynamic message passing framework that learns dynamically from data to generate distinctive messages reflecting sample-specific relevance among nodes in the graph. Extensive experiments on benchmark Human 3.6M and CMU Mocap datasets verify the effectiveness of our DD-GCN which obviously outperforms state-of-the-art GCN-based methods, especially when using long-term and our proposed extremely long-term protocol.
Xinshun Wang, Can Wang 0006, Yuan Gao 0008, Mengyuan Liu 0001
IEEE Trans. Image Process.3
2024 Temporal Decoupling Graph Convolutional Network for Skeleton-Based Gesture Recognition
abstract
Skeleton-based gesture recognition methods have achieved high success using Graph Convolutional Network (GCN), which commonly uses an adjacency matrix to model the spatial topology of skeletons. However, previous methods use the same adjacency matrix for skeletons from different frames, which limits the flexibility of GCN to model temporal information. To solve this problem, we propose a Temporal Decoupling Graph Convolutional Network (TD-GCN), which applies different adjacency matrices for skeletons from different frames. The main steps of each convolution layer in our proposed TD-GCN are as follows. To extract deep spatiotemporal information from skeleton joints, we first extract high-level spatiotemporal features from skeleton data. Then, channel-dependent and temporal-dependent adjacency matrices corresponding to different channels and frames are calculated to capture the spatiotemporal dependencies between skeleton joints. Finally, to fuse topology information from neighbor skeleton joints, spatiotemporal features of skeleton joints are fused based on channel-dependent and temporal-dependent adjacency matrices. To the best of our knowledge, we are the first to use temporal-dependent adjacency matrices for temporal-sensitive topology learning from skeleton joints. The proposed TD-GCN effectively improves the modeling ability of GCN and achieves state-of-the-art results on gesture datasets including SHREC'17 Track and DHG-14/28.
Xinshun Wang, Can Wang 0006, Yuan Gao 0008, Mengyuan Liu 0001
IEEE Trans. Multim.3
2023 Regress Before Construct: Regress Autoencoder for Point Cloud Self-supervised Learning
abstract
Masked Autoencoders (MAE) have demonstrated promising performance in self-supervised learning for both 2D and 3D computer vision. Nevertheless, existing MAE-based methods still have certain drawbacks. Firstly, the functional decoupling between the encoder and decoder is incomplete, which limits the encoder's representation learning ability. Secondly, downstream tasks solely utilize the encoder, failing to fully leverage the knowledge acquired through the encoder-decoder architecture in the pre-text task. In this paper, we propose Point Regress AutoEncoder (Point-RAE), a new scheme for regressive autoencoders for point cloud self-supervised learning. The proposed method decouples functions between the decoder and the encoder by introducing a mask regressor, which predicts the masked patch representation from the visible patch representation encoded by the encoder and the decoder reconstructs the target from the predicted masked patch representation. By doing so, we minimize the impact of decoder updates on the representation space of the encoder. Moreover, we introduce an alignment constraint to ensure that the representations for masked patches, predicted from the encoded representations of visible patches, are aligned with the masked patch presentations computed from the encoder. To make full use of the knowledge learned in the pre-training stage, we design a new finetune mode for the proposed Point-RAE. Extensive experiments demonstrate that our approach is efficient during pre-training and generalizes well on various downstream tasks. Specifically, our pre-trained models achieve a high accuracy of 90.28% on the ScanObjectNN hardest split and 94.1% accuracy on ModelNet40, surpassing all the other self-supervised learning methods. Our code and pretrained model are public available at: https://github.com/liuyyy111/Point-RAE.
Yang Liu 0264, Chen Chen 0001, Can Wang 0006, Xulin King, Mengyuan Liu 0001
ACM Multimedia3
2016 A new descriptor of gradients Self-Similarity for smile detection in unconstrained scenarios
Yuan Gao 0008, Hong Liu 0008, Can Wang 0006
Neurocomputing4
2016 Violence detection using Oriented VIolent Flows
Yuan Gao 0008, Hong Liu 0008, Xiaohu Sun, Can Wang 0006
Image Vis. Comput.4
2015 Body-structure based feature representation for person re-identification
abstract
Person re-identification is valuable for intelligent video surveillance and has drawn wide attention. Although person re-identification research is making progress, it still faces some challenges such as varying poses, illumination and viewpoints. As a major aspect of person re-identification, feature representation has been widely researched. Low-level descriptors are generally used in existing works, which do not take full advantage of body structure information and result in low discrimination. In this paper, body-structure based mid-level feature representation is proposed, which introduces body structure pyramid for codebook learning and feature pooling. Additionally, low computational LLC is used to encode mid-level features. Experimental results on two challenging datasets VIPeR and CUHK01 have demonstrated that our approach outperforms the state-of-the-art methods.
Hong Liu 0008, Liqian Ma, Can Wang 0006
ICASSP3
2015 A predictive model for narrow passage path planner by using Support Vector Machine in changing environments
abstract
Narrow passages in changing environments create huge difficulties, since locations and shapes of narrow passages in Configuration Space(C-space) change frequently. It is very important for a planner to identify narrow passages in real time and boost valid points within them effectively. A novel narrow passage predictive model for designing a path planner in changing environments is proposed in this paper. Firstly, an Expanded Dynamic Bridge Builder is presented to identify narrow passages rapidly with validity-toggle sampling points in C-space. Secondly, the predictive model is adopted to sample possibly free points within these narrow passages without invoking any collision detection in order to avoid intense computational complexity. The predictive model is obtained by the famous classification method of Support Vector Machine (SVM). A new feature, which includes a group of points' distance and validity information, is proposed in SVM training process to capture approximate structure of local narrow passages. Therefore, the predictive model can excavate the hidden similar structure of local narrow passages. Experiments carried out with two 6-DOFs manipulators show that our approach gain higher success rate of planning and time efficiency than other related methods.
Hong Liu 0008, Fang Xiao, Can Wang 0006
ICRA3
2014 Gender identification in unconstrained scenarios using Self-Similarity of Gradients features
abstract
Gender identification has been a hot research topic with wide application requirements from social life. In general, effective feature representation is the key to solving this problem. In this paper, a new feature named Self-Similarity of Gradients (GSS) is proposed, which captures pairwise statistics of localized gradient distributions. There are three contributions made by us to practical gender identification. First, GSS features are proposed for gender identification in the wild, which achieve good performance compared with baseline approaches. Second, we originally utilize 31-dimensional HOG for practical gender identification and its excellent results demonstrates that HOG with both contrast sensitive and insensitive information is a better fit for this topic than that with only contrast insensitive information. Last, feature combination and multi-classifier combination strategies are adopted and the best gender identification performance is achieved. Experimental results show that the combination of GSS, HOG and LBP using a linear SVM outperforms state-of-the-art on the LFW database, which meets the “wild” condition.
Hong Liu 0008, Yuan Gao 0008, Can Wang 0006
ICIP3
2014 Contact-free and pose-invariant hand-biometric-based personal identification system using RGB and depth data
abstract
Hand-biometric-based personal identification is considered to be an effective method for automatic recognition. However, existing systems require strict constraints during data acquisition, such as costly devices, specified postures, simple background, and stable illumination. In this paper, a contactless personal identification system is proposed based on matching hand geometry features and color features. An inexpensive Kinect sensor is used to acquire depth and color images of the hand. During image acquisition, no pegs or surfaces are used to constrain hand position or posture. We segment the hand from the background through depth images through a process which is insensitive to illumination and background. Then finger orientations and landmark points, like finger tips or finger valleys, are obtained by geodesic hand contour analysis. Geometric features are extracted from depth images and palmprint features from intensity images. In previous systems, hand features like finger length and width are normalized, which results in the loss of the original geometric features. In our system, we transform 2D image points into real world coordinates, so that the geometric features remain invariant to distance and perspective effects. Extensive experiments demonstrate that the proposed hand-biometric-based personal identification system is effective and robust in various practical situations.
Can Wang 0006, Hong Liu 0008
J. Zhejiang Univ. Sci. C1
2014 Scene-Adaptive Hierarchical Data Association for Multiple Objects Tracking
abstract
Obtaining reliable and discriminative target representation are two vital tasks for data association in multi-tracking. Pervious works always directly combine bunch of features for more discriminative target representation, but this is prone to error accumulation and unnecessary computational cost, which on the contrary may increase identity switches in data association. Moreover, reliability of a same feature in different scenes may vary a lot, especially for currently widespread network cameras, which have been settled in complex and various scenes, previous fixed feature selection scheme cannot meet general requirements. To address this problem, we propose a scene-adaptive hierarchical data association scheme, which adaptively selects features which have higher reliability on target representation in applied scene, and gradually combines features to the minimum requirements of discriminating ambiguous targets. Hierarchical feature space is constructed according to reliability of features in the multi-tracking system, and data association is conducted in different layers of the feature space adaptively. Our algorithm is validated on various challenging RGB-D and RGB datasets recorded in various indoor and outdoor scenes, for diversities of both features and scenes. Experimental results validate its effectiveness and efficiency.
Can Wang 0006, Hong Liu 0008, Yuan Gao 0008
IEEE Signal Process. Lett.1
2014 Depth Motion Detection - A Novel RS-Trigger Temporal Logic based Method
abstract
Recently, depth data is widely used in computer vision applications such as detection and tracking, which shows great promises in complicated environments due to its complementary natures to RGB data. However, previous works mostly use depth as an auxiliary cue of RGB data and overlook its inherent advantage on motion detection. Intrinsically different from RGB data, points in depth map essentially represents 3-D positions in the world, so depth video represents the variation of these “positions,” which is motion. Motivated by this, we proposed a novel motion detection scheme based on RS-Trigger temporal logic which best fits nature of depth data on motion detection. The proposed algorithm can fast detect motion regions in the scene without statistics of background and prior knowledge of objects to detect. In following refinement modules, a depth-invariant density-constant projection is proposed which contributes to a fast spatial clustering and accurate segmentation, for it transforms dense 3-D points cloud to depth-invariant 2-D map with density-constance, not only it overcomes depth-dependent sampling of depth sensor, but also overcomes the common ‘scale problem’ in 2-D image analysis, which makes it easy to set system parameters to de-noise and pop-out motion regions. Experimental results validate its effectiveness and efficiency.
Can Wang 0006, Hong Liu 0008, Liqian Ma
IEEE Signal Process. Lett.1
2013 Salient-motion-heuristic scheme for fast 3D optical flow estimation using RGB-D data
abstract
Optical flow is widely used for describing motion cues in the scene, but limited by slow estimating speed and illumination sensitivity. To handle both problems, this paper focuses on improving speed and accuracy of optical flow using RGB-D data and enhancing its robustness on motion description via fusing depth flow which is obtained only using depth data. First, salient motion regions (SMRs) are detected between depth frames which have good character on motion description for they all locate on moving objects. Then, depth flow is calculated to describe 3D motion for each SMR and directs fast orientation region growing on depth map. Thus larger motion regions are grown, and region-based optical flow estimation is conducted on grown regions. Estimation error is reduced and noise is inhibited due to depth constraints. Finally, a fusion scheme is adopted which combines depth flow and optical flow for better 3D motion description in the scene. Experiments on a RGB-D video data sets recorded in various complex scenes demonstrate the improved speed and robustness of the proposed method.
Can Wang 0006, Hong Liu 0008
ICASSP1
2013 Hierarchical data association and depth-invariant appearance model for indoor multiple objects tracking
abstract
Discriminative target representation is vital for data association in multi-tracking. In order to increase the discriminative power, pervious works always combine bunch of features for target representation. However, this is prone to error accumulation and unnecessary computational cost, which may increase identity switches in data association on the contrary. To address this problem, we propose a hierarchical data association scheme which gradually combines features to the minimum requirements of discriminating ambiguous targets. In addition, indoor multi-tracking is more challenging due to frequent occlusion, view-truncation, large scale and pose variation, which may bring considerable unreliability for target representation. To handle this a novel depth-invariant part-based appearance model using RGB-D data is proposed. The depth-invariant appearance have stable length metric proportional to the absolute length metric in the world coordinates, which increase its robustness to scale variation. The part-based nature makes it robust to partial occlusion and view-truncation. Our algorithm is validated on various challenging indoor environments and it demonstrates high processing speed up to 50 fps and competitive accuracy.
Hong Liu 0008, Can Wang 0006
ICIP2
2013 Unusual events detection based on multi-dictionary sparse representation using kinect
abstract
Unusual events detection plays a crucial role in surveillance applications, which is becoming more and more urgent need for public security. However, illumination and scale changing, lacking of sufficient training data and subjective of abnormality definition are some of the severe difficulties, which are hard to deal with by widely used traditional cameras. In order to solve these problems, first, a novel feature is proposed in this paper, which is named random local feature (RLF) to describe the spatial-temporal information of depth image detected by the Kinect sensor. Then, we expand the sparse representation framework to a multi-dictionary sparse representation framework, based on the intuition that that anomaly of a same event may vary a lot in different regions in a scene. We split the depth video into several regions and use detected RLF features in each region to train dictionary by K-SVD algorithm, and use the OMP algorithm to sparse-represent each feature. Finally, an objective function is introduced to evaluate the anomaly of features in each region according to reconstruction errors. Unusual events are defined as those incidences that occur very rarely in the entire video sequence in our system, which is tested on real data and demonstrates promising results in unusual events detection.
Can Wang 0006, Hong Liu 0008
ICIP1
2013 Maximally stable curvature regions for 3D hand tracking
abstract
Fast and robust hand detections and tracking is in increasing demand from areas such as natural Human Robot interaction(HRI) and surveillance systems. Previous works always use skin color or contour model to detection hand. However, they always fail for hands always exhibits drastic appearance change due to illumination change, non-rigid nature and hands are hard to discriminate from clutter background. Actually, the hand region has a specific nature that its curvature is relatively higher than other body parts and keeps stable whatever its poses and locations are, but none of pervious works exploit this nature for hand detection. In this work, a novel algorithm MSCR (Maximally Stable Curvature Regions) based on curvature nature to detect hands. It does not require manually initialization in the first frame, the hands are located by MSCR and skin color detector in the global image. 3D optical flow integrated Kalman Filter works to estimate the next location for local detector. Extensive experiments demonstrate that robust 3D tracking of hand articulations can be achieved in real-time with accurate results.
Can Wang 0006, Hong Liu 0008
ICIP1