EDBT 2026 Demo / reviewers in the wild / expert
Xu Zhao 0001
dblp:37/3580-1
· DBLP profile ↗
83ranked-venue papers
11as first author
29since 2021 · last 2026
0000-0002-8176-623XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 70 · 8 first-author · 22 since 2021Artificial intelligence and machine learning · 32 · 5 first-author · 14 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Diffusion-based Personalized Pathology Disentanglement for Impaired Gait AnalysisabstractIn the context of global population aging, the prevalence of neurodegenerative diseases is rapidly increasing. Vision-based impaired gait analysis emerges as a promising alternative for automatic and non-invasive diagnosis. While prior efforts have advanced either accuracy or interpretability of gait analysis, few have effectively addressed both aspects in a unified framework. To bridge this gap, we propose DPPD, a Diffusion-based Personalized Pathology Disentanglement model that jointly performs quantitative gait scoring, dementia subtyping, and qualitative anomaly highlighting. Motivated by the observation that pathological gait features exhibit stronger inter-class separability across different gait severity than raw features, DPPD is proposed based on the subject-specific pathology disentanglement perspective. Specifically, it comprises three key components: (1) a 3DmotionBERT for encoding gait representation from 3D human pose sequences estimated, (2) a latent diffusion-based Gait Denoiser for generating personalized normal gait features, and (3) a Dual Pathology Disentanglement mechanism that captures both static pose and dynamic motion pathological representation from the residual between raw and normal gait features. These disentangled pathologies further enable quantitative classification and qualitative anomaly highlighting. Experiments on the PDGait and 3DGait datasets demonstrate that DPPD outperforms state-of-the-art methods in classification accuracy while providing reliable and interpretable visualizations of gait anomalies. Xiaoyue Wan, Xu Zhao 0001 |
AAAI | 2 |
| 2026 | MESA: Effective Matching Redundancy Reduction by Semantic Area SegmentationabstractMatching redundancy, which refers to fine-grained feature comparison between irrelevant image areas, is a prevalent limitation in current feature matching approaches. It leads to unnecessary and error-prone computations, ultimately diminishing matching accuracy. To reduce matching redundancy, we propose MESA and DMESA, both leveraging advanced image understanding of Segment Anything Model (SAM) to establish semantic area matches prior to point matching. These informative area matches, then, can undergo effective internal feature comparison, facilitating precise inside-area point matching. Specifically, MESA adopts a sparse matching framework, while DMESA applies a dense one. Both of them first obtain candidate areas from SAM results through a novel Area Graph (AG). In MESA, matching the candidates is formulated as a graph energy minimization and solved by graphical models derived from AG. In contrast, DMESA performs area matching by generating dense matching distributions on the entire image, aiming at enhancing efficiency. The distributions are produced from off-the-shelf patch matching, modeled as the Gaussian Mixture Model, and refined via the Expectation Maximization. With less repetitive computation, DMESA showcases an area matching speed improvement of nearly five times compared to MESA, while maintaining competitive accuracy. Our methods are extensively evaluated on four different tasks across six datasets, encompassing both indoor and outdoor scenes. The results suggest that our method achieves notable accuracy improvements for nine baselines of point matching in most cases. Furthermore, our methods exhibit promise generalization and improved robustness against image resolution. Yesheng Zhang, Shuhan Shen, Xu Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Learning multi-scale spatial-frequency features for image denoising
Xu Zhao 0001, Chen Zhao 0002, Xiantao Hu, Hongliang Zhang 0002, Ying Tai, Jian Yang 0003 |
Pattern Recognit. | 1 |
| 2025 | Exploiting Multimodal Spatial-temporal Patterns for Video Object TrackingabstractMultimodal tracking has garnered widespread attention as a result of its ability to effectively address the inherent limitations of traditional RGB tracking. However, existing multimodal trackers mainly focus on the fusion and enhancement of spatial features or merely leverage the sparse temporal relationships between video frames. These approaches do not fully exploit the temporal correlations in multimodal videos, making it difficult to capture the dynamic changes and motion information of targets in complex scenarios. To alleviate this problem, we propose a unified multimodal spatial-temporal tracking approach named STTrack. In contrast to previous paradigms that solely relied on updating reference information, we introduced a temporal state generator (TSG) that continuously generates a sequence of tokens containing multimodal temporal information. These temporal information tokens are used to guide the localization of the target in the next time state, establish long-range contextual relationships between video frames, and capture the temporal trajectory of the target. Furthermore, at the spatial level, we introduced the mamba fusion and background suppression interactive (BSI) modules. These modules establish a dual-stage mechanism for coordinating information interaction and fusion between modalities. Extensive comparisons on five benchmark datasets illustrate that STTrack achieves state-of-the-art performance across various multimodal tracking scenarios. Xiantao Hu, Ying Tai, Xu Zhao 0001, Chen Zhao 0002, Zhenyu Zhang 0005, Jun Li 0027, Bineng Zhong 0001, Jian Yang 0003 |
AAAI | 3 |
| 2025 | Weakly-Supervised Video Highlight Detection by Characteristic and Commonality ModelingabstractVideo highlight detection is important for video understanding, as it localizes the attractive regions in the video automatically. Because the fully-supervised video highlight is expensive for its frame-level annotation, we present a novel network for weakly-supervised video highlight detection based on the characteristic and commonality of key moments, utilizing the intrinsic properties of highlight moments. Specifically, the characteristic of key moments refer to the fact that the attractive segments in a video often show significant differences from the ordinary segments, i.e., non-highlight segments. However, apart from the real highlight moment showing significant difference to the ordinary, the noisy frames are also different. Therefore, commonalities among the videos within the same class is proposed to ensure the rationality of the selected moments. These proposed properties of the highlight moment are implemented as two auxiliary modules with the commonly used weakly supervised video highlight detection architecture, injecting the proposed prior information to the network. In addition, an video-level importance aggregation module is proposed to estimate the importance scores for each segment in an unified model. Our method achieves the superior performance on the commonly used benchmarks. Chengze Zhao, Xu Zhao 0001 |
ICASSP | 3 |
| 2025 | Semantic-Guided Camera Ray Regression for Visual Localization
Yesheng Zhang, Xu Zhao 0001 |
ICCV | 2 |
| 2025 | Constructing Semantical Structure by Segmentation Integrated Video Embedding for Temporal Action DetectionabstractVideo embedding is the pivot in Temporal Action Detection (TAD). Once the video embedding can robustly capture the essence of actions and perceive activities in complex scenes, the TAD model can more accurately localize action boundaries. Currently, video embedding is typically based on rule-based pixel convolution or cube-based transformer, wherein structured semantic information is intertwined, leading to the submergence of crucial spatial semantic information, such as the intrinsic motion of key semantic objects and interactions among semantic objects. To address these limitations, it is imperative to explore alternative approaches. With the remarkable performance of general semantic segmentation models in visual representation, we introduce the general segmentation model SEEM into the video embedding paradigm, constructing a semantically structured representation from perceptual semantics to cognitive semantics. To more effectively utilize SEEM for structured video representation, we designed the Semantic Adapter (Sem-Adapter) as a bridge to connect the two models. Firstly, we design a Self-Motion Module (SMM) to pay attention to the self-motion of key semantic regions. Secondly, we propose a Mutual Relation Module (MRM) to construct the interactions between semantic regions. Extensive experiments on ActivityNet-1.3, THUMOS-14 and EPIC-Kitchens-100 reveal that our method significantly outperforms state-of-the-art methods under the same input modality, and our method improves the average mAP from 60.6% to 64.2% on THUMOS-14 with the same backbone. The code is available onhttps://github.com/shouxiaozixuan/semtad. Shuming Liu 0001, Chengze Zhao, Xu Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | RSB-Pose: Robust Short-Baseline Binocular 3D Human Pose Estimation With Occlusion HandlingabstractIn the domain of 3D Human Pose Estimation, which finds widespread daily applications, the requirement for convenient acquisition equipment continues to grow. To satisfy this demand, we focus on a short-baseline binocular setup that offers both portability and a geometric measurement capability that significantly reduces depth ambiguity. However, as the binocular baseline shortens, two serious challenges emerge: first, the robustness of 3D reconstruction against 2D errors deteriorates; second, occlusion reoccurs frequently due to the limited visual differences between two views. To address the first challenge, we propose the Stereo Co-Keypoints Estimation module to improve the view consistency of 2D keypoints and enhance the 3D robustness. In this module, the disparity is utilized to represent the correspondence of binocular 2D points, and the Stereo Volume Feature (SVF) is introduced to contain binocular features across different disparities. Through the regression of SVF, two-view 2D keypoints are simultaneously estimated in a collaborative way which restricts their view consistency. Furthermore, to deal with occlusions, a Pre-trained Pose Transformer module is introduced. Through this module, 3D poses are refined by perceiving pose coherence, a representation of joint correlations. This perception is injected by the Pose Transformer network and learned through a pre-training task that recovers iterative masked joints. Comprehensive experiments on H36M and MHAD datasets validate the effectiveness of our approach in the short-baseline binocular 3D Human Pose Estimation and occlusion handling. Xiaoyue Wan, Zhuo Chen 0028, Xu Zhao 0001 |
IEEE Trans. Image Process. | 3 |
| 2024 | FSGait: Fine-Grained Self-supervised Gait Abnormality Detection
Bingzhi Duan, Xiaoyue Wan, Xu Zhao 0001 |
ACCV (6) | 3 |
| 2024 | MESA: Matching Everything by Segmenting AnythingabstractFeature matching is a crucial task in the field of computer vision, which involves finding correspondences between images. Previous studies achieve remarkable performance using learning-based feature comparison. However, the pervasive presence of matching redundancy between images gives rise to unnecessary and error-prone computations in these methods, imposing limitations on their accuracy. To address this issue, we propose MESA, a novel approach to establish precise area (or region) matches for efficient matching redundancy reduction. MESA first leverages the advanced image understanding capability of SAM, a state-of-the-art foundation model for image segmentation, to obtain image areas with implicit semantic. Then, a multi-relational graph is proposed to model the spatial structure of these areas and construct their scale hierarchy. Based on graphical models derived from the graph, the area matching is reformulated as an energy minimization task and effectively resolved. Extensive experiments demonstrate that MESA yields substantial precision improvement for multiple point matchers in indoor and outdoor downstream tasks, e.g. +13.61% for DKM in indoor pose estimation. Yesheng Zhang, Xu Zhao 0001 |
CVPR | 2 |
| 2024 | Dual-Diffusion for Binocular 3D Human Pose EstimationabstractBinocular 3D human pose estimation (HPE), reconstructing a 3D pose from 2D poses of two views, offers practical advantages by combining multiview geometry with the convenience of a monocular setup. However, compared to a multiview setup, the reduction in the number of cameras increases uncertainty in 3D reconstruction. To address this issue, we leverage the diffusion model, which has shown success in monocular 3D HPE by recovering 3D poses from noisy data with high uncertainty. Yet, the uncertainty distribution of initial 3D poses remains unknown. Considering that 3D errors stem from 2D errors within geometric constraints, we recognize that the uncertainties of 3D and 2D are integrated in a binocular configuration, with the initial 2D uncertainty being well-defined. Based on this insight, we propose Dual-Diffusion specifically for Binocular 3D HPE, simultaneously denoising the uncertainties in 2D and 3D, and recovering plausible and accurate results. Additionally, we introduce Z-embedding as an additional condition for denoising and implement baseline-width-related pose normalization to enhance the model flexibility for various baseline settings. This is crucial as 3D error influence factors encompass depth and baseline width. Extensive experiments validate the effectiveness of our Dual-Diffusion in 2D refinement and 3D estimation. The code and models are available at https://github.com/sherrywan/Dual-Diffusion. Xiaoyue Wan, Zhuo Chen 0028, Bingzhi Duan, Xu Zhao 0001 |
NeurIPS | 4 |
| 2024 | M3Net: Movement Enhancement with Multi-Relation toward Multi-Scale video representation for Temporal Action Detection
Dongqi Wang 0006, Xu Zhao 0001 |
Pattern Recognit. | 3 |
| 2024 | An Embeddable Implicit IUVD Representation for Part-Based 3D Human Surface ReconstructionabstractTo reconstruct a 3D human surface from a single image, it is crucial to simultaneously consider human pose, shape, and clothing details. Recent approaches have combined parametric body models (such as SMPL), which capture body pose and shape priors, with neural implicit functions that flexibly learn clothing details. However, this combined representation introduces additional computation, e.g. signed distance calculation in 3D body feature extraction, leading to redundancy in the implicit query-and-infer process and failing to preserve the underlying body shape prior. To address these issues, we propose a novel IUVD-Feedback representation, consisting of an IUVD occupancy function and a feedback query algorithm. This representation replaces the time-consuming signed distance calculation with a simple linear transformation in the IUVD space, leveraging the SMPL UV maps. Additionally, it reduces redundant query points through a feedback mechanism, leading to more reasonable 3D body features and more effective query points, thereby preserving the parametric body prior. Moreover, the IUVD-Feedback representation can be embedded into any existing implicit human reconstruction pipeline without requiring modifications to the trained neural networks. Experiments on the THuman2.0 dataset demonstrate that the proposed IUVD-Feedback representation improves the robustness of results and achieves three times faster acceleration in the query-and-infer process. Furthermore, this representation holds potential for generative applications by leveraging its inherent semantic information from the parametric body model. Baoxing Li, Yehui Yang, Xu Zhao 0001 |
IEEE Trans. Image Process. | 4 |
| 2024 | Joint-Limb Compound Triangulation With Co-Fixing for Stereoscopic Human Pose EstimationabstractAs a special subset of multi-view settings for 3D human pose estimation, stereoscopic settings show promising applications in practice since they are not ill-posed but could be as mobile as monocular ones. However, when there are only two views, the problems of occlusions and “double counting” (ambiguity between symmetric joints) pose greater challenges that are not addressed by previous approaches. On this concern, we propose a novel framework to detect limb orientations in field form and incorporate them explicitly with joint features. Two modules are proposed to realize the fusion. At 3D level, we designcompound triangulationas an explicit module that produces the optimal pose using 2D joint locations and limb orientations. The module is derived from reformulating triangulation in 3D space, and expanding it with the optimization of limb orientations. At 2D level, we propose a parameter-free module namedco-fixingto enable joint and limb features to fix each other to alleviate the impact of “double counting.” Features from both parts are first used to infer each other via simple convolutions and then fixed by the inferred ones respectively. We test our method on two public benchmarks, Human3.6M and Total Capture, and our method achieves state-of-the-art performance on stereoscopic settings and comparable results on common 4-view benchmarks. Zhuo Chen 0028, Xiaoyue Wan, Yiming Bao, Xu Zhao 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Movement Enhancement toward Multi-Scale Video Feature Representation for Temporal Action DetectionabstractBoundary localization is a challenging problem in Temporal Action Detection (TAD), in which there are two main issues. First, the submergence of movement feature, i.e. the movement information in a snippet is covered by the scene information. Second, the scale of action, that is, the proportion of action segments in the entire video, is considerably variable. In this work, we first design a Movement Enhance Module (MEM) to highlight movement feature for better action location, and then, we propose a Scale Feature Pyramid Network (SFPN) to detect multi-scale actions in videos. For Movement Enhance Module, firstly, Movement Feature Extractor (MFE) is designed to get the movement feature. Secondly, we propose a Multi-Relation Enhance Module (MREM) to grasp valuable information correlation both locally and temporally. For Scale Feature Pyramid Network, we design a U-Shape Module to model different scale actions, moreover, we design the training and inference strategy of different scales, ensuring that each pyramid layer is only responsible for actions at a specific scale. These two innovations are integrated as the Movement Enhance Network (MENet), and extensive experiments conducted on two challenging benchmarks demonstrate its effectiveness. MENet outperforms other representative TAD methods on ActivityNet-1.3 and THUMOS-14. Dongqi Wang 0006, Xu Zhao 0001 |
ICCV | 3 |
| 2023 | View consistency aware holistic triangulation for 3D human pose estimationabstractThe rapid development of multi-view 3D human pose estimation (HPE) is attributed to the maturation of monocular 2D HPE and the geometry of 3D reconstruction. However, 2D detection outliers in occluded views due to neglect of view consistency, and 3D implausible poses due to lack of pose coherence, remain challenges. To solve this, we introduce a Multi-View Fusion module to refine 2D results by establishing view correlations. Then, Holistic Triangulation is proposed to infer the whole pose as an entirety, and anatomy prior is injected to maintain the pose coherence and improve the plausibility. Anatomy prior is extracted by PCA whose input is skeletal structure features, which can factor out global context and joint-by-joint relationship from abstract to concrete. Benefiting from the closed-form solution, the whole framework is trained end-to-end. Our method outperforms the state of the art in both precision and plausibility which is assessed by a new metric. Xiaoyue Wan, Zhuo Chen 0028, Xu Zhao 0001 |
Comput. Vis. Image Underst. | 3 |
| 2023 | Temporally consistent reconstruction of 3D clothed human surface with warp field
Baoxing Li, Yehui Yang, Xu Zhao 0001 |
Image Vis. Comput. | 4 |
| 2023 | FusePose: IMU-Vision Sensor Fusion in Kinematic Space for Parametric Human Pose EstimationabstractCommercial motion-capture systems produce excell- ent in-studio reconstructions, but offer no comparable solution for acquisition in everyday environments. We present a system for acquiring motions almost anywhere. This wearable system gathers ultrasonic time-of-flight and inertial measurements with a set of inexpensive miniature sensors worn on the garment. After recording, the information is combined using an Extended Kalman Filter to reconstruct joint configurations of a body. Experimental results show that even motions that are traditionally difficult to acquire are recorded with ease within their natural settings. Although our prototype does not reliably recover the global transformation, we show that the resulting motions are visually similar to the original ones, and that the combined acoustic and intertial system reduces the drift commonly observed in purely inertial systems. Our final results suggest that this system could become a versatile input device for a variety of augmented-reality applications. Yiming Bao, Xu Zhao 0001, Dahong Qian |
IEEE Trans. Multim. | 2 |
| 2022 | Structural Triangulation: A Closed-Form Solution to Constrained 3D Human Pose Estimation
Zhuo Chen 0028, Xu Zhao 0001, Xiaoyue Wan |
ECCV (5) | 2 |
| 2022 | Mask-Based Attention Parallel Network for in-the-Wild Facial Expression RecognitionabstractFacial expression recognition suffers big pose and occlusion in real world and attention mechanism is deployed widely to cope with these challenges. But most previous attention-based methods are inadequate in locating crucial expression-related regions precisely and capturing useful facial expression features comprehensively. For these reasons, we present a novel mask-based attention parallel network (MAPNet). Firstly, mask-based attention module that locates expression-related regions is constructed from binary mask extracted by key landmark detection. Secondly, the designed parallel network embeds mask-based attention modules into its different layers to acquire comprehensive facial expression features. Thirdly, the extracted parallel features are divided into several detached blocks from spatial dimension to predict facial expression independently. Finally, the expression label is acquired by combining two predictions of the parallel network and a new loss function is designed to weigh unbalanced facial expression distribution. We validate our method on three popular in-the-wild datasets and the results demonstrate that our MANPnet outperforms previous state-of-the-art methods among RAFDB, AffectNet and FEDRO. Lingzhao Ju, Xu Zhao 0001 |
ICASSP | 2 |
| 2022 | Spatio-Temporal Motion Aggregation Network for Video Action DetectionabstractRecognizing action patterns and detecting action instances are vital for spatial temporal action detection task, which aims to recognize the actions of interest in untrimmed videos and localize them in both space and time. The mainstream action tubelet detectors, how-ever, ignore the conflicts in features between localization and classification, and use localization features for temporal modeling, which leads to ineffective action classification. In this paper, we propose the Spatio-Temporal Motion Aggregation mechanism for integrating the local motion feature from a short term snippet and the longer spatio-temporal information to predict the action category. We design the Class-Agnostic Center Localization module to perform action instance center localization in the Class-Agnostic manner. Besides, Movement and Size Regression is proposed for movement estimation and spatial extent detection by using Gaussian kernels to encode training samples. These three modules work together to generate the tubelet detection results, which could be further linked to yield video-level tubes with a matching strategy. Our detector achieves the state-of-the-art performance in both frame-mAP and video-mAP metrics, on the UCF-24 and JHMDB datasets. Hongcheng Zhang, Xu Zhao 0001 |
ICASSP | 2 |
| 2022 | Mr.CAN: Class-Aware Network with Multi-Relations for Temporal Action DetectionabstractRecognizing action patterns and exploring multiple relations are vital for Temporal Action Detection (TAD) task, which aims at locating and classifying action segments in untrimmed videos. However, most existing methods attempt to build a general model to handle diverse actions, ignoring the huge difference between various classes. Besides, the exploration of temporal and semantic relations between different segments remains an ongoing challenge due to complex video content. In this paper, we contend that different action classes should be processed differently and thus design a new Class-Aware Mechanism to achieve accurate detection. Moreover, an effective module named Multi-relations Builder is proposed to establish temporal and semantic relations simultaneously. These two modules are integrated as Class-Aware Network with Multi-relations (MrCAN). In comprehensive experiments conducted on two benchmarks, it out-performs all other current methods and achieves state-of-the-art performance, improving the average mAP from 45.78% to 48.98% on THUMOS-14 and from 35.52% to 35.87% on ActivityNet-1.3 respectively. Furthermore, the well-designed Multi-relations Builder can also be used to boost some other existing methods. Dongqi Wang 0006, Xu Zhao 0001 |
ICME | 3 |
| 2022 | BACNet: Boundary-Anchor Complementary Network for Temporal Action DetectionabstractThe task of temporal action detection aims to locate and classify action segments in untrimmed videos. Most existing works usually consist of two components: snippet-level boundary segmentation and anchor-level action evaluation. These two components, however, are typically designed ir-relevantly, so the detection accuracy is undermined due to vague boundaries and complex video content. To tackle this problem, we design two supplementary modules. One mod-ule, termed as Anchor Aware Module (AAM), uses tem-poral and semantic related anchors to enhance snippet feature. The other module, named Boundary Aware Module (BAM), endows anchor feature with structured representation using intermediate supervision. Moreover, the ConvL-STM is applied to establish temporal relation in BAM with the structured representation. These two modules are in-tegrated as the Boundary-Anchor Complementary Network (BACNet), which achieves the state-of-the-art performance on both THUMOS-14 and ActivityNet-1.3 datasets. Dongqi Wang 0006, Xu Zhao 0001 |
ICME | 3 |
| 2022 | Semi-supervised Learning for Multi-label Video Action DetectionabstractSemi-supervised multi-label video action detection aims to locate all the persons and recognize their multiple action labels by leveraging both labeled and unlabeled videos. Compared to the single-label scenario, semi-supervised learning in multi-label video action detection is more challenging due to two significant issues: generation of multiple pseudo labels and class-imbalanced data distribution. In this paper, we propose an effective semi-supervised learning method to tackle these challenges. Firstly, to make full use of the informative unlabeled data for better training, we design an effective multiple pseudo labeling strategy by setting dynamic learnable threshold for each class. Secondly, to handle the long-tailed distribution for each class, we propose the unlabeled class balancing strategy. We select training samples according to the multiple pseudo labels generated during the training iteration, instead of the usual data re-sampling that requires label information before training. Then the balanced re-weighting is leveraged to mitigate the class imbalance caused by multi-label co-occurrence. Extensive experiments conducted on two challenging benchmarks, AVA and UCF101-24, demonstrate the effectiveness of our proposed designs. By using the unlabeled data effectively, our method achieves the state-of-the-art performance in video action detection on both AVA and UCF101-24 datasets. Besides, it can still achieve competitive performance compared with fully-supervised methods when using limited annotations on AVA dataset. Hongcheng Zhang, Xu Zhao 0001, Dongqi Wang 0006 |
ACM Multimedia | 2 |
| 2021 | Human Carving: A Parsing-Based Framework For 3d Human ReconstructionabstractHuman-centric computer vision tasks often benefit from each other. In this paper, we propose a novel framework called Human Carving to explore the relationships between human parsing and multi-view 3D human reconstruction, which is the first method to consider the two related tasks. It consists of three modules: 1) Pose-aware Multi-view Human Parsing, 2) Semantic Visual Hull Carving and 3) Hierarchical Human Model Fitting. Taking the sparse multi-view images as input, the framework automatically generates a Part-Aware Visual Hull (PAVH) of human body parts and then estimates the human shape and pose simultaneously. Experimental results on real scenes demonstrate the effectiveness of our framework. Baoxing Li, Xu Zhao 0001 |
ICIP | 2 |
| 2021 | Rgb-D Fusion For Point-Cloud-Based 3d Human Pose Estimationabstract3D human pose estimation is an important and challenging task in computer vision. In this paper, we propose a method to estimate 3D human pose from RGB-D images. We adopt a 2D pose estimator to extract color features from the RGB image. The color features are integrated with the depth image in the form of point cloud. To fully exploit geometric information, we design a 3D learning module to extract point-wise features. To take advantage of local information as well as facilitate the convergence of the model, we design a dense prediction module. It estimates the offset vectors and closeness scores from points to target keypoints. The point-wise estimations are weighted and summed up to a final 3D pose. Experimental results show that our method achieves state-of-the-art performance on MHAD and SURREAL datasets. Jiaming Ying, Xu Zhao 0001 |
ICIP | 2 |
| 2021 | Joint Intention and Trajectory Prediction Based on TransformerabstractAlthough autonomous driving technology has made tremendous progress in recent years, it is still challenging to predict the intentions and trajectories of pedestrians. The state-of-the-art methods suffer from two problems. (1) Existing works consider these two tasks separately, ignoring the connection between them. (2) The selection and integration of inputs for these tasks are not well designed. In this paper, these two tasks are taken into consideration in a unified model. In this way, the information provided by the labels of each other is shared, improving the performance of both tasks. Besides, in addition to the bounding boxes and speeds, orientation and road semantic segmentation features are taken into consideration to show the potential intention and road context of the pedestrian. And all the inputs are weighted by an attention module before integration. Meanwhile, a Transformer encoder is applied in our method to extract the temporal information from the fused feature sequence. Our method outperforms all previous models for both trajectory prediction and intention prediction tasks on the JAAD dataset and PIE dataset. Ze Sui, Yue Zhou 0005, Xu Zhao 0001, Ao Chen 0003, Yiyang Ni 0004 |
IROS | 3 |
| 2021 | Transferable Knowledge-Based Multi-Granularity Fusion Network for Weakly Supervised Temporal Action DetectionabstractDespite remarkable progress, temporal action detection is still limited for real application due to the great amount of manual annotations. This issue motivates interest in addressing this task under weak supervision, namely, locating the action instances using only video-level class labels. Many current works on this task are mainly based on the Class Activation Sequence (CAS), which is generated by the video classification network to describe the probability of each snippet being in a specific action class of the video. However, the CAS generated by a simple classification network can only focus on local discriminative parts instead of locating the entire interval of target actions. In this paper, we present a novel framework to handle this issue. Specifically, we propose to utilize convolutional kernels with varied dilation rates to enlarge the receptive fields, which can transfer the discriminative information to the surrounding non-discriminative regions. Then, we design a cascaded module with the proposed Online Adversarial Erasing (OAE) mechanism to further mine more relevant regions of target actions by feeding the erased-feature maps of discovered regions back into the system. In addition, inspired by the transfer learning method, we adopt an additional module to transfer the knowledge from trimmed videos to untrimmed videos to promote the classification performance on untrimmed videos. Finally, we employ a boundary regression module embedded with Outer-Inner-Contrastive (OIC) loss to automatically predict the boundaries based on the enhanced CAS. Extensive experiments are conducted on two challenging datasets, THUMOS14 and ActivityNet-1.3, and the experimental results clearly demonstrate the superiority of our unified framework. Haisheng Su, Xu Zhao 0001, Shuming Liu 0001, Zhilan Hu |
IEEE Trans. Multim. | 2 |
| 2021 | CAT: Corner Aided Tracking With Deep Regression NetworkabstractSingle object tracking in visual media is an important yet challenging task. Various challenges, especially target scale variation, shape deformation and occlusion, can have large effects on the performances of trackers. Current deep regression based trackers only pay close attention to regression on the center key point of the tracking target, meanwhile employ the image pyramid based multi-scale testing method to deal with scale estimation. Such procedure can not properly handle the three challenges. We address these challenges in a principled way by the aid of auxiliary regressions on the four bounding box corners of the tracking target. In this work, we propose the novel Corner Aided Tracker with deep regression network, abbreviated as CAT. Different from RPN-based trackers, in CAT, four corners along with the center key point of the bounding box for tracking target are simultaneously obtained by five corresponding response maps. Furthermore, to robustly and accurately generate tight bounding boxes for the tracking target and collect reliable samples for online training of the network, we propose an adaptive key point selection method to select the subset of reliable key points and drop the unreliable ones, based on the qualities of their corresponding response maps as well as the constraints from shape, scale and location. We demonstrate that the regressed corners can help naturally locate the tracking target with tight bounding boxes. The challenges of scale variation, shape deformation and occlusion can be handled explicitly. The commonly used time-consuming image pyramid based multi-scale testing method can also be discarded. Extensive experiments on OTB2013, OTB2015, UAV123, LaSOT, VOT2016 and VOT2018 datasets are conducted to report new state-of-the-art performances and demonstrate the effectiveness of CAT. Shiquan Zhang, Xu Zhao 0001, Liangji Fang |
IEEE Trans. Multim. | 2 |
| 2020 | Anatomy and Geometry Constrained One-Stage Framework for 3D Human Pose Estimation
Xu Zhao 0001 |
ACCV (1) | 2 |
| 2020 | TSI: Temporal Scale Invariant Network for Action Proposal Generation
Shuming Liu 0001, Xu Zhao 0001, Haisheng Su, Zhilan Hu |
ACCV (5) | 2 |
| 2020 | EdgeStereo: An Effective Multi-task Learning Network for Stereo Matching and Edge Detection
Xiao Song 0002, Xu Zhao 0001, Liangji Fang, Hanwen Hu, Yizhou Yu |
Int. J. Comput. Vis. | 2 |
| 2020 | Joint Learning of Local and Global Context for Temporal Action Proposal GenerationabstractTemporal action proposal generation is an important yet challenging problem, since temporal proposals with rich action content are indispensable for analysing real-world videos with long duration and high proportion irrelevant content. This problem requires methods not only generating proposals with precise temporal boundaries, but also retrieving proposals to cover ground truth action instances with high recall and high overlap using relatively fewer proposals. To address these difficulties, we introduce an effective and efficient proposal generation method, named Local-Global Network (LGN), by which local and global contexts are jointly learned to generate high quality proposals. Locally, LGN first locates temporal boundaries with high starting and ending probabilities separately, then directly combines these boundaries as proposals. Globally, LGN evaluates the actionness probability of multiple-durations temporal regions simultaneously using temporal convolutional layers and anchor mechanism. Finally, we combine the boundary probabilities of each proposal with actionness probability of matched temporal regions as the confidence score, which is used for retrieving proposals. We conduct experiments on two datasets: ActivityNet-1.3 and THUMOS-14, where LGN outperforms other state-of-the-art methods with both high recall and high temporal precision. Finally, further experiments demonstrate that by combining existing action classifiers, our method significantly improves the state-of-the-art temporal action detection performance. Xu Zhao 0001, Haisheng Su |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Oriented Spatial Transformer Network for Pedestrian Detection Using Fish-Eye CameraabstractPedestrian detection using fish-eye cameras is a principal research focus in computer vision. Lack of pedestrian datasets of fish-eye images and pedestrian distortion in fish-eye images are two primary challenges. In this paper, two approaches are proposed to deal with these two challenges, respectively. On the one hand, the projective model transformation (PMT) algorithm is proposed, which can transform normal images into fish-eye images. The PMT can be applied to most of the pedestrian datasets and generates corresponding fish-eye image datasets. In this way, enough training data can be provided through the PMT. On the other hand, the oriented spatial transformer network (OSTN) is designed to rectify warped pedestrian features using CNNs, so that pedestrians in fish-eye images are easier for detectors to recognize. The OSTN can be embedded into universal deep learning based detectors easily. Moreover, the new pedestrian detector, where the OSTN is embedded, can be trained end to end. Finally, the OSTN based fish-eye pedestrian detectors can be trained using fish-eye images, which are generated using the PMT. Experiments on ETH, KITTI, Citypersons, and real pedestrian datasets show the effectiveness of the PMT and accuracy improvement of pedestrian detection in fish-eye images using the OSTN. Yeqiang Qian, Ming Yang 0002, Xu Zhao 0001, Bing Wang 0006 |
IEEE Trans. Multim. | 3 |
| 2019 | Semantic Segmentation of Street Scenes Using Disparity Information
Hanwen Hu, Xu Zhao 0001 |
ICIG (1) | 2 |
| 2019 | ISDNet: Importance Guided Semi-supervised Adversarial Learning for Medical Image Segmentation
Qingtian Ning, Xu Zhao 0001, Dahong Qian |
ICIG (2) | 2 |
| 2019 | 3D Body Pose and Shape Estimation from Multi-View Images With Limb Geometric ConstraintabstractThe estimation of 3D body pose and shape has always been a challenging problem due to various reasons, such as the ambiguity in 2D images and complex articulated structure of the human body. In order to solve the ill-conditioned problems, in this paper, we bring up an end-to-end method to estimate 3D human shape and pose from multi-view RGB images. In the proposed framework, we first implement a CNN embedded with attention module to extract the image feature and design the view-pooling layer to combine the features from multiple views. Then we adopt a regression network with a novel geometric constraint of body limbs to estimate 3D human pose and shape. Additionally, during the training process, we employ the idea of adversarial learning in our model to help regress accurate pose and shape parameters. Extensive experiments are conducted on Human3.6M and MPI-INF-3DHP datasets, and our method achieves competitive results in the 3D pose and shape estimation task. Zixuan Gai, Xu Zhao 0001 |
ICIP | 2 |
| 2019 | Temporal Regularized Spatial Attention for Video-Based Person Re-IdentificationabstractVideo-based person re-identification aims at matching video sequences of a person across different camera views. How to explore the abundant appearance and motion information in a video sequence is crucial to tackle this problem. To this end, we first introduce a parameter-free spatial attention module to emphasize the importance of discriminative regions. Then we apply a temporal regularization term on spatial attention to refine corrupted region caused by occlusion and blur. This term allows the attention response at one position in a frame to be related to other frames of the same position. Extensive experiments are conducted on iLIDS-VID and PRID-2011 datasets. The experimental results demonstrate that our approach surpasses the existing state-of-the-art video-based person re-identification methods on iLIDS-VID and PRID-2011. Xu Zhao 0001 |
ICIP | 2 |
| 2019 | Small-objectness sensitive detection based on shifted single shot detector
Liangji Fang, Xu Zhao 0001, Shiquan Zhang |
Multim. Tools Appl. | 2 |
| 2019 | Discriminative representation combinations for accurate face spoofing detection
Xiao Song 0002, Xu Zhao 0001, Liangji Fang |
Pattern Recognit. | 2 |
| 2019 | Attention-Based Multiview Re-Observation Fusion Network for Skeletal Action RecognitionabstractAction recognition is an important and popular area in computer vision. Because of the helpfulness of action recognition of the skeleton and the development of related pose estimation techniques, action recognition based on skeleton data has drawn considerable attention and has been widely studied in recent years. In this paper, we propose an attention-based multiview re-observation fusion model for skeletal action recognition. The proposed model focuses on the factor of observation view of actions, which greatly influences action recognition. The model utilizes action information from multiple observation views to improve the recognition performance. In this method, we re-observe input skeleton data from several possible viewpoints, process these augmented observation data with a long short-term memory (LSTM) network separately, and, finally, fuse the outputs to generate the final recognition result. In the multiview fusion process, an attention mechanism is applied to regulate the fusion operation according to the helpfulness for the recognition of all views. In this way, the model can fuse information from multiple viewpoints to recognize actions and can learn to evaluate observation views to improve fusion performance. We also propose a multilayer feature attention method to improve the performance of the LSTM in our model. We utilize an attention mechanism to enhance the feature expression by finding and focusing on informative feature dimensions according to contextual action information. Moreover, we propose stacking multiple layers of attention operation in a multilayer LSTM network to further improve network performance. The final model is integrated into an end-to-end trainable network. Experiments conducted on two popular datasets, NTU RGB+D and SBU Kinect interaction, show that our model achieves state-of-the-art performance. Zhaoxuan Fan, Xu Zhao 0001, Haisheng Su |
IEEE Trans. Multim. | 2 |
| 2018 | Putting the Anchors Efficiently: Geometric Constrained Pedestrian Detection
Liangji Fang, Xu Zhao 0001, Xiao Song 0002, Shiquan Zhang, Ming Yang 0002 |
ACCV (5) | 2 |
| 2018 | Simultaneous Face Detection and Head Pose Estimation: A Fast and Unified Framework
Tingfeng Li, Xu Zhao 0001 |
ACCV (1) | 2 |
| 2018 | EdgeStereo: A Context Integrated Residual Pyramid Network for Stereo Matching
Xiao Song 0002, Xu Zhao 0001, Hanwen Hu, Liangji Fang |
ACCV (5) | 2 |
| 2018 | Cascaded Pyramid Mining Network for Weakly Supervised Temporal Action Localization
Haisheng Su, Xu Zhao 0001 |
ACCV (2) | 2 |
| 2018 | BSN: Boundary Sensitive Network for Temporal Action Proposal Generation
Xu Zhao 0001, Haisheng Su, Chongjing Wang, Ming Yang 0002 |
ECCV (4) | 2 |
| 2018 | Led: Localization-Quality Estimation Embedded DetectorabstractClassification subnetwork and box regression subnetwork are essential components in deep networks for object detection. However, we observe a contradiction that before NMS, some better localized detections do not correspond to higher classification confidences, and vice versa. This contradiction exists because classification confidences can not fully reflect the localization-quality (loc-quality) of each detection. In this work, we propose the Localization-quality Estimation embedded Detector abbreviated as LED, and a corresponding detection pipeline. In this detection pipeline, we first propose an accurate loc-quality estimation method for each detection, then combine the loc-quality with the corresponding classification confidence during inference to make each detection more reasonable and accurate. For efficiency, LED is designed as an one-stage network. Extensive experiments are conducted on Pascal VOC 2007 and KITTI car detection datasets to demonstrate the effectiveness of LED. Shiquan Zhang, Xu Zhao 0001, Liangji Fang, Haiping Fei |
ICIP | 2 |
| 2018 | Weakly Supervised Temporal Action Detection with Shot-Based Temporal Pooling Network
Haisheng Su, Xu Zhao 0001, Haiping Fei |
ICONIP (4) | 2 |
| 2018 | Measuring Crowd Collectiveness by Macroscopic and Microscopic Motion ConsistenciesabstractAs a scene-independent descriptor of crowd motions, crowd collectiveness quantifies the degree of constituent individuals moving as a union in a crowd scene. An effective measurement on crowd collectiveness is of great importance for applications in surveillance of public safety, human dynamics, and other areas. To this end, we propose a novel framework to measure crowd collectiveness by combining macroscopic and microscopic motion consistencies and define quantitatively the global and local consistency of crowd motions. The defined global consistency represents the likelihood of pairwise individuals belonging to the same collective group, whereas the local consistency reflects the degree of conformity in a local region. Based on the proposed collectiveness measure, a new algorithm, named group mining, is proposed to detect collective groups from a crowd. We validate the effectiveness of the proposed method on several synthetic particle systems and a real-world crowd database with human-labeled collectiveness. Experimental results show that, compared with the previous approaches, our collectiveness measure is more consistent with human perception, and the collective groups detected by our group mining algorithm are more accurate and robust. Xu Zhao 0001, Yuncai Liu |
IEEE Trans. Multim. | 2 |
| 2017 | An Online Approach for Gesture Recognition Toward Real-World Applications
Zhaoxuan Fan, Xu Zhao 0001, Wanli Jiang, Ming Yang 0002 |
ICIG (1) | 3 |
| 2017 | Cost efficient subcategory-aware CNN for object detectionabstractIn this paper, we propose an accurate and cost efficient deep CNN network for object detection. In contrast to the previous region-based methods like Sub-CNN [1], our detector is almost fully convolutional so that the computation cost can be reduced efficiently. By introducing position-sensitive score maps and exploiting subcategory information, our method is less time consuming while maintaining competitive performance on detecting objects with various scales. In addition, we remove image pyramid used in Sub-CNN to achieve further acceleration. The experimental results show that our approach is 1.3 times faster than Sub-CNN with only 14% number of parameters and archives comparable results on the challenging KITTI dataset. Compared with the state-of-the-art methods for object detection, our approach provides an efficient solution that takes into account both accuracy and speed. Tingfeng Li, Xu Zhao 0001 |
ICIP | 2 |
| 2017 | Temporal action localization with two-stream segment-based RNNabstractTemporal Action localization is a more challenging vision task than action recognition because videos to be analyzed are usually untrimmed and contain multiple action instances. In this paper, we investigate the potential of recurrent neural network, toward three critical aspects for solving this problem, namely, high-performance feature, high-quality temporal segments and effective recurrent neural network architecture. First of all, we introduce the two-stream (spatial and temporal) network for feature extraction. Then, we propose a novel temporal selective search method to generate temporal segments with variable lengths. Finally, we design a two-branch LSTM architecture for category prediction and confidence score computation. Our proposed approach to action localization, along with the key components, say, segments generation and classification architecture, are evaluated on the THUMOS'14 dataset and achieve promising performance by comparing with other state-of-the-art methods. Xu Zhao 0001, Zhaoxuan Fan |
ICIP | 2 |
| 2017 | Face spoofing detection by fusing binocular depth and spatial pyramid coding micro-texture featuresabstractRobust features are of vital importance to face spoofing detection, because various situations make feature space extremely complicated to partition. Thus in this paper, two novel and robust features for anti-spoofing are proposed. The first one is a binocular camera based depth feature called Template Face Matched Binocular Depth (TFBD) feature. The second one is a high-level micro-texture based feature called Spatial Pyramid Coding Micro-Texture (SPMT) feature. Novel template face registration algorithm and spatial pyramid coding algorithm are also introduced along with the two novel features. Multi-modal face spoofing detection is implemented based on these two robust features. Experiments are conducted on a widely used dataset and a comprehensive dataset constructed by ourselves. The results reveal that face spoofing detection with the fusion of our proposed features is of strong robustness and time efficiency, meanwhile outperforming other state-of-the-art traditional methods. Xiao Song 0002, Xu Zhao 0001 |
ICIP | 2 |
| 2017 | Single Shot Temporal Action DetectionabstractTemporal action detection is a very important yet challenging problem, since videos in real applications are usually long, untrimmed and contain multiple action instances. This problem requires not only recognizing action categories but also detecting start time and end time of each action instance. Many state-of-the-art methods adopt the "detection by classification" framework: first do proposal, and then classify proposals. The main drawback of this framework is that the boundaries of action instance proposals have been fixed during the classification step. To address this issue, we propose a novel Single Shot Action Detector (SSAD) network based on 1D temporal convolutional layers to skip the proposal generation step via directly detecting action instances in untrimmed video. On pursuit of designing a particular SSAD network that can work effectively for temporal action detection, we empirically search for the best network architecture of SSAD due to lacking existing models that can be directly adopted. Moreover, we investigate into input feature types and fusion strategies to further improve detection accuracy. We conduct extensive experiments on two challenging datasets: THUMOS 2014 and MEXaction2. When setting Intersection-over-Union threshold to 0.5 during evaluation, SSAD significantly outperforms other state-of-the-art systems by increasing mAP from $19.0%$ to $24.6%$ on THUMOS 2014 and from 7.4% to $11.0%$ on MEXaction2. Xu Zhao 0001 |
ACM Multimedia | 2 |
| 2017 | Online learning of dynamic multi-view gallery for person Re-identification
Yanna Zhao, Xu Zhao 0001, Zong Jie Xiang, Yuncai Liu |
Multim. Tools Appl. | 2 |
| 2017 | Context-Associative Hierarchical Memory Model for Human Activity Recognition and PredictionabstractHuman activity recognition is a challenging high-level vision task, for which multiple factors, such as subject, object, and their diverse interactions, have to be considered and modeled. Current learning-based methods are limited in the capability to integrate human-level concepts into an easily extensible computational framework. Inspired by the existing human memory model, we present a context-associative approach to recognize activity with human-object interaction. The proposed system can recognize incoming visual content based on the previous experienced activities. The high-level activity is parsed into consecutive subactivities, and we build a context cluster to model the temporal relations. The semantic attributes of the subactivity are organized by a concept hierarchy. Based on the hierarchy, a series of similarity functions are defined to turn the recognition computing into retrievals over the contextual memory, similar to the auto-associative characteristics of human memory. Partially matching in retrieval and stored memory make the activity prediction possible. The dynamical evolution of the brain memory is mimicked to allow decay and reinforcement of the input information, providing a natural way to maintain data and save computational time. We evaluate our approach on three data sets: CAD-120, MHOI, and OPPORTUNITY. The proposed method demonstrates promising results compared with other state-of-the-art techniques. Lei Wang 0060, Xu Zhao 0001, Yunfei Si, Liangliang Cao, Yuncai Liu |
IEEE Trans. Multim. | 2 |
| 2016 | Parallelized deformable part models with effective hypothesis pruningabstractAs a typical machine-learning based detection technique, deformable part models (DPM) achieve great success in detecting complex object categories. The heavy computational burden of DPM, however, severely restricts their utilization in many real world applications. In this work, we accelerate DPM via parallelization and hypothesis pruning. Firstly, we implement the original DPM approach on a GPU platform and parallelize it, making it 136 times faster than DPM release 5 without loss of detection accuracy. Furthermore, we use a mixture root template as a prefilter for hypothesis pruning, and achieve more than 200 times speedup over DPM release 5, apparently the fastest implementation of DPM yet. The performance of our method has been validated on the Pascal VOC 2007 and INRIA pedestrian datasets, and compared to other state-of-the-art techniques. Xu Zhao 0001 |
Comput. Vis. Media | 2 |
| 2016 | Reduce false positives for object detection by a priori probability in videos
Lei Wang 0060, Xu Zhao 0001, Yuncai Liu |
Neurocomputing | 2 |
| 2016 | Person Re-identification by encoding free energy feature maps
Yanna Zhao, Xu Zhao 0001, Ruotian Luo, Yuncai Liu |
Multim. Tools Appl. | 2 |
| 2015 | Adaptive appearance learning for human pose estimationabstractWe address the problem of pose estimation in videos. The part detectors play important roles, but traditional template-based detectors (e.g. Histogram of Gradient, HoG) fail at pose estimation due to the high variability in appearance. We present an adaptive representation of appearance and shape for articulated human body. The full representation of human body is based on the flexible mixture-of-parts model. We train a Naive Bayes classifier to obtain a confidence score of estimated pose by the basic mixture model, and based on the confidence we learn an instance-specific appearance model. For between-frame consistency, we design a time-efficient energy function for motion cues instead of complex motion models. We incorporate these models into a framework that allows for efficient inference. Quantitative evaluation of pose estimation conducted on two video datasets demonstrates the effectiveness of the proposed method. Lei Wang 0060, Xu Zhao 0001, Yuncai Liu |
ICIP | 2 |
| 2015 | A discriminative tracklets representation for crowd analysisabstractIn this work, we propose a discriminative tracklets representation for motion pattern extraction from crowded scene. The representation incorporates relative position, velocity, and direction information of tracklet into one compact form by shaping it within a rectangle. We adopt deep belief networks to extract low-dimensional features from this representation. It not only reduces the computational complexity for the following clustering, but also achieves more discriminative tracklets representation which is invariant to noises brought by tracking failures. To determine the spatio-temporal distribution of each motion pattern, a robust clustering scheme composed of three clustering procedures is proposed. Comprehensive experiments in multiple datasets validate the effectiveness of our approach. Chongjing Wang, Xu Zhao 0001, Yuncai Liu |
ICIP | 2 |
| 2015 | Detect coherent motions in crowd scenes based on tracklets associationabstractCoherent motion is a very common motion pattern in crowded scenes. Coherent Filter is a very effective and robust tool to detect coherent motions based on point trajectories, the performance of coherent filter depends on point trajectories' property. In this work, we present a two-stage strategy to extract dense, accurate and long-term point trajectories from crowded scenes. The method includes a tracklets acquisition procedure and a tracklets association procedure. We use LDOF tracker to acquire dense tracklets, and then formulate tracklets association as a linear assignment problem (LAP). Experiments conducted on challenging crowd datasets show that our trajectories are very suitable for detecting coherent motions in crowded scenes. Xu Zhao 0001, Yuncai Liu |
ICIP | 2 |
| 2014 | Person re-identification by free energy score space encodingabstractPerson re-identification is an important and challenging computer vision problem. Recent progress in this area is due to new visual features and models that deals with cross-view variations. Instead of working towards more complex models, we focus on low level features and their encoding. Low level features capturing the color and structural information are first extracted from human images. Gaussian Mixture Model (GMM) is then employed to approximate the distribution of the features, providing a relatively comprehensive statistical representation. Finally, low level features are mapped to a space by computing free energy score of the GMM. The mapped features are encoded into a fixed-length feature vector for person re-identification. Extensive experiments are conducted on several public datasets. Comparisons with benchmark person re-identification methods show the promising performance of our approach. Yanna Zhao, Xu Zhao 0001, Yuncai Liu |
ICIP | 2 |
| 2014 | Statistical background subtraction based on imbalanced learningabstractIn this paper, we study the class imbalance problem in statistical background subtraction. Firstly, we discuss the imbalance essence in background subtraction, and conclude that foreground and background are inherently imbalanced. Secondly, following the imbalanced learning strategy in machine learning, we present a spatio-temporal over-sampling method to resolve the class imbalance in background subtraction. Our method densely generate synthesized foreground samples in compact 3D spatio-temporal domain. Those generated samples could reduce the imbalance level between foreground and background from both quantity and quality, and therefore contribute to improvement of detection performance. We also define a new index to measure the change of imbalance level during over-sampling. Experiments are conducted on public datasets to demonstrate the effectiveness of our method. Xiang Zhang 0006, Xu Zhao 0001 |
ICME | 4 |
| 2013 | Motion pattern analysis in crowded scenes based on hybrid generative-discriminative feature mapsabstractCrowded scene analysis is becoming increasingly popular in computer vision field. In this paper, we propose a novel approach to analyze motion patterns by clustering the hybrid generative-discriminative feature maps using unsupervised hierarchical clustering algorithm. The hybrid generative-discriminative feature maps are derived by posterior divergence based on the tracklets which are captured by tracking dense points with three effective rules. The feature maps effectively associate low-level features with the semantical motion patterns by exploiting the hidden information in crowded scenes. Motion pattern analyzing is implemented in a completely unsupervised way and the feature maps are clustered automatically through hierarchical clustering algorithm building on the basis of graphic model. The experiment results precisely reveal the distributions of motion patterns in current crowded videos and demonstrate the effectiveness of our approach. Chongjing Wang, Xu Zhao 0001, Yuncai Liu |
ICIP | 2 |
| 2013 | Exploring discriminative pose sub-patterns for effective action classificationabstractArticulated configuration of human body parts is an essential representation of human motion, therefore is well suited for classifying human actions. In this work, we propose a novel approach to exploring the discriminative pose sub-patterns for effective action classification. These pose sub-patterns are extracted from a predefined set of 3D poses represented by hierarchical motion angles. The basic idea is motivated by the two observations: (1) There exist representative sub-patterns in each action class, from which the action class can be easily differentiated. (2) These sub-patterns frequently appear in the action class. By constructing a connection between frequent sub-patterns and the discriminative measure, we develop the SSPI, namely, the Support Sub-Pattern Induced learning algorithm for simultaneous feature selection and feature learning. Based on the algorithm, discriminative pose sub-patterns can be identified and used as a series of "magnetic centers" on the surface of normalized super-sphere for feature transform. The "attractive forces" from the sub-patterns determine the direction and step-length of the transform. This transformation makes a feature more discriminative while maintaining dimensionality invariance. Comprehensive experimental studies conducted on a large scale motion capture dataset demonstrate the effectiveness of the proposed approach for action classification and the superior performance over the state-of-the-art techniques. Xu Zhao 0001, Yuncai Liu, Yun Fu 0001 |
ACM Multimedia | 1 |
| 2013 | Multiple Subcategories Parts-Based Representation for One Sample Face IdentificationabstractSmall sample set, occlusion, and illumination variations are the critical obstacles for a face identification system towards practical application. In this paper, we propose a probabilistic generative model for parts-based data representation to address these difficulties. In our approach, multiple subcategories corresponding to the individual face parts, such as nose, mouth, eye, and so forth, are modeled within a probabilistic graphical model framework to mimic the process of generating a face image. The induced face representation, therefore, encodes rich discriminative information. Model training is totally unsupervised. Once the training is completed, a test sample from the face class can be recognized as a novel combination of learned parts. In summary, the main contributions of this work are threefold: 1) A novel hierarchical probabilistic generative model is proposed, which is capable of achieving an efficient parts-based representation for robust face identification. 2) A constrained variational EM algorithm is developed to learn the model parameters and infer the variables. 3) Two similarity metrics are specially designed for the novel parts-based feature representation, which are effective for matching score guided one sample face identification. The models and similarity metrics are validated on three face databases. Experimental results demonstrate the capabilities of the model to deal with small sample set, occlusions, and illumination variances. Xu Zhao 0001, Xiong Li 0004, Yun Fu 0001, Yuncai Liu |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2012 | Parallelized Annealed Particle Filter for real-time marker-less motion tracking via heterogeneous computing
Yatao Bian, Xu Zhao 0001, Yuncai Liu |
ICPR | 2 |
| 2012 | Hybrid generative-discriminative recognition of human action in 3D joint spaceabstractWe propose a novel human action recognition method based on the generative feature mapping over 3D human body joint sequences. The proposed method relies on Hidden Markov Model (HMM), but differs from the previous methods in the way of incorporating HMM and discriminative classifier, aiming to capture more discriminative information. Firstly, we use HMMs to model the joint sequences of human body. Then the Posterior Divergence is used to build feature mappings from the trained HMMs. The derived feature mappings map a variable-length joint sequence to a fixed-dimension feature vector which will be delivered to SVM for classification. We evaluate the proposed method and related methods on a large number of 3D joint sequences. The experimental results show its competitive performance, in comparison with other state-of-the-art methods. Xiong Li 0004, Xu Zhao 0001, Yuncai Liu |
ACM Multimedia | 3 |
| 2011 | Detecting Motion Patterns in Dynamic Crowd ScenesabstractDetecting motion pattern in dynamic crowd scenes is a challenging problem in computer vision field. In this paper, we propose a novel approach to detect the motion patterns from global perspective. To extract the discriminative spatial-temporal features, we introduce the Motion History Image (MHI) into the optical flow algorithm. Motion patterns are then detected by automatic clustering of optical flow vectors through hierarchical clustering. Experiment evaluation on some challenging videos shows reliable detection results and demonstrates the effectiveness of our proposed approach. Chongjing Wang, Xu Zhao 0001, Yuncai Liu |
ICIG | 2 |
| 2011 | A Method for Detection and Classification of Glass Defects in Low Resolution ImagesabstractThis paper presents a novel method for detection and recognition of glass defects in low resolution images. First, the defect region is located by the method of Canny edge detection, and thus the smallest connected region (rectangle) can be found. Then, the binary information of the core region can be obtained based on a specific filter. After noises are removed, a novel Binary Feature Histogram (BFH) is proposed to describe the characteristic of the glass defect. Finally, the AdaBoost method is adopted for classification. The classifiers are designed based on BFH. Experiments with 800 bubble images and 240 non-bubble images prove that the proposed method is effective and efficient for recognition of glass defects, such as bubbles and inclusions. Qing-Jie Kong, Xu Zhao 0001, Yuncai Liu |
ICIG | 3 |
| 2011 | Human Motion Tracking by Temporal-Spatial Local Gaussian Process ExpertsabstractHuman pose estimation via motion tracking systems can be considered as a regression problem within a discriminative framework. It is always a challenging task to model the mapping from observation space to state space because of the high-dimensional characteristic in the multimodal conditional distribution. In order to build the mapping, existing techniques usually involve a large set of training samples in the learning process which are limited in their capability to deal with multimodality. We propose, in this work, a novel online sparse Gaussian Process (GP) regression model to recover 3-D human motion in monocular videos. Particularly, we investigate the fact that for a given test input, its output is mainly determined by the training samples potentially residing in its local neighborhood and defined in the unified input-output space. This leads to a local mixture GP experts system composed of different local GP experts, each of which dominates a mapping behavior with the specific covariance function adapting to a local region. To handle the multimodality, we combine both temporal and spatial information therefore to obtain two categories of local experts. The temporal and spatial experts are integrated into a seamless hybrid system, which is automatically self-initialized and robust for visual tracking of nonlinear human motion. Learning and inference are extremely efficient as all the local experts are defined online within very small neighborhoods. Extensive experiments on two real-world databases, HumanEva and PEAR, demonstrate the effectiveness of our proposed model, which significantly improve the performance of existing models. Xu Zhao 0001, Yun Fu 0001, Yuncai Liu |
IEEE Trans. Image Process. | 1 |
| 2011 | Text From Corners: A Novel Approach to Detect Text and Caption in VideosabstractDetecting text and caption from videos is important and in great demand for video retrieval, annotation, indexing, and content analysis. In this paper, we present a corner based approach to detect text and caption from videos. This approach is inspired by the observation that there exist dense and orderly presences of corner points in characters, especially in text and caption. We use several discriminative features to describe the text regions formed by the corner points. The usage of these features is in a flexible manner, thus, can be adapted to different applications. Language independence is an important advantage of the proposed method. Moreover, based upon the text features, we further develop a novel algorithm to detect moving captions in videos. In the algorithm, the motion features, extracted by optical flow, are combined with text features to detect the moving caption patterns. The decision tree is adopted to learn the classification criteria. Experiments conducted on a large volume of real video shots demonstrate the efficiency and robustness of our proposed approaches and the real-world system. Our text and caption detection system was recently highlighted in a worldwide multimedia retrieval competition, Star Challenge, by achieving the superior performance with the top ranking. Xu Zhao 0001, Kai-Hsiang Lin, Yun Fu 0001, Yuxiao Hu 0001, Yuncai Liu, Thomas S. Huang |
IEEE Trans. Image Process. | 1 |
| 2010 | Sparse Coding on Local Spatial-Temporal Volumes for Human Action Recognition
Yan Zhu 0009, Xu Zhao 0001, Yun Fu 0001, Yuncai Liu |
ACCV (2) | 2 |
| 2010 | Bimodal gender recognition from face and fingerprintabstractThis paper focuses on multimodal gender recognition. To achieve a robust and discriminative performance for gender recognition, visual observations from both face and corresponding fingerprints are fused to serve for the task. The bag-of-words model is employed to structure the image representation. We propose a novel supervised method to construct the visual words, by which the redundant feature dimensions are discarded and the important dimensions for gender classification are highlighted. The dimension rearrangement is achieved by aligning the feature dimensions to a common normal vector of the hyperplane between categories. The Latent Dirichlet Allocation (LDA) model is extended to incorporate discriminative clues for supervised classification. We build the novel Discriminative LDA (D-LDA) model by maximizing the inter-class margins, which can significantly enhance the discriminative power of the whole model. Experiments on a large face and fingerprint database demonstrate the effectiveness of the proposed new feature and model. Complementary advantages benefited from face-fingerprint fusion to a robust gender recognition framework also get validated. Xiong Li 0004, Xu Zhao 0001, Yun Fu 0001, Yuncai Liu |
CVPR | 2 |
| 2010 | Multimodality gender estimation using Bayesian hierarchical modelabstractWe propose to estimate human gender from corresponding fingerprint and face information with the Bayesian hierarchical model. Different from previous works on fingerprint based gender estimation with specially designed features, our method extends to use general local image features. Furthermore, a novel word representation called latent word is designed to work with the Bayesian hierarchical model. The feature representation is embedded to our multimodality model, within which the information from fingerprint and face is fused at the decision level for gender estimation. Experiments on our internal database show the promising performance. Xiong Li 0004, Xu Zhao 0001, Huanxi Liu, Yun Fu 0001, Yuncai Liu |
ICASSP | 2 |
| 2010 | Weak Metric Learning for Feature Fusion towards Perception-Inspired Object Recognition
Xiong Li 0004, Xu Zhao 0001, Yun Fu 0001, Yuncai Liu |
MMM | 2 |
| 2010 | Human Pose Regression Through Multiview Visual FusionabstractWe consider the problem of estimating 3-D human body pose from visual signals within a discriminative framework. It is challenging because there is a wide gap between complex 3-D human motion and planar visual observation, which makes this a severely ill-conditioned problem. In this paper, we focus on three critical factors to tackle human body pose estimation, namely, feature extraction, learning algorithm, and camera utilization. On the feature level, we describe images using the salient interest points represented by scale-invariant feature transform (SIFT)-like descriptors, in which the position, appearance, and local structural information are encoded simultaneously. On the learning algorithm level, we propose to use Gaussian processes and multiple linear (ML) regression to model the mapping between poses and features. Fusing image information from multiple cameras in different views is of great interest to us on the camera level. We make a comprehensive evaluation on the HumanEva database and get two meaningful insights into the three crucial aspects for human pose estimation: 1) although the choice of feature is very important to the problem, once the learning algorithm becomes efficient, the choice of feature is no longer critical, and 2) the impact of information combination from multiple cameras on pose estimation is closely related to not only the quantity of image information, but also its quality. In most cases, it is true that the more information is involved, the better results can be achieved. But when the information quantity is the same, the differences in quality will lead to totally different performance. Furthermore, dense evaluations demonstrate that our approach is an accurate and robust solution to the human body pose estimation problem. Xu Zhao 0001, Yun Fu 0001, Huazhong Ning, Yuncai Liu, Thomas S. Huang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2009 | Temporal-Spatial Local Gaussian Process Experts for Human Pose Estimation
Xu Zhao 0001, Yun Fu 0001, Yuncai Liu |
ACCV (1) | 1 |
| 2008 | Discriminative estimation of 3D human pose using Gaussian processesabstractIn this paper, we present an efficient discriminative method for human pose estimation. This method learns a direct mapping from visual observations to human body configurations. The framework requires that the visual features should be powerful enough to discriminate the subtle differences between similar human poses. We propose to describe the image features using salient interest points that are represented by SIFT-like descriptors. The descriptor encode the position, appearance, and local structural information simultaneously. Bag-of-words representation is used to model the distribution of feature space. The descriptor can tolerate a range of illumination and position variations because it is computed on overlapped patches. We use Gaussian process regression to model the mapping from visual observations to human poses. This probabilistic regression algorithm is effective and robust to the pose estimation problem. We test our approach on the HumanEva data set. Experimental results demonstrate that our approach achieves the state of the art performance. Xu Zhao 0001, Huazhong Ning, Yuncai Liu, Thomas S. Huang |
ICPR | 1 |
| 2008 | Generative tracking of 3D human motion by hierarchical annealed genetic algorithm
Xu Zhao 0001, Yuncai Liu |
Pattern Recognit. | 1 |
| 2007 | Generative Estimation of 3D Human Pose Using Shape Contexts Matching
Xu Zhao 0001, Yuncai Liu |
ACCV (1) | 1 |
| 2007 | Tracking 3D Human Motion in Compact Base SpaceabstractIn this study, we present an efficient approach to recover 3D human motion from monocular image sequences in generative reconstruction framework. This approach is based on the extracting of motion base space. From the motion capture data with bothersome high dimension characteristic of human activity, we extract the motion base space in which human pose can be described essentially and concisely by a more controllable way. And then, the structure of this space corresponding to some special activities such as walking motion is explored with data clustering. For the single image, Gaussian mixture model is used to generate the candidates of 3D pose. The shape context is the common descriptor of image silhouette feature and synthetical feature of human model. We get the shortlist of 3D poses by measuring the shape contexts matching cost between image feature and the synthetical features. In tracking situation, an AR model trained by the example sequences produces almost accurate pose predictions. Experiments demonstrate that the proposed approach works well Xu Zhao 0001, Yuncai Liu |
WACV | 1 |