Zhenhua Wang 0003

dblp:214/9322 · DBLP profile ↗
← Back
36ranked-venue papers
7as first author
22since 2021 · last 2026
0000-0001-5473-1499ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 5 first-author · 10 since 2021Systems, architecture and hardware · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
YearPublicationVenuePosition
2026 LTSTrack: Visual tracking with long-term temporal sequence
Zhaochuan Zeng, Shilei Wang 0001, Yidong Song, Zhenhua Wang 0003, Jifeng Ning
Pattern Recognit.4
2026 Exploring Pruning-Based Efficient Object Tracking via Hybrid Knowledge Distillation
Yidong Song, Shilei Wang 0001, Zhaochuan Zeng, Jikai Zheng, Zhenhua Wang 0003, Jifeng Ning
IEEE Trans. Circuits Syst. Video Technol.5
2025 Graph Attention Network for Context-Aware Visual Tracking
abstract
Siamese-network-based trackers convert the general object tracking as a similarity matching task between a template and a search region. Using convolutional feature cross correlation (Xcorr) for similarity matching, a large number of Siamese trackers are proposed and achieved great success. However, due to the predefined size of the target feature, these trackers suffer from either retaining much background information or losing important foreground information. Moreover, the global matching between the target and search region also largely neglects the part-level structural information and the contextual information of the target. To tackle the aforementioned obstacles, in this article, we propose a simple context-aware Siamese graph attention network, which establishes part-to-part correspondence between the Siamese branches with a complete bipartite graph. The object information from the template is propagated to the search region via a graph attention mechanism. With such a design, a target-aware template input is enabled to replace the prefixed template region, which can adaptively fit the size and aspect ratio variations in different objects. Based on it, we further construct a context-aware feature matching mechanism to embed both the target and the contextual information in the search region. Experiments on challenging benchmarks including GOT-10k, TrackingNet, LaSOT, VOT2020, and OTB-100 demonstrate that the proposed SiamGAT* outperforms many state-of-the-art trackers and achieves leading performance. Code is available at: https://git.io/SiamGAT.
Yanyan Shao, Dongyan Guo, Zhenhua Wang 0003, Liyan Zhang 0001, Jianhua Zhang 0002
IEEE Trans. Neural Networks Learn. Syst.4
2024 Edge-Reserved Knowledge Distillation for Image Matting
abstract
Deep learning-based methods have made significant progress in natural image matting. However, mainstream approaches like CNN and ViT mainly focus on capturing the global features of images, but they lack specialized treatment for edges. They struggle to accurately distinguish between foregrounds and backgrounds that have similar colors and textures, which results in blurred edge areas between the foreground and the background. To solve the above problems, we propose Edge-reserved Knowledge Distillation Model (ERKD), which can reserving good edge features while distill multimodal semantic information from large models. In order to acquire multi-scale edge features, we design the Edge-Reserved Module and the Multimodal Feature Fusion Module. At the same time, to enhance the capture of edge features, we introduce CLIP for feature-level knowledge distillation operations. The comprehensive evaluation on Composition-1k and Distinction-646 datasets shows that the performance of this method surpasses existing techniques.
Zhenhua Wang 0003, Jifeng Ning
ICIP2
2024 Object discriminability re-extraction for distractor-aware visual object tracking
Dongyan Guo, Xiangjie Kong 0001, Zhenhua Wang 0003, Jianhua Zhang 0002
Comput. Vis. Image Underst.5
2024 Bidirectional Interaction of CNN and Transformer Feature for Visual Tracking
abstract
Empowered by the sophisticated long-range dependency modeling ability of Transformer, tracking performance has seen a dynamic increase in recent years. Approaches in this vein leverage the Transformer feature to integrate the information of target and search regions while neglecting the superior local representation extracted by their CNN backbone. To address this, we introduce a BIdirectional inTeraction mechanism between CNN and Transformer features for visual tracking, termed BIT-Tracker, which admits a comprehensive fusion of local and global representations, and thus boosts tracking performance. The first ingredient of BIT-Tracker is an aggregation of multi-level Transformer features to achieve a better global modeling ability. In order to combine the merits of both local and global representations, our second ingredient performs a bi-directional interaction between CNN and Transformer features, where the interaction is achieved via either querying the CNN feature from the Transformer feature or querying the Transformer feature from the CNN feature. Afterwards, the outputs from both directions are fused to predict the temporal locations of targets. Extensive experiments demonstrate the effectiveness of the proposed feature aggregation and bi-directional interaction modules. Impressively, BIT-Tracker achieves leading performance on eight tracking benchmarks and outperforms SOTA results by salient margins. Code will be made available.
Baozhen Sun, Zhenhua Wang 0003, Shilei Wang 0001, Yongkang Cheng, Jifeng Ning
IEEE Trans. Circuits Syst. Video Technol.2
2024 Modeling of Multiple Spatial-Temporal Relations for Robust Visual Object Tracking
abstract
Recently, one-stream trackers have achieved parallel feature extraction and relation modeling through the exploitation of Transformer-based architectures. This design greatly improves the performance of trackers. However, as one-stream trackers often overlook crucial tracking cues beyond the template, they prone to give unsatisfactory results against complex tracking scenarios. To tackle these challenges, we propose a multi-cue single-stream tracker, dubbed MCTrack here, which seamlessly integrates template information, historical trajectory, historical frame, and the search region for synchronized feature extraction and relation modeling. To achieve this, we employ two types of encoders to convert the template, historical frames, search region, and historical trajectory into tokens, which are then collectively fed into a Transformer architecture. To distill temporal and spatial cues, we introduce a novel adaptive update mechanism, which incorporates a thresholding component and a local multi-peak component to filter out less accurate and overly disturbed tracking cues. Empirically, MCTrack achieves leading performance on mainstream benchmark datasets, surpassing the most advanced SeqTrack by 2.0% in terms of the AO metric on GOT-10k. The code is available at https://github.com/wsumel/MCTrack.
Shilei Wang 0001, Zhenhua Wang 0003, Qianqian Sun, Gong Cheng 0003, Jifeng Ning
IEEE Trans. Image Process.2
2024 Self-Supervised Enhancement for Named Entity Disambiguation via Multimodal Graph Convolution
abstract
Named entity disambiguation (NED) finds the specific meaning of an entity mention in a particular context and links it to a target entity. With the emergence of multimedia, the modalities of content on the Internet have become more diverse, which poses difficulties for traditional NED, and the vast amounts of information make it impossible to manually label every kind of ambiguous data to train a practical NED model. In response to this situation, we present MMGraph, which uses multimodal graph convolution to aggregate visual and contextual language information for accurate entity disambiguation for short texts, and a self-supervised simple triplet network (SimTri) that can learn useful representations in multimodal unlabeled data to enhance the effectiveness of NED models. We evaluated these approaches on a new dataset, MMFi, which contains multimodal supervised data and large amounts of unlabeled data. Our experiments confirm the state-of-the-art performance of MMGraph on two widely used benchmarks and MMFi. SimTri further improves the performance of NED methods. The dataset and code are available at https://github.com/LanceZPF/NNED_MMGraph.
Kaining Ying, Zhenhua Wang 0003, Dongyan Guo, Cong Bai
IEEE Trans. Neural Networks Learn. Syst.3
2023 CTVIS: Consistent Training for Online Video Instance Segmentation
abstract
The discrimination of instance embeddings plays a vital role in associating instances across time for online video instance segmentation (VIS). Instance embedding learning is directly supervised by the contrastive loss computed upon the contrastive items (CIs), which are sets of anchor/positive/negative embeddings. Recent online VIS methods leverage CIs sourced from one reference frame only, which we argue is insufficient for learning highly discriminative embeddings. Intuitively, a possible strategy to enhance CIs is replicating the inference phase during training. To this end, we propose a simple yet effective training strategy, called Consistent Training for Online VIS (CTVIS), which devotes to aligning the training and inference pipelines in terms of building CIs. Specifically, CTVIS constructs CIs by referring inference the momentum-averaged embedding and the memory bank storage mechanisms, and adding noise to the relevant embeddings. Such an extension allows a reliable comparison between embeddings of current instances and the stable representations of historical instances, thereby conferring an advantage in modeling VIS challenges such as occlusion, re-identification, and deformation. Empirically, CTVIS outstrips the SOTA VIS models by up to +5.0 points on three VIS benchmarks, including YTVIS19 (55.1% AP), YTVIS21 (50.1% AP) and OVIS (35.5% AP). Furthermore, we find that pseudo-videos transformed from images can train robust models surpassing fully-supervised ones.
Kaining Ying, Weian Mao, Zhenhua Wang 0003, Hao Chen 0041, Lin Wu 0001, Yifan Liu 0001, Chengxiang Fan, Yunzhi Zhuge, Chunhua Shen
ICCV4
2023 Human-to-Human Interaction Detection
Zhenhua Wang 0003, Kaining Ying, Jiajun Meng, Jifeng Ning
ICONIP (4)1
2023 Human Interaction Understanding With Consistency-Aware Learning
abstract
Compared with the progress made on human activity classification, much less success has been achieved on human interaction understanding (HIU). Apart from the latter task is much more challenging, the main causation is that recent approaches learn human interactive relations via shallow graphical representations, which are inadequate to model complicated human interactive-relations. This paper proposes a deep consistency-aware framework aiming at tackling the grouping and labelling inconsistencies in HIU. This framework consists of three components, including a backbone CNN to extract image features, a factor graph network to implicitly learn higher-order consistencies among labelling and grouping variables, and a consistency-aware reasoning module to explicitly enforcing consistencies. The last module is inspired by our key observation that the consistency-aware reasoning bias can be embedded into an energy function or a particular loss function, minimizing which delivers consistent predictions. An efficient mean-field inference algorithm is proposed, such that all modules of our network could be trained in an end-to-end fashion. Experimental results demonstrate that the two proposed consistency-learning modules complement each other, and both make considerable contributions in achieving leading performance on three benchmarks of HIU. The effectiveness of the proposed approach is further validated by experiments on detecting human-object interactions.
Jiajun Meng, Zhenhua Wang 0003, Kaining Ying, Jianhua Zhang 0002, Dongyan Guo, Zhen Zhang 0008, Qinfeng Shi, Shengyong Chen
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 ISDA: Position-Aware Instance Segmentation with Deformable Attention
abstract
Most instance segmentation models are not end-to-end trainable due to either the incorporation of proposal estimation (RPN) as a pre-processing or non-maximum suppression (NMS) as a post-processing. Here we propose a novel end-to-end instance segmentation method termed ISDA. It reshapes the task into predicting a set of object masks, which are generated via traditional convolution operation with learned position-aware kernels and features of objects. Such kernels and features are learned by leveraging a deformable attention network with multi-scale representation. Thanks to the introduced set-prediction mechanism, the proposed method is NMS-free. Empirically, ISDA outperforms Mask R-CNN (the strong baseline) by 2.6 points on MS-COCO, and achieves leading performance compared with recent models. Code will be available soon.
Kaining Ying, Zhenhua Wang 0003, Cong Bai
ICASSP2
2022 Robust and Accurate Multi-Agent SLAM with Efficient Communication for Smart Mobiles
abstract
In a long-term large-scenario application, the multi-agent collaborative SLAM is expected to improve the robustness and efficiency of executing tasks for mobile agents. In this paper, a multi-agent collaborative visual-inertial SLAM system is proposed based on a centralized client-server (CS) architecture, where the clients run on smart mobiles. In general, multi-agent collaborative SLAM relies on robust and precise experience sharing and efficient communication among agents. The experience sharing requires the place recognition with a high recall and accuracy, the precise estimation of transformation between looping frames, and the map fusion with globally consistency. To this end, we devise an enhanced geometric verification, a re-projection optimization based on the error-aware weighting strategy, and a strategy of flexible fusion to meet these requirements. In addition, the multi-agent collaborative SLAM needs to exchange abundant information, which requires the efficient communication. Therefore, we design a CS collaborative loop detection mechanism which is more robust to network transmission. We perform extensive experiments on the EuRoc dataset and in real environments. Experimental results show that the proposed system achieves better results than state-of-the-art methods. Furthermore, we demonstrate the stability of the proposed collaborative SLAM in real environments with a bandwidth of 7.55Mbps.
Kaiqi Chen 0001, Ruyu Liu, Yanhong Yang, Zhenhua Wang 0003, Jianhua Zhang 0002
ICRA5
2022 Human Interaction Recognition with Skeletal Attention and Shift Graph Convolution
abstract
Human interaction recognition has wide applications including intelligent surveillance, intelligent transportation and the analysis of sports videos. In recent years, benefiting from the development of action recognition based on deep learning, the performance of human interaction recognition has been boosted. This paper tackles two vital issues in recognizing human interactions, namely target missing and inadequate feature expression. To this end, we first design a data preprocessing method using skeleton estimation and multi-object tracking, which effectively reduces the chance of missing detection. Second, we propose a two-stream network composing of an appearance branch and a pose branch. The appearance branch extracts features enhanced via part affinity maps and part confidences maps, while the pose branch trains a customized Shift-GCN to extract skeletal features from people-pairs. Appearance and pose features are then fused to generate a more powerful representation of human interactions. Extensive experiments on two existing benchmarks, UT and BIT-Interaction, as well as a new dataset crafted by us, namely Campus-Interaction (CI), demonstrate the superior performance of the proposed approach over the state-of-the-arts.
Zhenhua Wang 0003, Jiajun Meng, Sheng Liu 0002, Jianhua Zhang 0002, Shengyong Chen
IJCNN2
2022 Joint Classification and Regression for Visual Tracking with Fully Convolutional Siamese Networks
abstract
Abstract Visual tracking of generic objects is one of the fundamental but challenging problems in computer vision. Here, we propose a novel fully convolutional Siamese network to solve visual tracking by directly predicting the target bounding box in an end-to-end manner. We first reformulate the visual tracking task as two subproblems: a classification problem for pixel category prediction and a regression task for object status estimation at this pixel. With this decomposition, we design a simple yet effective Siamese architecture based classification and regression framework, termed SiamCAR, which consists of two subnetworks: a Siamese subnetwork for feature extraction and a classification-regression subnetwork for direct bounding box prediction. Since the proposed framework is both proposal- and anchor-free, SiamCAR can avoid the tedious hyper-parameter tuning of anchors, considerably simplifying the training. To demonstrate that a much simpler tracking framework can achieve superior tracking results, we conduct extensive experiments and comparisons with state-of-the-art trackers on a few challenging benchmarks. Without bells and whistles, SiamCAR achieves leading performance with a real-time speed. Furthermore, the ablation study validates that the proposed framework is effective with various backbone networks, and can benefit from deeper networks. Code is available at https://github.com/ohhhyeahhh/SiamCAR .
Dongyan Guo, Yanyan Shao, Zhenhua Wang 0003, Chunhua Shen, Liyan Zhang 0001, Shengyong Chen
Int. J. Comput. Vis.4
2022 SSA-Net: Spatial self-attention network for COVID-19 pneumonia infection segmentation with semi-supervised few-shot learning
Xiaoyan Wang 0007, Yiwen Yuan, Dongyan Guo, Ming Xia 0005, Zhenhua Wang 0003, Cong Bai, Shengyong Chen
Medical Image Anal.7
2022 Accurate Object Association and Pose Updating for Semantic SLAM
abstract
Current pandemic has caused the medical system to operate under high load. To relieve it, robots with high autonomy can be used to effectively execute contactless operations in hospitals and reduce cross-infection between medical staff and patients. Although semantic Simultaneous Localization and Mapping (SLAM) technology can improve the autonomy of robots, semantic object association is still a problem that is worthy of being studied. The key to solving this problem is to correctly associate multiple object measurements of one object landmark by using semantic information, and to refine the pose of object landmark in real time. To this end, we propose a hierarchical object association strategy and a pose-refinement approach. The former one consists of two levels, i.e., a short-term object association and a global one. In the first level, we employ the multiple-object-tracking for short-term object association, through which the incorrect association among objects whose locations are close and appearances are similar can be avoided. Moreover, the short-term object association can provide more abundant object appearance and more robust estimation of object pose for the global object association in the second level. To refine the object pose in the map, we develop an approach to choose the optimal object pose from all object measurements associated with an object landmark. The proposed method is comprehensively evaluated on seven simulated hospital sequences, a real hospital environment and the KITTI dataset. Experimental results show that our method has an obviously improvement in terms of robustness and accuracy for the object association and the trajectory estimation in the semantic SLAM.
Kaiqi Chen 0001, Qinying Chen, Zhenhua Wang 0003, Jianhua Zhang 0002
IEEE Trans. Intell. Transp. Syst.4
2021 Graph Attention Tracking
abstract
Siamese network based trackers formulate the visual tracking task as a similarity matching problem. Almost all popular Siamese trackers realize the similarity learning via convolutional feature cross-correlation between a target branch and a search branch. However, since the size of target feature region needs to be pre-fixed, these cross-correlation base methods suffer from either reserving much adverse background information or missing a great deal of foreground information. Moreover, the global matching be-tween the target and search region also largely neglects the target structure and part-level information.In this paper, to solve the above issues, we propose a simple target-aware Siamese graph attention network for general object tracking. We propose to establish part-to-part correspondence between the target and the search region with a complete bipartite graph, and apply the graph attention mechanism to propagate target information from the template feature to the search feature. Further, instead of using the pre-fixed region cropping for template-feature-area selection, we investigate a target-aware area selection mechanism to fit the size and aspect ratio variations of different objects. Experiments on challenging benchmarks including GOT-10k, UAV123, OTB-100 and LaSOT demonstrate that the proposed SiamGAT outperforms many state-of-the-art trackers and achieves leading performance. Code is available at: https://git.io/SiamGAT
Dongyan Guo, Yanyan Shao, Zhenhua Wang 0003, Liyan Zhang 0001, Chunhua Shen
CVPR4
2021 Consistency-Aware Graph Network for Human Interaction Understanding
abstract
Compared with the progress made on human activity classification, much less success has been achieved on human interaction understanding (HIU). Apart from the latter task is much more challenging, the main cause is that recent approaches learn human interactive relations via shallow graphical models, which is inadequate to model complicated human interactions. In this paper, we propose a consistency-aware graph network, which combines the representative ability of graph network and the consistency-aware reasoning to facilitate HIU. Our network consists of three components, a backbone CNN to extract image features, a factor graph network to learn third-order interactive relations among participants, and a consistency-aware reasoning module to enforce labeling and grouping consistencies. Our key observation is that the consistency-aware-reasoning bias for HIU can be embedded into an energy, minimizing which delivers consistent predictions. An efficient mean-field inference algorithm is proposed, such that all modules of our network could be trained jointly in an end-to-end manner. Experimental results show that our approach achieves leading performance on three benchmarks. Code is available at https://git.io/CAGNet.
Zhenhua Wang 0003, Jiajun Meng, Dongyan Guo, Jianhua Zhang 0002, Qinfeng Shi, Shengyong Chen
ICCV1
2021 End-to-end feature fusion Siamese network for adaptive visual tracking
abstract
Abstract According to observations, different visual objects have different salient features in different scenarios. Even for the same object, its salient shape and appearance features may change greatly from time to time in a long‐term tracking task. Motivated by them, an end‐to‐end feature fusion framework was proposed based on the Siamese network, named FF‐Siam, which can effectively fuse different features for adaptive visual tracking. The framework consists of four layers. A feature extraction layer is designed to extract the different features of the target region and search region. The extracted features are then put into a weight generation layer to obtain the channel weights, which indicate the importance of different feature channels. Both features and the channel weights are utilised in a template generation layer to generate a discriminative template. Finally, the corresponding response maps created by the convolution of the search region features and the template are applied with a fusion layer to obtain the final response map for locating the target. Experimental results demonstrate that the proposed framework achieves state‐of‐the‐art performance on the popular Temple‐Colour, OTB50 and UAV123 benchmarks.
Dongyan Guo, Weixuan Zhao, Zhenhua Wang 0003, Shengyong Chen
IET Image Process.5
2021 Detection and Segmentation of Unlearned Objects in Unknown Environment
abstract
Detecting and segmenting unlearned objects in unknown environment is a very important visual perception ability to enhance industrial intelligence. In this article, we present a novel conditional random field model integrating unimodal and cross-modal terms for detecting and segmenting object instances without knowing their categories and without sampling extra proposals. This model takes a paired image and point cloud as input, from which we first develop a set of novel category-independent features to distinguish objects. Then, a set of unary, pairwise, and higher order potentials are designed according to these category-independent features, and the cross-modal potential is introduced as a novel global constraints to keep the spatial consistency in both 2-D and 3-D modalities. In this novel model, the unlearned object detection and segmentation is treated as the process of pixel labeling. Thus, adjacent or occlusion object instances can also be separated efficiently from a labeled map. By comparison with the baseline methods, experimental results on a public RGB+D dataset show that the proposed model can obtain better performance with improved precision and recall rate. Moreover, we use the proposed method in a real industrial scene and achieve satisfactory performance.
Jianhua Zhang 0002, Jingbo Chen, Shengyong Chen, Zhenhua Wang 0003, Jianwei Zhang 0001
IEEE Trans. Ind. Informatics4
2021 Human Interaction Understanding With Joint Graph Decomposition and Node Labeling
abstract
The task of human interaction understanding involves both recognizing the action of each individual in the scene and decoding the interaction relationship among people, which is useful to a series of vision applications such as camera surveillance, video-based sports analysis and event retrieval. This paper divides the task into two problems including grouping people into clusters and assigning labels to each of them, and presents an approach to solving these problems in a joint manner. Our method does not assume the number of groups is known beforehand as this will substantially restrict its application. With the observation that the two challenges are highly correlated, the key idea is to model the pairwise interacting relations among people via a complete graph and its associated energy function such that the labeling and grouping problems are translated into the minimization of the energy function. We implement this joint framework by fusing both deep features and rich contextual cues, and learn the fusion parameters from data. An alternating search algorithm is developed in order to efficiently solve the associated inference problem. By combining the grouping and labeling results obtained with our method, we are able to achieve the semantic-level understanding of human interactions. Extensive experiments are performed to qualitatively and quantitatively evaluate the effectiveness of our approach, which outperforms state-of-the-art methods on several important benchmarks. An ablation study is also performed to verify the effectiveness of different modules within our approach.
Zhenhua Wang 0003, Jinchao Ge, Dongyan Guo, Jianhua Zhang 0002, Yanjing Lei, Shengyong Chen
IEEE Trans. Image Process.1
2020 SiamCAR: Siamese Fully Convolutional Classification and Regression for Visual Tracking
abstract
By decomposing the visual tracking task into two subproblems as classification for pixel category and regression for object bounding box at this pixel, we propose a novel fully convolutional Siamese network to solve visual tracking end-to-end in a per-pixel manner. The proposed framework SiamCAR consists of two simple subnetworks: one Siamese subnetwork for feature extraction and one classification-regression subnetwork for bounding box prediction. Different from state-of-the-art trackers like Siamese-RPN, SiamRPN++ and SPM, which are based on region proposal, the proposed framework is both proposal and anchor free. Consequently, we are able to avoid the tricky hyper-parameter tuning of anchors and reduce human intervention. The proposed framework is simple, neat and effective. Extensive experiments and comparisons with state-of-the-art trackers are conducted on challenging benchmarks including GOT-10K, LaSOT, UAV123 and OTB-50. Without bells and whistles, our SiamCAR achieves the leading performance with a considerable real-time speed. The code is available at https://github.com/ohhhyeahhh/SiamCAR.
Dongyan Guo, Zhenhua Wang 0003, Shengyong Chen
CVPR4
2020 SpSequenceNet: Semantic Segmentation Network on 4D Point Clouds
abstract
Point clouds are useful in many applications like autonomous driving and robotics as they provide natural 3D information of the surrounding environments. While there are extensive research on 3D point clouds, scene understanding on 4D point clouds, a series of consecutive 3D point clouds frames, is an emerging topic and yet under-investigated. With 4D point clouds (3D point cloud videos), robotic systems could enhance their robustness by leveraging the temporal information from previous frames. However, the existing semantic segmentation methods on 4D point clouds suffer from low precision due to the spatial and temporal information loss in their network structures. In this paper, we propose SpSequenceNet to address this problem. The network is designed based on 3D sparse convolution. And we introduce two novel modules, a cross-frame global attention module and a cross-frame local interpolation module, to capture spatial and temporal information in 4D point clouds. We conduct extensive experiments on SemanticKITTI, and achieve the state-of-the-art result of 43.1% on mIoU, which is 1.5% higher than the previous best approach.
Hanyu Shi 0002, Guosheng Lin, Hao Wang 0094, Tzu-Yi Hung, Zhenhua Wang 0003
CVPR5
2020 CalibRCNN: Calibrating Camera and LiDAR by Recurrent Convolutional Neural Network and Geometric Constraints
abstract
In this paper, we present Calibration Recurrent Convolutional Neural Network (CalibRCNN) to infer a 6 degrees of freedom (DOF) rigid body transformation between 3D LiDAR and 2D camera. Different from the existing methods, our 3D-2D CalibRCNN not only uses the LSTM network to extract the temporal features between 3D point clouds and RGB images of consecutive frames, but also uses the geometric loss and photometric loss obtained by the interframe constraint to refine the calibration accuracy of the predicted transformation parameters. The CalibRCNN aims at inferring the correspondence between projected depth image and RGB image to learn the underlying geometry of 2D-3D calibration. Thus, the proposed calibration model achieves a good generalization ability to adapt to unknown initial calibration error ranges, and other 3D LiDAR and 2D camera pairs with different intrinsic parameters from the training dataset. Extensive experiments have demonstrated that our CalibRCNN can achieve state-of-the-art accuracy by comparison with other CNN based methods.
Jieying Shi, Ziheng Zhu, Jianhua Zhang 0002, Ruyu Liu, Zhenhua Wang 0003, Shengyong Chen, Honghai Liu 0001
IROS5
2019 New Convex Relaxations for MRF Inference With Unknown Graphs
abstract
Treating graph structures of Markov random fields as unknown and estimating them jointly with labels have been shown to be useful for modeling human activity recognition and other related tasks. We propose two novel relaxations for solving this problem. The first is a linear programming (LP) relaxation, which is provably tighter than the existing LP relaxation. The second is a non-convex quadratic programming (QP) relaxation, which admits an efficient concave-convex procedure (CCCP). The CCCP algorithm is initialized by solving a convex QP relaxation of the problem, which is obtained by modifying the diagonal of the matrix that specifies the non-convex QP relaxation. We show that our convex QP relaxation is optimal in the sense that it minimizes the L1 norm of the diagonal modification vector. While the convex QP relaxation is not as tight as the existing and the new LP relaxations, when used in conjunction with the CCCP algorithm for the non-convex QP relaxation, it provides accurate solutions. We demonstrate the efficacy of our new relaxations for both synthetic data and human activity recognition.
Zhenhua Wang 0003, Qinfeng Shi, M. Pawan Kumar, Jianhua Zhang 0002
ICCV1
2019 Joint Grouping and Labeling via Complete Graph Decomposition
Jinchao Ge, Zhenhua Wang 0003, Jiajun Meng, Jianhua Zhang 0002, Shengyong Chen
ICONIP (5)2
2018 A Hierarchical Model for Action Recognition Based on Body Parts
abstract
As increasing attention is paid on human action recognition from skeleton data, this paper focuses on such tasks by proposing a hierarchical model to discover the structure information of body-parts involved in human actions. Considering human actions as simultaneous motions of different body-parts of the human skeleton, we propose a hierarchical model to simultaneously apply discriminative body-parts selection at a same scale and group coupling of bundles of body-parts at different scales, while we decompose the human skeleton into a hierarchy of body-parts of varying scales. To represent such hierarchy of body-parts, we accordingly build a hierarchical RRV (Rotation and Relative Velocity) descriptors. The hierarchical representations encoded by Fisher vectors of the hierarchical RRV descriptors are properly formulated into the hierarchical model via the proposed hierarchical mixed norm, to apply sparse selection of body-parts and regularize the structure of such hierarchy of body-parts. The extensive evaluations on three challenging datasets demonstrate the effectiveness of our proposed approach, which achieves superior performance compared to state-of-the-art results on different sizes of datasets, showing it is more widely applicable than existing approaches.
Zhanpeng Shao, Youfu Li 0001, Yao Guo 0002, Jianyu Yang 0002, Zhenhua Wang 0003
ICRA5
2018 Deep CRF-Graph Learning for Semantic Image Segmentation
Fuguang Ding, Zhenhua Wang 0003, Dongyan Guo, Shengyong Chen, Jianhua Zhang 0002, Zhanpeng Shao
PRICAI2
2018 Siamese Network Based Features Fusion for Adaptive Visual Tracking
Dongyan Guo, Weixuan Zhao, Zhenhua Wang 0003, Shengyong Chen, Jian Zhang 0002
PRICAI (1)4
2018 Understanding human activities in videos: A joint action and interaction learning approach
Zhenhua Wang 0003, Jiali Jin, Sheng Liu 0002, Jianhua Zhang 0002, Shengyong Chen, Zhen Zhang 0008, Dongyan Guo, Zhanpeng Shao
Neurocomputing1
2017 Joint label-interaction learning for human action recognition
abstract
Human interactions and their action categories preserve strong correlations, and the identification of the interaction configuration is of significant importance to improve the action recognition result. However, interactions are typically estimated using heuristics or treated as latent variables. The former usually produces incorrect interaction configuration while the latter introduces challenging training problem. Hence we propose a framework to jointly learn interactions and actions by designing a potential function using both features learned via deep neural networks and human interaction context. We propose an iterative approach to solve the associated inference problem efficiently and approximately. Experimental results on real datasets demonstrate that the proposed approach outperforms baselines by a large margin, and is competitive compared with the state-of-the-arts.
Jiali Jin, Zhenhua Wang 0003, Sheng Liu 0002, Jianhua Zhang 0002, Shengyong Chen, Qiu Guan
ICIP2
2017 Fast initialization for feature-based monocular slam
abstract
Initial map determines the effect of followed slam tracking. Most feature-based monocular slam initialize their map according to key points matching in close frames. Nevertheless, it will consume lots of computational resources and time. And it is easy to fail in some far scene or close scene. In this paper, we present a fast initialization method to reduce runtime and improve success rate of initialization for feature-based monocular slam. First, vanishing points detection based on line segment detector [1] is adopted. Second, we extract orb key points. And the coordinates of every key points are undistorted and normalized. Third, we generate the corresponding depth for each key point by normalizing its distance to the existing vanishing points or gaussian random number. We compare our method with state-of-the-art on public data sets and ours. The experiments show that our method outperforms on runtime and accuracy.
Shaobo Zhang 0005, Sheng Liu 0002, Jianhua Zhang 0002, Zhenhua Wang 0003, Xiaoyan Wang 0007
ICIP4
2017 A Spatio-Temporal CRF for Human Interaction Understanding
abstract
A better understanding of human interactions in videos can be achieved by simultaneously considering the coarse interactions between people, the action of each individual, and the activity of all people as a whole. We divide the recognition task into two stages. The first stage discriminates interactions and noninteractions, actions and activities based on local image information, while during the second stage, actions and activities are recognized in a global manner based on the local recognition results. A conditional random field (CRF) is designed to model human interactions in the spatio-temporal space. Different from most existing global models which cover either action or activity variables only, our model covers them both by considering the interactions between different types of variables. The graph structure of the CRF is predicted by a model learned from training data, which is different from traditional graph construction methods that typically rely on human heuristics. We learn the parameters of the CRF via structured support vector machine. We propose an efficient inference algorithm to tackle the estimation of labels in long videos containing many people. Our model admits both semantic-level understanding of human interactions in videos and competitive action and activity recognition performance.
Zhenhua Wang 0003, Sheng Liu 0002, Jianhua Zhang 0002, Shengyong Chen, Qiu Guan
IEEE Trans. Circuits Syst. Video Technol.1
2015 A Hybrid Loss for Multiclass and Structured Prediction
abstract
We propose a novel hybrid loss for multiclass and structured prediction problems that is a convex combination of a log loss for Conditional Random Fields (CRFs) and a multiclass hinge loss for Support Vector Machines (SVMs). We provide a sufficient condition for when the hybrid loss is Fisher consistent for classification. This condition depends on a measure of dominance between labels-specifically, the gap between the probabilities of the best label and the second best label. We also prove Fisher consistency is necessary for parametric consistency when learning models such as CRFs. We demonstrate empirically that the hybrid loss typically performs least as well as-and often better than-both of its constituent losses on a variety of tasks, such as human action recognition. In doing so we also provide an empirical comparison of the efficacy of probabilistic and margin based approaches to multiclass and structured prediction.
Qinfeng Shi, Mark D. Reid, Tibério S. Caetano, Anton van den Hengel, Zhenhua Wang 0003
IEEE Trans. Pattern Anal. Mach. Intell.5
2013 Bilinear Programming for Human Activity Recognition with Unknown MRF Graphs
abstract
Markov Random Fields (MRFs) have been successfully applied to human activity modelling, largely due to their ability to model complex dependencies and deal with local uncertainty. However, the underlying graph structure is often manually specified, or automatically constructed by heuristics. We show, instead, that learning an MRF graph and performing MAP inference can be achieved simultaneously by solving a bilinear program. Equipped with the bilinear program based MAP inference for an unknown graph, we show how to estimate parameters efficiently and effectively with a latent structural SVM. We apply our techniques to predict sport moves (such as serve, volley in tennis) and human activity in TV episodes (such as kiss, hug and Hi-Five). Experimental results show the proposed method outperforms the state-of-the-art.
Zhenhua Wang 0003, Qinfeng Shi, Chunhua Shen, Anton van den Hengel
CVPR1