Guijin Wang

dblp:37/6836 · DBLP profile ↗
← Back
88ranked-venue papers
14as first author
35since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 52 · 5 first-author · 21 since 2021Artificial intelligence and machine learning · 30 · 4 first-author · 12 since 2021Systems, architecture and hardware · 6 · 6 since 2021Computer networks · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSecurity and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 Visual Geometric 6-DoF Grasping: A Multi-View Framework with Sparse RGB Observations
Yixiang Dai, Kaiqin Yang, Yongjiang Zhao, Siang Chen, Guijin Wang
ISCAS9
2026 CoDiFSR: Code Diffusion paradigm with prior knowledge distillation for face super-resolution
Junhao Gu, Zhengguo Wang, Shuimu Chen, Wenming Yang, Guijin Wang
Pattern Recognit.6
2026 Rethinking 6-DoF grasp detection: A flexible framework for high-quality grasping
Pengwei Xie, Siang Chen, Kaiqin Yang, Guijin Wang
Pattern Recognit.5
2026 LVMF3D: Large Vision Model Boosting Multimodal Fusion for Indoor 3D Object Detection
abstract
3D object detection plays an important role in intelligent systems perceiving the world. Although many studies have been conducted to address this task, the detection accuracy is still limited by the network's learning capability. Therefore, we propose LVMF3D, a Large Vision Model (LVM) boosted multimodal fusion indoor 3D object detection framework, consisting of two branches. The pre-trained LVM is used as the RGB branch to better extract the image texture feature. The point branch is used to encode the spatial geometric feature. Furthermore, Point Fusion Module (PFM) and Multi-Scale Attention Fusion Module (MS-AFM) are specially designed in the 2D and 3D spaces, respectively, to realize more comprehensive and effective information fusion between the two branches. We conduct experiments on the indoor 3D object detection dataset SUN RGB-D and achieve state-of-the-art results compared to other 3D object detection methods.
Wenming Yang, Guijin Wang
IEEE Signal Process. Lett.4
2025 SAP-SLAM: Semantic-Assisted Perception SLAM with 3D Gaussian Splatting
abstract
The integration of 3D Gaussians has introduced a novel scene representation in Simultaneous Localization and Mapping (SLAM), characterized by explicit representation and differentiable rendering capabilities that enhance scene reconstruction and understanding. However, most current SLAM systems only exploit the basic representational capacity of 3D Gaussians, neglecting their potential to offer richer information and facilitate higher-dimensional scene comprehension. Furthermore, these systems often struggle with reconstruction when encountering rapid camera movements or depth missing. Drawing inspiration from 3D language field, which explores the intrinsic relationships among scene objects, we propose SAPSLAM, a dense SLAM system that combines high-fidelity reconstruction and advanced semantic understanding. Our approach leverages pre-trained visual models to extract semantic features, which are then fused, dimensionally reduced, and encoded into the 3D Gaussian model for optimization and rendering. The integration of these features improves the systems semantic comprehension and scene representation, ultimately enabling the creation of high-precision 3D semantic maps. Additionally, we introduce a semantic-guided Gaussian densification and pruning strategy, which uses semantic consistency to prioritize attention on poorly reconstructed areas, greatly improving performance in complex scenarios. SAP-SLAM achieves competitive results on both real-world and synthetic datasets, demonstrating superior capabilities in semantic understanding and reconstruction.
Yudong Lin, Wenming Yang, Guijin Wang, Qingmin Liao
ICRA4
2025 Region-Centric 6-Dof Grasp Detection: A Data-Efficient Solution for Cluttered Scenes
abstract
Robotic grasping, serving as the cornerstone of robot manipulation, is fundamental for embodied intelligence. Manipulation in challenging scenarios demands grasp detection algorithms with higher efficiency and generalizability. However, for general 6-Dof grasp detection, most data-driven methods directly extract scene-level features to generate grasp prediction, relying on a relatively heavy scene-level feature encoder and a significant amount of data with dense grasp labels for model training. In this letter, we propose a novel data-efficient 6-Dof grasp detection framework in cluttered scenes, named Region-Centric Grasp Detection (RCGD), consisting of an Iterative Search Module (ISM) and a Region Grasp Model (RGM). Concretely, ISM aims to retrieve potential region centers and aggregate multiple regions in a coarse-to-fine way. Then, RGM extracts aligned grasp-related embeddings and predicts grasps within these local regions. Benefiting from the region-centric paradigm and the training-free location strategy, RCGD significantly outperforms previous methods and shows minimal performance loss with even a very small portion of training data or labels. Furthermore, real-world robotic experiments in two distinct settings highlight the effectiveness of our method with a 95% success rate.
Siang Chen, Pengwei Xie, Dingchang Hu, Wenming Yang, Guijin Wang
IROS6
2025 FEG-VON: Frontier Embedding Graph for Efficient Visual Object Navigation
abstract
Visual object navigation, requiring agents to locate target objects in novel environments through egocentric visual observation, remains a critical challenge in Embodied AI. We propose FEG-VON, a training-free framework that constructs and maintains a Frontier Embedding Graph for efficient Visual Object Navigation. The graph initializes frontier embeddings using Vision Language Models (VLMs), where visual observations are encoded into spatially anchored semantic embeddings through cross-modal alignment with target text descriptors. We then update the graph by aggregating spatio-temporal semantic relations across frontiers, enabling online adaptation to new targets via similarity scoring without remapping. The evaluation results in public benchmarks demonstrate the superior performance of FEG-VON in both single- and multi-object navigation tasks compared with state-of-the-art methods. Crucially, FEG-VON eliminates dependency on task-specific training for exploration and advances the feasibility of zero-shot navigation in open-world environments.
Yingru Dai, Pengwei Xie, Yikai Liu, Siang Chen, Wenming Yang, Guijin Wang
IROS6
2025 Efficient End-to-End 6-Dof Grasp Detection Framework for Edge Devices with Hierarchical Heatmaps and Feature Propagation
abstract
6-DoF grasp detection is important for the advancement of intelligent embodied systems, as it provides feasible robot poses for object grasping. Various methods have been proposed to detect 6-DoF grasps through the extraction of 3D geometric features from RGBD or point cloud data. However, most of these approaches encounter challenges during real robot deployment due to their significant computational demands, which can be particularly problematic for mobile robot platforms, especially those reliant on edge computing devices. This paper presents an Efficient End-to-End Grasp Detection Network (E3GNet) for 6-DoF grasp detection utilizing hierarchical heatmap representations. E3GNet effectively identifies high-quality and diverse grasps in cluttered real-world environments. Benefiting from our end-to-end methodology and efficient network design, our approach surpasses previous methods in model inference efficiency and achieves real-time 6-Dof grasp detection on edge devices. Furthermore, real-world experiments validate the effectiveness of our method, achieving a satisfactory 94% object grasping success rate. More details can be found on our project page.
Kaiqin Yang, Yixiang Dai, Guijin Wang, Siang Chen
ISCAS3
2025 Butter: Frequency Consistency and Hierarchical Fusion for Autonomous Driving Object Detection
Xiaojian Lin, Wenxin Zhang 0005, Yuchu Jiang, Wangyu Wu, Kangxu Wang, Zongzheng Zhang, Guijin Wang, Lei Jin 0003, Hao Zhao 0002
ACM Multimedia8
2025 MDSI: Pluggable Multi-strategy Decoupling with Semantic Integration for RGB-D Gesture Recognition
Fengyi Fang, Zhehan Kan, Guijin Wang, Wenming Yang
Pattern Recognit.4
2025 Diffusion-Based Depth Inpainting for Transparent and Reflective Objects
abstract
Transparent and reflective objects, which are common in our everyday lives, present a significant challenge to 3D imaging techniques due to their unique visual and optical properties. Faced with these types of objects, RGB-D cameras fail to capture the real depth value with their accurate spatial information. To address this issue, we propose DITR, a diffusion-based Depth Inpainting framework specifically designed for Transparent and Reflective objects. This network consists of two stages, including a Region Proposal stage and a Depth Inpainting stage. DITR dynamically analyzes the optical and geometric depth loss and inpaints them automatically. Furthermore, comprehensive experimental results demonstrate that DITR is highly effective in depth inpainting tasks of transparent and reflective objects with robust adaptability.
Dingchang Hu, Yixiang Dai, Guijin Wang
IEEE Trans. Circuits Syst. Video Technol.4
2024 Category-Agnostic Pose Estimation for Point Clouds
abstract
The goal of object pose estimation is to visually determine the pose of a specific object in the RGB-D input. Unfortunately, when faced with new categories, both instance-based and category-based methods are unable to deal with unseen objects of unseen categories, which is a challenge for pose estimation. To address this issue, this paper proposes a method to introduce geometric features for pose estimation of point clouds without requiring category information. The method is based only on the patch feature of the point cloud, a geometric feature with rotation invariance. After training without category information, our method achieves as good results as other category-based methods. Our method successfully achieved pose annotation of no category information instances on the CAMERA25 dataset and ModelNet40 dataset.
Siang Chen, Pengwei Xie, Guijin Wang
ICIP5
2024 Agent-Guided Gaze Estimation Network by Two-Eye Asymmetry Exploration
abstract
Gaze estimation is an important task in understanding human visual attention. Despite the performance gain brought by recent algorithm development, the task remains challenging due to two-eye appearance asymmetry resulting from head pose variation and nonuniform illumination. In this paper, we propose a novel architecture, Agent-guided Gaze Estimation Network (AGE-Net), to make full and efficient use of two-eye features. By exploring the appearance asymmetry and the consequent feature space asymmetry, we devise a main branch and two agent regression tasks. The main branch extracts related features of the left and right eyes from low-level semantics. Meanwhile, the agent regression tasks extract asymmetric features of the left and right eyes from high-level semantics, so as to guide the main branch to learn more about the eye feature space. Experiments show that our method achieves state-of-the-art gaze estimation task performance on both MPIIGaze and EyeDiap datasets.
Wenming Yang, Guijin Wang
ICIP4
2024 Query-Guided Support Prototypes for Few-Shot 3D Indoor Segmentation
abstract
Few-shot 3D point cloud segmentation segments novel categories in point cloud scenes with only limited annotations. However, most current methods do not consider query content when exploring support prototypes, and thus suffer from intra-class variations between objects and incomplete representation of category information from annotated support samples. In this paper, we propose a novel Query-Guided support Prototype exploration Network (QGPNet) to tackle this challenge. Firstly, we present a point feature alignment module, which leverages geometry relationship between prototypes and query points, to tackle data misalignment caused by intra-class variations, and thus prevents incorrect label propagation from prototypes to query points. Secondly, we design a prototype feature mining strategy, which progressively harvests diverse support prototypes in the interaction with query features, to fully utilize the category information provided by annotated samples. Additionally, we introduce a semantic-aware data augmentation strategy for query samples in the training process, potentially improving the generalization ability of support prototypes on query samples. Extensive experiments on two indoor 3D datasets S3DIS and ScanNet demonstrate that QGPNet outperforms previous state-of-the-art methods by a large margin.
Dingchang Hu, Siang Chen, Huazhong Yang, Guijin Wang
IEEE Trans. Circuits Syst. Video Technol.4
2024 DeGCN: Deformable Graph Convolutional Networks for Skeleton-Based Action Recognition
abstract
Graph convolutional networks (GCN) have recently been studied to exploit the graph topology of the human body for skeleton-based action recognition. However, most of these methods unfortunately aggregate messages via an inflexible pattern for various action samples, lacking the awareness of intra-class variety and the suitableness for skeleton sequences, which often contain redundant or even detrimental connections. In this paper, we propose a novel Deformable Graph Convolutional Network (DeGCN) to adaptively capture the most informative joints. The proposed DeGCN learns the deformable sampling locations on both spatial and temporal graphs, enabling the model to perceive discriminative receptive fields. Notably, considering human action is inherently continuous, the corresponding temporal features are defined in a continuous latent space. Furthermore, we design an innovative multi-branch framework, which not only strikes a better trade-off between accuracy and model size, but also elevates the effect of ensemble between the joint and bone modalities remarkably. Extensive experiments show that our proposed method achieves state-of-the-art performances on three widely used datasets, NTU RGB+D, NTU RGB+D 120, and NW-UCLA.
Woomin Myung, Jing-Hao Xue, Guijin Wang
IEEE Trans. Image Process.4
2023 Hard Samples Based Margin Loss for Face Verification
abstract
Although softmax loss and its variants have achieved great success in face verification, the performance is still subject to the data imbalance and early saturation problems. In this paper, we define hard samples as minority class samples and early saturation samples, in order to address both issues, we propose a new loss function termed Hard-Samples based Margin (HSM) loss. Inspired by the class-variant margin normalized softmax loss, we add larger margin on minority classes, the proposed real-class margin overcomes the negative influence from the data imbalance via making the optimization more balanced, while by expanding the margin of early saturated samples, the proposed pseudo-class margin keeps the samples away from the saturation region. Comprehensive experiments show that our HSM loss consistently surpasses the state-of-the-art loss functions on four popular face verification benchmarks.
Xiaying Bai, Wenxian Zheng, Wenming Yang, Guijin Wang, Qingmin Liao
ICIP4
2023 EvenFace: Deep Face Recognition with Uniform Distribution of Identities
abstract
The development of loss functions over the past few years has brought great success to face recognition. Most algorithms focus on improving the intra-class compactness of face features but ignore the inter-class separability. In this paper, we propose a method named EvenFace, which introduces a regularization variance item and a mean term of inter-class separability to further promote the even distribution of class centers on the hypersphere, thereby increasing the inter-class distance. In order to evaluate the inter-class separability, a new index is proposed to better reflect the distribution of class centers and guide the classification. By penalizing the angle between each identity and its surrounding neighbors, the resulting uniform distribution of identities enables full exploitation of the feature space, leading to discriminative face representations. Our proposed loss function can effectively boost the performance of softmax loss variants. Quantitative comparisons with other state-of-the-art methods on several benchmarks demonstrate the superiority of EvenFace.
Yingfan Tao, Qiqi Bao 0001, Guijin Wang, Wenming Yang
ICME4
2023 Texture-Shape Optimized GAT for 3D Face Reconstruction
abstract
3D face reconstruction is widely used in face recognition research, online makeup, etc. However, texture and shape distortion regions usually exist in the reconstruction results. This paper proposes a novel framework named Texture-Shape optimized GAT for 3D face Reconstruction (TSGAT-3D), including data preprocessing and training phases. In the data preprocessing phase, we adopt the styleGAN2 to convert single-view images to multi-view images set. In the training phase, we design a novel Texture-Shape optimized Graph Attention Network, which can learn the facial prior knowledge from the multi-view images set, to improve the details of the initially reconstructed faces based on the auto-encoder module. This network can aggregate the features of face vertices according to the correlation of vertices, thereby improving the accuracy of the 3D reconstructed faces. Furthermore, we present a view loss function for this framework to constrain the shape and texture of the reconstructed face. Extensive experiments conducted on CelebA and Bosphorus show that the reconstruction results of our proposed method are closer to the real 3D faces.
Chao Hao, Guijin Wang
ISCAS3
2023 TROSD: A New RGB-D Dataset for Transparent and Reflective Object Segmentation in Practice
abstract
Transparent and reflective objects are omnipresent in our daily life, but their unique visual and optical characteristics are notoriously challenging even for state-of-the-art deep networks of semantic segmentation. To alleviate this challenge, we construct a new large-scale real-world RGB-D dataset called TROSD, which is more comprehensive than existing datasets for transparent and reflective object segmentation. Our TROSD dataset contains 11,060 RGB-D images with three semantic classes in terms of transparent objects, reflective objects, and others, covering a variety of daily scenes. Together with the dataset, we also introduce a novel network (TROSNet) as a high-standard baseline to assist other researchers to develop and benchmark their algorithms of transparent and reflective object segmentation. Moreover, extensive experiments also clearly show that the proposed TROSD dataset has an excellent capacity to facilitate the development of semantic segmentation algorithms with strong generalizability.
Guodong Zhang 0004, Wenming Yang, Jing-Hao Xue, Guijin Wang
IEEE Trans. Circuits Syst. Video Technol.5
2022 Pose-Invariant Face Recognition via Adaptive Angular Distillation
abstract
Pose-invariant face recognition is a practically useful but challenging task. This paper introduces a novel method to learn pose-invariant feature representation without normalizing profile faces to frontal ones or learning disentangled features. We first design a novel strategy to learn pose-invariant feature embeddings by distilling the angular knowledge of frontal faces extracted by teacher network to student network, which enables the handling of faces with large pose variations. In this way, the features of faces across variant poses can cluster compactly for the same person to create a pose-invariant face representation. Secondly, we propose a Pose-Adaptive Angular Distillation loss to mitigate the negative effect of uneven distribution of face poses in the training dataset to pay more attention to the samples with large pose variations. Extensive experiments on two challenging benchmarks (IJB-A and CFP-FP) show that our approach consistently outperforms the existing methods.
Zhenduo Zhang, Yongru Chen, Wenming Yang, Guijin Wang, Qingmin Liao
AAAI4
2022 SADG-Net: Sparse Adaptive Dynamic Guidance Network for Depth Completion
abstract
Existing depth completion methods with standard CNN and fixed guidance information often produce invalid value diffusion and mismatch between the guidance and the depth. To address this issue, we propose a SADG-Net for depth completion. Specifically, a sparse adaptive module is designed to infer the initial dense depth map and its confidence, as well as the affinity between pixels. Then we develop a dynamic guidance spatial propagation network to refine the initial depth map and dynamically update the guidance information with the inferred depth. In contrast to previous algorithms, our method effectively optimizes the processing for sparse depth and significantly alleviates the error accumulation issues in spatial propagation. Extensive experiments demonstrate that our model improves upon the state-of-the-art performance on NYUv2 and BIDCD datasets.
Guodong Zhang 0004, Chenchen Feng, Sifan Yang, Wenming Yang, Guijin Wang
ICME6
2022 Multi-scale Cross-Modal Transformer Network for RGB-D Object Detection
Pengwei Xie, Guijin Wang
MMM (1)3
2022 Distribution-aware Low-bit Quantization for 3D Point Cloud Networks
abstract
Various low-bit quantized methods have been widely exploited and shown decent performance on 2D vision tasks in recent years. Complemented with 2D images, 3D point clouds provide an opportunity to understand the surrounding environ-ment better. However, low-bit quantization methods designed for 2D vision tasks are not readily transferable to 3D point clouds due to the higher dimension of 3D data and the increased proportion of activations. In this work, we propose a novel quantization framework, DASCQ, for 3D point cloud processing. First, a new distribution-aware strategy (DA) is presented to decrease the deviation caused by extremely low-bit quantization through activation and weight distribution analysis. Second, a soft constraint manner (SC) is designed to smooth the training of quantized networks which suffer from backward propagation errors. We evaluate our approach on two 3D point cloud datasets, ModelNet40 and S3DIS. Results indicate that the performance of the proposed approach is superior to other state-of-the-art quantization methods on both shape classification and scene semantic segmentation tasks.
Dingchang Hu, Siang Chen, Huazhong Yang, Guijin Wang
VCIP4
2022 Frontal-Centers Guided Face: Boosting Face Recognition by Learning Pose-Invariant Features
abstract
In recent years, face recognition has made a remarkable breakthrough due to the emergence of deep learning. However, compared with frontal face recognition, plenty of deep face recognition models still suffer serious performance degradation when handling profile faces. To address this issue, we propose a novel Frontal-Centers Guided Loss (FCGFace) to obtain highly discriminative features for face recognition. Most existing discriminative feature learning approaches project features from the same class into a separated latent subspace. These methods only model the distribution at the identity-level but ignore the latent relationship between frontal and profile viewpoints. Different from these methods, FCGFace takes viewpoints into consideration by modeling the distribution at both the identity-level and the viewpoint-level. At the identity-level, a softmax-based loss is employed for a relatively rough classification. At the viewpoint-level, centers of frontal face features are defined to guide the optimization conducted in a more refined way. Specifically, our FCGFace is capable of adaptively adjusting the distribution of profile face features and narrowing the gap between them and frontal face features during different training stages to form compact identity clusters. Extensive experimental results on popular benchmarks, including cross-pose datasets (CFP-FP, CPLFW, VGGFace2-FP, and Multi-PIE) and non-cross-pose datasets (YTF, LFW, AgeDB-30, CALFW, IJB-B, IJB-C, and RFW), have demonstrated the superiority of our FCGFace over the SOTA competitors.
Yingfan Tao, Wenxian Zheng, Wenming Yang, Guijin Wang, Qingmin Liao
IEEE Trans. Inf. Forensics Secur.4
2021 FETNet: Feature Exchange Transformer Network for RGB-D Object Detection
Jing-Hao Xue, Pengwei Xie, Guijin Wang
BMVC4
2021 Ts-Unet: A Temporal Smoothed Unet for Video Anomaly Detection
Zhongliang Yang, Guijin Wang
ICIG (3)5
2021 Enhance Via Decoupling: Improving Multi-Label Classifiers With Variational Feature Augmentation
abstract
Multi-label classification remains a challenging problem due to the inherent label imbalance issue, which brings overfitting of minor categories to modern deep models. In this paper, to tackle this issue, we propose a novel method named Variational Feature Augmentation (VFA) to enhance the deep neural networks for multi-label classification. Our method decouples the feature vectors extracted by the backbone network into multiple low-dimensional spaces via a novely proposed Variational Feature Decoupling Module. The decoupled feature vectors are then re-combined with a shuffle operation and a Feature Augmentation Layer to enrich the minor co-occurrence relations, mitigating the label imbalance. Different from most other methods, VFA does not modify the network architecture or introduce extra computation cost in inference phase. We conduct comprehensive experiments on four benchmarks of two visual multi-label classification tasks, pedestrian attribute recognition and multi-label image recognition, and the results demonstrate the effectiveness and generality of the proposed VFA.
Guijin Wang, Jing-Hao Xue, Zijian Ding
ICIP2
2021 Triplet Angular Loss for Pose-Robust Face Recognition
abstract
Although face recognition has been widely applied in many areas, pose-robust face recognition is still a challenging topic due to the large pose variations in real scenes. In this paper, we propose to learn the pose-robust face representation by normalizing the profile face in feature level directly and jointly considering both intra-class compactness and inter-class separability. Our approach minimizes the angular distance between the profile face and the positive frontal anchor. And it maximizes the angular distance between the profile face and the negative frontal anchor simultaneously. Furthermore, we modify the Triplet loss and derive the Triplet Angular loss to guarantee the intra-class compactness and the inter-class separability in angular space. In this way, the faces under varying poses can cluster compactly to create a pose-robust feature representation. Extensive experiments on two challenging benchmarks (CFP-FP and IJB-A) illustrate that our approach achieves a competitive performance in the field of pose-robust face recognition.
Zhenduo Zhang, Yongru Chen, Wenming Yang, Guijin Wang, Qingmin Liao
IJCNN4
2021 Inter-patient ECG arrhythmia heartbeat classification based on unsupervised domain adaptation
Guijin Wang, Zijian Ding, Huazhong Yang
Neurocomputing1
2021 Generalisations of stochastic supervision models
Xiaoou Lu, Yangqi Qiao, Rui Zhu 0006, Guijin Wang, Zhanyu Ma, Jing-Hao Xue
Pattern Recognit.4
2021 CLECG: A Novel Contrastive Learning Framework for Electrocardiogram Arrhythmia Classification
abstract
Deep learning-based intelligent electrocardiogram (ECG) diagnosis algorithms heavily rely on large annotated datasets. Unfortunately, in the context of ECG diagnosis, privacy issues and the high cost of data annotations lead to a shortage of ECG datasets which severely limits the performance of the state-of-the-art ECG diagnosis algorithms. In this paper, we propose a novel instance-level contrastive learning scheme for ECG signals, namely CLECG, to mine effective information from unlabeled data. During the pre-training, CLECG encourages the representations of different augmented views of the same signal (positive samples) to be similar and increases the distance between representations of augmented views from the different signals (negative samples). The whole pre-training process does not require any form of labeling. Experimental results show that the proposed CLECG strategy outperforms other self-supervised methods and supervised transfer learning strategies.
Guijin Wang, Guodong Zhang 0004, Huazhong Yang
IEEE Signal Process. Lett.2
2021 Mixup Asymmetric Tri-Training for Heartbeat Classification Under Domain Shift
abstract
Due to the significant variability in waveforms and characteristics of ECG signals, developing fully automatic (i.e., requires no expert assistance) heartbeat classification algorithms with satisfactory performance on domain-shifted data remains challenging. In this letter, we propose a novel Mixup Asymmetric Tri-training (MIAT) method to improve the generalization ability of heartbeat classifiers in domain shift scenarios. First, we develop an ECG-based tri-branch CNN model, including one shared feature encoder followed by three branch networks. Next, to obtain target-discriminative features progressively, the tri-branch CNN is trained asymmetrically in each domain adaptation cycle, where two branches are used to assign pseudo-labels to the target domain samples and the third branch is trained on these pseudo-labeled target samples. Moreover, three kinds of mixup regularizations are incorporated into the training process. Experimental results on MITDB and SVDB show that the proposed MIAT outperforms the state-of-the-art methods in terms of F1-macro score and demonstrate the effectiveness of each mixup regularization.
Guijin Wang, Zijian Ding, Huazhong Yang
IEEE Signal Process. Lett.2
2021 Epipolar Geometry Guided Highly Robust Structured Light 3D Imaging
abstract
Structured light (SL) based three-dimensional (3D) imaging technology has been widely employed in many fields of computer vision. However, currently available SL based depth sensors are sensitive to imaging noises, which severely limits the performance of subsequent advanced vision tasks. In this letter, we propose a robust and practical SL illumination pattern coding method based on epipolar geometry. The proposed pattern can effectively alleviate coding redundancy in traditional global random speckle SL patterns and make the stereo matching more robust to noise. Meanwhile, this coding strategy supplies sufficient non-local similar blocks, which inspires us to propose a uni-direction block-stacking 3D filtering algorithm to further improve the 3D imaging quality. To verify the proposed algorithms, we developed a prototype using the off-the-shelf projector and camera. Both simulation and real scene experimental results show that the proposed methods can achieve high-quality 3D imaging performance under different noisy conditions.
Guijin Wang, Chenchen Feng, Xiaowei Hu 0004, Huazhong Yang
IEEE Signal Process. Lett.1
2021 Non-Local Aggregation for RGB-D Semantic Segmentation
abstract
Exploiting both RGB (2D appearance) and Depth (3D geometry) information can improve the performance of semantic segmentation. However, due to the inherent difference between the RGB and Depth information, it remains a challenging problem in how to integrate RGB-D features effectively. In this letter, to address this issue, we propose a Non-local Aggregation Network (NANet), with a well-designed Multi-modality Non-local Aggregation Module (MNAM), to better exploit the non-local context of RGB-D features at multi-stage. Compared with most existing RGB-D semantic segmentation schemes, which only exploit local RGB-D features, the MNAM enables the aggregation of non-local RGB-D information along both spatial and channel dimensions. The proposed NANet achieves comparable performances with state-of-the-art methods on popular RGB-D benchmarks, NYUDv2 and SUN-RGBD.
Guodong Zhang 0004, Jing-Hao Xue, Pengwei Xie, Sifan Yang, Guijin Wang
IEEE Signal Process. Lett.5
2021 Class-Variant Margin Normalized Softmax Loss for Deep Face Recognition
abstract
In deep face recognition, the commonly used softmax loss and its newly proposed variations are not yet sufficiently effective to handle the class imbalance and softmax saturation issues during the training process while extracting discriminative features. In this brief, to address both issues, we propose a class-variant margin (CVM) normalized softmax loss, by introducing a true-class margin and a false-class margin into the cosine space of the angle between the feature vector and the class-weight vector. The true-class margin alleviates the class imbalance problem, and the false-class margin postpones the early individual saturation of softmax. With negligible computational complexity increment during training, the new loss function is easy to implement in the common deep learning frameworks. Comprehensive experiments on the LFW, YTF, and MegaFace protocols demonstrate the effectiveness of the proposed CVM loss function.
Wanping Zhang, Yongru Chen, Wenming Yang, Guijin Wang, Jing-Hao Xue, Qingmin Liao
IEEE Trans. Neural Networks Learn. Syst.4
2020 Adaptive Region Aggregation Network: Unsupervised Domain Adaptation with Adversarial Training for ECG Delineation
abstract
Electrocardiogram (ECG) delineation, which provides clinically useful information for the diagnosis of cardiovascular disease, is an essential task in automated ECG analysis. The discrepancies among ECG signals from different datasets, namely domain shifts, may bring severe challenges to the cross-dataset performance of ECG delineation algorithms. The domain shifts are generally caused by the differences of conditions, collecting devices, and individual characteristics, and are inherent and non-negligible in ECG. In this work, we propose an unsupervised domain adaptation method called Adaptive Region Aggregation Network (ARAN) based on adversarial training to tackle domain shift problem in ECG delineation. The proposed algorithm promotes the state- of-the-art deep neural network RAN[1] to learn domain- invariant features and achieve improving performance on both source and target domain. The experiments results on two public datasets, LUDB and QT database, prove that our approach can effectively improve the cross-dataset performance of the state-of-the-art deep learning model.
Guijin Wang, Zijian Ding
ICASSP2
2020 Weakly Supervised Segmentation Guided Hand Pose Estimation During Interaction with Unknown Objects
abstract
Hand pose estimation is important for human computer interaction, but the performance is not satisfying when the hand is interacting with objects. To alleviate the influence of unknown objects, we propose a novel weakly supervised segmentation guided scheme to estimate hand poses. Approximate hand masks generated from annotations of sparse hand joints are used to supervise the segmentation task. Better features can be extracted since they are shared between the two tasks of hand segmentation and hand pose estimation. With the guidance of weakly supervised segmentation, the network can learn intermediate features balanced between focusing on the foreground and preserving contextual information. Finally the xy and z coordinates are estimated in different branches but utilizing shared feature maps. Experimental results of three different tasks on the publicly available FHAD dataset demonstrate the effectiveness of the proposed architecture.
Cairong Zhang, Guijin Wang, Xinghao Chen 0001, Pengwei Xie, Toshihiko Yamasaki
ICASSP2
2020 Emotion Recognition with Facial Landmark Heatmaps
Siyi Mo, Wenming Yang, Guijin Wang, Qingmin Liao
MMM (1)3
2020 Pose guided structured region ensemble network for cascaded hand pose estimation
Xinghao Chen 0001, Guijin Wang, Hengkai Guo, Cairong Zhang
Neurocomputing2
2020 Deep learning for image super-resolution
Wenming Yang, Fei Zhou 0001, Rui Zhu 0006, Kazuhiro Fukui, Guijin Wang, Jing-Hao Xue
Neurocomputing5
2020 Bi-Stream Pose-Guided Region Ensemble Network for Fingertip Localization From Stereo Images
abstract
In human-computer interaction, it is important to accurately estimate the hand pose, especially fingertips. However, traditional approaches to fingertip localization mainly rely on depth images and thus suffer considerably from noise and missing values. Instead of depth images, stereo images can also provide 3-D information of hands. There are nevertheless limitations on the dataset size, global viewpoints, hand articulations, and hand shapes in publicly available stereo-based hand pose datasets. To mitigate these limitations and promote further research on hand pose estimation from stereo images, we build a new large-scale binocular hand pose dataset called THU-Bi-Hand, offering a new perspective for fingertip localization. In the THU-Bi-Hand dataset, there are 447k pairs of stereo images of different hand shapes from ten subjects with accurate 3-D location annotations of the wrist and five fingertips. Captured with minimal restriction on the range of hand motion, the dataset covers a large global viewpoint space and hand articulation space. To better present the performance of fingertip localization on THU-Bi-Hand, we propose a novel scheme termed bi-stream pose-guided region ensemble network (Bi-Pose-REN). It extracts more representative feature regions around joints in the feature maps under the guidance of the previously estimated pose. The feature regions are integrated hierarchically according to the topology of hand joints to regress a refined hand pose. Bi-Pose-REN and several existing methods are evaluated on THU-Bi-Hand so that benchmarks are provided for further research. Experimental results show that our Bi-Pose-REN has achieved the best performance on THU-Bi-Hand.
Guijin Wang, Cairong Zhang, Xinghao Chen 0001, Xiangyang Ji, Jing-Hao Xue
IEEE Trans. Neural Networks Learn. Syst.1
2019 Blurring-Effect-Free CNN for Optimization of Structural Edges in Focus Stacking
abstract
Focus stacking is a computational technique to extend the Depth of Field (DOF) through combining multiple images taken at various focus distances. In depth-estimation based focus stacking approaches, researchers reconstruct the all-in-focus image by extracting pixels from focal stack pixel-by-pixel based on estimated depthmap. However, existing methods could not cope with the blurring-effect problem in structural edges, where depth values change abruptly. In this work, we propose a novel convolutional neural network (BEF-CNN) to reconstruct blurring-effect-free patches from focal stack to enhance performance of all-in-focus image. To the best of our knowledge, it is the first work to utilize CNN to generate all-in-focus RGB images directly instead of pixel-to-pixel correspondence from depthmap. Experimental results validate that the proposed algorithm could achieve best all-in-focus reconstruction performance.
Guijin Wang, Xinghao Chen 0001, Xuanwu Yin, Xiaowei Hu 0004
ICIP2
2019 A global and updatable ECG beat classification system based on recurrent neural networks and active learning
Guijin Wang, Chenshuang Zhang, Yongpan Liu, Huazhong Yang, Dapeng Fu
Inf. Sci.1
2018 Bi-stream Region Ensemble Network: Promoting Accuracy in Fingertip Localization from Stereo Images
Cairong Zhang, Guijin Wang, Xinghao Chen 0001, Huazhong Yang
BMVC2
2018 Depth-Based 3D Hand Pose Estimation: From Current Achievements to Future Goals
abstract
In this paper, we strive to answer two questions: What is the current state of 3D hand pose estimation from depth images? And, what are the next challenges that need to be tackled? Following the successful Hands In the Million Challenge (HIM2017), we investigate the top 10 state-of-the-art methods on three tasks: single frame 3D pose estimation, 3D hand tracking, and hand pose estimation during object interaction. We analyze the performance of different CNN structures with regard to hand shape, joint visibility, view point and articulation distributions. Our findings include: (1) isolated 3D hand pose estimation achieves low mean errors (10 mm) in the view point range of [70, 120] degrees, but it is far from being solved for extreme view points; (2) 3D volumetric representations outperform 2D CNNs, better capturing the spatial structure of the depth data; (3) Discriminative methods still generalize poorly to unseen hand shapes; (4) While joint occlusions pose a challenge for most methods, explicit modeling of structure constraints can significantly narrow the gap between errors on visible and occluded joints.
Shanxin Yuan, Guillermo Garcia-Hernando, Björn Stenger, Gyeongsik Moon, Ju Yong Chang, Kyoung Mu Lee, Pavlo Molchanov 0001, Jan Kautz, Sina Honari, Liuhao Ge, Junsong Yuan 0001, Xinghao Chen 0001, Guijin Wang, Fan Yang 0032, Kai Akiyama, Yang Wu 0001, Qingfu Wan, Meysam Madadi, Sergio Escalera, Shile Li, Dongheui Lee, Iasonas Oikonomidis, Antonis A. Argyros, Tae-Kyun Kim 0001
CVPR13
2018 Scene-Adaptive Image Acquisition for Focus Stacking
abstract
Focus stacking is a promising technique to extend depth of field in general photography by fusing images captured at different focusing distances. In this paper, we propose a round-trip scene-adaptive image acquisition system to automatically capture focal stack and fuse a high quality all-in-focus image. Based on scene analysis, we cover entire depth range of the scene in the forward optical scanning and refine all objects' focusing positions accurately in the backward scanning. With captured images, we firstly extract depthmap and all-in-focus image with combination of max-gradient flow and blur kernel estimation. Secondly, a superpixel-level Gaussian Fitting is proposed to determine the next location to capture. Experiments on simulated data show that our method attain high quality all-in-focus image with fewer captured images.
Guijin Wang, Xiaowei Hu 0004, Huazhong Yang
ICIP2
2018 Region ensemble network: Towards good practices for deep 3D hand pose estimation
Guijin Wang, Xinghao Chen 0001, Hengkai Guo, Cairong Zhang
J. Vis. Commun. Image Represent.1
2018 All-in-focus with directional-max-gradient flow and labeled iterative depth propagation
Guijin Wang, Xuanwu Yin, Huazhong Yang
Pattern Recognit.1
2018 LRID: A new metric of multi-class imbalance degree based on likelihood-ratio test
abstract
In this paper, we introduce a new likelihood ratio imbalance degree (LRID) to measure the class-imbalance extent of multi-class data. Imbalance ratio (IR) is usually used to measure class-imbalance extent in imbalanced learning problems. However, IR cannot capture the detailed information in the class distribution of multi-class data, because it only utilises the information of the largest majority class and the smallest minority class. Imbalance degree (ID) has been proposed to solve the problem of IR for multi-class data. However, we note that improper use of distance metric in ID can have harmful effect on the results. In addition, ID assumes that data with more minority classes are more imbalanced than data with less minority classes, which is not always true in practice. Thus ID cannot provide reliable measurement when the assumption is violated. In this paper, we propose a new metric based on the likelihood-ratio test, LRID, to provide a more reliable measurement of class-imbalance extent for multi-class data. Experiments on both simulated and real data show that LRID is competitive with IR and ID, and can reduce the negative correlation with F1 scores by up to 0.55.
Rui Zhu 0006, Ziyu Wang 0003, Zhanyu Ma, Guijin Wang, Jing-Hao Xue
Pattern Recognit. Lett.4
2017 Depth-Based Focus Stacking with Labeled-Laplacian Propagation
Guijin Wang, Xuanwu Yin, Xiaowei Hu 0004, Huazhong Yang
ICIG (3)2
2017 Motion feature augmented recurrent neural network for skeleton-based dynamic hand gesture recognition
abstract
Dynamic hand gesture recognition has attracted increasing interests because of its importance for human computer interaction. In this paper, we propose a new motion feature augmented recurrent neural network for skeleton-based dynamic hand gesture recognition. Finger motion features are extracted to describe finger movements and global motion features are utilized to represent the global movement of hand skeleton. These motion features are then fed into a bidirectional recurrent neural network (RNN) along with the skeleton sequence, which can augment the motion features for RNN and improve the classification performance. Experiments demonstrate that our proposed method is effective and outperforms start-of-the-art methods.
Xinghao Chen 0001, Hengkai Guo, Guijin Wang, Li Zhang 0023
ICIP3
2017 Region ensemble network: Improving convolutional network for hand pose estimation
abstract
Hand pose estimation from monocular depth images is an important and challenging problem for human-computer interaction. Recently deep convolutional networks (ConvNet) with sophisticated design have been employed to address it, but the improvement over traditional methods is not so apparent. To promote the performance of directly 3D coordinate regression, we propose a tree-structured Region Ensemble Network (REN), which partitions the convolution outputs into regions and integrates the results from multiple regressors on each regions. Compared with multi-model ensemble, our model is completely end-to-end training. The experimental results demonstrate that our approach achieves the best performance among state-of-the-arts on two public datasets.
Hengkai Guo, Guijin Wang, Xinghao Chen 0001, Cairong Zhang, Fei Qiao, Huazhong Yang
ICIP2
2017 A Robust Static Sign Language Recognition System Based on Hand Key Points Estimation
Feng Chen 0007, Guijin Wang, Jinsheng Ren, Jianwu Dong
ISDA3
2017 Two-stream binocular network: Accurate near field finger detection based on binocular images
abstract
Fingertip detection plays an important role in human computer interaction. Previous works transform binocular images into depth images. Then depth-based hand pose estimation methods are used to predict 3D positions of fingertips. Different from previous works, we propose a new framework, named Two-Stream Binocular Network (TSBnet) to detect fingertips from binocular images directly. TSBnet first shares convolutional layers for low level features of right and left images. Then it extracts high level features in two-stream convolutional networks separately. Further, we add a new layer: binocular distance measurement layer to improve performance of our model. To verify our scheme, we build a binocular hand image dataset, containing about 117k pairs of images in training set and 10k pairs of images in test set. Our methods achieve an average error of 10.9mm on our test set, outperforming previous work by 5.9mm (relatively 35.1%).
Guijin Wang, Cairong Zhang, Hengkai Guo, Xinghao Chen 0001, Huazhong Yang
VCIP2
2016 Accurate fingertip detection from binocular mask images
abstract
Accurate fingertip detection is important for hand-based human computer interaction. Different from prior methods based on depth image, in this paper we propose a novel scheme for accurate fingertip detection from binocular mask images without explicitly computing the depth map. To demonstrate our proposed scheme, we build a new hand dataset containing synthetic and real binocular images. A deep convolutional neural network (CNN) is utilized as a baseline method to demonstrate the proposed scheme. The mask images are extracted from binocular images and fed into the CNN to predict the 3D positions of fingertips and palm center. Experiments show that this method achieves the mean error of 8.30mm on synthetic data and 4.64mm on real data, which is promosing for accurate 3D interaction. The proposed scheme runs at 130fps on a CPU and is promising for real-time applications.
Xinghao Chen 0001, Guijin Wang, Hengkai Guo
VCIP2
2016 Latent variable pictorial structure for human pose estimation on depth images
Guijin Wang, Qingmin Liao, Jing-Hao Xue
Neurocomputing2
2016 A novel hierarchical framework for human action recognition
Hongzhao Chen, Guijin Wang, Jing-Hao Xue
Pattern Recognit.2
2015 Depth-images-based pose estimation using regression forests and graphical models
Guijin Wang, Qingmin Liao, Jing-Hao Xue
Neurocomputing2
2015 Embedding metric learning into set-based face recognition for video surveillance
Guijin Wang, Chenbo Shi, Jing-Hao Xue
Neurocomputing1
2015 Subcategory Clustering with Latent Feature Alignment and Filtering for Object Detection
abstract
For objects with large appearance variations, it has been proved that their detection performance can be effectively improved by clustering positive training instances into subcategories and learning multi-component models for the subcategories. However, it is not trivial to generate subcategories of high quality, due to the difficulty in measuring the similarity between positive instances. In this letter we propose a new weakly supervised clustering method to achieve better sub-categorization. Our method provides a more precise measurement of the similarity by aligning the positive instances through latent variables and filtering the aligned features. As a better alternative to the initialization step of the latent-SVM algorithm for the learning of the multi-component models, our method can lead to a superior performance gain for object detection. We demonstrate this on various real-world datasets.
Zhiwei Ruan, Guijin Wang, Jing-Hao Xue, Xinggang Lin
IEEE Signal Process. Lett.2
2015 High-Accuracy Stereo Matching Based on Adaptive Ground Control Points
abstract
This paper proposes a novel high-accuracy stereo matching scheme based on adaptive ground control points (AdaptGCP). Different from traditional fixed GCP-based methods, we consider color dissimilarity, spatial relation, and the pixel-matching reliability to select GCP adaptively in each local support window. To minimize the global energy, we propose a practical solution, named as alternating updating scheme of disparity and confidence map, which can effectively eliminate the redundant and interfering information of unreliable pixels. The disparity values of those unreliable pixels are reassigned with the information provided by local plane model, which is fitted with GCPs. Then, the confidence map is updated according to the disparity reassignment and the left-right consistency. Finally, the disparity map is refined by multistep filers. Quantitative evaluations demonstrate the effectiveness of our AdaptGCP scheme for regularizing the ill-posed matching problem. The top ranks on Middlebury benchmark with different error thresholds show that our algorithm achieves the state-of-the-art performance among the latest stereo matching algorithms. This paper provides a new insight toward high-accuracy stereo matching.
Chenbo Shi, Guijin Wang, Xuanwu Yin, Xiaokang Pei, Bei He, Xinggang Lin
IEEE Trans. Image Process.2
2014 Detection of user-registered dog faces
Zhiwei Ruan, Guijin Wang, Jing-Hao Xue, Xinggang Lin
Neurocomputing2
2014 Iterative transductive learning for automatic image segmentation and matting with RGB-D data
Bei He, Guijin Wang, Cha Zhang
J. Vis. Commun. Image Represent.2
2013 POP: Person Re-identification Post-rank Optimisation
abstract
Owing to visual ambiguities and disparities, person re-identification methods inevitably produce sub optimal rank-list, which still requires exhaustive human eyeballing to identify the correct target from hundreds of different likely-candidates. Existing re-identification studies focus on improving the ranking performance, but rarely look into the critical problem of optimising the time-consuming and error-prone post-rank visual search at the user end. In this study, we present a novel one-shot Post-rank Optimization (POP) method, which allows a user to quickly refine their search by either "one-shot" or a couple of sparse negative selections during a re-identification process. We conduct systematic behavioural studies to understand user's searching behaviour and show that the proposed method allows correct re-identification to converge 2.6 times faster than the conventional exhaustive search. Importantly, through extensive evaluations we demonstrate that the method is capable of achieving significant improvement over the state-of-the-art distance metric learning based ranking models, even with just "one shot" feedback optimisation, by as much as over 30% performance improvement for rank 1 re-identification on the VIPeR and i-LIDS datasets.
Chen Change Loy, Shaogang Gong, Guijin Wang
ICCV4
2013 Iterative transductive learning for alpha matting
abstract
In this paper, we propose a matting algorithm based on iterative transductive learning (for short: ITM). To avoid over-smooth results of recent methods, we introduce the influence of unlabeled regions as well as the consistency of neighboring pixels to re-design the optimization for alpha matting. A novel asymmetric Laplacian matrix is also proposed to further relieve the over-smoothness. To optimize the matting problem, we adjust the constrain coefficients between the initialized alpha matte and the asymmetric Laplacian matrix iteratively to achieve accurate alpha mattes. Consequently, during the iteration, high confidence pixels maintain their refined alpha values, whereas low confidence ones are updated by their neighbors gradually. Experimental results demonstrate that our algorithm is more precise than many state-of-the-art methods in terms of the accuracy.
Bei He, Guijin Wang, Chenbo Shi, Xuanwu Yin, Xinggang Lin
ICIP2
2012 Data level object detector adaptation with online multiple instance samples
abstract
In object detection, the offline trained detector's performance may be degraded in a particular deployed environment, because of the large variation of different environments. In this work, we propose a data level object detector adaptation method to new environments. By recording a small amount of offline data, it's fully compatible with offline training method and easy to implement. We re-derive an efficient MILBoost by eliminating line search in optimization and introduce it to collect online multiple instance samples, which don't require strict sample alignment. Experiment results with the human detector on public datasets illustrate the effectiveness of the proposed adaptation method. The adapted detector has good adaptation ability, while maintaining its generalization ability as well.
Bobo Zeng, Guijin Wang, Zhiwei Ruan, Xinggang Lin
ICASSP2
2012 Local matting based on sample-pair propagation and iterative refinement
abstract
This paper proposes a novel local matting algorithm based on sample-pair propagation and iterative refinement. Since sample-pairs of the foreground and background in the neighborhood are limited, they fail to fit the linear model well. We propose a sample-pair propagation scheme which propagates the confident sample-pair of each pixel to its neighbors so that they can collect more confident sample-pairs to estimate alpha values accurately. To avoid high time and space complexity of the global optimization, we convert matting into a de-noising problem and refine alpha values via fitting the linear model and smoothing the alpha matte locally and iteratively. Experimental results demonstrate that our algorithm produces more accurate results than the state-of-the-art of local matting.
Bei He, Guijin Wang, Zhiwei Ruan, Xuanwu Yin, Xiaokang Pei, Xinggang Lin
ICIP2
2012 A compact association of particle filtering and kernel based object tracking
Anbang Yao, Xinggang Lin, Guijin Wang
Pattern Recognit.3
2011 Multiple instance tracking based on hierarchical maximizing bag's margin boosting
abstract
In online tracking, the tracker evolves to reflect variations in object appearance and surroundings. This updating process is formulated as a supervised learning problem, thus a slight inaccuracy of the tracker will degrade the updating. Multiple Instance Learning (MIL) is used to alleviate such a problem by representing training samples in bags of image patches (or called instances). Difficulties are then passed on to the learning method to train a classifier that discovers the most accurate instance. This paper proposes a Maximizing Bag's Margin (MBM) criteria for MIL. Combined with MBM, a hierarchical boosting is proposed for updating, in which bag and instance weights are introduced to guide classifier retrain ing. Our approach effectively improves the updating's efficiency with less computation cost. Experiments demonstrate the benefits of our method.
Guijin Wang, Xinggang Lin, Bobo Zeng
ICASSP2
2011 A new framework for on-line object tracking based on SURF
Quan Miao, Guijin Wang, Chenbo Shi, Xinggang Lin, Zhiwei Ruan
Pattern Recognit. Lett.2
2010 Anomaly detection in surveillance video using motion direction statistics
abstract
A novel approach for detecting anomaly in visual surveillance system is proposed in this paper. It is composed of three parts:(a) a dense motion field and motion statistics method, (b) one-class SVM for one-class classification, (c) motion directional PCA for feature dimensionality reduction. Experiments demonstrate the effectiveness of proposed algorithm in detecting abnormal events in surveillance video, while keeping a low false alarm rate. Moreover, it works well in complicated situation where the common tracking or detection module won't work.
Guijin Wang, Wenxin Ning, Xinggang Lin
ICIP2
2010 Scale and rotation invariant feature-based object tracking via modified on-line boosting
abstract
Object tracking is a major technique in image processing and computer vision. In this paper, we propose a new robust feature-based tracking scheme by employing adaptive classifiers to match the detected keypoints in consecutive frames. The novelty of this paper is that the design of online boosting is combined with the invariance of local features so that the classifier-based descriptions are formed in association with the scale and rotation information. Furthermore, we introduce a sample weighting mechanism in the on-line classifier updating, for the subsequent tracking. Experimental results demonstrate the robustness and accuracy of our proposed technique.
Quan Miao, Guijin Wang, Xinggang Lin, Chenbo Shi, Chao Liao
ICIP2
2010 Topology based affine invariant descriptor for MSERs
abstract
This paper introduces a topology based affine invariant descriptor for maximally stable extremal regions (MSERs). The popular SIFT descriptor computes the texture information on a grey-scale patch. Instead our descriptor use only the topology and geometric information among MSERs so that features can be rapidly matched regardless of the texture in the image patch. Based on the ellipses fitting for the detected MSERs, geometric affine invariants between ellipses pair are extracted as the descriptors. Finally topology based voting selector is designed to achieve the best correspondences. Experiment shows that our descriptor is not only computational faster than SIFT descriptor, but also has better performance on wide angle of view and nonlinear illumination change. In addition, our descriptor shows a good result on multi sensor images registration.
Chenbo Shi, Guijin Wang, Xinggang Lin, Chao Liao, Quan Miao
ICIP2
2010 An incremental Bhattacharyya dissimilarity measure for particle filtering
Anbang Yao, Guijin Wang, Xinggang Lin, Xiujuan Chai
Pattern Recognit.2
2009 Hand posture recognition in video using multiple cues
abstract
Hand posture conveys profound information for computer vision applications, but the articulated hand structure and restraint capture condition cast a tough obstacle on practical implementation, especially in real time video. This paper presents a framework to recognize hand postures in consecutive video frames. Mixture of Gaussian skin/non skin models is constructed for hand region detection, followed by particle filter to track hand. Then a soft-decision scheme based on extended Histogram of Orientated gradient is proposed to refine the best posture region and recognize it from pre-defined posture set. Experimental result shows promising performance under various capture conditions.
Liang Sha, Guijin Wang, Anbang Yao, Xinggang Lin, Xiujuan Chai
ICME2
2009 Recovery of upper body poses in static images based on joints detection
Zhilan Hu, Guijin Wang, Xinggang Lin, Hong Yan 0001
Pattern Recognit. Lett.2
2008 Kernel based articulated object tracking with scale adaptation and model update
abstract
Kernel based object tracking (KBOT) is one of the most popular and effective techniques for tracking task. However the constancy of the target model and unsound scale adaptation method are two main limitations. In this paper, we present a kernel based approach incorporated with scale estimation and target model update for articulated object tracking task. After predicating the object center with scale fixed KBOT, we extend scale selection theory to estimate the local optimal object scale. Once the object scale has been estimated, a kernel density estimation based strategy is developed to update the target model. Experimental results show that our approach is superior to traditional KBOT in the following two aspects: 1) it is less affected by the object scale change; 2) it is less prone to appearance variation.
Anbang Yao, Guijin Wang, Xinggang Lin
ICASSP2
2007 Enhanced Shot Change Detection using Motion Features for Soccer Video Analysis
abstract
An enhanced shot change detection algorithm for soccer video analysis is proposed in this paper. Features extracted from reliable motion vectors (MVs) are used to enhance the accuracy and efficiency of color-based method. In order to eliminate unreliable MVs, a MV filtration scheme is designed based on the principle of block-based motion search. Then two features, named as proportion of reliable MVs and centrality of reliable MVs respectively, are extracted from the filtration results in each frame. The experiments on 180 minutes' soccer video demonstrate that our algorithm achieved a considerable improvement over the widely-adopted color-based method.
Yichuan Hu, Guijin Wang, Xinggang Lin
ICME3
2006 Practical rate control for video over WLAN
Guijin Wang, Satoshi Futemma, Masato Kawada, Eisaburo Itakura
CCNC1
2005 JPEG2000 based real-time scalable video communication system over the Internet
abstract
This paper presents a JPEG2000 based real-time scalable video communication system developed in Sony, named as "BEAM" video system. Two practical objectives are emphasized in this system: scalable video communication and real-time communication. We extend IETF RTP and RTSP streaming protocol to take care of scalable delivery of video data from a single layered coding data to heterogeneous devices and different resolution display such as HDTV, standard TV, PDA and mobile phone. Utilizing an advantage of low coding delay JPEG2000 codec, real-time ARQ (RT-ARQ) is proposed to resolve the QoS issues for real-time applications. In addition to this protocol, various network adaptive control techniques, including rate control and error control, are presented to achieve high quality video transmission.
Eisaburo Itakura, Satoshi Futemma, Guijin Wang, Kenji Yamane
CCNC3
2005 FEC-based scalable multiple description coding for overlay network streaming
abstract
We consider the problem of distributing video data from a server to a population of interested clients. Deploying overlay network with application-layer multicast, we propose a FEC-based scalable multiple description coding (SMDC) scheme to arrive at the solution to the problem. Without the performance penalty on each receiver, our SMDC can achieve optimal performance of each client with different network conditions as simply as additional assembling and truncating operations. In the mean time, a R-D allocation algorithm is presented to decide the rates of the source and channel. Simulation shows that our SMDC have better quality for heterogeneous receivers than traditional scheme.
Guijin Wang, Satoshi Futemma, Eisaburo Itakura
CCNC1
2004 An end-to-end robust approach for scalable video over the internet
abstract
This paper introduces an end-to-end robust approach for scalable video over the Internet. The traditional method only considers congestion control, error control and is unable to achieve end-to-end high-quality video transmission in the error-prone environment like the Internet since it does not consider the packetization behavior, network conditions and the media characteristics simultaneously. This paper presents an end-to-end approach for scalable video over the Internet, combining network adaptive congestion control and unequal error control. Considering requirements of multimedia transmission, this paper introduces multimedia congestion control to estimate available bandwidth and smooth the media sending rate. Specially in the transport layer we propose unequal interleaving packetization method and unequal error protection scheme, which can alleviate the effect of the packet loss well. Further we develop the rate-distortion theory for the scalable video over the Internet. Thereafter the optimal bit allocation is presented to determine the bits budgets for the source part and error control part. Simulation shows our scheme can achieve good performance for scalable video over the Internet. Copyright by Science in China Press 2004.
Guijin Wang, Qian Zhang 0001, Wenwu Zhu 0001, Xinggang Lin
Sci. China Ser. F Inf. Sci.1
2004 Error robust scalable audio streaming over wireless IP networks
abstract
Streaming high-fidelity audio over wireless Internet protocol (IP) networks is a challenging task because the networks present not only packet losses, but also residual bit errors. These losses and errors have severe adverse effect on the compressed audio bitstream. To solve this problem, this paper introduces error resilience in conjunction with error protection for scalable audio streaming over wireless networks. Specifically, error resilience is achieved by performing bitstream data partitioning and reversible variable length coding in the audio coder. Error protection is provided by layered product channel code to simultaneously handle packet losses and residual bit errors. Both the row and column codes of the product code provide unequal error protection for different layers of the audio bitstream by considering the characteristics of the scalable audio. Rate-distortion optimization is performed to determine the best source-channel coding tradeoff that minimizes the average expected end-to-end distortion. Simulation results demonstrate the effectiveness of our proposed approach.
Qian Zhang 0001, Guijin Wang, Zixiang Xiong, Jianping Zhou 0001, Wenwu Zhu 0001
IEEE Trans. Multim.2
2002 Error protection for scalable image over 3G-IP network
abstract
Digital media, like image and video, over third-generation wireless networks is a challenging task because the wireless networks present not only packet loss, but also bit errors. To address this problem, this paper proposes a novel error protection scheme for scalable image over 3G-IP networks. Taking into consideration the scalable nature of the image data, error protection is provided by layered product channel codes to mitigate the effect of the packet loss and bit errors. Meanwhile, rate-distortion optimization is performed to determine the protection levels of both the row channel codes and the column codes so as to minimize the expected end-to-end distortion. Simulation results demonstrate the effectiveness of our proposed approach.
Guijin Wang, Xinggang Lin
ICIP (2)1
2002 Qos-guarantee error control for scalable image over wireless fading channel
Guijin Wang, Xinggang Lin
VCIP1
2001 Channel-adaptive error protection for scalable audio streaming over wireless Internet
abstract
Streaming high-fidelity audio over networks with both bit errors and packet erasures is becoming increasingly important due to the emerging of wireless Internet. In this paper, we propose an end-to-end architecture for scalable audio streaming over wireless Internet. Considering the characteristic of scalable audio, a novel layered product code is presented to handle bit errors and packet losses simultaneously. Specifically, unequal row channel code and unequal column code are adopted for different layers of scalable audio based on their quality impacts. Moreover, rate-distortion based bit allocation is proposed to determine each channel-coding rate and the source-coding rate so as to minimize the expected end-to-end distortion. The simulation results demonstrate the effectiveness of our proposed error protection scheme.
Guijin Wang, Qian Zhang 0001, Wenwu Zhu 0001, Jianping Zhou 0001
GLOBECOM1
2001 Channel-adaptive unequal error protection for scalable video transmission over wireless channel
Guijin Wang, Qian Zhang 0001, Wenwu Zhu 0001, Ya-Qin Zhang
VCIP1
2000 Resource allocation with adaptive QoS for multimedia transmission over W-CDMA channels
abstract
This paper addresses the important issues of resource allocation and rate adaptation for multiple media, such as audio, video, email, and Web traffic, transmitted over W-CDMA (wideband code division multiple access) channel with adaptive QoS (quality of service) support. In order to have QoS support for different types of media, we develop an architecture combining the link layer with application layer controls. It consists of the following contributions: (1) an appropriate model to estimate the varying fading channel is proposed; (2) a hybrid delay-constrained ARQ (automatic repeat request) and UEP (unequal error protection) mechanism that dynamically adapt to the time-varying channel is presented to meet the QoS requirements for different applications; (3) a new resource allocation scheme that considers varying media characteristics is described to be adapted to changing bit error rate (BER) conditions. Simulation results demonstrate the effectiveness of our proposed scheme.
Qian Zhang 0001, Wenwu Zhu 0001, Guijin Wang, Ya-Qin Zhang
WCNC3