Baojin Huang

dblp:261/3660 · DBLP profile ↗
← Back
35ranked-venue papers
9as first author
31since 2021 · last 2025
0000-0002-4882-5787ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 5 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 4 first-author · 15 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 ProJudge: A Multi-Modal Multi-Discipline Benchmark and Instruction-Tuning Dataset for Mllm-Based Process Judges
Jiaxin Ai, Zhaopan Xu, Fanrui Zhang, Zizhen Li, Yukang Feng, Baojin Huang, Zhongyuan Wang 0001, Kaipeng Zhang
ICCV9
2025 Adversarial intensity awareness for robust object detection
Jikang Cheng, Baojin Huang, Zhen Han 0002, Zhongyuan Wang 0001
Comput. Vis. Image Underst.2
2025 Luminance decomposition and reconstruction for high dynamic range Video Quality Assessment
Jifan Yang, Zhongyuan Wang 0001, Baojin Huang, Jiaxin Ai, Yuhong Yang 0001, Jing Xiao 0004, Zixiang Xiong
Pattern Recognit.3
2025 DALFace: Dynamic Association Learning for Face Recognition
abstract
Face recognition owes its success to the availability of large-scale training data. Recent adaptive margin-based loss functions pay more attention to hard (misclassified) samples, resulting in more discriminative face embeddings. However, large-scale datasets inevitably include open-set noise samples, which are usually mistaken for hard samples by mining-based methods and thus mislead the training of the model. In this work, we redefine hard samples and further design a dynamic association learning strategy for mining hard samples while ignoring noise. We argue that the difficulty of recognizing a sample depends on both identity-related and objective factors. On one hand, intrinsic attributes such as facial structure and face shape inherently influence the ease of identity recognition. On the other hand, external factors, including pose, occlusion, and resolution, directly affect the recognizability of a sample. Particularly in the case of noise samples, although they pose challenges for the deep network similar to hard samples, should not be regarded as hard samples. To this end, we propose an associated prototype learning method to achieve an approximation of face identity difficulty by exploring the fitting trends of identity prototype. Furthermore, we design a dynamic sample learning method to distinguish noise samples from hard samples by observing the distance fluctuation from the class center during sample learning. All observations are integrated into the loss function through adaptive margins and sample weights. Extensive experiments and visualizations on several datasets demonstrate that our method significantly outperforms state-of-the-art counterparts.
Baojin Huang, Guangcheng Wang, Kui Jiang, Zhongyuan Wang 0001
IEEE Trans. Multim.1
2024 Unconventional Face Adversarial Attack
Baojin Huang, Zhen Han 0002, Dengshi Li
ICANN (2)2
2024 ICPR 2024 Competition on Resource-Limited Infrared Small Target Detection Challenge: Methods and Results
Boyang Li 0007, Xinyi Ying, Ruojing Li, Yongxian Liu, Yangsi Shi, Xin Zhang 0170, Mingyuan Hu, Yukai Zhang, Dongli Tang, Qiang Ling 0002, Zaiping Lin, Weidong Sheng, Chenxu Peng, Huoren Yang, Lingjie Liu, Zelin Shi, Yunpeng Liu 0001, Chuang Yu 0003, Jinmiao Zhao, Heng Xiang, Tianyu Li 0005, Minghang Zhou, Chenxi Lan, Dongyu Xi, Chaofan Qiao, Yupeng Gao, Yongxu Liu 0006, Deping Chen, Xiaopeng Song, Jiuping Yang, Zhaobing Qiu, Rixiang Ni, Changhai Luo, Shuyuan Zheng, Baojin Huang, Xiaoqi Zhou, Qingshan Guo, Dangxuan Wu, Haodong Zeng, Qiang Fu 0017, Yimian Dai, Renke Kou, Jian Song 0007, Changfeng Feng, Zihao Xiong, Mengxuan Xiao, Yingxu Liu, Quanyi Zhao
ICPR (34)51
2024 Unlabeled Data Assistant: Improving Mask Robustness for Face Recognition
abstract
The existing masked face recognition algorithms almost tend to adopt synthetic masked face datasets for training. However, these models are limited as they rely on existing mask augmentation methods, which contain few mask patterns and cannot simulate shadows and textures in realistic scenes. To overcome this limitation, we propose a semi-supervised face recognition framework to fully exploit unlabeled real masked face samples, improving the mask robustness of the recognition model. More specifically, unlike the original face embedding network, we design a part-aware network to explore multi-region face representation based on the face structure. In this way, we obtain multiple face sub-embeddings, which correspond to different regions of the face, including the upper half, the lower half and the whole. Crucially, we use the norm of the sub-embedding to represent the activation state of the facial region features. For the input unlabeled masked face image, we restrict the sub-embedding norm of its lower half to weaken the face feature representation of the occluded area. For normal face samples, their partial features are kept activated by maintaining the sub-embedding norm, which guides the deep network does not ignore the available information. Moreover, we employ the margin-based recognition loss for normal samples to ensure that the model is sufficiently discriminative for normal facial features. Extensive experimental results on both normal and real masked face datasets show that our approach significantly outperforms the state-of-the-arts. Code is available at https://github.com/Baojin-Huang/UFace.
Baojin Huang, Zhongyuan Wang 0001, Jifan Yang, Zhen Han 0002, Chao Liang 0001
IEEE Trans. Inf. Forensics Secur.1
2024 Deep Hashing Network With Hybrid Attention and Adaptive Weighting for Image Retrieval
abstract
Due to the low computational cost of Hamming distance, hashing-based image retrieval has been universally acknowledged. Therefore, it is becoming increasingly important to quickly generate high-precision hash codes (also hash features) from images. However, the existing deep hashing methods are vulnerable to image content variations; that is, it is difficult to generate stable and consistent hash codes for similar images. In addition, generating hash codes of different lengths requires retraining the model, which is expensive in training time. To address these problems, this paper proposes a deep hashing network (DHN) with a hybrid attention mechanism and adaptive weighting (HAAW) learning. It mainly consists of a feature extraction module, feature refinement module, classification layer, hash layer and an adaptive weight layer. In particular, the hybrid attention mechanism combines bottom-up pixel saliency and top-down semantic constraints, in which the former is achieved through channel and spatial attention (CSA) and the latter is supervised by classification labels. In this way, it encourages the network to focus on dominant semantic features without being disturbed by irrelevant objects so that semantically similar images can be mapped to approximate hash codes. We further propose an adaptive weighting learning algorithm to generate weights for each bit of the hash code generated by the deep network. Then, we directly generate shorter hash codes from the available long hash code according to the importance of bits represented by the weights. This avoids retraining the network for learning hash codes of different lengths. Extensive experiments on public CIFAR-10, NUS_WIDE and ImageNet datasets show that our method has achieved substantial improvements over the counterparts in terms of precision and speed.
Yingjiao Pei, Zhongyuan Wang 0001, Heling Chen, Baojin Huang, Weiping Tu
IEEE Trans. Multim.5
2024 Person-action Instance Search in Story Videos: An Experimental Study
abstract
Person-Action instance search (P-A INS) aims to retrieve the instances of a specific person doing a specific action, which appears in the 2019–2021 INS tasks of the world-famous TREC Video Retrieval Evaluation (TRECVID). Most of the top-ranking solutions can be summarized with a Division-Fusion-Optimization (DFO) framework, in which person and action recognition scores are obtained separately, then fused, and, optionally, further optimized to generate the final ranking. However, TRECVID only evaluates the final ranking results, ignoring the effects of intermediate steps and their implementation methods. We argue that conducting the fine-grained evaluations of intermediate steps of DFO framework will (1) provide a quantitative analysis of the different methods’ performance in intermediate steps; (2) find out better design choices that contribute to improving retrieval performance; and (3) inspire new ideas for future research from the limitation analysis of current techniques. Particularly, we propose an indirect evaluation method motivated by the leave-one-out strategy, which finds an optimal solution surpassing the champion teams in 2020–2021 INS tasks. Moreover, to validate the generalizability and robustness of the proposed solution under various scenarios, we specifically construct a new large-scale P-A INS dataset and conduct comparative experiments with both the leading NIST TRECVID INS solution and the state-of-the-art P-A INS method. Finally, we discuss the limitations of our evaluation work and suggest future research directions.
Yanrui Niu, Chao Liang 0001, Ankang Lu, Baojin Huang, Zhongyuan Wang 0001
ACM Trans. Inf. Syst.4
2024 Joint Distortion Restoration and Quality Feature Learning for No-reference Image Quality Assessment
abstract
No-reference image quality assessment (NR-IQA) methods, inspired by the free energy principle, improve the accuracy of image quality prediction by simulating the human brain’s repair process for distorted images. However, existing methods use separate optimization schemes for distortion restoration and quality prediction, which undermines the accurate mapping of feature representations to quality scores. To address this issue, we propose a joint restoration and quality feature learning NR-IQA (RQFL-IQA) method to jointly tackle distortion image restoration and quality prediction within a unified framework. To accurately establish the quality reconstruction relationship between distorted and restored images, a hybrid loss function based on pixel-wise and structure-wise representations is used to improve the restoration capability of the image restoration network. The proposed RQFL-IQA exploits rich labels, including restored images and quality scores, to enable the model to learn more discriminative features and establish a more accurate mapping from feature representation to quality scores. In addition, to avoid the impact of poor restoration on quality prediction, we propose a module with a cleaning function to reweight the fusion of restored and primitive features to achieve more perceptual consistency in feature fusion. Experimental results on public IQA datasets show that the proposed RQFL-IQA is superior over existing methods.
Jifan Yang, Zhongyuan Wang 0001, Baojin Huang, Jiaxin Ai, Yuhong Yang 0001, Zixiang Xiong
ACM Trans. Multim. Comput. Commun. Appl.3
2024 Auxiliary Information Guided Self-attention for Image Quality Assessment
abstract
Image quality assessment (IQA) is an important problem in computer vision with many applications. We propose a transformer-based multi-task learning framework for the IQA task. Two subtasks: constructing an auxiliary information error map and completing image quality prediction, are jointly optimized using a shared feature extractor. We use visual transformers (ViT) as a feature extractor for feature extraction and guide ViT to focus on image quality-related features by building auxiliary information error map subtask. In particular, we propose a fusion network that includes a channel focus module. Unlike the fusion methods commonly used in previous IQA methods, we use the fusion network, including the channel attention module, to fuse the auxiliary information error map features with the image features, which facilitates the model to mine the image quality features for more accurate image quality assessment. And by jointly optimizing the two subtasks, ViT focuses more on extracting image quality features and building a more precise mapping from feature representation to quality score. With slight adjustments to the model, our approach can be used in both no-reference (NR) and full-reference (FR) IQA environments. We evaluate the proposed method in multiple IQA databases, showing better performance than state-of-the-art FR and NR IQA methods.
Jifan Yang, Zhongyuan Wang 0001, Guangcheng Wang, Baojin Huang, Yuhong Yang 0001, Weiping Tu
ACM Trans. Multim. Comput. Commun. Appl.4
2023 Implicit Identity Driven Deepfake Face Swapping Detection
abstract
In this paper, we consider the face swapping detection from the perspective of face identity. Face swapping aims to replace the target face with the source face and generate the fake face that the human cannot distinguish between real and fake. We argue that the fake face contains the explicit identity and implicit identity, which respectively corresponds to the identity of the source face and target face during face swapping. Note that the explicit identities of faces can be extracted by regular face recognizers. Particularly, the implicit identity of real face is consistent with the its explicit identity. Thus the difference between explicit and implicit identity of face facilitates face swapping detection. Following this idea, we propose a novel implicit identity driven framework for face swapping detection. Specifically, we design an explicit identity contrast (EIC) loss and an implicit identity exploration (IIE) loss, which supervises a CNN backbone to embed face images into the implicit identity space. Under the guidance of EIC, real samples are pulled closer to their explicit identities, while fake samples are pushed away from their explicit identities. More-over, IIE is derived from the margin-based classification loss function, which encourages the fake faces with known target identities to enjoy intra-class compactness and inter-class diversity. Extensive experiments and visualizations on several datasets demonstrate the generalization of our method against the state-of-the-art counterparts.
Baojin Huang, Zhongyuan Wang 0001, Jifan Yang, Jiaxin Ai, Qin Zou 0001, Qian Wang 0002, Dengpan Ye
CVPR1
2023 Continuous Learning for Blind Image Quality Assessment with Contrastive Transformer
abstract
Most existing blind image quality assessment (BIQA) models focus on improving performance on existing datasets and are weak in adapting to unknown distortion or degradation types. In this paper, we propose a Transformer-based BIQA contrastive continual learning approach to improve model transfer performance. The basic idea is that the model continuously learns from the IQA data stream, integrating new knowledge from the current dataset. At the same time, limited access to previous data using a limited memory budget prevents forgetting the knowledge gained from the dataset of old tasks. We design an attentional contrastive learning strategy based on the Transformer architecture with a designed attentional focus contrastive loss to rebalance the contrastive learning between the new and old tasks, which can consolidate the previously learned representations. In addition, we used a structure similar to the cumulative classifier, balancing the learning of the current task quality score with the quality scores of all observed tasks. Extensive experiments demonstrate the feasibility of the proposed continuous learning approach compared to the standard training techniques of BIQA.
Jifan Yang, Zhongyuan Wang 0001, Baojin Huang
ICASSP3
2023 Deepfake Face Provenance for Proactive Forensics
abstract
Malicious deepfake face not only violates the privacy of personal identities, but also confuses the public and causes huge social harm. The current deepfake detection only stays at the level of distinguishing between true and false, but cannot trace the original genuine face corresponding to the fake face, that is, it does not have the ability to trace the source of evidence. The deepfake countermeasure technology for judicial forensics urgently calls for deepfake inversion. This paper pioneers an interesting question about face deepfake, active forensics that "know what it is and how it happened". Given that deepfake faces do not completely discard the features of original faces, especially facial expressions and poses, we argue that original faces can be approximately speculated from their deepfake counterparts. Correspondingly, we design a disentangling reversing network that decouples latent space features of deepfake faces under the supervision of real-fake face pair samples to infer original faces in reverse.
Jiaxin Ai, Zhongyuan Wang 0001, Baojin Huang, Zhen Han 0002, Qin Zou 0001
ICIP3
2023 DeepReversion: Reversely Inferring the Original Face from the DeepFake Face
abstract
Deepfake techniques can generate realistic fake images and videos. Malicious fake facial images quickly spread through the Internet, posing a potential threat to personal privacy and judicial forensics. However, the defense methods against deepfake proposed so far mainly focus on the discrimination of authenticity, but cannot identify the true source of the forged face, i.e., the original genuine face corresponding to the face-swapped fake face. This paper poses an interesting issue for face deepfake, which is the proactive forensics of “knowing what and knowing how”. In view of the fact that the fake face exhibits high similarity with the original face, especially the facial expression and pose, we argue that the original face can be approximately estimated from the deepfake counterpart. Accordingly, we advocate a deep-learning-based face inversion approach, so-called DeepReversion, which learns the inverse mapping from the deepfake face to the original face. Based on UNet, we design a specific end-to-end DeepReversion network, and conduct comprehensive experiments on public deepfake datasets. The experimental results show that the speculated face is highly consistent with the original face in terms of visual effects, PSNR, SSIM and similarity given by face recognizers.
Jiaxin Ai, Zhongyuan Wang 0001, Baojin Huang, Zhen Han 0002
IJCNN3
2023 A Spatio-Temporal Identity Verification Method for Person-Action Instance Search in Movies
Yanrui Niu, Jingyao Yang, Chao Liang 0001, Baojin Huang, Zhongyuan Wang 0001
MMM (1)4
2023 Depth map guided triplet network for deepfake face detection
Buyun Liang 0002, Zhongyuan Wang 0001, Baojin Huang, Qin Zou 0001, Qian Wang 0002
Neural Networks3
2023 Deep Constraints Space of Medium Modality for RGB-Infrared Person Re-identification
Baojin Huang, Wen-cheng Qin
Neural Process. Lett.1
2023 PLFace: Progressive Learning for Face Recognition with Mask Bias
Baojin Huang, Zhongyuan Wang 0001, Guangcheng Wang, Kui Jiang, Zhen Han 0002, Tao Lu 0001, Chao Liang 0001
Pattern Recognit.1
2023 HeadPose-Softmax: Head pose adaptive curriculum learning loss for deep face recognition
Jifan Yang, Zhongyuan Wang 0001, Baojin Huang, Jinsheng Xiao, Chao Liang 0001, Zhen Han 0002, Hua Zou 0002
Pattern Recognit.3
2023 Joint Segmentation and Identification Feature Learning for Occlusion Face Recognition
abstract
The existing occlusion face recognition algorithms almost tend to pay more attention to the visible facial components. However, these models are limited because they heavily rely on existing face segmentation approaches to locate occlusions, which is extremely sensitive to the performance of mask learning. To tackle this issue, we propose a joint segmentation and identification feature learning framework for end-to-end occlusion face recognition. More particularly, unlike employing an external face segmentation model to locate the occlusion, we design an occlusion prediction module supervised by known mask labels to be aware of the mask. It shares underlying convolutional feature maps with the identification network and can be collaboratively optimized with each other. Furthermore, we propose a novel channel refinement network to cast the predicted single-channel occlusion mask into a multi-channel mask matrix with each channel owing a distinct mask map. Occlusion-free feature maps are then generated by projecting multi-channel mask probability maps onto original feature maps. Thus, it can suppress the representation of occlusion elements in both the spatial and channel dimensions under the guidance of the mask matrix. Moreover, in order to avoid misleading aggressively predicted mask maps and meanwhile actively exploit usable occlusion-robust features, we aggregate the original and occlusion-free feature maps to distill the final candidate embeddings by our proposed feature purification module. Lastly, to alleviate the scarcity of real-world occlusion face recognition datasets, we build large-scale synthetic occlusion face datasets, totaling up to 980193 face images of 10574 subjects for the training dataset and 36721 face images of 6817 subjects for the testing dataset, respectively. Extensive experimental results on the synthetic and real-world occlusion face datasets show that our approach significantly outperforms the state-of-the-art in both 1:1 face verification and 1:N face identification.
Baojin Huang, Zhongyuan Wang 0001, Kui Jiang, Qin Zou 0001, Xin Tian 0006, Tao Lu 0001, Zhen Han 0002
IEEE Trans. Neural Networks Learn. Syst.1
2023 Local Eyebrow Feature Attention Network for Masked Face Recognition
abstract
During the COVID-19 coronavirus epidemic, wearing masks has become increasingly popular. Traditional occlusion face recognition algorithms are almost ineffective for such heavy mask occlusion. Therefore, it is urgent to improve the recognition performance of the existing face recognition technology on masked faces. Due to the limited visible feature points of the masked face image relative to the normal face image, we have to exploit the identification potential of eyebrow (referring to eyes and brows) features. This article proposes a local eyebrow feature attention network for masked face recognition, which consists of feature extraction, eyebrow region pooling, and feature fusion. To highlight the eyebrow region, we first use the eyebrow region pooling to separate the local features of eyebrows from the learned overall facial features. We then make full use of the symmetry of left and right eyebrows to emphasize their discriminant ability, due to the inadequate fine information of the low-resolution eyebrows. In particular, in view of the symmetrical similarity between eyebrow pairs and the subordinate relationship between facial components and the whole, we propose a feature fusion model based on graph convolutional network (GCN) to learn the feature association structure of eye features, brow features, and global facial features. We construct the benchmark datasets for masked face recognition to validate our approach, including real-world masked face recognition dataset (RMFRD) and synthetic masked face recognition dataset (SMFRD). Extensive experimental results on both public datasets and our built masked face datasets show that our approach significantly outperforms the state-of-the-arts.
Baojin Huang, Zhongyuan Wang 0001, Guangcheng Wang, Zhen Han 0002, Kui Jiang
ACM Trans. Multim. Comput. Commun. Appl.1
2022 Two-stage unsupervised facial image quality measurement
Guangcheng Wang, Zhongyuan Wang 0001, Baojin Huang, Kui Jiang, Zheng He 0001, Hancheng Zhu, Jinsheng Xiao, Xin Tian 0006
Inf. Sci.3
2022 Learning diverse and deep clues for person reidentification
Wen-cheng Qin, Baojin Huang, Pinzhong Qin, Zhiyong Huang 0004, Daidi Zhong
Image Vis. Comput.2
2022 Realistic frontal face reconstruction using coupled complementarity of far-near-sighted face images
Kangli Zeng, Zhongyuan Wang 0001, Tao Lu 0001, Jianyu Chen 0008, Baojin Huang, Zhen Han 0002, Xin Tian 0006
Pattern Recognit.5
2022 Deep Constraints Space via Channel Alignment for Visible-Infrared Person Re-identification
abstract
Reducing the inter-modality gap has been the core of visible-infrared person re-identification (VI-ReID). Most of the existing methods directly constrain the features to reduce the gap between the different modalities. However, the consistency of channel semantic information across different modalities is ignored. In this letter, channel alignment along with severe semantic constraints is proposed to alleviate the significant distribution differences between modalities. To obtain the optimum channel matching mode, a novel channel alignment mechanism termed Channel Instance Level Alignment (CILA) is proposed at the shallow layer, and matched channel features are constrained by the proposed Channel Instance Alignment (CIA) Loss. In the middle layer, we divide the features into multiple hierarchically-aware channel clusters and align the channel clusters by Channel Cluster Alignment (CCA) Loss. The proposed method is validated on SYSU-MM01 and RegDB, and extensive experiments show that the proposed method achieves competitive performance compared with the state of the arts.
Wen-cheng Qin, Baojin Huang, Zhiyong Huang 0004, Lamia Tahsin, Daming Sun
IEEE Signal Process. Lett.2
2021 When Face Recognition Meets Occlusion: A New Benchmark
abstract
The existing face recognition datasets usually lack occlusion samples, which hinders the development of face recognition. Especially during the COVID-19 coronavirus epidemic, wearing a mask has become an effective means of preventing the virus spread. Traditional CNN-based face recognition models trained on existing datasets are almost ineffective for heavy occlusion. To this end, we pioneer a simulated occlusion face recognition dataset. In particular, we first collect a variety of glasses and masks as occlusion, and randomly combine the occlusion attributes (occlusion objects, textures,and colors) to achieve a large number of more realistic occlusion types. We then cover them in the proper position of the face image with the normal occlusion habit. Furthermore, we reasonably combine original normal face images and occluded face images to form our final dataset, termed as Webface-OCC. It covers 804,704 face images of 10,575 subjects, with diverse occlusion types to ensure its diversity and stability. Extensive experiments on public datasets show that the ArcFace retrained by our dataset significantly outperforms the state-of-the-arts. Webface-OCC is available at https://github.com/Baojin-Huang/Webface-OCC.
Baojin Huang, Zhongyuan Wang 0001, Guangcheng Wang, Kui Jiang, Kangli Zeng, Zhen Han 0002, Xin Tian 0006, Yuhong Yang 0001
ICASSP1
2021 A Tilt-Angle Face Dataset And Its Validation
abstract
Since the surveillance cameras are usually mounted at a high position to overlook targets, tilt-angle faces on overhead view are common in the public video surveillance environment. Face recognition approaches based on deep learning models have achieved excellent performance, but there remains a large gap for the overlooking surveillance scenarios. The results of face recognition depend not only on the structure of the model, but also on the completeness and diversity of the training samples. The existing multi-pose face datasets do not cover complete top-view face samples, and the models trained by them thus cannot provide satisfactory accuracy. To this end, this paper pioneers a multi-view tilt-angle face dataset (TFD), which is collected with an elaborately devised overhead capture equipment. TFD contains 11,124 face images from 927 subjects, covering a variety of tilt angles on the overhead view. To verify the validity of the constructed dataset, we further conduct comprehensive face detection and recognition experiments using the corresponding models trained by WiderFace, Webface and our TFD, respectively. Experimental results show that our TFD substantially promotes the face detection and recognition accuracy under the top-view situation. TFD is available at https://github.com/huang1204510135/D FD.
Nanxi Wang, Zhongyuan Wang 0001, Zheng He 0001, Baojin Huang, Liguo Zhou, Zhen Han 0002
ICIP4
2021 Cross-View Gait Recognition Based on Feature Fusion
abstract
Compared to face recognition, gait recognition is one of the most promising video biometric recognition technologies given that gait images can be readily captured at a distance and gait characteristics are robust to appearance camouflage. A lot of existing gait recognition methods aim at a single scene such as fixed cameras, but the recognition accuracy will decrease sharply if the viewpoints are changed. In this paper, we improve the existing methods and propose a cross-view gait recognition method based on feature fusion. Firstly, a multi-scale feature fusion module is proposed to extract the features of gait sequences with different granularities. Then, a dual-path structure is introduced to learn global appearance features and fine-grained local features, respectively. The features of two paths are gradually merged as the network deepens to obtain the complementary information. In the last feature mapping stage, the Generalized-Mean pooling is used to favour discriminative representation. Extensive experiments on the public dataset CASIA-B show that our method can achieve state-of-the-art recognition performance.
Zhongyuan Wang 0001, Jianyu Chen 0008, Baojin Huang
ICTAI4
2021 Silicone mask face anti-spoofing detection based on visual saliency and facial motion
Guangcheng Wang, Zhongyuan Wang 0001, Kui Jiang, Baojin Huang, Zheng He 0001, Ruimin Hu
Neurocomputing4
2021 Decomposition Makes Better Rain Removal: An Improved Attention-Guided Deraining Network
abstract
Rain streaks in the air show diverse characteristics with different shapes, directions, densities, even the complex overlapped phenomenon, causing great challenges for the deraining task. Recently, deep learning based image deraining methods have been extensively investigated due to their excellent performance. However, most of the existing algorithms still have limitations in removing rain streaks while preserving rich textural details under complicated rain conditions. To this end, we propose to decompose rain streaks into multiple rain layers and individually estimate each of them along the network stages to cope with the increasing abstracts. To better characterize rain layers, an improved non-local block is designed to exploit the self-similarity of rain information by learning the holistic spatial feature correlations while reducing the calculation complexity. Moreover, a mixed attention mechanism is applied to guide the fusion of rain layers by focusing on the local and global overlaps among these rain layers. Extensive experiments on both synthetic rainy/rain-haze/raindrop datasets, real-world samples, the haze, and low-light scenarios show substantial improvements both on quantitative indicators and visual effects over the current state-of-the-art technologies. The source code is available athttps://github.com/kuihua/IADN.
Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Chen Chen 0001, Zhen Han 0002, Tao Lu 0001, Baojin Huang, Junjun Jiang
IEEE Trans. Circuits Syst. Video Technol.7
2020 Multi-Scale Progressive Fusion Network for Single Image Deraining
abstract
Rain streaks in the air appear in various blurring degrees and resolutions due to different distances from their positions to the camera. Similar rain patterns are visible in a rain image as well as its multi-scale (or multi-resolution) versions, which makes it possible to exploit such complementary information for rain streak representation. In this work, we explore the multi-scale collaborative representation for rain streaks from the perspective of input image scales and hierarchical deep features in a unified framework, termed multi-scale progressive fusion network (MSPFN) for single image rain streak removal. For the similar rain streaks at different positions, we employ recurrent calculation to capture the global texture, thus allowing to explore the complementary and redundant information at the spatial dimension to characterize target rain streaks. Besides, we construct multi-scale pyramid structure, and further introduce the attention mechanism to guide the fine fusion of these correlated information from different scales. This multi-scale progressive fusion strategy not only promotes the cooperative representation, but also boosts the end-to-end training. Our proposed method is extensively evaluated on several benchmark datasets and achieves the state-of-the-art results. Moreover, we conduct experiments on joint deraining, detection, and segmentation tasks, and inspire a new research direction of vision task driven image deraining. The source code is available at https://github.com/kuihua/MSPFN.
Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Chen Chen 0001, Baojin Huang, Yimin Luo, Jiayi Ma 0001, Junjun Jiang
CVPR5
2020 Lightweight Progressive Residual Clique Network for Image Super-Resolution
abstract
Deeper and wider convolutional neural networks (CNN) hava been widely applied to the single image super-resolution (SR) task for its appealing performance. However, enormous parametric memory footprint hinders its real-time application on mobile devices, especially in the energy-sensitive environment. In this work, we take both the reconstruction performance and efficiency into consideration and propose a lightweight progressive residual clique network (PRCN) for image SR. PRCN is built on the two-stage residual channel separation block (RCSB) and long-skip connections. First, we divide the input into four channel groups to differently learn texture details, immediately followed by a primary fusion to establish cross-channel correspondence in the first stage. Then we perform a further fusion on the outputs of the first stage to constitute a clique for the refinement in the second stage. Meanwhile, we employ SENet to improve the outputs of the second stage with the separate features of the first stage. This design not only enforces the correlation across channels, but also allows fewer densely connected blocks. Experimental results on public datasets show that PRCN outperforms state-of-the-art methods in terms of performance and complexity.
Baojin Huang, Zheng He 0001, Zhongyuan Wang 0001, Kui Jiang, Guangcheng Wang
ICTAI1
2020 Story segmentation for news broadcast based on primary caption
abstract
In the information explosion era, people only want to access the news information that they are interested in. News broadcast story segmentation is strongly needed, which is an essential basis for personalized delivery and short video. The existing advanced story boundary segmentation methods utilize semantic similarity of subtitles, thus entailing complex semantic computation. The title texts of news broadcast programs include headline (or primary) captions, dialogue captions and the channel logo, while the same story clips only render one primary caption in most news broadcast. Inspired by this fact, we propose a simple method for story segmentation based on the primary caption, which combines YOLOv3 based primary caption extraction and preliminary location of boundaries. In particular, we introduce mean hash to achieve the fast and reliable comparison for detected small-size primary caption blocks. We further incorporate scene recognition to exact the preliminary boundaries, because the primary captions always appear later than the story boundary. Experimental results on two Chinese news broadcast datasets show that our method enjoys high accuracy in terms of R, P and F1-measures.
Heling Chen, Zhongyuan Wang 0001, Yingjiao Pei, Baojin Huang, Weiping Tu
MMAsia4
2020 Video scene detection based on link prediction using graph convolution network
abstract
With the development of the Internet, multimedia data grows by an exponential level. The demand for video organization, summarization and retrieval has been increasing where scene detection plays an essential role. Existing shot clustering algorithms for scene detection usually treat temporal shot sequence as unconstrained data. The graph based scene detection methods can locate the scene boundaries by taking the temporal relation among shots into account, while most of them only rely on low-level features to determine whether the connected shot pairs are similar or not. The optimized algorithms considering temporal sequence of shots or combining multi-modal features will bring parameter trouble and computational burden. In this paper, we propose a novel temporal clustering method based on graph convolution network and the link transitivity of shot nodes, without involving complicated steps and prior parameter setting such as the number of clusters. In particular, the graph convolution network is used to predict the link possibility of node pairs that are close in temporal sequence. The shots are then clustered into scene segments by merging all possible links. Experimental results on BBC and OVSD datasets show that our approach is more robust and effective than the comparison methods in terms of F1-score.
Yingjiao Pei, Zhongyuan Wang 0001, Heling Chen, Baojin Huang, Weiping Tu
MMAsia4