Bolun Zheng

dblp:225/9602 · DBLP profile ↗
← Back
58ranked-venue papers
7as first author
56since 2021 · last 2026
0000-0001-8788-1725ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 34 · 5 first-author · 32 since 2021Artificial intelligence and machine learning · 23 · 4 first-author · 22 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multi-scale sampling and feature fusion for dynamic human rendering
Kainan Yu, Bolun Zheng, Qianyu Zhang 0002, Fangni Chen, Jiyong Zhang 0001, Canjin Wang
J. Vis. Commun. Image Represent.2
2026 Reference-aware image harmonization
Hongling Gu, Bolun Zheng, Qianyu Zhang 0002, Canjin Wang, Yayun Wang, Zongpeng Li
Neural Networks3
2026 Position-Sensitive painterly image harmonization
Bolun Zheng, Qianyu Zhang 0002, Canjin Wang, Yayun Wang, Heng Jin, Qiankun Li 0005, Guodao Zhang, Zongpeng Li
Neural Networks2
2026 Enhancing out-of-distribution detection with bilateral distribution score
Bolun Zheng, Yao Zhu 0003, Canjin Wang
Neural Networks1
2026 A benchmark for robust salient object detection in adverse weather conditions
Bolun Zheng, Rongfeng Lu, Xiaokai Yang, Qianyu Zhang 0002, Yu Liu 0005, Xiaofei Zhou 0003
Pattern Recognit.3
2026 Video Demoiréing With Spatial-Temporal Filtering in Frequency Domain
abstract
When acquiring images or videos of electronic displays, moiré patterns often arise due to aliasing between overlapping pixel grids, substantially compromising the perceptual quality of the captured content. Although frequency domain techniques have demonstrated high efficacy in image demoiréing, existing video approaches often overlook inter-frame frequency contextual relationships. This limitation restricts their capacity to achieve consistent temporal coherence and reconstruction fidelity. To overcome these challenges, we introduce a novel network (STFNet) with spatial-temporal filtering in frequency domain for video demoiréing. The proposed architecture comprises two dedicated stages: (1) Temporal-Guided Filtering (TGF), aims to adaptively incorporate temporal cues into learnable bandpass filters; and (2) Joint Filtering with Partially Shared Passbands (JFPS), which enhances representation learning of low-frequency moiré textures through strategic parameter sharing. Comprehensive evaluations on multiple public benchmarks confirm the superiority of our method. STFNet consistently outperforms state-of-the-art alternatives across both quantitative metrics and perceptual quality assessments, demonstrating robust performance in dynamic moiré suppression and detail preservation.
Zhongqi Liu, Bolun Zheng, Qianyu Zhang 0002, Heng Jin, Qiankun Li 0005, Xu Jia 0012, Jiyong Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2026 MEF-GD: Multimodal Enhancement and Fusion Network for Garment Designer
abstract
In recent years, with advancements in generative models, an increasing number of garment design methods have been proposed. A generative model capable of generating garment images from text and sketches can provide designers with valuable visual references and creative inspiration to aid in the design process. Existing multimodal garment design methods face the challenge of lacking precise control over the generated results in relation to both sketches and text. In this paper, we propose Multimodal Enhancement and Fusion Network for Garment Design (MEF-GD). Our model inputs image conditions into Stable Diffusion based on ControlNet. On one hand, directly inputting image conditions can lead to feature forgetting, defined as the phenomenon in deep neural networks where previously learned feature representations are lost. To address this issue, we propose a multiple feature injection module to more effectively enhance image condition features. On the other hand, ControlNet fuses control features into Stable Diffusion through pointwise addition, which ignores the interaction between multimodal features and results in the fused features being biased towards the control features, overlooking Stable Diffusion features. To address this limitation, we introduce content-guided attention for more effective feature fusion and improve the expression of text features. Additionally, existing datasets often contain vague textual descriptions of garments. It is difficult to train the model on such a dataset to learn accurate alignment between generated image and the textual descriptions. To address this issue, we have designed a multimodal large model text optimization module to improve the quality and clarity of text generation. Compared to existing multimodal garment design methods, MEF-GD achieves more effective alignment with both textual and sketch-based inputs in generating garment images. Compared to MGD, MEF-GD achieves a decrease of 2.44 in FID and an increase of 0.83 in CLIP Score on Multi-VITON-HD dataset. The code will be available at https://github.com/fengyun691340/MEF-GD.
Dan Song 0006, Jianhao Zeng, Hongshuo Tian, Bolun Zheng, Rongbao Kang, Anan Liu
IEEE Trans. Circuits Syst. Video Technol.5
2026 LLFeat: Noise-Aware Feature Matching Under Various Low-Light Conditions
Longjian Zeng, Zunjie Zhu, Ming Lu 0002, Bolun Zheng, Rongfeng Lu, Tingyu Wang 0002, Zhongtian Zheng, Yaoqi Sun, Chenggang Yan 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 Constituency-Tree-Induced Vision-Language Alignment for Multimodal Large Language Models
abstract
Multimodal large language models (MLLMs) integrate sophisticated large vision models (LVMs) to empower large language models (LLMs) with vision ability to perceive, reason, and interact in vision-language (V-L) tasks, while the modality bridge between two specialists becomes the bottleneck that translates visual signals into linguistic representations. However, most of the existing methods train the modality bridge with coarse-grained image-text pairs, neglecting the structural mapping between V-L semantics that facilitates modality translation from LVMs to LLMs. To mitigate this, we propose a Constituency-Tree-Induced Multimodal Bridging mechanism (CTIMB) that learns the fine-grained connection from LVMs to LLMs by the structural guidance from multi-modal constituency tree. Our approach consists of: 1) the multi-modal constituency-tree parser that jointly exploits the semantic structure of vision and language; 2) the lightweight connector that translates visual signals into linguistic representation and re-arranges them according to the constituency-tree structure; 3) the dynamic construction loss that aids in aligning the semantic structures derived from the tree parser and the connector. The CTIMB can learn the fine-grained mapping between visual and linguistic semantics, seamlessly bridge the LVMs and LLMs to enhance V-L tasks, and is more cost-efficient compared with current methods. Extensive experiments have demonstrated that our method more accurately interprets the visual features, enabling LLMs to conduct downstream tasks more effectively, and achieve superior performance with less training cost.
Yingchen Zhai, Ning Xu 0003, Hongshuo Tian, Bolun Zheng, Chenggang Yan 0001, Jinbo Cao, Rongbao Kang, Anan Liu
IEEE Trans. Circuits Syst. Video Technol.4
2026 MMToT: Multi-Modal Token-of-Thought Reasoning for Large Models
abstract
With the development of Large models (LMs), recent methods tend to leverage them to complete various downstream tasks, like VQA and image caption. A typical method is Chain-of-Thought (CoT) prompting, which improves the reasoning abilities of LMs by providing intermediate steps. However, existing CoT prompts have two main drawbacks: 1) They typically represent CoT through multiple sentences, which often introduce irrelevant textual or visual context that may confuse LMs. 2) The current CoT methods fail to consider how the contributions of different tokens vary for answer inference. For this, we propose the Multi-Modal Token-of-Thought (MMToT), a novel token-level prompt method to improve LMs' multi-modal reasoning capabilities. Furthermore, MMToT stress on two strengths against CoT: 1) To prevent LMs from being affected by irrelevant contexts, we propose to extract explicit multi-modal tokens rather than sentences to construct MMToT, enhancing the reliability of generated answers. 2) To ensure that LMs prioritize tokens with high contribution scores during answer generation, we propose a confident decision-making module to evaluate and integrate each token's contribution in MMToT. Compared to existing methods, the proposed MMToT demonstrates superior performance on Science-QA, MATH, OKVQA, and VQA-introspect datasets. Furthermore, ablation studies and visualization results validate the effectiveness and interpretability of MMToT.
Ning Xu 0003, Zimu Lu, Hongshuo Tian, Bolun Zheng, Jinbo Cao, Anan Liu
IEEE Trans. Multim.4
2026 Knowledge and multi-detail enhanced GAN for human-driven text-to-image synthesis
abstract
Human-driven text-to-image synthesis aims to create controllable images, which not only adhere to the semantic of given text but also incorporate the visual characteristics of given human. For example, given “a man on the beach” (text) along with a photo of human, the model aims to generate an image depicting the human on the beach. Although current diffusion-based methods have shown promise in this task, they face two major limitations: (1) The generated images appear to be a bit stiff and unnatural, almost like collages of human and backgrounds; (2) The details of human in the generated image are inconsistent with those in the input, losing the original identity. To address these issues, we present the Knowledge and Multi-Detail Enhanced GAN for the task of human-driven text-to-image synthesis. It employs external knowledge as references to improve the harmony between human and backgrounds, and uses CLIP’s multi-layer features to intensify human details. First, we search the database to retrieve external images that are similar to the given text, serving as our knowledge. Second, to preserve the human details, we present the Multi-Detail Enhancer, which uses the image encoder of CLIP to extract human representation at multiple levels. Third, to enhance the human-background naturalness, we present the Knowledge Attention Enhancer, which can seamlessly blend human, text, and knowledge by attentively retain useful information and filter out noise from knowledge. Finally, we introduce the dual discriminators to guide the entire network, which can facilitate the accurate capture of human details and generation of images. Extensive experiments demonstrate the superiority of our method with its efficiency and lower computational demands. It is about 300 times faster than diffusion-based models, uses only 5% of the parameters, and completes training in just two days on three V100 GPUs.
Ning Xu 0003, Zhewen Shen, Hongshuo Tian, Bolun Zheng, Chenggang Yan 0001, Jinbo Cao, Rongbao Kang, Anan Liu
Vis. Informatics4
2025 SMTPD: A New Benchmark for Temporal Prediction of Social Media Popularity
abstract
Social media popularity prediction task aims to predict the popularity of posts on social media platforms, which has a positive driving effect on application scenarios such as content optimization, digital marketing and online advertising. Though many studies have made significant progress, few of them pay much attention to the integration between popularity prediction with temporal alignment. In this paper, with exploring YouTube’s multilingual and multi-modal content, we construct a new social media temporal popularity prediction benchmark, namely SMTPD, and suggest a baseline framework for temporal popularity prediction. Through data analysis and experiments, we verify that temporal alignment and early popularity play crucial roles in social media popularity prediction for not only deepening the understanding of temporal dynamics of popularity in social media but also offering a suggestion about developing more effective prediction models in this field. Code is available at https://github.com/zhuwei321/SMTPD
Yijie Xu, Bolun Zheng, Hangjia Pan, Yuchen Yao, Ning Xu 0003, Anan Liu, Chenggang Yan 0001
CVPR2
2025 DepthDark: Robust Monocular Depth Estimation for Low-Light Environments
abstract
In recent years, foundation models for monocular depth estimation have received increasing attention. Current methods mainly address typical daylight conditions, but their effectiveness notably decreases in low-light environments. There is a lack of robust foundational models for monocular depth estimation specifically designed for low-light scenarios. This largely stems from the absence of large-scale, high-quality paired depth datasets for low-light conditions and the effective parameter-efficient fine-tuning (PEFT) strategy. To address these challenges, we propose DepthDark, a robust foundation model for low-light monocular depth estimation. We first introduce a flare-simulation module and a noise-simulation module to accurately simulate the imaging process under nighttime conditions, producing high-quality paired depth datasets for low-light conditions. Additionally, we present an effective low-light PEFT strategy that utilizes illumination guidance and multiscale feature fusion to enhance the model's capability in low-light environments. Our method achieves state-of-the-art depth estimation performance on the challenging nuScenes-Night and RobotCar-Night datasets, validating its effectiveness using limited training data and computing resources.
Longjian Zeng, Zunjie Zhu, Rongfeng Lu, Ming Lu 0002, Bolun Zheng, Chenggang Yan 0001, Anke Xue
ACM Multimedia5
2025 EHPE: A Segmented Architecture for Enhanced Hand Pose Estimation
abstract
3D hand pose estimation has garnered great attention in recent years due to its critical applications in human-computer interaction, virtual reality, and related fields. Accurate estimation of hand joints is essential for high-quality hand pose estimation. However, existing methods neglect the importance of Distal Phalanx Tip (TIP) and Wrist in predicting hand joints overall and often fail to account for the phenomenon of error accumulation for distal joints in gesture estimation, which can cause certain joints to incur larger errors, resulting in misalignments and artifacts in pose estimation and degrading the overall reconstruction quality. To address this challenge, we propose a novel segmented architecture for enhanced hand pose estimation (EHPE). We perform a local extraction of the TIP and wrist, thus alleviating the effect of error accumulation on the prediction of the TIP and further reduce the predictive errors for all joints on this basis. EHPE consists of two key stages: In the TIP and Wrist Joints Extraction stage (TW-stage), the positions of the TIP and wrist joints are estimated to provide an initial accurate joint configuration; In the Prior Guided Joints Estimation stage (PG-stage), a dual-branch interaction network is employed to refine the positions of the remaining joints. Extensive experiments on two widely used benchmarks demonstrate that EHPE achieves state-of-the-art performance.
Bolun Zheng, Xinjie Liu, Qianyu Zhang 0002, Canjin Wang, Fangni Chen, Mingen Xu
ACM Multimedia1
2025 Mixture of causal experts: A causal perspective to build dual-level mixture-of-experts models
Ning Xu 0003, Hongshuo Tian, Bolun Zheng, Jinbo Cao, Anan Liu
Expert Syst. Appl.4
2025 Counterfactual GAN for debiased text-to-image synthesis
Xianghua Kong, Ning Xu 0003, Zefang Sun, Zhewen Shen, Bolun Zheng, Chenggang Yan 0001, Jinbo Cao, Rongbao Kang, Anan Liu
Multim. Syst.5
2025 Pyramid Learnable Bandpass Filters for Ultra-High-Definition Image Demoiréing
abstract
Moiré patterns usually depend on the style of display grids and the position of shooting camera, appearing in the form of stripes, meshes or ripples, with various and irregular colors. Compared with low-resolution moiré images, high-definition (HD) and ultra-high-definition (UHD) moiré images exhibit more complex moiré patterns, e.g., wider distribution of moiré frequencies and higher coupling degree of moirés of different scales, which poses a greater challenge to the modeling capabilities of the model. To address these challenges, we propose a novel Pyramid Learnable Bandpass Filtering Network (PBNet) for demoiréing UHD images. Specifically, we propose a pyramid learnable bandpass filter (P-LBF) to perform multi-scale filtering in the same semantic context to obtain richer frequency domain information. The P-LBF contains three stages: aligning, filtering and fusing. First, we introduce a pyramid alignment (DA) to align neighbor pixels for eliminating the deviations raised by different styles of display grids and relative position of the shooting camera. Then, a pyramid filtering (PF) is conducted to model the complex and variable moiré patterns with aligned neighbor pixels. Finally, the frequency domain responses of these different scales are fused with a multi-dimensional feature fusion (MFF). The PBNet is constructed based on the P-LBF, incorporating a cross-layer feature fusion (CLF) module to facilitate more effective information interaction between features at different depths. Extensive experiments on four public datasets show that our model achieves state-of-the-art performance for both high- and low-resolution moiré images. The code is publicly available at:https://github.com/liuzhongqi1/PBNet.
Zhongqi Liu, Bolun Zheng, Qianyu Zhang 0002, Xu Jia 0012, Jiyong Zhang 0001, Chenggang Yan 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 Hierarchical Frequency-Based Upsampling and Refining for HEVC Compressed Video Enhancement
abstract
Video compression artifacts arise from quantization applied in the frequency domain. Video quality enhancement aims to reduce such compression artifacts and reconstruct a visually pleasant result. While existing methods effectively reduce artifacts in the spatial domain, they often overlook the rich frequency domain information, especially in addressing multi-scale compression artifacts. This work introduces a frequency-domain upsampling strategy within a multi-scale framework, specifically designed to focus on high-frequency details rather than simply blending neighboring pixels during the upsampling process. Our proposed hierarchical frequency-based upsampling and refinement neural network (HFUR) consists of two modules: implicit frequency upsampling (ImpFreqUp) and hierarchical and iterative refinement (HIR). ImpFreqUp exploits the DCT-domain prior derived through an implicit DCT transform, and accurately reconstructs the DCT-domain signal via a coarse-to-fine transfer. Additionally, HIR is introduced to facilitate cross-collaboration and information compensation between the scales, further refining the feature maps and promoting the visual quality of the final output. We demonstrate the effectiveness of the proposed modules via ablation experiments and visualized results. Experimental results demonstrate that HFUR outperforms the state-of-the-art methods up to 0.13dB/0.17dB on both constant bit rate and constant QP modes. The code is available athttps://github.com/zqqqyu/HFUR.
Qianyu Zhang 0002, Bolun Zheng, Xingying Chen, Zunjie Zhu, Canjin Wang, Zongpeng Li, Xu Jia 0012, Chengang Yan
IEEE Trans. Circuits Syst. Video Technol.2
2025 Target Detection Based on Regional Feature Difference in Synthetic Aperture Interferometric Radiometer
Bo Fang 0008, Fei Hu 0002, Yanyu Xu 0002, Yakai Hao, Jingyu Tao, Jiale Min, Bolun Zheng
IEEE Trans. Geosci. Remote. Sens.9
2025 GCBF: Grouped Cross-Band Fusion Network for Multispectral Scene Classification
abstract
Remote sensing scene classification is a crucial task for remote sensing image interpretation. Existing multispectral scene classification methods have overlooked the interrelationships between different spectral bands, which limits the mining of complementary information within the images. Addressing this issue, we propose a grouped cross-band fusion (GCBF) network for remote sensing multispectral scene classification to take full advantage of complementary information between various spectral bands. Firstly, we separate the various bands of the given multispectral image into different groups to better capture the characteristics of each spectral band. Then, we use the existing UniFormer as a feature extractor to learn the representations of red, green, and blue (RGB) bands. For the spectral bands other than RGB, we propose a new network called multi-stage grouped spectral feature extraction (MGSFE) network to learn discriminative representations. We also draw inspiration from the band combination in the field of remote sensing and introduce a cross-band attention fusion (CBAF) module designed to adaptively merge features from both the RGB bands and other spectral bands. Extensive experiments on three widely used remote sensing multispectral scene classification datasets of BigEarthNet, SEN12MS, and EuroSAT demonstrate the superiority of our proposed method compared with several state-of-the-art (SOTA) methods.
Jin Li 0069, Yu Liu 0005, Wenda Zhao 0003, Zhizhuo Jiang, Xueqian Wang 0002, Bolun Zheng
IEEE Trans. Geosci. Remote. Sens.6
2025 IFENet: Interaction, Fusion, and Enhancement Network for V-D-T Salient Object Detection
abstract
Visible-depth-thermal (VDT) salient object detection (SOD) aims to highlight the most visually attractive object by utilizing the triple-modal cues. However, existing models don't give sufficient exploration of the multi-modal correlations and differentiation, which leads to unsatisfactory detection performance. In this paper, we propose an interaction, fusion, and enhancement network (IFENet) to conduct the VDT SOD task, which contains three key steps including the multi-modal interaction, the multi-modal fusion, and the spatial enhancement. Specifically, embarking on the Transformer backbone, our IFENet can acquire multi-scale multi-modal features. Firstly, the inter-modal and intra-modal graph-based interaction (IIGI) module is deployed to explore inter-modal channel correlation and intra-modal long-term spatial dependency. Secondly, the gated attention-based fusion (GAF) module is employed to purify and aggregate the triple-modal features, where multi-modal features are filtered along spatial, channel, and modality dimensions, respectively. Lastly, the frequency split-based enhancement (FSE) module separates the fused feature into high-frequency and low-frequency components to enhance spatial information (i.e., boundary details and object location) of the salient object. Extensive experiments are performed on VDT-2048 dataset, and the results show that our saliency model consistently outperforms 13 state-of-the-art models. Our code and results are available at https://github.com/Lx-Bao/IFENet.
Liuxin Bao, Xiaofei Zhou 0003, Bolun Zheng, Runmin Cong, Haibing Yin, Jiyong Zhang 0001, Chenggang Yan 0001
IEEE Trans. Image Process.3
2024 Quad Bayer Joint Demosaicing and Denoising Based on Dual Encoder Network with Joint Residual Learning
abstract
The recent imaging technology Quad Bayer CFA brings better imaging PSNR and higher visual quality compared to traditional Bayer CFA, but also serious challenges for demosaicing and denoising during the ISP pipeline. In this paper, we propose a novel dual encoder network, namely DRNet, to achieve joint demosaicing and denoising for Quad Bayer CFA. The dual encoders are carefully designed in that one is mainly constructed by a joint residual block to jointly estimate the residuals for demosaicing and denoising separately. In contrast, the other one is started with a pixel modulation block which is specially designed to match the characteristics of Quad Bayer pattern for better feature extraction. We demonstrate the effectiveness of each proposed component through detailed ablation investigations. The comparison results on public benchmarks illustrate that our DRNet achieves an apparent performance gain~(0.38dB to the 2nd best) from the state-of-the-art method and balances performance and efficiency well. The experiments on real-world images show that the proposed method could enhance the reconstruction quality from the native ISP algorithm.
Bolun Zheng, Haoran Li 0025, Tingyu Wang 0002, Xiaofei Zhou 0003, Chenggang Yan 0001
AAAI1
2024 Infrared Small Target Detection with Scale and Location Sensitivity
abstract
Recently, infrared small target detection (IRSTD) has been dominated by deep-learning-based methods. However, these methods mainly focus on the design of complex model structures to extract discriminative features, leaving the loss functions for IRSTD under-explored. For ex-ample, the widely used Intersection over Union (IoU) and Dice losses lack sensitivity to the scales and locations of targets, limiting the detection performance of detectors. In this paper, we focus on boosting detection performance with a more effective loss but a simpler model structure. Specifically, we first propose a novel Scale and Location Sensitive (SLS) loss to handle the limitations of existing losses: 1) for scale sensitivity, we compute a weight for the IoU loss based on target scales to help the detector distinguish targets with different scales: 2) for location sensitivity, we introduce a penalty term based on the center points of targets to help the detector localize targets more precisely. Then, we design a simple Multi-Scale Head to the plain U-Net (MSHNet). By applying SLS loss to each scale of the predictions, our MSHNet outperforms existing state-of-the-art methods by a large margin. In addition, the detection performance of existing detectors can be further improved when trained with our SLS loss, demonstrating the effectiveness and generalization of our SLS loss. The code is available at https://github.com/ying-fu/MSHNet.
Qiankun Liu 0001, Rui Liu 0039, Bolun Zheng, Hongkui Wang, Ying Fu 0001
CVPR3
2024 Aerial-view geo-localization based on multi-layer local pattern cross-attention network
Haoran Li 0025, Tingyu Wang 0002, Qiang Zhao 0005, Shaowei Jiang, Chenggang Yan 0001, Bolun Zheng
Appl. Intell.7
2024 Enhanced local distribution learning for real image super-resolution
Yaoqi Sun, Aiai Huang, Chenggang Yan 0001, Bolun Zheng
Comput. Vis. Image Underst.6
2024 Neural image re-exposure
Xinyu Zhang 0017, Hefei Huang, Xu Jia 0012, Dong Wang 0004, Lihe Zhang, Bolun Zheng, Wei Zhou 0021, Huchuan Lu
Comput. Vis. Image Underst.6
2024 Rethinking Out-of-Distribution Detection From a Human-Centric Perspective
Yao Zhu 0003, Yuefeng Chen, Rong Zhang 0006, Hui Xue 0001, Xiang Tian 0002, Rongxin Jiang 0001, Bolun Zheng, Yaowu Chen
Int. J. Comput. Vis.8
2024 GoLDFormer: A global-local deformable window transformer for efficient image restoration
Bolun Zheng, Chenggang Yan 0001, Zunjie Zhu, Tingyu Wang 0002, Gregory Slabaugh, Shanxin Yuan
J. Vis. Commun. Image Represent.2
2024 Feature rectification and enhancement for no-reference image quality assessment
Daoquan Huang, Zhuonan Shen, Chenggang Yan 0001, Bolun Zheng
J. Vis. Commun. Image Represent.7
2024 Learning degradation priors for reliable no-reference image quality assessment
Zhuonan Shen, Bolun Zheng, Dingguo Yu, Chenggang Yan 0001
J. Vis. Commun. Image Represent.3
2024 Body Joint Boundary Prototype Match for Few-Shot Remote Sensing Semantic Segmentation
abstract
Deep networks require a large number of samples for optimization, so few-shot segmentation in remote sensing scenes is still an open problem. However, this challenge is exacerbated by the feature blurring and aliasing of bodies (low frequency) and boundaries (high frequency). The existing methods usually only focus on the body part of the class, that is, the low-frequency part, and ignore the critical role of boundary information, that is, high-frequency details, on feature representation. In this letter, we propose a novel body joint boundary prototype match (B2PM) approach that aims to enable prior learning of low- and high-frequency information by explicitly modeling the body and boundary features of objects. First, body-aware prototype learning (BodyPL) realizes the adaptive modeling of the body part of the object through a precise farthest point sampling (FPS) initialization algorithm and an adaptive part shift (APS) strategy, which alleviates the feature ambiguity of the body. Second, boundary-aware prototype learning (BoundPL) explicitly models boundary prototypes by building a patch division and assignment strategy to alleviate feature aliasing at boundaries. Finally, prototype match performs prior knowledge aggregation by computing the affinity between query features and support prototypes. Extensive experiments on commonly used benchmarks (iSAID and PASCAL VOC) demonstrate that B2PM improves the state of the art by significant margins.
Yongqiang Mao, Zhizhuo Jiang, Yu Liu 0005, Yaowen Li, Chenggang Yan 0001, Bolun Zheng
IEEE Geosci. Remote. Sens. Lett.7
2024 Non-local degradation modeling for spatially adaptive single image super-resolution
Qianyu Zhang 0002, Bolun Zheng, Zongpeng Li, Yu Liu 0005, Zunjie Zhu, Gregory Slabaugh, Shanxin Yuan
Neural Networks2
2024 A Unified Asymmetric Knowledge Distillation Framework for Image Classification
abstract
Abstract Knowledge distillation is a model compression technique that transfers knowledge learned by teacher networks to student networks. Existing knowledge distillation methods greatly expand the forms of knowledge, but also make the distillation models complex and symmetric. However, few studies have explored the commonalities among these methods. In this study, we propose a concise distillation framework to unify these methods and a method to construct asymmetric knowledge distillation under the framework. Asymmetric distillation aims to enable differentiated knowledge transfers for different distillation objects. We designed a multi-stage shallow-wide branch bifurcation method to distill different knowledge representations and a grouping ensemble strategy to supervise the network to teach and learn selectively. Consequently, we conducted experiments using image classification benchmarks to verify the proposed method. Experimental results show that our implementation can achieve considerable improvements over existing methods, demonstrating the effectiveness of the method and the potential of the framework.
Xin Ye 0010, Xiang Tian 0002, Bolun Zheng, Fan Zhou 0007, Yaowu Chen
Neural Process. Lett.3
2024 SDPL: Shifting-Dense Partition Learning for UAV-View Geo-Localization
abstract
Cross-view geo-localization aims to match images of the same target from different platforms, e.g., drone and satellite. It is a challenging task due to the changing appearance of targets and environmental content from different views. Most methods focus on obtaining more comprehensive information through feature map segmentation, while inevitably destroying the image structure, and are sensitive to the shifting and scale of the target in the query. To address the above issues, we introduce simple yet effective part-based representation learning, shifting-dense partition learning (SDPL). We propose a dense partition strategy (DPS), dividing the image into multiple parts to explore contextual information while explicitly maintaining the global structure. To handle scenarios with non-centered targets, we further propose the shifting-fusion strategy, which generates multiple sets of parts in parallel based on various segmentation centers, and then adaptively fuses all features to integrate their anti-offset ability. Extensive experiments show that SDPL is robust to position shifting, and performs competitively on two prevailing benchmarks, University-1652 and SUES-200. In addition, SDPL shows satisfactory compatibility with a variety of backbone networks (e.g., ResNet and Swin).https://github.com/C-water/SDPL_release.
Tingyu Wang 0002, Haoran Li 0025, Rongfeng Lu, Yaoqi Sun, Bolun Zheng, Chenggang Yan 0001
IEEE Trans. Circuits Syst. Video Technol.7
2024 Multi-stage reasoning on introspecting and revising bias for visual question answering
abstract
Visual Question Answering (VQA) is a task that involves predicting an answer to a question depending on the content of an image. However, recent VQA methods have relied more on language priors between the question and answer rather than the image content. To address this issue, many debiasing methods have been proposed to reduce language bias in model reasoning. However, the bias can be divided into two categories: good bias and bad bias. Good bias can benefit to the answer prediction, while the bad bias may associate the models with the unrelated information. Therefore, instead of excluding good and bad bias indiscriminately in existing debiasing methods, we proposed a bias discrimination module to distinguish them. Additionally, bad bias may reduce the model’s reliance on image content during answer reasoning and thus attend little on image features updating. To tackle this, we leverage Markov theory to construct a Markov field with image regions and question words as nodes. This helps with feature updating for both image regions and question words, thereby facilitating more accurate and comprehensive reasoning about both the image content and question. To verify the effectiveness of our network, we evaluate our network on VQA v2 and VQA cp v2 datasets and conduct extensive quantity and quality studies to verify the effectiveness of our proposed network. Experimental resu- lts show that our network achieves significant performance against the previous state-of-the-art methods.
Anan Liu, Zimu Lu, Ning Xu 0003, Min Liu 0008, Chenggang Yan 0001, Bolun Zheng, Yulong Duan, Xuanya Li
ACM Trans. Web6
2023 Improving Dynamic HDR Imaging with Fusion Transformer
abstract
Reconstructing a High Dynamic Range (HDR) image from several Low Dynamic Range (LDR) images with different exposures is a challenging task, especially in the presence of camera and object motion. Though existing models using convolutional neural networks (CNNs) have made great progress, challenges still exist, e.g., ghosting artifacts. Transformers, originating from the field of natural language processing, have shown success in computer vision tasks, due to their ability to address a large receptive field even within a single layer. In this paper, we propose a transformer model for HDR imaging. Our pipeline includes three steps: alignment, fusion, and reconstruction. The key component is the HDR transformer module. Through experiments and ablation studies, we demonstrate that our model outperforms the state-of-the-art by large margins on several popular public datasets.
Rufeng Chen, Bolun Zheng, Chenggang Yan 0001, Gregory Slabaugh, Shanxin Yuan
AAAI2
2023 Towards Confidence-Aware Commonsense Knowledge Integration for Scene Graph Generation
abstract
Commonsense knowledge has been widely explored to improve Scene Graph Generation (SGG). Existing methods simply incorporate the described relations of knowledge bases into each part of the scene for a concrete understanding. However, they ignore the discussion about whether a visual scene needs to associate commonsense knowledge for making inferences. Specifically, the difficulty of relation recognition varies from its type. Some frequent spatial relations (e.g. on) usually produce less perception error even without any prior information, while others involved many rules and patterns (e.g. throwing) possess few samples and require to combine with some commonsense knowledge as supplementary. In this paper, we propose a novel confidence-aware commonsense knowledge integration for SGG. Firstly, we depend on mutual information maximization to design a hybrid-attention module, which decreases the uncertainty in representation learning given external knowledge. Second, we introduce an extra branch for SGG network to perform confidence estimation independent of any ground truth labels, in which the output scalar explicitly reflects the difficulty of visual recognition. This value is equipped with the ability to balance the demand for commonsense knowledge in a given scene. Experiments are conducted with the backbone of MOTIFS on Visual Genome (VG) and our method effectively promotes the metric of mRecall with little performance hit for metric Recall, especially for predicting unseen relations.
Hongshuo Tian, Ning Xu 0003, Yanhui Wang 0001, Chenggang Yan 0001, Bolun Zheng, Xuanya Li, Anan Liu
ICME5
2023 Semantic Embedding Uncertainty Learning for Image and Text Matching
abstract
Image and text matching measures the semantic similarity for cross-modal retrieval. The core of this task is semantic embedding, which mines the intrinsic characteristics of visual and textual for discriminative representation. However, cross-modal ambiguity of image and text (the existence of one-to-many associations) is prone to semantic diversity. The mainstream approaches utilized the fixed point embedding to represent semantics, which ignored the embedding uncertainty caused by semantic diversity leading to incorrect results. To address this issue, we propose a novel Semantic Embedding Uncertainty Learning (SEUL), which represents the embedding uncertainty of image and text as Gaussian distributions and simultaneously learns the salient embedding (mean) and uncertainty (variance) in the common space. We design semantic uncertainty embedding for facilitating the robustness of the representation in the semantic diversity context. A combined objective function is proposed, which optimizes the semantic uncertainty and maintains discriminability to enhance cross-modal associations. Extended experiments are performed on two datasets to demonstrate advanced performance.
Yan Wang 0114, Yuting Su 0001, Wenhui Li 0001, Chenggang Yan 0001, Bolun Zheng, Xuanya Li, Anan Liu
ICME5
2023 Aggregating transformers and CNNs for salient object detection in optical remote sensing images
Liuxin Bao, Xiaofei Zhou 0003, Bolun Zheng, Haibing Yin, Zunjie Zhu, Jiyong Zhang 0001, Chenggang Yan 0001
Neurocomputing3
2023 SRI-Net: Similarity retrieval-based inference network for light field salient object detection
Chengtao Lv, Xiaofei Zhou 0003, Deyang Liu, Bolun Zheng, Jiyong Zhang 0001, Chenggang Yan 0001
J. Vis. Commun. Image Represent.5
2023 CANet: Context-aware Aggregation Network for Salient Object Detection of Surface Defects
Bin Wan, Xiaofei Zhou 0003, Mang Xiao, Yaoqi Sun, Bolun Zheng, Jiyong Zhang 0001, Chenggang Yan 0001
J. Vis. Commun. Image Represent.6
2023 Caps-SSENet: An Improved Estimation Method for SAR Ship Size
abstract
Accurate estimation of the sizes of ship targets plays a critical role in the task of ship classification in synthetic aperture radar (SAR) images. Existing deep neural networks (DNNs)-based methods for SAR ship size estimation (SSE) often adopt a fully connected structure that has limited capability in accurately modeling the relationships of features extracted from SAR images, leading to degraded performance of size estimation. It has been demonstrated that capsule networks provide new guidelines to capture relationships of image features by replacing traditional neurons with capsules, where the dynamic routing strategy is used to calculate correlations among capsules. In this letter, we propose an improved method for SAR SSE based on the capsule network named Caps-SSE network (SSENet). In our Caps-SSENet, a capsule-neural-mixing size mapping module is designed to transform the extracted image features into capsules and complete the estimation of ship sizes using informative feature correlations from dynamic routing. In addition, an average scaled mean square error (ASMSE) loss is proposed to improve the size estimation performance of small ships. Experimental results based on measured SAR data show that the proposed method reduces the estimation error of ship sizes in SAR images in comparison with the existing state-of-the-art method.
Yu Liu 0005, Xueqian Wang 0002, Zhizhuo Jiang, Gang Li 0008, Bolun Zheng, Jiyong Zhang 0001, You He 0003
IEEE Geosci. Remote. Sens. Lett.7
2023 Depth-guided deep filtering network for efficient single image bokeh rendering
Bolun Zheng, Xiaofei Zhou 0003, Aiai Huang, Yaoqi Sun, Chuqiao Chen, Chenggang Yan 0001, Shanxin Yuan
Neural Comput. Appl.2
2023 Transformer-Based Multi-Scale Feature Integration Network for Video Saliency Prediction
abstract
Most cutting-edge video saliency prediction models rely on spatiotemporal features extracted by 3D convolutions due to its local contextual cues acquirement ability. However, the shortage of 3D convolutions is that it cannot effectively capture long-term spatiotemporal dependencies in videos. To address this limitation, we propose a novel Transformer-based Multi-scale Feature Integration Network (TMFI-Net) for video saliency prediction, where the proposed TMFI-Net consists of a semantic-guided encoder and a hierarchical decoder. Firstly, embarking on the Transformer-based multi-level spatiotemporal features, the semantic-guided encoder enhances the features by inserting the high-level feature into each level feature via a top-down pathway and a longitudinal connection, which endows the multi-level spatiotemporal features with rich contextual information. In this way, the features are steered to give more concerns to saliency regions. Secondly, the hierarchical decoder employs a multi-dimensional attention (MA) module to elevate features along channel, temporal, and spatial dimensions jointly. Successively, the hierarchical decoder deploys a progressive decoding block to conduct an initial saliency prediction, which provides a coarse localization of saliency regions. Lastly, considering the complementarity of different saliency predictions, we integrate all initial saliency prediction results into the final saliency map. Comprehensive experimental results on four video saliency datasets firmly demonstrate that our model achieves superior performance when compared with the state-of-the-art video saliency models. The code is available athttps://github.com/wusonghe/TMFI-Net.
Xiaofei Zhou 0003, Songhe Wu, Bolun Zheng, Shuai Wang 0003, Haibing Yin, Jiyong Zhang 0001, Chenggang Yan 0001
IEEE Trans. Circuits Syst. Video Technol.4
2023 Edge-Guided Recurrent Positioning Network for Salient Object Detection in Optical Remote Sensing Images
abstract
Optical remote sensing images (RSIs) have been widely used in many applications, and one of the interesting issues about optical RSIs is the salient object detection (SOD). However, due to diverse object types, various object scales, numerous object orientations, and cluttered backgrounds in optical RSIs, the performance of the existing SOD models often degrade largely. Meanwhile, cutting-edge SOD models targeting optical RSIs typically focus on suppressing cluttered backgrounds, while they neglect the importance of edge information which is crucial for obtaining precise saliency maps. To address this dilemma, this article proposes an edge-guided recurrent positioning network (ERPNet) to pop-out salient objects in optical RSIs, where the key point lies in the edge-aware position attention unit (EPAU). First, the encoder is used to give salient objects a good representation, that is, multilevel deep features, which are then delivered into two parallel decoders, including: 1) an edge extraction part and 2) a feature fusion part. The edge extraction module and the encoder form a U-shape architecture, which not only provides accurate salient edge clues but also ensures the integrality of edge information by extra deploying the intraconnection. That is to say, edge features can be generated and reinforced by incorporating object features from the encoder. Meanwhile, each decoding step of the feature fusion module provides the position attention about salient objects, where position cues are sharpened by the effective edge information and are used to recurrently calibrate the misaligned decoding process. After that, we can obtain the final saliency map by fusing all position attention cues. Extensive experiments are conducted on two public optical RSIs datasets, and the results show that the proposed ERPNet can accurately and completely pop-out salient objects, which consistently outperforms the state-of-the-art SOD models.
Xiaofei Zhou 0003, Kunye Shen, Li Weng, Runmin Cong, Bolun Zheng, Jiyong Zhang 0001, Chenggang Yan 0001
IEEE Trans. Cybern.5
2023 Information-Containing Adversarial Perturbation for Combating Facial Manipulation Systems
abstract
With the development of deep learning technology, the facial manipulation system has become powerful and easy to use. Such systems can modify the attributes of the given facial images, such as hair color, gender, and age. Malicious applications of such systems pose a serious threat to individuals’ privacy and reputation. Existing studies have proposed various approaches to protect images against facial manipulations. Passive defense methods aim to detect whether the face is real or fake, which works for posterior forensics but can not prevent malicious manipulation. Initiative defense methods protect images upfront by injecting adversarial perturbations into images to disrupt facial manipulation systems but can not identify whether the image is fake. To address the limitation of existing methods, we propose a novel two-tier protection method named Information-containing Adversarial Perturbation (IAP), which provides more comprehensive protection for facial images. We use an encoder to map a facial image and its identity message to a cross-model adversarial example which can disrupt multiple facial manipulation systems to achieve initiative protection. Recovering the message in adversarial examples with a decoder serves passive protection, contributing to provenance tracking and fake image detection. We introduce a feature-level correlation measurement that is more suitable to measure the difference between the facial images than the commonly used mean squared error. Moreover, we propose a spectral diffusion method to spread messages to different frequency channels, thereby improving the robustness of the message against facial manipulation. Extensive experimental results demonstrate that our proposed IAP can recover the messages from the adversarial examples with high average accuracy and effectively disrupt the facial manipulation systems.
Yao Zhu 0003, Yuefeng Chen, Rong Zhang 0006, Xiang Tian 0002, Bolun Zheng, Yaowu Chen
IEEE Trans. Inf. Forensics Secur.6
2022 DomainPlus: Cross Transform Domain Learning towards High Dynamic Range Imaging
abstract
High dynamic range (HDR) imaging by combining multiple low dynamic range (LDR) images of different exposures provides a promising way to produce high quality photographs. However, the misalignment between the input images leads to ghosting artifacts in the reconstructed HDR image. In this paper, we propose a cross-transform domain neural network for efficient HDR imaging. Our approach consists of two modules: a merging module and a restoration module. For the merging module, we propose a Multiscale Attention with Fronted Fusion (MAFF) mechanism to achieve coarse-to-fine spatial fusion. For the restoration module, we propose fronted Discrete Wavelet Transform (DWT) and Discrete Cosine Transform (DCT)-based learnable bandpass filters to formulate a cross-transform domain learning block, dubbed DomainPlus Block (DPB) for effective ghosting removal. Our ablation study and comprehensive experiments show that DomainPlus outperforms the existing state-of-the-art on several datasets.
Bolun Zheng, Xiaokai Pan, Xiaofei Zhou 0003, Gregory Slabaugh, Chenggang Yan 0001, Shanxin Yuan
ACM Multimedia1
2022 Boosting Out-of-distribution Detection with Typical Features
abstract
Out-of-distribution (OOD) detection is a critical task for ensuring the reliability and safety of deep neural networks in real-world scenarios. Different from most previous OOD detection methods that focus on designing OOD scores or introducing diverse outlier examples to retrain the model, we delve into the obstacle factors in OOD detection from the perspective of typicality and regard the feature's high-probability region of the deep model as the feature's typical set. We propose to rectify the feature into its typical set and calculate the OOD score with the typical features to achieve reliable uncertainty estimation. The feature rectification can be conducted as a plug-and-play module with various OOD scores. We evaluate the superiority of our method on both the commonly used benchmark (CIFAR) and the more challenging high-resolution benchmark with large label space (ImageNet). Notably, our approach outperforms state-of-the-art methods by up to 5.11% in the average FPR95 on the ImageNet benchmark.
Yao Zhu 0003, Yuefeng Chen, Chuanlong Xie, Rong Zhang 0006, Hui Xue 0001, Xiang Tian 0002, Bolun Zheng, Yaowu Chen
NeurIPS8
2022 Bidirectional difference locating and semantic consistency reasoning for change captioning
abstract
Change captioning is an emerging task to describe the changes between a pair of images. The difficulty in this task is to discover the differences between the two images. Recently, some methods have been proposed to address this problem. However, they all employ unidirectional difference localization to identify the changes. This can lead to ambiguity about the nature of the changes. Instead, we propose a framework with bidirectional difference localization and semantic consistency reasoning to describe the image changes. First, we locate the changes in the two images by capturing bidirectional differences. Then we design a decoder with spatial-channel attention to generate the change caption. Finally, we introduce semantic consistency reasoning to constrain our bidirectional difference localization module and spatial-channel attention module. Extensive experiments on three public data sets show that the performance of our proposed model outperforms the state-of-the-art change captioning models by a large margin.
Yaoqi Sun, Liang Li 0003, Tongyv Lu, Bolun Zheng, Chenggang Yan 0001, Yongjun Bao, Guiguang Ding, Gregory Slabaugh
Int. J. Intell. Syst.5
2022 Learning Frequency Domain Priors for Image Demoireing
abstract
Image demoireing is a multi-faceted image restoration task involving both moire pattern removal and color restoration. In this paper, we raise a general degradation model to describe an image contaminated by moire patterns, and propose a novel multi-scale bandpass convolutional neural network (MBCNN) for single image demoireing. For moire pattern removal, we propose a multi-block-size learnable bandpass filters (M-LBFs), based on a block-wise frequency domain transform, to learn the frequency domain priors of moire patterns. We also introduce a new loss function named Dilated Advanced Sobel loss (D-ASL) to better sense the frequency information. For color restoration, we propose a two-step tone mapping strategy, which first applies a global tone mapping to correct for a global color shift, and then performs local fine tuning of the color per pixel. To determine the most appropriate frequency domain transform, we investigate several transforms including DCT, DFT, DWT, learnable non-linear transform and learnable orthogonal transform. We finally adopt the DCT. Our basic model won the AIM2019 demoireing challenge. Experimental results on three public datasets show that our method outperforms state-of-the-art methods by a large margin.
Bolun Zheng, Shanxin Yuan, Chenggang Yan 0001, Xiang Tian 0002, Jiyong Zhang 0001, Yaoqi Sun, Lin Liu 0016, Ales Leonardis, Gregory Slabaugh
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Each Part Matters: Local Patterns Facilitate Cross-View Geo-Localization
abstract
Cross-view geo-localization is to spot images of the same geographic target from different platforms,e.g., drone-view cameras and satellites. It is challenging in the large visual appearance changes caused by extreme viewpoint variations. Existing methods usually concentrate on mining the fine-grained feature of the geographic target in the image center, but underestimate the contextual information in neighbor areas. In this work, we argue that neighbor areas can be leveraged as auxiliary information, enriching discriminative clues for geo-localization. Specifically, we introduce a simple and effective deep neural network, called Local Pattern Network (LPN), to take advantage of contextual information in an end-to-end manner. Without using extra part estimators, LPN adopts a square-ring feature partition strategy, which provides the attention according to the distance to the image center. It eases the part matching and enables the part-wise representation learning. Owing to the square-ring partition design, the proposed LPN has good scalability to rotation variations and achieves competitive results on three prevailing benchmarks,i.e., University-1652, CVUSA and CVACT. Besides, we also show the proposed LPN can be easily embedded into other frameworks to further boost performance.
Tingyu Wang 0002, Zhedong Zheng, Chenggang Yan 0001, Jiyong Zhang 0001, Yaoqi Sun, Bolun Zheng, Yi Yang 0001
IEEE Trans. Circuits Syst. Video Technol.6
2022 CBREN: Convolutional Neural Networks for Constant Bit Rate Video Quality Enhancement
abstract
Constant bit rate (CBR) videos are widely used in streaming playback applications. However, the image quality of the CBR video is often unstable, especially for scenes with large motion. To this end, we design a new model to represent the distortion of High Efficiency Video Coding (HEVC) constant bit rate video, and propose a neural network for a constant bit rate video quality enhancement (CBREN). We propose a dual-domain restoration module (DRM) to jointly learn the prior knowledge in the pixel domain and the frequency domain. To address the degradation resulting from compression, we propose a two-step quantization degradation estimation strategy. The Inverse DCT (IDCT) Translation Unit (ITU) is used to constrain the quantization table of the constant bit rate video to a suitable range, and the Dynamic Alpha Unit (DAU) is used to fine-tune the quantization table according to the content of each frame. In order to effectively reduce the block distortion of different sizes produced in the compression process, we adopt a multi-scale network. Extensive experiments show that our approach can greatly enhance the quality of CBR compressed video. Moreover, our method can also be applied to constant quantization parameter (CQP) video enhancement tasks, and is certainly superior to existing methods.
Hengrun Zhao, Bolun Zheng, Shanxin Yuan, Chenggang Yan 0001, Liang Li 0003, Gregory Slabaugh
IEEE Trans. Circuits Syst. Video Technol.2
2022 Toward Understanding and Boosting Adversarial Transferability From a Distribution Perspective
abstract
Transferable adversarial attacks against Deep neural networks (DNNs) have received broad attention in recent years. An adversarial example can be crafted by a surrogate model and then attack the unknown target model successfully, which brings a severe threat to DNNs. The exact underlying reasons for the transferability are still not completely understood. Previous work mostly explores the causes from the model perspective, e.g., decision boundary, model architecture, and model capacity. Here, we investigate the transferability from the data distribution perspective and hypothesize that pushing the image away from its original distribution can enhance the adversarial transferability. To be specific, moving the image out of its original distribution makes different models hardly classify the image correctly, which benefits the untargeted attack, and dragging the image into the target distribution misleads the models to classify the image as the target class, which benefits the targeted attack. Towards this end, we propose a novel method that crafts adversarial examples by manipulating the distribution of the image. We conduct comprehensive transferable attacks against multiple DNNs to demonstrate the effectiveness of the proposed method. Our method can significantly improve the transferability of the crafted attacks and achieves state-of-the-art performance in both untargeted and targeted scenarios, surpassing the previous best method by up to 40% in some cases. In summary, our work provides new insight into studying adversarial transferability and provides a strong counterpart for future research on adversarial defense.
Yao Zhu 0003, Yuefeng Chen, Kejiang Chen, Yuan He 0011, Xiang Tian 0002, Bolun Zheng, Yaowu Chen, Qingming Huang
IEEE Trans. Image Process.7
2022 Age-Invariant Face Recognition by Multi-Feature Fusionand Decomposition with Self-attention
abstract
Different from general face recognition, age-invariant face recognition (AIFR) aims at matching faces with a big age gap. Previous discriminative methods usually focus on decomposing facial feature into age-related and age-invariant components, which suffer from the loss of facial identity information. In this article, we propose a novel Multi-feature Fusion and Decomposition (MFD) framework for age-invariant face recognition, which learns more discriminative and robust features and reduces the intra-class variants. Specifically, we first sample multiple face images of different ages with the same identity as a face time sequence. Then, the multi-head attention is employed to capture contextual information from facial feature series, extracted by the backbone network. Next, we combine feature decomposition with fusion based on the face time sequence to ensure that the final age-independent features effectively represent the identity information of the face and have stronger robustness against the aging process. Besides, we also mitigate imbalanced age distribution in the training data by a re-weighted age loss. We experimented with the proposed MFD over the popular CACD and CACD-VS datasets, where we show that our approach improves the AIFR performance than previous state-of-the-art methods. We simultaneously show the performance of MFD on LFW dataset.
Chenggang Yan 0001, Lixuan Meng, Liang Li 0003, Jian Yin 0003, Jiyong Zhang 0001, Yaoqi Sun, Bolun Zheng
ACM Trans. Multim. Comput. Commun. Appl.9
2021 Evolution of ICTs-empowered-identification: A general re-ranking method for person re-identification
Tongkun Xu, Bolun Zheng, Yaoqi Sun, Anan Liu, Zhendong Mao 0001, Chenggang Yan 0001
Pattern Recognit. Lett.3
2021 Dynamic Selective Network for RGB-D Salient Object Detection
abstract
RGB-D saliency detection is receiving more and more attention in recent years. There are many efforts have been devoted to this area, where most of them try to integrate the multi-modal information, i.e. RGB images and depth maps, via various fusion strategies. However, some of them ignore the inherent difference between the two modalities, which leads to the performance degradation when handling some challenging scenes. Therefore, in this paper, we propose a novel RGB-D saliency model, namely Dynamic Selective Network (DSNet), to perform salient object detection (SOD) in RGB-D images by taking full advantage of the complementarity between the two modalities. Specifically, we first deploy a cross-modal global context module (CGCM) to acquire the high-level semantic information, which can be used to roughly locate salient objects. Then, we design a dynamic selective module (DSM) to dynamically mine the cross-modal complementary information between RGB images and depth maps, and to further optimize the multi-level and multi-scale information by executing the gated and pooling based selection, respectively. Moreover, we conduct the boundary refinement to obtain high-quality saliency maps with clear boundary details. Extensive experiments on eight public RGB-D datasets show that the proposed DSNet achieves a competitive and excellent performance against the current 17 state-of-the-art RGB-D SOD models.
Hongfa Wen, Chenggang Yan 0001, Xiaofei Zhou 0003, Runmin Cong, Yaoqi Sun, Bolun Zheng, Jiyong Zhang 0001, Yongjun Bao, Guiguang Ding
IEEE Trans. Image Process.6
2020 Image Demoireing with Learnable Bandpass Filters
abstract
Image demoireing is a multi-faceted image restoration task involving both texture and color restoration. In this paper, we propose a novel multiscale bandpass convolutional neural network (MBCNN) to address this problem. As an end-to-end solution, MBCNN respectively solves the two sub-problems. For texture restoration, we propose a learnable bandpass filter (LBF) to learn the frequency prior for moire texture removal. For color restoration, we propose a two-step tone mapping strategy, which first applies a global tone mapping to correct for a global color shift, then performs local fine tuning of the color per pixel. Through an ablation study, we demonstrate the effectiveness of the different components of MBCNN. Experimental results on two public datasets show that our method outperforms state-of-the-art methods by a large margin (more than 2dB in terms of PSNR).
Bolun Zheng, Shanxin Yuan, Gregory Slabaugh, Ales Leonardis
CVPR1
2020 Implicit Dual-Domain Convolutional Network for Robust Color Image Compression Artifact Reduction
abstract
Several dual-domain convolutional neural network-based methods show outstanding performance in reducing image compression artifacts. However, they are unable to handle color images as the compression processes for gray scale and color images are different. Moreover, these methods train a specific model for each compression quality, and they require multiple models to achieve different compression qualities. To address these problems, we proposed an implicit dual-domain convolutional network (IDCN) with a pixel position labeling map and quantization tables as inputs. We proposed an extractor-corrector framework-based dual-domain correction unit (DCU) as the basic component to formulate the IDCN; the implicit dual-domain translation allows the IDCN to handle color images with discrete cosine transform (DCT)-domain priors. A flexible version of IDCN (IDCN-f) was also developed to handle a wide range of compression qualities. Experiments for both objective and subjective evaluations on benchmark datasets show that IDCN is superior to the state-of-the-art methods and IDCN-f exhibits excellent abilities to handle a wide range of compression qualities with a little trade-off against performance; further, it demonstrates great potential for practical applications.
Bolun Zheng, Yaowu Chen, Xiang Tian 0002, Fan Zhou 0007
IEEE Trans. Circuits Syst. Video Technol.1