Bo Han 0004

dblp:241/0472-0004 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0001-6282-5428ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Efficient and Accurate Object Detection With Asymmetric Progressive Semi-Decoupled Head and Harmonic Focal Loss
abstract
Efficiently and accurately recognizing interesting objects within the image and regressing bounding boxes to enclose them has been a persistent pursuit in object detection. However, existing detectors fail to achieve both aspects simultaneously due to insufficient task interaction and suboptimal classification behavior. To solve the problem, this paper proposes a novel detector with Efficient Asymmetric Progressive Semi-Decoupled Head (EAPSDH) and Harmonic Focal Loss (HFL). Specifically, we generalize the detection head into a progressive asymmetric paradigm that performs hierarchical and dynamically recalibrated interaction between classification and localization, enabling iterative mutual enhancement in an efficient manner beyond the prior designs. Meanwhile, HFL is proposed to improve classifier optimization by addressing the imbalance between positive and negative samples. HFL dynamically increases the loss weights of positive samples, amplifying their gradient contributions during classifier training, which significantly reduces classification error. By jointly improving task-specific feature representation and classification optimization, EAPSDH and HFL complement each other to alleviate the inconsistency between classification and localization performance, resulting in an efficient and accurate one-stage detector termed EADet. Experimental results on the MS COCO database demonstrate that EADet effectively mitigates the inconsistency between classification and localization performance. Furthermore, EADet achieves a strong trade-off between accuracy and speed, reaching 47.4 AP at 33.2 FPS on the MS COCO with ResNet-101 under the $2\times $ training schedule, demonstrating its effectiveness compared with recent state-of-the-art detectors. Code will be available at https://github.com/HB-X/EADet.
Bo Han 0004, Lihuo He, Junjie Ke, Jiehao Tang, Di Wang 0011, Xinbo Gao 0001
IEEE Trans. Image Process.1
2025 A Two-Stage AIGC Image Quality Assessment with T2I Correspondence and Visual Perception
abstract
Image quality assessment (IQA) of artificial intelligence-generated content (AIGC) has recently attracted significant research attention. Unlike general-purpose IQA, which primarily focuses on evaluating image content, AIGCIQA often requires addressing both the Text-to-Image (T2I) correspondence and the perceptual quality of images. To address this requirement, this paper proposes a novel two-stage AIGCIQA method. The first stage evaluates the alignment of the AI-generated images (AIGIs) with their corresponding descriptions, serving as an indicator of overall image quality. Specifically, positive and negative prompts are constructed to describe the T2I correspondence degree, and then a CLIP model is employed to predict the degree based on these image-prompt pairs. The second stage refines the perceptual quality assessment by integrating both global and local degradation features of AIGIs. Importantly, the contribution of local features is measured according to their correlation with the overall image, ensuring key regions are adequately represented in the quality prediction. Experimental results on AGIQA-1K, AGIQA-3K, and AIGCIQA2023 demonstrate the superior performance of the proposed method.
Jili Xia, Lihuo He, Bo Hu 0008, Bo Han 0004, Xinbo Gao 0001
ICASSP4
2025 Progressive Semi-Decoupled Detector for Accurate Object Detection
abstract
Inconsistent accuracy between classification and localization tasks is a common challenge in modern object detection. Task decoupling, which employs distinct features or labeling strategies for each task, is a widely used approach to address this issue. Although it has led to noteworthy advancements, this approach is insufficient as it neglects task interdependence and lacks an explicit consistency constraint. To bridge this gap, this paper proposes the Progressive Semi-Decoupled Detector (ProSDD) to enhance both classification and localization accuracy. Specifically, a new detection head is designed that incorporates feature suppression and enhancement mechanism (FSEM) and bidirectional interaction module (BIM). Compared with the decoupled head, it not only filters out task-irrelevant information and enhances task-related information, but also avoids excessive decoupling at the feature level. Moreover, both FSEM and BIM are used multiple times, thus forming a progressive semi-decoupled head. Then, a novel consistency loss is proposed and integrated into the loss function of object detection, ensuring harmonic performance in classification and localization. Experimental results demonstrate that the proposed ProSDD effectively alleviates inconsistent accuracy and achieves high-quality object detection. Taking the pretrained ResNet-50 as the backbone, ProSDD achieves a remarkable 43.3 AP on the MS COCO dataset, surpassing contemporary state-of-the-art detectors by a substantial margin under the equivalent configurations. Code is available athttps://github.com/HB-X/ProSDD.
Bo Han 0004, Lihuo He, Junjie Ke, Jinjian Wu, Xinbo Gao 0001
IEEE Trans. Multim.1
2024 Weighted parallel decoupled feature pyramid network for object detection
Bo Han 0004, Lihuo He, Junjie Ke, Chenwei Tang, Xinbo Gao 0001
Neurocomputing1
2024 ProFPN: Progressive feature pyramid network with soft proposal assignment for object detection
Junjie Ke, Lihuo He, Bo Han 0004, Jie Li 0001, Xinbo Gao 0001
Knowl. Based Syst.3
2024 General Deformable RoI Pooling and Semi-Decoupled Head for Object Detection
abstract
Object detection aims to classify interest objects within an image and pinpoint their positions using predicted rectangular bounding boxes. However, classification and localization tasks are heterogeneous, not only spatially misaligned but also differing in properties and feature requirements. Modern detectors commonly share the spatial region and detection head for both tasks, making them challenging to achieve optimal performance altogether, resulting in inconsistent accuracy. Specifically, the predicted bounding box may have higher classification confidence but lower localization quality, or vice versa. To tackle this issue, the spatial decoupling mechanism via general deformable RoI pooling is first proposed. This mechanism separately pursues the favorable regions for classification and localization, and subsequently extracts the corresponding features. Then, the semi-decoupled head is designed. Compared to the decoupled head that utilizes independent classification and localization networks, potentially leading to excessive decoupling and compromised detection performance, the semi-decoupled head enables the networks to mutually enhance each other while concentrating on their respective tasks. In addition, the semi-decoupled head also introduces a redundancy suppression module to filter out redundant task-irrelevant information of features extracted by separate networks and reinforce task-related information. By combining the spatial decoupling mechanism with the semi-decoupled head, the proposed detector achieves an impressive 43.7 AP in Faster R-CNN framework with ResNet-101 as backbone network. Without bells and whistles, extensive experimental results on the popular MS COCO dataset demonstrate that the proposed detector suppresses the baseline by a significant margin and outperforms some state-of-the-art detectors. Code is available athttps://github.com/HB-X/gdpool_semi_dehead.
Bo Han 0004, Lihuo He, Wen Lu 0004, Xinbo Gao 0001
IEEE Trans. Multim.1
2024 VLDadaptor: Domain Adaptive Object Detection With Vision-Language Model Distillation
abstract
Domain adaptive object detection (DAOD) aims to develop a detector trained on labeled source domains to identify objects in unlabeled target domains. A primary challenge in DAOD is the domain shift problem. Most existing methods learn domain-invariant features within single domain embedding space, often resulting in heavy model biases due to the intrinsic data properties of source domains. To mitigate the model biases, this paper proposes VLDadaptor, a domain adaptive object detector based on vision-language models (VLMs) distillation. Firstly, the proposed method integrates domain-mixed contrastive knowledge distillation between the visual encoder of CLIP and the detector by transferring category-level instance features, which guarantees the detector can extract domain-invariant visual instance features across domains. Then, VLDadaptor employs domain-mixed consistency distillation between the text encoder of CLIP and detector by aligning text prompt embeddings with visual instance features, which helps to maintain the category-level feature consistency among the detector, text encoder and the visual encoder of VLMs. Finally, the proposed method further promotes the adaptation ability by adopting a prompt-based memory bank to generate semantic-complete features for graph matching. These contributions enable VLDadaptor to extract visual features into the visual-language embedding space without any evident model bias towards specific domains. Extensive experimental results demonstrate that the proposed method achieves state-of-the-art performance on Pascal VOC to Clipart adaptation tasks and exhibits high accuracy on driving scenario tasks with significantly less training time.
Junjie Ke, Lihuo He, Bo Han 0004, Jie Li 0001, Di Wang 0011, Xinbo Gao 0001
IEEE Trans. Multim.3
2020 Weighted Guided Image Filtering With Steering Kernel
abstract
Due to its local property, guided image filter (GIF) generally suffers from halo artifacts near edges. To make up for the deficiency, a weighted guided image filter (WGIF) was proposed recently by incorporating an edge-aware weighting into the filtering process. It takes the advantages of local and global operations, and achieves better performance in edge-preserving. However, edge direction, a vital property of the guidance image, is not considered fully in these guided filters. In order to overcome the drawback, we propose a novel version of GIF, which can leverage the edge direction more sufficiently. In particular, we utilize the steering kernel to adaptively learn the direction and incorporate the learning results into the filtering process to improve the filter's behavior. Theoretical analysis shows that the proposed method can get more powerful performance with preserving edges and reducing halo artifacts effectively. Similar conclusions are also reached through the thorough experiments including edge-aware smoothing, detail enhancement, denoising and dehazing.
Zhonggui Sun, Bo Han 0004, Jie Li 0001, Xinbo Gao 0001
IEEE Trans. Image Process.2