Junfei Yi

dblp:274/4774 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
17since 2021 · last 2026
0000-0003-1737-0078ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Distilling Object Detectors via Monte Carlo Dropout
abstract
Knowledge distillation (KD) has become a fundamental technique for model compression in object detection tasks. The data noise and training randomness may cause the knowledge of the teacher model to be unreliable, referred to as knowledge uncertainty. Existing methods neglect this uncertainty, potentially hindering the student's capacity to capture and understand latent "dark knowledge". In this work, we introduce a novel strategy that explicitly incorporates knowledge uncertainty, named Uncertainty-Driven Knowledge Extraction and Transfer (UET). Given the unknown, high-dimensional nature of the knowledge distribution, we employ Monte Carlo dropout to effectively estimate the teacher's uncertainty. Leveraging information theory, we combine uncertainty with deterministic knowledge, enabling the student to benefit from both precision and diversity. UET is a plug-and-play method that integrates seamlessly with existing distillation techniques. We validate our approach through comprehensive experiments across various distillation strategies, detectors, and backbones. Specifically, UET achieves state-of-the-art results, with a ResNet50-based GFL detector obtaining 44.1% mAP on the COCO dataset-surpassing baseline performance by 3.9%.
Junfei Yi, Hui Zhang 0023, Jianxu Mao, Tengfei Liu 0005, Mingjie Li 0006, Sihao Lin, Hanyu Gu, Zhihui Li 0001, Xiaojun Chang, Yaonan Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 MGLD-TLNet: Multigeometric and Long-Distance Representation Network for Transmission Line Inspection
abstract
Effective transmission line (TL) inspection in complex corridor environments is essential for ensuring reliable power delivery. This work presents a 3D-based perception method for this task. The proposed method is designed by considering two key characteristics of TL inspection. First, the point cloud data are sparse and class distributions are highly imbalanced, which weakens the signals from thin conductors and tower components. To address this issue, we model long-range spatial relations along the corridor to mitigate data sparsity and imbalance. Second, strong structural correlations exist between conductors and towers, which can be leveraged to improve perception performance. To exploit this property, we construct a unified 3-D representation that jointly models towers, conductors, and vegetation, while fusing Cartesian and polar geometries through geometry-aware alignment. Experiments on real-world corridor datasets demonstrate that the proposed method, termed multigeometric and long-distance TL perception Network (MGLD-TLNet), consistently improves stability and accuracy under conditions of sparsity, occlusion, and complex environmental interactions.
Hui Zhang 0023, Kaining Zhang, Baheti Biekezat, Hang Zhong, Junfei Yi, Jianxu Mao, Yaonan Wang 0001
IEEE Trans. Cybern.6
2026 Investigating a Unified 3-D Object Detection Method for Different Multibeam LiDAR
Ziming Tao, Jianxu Mao, Yaonan Wang 0001, Caiping Liu, Junfei Yi, Zhenyu He 0015, Xiaojun Chang, Hui Zhang 0023
IEEE Trans. Ind. Informatics5
2026 You Can Only Tune Normalization: A Simple and Effective Approach to Parameter-Efficient Fine-Tuning
abstract
To tackle the issue of excessive parameter volumes during fine-tuning of large-scale pre-trained models with full parameters, Parameter-Efficient Fine-Tuning (PEFT) methods have been introduced. The core concept involves freezing the backbone network of the model and updating only a small subset of parameters. This strategy not only decreases the number of parameters needed for training but also delivers performance comparable to Full-Tuning, even surpassing it on certain datasets. However, most popular PEFT methods introduce extra parameters or modules for fine-tuning, which come with inherent limitations. In response, we propose a straightforward and efficient PEFT method called You Can Only Tune Normalization (YONO). YONO focuses solely on tuning the normalization layer and the final classification layer of the model. This method avoids adding extra modules, making it easily applicable to any model without causing inference delays. We extensively tested YONO on 28 benchmark datasets, and the results indicate that it requires significantly fewer parameters compared to other advanced PEFT methods. Additionally, we validated YONO’s efficiency and generalizability across various vision models. Finally, we further explore the essence of PEFT methods, whether they learn new knowledge or expose the capabilities that a model has already learned. Our findings suggest that YONO is more sensitive to improvements in dataset quality, making it a promising candidate for future scaling to larger models.
Lingyun Huang, Jianxu Mao, Junfei Yi, Ziming Tao, Ziyang Peng, Wei He 0001, Rui Liu 0028, Yaonan Wang 0001
ACM Trans. Intell. Syst. Technol.3
2026 Image-Quality-Guided Consistency-Alignment Network for Fault Diagnosis
abstract
Signal-to-image transformation has been widely used in mechanical fault diagnosis to provide unified visual representations for downstream diagnostic models. However, the resulting diagnostic images can exhibit substantial quality variations caused by measurement noise, sensor degradation, and partial signal loss. Most existing methods either discard low-quality samples or implicitly assume equal reliability across samples, leading to information loss or degraded robustness. This paper proposes an image-quality-guided consistency-alignment network that explicitly estimates sample quality and leverages degraded data during training. First, vibration signals are decomposed into sub-bands, and energy-ratio criteria are used to select informative components for reconstruction. The reconstructed signals are subsequently fused into fixed-layout two-dimensional images via a sector-allocation strategy. Next, a diagnosis-aware quality assessment module assigns a quality score to each image to guide training. Low-quality samples are further regularized via cross-quality feature alignment using a quality-weighted supervised pairwise loss, encouraging them to align with high-quality counterparts from the same class. Finally, a Swin Transformer backbone performs classification. Experiments on the proprietary RGFD dataset and the public SEUB dataset demonstrate consistent gains under controlled mixed-quality settings, with accuracies of up to 99.4% and 98.7%, respectively, while additional evaluations under reproducible degradations further confirm the robustness of the proposed quality-aware learning mechanism.
Zhuowei Li 0011, Jianxu Mao, Yaonan Wang 0001, Junfei Yi, Caiping Liu, Hui Zhang 0023
IEEE Trans. Reliab.5
2025 HC-LLM: Historical-Constrained Large Language Models for Radiology Report Generation
abstract
Radiology report generation (RRG) models typically focus on individual exams, often overlooking the integration of historical visual or textual data, which is crucial for patient follow-ups. Traditional methods usually struggle with long sequence dependencies when incorporating historical information, but large language models (LLMs) excel at in-context learning, making them well-suited for analyzing longitudinal medical data. In light of this, we propose a novel Historical-Constrained Large Language Models (HC-LLM) framework for RRG, empowering LLMs with longitudinal report generation capabilities by constraining the consistency and differences between longitudinal images and their corresponding reports. Specifically, our approach extracts both time-shared and time-specific features from longitudinal chest X-rays and diagnostic reports to capture disease progression. Then, we ensure consistent representation by applying intra-modality similarity constraints and aligning various features across modalities with multimodal contrastive and structural constraints. These combined constraints effectively guide the LLMs in generating diagnostic reports that accurately reflect the progression of the disease, achieving state-of-the-art results on the Longitudinal-MIMIC dataset. Notably, our approach performs well even without historical data during testing and can be easily adapted to other multimodal large models, enhancing its versatility.
Tengfei Liu 0005, Jiapu Wang, Yongli Hu, Mingjie Li 0006, Junfei Yi, Xiaojun Chang, Junbin Gao
AAAI5
2025 CVPT: Cross Visual Prompt Tuning
Lingyun Huang, Jianxu Mao, Junfei Yi, Ziming Tao, Yaonan Wang 0001
ICCV3
2025 Wavelet-Based Distillation with Structured Frequency Alignment
Pengyu Lu, Junfei Yi, Jianxu Mao, Junlong Yu, Shuohao Xiao, Zhenyu He 0015, Yaonan Wang 0001
ICIG (2)2
2025 FFTA-Net: A Frequency-Domain Fusion and Temporal Alignment Network for Transmission Line Defect Detection
Jianxu Mao, Yaonan Wang 0001, Junlong Yu, Junfei Yi, Zhenyu He 0015, Ziming Tao, Hui Zhang 0023
ICIG (2)5
2025 Foreground-Aware Enhancement-Based Multimodal 3D Object Detection
abstract
LiDAR is one of the most widely used 3D detection sensors in applications such as autonomous driving and unmanned inspection. However, when uniformly sampling the entire scene to generate point cloud data, the number of foreground points reflected by target objects is often significantly lower than that of the background points, which have a larger coverage area. This imbalance poses considerable challenges to the performance of object detection models, especially in detecting small or distant objects. To overcome this challenge, this paper presents a Foreground-Aware Enhancement-based Multimodal 3D Object Detection Method (PFA), which effectively mitigates the low detection accuracy of small and distant objects caused by insufficient foreground points. The proposed method incorporates a Foreground-Aware Enhancement Module (FAEM) and a Region-Focused Attention Module (RFAM). The FAEM module enhances the model’s focus on foreground regions, while the RFAM module strengthens multimodal fused features. Together, these components significantly improve the detection accuracy of small and distant objects. Experimental results on the KITTI dataset demonstrate that the proposed method achieves 3D detection accuracies of 84.65%, 59.56%, and 71.48% for cars, pedestrians, and cyclists, respectively, under the hard evaluation level. Furthermore, the model also shows significant advantages on the KITTI public test set and validation set for both easy and moderate samples, fully validating its effectiveness and generalizability in enhancing multimodal 3D object detection accuracy.
Ziyang Peng, Wei He 0001, Jianxu Mao, Ziming Tao, Junfei Yi, Yaonan Wang 0001
IJCNN5
2025 LDFCDet: Boosting 3D Object Detectors With Low-High Level Feature Crosses Using Laplace Distribution
abstract
Highly accurate 3D object detection is critical for autonomous driving and robotic sensing system. However, some objects with few foreground points significantly affect the accuracy of 3D object detection. As the network depth increases, the low-level features of these objects are gradually lost, especially for the hard object. Due to this issue, current LiDAR-only based and multimodal methods often misclassify background as foreground. Therefore, how to leverage the low-level feature that contain information about these objects in the high layer of the network becomes the key to optimizing the issue. In this paper, we propose LDFCDet, a framework boosting 3D object detectors with low-high level feature crosses using Laplace distribution(LD). In our proposed method, we design a low-high level feature crosses module(LHFCM) to embed low-level feature into high-level feature in the deeper layer of the network, and use Laplace distribution to obtain a new low-high level feature that includes information about these objects with few foreground points. In addition, we propose a res-gated feature aggregation module(RGFAM) to fuse the mutli-scale features. Our approach is well-suited for both LiDAR-based and multimodal methods.We evaluate the LDFCDet on the widely used KITTI dataset, and our method outperforms almost current 3D object detection methods on the challenging KITTI test set. Moreover, we conducted comparative experiments on the ONCE dataset, and the results further demonstrate the effectiveness and superiority of our method.
Zhenyu He 0015, Jianxu Mao, Yaonan Wang 0001, Junlong Yu, Ziming Tao, Junfei Yi, Hui Zhang 0023, Shaoyuan Wang
IEEE Trans. Circuits Syst. Video Technol.6
2025 Tackling Real-World Complexity: Hierarchical Modeling and Dynamic Prompting for Multimodal Long Document Classification
abstract
With the rapid growth of internet content, multimodal long document data has become increasingly prominent, drawing significant attention from researchers. However, most existing methods primarily focus on scenarios where all modalities are present, often overlooking more challenging and realistic cases involving missing image modality. To address this limitation, we propose a robust multimodal long document classification (MLDC) framework that integrates hierarchical modeling and dynamic prompting to handle complex multimodal long document data. Our approach begins by leveraging hierarchical modeling combined with an Adaptive Correlation Multimodal Transformer (ACMT) to effectively capture relationships between text and images at both section and sentence levels. We also introduce a Dynamic Prompt Generation (DPG) module at both levels to enhance the model’s robustness in handling missing image data. By evaluating sample uncertainty, the DPG module dynamically adjusts both the number of prompts and the prompts themselves, allowing the model to better adapt to the varying needs of different samples. Finally, a Hierarchical Heterogeneous Graph (HHG) is introduced to enhance feature interactions across levels, further improving the coherence and accuracy of the model. Extensive experiments on four multi-modal long document datasets demonstrate that our model shows superior performance compared to existing state-of-the-art MLDC classification methods in various conditions.
Tengfei Liu 0005, Yongli Hu, Mingjie Li 0006, Junfei Yi, Xiaojun Chang, Junbin Gao
IEEE Trans. Circuits Syst. Video Technol.4
2025 FMSD: Focal Multi-Scale Shape-Feature Distillation Network for Small Fasteners Detection in Electric Power Scene
abstract
In the electric power scene, fasteners play a pivotal role in securing and connecting electrical equipment, with small fastener detection (SFD) being crucial for ensuring operational stability. Despite the replacement of manual inspection methods by non-destructive techniques employing deep learning, these approaches often demand substantial computational resources and involve numerous parameters. While knowledge distillation (KD) can be a viable solution, existing KD methods may often fail to achieve satisfactory performance when dealing with small object presentation and little inter-class variability in SFD tasks. To alleviate this, we propose a Focal Multi-scale Shape-feature Distillation Network (FMSD) to achieve efficient and precise fastener detection in electric power scenarios. Specifically, we propose a novel Multi-Scale Shape-Aware Feature Aggregation module (MSFA) to augment the network's perception of object shape and scale during the KD process. Additionally, we propose a Contour-Guided Distillation (CGD) module to optimize the transfer of the extracted shape-sensitive knowledge between the teacher and student models. Through a series of experiments compared with existing state-of-the-art (SOTA) methods, our method demonstrates superior performance over existing SOTA techniques, both efficiently and effectively. Furthermore, validation on publicly available power scene datasets confirms the generalizability and adaptability of our proposed FMSD across various settings.
Junfei Yi, Jianxu Mao, Hui Zhang 0023, Mingjie Li 0006, Kai Zeng 0010, Mingtao Feng, Xiaojun Chang, Yaonan Wang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2025 Toward Efficient Power Scene Detection via Topology-Preserved Knowledge Distillation
abstract
The power industry relies on efficient inspection systems to ensure stability and safety. While deep learning has advanced automated inspection, its reliance on custom modules for specific tasks can impact efficiency. Knowledge distillation (KD) offers a balanced solution, but the complex textures and structures of power equipment challenge conventional KD methods, which often fail to capture essential local semantic and topological relationships. To address this, we proposeTopNet, a novel topology-preserved KD framework for power scene detection tasks. Specifically, we model the teacher’s knowledge as a graph, where nodes encode local fine-grained features and edges capture global topological relationships. Based on this, we introduce node feature distillation and edge feature distillation to transfer local–global structural knowledge, which can enhance the student’s ability to perceive objects. Furthermore, we also introduce aggregated feature distillation to incorporate and transfer contextual semantic knowledge. Comprehensive experiments are conducted on two different benchmark datasets to demonstrate that TopNet achieves state-of-the-art detection performance with high efficiency, offering a robust solution for automated power equipment inspection.
Junfei Yi, Tengfei Liu 0005, Jianxu Mao, Yaonan Wang 0001, Hui Zhang 0023, He Xie, Hang Zhong, Xiaojun Chang
IEEE Trans. Ind. Informatics1
2025 Balancing Accuracy and Efficiency With a Multiscale Uncertainty-Aware Knowledge-Based Network for Transmission Line Inspection
abstract
Real-world transmission line inspections (RTLIs) ensure power stability and safety. Deep learning (DL) models have become prevalent approaches for performing RTLI tasks. However, the high computational demands and substantial parameter requirements of DL models limit their real-world applicability. This article introduces a novel approach, a multiscale uncertainty-aware knowledge-based network, which is designed to balance the accuracy and efficiency in RTLI tasks. Specifically, we propose an uncertainty-aware knowledge distillation method that incorporates pixel-level uncertainty into the knowledge transfer process, mitigating the impact of noisy knowledge derived from extra background information contained in ground truths. In addition, our method integrates a multiscale relationship distillation technique, thus enhancing the transfer of multiscale information between the teacher and student models. Consequently, RTLI tasks can be efficiently accomplished using the well-learned lightweight student model. Comprehensive experiments conducted on a real-world dataset collected via uncrewedaerial vehicles demonstrate the efficacy of our proposed approach in terms of achieving high detection accuracy with reduced computational costs.
Junfei Yi, Jianxu Mao, Hui Zhang 0023, Yurong Chen 0003, Tengfei Liu 0005, Kai Zeng 0010, He Xie, Yaonan Wang 0001
IEEE Trans. Ind. Informatics1
2024 MSRN: Multilevel Spatial Refinement Network for Transmission Line Fastener Defect Detection
abstract
Transmission line (TL) fasteners play the role of connecting components in smart-grid transmission processes with abnormal TL fastener states, seriously impacting the power supply. Therefore, regular detection of TL fasteners is significant. However, the images taken by unmanned aerial vehicle have problems, such as small size and complex background, which bring great challenges to the existing object detection models. Based on this, this article proposes a multilevel spatial refinement network (MSRN), including an attention-guided receptive field enhanced feature pyramid network (ARFE-FPN) and a double refinement head (DR-Head). For the small target problem, ARFE-FPN first uses dilated convolution to expand the receptive field, and uses global average pooling to extract background activation values. Then, it performs channel weighting on TL fastener features under different fields of view. For the problem of complex background, DR-Head first constructs a semantic prediction task to realize the preseparation of foreground and background, and then combines the high-resolution feature map to further highlight the fastener features in the low-resolution feature map. Experiments on the TL fastener dataset show that MSRN has the best detection accuracy, and its AP can reach 92$\%$.
Jianxu Mao, Qingxian Liu, Yaonan Wang 0001, Weixing Peng, Junfei Yi, Ziming Tao, Hui Zhang 0023, Caiping Liu
IEEE Trans. Ind. Informatics5
2022 Review on the COVID-19 pandemic prevention and control system based on AI
Junfei Yi, Hui Zhang 0023, Jianxu Mao, Yurong Chen 0003, Hang Zhong, Yaonan Wang 0001
Eng. Appl. Artif. Intell.1