VLDB 2026 Research / reviewers in the wild / expert
Junfei Yi
dblp:274/4774
· DBLP profile ↗
17ranked-venue papers
5as first author
17since 2021 · last 2026
0000-0003-1737-0078ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Distilling Object Detectors via Monte Carlo DropoutabstractKnowledge distillation (KD) has become a fundamental technique for model compression in object detection tasks. The data noise and training randomness may cause the knowledge of the teacher model to be unreliable, referred to as knowledge uncertainty. Existing methods neglect this uncertainty, potentially hindering the student's capacity to capture and understand latent "dark knowledge". In this work, we introduce a novel strategy that explicitly incorporates knowledge uncertainty, named Uncertainty-Driven Knowledge Extraction and Transfer (UET). Given the unknown, high-dimensional nature of the knowledge distribution, we employ Monte Carlo dropout to effectively estimate the teacher's uncertainty. Leveraging information theory, we combine uncertainty with deterministic knowledge, enabling the student to benefit from both precision and diversity. UET is a plug-and-play method that integrates seamlessly with existing distillation techniques. We validate our approach through comprehensive experiments across various distillation strategies, detectors, and backbones. Specifically, UET achieves state-of-the-art results, with a ResNet50-based GFL detector obtaining 44.1% mAP on the COCO dataset-surpassing baseline performance by 3.9%. Junfei Yi, Hui Zhang 0023, Jianxu Mao, Tengfei Liu 0005, Mingjie Li 0006, Sihao Lin, Hanyu Gu, Zhihui Li 0001, Xiaojun Chang, Yaonan Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | MGLD-TLNet: Multigeometric and Long-Distance Representation Network for Transmission Line InspectionabstractEffective transmission line (TL) inspection in complex corridor environments is essential for ensuring reliable power delivery. This work presents a 3D-based perception method for this task. The proposed method is designed by considering two key characteristics of TL inspection. First, the point cloud data are sparse and class distributions are highly imbalanced, which weakens the signals from thin conductors and tower components. To address this issue, we model long-range spatial relations along the corridor to mitigate data sparsity and imbalance. Second, strong structural correlations exist between conductors and towers, which can be leveraged to improve perception performance. To exploit this property, we construct a unified 3-D representation that jointly models towers, conductors, and vegetation, while fusing Cartesian and polar geometries through geometry-aware alignment. Experiments on real-world corridor datasets demonstrate that the proposed method, termed multigeometric and long-distance TL perception Network (MGLD-TLNet), consistently improves stability and accuracy under conditions of sparsity, occlusion, and complex environmental interactions. Hui Zhang 0023, Kaining Zhang, Baheti Biekezat, Hang Zhong, Junfei Yi, Jianxu Mao, Yaonan Wang 0001 |
IEEE Trans. Cybern. | 6 |
| 2026 | Investigating a Unified 3-D Object Detection Method for Different Multibeam LiDAR
Ziming Tao, Jianxu Mao, Yaonan Wang 0001, Caiping Liu, Junfei Yi, Zhenyu He 0015, Xiaojun Chang, Hui Zhang 0023 |
IEEE Trans. Ind. Informatics | 5 |
| 2026 | You Can Only Tune Normalization: A Simple and Effective Approach to Parameter-Efficient Fine-TuningabstractTo tackle the issue of excessive parameter volumes during fine-tuning of large-scale pre-trained models with full parameters, Parameter-Efficient Fine-Tuning (PEFT) methods have been introduced. The core concept involves freezing the backbone network of the model and updating only a small subset of parameters. This strategy not only decreases the number of parameters needed for training but also delivers performance comparable to Full-Tuning, even surpassing it on certain datasets. However, most popular PEFT methods introduce extra parameters or modules for fine-tuning, which come with inherent limitations. In response, we propose a straightforward and efficient PEFT method called You Can Only Tune Normalization (YONO). YONO focuses solely on tuning the normalization layer and the final classification layer of the model. This method avoids adding extra modules, making it easily applicable to any model without causing inference delays. We extensively tested YONO on 28 benchmark datasets, and the results indicate that it requires significantly fewer parameters compared to other advanced PEFT methods. Additionally, we validated YONO’s efficiency and generalizability across various vision models. Finally, we further explore the essence of PEFT methods, whether they learn new knowledge or expose the capabilities that a model has already learned. Our findings suggest that YONO is more sensitive to improvements in dataset quality, making it a promising candidate for future scaling to larger models. Lingyun Huang, Jianxu Mao, Junfei Yi, Ziming Tao, Ziyang Peng, Wei He 0001, Rui Liu 0028, Yaonan Wang 0001 |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2026 | Image-Quality-Guided Consistency-Alignment Network for Fault DiagnosisabstractSignal-to-image transformation has been widely used in mechanical fault diagnosis to provide unified visual representations for downstream diagnostic models. However, the resulting diagnostic images can exhibit substantial quality variations caused by measurement noise, sensor degradation, and partial signal loss. Most existing methods either discard low-quality samples or implicitly assume equal reliability across samples, leading to information loss or degraded robustness. This paper proposes an image-quality-guided consistency-alignment network that explicitly estimates sample quality and leverages degraded data during training. First, vibration signals are decomposed into sub-bands, and energy-ratio criteria are used to select informative components for reconstruction. The reconstructed signals are subsequently fused into fixed-layout two-dimensional images via a sector-allocation strategy. Next, a diagnosis-aware quality assessment module assigns a quality score to each image to guide training. Low-quality samples are further regularized via cross-quality feature alignment using a quality-weighted supervised pairwise loss, encouraging them to align with high-quality counterparts from the same class. Finally, a Swin Transformer backbone performs classification. Experiments on the proprietary RGFD dataset and the public SEUB dataset demonstrate consistent gains under controlled mixed-quality settings, with accuracies of up to 99.4% and 98.7%, respectively, while additional evaluations under reproducible degradations further confirm the robustness of the proposed quality-aware learning mechanism. Zhuowei Li 0011, Jianxu Mao, Yaonan Wang 0001, Junfei Yi, Caiping Liu, Hui Zhang 0023 |
IEEE Trans. Reliab. | 5 |
| 2025 | HC-LLM: Historical-Constrained Large Language Models for Radiology Report GenerationabstractRadiology report generation (RRG) models typically focus on individual exams, often overlooking the integration of historical visual or textual data, which is crucial for patient follow-ups. Traditional methods usually struggle with long sequence dependencies when incorporating historical information, but large language models (LLMs) excel at in-context learning, making them well-suited for analyzing longitudinal medical data. In light of this, we propose a novel Historical-Constrained Large Language Models (HC-LLM) framework for RRG, empowering LLMs with longitudinal report generation capabilities by constraining the consistency and differences between longitudinal images and their corresponding reports. Specifically, our approach extracts both time-shared and time-specific features from longitudinal chest X-rays and diagnostic reports to capture disease progression. Then, we ensure consistent representation by applying intra-modality similarity constraints and aligning various features across modalities with multimodal contrastive and structural constraints. These combined constraints effectively guide the LLMs in generating diagnostic reports that accurately reflect the progression of the disease, achieving state-of-the-art results on the Longitudinal-MIMIC dataset. Notably, our approach performs well even without historical data during testing and can be easily adapted to other multimodal large models, enhancing its versatility. Tengfei Liu 0005, Jiapu Wang, Yongli Hu, Mingjie Li 0006, Junfei Yi, Xiaojun Chang, Junbin Gao |
AAAI | 5 |
| 2025 | CVPT: Cross Visual Prompt Tuning
Lingyun Huang, Jianxu Mao, Junfei Yi, Ziming Tao, Yaonan Wang 0001 |
ICCV | 3 |
| 2025 | Wavelet-Based Distillation with Structured Frequency Alignment
Pengyu Lu, Junfei Yi, Jianxu Mao, Junlong Yu, Shuohao Xiao, Zhenyu He 0015, Yaonan Wang 0001 |
ICIG (2) | 2 |
| 2025 | FFTA-Net: A Frequency-Domain Fusion and Temporal Alignment Network for Transmission Line Defect Detection
Jianxu Mao, Yaonan Wang 0001, Junlong Yu, Junfei Yi, Zhenyu He 0015, Ziming Tao, Hui Zhang 0023 |
ICIG (2) | 5 |
| 2025 | Foreground-Aware Enhancement-Based Multimodal 3D Object DetectionabstractLiDAR is one of the most widely used 3D detection sensors in applications such as autonomous driving and unmanned inspection. However, when uniformly sampling the entire scene to generate point cloud data, the number of foreground points reflected by target objects is often significantly lower than that of the background points, which have a larger coverage area. This imbalance poses considerable challenges to the performance of object detection models, especially in detecting small or distant objects. To overcome this challenge, this paper presents a Foreground-Aware Enhancement-based Multimodal 3D Object Detection Method (PFA), which effectively mitigates the low detection accuracy of small and distant objects caused by insufficient foreground points. The proposed method incorporates a Foreground-Aware Enhancement Module (FAEM) and a Region-Focused Attention Module (RFAM). The FAEM module enhances the model’s focus on foreground regions, while the RFAM module strengthens multimodal fused features. Together, these components significantly improve the detection accuracy of small and distant objects. Experimental results on the KITTI dataset demonstrate that the proposed method achieves 3D detection accuracies of 84.65%, 59.56%, and 71.48% for cars, pedestrians, and cyclists, respectively, under the hard evaluation level. Furthermore, the model also shows significant advantages on the KITTI public test set and validation set for both easy and moderate samples, fully validating its effectiveness and generalizability in enhancing multimodal 3D object detection accuracy. Ziyang Peng, Wei He 0001, Jianxu Mao, Ziming Tao, Junfei Yi, Yaonan Wang 0001 |
IJCNN | 5 |
| 2025 | LDFCDet: Boosting 3D Object Detectors With Low-High Level Feature Crosses Using Laplace DistributionabstractHighly accurate 3D object detection is critical for autonomous driving and robotic sensing system. However, some objects with few foreground points significantly affect the accuracy of 3D object detection. As the network depth increases, the low-level features of these objects are gradually lost, especially for the hard object. Due to this issue, current LiDAR-only based and multimodal methods often misclassify background as foreground. Therefore, how to leverage the low-level feature that contain information about these objects in the high layer of the network becomes the key to optimizing the issue. In this paper, we propose LDFCDet, a framework boosting 3D object detectors with low-high level feature crosses using Laplace distribution(LD). In our proposed method, we design a low-high level feature crosses module(LHFCM) to embed low-level feature into high-level feature in the deeper layer of the network, and use Laplace distribution to obtain a new low-high level feature that includes information about these objects with few foreground points. In addition, we propose a res-gated feature aggregation module(RGFAM) to fuse the mutli-scale features. Our approach is well-suited for both LiDAR-based and multimodal methods.We evaluate the LDFCDet on the widely used KITTI dataset, and our method outperforms almost current 3D object detection methods on the challenging KITTI test set. Moreover, we conducted comparative experiments on the ONCE dataset, and the results further demonstrate the effectiveness and superiority of our method. Zhenyu He 0015, Jianxu Mao, Yaonan Wang 0001, Junlong Yu, Ziming Tao, Junfei Yi, Hui Zhang 0023, Shaoyuan Wang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Tackling Real-World Complexity: Hierarchical Modeling and Dynamic Prompting for Multimodal Long Document ClassificationabstractWith the rapid growth of internet content, multimodal long document data has become increasingly prominent, drawing significant attention from researchers. However, most existing methods primarily focus on scenarios where all modalities are present, often overlooking more challenging and realistic cases involving missing image modality. To address this limitation, we propose a robust multimodal long document classification (MLDC) framework that integrates hierarchical modeling and dynamic prompting to handle complex multimodal long document data. Our approach begins by leveraging hierarchical modeling combined with an Adaptive Correlation Multimodal Transformer (ACMT) to effectively capture relationships between text and images at both section and sentence levels. We also introduce a Dynamic Prompt Generation (DPG) module at both levels to enhance the model’s robustness in handling missing image data. By evaluating sample uncertainty, the DPG module dynamically adjusts both the number of prompts and the prompts themselves, allowing the model to better adapt to the varying needs of different samples. Finally, a Hierarchical Heterogeneous Graph (HHG) is introduced to enhance feature interactions across levels, further improving the coherence and accuracy of the model. Extensive experiments on four multi-modal long document datasets demonstrate that our model shows superior performance compared to existing state-of-the-art MLDC classification methods in various conditions. Tengfei Liu 0005, Yongli Hu, Mingjie Li 0006, Junfei Yi, Xiaojun Chang, Junbin Gao |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | FMSD: Focal Multi-Scale Shape-Feature Distillation Network for Small Fasteners Detection in Electric Power SceneabstractIn the electric power scene, fasteners play a pivotal role in securing and connecting electrical equipment, with small fastener detection (SFD) being crucial for ensuring operational stability. Despite the replacement of manual inspection methods by non-destructive techniques employing deep learning, these approaches often demand substantial computational resources and involve numerous parameters. While knowledge distillation (KD) can be a viable solution, existing KD methods may often fail to achieve satisfactory performance when dealing with small object presentation and little inter-class variability in SFD tasks. To alleviate this, we propose a Focal Multi-scale Shape-feature Distillation Network (FMSD) to achieve efficient and precise fastener detection in electric power scenarios. Specifically, we propose a novel Multi-Scale Shape-Aware Feature Aggregation module (MSFA) to augment the network's perception of object shape and scale during the KD process. Additionally, we propose a Contour-Guided Distillation (CGD) module to optimize the transfer of the extracted shape-sensitive knowledge between the teacher and student models. Through a series of experiments compared with existing state-of-the-art (SOTA) methods, our method demonstrates superior performance over existing SOTA techniques, both efficiently and effectively. Furthermore, validation on publicly available power scene datasets confirms the generalizability and adaptability of our proposed FMSD across various settings. Junfei Yi, Jianxu Mao, Hui Zhang 0023, Mingjie Li 0006, Kai Zeng 0010, Mingtao Feng, Xiaojun Chang, Yaonan Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Toward Efficient Power Scene Detection via Topology-Preserved Knowledge DistillationabstractThe power industry relies on efficient inspection systems to ensure stability and safety. While deep learning has advanced automated inspection, its reliance on custom modules for specific tasks can impact efficiency. Knowledge distillation (KD) offers a balanced solution, but the complex textures and structures of power equipment challenge conventional KD methods, which often fail to capture essential local semantic and topological relationships. To address this, we proposeTopNet, a novel topology-preserved KD framework for power scene detection tasks. Specifically, we model the teacher’s knowledge as a graph, where nodes encode local fine-grained features and edges capture global topological relationships. Based on this, we introduce node feature distillation and edge feature distillation to transfer local–global structural knowledge, which can enhance the student’s ability to perceive objects. Furthermore, we also introduce aggregated feature distillation to incorporate and transfer contextual semantic knowledge. Comprehensive experiments are conducted on two different benchmark datasets to demonstrate that TopNet achieves state-of-the-art detection performance with high efficiency, offering a robust solution for automated power equipment inspection. Junfei Yi, Tengfei Liu 0005, Jianxu Mao, Yaonan Wang 0001, Hui Zhang 0023, He Xie, Hang Zhong, Xiaojun Chang |
IEEE Trans. Ind. Informatics | 1 |
| 2025 | Balancing Accuracy and Efficiency With a Multiscale Uncertainty-Aware Knowledge-Based Network for Transmission Line InspectionabstractReal-world transmission line inspections (RTLIs) ensure power stability and safety. Deep learning (DL) models have become prevalent approaches for performing RTLI tasks. However, the high computational demands and substantial parameter requirements of DL models limit their real-world applicability. This article introduces a novel approach, a multiscale uncertainty-aware knowledge-based network, which is designed to balance the accuracy and efficiency in RTLI tasks. Specifically, we propose an uncertainty-aware knowledge distillation method that incorporates pixel-level uncertainty into the knowledge transfer process, mitigating the impact of noisy knowledge derived from extra background information contained in ground truths. In addition, our method integrates a multiscale relationship distillation technique, thus enhancing the transfer of multiscale information between the teacher and student models. Consequently, RTLI tasks can be efficiently accomplished using the well-learned lightweight student model. Comprehensive experiments conducted on a real-world dataset collected via uncrewedaerial vehicles demonstrate the efficacy of our proposed approach in terms of achieving high detection accuracy with reduced computational costs. Junfei Yi, Jianxu Mao, Hui Zhang 0023, Yurong Chen 0003, Tengfei Liu 0005, Kai Zeng 0010, He Xie, Yaonan Wang 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2024 | MSRN: Multilevel Spatial Refinement Network for Transmission Line Fastener Defect DetectionabstractTransmission line (TL) fasteners play the role of connecting components in smart-grid transmission processes with abnormal TL fastener states, seriously impacting the power supply. Therefore, regular detection of TL fasteners is significant. However, the images taken by unmanned aerial vehicle have problems, such as small size and complex background, which bring great challenges to the existing object detection models. Based on this, this article proposes a multilevel spatial refinement network (MSRN), including an attention-guided receptive field enhanced feature pyramid network (ARFE-FPN) and a double refinement head (DR-Head). For the small target problem, ARFE-FPN first uses dilated convolution to expand the receptive field, and uses global average pooling to extract background activation values. Then, it performs channel weighting on TL fastener features under different fields of view. For the problem of complex background, DR-Head first constructs a semantic prediction task to realize the preseparation of foreground and background, and then combines the high-resolution feature map to further highlight the fastener features in the low-resolution feature map. Experiments on the TL fastener dataset show that MSRN has the best detection accuracy, and its AP can reach 92$\%$. Jianxu Mao, Qingxian Liu, Yaonan Wang 0001, Weixing Peng, Junfei Yi, Ziming Tao, Hui Zhang 0023, Caiping Liu |
IEEE Trans. Ind. Informatics | 5 |
| 2022 | Review on the COVID-19 pandemic prevention and control system based on AI
Junfei Yi, Hui Zhang 0023, Jianxu Mao, Yurong Chen 0003, Hang Zhong, Yaonan Wang 0001 |
Eng. Appl. Artif. Intell. | 1 |