Linfeng Jiang

dblp:231/1781 · DBLP profile ↗
← Back
23ranked-venue papers
6as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Closed-Loop Hybrid Digital Twin Platform for Connected and Automated Vehicle Validation
Kanglong Quan, Zhebing Xia, Linfeng Jiang, Ziheng Qiao, Dapeng Dong, Dongyao Jia
IWCMC3
2026 Occluded person re-identification via feature enhancement and contextual refinement
Dongcan Liu, Linfeng Jiang
Expert Syst. Appl.2
2026 The Landscape of AI Alignment: A Comprehensive Review of Theories and Methods
abstract
This paper delves into the crucial domain of AI alignment, aiming to provide a comprehensive understanding of its fundamental concepts, motivations, and the alignment cycle framework. With the increasing capabilities of AI systems, the risks associated with misalignment have become more prominent. We first analyze the risks and causes of misalignment, highlighting the importance of alignment research. Then, we introduce the alignment cycle, which consists of forward alignment and backward alignment, and discuss how it serves as a framework for ensuring AI systems adhere to human intentions and values. This research contributes to the growing body of knowledge in AI safety, offering insights for future research and development in the field.
Linfeng Jiang
Int. J. Pattern Recognit. Artif. Intell.3
2026 Mutual Distillation Driven Dual-Space Matching for Visible-Infrared Person Re-Identification
abstract
Visible–infrared person re-identification (VI-ReID) aims to match pedestrian images across heterogeneous modalities. As a key technology in intelligent transportation systems, VI-ReID supports cross-camera tracking, behavior analysis, and security monitoring, particularly in nighttime or low-illumination scenarios. Despite recent advances, existing methods still encounter two critical challenges: (i) semantic misalignment between low-level and high-level features across modalities, and (ii) distribution discrepancies between visible and infrared images. To address these challenges, we propose a novel framework, Mutual Distillation Driven Dual-Space Matching (MDDM), which performs modality alignment in two complementary spaces. For challenge (i), we design a Dual Level Fusion (DLF) module to capture and adaptively fuse hierarchical features, aligning modalities by integrating both low- and high-level semantics across spatial and channel dimensions. In addition, a Modality Invariant Augmentation (MIA) module is developed to extract fine-grained semantic cues and enhance identity discrimination, thereby reinforcing the correlation between visible and infrared modalities and facilitating the learning of robust shared representations. For challenge (ii), we introduce Dual-Space Matching (DSM), which aligns features in both Hilbert and Euclidean spaces. Furthermore, a mutual distillation strategy is incorporated to promote cross-space consistency and alleviate modality-specific discrepancies. Extensive experiments on widely used VI-ReID benchmarks demonstrate the superiority and flexibility of the proposed method, which consistently achieves competitive performance across multiple datasets. Our code is available at https://github.com/lfjiang-cn/MDDM.
Linfeng Jiang, Dongcan Liu, Jinsheng Ji, Ting Bai 0001
IEEE Trans. Circuits Syst. Video Technol.1
2026 Feature-Preserving Fuzzy Clustering for Blurred G-Image Segmentation
abstract
G-images, defined as graph-structured data with complex topologies, have played a significant role in various fields. Current research mainly focuses on denoising and segmenting observed G-images. However, due to sampling or information degradation, they often contain blurred texture information. Therefore, reconstructing and segmenting them accurately is a critical challenge. To address this issue, this work elaborates a feature-preserving FuzzyC-Means (FCM) algorithm by the aid of a wavelet frame transform, which is aimed at segmenting observed G-images with noise and blur. Given the superior performance of tight wavelet frames in feature extraction, this work leverages the proposed algorithm in a wavelet space, thus incorporating both the original and rectified features of G-images to maintain high robustness. To improve segmentation accuracy, it also uses the local binary pattern code to identify and enhance the blurred features. Additionally, to preserve the similarity between any vertex and its adjacent nodes, it introduces a Kullback-Leibler divergence term as a part of FCM’s objective function. Moreover, convergence analysis establishes that the entire sequence of iterates generated by the algorithm is globally convergent to a critical point. Finally, numerical experiments are conducted by comparing the proposed algorithm with other peers on both synthetic and real-world G-images with noise and blur of different levels. Experimental evidence shows that the proposed algorithm exhibits superior effectiveness and robustness compared to existing peers.
Linfeng Jiang, Cong Wang 0033, Jianbin Yang, Xuelong Li 0001
IEEE Trans. Circuits Syst. Video Technol.1
2026 YOLO-MAFF: A Traffic Sign Detection Method Based on Multi-Scale Attention and Adaptive Feature Fusion
abstract
Traffic sign detection is a vital component of intelligent transportation systems. However, in real-world driving scenarios, challenges such as illumination variations, occlusions, and low resolution of small objects can significantly reduce detection accuracy. To overcome these challenges, we propose YOLO-MAFF, a traffic sign detection network that integrates a multi-scale attention mechanism and adaptive feature fusion. Firstly, a backbone network incorporating a multi-scale channel attention mechanism is designed. By integrating multi-scale contextual information with channel attention, efficient feature extraction and representation learning are facilitated. Secondly, a pyramid network based on adaptive feature fusion is developed to learn spatial attention maps. By fusing feature maps at various scales and emphasizing or suppressing region-specific features, the network can alleviate inconsistencies in feature representations. Finally, a small object detection layer is designed to preserve shallow-level detail information in the feature maps, enabling the network to detect small traffic signs. In the experimental section, YOLO-MAFF is evaluated on four datasets, i.e., TT100K, CCTSDB2021, CURE-TSD, and COCO. The experimental results show that YOLO-MAFF exhibits superior performance in traffic sign detection tasks. Compared to the baseline YOLOv8s, our method improves the mAP by 4.8% on TT100k (reaching 90.2%), 2.8% on CCTSDB2021 (reaching 86.0%), 2.7% (reaching 53.7%) on CURE-TSD, and 2.0% (reaching 72.5%) on the COCO dataset. The source code is available athttps://github.com/lfjiang-cn/yolo-maff
Linfeng Jiang, Peidong Zhan, Ting Bai 0001
IEEE Trans. Intell. Transp. Syst.1
2025 TKD: An Efficient Deep Learning Compiler with Cross-Device Knowledge Distillation
abstract
Generating high-performance tensor programs on resource-constrained devices is challenging for current Deep Learning (DL) compilers that use learning-based cost models to predict the performance of tensor programs. Due to the inability of cost models to leverage cross-device information, it is extremely time-consuming to collect data and train a new cost model. To address this problem, this paper proposes TKD, a novel DL compiler that can be efficiently adapted to devices that are resource-constrained. TKD reduces the time budget by over 11x through an adaptive tensor program filter that eliminates redundant and unimportant measurements of tensor programs. Furthermore, by refining the cost model architecture with a multi-head attention module and distilling transferable knowledge from source devices, TKD outperforms state-of-the-art methods in prediction accuracy, compilation time, and compilation quality. We conducted experiments on the edge GPU, NVIDIA Jetson TX2, and the results show that compared to TenSet and TLP, TKD reduces compilation time by 1.58x and 1.16x, while achieving 1.40x and 1.27x speedups of the tensor programs, respectively.
Chaoyao Shen, Linfeng Jiang, Meng Zhang 0010
DATE3
2025 Real-Time Semantic Segmentation for UAV Perspectives on Embedded Platforms
Chaoyao Shen, Yuning Ji, Linfeng Jiang, Meng Zhang 0010
ICIC (1)5
2025 PocketSR: The Super-Resolution Expert in Your Pocket Mobiles
abstract
Real-world image super-resolution (RealSR) aims to enhance the visual quality of in-the-wild images, such as those captured by mobile phones. While existing methods leveraging large generative models demonstrate impressive results, the high computational cost and latency make them impractical for edge deployment. In this paper, we introduce PocketSR, an ultra-lightweight, single-step model that brings generative modeling capabilities to RealSR while maintaining high fidelity. To achieve this, we design LiteED, a highly efficient alternative to the original computationally intensive VAE in SD, reducing parameters by 97.5\% while preserving high-quality encoding and decoding. Additionally, we propose online annealing pruning for the U-Net, which progressively shifts generative priors from heavy modules to lightweight counterparts, ensuring effective knowledge transfer and further optimizing efficiency. To mitigate the loss of prior knowledge during pruning, we incorporate a multi-layer feature distillation loss. Through an in-depth analysis of each design component, we provide valuable insights for future research. PocketSR, with a model size of 146M parameters, processes 4K images in just 0.8 seconds, achieving a remarkable speedup over previous methods. Notably, it delivers performance on par with state-of-the-art single-step and even multi-step RealSR models, making it a highly practical solution for edge-device applications.
Haoze Sun, Linfeng Jiang, Renjing Pei, Zhixin Wang, Haoyu Chen 0003, Fenglong Song, Yujiu Yang 0001, Wenbo Li 0002
NeurIPS2
2025 Multi-scale based cross-modal semantic alignment network for radiology report generation
abstract
The automatic generation of radiology reports draws attention for easing radiologists’ workload and aiding diagnosis. Cross-modal alignment between images and text is critical for high-quality reporting, but cross-modal alignment has not been fully explored at this time due to a lack of annotation. Meanwhile, existing alignment methods utilize single scale image region features for alignment and cannot accommodate the different sizes of anatomical structures in radiology images. To address these problems, we propose a Multi-scale based Cross-modal Semantic Alignment Network (MCSANet). It includes three modules: a multi-scale visual feature extraction module, capturing key image information in windows of different sizes; a cross-modal semantic alignment module, achieving semantic alignment between the two modalities without relying on additional auxiliary information; and a transformer report generator, which generates radiology reports using final features. Experimental results show that MCSANet surpasses other leading approaches on the IU-Xray and MIMIC-CXR datasets.
Dun Lan, Linfeng Jiang
SMC5
2025 A novel approach for cluster detection in trajectory data with low cluster-to-noise density ratio
abstract
A spatial cluster of trajectories refers to objects that follow similar paths, revealing shared movement trends and aiding in anomaly detection. However, detecting clusters in trajectory data becomes challenging when the cluster-to-noise density ratio (CNDR) is low. For example, clusters in free-range sheep movements are easily seen due to their group behaviour, whereas the diversity of human movement introduces significant noise, making clustering difficult. The L-function, widely used for clustering detection in various data types (e.g. point or OD flow data), captures aggregation changes across scales without relying on predefined thresholds, offering potential for low CNDR trajectory data. Thus, we define a trajectory space to derive the Trajectory L (TL)-function for multipoint trajectories. Then we use the second derivative of the TL-function and the local TL-function to identify cluster sizes and extract clusters. Inflection points in the second derivatives enable the detection of subtle changes in aggregation, allowing for precise and sensitive cluster identification. Simulation experiments show that our method outperforms four state-of-the-art approaches in detecting clusters under low CNDR conditions while avoiding parameter dependency. We validated the generality and robustness of our method using both taxi GPS trajectories and mobile phone signalling trajectories. Furthermore, our work lays a rigorous and extensible foundation for the future formulation of spatiotemporal statistical frameworks tailored to trajectory data.
Zidong Fang, Tao Pei, Xiaorui Yan, Linfeng Jiang, Hua Shu 0001, Jie Chen 0077
Int. J. Geogr. Inf. Sci.6
2025 Identification of indoor states of individuals based on mobile phone data
abstract
In urban regions, individuals predominantly spend their time indoors. Accurately identifying these indoor states is essential for many fields, such as public health and urban planning. Existing methods generally rely on sensors placed at specific locations or on volunteers’ mobile phones, limiting their applicability to those locations and a fraction of the population. To overcome these limitations, we propose a novel framework that leverages cellular signaling data—providing extensive spatio-temporal and population-wide coverage—to identify individuals’ indoor states comprehensively. We extract three types of features: interaction between individuals’ mobile phones and cells, individuals’ moving and stationary, and environmental context. Using these features, we apply three machine learning models—CatBoost, Random Forest (RF) and Support Vector Machine (SVM)—along with an interpretable machine learning model, Associative Tree (AT), to identify the indoor states. Evaluation with a ground truth dataset shows that CatBoost outperforms the other models, with an F1 score of 97.21% in quantifying the time individuals spend indoors. To our knowledge, this is the first study to identify the indoor states of individuals using cellular signaling data. We argue that this study can contribute to advancements in areas such as public health and urban planning.
Linfeng Jiang, Tao Pei, Mingbo Wu, Zidong Fang, Meng Gao 0001, Xiaorui Yan, Dasheng Ge
Int. J. Geogr. Inf. Sci.1
2025 D3Impute: Dropout-aware discrimination, distribution-aware modeling, and density-guide imputation for scRNA-seq data
abstract
Single-cell RNA sequencing (scRNA-seq) has revolutionized the study of cellular heterogeneity. A major challenge, however, lies in the prevalence of non-biological zeros-false measurements caused by technical limitations that mask a cell's true transcriptome. This fundamental issue of distinguishing these artifacts from true biological zeros, where a gene is genuinely absent, remains a key hurdle for computational methods, as misclassification can distort biological signals during data recovery. To overcome this, we introduce D3Impute, a discriminative imputation framework built on three key innovations: (1) a distribution-aware normalization step that adapts to dataset-specific characteristics while preserving meaningful biological variation; (2) a dual-network discriminator that uses bulk RNA-seq data as a biological reference to accurately identify non-biological zeros while retaining the true biological zeros; and (3) a density-guided imputation engine that recovers expression values while maintaining local cellular neighborhood structures. Through comprehensive benchmarking against 12 state-of-the-art methods across six diverse datasets, D3Impute demonstrates consistent and significant improvements in essential downstream analyses, including cell clustering, trajectory inference, and differential expression detection. Furthermore, we provide an extensive practical evaluation of D3Impute, demonstrating its robustness across varying data qualities and providing clear guidelines for optimal application. By offering a robust, biologically informed, and user-oriented solution, D3Impute not only enhances scRNA-seq data analysis but also offers a generalizable framework for handling zero-inflated data in computational biology.
Linfeng Jiang, Yuan Zhu 0005
PLoS Comput. Biol.2
2024 Spatiotemporal mobility network of global scientists, 1970-2020
abstract
The mobility of scientists, manifested by movements to new academic institutions, grows with globalization and plays a crucial role in individual careers, institutional productivity, and knowledge dissemination. Current research on scientists’ mobility focuses on aggregated levels such as inter-country mobility, with little attention paid to fine-grained institutional level, leading to a simplified spatial portrayal of the mobility. To fill the gap, we take scientists in geography as examples, and reconstructed their dynamic mobility network among institutions from 1970 to 2020 based on massive literature metadata. Our findings reveal the spatial mobility pattern that is now dominated by North America, Western and Northern Europe, East Asia, and Oceania, with the trend of intensification, multipolarity, and inequality over time. Specifically, the mobility network exhibits clear community structure largely constrained by spatial proximity and national borders. We also uncovered a universal downward mobility pattern embedded in the hierarchical structure. Our quantitative analysis further suggest that mobility is facilitated by multiple realities, including spatial, cultural, and scientific proximity, institutional rankings and national economic levels, cooperation, and visa-free policies, with varying dynamics. These results contribute to spatiotemporal insights into the mechanisms of scientific development in theory, and the basis for talent policymaking in practice.
Tao Pei, Zidong Fang, Mingbo Wu, Xiaorui Yan, Jingyu Jiang, Linfeng Jiang, Jie Chen 0077
Int. J. Geogr. Inf. Sci.9
2024 DSFPAP-Net: Deeper and Stronger Feature Path Aggregation Pyramid Network for Object Detection in Remote Sensing Images
abstract
Rapid detection of small objects in remote sensing (RS) images is crucial for intelligence acquisition, for instance, enemy ship detection. Instead of employing images with high resolution, low-resolution images of the same size typically cover a wider area and thus facilitate efficient object detection. However, accurately detecting small objects in such images remains a challenge due to their limited visual information and the difficulty in distinguishing them from the background. To address this issue, we propose a small object detection method called the Deeper and Stronger Feature Path Aggregation Pyramid Network for low-resolution remote sensing images. First, our approach involves designing aggregation networks with deeper paths and utilizing feature layers closer to the shallow layers to enhance the acquisition of information about small objects. Second, to enhance the network’s focus on small objects, we propose a Resolution-Adjustable 3-D Weighted Attention (RA3-DWA) mechanism. This mechanism enables independent learning of spatial feature information and assigns 3-D weights specifically to small objects, resulting in improved detection accuracy for small objects. Finally, we propose the Fast-EIoU loss function to accelerate the regression of the model boundary. This loss function assigns an acceleration factor to the length loss and width loss, respectively, thereby improving the detection accuracy of small objects. Experiments on Levir-Ship and DOTA demonstrate the effectiveness and efficiency of the proposed method. Compared to the baseline YOLOv5, our method has improved the average detection accuracy of the Levir-Ship dataset by 6.7% (reaching up to 82.6%) and the accuracy of the DOTA dataset by 6.4% (reaching up to 73.7%).
Linfeng Jiang, Yahao Li, Ting Bai 0001
IEEE Geosci. Remote. Sens. Lett.1
2021 Adversarial erasing attention for fine-grained image classification
Jinsheng Ji, Linfeng Jiang, Tao Zhang 0027, Weilin Zhong, Huilin Xiong
Multim. Tools Appl.2
2020 Combining Multilevel Features for Remote Sensing Image Scene Classification With Attention Model
abstract
Remote sensing (RS) image scene classification is a challenging task due to its intraclass variety and the interclass similarity. Recently, many convolutional neural network (CNN)-based methods explore the network to handle this task. However, RS images usually have confusing background in addition to the relevant objects, and features only derived from the whole RS images cannot achieve satisfying results. To solve the problem, this letter proposed a method of utilizing the attention network to localize multiscale discriminative regions of the RS scene images and combining features learned from the localized regions by a classification network. Specifically, the classification network is composed of three subnetworks, which are trained by certain scaled regions separately. To learn more discriminative feature representations, feature fusion module is introduced to fuse the features of the three subnetworks in a more effective way. Experiments conducted on the AID and NWPU-RESISC45 data sets evaluate the effectiveness of the proposed method.
Jinsheng Ji, Tao Zhang 0027, Linfeng Jiang, Weilin Zhong, Huilin Xiong
IEEE Geosci. Remote. Sens. Lett.3
2020 Exploiting context based on CNN and coding representations for pedestrian co-detection
Linfeng Jiang, Jinsheng Ji, Weilin Zhong, Tao Zhang 0027, Huilin Xiong
Multim. Tools Appl.1
2020 A part-based attention network for person re-identification
Weilin Zhong, Linfeng Jiang, Tao Zhang 0027, Jinsheng Ji, Huilin Xiong
Multim. Tools Appl.2
2019 Aircraft Detection from Remote Sensing Image Based on A Weakly Supervised Attention Model
abstract
Aircraft detection from high resolution remote sensing image is a challenging task due to the lack of annotation information, large-scale image size, and sparse distribution of aircraft. Recently, some convolutional neural network(CNN) based methods explore the attention based weakly supervised way to localize the aircraft without manual annotation information. However, the detection results are not satisfied with high false detection ratio. In this paper, a method of utilizing weakly supervised attention model to localize the multi-scale aircrafts is presented, in which the attention model is carried out in a weakly supervised way. Compared with other CNN based method, the proposed attention model can obtain more accurate attention map and localize the aircrafts more precisely. The experimental results on two challenging datasets demonstrate that the proposed method achieves higher detection accuracy and lower false detection ratio than other methods.
Jinsheng Ji, Tao Zhang 0027, Zhen Yang 0012, Linfeng Jiang, Weilin Zhong, Huilin Xiong
IGARSS4
2019 Combining multilevel feature extraction and multi-loss learning for person re-identification
Weilin Zhong, Linfeng Jiang, Tao Zhang 0027, Jinsheng Ji, Huilin Xiong
Neurocomputing2
2019 Discriminative representation learning for person re-identification via multi-loss training
Weilin Zhong, Tao Zhang 0027, Linfeng Jiang, Jinsheng Ji, Zenghui Zhang, Huilin Xiong
J. Vis. Commun. Image Represent.3
2018 A Multi-part Convolutional Attention Network for Fine-Grained Image Recognition
abstract
The goal of fine-grained image recognition is to recognize hundreds of sub-categories affiliating to the same basic-level category (e.g., bird species). It is a highly challenging task due to the large intra-class variance and small inter-class variance. Existing approaches deal with the subtle difference among object classes via learning and localizing discriminative parts. However, most of the part localization methods follow a step-to-step manner that first localizes larger parts and then generates smaller parts from the larger ones, which is not efficient. In this paper, we present a Multi-part Convolutional Attention Network (M-CAN), which simultaneously focuses on the discriminative image parts at multiple scales. In specific, a convolutional attention based part localization network is presented to localize multi-scale parts from different layers of the deep Convolutional Neural Networks (CNN). Importantly, our part localization network requires no part annotations but only the image labels, which avoids the heavy labor of complex part labeling. We conduct comprehensive experiments and the experimental results show that, our method outperforms the state-of-the-art approaches on three challenging fine-grained datasets, including CUB-Birds, Stanford-Dogs and Stanford-Cars.
Weilin Zhong, Linfeng Jiang, Tao Zhang 0027, Jinsheng Ji, Huilin Xiong
ICPR2