Wei Li 0075

dblp:64/6025-75 · DBLP profile ↗
← Back
20ranked-venue papers
2as first author
18since 2021 · last 2026
0000-0002-3786-4959ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 6 · 5 since 2021Computer networks · 4 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multi-Task Learning for Airport Surface Surveillance: A Review
abstract
ABSTRACT The rapid growth of air transportation has surpassed the capabilities of traditional airport surveillance methods, such as visual observation and auxiliary equipment (e.g., ADS‐B, MLAT, radar), which struggle to provide all‐area, all‐weather situation awareness. Vision‐based deep learning methods, being cost‐effective and scalable, present promising alternatives but often fall short in delivering comprehensive awareness. Multi‐task learning (MTL) addresses these gaps by enabling models to simultaneously learn multiple related tasks, improving overall perception and decision‐making. Thus, this review reviews MTL systems for airport surface surveillance, categorising tasks into scene perception for intensive estimation and monitoring for non‐intensive estimation. This review identifies three key challenges: (1) efficient information sharing across tasks, (2) balancing multiple tasks, and (3) enhancing model training efficiency. The review examines these challenges through three lenses: neural network architecture, loss function optimization, and learning paradigms, proposing optimization strategies for each. As the first review to focus on MTL in airport surveillance, this review provides valuable insights into model design, task balancing, and training strategies, offering guidance for the future development of intelligent airport monitoring systems.
Daoyong Fu, Xiangtong Wang, Fangrui Wu, Songchen Han, Binbin Liang, Wei Li 0075
Expert Syst. J. Knowl. Eng.6
2026 DWSF-Net: A Dynamic Wavelet-Based Spatial-Frequency Fusion Network for Multispectral Object Detection
abstract
Multispectral object detection aims to identify tar gets under diverse illumination conditions by leveraging complementary information from multiple spectral modalities. A major challenge in this field lies in effectively fusing multispectral features while accounting for both spatial and frequency domain characteristics. Existing methods primarily focus on spatial fusion, often neglecting critical frequency domain cues and treating all spectral channels equally, despite their distinct properties. In particular, RGB images capture high-frequency texture and color, whereas infrared (IR) images focus on low frequency thermal signatures—rendering conventional spatial only fusion suboptimal. To address these challenges, we propose a novel Dynamic Wavelet-based Spatial-Frequency Fusion Net work (DWSF-Net) that integrates both spatial and frequency information for enhanced multispectral representation. DWSF Net introduces a learnable wavelet encoder to adaptively extract frequency-aware features, a wavelet modulation fusion module to selectively combine informative sub-bands across spectra, and a frequency-domain sub-band fusion scheme with adaptive weight learning to refine cross-spectral integration. Finally, modulated spatial features and adaptively fused frequency components are aggregated to form the final representation. Extensive experiments conducted on three public datasets demonstrate that the proposed DWSF-Net achieves state-of-the-art performance, highlighting its effectiveness and potential for improving the accuracy of multispectral object detection.
Fan Yang 0104, Wei Li 0075, Lei Li 0020, Jianwei Zhang 0013
IEEE Trans. Multim.2
2025 ISPDiffuser: Learning RAW-to-sRGB Mappings with Texture-Aware Diffusion Models and Histogram-Guided Color Consistency
abstract
RAW-to-sRGB mapping, or the simulation of the traditional camera image signal processor (ISP), aims to generate DSLR-quality sRGB images from raw data captured by smartphone sensors. Despite achieving comparable results to sophisticated handcrafted camera ISP solutions, existing learning-based methods still struggle with detail disparity and color distortion. In this paper, we present ISPDiffuser, a diffusion-based decoupled framework that separates the RAW-to-sRGB mapping into detail reconstruction in grayscale space and color consistency mapping from grayscale to sRGB. Specifically, we propose a texture-aware diffusion model that leverages the generative ability of diffusion models to focus on local detail recovery, in which a texture enrichment loss is further proposed to prompt the diffusion model to generate more intricate texture details. Subsequently, we introduce a histogram-guided color consistency module that utilizes color histogram as guidance to learn precise color information for grayscale to sRGB color consistency mapping, with a color consistency loss designed to constrain the learned color information. Extensive experimental results show that the proposed ISPDiffuser outperforms state-of-the-art competitors both quantitatively and visually.
Yang Ren 0001, Hai Jiang 0006, Menglong Yang, Wei Li 0075, Shuaicheng Liu
AAAI4
2025 Learning Arbitrary-Scale RAW Image Downscaling with Wavelet-based Recurrent Reconstruction
abstract
Image downscaling is critical for efficient storage and transmission of high-resolution (HR) images. Existing learning-based methods focus on performing downscaling within the sRGB domain, which typically suffers from blurred details and unexpected artifacts. RAW images, with their unprocessed photonic information, offer greater flexibility but lack specialized downscaling frameworks. In this paper, we propose a wavelet-based recurrent reconstruction framework that leverages the information lossless attribute of wavelet transformation to fulfill the arbitrary-scale RAW image downscaling in a coarse-to-fine manner, in which the Low-Frequency Arbitrary-Scale Downscaling Module (LASDM) and the High-Frequency Prediction Module (HFPM) are proposed to preserve structural and textural integrity of the reconstructed low-resolution (LR) RAW images, alongside an energy-maximization loss to align high-frequency energy between HR and LR domain. Furthermore, we introduce the Realistic Non-Integer RAW Downscaling (Real-NIRD) dataset, featuring a non-integer downscaling factor of 1.3×, and incorporate it with publicly available datasets with integer factors (2×, 3×, 4×) for comprehensive benchmarking arbitrary-scale image downscaling purposes. Extensive experiments demonstrate that our method outperforms existing state-of-the-art competitors both quantitatively and visually. The code and dataset will be released at https://github.com/RenYangSCU/ASRD.
Yang Ren 0001, Hai Jiang 0006, Wei Li 0075, Menglong Yang, Heng Zhang 0042, Zehua Sheng, Qingsheng Ye, Shuaicheng Liu
ACM Multimedia3
2025 Multidimensional Fusion Network for Multispectral Object Detection
abstract
Multispectral object detection has attracted increasing attention recently due to its superior detection capacity under various illumination conditions. The key challenge lies in the effective aggregation of multi-spectral features to derive highly discriminative representations. To address this challenge, we propose a novel Multidimensional Fusion Network (MMFN) to explore multi-modal information from local, global, and channel perspectives. Specifically, at the local level, local features of different modalities and their inter-relationships are captured by a window-shifted fusion. As a complement to the local information, we designed a global interaction module that facilitates the fusion of holistic, high-level semantic information spanning the entire image. We distillate the channel dependencies and complementarities between different modalities through cross-channel learning and generate the final fused representation. Comprehensive experiments conducted on three publicly available datasets provide compelling evidence validating the superiority of the proposed methodology. The results exhibit notable performance gains over state-of-the-art multispectral object detectors. Our code will be released.
Fan Yang 0104, Binbin Liang, Wei Li 0075, Jianwei Zhang 0013
IEEE Trans. Circuits Syst. Video Technol.3
2024 Multi-Modal Disordered Representation Learning Network for Description-Based Person Search
abstract
Description-based person search aims to retrieve images of the target identity via textual descriptions. One of the challenges for this task is to extract discriminative representation from images and descriptions. Most existing methods apply the part-based split method or external models to explore the fine-grained details of local features, which ignore the global relationship between partial information and cause network instability. To overcome these issues, we propose a Multi-modal Disordered Representation Learning Network (MDRL) for description-based person search to fully extract the visual and textual representations. Specifically, we design a Cross-modality Global Feature Learning Architecture to learn the global features from the two modalities and meet the demand of the task. Based on our global network, we introduce a Disorder Local Learning Module to explore local features by a disordered reorganization strategy from both visual and textual aspects and enhance the robustness of the whole network. Besides, we introduce a Cross-modality Interaction Module to guide the two streams to extract visual or textual representations considering the correlation between modalities. Extensive experiments are conducted on two public datasets, and the results show that our method outperforms the state-of-the-art methods on CUHK-PEDES and ICFG-PEDES datasets and achieves superior performance.
Fan Yang 0104, Wei Li 0075, Menglong Yang, Binbin Liang
AAAI2
2024 Space Networking Kit: A Novel Simulation Platform for Emerging LEO Mega-constellations
abstract
Futuristic Space Networks (SN) present unprece-dented prospects for ubiquitous, low-latency Internet services. Yet, these networks also encounter unique challenges arising from the dynamic nature of satellites on a global scale. To comprehensively address emerging issues in SNs, researchers require the capability to conduct a diverse array of experiments. However, existing experimental approaches either implement visualization functionality but lack network functionality (e.g., the space simulator), or implement network functionality but lack visualization (e.g., the networking simulator). In this paper, we present SNK, a novel simulation platform with visualization and networking capabilities for evaluating the space network performance of global Internet services. SNK offers real-time communication visualization and supports the simulation of routing between edge node of network. The platform enables the evaluation of routing and network performance metrics such as latency, stretch, network capacity, and throughput under different network structures and density. The effectiveness of SNK is demonstrated through various simulation cases, including the routing between fixed edge stations or mobile edge stations and analysis of snace network structures.
Xiangtong Wang, Xiaodong Han, Menglong Yang, Songchen Han, Wei Li 0075
ICC5
2023 Enabling High-Connectivity LEO Satellite Networks Via Encountering Inter-Satellite Links
abstract
The use of a mega-constellation comprised of thousands of Low Earth Orbit (LEO) satellites for global internet service has garnered significant attention. In the network layer of this system, geographical routing has been found to outperform centralized strategies due to the lower complexity and overhead. However, geographical routing can still result in “dead-ends” due to network gaps. To address this issue, we propose the use of encountering inter-satellite links (eISLs) to improve network connectivity and routing reachability. We further present a system model and analysis of eISLs, as well as our Dynamic eISLs Configuration (DeC) algorithm for establishing eISLs between encountering satellites. Our experimental results demonstrate that our proposed DeC under eISLs enabling in satellite networks can significantly reduce propagation latency by 22% and path stretch by 15% in centralized routing algorithms. Moreover, in geographical routing, DeC can effectively improve the reachable ratio from 55% to 100% while maintaining a 28% increase in throughput, outperforming schemes without eISLs. Our proposed eISL-enabled satellite network architecture shows promising results in improving routing efficiency and connectivity in LEO satellite systems.
Xiangtong Wang, Wei Li 0075, Songchen Han, Menglong Yang, Zhiyun Jiang
GLOBECOM2
2023 Dynamic Graph Neural Networks-Based Alert Link Prediction for Online Service Systems
abstract
A fault in large online service systems often triggers numerous alerts due to the complex business and component dependencies among services, which is known as “alert storm”. In a short time, an online service system may generate a huge amount of alert data. This poses a challenge for on-call engineers to identify alerts that are associated with a system failure for root cause analysis. In this paper, we propose DyAlert, a dynamic graph neural networks-based approach for linking alerts that might be triggered by a same fault to reduce the burden of on-call engineers in the fault analysis. Our insight is that alerts are often triggered by alert propagation when a system failure occurs, e.g., alert$a$would lead to the occurrence of alert$b$. Whether two alerts should be linked depends on if one alert is triggered by the propagation of the other. Leveraging this insight, we design a dynamic graph (namely Alert-Metric Dynamic Graph) that describes the propagation process of alerts. Based on the dynamic graph, we train a neural networks-based model to predict alert links. We evaluate DyAlert with real-world data collected from an online service system running 85 business units and about 30,000 different services in a large enterprise. The results show that DyAlert is effective in predicting alert links and it outperforms the state-of-the-art approaches with an average increase of 0.259 in F1-score.
Chenxi Zhang 0003, Dingyu Yang, Xin Peng 0001, Jiayu Ou, Zheshun Wu, Xiaojun Qu, Wei Li 0075
ASE10
2023 Spatiotemporal Interaction Transformer Network for Video-Based Person Reidentification in Internet of Things
abstract
Video-based person reidentification, which is a significant application in the Internet of Things, aims to identify the same person in different video sequences across nonoverlapping cameras. Existing methods usually utilize temporal cues to enhance spatial features. However, these methods learn the temporal and spatial information separately, which breaks the relationship between them and ignores the positive role of temporal information for learning frame-level spatial representation in the process of spatial representation learning. In this article, we propose a novel spatiotemporal interaction transformer network (SITN) to solve this problem. To model the temporal information and the relationship between frames, we introduce a temporal interaction module (TIM) to interact between frame information. Meanwhile, we combine TIM with spatial transformer encoder to explore the positive role of temporal information in the learning procedure of the frame-level spatial feature. Moreover, we propose a transformer local learning scheme by reconstructing the 2-D spatial information of the frame patch sequences and extracting local features in a striped manner to strengthen the discriminative capability of our model. Extensive experiments are conducted on four public benchmarks. The results show that our model is superior compared with state-of-the-art methods.
Fan Yang 0104, Wei Li 0075, Binbin Liang, Jianwei Zhang 0013
IEEE Internet Things J.2
2023 Spatial-temporal alignment of time series with different sampling rates based on cellular multi-objective whale optimization
Binbin Liang, Songchen Han, Wei Li 0075, Guoxin Huang, Ruliang He
Inf. Process. Manag.3
2023 Similarity Measure of Time Series With Different Sampling Frequencies Based on Context Density Consistency and Dynamic Time Warping
abstract
Similarity measure of time series with different sampling frequencies is vitally important for many signal processing applications. Dynamic Time Warping (DTW) is one of the most popular methods for similarity measure of time series. However, conventional DTW algorithms have limitations when dealing with time series of different sampling frequencies due to context density inconsistency such as different internal change frequencies of derivatives, shapes, events and distances in local neighborhoods. In light of this, we propose a novel Context Density Consistency Dynamic Time Warping (CDC-DTW) algorithm. It firstly designs local context windows adaptive to the lengths of time series. Then it proposed a local spatial-temporal context density consistency technique by down-sampling and interpolation compensating the high-frequency time series following the context density of low-frequency time series. Besides, a normalized Hamming window weighting function is embedded into the local contexts to create robust weighted cost measure. Extensive experimental results on 128 gold-standard UCR datasets showed that CDC-DTW increased the similarity measure accuracy by 70.53% in average comparing with other 6 classic and state-of-the-art DTW baseline algorithms.
Wei Li 0075, Ruliang He, Binbin Liang, Fan Yang 0104, Songchen Han
IEEE Signal Process. Lett.1
2023 The 6D Pose Estimation of the Aircraft Using Geometric Property
abstract
The take-off and landing activities of aircraft must operate daily in a safe and orderly manner to provide a safe and efficient transition of passengers and goods. The 6D pose estimation of the aircraft is the key to guaranteeing the safe take-off and landing of the aircraft. The vision-based 6D pose estimation method, an important method when GPS and gyroscope are unavailable, faces the problem of poor estimation accuracy due to large scenes and large depths. An end-to-end 6D pose estimation method of the aircraft is proposed to solve this problem. Firstly, this paper combines the rigid structure characteristic of the aircraft with the direction property of arrows to build an aircraft 3D skeleton with reconstruction ability, simplicity, and direction properties. Secondly, this paper reconstructs the predesigned 3D skeleton of the aircraft from an RGB image and explores the 6D pose information in the reconstructed 3D skeleton. A 3D matrix is used to show the 3D skeleton and improve the encoding of the spatial information in the 3D skeleton. The experimental results show that the proposed method outperforms Wide-Depth-Range by 199% and 105% on the metric ADD and Rete, respectively. Compared with YOLO6D, the proposed method is 58.9% faster.
Daoyong Fu, Songchen Han, Binbin Liang, Wei Li 0075
IEEE Trans. Circuits Syst. Video Technol.4
2023 Nested Densely Atrous Spatial Pyramid Pooling and Deep Dense Short Connection for Skeleton Detection
abstract
The skeleton shows the local symmetry and the shape/topology of the object, and it is utilized for human pose recognition, road detection, text detection, and the representation of industrial parts. However, the size of the skeleton is variable, which makes high-level feature representation difficult. Existing methods only attempt to integrate multilevel features but ignore the extraction of high-level features and multiscales of contextual information that are helpful for the skeleton detection task. Thus, the contributions of this article include two aspects. The first contribution is to propose a nested densely atrous spatial pyramid pooling method that connects the atrous convolutions with different dilation rates in a nested cascade mode, which can provide the multiscale denser contextual information, a larger receptive field, and more local features. The second one is to propose a deep dense short connection (DDSC) that explores the role of features at different levels in the task of skeleton detection. DDSC adopts concatenation to fuse high-level semantic information with shape information. The proposed method is evaluated on four common datasets, and the experimental results show the effectiveness of the proposed method.
Daoyong Fu, Xiaofei Zeng, Songchen Han, Hanren Lin, Wei Li 0075
IEEE Trans. Hum. Mach. Syst.5
2022 Multi-stage attention network for video-based person re-identification
abstract
Abstract Video‐based person re‐identification (Re‐ID) has received increasing attention in video surveillance analysis in recent years. To extract relevant information of the target, many existing methods utilise the attention mechanism in the residual block of the ResNet. However, these methods only focus on the residual block and ignore the output of the shortcut part, which also contains rich information about the person. To solve this problem, a different aspect of network design is investigated: the insert position of the attention module. To simultaneously explore the discriminative information in both the residual block and the shortcut, a novel multi‐stage attention method is proposed by inserting the attention mechanism between stages of ResNet. Using this method can effectively extract the rich discriminative features of the target to better distinguish different pedestrians and improve the feature extraction capabilities of the model. Extensive experiments are conducted on four popular video‐based person Re‐ID datasets to demonstrate the effectiveness of the authors’ proposed method and display its superiority with the existing video‐based person Re‐ID methods.
Fan Yang 0104, Wei Li 0075, Binbin Liang, Songchen Han, Xuan Zhu 0007
IET Comput. Vis.2
2022 Airport small object detection based on feature enhancement
abstract
Abstract Video object detection is essential for airport surface surveillance, but the objects on the scene are mostly small objects with low resolution, they have no obvious feature information. Due to the scale differences of the objects and the fixed receptive field on the feature maps, detectors cannot model multi‐scale context information and cover all objects. In addition, although the video detection algorithm can be used as a method to solve the problem of small object detection, the temporal feature fusion method of current video detection is very dependent on the quality of a single feature map. Therefore, this paper aims to enhance the features of small objects of a single image. First, an attentional multi‐scale feature fusion enhancement (A‐MSFFE) network is built on the memory‐enhanced global‐local aggregation (MEGA) to supplement semantic and spatial information of small objects. Then, a context feature enhancement (CFE) module is designed for obtaining different receptive fields through different dilated convolutions. Meanwhile, a video detection dataset about the airport is established. Finally, the experimental results show that the proposed method can improve the detection accuracies of small objects and outperform other state‐of‐the‐art video object detection algorithms in self‐built airport dataset.
Xuan Zhu 0007, Binbin Liang, Daoyong Fu, Guoxin Huang, Fan Yang 0104, Wei Li 0075
IET Image Process.6
2022 Relation-based global-partial feature learning network for video-based person re-identification
Fan Yang 0104, Xiangtong Wang, Xuan Zhu 0007, Binbin Liang, Wei Li 0075
Neurocomputing5
2021 RLStereo: Real-Time Stereo Matching Based on Reinforcement Learning
abstract
Many state-of-the-art stereo matching algorithms based on deep learning have been proposed in recent years, which usually construct a cost volume and adopt cost filtering by a series of 3D convolutions. In essence, the possibility of all the disparities is exhaustively represented in the cost volume, and the estimated disparity holds the maximal possibility. The cost filtering could learn contextual information and reduce mismatches in ill-posed regions. However, this kind of methods has two main disadvantages: 1) cost filtering is very time-consuming, and it is thus difficult to simultaneously satisfy the requirements for both speed and accuracy; 2) thickness of the cost volume determines the disparity range which can be estimated, and the pre-defined disparity range may not meet the demand of practical application. This paper proposes a novel real-time stereo matching method called RLStereo, which is based on reinforcement learning and abandons the cost volume or the routine of exhaustive search. The trained RLStereo makes only a few actions iteratively to search the value of the disparity for each pair of stereo images. Experimental results show the effectiveness of the proposed method, which achieves comparable performances to state-of-the-art algorithms with real-time speed on the public large-scale testset, i.e., Scene Flow.
Menglong Yang, Fangrui Wu, Wei Li 0075
IEEE Trans. Image Process.3
2020 WaveletStereo: Learning Wavelet Coefficients of Disparity Map in Stereo Matching
abstract
Some stereo matching algorithms based on deep learning have been proposed and achieved state-of-the-art performances since some public large-scale datasets were put online. However, the disparity in smooth regions and detailed regions is still difficult to accurately estimate simultaneously. This paper proposes a novel stereo matching method called WaveletStereo, which learns the wavelet coefficients of the disparity rather than the disparity itself. The WaveletStereo consists of several sub-modules, where the low-frequency sub-module generates the low-frequency wavelet coefficients, which aims at learning global context information and well handling the low-frequency regions such as textureless surfaces, and the others focus on the details. In addition, a densely connected atrous spatial pyramid block is introduced for better learning the multi-scale image features. Experimental results show the effectiveness of the proposed method, which achieves state-of-the-art performance on the large-scale test dataset Scene Flow.
Menglong Yang, Fangrui Wu, Wei Li 0075
CVPR3
2018 A heading adjustment method in wireless directional sensor networks
Wei Li 0075, Changxin Huang, Songchen Han
Comput. Networks1