Binbin Liang

dblp:133/5925 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Multi-Task Learning for Airport Surface Surveillance: A Review
abstract
ABSTRACT The rapid growth of air transportation has surpassed the capabilities of traditional airport surveillance methods, such as visual observation and auxiliary equipment (e.g., ADS‐B, MLAT, radar), which struggle to provide all‐area, all‐weather situation awareness. Vision‐based deep learning methods, being cost‐effective and scalable, present promising alternatives but often fall short in delivering comprehensive awareness. Multi‐task learning (MTL) addresses these gaps by enabling models to simultaneously learn multiple related tasks, improving overall perception and decision‐making. Thus, this review reviews MTL systems for airport surface surveillance, categorising tasks into scene perception for intensive estimation and monitoring for non‐intensive estimation. This review identifies three key challenges: (1) efficient information sharing across tasks, (2) balancing multiple tasks, and (3) enhancing model training efficiency. The review examines these challenges through three lenses: neural network architecture, loss function optimization, and learning paradigms, proposing optimization strategies for each. As the first review to focus on MTL in airport surveillance, this review provides valuable insights into model design, task balancing, and training strategies, offering guidance for the future development of intelligent airport monitoring systems.
Daoyong Fu, Xiangtong Wang, Fangrui Wu, Songchen Han, Binbin Liang, Wei Li 0075
Expert Syst. J. Knowl. Eng.5
2025 Incorporating Fourier Transformation With Diffusion Models for Low-Light Image Enhancement
abstract
In this letter, we propose a diffusion-based framework that leverages the generative ability of diffusion models and the advantages of the physically explainable Fourier transformation for visually satisfactory low-light image enhancement. Specifically, we first employ an encoder to convert the paired low-light and normal-light images into latent features and transform the features into the frequency domain through Fourier transformation, resulting in amplitude components that contain illumination information and phase components that represent details information. Subsequently, we present the latent-Fourier diffusion model which performs diffusion operations on the phase components for details reconstruction. Furthermore, we propose a lightness boost module to reconstruct amplitude aiming to improve the contrast in the frequency domain, and the restored feature obtained by performing inverse Fourier transformation on the reconstructed phase and amplitude components is further refined by the proposed latent feature fusion module to achieve better visual perception. Finally, the refined feature is taken as input to a decoder to produce the final restored image. Extensive experiments on publicly available benchmarks demonstrate our proposed method outperforms state-of-the-art competitors.
Ailin Ma, Hai Jiang 0006, Binbin Liang, Songchen Han
IEEE Signal Process. Lett.3
2025 Multidimensional Fusion Network for Multispectral Object Detection
abstract
Multispectral object detection has attracted increasing attention recently due to its superior detection capacity under various illumination conditions. The key challenge lies in the effective aggregation of multi-spectral features to derive highly discriminative representations. To address this challenge, we propose a novel Multidimensional Fusion Network (MMFN) to explore multi-modal information from local, global, and channel perspectives. Specifically, at the local level, local features of different modalities and their inter-relationships are captured by a window-shifted fusion. As a complement to the local information, we designed a global interaction module that facilitates the fusion of holistic, high-level semantic information spanning the entire image. We distillate the channel dependencies and complementarities between different modalities through cross-channel learning and generate the final fused representation. Comprehensive experiments conducted on three publicly available datasets provide compelling evidence validating the superiority of the proposed methodology. The results exhibit notable performance gains over state-of-the-art multispectral object detectors. Our code will be released.
Fan Yang 0104, Binbin Liang, Wei Li 0075, Jianwei Zhang 0013
IEEE Trans. Circuits Syst. Video Technol.2
2024 Multi-Modal Disordered Representation Learning Network for Description-Based Person Search
abstract
Description-based person search aims to retrieve images of the target identity via textual descriptions. One of the challenges for this task is to extract discriminative representation from images and descriptions. Most existing methods apply the part-based split method or external models to explore the fine-grained details of local features, which ignore the global relationship between partial information and cause network instability. To overcome these issues, we propose a Multi-modal Disordered Representation Learning Network (MDRL) for description-based person search to fully extract the visual and textual representations. Specifically, we design a Cross-modality Global Feature Learning Architecture to learn the global features from the two modalities and meet the demand of the task. Based on our global network, we introduce a Disorder Local Learning Module to explore local features by a disordered reorganization strategy from both visual and textual aspects and enhance the robustness of the whole network. Besides, we introduce a Cross-modality Interaction Module to guide the two streams to extract visual or textual representations considering the correlation between modalities. Extensive experiments are conducted on two public datasets, and the results show that our method outperforms the state-of-the-art methods on CUHK-PEDES and ICFG-PEDES datasets and achieves superior performance.
Fan Yang 0104, Wei Li 0075, Menglong Yang, Binbin Liang
AAAI4
2023 Spatiotemporal Interaction Transformer Network for Video-Based Person Reidentification in Internet of Things
abstract
Video-based person reidentification, which is a significant application in the Internet of Things, aims to identify the same person in different video sequences across nonoverlapping cameras. Existing methods usually utilize temporal cues to enhance spatial features. However, these methods learn the temporal and spatial information separately, which breaks the relationship between them and ignores the positive role of temporal information for learning frame-level spatial representation in the process of spatial representation learning. In this article, we propose a novel spatiotemporal interaction transformer network (SITN) to solve this problem. To model the temporal information and the relationship between frames, we introduce a temporal interaction module (TIM) to interact between frame information. Meanwhile, we combine TIM with spatial transformer encoder to explore the positive role of temporal information in the learning procedure of the frame-level spatial feature. Moreover, we propose a transformer local learning scheme by reconstructing the 2-D spatial information of the frame patch sequences and extracting local features in a striped manner to strengthen the discriminative capability of our model. Extensive experiments are conducted on four public benchmarks. The results show that our model is superior compared with state-of-the-art methods.
Fan Yang 0104, Wei Li 0075, Binbin Liang, Jianwei Zhang 0013
IEEE Internet Things J.3
2023 Spatial-temporal alignment of time series with different sampling rates based on cellular multi-objective whale optimization
Binbin Liang, Songchen Han, Wei Li 0075, Guoxin Huang, Ruliang He
Inf. Process. Manag.1
2023 Similarity Measure of Time Series With Different Sampling Frequencies Based on Context Density Consistency and Dynamic Time Warping
abstract
Similarity measure of time series with different sampling frequencies is vitally important for many signal processing applications. Dynamic Time Warping (DTW) is one of the most popular methods for similarity measure of time series. However, conventional DTW algorithms have limitations when dealing with time series of different sampling frequencies due to context density inconsistency such as different internal change frequencies of derivatives, shapes, events and distances in local neighborhoods. In light of this, we propose a novel Context Density Consistency Dynamic Time Warping (CDC-DTW) algorithm. It firstly designs local context windows adaptive to the lengths of time series. Then it proposed a local spatial-temporal context density consistency technique by down-sampling and interpolation compensating the high-frequency time series following the context density of low-frequency time series. Besides, a normalized Hamming window weighting function is embedded into the local contexts to create robust weighted cost measure. Extensive experimental results on 128 gold-standard UCR datasets showed that CDC-DTW increased the similarity measure accuracy by 70.53% in average comparing with other 6 classic and state-of-the-art DTW baseline algorithms.
Wei Li 0075, Ruliang He, Binbin Liang, Fan Yang 0104, Songchen Han
IEEE Signal Process. Lett.3
2023 The 6D Pose Estimation of the Aircraft Using Geometric Property
abstract
The take-off and landing activities of aircraft must operate daily in a safe and orderly manner to provide a safe and efficient transition of passengers and goods. The 6D pose estimation of the aircraft is the key to guaranteeing the safe take-off and landing of the aircraft. The vision-based 6D pose estimation method, an important method when GPS and gyroscope are unavailable, faces the problem of poor estimation accuracy due to large scenes and large depths. An end-to-end 6D pose estimation method of the aircraft is proposed to solve this problem. Firstly, this paper combines the rigid structure characteristic of the aircraft with the direction property of arrows to build an aircraft 3D skeleton with reconstruction ability, simplicity, and direction properties. Secondly, this paper reconstructs the predesigned 3D skeleton of the aircraft from an RGB image and explores the 6D pose information in the reconstructed 3D skeleton. A 3D matrix is used to show the 3D skeleton and improve the encoding of the spatial information in the 3D skeleton. The experimental results show that the proposed method outperforms Wide-Depth-Range by 199% and 105% on the metric ADD and Rete, respectively. Compared with YOLO6D, the proposed method is 58.9% faster.
Daoyong Fu, Songchen Han, Binbin Liang, Wei Li 0075
IEEE Trans. Circuits Syst. Video Technol.3
2022 Design of Physical Layer Coding for Intermittent-Resistant Backscatter Communications Using Polar Codes
Binbin Liang
WASA (3)2
2022 Multi-stage attention network for video-based person re-identification
abstract
Abstract Video‐based person re‐identification (Re‐ID) has received increasing attention in video surveillance analysis in recent years. To extract relevant information of the target, many existing methods utilise the attention mechanism in the residual block of the ResNet. However, these methods only focus on the residual block and ignore the output of the shortcut part, which also contains rich information about the person. To solve this problem, a different aspect of network design is investigated: the insert position of the attention module. To simultaneously explore the discriminative information in both the residual block and the shortcut, a novel multi‐stage attention method is proposed by inserting the attention mechanism between stages of ResNet. Using this method can effectively extract the rich discriminative features of the target to better distinguish different pedestrians and improve the feature extraction capabilities of the model. Extensive experiments are conducted on four popular video‐based person Re‐ID datasets to demonstrate the effectiveness of the authors’ proposed method and display its superiority with the existing video‐based person Re‐ID methods.
Fan Yang 0104, Wei Li 0075, Binbin Liang, Songchen Han, Xuan Zhu 0007
IET Comput. Vis.3
2022 Airport small object detection based on feature enhancement
abstract
Abstract Video object detection is essential for airport surface surveillance, but the objects on the scene are mostly small objects with low resolution, they have no obvious feature information. Due to the scale differences of the objects and the fixed receptive field on the feature maps, detectors cannot model multi‐scale context information and cover all objects. In addition, although the video detection algorithm can be used as a method to solve the problem of small object detection, the temporal feature fusion method of current video detection is very dependent on the quality of a single feature map. Therefore, this paper aims to enhance the features of small objects of a single image. First, an attentional multi‐scale feature fusion enhancement (A‐MSFFE) network is built on the memory‐enhanced global‐local aggregation (MEGA) to supplement semantic and spatial information of small objects. Then, a context feature enhancement (CFE) module is designed for obtaining different receptive fields through different dilated convolutions. Meanwhile, a video detection dataset about the airport is established. Finally, the experimental results show that the proposed method can improve the detection accuracies of small objects and outperform other state‐of‐the‐art video object detection algorithms in self‐built airport dataset.
Xuan Zhu 0007, Binbin Liang, Daoyong Fu, Guoxin Huang, Fan Yang 0104, Wei Li 0075
IET Image Process.2
2022 Relation-based global-partial feature learning network for video-based person re-identification
Fan Yang 0104, Xiangtong Wang, Xuan Zhu 0007, Binbin Liang, Wei Li 0075
Neurocomputing4