Zhen Lan

dblp:283/6263 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0001-6112-9279ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Image recognition and object detection · 67% Efficient and distributed learning · 33%
Human-computer interaction and pervasive computing
1 paper
Wearable and physiological sensing · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection › object detection
dark object detection
0.912025
Adaptive Knowledge Distillation With Attention-Based Multi-Modal Fusion for Robust Dim Object Detection · IEEE Trans. Multim. 2025
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.912025
Adaptive Knowledge Distillation With Attention-Based Multi-Modal Fusion for Robust Dim Object Detection · IEEE Trans. Multim. 2025
Computer vision › Image recognition and object detection
object detection
0.912025
Adaptive Knowledge Distillation With Attention-Based Multi-Modal Fusion for Robust Dim Object Detection · IEEE Trans. Multim. 2025
Wearable and physiological sensing
brain-computer interface
0.312025
Adaptive Knowledge Distillation With Attention-Based Multi-Modal Fusion for Robust Dim Object Detection · IEEE Trans. Multim. 2025

Methods — techniques the papers use, named apart from their topics

eye-tracking-based slow serial visual presentation · 1.7attention-based multimodal fusion · 1.7
YearPublicationVenuePosition
2026 Adaptive Modality Balanced Online Knowledge Distillation for Brain-Eye-Computer-Based Dim Object Detection
abstract
Advanced cognition can be measured from the human brain using brain-computer interfaces (BCIs). Integrating these interfaces with computer vision techniques, which possess efficient feature extraction capabilities, can achieve more robust and accurate detection of dim targets in aerial images. However, existing target detection methods primarily concentrate on homogeneous data, lacking efficient and versatile processing capabilities for heterogeneous multimodal data. In this article, we first build a brain-eye-computer-based object detection system for aerial images under few-shot conditions. This system detects suspicious targets using region proposal networks (RPNs), evokes the event-related potential (ERP) signal in electroencephalogram (EEG) through the eye-tracking-based slow serial visual presentation (ESSVP) paradigm, and constructs the EEG-image data pairs with eye movement data. Then, an adaptive modality balanced online knowledge distillation (AMBOKD) method is proposed to recognize dim objects with the EEG-image data. AMBOKD fuses EEG and image features using a multihead attention module, establishing a new modality with comprehensive features. To enhance the performance and robust capability of the fusion modality, simultaneous training and mutual learning between modalities are enabled by end-to-end online KD (OKD). During the learning process, an adaptive modality balancing module is proposed to ensure multimodal equilibrium by dynamically adjusting the weights of the importance and the training gradients across various modalities. The effectiveness and superiority of our method are demonstrated by comparing it with existing state-of-the-art methods. Additionally, experiments conducted on public datasets and real-world scenarios demonstrate the reliability and practicality of the proposed system and the designed method. The dataset and the source code can be found at: https://github.com/lizixing23/AMBOKD.
Zixing Li, Zhen Lan, Xiaojia Xiang, Jun Lai, Dengqing Tang
IEEE Trans. Neural Networks Learn. Syst.3
2025 RMKD: Relaxed matching knowledge distillation for short-length SSVEP-based brain-computer interfaces
Zhen Lan, Zixing Li, Xiaojia Xiang, Dengqing Tang, Min Wu 0008, Zhenghua Chen
Neural Networks1
2025 MTSNet: Convolution-Based Transformer Network With Multi-Scale Temporal-Spectral Feature Fusion for SSVEP Signal Decoding
abstract
Improving the decoding performance of steady-state visual evoked (SSVEP) signals is crucial for the practical application of SSVEP-based brain-computer interface (BCI) systems. Although numerous methods have achieved impressive results in decoding SSVEP signals, most of them focus only on the temporal or spectral domain information or concatenate them directly, which may ignore the complementary relationship between different features. To address this issue, we propose a dual-branch convolution-based Transformer network with multi-scale temporal-spectral feature fusion, termed MTSNet, to improve the decoding performance of SSVEP signals. Specifically, the temporal branch extracts temporal features from the SSVEP signals using the multi-level convolution- based Transformer (Convformer) that can adapt to the dynamic fluctuations of SSVEP signals. In parallel, the spectral branch takes the complex spectrum converted from temporal signals by the zero-padding fast Fourier transform as input and uses the Convformer to extract spectral features. These extracted temporal and spectral features are then integrated by the multi-scale feature fusion module to obtain comprehensive features with different scale information, thereby enhancing the interactions between the features and improving the effectiveness and robustness. Extensive experimental results on two widely used public SSVEP datasets, Benchmark and BETA, show that the proposed MTSNet significantly outperforms the state-of-the-art calibration-free methods in terms of accuracy and ITR. The superior performance demonstrates the effectiveness of our method in decoding SSVEP signals, which may facilitate the practical application of SSVEP-based BCI systems.
Zhen Lan, Zixing Li, Xiaojia Xiang, Dengqing Tang, Min Wu 0008, Zhenghua Chen
IEEE J. Biomed. Health Informatics1
2025 Adaptive Knowledge Distillation With Attention-Based Multi-Modal Fusion for Robust Dim Object Detection
abstract
Automated object detection in aerial images is crucial in both civil and military applications. Existing computer vision-based object detection methods are not robust enough to precisely detect dim objects in aerial images due to the cluttered backgrounds, various observing angles, small object scales, and severe occlusions. Recently, electroencephalography (EEG)-based object detection methods have received increasing attention owing to the advanced cognitive capabilities of human vision. However, how to combine the human intelligence with computer intelligence to achieve robust dim object detection is still an open question. In this paper, we propose a novel approach to efficiently fuse and exploit the properties of multi-modal data for dim object detection. Specifically, we first design a brain-computer interface (BCI) paradigm called eye-tracking-based slow serial visual presentation (ESSVP) to simultaneously collect the paired EEG and image data when subjects search for the dim objects in aerial images. Then, we develop an attention-based multi-modal fusion network to selectively aggregate the learned features of EEG and image modalities. Furthermore, we propose an adaptive multi-teacher knowledge distillation method to efficiently train the multi-modal dim object detector for better performance. To evaluate the effectiveness of our method, we conduct extensive experiments on the collected dataset in subject-dependent and subject-independent tasks. The experimental results demonstrate that the proposed dim object detection method exhibits superior effectiveness and robustness compared to the baselines and the state-of-the-art methods.
Zhen Lan, Zixing Li, Xiaojia Xiang, Dengqing Tang, Jun Lai
IEEE Trans. Multim.1
2024 M2KD: Multi-Teacher Multi-Modal Knowledge Distillation for Aerial View Object Classification
abstract
Object classification in aerial images is expected to play an important role in a wide range of applications. Multi-modal methods have emerged as a promising approach in aerial image classification due to the differences and comple-mentarities between different modalities. However, most existing methods simply combine multi-modal features or directly use a single optimization strategy for joint training, which is not comprehensive and usually constrains the classification accuracy. To mitigate this problem, we propose a multi-teacher multi-modal knowledge distillation (M2KD) method for aerial view object classification tasks. Specifically, the attention-based feature fusion network is first constructed to extract and merge more discriminative features from multi-modal data, i.e., the aerial images and electroencephalography (EEG) signals. To further improve the classification performance, the multi-teacher knowledge distillation framework is designed to assist the training of the student by leveraging the complementary multi-modal knowledge. Extensive experiments on the collected multi-modal dataset demonstrate the contribution and effectiveness of our M2KD method for aerial view object classification.
Zhen Lan, Zixing Li, Xiaojia Xiang, Dengqing Tang
IJCNN1
2024 Multimodal Mutual Learning with Online Knowledge Distillation for Dim Object Recognition in Aerial Images
abstract
Deep learning methods have shown promise in various visual tasks such as object recognition. However, achieving robust and accurate performance in dim object recognition for remote sensing images remains challenging in the field of computer vision. This challenge can be attributed to factors such as cluttered backgrounds, varying observing angles, and limited availability of labeled data. In contrast, the human brain exhibits robust and efficient recognition of sensitive targets. To leverage the strengths of both computer calculation and human cognition, we propose a multimodal mutual learning with online knowledge distillation method (MMOKD) for object recognition. Our approach enables simultaneous training and mutual learning between modalities, where each modality serves as both a teacher and a student. A series of experiments are conducted to verify the potential of multimodal learning for object recognition. The results demonstrate that our approach not only enhances the robustness of multimodal fusion model, but also improves the accuracy of visual modality.
Zixing Li, Zhen Lan, Xiaojia Xiang, Dengqing Tang
SMC2
2022 Deep Reinforcement Learning of Collision-Free Flocking Policies for Multiple Fixed-Wing UAVs Using Local Situation Maps
abstract
The evolution of artificial intelligence and Internet of Things (IoT) envision a highly integrated artificial IoT (AIoT) network. Flocking and cooperation with multiple unmanned aerial vehicles (UAVs) are expected to play a vital role in industrial AIoT networks. In this article, we formulate the collision-free flocking problem of fixed-wing UAVs as a Markov decision process and solve it in the deep reinforcement learning (DRL) framework. Our method can deal with a variable number of followers by encoding the dynamic environmental state into a fixed-length embedding tensor. Specifically, each follower constructs a fixed-size local situation map that describes the collision risks with other followers nearby. The local situation maps are used by a proposed DRL algorithm to learn the collision-free flocking behavior. To further improve the learning efficiency, we design a reference-point-based action selection strategy and an adaptive mechanism. We compare the proposed MA2D3QN algorithm with several benchmark DRL algorithms through numerical simulation, and we verify its advantages in learning efficiency and performance. Finally, we demonstrate the scalability and adaptability of MA2D3QN in a semiphysical simulation experiment.
Chang Wang 0005, Xiaojia Xiang, Zhen Lan, Yuna Jiang
IEEE Trans. Ind. Informatics4
2021 Flocking and Collision Avoidance for a Dynamic Squad of Fixed-Wing UAVs Using Deep Reinforcement Learning
abstract
Developing the flocking behavior for a dynamic squad of fixed-wing UAVs is still a challenge due to kinematic complexity and environmental uncertainty. In this paper, we deal with the decentralized flocking and collision avoidance problem through deep reinforcement learning (DRL). Specifically, we formulate a decentralized DRL-based decision making framework from the perspective of every follower, where a collision avoidance mechanism is integrated into the flocking controller. Then, we propose a novel reinforcement learning algorithm PS-CACER for training a shared control policy for all the followers. Besides, we design a plug-n-play embedding module based on convolutional neural networks and the attention mechanism. As a result, the variable-length system state can be encoded into a fixed-length embedding vector, which makes the learned DRL policy independent with the number and the order of followers. Finally, numerical simulation results demonstrate the effectiveness of the proposed method, and the learned policies can be directly transferred to semi-physical simulation without any parameter finetuning.
Xiaojia Xiang, Chang Wang 0005, Zhen Lan
IROS4
2021 MACRO: Multi-Attention Convolutional Recurrent Model for Subject-Independent ERP Detection
abstract
Due to the low signal-to-noise ratio, limited training samples, and large inter-subject variabilities in electroencephalogram (EEG) signals, developing a subject-independent brain-computer interface (BCI) system used for new users without any calibration is still challenging. In this letter, we propose a novel Multi-Attention Convolutional Recurrent mOdel (MACRO) for EEG-based event-related potential (ERP) detection in the subject-independent scenario. Specifically, the convolutional recurrent network is designed to capture the spatial-temporal features, while the multi-attention mechanism is integrated to focus on the most discriminative channels and temporal periods of EEG signals. Comprehensive experiments conducted on a benchmark dataset for RSVP-based BCIs show that our method achieves the best performance compared with the five state-of-the-art baseline methods. This result indicates that our method is able to extract the underlying subject-invariant EEG features and generalize to unseen subjects. Finally, the ablation studies verify the effectiveness of the designed multi-attention mechanism in MACRO for EEG-based ERP detection.
Zhen Lan, Zixing Li, Dengqing Tang, Xiaojia Xiang
IEEE Signal Process. Lett.1