VLDB 2026 Research / reviewers in the wild / expert
Lefan Wang
dblp:199/6044
· DBLP profile ↗
8ranked-venue papers
5as first author
8since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language ModelsabstractIn recent years, vision language models (VLMs) have made significant advancements in video understanding. However, a crucial capability — fine-grained motion comprehension — remains under-explored in current benchmarks. To address this gap, we propose MotionBench, a comprehensive evaluation benchmark designed to assess the fine-grained motion comprehension of video understanding models. MotionBench evaluates models’ motion-level perception through six primary categories of motion-oriented question types and includes data collected from diverse sources, ensuring a broad representation of real-world video content. Experimental results reveal that existing VLMs perform poorly in understanding fine-grained motions. To enhance VLM’s ability to perceive fine-grained motion within a limited sequence length of LLM, we conduct extensive experiments reviewing VLM architectures optimized for video feature compression and propose a novel and efficient Through-Encoder (TE) Fusion method. Experiments show that higher frame rate inputs and TE Fusion yield improvements in motion understanding, yet there is still substantial room for enhancement. Our benchmark aims to guide and motivate the development of more capable video understanding models, emphasizing the importance of fine-grained motion comprehension. Project page: https://motion-bench.github.io. Wenyi Hong, Yean Cheng, Zhuoyi Yang, Lefan Wang, Xiaotao Gu, Shiyu Huang 0001, Yuxiao Dong, Jie Tang 0001 |
CVPR | 5 |
| 2025 | Semantic Representation Attack against Aligned Large Language ModelsabstractLarge Language Models (LLMs) increasingly employ alignment techniques to prevent harmful outputs. Despite these safeguards, attackers can circumvent them by crafting prompts that induce LLMs to generate harmful content. Current methods typically target exact affirmative responses, suffering from limited convergence, unnatural prompts, and high computational costs. We introduce semantic representation attacks, a novel paradigm that fundamentally reconceptualizes adversarial objectives against aligned LLMs. Rather than targeting exact textual patterns, our approach exploits the semantic representation space that can elicit diverse responses that share equivalent harmful meanings. This innovation resolves the inherent trade-off between attack effectiveness and prompt naturalness that plagues existing methods. Our Semantic Representation Heuristic Search (SRHS) algorithm efficiently generates semantically coherent adversarial prompts by maintaining interpretability during incremental search. We establish rigorous theoretical guarantees for semantic convergence and demonstrate that SRHS achieves unprecedented attack success rates (89.4% averaged across 18 LLMs, including 100% on 11 models) while significantly reducing computational requirements. Extensive experiments show that our method consistently outperforms existing approaches. Jiawei Lian, Jianhong Pan, Lefan Wang, Yi Wang 0068, Shaohui Mei, Lap-Pui Chau |
NeurIPS | 3 |
| 2025 | PADetBench: Towards benchmarking texture- and patch-based physical attacks against object detection
Jiawei Lian, Jianhong Pan, Lefan Wang, Yi Wang 0068, Shaohui Mei, Lap-Pui Chau |
Knowl. Based Syst. | 3 |
| 2025 | CAMCFormer: Cross-Attention and Multicorrelation Aided Transformer for Few-Shot Object Detection in Optical Remote Sensing ImagesabstractFew-shot object detection (FSOD) enables the detection of novel-class objects in remote sensing images (RSIs) with limited labeled samples. Although convolutional neural networks (CNNs) are commonly used for this task, they suffer from two inherent constraints. First, their limited local receptive field fails to capture global context within a single image and the relational dependencies between query and support images. Second, an additional feature alignment mechanism is typically required to bridge the gap between query and support images. To address these challenges, this work introduces a novel cross-attention and multicorrelation aided transformer (CAMCFormer) FSOD framework tailored for global feature representation and multicorrelation modeling in complex and large-scale RSIs. Specifically, a long-distance cross-attention module (LDCAM) is devised to capture dependencies between distant elements across query and support images at each feature extraction layer. This module facilitates the exchange of contextual information between images, resulting in more comprehensive feature representations and eliminating the need for separate feature alignment and fusion modules. Multicorrelation aided heads (MAHs) are constructed to enhance detection performance further to model various relational aspects, i.e., channel-correlation detection head (CCDH), spatial-correlation detection head (SCDH), and cross-attention detection head (CADH). These aided heads contribute to more robust and accurate classification and localization. Comprehensive experiments have been conducted, demonstrating the superiority of the proposed framework compared to several state-of-the-art detectors, highlighting its potential as an effective solution for FSOD in remote sensing scenarios. Lefan Wang, Shaohui Mei, Yi Wang 0068, Jiawei Lian, Zonghao Han, Yan Feng 0005 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Ankle Kinematics Estimation Using Artificial Neural Network and Multimodal IMU DataabstractInertial measurement units (IMUs) have become attractive for monitoring joint kinematics due to their portability and versatility. However, their limited accuracy, inability to analyze data in real-time, and complex data fusion algorithms requiring precise sensor-to-segment calibrations hinder their clinical and daily use. This paper introduces KEEN (KinEmatics Estimation Network), an innovative framework that exploits lightweight artificial neural networks (ANNs) to provide real-time predictions of multi-plane ankle kinematics using a minimal number of IMUs, without calibration requirements. Five ANN algorithms were developed and evaluated using 42 inputs derived from four IMUs in both intra-subject and inter-subject tasks. Extensive experimental results yielded exciting findings: even a single IMU located at the heel can provide clinically acceptable estimations of ankle kinematics, implying significant potential for cost and energy savings. Statistical analysis demonstrated the superiority of the developed Long Short-Term Memory (LSTM) network over the other models in intra-subject tasks, achieving impressive accuracy (RMSE: 1.88$\mathrm{^{\circ }}$$\pm$0.02$\mathrm{^{\circ }}$, MAE: 1.41$\mathrm{^{\circ }}$$\pm$0.01$\mathrm{^{\circ }}$, and r2 score: 0.93$\pm$0.01), indicating strong generalization within the same subject. In inter-subject tasks, the convolutional neural network (CNN) and the CNN-LSTM models showed comparable performance but statistically outperformed the other models in terms of estimation accuracy across various inputs. When using a single IMU, the CNN model achieved the lowest error (RMSE: 4.13$\mathrm{^{\circ }}$$\pm$0.55$\mathrm{^{\circ }}$, MAE: 3.33$\mathrm{^{\circ }}$$\pm$0.48$\mathrm{^{\circ }}$, and r2 score: 0.50$\pm$0.21), showcasing its effective generalization to new subjects. Furthermore, deploying the CNN into a microcontroller, with a sinlge IMU at the heel, resulted in promising real-time ankle kinematics estimations (RMSE: 3.34$\mathrm{^{\circ }}$$\pm$0.48$\mathrm{^{\circ }}$, MAE: 2.68$\mathrm{^{\circ }}$$\pm$0.46$\mathrm{^{\circ }}$ and r2 score: 0.63$\pm$0.07). Overall, this research highlights the potential of combining IMUs with ANNs as reliable and practical tools for early prevention and rehabilitation of ankle injuries. Lefan Wang, Pingfan Song, Thomas Stone, Adrian Weller, Sebastian W. Pattinson |
IEEE J. Biomed. Health Informatics | 1 |
| 2024 | Transformer-Based Few-Shot Object Detection with Multi-Relation Matching for Remote Sensing ImagesabstractFew-shot object detection (FSOD) on remote sensing images (RSIs) has garnered significant research interest due to its ability to detect novel classes using very few training examples from challenging remote sensing scenarios. Meta-learning FSOD methods, based on Faster R-CNN and YOLO structures, utilize a two-branch Siamese network as the backbone and compute the similarity between image regions for effective detection. However, almost all methods rely on extracting features using convolutional neural networks (CNNs). Inspired by the improved performance of transformer backbones for downstream tasks, a transformer-based FSOD method is proposed, which employs a transformer backbone with asymmetric-batched cross-attention for the two-branch feature extraction. Our model can improve the classification performance by introducing a Multi-Relation Matching (MRM) head for FSOD to enhance the similarity relation matching learning between two branches. Comprehensive experiments on DIOR benchmarks demonstrate the effectiveness of our model. Lefan Wang, Jiawei Lian, Yan Feng 0005, Shaohui Mei |
IGARSS | 1 |
| 2024 | Few-Shot Object Detection With Multilevel Information Interaction for Optical Remote Sensing ImagesabstractMetalearning has been widely applied to solve the few-shot object detection (FSOD) problem in natural scenes, which performs similarity measurement and information aggregation of the support set and the query set. However, regarding remote sensing images (RSIs), many difficulties caused by their disparities need to be further addressed, such as inconsistencies in imaging scale, direction, and background between support and query images. These result in feature misalignment and attention bias, interfering with model performance. In this article, a multilevel information interaction (MLII) strategy is proposed for FSOD to alleviate feature misalignment and attention bias. Information interactions are conducted within multiple scales of features and highlight similar regions of query and support features. A semantic enhancement module (SEM) is proposed to assist MLII in extracting key information and achieving more discriminative feature representation. Moreover, a feature cross-aggregation module (FCM) with separate classification losses is designed to train the detector to identify objects that coexist in query and support images. Extensive experiments demonstrate that the proposed method outperforms several state-of-the-art few-shot object detectors over commonly used benchmark datasets, i.e., DIOR and NWPU-10. Lefan Wang, Shaohui Mei, Yi Wang 0068, Jiawei Lian, Zonghao Han |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Diversity Measurement-Based Meta-Learning for Few-Shot Object Detection of Remote Sensing ImagesabstractMost object detection methods based on deep learning require large amounts of labeled data and can detect only the categories in the training set. Such issues significantly limit applications in remote sensing scenarios where it usually needs to recognize novel, unseen objects given very few training examples. To address these limitations, a novel meta-learning-based object detection method using Faster R-CNN framework is proposed for optical remote sensing image. Specifically, a diversity measurement module is proposed to measure diversity information between support images and query images on base classes so as to acquire more meta-knowledge. Experiments on DIOR dataset demonstrate our method has achieved superior performance than state-of-the-art meta-learning detection models in the field of remote sensing. Lefan Wang, Zonghao Han, Yan Feng 0005, Jiang Wei, Shaohui Mei |
IGARSS | 1 |