VLDB 2026 Research / reviewers in the wild / expert
Zefeng Li
dblp:189/8685
· DBLP profile ↗
12ranked-venue papers
4as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Loyalty-SMOTE: Data synthesis algorithm for effective imbalanced data classification
Shengquan Hu, Junfei Li, Zefeng Li, Ka Lun Eddie Law |
Neural Networks | 3 |
| 2025 | Denoised Recommendation Model with Collaborative Signal Decoupling
Zefeng Li, Ning Yang 0001 |
IEEE Big Data | 1 |
| 2025 | A Two-Stage Method for Specular Highlight Detection and Removal in Medical Images
Zefeng Li, Mingyue Cui, Daosong Hu, Jin Gong, Jingchong Weng, Lele Tian, Kai Huang 0001 |
MICCAI (10) | 1 |
| 2025 | GaussianFusion: Gaussian-Based Multi-Sensor Fusion for End-to-End Autonomous DrivingabstractMulti-sensor fusion is crucial for improving the performance and robustness of end-to-end autonomous driving systems. Existing methods predominantly adopt either attention-based flatten fusion or bird’s eye view fusion through geometric transformations. However, these approaches often suffer from limited interpretability or dense computational overhead. In this paper, we introduce GaussianFusion, a Gaussian-based multi-sensor fusion framework for end-to-end autonomous driving. Our method employs intuitive and compact Gaussian representations as intermediate carriers to aggregate information from diverse sensors. Specifically, we initialize a set of 2D Gaussians uniformly across the driving scene, where each Gaussian is parameterized by physical attributes and equipped with explicit and implicit features. These Gaussians are progressively refined by integrating multi-modal features. The explicit features capture rich semantic and spatial information about the traffic scene, while the implicit features provide complementary cues beneficial for trajectory planning. To fully exploit rich spatial and semantic information in Gaussians, we design a cascade planning head that iteratively refines trajectory predictions through interactions with Gaussians. Extensive experiments on the NAVSIM and Bench2Drive benchmarks demonstrate the effectiveness and robustness of the proposed GaussianFusion framework. The source code is included in the supplementary material and will be released publicly. Shuai Liu 0009, Quanmin Liang, Zefeng Li, Boyang Li 0009, Kai Huang 0001 |
NeurIPS | 3 |
| 2025 | Sometimes Simpler is Better: A Comprehensive Analysis of State-of-the-Art Provenance-Based Intrusion Detection Systems
Tristan Bilot, Baoxiang Jiang, Zefeng Li, Nour El Madhoun, Khaldoun Al Agha, Anis Zouaoui, Thomas Pasquier |
USENIX Security Symposium | 3 |
| 2024 | GroundingGPT: Language Enhanced Multi-modal Grounding ModelabstractZhaowei Li, Qi Xu, Dong Zhang, Hang Song, YiQing Cai, Qi Qi, Ran Zhou, Junting Pan, Zefeng Li, Vu Tu, Zhida Huang, Tao Wang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yiqing Cai, Junting Pan, Zefeng Li, Vu Tu, Zhida Huang |
ACL (1) | 9 |
| 2024 | Leveraging Social Context for Humor Recognition and Sense of Humor Evaluation in Social Media with a New Chinese Humor Corpus - HumorWBabstractWith the development of the Internet, social media has produced a large amount of user-generated data, which brings new challenges for humor computing. Traditional humor computing research mainly focuses on the content, while neglecting the information of interaction relationships in social media. In addition, both content and users are important in social media, while existing humor computing research mainly focuses on content rather than people. To address these problems, we model the information transfer and entity interactions in social media as a heterogeneous graph, and create the first dataset which introduces the social context information - HumorWB, which is collected from Chinese social media - Weibo. Two humor-related tasks are designed in the dataset. One is a content-oriented humor recognition task, and the other is a novel humor evaluation task. For the above tasks, we purpose a graph-based model called SCOG, which uses heterogeneous graph neural networks to optimize node representation for downstream tasks. Experimental results demonstrate the effectiveness of feature extraction and graph representation learning methods in the model, as well as the necessity of introducing social context information. Zeyuan Zeng, Zefeng Li, Liang Yang 0003, Hongfei Lin |
LREC/COLING | 2 |
| 2024 | Image Caption Method from Coarse to Fine Based On Dual Encoder-Decoder FrameworkabstractEncoders are widely used in the field of image caption, but the statements generated by the current image caption method may miss the target and the generated description statements are not appropriate enough for the image content. In order to solve the above problems, we propose a coarse-fine image caption method based on dual encoder-decoder framework, which provides a mechanism for discovering and correcting omissions and enables the model to generate a complete image description. Firstly, an image feature extractor based on global and local information is designed, which can extract global information and local information of image and obtain more abundant image representation. Secondly, a dual encoder-decoder framework is designed, which consists of a coarse-grained encoder-decoder and a fine-grained encoder-decoder. Coarse-grained encoder-decoder requires only the original image features as input, which is processed by transformer to produce a coarse text description. In addition, an image feature auto-enhancement module is proposed to detect missing objects in coarse text and enhance their feature expression. Finally, the fine-grained encoder-decoder uses both the image feature and the coarse text caption as input, and generates the final fine-grained caption after multi-modal information fusion. Experimental results on MSCOCO datasets show that our proposed method outperforms previous image caption methods and achieves a performance of 39.7 BLEU-4 score and 121.6 CIDEr-A score. Zefeng Li, Yuehu Liu, Yonghong Song |
IJCNN | 1 |
| 2024 | MASTERKEY: Automated Jailbreaking of Large Language Model Chatbots
Gelei Deng, Yi Liu 0069, Yuekang Li, Kailong Wang 0001, Ying Zhang 0066, Zefeng Li, Haoyu Wang 0001, Tianwei Zhang 0004, Yang Liu 0003 |
NDSS | 6 |
| 2024 | SeisCLIP: A Seismology Foundation Model Pre-Trained by Multimodal Data for Multipurpose Seismic Feature ExtractionabstractIn seismology, while training a specific deep learning model for each task is common, it often faces challenges such as the scarcity of labeled data and limited regional generalization. Addressing these issues, we introduce SeisCLIP: a foundation model for seismology, leveraging contrastive learning during pre-training on multi-modal data of seismic waveform spectra and the corresponding local and global event information. SeisCLIP consists of a transformer-based spectrum encoder and an MLP-based information encoder that are jointly pre-trained on massive data. During pre-training, contrastive learning aims to enhance representations by training two encoders to bring corresponding waveform spectra and event information closer in the feature space, while distancing uncorrelated pairs. Remarkably, the pre-trained spectrum encoder offers versatile features, enabling its application across diverse tasks and regions. Thus, it requires only modest datasets for fine-tuning to specific downstream tasks. Our evaluations demonstrate SeisCLIP’s superior performance over baseline methods in tasks like event classification, localization, and focal mechanism analysis, even when using distinct datasets from various regions. In essence, SeisCLIP emerges as a promising foundational model for seismology, potentially revolutionizing foundation-model-based research in the domain. Xu Si, Xinming Wu, Hanlin Sheng, Zefeng Li |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | BloodNet: An attention-based deep network for accurate, efficient, and costless bloodstain time since deposition inferenceabstractThe time since deposition (TSD) of a bloodstain, i.e., the time of a bloodstain formation is an essential piece of biological evidence in crime scene investigation. The practical usage of some existing microscopic methods (e.g., spectroscopy or RNA analysis technology) is limited, as their performance strongly relies on high-end instrumentation and/or rigorous laboratory conditions. This paper presents a practically applicable deep learning-based method (i.e., BloodNet) for efficient, accurate, and costless TSD inference from a macroscopic view, i.e., by using easily accessible bloodstain photos. To this end, we established a benchmark database containing around 50,000 photos of bloodstains with varying TSDs. Capitalizing on such a large-scale database, BloodNet adopted attention mechanisms to learn from relatively high-resolution input images the localized fine-grained feature representations that were highly discriminative between different TSD periods. Also, the visual analysis of the learned deep networks based on the Smooth Grad-CAM tool demonstrated that our BloodNet can stably capture the unique local patterns of bloodstains with specific TSDs, suggesting the efficacy of the utilized attention mechanism in learning fine-grained representations for TSD inference. As a paired study for BloodNet, we further conducted a microscopic analysis using Raman spectroscopic data and a machine learning method based on Bayesian optimization. Although the experimental results show that such a new microscopic-level approach outperformed the state-of-the-art by a large margin, its inference accuracy is significantly lower than BloodNet, which further justifies the efficacy of deep learning techniques in the challenging task of bloodstain TSD inference. Our code is publically accessible via https://github.com/shenxiaochenn/BloodNet. Our datasets and pre-trained models can be freely accessed via https://figshare.com/articles/dataset/21291825. Gongji Wang, Qinru Sun, Zefeng Li, Xinggong Liang, Run Chen, Fan Wang 0023, Zhenyuan Wang, Chunfeng Lian |
Briefings Bioinform. | 6 |
| 2022 | Memeplate: A Chinese Multimodal Dataset for Humor Understanding in Meme Templates
Zefeng Li, Hongfei Lin, Liang Yang 0003, Bo Xu 0009, Shaowu Zhang 0002 |
NLPCC (1) | 1 |