VLDB 2026 Research / reviewers in the wild / expert
Yongmei Zhou
dblp:96/5435
· DBLP profile ↗
16ranked-venue papers
0as first author
13since 2021 · last 2025
0000-0003-2661-3078ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 8 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1Computer networks · 1Security and privacy · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Generative modeling · 61% Language models and text generation · 30% Trustworthy machine learning · 9% | |
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 54% Web and social media mining · 46% | |
| Computer graphics and multimedia
1 paper |
Multimedia analysis and retrieval · 50% Virtual and augmented reality · 50% | |
| Network and information security
2 papers |
Security and privacy of machine learning · 64% Malware analysis · 28% Network security · 8% |
Topics — the 12 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling › diffusion model › score-based generative model
denoising diffusion |
0.9 | 1 | 2025 | DiffTMR: Diffusion-based Hierarchical Alignment for Text-Molecule Retrieval · ACM Multimedia 2025 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | DiffTMR: Diffusion-based Hierarchical Alignment for Text-Molecule Retrieval · ACM Multimedia 2025 |
Natural language and speech › Language models and text generation
large language model safety |
0.9 | 1 | 2025 | Jailbreaking? One Step Is Enough! · ACL (1) 2025 |
Information retrieval
cross-modal retrieval |
0.9 | 1 | 2025 | DiffTMR: Diffusion-based Hierarchical Alignment for Text-Molecule Retrieval · ACM Multimedia 2025 |
Security and privacy of machine learning › adversarial attack
jailbreak attack |
0.9 | 1 | 2025 | Jailbreaking? One Step Is Enough! · ACL (1) 2025 |
Virtual and augmented reality › immersive interaction › multimodal interaction
cross-modal interaction |
0.8 | 1 | 2024 | Cross-Modal Interaction via Reinforcement Feedback for Audio-Lyrics Retrieval · IEEE ACM Trans. Audio Speech Lang. Process. 2024 |
Multimedia analysis and retrieval
cross-modal retrieval |
0.8 | 1 | 2024 | Cross-Modal Interaction via Reinforcement Feedback for Audio-Lyrics Retrieval · IEEE ACM Trans. Audio Speech Lang. Process. 2024 |
Web and social media mining
social influence analysis |
0.4 | 1 | 2019 | An Immunization Framework for Social Networks Through Big Data Based Influence Modeling · IEEE Trans. Dependable Secur. Comput. 2019 |
Web and social media mining
social network analysis |
0.4 | 1 | 2019 | An Immunization Framework for Social Networks Through Big Data Based Influence Modeling · IEEE Trans. Dependable Secur. Comput. 2019 |
Malware analysis › malware defense
malware propagation containment |
0.4 | 1 | 2019 | An Immunization Framework for Social Networks Through Big Data Based Influence Modeling · IEEE Trans. Dependable Secur. Comput. 2019 |
Machine learning › Trustworthy machine learning
out-of-distribution generalization |
0.3 | 1 | 2025 | DiffTMR: Diffusion-based Hierarchical Alignment for Text-Molecule Retrieval · ACM Multimedia 2025 |
Network security › intrusion detection and prevention
intrusion detection |
0.1 | 1 | 2019 | An Immunization Framework for Social Networks Through Big Data Based Influence Modeling · IEEE Trans. Dependable Secur. Comput. 2019 |
Methods — techniques the papers use, named apart from their topics
hierarchical diffusion alignment · 1.7dynamic perturbation embedding · 1.7adversarial prompting · 1.7reinforcement learning · 0.8influence spreading tree · 0.8cross-modal attention · 0.8breadth-first search · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Jailbreaking? One Step Is Enough!abstractWeixiong Zheng, Peijian Zeng, YiWei Li, Hongyan Wu, Nankai Lin, Junhao Chen, Aimin Yang, Yongmei Zhou. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Weixiong Zheng, Peijian Zeng, Nankai Lin, Aimin Yang 0002, Yongmei Zhou |
ACL (1) | 8 |
| 2025 | LLM-Driven Effective Knowledge Tracing by Integrating Dual-Channel Difficulty
Jiahui Cen, Jianghao Lin, Dong Zhou 0001, Weixuan Zhong, Aimin Yang 0002, Yongmei Zhou |
IEEE Big Data | 7 |
| 2025 | Central-Guided Convolutional Dual Attention for Document-Level Event Argument Extraction
Chengdong Lin, Jianghao Lin, Dong Zhou 0001, Yongmei Zhou, Aimin Yang 0002 |
IEEE Big Data | 4 |
| 2025 | WATER: A Two-Stage in-Context Learning Debiasing Framework for Multilingual Text ClassificationabstractRecently, Large Language Models (LLMs) have shown remarkable success across a variety of tasks, with rapid advancements in supporting multilingual capabilities. However, these models exhibit varying degrees of demographic biases in text classification tasks. Most existing research focuses on debiasing pre-trained models or addressing biases in monolingual text classification, resulting in limited exploration in multilingual contexts. To solve the above problems, this paper introduces a tWo-stAge in-conText learning dEbiasing fRamework (WATER). Our approach does not require updating the model's parameters and is adaptable to any language. It includes three key modules: sample selection, sample filtering, and template filling and prediction. In the first stage, we leverage a sample selection module to identify text that closely matches the model embeddings. In the second stage, we introduce an innovative Contextual Disparity Measure (CDM) in the sample filtering module to filter out samples that effectively address the bias associated with specific attributes. Finally, the template filling and prediction module is used to fill the selected samples into the template and input them into the model to complete the multilingual text classification task. Our experimental results verify the effectiveness of our method in mitigating biases related to four sensitive attributes of gender, age, race, and country, demonstrating its potential to improve the fairness and accuracy of LLMs in multilingual classification tasks. Zeyong Long, Dong Zhou 0001, Zhijin Chen, Yongmei Zhou, Nankai Lin, Aimin Yang 0002 |
CSCWD | 5 |
| 2025 | FairTriplet: Balancing Fairness and Accuracy in Contextual Pre-Trained Models Through Prefix TuningabstractNatural language processing models learn powerful language representation abilities from vast amounts of data, but they also inherit societal biases embedded in that data. Current research on debiasing often struggles to balance the removal of model bias with the preservation of model performance. Most existing approaches depend on fine-tuning model parameters, which can introduce uncertainties in model performance due to the modifications made to these parameters. In this paper, we propose a novel debiasing framework called FairTriplet. First, this framework employs prefix tuning to freeze the parameters of the original pre-trained model. Then, it optimizes the prefix parameters through two debiasing terms. These two debiasing terms function by reducing the semantic distance between social groups (e.g., male and female) and increasing the semantic distance between social groups and neutral attributes (e.g., family and occupation) in the semantic space. This approach not only removes bias from the model but also preserves its performance. Experimental results demonstrate that FairTriplet achieves state-of-the-art (SOTA) levels in debiasing while maintaining model performance on GLUE downstream tasks. Zeyong Long, Weixiong Zheng, Dong Zhou 0001, Yongmei Zhou, Nankai Lin, Aimin Yang 0002 |
CSCWD | 4 |
| 2025 | Enhancing Cross-Lingual Aspect-Based Sentiment Analysis with Code-Mixed In-Context Demonstrations and Language-Specific TagsabstractCross-lingual Aspect-based Sentiment Analysis (XABSA) aims to extract aspect-level sentiments across multiple languages. This task typically relies on source language data to train models and transfer them to target languages, so it faces significant challenges such as data scarcity and language disparities. To this end, this study proposes a code-Mixed In-conteXt lEaRning (MIXER). We design four kinds of demonstration retrieval libraries to introduce Code-mixed In-Context Demonstrations (CICD), which use the code-mixed mechanism to integrate the features of the target language and enrich the target language's knowledge while retaining the source language's knowledge. Language-Specific tags (LST) are introduced to enhance the model's understanding of multilingual demonstrations. To validate the effectiveness of MIXER, we conduct extensive experiments on the SemEval-2016 dataset, comparing its performance against existing XABSA methods. The experimental results show that MIXER performs better than existing XABSA methods with average F1 scores on the Mistral and Llama3 improved by 1.59% and 1.44%, respectively, highlighting its potential for broader multilingual applications. Meiyu Zeng, Xingming Liao, Yongmei Zhou, Nankai Lin, Aimin Yang 0002 |
CSCWD | 3 |
| 2025 | DomainDiff: Unified Two-Stage Optimization for Text-Video RetrievalabstractThe primary challenge in text-video retrieval lies in achieving cross-modal semantic alignment, particularly the discrepancy between the conciseness of textual descriptions, which often fail to fully encapsulate the breadth of video content, and the redundancy in video data, which introduces noise and masks important semantic features. Current methods align text and video by mapping them into a shared feature space. Despite notable advancements, the inherent differences in modality-specific representations create a bottleneck for fixed-point embedding techniques, making models highly sensitive to dataset distribution and hindering their generalization ability. In this paper, we present DomainDiff, a framework that enhances the embedding space through a two-stage process. In the first stage, stochastic domain modeling, we semantically expand text embeddings to explore potential regions aligned with video content. Simultaneously, we filter video segments to reduce redundancy and highlight key frames. In the second stage, the dynamic agent attention diffusion network, we leverage the generative properties of diffusion models to optimize the embedding space by viewing it from a joint probability distribution perspective. An agent attention mechanism dynamically integrates text and video features, ensuring accurate cross-modal alignment. Experimental results demonstrate that DomainDiff significantly improves retrieval performance across five benchmark datasets, with R@1 improvements ranging from 3% to 7.4%. Moreover, DomainDiff outperforms existing methods in handling long videos and complex textual descriptions, showcasing superior semantic robustness and generalization across varying distributions. Chenxu Wang 0019, Dong Zhou 0001, Jianghao Lin, Yongmei Zhou, Aimin Yang 0002 |
ICMR | 4 |
| 2025 | DiffTMR: Diffusion-based Hierarchical Alignment for Text-Molecule RetrievalabstractMolecular retrieval is critical in drug discovery and molecular design. Traditional discriminative methods often model the conditional probability distribution of retrieving candidates, treating the query text as a deterministic input. However, these approaches have notable limitations: (1) They often overlook the statistical properties of the original data distributions of queries and candidates, preventing the recognition of out-of-distribution data. (2) They struggle to balance retrieval accuracy and diversity when processing open-ended semantic queries. To address these challenges, we introduce DiffTMR, a novel framework that reformulates text-molecule retrieval as a reverse denoising process, progressively generating the joint distribution of candidates and queries from noises. DiffTMR uniquely integrates hierarchical diffusion alignment with dynamic perturbation embedding mechanisms. By employing text-anchored perturbations, it enhances the diversity of molecular representations, and through global-local progressive denoising, it achieves cross-modal hierarchical alignment. This leads to significant improvements in retrieval accuracy and out-of-domain generalization. Evaluations on benchmark datasets ChEBI-20 and PCdes demonstrate that DiffTMR surpasses current leading baselines by 4.2%-5.4% in Hits@1 metrics and exhibits superior performance in out-of-domain retrieval tasks. Chenxu Wang 0019, Dong Zhou 0001, Jianghao Lin, Yongmei Zhou, Aimin Yang 0002 |
ACM Multimedia | 5 |
| 2025 | GS2F: Multimodal Fake News Detection Utilizing Graph Structure and Guided Semantic FusionabstractThe prevalence of fake news online has become a significant societal concern. To combat this, multimodal detection techniques based on images and text have shown promise. Yet, these methods struggle to analyze complex relationships within and between modalities due to the diverse discriminative elements in the news content. In addition, research on multimodal and multi-class fake news detection remains insufficient. To address the above challenges, in this article, we propose a novel detection model, GS 2 F, leveraging g raph s tructure and g uided s emantic f usion. Specifically, we construct a multimodal graph structure to align two modalities and employ graph contrastive learning for refined fusion representations. Furthermore, a guided semantic fusion module is introduced to maximize the utilization of single-modal information and a dynamic contribution assignment layer is designed to weigh the importance of image, text, and multimodal features. Experimental results on Fakeddit demonstrate that our model outperforms existing methods, marking a step forward in the multimodal and multi-class fake news detection. Dong Zhou 0001, Qiang Ouyang, Nankai Lin, Yongmei Zhou, Aimin Yang 0002 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2024 | A Retrieval-Augmented Contrastive Framework for Legal Case Retrieval Based on Event Information
Changyong Fan, Nankai Lin, Dong Zhou 0001, Yongmei Zhou, Aimin Yang 0002 |
ACML | 4 |
| 2024 | HiRAG: A Historical Information-Driven Retrieval-Augmented Generation Framework for Background Summarization
Dong Zhou 0001, Binli Zeng, Nankai Lin, Yongmei Zhou, Aimin Yang 0002 |
ACML | 4 |
| 2024 | Cross-Modal Interaction via Reinforcement Feedback for Audio-Lyrics RetrievalabstractThe task of retrieving audio content relevant to lyric queries and vice versa plays a critical role in music-oriented applications. In this process, robust feature representations have to be learned for two modalities. Furthermore, interactions between different modalities should be properly captured at a fine-grained level. Existing approaches can effectively extract modal representations and perform retrieving between different modalities through alignment. However, these approaches model interactions between audio and lyrics in a coarse-grained manner. Especially the input features and interactions between enhanced representations produced by the alignment module are largely ignored, resulting in low-quality modality representations for final retrieval. This paper presents a novel method named CMRF that accomplishes cross-modal interactions via a reinforcement feedback procedure to learn high-quality multi-modal embeddings. Initially, we implicitly assimilate representations across distinct modalities via directional pairwise cross-modal attention. Subsequently, our approach recurrently identifies pivotal constituents within these elevated-level attributes to engage with the primary input features via reinforcement learning, thus augmenting the quality of multi-modal embeddings. In addition, we introduce a novel audio-lyrics datasetAL-song, which consists of paired audio with corresponding lyrics for the audio-lyrics retrieval task. The empirical findings derived from theAL-songdataset and the benchmark datasetSounddescssubstantiate the efficacy and efficiency of CMRF when juxtaposed with state-of-the-art methodologies. Dong Zhou 0001, Fang Lei, Lin Li 0001, Yongmei Zhou, Aimin Yang 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | Deep transfer learning mechanism for fine-grained cross-domain sentiment classificationabstractThe goal of cross-domain sentiment classification is to utilise useful information in the source domain to help classify sentiment polarity in the target domain, which has a large number of unlabelled data. Most of the existing methods focus on extracting the invariant features between two domains. But they cannot make better use of the unlabelled data in the target domain. To solve this problem, we present a deep transfer learning mechanism (DTLM) for fine-grained cross-domain sentiment classification. DTLM provides a transfer mechanism to better transfer sentiment across domains by incorporating BERT(Bidirextional Encoder Representations from Transformers) and KL (Kullback-Leibler) divergence. We introduce BERT as a feature encoder to map the text data of different domains into a shared feature space. Then, we design a domain adaptive model using KL divergence to eliminate the difference of feature distribution between the source domain and target domain. In addition, we introduce the entropy minimisation and consistency regularisation to process unlabelled samples in the target domain. Extensive experiments on the datasets from YelpAspect, SemEval 2014 task 4 and Twitter not only demonstrate the effectiveness of our proposed method but also provide a better way for cross-domain sentiment classification. Zixuan Cao, Yongmei Zhou, Aimin Yang 0002, Sancheng Peng |
Connect. Sci. | 2 |
| 2019 | An Immunization Framework for Social Networks Through Big Data Based Influence ModelingabstractSocial networks are critical in terms of information or malware propagation. However, how to contain the spreading of malware in social networks is still an open and challenging issue. In this paper, we propose a novel defending method through big data based influence modeling. We first establish a social interaction graph based on big data sets of the studied object. Based on the graph, we are able to measure direct influence of individuals by computing each node's strength, which includes the degree of the node and the total number of messages sent by each user to her friends. Then, we design an algorithm to construct influence spreading tree using the breadth first search strategy, and measure indirect influence of individuals by traversing the tree. We identify the top k influential nodes among all the nodes via the social influence strength, and propose an immunization algorithm to defend social networks against various attacks. The extensive experiments show that influence can spread easily in social networks, and the greater the influence of initial spread node is, the more impact it is on the malware propagation in social networks. The proposed method provides an effective solution to the prevention of malware or malicious messages propagation in social networks. Sancheng Peng, Guojun Wang 0001, Yongmei Zhou, Cong Wan, Cong Wang 0009, Shui Yu 0001, Jianwei Niu 0002 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2018 | Influence analysis in social networks: A survey
Sancheng Peng, Yongmei Zhou, Lihong Cao, Shui Yu 0001, Jianwei Niu 0002, Weijia Jia 0001 |
J. Netw. Comput. Appl. | 2 |
| 2017 | Throughput-Optimized FPGA Accelerator for Deep Convolutional Neural NetworksabstractDeep convolutional neural networks (CNNs) have gained great success in various computer vision applications. State-of-the-art CNN models for large-scale applications are computation intensive and memory expensive and, hence, are mainly processed on high-performance processors like server CPUs and GPUs. However, there is an increasing demand of high-accuracy or real-time object detection tasks in large-scale clusters or embedded systems, which requires energy-efficient accelerators because of the green computation requirement or the limited battery restriction. Due to the advantages of energy efficiency and reconfigurability, Field-Programmable Gate Arrays (FPGAs) have been widely explored as CNN accelerators. In this article, we present an in-depth analysis of computation complexity and the memory footprint of each CNN layer type. Then a scalable parallel framework is proposed that exploits four levels of parallelism in hardware acceleration. We further put forward a systematic design space exploration methodology to search for the optimal solution that maximizes accelerator throughput under the FPGA constraints such as on-chip memory, computational resources, external memory bandwidth, and clock frequency. Finally, we demonstrate the methodology by optimizing three representative CNNs (LeNet, AlexNet, and VGG-S) on a Xilinx VC709 board. The average performance of the three accelerators is 424.7, 445.6, and 473.4GOP/s under 100MHz working frequency, which outperforms the CPU and previous work significantly. Yong Dou, Jingfei Jiang, Jinwei Xu, Shijie Li 0002, Yongmei Zhou, Yingnan Xu |
ACM Trans. Reconfigurable Technol. Syst. | 6 |