VLDB 2026 Research / reviewers in the wild / expert
Xiaowei Yuan
dblp:46/957
· DBLP profile ↗
14ranked-venue papers
8as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 7 first-author · 8 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Language models and text generation · 74% Probabilistic and Bayesian machine learning · 10% Face, body and person analysis · 5% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 13 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
retrieval-augmented generation |
1.2 | 2 | 2026 | Efficient Prior-Guided Reasoning for Robust Retrieval-Augmented Generation under Conflicts · ACL (1) 2026 Improving Zero-shot LLM Re-Ranker with Risk Minimization · EMNLP 2024 |
Natural language and speech › Language models and text generation › retrieval-augmented generation
robust retrieval-augmented generation |
1.0 | 1 | 2026 | Efficient Prior-Guided Reasoning for Robust Retrieval-Augmented Generation under Conflicts · ACL (1) 2026 |
Natural language and speech › Language models and text generation
large language model |
0.9 | 1 | 2025 | Exploiting Contextual Knowledge in LLMs through V-usable Information based Layer Enhancement · ACL (1) 2025 |
Machine learning › Probabilistic and Bayesian machine learning
bayesian decision theory |
0.8 | 1 | 2024 | Improving Zero-shot LLM Re-Ranker with Risk Minimization · EMNLP 2024 |
Natural language and speech › Language models and text generation
in-context generation |
0.8 | 1 | 2024 | On the In-context Generation of Language Models · EMNLP 2024 |
Natural language and speech › Language models and text generation
in-context learning |
0.8 | 1 | 2024 | On the In-context Generation of Language Models · EMNLP 2024 |
Natural language and speech › Language models and text generation › large language model training
language model pretraining |
0.8 | 1 | 2024 | On the In-context Generation of Language Models · EMNLP 2024 |
Information retrieval
reranking |
0.8 | 1 | 2024 | Improving Zero-shot LLM Re-Ranker with Risk Minimization · EMNLP 2024 |
Information retrieval
retrieval models and ranking |
0.8 | 1 | 2024 | Improving Zero-shot LLM Re-Ranker with Risk Minimization · EMNLP 2024 |
Computer vision › 3D vision
3d face reconstruction |
0.4 | 1 | 2019 | Face De-Occlusion Using 3D Morphable Model and Generative Adversarial Network · ICCV 2019 |
Computer vision › Face, body and person analysis › face restoration
face de-occlusion |
0.4 | 1 | 2019 | Face De-Occlusion Using 3D Morphable Model and Generative Adversarial Network · ICCV 2019 |
Machine learning › Trustworthy machine learning
robustness |
0.3 | 1 | 2026 | Efficient Prior-Guided Reasoning for Robust Retrieval-Augmented Generation under Conflicts · ACL (1) 2026 |
Machine learning › Generative modeling
generative adversarial network |
0.1 | 1 | 2019 | Face De-Occlusion Using 3D Morphable Model and Generative Adversarial Network · ICCV 2019 |
Methods — techniques the papers use, named apart from their topics
risk minimization · 1.5query likelihood model · 1.5reinforcement learning · 1.0bernoulli-gated dropout · 1.0v-usable information · 0.9synthetic dataset generation · 0.8latent variable model · 0.8adversarial learning · 0.43d morphable model · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Prior-Guided Reasoning for Robust Retrieval-Augmented Generation under ConflictsabstractRetrieval-Augmented Generation (RAG) has become a standard paradigm for grounding Large Language Models (LLMs) with external knowledge.However, RAG performance often degrades substantially when faced with noisy, outdated, or conflicting retrieved information.In this work, we empirically demonstrate that Prior-Guided Reasoning-a strategy that explicitly elicits the model's parametric knowledge as prior information to guide reasoning on retrieved documents-effectively mitigates the impact of external conflicts.Building on this, we propose BrPr (Bernoulligated reinforcement learning for Prior-Guided reasoning), a framework that achieves robust performance across varying degrees of external inconsistency.Furthermore, by employing a Bernoulli-gated dropout mechanism during training, BrPr distills the prior-driven reasoning capability into the model parameters, enabling efficient latent reasoning without explicit prior generation.The experimental results demonstrate that BrPr consistently exhibits superior robustness to external conflicts and noise. Xiaowei Yuan, Ziyang Huang 0005, Zhao Yang 0004, Yequan Wang, Jun Zhao 0001, Kang Liu 0001 |
ACL (1) | 1 |
| 2025 | Exploiting Contextual Knowledge in LLMs through V-usable Information based Layer EnhancementabstractXiaowei Yuan, Zhao Yang, Ziyang Huang, Yequan Wang, Siqi Fan, Yiming Ju, Jun Zhao, Kang Liu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xiaowei Yuan, Zhao Yang 0004, Ziyang Huang 0005, Yequan Wang, Siqi Fan 0001, Yiming Ju, Jun Zhao 0001, Kang Liu 0001 |
ACL (1) | 1 |
| 2024 | On the In-context Generation of Language ModelsabstractLarge language models (LLMs) are found to have the ability of in-context generation (ICG): when they are fed with an in-context prompt concatenating a few somehow similar examples, they can implicitly recognize the pattern of them and then complete the prompt in the same pattern.ICG is curious, since language models are usually not explicitly trained in the same way as the in-context prompt, and the distribution of examples in the prompt differs from that of sequences in the pretrained corpora.This paper provides a systematic study of the ICG ability of language models, covering discussions about its source and influential factors, in the view of both theory and empirical experiments.Concretely, we first propose a plausible latent variable model to model the distribution of the pretrained corpora, and then formalize ICG as a problem of next topic prediction.With this framework, we can prove that the repetition nature of a few topics ensures the ICG ability on them theoretically.Then, we use this controllable pretrained distribution to generate several medium-scale synthetic datasets (token scale: 2.1B~3.9B)and experiment with different settings of Transformer architectures (parameter scale: 4M~234M).Our experimental results further offer insights into how the data and model architectures influence ICG. Zhongtao Jiang, Yuanzhe Zhang, Xiaowei Yuan, Jun Zhao 0001, Kang Liu 0001 |
EMNLP | 4 |
| 2024 | Improving Zero-shot LLM Re-Ranker with Risk MinimizationabstractIn the Retrieval-Augmented Generation (RAG) system, advanced Large Language Models (LLMs) have emerged as effective Query Likelihood Models (QLMs) in an unsupervised way, which re-rank documents based on the probability of generating the query given the content of a document.However, directly prompting LLMs to approximate QLMs inherently is biased, where the estimated distribution might diverge from the actual document-specific distribution.In this study, we introduce a novel framework, UR 3 , which leverages Bayesian decision theory to both quantify and mitigate this estimation bias.Specifically, UR 3 reformulates the problem as maximizing the probability of document generation, thereby harmonizing the optimization of query and document generation probabilities under a unified risk minimization objective.Our empirical results indicate that UR 3 significantly enhances re-ranking, particularly in improving the Top-1 accuracy.It benefits the QA tasks by achieving higher accuracy with fewer input documents. Xiaowei Yuan, Zhao Yang 0004, Yequan Wang, Jun Zhao 0001, Kang Liu 0001 |
EMNLP | 1 |
| 2024 | Wireless Channel Key Generation Based on Multisubcarrier Phase DifferenceabstractWireless channel key generation technology is an important mechanism to guarantee the security of wireless network, but influenced by the key length and the actual electromagnetic environment, wireless channel key generation technology is faced with the challenge of high-key generation rate (KGR) and low-key disagreement rate (KDR). The existing key generation methods also lack the full use of the channel state information (CSI). We propose a key generation method based on multisubcarrier phase difference to expand the randomness source dimension, eliminate the phase bias, offset part of the noise influence, and set the threshold screening data to reduce the influence of measurement error. We further propose a key generation method based on resampling of kernel density estimation (KDE), which yields highly reciprocal randomness sources by resampling the results of KDE of phase difference values. To fill the metric gap of whether a method keeps low KDR while increasing the KGR, the evaluation metric of effective improvement ratio (EIR) is proposed. The two methods we proposed have a higher EIR than the method of using multiple-input and multiple-output (MIMO) and increasing the quantization level, achieving the goal of increasing the KGR while maintaining the low KDR. The KGR can reach about 12146 bits/s, and the KDR is 1.83%. The keys obtained by both methods can effectively prevent passive eavesdropping and meet the randomness requirements. Xiaowei Yuan, Yu Jiang 0020, Guyue Li, Aiqun Hu |
IEEE Internet Things J. | 1 |
| 2024 | Contrastive Language-knowledge Graph Pre-trainingabstractRecent years have witnessed a surge of academic interest in knowledge-enhanced pre-trained language models (PLMs) that incorporate factual knowledge to enhance knowledge-driven applications. Nevertheless, existing studies primarily focus on shallow, static, and separately pre-trained entity embeddings, with few delving into the potential of deep contextualized knowledge representation for knowledge incorporation. Consequently, the performance gains of such models remain limited. In this article, we introduce a simple yet effective knowledge-enhanced model, College ( Co ntrastive L anguage-Know le dge G raph Pr e -training), which leverages contrastive learning to incorporate factual knowledge into PLMs. This approach maintains the knowledge in its original graph structure to provide the most available information and circumvents the issue of heterogeneous embedding fusion. Experimental results demonstrate that our approach achieves more effective results on several knowledge-intensive tasks compared to previous state-of-the-art methods. Our code and trained models are available at https://github.com/Stacy027/COLLEGE . Xiaowei Yuan, Kang Liu 0001, Yequan Wang |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2023 | Learning Representation for Anomaly Detection of Vehicle TrajectoriesabstractPredicting the future trajectories of surrounding vehicles based on their history trajectories is a critical task in autonomous driving. However, when small crafted perturbations are introduced to those history trajectories, the resulting anomalous (or adversarial) trajectories can significantly mislead the future trajectory prediction module of the ego vehicle, which may result in unsafe planning and even fatal accidents. Therefore, it is of great importance to detect such anomalous trajectories of the surrounding vehicles for system safety, but few works have addressed this issue. In this work, we propose two novel methods for learning effective and efficient representations for online anomaly detection of vehicle trajectories. Different from general time-series anomaly detection, anomalous vehicle trajectory detection deals with much richer contexts on the road and fewer observable patterns on the anomalous trajectories themselves. To address these challenges, our methods exploit contrastive learning techniques and trajectory semantics to capture the patterns underlying the driving scenarios for effective anomaly detection under supervised and unsupervised settings, respectively. We conduct extensive experiments to demonstrate that our supervised method based on contrastive learning and unsupervised method based on reconstruction with semantic latent space can significantly improve the performance of anomalous trajectory detection in their corresponding settings over various baseline methods. We also demonstrate our methods' generalization ability to detect unseen patterns of anomalies. Ruochen Jiao, Juyang Bai, Xiangguo Liu, Takami Sato, Xiaowei Yuan, Qi Alfred Chen, Qi Zhu 0002 |
IROS | 5 |
| 2023 | Attribute-based anonymous credential: Delegation, traceability, and revocation
Peng Li 0059, Junzuo Lai, Wei Wu 0001, Xiaowei Yuan |
Comput. Networks | 7 |
| 2022 | Design of an Autoencoder-based Anomaly Detection for the DoH traffic SystemabstractDNS has encountered complex and diversified attacks over the years due to its special status on the Internet. The concept of DNS-over-HTTPS (DoH) has been proposed to protect user privacy by encapsulating DNS into HTTPS, which increases the difficulty of DNS tunnel detection but also faces some new attacks. In recent years, many researchers have discussed the detection methods of DoH tunnel. However, most of them need large-scale labeled datasets and extract statistical features, which is time-consuming and costs immense manpower, so it is impractical to be used in the real-world. In this paper, we developed a system called AADDS: an Autoencoder-based Anomaly Detection for the DoH traffic System consists of Traffic Capture module and Anomaly Detection module. The Traffic Capture module is developed based on nff-go, which can collect features stably in a high-speed Ethernet environment and greatly reduce the workload. For the Anomaly Detection module, we used bidirectional Long and Short-Term Memory (Bi-LSTM) to build an autoencoder network. Several essential experiments proved that our method has fewer parameters while ensuring higher accuracy, and it outperforms the state-of-the-art methods. Xinhui Du, Dongxin Liu, Zhongji Liu, Xiaowei Yuan, Tong Li 0012, Haojiang Deng |
CSCWD | 5 |
| 2022 | Pay attention to emoji: Feature Fusion Network with EmoGraph2vec Model for Sentiment AnalysisabstractWith the explosive growth of social media, opinionated postings with emojis have increased explosively. Many emojis are used to express emotions, attitudes, and opinions. Emoji representation learning can be helpful to improve the performance of emoji-related natural language processing tasks, especially in text sentiment analysis. However, most studies have only utilized the fixed descriptions provided by the Unicode Consortium without consideration of actual usage scenarios. As for the sentiment analysis task, many researchers ignore the emotional impact of the interaction between text and emojis. It results that the emotional semantics of emojis cannot be fully explored. In this work, we propose a method called EmoGraph2vec to learn emoji representations by constructing a co-occurrence graph network from social data and enriching the semantic information based on an external knowledge base EmojiNet to embed emoji nodes. Based on EmoGraph2vec model, we design a novel neural network to incorporate text and emoji information into sentiment analysis, which uses a hybrid-attention module combined with TextCNN-based classifier to improve performance. Experimental results show that the proposed model can outperform several baselines for sentiment analysis on benchmark datasets. Additionally, we conduct a series of ablation and comparison experiments to investigate the effectiveness and interpretability of our model. Xiaowei Yuan, Xiaodan Zhang 0004, Honglei Lv |
ICPR | 1 |
| 2022 | Practical Federated Learning for Samples with Different IDs
Junzuo Lai, Xiaowei Yuan, Beibei Song |
ProvSec | 3 |
| 2021 | Emoji-Based Co-Attention Network for Microblog Sentiment Analysis
Xiaowei Yuan, Honglei Lv |
ICONIP (5) | 1 |
| 2019 | Face De-Occlusion Using 3D Morphable Model and Generative Adversarial NetworkabstractIn recent decades, 3D morphable model (3DMM) has been commonly used in image-based photorealistic 3D face reconstruction. However, face images are often corrupted by serious occlusion by non-face objects including eyeglasses, masks, and hands. Such objects block the correct capture of landmarks and shading information. Therefore, the reconstructed 3D face model is hardly reusable. In this paper, a novel method is proposed to restore de-occluded face images based on inverse use of 3DMM and generative adversarial network. We utilize the 3DMM prior to the proposed adversarial network and combine a global and local adversarial convolutional neural network to learn face de-occlusion model. The 3DMM serves not only as geometric prior but also proposes the face region for the local discriminator. Experiment results confirm the effectiveness and robustness of the proposed algorithm in removing challenging types of occlusions with various head poses and illumination. Furthermore, the proposed method reconstructs the correct 3D face model with de-occluded textures. Xiaowei Yuan, In Kyu Park |
ICCV | 1 |
| 2000 | How to Build Up an Infrastructure for Intercultural Usability EngineeringabstractSiemens is a global enterprise that sells its products in more than 190 countries throughout the world. This internationalization of products means more than just translating the operating instructions and making changes to formats. True adaptation goes much deeper and takes into account different requirements in terms of functionality. This article starts by defining the term culture and then considers the requirements that have to be met in the context of intercultural usability engineering. Taking the establishment of usability laboratories in Beijing, China and Princeton, New Jersey as examples, the article then presents the challenges facing international cooperation and possible solutions, taking a detailed look at differences in the infrastructure, at key qualifications in international cooperation, and at the development of appropriate test methods. Andreas Beu, Pia Honold, Xiaowei Yuan |
Int. J. Hum. Comput. Interact. | 3 |