VLDB 2026 Research / reviewers in the wild / expert
Xiao Zhou 0004
dblp:267/2864-4
· DBLP profile ↗
6ranked-venue papers
0as first author
4since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Language models and text generation · 30% Efficient and distributed learning · 26% Question answering and dialogue systems · 17% |
Topics — the 11 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning › inference efficiency
context compression |
1.0 | 1 | 2026 | SADA: Bridging In-Context Learning and Fine-Tuning via State-Aligned Distillation Adapters · ACL (1) 2026 |
Natural language and speech › Language models and text generation
in-context learning |
1.0 | 1 | 2026 | SADA: Bridging In-Context Learning and Fine-Tuning via State-Aligned Distillation Adapters · ACL (1) 2026 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
1.0 | 1 | 2026 | SADA: Bridging In-Context Learning and Fine-Tuning via State-Aligned Distillation Adapters · ACL (1) 2026 |
Natural language and speech › Question answering and dialogue systems
dialogue generation |
0.9 | 2 | 2020 | MovieChats: Chat like Humans in a Closed Domain · EMNLP (1) 2020 Diversifying Dialogue Generation with Non-Conversational Text · ACL 2020 |
Machine learning › Trustworthy machine learning › robustness
adversarial robustness |
0.6 | 1 | 2022 | RoCBert: Robust Chinese Bert with Multimodal Contrastive Pretraining · ACL (1) 2022 |
Natural language and speech › Language models and text generation
pre-trained language model |
0.6 | 1 | 2022 | RoCBert: Robust Chinese Bert with Multimodal Contrastive Pretraining · ACL (1) 2022 |
Machine learning › Trustworthy machine learning
robustness |
0.6 | 1 | 2022 | RoCBert: Robust Chinese Bert with Multimodal Contrastive Pretraining · ACL (1) 2022 |
Natural language and speech › Language models and text generation › large language model training
robust pretraining |
0.6 | 1 | 2022 | RoCBert: Robust Chinese Bert with Multimodal Contrastive Pretraining · ACL (1) 2022 |
Natural language and speech › Question answering and dialogue systems › dialogue generation › dialogue response generation
response diversity |
0.4 | 1 | 2020 | Diversifying Dialogue Generation with Non-Conversational Text · ACL 2020 |
Machine learning › Representation and self-supervised learning › pre-training
multimodal pretraining |
0.2 | 1 | 2022 | RoCBert: Robust Chinese Bert with Multimodal Contrastive Pretraining · ACL (1) 2022 |
Natural language and speech › Information extraction and text analysis › information retrieval
knowledge retrieval |
0.1 | 1 | 2020 | MovieChats: Chat like Humans in a Closed Domain · EMNLP (1) 2020 |
Methods — techniques the papers use, named apart from their topics
knowledge distillation · 1.0in-context learning · 1.0adapter · 1.0multimodal pretraining · 0.6contrastive learning · 0.6seq2seq · 0.4pre-training · 0.4neural approaches · 0.4iterative back-translation · 0.4fine-tuning · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SADA: Bridging In-Context Learning and Fine-Tuning via State-Aligned Distillation AdaptersabstractPrompt-based in-context learning (ICL) and parameter fine-tuning are two dominant paradigms for incorporating external information into large language models (LLMs), but they incur high inference costs or require expensive retraining.To bridge this gap, context-to-parameter mapping converts prompts into temporary adapter weights.However, we identify a critical failure mode in existing methods: hiddenstate collapse, where the adapter-augmented model's internal states diverge sharply from the full-context oracle in deeper layers.We trace this failure to two coupled gaps: suboptimal Input-Selection and inadequate Supervision-Signal.To address these issues, we propose SADA (State-Aligned Distillation Adapters).We establish the attention-block output as a principled feature interface to improve input selection and introduce statealignment distillation to enforce consistency between the adapter-augmented model and the full-context oracle.Experiments on long-context language modeling (PG19) and downstream NLU and summarization benchmarks show that SADA consistently outperforms strong baselines like StreamAdapter and GenerativeAdapter, achieving performance comparable to ICL while significantly reducing memory footprint and latency.We further analyze when parameterized context compression is effective and when explicit context retention remains preferable. Tianlong Wang, Linhao Zhang, Aiwei Liu, Xiao Zhou 0004 |
ACL (1) | 7 |
| 2025 | RadIR: A Scalable Framework for Multi-grained Medical Image Retrieval via Radiology Report Mining
Chaoyi Wu, Xiao Zhou 0004, Ya Zhang 0002, Yanfeng Wang 0001, Weidi Xie |
MICCAI (5) | 4 |
| 2023 | Personality Understanding of Fictional Characters during Book ReadingabstractMo Yu, Jiangnan Li, Shunyu Yao, Wenjie Pang, Xiaochen Zhou, Zhou Xiao, Fandong Meng, Jie Zhou. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Mo Yu, Wenjie Pang, Xiaochen Zhou, Xiao Zhou 0004, Fandong Meng, Jie Zhou 0016 |
ACL (1) | 6 |
| 2022 | RoCBert: Robust Chinese Bert with Multimodal Contrastive PretrainingabstractLarge-scale pretrained language models have achieved SOTA results on NLP tasks.However, they have been shown vulnerable to adversarial attacks especially for logographic languages like Chinese.In this work, we propose ROCBERT: a pretrained Chinese Bert that is robust to various forms of adversarial attacks like word perturbation, synonyms, typos, etc.It is pretrained with the contrastive learning objective which maximizes the label consistency under different synthesized adversarial examples.The model takes as input multimodal information including the semantic, phonetic and visual features.We show all these features are important to the model robustness since the attack can be performed in all the three forms.Across 5 Chinese NLU tasks, ROCBERT outperforms strong baselines under three blackbox adversarial algorithms without sacrificing the performance on clean testset.It also performs the best in the toxic content detection task under human-made attacks. * Equal contribution. Hui Su, Xiaoyu Shen 0001, Xiao Zhou 0004, Tuo Ji, Jiarui Fang, Jie Zhou 0016 |
ACL (1) | 4 |
| 2020 | Diversifying Dialogue Generation with Non-Conversational TextabstractNeural network-based sequence-to-sequence (seq2seq) models strongly suffer from the lowdiversity problem when it comes to opendomain dialogue generation.As bland and generic utterances usually dominate the frequency distribution in our daily chitchat, avoiding them to generate more interesting responses requires complex data filtering, sampling techniques or modifying the training objective.In this paper, we propose a new perspective to diversify dialogue generation by leveraging non-conversational text.Compared with bilateral conversations, nonconversational text are easier to obtain, more diverse and cover a much broader range of topics.We collect a large-scale nonconversational corpus from multi sources including forum comments, idioms and book snippets.We further present a training paradigm to effectively incorporate these text via iterative back translation.The resulting model is tested on two conversational datasets and is shown to produce significantly more diverse responses without sacrificing the relevance with context. Hui Su, Xiaoyu Shen 0001, Sanqiang Zhao, Xiao Zhou 0004, Pengwei Hu 0001, Randy Zhong, Cheng Niu, Jie Zhou 0016 |
ACL | 4 |
| 2020 | MovieChats: Chat like Humans in a Closed DomainabstractBeing able to perform in-depth chat with humans in a closed domain is a precondition before an open-domain chatbot can ever be claimed.In this work, we take a close look at the movie domain and present a large-scale high-quality corpus with fine-grained annotations in hope of pushing the limit of moviedomain chatbots.We propose a unified, readily scalable neural approach which reconciles all subtasks like intent prediction and knowledge retrieval.The model is first pretrained on the huge general-domain data, then finetuned on our corpus.We show this simple neural approach trained on high-quality data is able to outperform commercial systems replying on complex rules.On both the static and interactive tests, we find responses generated by our system exhibits remarkably good engagement and sensibleness close to human-written ones.We further analyze the limits of our work and point out potential directions for future work 1 . Hui Su, Xiaoyu Shen 0001, Xiao Zhou 0004, Ernie Chang, Cheng Niu, Jie Zhou 0016 |
EMNLP (1) | 3 |