Xiao Zhou 0004

dblp:267/2864-4 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
4since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Language models and text generation · 30% Efficient and distributed learning · 26% Question answering and dialogue systems · 17%

Topics — the 11 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › inference efficiency
context compression
1.012026
SADA: Bridging In-Context Learning and Fine-Tuning via State-Aligned Distillation Adapters · ACL (1) 2026
Natural language and speech › Language models and text generation
in-context learning
1.012026
SADA: Bridging In-Context Learning and Fine-Tuning via State-Aligned Distillation Adapters · ACL (1) 2026
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
1.012026
SADA: Bridging In-Context Learning and Fine-Tuning via State-Aligned Distillation Adapters · ACL (1) 2026
Natural language and speech › Question answering and dialogue systems
dialogue generation
0.922020
MovieChats: Chat like Humans in a Closed Domain · EMNLP (1) 2020
Diversifying Dialogue Generation with Non-Conversational Text · ACL 2020
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
0.612022
RoCBert: Robust Chinese Bert with Multimodal Contrastive Pretraining · ACL (1) 2022
Natural language and speech › Language models and text generation
pre-trained language model
0.612022
RoCBert: Robust Chinese Bert with Multimodal Contrastive Pretraining · ACL (1) 2022
Machine learning › Trustworthy machine learning
robustness
0.612022
RoCBert: Robust Chinese Bert with Multimodal Contrastive Pretraining · ACL (1) 2022
Natural language and speech › Language models and text generation › large language model training
robust pretraining
0.612022
RoCBert: Robust Chinese Bert with Multimodal Contrastive Pretraining · ACL (1) 2022
Natural language and speech › Question answering and dialogue systems › dialogue generation › dialogue response generation
response diversity
0.412020
Diversifying Dialogue Generation with Non-Conversational Text · ACL 2020
Machine learning › Representation and self-supervised learning › pre-training
multimodal pretraining
0.212022
RoCBert: Robust Chinese Bert with Multimodal Contrastive Pretraining · ACL (1) 2022
Natural language and speech › Information extraction and text analysis › information retrieval
knowledge retrieval
0.112020
MovieChats: Chat like Humans in a Closed Domain · EMNLP (1) 2020

Methods — techniques the papers use, named apart from their topics

knowledge distillation · 1.0in-context learning · 1.0adapter · 1.0multimodal pretraining · 0.6contrastive learning · 0.6seq2seq · 0.4pre-training · 0.4neural approaches · 0.4iterative back-translation · 0.4fine-tuning · 0.4
YearPublicationVenuePosition
2026 SADA: Bridging In-Context Learning and Fine-Tuning via State-Aligned Distillation Adapters
abstract
Prompt-based in-context learning (ICL) and parameter fine-tuning are two dominant paradigms for incorporating external information into large language models (LLMs), but they incur high inference costs or require expensive retraining.To bridge this gap, context-to-parameter mapping converts prompts into temporary adapter weights.However, we identify a critical failure mode in existing methods: hiddenstate collapse, where the adapter-augmented model's internal states diverge sharply from the full-context oracle in deeper layers.We trace this failure to two coupled gaps: suboptimal Input-Selection and inadequate Supervision-Signal.To address these issues, we propose SADA (State-Aligned Distillation Adapters).We establish the attention-block output as a principled feature interface to improve input selection and introduce statealignment distillation to enforce consistency between the adapter-augmented model and the full-context oracle.Experiments on long-context language modeling (PG19) and downstream NLU and summarization benchmarks show that SADA consistently outperforms strong baselines like StreamAdapter and GenerativeAdapter, achieving performance comparable to ICL while significantly reducing memory footprint and latency.We further analyze when parameterized context compression is effective and when explicit context retention remains preferable.
Tianlong Wang, Linhao Zhang, Aiwei Liu, Xiao Zhou 0004
ACL (1)7
2025 RadIR: A Scalable Framework for Multi-grained Medical Image Retrieval via Radiology Report Mining
Chaoyi Wu, Xiao Zhou 0004, Ya Zhang 0002, Yanfeng Wang 0001, Weidi Xie
MICCAI (5)4
2023 Personality Understanding of Fictional Characters during Book Reading
abstract
Mo Yu, Jiangnan Li, Shunyu Yao, Wenjie Pang, Xiaochen Zhou, Zhou Xiao, Fandong Meng, Jie Zhou. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Mo Yu, Wenjie Pang, Xiaochen Zhou, Xiao Zhou 0004, Fandong Meng, Jie Zhou 0016
ACL (1)6
2022 RoCBert: Robust Chinese Bert with Multimodal Contrastive Pretraining
abstract
Large-scale pretrained language models have achieved SOTA results on NLP tasks.However, they have been shown vulnerable to adversarial attacks especially for logographic languages like Chinese.In this work, we propose ROCBERT: a pretrained Chinese Bert that is robust to various forms of adversarial attacks like word perturbation, synonyms, typos, etc.It is pretrained with the contrastive learning objective which maximizes the label consistency under different synthesized adversarial examples.The model takes as input multimodal information including the semantic, phonetic and visual features.We show all these features are important to the model robustness since the attack can be performed in all the three forms.Across 5 Chinese NLU tasks, ROCBERT outperforms strong baselines under three blackbox adversarial algorithms without sacrificing the performance on clean testset.It also performs the best in the toxic content detection task under human-made attacks. * Equal contribution.
Hui Su, Xiaoyu Shen 0001, Xiao Zhou 0004, Tuo Ji, Jiarui Fang, Jie Zhou 0016
ACL (1)4
2020 Diversifying Dialogue Generation with Non-Conversational Text
abstract
Neural network-based sequence-to-sequence (seq2seq) models strongly suffer from the lowdiversity problem when it comes to opendomain dialogue generation.As bland and generic utterances usually dominate the frequency distribution in our daily chitchat, avoiding them to generate more interesting responses requires complex data filtering, sampling techniques or modifying the training objective.In this paper, we propose a new perspective to diversify dialogue generation by leveraging non-conversational text.Compared with bilateral conversations, nonconversational text are easier to obtain, more diverse and cover a much broader range of topics.We collect a large-scale nonconversational corpus from multi sources including forum comments, idioms and book snippets.We further present a training paradigm to effectively incorporate these text via iterative back translation.The resulting model is tested on two conversational datasets and is shown to produce significantly more diverse responses without sacrificing the relevance with context.
Hui Su, Xiaoyu Shen 0001, Sanqiang Zhao, Xiao Zhou 0004, Pengwei Hu 0001, Randy Zhong, Cheng Niu, Jie Zhou 0016
ACL4
2020 MovieChats: Chat like Humans in a Closed Domain
abstract
Being able to perform in-depth chat with humans in a closed domain is a precondition before an open-domain chatbot can ever be claimed.In this work, we take a close look at the movie domain and present a large-scale high-quality corpus with fine-grained annotations in hope of pushing the limit of moviedomain chatbots.We propose a unified, readily scalable neural approach which reconciles all subtasks like intent prediction and knowledge retrieval.The model is first pretrained on the huge general-domain data, then finetuned on our corpus.We show this simple neural approach trained on high-quality data is able to outperform commercial systems replying on complex rules.On both the static and interactive tests, we find responses generated by our system exhibits remarkably good engagement and sensibleness close to human-written ones.We further analyze the limits of our work and point out potential directions for future work 1 .
Hui Su, Xiaoyu Shen 0001, Xiao Zhou 0004, Ernie Chang, Cheng Niu, Jie Zhou 0016
EMNLP (1)3