EDBT 2026 Demo / reviewers in the wild / expert
Xiaoze Jiang
dblp:254/1470
· DBLP profile ↗
8ranked-venue papers
4as first author
5since 2021 · last 2026
0000-0002-5463-7176ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Language models and text generation · 32% Question answering and dialogue systems · 31% Knowledge representation and reasoning · 14% | |
| Databases, data mining, and information retrieval
3 papers |
Information retrieval · 71% Data mining · 14% Recommender systems · 14% |
Topics — the 19 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Question answering and dialogue systems
visual dialog |
1.8 | 4 | 2021 | Learning Dual Encoding Model for Adaptive Visual Understanding in Visual Dialogue · IEEE Trans. Image Process. 2021 KBGN: Knowledge-Bridge Graph Network for Adaptive Vision-Text Reasoning in Visual Dialogue · ACM Multimedia 2020 DAM: Deliberation, Abandon and Memory Networks for Generating Detailed and Non-repetitive Responses in Visual Dialogue · IJCAI 2020 |
Information retrieval › retrieval models › neural retrieval
dense retrieval |
1.0 | 1 | 2026 | CroPS: Improving Dense Retrieval with Cross-Perspective Positive Samples in Short-Video Search · AAAI 2026 |
Data mining › predictive modeling › classification › structured classification
hierarchical labeling |
1.0 | 1 | 2026 | CroPS: Improving Dense Retrieval with Cross-Perspective Positive Samples in Short-Video Search · AAAI 2026 |
Information retrieval
ranking |
1.0 | 1 | 2026 | SID-Coord: Coordinating Semantic IDs for ID-based Ranking in Short-Video Search · SIGIR 2026 |
Information retrieval › retrieval models
retrieval model training |
1.0 | 1 | 2026 | CroPS: Improving Dense Retrieval with Cross-Perspective Positive Samples in Short-Video Search · AAAI 2026 |
Recommender systems › generative recommendation
semantic ID |
1.0 | 1 | 2026 | SID-Coord: Coordinating Semantic IDs for ID-based Ranking in Short-Video Search · SIGIR 2026 |
Natural language and speech › Language models and text generation › large language model safety
detoxification |
0.9 | 1 | 2025 | Personalized Query Auto-Completion for Long and Short-Term Interests with Adaptive Detoxification Generation · KDD (2) 2025 |
Information retrieval › query suggestion › query auto-completion
personalized query auto-completion |
0.9 | 1 | 2025 | Personalized Query Auto-Completion for Long and Short-Term Interests with Adaptive Detoxification Generation · KDD (2) 2025 |
Information retrieval › query suggestion
query auto-completion |
0.9 | 1 | 2025 | Personalized Query Auto-Completion for Long and Short-Term Interests with Adaptive Detoxification Generation · KDD (2) 2025 |
Natural language and speech › Language models and text generation › multilingual language models
multilingual pretrained language model |
0.6 | 1 | 2022 | XLM-K: Improving Cross-Lingual Language Model Pre-training with Multilingual Knowledge · AAAI 2022 |
Machine learning › Representation and self-supervised learning › representation learning
visual representation learning |
0.5 | 1 | 2021 | Learning Dual Encoding Model for Adaptive Visual Understanding in Visual Dialogue · IEEE Trans. Image Process. 2021 |
Natural language and speech › Language models and text generation
decoding |
0.4 | 1 | 2020 | DAM: Deliberation, Abandon and Memory Networks for Generating Detailed and Non-repetitive Responses in Visual Dialogue · IJCAI 2020 |
Natural language and speech › Question answering and dialogue systems › dialogue generation
dialogue response generation |
0.4 | 1 | 2020 | DAM: Deliberation, Abandon and Memory Networks for Generating Detailed and Non-repetitive Responses in Visual Dialogue · IJCAI 2020 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge graph reasoning |
0.4 | 1 | 2020 | KBGN: Knowledge-Bridge Graph Network for Adaptive Vision-Text Reasoning in Visual Dialogue · ACM Multimedia 2020 |
Computer vision › Vision and language
multimodal reasoning |
0.4 | 1 | 2020 | KBGN: Knowledge-Bridge Graph Network for Adaptive Vision-Text Reasoning in Visual Dialogue · ACM Multimedia 2020 |
Natural language and speech › Language models and text generation
text generation |
0.4 | 1 | 2020 | DAM: Deliberation, Abandon and Memory Networks for Generating Detailed and Non-repetitive Responses in Visual Dialogue · IJCAI 2020 |
Computer vision › Vision and language
visual question answering |
0.3 | 2 | 2021 | Learning Dual Encoding Model for Adaptive Visual Understanding in Visual Dialogue · IEEE Trans. Image Process. 2021 DualVD: An Adaptive Dual Encoding Model for Deep Visual Understanding in Visual Dialogue · AAAI 2020 |
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer |
0.2 | 1 | 2022 | XLM-K: Improving Cross-Lingual Language Model Pre-training with Multilingual Knowledge · AAAI 2022 |
Natural language and speech › Information extraction and text analysis
named entity recognition |
0.2 | 1 | 2022 | XLM-K: Improving Cross-Lingual Language Model Pre-training with Multilingual Knowledge · AAAI 2022 |
Methods — techniques the papers use, named apart from their topics
attention mechanism · 1.8reject preference optimization · 1.7hierarchical user representation · 1.7large language model knowledge synthesis · 1.0gating mechanism · 1.0contrastive learning · 1.0attention-based fusion · 1.0dual encoding · 0.9object entailment · 0.6masked entity prediction · 0.6memory network · 0.5recurrent neural network · 0.4graph neural network · 0.4feature selection · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CroPS: Improving Dense Retrieval with Cross-Perspective Positive Samples in Short-Video SearchabstractDense retrieval has become a foundational paradigm in modern search systems, especially on short-video platforms. However, most industrial systems adopt a self-reinforcing training pipeline that relies on historically exposed user interactions for supervision. This paradigm inevitably leads to a filter bubble effect, where potentially relevant but previously unseen content is excluded from the training signal, biasing the model toward narrow and conservative retrieval. In this paper, we present CroPS (Cross-Perspective Positive Samples), a novel retrieval data engine designed to alleviate this problem by introducing diverse and semantically meaningful positive examples from multiple perspectives. CroPS enhances training with positive signals derived from user query reformulation behavior (query-level), engagement data in recommendation streams (system-level), and world knowledge synthesized by large language models (knowledge-level). To effectively utilize these heterogeneous signals, we introduce a Hierarchical Label Assignment (HLA) strategy and a corresponding H-InfoNCE loss that together enable fine-grained, relevance-aware optimization. Extensive experiments conducted on Kuaishou Search, a large-scale commercial short-video search platform, demonstrate that CroPS significantly outperforms strong baselines both offline and in live A/B tests, achieving superior retrieval performance and reducing query reformulation rates. CroPS is now fully deployed in Kuaishou Search, serving hundreds of millions of users daily. Ao Xie, Quanzhi Zhu, Xiaoze Jiang, Zhiheng Qin, Enyun Yu |
AAAI | 4 |
| 2026 | SID-Coord: Coordinating Semantic IDs for ID-based Ranking in Short-Video SearchabstractLarge-scale short-video search ranking models are typically trained on sparse co-occurrence signals over hashed item identifiers (HIDs). While effective at memorizing frequent interactions, such ID-based models struggle to generalize to long-tailed items with limited exposure. This memorization–generalization trade-off remains a longstanding challenge in such industrial systems. We propose SID-Coord, a lightweight Semantic ID framework that incorporates discrete, trainable semantic IDs (SIDs) directly into ID-based ranking models. Instead of treating semantic signals as auxiliary dense features, SID-Coord represents semantics as structured identifiers and coordinates HID-based memorization with SID-based generalization within a unified modeling framework. To enable effective coordination, SID-Coord introduces three components: (1) an attention-based fusion module over hierarchical SIDs to capture multi-level semantics, (2) a target-aware HID–SID gating mechanism that adaptively balances memorization and generalization, and (3) a SID-driven interest alignment module that models the semantic similarity distribution between target items and user histories. SID-Coord can be integrated into existing production ranking systems without modifying the backbone model. Online A/B experiments in a real-world production environment show statistically significant improvements, with a +0.664% gain in long-play rate in search and a +0.369% increase in search playback duration. Shunyu Zhang, Xiaoze Jiang, Jingwei Zhuo |
SIGIR | 5 |
| 2025 | Personalized Query Auto-Completion for Long and Short-Term Interests with Adaptive Detoxification GenerationabstractQuery auto-completion (QAC) plays a crucial role in modern search systems.However, in real-world applications, there are two pressing challenges that still need to be addressed.First, there is a need for hierarchical personalized representations for users.Previous approaches have typically used users' search behavior as a single, overall representation, which proves inadequate in more nuanced generative scenarios.Additionally, query prefixes are typically short and may contain typos or sensitive information, increasing the likelihood of generating toxic content compared to traditional text generation tasks.Such toxic content can degrade user experience and lead to public relations issues.Therefore, the second critical challenge is detoxifying QAC systems.Recent efforts to mitigate toxicity have involved generating queries unrelated to the given prefix, leading this approach still negatively impacts user experience.To address these two limitations, we propose a novel model (LaD) that captures personalized information from both long-term and short-term interests, incorporating adaptive detoxification.In LaD, personalized information is captured hierarchically at both coarse-grained and fine-grained levels.This approach preserves as much personalized information as possible while enabling online generation within time constraints.To move a further step, we propose an online training method based on Reject Preference Optimization (RPO).By incorporating a special token [Reject] during both the training and inference processes, the model achieves adaptive detoxification.Consequently, the generated text presented to users is both non-toxic and relevant to the given prefix.We conduct comprehensive experiments on industrial-scale datasets and perform online A/B tests, delivering the largest single-experiment metric improvement in nearly two years of our product.Our model has been deployed on Kuaishou search, driving the primary traffic Xiaoze Jiang, Zhiheng Qin, Enyun Yu |
KDD (2) | 2 |
| 2022 | XLM-K: Improving Cross-Lingual Language Model Pre-training with Multilingual KnowledgeabstractCross-lingual pre-training has achieved great successes using monolingual and bilingual plain text corpora. However, most pre-trained models neglect multilingual knowledge, which is language agnostic but comprises abundant cross-lingual structure alignment. In this paper, we propose XLM-K, a cross-lingual language model incorporating multilingual knowledge in pre-training. XLM-K augments existing multilingual pre-training with two knowledge tasks, namely Masked Entity Prediction Task and Object Entailment Task. We evaluate XLM-K on MLQA, NER and XNLI. Experimental results clearly demonstrate significant improvements over existing multilingual language models. The results on MLQA and NER exhibit the superiority of XLM-K in knowledge related tasks. The success in XNLI shows a better cross-lingual transferability obtained in XLM-K. What is more, we provide a detailed probing analysis to confirm the desired knowledge captured in our pre-training regimen. The code is available at https://github.com/microsoft/Unicoder/tree/master/pretraining/xlmk. Xiaoze Jiang, Yaobo Liang, Weizhu Chen |
AAAI | 1 |
| 2021 | Learning Dual Encoding Model for Adaptive Visual Understanding in Visual DialogueabstractDifferent from Visual Question Answering task that requires to answer only one question about an image, Visual Dialogue task involves multiple rounds of dialogues which cover a broad range of visual content that could be related to any objects, relationships or high-level semantics. Thus one of the key challenges in Visual Dialogue task is to learn a more comprehensive and semantic-rich image representation that can adaptively attend to the visual content referred by variant questions. In this paper, we first propose a novel scheme to depict an image from both visual and semantic views. Specifically, the visual view aims to capture the appearance-level information in an image, including objects and their visual relationships, while the semantic view enables the agent to understand high-level visual semantics from the whole image to the local regions. Furthermore, on top of such dual-view image representations, we propose a Dual Encoding Visual Dialogue (DualVD) module, which is able to adaptively select question-relevant information from the visual and semantic views in a hierarchical mode. To demonstrate the effectiveness of DualVD, we propose two novel visual dialogue models by applying it to the Late Fusion framework and Memory Network framework. The proposed models achieve state-of-the-art results on three benchmark datasets. A critical advantage of the DualVD module lies in its interpretability. We can analyze which modality (visual or semantic) has more contribution in answering the current question by explicitly visualizing the gate values. It gives us insights in understanding of information selection mode in the Visual Dialogue task. The code is available at https://github.com/JXZe/Learning_DualVD. Jing Yu 0007, Xiaoze Jiang, Zengchang Qin, Weifeng Zhang 0002, Yue Hu 0002, Qi Wu 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | DualVD: An Adaptive Dual Encoding Model for Deep Visual Understanding in Visual DialogueabstractDifferent from Visual Question Answering task that requires to answer only one question about an image, Visual Dialogue involves multiple questions which cover a broad range of visual content that could be related to any objects, relationships or semantics. The key challenge in Visual Dialogue task is thus to learn a more comprehensive and semantic-rich image representation which may have adaptive attentions on the image for variant questions. In this research, we propose a novel model to depict an image from both visual and semantic perspectives. Specifically, the visual view helps capture the appearance-level information, including objects and their relationships, while the semantic view enables the agent to understand high-level visual semantics from the whole image to the local regions. Futhermore, on top of such multi-view image features, we propose a feature selection framework which is able to adaptively capture question-relevant information hierarchically in fine-grained level. The proposed method achieved state-of-the-art results on benchmark Visual Dialogue datasets. More importantly, we can tell which modality (visual or semantic) has more contribution in answering the current question by visualizing the gate values. It gives us insights in understanding of human cognition in Visual Dialogue. Xiaoze Jiang, Jing Yu 0007, Zengchang Qin, Yingying Zhuang, Yue Hu 0002, Qi Wu 0001 |
AAAI | 1 |
| 2020 | DAM: Deliberation, Abandon and Memory Networks for Generating Detailed and Non-repetitive Responses in Visual DialogueabstractVisual Dialogue task requires an agent to be engaged in a conversation with human about an image. The ability of generating detailed and non-repetitive responses is crucial for the agent to achieve human-like conversation. In this paper, we propose a novel generative decoding architecture to generate high-quality responses, which moves away from decoding the whole encoded semantics towards the design that advocates both transparency and flexibility. In this architecture, word generation is decomposed into a series of attention-based information selection steps, performed by the novel recurrent Deliberation, Abandon and Memory (DAM) module. Each DAM module performs an adaptive combination of the response-level semantics captured from the encoder and the word-level semantics specifically selected for generating each word. Therefore, the responses contain more detailed and non-repetitive descriptions while maintaining the semantic accuracy. Furthermore, DAM is flexible to cooperate with existing visual dialogue encoders and adaptive to the encoder structures by constraining the information selection mode in DAM. We apply DAM to three typical encoders and verify the performance on the VisDial v1.0 dataset. Experimental results show that the proposed models achieve new state-of-the-art performance with high-quality responses. The code is available at https://github.com/JXZe/DAM. Xiaoze Jiang, Jing Yu 0007, Yajing Sun, Zengchang Qin, Yue Hu 0002, Qi Wu 0001 |
IJCAI | 1 |
| 2020 | KBGN: Knowledge-Bridge Graph Network for Adaptive Vision-Text Reasoning in Visual DialogueabstractVisual dialogue is a challenging task that needs to extract implicit information from both visual (image) and textual (dialogue history) contexts. Classical approaches pay more attention to the integration of the current question, vision knowledge and text knowledge, despising the heterogeneous semantic gaps between the cross-modal information. In the meantime, the concatenation operation has become de-facto standard to the cross-modal information fusion, which has a limited ability in information retrieval. In this paper, we propose a novel Knowledge-Bridge Graph Network (KBGN) model by using graph to bridge the cross-modal semantic relations between vision and text knowledge in fine granularity, as well as retrieving required knowledge via an adaptive information selection mode. Moreover, the reasoning clues for visual dialogue can be clearly drawn from intra-modal entities and inter-modal bridges. Experimental results on VisDial v1.0 and VisDial-Q datasets demonstrate that our model outperforms existing models with state-of-the-art results. Xiaoze Jiang, Siyi Du, Zengchang Qin, Yajing Sun, Jing Yu 0007 |
ACM Multimedia | 1 |