EDBT 2026 Demo / reviewers in the wild / expert
Xiao Xu 0005
dblp:64/4216-5
· DBLP profile ↗
17ranked-venue papers
4as first author
16since 2021 · last 2025
0000-0001-5774-1094ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 3 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Can Large Language Models Understand You Better? An MBTI Personality Detection Dataset Aligned with Population TraitsabstractThe Myers-Briggs Type Indicator (MBTI) is one of the most influential personality theories reflecting individual differences in thinking, feeling, and behaving. MBTI personality detection has garnered considerable research interest and has evolved significantly over the years. However, this task tends to be overly optimistic, as it currently does not align well with the natural distribution of population personality traits. Specifically, the self-reported labels in existing datasets result in data quality issues and the hard labels fail to capture the full range of population personality distributions. In this paper, we identify the task by constructing MBTIBench, the first manually annotated MBTI personality detection dataset with soft labels, under the guidance of psychologists. Our experimental results confirm that soft labels can provide more benefits to other psychological tasks than hard labels. We highlight the polarized predictions and biases in LLMs as key directions for future research. Bohan Li 0010, Jiannan Guan, Longxu Dou, Yunlong Feng, Dingzirui Wang, Yang Xu 0049, Enbo Wang, Qiguang Chen, Bichen Wang, Xiao Xu 0005, Libo Qin 0001, Qingfu Zhu, Wanxiang Che |
COLING | 10 |
| 2025 | Chart2Code53: A Large-Scale Diverse and Complex Dataset for Enhancing Chart-to-Code GenerationabstractChart2Code has recently received significant attention in the multimodal community due to its potential to reduce the burden of visualization and promote a more detailed understanding of charts.However, existing Chart2Coderelated training datasets suffer from at least one of the following issues: (1) limited scale, (2) limited type coverage, and ( 3) inadequate complexity.To address these challenges, we seek more diverse sources that better align with real-world user distributions and propose dual data synthesis pipelines: (1) Synthesize based on online plotting code.(2) Synthesize based on the chart images in the academic paper.We create a large-scale Chart2Code training dataset Chart2Code53, including 53 chart types, 130K Chart-code pairs based on the pipeline.Experimental results demonstrate that even with few parameters, the model finetuned on Chart2Code53 achieves state-ofthe-art performance on multiple Chart2Code benchmarks within open-source models 1 . Tianhao Niu, Yiming Cui 0001, Baoxin Wang, Xiao Xu 0005, Qingfu Zhu, Dayong Wu, Shijin Wang 0001, Wanxiang Che |
EMNLP | 4 |
| 2025 | Improving Consistency Identification in Task-oriented Dialogue Through Multi-Agent CollaborationabstractConsistency identification in task-oriented dialog (CI-ToD) typically consists of three sub-tasks: User Query Inconsistency (QI) identification, Dialogue History Inconsistency (HI) identification, and Knowledge Base Inconsistency (KBI) identification, which aim to determine inconsistent relationships between system response and user query, dialogue history, and knowledge base. Previous approaches focus on the exploration of deep learning models for CI-ToD. While these models achieve remarkable progress, they still rely on large amounts of labeled data, which is hard to achieve in real-world scenarios. Motivated by this, in the paper, we aim to explore large language models for CI-ToD, which do not require any training data. In addition, we further introduce a multi-agent collaboration framework (MAC-CIToD) to model the interaction across three sub-tasks in CI-ToD, including (1) Full Connection paradigm, (2) Cycle Connection paradigm, and (3) Central Connection paradigm, which effectively builds interaction across QI, HI, and KBI. Experiments on the standard benchmark reveal that our framework achieves superior performance. Additionally, we compare MAC-CIToD with the most advanced trained approaches and find that its zero-shot performance on most metrics even surpasses that of models after training on the CI-ToD dataset. Ruoxi Zhou, Qiguang Chen, Xiao Xu 0005, Hao Fei 0003, Dagang Li 0001, Wanxiang Che, Libo Qin 0001 |
IJCAI | 5 |
| 2025 | Visual Thoughts: A Unified Perspective of Understanding Multimodal Chain-of-ThoughtabstractLarge Vision-Language Models (LVLMs) have achieved significant success in multimodal tasks, with multimodal chain-of-thought (MCoT) further enhancing performance and interpretability. Recent MCoT methods fall into two categories: (i) Textual-MCoT (T-MCoT), which takes multimodal input and produces textual output; and (ii) Interleaved-MCoT (I-MCoT), which generates interleaved image-text outputs. Despite advances in both approaches, the mechanisms driving these improvements are not fully understood. To fill this gap, we first reveal that MCoT boosts LVLMs by incorporating $\textit{visual thoughts}$, which convey image information to the reasoning process regardless of the MCoT format, depending only on clarity and conciseness of expression. Furthermore, to explore visual thoughts systematically, we define four distinct forms of visual thought expressions and analyze them comprehensively. Our findings demonstrate that these forms differ in clarity and conciseness, yielding varying levels of MCoT improvement. Additionally, we explore the internal nature of visual thoughts, finding that visual thoughts serve as intermediaries between the input image and reasoning to deeper transformer layers, enabling more advanced visual information transmission. We hope that the visual thoughts can inspire further breakthroughs for future MCoT research. Zihui Cheng, Qiguang Chen, Xiao Xu 0005, Jiaqi Wang 0012, Weiyun Wang, Hao Fei 0003, Yidong Wang 0003, Alex Jinpeng Wang, Zhi Chen 0006, Wanxiang Che, Libo Qin 0001 |
NeurIPS | 3 |
| 2025 | Manager: Aggregating Insights From Unimodal Experts in Two-Tower VLMs and MLLMsabstractTwo-Tower Vision–Language Models (VLMs) have demonstrated strong performance across various downstream VL tasks. While BridgeTower further enhances performance by building bridges between encoders, it(i)suffers from ineffective layer-by-layer utilization of unimodal representations,(ii)restricts the flexible exploitation of different levels of unimodal semantic knowledge, and(iii)is limited to the evaluation on traditional low-resolution datasets only with the Two-Tower VLM architecture. In this work, we propose Manager, a lightweight, efficient and effective plugin that adaptively aggregates insights from different levels of pre-trained unimodal experts to facilitate more comprehensive VL alignment and fusion. First, under the Two-Tower VLM architecture, we introduce ManagerTower, a novel VLM that introduces the manager in each cross-modal layer. Whether with or without VL pre-training, ManagerTower outperforms previous strong baselines and achieves superior performance on 4 downstream VL tasks. Moreover, we extend our exploration to the latest Multimodal Large Language Model (MLLM) architecture.We demonstrate that LLaVA-OV-Manager significantly boosts the zero-shot performance of LLaVA-OV across different categories of capabilities, images, and resolutions on 20 downstream datasets, whether the multi-grid algorithm is enabled or not. In-depth analysis reveals that both our manager and the multi-grid algorithm can be viewed as a plugin that improves the visual representation by capturing more diverse visual details from two orthogonal perspectives (depth and width). Their synergy can mitigate the semantic ambiguity caused by the multi-grid algorithm and further improve performance. Code and models are available at https://github.com/LooperXX/ManagerTower. Xiao Xu 0005, Libo Qin 0001, Wanxiang Che, Min-Yen Kan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Semantic-Guided Generative Image Augmentation Method with Diffusion Models for Image ClassificationabstractExisting image augmentation methods consist of two categories: perturbation-based methods and generative methods. Perturbation-based methods apply pre-defined perturbations to augment an original image, but only locally vary the image, thus lacking image diversity. In contrast, generative methods bring more image diversity in the augmented images but may not preserve semantic consistency, thus may incorrectly change the essential semantics of the original image. To balance image diversity and semantic consistency in augmented images, we propose SGID, a Semantic-guided Generative Image augmentation method with Diffusion models for image classification. Specifically, SGID employs diffusion models to generate augmented images with good image diversity. More importantly, SGID takes image labels and captions as guidance to maintain semantic consistency between the augmented and original images. Experimental results show that SGID outperforms the best augmentation baseline by 1.72% on ResNet-50 (from scratch), 0.33% on ViT (ImageNet-21k), and 0.14% on CLIP-ViT (LAION-2B). Moreover, SGID can be combined with other image augmentation baselines and further improves the overall performance. We demonstrate the semantic consistency and image diversity of SGID through quantitative human and automated evaluations, as well as qualitative case studies. Bohan Li 0010, Xiao Xu 0005, Yutai Hou, Yunlong Feng, Xuanliang Zhang, Qingfu Zhu, Wanxiang Che |
AAAI | 2 |
| 2024 | M³CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-ThoughtabstractMulti-modal Chain-of-Thought (MCoT) requires models to leverage knowledge from both textual and visual modalities for step-bystep reasoning, which gains increasing attention.Nevertheless, the current MCoT benchmark still faces some challenges: (1) absence of visual modal reasoning, (2) single-step visual modal reasoning, and (3) Domain missing, thereby hindering the development of MCoT.Motivated by this, we introduce a novel benchmark (M 3 CoT) to address the above challenges, advancing the multi-domain, multi-step, and multi-modal CoT.Additionally, we conduct a thorough evaluation involving abundant MCoT approaches on Vision Large Language Models (VLLMs).In addition, we highlight that the current VLLMs still struggle to correctly reason in M 3 CoT and there remains a large gap between existing VLLMs and human performance in M 3 CoT, despite their superior results on previous MCoT benchmarks.To our knowledge, we take the first meaningful step toward the multi-domain, multi-step, and multi-modal scenario in MCoT.We hope that M 3 CoT can serve as a valuable resource, providing a pioneering foundation in multi-domain, multi-step, multi-modal chain-of-thought research. Qiguang Chen, Libo Qin 0001, Zhi Chen 0006, Xiao Xu 0005, Wanxiang Che |
ACL (1) | 5 |
| 2024 | A Two-Stage Framework with Self-Supervised Distillation for Cross-Domain Text ClassificationabstractCross-domain text classification is a crucial task as it enables models to adapt to a target domain that lacks labeled data. It leverages or reuses rich labeled data from the different but related source domain(s) and unlabeled data from the target domain. To this end, previous work focuses on either extracting domain-invariant features or task-agnostic features, ignoring domain-aware features that may be present in the target domain and could be useful for the downstream task. In this paper, we propose a two-stage framework for cross-domain text classification. In the first stage, we finetune the model with mask language modeling (MLM) and labeled data from the source domain. In the second stage, we further fine-tune the model with self-supervised distillation (SSD) and unlabeled data from the target domain. We evaluate its performance on a public cross-domain text classification benchmark and the experiment results show that our method achieves new state-of-the-art results for both single-source domain adaptations (94.17% +1.03%) and multi-source domain adaptations (95.09% +1.34%). Yunlong Feng, Bohan Li 0010, Libo Qin 0001, Xiao Xu 0005, Wanxiang Che |
LREC/COLING | 4 |
| 2024 | Pro-HAN: A Heterogeneous Graph Attention Network for Profile-based Spoken Language UnderstandingabstractRecently, Profile-based Spoken Language Understanding (SLU) has gained increasing attention, which aims to incorporate various types of supplementary profile information (i.e., Knowledge Graph, User Profile, Context Awareness) to eliminate the prevalent ambiguities in user utterances. However, existing approaches can only separately model different profile information, without considering their interrelationships or excluding irrelevant and conflicting information within them. To address the above issues, we introduce a Heterogeneous Graph Attention Network to perform reasoning across multiple Profile information, called Pro-HAN. Specifically, we design three types of edges, denoted as intra-Pro, inter-Pro, and utterance-Pro, to capture interrelationships among multiple Pros. We establish a new state-of-the-art on the ProSLU dataset, with an improvement of approximately 8% across all three metrics. Further analysis experiments also confirm the effectiveness of our method in modeling multi-source profile information. Dechuan Teng, Chunlin Lu, Xiao Xu 0005, Wanxiang Che, Libo Qin 0001 |
ICASSP | 3 |
| 2023 | BridgeTower: Building Bridges between Encoders in Vision-Language Representation LearningabstractVision-Language (VL) models with the Two-Tower architecture have dominated visual-language representation learning in recent years. Current VL models either use lightweight uni-modal encoders and learn to extract, align and fuse both modalities simultaneously in a deep cross-modal encoder, or feed the last-layer uni-modal representations from the deep pre-trained uni-modal encoders into the top cross-modal encoder. Both approaches potentially restrict vision-language representation learning and limit model performance. In this paper, we propose BridgeTower, which introduces multiple bridge layers that build a connection between the top layers of uni-modal encoders and each layer of the cross-modal encoder. This enables effective bottom-up cross-modal alignment and fusion between visual and textual representations of different semantic levels of pre-trained uni-modal encoders in the cross-modal encoder. Pre-trained with only 4M images, BridgeTower achieves state-of-the-art performance on various downstream vision-language tasks. In particular, on the VQAv2 test-std set, BridgeTower achieves an accuracy of 78.73%, outperforming the previous state-of-the-art model METER by 1.09% with the same pre-training data and almost negligible additional parameters and computational costs. Notably, when further scaling the model, BridgeTower achieves an accuracy of 81.15%, surpassing models that are pre-trained on orders-of-magnitude larger datasets. Code and checkpoints are available at https://github.com/microsoft/BridgeTower. Xiao Xu 0005, Chenfei Wu, Shachar Rosenman, Vasudev Lal, Wanxiang Che, Nan Duan 0001 |
AAAI | 1 |
| 2023 | ManagerTower: Aggregating the Insights of Uni-Modal Experts for Vision-Language Representation LearningabstractXiao Xu, Bei Li, Chenfei Wu, Shao-Yen Tseng, Anahita Bhiwandiwalla, Shachar Rosenman, Vasudev Lal, Wanxiang Che, Nan Duan. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Xiao Xu 0005, Chenfei Wu, Shao-Yen Tseng, Anahita Bhiwandiwalla, Shachar Rosenman, Vasudev Lal, Wanxiang Che, Nan Duan 0001 |
ACL (1) | 1 |
| 2023 | Modularized Pre-Training for End-to-End Task-Oriented DialogueabstractPre-training forend-to-endtask-orienteddialoguesystems (EToDs) is a challenging task due to its unique knowledge base query (accuracy) need and lack of sufficient training data (fluency). In this paper, we try to mitigate the above challenges by introducing a modularized pre-training framework for EToDs, which achieves to effectively improve both accuracy and fluency of EToDs through a pre-training paradigm. The core insight is a modular design by decomposing EToDs into ageneration (fluency)module and aknowledge-retriever (accuracy)module, which allows us to optimize each module by pre-training these two sub-modules with different well-designed pre-training tasks, respectively. In addition, such a modularized paradigm enables us to make full use of large amounts of KB-free dialogue corpus for the pre-traininggenerationmodule, which can alleviate the insufficient training problem. Furthermore, we introduce a newconsistency-guideddata augmentation (CGDA) strategy to cope with the data scarcity problem to better pre-train theknowledge-retrievermodule. Finally, we fine-tune the pre-trainedgenerationmodule andknowledge-retrievermodule jointly. Experimental results on three datasets show that our model achieve superior performance in terms of both fluency and accuracy. To our knowledge, this is the first work to explore modularized pre-training methods for EToDs. Libo Qin 0001, Xiao Xu 0005, Lehan Wang, Yue Zhang 0004, Wanxiang Che |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Text Is No More Enough! A Benchmark for Profile-Based Spoken Language UnderstandingabstractCurrent researches on spoken language understanding (SLU) heavily are limited to a simple setting: the plain text-based SLU that takes the user utterance as input and generates its corresponding semantic frames (e.g., intent and slots). Unfortunately, such a simple setting may fail to work in complex real-world scenarios when an utterance is semantically ambiguous, which cannot be achieved by the text-based SLU models. In this paper, we first introduce a new and important task, Profile-based Spoken Language Understanding (ProSLU), which requires the model that not only relies on the plain text but also the supporting profile information to predict the correct intents and slots. To this end, we further introduce a large-scale human-annotated Chinese dataset with over 5K utterances and their corresponding supporting profile information (Knowledge Graph (KG), User Profile (UP), Context Awareness (CA)). In addition, we evaluate several state-of-the-art baseline models and explore a multi-level knowledge adapter to effectively incorporate profile information. Experimental results reveal that all existing text-based SLU models fail to work when the utterances are semantically ambiguous and our proposed framework can effectively fuse the supporting information for sentence-level intent detection and token-level slot filling. Finally, we summarize key challenges and provide new points for future directions, which hopes to facilitate the research. Xiao Xu 0005, Libo Qin 0001, Kaiji Chen, Guoxing Wu, Linlin Li 0001, Wanxiang Che |
AAAI | 1 |
| 2022 | IPGAN: Generating Informative Item Pairs by Adversarial SamplingabstractNegative sampling plays an important role in ranking-based recommender models. However, most existing sampling methods cannot generate informative item pairs with positive and negative instances due to two limitations: 1) they merely treat observed items as positive instances, ignoring the existence of potential positive items (i.e., nonobserved items users may prefer) and the probability of observed but noisy items and 2) they fail to capture the relationship between positive and negative items during negative sampling, which may cause the unexpected selection of potential positive items. In this article, we introduce a dynamic sampling strategy to search informative item pairs. Specifically, we first sample a positive instance from all the items by leveraging the overall features of user's observed items. Then, we strategically select a negative instance by considering its correlation with the sampled positive one. Formally, we propose an item pair generative adversarial network named IPGAN, where our sampling strategy is realized in two generative models for positive and negative instances, respectively. In addition, IPGAN can also ensure that the sampled item pairs are informative relative to the ground truth by a discriminative model. What is more, we propose a batch-training approach to further enhance both user and item modeling by alleviating the special bias (noise) from different users. This approach can also significantly accelerate the process of model training compared with classical GAN method for recommendation. Experimental results on three real data sets show that our approach outperforms other state-of-the-art approaches in terms of recommendation accuracy. Guibing Guo, Bowei Chen 0004, Xiao Xu 0005, Xu Chen 0017, Zhenhua Dong, Xiuqiang He 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2021 | GL-GIN: Fast and Accurate Non-Autoregressive Model for Joint Multiple Intent Detection and Slot FillingabstractLibo Qin, Fuxuan Wei, Tianbao Xie, Xiao Xu, Wanxiang Che, Ting Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Libo Qin 0001, Fuxuan Wei, Tianbao Xie, Xiao Xu 0005, Wanxiang Che, Ting Liu 0001 |
ACL/IJCNLP (1) | 4 |
| 2021 | Don't be Contradicted with Anything! CI-ToD: Towards Benchmarking Consistency for Task-oriented Dialogue SystemabstractConsistency Identification has obtained remarkable success on open-domain dialogue, which can be used for preventing inconsistent response generation.However, in contrast to the rapid development in open-domain dialogue, few efforts have been made to the task-oriented dialogue direction.In this paper, we argue that consistency problem is more urgent in task-oriented domain.To facilitate the research, we introduce CI-ToD, a novel dataset for Consistency Identification in Taskoriented Dialog system.In addition, we not only annotate the single label to enable the model to judge whether the system response is contradictory, but also provide more finegrained labels (i.e., Dialogue History Inconsistency, User Query Inconsistency and Knowledge Base Inconsistency) to encourage model to know what inconsistent sources lead to it.Empirical results show that state-of-the-art methods only achieve 51.3%, which is far behind the human performance of 93.2%, indicating that there is ample room for improving consistency identification ability.Finally, we conduct exhaustive experiments and qualitative analysis to comprehend key challenges and provide guidance for future directions.All datasets and models are publicly available at https://github.com/yizhen20133868/CI-ToD. * Email corresponding. User: Give me directions to the closest grocery store.System: There is a whole foods 2 miles away and their address is 880_ames_ct.User: I need a route that avoids all heavy traffic.System: mandarin roots is located at 271 springer street. Libo Qin 0001, Tianbao Xie, Shijue Huang, Qiguang Chen, Xiao Xu 0005, Wanxiang Che |
EMNLP (1) | 5 |
| 2020 | Dynamic Fusion Network for Multi-Domain End-to-end Task-Oriented DialogabstractRecent studies have shown remarkable success in end-to-end task-oriented dialog system. However, most neural models rely on large training data, which are only available for a certain number of task domains, such as navigation and scheduling. This makes it difficult to scalable for a new domain with limited labeled data. However, there has been relatively little research on how to effectively use data from all domains to improve the performance of each domain and also unseen domains. To this end, we investigate methods that can make explicit use of domain knowledge and introduce a shared-private network to learn shared and specific knowledge. In addition, we propose a novel Dynamic Fusion Network (DF-Net) which automatically exploit the relevance between the target domain and each domain. Results show that our models outperforms existing methods on multi-domain dialogue, giving the state-of-the-art in the literature. Besides, with little training data, we show its transferability by outperforming prior best model by 13.9% on average. Libo Qin 0001, Xiao Xu 0005, Wanxiang Che, Yue Zhang 0004, Ting Liu 0001 |
ACL | 2 |