EDBT 2026 Demo / reviewers in the wild / expert
Jiaqi Bai 0001
dblp:219/9777-1
· DBLP profile ↗
20ranked-venue papers
7as first author
20since 2021 · last 2025
0000-0001-8312-6992ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 7 first-author · 17 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Mitigating Hallucinations in Large Vision-Language Models by Adaptively Constraining Information FlowabstractLarge vision-language models show tremendous potential in understanding visual information through human languages. However, they are prone to suffer from object hallucination, i.e., the generated image descriptions contain objects that do not exist in the image. In this paper, we reveal that object hallucination can be attributed to overconfidence in irrelevant visual features when soft visual tokens map to the LLM's word embedding space. Specifically, by figuring out the semantic similarity between visual tokens and LLM's word embedding, we observe that the smoothness of similarity distribution strongly correlates with the emergence of object hallucinations. To mitigate hallucinations, we propose using the Variational Information Bottleneck (VIB) to alleviate overconfidence by introducing stochastic noise, facilitating the constraining of irrelevant information. Furthermore, we propose an entropy-based noise-controlling strategy to enable the injected noise to be adaptively constrained regarding the smoothness of the similarity distribution. We adapt the proposed AdaVIB across distinct model architectures. Experimental results demonstrate that the proposed AdaVIB mitigates object hallucinations by effectively alleviating the overconfidence in irrelevant visual features, with consistent improvements on two object hallucination benchmarks. Jiaqi Bai 0001, Hongcheng Guo, Zhongyuan Peng, Jian Yang 0030, Zhoujun Li 0001, Mohan Li, Zhihong Tian |
AAAI | 1 |
| 2025 | XCOT: Cross-lingual Instruction Tuning for Cross-lingual Chain-of-Thought ReasoningabstractChain-of-thought (CoT) has emerged as a powerful technique to elicit reasoning in large language models and improve a variety of downstream tasks. CoT mainly demonstrates excellent performance in English, but its usage in low-resource languages is constrained due to poor language generalization. To bridge the gap among different languages, we propose a cross-lingual instruction fine-tuning framework (xCoT) to transfer knowledge from high-resource languages to low-resource languages. Specifically, the multilingual instruction training data (xCoT-Instruct) is created to encourage the semantic alignment of multiple languages. We introduce cross-lingual in-context few-shot learning (xICL) to accelerate multilingual agreement in instruction tuning, where some fragments of source languages in examples are randomly substituted by their counterpart translations of target languages. During multilingual instruction tuning, we adopt the randomly online CoT strategy to enhance the multilingual reasoning ability of the large language model by first translating the query to another language and then answering in English. To further facilitate the language transfer, we leverage the high-resource CoT to supervise the training of low-resource languages with cross-lingual distillation. Experimental results demonstrate the superior performance of xCoT in reducing the gap among different languages, highlighting its potential to reduce the cross-lingual gap. Linzheng Chai, Jian Yang 0030, Tao Sun 0016, Hongcheng Guo, Xinnian Liang, Jiaqi Bai 0001, Tongliang Li, Qiyao Peng 0001, Zhoujun Li 0001 |
AAAI | 8 |
| 2025 | MAC-SQL: A Multi-Agent Collaborative Framework for Text-to-SQLabstractRecent LLM-based Text-to-SQL methods usually suffer from significant performance degradation on “huge” databases and complex user questions that require multi-step reasoning. Moreover, most existing methods neglect the crucial significance of LLMs utilizing external tools and model collaboration. To address these challenges, we introduce MAC-SQL, a novel LLM-based multi-agent collaborative framework. Our framework comprises a core decomposer agent for Text-to-SQL generation with few-shot chain-of-thought reasoning, accompanied by two auxiliary agents that utilize external tools or models to acquire smaller sub-databases and refine erroneous SQL queries. The decomposer agent collaborates with auxiliary agents, which are activated as needed and can be expanded to accommodate new features or tools for effective Text-to-SQL parsing. In our framework, We initially leverage GPT-4 as the strong backbone LLM for all agent tasks to determine the upper bound of our framework. We then fine-tune an open-sourced instruction-followed model, SQL-Llama, by leveraging Code Llama 7B, to accomplish all tasks as GPT-4 does. Experiments show that SQL-Llama achieves a comparable execution accuracy of 43.94, compared to the baseline accuracy of 46.35 for vanilla GPT-4. At the time of writing, MAC-SQL+GPT-4 achieves an execution accuracy of 59.59 when evaluated on the BIRD benchmark, establishing a new state-of-the-art (SOTA) on its holdout test set. Changyu Ren, Jian Yang 0030, Xinnian Liang, Jiaqi Bai 0001, Linzheng Chai, Qian-Wen Zhang, Xing Sun 0001, Zhoujun Li 0001 |
COLING | 5 |
| 2024 | LogFormer: A Pre-train and Tuning Pipeline for Log Anomaly DetectionabstractLog anomaly detection is a key component in the field of artificial intelligence for IT operations (AIOps). Considering log data of variant domains, retraining the whole network for unknown domains is inefficient in real industrial scenarios. However, previous deep models merely focused on extracting the semantics of log sequences in the same domain, leading to poor generalization on multi-domain logs. To alleviate this issue, we propose a unified Transformer-based framework for Log anomaly detection (LogFormer) to improve the generalization ability across different domains, where we establish a two-stage process including the pre-training and adapter-based tuning stage. Specifically, our model is first pre-trained on the source domain to obtain shared semantic knowledge of log data. Then, we transfer such knowledge to the target domain via shared parameters. Besides, the Log-Attention module is proposed to supplement the information ignored by the log-paring. The proposed method is evaluated on three public datasets and one real-world dataset. Experimental results on multiple benchmarks demonstrate the effectiveness of our LogFormer with fewer trainable parameters and lower training costs. Hongcheng Guo, Jian Yang 0030, Jiaqi Bai 0001, Boyang Wang 0006, Zhoujun Li 0001, Tieqiao Zheng, Bo Zhang 0096, Junran Peng |
AAAI | 4 |
| 2024 | Towards Real-world Scenario: Imbalanced New Intent DiscoveryabstractShun Zhang, Yan Chaoran, Jian Yang, Jiaheng Liu, Ying Mo, Jiaqi Bai, Tongliang Li, Zhoujun Li. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Chaoran Yan, Jian Yang 0030, Ying Mo, Jiaqi Bai 0001, Tongliang Li, Zhoujun Li 0001 |
ACL (1) | 6 |
| 2024 | m3P: Towards Multimodal Multilingual Translation with Multimodal PromptabstractMultilingual translation supports multiple translation directions by projecting all languages in a shared space, but the translation quality is undermined by the difference between languages in the text-only modality, especially when the number of languages is large. To bridge this gap, we introduce visual context as the universal language-independent representation to facilitate multilingual translation. In this paper, we propose a framework to leverage the multimodal prompt to guide the Multimodal Multilingual Neural Machine Translation (m3P), which aligns the representations of different languages with the same meaning and generates the conditional vision-language memory for translation. We construct a multilingual multimodal instruction dataset (InstrMulti102) to support 102 languages Our method aims to minimize the representation distance of different languages by regarding the image as a central language. Experimental results show that m3P outperforms previous text-only baselines and multilingual multimodal methods by a large margin. Furthermore, the probing experiments validate the effectiveness of our method in enhancing translation under the low-resource and massively multilingual scenario. Jian Yang 0030, Hongcheng Guo, Yuwei Yin, Jiaqi Bai 0001, Xinnian Liang, Linzheng Chai, Liqun Yang, Zhoujun Li 0001 |
LREC/COLING | 4 |
| 2024 | New Intent Discovery with Attracting and Dispersing PrototypeabstractNew Intent Discovery (NID) aims to recognize known and infer new intent categories with the help of limited labeled and large-scale unlabeled data. The task is addressed as a feature-clustering problem and recent studies augment instance representation. However, existing methods fail to capture cluster-friendly representations, since they show less capability to effectively control and coordinate within-cluster and between-cluster distances. Tailored to the NID problem, we propose a Robust and Adaptive Prototypical learning (RAP) framework for globally distinct decision boundaries for both known and new intent categories. Specifically, a robust prototypical attracting learning (RPAL) method is designed to compel instances to gravitate toward their corresponding prototype, achieving greater within-cluster compactness. To attain larger between-cluster separation, another adaptive prototypical dispersing learning (APDL) method is devised to maximize the between-cluster distance from the prototype-to-prototype perspective. Experimental results evaluated on three challenging benchmarks (CLINC, BANKING, and StackOverflow) of our method with better cluster-friendly representation demonstrate that RAP brings in substantial improvements over the current state-of-the-art methods (even large language model) by a large margin (average 5.5% improvement). Jian Yang 0030, Jiaqi Bai 0001, Chaoran Yan, Tongliang Li, Zhoujun Li 0001 |
LREC/COLING | 3 |
| 2024 | RoNID: New Intent Discovery with Generated-Reliable Labels and Cluster-friendly Representations
Chaoran Yan, Jian Yang 0030, Changyu Ren, Jiaqi Bai 0001, Tongliang Li, Zhoujun Li 0001 |
DASFAA (5) | 5 |
| 2024 | OWL: A Large Language Model for IT OperationsabstractWith the rapid advancement of IT operations, managing and analyzing large data volumes efficiently for practical applications has become increasingly critical. Natural Language Processing (NLP) techniques have demonstrated remarkable capabilities in various tasks, including named entity recognition, machine translation, and dialogue systems. Recently, Large Language Models (LLMs) have achieved significant improvements across various domain-specific areas. However, there is a noticeable gap in the development of specialized Large Language Models (LLMs) tailored for IT operations. In this paper, we introduce the OWL, a large language model trained on our constructed Owl-Instruct with a wide range of IT-related information. Specifically, limited by the maximum input length, we propose the \textbf{H}omogeneous \textbf{M}arkov \textbf{C}ontext \textbf{E}xtension method (HMCE). The mixture-of-adapter strategy is leveraged to improve the parameter-efficient tuning across different domains or tasks.
Further, we evaluate the performance of OWL on the Owl-Bench established by us and open IT-related benchmarks. OWL demonstrates superior performance results on IT tasks, which outperforms existing models by significant margins. Moreover, we hope that the findings of our work will provide more insights to revolutionize the techniques of IT operations with specialized LLMs. Hongcheng Guo, Jian Yang 0030, Liqun Yang, Linzheng Chai, Jiaqi Bai 0001, Junran Peng, Xiaorong Hu, Dongfeng Zhang, Xu Shi 0005, Tieqiao Zheng, Liangfan Zheng, Bo Zhang 0096, Ke Xu 0001, Zhoujun Li 0001 |
ICLR | 6 |
| 2024 | TiNID: A Transfer and Interpretable LLM-Enhanced Framework for New Intent Discovery
Chaoran Yan, Jian Yang 0030, Wei Zhang 0384, Changyu Ren, Tongliang Li, Jiaqi Bai 0001, Zhoujun Li 0001 |
ECML/PKDD (5) | 7 |
| 2024 | mt4CrossOIE: Multi-stage tuning for cross-lingual open information extraction
Tongliang Li, Linzheng Chai, Jian Yang 0030, Jiaqi Bai 0001, Yuwei Yin, Hongcheng Guo, Liqun Yang, Hebboul Zine El Abidine, Zhoujun Li 0001 |
Expert Syst. Appl. | 5 |
| 2024 | Infusing internalized knowledge of language models into hybrid prompts for knowledgeable dialogue generation
Jiaqi Bai 0001, Jian Yang 0030, Hongcheng Guo, Zhoujun Li 0001 |
Knowl. Based Syst. | 1 |
| 2023 | GripRank: Bridging the Gap between Retrieval and Generation via the Generative Knowledge Improved Passage RankingabstractRetrieval-enhanced text generation has shown remarkable progress on knowledge-intensive language tasks, such as open-domain question answering and knowledge-enhanced dialogue generation, by leveraging passages retrieved from a large passage corpus for delivering a proper answer given the input query. However, the retrieved passages are not ideal for guiding answer generation because of the discrepancy between retrieval and generation, i.e., the candidate passages are all treated equally during the retrieval procedure without considering their potential to generate a proper answer. This discrepancy makes a passage retriever deliver a sub-optimal collection of candidate passages to generate the answer. In this paper, we propose the GeneRative Knowledge Improved Passage Ranking (GripRank) approach, addressing the above challenge by distilling knowledge from a generative passage estimator (GPE) to a passage ranker, where the GPE is a generative language model used to measure how likely the candidate passages can generate the proper answer. We realize the distillation procedure by teaching the passage ranker learning to rank the passages ordered by the GPE. Furthermore, we improve the distillation quality by devising a curriculum knowledge distillation mechanism, which allows the knowledge provided by the GPE can be progressively distilled to the ranker through an easy-to-hard curriculum, enabling the passage ranker to correctly recognize the provenance of the answer from many plausible candidates. We conduct extensive experiments on four datasets across three knowledge-intensive language tasks. Experimental results show advantages over the state-of-the-art methods for both passage ranking and answer generation on the KILT benchmark. Jiaqi Bai 0001, Hongcheng Guo, Jian Yang 0030, Xinnian Liang, Zhoujun Li 0001 |
CIKM | 1 |
| 2023 | Modeling Intra-class and Inter-class Constraints for Out-of-Domain Detection
Jiaqi Bai 0001, Tongliang Li, Zhoujun Li 0001 |
DASFAA (4) | 2 |
| 2023 | Enhancing Dialogue Summarization with Topic-Aware Global- and Local- Level CentralityabstractDialogue summarization aims to condense a given dialogue into a simple and focused summary text.Typically, both the roles' viewpoints and conversational topics change in the dialogue stream.Thus how to effectively handle the shifting topics and select the most salient utterance becomes one of the major challenges of this task.In this paper, we propose a novel topic-aware Global-Local Centrality (GLC) model to help select the salient context from all sub-topics.The centralities are constructed at both the global and local levels.The global one aims to identify vital sub-topics in the dialogue and the local one aims to select the most important context in each sub-topic.Specifically, the GLC collects sub-topic based on the utterance representations.And each utterance is aligned with one sub-topic.Based on the sub-topics, the GLC calculates globaland local-level centralities.Finally, we combine the two to guide the model to capture both salient context and sub-topics when generating summaries.Experimental results show that our model outperforms strong baselines on three public dialogue summarization datasets: CSDS, MC, and SAMSUM.Further analysis demonstrates that our GLC can exactly identify vital contents from sub-topics. 1 Xinnian Liang, Shuangzhi Wu, Chenhao Cui, Jiaqi Bai 0001, Chao Bian 0006, Zhoujun Li 0001 |
EACL | 4 |
| 2023 | Label-Guided Contrastive Learning for Out-of-Domain DetectionabstractOut-of-Domain (OOD) detection from user utterances plays an important part in task-oriented dialogue systems. Recent studies utilize supervised or self-supervised contrastive learning (CL) to learn discriminative semantic features for OOD detection. The supervised contrastive learning (SCL) methods only model class-level features of different in-domain (IND) intents while self-supervised CL (SSCL) methods can model instance-level features. However, the SSCL methods require complex data augmentation and are vulnerable to intrinsic false-negative pairs. To address the issues above and leverage both types of CL, we propose a novel Label-Guided Contrastive Learning (LGCL) framework. LGCL models both instance-level and class-level discriminative semantic representations by employing the sample and its corresponding label simultaneously, such that the prior knowledge of IND can be fully leveraged. Experiment and analysis results on two public benchmark datasets show that the proposed method significantly outperforms baselines on the OOD detection task. Tongliang Li, Jiaqi Bai 0001, Zhoujun Li 0001 |
ICASSP | 3 |
| 2023 | KnowPrefix-Tuning: A Two-Stage Prefix-Tuning Framework for Knowledge-Grounded Dialogue Generation
Jiaqi Bai 0001, Ze Yang 0001, Jian Yang 0030, Xinnian Liang, Hongcheng Guo, Zhoujun Li 0001 |
ECML/PKDD (2) | 1 |
| 2023 | KINet: Incorporating Relevant Facts Into Knowledge-Grounded Dialog GenerationabstractKnowledge-grounded conversation has led to great progress in producing informative dialog responses by leveraging external knowledge. This work focuses on two affiliated knowledge grounded conversation tasks:Knowledge SelectionandResponse Generation. Previous work followed the paradigm of selecting the most optimal knowledge piece to guide the conversation towards generating the proper response. However, some knowledge pieces, which are not recognized as optimal, may still benefit response generation. How to effectively leverage these relevant knowledge pieces for response generation still remain a tricky issue. To address this problem, we proposeKINet, aKnowledgeIncorporationNetwork, which deals with the problem by boosting both the knowledge selection and the response generation. The proposed model contains a negative enhanced knowledge approximator which improves knowledge selection by enhancing the dense representation of knowledge pieces, and a curriculum knowledge sampler which improves generated responses by incorporating more relevant knowledge pieces in an easy-to-hard manner. We conduct the experiment on two datasets of knowledge-grounded conversations, the results show that the proposed model significantly outperforms state-of-the-art methods in terms of both automatic and human evaluations. Jiaqi Bai 0001, Ze Yang 0001, Jian Yang 0030, Hongcheng Guo, Zhoujun Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2021 | Learning to Copy Coherent Knowledge for Response GenerationabstractKnowledge-driven dialog has shown remarkable performance to alleviate the problem of generating uninformative responses in the dialog system. However, incorporating knowledge coherently and accurately into response generation is still far from being solved. Previous works dropped into the paradigm of non-goal-oriented knowledge-driven dialog, they are prone to ignore the effect of dialog goal, which has potential impacts on knowledge exploitation and response generation. To address this problem, this paper proposes a Goal-Oriented Knowledge Copy network, GOKC. Specifically, a goal-oriented knowledge discernment mechanism is designed to help the model discern the knowledge facts that are highly correlated to the dialog goal and the dialog context. Besides, a context manager is devised to copy facts not only from the discerned knowledge but also from the dialog goal and the dialog context, which allows the model to accurately restate the facts in the generated response. The empirical studies are conducted on two benchmarks of goal-oriented knowledge-driven dialog generation. The results show that our model can significantly outperform several state-of-the-art models in terms of both automatic evaluation and human judgments. Jiaqi Bai 0001, Ze Yang 0001, Xinnian Liang, Wei Wang 0301, Zhoujun Li 0001 |
AAAI | 1 |
| 2021 | Jointly Learning to Repair Code and Generate Commit MessageabstractWe propose a novel task of jointly repairing program codes and generating commit messages.Code repair and commit message generation are two essential and related tasks for software development.However, existing work usually performs the two tasks independently.We construct a multilingual triple dataset including buggy code, fixed code, and commit messages for this novel task.We provide the cascaded models as baseline, which are enhanced with different training approaches, including the teacher-student method, the multi-task method, and the backtranslation method.To deal with the error propagation problem of the cascaded method, the joint model is proposed that can both repair the code and generate the commit message in a unified framework.Experimental results show that the enhanced cascaded model with teacher-student method and multitask-learning method achieves the best score on different metrics of automated code repair, and the joint model behaves better than the cascaded model on commit message generation. Jiaqi Bai 0001, Ambrosio Blanco, Shujie Liu 0001, Furu Wei, Ming Zhou 0001, Zhoujun Li 0001 |
EMNLP (1) | 1 |