EDBT 2026 Demo / reviewers in the wild / expert
Songlin Zhai
dblp:213/7563
· DBLP profile ↗
13ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0003-2666-9952ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TaxReasoning: Benchmarking Knowledge-Intensive Mathematical Reasoning with Evolving Tax LawsabstractRecent studies have explored the capabilities of large language models (LLMs) in solving knowledge-intensive mathematical reasoning problems. However, existing benchmarks predominantly involve static theorems that LLMs have encountered during pretraining, failing to assess dynamic knowledge integration. In this work, we introduce TaxReasoning, a novel benchmark designed to evaluate LLMs’ abilities in real-world tax calculation scenarios. These tasks require not only mathematical reasoning and numerical computation, but also the extraction and application of complex, frequently updated tax regulations. Through extensive experiments with state-of-the-art LLMs using diverse prompting strategies and knowledge augmentation techniques, we uncover substantial limitations in their ability to handle dynamic, knowledge-intensive questions—primarily due to missing domain-specific knowledge and ineffective retrieval. Even the best-performing models fall significantly short of human-level performance. Our analysis points to key avenues for improvement, including enhancing LLMs' reasoning capabilities, developing more effective knowledge summarization techniques, and improving retrieval strategies. TaxReasoning offers a critical testbed for advancing LLMs in dynamic knowledge-intensive domains. Nan Hu 0004, Huikang Hu, Guilin Qi, Songlin Zhai, Yongrui Chen 0002, Tianxing Wu 0001, Tongtong Wu, Jiaoyan Chen 0001, Jeff Z. Pan |
AAAI | 6 |
| 2026 | CRAG: Causality-Aware Retrieval-Augmented Generation for Budget Auditing QA
Guilin Qi, Xiaolong Ye, Songlin Zhai, Yongrui Chen 0002, Shenwen Zhong |
DASFAA (6) | 6 |
| 2026 | IGen: Redefining long-term event prediction with iterative generation and dynamic balancing
Yan Wang 0124, Songlin Zhai, Yongrui Chen 0002, Shenyu Zhang 0002, Zhihua Chai, Guilin Qi |
Inf. Process. Manag. | 3 |
| 2025 | Parameter-Aware Contrastive Knowledge Editing: Tracing and Rectifying based on Critical Transmission PathsabstractLarge language models (LLMs) have encoded vast amounts of knowledge in their parameters, but the acquired knowledge can sometimes be incorrect or outdated over time, necessitating rectification after pre-training.Traditional localized methods in knowledge-based model editing (KME) typically assume that knowledge is stored in particular intermediate layers.However, recent research suggests that these methods do not identify the optimal locations for parameter editing, as knowledge gradually accumulates across all layers in LLMs during the forward pass rather than being stored in specific layers.This paper, for the first time, introduces the concept of critical transmission paths into KME for parameter updating.Specifically, these paths capture the key information flows that significantly influence the model predictions for the editing process.To facilitate this process, we also design a parameter-aware contrastive rectifying algorithm that considers less important paths as contrastive examples.Experiments on two prominent datasets and three widely used LLMs demonstrate the superiority of our method in editing performance. Songlin Zhai, Guilin Qi |
ACL (1) | 1 |
| 2025 | TEF: Causality-Aware Taxonomy Expansion via Front-Door CriterionabstractTaxonomy expansion is a primary method for enriching taxonomies, involving appending a large number of additional nodes (i.e., queries) to an existing taxonomy (i.e., seed), with the crucial step being the identification of the appropriate anchor (parent node) for each query by incorporating the structural information of the seed. Despite advancements, existing research still faces an inherent challenge of spurious query-anchor matching, often due to various interference factors (e.g., the consistency of sibling nodes), resulting in biased identifications. To address the bias in taxonomy expansion caused by unobserved factors, we introduce the Structural Causal Model (SCM), known for its bias elimination capabilities, to prevent these factors from confounding the task through backdoor paths. Specifically, we employ the Front-Door Criterion, which guides the decomposition of the expansion process into a parser module and a connector. This enables the proposed causal-aware Taxonomy Expansion model to isolate confounding effects and reveal the true causal relationship between the query and the anchor. Extensive experiments on three benchmarks validate the effectiveness of TEF, with a notable 6.1% accuracy improvement over the state-of-the-art on the SemEval16-Environment dataset. Songlin Zhai, Zhongjian Hu, Guilin Qi |
COLING | 2 |
| 2025 | Harnessing Diverse Perspectives: A Multi-agent Framework for Enhanced Error Detection in Knowledge Graphs
Yu Li 0021, Yi Huang 0017, Guilin Qi, Junlan Feng, Nan Hu 0004, Songlin Zhai, Haohan Xue, Yongrui Chen 0002, Ruoyan Shen, Tongtong Wu |
DASFAA (6) | 6 |
| 2025 | Peripheral Memory for LLMs: Integration of Sequential Memory Banks with Adaptive QueryingabstractLarge Language Models (LLMs) have revolutionized various natural language processing tasks with their remarkable capabilities.
However, challenges persist in effectively integrating new knowledge into LLMs without compromising their performance, particularly in the Large Language Models (LLMs) have revolutionized various natural language processing tasks with their remarkable capabilities.
However, a challenge persists in effectively processing new information, particularly in the area of long-term knowledge updates without compromising model performance.
To address this challenge, this paper introduces a novel memory augmentation framework that conceptualizes memory as a peripheral component (akin to physical RAM), with the LLM serving as the information processor (analogous to a CPU).
Drawing inspiration from RAM architecture, we design memory as a sequence of memory banks, each modeled using Kolmogorov-Arnold Network (KAN) to ensure smooth state transitions.
Memory read and write operations are dynamically controlled by query signals derived from the LLMs' internal states, closely mimicking the interaction between a CPU and RAM.
Furthermore, a dedicated memory bank is used to generate a mask value that indicates the relevance of the retrieved data, inspired by the sign bit in binary coding schemes.
The retrieved memory feature is then integrated as a prefix to enhance the model prediction.
Extensive experiments on knowledge-based model editing validate the effectiveness and efficiency of our peripheral memory. Songlin Zhai, Yongrui Chen 0002, Guilin Qi |
ICML | 1 |
| 2025 | DST: Continual event prediction by decomposing and synergizing the task commonality and specificity
Songlin Zhai, Yongrui Chen 0002, Shenyu Zhang 0002, Guilin Qi |
Inf. Process. Manag. | 2 |
| 2024 | Attributed Triple Extraction by Combination Under Contrastive Learning
Runzhe Wang, Guilin Qi, Yongrui Chen 0002, Songlin Zhai, Rihui Jin, Nijun Li, Qianren Wang |
DASFAA (7) | 6 |
| 2024 | Event is more valuable than you think: Improving the Similar Legal Case Retrieval via event knowledge
Songlin Zhai, Yongrui Chen 0002, Guilin Qi |
Inf. Process. Manag. | 2 |
| 2024 | Which is better? Taxonomy induction with learning the optimal structure via contrastive learning
Songlin Zhai, Zhihua Chai, Tianxing Wu 0001, Guilin Qi |
Knowl. Based Syst. | 2 |
| 2023 | DNG: Taxonomy Expansion by Exploring the Intrinsic Directed Structure on Non-gaussian SpaceabstractTaxonomy expansion is the process of incorporating a large number of additional nodes (i.e., ''queries'') into an existing taxonomy (i.e., ''seed''), with the most important step being the selection of appropriate positions for each query. Enormous efforts have been made by exploring the seed's structure. However, existing approaches are deficient in their mining of structural information in two ways: poor modeling of the hierarchical semantics and failure to capture directionality of the is-a relation. This paper seeks to address these issues by explicitly denoting each node as the combination of inherited feature (i.e., structural part) and incremental feature (i.e., supplementary part). Specifically, the inherited feature originates from ''parent'' nodes and is weighted by an inheritance factor. With this node representation, the hierarchy of semantics in taxonomies (i.e., the inheritance and accumulation of features from ''parent'' to ''child'') could be embodied. Additionally, based on this representation, the directionality of the is-a relation could be easily translated into the irreversible inheritance of features. Inspired by the Darmois-Skitovich Theorem, we implement this irreversibility by a non-Gaussian constraint on the supplementary feature. A log-likelihood learning objective is further utilized to optimize the proposed model (dubbed DNG), whereby the required non-Gaussianity is also theoretically ensured. Extensive experimental results on two real-world datasets verify the superiority of DNG relative to several strong baselines. Songlin Zhai, Weiqing Wang 0001, Yuan-Fang Li |
AAAI | 1 |
| 2018 | VSE-ens: Visual-Semantic Embeddings with Efficient Negative SamplingabstractJointing visual-semantic embeddings (VSE) have become a research hotpot for the task of image annotation, which suffers from the issue of semantic gap, i.e., the gap between images' visual features (low-level) and labels' semantic features (high-level). This issue will be even more challenging if visual features cannot be retrieved from images, that is, when images are only denoted by numerical IDs as given in some real datasets. The typical way of existing VSE methods is to perform a uniform sampling method for negative examples that violate the ranking order against positive examples, which requires a time-consuming search in the whole label space. In this paper, we propose a fast adaptive negative sampler that can work well in the settings of no figure pixels available. Our sampling strategy is to choose the negative examples that are most likely to meet the requirements of violation according to the latent factors of images. In this way, our approach can linearly scale up to large datasets. The experiments demonstrate that our approach converges 5.02x faster than the state-of-the-art approaches on OpenImages, 2.5x on IAPR-TCI2 and 2.06x on NUS-WIDE datasets, as well as better ranking accuracy across datasets. Guibing Guo, Songlin Zhai, Fajie Yuan, Yuan Liu 0002, Xingwei Wang 0001 |
AAAI | 2 |