Shulin Cao

dblp:229/2976 · DBLP profile ↗
← Back
17ranked-venue papers
2as first author
17since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 2 first-author · 15 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 SeaKR: Self-aware Knowledge Retrieval for Adaptive Retrieval Augmented Generation
abstract
Zijun Yao, Weijian Qi, Liangming Pan, Shulin Cao, Linmei Hu, Liu Weichuan, Lei Hou, Juanzi Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zijun Yao 0002, Weijian Qi, Liangming Pan, Shulin Cao, Linmei Hu, Weichuan Liu, Lei Hou 0001, Juan-Zi Li
ACL (1)4
2025 LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks
abstract
Yushi Bai, Shangqing Tu, Jiajie Zhang, Hao Peng, Xiaozhi Wang, Xin Lv, Shulin Cao, Jiazheng Xu, Lei Hou, Yuxiao Dong, Jie Tang, Juanzi Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yushi Bai, Shangqing Tu, Hao Peng 0015, Xiaozhi Wang, Shulin Cao, Jiazheng Xu, Lei Hou 0001, Yuxiao Dong, Jie Tang 0001, Juan-Zi Li
ACL (1)7
2025 LongReward: Improving Long-context Large Language Models with AI Feedback
abstract
Jiajie Zhang, Zhongni Hou, Xin Lv, Shulin Cao, Zhenyu Hou, Yilin Niu, Lei Hou, Yuxiao Dong, Ling Feng, Juanzi Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zhongni Hou, Shulin Cao, Yilin Niu, Lei Hou 0001, Yuxiao Dong, Juan-Zi Li
ACL (1)4
2025 AtomR: Atomic Operator-Empowered Large Language Models for Heterogeneous Knowledge Reasoning
abstract
Despite the outstanding capabilities of large language models (LLMs), knowledge-intensive reasoning still remains a challenging task due to LLMs' limitations in compositional reasoning and the hallucination problem. A prevalent solution is to employ chain-of-thought (CoT) with retrieval-augmented generation (RAG), which first formulates a reasoning plan by decomposing complex questions into simpler sub-questions, and then applies iterative RAG at each sub-question. However, prior works exhibit two crucial problems: inadequate reasoning planning and poor incorporation of heterogeneous knowledge. In this paper, we introduce AtomR, a framework for LLMs to conduct accurate heterogeneous knowledge reasoning at the atomic level. Inspired by how knowledge graph query languages model compositional reasoning through combining predefined operations, we propose three atomic knowledge operators, a unified set of operators for LLMs to retrieve and manipulate knowledge from heterogeneous sources. First, in the reasoning planning stage, AtomR decomposes a complex question into a reasoning tree where each leaf node corresponds to an atomic knowledge operator, achieving question decomposition that is highly fine-grained and orthogonal. Subsequently, in the reasoning execution stage, AtomR executes each atomic knowledge operator, which flexibly selects, retrieves, and operates atomic level knowledge from heterogeneous sources. We also introduce BlendQA, a challenging benchmark specially tailored for heterogeneous knowledge reasoning. Experiments on three single-source and two multi-source datasets show that AtomR outperforms state-of-the-art baselines by a large margin, with absolute F1 score improvements of 9.4% on 2WikiMultihop and 9.5% on BlendQA. We release our code and data https://github.com/THU-KEG/AtomR.git.
Amy Xin, Jinxin Liu 0002, Zijun Yao 0002, Zhicheng Lee, Shulin Cao, Lei Hou 0001, Juan-Zi Li
KDD (2)5
2025 Dynamic multi teacher knowledge distillation for semantic parsing in KBQA
Ao Zou, Shulin Cao, Jinxin Liu 0002, Lei Hou 0001
Expert Syst. Appl.3
2024 Untangle the KNOT: Interweaving Conflicting Knowledge and Reasoning Skills in Large Language Models
abstract
Providing knowledge documents for large language models (LLMs) has emerged as a promising solution to update the static knowledge inherent in their parameters. However, knowledge in the document may conflict with the memory of LLMs due to outdated or incorrect knowledge in the LLMs’ parameters. This leads to the necessity of examining the capability of LLMs to assimilate supplemental external knowledge that conflicts with their memory. While previous studies have explained to what extent LLMs extract conflicting knowledge from the provided text, they neglect the necessity to <b>reason</b> with conflicting knowledge. Furthermore, there lack a detailed analysis on strategies to enable LLMs to resolve conflicting knowledge via prompting, decoding strategy, and supervised fine-tuning. To address these limitations, we construct a new dataset, dubbed KNOT, for knowledge conflict resolution examination in the form of question answering. KNOT facilitates in-depth analysis by dividing reasoning with conflicting knowledge into three levels: (1) Direct Extraction, which directly extracts conflicting knowledge to answer questions. (2) Explicit Reasoning, which reasons with conflicting knowledge when the reasoning path is explicitly provided in the question. (3) Implicit Reasoning, where reasoning with conflicting knowledge requires LLMs to infer the reasoning path independently to answer questions. We also conduct extensive experiments on KNOT to establish empirical guidelines for LLMs to utilize conflicting knowledge in complex circumstances. Dataset and associated codes can be accessed at our <a href=https://github.com/THU-KEG/KNOT>GitHub repository</a> .
Yantao Liu, Zijun Yao 0002, Shulin Cao, Jifan Yu, Lei Hou 0001, Juan-Zi Li
LREC/COLING5
2024 KB-Plugin: A Plug-and-play Framework for Large Language Models to Induce Programs over Low-resourced Knowledge Bases
abstract
Program induction (PI) has become a promising paradigm for using knowledge bases (KBs) to help large language models (LLMs) answer complex knowledge-intensive questions.Nonetheless, PI typically relies on a large number of parallel question-program pairs to make the LLM aware of the schema of a given KB, and is thus challenging for many low-resourced KBs that lack annotated data.To this end, we propose KB-Plugin, a plug-and-play framework that enables LLMs to induce programs over any low-resourced KB.Firstly, KB-Plugin adopts self-supervised learning to encode the detailed schema information of a given KB into a pluggable module, namely schema plugin.Secondly, KB-Plugin utilizes abundant annotated data from a rich-resourced KB to train another pluggable module, namely PI plugin, which can help the LLM extract questionrelevant schema information from the schema plugin of any KB and utilize the information to induce programs over this KB.Experiments show that KB-Plugin outperforms SoTA lowresourced PI methods with 25× smaller backbone LLM on both large-scale and domainspecific KBs, and even approaches the performance of supervised methods.
Shulin Cao, Linmei Hu, Lei Hou 0001, Juan-Zi Li
EMNLP2
2024 KoLA: Carefully Benchmarking World Knowledge of Large Language Models
abstract
The unprecedented performance of large language models (LLMs) necessitates improvements in evaluations. Rather than merely exploring the breadth of LLM abilities, we believe meticulous and thoughtful designs are essential to thorough, unbiased, and applicable evaluations. Given the importance of world knowledge to LLMs, we construct a Knowledge-oriented LLM Assessment benchmark (KoLA), in which we carefully design three crucial factors: (1) For ability modeling, we mimic human cognition to form a four-level taxonomy of knowledge-related abilities, covering 19 tasks. (2) For data, to ensure fair comparisons, we use both Wikipedia, a corpus prevalently pre-trained by LLMs, along with continuously collected emerging corpora, aiming to evaluate the capacity to handle unseen data and evolving knowledge. (3) For evaluation criteria, we adopt a contrastive system, including overall standard scores for better numerical comparability across tasks and models, and a unique self-contrast metric for automatically evaluating knowledge-creating ability. We evaluate 21 open-source and commercial LLMs and obtain some intriguing findings. The KoLA dataset will be updated every three months to provide timely references for developing LLMs and knowledge-related systems.
Jifan Yu, Xiaozhi Wang, Shangqing Tu, Shulin Cao, Daniel Zhang-Li, Hao Peng 0015, Zijun Yao 0002, Hanming Li, Zheyuan Zhang 0002, Yushi Bai, Yantao Liu, Amy Xin, Kaifeng Yun, Linlu Gong, Nianyi Lin, Zhi-Li Wu, Yunjia Qi, Weikai Li 0002, Kaisheng Zeng, Ji Qi 0003, Hailong Jin, Jinxin Liu 0002, Yu Gu 0029, Yuan Yao 0011, Ning Ding 0002, Lei Hou 0001, Zhiyuan Liu 0001, Bin Xu 0001, Jie Tang 0001, Juan-Zi Li
ICLR4
2023 Reasoning over Hierarchical Question Decomposition Tree for Explainable Question Answering
abstract
Jiajie Zhang, Shulin Cao, Tingjian Zhang, Xin Lv, Juanzi Li, Lei Hou, Jiaxin Shi, Qi Tian. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Shulin Cao, Tingjian Zhang, Juan-Zi Li, Lei Hou 0001, Jiaxin Shi
ACL (1)2
2023 FC-KBQA: A Fine-to-Coarse Composition Framework for Knowledge Base Question Answering
abstract
Lingxi Zhang, Jing Zhang, Yanling Wang, Shulin Cao, Xinmei Huang, Cuiping Li, Hong Chen, Juanzi Li. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Lingxi Zhang, Jing Zhang 0001, Shulin Cao, Xinmei Huang, Cuiping Li 0001, Hong Chen 0001, Juan-Zi Li
ACL (1)4
2022 Program Transfer for Answering Complex Questions over Knowledge Bases
abstract
Shulin Cao, Jiaxin Shi, Zijun Yao, Xin Lv, Jifan Yu, Lei Hou, Juanzi Li, Zhiyuan Liu, Jinghui Xiao. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Shulin Cao, Jiaxin Shi, Zijun Yao 0002, Jifan Yu, Lei Hou 0001, Juan-Zi Li, Zhiyuan Liu 0001, Jinghui Xiao
ACL (1)1
2022 KQA Pro: A Dataset with Explicit Compositional Programs for Complex Question Answering over Knowledge Base
abstract
Shulin Cao, Jiaxin Shi, Liangming Pan, Lunyiu Nie, Yutong Xiang, Lei Hou, Juanzi Li, Bin He, Hanwang Zhang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Shulin Cao, Jiaxin Shi, Liangming Pan, Lunyiu Nie, Yutong Xiang, Lei Hou 0001, Juan-Zi Li, Hanwang Zhang
ACL (1)1
2022 Triple-as-Node Knowledge Graph and Its Embeddings
Jiaxin Shi, Shulin Cao, Lei Hou 0001, Juan-Zi Li
DASFAA (1)3
2022 GraphQ IR: Unifying the Semantic Parsing of Graph Query Languages with One Intermediate Representation
abstract
Subject to the huge semantic gap between natural and formal languages, neural semantic parsing is typically bottlenecked by its complexity of dealing with both input semantics and output syntax.Recent works have proposed several forms of supplementary supervision but none is generalized across multiple formal languages.This paper proposes a unified intermediate representation (IR) for graph query languages, named GraphQ IR.It has a natural-language-like expression that bridges the semantic gap and formally defined syntax that maintains the graph structure.Therefore, a neural semantic parser can more precisely convert user queries into GraphQ IR, which can be later losslessly compiled into various downstream graph query languages.Extensive experiments on several benchmarks including KQA PRO, OVERNIGHT, GRAILQA and METAQA-Cypher under standard i.i.d., out-of-distribution and low-resource settings validate GraphQ IR's superiority over the previous state-of-the-arts with a maximum 11% accuracy improvement.
Lunyiu Nie, Shulin Cao, Jiaxin Shi, Jiuding Sun, Lei Hou 0001, Juan-Zi Li, Jidong Zhai
EMNLP2
2021 MOOCCubeX: A Large Knowledge-centered Repository for Adaptive Learning in MOOCs
abstract
The prosperity of massive open online courses provides fodder for plentiful research efforts on adaptive learning. However, current open-access educational datasets are still far from sufficient to meet the need for various topics of adaptive learning. Existing released datasets often cover only small-scale data, lack fine-grained knowledge concepts. They are even difficult to curate and supplement due to platform limitations. In this work, we construct MOOCCubeX, a large, knowledge-centered repository consisting of 4,216 courses, 230,263 videos, 358,265 exercises, 637,572 fine-grained concepts and over 296 million behavioral data of 3,330,294 students, for supporting the research topics on adaptive learning in MOOCs. Licensed by XuetangX, one of the largest MOOC websites in China, we obtain abundant and diverse course resources and student behavioral data and are permitted to make subsequent periodic updates. We propose a framework to accomplish data processing, weakly supervised fine-grained concept graph mining, and data curation to improve usability and richness. Based on the fine-grained concepts, we re-organize the data from the knowledge perspective and acquire more external learning resources from the web. Our repository is now available at https://github.com/THU-KEG/MOOCCubeX.
Jifan Yu, Yuquan Wang, Qingyang Zhong, Gan Luo, Yiming Mao 0005, Wenzheng Feng, Wei Xu 0017, Shulin Cao, Kaisheng Zeng, Zijun Yao 0002, Lei Hou 0001, Yankai Lin 0001, Peng Li 0030, Jie Zhou 0016, Bin Xu 0001, Juan-Zi Li, Jie Tang 0001, Maosong Sun 0001
CIKM9
2021 TransferNet: An Effective and Transparent Framework for Multi-hop Question Answering over Relation Graph
abstract
Multi-hop Question Answering (QA) is a challenging task because it requires precise reasoning with entity relations at every step towards the answer.The relations can be represented in terms of labels in knowledge graph (e.g., spouse) or text in text corpus (e.g., they have been married for 26 years).Existing models usually infer the answer by predicting the sequential relation path or aggregating the hidden graph features.The former is hard to optimize, and the latter lacks interpretability.In this paper, we propose Trans-ferNet, an effective and transparent model for multi-hop QA, which supports both label and text relations in a unified framework.Trans-ferNet jumps across entities at multiple steps.At each step, it attends to different parts of the question, computes activated scores for relations, and then transfer the previous entity scores along activated relations in a differentiable way.We carry out extensive experiments on three datasets and demonstrate that TransferNet surpasses the state-of-the-art models by a large margin.In particular, on MetaQA, it achieves 100% accuracy in 2-hop and 3-hop questions.By qualitative analysis, we show that TransferNet has transparent and interpretable intermediate results.
Jiaxin Shi, Shulin Cao, Lei Hou 0001, Juan-Zi Li, Hanwang Zhang
EMNLP (1)2
2021 Bisecting k-means based fingerprint indoor localization
Haojie Zhao, Shulin Cao, Shasha Fu, Dingde Jiang
Wirel. Networks4