EDBT 2026 Demo / reviewers in the wild / expert
Shuofei Qiao
dblp:333/2421
· DBLP profile ↗
15ranked-venue papers
6as first author
15since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 6 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Why Do Open-Source LLMs Struggle with Data Analysis? A Systematic Empirical StudyabstractLarge Language Models (LLMs) hold promise in automating data analysis tasks, yet open-source models face significant limitations in these kinds of reasoning-intensive scenarios. In this work, we investigate strategies to enhance the data analysis capabilities of open-source LLMs. By curating a seed dataset of diverse, realistic scenarios, we evaluate models across three dimensions: data understanding, code generation, and strategic planning. Our analysis reveals three key findings: (1) Strategic planning quality serves as the primary determinant of model performance; (2) Interaction design and task complexity significantly influence reasoning capabilities; (3) Data quality demonstrates a greater impact than diversity in achieving optimal performance. We leverage these insights to develop a data synthesis methodology, demonstrating significant improvements in open-source LLMs' analytical reasoning capabilities. Jintian Zhang, Shuofei Qiao, Lun Du, Da Zheng 0004, Ningyu Zhang 0001, Huajun Chen |
AAAI | 5 |
| 2026 | KnowRL: Exploring Knowledgeable Reinforcement Learning for FactualityabstractSlow-thinking Large Language Models (LLMs) have demonstrated strong reasoning capabilities but often suffer from severe hallucinations due to an inability to recognize their knowledge boundaries.Existing Reinforcement Learning (RL) approaches typically rely on outcomeoriented rewards, which can inadvertently reinforce fabricated reasoning paths when the final answer is correct.To address this, we propose Knowledge-enhanced RL, KnowRL, a framework that integrates factual supervision directly into the reasoning process.By decomposing the chain of thought into atomic facts and verifying them against the corresponding ground-truth knowledge, KnowRL performs fine-grained checks to encourage models to reason faithfully.Crucially, this processoriented supervision teaches the model to identify its knowledge boundaries, learning to say "I don't know" instead of fabricating answers when information is missing.Experimental results demonstrate that KnowRL effectively mitigates hallucinations-reducing the Incorrect Rate on SimpleQA by 20.3% for distillationbased slow-thinking models while maintaining strong performance on complex reasoning benchmarks like GPQA and AIME 2025.Furthermore, our method shows robust transferability to out-of-distribution tasks, indicating that the model learns a generalizable verification behavior 1 . Baochang Ren, Shuofei Qiao, Ningyu Zhang 0001, Da Zheng 0004, Huajun Chen |
ACL (1) | 2 |
| 2026 | Mitigating Context Interference for Reliable and Efficient Search AgentsabstractBoyang Xue, Bin Wu, Shuofei Qiao, Sheng Wang, Rui Wang, Yiming Du, Hongru Wang, Jeff Z. Pan, Emine Yilmaz, Kam-Fai Wong, Aldo Lipani. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Boyang Xue, Bin Wu 0025, Shuofei Qiao, Rui Wang 0092, Yiming Du, Hongru Wang 0003, Jeff Z. Pan, Emine Yilmaz, Kam-Fai Wong, Aldo Lipani |
ACL (1) | 3 |
| 2025 | OS Agents: A Survey on MLLM-based Agents for Computer, Phone and Browser UseabstractXueyu Hu, Tao Xiong, Biao Yi, Zishu Wei, Ruixuan Xiao, Yurun Chen, Jiasheng Ye, Meiling Tao, Xiangxin Zhou, Ziyu Zhao, Yuhuai Li, Shengze Xu, Shenzhi Wang, Xinchen Xu, Shuofei Qiao, Zhaokai Wang, Kun Kuang, Tieyong Zeng, Liang Wang, Jiwei Li, Yuchen Eleanor Jiang, Wangchunshu Zhou, Guoyin Wang, Keting Yin, Zhou Zhao, Hongxia Yang, Fan Wu, Shengyu Zhang, Fei Wu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xueyu Hu, Biao Yi, Zishu Wei, Ruixuan Xiao, Yurun Chen 0004, Jiasheng Ye, Meiling Tao, Xiangxin Zhou, Ziyu Zhao 0001, Yuhuai Li, Shengze Xu, Shenzhi Wang, Shuofei Qiao, Zhaokai Wang, Kun Kuang 0001, Tieyong Zeng, Liang Wang 0001, Jiwei Li 0001, Yuchen Eleanor Jiang, Wangchunshu Zhou, Guoyin Wang 0002, Keting Yin, Zhou Zhao 0001, Hongxia Yang, Fan Wu 0006, Shengyu Zhang 0001, Fei Wu 0001 |
ACL (1) | 15 |
| 2025 | Agentic Knowledgeable Self-awarenessabstractShuofei Qiao, Zhisong Qiu, Baochang Ren, Xiaobin Wang, Xiangyuan Ru, Ningyu Zhang, Xiang Chen, Yong Jiang, Pengjun Xie, Fei Huang, Huajun Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Shuofei Qiao, Zhisong Qiu, Baochang Ren, Xiaobin Wang, Xiangyuan Ru, Ningyu Zhang 0001, Xiang Chen 0016, Yong Jiang 0005, Pengjun Xie, Fei Huang 0002, Huajun Chen |
ACL (1) | 1 |
| 2025 | LightThinker: Thinking Step-by-Step CompressionabstractJintian Zhang, Yuqi Zhu, Mengshu Sun, Yujie Luo, Shuofei Qiao, Lun Du, Da Zheng, Huajun Chen, Ningyu Zhang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Jintian Zhang, Mengshu Sun, Shuofei Qiao, Lun Du, Da Zheng 0004, Huajun Chen, Ningyu Zhang 0001 |
EMNLP | 5 |
| 2025 | Benchmarking Agentic Workflow GenerationabstractLarge Language Models (LLMs), with their exceptional ability to handle a wide range of tasks, have driven significant advancements in tackling reasoning and planning tasks, wherein decomposing complex problems into executable workflows is a crucial step in this process. Existing workflow evaluation frameworks either focus solely on holistic performance or suffer from limitations such as restricted scenario coverage, simplistic workflow structures, and lax evaluation standards. To this end, we introduce WorfBench, a unified workflow generation benchmark with multi-faceted scenarios and intricate graph workflow structures. Additionally, we present WorfEval, a systemic evaluation protocol utilizing subsequence and subgraph matching algorithms to accurately quantify the LLM agent's workflow generation capabilities. Through comprehensive evaluations across different types of LLMs, we discover distinct gaps between the sequence planning capabilities and graph planning capabilities of LLM agents, with even GPT-4 exhibiting a gap of around 15%. We also train two open-source models and evaluate their generalization abilities on held-out tasks. Furthermore, we observe that the generated workflows can enhance downstream tasks, enabling them to achieve superior performance with less time during inference. Code and dataset are available at https://github.com/zjunlp/WorfBench. Shuofei Qiao, Runnan Fang, Zhisong Qiu, Xiaobin Wang, Ningyu Zhang 0001, Yong Jiang 0005, Pengjun Xie, Fei Huang 0002, Huajun Chen |
ICLR | 1 |
| 2024 | AutoAct: Automatic Agent Learning from Scratch for QA via Self-PlanningabstractShuofei Qiao, Ningyu Zhang, Runnan Fang, Yujie Luo, Wangchunshu Zhou, Yuchen Jiang, Chengfei Lv, Huajun Chen. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Shuofei Qiao, Ningyu Zhang 0001, Runnan Fang, Wangchunshu Zhou, Yuchen Eleanor Jiang, Chengfei Lv, Huajun Chen |
ACL (1) | 1 |
| 2024 | Making Language Models Better Tool Learners with Execution FeedbackabstractShuofei Qiao, Honghao Gui, Chengfei Lv, Qianghuai Jia, Huajun Chen, Ningyu Zhang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Shuofei Qiao, Honghao Gui, Chengfei Lv, Qianghuai Jia, Huajun Chen, Ningyu Zhang 0001 |
NAACL-HLT | 1 |
| 2024 | Agent Planning with World Knowledge ModelabstractRecent endeavors towards directly using large language models (LLMs) as agent models to execute interactive planning tasks have shown commendable results. Despite their achievements, however, they still struggle with brainless trial-and-error in global planning and generating hallucinatory actions in local planning due to their poor understanding of the "real" physical world. Imitating humans' mental world knowledge model which provides global prior knowledge before the task and maintains local dynamic knowledge during the task, in this paper, we introduce parametric World Knowledge Model (WKM) to facilitate agent planning. Concretely, we steer the agent model to self-synthesize knowledge from both expert and sampled trajectories. Then we develop WKM, providing prior task knowledge to guide the global planning and dynamic state knowledge to assist the local planning. Experimental results on three real-world simulated datasets with Mistral-7B, Gemma-7B, and Llama-3-8B demonstrate that our method can achieve superior performance compared to various strong baselines. Besides, we analyze to illustrate that our WKM can effectively alleviate the blind trial-and-error and hallucinatory action issues, providing strong support for the agent's understanding of the world. Other interesting findings include: 1) our instance-level task knowledge can generalize better to unseen tasks, 2) weak WKM can guide strong agent model planning, and 3) unified WKM training has promising potential for further development. Shuofei Qiao, Runnan Fang, Ningyu Zhang 0001, Xiang Chen 0016, Shumin Deng, Yong Jiang 0005, Pengjun Xie, Fei Huang 0002, Huajun Chen |
NeurIPS | 1 |
| 2024 | InstructIE: A Bilingual Instruction-based Information Extraction Dataset
Honghao Gui, Shuofei Qiao, Jintian Zhang, Hongbin Ye, Mengshu Sun, Lei Liang 0002, Jeff Z. Pan, Huajun Chen, Ningyu Zhang 0001 |
ISWC (3) | 2 |
| 2024 | LLMs for knowledge graph construction and reasoning: recent capabilities and future opportunities
Jing Chen 0060, Shuofei Qiao, Yixin Ou, Yunzhi Yao, Shumin Deng, Huajun Chen, Ningyu Zhang 0001 |
World Wide Web (WWW) | 4 |
| 2023 | On Analyzing the Role of Image for Visual-Enhanced Relation Extraction (Student Abstract)abstractMultimodal relation extraction is an essential task for knowledge graph construction. In this paper, we take an in-depth empirical analysis that indicates the inaccurate information in the visual scene graph leads to poor modal alignment weights, further degrading performance. Moreover, the visual shuffle experiments illustrate that the current approaches may not take full advantage of visual information. Based on the above observation, we further propose a strong baseline with an implicit fine-grained multimodal alignment based on Transformer for multimodal relation extraction. Experimental results demonstrate the better performance of our method. Codes are available at https://github.com/zjunlp/DeepKE/tree/main/example/re/multimodal. Lei Li 0040, Xiang Chen 0016, Shuofei Qiao, Feiyu Xiong, Huajun Chen, Ningyu Zhang 0001 |
AAAI | 3 |
| 2023 | Reasoning with Language Model Prompting: A SurveyabstractShuofei Qiao, Yixin Ou, Ningyu Zhang, Xiang Chen, Yunzhi Yao, Shumin Deng, Chuanqi Tan, Fei Huang, Huajun Chen. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Shuofei Qiao, Yixin Ou, Ningyu Zhang 0001, Xiang Chen 0016, Yunzhi Yao, Shumin Deng, Chuanqi Tan, Fei Huang 0002, Huajun Chen |
ACL (1) | 1 |
| 2023 | One Model for All Domains: Collaborative Domain-Prefix Tuning for Cross-Domain NERabstractCross-domain NER is a challenging task to address the low-resource problem in practical scenarios. Previous typical solutions mainly obtain a NER model by pre-trained language models (PLMs) with data from a rich-resource domain and adapt it to the target domain. Owing to the mismatch issue among entity types in different domains, previous approaches normally tune all parameters of PLMs, ending up with an entirely new NER model for each domain. Moreover, current models only focus on leveraging knowledge in one general source domain while failing to successfully transfer knowledge from multiple sources to the target. To address these issues, we introduce Collaborative Domain-Prefix Tuning for cross-domain NER (CP-NER) based on text-to-text generative PLMs. Specifically, we present text-to-text generation grounding domain-related instructors to transfer knowledge to new domain NER tasks without structural modifications. We utilize frozen PLMs and conduct collaborative domain-prefix tuning to stimulate the potential of PLMs to handle NER tasks across various domains. Experimental results on the Cross-NER benchmark show that the proposed approach has flexible transfer ability and performs better on both one-source and multiple-source cross-domain NER tasks. Xiang Chen 0016, Lei Li 0040, Shuofei Qiao, Ningyu Zhang 0001, Chuanqi Tan, Yong Jiang 0005, Fei Huang 0002, Huajun Chen |
IJCAI | 3 |