EDBT 2026 Demo / reviewers in the wild / expert
Xihe Qiu
dblp:258/7989
· DBLP profile ↗
8ranked-venue papers in the field
1as first author
8since 2021 · last 2026
0000-0003-4024-925XORCID · verified
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 4Information Retrieval & Web Search · 3 (1 first)Knowledge Engineering, Semantic Web & Information Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AURORA: Automated Training Framework of Universal Process Reward Models via Ensemble Prompting and Reverse VerificationabstractThe reasoning capabilities of advanced large language models (LLMs) like o1 have revolutionized artificial intelligence applications. Nevertheless, evaluating and optimizing complex reasoning processes remain significant challenges due to diverse policy distributions and the inherent limitations of human effort and accuracy. In this paper, we present AURORA, a novel automated framework for training universal process reward models (PRMs) using ensemble prompting and reverse verification. The framework employs a two-phase approach: First, it uses diverse prompting strategies and ensemble methods to perform automated annotation and evaluation of processes, ensuring robust assessments for reward learning. Second, it leverages practical reference answers for reverse verification, enhancing the model's ability to validate outputs and improving training accuracy. To assess the framework's performance, we extend beyond the existing ProcessBench benchmark by introducing UniversalBench, which evaluates reward predictions across full trajectories under diverse policy distribtion with long Chain-of-Thought (CoT) outputs. Experimental results demonstrate that AURORA enhances process evaluation accuracy, improves PRMs' accuracy for diverse policy distributions and long-CoT responses. Xiaoyu Tan, Tianchu Yao, Chao Qu, Bin Li 0091, Dakuan Lu, Haozhe Wang 0002, Yinghui Xu 0001, Xihe Qiu |
KDD (1) | 9 |
| 2026 | Curiosity Driven Knowledge Retrieval for Mobile AgentsabstractMobile agents have made progress toward reliable smartphone automation, yet performance in complex applications remains limited by incomplete knowledge and weak generalization to unseen environments. We introduce a curiosity driven knowledge retrieval framework that formalizes uncertainty during execution as a curiosity score. When this score exceeds a threshold, the system retrieves external information from documentation, code repositories, and historical trajectories. Retrieved content is organized into structured AppCards, which encode functional semantics, parameter conventions, interface mappings, and interaction patterns. During execution, an enhanced agent selectively integrates relevant AppCards into its reasoning process, thereby compensating for knowledge blind spots and improving planning reliability. Evaluation on the AndroidWorld benchmark shows consistent improvements across backbones, with an average gain of six percentage points and a new state of the art success rate of 88.8% when combined with GPT-5. Analysis indicates that AppCards are particularly effective for multi step and cross application tasks, while improvements depend on the backbone model. Case studies further confirm that AppCards reduce ambiguity, shorten exploration, and support stable execution trajectories. Task trajectories are publicly available at https://lisalsj.github.io/Droidrun-appcard/. Xiaoyu Tan, Shahir Ali, Niels Schmidt, Gengchen Ma, Xihe Qiu |
WWW | 6 |
| 2025 | Struct-X: Enhancing the Reasoning Capabilities of Large Language Models in Structured Data Scenarios
Xiaoyu Tan, Haoyu Wang 0011, Xihe Qiu, Leijun Cheng, Yinghui Xu 0001, Yuan Qi 0001 |
KDD (1) | 3 |
| 2024 | ILTS: Inducing Intention Propagation in Decentralized Multi-Agent Tasks with Large Language Models
Xihe Qiu, Haoyu Wang 0011, Xiaoyu Tan, Chao Qu |
CIKM | 1 |
| 2024 | Enhancing Personalized Headline Generation via Offline Goal-conditioned Reinforcement Learning with Large Language ModelsabstractRecently, significant advancements have been made in Large Language Models (LLMs) through the implementation of various alignment techniques. These techniques enable LLMs to generate highly tailored content in response to diverse user instructions. Consequently, LLMs have the potential to serve as robust, customizable recommendation systems in the field of content recommendation. However, using LLMs with user individual information and online exploration remains a challenge, which are important perspectives in developing personalized news headline generation algorithms. In this paper, we propose a novel framework to generate personalized news headlines using LLMs with extensive online exploration. The proposed approach involves initially training an offline goal-conditioned policy using supervised learning. Subsequently, online exploration is employed to collect new data for the next training iteration. Results from simulations, experiments, and real-word scenario demonstrate that our framework achieves outstanding performance on established benchmarks and can effectively generate personalized headlines under different reward settings. By treating the LLM as a goal-conditioned agent, the model can perform online exploration by modifying the goals without frequently retraining the model. To the best of our knowledge, this work represents the first investigation into the capability of LLMs to generate customized news headlines with goal-conditioned reinforcement learning via supervised learning within LLMs. Xiaoyu Tan, Leijun Cheng, Xihe Qiu, Shaojie Shi, Yinghui Xu 0001, Yuan Qi 0001 |
KDD | 3 |
| 2024 | Enhancing Task Performance in Continual Instruction Fine-tuning Through Format UniformityabstractIn recent advancements, large language models (LLMs) have demonstrated remarkable capabilities in diverse tasks, primarily through interactive question-answering with humans. This development marks significant progress towards artificial general intelligence (AGI). Despite their superior performance, LLMs often exhibit limitations when adapted to domain-specific tasks through instruction fine-tuning (IF). The primary challenge lies in the discrepancy between the data distribution in general and domain-specific contexts, leading to suboptimal accuracy in specialized tasks. To address this, continual instruction fine-tuning (CIF), particularly supervised fine-tuning (SFT), on targeted domain-specific instruction datasets is necessary. Our ablation study reveals that the structure of these instruction datasets critically influences CIF performance, with substantial data distributional shifts resulting in notable performance degradation. In this paper, we introduce a novel framework that enhances CIF by promoting format uniformity. We assess our approach using the Llama2 chat model across various domain-specific instruction datasets. The results demonstrate not only an improvement in task-specific performance under CIF but also a reduction in catastrophic forgetting (CF). This study contributes to the optimization of LLMs for domain-specific applications, highlighting the significance of data structure and distribution in CIF. Xiaoyu Tan, Leijun Cheng, Xihe Qiu, Shaojie Shi, Yinghui Xu 0001, Yuan Qi 0001 |
SIGIR | 3 |
| 2023 | CRNN-SA: A Network Intrusion Detection Method Based on Deep Learning
Wanxiao Liu, Jue Chen 0001, Xihe Qiu |
ADMA (2) | 3 |
| 2022 | A model-based hybrid soft actor-critic deep reinforcement learning algorithm for optimal ventilator settings
Shaotao Chen, Xihe Qiu, Xiaoyu Tan, Zhijun Fang 0001, Yaochu Jin |
Inf. Sci. | 2 |