EDBT 2026 Demo / reviewers in the wild / expert
Yinghui Xu 0001
dblp:15/2775-1
· DBLP profile ↗
4ranked-venue papers in the field
0as first author
4since 2021 · last 2026
0009-0002-7346-2794ORCID · conflict
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 3Information Retrieval & Web Search · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AURORA: Automated Training Framework of Universal Process Reward Models via Ensemble Prompting and Reverse VerificationabstractThe reasoning capabilities of advanced large language models (LLMs) like o1 have revolutionized artificial intelligence applications. Nevertheless, evaluating and optimizing complex reasoning processes remain significant challenges due to diverse policy distributions and the inherent limitations of human effort and accuracy. In this paper, we present AURORA, a novel automated framework for training universal process reward models (PRMs) using ensemble prompting and reverse verification. The framework employs a two-phase approach: First, it uses diverse prompting strategies and ensemble methods to perform automated annotation and evaluation of processes, ensuring robust assessments for reward learning. Second, it leverages practical reference answers for reverse verification, enhancing the model's ability to validate outputs and improving training accuracy. To assess the framework's performance, we extend beyond the existing ProcessBench benchmark by introducing UniversalBench, which evaluates reward predictions across full trajectories under diverse policy distribtion with long Chain-of-Thought (CoT) outputs. Experimental results demonstrate that AURORA enhances process evaluation accuracy, improves PRMs' accuracy for diverse policy distributions and long-CoT responses. Xiaoyu Tan, Tianchu Yao, Chao Qu, Bin Li 0091, Dakuan Lu, Haozhe Wang 0002, Yinghui Xu 0001, Xihe Qiu |
KDD (1) | 8 |
| 2025 | Struct-X: Enhancing the Reasoning Capabilities of Large Language Models in Structured Data Scenarios
Xiaoyu Tan, Haoyu Wang 0011, Xihe Qiu, Leijun Cheng, Yinghui Xu 0001, Yuan Qi 0001 |
KDD (1) | 7 |
| 2024 | Enhancing Personalized Headline Generation via Offline Goal-conditioned Reinforcement Learning with Large Language ModelsabstractRecently, significant advancements have been made in Large Language Models (LLMs) through the implementation of various alignment techniques. These techniques enable LLMs to generate highly tailored content in response to diverse user instructions. Consequently, LLMs have the potential to serve as robust, customizable recommendation systems in the field of content recommendation. However, using LLMs with user individual information and online exploration remains a challenge, which are important perspectives in developing personalized news headline generation algorithms. In this paper, we propose a novel framework to generate personalized news headlines using LLMs with extensive online exploration. The proposed approach involves initially training an offline goal-conditioned policy using supervised learning. Subsequently, online exploration is employed to collect new data for the next training iteration. Results from simulations, experiments, and real-word scenario demonstrate that our framework achieves outstanding performance on established benchmarks and can effectively generate personalized headlines under different reward settings. By treating the LLM as a goal-conditioned agent, the model can perform online exploration by modifying the goals without frequently retraining the model. To the best of our knowledge, this work represents the first investigation into the capability of LLMs to generate customized news headlines with goal-conditioned reinforcement learning via supervised learning within LLMs. Xiaoyu Tan, Leijun Cheng, Xihe Qiu, Shaojie Shi, Yinghui Xu 0001, Yuan Qi 0001 |
KDD | 7 |
| 2024 | Enhancing Task Performance in Continual Instruction Fine-tuning Through Format UniformityabstractIn recent advancements, large language models (LLMs) have demonstrated remarkable capabilities in diverse tasks, primarily through interactive question-answering with humans. This development marks significant progress towards artificial general intelligence (AGI). Despite their superior performance, LLMs often exhibit limitations when adapted to domain-specific tasks through instruction fine-tuning (IF). The primary challenge lies in the discrepancy between the data distribution in general and domain-specific contexts, leading to suboptimal accuracy in specialized tasks. To address this, continual instruction fine-tuning (CIF), particularly supervised fine-tuning (SFT), on targeted domain-specific instruction datasets is necessary. Our ablation study reveals that the structure of these instruction datasets critically influences CIF performance, with substantial data distributional shifts resulting in notable performance degradation. In this paper, we introduce a novel framework that enhances CIF by promoting format uniformity. We assess our approach using the Llama2 chat model across various domain-specific instruction datasets. The results demonstrate not only an improvement in task-specific performance under CIF but also a reduction in catastrophic forgetting (CF). This study contributes to the optimization of LLMs for domain-specific applications, highlighting the significance of data structure and distribution in CIF. Xiaoyu Tan, Leijun Cheng, Xihe Qiu, Shaojie Shi, Yinghui Xu 0001, Yuan Qi 0001 |
SIGIR | 7 |