Yinghui Xu 0001

dblp:15/2775-1 · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
14since 2021 · last 2026
0009-0002-7346-2794ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AURORA: Automated Training Framework of Universal Process Reward Models via Ensemble Prompting and Reverse Verification
abstract
The reasoning capabilities of advanced large language models (LLMs) like o1 have revolutionized artificial intelligence applications. Nevertheless, evaluating and optimizing complex reasoning processes remain significant challenges due to diverse policy distributions and the inherent limitations of human effort and accuracy. In this paper, we present AURORA, a novel automated framework for training universal process reward models (PRMs) using ensemble prompting and reverse verification. The framework employs a two-phase approach: First, it uses diverse prompting strategies and ensemble methods to perform automated annotation and evaluation of processes, ensuring robust assessments for reward learning. Second, it leverages practical reference answers for reverse verification, enhancing the model's ability to validate outputs and improving training accuracy. To assess the framework's performance, we extend beyond the existing ProcessBench benchmark by introducing UniversalBench, which evaluates reward predictions across full trajectories under diverse policy distribtion with long Chain-of-Thought (CoT) outputs. Experimental results demonstrate that AURORA enhances process evaluation accuracy, improves PRMs' accuracy for diverse policy distributions and long-CoT responses.
Xiaoyu Tan, Tianchu Yao, Chao Qu, Bin Li 0091, Dakuan Lu, Haozhe Wang 0002, Yinghui Xu 0001, Xihe Qiu
KDD (1)8
2025 CogniDual Framework: Self-Training Large Language Models within a Dual-System Theoretical Framework for Improving Cognitive Tasks
abstract
Cognitive psychology investigates perception, attention, memory, language, problem-solving, decision-making, and reasoning. Kahneman’s dual-system theory elucidates the human decision-making process, distinguishing between the rapid, intuitive System 1 and the deliberative, rational System 2. Recent advancements have positioned large language Models (LLMs) as formidable tools nearing human-level proficiency in various cognitive tasks. Nonetheless, the presence of a dual-system framework analogous to human cognition in LLMs remains unexplored. This study introduces the CogniDual Framework for LLMs (CFLLMs), designed to assess whether LLMs can, through self-training, evolve from deliberate deduction to intuitive responses, thereby emulating the human process of acquiring and mastering new information. Our findings reveal the cognitive mechanisms behind LLMs’ response generation, enhancing our understanding of their capabilities in cognitive psychology. Practically, self-trained models can provide faster responses to certain queries, reducing computational demands during inference.
Yongxin Deng, Xihe Qiu, Xiaoyu Tan, Chao Qu, Yinghui Xu 0001
ICASSP7
2025 An Attentive Dual-Encoder Framework Leveraging Multimodal Visual and Semantic Information for Automatic OSAHS Diagnosis
abstract
Obstructive sleep apnea-hypopnea syndrome (OS-AHS) is a common sleep disorder caused by upper airway blockage, leading to oxygen deprivation and disrupted sleep. Traditional diagnosis using polysomnography (PSG) is expensive, time-consuming, and uncomfortable. Existing deep learning methods using facial image analysis lack accuracy due to poor facial feature capture and limited sample sizes. To address this, we propose a multimodal dual encoder model that integrates visual and language inputs for automated OSAHS diagnosis. The model balances data using randomOverSampler (ROS), extracts key facial features with attention grids, and converts basic physiological data into meaningful text. Cross attention combines image and text data for better feature extraction, and ordered regression loss ensures stable learning. Our approach improves diagnostic efficiency and accuracy, achieving 91.3% top-1 accuracy in a four class severity classification task, demonstrating state of the art performance. Code is available at https://github.com/luboyan6/VTA-OSAHS.
Yingchen Wei, Xihe Qiu, Xiaoyu Tan, Yinghui Xu 0001, Yuan Qi 0001
ICASSP6
2025 Leave the Bias in Bias: Mitigating the Label Noise Effects in Continual Visual Instruction Fine-Tuning
abstract
In recent years, multimodal large language models (MLLMs) with vision processing capability have shown substantial advancements, excelling particularly in interpreting general images. Their application in domain-specific tasks, like those in the medical fields, is further enhanced through continuous visual instruction fine-tuning (CVIF). Despite these advancements, a significant challenge arises from label noise encountered during the collection of domain-specific data. Our studies reveal that this label noise can adversely affect the learning of vision projection embeddings and contribute to inaccuracies in LLMs’ fine-tuning, often leading to hallucinations. In this paper, we introduce a novel framework designed to minimize the impact of label noise. Our approach focuses on stabilizing the learning of vision embeddings and reducing the effect of label noise through the inherent semantic understanding of uncertainty in LLMs. Extensive experiments demonstrate that our framework maintains robust performance in general visual question-answer (VQA) tasks while showing significant effectiveness in medical VQA tasks. To the best of our knowledge, this is the first study to specifically address and analyze the impact of label noise in CVIF.
Xiaoyu Tan, Teqi Hao, Xihe Qiu, Shaojie Shi, Yinghui Xu 0001, Yuan Qi 0001
ICME7
2025 HeStIa: Asynchronous Embodied Dynamic Locomotion Learning for Walking Robots through Multimodal Large Language Models
abstract
The control of locomotion in walking robots with various architectural designs presents significant challenges. While existing approaches primarily rely on low-level state information and isolated visual features, lacking the high-level semantic understanding that humans use to reason about movement and posture, we propose HeStIa, a novel framework that bridges visual perception, natural language understanding, and robotic control through multimodal learning. Our framework leverages multimodal large language models (MLLMs) to establish a semantic bridge between visual observations and motion control, enabling robots to understand and adjust their locomotion through both visual and linguistic modalities. By leveraging multimodal large language models (MLLMs), HeStIa establishes a semantic connection between visual observations and motion control, enabling robots to comprehend and adapt their locomotion through both visual and linguistic modalities. Our approach extracts spatiotemporal visual features from robot movements and transforms them into a cross-modal embedding space shared with textual descriptions. HeStIa incorporates an innovative vision-language-motion fusion mechanism to provide informed, context-aware feedback during the dynamic learning process. Through an asynchronous design, HeStIa effectively mitigates the inference delays typically associated with MLLMs while maintaining real-time performance in dynamic scenarios. The cross-modal representations learned by HeStIa facilitate more intuitive and efficient locomotion learning by grounding visual observations in natural language descriptions. Our comprehensive evaluation shows substantial improvements in motion naturalness, stability, and adaptability across diverse environmental conditions.
Xiaoyu Tan, Haoyu Wang 0011, Yinghui Xu 0001, Xihe Qiu
IROS4
2025 Struct-X: Enhancing the Reasoning Capabilities of Large Language Models in Structured Data Scenarios
Xiaoyu Tan, Haoyu Wang 0011, Xihe Qiu, Leijun Cheng, Yinghui Xu 0001, Yuan Qi 0001
KDD (1)7
2025 MVP-LLMs: Optimizing Intervention Timing and Subsequent Decision Support for Mechanical Ventilation Parameter Control Using Large Language Models
Teqi Hao, Xiaoyu Tan, Bin Li 0091, Chao Qu, Yinghui Xu 0001, Xihe Qiu
MICCAI (5)6
2025 Prolog-Driven Rule-Based Diagnostics with Large Language Models for Precise Clinical Decision Support
Xiaoyu Tan, Bin Li 0091, Weidi Xu, Chao Qu, Yinghui Xu 0001, Yuan Qi 0001, Xihe Qiu
MICCAI (10)6
2025 ORIGAMISPACE: Benchmarking Multimodal LLMs in Multi-Step Spatial Reasoning with Mathematical Constraints
abstract
Spatial reasoning is a key capability in the field of artificial intelligence, especially crucial in areas such as robotics, computer vision, and natural language understanding. However, evaluating the ability of multimodal large language models (MLLMs) in complex spatial reasoning still faces challenges, particularly in scenarios requiring multi-step reasoning and precise mathematical constraints. This paper introduces ORIGAMISPACE, a new dataset and benchmark designed to evaluate the multi-step spatial reasoning ability and the capacity to handle mathematical constraints of MLLMs through origami tasks. The dataset contains 350 data instances, each comprising a strictly formatted crease pattern (CP diagram), the Compiled Flat Pattern, the complete Folding Process, and the final Folded Shape Image. We propose four evaluation tasks: Pattern Prediction, Multi-step Spatial Reasoning, Spatial Relationship Prediction, and End-to-End CP Code Generation. For the CP code generation task, we design an interactive environment and explore the possibility of using reinforcement learning methods to train MLLMs. Through experiments on existing MLLMs, we initially reveal the strengths and weaknesses of these models in handling complex spatial reasoning tasks.
Rui Xu 0026, Dakuan Lu, Zicheng Zhao, Xiaoyu Tan, Xintao Wang 0001, Jiangjie Chen, Yinghui Xu 0001
NeurIPS8
2024 Robust Deep Hawkes Process Under Label Noise of Both Event and Occurrence
abstract
Integrating deep neural networks with the Hawkes process has significantly improved predictive capabilities in finance, health informatics, and information technology. Nevertheless, these models often face challenges in real-world settings, particularly due to substantial label noise. This issue is of significant concern in the medical field, where label noise can arise from delayed updates in electronic medical records or misdiagnoses, leading to increased prediction risks. Our research indicates that deep Hawkes process models exhibit reduced robustness when dealing with label noise, particularly when it affects both event types and timing. To address these challenges, we first investigate the influence of label noise in approximated intensity functions and present a novel framework, the Robust Deep Hawkes Process (RDHP), to overcome the impact of label noise on the intensity function of Hawkes models, considering both the events and their occurrences. We tested RDHP using multiple open-source benchmarks with synthetic noise and conducted a case study on obstructive sleep apnea-hypopnea syndrome (OSAHS) in a real-world setting with inherent label noise. The results demonstrate that RDHP can effectively perform classification and regression tasks, even in the presence of noise related to events and their timing. To the best of our knowledge, this is the first study to successfully address both event and time label noise in deep Hawkes process models, offering a promising solution for medical applications, specifically in diagnosing OSAHS.
Xiaoyu Tan, Bin Li 0091, Xihe Qiu, Yinghui Xu 0001
ECAI5
2024 Hybrid Directional Graph Neural Network for Molecules
abstract
Equivariant message passing neural networks have emerged as the prevailing approach for predicting chemical properties of molecules due to their ability to leverage translation and rotation symmetries, resulting in a strong inductive bias. However, the equivariant operations in each layer can impose excessive constraints on the function form and network flexibility. To address these challenges, we introduce a novel network called the Hybrid Directional Graph Neural Network (HDGNN), which effectively combines strictly equivariant operations with learnable modules. We evaluate the performance of HDGNN on the QM9 dataset and the IS2RE dataset of OC20, demonstrating its state-of-the-art performance on several tasks and competitive performance on others. Our code is anonymously released on https://github.com/ajy112/HDGNN.
Junyi An, Chao Qu, Fenglei Cao, Yinghui Xu 0001, Yuan Qi 0001, Furao Shen
ICLR5
2024 Enhancing Personalized Headline Generation via Offline Goal-conditioned Reinforcement Learning with Large Language Models
abstract
Recently, significant advancements have been made in Large Language Models (LLMs) through the implementation of various alignment techniques. These techniques enable LLMs to generate highly tailored content in response to diverse user instructions. Consequently, LLMs have the potential to serve as robust, customizable recommendation systems in the field of content recommendation. However, using LLMs with user individual information and online exploration remains a challenge, which are important perspectives in developing personalized news headline generation algorithms. In this paper, we propose a novel framework to generate personalized news headlines using LLMs with extensive online exploration. The proposed approach involves initially training an offline goal-conditioned policy using supervised learning. Subsequently, online exploration is employed to collect new data for the next training iteration. Results from simulations, experiments, and real-word scenario demonstrate that our framework achieves outstanding performance on established benchmarks and can effectively generate personalized headlines under different reward settings. By treating the LLM as a goal-conditioned agent, the model can perform online exploration by modifying the goals without frequently retraining the model. To the best of our knowledge, this work represents the first investigation into the capability of LLMs to generate customized news headlines with goal-conditioned reinforcement learning via supervised learning within LLMs.
Xiaoyu Tan, Leijun Cheng, Xihe Qiu, Shaojie Shi, Yinghui Xu 0001, Yuan Qi 0001
KDD7
2024 Enhancing Task Performance in Continual Instruction Fine-tuning Through Format Uniformity
abstract
In recent advancements, large language models (LLMs) have demonstrated remarkable capabilities in diverse tasks, primarily through interactive question-answering with humans. This development marks significant progress towards artificial general intelligence (AGI). Despite their superior performance, LLMs often exhibit limitations when adapted to domain-specific tasks through instruction fine-tuning (IF). The primary challenge lies in the discrepancy between the data distribution in general and domain-specific contexts, leading to suboptimal accuracy in specialized tasks. To address this, continual instruction fine-tuning (CIF), particularly supervised fine-tuning (SFT), on targeted domain-specific instruction datasets is necessary. Our ablation study reveals that the structure of these instruction datasets critically influences CIF performance, with substantial data distributional shifts resulting in notable performance degradation. In this paper, we introduce a novel framework that enhances CIF by promoting format uniformity. We assess our approach using the Llama2 chat model across various domain-specific instruction datasets. The results demonstrate not only an improvement in task-specific performance under CIF but also a reduction in catastrophic forgetting (CF). This study contributes to the optimization of LLMs for domain-specific applications, highlighting the significance of data structure and distribution in CIF.
Xiaoyu Tan, Leijun Cheng, Xihe Qiu, Shaojie Shi, Yinghui Xu 0001, Yuan Qi 0001
SIGIR7
2023 Provably Invariant Learning without Domain Information
abstract
Typical machine learning applications always assume the data follows independent and identically distributed (IID) assumptions. In contrast, this assumption is frequently violated in real-world circumstances, leading to the Out-of-Distribution (OOD) generalization problem and a major drop in model robustness. To mitigate this issue, the invariant learning technique is leveraged to distinguish between spurious features and invariant features among all input features and to train the model purely on the basis of the invariant features. Numerous invariant learning strategies imply that the training data should contain domain information. Such information includes the environment index or auxiliary information acquired from prior knowledge. However, acquiring these information is typically impossible in practice. In this study, we present TIVA for environment-independent invariance learning, which requires no environment-specific information in training data. We discover and prove that, given certain mild data conditions, it is possible to train an environment partitioning policy based on attributes that are independent of the targets and then conduct invariant risk minimization. We examine our method in comparison to other baseline methods, which demonstrate superior performance and excellent robustness under OOD, using multiple benchmarks.
Xiaoyu Tan, Lin Yong, Shengyu Zhu 0001, Chao Qu, Xihe Qiu, Yinghui Xu 0001, Peng Cui 0001, Yuan Qi 0001
ICML6