Xiaoyu Tan

dblp:194/9401 · DBLP profile ↗
← Back
41ranked-venue papers
10as first author
39since 2021 · last 2026
0000-0003-3555-7143ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 6 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 10 since 2021Databases, data management, data science and information retrieval · 8 · 4 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 5 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author
YearPublicationVenuePosition
2026 From Answers to Arguments: Toward Trustworthy Clinical Diagnostic Reasoning with Toulmin-Guided Curriculum Goal-Conditioned Learning
abstract
The integration of Large Language Models (LLMs) into clinical decision support is critically obstructed by their opaque and often unreliable reasoning.In the high-stakes domain of healthcare, correct answers alone are insufficient; clinical practice demands full transparency to ensure patient safety and enable professional accountability.A pervasive and dangerous weakness of current LLMs is their tendency to produce "correct answers through flawed reasoning."This issue is far more than a minor academic flaw; such process errors signal a fundamental lack of robust understanding, making the model prone to broader hallucinations and unpredictable failures when faced with real-world clinical complexity.In this paper, we establish a framework for trustworthy clinical argumentation by adapting the Toulmin model to the diagnostic process.We propose a novel training pipeline: Curriculum Goal-Conditioned Learning (CGCL), designed to progressively train LLM to generate diagnostic arguments that explicitly follow this Toulmin structure.CGCL's progressive three-stage curriculum systematically builds a solid clinical argument: (1) extracting facts and generating differential diagnoses; (2) justifying a core hypothesis while rebutting alternatives; and (3) synthesizing the analysis into a final, qualified conclusion.We validate CGCL using T-Eval, a quantitative framework measuring the integrity of the diagnosis reasoning.Experiments show that our method achieves diagnostic accuracy and reasoning quality comparable to resourceintensive Reinforcement Learning (RL) methods, while offering a more stable and efficient training pipeline.1
Chen Zhan, Xiaoyu Tan, Gengchen Ma, Yujie Xiong, Xihe Qiu
ACL (1)2
2026 AURORA: Automated Training Framework of Universal Process Reward Models via Ensemble Prompting and Reverse Verification
abstract
The reasoning capabilities of advanced large language models (LLMs) like o1 have revolutionized artificial intelligence applications. Nevertheless, evaluating and optimizing complex reasoning processes remain significant challenges due to diverse policy distributions and the inherent limitations of human effort and accuracy. In this paper, we present AURORA, a novel automated framework for training universal process reward models (PRMs) using ensemble prompting and reverse verification. The framework employs a two-phase approach: First, it uses diverse prompting strategies and ensemble methods to perform automated annotation and evaluation of processes, ensuring robust assessments for reward learning. Second, it leverages practical reference answers for reverse verification, enhancing the model's ability to validate outputs and improving training accuracy. To assess the framework's performance, we extend beyond the existing ProcessBench benchmark by introducing UniversalBench, which evaluates reward predictions across full trajectories under diverse policy distribtion with long Chain-of-Thought (CoT) outputs. Experimental results demonstrate that AURORA enhances process evaluation accuracy, improves PRMs' accuracy for diverse policy distributions and long-CoT responses.
Xiaoyu Tan, Tianchu Yao, Chao Qu, Bin Li 0091, Dakuan Lu, Haozhe Wang 0002, Yinghui Xu 0001, Xihe Qiu
KDD (1)1
2026 Curiosity Driven Knowledge Retrieval for Mobile Agents
abstract
Mobile agents have made progress toward reliable smartphone automation, yet performance in complex applications remains limited by incomplete knowledge and weak generalization to unseen environments. We introduce a curiosity driven knowledge retrieval framework that formalizes uncertainty during execution as a curiosity score. When this score exceeds a threshold, the system retrieves external information from documentation, code repositories, and historical trajectories. Retrieved content is organized into structured AppCards, which encode functional semantics, parameter conventions, interface mappings, and interaction patterns. During execution, an enhanced agent selectively integrates relevant AppCards into its reasoning process, thereby compensating for knowledge blind spots and improving planning reliability. Evaluation on the AndroidWorld benchmark shows consistent improvements across backbones, with an average gain of six percentage points and a new state of the art success rate of 88.8% when combined with GPT-5. Analysis indicates that AppCards are particularly effective for multi step and cross application tasks, while improvements depend on the backbone model. Case studies further confirm that AppCards reduce ambiguity, shorten exploration, and support stable execution trajectories. Task trajectories are publicly available at https://lisalsj.github.io/Droidrun-appcard/.
Xiaoyu Tan, Shahir Ali, Niels Schmidt, Gengchen Ma, Xihe Qiu
WWW2
2026 What the doodle: Sketching mental pictures with decoding brain signals, language understanding, and AI creativity
Gengchen Ma, Haoyu Wang 0011, Fenghao Sun, Chen Zhan, Xiaoyu Tan, Xihe Qiu
Expert Syst. Appl.5
2026 MIMAR-OSA: Enhancing obstructive sleep apnea diagnosis through multimodal data integration and missing modality reconstruction
Xihe Qiu, Yingchen Wei, Xiaoyu Tan, Weidi Xu, Jingru Ma, Zhijun Fang 0001
Pattern Recognit.3
2026 MLLoRA: Leveraging meta-learning and LoRA for efficient multi-task fine-tuning in LLM-driven medical applications
Xihe Qiu, Siyue Shao, Qian Du, Xiaoyu Tan
Pattern Recognit. Lett.5
2026 ReMALIS: Inference-Guided Intention Propagation for Multiagent Stochastic Task Coordination With Large Language Models in Complex Networks Domains
Xihe Qiu, Haoyu Wang 0011, Xiaoyu Tan, Yujie Xiong, Zhijun Fang 0001
IEEE Trans. Comput. Soc. Syst.3
2025 CogniDual Framework: Self-Training Large Language Models within a Dual-System Theoretical Framework for Improving Cognitive Tasks
abstract
Cognitive psychology investigates perception, attention, memory, language, problem-solving, decision-making, and reasoning. Kahneman’s dual-system theory elucidates the human decision-making process, distinguishing between the rapid, intuitive System 1 and the deliberative, rational System 2. Recent advancements have positioned large language Models (LLMs) as formidable tools nearing human-level proficiency in various cognitive tasks. Nonetheless, the presence of a dual-system framework analogous to human cognition in LLMs remains unexplored. This study introduces the CogniDual Framework for LLMs (CFLLMs), designed to assess whether LLMs can, through self-training, evolve from deliberate deduction to intuitive responses, thereby emulating the human process of acquiring and mastering new information. Our findings reveal the cognitive mechanisms behind LLMs’ response generation, enhancing our understanding of their capabilities in cognitive psychology. Practically, self-trained models can provide faster responses to certain queries, reducing computational demands during inference.
Yongxin Deng, Xihe Qiu, Xiaoyu Tan, Chao Qu, Yinghui Xu 0001
ICASSP3
2025 An Attentive Dual-Encoder Framework Leveraging Multimodal Visual and Semantic Information for Automatic OSAHS Diagnosis
abstract
Obstructive sleep apnea-hypopnea syndrome (OS-AHS) is a common sleep disorder caused by upper airway blockage, leading to oxygen deprivation and disrupted sleep. Traditional diagnosis using polysomnography (PSG) is expensive, time-consuming, and uncomfortable. Existing deep learning methods using facial image analysis lack accuracy due to poor facial feature capture and limited sample sizes. To address this, we propose a multimodal dual encoder model that integrates visual and language inputs for automated OSAHS diagnosis. The model balances data using randomOverSampler (ROS), extracts key facial features with attention grids, and converts basic physiological data into meaningful text. Cross attention combines image and text data for better feature extraction, and ordered regression loss ensures stable learning. Our approach improves diagnostic efficiency and accuracy, achieving 91.3% top-1 accuracy in a four class severity classification task, demonstrating state of the art performance. Code is available at https://github.com/luboyan6/VTA-OSAHS.
Yingchen Wei, Xihe Qiu, Xiaoyu Tan, Yinghui Xu 0001, Yuan Qi 0001
ICASSP3
2025 Refine Knowledge of Large Language Models via Adaptive Contrastive Learning
abstract
How to alleviate the hallucinations of Large Language Models (LLMs) has always been the fundamental goal pursued by the LLMs research community. Looking through numerous hallucination-related studies, a mainstream category of methods is to reduce hallucinations by optimizing the knowledge representation of LLMs to change their output. Considering that the core focus of these works is the knowledge acquired by models, and knowledge has long been a central theme in human societal progress, we believe that the process of models refining knowledge can greatly benefit from the way humans learn. In our work, by imitating the human learning process, we design an Adaptive Contrastive Learning strategy. Our method flexibly constructs different positive and negative samples for contrastive learning based on LLMs' actual mastery of knowledge. This strategy helps LLMs consolidate the correct knowledge they already possess, deepen their understanding of the correct knowledge they have encountered but not fully grasped, forget the incorrect knowledge they previously learned, and honestly acknowledge the knowledge they lack. Extensive experiments and detailed analyses on widely used datasets demonstrate the effectiveness and competitiveness of our method.
Haojing Huang 0001, Jiayi Kuang, Yangning Li, Shu-Yu Guo, Chao Qu, Xiaoyu Tan, Hai-Tao Zheng 0002, Ying Shen 0001, Philip S. Yu
ICLR7
2025 Leave the Bias in Bias: Mitigating the Label Noise Effects in Continual Visual Instruction Fine-Tuning
abstract
In recent years, multimodal large language models (MLLMs) with vision processing capability have shown substantial advancements, excelling particularly in interpreting general images. Their application in domain-specific tasks, like those in the medical fields, is further enhanced through continuous visual instruction fine-tuning (CVIF). Despite these advancements, a significant challenge arises from label noise encountered during the collection of domain-specific data. Our studies reveal that this label noise can adversely affect the learning of vision projection embeddings and contribute to inaccuracies in LLMs’ fine-tuning, often leading to hallucinations. In this paper, we introduce a novel framework designed to minimize the impact of label noise. Our approach focuses on stabilizing the learning of vision embeddings and reducing the effect of label noise through the inherent semantic understanding of uncertainty in LLMs. Extensive experiments demonstrate that our framework maintains robust performance in general visual question-answer (VQA) tasks while showing significant effectiveness in medical VQA tasks. To the best of our knowledge, this is the first study to specifically address and analyze the impact of label noise in CVIF.
Xiaoyu Tan, Teqi Hao, Xihe Qiu, Shaojie Shi, Yinghui Xu 0001, Yuan Qi 0001
ICME1
2025 One Example Shown, Many Concepts Known! Counterexample-Driven Conceptual Reasoning in Mathematical LLMs
abstract
Leveraging mathematical Large Language Models (LLMs) for proof generation is a fundamental topic in LLMs research. We argue that the ability of current LLMs to prove statements largely depends on whether they have encountered the relevant proof process during training. This reliance limits their deeper understanding of mathematical theorems and related concepts. Inspired by the pedagogical method of "proof by counterexamples" commonly used in human mathematics education, our work aims to enhance LLMs’ ability to conduct mathematical reasoning and proof through counterexamples. Specifically, we manually create a high-quality, university-level mathematical benchmark, COUNTERMATH, which requires LLMs to prove mathematical statements by providing counterexamples, thereby assessing their grasp of mathematical concepts. Additionally, we develop a data engineering framework to automatically obtain training data for further model improvement. Extensive experiments and detailed analyses demonstrate that COUNTERMATH is challenging, indicating that LLMs, such as OpenAI o1, have insufficient counterexample-driven proof capabilities. Moreover, our exploration into model training reveals that strengthening LLMs’ counterexample-driven conceptual reasoning abilities is crucial for improving their overall mathematical capabilities. We believe that our work offers new perspectives on the community of mathematical LLMs.
Jiayi Kuang, Haojing Huang 0001, Zhikun Xu, Xinnian Liang, Wenlian Lu, Yangning Li, Xiaoyu Tan, Chao Qu, Ying Shen 0001, Hai-Tao Zheng 0002, Philip S. Yu
ICML9
2025 HeStIa: Asynchronous Embodied Dynamic Locomotion Learning for Walking Robots through Multimodal Large Language Models
abstract
The control of locomotion in walking robots with various architectural designs presents significant challenges. While existing approaches primarily rely on low-level state information and isolated visual features, lacking the high-level semantic understanding that humans use to reason about movement and posture, we propose HeStIa, a novel framework that bridges visual perception, natural language understanding, and robotic control through multimodal learning. Our framework leverages multimodal large language models (MLLMs) to establish a semantic bridge between visual observations and motion control, enabling robots to understand and adjust their locomotion through both visual and linguistic modalities. By leveraging multimodal large language models (MLLMs), HeStIa establishes a semantic connection between visual observations and motion control, enabling robots to comprehend and adapt their locomotion through both visual and linguistic modalities. Our approach extracts spatiotemporal visual features from robot movements and transforms them into a cross-modal embedding space shared with textual descriptions. HeStIa incorporates an innovative vision-language-motion fusion mechanism to provide informed, context-aware feedback during the dynamic learning process. Through an asynchronous design, HeStIa effectively mitigates the inference delays typically associated with MLLMs while maintaining real-time performance in dynamic scenarios. The cross-modal representations learned by HeStIa facilitate more intuitive and efficient locomotion learning by grounding visual observations in natural language descriptions. Our comprehensive evaluation shows substantial improvements in motion naturalness, stability, and adaptability across diverse environmental conditions.
Xiaoyu Tan, Haoyu Wang 0011, Yinghui Xu 0001, Xihe Qiu
IROS1
2025 Struct-X: Enhancing the Reasoning Capabilities of Large Language Models in Structured Data Scenarios
Xiaoyu Tan, Haoyu Wang 0011, Xihe Qiu, Leijun Cheng, Yinghui Xu 0001, Yuan Qi 0001
KDD (1)1
2025 MVP-LLMs: Optimizing Intervention Timing and Subsequent Decision Support for Mechanical Ventilation Parameter Control Using Large Language Models
Teqi Hao, Xiaoyu Tan, Bin Li 0091, Chao Qu, Yinghui Xu 0001, Xihe Qiu
MICCAI (5)2
2025 Robust Sleep Stage Prediction from Electroencephalogram with Label Noise Using Multimodal Large Language Models
Xihe Qiu, Chen Zhan, Gengchen Ma, Xiaoyu Tan
MICCAI (6)5
2025 Prolog-Driven Rule-Based Diagnostics with Large Language Models for Precise Clinical Decision Support
Xiaoyu Tan, Bin Li 0091, Weidi Xu, Chao Qu, Yinghui Xu 0001, Yuan Qi 0001, Xihe Qiu
MICCAI (10)1
2025 Atomic Thinking of LLMs: Decoupling and Exploring Mathematical Reasoning Abilities
abstract
Large Language Models (LLMs) have demonstrated outstanding performance in mathematical reasoning capabilities. However, we argue that current large-scale reasoning models primarily rely on scaling up training datasets with diverse mathematical problems and long thinking chains, which raises questions about whether LLMs genuinely acquire mathematical concepts and reasoning principles or merely remember the training data. In contrast, humans tend to break down complex problems into multiple fundamental atomic capabilities. Inspired by this, we propose a new paradigm for evaluating mathematical atomic capabilities. Our work categorizes atomic abilities into two dimensions: (1) field-specific abilities across four major mathematical fields, algebra, geometry, analysis, and topology, and (2) logical abilities at different levels, including conceptual understanding, forward multi-step reasoning with formal math language, and counterexample-driven backward reasoning. We propose corresponding training and evaluation datasets for each atomic capability unit, and conduct extensive experiments about how different atomic capabilities influence others, to explore the strategies to elicit the required specific atomic capability. Evaluation and experimental results on advanced models show many interesting discoveries and inspirations about the different performances of models on various atomic capabilities and the interactions between atomic capabilities. Our findings highlight the importance of decoupling mathematical intelligence into atomic components, providing new insights into model cognition and guiding the development of training strategies toward a more efficient, transferable, and cognitively grounded paradigm of "atomic thinking".
Jiayi Kuang, Haojing Huang 0001, Xinnian Liang, Zhikun Xu, Yangning Li, Xiaoyu Tan, Chao Qu, Meishan Zhang, Ying Shen 0001, Philip S. Yu
NeurIPS7
2025 ORIGAMISPACE: Benchmarking Multimodal LLMs in Multi-Step Spatial Reasoning with Mathematical Constraints
abstract
Spatial reasoning is a key capability in the field of artificial intelligence, especially crucial in areas such as robotics, computer vision, and natural language understanding. However, evaluating the ability of multimodal large language models (MLLMs) in complex spatial reasoning still faces challenges, particularly in scenarios requiring multi-step reasoning and precise mathematical constraints. This paper introduces ORIGAMISPACE, a new dataset and benchmark designed to evaluate the multi-step spatial reasoning ability and the capacity to handle mathematical constraints of MLLMs through origami tasks. The dataset contains 350 data instances, each comprising a strictly formatted crease pattern (CP diagram), the Compiled Flat Pattern, the complete Folding Process, and the final Folded Shape Image. We propose four evaluation tasks: Pattern Prediction, Multi-step Spatial Reasoning, Spatial Relationship Prediction, and End-to-End CP Code Generation. For the CP code generation task, we design an interactive environment and explore the possibility of using reinforcement learning methods to train MLLMs. Through experiments on existing MLLMs, we initially reveal the strengths and weaknesses of these models in handling complex spatial reasoning tasks.
Rui Xu 0026, Dakuan Lu, Zicheng Zhao, Xiaoyu Tan, Xintao Wang 0001, Jiangjie Chen, Yinghui Xu 0001
NeurIPS4
2025 An end-to-end audio classification framework with diverse features for obstructive sleep apnea-hypopnea syndrome diagnosis
Bin Li 0091, Xihe Qiu, Xiaoyu Tan, Zhijun Fang 0001
Appl. Intell.3
2025 An innovative contrastive learning approach to improve image recognition robustness and interpretability via simulated environmental perturbations
Leijun Cheng, Xihe Qiu, Xiaoyu Tan, Haoyu Wang 0011, Yujie Xiong
Eng. Appl. Artif. Intell.3
2025 Reward guidance for reinforcement learning tasks based on large language models: The LMGT framework
Yongxin Deng, Xihe Qiu, Jue Chen 0001, Xiaoyu Tan
Knowl. Based Syst.4
2025 LLM-GAODE: Large-language-model augmented neural ordinary differential equation network for video nystagmography classification
abstract
Benign paroxysmal positional vertigo (BPPV), a common type of vertigo with complex etiologies, is traditionally diagnosed using video nystagmography (VNG). Current automated methods lack diagnostic precision owing to subjective interpretation of eye movement characteristics. To address these challenges, we introduce a l arge l anguage m odel-augmented G ram-based a ttentive neural o rdinary d ifferential e quation ( LLM-GAODE ), an innovative and data-driven framework integrating eye-tracking technology with a Gram-based attention mechanism and a neural ordinary differential equation network to improve BPPV classification. Furthermore, when the neural network exhibits low confidence in its predictions, an LLM can supplement the process with advanced reasoning in natural language. LLM-GAODE was evaluated using an extensive VNG dataset provided by a collaborative university hospital. Results suggest that LLM-GAODE significantly outperforms existing benchmarks in trajectory classification for BPPV diagnosis. The framework enhances BPPV diagnostic accuracy and achieves state-of-the-art performance in open-source trajectory classification benchmarks. The code is available at https://github.com/XiheQiu/LLM-GAODE .
Xihe Qiu, Shaojie Shi, Bin Li 0091, Xiaoyu Tan, Yongbin Gao, Shuo Li 0001
Knowl. Based Syst.4
2024 ILTS: Inducing Intention Propagation in Decentralized Multi-Agent Tasks with Large Language Models
Xihe Qiu, Haoyu Wang 0011, Xiaoyu Tan, Chao Qu
CIKM3
2024 Robust Deep Hawkes Process Under Label Noise of Both Event and Occurrence
abstract
Integrating deep neural networks with the Hawkes process has significantly improved predictive capabilities in finance, health informatics, and information technology. Nevertheless, these models often face challenges in real-world settings, particularly due to substantial label noise. This issue is of significant concern in the medical field, where label noise can arise from delayed updates in electronic medical records or misdiagnoses, leading to increased prediction risks. Our research indicates that deep Hawkes process models exhibit reduced robustness when dealing with label noise, particularly when it affects both event types and timing. To address these challenges, we first investigate the influence of label noise in approximated intensity functions and present a novel framework, the Robust Deep Hawkes Process (RDHP), to overcome the impact of label noise on the intensity function of Hawkes models, considering both the events and their occurrences. We tested RDHP using multiple open-source benchmarks with synthetic noise and conducted a case study on obstructive sleep apnea-hypopnea syndrome (OSAHS) in a real-world setting with inherent label noise. The results demonstrate that RDHP can effectively perform classification and regression tasks, even in the presence of noise related to events and their timing. To the best of our knowledge, this is the first study to successfully address both event and time label noise in deep Hawkes process models, offering a promising solution for medical applications, specifically in diagnosing OSAHS.
Xiaoyu Tan, Bin Li 0091, Xihe Qiu, Yinghui Xu 0001
ECAI1
2024 Subequivariant Reinforcement Learning Framework for Coordinated Motion Control
abstract
Effective coordination is crucial for motion control with reinforcement learning, especially as the complexity of agents and their motions increases. However, many existing methods struggle to account for the intricate dependencies between joints. We introduce CoordiGraph, a novel architecture that leverages subequivariant principles from physics to enhance coordination of motion control with reinforcement learning. This method embeds the principles of equivariance as inherent patterns in the learning process under gravity influence, which aids in modeling the nuanced relationships between joints vital for motion control. Through extensive experimentation with sophisticated agents in diverse environments, we highlight the merits of our approach. Compared to current leading methods, CoordiGraph notably enhances generalization and sample efficiency.
Haoyu Wang 0011, Xiaoyu Tan, Xihe Qiu, Chao Qu
ICRA2
2024 Enhancing Personalized Headline Generation via Offline Goal-conditioned Reinforcement Learning with Large Language Models
abstract
Recently, significant advancements have been made in Large Language Models (LLMs) through the implementation of various alignment techniques. These techniques enable LLMs to generate highly tailored content in response to diverse user instructions. Consequently, LLMs have the potential to serve as robust, customizable recommendation systems in the field of content recommendation. However, using LLMs with user individual information and online exploration remains a challenge, which are important perspectives in developing personalized news headline generation algorithms. In this paper, we propose a novel framework to generate personalized news headlines using LLMs with extensive online exploration. The proposed approach involves initially training an offline goal-conditioned policy using supervised learning. Subsequently, online exploration is employed to collect new data for the next training iteration. Results from simulations, experiments, and real-word scenario demonstrate that our framework achieves outstanding performance on established benchmarks and can effectively generate personalized headlines under different reward settings. By treating the LLM as a goal-conditioned agent, the model can perform online exploration by modifying the goals without frequently retraining the model. To the best of our knowledge, this work represents the first investigation into the capability of LLMs to generate customized news headlines with goal-conditioned reinforcement learning via supervised learning within LLMs.
Xiaoyu Tan, Leijun Cheng, Xihe Qiu, Shaojie Shi, Yinghui Xu 0001, Yuan Qi 0001
KDD1
2024 Enhancing Task Performance in Continual Instruction Fine-tuning Through Format Uniformity
abstract
In recent advancements, large language models (LLMs) have demonstrated remarkable capabilities in diverse tasks, primarily through interactive question-answering with humans. This development marks significant progress towards artificial general intelligence (AGI). Despite their superior performance, LLMs often exhibit limitations when adapted to domain-specific tasks through instruction fine-tuning (IF). The primary challenge lies in the discrepancy between the data distribution in general and domain-specific contexts, leading to suboptimal accuracy in specialized tasks. To address this, continual instruction fine-tuning (CIF), particularly supervised fine-tuning (SFT), on targeted domain-specific instruction datasets is necessary. Our ablation study reveals that the structure of these instruction datasets critically influences CIF performance, with substantial data distributional shifts resulting in notable performance degradation. In this paper, we introduce a novel framework that enhances CIF by promoting format uniformity. We assess our approach using the Llama2 chat model across various domain-specific instruction datasets. The results demonstrate not only an improvement in task-specific performance under CIF but also a reduction in catastrophic forgetting (CF). This study contributes to the optimization of LLMs for domain-specific applications, highlighting the significance of data structure and distribution in CIF.
Xiaoyu Tan, Leijun Cheng, Xihe Qiu, Shaojie Shi, Yinghui Xu 0001, Yuan Qi 0001
SIGIR1
2024 Multivariate graph neural networks on enhancing syntactic and semantic for aspect-based sentiment analysis
Haoyu Wang 0011, Xihe Qiu, Xiaoyu Tan
Appl. Intell.3
2024 Balancing therapeutic effect and safety in ventilator parameter recommendation: An offline reinforcement learning approach
Xihe Qiu, Xiaoyu Tan
Eng. Appl. Artif. Intell.3
2024 Chain-of-LoRA: Enhancing the Instruction Fine-Tuning Performance of Low-Rank Adaptation on Diverse Instruction Set
abstract
Recently, large language models (LLMs) with conversational-style interaction, such as ChatGPT and Claude, have gained significant importance in the advancement of artificial general intelligence (AGI). However, the extensive resource requirements during pre-training, instruction fine-tuning (IF), and reinforcement learning through human feedback (RLHF) pose challenges, particularly for individuals and studios with limited resources. Moreover, sensitive data that cannot be deployed on remote training platforms or queried through APIs further exacerbates this issue. To address these limitations, researchers have introduced a parameter-efficient framework called low-rank adaptation (LoRA) for IF on LLMs. However, training individual LoRA networks faces capacity constraints and struggles to adapt to large domains with significant distributional shifts across different tasks. In this paper, we propose a novel framework called chain-of-LoRA to enhance the IF performance of LoRA. Our approach involves training a LoRA network to classify the instruction type and then utilizing task-specific LoRA networks to accomplish the respective tasks. By training multiple task-specific LoRA networks, we exploit a trade-off between performance and disk storage, leveraging the easily expandable and cost-effective nature of disk storage compared to precious graphical resources. Our experimental results demonstrate that our proposed framework achieves comparable performance to typical direct IF on LLMs.
Xihe Qiu, Teqi Hao, Shaojie Shi, Xiaoyu Tan, Yujie Xiong
IEEE Signal Process. Lett.4
2023 Bellman Meets Hawkes: Model-Based Reinforcement Learning via Temporal Point Processes
abstract
We consider a sequential decision making problem where the agent faces the environment characterized by the stochastic discrete events and seeks an optimal intervention policy such that its long-term reward is maximized. This problem exists ubiquitously in social media, finance and health informatics but is rarely investigated by the conventional research in reinforcement learning. To this end, we present a novel framework of the model-based reinforcement learning where the agent's actions and observations are asynchronous stochastic discrete events occurring in continuous-time. We model the dynamics of the environment by Hawkes process with external intervention control term and develop an algorithm to embed such process in the Bellman equation which guides the direction of the value gradient. We demonstrate the superiority of our method in both synthetic simulator and real-data experiments.
Chao Qu, Xiaoyu Tan, Siqiao Xue, Xiaoming Shi 0001, James Zhang, Hongyuan Mei
AAAI2
2023 Gram-based Attentive Neural Ordinary Differential Equations Network for Video Nystagmography Classification
abstract
Video nystagmography (VNG) is the diagnostic gold standard of benign paroxysmal positional vertigo (BPPV), which requires medical professionals to examine the direction, frequency, intensity, duration, and variation in the strength of nystagmus on a VNG video. This is a tedious process heavily influenced by the doctor’s experience, which is error-prone. Recent automatic VNG classification methods approach this problem from the perspective of video analysis without considering medical prior knowledge, resulting in unsatisfactory accuracy and limited diagnostic capability for nystagmographic types, thereby preventing their clinical application. In this paper, we propose an end-to-end data-driven novel BPPV diagnosis framework (TC-BPPV) by considering this problem as an eye trajectory classification problem due to the disease’s symptoms and experts’ prior knowledge. In this framework, we utilize an eye movement tracking system to capture the eye trajectory and propose the Gram-based attentive neural ordinary differential equations network (Gram-AODE) to perform classification. We validate our framework using the VNG dataset provided by the collaborative university hospital and achieve state-of-the-art performance. We also evaluate Gram-AODE on multiple open-source benchmarks to demonstrate its effectiveness in trajectory classification. Code is available at https://github.com/XiheQiu/Gram-AODE.
Xihe Qiu, Shaojie Shi, Xiaoyu Tan, Chao Qu, Zhijun Fang 0001, Yongbin Gao, Peixia Wu
ICCV3
2023 Provably Invariant Learning without Domain Information
abstract
Typical machine learning applications always assume the data follows independent and identically distributed (IID) assumptions. In contrast, this assumption is frequently violated in real-world circumstances, leading to the Out-of-Distribution (OOD) generalization problem and a major drop in model robustness. To mitigate this issue, the invariant learning technique is leveraged to distinguish between spurious features and invariant features among all input features and to train the model purely on the basis of the invariant features. Numerous invariant learning strategies imply that the training data should contain domain information. Such information includes the environment index or auxiliary information acquired from prior knowledge. However, acquiring these information is typically impossible in practice. In this study, we present TIVA for environment-independent invariance learning, which requires no environment-specific information in training data. We discover and prove that, given certain mild data conditions, it is possible to train an environment partitioning policy based on attributes that are independent of the targets and then conduct invariant risk minimization. We examine our method in comparison to other baseline methods, which demonstrate superior performance and excellent robustness under OOD, using multiple benchmarks.
Xiaoyu Tan, Lin Yong, Shengyu Zhu 0001, Chao Qu, Xihe Qiu, Yinghui Xu 0001, Peng Cui 0001, Yuan Qi 0001
ICML1
2023 An Attentive LSTM based approach for adverse drug reactions prediction
Jiahui Qian, Xihe Qiu, Xiaoyu Tan, Jue Chen 0001
Appl. Intell.3
2022 A Meta Reinforcement Learning Approach for Predictive Autoscaling in the Cloud
abstract
Predictive autoscaling (autoscaling with workload forecasting) is an important mechanism that supports autonomous adjustment of computing resources in accordance with fluctuating workload demands in the Cloud. In recent works, Reinforcement Learning (RL) has been introduced as a promising approach to learn the resource management policies to guide the scaling actions under the dynamic and uncertain cloud environment. However, RL methods face the following challenges in steering predictive autoscaling, such as lack of accuracy in decision-making, inefficient sampling and significant variability in workload patterns that may cause policies to fail at test time. To this end, we propose an end-to-end predictive meta model-based RL algorithm, aiming to optimally allocate resource to maintain a stable CPU utilization level, which incorporates a specially-designed deep periodic workload prediction model as the input and embeds the Neural Process [11, 16] to guide the learning of the optimal scaling actions over numerous application services in the Cloud. Our algorithm not only ensures the predictability and accuracy of the scaling strategy, but also enables the scaling decisions to adapt to the changing workloads with high sample efficiency. Our method has achieved significant performance improvement compared to the existing algorithms and has been deployed online at Alipay, supporting the autoscaling of applications for the world-leading payment platform.
Siqiao Xue, Chao Qu, Xiaoming Shi 0001, Cong Liao, Shiyi Zhu, Xiaoyu Tan, Lintao Ma, Shiyu Wang 0001, Yun Hu 0001, Lei Lei 0001, Yangfei Zheng, James Zhang
KDD6
2022 Sparse-attentive meta temporal point process for clinical decision support
Yajun Ru, Xihe Qiu, Xiaoyu Tan, Yongbin Gao, Yaochu Jin
Neurocomputing3
2022 A model-based hybrid soft actor-critic deep reinforcement learning algorithm for optimal ventilator settings
Shaotao Chen, Xihe Qiu, Xiaoyu Tan, Zhijun Fang 0001, Yaochu Jin
Inf. Sci.3
2022 A latent batch-constrained deep reinforcement learning approach for precision dosing clinical decision support
Xihe Qiu, Xiaoyu Tan, Shaotao Chen, Yajun Ru, Yaochu Jin
Knowl. Based Syst.2
2019 Simulation of Robot-Assisted Flexible Needle Insertion Using Deep Q-Network
abstract
Flexible needle insertion with bevel tips is becoming a preferred method for approaching targets in the human body in the least invasive manner. However, to successfully implement needle insertion, surgeons require prolonged training processes and long-term experience to develop essential handling skills. This paper presents a new path planning approach with Deep Reinforcement Learning (DRL) to implement automatic needle insertion using a surgical robot. In this paper, Deep Q-Network (DQN) algorithm is utilized to learn the control policy for flexible needle steering with needle-tissue interaction. As the human body is composed of a complex environment such as tissues, blood vessels, bones, and muscles, the uncertainty of the needle-tissue interaction should be considered during insertion. To model this complex interaction in path planning, utilizing a neural network to approximate the action-value function is more efficient than using traditional array methods in terms of time and accuracy. In our simulation, the agent (needle) can be controlled with 2 degrees of freedom (bevel direction rotation and insertion) and received negative rewards when it collides with obstacles, goes out of range, or exceeds a predefined number of rotations. During the training, the agent demonstrates the accuracy and efficiency of the learned policy through feedback scores in every episode. In addition, this system incorporates the uncertainty within flexible needle-tissue interaction using a stochastic environment. Compared with other traditional methods for flexible needle path planning, we demonstrated that motion planning of bevel-tip flexible needles in complex human bodies using DRL has better efficiency and accuracy.
Yonggu Lee, Xiaoyu Tan, Chin-Boon Chng, Chee-Kong Chui
SMC2
2016 Design and implementation of a patient-specific cognitive engine for robotic needle insertion
abstract
In order to develop an effective and user-friendly control method for surgical robotic system, we propose a new framework of cognitive engine to supervise and regulate the surgical processes. The framework aims to make the surgical processes understandable by both human operators and robots. A prototype cognitive engine was implemented using ontology and SPARQL query language on JAVA and tested in ex-vivo phantom experiments with a robotic RF needle insertion system. The prototype cognitive engine has successfully guided the robot in execution of surgical procedures.
Xiaoyu Tan, Chin-Boon Chng, Yvonne Ho, Rong Wen, Kah-Bin Lim 0001, Chee-Kong Chui
SMC1