VLDB 2026 Research / reviewers in the wild / expert
Xihe Qiu
dblp:258/7989
· DBLP profile ↗
48ranked-venue papers
9as first author
48since 2021 · last 2026
0000-0003-4024-925XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 6 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 8 since 2021Systems, architecture and hardware · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 6 since 2021Computer networks · 4 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Answers to Arguments: Toward Trustworthy Clinical Diagnostic Reasoning with Toulmin-Guided Curriculum Goal-Conditioned LearningabstractThe integration of Large Language Models (LLMs) into clinical decision support is critically obstructed by their opaque and often unreliable reasoning.In the high-stakes domain of healthcare, correct answers alone are insufficient; clinical practice demands full transparency to ensure patient safety and enable professional accountability.A pervasive and dangerous weakness of current LLMs is their tendency to produce "correct answers through flawed reasoning."This issue is far more than a minor academic flaw; such process errors signal a fundamental lack of robust understanding, making the model prone to broader hallucinations and unpredictable failures when faced with real-world clinical complexity.In this paper, we establish a framework for trustworthy clinical argumentation by adapting the Toulmin model to the diagnostic process.We propose a novel training pipeline: Curriculum Goal-Conditioned Learning (CGCL), designed to progressively train LLM to generate diagnostic arguments that explicitly follow this Toulmin structure.CGCL's progressive three-stage curriculum systematically builds a solid clinical argument: (1) extracting facts and generating differential diagnoses; (2) justifying a core hypothesis while rebutting alternatives; and (3) synthesizing the analysis into a final, qualified conclusion.We validate CGCL using T-Eval, a quantitative framework measuring the integrity of the diagnosis reasoning.Experiments show that our method achieves diagnostic accuracy and reasoning quality comparable to resourceintensive Reinforcement Learning (RL) methods, while offering a more stable and efficient training pipeline.1 Chen Zhan, Xiaoyu Tan, Gengchen Ma, Yujie Xiong, Xihe Qiu |
ACL (1) | 6 |
| 2026 | AURORA: Automated Training Framework of Universal Process Reward Models via Ensemble Prompting and Reverse VerificationabstractThe reasoning capabilities of advanced large language models (LLMs) like o1 have revolutionized artificial intelligence applications. Nevertheless, evaluating and optimizing complex reasoning processes remain significant challenges due to diverse policy distributions and the inherent limitations of human effort and accuracy. In this paper, we present AURORA, a novel automated framework for training universal process reward models (PRMs) using ensemble prompting and reverse verification. The framework employs a two-phase approach: First, it uses diverse prompting strategies and ensemble methods to perform automated annotation and evaluation of processes, ensuring robust assessments for reward learning. Second, it leverages practical reference answers for reverse verification, enhancing the model's ability to validate outputs and improving training accuracy. To assess the framework's performance, we extend beyond the existing ProcessBench benchmark by introducing UniversalBench, which evaluates reward predictions across full trajectories under diverse policy distribtion with long Chain-of-Thought (CoT) outputs. Experimental results demonstrate that AURORA enhances process evaluation accuracy, improves PRMs' accuracy for diverse policy distributions and long-CoT responses. Xiaoyu Tan, Tianchu Yao, Chao Qu, Bin Li 0091, Dakuan Lu, Haozhe Wang 0002, Yinghui Xu 0001, Xihe Qiu |
KDD (1) | 9 |
| 2026 | Curiosity Driven Knowledge Retrieval for Mobile AgentsabstractMobile agents have made progress toward reliable smartphone automation, yet performance in complex applications remains limited by incomplete knowledge and weak generalization to unseen environments. We introduce a curiosity driven knowledge retrieval framework that formalizes uncertainty during execution as a curiosity score. When this score exceeds a threshold, the system retrieves external information from documentation, code repositories, and historical trajectories. Retrieved content is organized into structured AppCards, which encode functional semantics, parameter conventions, interface mappings, and interaction patterns. During execution, an enhanced agent selectively integrates relevant AppCards into its reasoning process, thereby compensating for knowledge blind spots and improving planning reliability. Evaluation on the AndroidWorld benchmark shows consistent improvements across backbones, with an average gain of six percentage points and a new state of the art success rate of 88.8% when combined with GPT-5. Analysis indicates that AppCards are particularly effective for multi step and cross application tasks, while improvements depend on the backbone model. Case studies further confirm that AppCards reduce ambiguity, shorten exploration, and support stable execution trajectories. Task trajectories are publicly available at https://lisalsj.github.io/Droidrun-appcard/. Xiaoyu Tan, Shahir Ali, Niels Schmidt, Gengchen Ma, Xihe Qiu |
WWW | 6 |
| 2026 | CNEKA: An Algorithm for SDN Controller Placement Based on Graph Convolutional NetworksabstractABSTRACT With the rapid development of software‐defined networking (SDN), the single‐controller architecture is unable to meet the performance and reliability requirements of the whole system. Consequently, a distributed multicontroller architecture has been proposed, in which the number and locations of controllers must be determined rigorously, formulating the controller placement problem (CPP). In order to solve the CPP by optimizing the propagation latency, we propose a convolutional node embedding and K‐means algorithm (CNEKA), which integrates information propagation among adjacent nodes and calculation of embedding vectors by graph convolutional networks (GCNs) with the graph segmentation by K‐means algorithm. As far as we know, we are the first paper to apply GCN to solving CPP. The study demonstrates that the CNEKA algorithm significantly enhances performance in optimizing average and worst‐case latency between controllers and switches, as well as the propagation latency of the whole network, in varied experimental conditions. CNEKA achieves up to a 64.18% reduction in propagation latency when compared to other algorithms, and maintains a high stability with a fluctuation less than 7.6% when repeating the same experiments for several times. Moreover, CNEKA can always find the optimal or near‐optimal solution with an error less than 11.15% when compared with the global optimal solution. Yirui Rao, Jue Chen 0001, Xihe Qiu |
Concurr. Comput. Pract. Exp. | 3 |
| 2026 | What the doodle: Sketching mental pictures with decoding brain signals, language understanding, and AI creativity
Gengchen Ma, Haoyu Wang 0011, Fenghao Sun, Chen Zhan, Xiaoyu Tan, Xihe Qiu |
Expert Syst. Appl. | 6 |
| 2026 | PAST: Pairwise attention swin transformer for offline signature verification
Yujie Xiong, Jian-Xin Ren, Dong-Hai Zhu, Xijiong Xie, Xihe Qiu |
Int. J. Document Anal. Recognit. | 5 |
| 2026 | MIF-gaus: Monocular implicit feature-driven generalizable Gaussian splatting reconstruction
Ying Li 0020, Chenmou Wu, Jiuqing Dong, Xihe Qiu, Yongbin Gao |
Neurocomputing | 6 |
| 2026 | SCV-IDS: CNN-ViT Crossattention on Serialized Traffic Images for Spatiotemporal Intrusion DetectionabstractWith the growing sophistication of cyber threats, IDSs (Intrusion Detection Systems) have become increasingly crucial for network security. However, current ViT (Vision Transformer)-based methods often fail to capture subtle local attack patterns, while existing image-based approaches typically generate static representations that cannot effectively model temporal relationships in network traffic. To address these limitations, we propose SCV-IDS (Serialized Traffic Images, CNN, and ViT-Intrusion Detection System), a novel IDS featuring: (1) a dual-branch CNN-ViT architecture with cross-attention mechanism that simultaneously captures local anomalies and global attack patterns while enabling deep feature fusion; (2) a serialized 24-bit RGB encoding scheme with sliding windows that preserves spatiotemporal relationships in network traffic. Extensive experiments on the UNSW-NB15 dataset demonstrate SCV-IDS’s superior performance, achieving 85.77% accuracy and 84.59% F1 in multiple classification, outperforming state-of-the-art methods by up to 7.37% and 6.59%, respectively. Cross-dataset evaluation on CIC-IDS2017 further validates its generalization capability, with 99.81% accuracy for multi-class classification. The framework’s exceptional spatiotemporal learning capability, combined with its ability to effectively integrate local anomaly detection and global pattern recognition, makes it particularly powerful for identifying sophisticated attacks in modern network environments. Jue Chen 0001, Haidong Peng, Henghua Zhang, Xihe Qiu |
IEEE Internet Things J. | 5 |
| 2026 | MIMAR-OSA: Enhancing obstructive sleep apnea diagnosis through multimodal data integration and missing modality reconstruction
Xihe Qiu, Yingchen Wei, Xiaoyu Tan, Weidi Xu, Jingru Ma, Zhijun Fang 0001 |
Pattern Recognit. | 1 |
| 2026 | MLLoRA: Leveraging meta-learning and LoRA for efficient multi-task fine-tuning in LLM-driven medical applications
Xihe Qiu, Siyue Shao, Qian Du, Xiaoyu Tan |
Pattern Recognit. Lett. | 1 |
| 2026 | ReMALIS: Inference-Guided Intention Propagation for Multiagent Stochastic Task Coordination With Large Language Models in Complex Networks Domains
Xihe Qiu, Haoyu Wang 0011, Xiaoyu Tan, Yujie Xiong, Zhijun Fang 0001 |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2026 | Multiagent Fuzzy Reinforcement Learning With LLM for Cooperative Navigation of Endovascular RoboticsabstractEndovascular interventions require precise, cooperative control of multiple instruments, such as guidewires and catheters, to navigate complex vascular anatomies. Current robotic systems, reliant on leader-follower control, depend heavily on operator expertise and lack intelligence. Learning-based methods, often limited to single-instrument control, fall short in complex clinical scenarios requiring multi-instrument coordination. This study proposes a Multi-Agent Fuzzy Reinforcement Learning (MAFRL) framework, guided by large language models (LLMs), for task-level autonomous, cooperative navigation in endovascular robotics. LLMs provide procedural priors and context-aware policy guidance, enabling adaptive decision-making for collaborative guidewire and catheter agents. Central to the framework, fuzzy reinforcement learning mitigates LLM-induced uncertainties by adaptively embedding clinical constraints into reward functions, ensuring strict adherence to procedural safety and precise alignment with the complexities of real-world endovascular interventions. Validated in a 3D vascular simulation, this approach achieves superior navigation performance and procedural efficiency compared to conventional methods, underscoring the transformative potential of fuzzy reinforcement learning in advancing LLM-guided MARL for endovascular robotics. Tianliang Yao, Yueqi Xu, Haoyu Wang 0011, Xihe Qiu, Kaspar Althoefer, Peng Qi 0001 |
IEEE Trans. Fuzzy Syst. | 4 |
| 2025 | Parameter-Efficient Fine-Tuning of Large Language Models via Deconvolution in SubspaceabstractThis paper proposes a novel parameter-efficient fine-tuning method that combines the knowledge completion capability of deconvolution with the subspace learning ability, reducing the number of parameters required for fine-tuning by 8 times . Experimental results demonstrate that our method achieves superior training efficiency and performance compared to existing models. Jia-Chen Zhang, Yujie Xiong, Chun-Ming Xia, Dong-Hai Zhu, Xihe Qiu |
COLING | 5 |
| 2025 | CogniDual Framework: Self-Training Large Language Models within a Dual-System Theoretical Framework for Improving Cognitive TasksabstractCognitive psychology investigates perception, attention, memory, language, problem-solving, decision-making, and reasoning. Kahneman’s dual-system theory elucidates the human decision-making process, distinguishing between the rapid, intuitive System 1 and the deliberative, rational System 2. Recent advancements have positioned large language Models (LLMs) as formidable tools nearing human-level proficiency in various cognitive tasks. Nonetheless, the presence of a dual-system framework analogous to human cognition in LLMs remains unexplored. This study introduces the CogniDual Framework for LLMs (CFLLMs), designed to assess whether LLMs can, through self-training, evolve from deliberate deduction to intuitive responses, thereby emulating the human process of acquiring and mastering new information. Our findings reveal the cognitive mechanisms behind LLMs’ response generation, enhancing our understanding of their capabilities in cognitive psychology. Practically, self-trained models can provide faster responses to certain queries, reducing computational demands during inference. Yongxin Deng, Xihe Qiu, Xiaoyu Tan, Chao Qu, Yinghui Xu 0001 |
ICASSP | 2 |
| 2025 | An Attentive Dual-Encoder Framework Leveraging Multimodal Visual and Semantic Information for Automatic OSAHS DiagnosisabstractObstructive sleep apnea-hypopnea syndrome (OS-AHS) is a common sleep disorder caused by upper airway blockage, leading to oxygen deprivation and disrupted sleep. Traditional diagnosis using polysomnography (PSG) is expensive, time-consuming, and uncomfortable. Existing deep learning methods using facial image analysis lack accuracy due to poor facial feature capture and limited sample sizes. To address this, we propose a multimodal dual encoder model that integrates visual and language inputs for automated OSAHS diagnosis. The model balances data using randomOverSampler (ROS), extracts key facial features with attention grids, and converts basic physiological data into meaningful text. Cross attention combines image and text data for better feature extraction, and ordered regression loss ensures stable learning. Our approach improves diagnostic efficiency and accuracy, achieving 91.3% top-1 accuracy in a four class severity classification task, demonstrating state of the art performance. Code is available at https://github.com/luboyan6/VTA-OSAHS. Yingchen Wei, Xihe Qiu, Xiaoyu Tan, Yinghui Xu 0001, Yuan Qi 0001 |
ICASSP | 2 |
| 2025 | Leave the Bias in Bias: Mitigating the Label Noise Effects in Continual Visual Instruction Fine-TuningabstractIn recent years, multimodal large language models (MLLMs) with vision processing capability have shown substantial advancements, excelling particularly in interpreting general images. Their application in domain-specific tasks, like those in the medical fields, is further enhanced through continuous visual instruction fine-tuning (CVIF). Despite these advancements, a significant challenge arises from label noise encountered during the collection of domain-specific data. Our studies reveal that this label noise can adversely affect the learning of vision projection embeddings and contribute to inaccuracies in LLMs’ fine-tuning, often leading to hallucinations. In this paper, we introduce a novel framework designed to minimize the impact of label noise. Our approach focuses on stabilizing the learning of vision embeddings and reducing the effect of label noise through the inherent semantic understanding of uncertainty in LLMs. Extensive experiments demonstrate that our framework maintains robust performance in general visual question-answer (VQA) tasks while showing significant effectiveness in medical VQA tasks. To the best of our knowledge, this is the first study to specifically address and analyze the impact of label noise in CVIF. Xiaoyu Tan, Teqi Hao, Xihe Qiu, Shaojie Shi, Yinghui Xu 0001, Yuan Qi 0001 |
ICME | 3 |
| 2025 | PTQ4RIS: Post-Training Quantization for Referring Image SegmentationabstractReferring Image Segmentation (RIS), aims to segment the object referred by a given sentence in an image by understanding both visual and linguistic information. However, existing RIS methods tend to explore top-performance models, disregarding considerations for practical applications on resources-limited edge devices. This oversight poses a significant challenge for on-device RIS inference. To this end, we propose an effective and efficient post-training quantization framework termed PTQ4RIS. Specifically, we first conduct an in-depth analysis of the root causes of performance degradation in RIS model quantization and propose dual-region quantization (DRQ) and reorder-based outlier-retained quantization (RORQ) to address the quantization difficulties in visual and text encoders. Extensive experiments on three benchmarks with different bits settings (from 8 to 4 bits) demonstrates its superior performance. Importantly, we are the first PTQ method specifically designed for the RIS task, highlighting the feasibility of PTQ in RIS applications. The code is available at https://github.com/gugu511yy/PTQ4RIS. Kaiying Zhu, Xihe Qiu, Shibo Zhao, Sifan Zhou |
ICRA | 4 |
| 2025 | HeStIa: Asynchronous Embodied Dynamic Locomotion Learning for Walking Robots through Multimodal Large Language ModelsabstractThe control of locomotion in walking robots with various architectural designs presents significant challenges. While existing approaches primarily rely on low-level state information and isolated visual features, lacking the high-level semantic understanding that humans use to reason about movement and posture, we propose HeStIa, a novel framework that bridges visual perception, natural language understanding, and robotic control through multimodal learning. Our framework leverages multimodal large language models (MLLMs) to establish a semantic bridge between visual observations and motion control, enabling robots to understand and adjust their locomotion through both visual and linguistic modalities. By leveraging multimodal large language models (MLLMs), HeStIa establishes a semantic connection between visual observations and motion control, enabling robots to comprehend and adapt their locomotion through both visual and linguistic modalities. Our approach extracts spatiotemporal visual features from robot movements and transforms them into a cross-modal embedding space shared with textual descriptions. HeStIa incorporates an innovative vision-language-motion fusion mechanism to provide informed, context-aware feedback during the dynamic learning process. Through an asynchronous design, HeStIa effectively mitigates the inference delays typically associated with MLLMs while maintaining real-time performance in dynamic scenarios. The cross-modal representations learned by HeStIa facilitate more intuitive and efficient locomotion learning by grounding visual observations in natural language descriptions. Our comprehensive evaluation shows substantial improvements in motion naturalness, stability, and adaptability across diverse environmental conditions. Xiaoyu Tan, Haoyu Wang 0011, Yinghui Xu 0001, Xihe Qiu |
IROS | 5 |
| 2025 | Struct-X: Enhancing the Reasoning Capabilities of Large Language Models in Structured Data Scenarios
Xiaoyu Tan, Haoyu Wang 0011, Xihe Qiu, Leijun Cheng, Yinghui Xu 0001, Yuan Qi 0001 |
KDD (1) | 3 |
| 2025 | MVP-LLMs: Optimizing Intervention Timing and Subsequent Decision Support for Mechanical Ventilation Parameter Control Using Large Language Models
Teqi Hao, Xiaoyu Tan, Bin Li 0091, Chao Qu, Yinghui Xu 0001, Xihe Qiu |
MICCAI (5) | 7 |
| 2025 | Robust Sleep Stage Prediction from Electroencephalogram with Label Noise Using Multimodal Large Language Models
Xihe Qiu, Chen Zhan, Gengchen Ma, Xiaoyu Tan |
MICCAI (6) | 1 |
| 2025 | Prolog-Driven Rule-Based Diagnostics with Large Language Models for Precise Clinical Decision Support
Xiaoyu Tan, Bin Li 0091, Weidi Xu, Chao Qu, Yinghui Xu 0001, Yuan Qi 0001, Xihe Qiu |
MICCAI (10) | 8 |
| 2025 | An end-to-end audio classification framework with diverse features for obstructive sleep apnea-hypopnea syndrome diagnosis
Bin Li 0091, Xihe Qiu, Xiaoyu Tan, Zhijun Fang 0001 |
Appl. Intell. | 2 |
| 2025 | An innovative contrastive learning approach to improve image recognition robustness and interpretability via simulated environmental perturbations
Leijun Cheng, Xihe Qiu, Xiaoyu Tan, Haoyu Wang 0011, Yujie Xiong |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | MOOO-RDQN: A deep reinforcement learning based method for multi-objective optimization of controller placement and traffic monitoring in SDN
Jue Chen 0001, Yurui Ma, Wenjing Lv, Xihe Qiu |
J. Netw. Comput. Appl. | 4 |
| 2025 | CRGT-SA: an interlaced and spatiotemporal deep learning model for network intrusion detectionabstractTo address the challenge of cyberattacks, intrusion detection systems (IDSs) are introduced to recognize intrusions and protect computer networks. Among all these IDSs, conventional machine learning methods rely on shallow learning and have unsatisfactory performance. Unlike machine learning methods, deep learning methods are the mainstream methods because of their capability to handle mass data without prior knowledge of specific domain expertise. Concerning deep learning, long short-term memory (LSTM) and temporal convolutional networks (TCNs) can be used to extract temporal features from different angles, while convolutional neural networks (CNNs) are valuable for learning spatial properties. Based on the above, this paper proposes a novel interlaced and spatiotemporal deep learning model called CRGT-SA, which combines CNN with gated TCN and recurrent neural network (RNN) modules to learn spatiotemporal properties, and imports the self-attention mechanism to select significant features. More specifically, our proposed model splits the feature extraction into multiple steps with a gradually increasing granularity, and executes each step with a combined CNN, LSTM, and gated TCN module. Our proposed CRGT-SA model is validated using the UNSW-NB15 dataset and is compared with other compelling techniques, including traditional machine learning and deep learning models as well as state-of-the-art deep learning models. According to the simulation results, our proposed model exhibits the highest accuracy and F1-score among all the compared methods. More specifically, our proposed model achieves 91.5% and 90.5% accuracy for binary and multi-class classifications respectively, and demonstrates its ability to protect the Internet from complicated cyberattacks. Moreover, we conduct another series of simulations on the NSL-KDD dataset; the simulation results of comparison with other models further prove the generalization ability of our proposed model. Jue Chen 0001, Wanxiao Liu, Xihe Qiu, Wenjing Lv, Yujie Xiong |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2025 | Reward guidance for reinforcement learning tasks based on large language models: The LMGT framework
Yongxin Deng, Xihe Qiu, Jue Chen 0001, Xiaoyu Tan |
Knowl. Based Syst. | 2 |
| 2025 | LLM-GAODE: Large-language-model augmented neural ordinary differential equation network for video nystagmography classificationabstractBenign paroxysmal positional vertigo (BPPV), a common type of vertigo with complex etiologies, is traditionally diagnosed using video nystagmography (VNG). Current automated methods lack diagnostic precision owing to subjective interpretation of eye movement characteristics. To address these challenges, we introduce a l arge l anguage m odel-augmented G ram-based a ttentive neural o rdinary d ifferential e quation ( LLM-GAODE ), an innovative and data-driven framework integrating eye-tracking technology with a Gram-based attention mechanism and a neural ordinary differential equation network to improve BPPV classification. Furthermore, when the neural network exhibits low confidence in its predictions, an LLM can supplement the process with advanced reasoning in natural language. LLM-GAODE was evaluated using an extensive VNG dataset provided by a collaborative university hospital. Results suggest that LLM-GAODE significantly outperforms existing benchmarks in trajectory classification for BPPV diagnosis. The framework enhances BPPV diagnostic accuracy and achieves state-of-the-art performance in open-source trajectory classification benchmarks. The code is available at https://github.com/XiheQiu/LLM-GAODE . Xihe Qiu, Shaojie Shi, Bin Li 0091, Xiaoyu Tan, Yongbin Gao, Shuo Li 0001 |
Knowl. Based Syst. | 1 |
| 2025 | Adaptive heterogeneous graph reasoning for relational understanding in interconnected systems
Bin Li 0091, Haoyu Wang 0011, Xaoyu Tan, Jue Chen 0001, Xihe Qiu |
J. Supercomput. | 6 |
| 2024 | ILTS: Inducing Intention Propagation in Decentralized Multi-Agent Tasks with Large Language Models
Xihe Qiu, Haoyu Wang 0011, Xiaoyu Tan, Chao Qu |
CIKM | 1 |
| 2024 | Robust Deep Hawkes Process Under Label Noise of Both Event and OccurrenceabstractIntegrating deep neural networks with the Hawkes process has significantly improved predictive capabilities in finance, health informatics, and information technology. Nevertheless, these models often face challenges in real-world settings, particularly due to substantial label noise. This issue is of significant concern in the medical field, where label noise can arise from delayed updates in electronic medical records or misdiagnoses, leading to increased prediction risks. Our research indicates that deep Hawkes process models exhibit reduced robustness when dealing with label noise, particularly when it affects both event types and timing. To address these challenges, we first investigate the influence of label noise in approximated intensity functions and present a novel framework, the Robust Deep Hawkes Process (RDHP), to overcome the impact of label noise on the intensity function of Hawkes models, considering both the events and their occurrences. We tested RDHP using multiple open-source benchmarks with synthetic noise and conducted a case study on obstructive sleep apnea-hypopnea syndrome (OSAHS) in a real-world setting with inherent label noise. The results demonstrate that RDHP can effectively perform classification and regression tasks, even in the presence of noise related to events and their timing. To the best of our knowledge, this is the first study to successfully address both event and time label noise in deep Hawkes process models, offering a promising solution for medical applications, specifically in diagnosing OSAHS. Xiaoyu Tan, Bin Li 0091, Xihe Qiu, Yinghui Xu 0001 |
ECAI | 3 |
| 2024 | Subequivariant Reinforcement Learning Framework for Coordinated Motion ControlabstractEffective coordination is crucial for motion control with reinforcement learning, especially as the complexity of agents and their motions increases. However, many existing methods struggle to account for the intricate dependencies between joints. We introduce CoordiGraph, a novel architecture that leverages subequivariant principles from physics to enhance coordination of motion control with reinforcement learning. This method embeds the principles of equivariance as inherent patterns in the learning process under gravity influence, which aids in modeling the nuanced relationships between joints vital for motion control. Through extensive experimentation with sophisticated agents in diverse environments, we highlight the merits of our approach. Compared to current leading methods, CoordiGraph notably enhances generalization and sample efficiency. Haoyu Wang 0011, Xiaoyu Tan, Xihe Qiu, Chao Qu |
ICRA | 3 |
| 2024 | Enhancing Personalized Headline Generation via Offline Goal-conditioned Reinforcement Learning with Large Language ModelsabstractRecently, significant advancements have been made in Large Language Models (LLMs) through the implementation of various alignment techniques. These techniques enable LLMs to generate highly tailored content in response to diverse user instructions. Consequently, LLMs have the potential to serve as robust, customizable recommendation systems in the field of content recommendation. However, using LLMs with user individual information and online exploration remains a challenge, which are important perspectives in developing personalized news headline generation algorithms. In this paper, we propose a novel framework to generate personalized news headlines using LLMs with extensive online exploration. The proposed approach involves initially training an offline goal-conditioned policy using supervised learning. Subsequently, online exploration is employed to collect new data for the next training iteration. Results from simulations, experiments, and real-word scenario demonstrate that our framework achieves outstanding performance on established benchmarks and can effectively generate personalized headlines under different reward settings. By treating the LLM as a goal-conditioned agent, the model can perform online exploration by modifying the goals without frequently retraining the model. To the best of our knowledge, this work represents the first investigation into the capability of LLMs to generate customized news headlines with goal-conditioned reinforcement learning via supervised learning within LLMs. Xiaoyu Tan, Leijun Cheng, Xihe Qiu, Shaojie Shi, Yinghui Xu 0001, Yuan Qi 0001 |
KDD | 3 |
| 2024 | A Medical Image Segmentation Method based on Multi-scale Features and Contour Loss Constrain (S)abstractIn recent years, deep learning has made breakthroughs in medical image segmentation, especially the U-Net architecture, which is becoming a benchmark for various medical image segmentation tasks due to the accuracy of its segmentation results.Although U-Net has achieved great success in many medical image segmentation tasks, it is still unsatisfactory in segmenting object boundaries and small objects.This is due to the fact that segmentation networks gradually lose information, especially edge information and small object information, during the process of convolution and downsampling of features.In order to solve the above problems, we design a novel method that uses a multi-scale module as a feature extractor in the contraction path of the Ushaped structure, which better captures the scale changes of the target object by acquiring image features at different scales; in the prediction stage, a contour prediction branch is constructed to constrain the loss of the target's contour, so that the segmentation network pays more attention to the boundaries of the target.We have validated the performance of our method on the Automated Cardiac Diagnosis Challenge (ACDC) and the spleen segmentation tasks of the Medical Segmentation Decathlon (MSD).The results show that our method obtained the best 95% Hausdorff Distance (HD) metrics on both the ACDC dataset and the Spleen dataset, as well as being quite competitive with other state-of-the-art methods in terms of Dice scores. Jian Niu, Zhijun Fang 0001, Xihe Qiu |
SEKE | 4 |
| 2024 | Enhancing Task Performance in Continual Instruction Fine-tuning Through Format UniformityabstractIn recent advancements, large language models (LLMs) have demonstrated remarkable capabilities in diverse tasks, primarily through interactive question-answering with humans. This development marks significant progress towards artificial general intelligence (AGI). Despite their superior performance, LLMs often exhibit limitations when adapted to domain-specific tasks through instruction fine-tuning (IF). The primary challenge lies in the discrepancy between the data distribution in general and domain-specific contexts, leading to suboptimal accuracy in specialized tasks. To address this, continual instruction fine-tuning (CIF), particularly supervised fine-tuning (SFT), on targeted domain-specific instruction datasets is necessary. Our ablation study reveals that the structure of these instruction datasets critically influences CIF performance, with substantial data distributional shifts resulting in notable performance degradation. In this paper, we introduce a novel framework that enhances CIF by promoting format uniformity. We assess our approach using the Llama2 chat model across various domain-specific instruction datasets. The results demonstrate not only an improvement in task-specific performance under CIF but also a reduction in catastrophic forgetting (CF). This study contributes to the optimization of LLMs for domain-specific applications, highlighting the significance of data structure and distribution in CIF. Xiaoyu Tan, Leijun Cheng, Xihe Qiu, Shaojie Shi, Yinghui Xu 0001, Yuan Qi 0001 |
SIGIR | 3 |
| 2024 | Multivariate graph neural networks on enhancing syntactic and semantic for aspect-based sentiment analysis
Haoyu Wang 0011, Xihe Qiu, Xiaoyu Tan |
Appl. Intell. | 2 |
| 2024 | An improved artificial bee colony algorithm to minimum propagation latency and balanced load for controller placement in Software Defined Network
Yurui Ma, Jue Chen 0001, Wenjing Lv, Xihe Qiu, Wanxiao Liu |
Comput. Networks | 4 |
| 2024 | Balancing therapeutic effect and safety in ventilator parameter recommendation: An offline reinforcement learning approach
Xihe Qiu, Xiaoyu Tan |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | Chain-of-LoRA: Enhancing the Instruction Fine-Tuning Performance of Low-Rank Adaptation on Diverse Instruction SetabstractRecently, large language models (LLMs) with conversational-style interaction, such as ChatGPT and Claude, have gained significant importance in the advancement of artificial general intelligence (AGI). However, the extensive resource requirements during pre-training, instruction fine-tuning (IF), and reinforcement learning through human feedback (RLHF) pose challenges, particularly for individuals and studios with limited resources. Moreover, sensitive data that cannot be deployed on remote training platforms or queried through APIs further exacerbates this issue. To address these limitations, researchers have introduced a parameter-efficient framework called low-rank adaptation (LoRA) for IF on LLMs. However, training individual LoRA networks faces capacity constraints and struggles to adapt to large domains with significant distributional shifts across different tasks. In this paper, we propose a novel framework called chain-of-LoRA to enhance the IF performance of LoRA. Our approach involves training a LoRA network to classify the instruction type and then utilizing task-specific LoRA networks to accomplish the respective tasks. By training multiple task-specific LoRA networks, we exploit a trade-off between performance and disk storage, leveraging the easily expandable and cost-effective nature of disk storage compared to precious graphical resources. Our experimental results demonstrate that our proposed framework achieves comparable performance to typical direct IF on LLMs. Xihe Qiu, Teqi Hao, Shaojie Shi, Xiaoyu Tan, Yujie Xiong |
IEEE Signal Process. Lett. | 1 |
| 2023 | CRNN-SA: A Network Intrusion Detection Method Based on Deep Learning
Wanxiao Liu, Jue Chen 0001, Xihe Qiu |
ADMA (2) | 3 |
| 2023 | Gram-based Attentive Neural Ordinary Differential Equations Network for Video Nystagmography ClassificationabstractVideo nystagmography (VNG) is the diagnostic gold standard of benign paroxysmal positional vertigo (BPPV), which requires medical professionals to examine the direction, frequency, intensity, duration, and variation in the strength of nystagmus on a VNG video. This is a tedious process heavily influenced by the doctor’s experience, which is error-prone. Recent automatic VNG classification methods approach this problem from the perspective of video analysis without considering medical prior knowledge, resulting in unsatisfactory accuracy and limited diagnostic capability for nystagmographic types, thereby preventing their clinical application. In this paper, we propose an end-to-end data-driven novel BPPV diagnosis framework (TC-BPPV) by considering this problem as an eye trajectory classification problem due to the disease’s symptoms and experts’ prior knowledge. In this framework, we utilize an eye movement tracking system to capture the eye trajectory and propose the Gram-based attentive neural ordinary differential equations network (Gram-AODE) to perform classification. We validate our framework using the VNG dataset provided by the collaborative university hospital and achieve state-of-the-art performance. We also evaluate Gram-AODE on multiple open-source benchmarks to demonstrate its effectiveness in trajectory classification. Code is available at https://github.com/XiheQiu/Gram-AODE. Xihe Qiu, Shaojie Shi, Xiaoyu Tan, Chao Qu, Zhijun Fang 0001, Yongbin Gao, Peixia Wu |
ICCV | 1 |
| 2023 | Provably Invariant Learning without Domain InformationabstractTypical machine learning applications always assume the data follows independent and identically distributed (IID) assumptions. In contrast, this assumption is frequently violated in real-world circumstances, leading to the Out-of-Distribution (OOD) generalization problem and a major drop in model robustness. To mitigate this issue, the invariant learning technique is leveraged to distinguish between spurious features and invariant features among all input features and to train the model purely on the basis of the invariant features. Numerous invariant learning strategies imply that the training data should contain domain information. Such information includes the environment index or auxiliary information acquired from prior knowledge. However, acquiring these information is typically impossible in practice. In this study, we present TIVA for environment-independent invariance learning, which requires no environment-specific information in training data. We discover and prove that, given certain mild data conditions, it is possible to train an environment partitioning policy based on attributes that are independent of the targets and then conduct invariant risk minimization. We examine our method in comparison to other baseline methods, which demonstrate superior performance and excellent robustness under OOD, using multiple benchmarks. Xiaoyu Tan, Lin Yong, Shengyu Zhu 0001, Chao Qu, Xihe Qiu, Yinghui Xu 0001, Peng Cui 0001, Yuan Qi 0001 |
ICML | 5 |
| 2023 | An Attentive LSTM based approach for adverse drug reactions prediction
Jiahui Qian, Xihe Qiu, Xiaoyu Tan, Jue Chen 0001 |
Appl. Intell. | 2 |
| 2023 | A density algorithm for controller placement problem in software defined wide area networks
Dun He, Jue Chen 0001, Xihe Qiu |
J. Supercomput. | 3 |
| 2022 | A cross entropy based approach to minimum propagation latency for controller placement in Software Defined Network
Jue Chen 0001, Yujie Xiong, Xihe Qiu, Dun He, Hanmin Yin, Changwei Xiao |
Comput. Commun. | 3 |
| 2022 | Sparse-attentive meta temporal point process for clinical decision support
Yajun Ru, Xihe Qiu, Xiaoyu Tan, Yongbin Gao, Yaochu Jin |
Neurocomputing | 2 |
| 2022 | A model-based hybrid soft actor-critic deep reinforcement learning algorithm for optimal ventilator settings
Shaotao Chen, Xihe Qiu, Xiaoyu Tan, Zhijun Fang 0001, Yaochu Jin |
Inf. Sci. | 2 |
| 2022 | A latent batch-constrained deep reinforcement learning approach for precision dosing clinical decision support
Xihe Qiu, Xiaoyu Tan, Shaotao Chen, Yajun Ru, Yaochu Jin |
Knowl. Based Syst. | 1 |