VLDB 2026 Research / reviewers in the wild / expert
Shuyang Jiang
dblp:153/1949
· DBLP profile ↗
11ranked-venue papers
5as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 5 first-author · 11 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MedS³: Towards Medical Slow Thinking with Self-Evolved Soft Dual-sided Process SupervisionabstractMedical language models face critical barriers to real-world clinical reasoning applications. However, mainstream efforts, which fall short in task coverage, lack fine-grained supervision for intermediate reasoning steps, and rely on proprietary systems, are still far from a versatile, credible and efficient language model for clinical reasoning usage. To this end, we propose MedS3, a self-evolving framework that imparts robust reasoning capabilities to small, deployable models. Starting with 8,000 curated instances sampled via a curriculum strategy across five medical domains and 16 datasets, we use a small base policy model to conduct Monte Carlo Tree Search (MCTS) for constructing rule-verifiable reasoning trajectories. Self-explored reasoning trajectories ranked by node values are used to bootstrap the policy model via reinforcement fine-tuning and preference learning. Moreover, we introduce a soft dual process reward model that incorporates value dynamics: steps that degrade node value are penalized, enabling fine-grained identification of reasoning errors even when the final answer is correct. Experiments on eleven benchmarks show that MedS3 outperforms the previous state-of-the-art medical model by +6.45 accuracy points and surpasses 32B-scale general-purpose reasoning models by +8.57 points. Additional empirical analysis further demonstrates that MedS3 achieves robust and faithful reasoning behavior. Shuyang Jiang, Yusheng Liao, Zhe Chen 0024, Ya Zhang 0002, Yanfeng Wang 0001, Yu Wang 0027 |
AAAI | 1 |
| 2026 | Miner: Mining Intrinsic Mastery for Data-Efficient RL in Large Reasoning ModelsabstractCurrent critic-free RL methods for large reasoning models suffer from severe inefficiency when training on positive homogeneous prompts (where all rollouts are correct), resulting in waste of rollouts due to zero advantage estimates.We introduce a radically simple yet powerful solution to Mine intrinsic mastery (MINER), that repurposes the policy's intrinsic uncertainty as a self-supervised reward signal, with no external supervision, auxiliary models, or additional inference cost.Our method pioneers two key innovations: (1) a token-level focal credit assignment mechanism that dynamically amplifies gradients on critical uncertain tokens while suppressing overconfident ones, and (2) adaptive advantage calibration to seamlessly integrate intrinsic and verifiable rewards.Evaluated across six reasoning benchmarks on Qwen3-4B and Qwen3-8B base models, MINER achieves state-of-theart performance among the other four algorithms, yielding up to 4.58 absolute gains in Pass@1 and 6.66 gains in Pass@K compared to GRPO.Comparison with other methods targeted at exploration enhancement further discloses the superiority of the two newly proposed innovations.This demonstrates that latent uncertainty exploitation is both necessary and sufficient for efficient and scalable RL training of reasoning models.Code is available at https://github.com/pixas/Miner. Shuyang Jiang, Ya Zhang 0002, Yanfeng Wang 0001, Yu Wang 0027 |
ACL (1) | 1 |
| 2025 | Towards Omni-RAG: Comprehensive Retrieval-Augmented Generation for Large Language Models in Medical ApplicationsabstractLarge language models hold promise for addressing medical challenges, such as medical diagnosis reasoning, research knowledge acquisition, clinical decision-making, and consumer health inquiry support. However, they often generate hallucinations due to limited medical knowledge. Incorporating external knowledge is therefore critical, which necessitates multi-source knowledge acquisition. We address this challenge by framing it as a source planning problem, which is to formulate context-appropriate queries tailored to the attributes of diverse sources. Existing approaches either overlook source planning or fail to achieve it effectively due to misalignment between the model’s expectation of the sources and their actual content. To bridge this gap, we present MedOmniKB, a repository comprising multigenre and multi-structured medical knowledge sources. Leveraging these sources, we propose the Source Planning Optimisation method, which enhances multi-source utilisation. Our approach involves enabling an expert model to explore and evaluate potential plans while training a smaller model to learn source alignment. Experimental results demonstrate that our method substantially improves multi-source planning performance, enabling the optimised small model to achieve state-of-the-art results in leveraging diverse medical knowledge sources. Zhe Chen 0024, Yusheng Liao, Shuyang Jiang, Pingjie Wang, Yiqiu Guo, Yanfeng Wang 0001, Yu Wang 0027 |
ACL (1) | 3 |
| 2025 | ReflecTool: Towards Reflection-Aware Tool-Augmented Clinical AgentsabstractLarge Language Models (LLMs) have shown promising potential in the medical domain, assisting with tasks like clinical note generation and patient communication.However, current LLMs are limited to text-based communication, hindering their ability to interact with diverse forms of information in clinical environments.Despite clinical agents succeeding in diverse signal interaction, they are oriented to a single clinical scenario and hence fail for broader applications.To evaluate clinical agents holistically, we propose ClinicalAgent Bench (CAB), a comprehensive medical agent benchmark consisting of 18 tasks across five key realistic clinical dimensions.Building on this, we introduce REFLECTOOL, a novel framework that excels at utilizing domain-specific tools within two stages.The first optimization stage progressively enlarges a long-term memory by saving successful solving processes and toolwise experience of agents in a tiny pre-defined training set.In the following inference stage, REFLECTOOL can search for supportive successful demonstrations from already built longterm memory to guide the tool selection strategy, and a verifier improves the tool usage according to the tool-wise experience with two verification methods-iterative refinement and candidate selection.Extensive experiments on CAB demonstrate that REFLECTOOL surpasses the pure LLMs with more than 10 points and the well-established agent-based methods with 3 points, highlighting its adaptability and effectiveness in solving complex clinical tasks.Our code and datasets are available at https: //github.com/BlueZeros/ReflecTool. Yusheng Liao, Shuyang Jiang, Yanfeng Wang 0001, Yu Wang 0027 |
ACL (1) | 2 |
| 2025 | SeMob: Semantic Synthesis for Dynamic Urban Mobility PredictionabstractHuman mobility prediction is vital for urban services, but often fails to account for abrupt changes from external events.Existing spatiotemporal models struggle to leverage textual descriptions detailing these events.We propose SeMob, an LLM-powered semantic synthesis pipeline for dynamic mobility prediction.Specifically, SeMob employs a multiagent framework where LLM-based agents automatically extract and reason about spatiotemporally related text from complex online texts.Fine-grained relevant contexts are then incorporated with spatiotemporal data through our proposed innovative progressive fusion architecture.The rich pre-trained event prior contributes enriched insights about event-driven prediction, and hence results in a more aligned forecasting model.Evaluated on a dataset constructed through our pipeline, SeMob achieves maximal reductions of 13.92% in MAE and 11.12% in RMSE compared to the spatiotemporal model.Notably, the framework exhibits pronounced superiority especially within spatiotemporal regions close to an event's location and time of occurrence 1 . Runfei Chen, Shuyang Jiang |
EMNLP | 2 |
| 2025 | Fine-tuning with Reserved Majority for Noise ReductionabstractParameter-efficient fine-tuning (PEFT) has revolutionized supervised fine-tuning, where LoRA and its variants gain the most popularity due to their low training costs and zero inference latency.
However, LoRA tuning not only injects knowledgeable features but also noisy hallucination during fine-tuning, which hinders the utilization of tunable parameters with the increasing LoRA rank.
In this work, we first investigate in-depth the redundancies among LoRA parameters with substantial empirical studies.
Aiming to resemble the learning capacity of high ranks from the findings, we set up a new fine-tuning framework, \textbf{P}arameter-\textbf{Re}dundant \textbf{F}ine-\textbf{T}uning (\preft), which follows the vanilla LoRA tuning process but is required to reduce redundancies before merging LoRA parameters back to pre-trained models.
Based on this framework, we propose \textbf{No}ise reduction with \textbf{R}eserved \textbf{M}ajority~(\norm), which decomposes the LoRA parameters into majority parts and redundant parts with random singular value decomposition.
The major components are determined by the proposed \search method, specifically employing subspace similarity to confirm the parameter groups that share the highest similarity with the base weight.
By employing \norm, we enhance both the learning capacity and benefits from larger ranks, which consistently outperforms both LoRA and other \preft-based methods on various downstream tasks, such as general instruction tuning, math reasoning and code generation.
Code is available at \url{https://github.com/pixas/NoRM}. Shuyang Jiang, Yusheng Liao, Ya Zhang 0002, Yanfeng Wang 0001, Yu Wang 0027 |
ICLR | 1 |
| 2024 | TAIA: Large Language Models are Out-of-Distribution Data LearnersabstractFine-tuning on task-specific question-answer pairs is a predominant method for enhancing the performance of instruction-tuned large language models (LLMs) on downstream tasks. However, in certain specialized domains, such as healthcare or harmless content generation, it is nearly impossible to obtain a large volume of high-quality data that matches the downstream distribution. To improve the performance of LLMs in data-scarce domains with domain-mismatched data, we re-evaluated the Transformer architecture and discovered that not all parameter updates during fine-tuning contribute positively to downstream performance. Our analysis reveals that within the self-attention and feed-forward networks, only the fine-tuned attention parameters are particularly beneficial when the training set's distribution does not fully align with the test set. Based on this insight, we propose an effective inference-time intervention method: \uline{T}raining \uline{A}ll parameters but \uline{I}nferring with only \uline{A}ttention (TAIA). We empirically validate TAIA using two general instruction-tuning datasets and evaluate it on seven downstream tasks involving math, reasoning, and knowledge understanding across LLMs of different parameter sizes and fine-tuning techniques. Our comprehensive experiments demonstrate that TAIA achieves superior improvements compared to both the fully fine-tuned model and the base model in most scenarios, with significant performance gains. The high tolerance of TAIA to data mismatches makes it resistant to jailbreaking tuning and enhances specialized tasks using general data. Code is available in \url{https://github.com/pixas/TAIA_LLM}. Shuyang Jiang, Yusheng Liao, Ya Zhang 0002, Yanfeng Wang 0001, Yu Wang 0027 |
NeurIPS | 1 |
| 2023 | Self-Improvement of Non-autoregressive Model via Sequence-Level DistillationabstractAlthough Non-autoregressive Transformer (NAT) models have achieved great success in terms of fast inference speed, this speedup comes with a performance drop due to the inherent multi-modality problem of the NAT model.Previous works commonly alleviate this problem by replacing the target side of the raw data with distilled data generated by Autoregressive Transformer (AT) models.However, the multimodality problem in the distilled data is still significant and thus limits further improvement of the NAT models.In this paper, we propose a method called Sequence-Level Self-Distillation (SLSD), which aims to generate distilled data by the NAT model itself, eliminating the need for additional teacher networks.Furthermore, SLSD can adapt to different NAT models without precise adjustments since the self-distilled data is generated from the same types of NAT models.We conduct extensive experiments on WMT14 EN↔DE and WMT16 EN↔RO and choose five classic NAT models as the backbones to validate the generality and effectiveness of SLSD.The results show that our approach can consistently improve all models on both raw data and distilled data without sacrificing the inference speed. Yusheng Liao, Shuyang Jiang, Yu Wang 0027, Yanfeng Wang 0001 |
EMNLP | 2 |
| 2023 | CAB: Comprehensive Attention Benchmarking on Long Sequence ModelingabstractTransformer has achieved remarkable success in language, image, and speech processing. Recently, various efficient attention architectures have been proposed to improve transformer’s efficiency while largely preserving its efficacy, especially in modeling long sequences. A widely-used benchmark to test these efficient methods’ capability on long-range modeling is Long Range Arena (LRA). However, LRA only focuses on the standard bidirectional (or noncausal) self attention, and completely ignores cross attentions and unidirectional (or causal) attentions, which are equally important to downstream applications. In this paper, we propose Comprehensive Attention Benchmark (CAB) under a fine-grained attention taxonomy with four distinguishable attention patterns, namely, noncausal self, causal self, noncausal cross, and causal cross attentions. CAB collects seven real-world tasks from different research areas to evaluate efficient attentions under the four attention patterns. Among these tasks, CAB validates efficient attentions in eight backbone networks to show their generalization across neural architectures. We conduct exhaustive experiments to benchmark the performances of nine widely-used efficient attention architectures designed with different philosophies on CAB. Extensive experimental results also shed light on the fundamental problems of efficient attentions, such as efficiency length against vanilla attention, performance consistency across attention patterns, the benefit of attention mechanisms, and interpolation/extrapolation on long-context language modeling. Shuyang Jiang, Jiangtao Feng, Lingpeng Kong |
ICML | 2 |
| 2023 | Attentive Multi-Layer Perceptron for Non-autoregressive Generation
Shuyang Jiang, Jiangtao Feng, Lingpeng Kong |
ECML/PKDD (2) | 1 |
| 2023 | Efficient 3D Deep LiDAR OdometryabstractAn efficient 3D point cloud learning architecture, named EfficientLO-Net, for LiDAR odometry is first proposed in this article. In this architecture, the projection-aware representation of the 3D point cloud is proposed to organize the raw 3D point cloud into an ordered data form to achieve efficiency. The Pyramid, Warping, and Cost volume (PWC) structure for the LiDAR odometry task is built to estimate and refine the pose in a coarse-to-fine approach. A projection-aware attentive cost volume is built to directly associate two discrete point clouds and obtain embedding motion patterns. Then, a trainable embedding mask is proposed to weigh the local motion patterns to regress the overall pose and filter outlier points. The trainable pose warp-refinement module is iteratively used with embedding mask optimized hierarchically to make the pose estimation more robust for outliers. The entire architecture is holistically optimized end-to-end to achieve adaptive learning of cost volume and mask, and all operations involving point cloud sampling and grouping are accelerated by projection-aware 3D feature learning methods. The superior performance and effectiveness of our LiDAR odometry architecture are demonstrated on KITTI, M2DGR, and Argoverse datasets. Our method outperforms all recent learning-based methods and even the geometry-based approach, LOAM with mapping optimization, on most sequences of KITTI odometry dataset. We open sourced our codes at: https://github.com/IRMVLab/EfficientLO-Net. Guangming Wang 0001, Xinrui Wu, Shuyang Jiang, Zhe Liu 0022, Hesheng Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |