VLDB 2026 Research / reviewers in the wild / expert
Shaoyuan Chen
dblp:126/1575
· DBLP profile ↗
8ranked-venue papers
4as first author
4since 2021 · last 2026
0000-0003-3526-3241ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorComputer networks · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Efficient and distributed learning · 63% Language models and text generation · 20% Deep learning architectures and training · 17% | |
| Databases, data mining, and information retrieval
1 paper |
Graph data management · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Hardware accelerators and domain-specific architectures · 62% Distributed systems · 19% GPUs and heterogeneous computing · 19% |
Topics — the 10 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
large language model |
1.0 | 1 | 2026 | OTARo: Once Tuning for All Precisions Toward Robust On-Device LLMs · AAAI 2026 |
Machine learning › Efficient and distributed learning › model compression › quantization
mixed-precision quantization |
1.0 | 1 | 2026 | OTARo: Once Tuning for All Precisions Toward Robust On-Device LLMs · AAAI 2026 |
Machine learning › Efficient and distributed learning
model quantization |
1.0 | 1 | 2026 | OTARo: Once Tuning for All Precisions Toward Robust On-Device LLMs · AAAI 2026 |
Machine learning › Deep learning architectures and training › mixture of experts
mixture-of-experts inference |
0.9 | 1 | 2025 | KTransformers: Unleashing the Full Potential of CPU/GPU Hybrid Inference for MoE Models · SOSP 2025 |
Machine learning › Efficient and distributed learning
model inference |
0.9 | 1 | 2025 | KTransformers: Unleashing the Full Potential of CPU/GPU Hybrid Inference for MoE Models · SOSP 2025 |
Graph data management › graph query processing
distributed graph queries |
0.9 | 1 | 2025 | Scaling Asynchronous Graph Query Processing via Partitioned Stateful Traversal Machines · ICDE 2025 |
Graph data management
graph query processing |
0.9 | 1 | 2025 | Scaling Asynchronous Graph Query Processing via Partitioned Stateful Traversal Machines · ICDE 2025 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.9 | 1 | 2025 | KTransformers: Unleashing the Full Potential of CPU/GPU Hybrid Inference for MoE Models · SOSP 2025 |
Machine learning › Efficient and distributed learning
on-device inference |
0.3 | 1 | 2026 | OTARo: Once Tuning for All Precisions Toward Robust On-Device LLMs · AAAI 2026 |
GPUs and heterogeneous computing
CPU-GPU heterogeneous computing |
0.3 | 1 | 2025 | KTransformers: Unleashing the Full Potential of CPU/GPU Hybrid Inference for MoE Models · SOSP 2025 |
Methods — techniques the papers use, named apart from their topics
query memoranda · 1.7partitioned stateful traversal machine · 1.7CPU/GPU hybrid inference · 1.7quantization-aware fine-tuning · 1.0bit-width path search · 1.0asynchronous gradient accumulation · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OTARo: Once Tuning for All Precisions Toward Robust On-Device LLMsabstractLarge Language Models (LLMs) fine-tuning techniques not only improve the adaptability to diverse downstream tasks, but also mitigate adverse effects of model quantization. Despite this, conventional quantization suffers from its structural limitation that hinders flexibility during the fine-tuning and deployment stages. Practical on-device tasks demand different quantization precisions (i.e. different bit-widths), e.g., understanding tasks tend to exhibit higher tolerance to reduced precision compared to generation tasks. Conventional quantization, typically relying on scaling factors that are incompatible across bit-widths, fails to support the on-device switching of precisions when confronted with complex real‑world scenarios. To overcome the dilemma, we propose OTARo, a novel method that enables on-device LLMs to flexibly switch quantization precisions while maintaining performance robustness through once fine-tuning. OTARo introduces Shared Exponent Floating Point (SEFP), a distinct quantization mechanism, to produce different bit-widths through simple mantissa truncations of a single model. Moreover, to achieve bit-width robustness in downstream applications, OTARo performs a learning process toward losses induced by different bit-widths. The method involves two critical strategies: (1) Exploitation-Exploration Bit-Width Path Search (BPS), which iteratively updates the search path via a designed scoring mechanism; (2) Low-Precision Asynchronous Accumulation (LAA), which performs asynchronous gradient accumulations and delayed updates under low bit-widths. Experiments on popular LLMs, e.g., LLaMA3.2-1B, LLaMA3-8B, demonstrate that OTARo achieves consistently strong and robust performance for all precisions. Shaoyuan Chen, Zhixuan Chen, Zhihang Yuan, Qiang Wu 0012 |
AAAI | 1 |
| 2026 | DualPath: Accelerating Agentic LLM Inference by Harvesting Disaggregated KV-Cache Storage I/OabstractThe performance of multi-turn, agentic LLM inference is increasingly dominated by KV-Cache storage I/O rather than computation. In prevalent disaggregated architectures, loading the massive KV-Cache from external storage creates a fundamental imbalance: storage NICs on prefill engines become bandwidth-saturated, while those on decoding engines remain idle. This asymmetry severely constrains overall system throughput. Yongtong Wu, Shaoyuan Chen, Rilin Huang, Yixuan Tan, Yinmin Zhong, Xin Jin 0008, Panpan Huang |
SIGCOMM | 2 |
| 2025 | Scaling Asynchronous Graph Query Processing via Partitioned Stateful Traversal MachinesabstractDue to the escalating demand to analyze large graphs, many organizations are now collecting billion-level property graph datasets, concurrently executing many complex graph queries against them, and expecting interactive-level response latency. However, such requirements are particularly challenging because of the notoriously irregular data access pattern and complex dependencies between heterogeneous subtasks. Despite the widespread availability of many-core CPUs and high-speed networking in modern datacenters, existing distributed graph query systems struggle with their inherent inefficiencies, resulting in low hardware utilization and poor query performance on these state-of-the-art hardware. To address these challenges, we introduce the Partitioned Stateful Traversal Machine (PSTM), which extends the Gremlin graph traversal machine. PSTM retains the expressive power of the Gremlin query language, enabling it to accommodate a wide range of graph query tasks, including traversal, pattern matching, filtering, and result aggregation. It additionally introduces query memoranda, allowing for more efficient implementation and execution of numerous graph queries in distributed environments. Moreover, PSTM facilitates various system-level optimizations, such as massively parallel execution, overlapping computation with communication, locality-aware data access, and lightweight progress tracking. Building upon PSTM, we develop GraphDance, a distributed graph database featuring an efficient asynchronous PSTM run-time. Our evaluations, conducted on an 8-node cluster, show that GraphDance achieves millisecond-level query latency for complex queries on terabyte-scale graphs, with an average latency reduction of 89.2% across all interactive complex queries in the LDBC SNB benchmark compared to existing distributed graph query systems. Shaoyuan Chen, Hongtao Chen, Shaonan Ma, Yajie Qin, Weiyu Xie, Kang Chen 0001, Xia Liao, Yingdi Shan, Jinlei Jiang, Yongwei Wu 0001 |
ICDE | 1 |
| 2025 | KTransformers: Unleashing the Full Potential of CPU/GPU Hybrid Inference for MoE ModelsabstractDue to the sparse nature of Mixture-of-Experts (MoE) models, they are particularly suitable for hybrid CPU/GPU inference, especially in low-concurrency scenarios. This hybrid approach leverages both the large, cost-effective memory capacity of CPU/DRAM and the high bandwidth of GPU/VRAM. However, existing hybrid solutions remain bottlenecked by CPU computation limits and CPU-GPU synchronization overheads, severely restricting their ability to efficiently run state-of-the-art large MoE models, such as the 671B DeepSeek-V3/R1. Hongtao Chen, Weiyu Xie, Boxin Zhang, Jingqi Tang, Shaoyuan Chen, Ziwei Yuan, Chengyu Qiu, Yuening Zhu, Qingliang Ou, Jiaqi Liao, Xianglin Chen, Zhiyuan Ai, Yongwei Wu 0001 |
SOSP | 7 |
| 2020 | Leaf Aging Affects the Variability of Canopy Reflectance with Stand Development in Evergreen Chinese FIR PlantationabstractDespite the long-term records of satellite observations, factors controlling the seasonal and interannual variations in canopy reflectance remain poorly understood. Leaf optical properties (LOP, including leaf reflectance and transmittance) changes as leaves age, and thus impact the seasonal pattern of canopy reflectance, i.e., the “leaf age effect”. Here, we combined the Geometric Optical Radiative Transfer (GORT) model with continuous field measurements of leaf- to stand-scale characteristics to simulate canopy reflectance in a Chinese fir plantation with stand development (1-33 yr). We found that canopy structure controls the variations in canopy reflectance during young stages (<; 10 yr) and that leaf age controls the variations in canopy reflectance after canopy closure. Moreover, we found that the “leaf age effect” get enhanced with stand development, with R2 increased from about 0.1 to 0.56, 0.67, 0.92, and 0.82 for young, half-mature, near-mature, and mature stages, respectively. This study reveals the stand age dependence of leaf age effect on canopy reflectance which improves our interpretation and understanding of satellite observations to study the ecosystem function of forests. Qiaoli Wu, Jinling Song, Jindi Wang, Conghe Song, Shaoyuan Chen, Lei Yang 0046 |
IGARSS | 5 |
| 2016 | A Multiantenna RFID Reader With Blind Adaptive BeamformingabstractA passive ultra-high frequency (UHF) RF identification (RFID) system is proposed that employs multiple antennas at the reader and single antenna at each tag. The gain due to multiple antennas in terms of the maximum reader interrogation range is first quantified. Then, a blind adaptive beamforming (BABF) algorithm that does not require the channel state information (CSI) is presented to improve both the interrogation range and the data transmission performance of the system in both the full- and half-duplex configurations. Moreover, an estimator for the number of tags within the reader interrogation range is proposed to enhance the system throughput. Simulation results show that both the interrogation range and the packet error rate (PER) performance for data transmission can be improved significantly using multiple antennas at the reader. A Universal Software Radio Peripheral (USRP)-based prototype system is implemented to validate the interrogation range improvement by employing multiple-antenna readers with several beamforming schemes. The proposed multiantenna RFID reader with the BABF complies with the existing RFID standard and can be readily implemented in existing systems. Shaoyuan Chen, Xiaodong Wang 0001 |
IEEE Internet Things J. | 1 |
| 2015 | Extraction and application of leaf area index's priori knowledge in time series for typical cropsabstractThe ill-posed inversion problem is to be solved urgently. Priori knowledge is introduced to increase the inversion information to improve the inversion quality. MODIS time-series leaf area index (LAI) data is used to extract the priori knowledge. Savitzky-Golay (SG) filter and bi-Gaussian curve-fit are performed to reconstruct the original time-series LAI data, then, the upper envelope of the smoothed time-series curve is achieved to get a fixed range of LAI in a certain growing season. The variation of LAI was restricted with the range based on Look-up Table (LUT) method with the PROSAIL model. The validation results show that the priori knowledge extracted from LAI time-series data is efficiency on improving the LAI inversion precision. Shaoyuan Chen, Hua Yang 0005, Jingjing Pan |
IGARSS | 1 |
| 2014 | Application of multi-output support vector regression in remore sensing inversionsabstractThis paper extends the standard support vector regression (SVR) into multi-dimensional case to estimate different biophysical parameters simultaneously. The improvement is made by the Vapnik loss function of L2 form. The proposed multi-output SVR (MO-SVR) is implemented in the joint inversion of Leaf Area Index (LAI) and vegetation cover fraction (fCover) over SPOT2/HRV1 data in the Fundulea site (VALERI), Romania. Comparison between standard SVR and MO-SVR, and the validation using biophysical maps both indicate the better fitness and accuracy of MO-SVR than the traditional form. Jingjing Pan, Hua Yang 0005, Peipei Xu, Shaoyuan Chen |
IGARSS | 4 |