Chenyang Shao

dblp:306/7426 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Diffuse Thinking: Exploring Diffusion Language Models as Efficient Thought Proposers for Reasoning
abstract
Large language models (LLMs) have demonstrated strong capabilities in complex reasoning tasks, yet the multi-step reasoning processes often lead to expensive computation cost.The recent advance of diffusion language models (DLMs) adopts a parallel, non-autoregressive generation mechanism, which enables the efficient production of multiple outputs.In this paper, we explore a collaborative reasoning framework that combines diffusion-based generation with autoregressive evaluation.Specifically, we leverage DLMs to efficiently propose stepwise reasoning thoughts, and employ LLMs as evaluators to assess and select candidates based on their plausibility and correctness.By decoupling proposal generation from evaluation, our framework exploits the strengths of both models: efficient exploration from diffusion models and causally grounded assessment from autoregressive models, which naturally aligns with the divergent-convergent thinking framework in cognitive psychology.Experiments across various mathematical and logical reasoning benchmarks demonstrate that, our framework improves inference efficiency while maintaining competitive or superior reasoning accuracy, laying the groundwork for building efficient reasoning architectures.Our code is open-source at https://anonymous.4open. science/r/
Chenyang Shao, Sijian Ren, Fengli Xu, Yong Li 0008
ACL (1)1
2026 Generating Realistic Human Mobility Data with Hybrid Large Language Model Agent
Chenyang Shao, Bingbing Fan, Jingtao Ding, Yuan Yuan 0032, Meng Wang 0001, Fengli Xu
KDD (1)1
2026 Route-and-Reason: Energy-Efficient Scaling of LLM Reasoning via Reinforced Model Routing
Chenyang Shao, Fengli Xu, Yong Li 0008
WWW1
2025 Division-of-Thoughts: Harnessing Hybrid Language Model Synergy for Efficient On-Device Agents
abstract
The rapid expansion of web content has made on-device AI assistants indispensable for helping users manage the increasing complexity of online tasks. The emergent reasoning ability in large language models offer a promising path for next-generation on-device AI agents. However, deploying full-scale Large Language Models (LLMs) on resource-limited local devices is challenging. In this paper, we propose Division-o f-Thoughts (DoT), a collaborative reasoning framework leveraging the synergy between locally deployed Smaller-scale Language Models (SLMs) and cloud-based LLMs. DoT leverages a Task Decomposer to elicit the inherent planning abilities in language models to decompose user queries into smaller sub-tasks, which allows hybrid language models to fully exploit their respective strengths. Besides, DoT employs a Task Scheduler to analyze the pair-wise dependency of sub-tasks and create a dependency graph, facilitating parallel reasoning of sub-tasks and the identification of key steps. To allocate the appropriate model based on the difficulty of sub-tasks, DoT leverages a Plug-and-Play Adapter, which is an additional task head attached to the SLM that does not alter the SLM's parameters. To boost adapter's task allocation capability, we propose a self-reinforced training method that relies solely on task execution feedback. Extensive experiments on various benchmarks demonstrate that our DoT significantly reduces LLM costs while maintaining competitive reasoning accuracy. Specifically, DoT reduces the average reasoning time and API costs by 66.12% and 83.57%, while achieving comparable reasoning accuracy with the best baseline methods.
Chenyang Shao, Fengli Xu
WWW1
2024 Spatio-Temporal Few-Shot Learning via Diffusive Neural Network Generation
abstract
Spatio-temporal modeling is foundational for smart city applications, yet it is often hindered by data scarcity in many cities and regions. To bridge this gap, we propose a novel generative pre-training framework, GPD, for spatio-temporal few-shot learning with urban knowledge transfer. Unlike conventional approaches that heavily rely on common feature extraction or intricate few-shot learning designs, our solution takes a novel approach by performing generative pre-training on a collection of neural network parameters optimized with data from source cities. We recast spatio-temporal few-shot learning as pre-training a generative diffusion model, which generates tailored neural networks guided by prompts, allowing for adaptability to diverse data distributions and city-specific characteristics. GPD employs a Transformer-based denoising diffusion model, which is model-agnostic to integrate with powerful spatio-temporal neural networks. By addressing challenges arising from data gaps and the complexity of generalizing knowledge across cities, our framework consistently outperforms state-of-the-art baselines on multiple real-world datasets for tasks such as traffic speed prediction and crowd flow prediction. The implementation of our approach is available: https://github.com/tsinghua-fib-lab/GPD.
Yuan Yuan 0032, Chenyang Shao, Jingtao Ding, Depeng Jin, Yong Li 0008
ICLR2
2024 GFSPX: an efficient lightweight block cipher for resource-constrained IoT nodes
Xing Zhang 0007, Chenyang Shao, Tianning Li, Changda Wang 0001
J. Supercomput.2
2023 Spatio-temporal Diffusion Point Processes
abstract
Spatio-temporal point process (STPP) is a stochastic collection of events accompanied with time and space. Due to computational complexities, existing solutions for STPPs compromise with conditional independence between time and space, which consider the temporal and spatial distributions separately. The failure to model the joint distribution leads to limited capacities in characterizing the spatio-temporal entangled interactions given past events. In this work, we propose a novel parameterization framework for STPPs, which leverages diffusion models to learn complex spatio-temporal joint distributions. We decompose the learning of the target joint distribution into multiple steps, where each step can be faithfully described by a Gaussian distribution. To enhance the learning of each step, an elaborated spatio-temporal co-attention module is proposed to capture the interdependence between the event time and space adaptively. For the first time, we break the restrictions on spatio-temporal dependencies in existing solutions, and enable a flexible and accurate modeling paradigm for STPPs. Extensive experiments from a wide range of fields, such as epidemiology, seismology, crime, and urban mobility, demonstrate that our framework outperforms the state-of-the-art baselines remarkably. Further in-depth analyses validate its ability to capture spatio-temporal interactions, which can learn adaptively for different scenarios. The datasets and source code are available online: https://github.com/tsinghua-fib-lab/Spatio-temporal-Diffusion-Point-Processes.
Yuan Yuan 0032, Jingtao Ding, Chenyang Shao, Depeng Jin, Yong Li 0008
KDD3