Tianyang Zheng

dblp:48/8401 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Lethe: Layer- and Time-Adaptive KV Cache Pruning for Reasoning-Intensive LLM Serving
abstract
Generative reasoning with large language models (LLMs) often involves long decoding sequences, leading to substantial memory and latency overheads from accumulating key-value (KV) caches. While existing KV compression methods primarily focus on reducing prefill memory from long input sequences, they fall short in addressing the dynamic and layer-sensitive nature of long-form generation, which is central to reasoning tasks. We propose Lethe, a dynamic KV cache management framework that introduces adaptivity along both the spatial and temporal dimensions of decoding. Along the spatial dimension, Lethe performs layerwise sparsity-aware allocation, assigning token pruning budgets to each transformer layer based on estimated attention redundancy. Along the temporal dimension, Lethe conducts multi-round token pruning during generation, driven by a Recency-Aware Selective Retention (RASR) mechanism. RASR extends traditional recency-based heuristics by also considering token relevance derived from evolving attention patterns, enabling informed decisions about which tokens to retain or evict. Empirical results demonstrate that Lethe achieves a favorable balance between efficiency and generation quality across diverse models and tasks, increases throughput by up to 2.56×.
Daming Zhao, WenXuan Hou, Tianyang Zheng, Weiye Ji, Jidong Zhai
AAAI5
2026 CHIME: Cost-Constrained Hybrid Popularity-Aware Intelligent Service Caching Framework for MEC
Tianyang Zheng, Pengfei Yang 0001, Chenlu Zhai, Wenkai Lv, Yueli Ding, Quan Wang 0006
IEEE Internet Things J.1
2025 Alternating optimization for energy consumption-oriented task offloading in SAGIN
Pengfei Yang 0001, Tianyang Zheng, Weidi Su, Bijie Yi, Wenkai Lv, Quan Wang 0006
Comput. Networks2
2025 Cortex: Enhancing Resource Utilization in Edge Clusters Through Efficient Co-Location of LC and BE Workloads
abstract
In edge computing environments, the co-location of latency-critical (LC) services and best-effort (BE) jobs is a key strategy for enhancing resource utilization. However, existing analysis-based co-location strategies incur high analytical costs and struggle to rapidly adapt to the evolving fields of edge computing and microservice architectures, often failing to effectively meet the demands of edge computing environments. Feedback-based co-location strategies, while reducing analytical overhead, lack sufficient research in multi-node environments, resulting in overly coarse-grained deployment strategies. These strategies do not adequately consider the dynamic workloads and resource constraints inherent in edge computing, leading to improper resource allocation and degraded performance. This paper introduces Cortex, a Kubernetes-based co-location framework for edge device clusters that addresses these challenges by innovatively transforming the co-location deployment problem into a Minimum Cost Maximum Flow (MCMF) problem and employing the Network Simplex Algorithm (NSA) to optimize resource allocation and ensure QoS of LC services. Cortex also features a dynamic adjustment mechanism that adapts to changes in the request load of LC services, thereby minimizing the performance loss of BE jobs and reducing resource wastage. Our experiments in real edge device clusters demonstrate that Cortex significantly improves system resource utilization by 12.81%, increases the QoS satisfaction rate by 17.86%, and boosts the number of BE jobs by 51.96% compared to existing methods.
Tianyang Zheng, Pengfei Yang 0001, Quan Wang 0006, Wenkai Lv
IEEE Internet Things J.1
2024 Stable Matching with Approval Preferences Under Partial Information
Yaqin Chu, Junjie Luo 0001, Tianyang Zheng
AAIM (2)3
2024 Graph-Reinforcement-Learning-Based Dependency-Aware Microservice Deployment in Edge Computing
abstract
Microservice architecture is a design philosophy that achieves decoupling by decomposing a monolithic application into multiple lightweight microservices. Meanwhile, edge computing can significantly reduce service latency and network congestion by extending computation and storage resources to the network edge. Therefore, in the microservice-oriented edge computing platform, a fundamental problem is how to efficiently deploy microservices with complex dependencies on the resource-constrained edge servers to satisfy the Quality of Service (QoS) constraints of users. Most of the existing studies ignore multiple call graphs with differentiated dependencies for an application, which often result in the violation of QoS. To address this issue, in this article, we first model the request response time of multiple instances and multiple call graphs scenario with service conflicts. Then, different from the existing heuristic or approximation algorithms which rely heavily on expert knowledge, we propose a graph-reinforcement-learning-based deployment (GRLD) framework. GRLD uses a graph convolutional network (GCN) to extract the graph data required for multiple call graphs with messages passing and aggregation, and the generated feature is fed into the underlying network of deep-reinforcement-learning (DRL). Experimental results show that GRLD outperforms counterparts in reducing service deployment overhead while satisfying QoS constraints of multiple call graphs.
Wenkai Lv, Pengfei Yang 0001, Tianyang Zheng, Chengmin Lin, Minwen Deng, Quan Wang 0006
IEEE Internet Things J.3
2023 Energy Consumption and QoS-Aware Co-Offloading for Vehicular Edge Computing
abstract
By deploying computing, storage, and bandwidth resources at the user side, vehicular edge computing (VEC) provides low-delay services for vehicle users. However, due to the limited resources of edge servers, how to efficiently meet the Quality-of-Service (QoS) requirements of multiple tasks and save the total energy consumption in a dynamic environment is an important issue in VEC. In this article, we first propose an energy consumption and QoS-aware co-offloading model. Unlike most previous studies, our goal is to minimize the total energy consumption while guaranteeing the QoS constraints of tasks, thus avoiding the overallocation of resources and high energy consumption caused by the one-sided pursuit of delay minimization. Then, without the requirements for domain experts, we propose Bayesian optimization-based computation offloading (BOCO) method to find the optimal offloading decision. To the best of our knowledge, this work is the first to apply Bayesian optimization to computation offloading in VEC. Furthermore, we conduct a series of experiments and comparisons with other offloading methods to analyze the effectiveness and performance of the proposed algorithm. Experimental results verify that our proposed BOCO outperforms counterparts.
Wenkai Lv, Pengfei Yang 0001, Tianyang Zheng, Bijie Yi, Yunqing Ding, Quan Wang 0006, Minwen Deng
IEEE Internet Things J.3