Lianjie Cao

dblp:127/2812 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0003-3408-9050ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 7 · 4 first-author · 4 since 2021Systems, architecture and hardware · 5 · 5 since 2021Security and privacy · 1 · 1 first-author
YearPublicationVenuePosition
2026 Scaling Attention Beyond GPUs for LLM Inference
abstract
Scaling inference for large language models is increasingly constrained by limited GPU memory, primarily due to the expanding intermediate states (KV caches) required for long-context generation and multi-user workloads. Once the KV cache exceeds the capacity of high-bandwidth memory, it must be offloaded to host memory and reloaded on demand, a workflow severely bottlenecked by the CPU–GPU interconnect, typically PCIe. Existing approaches exploiting offload KV caches to CPU memory and selectively reload partial segments for attention computation often underutilize CPU compute resources and suffer from accuracy degradation. We present Beyond, a drop-in runtime that integrates a smart offloading scheme to selectively identify and retain salient KV entries across continuous decoding sessions, together with a hybrid CPU–GPU attention mechanism for scalable inference. Beyond executes dense attention over recent KV entries stored in GPU memory while performing parallel, per-head sparse attention on salient contextual KV entries residing in CPU memory. The outputs are fused efficiently through a log-sum-exp scheme. During the bandwidth-constrained decoding phase, oversized KV caches are processed cooperatively by the aggregated CPU and GPU memory bandwidth, with only minimal PCIe data movement. Experiments across diverse models and workloads demonstrate that Beyond improves scalability, supports longer sequences and larger batch sizes, and outperforms existing sparse attention baselines in both efficiency and accuracy—all on commodity GPU hardware.
Weishu Deng, Peiran Du, Lingfeng Xiang, Chen Zhong 0002, Faraz Ahmed, Lianjie Cao, Puneet Sharma 0001, Song Jiang 0001, Hui Lu 0001, Jia Rao
HPDC8
2026 Griffin: Coherency-Aware Task Scheduling and Memory Allocation for CXL Interconnects
abstract
CXL is an emerging interconnect that has the potential to efficiently realize memory disaggregation. This is because CXL enables the expansion of memory beyond individual hosts, and supports coherent memory sharing among multiple hosts. However, CXL introduces several performance overheads due to the cache coherency protocol for memory sharing, as well as placement constraints for shared data, which, if ignored, can lead to correctness issues. This paper presents the first analysis of the impact of CXL memory sharing and shows that the overheads of hardware-based coherency in CXL interconnects are substantial. We then propose Griffin, a new coherency-aware task and memory allocator for CXL disaggregated memory systems. Griffin introduces new abstractions and algorithms that allow it to prioritize which data is allocated remotely and to which memory node, to efficiently reduce the coherence overheads associated with both the amount of shared data and the load on CXL coherence resources. Our simulation results show that Griffin reduces the total memory time by up to 4.29 × compared to a standard baseline and 1.71 × compared to an advanced baseline.
Suyeon Lee, Khaled Diab 0001, Diman Zad Tootaghaj, Lianjie Cao, Puneet Sharma 0001, Ada Gavrilovska
ICS4
2026 Beyond Monoliths: Enabling Flexible and Composable AI Systems via Memory Disaggregation
abstract
Modern state-of-the-art AI systems are increasingly built as monolithic supernodes integrating large numbers of specialized accelerators with proprietary high-bandwidth interconnects. These systems provision compute, memory, and networking resources in fixed ratios at design time. As AI workloads evolve, their resource demands increasingly diverge from these static configurations, leading to underutilization, limited scalability, and high operational cost. Composable systems based on disaggregated resources offer a more flexible alternative by allowing memory and compute capacity to be scaled independently without replicating an entire supernode.
Divya Kiran Kadiyala, Lianjie Cao, Jinsun Yoo, Puneet Sharma 0001, Samantika Sury, Alexandros Daglis
SIGCOMM2
2026 CCSwitch: A Scalable Data Plane for Non-Blocking In-Network Collective Communication
abstract
Collective communication operations in AI and HPC workloads generate heavy network traffic. Offloading these operations to network switches reduces latency, but performing arithmetic and replication at line rate is difficult, especially as port counts and link speeds grow. Existing in-network approaches rely on accumulation buffers that not only limit throughput but also require complex state management to handle stragglers and congestion. We present CCSwitch, a modular switching fabric built from 4×4 non-blocking Collective Engines (CEs). Each CE combines spatial and temporal parallelism to perform reductions without accumulation buffers. CEs compose into k-ary n-tree topologies, scaling to 32- and 256-port switches while preserving non-blocking throughput. Source routing and flit-level synchronization keep per-switch state minimal. Our FPGA implementation shows that CCSwitch's quaternary-tree reduction fabric uses up to 23% fewer LUTs and 12–30% fewer flip-flops than a comparable Clos-based design at equal throughput. Enabling the full feature set—source routing, replication, and time-multiplexed VCs—uses 1.4–1.8× more LUTs than the circuit-switched baseline, well below the 3–5× overhead typical of packet-switched NoC routers, while supporting concurrent collectives on shared links.
Sumukh Pinge, Hardik Soni 0001, Bob Lantz, Khaled Diab 0001, Lianjie Cao, Tajana Rosing, Puneet Sharma 0001
SIGCOMM5
2025 Can Hardware Outsmart Software in Tiered Memory Management? A CMM-H Case Study
abstract
With the advent of Compute Express Link (CXL), hardware-managed memory tiering has become a reality. In this paper, we investigate Samsung's CXL Memory Module-Hybrid (CMM-H), a CXL Type 3 device integrating DRAM and NAND flash managed by an FPGA-based controller and providing byte-addressable memory interface via the cxl.mem protocol. We perform a detailed evaluation of CMM-H and compare its performance with OS-level and block-level tiering solutions. Our results highlight the performance benefits of CMM-H for cache-hit scenarios and identify key limitations for cache-miss situations, offering insights into the trade-offs involved in adopting hardware-managed memory tiering in emerging CXL-based systems.
Lingfeng Xiang, Lianjie Cao, Faraz Ahmed, Jia Rao, Hui Lu 0001, Puneet Sharma 0001
SYSTOR4
2024 Accelerating Containerized Machine Learning Workloads
abstract
To facilitate various Machine Learning (ML) training and inference tasks, enterprises tend to build large and expensive clusters and share them among different teams for diverse ML workloads. Virtualized platforms (containers/VMs) and schedulers are typically deployed to allow such access, manage heterogeneous resources and schedule ML jobs in these clusters. However, allocating resource budgets for different ML jobs to achieve best performance and cluster resource efficiency remains a significant challenge. This work proposes Nearchus to accelerate distributed ML training while ensuring high resource efficiency by using adaptive resource allocation. Nearchus automatically identifies potential performance bottlenecks for running jobs and re-allocates resources to provide optimized run-time performance with high resource efficiency. Nearchus’s resource configuration significantly improves the training speed of individual jobs up to 71.4%–129.1% against state-of-the-art resource schedulers, and reduces job completion and queuing time by 35.6% and 67.8%, respectively.
Ali Tariq, Lianjie Cao, Faraz Ahmed, Eric Rozner, Puneet Sharma 0001
NOMS2
2024 Conspirator: SmartNIC-Aided Control Plane for Distributed ML Workloads
Yunming Xiao, Diman Zad Tootaghaj, Aditya Dhakal, Lianjie Cao, Puneet Sharma 0001, Aleksandar Kuzmanovic
USENIX ATC4
2023 When Caching Systems Meet Emerging Storage Devices: A Case Study
abstract
Block-layer caching systems improve the I/O performance by using hybrid storage devices; the advent of fast, byte-addressable storage enables caching systems to further leverage new storage tiers (e.g., with persistent memory as the cache device and SSD as the backend device) to achieve better caching performance. However, the new storage devices also challenge the design and implementation of existing block-based caching systems. This paper conducts a comprehensive performance study of a popular caching system, Open CAS, and identifies new, unrevealed software bottlenecks. Our observations and root cause analysis cast light on optimizing the software stack of caching systems to incorporate emerging storage technologies.
Lianjie Cao, Faraz Ahmed, Hui Lu 0001, Puneet Sharma 0001
HotStorage2
2022 Metered Boot: Trusted Framework for Application Usage Rights Management in Virtualized Ecosystems
abstract
The adoption of virtualization and cloud computing technologies have revolutionized how services and applications can be developed, deployed, and operated to achieve better elasticity, flexibility, and scalability. Multiple stakeholders can be involved for providing online services; each of them plays one or more roles (i.e., service operator, application vendor, and infrastructure provider) to create a customized operating model based on the business requirements. The operating model changes from one business to another, and it may even change at different stages of the same business. A trusted relationship among stakeholders for secure information exchange is the key to enable such flexibility. However, traditional usage compliance methods (e.g., in-person audit, dynamic licensing, and subscription) lack explicit trust among involved parties and the flexibility and scalability to support dynamic sizing of services and applications with low overhead. In this work, we argue the need for a new trust framework to manage application usage rights and propose Metered Boot to provide trusted, capacity/usage-based usage rights management for services and applications deployed in virtualized environments. Metered Boot decouples application workload instantiation for service operators, usage rights governance for application vendors, and resource provisioning for infrastructure providers. We leverage cryptoprocessors (e.g., Trusted Platform Module (TPM)) on commodity servers to generate trusted proofs which are managed by efficient cryptographic construction, Merkle hash tree, for usage rights compliance. We integrated our framework with OpenStack and demonstrate that Metered Boot is able to achieve high scalability and low overhead for instantiating virtual network functions (VNFs).
Arun Raghuramu, Lianjie Cao, Puneet Sharma 0001, Joon-Myung Kang, Chen-Nee Chuah, Vinay Saxena
IEEE Trans. Netw. Serv. Manag.2
2021 Co-locating containerized workload using service mesh telemetry
abstract
The cloud-native architecture and container-based technologies are revolutionizing how online services and applications are designed, developed, and managed by offering better elasticity and flexibility to developers and operators. However, the increasing adoption of microservice and serverless designs makes application workload more decomposed and transient at a larger scale. Most existing container orchestration systems still manage application workload based on simple system-level resource usage and policies manually created by operators, leading to ineffective application-agnostic scheduling and extra management burden for operators.
Lianjie Cao, Puneet Sharma 0001
CoNEXT1
2019 Data-driven Resource Allocation in Virtualized Environments
Lianjie Cao, Sonia Fahmy, Puneet Sharma 0001
IM1
2018 Data-driven resource flexing for network functions visualization
abstract
Resource flexing is the notion of allocating resources on-demand as workload changes. This is a key advantage of Virtualized Network Functions (VNFs) over their non-virtualized counterparts. However, it is difficult to balance the timeliness and resource efficiency when making resource flexing decisions due to unpredictable workloads and complex VNF processing logic.
Lianjie Cao, Sonia Fahmy, Puneet Sharma 0001, Shandian Zhe
ANCS1
2017 Towards High Fidelity Network Emulation
abstract
Instantiating a distributed application that involves extensive inter-node communication onto a network is a challenging task. In this work, we focus on the special case of mapping a network emulation experiment onto a cluster comprising several (possibly heterogeneous) physical machines. We automatically profile the available physical machine resources, and use this information, together with the characteristics of the experimental topology, to determine an efficient mapping that preserves performance fidelity. We design an algorithm, which we call the “Waterfall” algorithm, and integrate it into a complete framework for profiling and mapping. We demonstrate the effectiveness of our framework via simulations and two sets of Crossfire Distributed Denial of Service attack testbed experiments.
Lianjie Cao, Xiangyu Bu, Sonia Fahmy, Siyuan Cao
ICCCN1
2013 PhishLive: A View of Phishing and Malware Attacks from an Edge Router
Lianjie Cao, Thibaut Probst, Ramana Rao Kompella
PAM1