VLDB 2026 Research / reviewers in the wild / expert
Juncheng Yang
dblp:183/0795
· DBLP profile ↗
30ranked-venue papers
14as first author
22since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 8 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 6 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 4 first-author · 4 since 2021Computer networks · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ZipLLM: Efficient LLM Storage via Model-Aware Synergistic Data Deduplication and Compression
Tingfeng Lan, Zhaoyuan Su, Juncheng Yang, Yue Cheng 0001 |
NSDI | 4 |
| 2025 | Helix: Serving Large Language Models over Heterogeneous GPUs and Network via Max-FlowabstractThis paper introduces Helix, a distributed system for high-throughput, low-latency large language model (LLM) serving in heterogeneous GPU clusters. The key idea behind Helix is to formulate inference computation of LLMs over heterogeneous GPUs and network connections as a max-flow problem on directed, weighted graphs, whose nodes represent GPU instances and edges capture both GPU and network heterogeneity through their capacities. Helix then uses a mixed integer linear programming (MILP) algorithm to discover highly optimized strategies to serve LLMs on heterogeneous GPUs. This approach allows Helix to jointly optimize model placement and request scheduling, two highly entangled tasks in heterogeneous LLM serving. Our evaluation on several heterogeneous clusters ranging from 24 to 42 GPU nodes shows that Helix improves serving throughput by up to 3.3x and reduces prompting and decoding latency by up to 66% and 24%, respectively, compared to existing approaches. Helix is available at https://github.com/Thesys-lab/Helix-ASPLOS25. Yixuan Mei, Yonghao Zhuang 0001, Xupeng Miao, Juncheng Yang, K. V. Rashmi |
ASPLOS (1) | 4 |
| 2025 | Demystifying and Improving Lazy Promotion in Cache Eviction
Qinghan Chen, Muhammad Haekal Muhyidin Al-Araby, Ziyue Qiu, K. V. Rashmi, Juncheng Yang |
Proc. VLDB Endow. | 6 |
| 2024 | Soft-Prompting with Graph-of-Thought for Multi-modal Representation LearningabstractThe chain-of-thought technique has been received well in multi-modal tasks. It is a step-by-step linear reasoning process that adjusts the length of the chain to improve the performance of generated prompts. However, human thought processes are predominantly non-linear, as they encompass multiple aspects simultaneously and employ dynamic adjustment and updating mechanisms. Therefore, we propose a novel Aggregation-Graph-of-Thought (AGoT) mechanism for soft-prompt tuning in multi-modal representation learning. The proposed AGoT models the human thought process not only as a chain but also models each step as a reasoning aggregation graph to cope with the overlooked multiple aspects of thinking in single-step reasoning. This turns the entire reasoning process into prompt aggregation and prompt flow operations. Experiments show that our multi-modal model enhanced with AGoT soft-prompting achieves good results in several tasks such as text-image retrieval, visual question answering, and image recognition. In addition, we demonstrate that it has good domain generalization performance due to better reasoning. Juncheng Yang, Zuchao Li, Shuai Xie, Wei Yu 0009, Shijun Li 0001, Bo Du 0001 |
LREC/COLING | 1 |
| 2024 | Cross-Modal Adapter: Parameter-Efficient Transfer Learning Approach for Vision-Language ModelsabstractAdapter-based parameter-efficient transfer learning has achieved exciting results in vision-language models. Traditional adapter methods often require training or fine-tuning, facing challenges such as insufficient samples or resource limitations. While some methods overcome the need for training by leveraging image modality cache and retrieval, they overlook the text modality’s importance and cross-modal cues for the efficient adaptation of parameters in visual-language models. This work introduces a cross-modal parameter-efficient approach named XMAdapter. XMAdapter establishes cache models for both text and image modalities. It then leverages retrieval through visual-language bimodal information to gather clues for inference. By dynamically adjusting the affinity ratio, it achieves cross-modal fusion, decoupling different modal similarities to assess their respective contributions. Additionally, it explores hard samples based on differences in cross-modal affinity and enhances model performance through adaptive adjustment of sample learning intensity. Extensive experimental results on benchmark datasets demonstrate that XMAdapter outperforms previous adapter-based methods significantly regarding accuracy, generalization, and efficiency. Juncheng Yang, Zuchao Li, Shuai Xie, Weiping Zhu 0004, Wei Yu 0009, Shijun Li 0001 |
ICME | 1 |
| 2024 | DSDRNet: Disentangling Representation and Reconstruct Network for Domain GeneralizationabstractDomain generalization faces challenges due to the distribution shift between training and testing sets, and the presence of unseen target domains. Common solutions include domain alignment, meta-learning, data augmentation, or ensemble learning, all of which rely on domain labels or domain adversarial techniques. In this paper, we propose a Dual-Stream Separation and Reconstruction Network, dubbed DSDRNet. It is a disentanglement-reconstruction approach that integrates features of both inter-instance and intra-instance through dual-stream fusion. The method introduces novel supervised signals by combining inter-instance semantic distance and intra-instance similarity. Incorporating Adaptive Instance Normalization (AdaIN) into a two-stage cyclic reconstruction process enhances self-disentangled reconstruction signals to facilitate model convergence. Extensive experiments on four benchmark datasets demonstrate that DSDRNet outperforms other popular methods in terms of domain generalization capabilities. Juncheng Yang, Qiwei Li 0002, Shuai Xie |
IJCNN | 1 |
| 2024 | SIEVE is Simpler than LRU: an Efficient Turn-Key Eviction Algorithm for Web Caches
Yazhuo Zhang, Juncheng Yang, Yao Yue, Ymir Vigfusson, K. V. Rashmi |
NSDI | 2 |
| 2024 | Generalizing to unseen domains via PatchMix
Juncheng Yang, Zuchao Li, Shuai Xie, Wei Yu 0009, Shijun Li 0001 |
Multim. Syst. | 1 |
| 2024 | A Time-Series-Based Sample Amplification Model for Data Stream with Sparse SamplesabstractAbstract The data stream is a dynamic collection of data that changes over time, and predicting the data class can be challenging due to sparse samples, complex interdependent characteristics between data, and random fluctuations. Accurately predicting the data stream in sparse data can create complex challenges. Due to its incremental learning nature, the neural networks suitable approach for streaming visualization. However, the high computational cost limits their applicability to high-speed streams, which has not yet been fully explored in the existing approaches. To solve these problems, this paper proposes an end-to-end dynamic separation neural network (DSN) approach based on the characteristics of data stream fluctuations, which expands the static sample at a given moment into a sequence of sample streams in the time dimension, thereby increasing the sparse samples. The Temporal Augmentation Module (TAM) can overcome these challenges by modifying the sparse data stream and reducing time complexity. Moreover, a neural network that uses a Variance Detection Module (VDM) can effectively detect the variance of the input data stream through the network and dynamically adjust the degree of differentiation between samples to enhance the accuracy of forecasts. The proposed method adds significant information regarding the data sparse samples and enhances low dimensional samples to high data samples to overcome the sparse data stream problem. In VDM the preprocessed data achieve data augmentation and the samples are transmitted to VDM. The proposed method is evaluated using different types of data streaming datasets to predict the sparse data stream. Experimental results demonstrate that the proposed method achieves a high prediction accuracy and that the data stream has significant effects and strong robustness compared to other existing approaches. Juncheng Yang, Wei Yu 0009, Shijun Li 0001 |
Neural Process. Lett. | 1 |
| 2023 | LatenSeer: Causal Modeling of End-to-End Latency Distributions by Harnessing Distributed TracingabstractEnd-to-end latency estimation in web applications is crucial for system operators to foresee the effects of potential changes, helping ensure system stability, optimize cost, and improve user experience. However, estimating latency in microservices-based architectures is challenging due to the complex interactions between hundreds or thousands of loosely coupled microservices. Current approaches either track only latency-critical paths or require laborious bespoke instrumentation, which is unrealistic for end-to-end latency estimation in complex systems. Yazhuo Zhang, Rebecca Isaacs, Yao Yue, Juncheng Yang, Lei Zhang 0223, Ymir Vigfusson |
SoCC | 4 |
| 2023 | FrozenHot Cache: Rethinking Cache Management for Modern HardwareabstractCaching is crucial for accelerating data access, employed as a ubiquitous design in modern systems at many parts of computer systems. With increasing core count, and shrinking latency gap between cache and modern storage devices, hit-path scalability becomes increasingly critical. However, existing production in-memory caches often use list-based management with promotion on each cache hit, which requires extensive locking and poses a significant overhead for scaling beyond a few cores. Moreover, existing techniques for improving scalability either (1) only focus on the indexing structure and do not improve cache management scalability, or (2) sacrifice efficiency or miss-path scalability. Ziyue Qiu, Juncheng Yang, Juncheng Zhang, Cheng Li 0001, Xiaosong Ma, Qi Chen 0009, Mao Yang 0004, Yinlong Xu 0001 |
EuroSys | 2 |
| 2023 | GL-Cache: Group-level learning for efficient and high-performance caching
Juncheng Yang, Ziming Mao, Yao Yue, K. V. Rashmi |
FAST | 1 |
| 2023 | FIFO can be Better than LRU: the Power of Lazy Promotion and Quick DemotionabstractLRU has been the basis of cache eviction algorithms for decades, with a plethora of innovations on improving LRU's miss ratio and throughput. While it is well-known that FIFO-based eviction algorithms provide significantly better throughput and scalability, they lag behind LRU on miss ratio, thus, cache efficiency. Juncheng Yang, Ziyue Qiu, Yazhuo Zhang, Yao Yue, K. V. Rashmi |
HotOS | 1 |
| 2023 | MAPLE: Semi-Supervised Learning with Multi-Alignment and Pseudo-LearningabstractData augmentation has undoubtedly enabled a significant leap forward in training a high-accuracy deep network. Besides the commonly used augmentation to target data, e.g., random cropping, flipping, and rotation, recent works have been dedicated to mining generalized knowledge by using multiple sources. However, along with plentiful data comes the huge data distribution gap between the target and different sources (hybrid shift). To mitigate this problem, existing methods tend to manually annotate more data. Unlike previous methods, this paper focuses on the study of learning deep models by gathering knowledge from multiple sources in a labor-free fashion and further proposes the "Multi-Alignment and Pseudo-Learning'' method, dubbed MAPLE. MAPLE constructs the multi-alignment module, which consists of multiple discriminators to align different data distributions via an adversarial process. In addition, a novel semi-supervised learning (SSL) manner is introduced to further facilitate the utility of our MAPLE. Extensive evaluations conducted on four benchmarks show the effectiveness of the proposed MAPLE, which achieves state-of-the-art performance outperforming existing methods by an obvious margin. Juncheng Yang, Zuchao Li, Wei Yu 0009, Bo Du 0001, Shijun Li 0001 |
KDD | 1 |
| 2023 | FIFO queues are all you need for cache evictionabstractAs a cache eviction algorithm, FIFO has a lot of attractive properties, such as simplicity, speed, scalability, and flash-friendliness. The most prominent criticism of FIFO is its low efficiency (high miss ratio). Juncheng Yang, Yazhuo Zhang, Ziyue Qiu, Yao Yue, K. V. Rashmi |
SOSP | 1 |
| 2023 | Efficient Fault Tolerance for Recommendation Model Training via Erasure CodingabstractDeep-learning-based recommendation models (DLRMs) are widely deployed to serve personalized content. In addition to using neural networks, DLRMs have large, sparsely-accessed embedding tables, which map categorical features to a learned dense representation. Due to the large sizes of embedding tables, DLRM training is typically distributed across the memory of tens or hundreds of nodes. Node failures are common in such large systems and must be mitigated to enable training to complete within production deadlines. Checkpointing is the primary approach used for fault tolerance in these systems, but incurs significant time overhead both during normal operation and when recovering from failures. As these overheads increase with DLRM size, checkpointing is slated to become an even larger overhead for future DLRMs, which are expected to grow. This calls for rethinking fault tolerance in DLRM training. We present ECRec, a DLRM training system that achieves efficient fault tolerance by coupling erasure coding with the unique characteristics of DLRM training. ECRec takes a hybrid approach between erasure coding and replicating different DLRM parameters, correctly and efficiently updates redundant parameters, and enables training to proceed without pauses, while maintaining the consistency of the recovered parameters. We implement ECRec atop XDL, an open-source, industrial-scale DLRM training system. Compared to checkpointing, ECRec reduces training-time overhead on large DLRMs by up to 66%, recovers from failure up to 9.8× faster, and continues training during recovery with only a 7--13% drop in throughput (whereas checkpointing must pause). Jack Kosaian, Juncheng Yang, K. V. Rashmi |
Proc. VLDB Endow. | 4 |
| 2022 | C2DN: How to Harness Erasure Codes at the Edge for Efficient Content Delivery
Juncheng Yang, Anirudh Sabnis, Daniel S. Berger, K. V. Rashmi, Ramesh K. Sitaraman |
NSDI | 1 |
| 2022 | Kangaroo: Theory and Practice of Caching Billions of Tiny Objects on FlashabstractMany social-media and IoT services have very large working sets consisting of billions of tiny (≈100 B) objects. Large, flash-based caches are important to serving these working sets at acceptable monetary cost. However, caching tiny objects on flash is challenging for two reasons: (i) SSDs can read/write data only in multi-KB “pages” that are much larger than a single object, stressing the limited number of times flash can be written; and (ii) very few bits per cached object can be kept in DRAM without losing flash’s cost advantage. Unfortunately, existing flash-cache designs fall short of addressing these challenges: write-optimized designs require too much DRAM, and DRAM-optimized designs require too many flash writes. We present Kangaroo , a new flash-cache design that optimizes both DRAM usage and flash writes to maximize cache performance while minimizing cost. Kangaroo combines a large, set-associative cache with a small, log-structured cache. The set-associative cache requires minimal DRAM, while the log-structured cache minimizes Kangaroo’s flash writes. Experiments using traces from Meta and Twitter show that Kangaroo achieves DRAM usage close to the best prior DRAM-optimized design, flash writes close to the best prior write-optimized design, and miss ratios better than both. Kangaroo’s design is Pareto-optimal across a range of allowed write rates, DRAM sizes, and flash sizes, reducing misses by 29% over the state of the art. These results are corroborated by analytical models presented herein and with a test deployment of Kangaroo in a production flash cache at Meta. Sara McAllister, Benjamin Berg, Julian Tutuncu-Macias, Juncheng Yang, Sathya Gunasekar, Jimmy Lu, Daniel S. Berger, Nathan Beckmann, Gregory R. Ganger |
ACM Trans. Storage | 4 |
| 2021 | Segcache: a memory-efficient and scalable in-memory key-value cache for small objects
Juncheng Yang, Yao Yue, K. V. Rashmi |
NSDI | 1 |
| 2021 | Kangaroo: Caching Billions of Tiny Objects on FlashabstractMany social-media and IoT services have very large working sets consisting of billions of tiny (≈100 B) objects. Large, flash-based caches are important to serving these working sets at acceptable monetary cost. However, caching tiny objects on flash is challenging for two reasons: (i) SSDs can read/write data only in multi-KB "pages" that are much larger than a single object, stressing the limited number of times flash can be written; and (ii) very few bits per cached object can be kept in DRAM without losing flash's cost advantage. Unfortunately, existing flash-cache designs fall short of addressing these challenges: write-optimized designs require too much DRAM, and DRAM-optimized designs require too many flash writes. Sara McAllister, Benjamin Berg, Julian Tutuncu-Macias, Juncheng Yang, Sathya Gunasekar, Jimmy Lu, Daniel S. Berger, Nathan Beckmann, Gregory R. Ganger |
SOSP | 4 |
| 2021 | Skyline Diagram: Efficient Space Partitioning for Skyline QueriesabstractSkyline queries are important in many application domains. In this paper, we propose a novel structure Skyline Diagram, which given a set of points, partitions the plane into a set of regions, referred to as skyline polyominos. All query points in the same skyline polyomino have the same skyline query results. Similar to kth-order Voronoi diagram commonly used to facilitate k nearest neighbor (kNN) queries, skyline diagram can be used to facilitate skyline queries and many other applications. However, it may be computationally expensive to build the skyline diagram. By exploiting some interesting properties of skyline, we present several efficient algorithms for building the diagram with respect to three kinds of skyline queries, quadrant, global, and dynamic skylines. In addition, we propose an approximate skyline diagram which can significantly reduce the space cost. Experimental results on both real and synthetic datasets show that our algorithms are efficient and scalable. Jinfei Liu, Juncheng Yang, Li Xiong 0001, Jian Pei 0001, Jun Luo 0007, Yuzhang Guo, Shuaicheng Ma, Chenglin Fan |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2021 | A Large-scale Analysis of Hundreds of In-memory Key-value Cache Clusters at TwitterabstractModern web services use in-memory caching extensively to increase throughput and reduce latency. There have been several workload analyses of production systems that have fueled research in improving the effectiveness of in-memory caching systems. However, the coverage is still sparse considering the wide spectrum of industrial cache use cases. In this work, we significantly further the understanding of real-world cache workloads by collecting production traces from 153 in-memory cache clusters at Twitter, sifting through over 80 TB of data, and sometimes interpreting the workloads in the context of the business logic behind them. We perform a comprehensive analysis to characterize cache workloads based on traffic pattern, time-to-live (TTL), popularity distribution, and size distribution. A fine-grained view of different workloads uncover the diversity of use cases: many are far more write-heavy or more skewed than previously shown and some display unique temporal patterns. We also observe that TTL is an important and sometimes defining parameter of cache working sets. Our simulations show that ideal replacement strategy in production caches can be surprising, for example, FIFO works the best for a large number of workloads. Juncheng Yang, Yao Yue, K. V. Rashmi |
ACM Trans. Storage | 1 |
| 2020 | PACEMAKER: Avoiding HeART attacks in storage clusters with disk-adaptive redundancy
Saurabh Kadekodi, Francisco Maturana, Suhas Jayaram Subramanya, Juncheng Yang, K. V. Rashmi, Gregory R. Ganger |
OSDI | 4 |
| 2020 | A large scale analysis of hundreds of in-memory cache clusters at Twitter
Juncheng Yang, Yao Yue, K. V. Rashmi |
OSDI | 1 |
| 2019 | Secure and Efficient Skyline Queries on Encrypted DataabstractOutsourcing data and computation to cloud server provides a cost-effective way to support large scale data storage and query processing. However, due to security and privacy concerns, sensitive data (e.g., medical records) need to be protected from the cloud server and other unauthorized users. One approach is to outsource encrypted data to the cloud server and have the cloud server perform query processing on the encrypted data only. It remains a challenging task to support various queries over encrypted data in a secure and efficient way such that the cloud server does not gain any knowledge about the data, query, and query result. In this paper, we study the problem of secure skyline queries over encrypted data. The skyline query is particularly important for multi-criteria decision making but also presents significant challenges due to its complex computations. We propose a fully secure skyline query protocol on data encrypted using semantically-secure encryption. As a key subroutine, we present a new secure dominance protocol, which can be also used as a building block for other queries. Furthermore, we demonstrate two optimizations, data partitioning and lazy merging, to further reduce the computation load. Finally, we provide both serial and parallelized implementations and empirically study the protocols in terms of efficiency and scalability under different parameter settings, verifying the feasibility of our proposed solutions. Jinfei Liu, Juncheng Yang, Li Xiong 0001, Jian Pei 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2018 | Mutant: Balancing Storage Cost and Latency in LSM-Tree Data StoresabstractToday's cloud database systems are not designed for seamless cost-performance trade-offs for changing SLOs. Database engineers have a limited number of trade-offs due to the limited storage types offered by cloud vendors, and switching to a different storage type requires a time-consuming data migration to a new database. We propose Mutant, a new storage layer for log-structured merge tree (LSM-tree) data stores that dynamically balances database cost and performance by organizing SSTables (files that store a subset of records) into different storage types based on SSTable access frequencies. We implemented Mutant by extending RocksDB and found in our evaluation that Mutant delivers seamless cost-performance trade-offs with the YCSB workload and a real-world workload trace. Moreover, through additional optimizations, Mutant lowers the user-perceived latency significantly compared with the unmodified database. Hobin Yoon, Juncheng Yang, Sveinn Fannar Kristjansson, Steinn E. Sigurðarson, Ymir Vigfusson, Ada Gavrilovska |
SoCC | 2 |
| 2018 | Skyline Diagram: Finding the Voronoi Counterpart for Skyline QueriesabstractSkyline queries are important in many application domains. In this paper, we propose a novel structure Skyline Diagram, which given a set of points, partitions the plane into a set of regions, referred to as skyline polyominos. All query points in the same skyline polyomino have the same skyline query results. Similar to k^th-order Voronoi diagram commonly used to facilitate k nearest neighbor (kNN) queries, skyline diagram can be used to facilitate skyline queries and many other applications. However, it may be computationally expensive to build the skyline diagram. By exploiting some interesting properties of skyline, we present several efficient algorithms for building the diagram with respect to three kinds of skyline queries, quadrant, global, and dynamic skylines. Experimental results on both real and synthetic datasets show that our algorithms are efficient and scalable. Jinfei Liu, Juncheng Yang, Li Xiong 0001, Jian Pei 0001, Jun Luo 0007 |
ICDE | 2 |
| 2017 | Mithril: mining sporadic associations for cache prefetchingabstractThe growing pressure on cloud application scalability has accentuated storage performance as a critical bottleneck. Although cache replacement algorithms have been extensively studied, cache prefetching - reducing latency by retrieving items before they are actually requested - remains an underexplored area. Existing approaches to history-based prefetching, in particular, provide too few benefits for real systems for the resources they cost. Juncheng Yang, Reza Karimi, Trausti Saemundsson, Avani Wildani, Ymir Vigfusson |
SoCC | 1 |
| 2017 | Secure Skyline Queries on Cloud PlatformabstractOutsourcing data and computation to cloud server provides a cost-effective way to support large scale data storage and query processing. However, due to security and privacy concerns, sensitive data (e.g., medical records) need to be protected from the cloud server and other unauthorized users. One approach is to outsource encrypted data to the cloud server and have the cloud server perform query processing on the encrypted data only. It remains a challenging task to support various queries over encrypted data in a secure and efficient way such that the cloud server does not gain any knowledge about the data, query, and query result. In this paper, we study the problem of secure skyline queries over encrypted data. The skyline query is particularly important for multi-criteria decision making but also presents significant challenges due to its complex computations. We propose a fully secure skyline query protocol on data encrypted using semantically-secure encryption. As a key subroutine, we present a new secure dominance protocol, which can be also used as a building block for other queries. Finally, we provide both serial and parallelized implementations and empirically study the protocols in terms of efficiency and scalability under different parameter settings, verifying the feasibility of our proposed solutions. Jinfei Liu, Juncheng Yang, Li Xiong 0001, Jian Pei 0001 |
ICDE | 2 |
| 2016 | Enabling Space Elasticity in Storage SystemsabstractStorage systems are designed to never lose data. However, modern applications increasingly use local storage to improve performance by storing soft state such as cached, prefetched or precomputed results. Required is elastic storage, where cloud providers can alter the storage footprint of applications by removing and regenerating soft state based on resource availability and access patterns. We propose a new abstraction called a motif that enables storage elasticity by allowing applications to describe how soft state can be regenerated. Carillon is a system that uses motifs to dynamically change the storage space used by applications. Carillon is implemented as a runtime and a collection of shim layers that interpose between applications and specific storage APIs; we describe shims for a filesystem (Carillon-FS) and a key-value store (Carillon-KV). We show that Carillon-FS allows us to dynamically alter the storage footprint of a VM, while Carillon-KV enables a graph database that accelerates performance based on available storage space. Helgi Sigurbjarnarson, Pétur Orri Ragnarsson, Juncheng Yang, Ymir Vigfusson, Mahesh Balakrishnan 0001 |
SYSTOR | 3 |