VLDB 2026 Research / reviewers in the wild / expert
Ziheng Jiang
dblp:14/8980
· DBLP profile ↗
18ranked-venue papers
5as first author
9since 2021 · last 2026
0009-0001-7732-4391ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 4 first-author · 2 since 2021Software engineering, systems software and programming languages · 4 · 3 since 2021Systems, architecture and hardware · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SwiftSpec: Disaggregated Speculative Decoding and Fused Kernels for Low-Latency LLM InferenceabstractLow-latency, single-request decoding of large language models is critical for interactive systems with tight SLA demands. Prior work reduces latency through speculative decoding (combining a small draft model with a larger target model), but the draft model remains on the critical path, and communication overhead limits scaling across GPUs due to the small batch size associated with single-request decoding. To address these limitations, this paper introduces SwiftSpec: a system architecture that disaggregates draft and target models across homogeneous GPUs within a single node and utilizes NCCL-low-latency primitives directly to improve the performance of core GEMM and attention kernels. Our implementation includes 3k lines of custom CUDA for fused kernels and an evolving tree cache for KV-cache consistency and maximized reuse between draft and target models. On a single 8×H800 GPU node, SwiftSpec achieves 347 tokens/s for Llama-3-70B---1.3× faster than NVIDIA's own benchmarks on a higher-performance 8×H200 setup---and averages 1.75× faster decoding than state-of-the-art speculative decoding across five model families and six datasets. Specifically, we find that for Llama-3-70B SwiftSpec is significantly faster across all 480 tested queries, showing 1.7× speedup over the best open-source baseline for 95th percentile requests. Code for SwiftSpec will be available at https://github.com/ByteDance-Seed/SwiftSpec Ziheng Jiang, Chengquan Jiang, Menghan Yu, Size Zheng 0001, Haibin Lin, Xin Liu 0086, Henry Hoffmann |
ASPLOS (2) | 2 |
| 2026 | MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in ProductionabstractWe present MegaScale-MoE, a production system tailored for the efficient training of large-scale mixture-of-experts (MoE) models. MoE emerges as a promising architecture to scale large language models (LLMs) to unprecedented sizes, thereby enhancing model performance. However, existing MoE training systems experience a degradation in training efficiency, exacerbated by the escalating scale of MoE models and the continuous evolution of hardware. Chao Jin 0007, Ziheng Jiang, Zhihao Bai, Juncai Liu, Xiang Li 0067, Ningxin Zheng, Qi Huang 0001, Wen Heng, Yiyuan Ma, Wenlei Bao, Size Zheng 0001, Xuegui Zheng, Yanghua Peng, Haibin Lin, Xuanzhe Liu, Xin Jin 0008, Xin Liu 0086 |
EuroSys | 2 |
| 2025 | Relax: Composable Abstractions for End-to-End Dynamic Machine LearningabstractDynamic shape computations have become critical in modern machine learning workloads, especially in emerging large language models. The success of these models has driven the demand for their universal deployment across a diverse set of backend environments. In this paper, we present Relax, a compiler abstraction for optimizing end-to-end dynamic machine learning workloads. Relax introduces a cross-level abstraction that encapsulates computational graphs, loop-level tensor programs, and external library calls in a single representation. Relax also introduces first-class symbolic shape annotations to track dynamic shape computations globally across the program, enabling dynamic shape-aware cross-level optimizations. We build an end-to-end compilation framework using the proposed approach to optimize dynamic shape models. Experimental results on LLMs show that Relax delivers performance competitive with state-of-the-art systems across various GPUs and enables deployment of emerging models to a broader set of emerging environments, including mobile phones, embedded devices, and web browsers. Ruihang Lai, Junru Shao, Siyuan Feng 0007, Steven Lyubomirsky, Bohan Hou, Wuwei Lin, Zihao Ye 0001, Hongyi Jin, Jiawei Liu 0004, Lesheng Jin, Yaxing Cai, Ziheng Jiang, Sunghyun Park 0004, Prakalp Srivastava, Jared Roesch, Todd C. Mowry, Tianqi Chen 0001 |
ASPLOS (2) | 13 |
| 2025 | Understanding Stragglers in Large Model Training Using What-if Analysis
Jinkun Lin, Ziheng Jiang, Zuquan Song, Sida Zhao, Menghan Yu, Zhanghan Wang, Zuocheng Shi, Zherui Liu, Shuguang Wang, Haibin Lin, Xin Liu 0086, Aurojit Panda, Jinyang Li 0001 |
OSDI | 2 |
| 2025 | MegaScale-Infer: Efficient Mixture-of-Experts Model Serving with Disaggregated Expert ParallelismabstractMixture-of-Experts (MoE) showcases tremendous potential to scale large language models (LLMs) with enhanced performance and reduced computational complexity. However, its sparsely activated architecture shifts feed-forward networks (FFNs) from being compute-intensive to memory-intensive during inference, leading to substantially lower GPU utilization and increased operational costs. Ruidong Zhu, Ziheng Jiang, Chao Jin 0007, Cesar A. Stuardo, Huaping Zhou, Jianzhe Xiao, Lingjun Liu, Haibin Lin, Li-Wen Chang, Jianxi Ye, Xuanzhe Liu, Xin Jin 0008, Xin Liu 0086 |
SIGCOMM | 2 |
| 2024 | Towards Automated Chinese Ancient Character Restoration: A Diffusion-Based Method with a New DatasetabstractAutomated Chinese ancient character restoration (ACACR) remains a challenging task due to its historical significance and aesthetic complexity. Existing methods are constrained by non-professional masks and even overfitting when training on small-scale datasets, which hinder their interdisciplinary application to traditional fields. In this paper, we are proud to introduce the Chinese Ancient Rubbing and Manuscript Character Dataset (ARMCD), which consists of 15,553 real-world ancient single-character images with 42 rubbings and manuscripts, covering the works of over 200 calligraphy artists spanning from 200 to 1,800 AD. We are also dedicated to providing professional synthetic masks by extracting localized erosion from real eroded images. Moreover, we propose DiffACR (Diffusion model for automated Chinese Ancient Character Restoration), a diffusion-based method for the ACACR task. Specifically, we regard the synthesis of eroded images as a special form of cold diffusion on uneroded ones and extract the prior mask directly from the eroded images. Our experiments demonstrate that our method comprehensively outperforms most existing methods on the proposed ARMCD. Dataset and code are available at https://github.com/lhl322001/DiffACR. Chenghao Du, Ziheng Jiang, Jiawei Ma, Chen Ye 0002 |
AAAI | 3 |
| 2024 | MegaScale: Scaling Large Language Model Training to More Than 10, 000 GPUs
Ziheng Jiang, Haibin Lin, Yinmin Zhong, Qi Huang 0001, Yangrui Chen, Zhi Zhang 0005, Yanghua Peng, Xiang Li 0067, Shibiao Nong, Yulu Jia, Sun He, Hongmin Chen, Zhihao Bai, Qi Hou, Shipeng Yan, Yiyao Sheng, Zhuo Jiang, Haohan Xu, Zhang Zhang 0003, Pengfei Nie, Leqi Zou, Sida Zhao, Zherui Liu, Xiaoying Jia 0001, Jianxi Ye, Xin Jin 0008, Xin Liu 0086 |
NSDI | 1 |
| 2023 | EfficientPhys: Enabling Simple, Fast and Accurate Camera-Based Cardiac MeasurementabstractCamera-based physiological measurement is a growing field with neural models providing state-of-the-art performance. Prior research has explored various "end-to-end" architectures; however these methods still require several preprocessing steps and are not able to run directly on mobile and edge devices. The operations are often non-trivial to implement, making replication and deployment difficult and can even have a higher computational budget than the "core" network itself. In this paper, we propose two novel and efficient neural models for camera-based physiological measurement called EfficientPhys that remove the need for face detection, segmentation, normalization, color space transformation or any other preprocessing steps. Using an input of raw video frames, our models achieve strong accuracy on three public datasets. We show that this is the case whether using a transformer or convolutional backbone. We further evaluate the latency of the proposed networks and show that our most lightweight network also achieves a 33% improvement in efficiency. Xin Liu 0034, Brian L. Hill, Ziheng Jiang, Shwetak N. Patel, Daniel McDuff |
WACV | 3 |
| 2021 | Characterizing Structural Regularities of Labeled Data in Overparameterized ModelsabstractHumans are accustomed to environments that contain both regularities and exceptions. For example, at most gas stations, one pays prior to pumping, but the occasional rural station does not accept payment in advance. Likewise, deep neural networks can generalize across instances that share common patterns or structures, yet have the capacity to memorize rare or irregular forms. We analyze how individual instances are treated by a model via a consistency score. The score characterizes the expected accuracy for a held-out instance given training sets of varying size sampled from the data distribution. We obtain empirical estimates of this score for individual instances in multiple data sets, and we show that the score identifies out-of-distribution and mislabeled examples at one end of the continuum and strongly regular examples at the other end. We identify computationally inexpensive proxies to the consistency score using statistics collected during training. We apply the score toward understanding the dynamics of representation learning and to filter outliers during training. Ziheng Jiang, Chiyuan Zhang, Kunal Talwar, Michael C. Mozer |
ICML | 1 |
| 2018 | Learning to Optimize Tensor ProgramsabstractWe introduce a learning-based framework to optimize tensor programs for deep learning workloads. Efficient implementations of tensor operators, such as matrix multiplication and high dimensional convolution are key enablers of effective deep learning systems. However, existing systems rely on manually optimized libraries such as cuDNN where only a narrow range of server class GPUs are well-supported. The reliance on hardware specific operator libraries limits the applicability of high-level graph optimizations and incurs significant engineering costs when deploying to new hardware targets. We use learning to remove this engineering burden. We learn domain specific statistical cost models to guide the search of tensor operator implementations over billions of possible program variants. We further accelerate the search by effective model transfer across workloads. Experimental results show that our framework delivers performance competitive with state-of-the-art hand-tuned libraries for low-power CPU, mobile GPU, and server-class GPU. Tianqi Chen 0001, Lianmin Zheng, Eddie Q. Yan, Ziheng Jiang, Thierry Moreau, Luis Ceze, Carlos Guestrin, Arvind Krishnamurthy |
NeurIPS | 4 |
| 2018 | TVM: An Automated End-to-End Optimizing Compiler for Deep Learning
Tianqi Chen 0001, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Q. Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Luis Ceze, Carlos Guestrin, Arvind Krishnamurthy |
OSDI | 3 |
| 2018 | Multi-Dimensional Network Embedding with Hierarchical StructureabstractInformation networks are ubiquitous in many applications. A popular way to facilitate the information in a network is to embed the network structure into low-dimension spaces where each node is represented as a vector. The learned representations have been proven to advance various network analysis tasks such as link prediction and node classification. The majority of existing embedding algorithms are designed for the networks with one type of nodes and one dimension of relations among nodes. However, many networks in the real-world complex systems have multiple types of nodes and multiple dimensions of relations. For example, an e-commerce network can have users and items, and items can be viewed or purchased by users, corresponding to two dimensions of relations. In addition, some types of nodes can present hierarchical structure. For example, authors in publication networks are associated to affiliations; and items in e-commerce networks belong to categories. Most of existing methods cannot be naturally applicable to these networks. In this paper, we aim to learn representations for networks with multiple dimensions and hierarchical structure. In particular, we provide an approach to capture independent information from each dimension and dependent information across dimensions and propose a framework MINES, which performs Multi-dImension Network Embedding with hierarchical Structure. Experimental results on a network from a real-world e-commerce website demonstrate the effectiveness of the proposed framework. Yao Ma 0001, Zhaochun Ren, Ziheng Jiang, Jiliang Tang, Dawei Yin 0001 |
WSDM | 3 |
| 2018 | A Path-constrained Framework for Discriminating Substitutable and Complementary Products in E-commerceabstractIn personalized recommendation, candidate generation plays an infrastructural role by retrieving candidates out of billions of items. During this process, substitutes and complements constitute two main classes of retrieved candidates: substitutable products are interchangeable, whereas complementary products might be purchased together by users. Discriminating substitutable and complementary products is playing an increasingly important role in e-commerce portals by affecting the performance of candidate generation, e.g., when a user has browsed a t-shirt, it is reasonable to retrieve similar t-shirts, i.e., substitutes; whereas if the user has already purchased one, it would be better to retrieve trousers, hats or shoes, as complements of t-shirts. In this paper, we propose a path-constrained framework (PMSC) for discriminating substitutes and complements. Specifically, for each product, we first learn its embedding representations in a general semantic space. Thereafter, we project the embedding vectors into two separate spaces via a novel mapping function. In the end, we incorporate each embedding with path-constraints to further boost the discriminative ability of the model. Extensive experiments conducted on two e-commerce datasets show the effectiveness of our proposed method. Zihan Wang 0002, Ziheng Jiang, Zhaochun Ren, Jiliang Tang, Dawei Yin 0001 |
WSDM | 2 |
| 2014 | Locality-Constrained Low-Rank Coding for Image ClassificationabstractLow-rank coding (LRC), originated from matrix decomposition, is recently introduced into image classification. Following the standard bag-of-words (BOW) pipeline, when coding the data matrix in the sense of low-rankness incorporates contextual information into the traditional BOW model, this can capture the dependency relationship among neighbor patches. It differs from the traditional sparse coding paradigms which encode patches independently. Current LRC-based methods use l_1 norm to increase the discrimination and sparseness of the learned codes. However, such methods fail to consider the local manifold structure between dataspace and dictionary space. To solve this problem, we propose a locality-constrained low-rank coding (LCLR) algorithm for image representations. By using the geometric structure information as a regularization term,we can obtain more discriminative representations. In addition, we present a fast and stable online algorithmto solve the optimization problem. In the experiments,we evaluate LCLR with four benchmarks, including one face recognition dataset (extended Yale B), one handwrittendigit recognition dataset (USPS), and two image datasets (Scene13 for scene recognition and Caltech101 for object recognition). Experimental results show thatour approach outperforms many state-of-the-art algorithmseven with a linear classifier. Ziheng Jiang, Lihong Peng |
AAAI | 1 |
| 2013 | Learning open-domain comparable entity graphs from user search queriesabstractA frequent behavior of internet users is to compare among various comparable entities for decision making. As an instance, a user may compare among iPhone 5, Lumia 920 etc. products before deciding which cellphone to buy. However, it is a challenging problem to know what entities are generally comparable from the users' viewpoints in the open domain Web. In this paper, we propose a novel solution, which is known as Comparable Entity Graph Mining (CEGM), to learn an open-domain comparable entity graph from the user search queries. CEGM firstly mine seed comparable entity pairs from user search queries automatically using predefined query patterns. Next, it discovers more entity pairs with a confidence classifier in a bootstrapping fashion. Newly discovered entity pairs are organized into an open-domain comparable entity graph. Based on our empirical study over 1 billion queries of a commercial search engine, we build a comparable entity graph which covers 73.4% queries in the top 50 million unique queries of a commercial search engine. Through manual labeling in sampled sub-graphs, the average precision of comparable entities is 89.4%. As applications of the learned entity graph, the entity recommendation in Web search is empirically studied. Ziheng Jiang, Lei Ji 0001, Jun Yan 0001, Ping Guo 0002, Ning Liu 0001 |
CIKM | 1 |
| 2012 | Combining LVQ with SVM technique for image semantic annotation
Ping Guo 0002, Ziheng Jiang, Yao Yao 0005 |
Neural Comput. Appl. | 2 |
| 2011 | A study of block-global feature based supervised image annotationabstractIn order to get better semantic annotation performance, block-global features are extracted as low-level visual features for image semantic annotation. Specifically, wellknown global feature extraction method, namely two-dimensional principal component analysis (2DPCA) is applied to extract the image block-global features. Unlike typical image annotation methods which use local features or global features separately, we propose to extract global features from image local regions (block) with the expectation of: a) combining the advantages of local and global features; b) discovering multiple semantic meanings in one image. In the experiment, comparative studies have been done for the performance of block-global feature extraction methods with widely used local feature extraction method such as scale invariant feature transform. The results show that 2DPCA has a significantly better performance than the performance of other methods. Ziheng Jiang, Ping Guo 0002, Lixiong Liu |
SMC | 2 |
| 2010 | Feature data optimization with LVQ technique in semantic image annotationabstractIn order to improve the classifier performance in semantic image annotation, we propose a novel method which adopts learning vector quantization (LVQ) technique to optimize low level feature data extracted from given image. Some representative vectors are selected with LVQ to train support vector machine (SVM) classifier instead of using all feature data. Performance is compared between the methods with and without feature data optimization when SVM is applied to semantic image annotation. Experiment results show that the proposed method has a better performance than that without using LVQ technique. Ziheng Jiang, Ping Guo 0002 |
ISDA | 1 |