EDBT 2026 Demo / reviewers in the wild / expert
Zhekai Zhang
dblp:213/8176
· DBLP profile ↗
11ranked-venue papers
1as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 4 since 2021Systems, architecture and hardware · 5 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LEGO: Spatial Accelerator Generation and Optimization for Tensor ApplicationsabstractModern tensor applications, especially foundation models and generative AI applications require multiple input modalities (both vision and language), which increases the demand for flexible accelerator architecture. Existing frameworks suffer from the trade-off between design flexibility and productivity of RTL generation: either limited to very few hand-written templates or cannot automatically generate the RTL. To address this challenge, we propose the LEGO framework, which targets tensor applications and automatically generates spatial architecture design and outputs synthesizable RTL code without handwritten RTL design templates. Leveraging the affine-transformation-based architecture representation, LEGO front end finds interconnections between function units, synthesizes the memory system, and fuses different spatial dataflow designs based on data reuse analysis. LEGO back end then translates the hardware in a primitive-level graph to perform lower-level optimizations, and applies a set of linear-programming algorithms to optimally insert pipeline registers and reduce the overhead of unused logic when switching spatial dataflows. Our evaluation demonstrates that LEGO can achieve $3.2 \times$ speedup and $2.4 \times$ energy efficiency compared to previous work Gemmini, and can generate one architecture for diverse modern foundation models in generative AI applications. Yujun Lin 0001, Zhekai Zhang, Song Han 0003 |
HPCA | 2 |
| 2025 | SVDQuant: Absorbing Outliers by Low-Rank Component for 4-Bit Diffusion ModelsabstractDiffusion models can effectively generate high-quality images. However, as they scale, rising memory demands and higher latency pose substantial deployment challenges. In this work, we aim to accelerate diffusion models by quantizing their weights and activations to 4 bits. At such an aggressive level, both weights and activations are highly sensitive, where existing post-training quantization methods like smoothing become insufficient. To overcome this limitation, we propose *SVDQuant*, a new 4-bit quantization paradigm. Different from smoothing, which redistributes outliers between weights and activations, our approach *absorbs* these outliers using a low-rank branch. We first consolidate the outliers by shifting them from activations to weights. Then, we use a high-precision, low-rank branch to take in the weight outliers with Singular Value Decomposition (SVD), while a low-bit quantized branch handles the residuals. This process eases the quantization on both sides. However, naively running the low-rank branch independently incurs significant overhead due to extra data movement of activations, negating the quantization speedup. To address this, we co-design an inference engine *Nunchaku* that fuses the kernels of the low-rank branch into those of the low-bit branch to cut off redundant memory access. It can also seamlessly support off-the-shelf low-rank adapters (LoRAs) without re-quantization. Extensive experiments on SDXL, PixArt-$\Sigma$, and FLUX.1 validate the effectiveness of SVDQuant in preserving image quality. We reduce the memory usage for the 12B FLUX.1 models by 3.5×, achieving 3.0× speedup over the 4-bit weight-only quantization (W4A16) baseline on the 16GB laptop 4090 GPU with INT4 precision. On the latest RTX 5090 desktop with Blackwell architecture, we achieve a 3.1× speedup compared to the W4A16 model using NVFP4 precision. Our quantization library and inference engine are available at https://github.com/mit-han-lab/deepcompressor/ and https://github.com/mit-han-lab/nunchaku/, correspondingly. Yujun Lin 0001, Zhekai Zhang, Tianle Cai, Xiuyu Li, Junxian Guo, Enze Xie, Chenlin Meng, Jun-Yan Zhu, Song Han 0003 |
ICLR | 3 |
| 2025 | SANA: Efficient High-Resolution Text-to-Image Synthesis with Linear Diffusion TransformersabstractWe introduce Sana, a text-to-image framework that can efficiently generate images up to 4096$\times$4096 resolution. Sana can synthesize high-resolution, high-quality images with strong text-image alignment at a remarkably fast speed, deployable on laptop GPU. Core designs include: (1) Deep compression autoencoder: unlike traditional AEs, which compress images only 8$\times$, we trained an AE that can compress images 32$\times$, effectively reducing the number of latent tokens. (2) Linear DiT: we replace all vanilla attention in DiT with linear attention, which is more efficient at high resolutions without sacrificing quality. (3) Decoder-only text encoder: we replaced T5 with modern decoder-only small LLM as the text encoder and designed complex human instruction with in-context learning to enhance the image-text alignment. (4) Efficient training and sampling: we propose Flow-DPM-Solver to reduce sampling steps, with efficient caption labeling and selection to accelerate convergence. As a result, Sana-0.6B is very competitive with modern giant diffusion model (e.g. Flux-12B), being 20 times smaller and 100+ times faster in measured throughput. Moreover, Sana-0.6B can be deployed on a 16GB laptop GPU, taking less than 1 second to generate a 1024$\times$1024 resolution image. Sana enables content creation at low cost. Code and model will be publicly released upon publication. Enze Xie, Junsong Chen, Junyu Chen 0003, Han Cai, Haotian Tang, Yujun Lin 0001, Zhekai Zhang, Ligeng Zhu, Yao Lu 0006, Song Han 0003 |
ICLR | 7 |
| 2025 | SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion TransformerabstractThis paper presents SANA-1.5, a linear Diffusion Transformer for efficient scaling in text-to-image generation. Building upon SANA-1.0, we introduce three key innovations: (1) Efficient Training Scaling: A depth-growth paradigm that enables scaling from 1.6B to 4.8B parameters with significantly reduced computational resources, combined with a memory-efficient 8-bit optimizer. (2) Model Depth Pruning: A block importance analysis technique for efficient model compression to arbitrary sizes with minimal quality loss. (3) Inference-time Scaling: A repeated sampling strategy that trades computation for model capacity, enabling smaller models to match larger model quality at inference time. Through these strategies, SANA-1.5 achieves a text-image alignment score of 0.72 on GenEval, which can be further improved to 0.80 through inference scaling, establishing a new SoTA on GenEval benchmark. These innovations enable efficient model scaling across different compute budgets while maintaining high quality, making high-quality image generation more accessible. Enze Xie, Junsong Chen, Ligeng Zhu, Yujun Lin 0001, Zhekai Zhang, Junyu Chen 0003, Han Cai, Daquan Zhou, Song Han 0003 |
ICML | 7 |
| 2025 | Exploring the Effectiveness of Open-Source Donation Platform: An Empirical Study on OpencollectiveabstractABSTRACT In recent years, with the development of the open‐source community, various open‐source donation platforms have emerged. These platforms effectively alleviate the financial pressures faced by open‐source projects through diversified funding sources and flexible donation methods. As one of the most representative open‐source donation platforms, Opencollective has garnered widespread attention from both the open‐source community and academia. Although Opencollective claims to provide more funding opportunities for open‐source projects, the extent to which it effectively addresses the financial challenges faced by these projects remains unclear. While there have been studies on the effectiveness of traditional donation models, research on the effectiveness of emerging donation platforms such as Opencollective is still limited. Given that a large number of open‐source projects are urgently seeking donations, understanding the effectiveness of donations through Opencollective is crucial for these projects. To address this gap, we have made an early step in this direction. This paper conducts a comprehensive study on the effectiveness of donations through the Opencollective, employing a combination of quantitative and qualitative analysis and identifies the following key findings: (1) Opencollective attracts a diverse group of participants, including individual donors, sponsors, contributors, and project managers, with individual donors constituting the largest group. Most donations are concentrated in the range of $5 to $10, indicating that the platform largely relies on small but frequent donations from individuals. (2) Only about 26.61% of open‐source projects receive donations through Opencollective, with approximately 64.38% of these projects receiving a total donation amount of less than $50,000. The likelihood of receiving donations increases with project scale, maturity and the number of stars. Among projects that have received donations, larger projects with stronger social media promotion, greater attention and more issues are more likely to receive additional donations. (3) The positive impact of donations on project development and spend activities is significant only in the short term, with no notable long‐term effects. In contrast, donations do not have a significant short‐term impact on community engagement. Although the long‐term effect is slightly positive, it is not statistically significant. (4) The main shortcomings of Opencollective include insufficient project management and collaboration features, inadequate user experience and interface design, high transaction fees, and a lack of transparency in fund allocation and usage. Our findings provide significant theoretical support and practical recommendations for the effectiveness of emerging donation platforms and the sustainable development of open‐source projects. Shuoxiao Zhang, Enyi Tang, Zhekai Zhang, Yixiao Shan, Haofeng Zhang 0001, Xuandong Li |
J. Softw. Evol. Process. | 4 |
| 2024 | Lightening-Transformer: A Dynamically-Operated Optically-Interconnected Photonic Transformer AcceleratorabstractThe wide adoption and significant computing resource cost of attention-based transformers, e.g., Vision Transformers and large language models, have driven the demand for efficient hardware accelerators. While electronic accelerators have been commonly used, there is a growing interest in exploring photonics as an alternative technology due to its high energy efficiency and ultra-fast processing speed. Photonic accelerators have demonstrated promising results for convolutional neural networks (CNNs) workloads, which predominantly rely on weight-static linear operations. However, they encounter challenges when it comes to efficiently supporting attention-based Transformer architectures, raising questions about the applicability of photonics to advanced machine-learning tasks. The primary hurdle lies in their inefficiency in handling the unique workloads inherent to Transformers, i.e., dynamic and full-range tensor multiplication. In this work, we propose Lightening-Transformer, the first light-empowered, high-performance, and energy-efficient photonic Transformer accelerator. To overcome the fundamental limitation of existing photonic tensor core designs, we introduce a novel dynamically-operated photonic tensor core, DPTC, consisting of a crossbar array of interference-based optical vector dot-product engines, supporting highly parallel, dynamic, and full-range matrix multiplication. Furthermore, we design a dedicated accelerator that integrates our novel photonic computing cores with photonic interconnects for inter-core data broadcast, fully unleashing the power of optics. The comprehensive evaluation demonstrates that Lightening-Transformer achieves >2.6x energy and > 12 x latency reductions compared to prior photonic accelerators and delivers the lowest energy cost and 2 to 3 orders of magnitude lower energy-delay product compared to the electronic Transformer accelerator, all while maintaining digital-comparable accuracy. Our work highlights the immense potential of photonics for efficient hardware accelerators, particularly for advanced machine-learning workloads, such as Transformer-backboned large language models (LLM). Our implementation is available at https://github.com/zhuhanqing/Lightening-Transformer. Hanqing Zhu, Jiaqi Gu 0002, Hanrui Wang 0002, Zixuan Jiang, Zhekai Zhang, Rongxing Tang, Chenghao Feng, Song Han 0003, Ray T. Chen, David Z. Pan |
HPCA | 5 |
| 2024 | Interactive learning for multi-finger dexterous hand: A model-free hierarchical deep reinforcement learning approach
Baojiang Li, Shengjie Qiu, Jibo Bai, Bin Wang 0072, Zhekai Zhang, Haiyan Wang 0013, Xichao Wang |
Knowl. Based Syst. | 5 |
| 2021 | SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head PruningabstractThe attention mechanism is becoming increasingly popular in Natural Language Processing (NLP) applications, showing superior performance than convolutional and recurrent architectures. However, general-purpose platforms such as CPUs and GPUs are inefficient when performing attention inference due to complicated data movement and low arithmetic intensity. Moreover, existing NN accelerators mainly focus on optimizing convolutional or recurrent models, and cannot efficiently support attention. In this paper, we present SpAtten, an efficient algorithm-architecture co-design that leverages token sparsity, head sparsity, and quantization opportunities to reduce the attention computation and memory access. Inspired by the high redundancy of human languages, we propose the novel cascade token pruning to prune away unimportant tokens in the sentence. We also propose cascade head pruning to remove unessential heads. Cascade pruning is fundamentally different from weight pruning since there is no trainable weight in the attention mechanism, and the pruned tokens and heads are selected on the fly. To efficiently support them on hardware, we design a novel top-k engine to rank token and head importance scores with high throughput. Furthermore, we propose progressive quantization that first fetches MSBs only and performs the computation; if the confidence is low, it fetches LSBs and recomputes the attention outputs, trading computation for memory reduction.Extensive experiments on 30 benchmarks show that, on average, SpAtten reduces DRAM access by 10.0× with no accuracy loss, and achieves 1.6×, 3.0×, 162×, 347× speedup, and 1.4×, 3.2×, 1193×, 4059× energy savings over A3accelerator, MNNFast accelerator, TITAN Xp GPU, Xeon CPU, respectively. Hanrui Wang 0002, Zhekai Zhang, Song Han 0003 |
HPCA | 2 |
| 2021 | PointAcc: Efficient Point Cloud AcceleratorabstractDeep learning on point clouds plays a vital role in a wide range of applications such as autonomous driving and AR/VR. These applications interact with people in real time on edge devices and thus require low latency and low energy. Compared to projecting the point cloud to 2D space, directly processing 3D point cloud yields higher accuracy and lower #MACs. However, the extremely sparse nature of point cloud poses challenges to hardware acceleration. For example, we need to explicitly determine the nonzero outputs and search for the nonzero neighbors (mapping operation), which is unsupported in existing accelerators. Furthermore, explicit gather and scatter of sparse features are required, resulting in large data movement overhead. Yujun Lin 0001, Zhekai Zhang, Haotian Tang, Hanrui Wang 0002, Song Han 0003 |
MICRO | 2 |
| 2020 | SpArch: Efficient Architecture for Sparse Matrix MultiplicationabstractGeneralized Sparse Matrix-Matrix Multiplication (SpGEMM) is a ubiquitous task in various engineering and scientific applications. However, inner product based SpGEMM introduces redundant input fetches for mismatched nonzero operands, while outer product based approach suffers from poor output locality due to numerous partial product matrices. Inefficiency in the reuse of either inputs or outputs data leads to extensive and expensive DRAM access. To address this problem, this paper proposes an efficient sparse matrix multiplication accelerator architecture, SpArch, which jointly optimizes the data locality for both input and output matrices. We first design a highly parallelized streaming-based merger to pipeline the multiply and merge stage of partial matrices so that partial matrices are merged on chip immediately after produced. We then propose a condensed matrix representation that reduces the number of partial matrices by three orders of magnitude and thus reduces DRAM access by 5.4x. We further develop a Huffman tree scheduler to improve the scalability of the merger for larger sparse matrices, which reduces the DRAM access by another 1.8x. We also resolve the increased input matrix read induced by the new representation using a row prefetcher with near-optimal buffer replacement policy, further reducing the DRAM access by 1.5x. Evaluated on 20 benchmarks, SpArch reduces the total DRAM access by 2.8x over previous state-of-the-art. On average, SpArch achieves 4x, 19x, 18x, 17x, 1285x speedup and 6x, 164x, 435x, 307x, 62x energy savings over OuterSpace, MKL, cuSPARSE, CUSP, and ARM Armadillo, respectively. Zhekai Zhang, Hanrui Wang 0002, Song Han 0003, William J. Dally |
HPCA | 1 |
| 2020 | Once-for-All: Train One Network and Specialize it for Efficient Deployment
Han Cai, Chuang Gan 0001, Tianzhe Wang, Zhekai Zhang, Song Han 0003 |
ICLR | 4 |