EDBT 2026 Demo / reviewers in the wild / expert
Shang Yang
dblp:79/9960
· DBLP profile ↗
29ranked-venue papers
1as first author
26since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 1 first-author · 15 since 2021Systems, architecture and hardware · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive DrafterabstractThe emergence of Large Language Models (LLMs) with strong reasoning capabilities marks a significant milestone, unlocking new frontiers in complex problem-solving. However, training these reasoning models, typically using Reinforcement Learning (RL), encounters critical efficiency bottlenecks: response generation during RL training exhibits a persistent long-tail distribution, where a few very long responses dominate execution time, wasting resources and inflating costs. To address this, we propose TLT, a system that accelerates reasoning RL training losslessly by integrating adaptive speculative decoding. Applying speculative decoding in RL is challenging due to the dynamic workloads, evolving target model, and draft model training overhead. TLT overcomes these obstacles with two synergistic components: (1) Adaptive Drafter, a lightweight draft model trained continuously on idle GPUs during long-tail generation to maintain alignment with the target model at no extra cost; and (2) Adaptive Rollout Engine, which maintains a memory-efficient pool of pre-captured CUDAGraphs and adaptively select suitable SD strategies for each input batch. Evaluations demonstrate that TLT achieves over 1.7x end-to-end RL training speedup over state-of-the-art systems, preserves the model accuracy, and yields a high-quality draft model as a free byproduct suitable for efficient deployment. Code is released at https://github.com/mit-han-lab/fastrl. Qinghao Hu 0004, Shang Yang, Junxian Guo, Xiaozhe Yao, Yujun Lin 0001, Yuxian Gu, Han Cai, Chuang Gan 0001, Ana Klimovic, Song Han 0001 |
ASPLOS (2) | 2 |
| 2026 | 2PADMS: Two-stage prediction and data migration strategy based on hard disk failure time
Huiyuan Qiang, Yuequan Li, Hongzhang Yang, Yaofeng Tu, Shang Yang |
Expert Syst. Appl. | 6 |
| 2026 | GFPP: A confidence-aware file system prefetching method based on deep graph networks
Hongzhang Yang, Shang Yang |
Expert Syst. Appl. | 4 |
| 2025 | NVILA: Efficient Frontier Visual Language ModelsabstractVisual language models (VLMs) have made significant advances in accuracy in recent years. However, their efficiency has received much less attention. This paper introduces NVILA, a family of open VLMs designed to optimize both efficiency and accuracy. Building on top of VILA, we improve its model architecture by first scaling up the spatial and temporal resolutions, and then compressing visual tokens. This "scale-then-compress" approach enables NVILA to efficiently process high-resolution images and long videos. We also conduct a systematic investigation to enhance the efficiency of NVILA throughout its entire lifecycle, from training to deployment. NVILA matches or surpasses the accuracy of many leading open and proprietary VLMs across a wide range of image and video benchmarks. At the same time, it reduces training costs by 1.9-5.1×, prefilling latency by 1.6-2.2×, and decoding latency by 1.2-2.8×. Ligeng Zhu, Baifeng Shi, Zhuoyang Zhang, Yuming Lou, Shang Yang, Haocheng Xi, Shiyi Cao, Yuxian Gu, Dacheng Li, Xiuyu Li, Haotian Tang, Yunhao Fang, Yukang Chen, Cheng-Yu Hsieh, De-An Huang, An-Chieh Cheng, Jinyi Hu, Sifei Liu, Ranjay Krishna, Pavlo Molchanov 0001, Jan Kautz, Hongxu Yin, Song Han 0003, Yao Lu 0006 |
CVPR | 6 |
| 2025 | SparseVILA: Decoupling Visual Sparsity for Efficient VLM InferenceabstractVision Language Models (VLMs) have rapidly advanced in integrating visual and textual reasoning, powering applications across high-resolution image understanding, long-video analysis, and multi-turn conversation. However, their scalability remains limited by the growing number of visual tokens that dominate inference latency. We present SparseVILA, a new paradigm for efficient VLM inference that decouples visual sparsity across the prefilling and decoding stages. SparseVILA distributes sparsity across stages by pruning redundant visual tokens during prefill and retrieving only query-relevant tokens during decoding. This decoupled design matches leading prefill pruning methods while preserving multi-turn fidelity by retaining most of the visual cache so that query-aware tokens can be retrieved at each conversation round. Built on an AWQ-optimized inference pipeline, SparseVILA achieves up to 4.0 times faster prefilling, 2.5 times faster decoding, and an overall 2.6 times end-to-end speedup on long-context video tasks -- while improving accuracy on document-understanding and reasoning tasks. By decoupling query-agnostic pruning and query-aware retrieval, SparseVILA establishes a new direction for efficient multimodal inference, offering a training-free, architecture-agnostic framework for accelerating large VLMs without sacrificing capability. Samir Khaki, Junxian Guo, Shang Yang, Yukang Chen, Konstantinos N. Plataniotis, Yao Lu 0006, Song Han 0003 |
ICCV | 4 |
| 2025 | Deep Compression Autoencoder for Efficient High-Resolution Diffusion ModelsabstractWe present Deep Compression Autoencoder (DC-AE), a new family of autoencoders for accelerating high-resolution diffusion models. Existing autoencodes have demonstrated impressive results at a moderate spatial compression ratio (e.g., 8x), but fail to maintain satisfactory reconstruction accuracy for high spatial compression ratios (e.g., 64x). We address this challenge by introducing two key techniques: (1) Residual Autoencoding, where we design our models to learn residuals based on the space-to-channel transformed features to alleviate the optimization difficulty of high spatial-compression autoencoders; (2) Decoupled High-Resolution Adaptation, an efficient decoupled three-phase training strategy for mitigating the generalization penalty of high spatial-compression autoencoders. With these designs, we improve the autoencoder's spatial compression ratio up to 128 while maintaining the reconstruction quality. Applying our DC-AE to latent diffusion models, we achieve significant speedup without accuracy drop. For example, on ImageNet 512x512, our DC-AE provides 19.1x inference speedup and 17.9x training speedup on H100 GPU for UViT-H while achieving a better FID, compared with the widely used SD-VAE-f8 autoencoder. Junyu Chen 0003, Han Cai, Junsong Chen, Enze Xie, Shang Yang, Haotian Tang, Song Han 0003 |
ICLR | 5 |
| 2025 | LongVILA: Scaling Long-Context Visual Language Models for Long VideosabstractLong-context capability is critical for multi-modal foundation models, especially for long video understanding. We introduce LongVILA, a full-stack solution for long-context visual-language models by co-designing the algorithm and system. For model training, we upgrade existing VLMs to support long video understanding by incorporating two additional stages, i.e., long context extension and long video supervised fine-tuning. However, training on long video is computationally and memory intensive. We introduce the long-context Multi-Modal Sequence Parallelism (MM-SP) system that efficiently parallelizes long video training and inference, enabling 2M context length training on 256 GPUs without any gradient checkpointing. LongVILA efficiently extends the number of video frames of VILA from 8 to 2048, achieving 99.8% accuracy in 6,000-frame (more than 1 million tokens) video needle-in-a-haystack. LongVILA-7B demonstrates strong accuracy on 9 popular video benchmarks, e.g., 65.1% VideoMME with subtitle. Besides, MM-SP is 2.1x - 5.7x faster than ring style sequence parallelism and 1.1x - 1.4x faster than Megatron with a hybrid context and tensor parallelism. Moreover, it seamlessly integrates with Hugging Face Transformers. Yukang Chen, Fuzhao Xue, Dacheng Li, Qinghao Hu 0004, Ligeng Zhu, Xiuyu Li, Yunhao Fang, Haotian Tang, Shang Yang, Yihui He, Hongxu Yin, Pavlo Molchanov 0001, Jan Kautz, Linxi Fan, Yuke Zhu, Yao Lu 0006, Song Han 0003 |
ICLR | 9 |
| 2025 | HART: Efficient Visual Generation with Hybrid Autoregressive TransformerabstractWe introduce Hybrid Autoregressive Transformer (HART), the first autoregressive (AR) visual generation model capable of directly generating 1024x1024 images, rivaling diffusion models in image generation quality. Existing AR models face limitations due to the poor image reconstruction quality of their discrete tokenizers and the prohibitive training costs associated with generating 1024px images. To address these challenges, we present the hybrid tokenizer, which decomposes the continuous latents from the autoencoder into two components: discrete tokens representing the big picture and continuous tokens representing the residual components that cannot be represented by the discrete tokens. The discrete component is modeled by a scalable-resolution discrete AR model, while the continuous component is learned with a lightweight residual diffusion module with only 37M parameters. Compared with the discrete-only VAR tokenizer, our hybrid approach improves reconstruction FID from 2.11 to 0.30 on MJHQ-30K, leading to a 31% generation FID improvement from 7.85 to 5.38. HART also outperforms state-of-the-art diffusion models in both FID and CLIP score, with 4.5-7.7$\times$ higher throughput and 6.9-13.4$\times$ lower MACs. Our code is open sourced at https://github.com/mit-han-lab/hart. Haotian Tang, Yecheng Wu, Shang Yang, Enze Xie, Junsong Chen, Junyu Chen 0003, Zhuoyang Zhang, Han Cai, Yao Lu 0006, Song Han 0003 |
ICLR | 3 |
| 2025 | DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming HeadsabstractDeploying long-context large language models (LLMs) is essential but poses significant computational and memory challenges.
Caching all Key and Value (KV) states across all attention heads consumes substantial memory.
Existing KV cache pruning methods either damage the long-context capabilities of LLMs or offer only limited efficiency improvements.
In this paper, we identify that only a fraction of attention heads, a.k.a, Retrieval Heads, are critical for processing long contexts and require full attention across all tokens.
In contrast, all other heads, which primarily focus on recent tokens and attention sinks—referred to as Streaming Heads—do not require full attention.
Based on this insight, we introduce DuoAttention, a framework that only applies a full KV cache to retrieval heads while using a light-weight, constant-length KV cache for streaming heads, which reduces both LLM's decoding and pre-filling memory and latency without compromising its long-context abilities.
DuoAttention uses a lightweight, optimization-based algorithm with synthetic data to identify retrieval heads accurately.
Our method significantly reduces long-context inference memory by up to 2.55$\times$ for MHA and 1.67$\times$ for GQA models while speeding up decoding by up to 2.18$\times$ and 1.50$\times$ and accelerating pre-filling by up to 1.73$\times$ and 1.63$\times$ for MHA and GQA models, respectively, with minimal accuracy loss compared to full attention.
Notably, combined with quantization, DuoAttention enables Llama-3-8B decoding with 3.33 million context length measured on a single A100 GPU. Code is provided in https://github.com/mit-han-lab/duo-attention. Guangxuan Xiao, Jingwei Zuo, Junxian Guo, Shang Yang, Haotian Tang, Song Han 0003 |
ICLR | 5 |
| 2025 | Jet-Nemotron: Efficient Language Model with Post Neural Architecture SearchabstractWe present Jet-Nemotron, a new family of hybrid-architecture language models, which matches or exceeds the accuracy of leading full-attention models while significantly improving generation throughput. Jet-Nemotron is developed using Post Neural Architecture Search (PostNAS), a novel neural architecture exploration pipeline that enables efficient model design. Unlike prior approaches, PostNAS begins with a pre-trained full-attention model and freezes its MLP weights, allowing efficient exploration of attention block designs. The pipeline includes four key components: (1) learning optimal full-attention layer placement and elimination, (2) linear attention block selection, (3) designing new attention blocks, and (4) performing hardware-aware hyperparameter search. Our Jet-Nemotron-2B model achieves comparable or superior accuracy to Qwen3, Qwen2.5, Gemma3, and Llama3.2 across a comprehensive suite of benchmarks while delivering up to 53.6× generation throughput speedup and 6.1× prefilling speedup. It also achieves higher accuracy on MMLU and MMLU-Pro than recent advanced MoE full-attention models, such as DeepSeek-V3-Small and Moonlight, despite their larger scale with 15B total and 2.2B activated parameters. Yuxian Gu, Qinghao Hu 0004, Haocheng Xi, Junyu Chen 0003, Shang Yang, Song Han 0003, Han Cai |
NeurIPS | 5 |
| 2025 | Accelerating complex graph queries by summary-based hybrid partitioning for discovering vulnerabilities of distribution equipment
Shang Yang, Yinglong Ma 0001 |
Future Gener. Comput. Syst. | 3 |
| 2025 | DERAID: A Decryption and Encryption Integrated Redundant Array of Independent Disks TechnologyabstractIn the era of big data, ensuring both data confidentiality and reliability has become a critical concern for data owners. While encryption algorithms and erasure coding techniques can independently guarantee confidentiality and reliability, respectively, traditional methods that apply encryption before or after erasure coding often lead to significant performance degradation and increased storage overhead. To address these limitations, this paper proposes DERAID, a novel RAID-based technology that integrates encryption directly into the coding process. DERAID is designed to improve the efficiency of encoding, decoding, updating and reconstruction operations, while simultaneously ensuring data confidentiality and reliability and minimizing the data expansion rate. The scheme employs a unified coding structure based on the ECB encryption mode, where conventional parity blocks are replaced by a combination of cryptographic keys and verification blocks. Experimental results demonstrate that, compared to traditional methods and QS-code, DERAID improves encoding performance by 44.7–95.3%, updating performance by 48.2–99.6% and reconstruction performance by 63.0–99.9%, while reducing the data expansion rate by 3.3–36.9%. These results indicate that DERAID achieves a practical balance between security, reliability and performance. Hongzhang Yang, Sen Yuan, Ping Wang 0003, Shang Yang |
Int. J. Softw. Eng. Knowl. Eng. | 5 |
| 2025 | BAQoS: A Burst I/O Aware Quality of Service Optimization for Cloud Storage ServiceabstractIn cloud storage services, burst I/O workloads from data analytics and artificial intelligence/machine learning (AI/ML) applications present significant challenges to Quality of Service (QoS) management. Existing scheduling models like dmClock ensure fair and stable I/O bandwidth allocation in typical scenarios. However, they falter under frequent burst traffic, leading to lower resource utilization and higher task latency. To address this, we propose BAQoS, a Burst I/O Aware Quality of Service Optimization for Cloud Storage Service. BAQoS employs refined request classification, a burst-aware hierarchical scheduling algorithm, and a high-performance scheduler architecture (HPSA). These features enable dynamic resource allocation and efficient scheduling for both burst and regular requests. Experiments show that BAQoS markedly enhances performance under burst workloads, accelerating burst request processing by up to 7.09 and improving overall system performance by 48.86%. Furthermore, BAQoS ensures superior performance for non-burst users, achieving a 51.15% performance boost for high-priority users, a 5.34% increase in system throughput, and an over 50% reduction in IOPS standard deviation among same-priority users. Jingzhe Zhao, Hongzhang Yang, Guangping Xu, Ping Wang 0003, Shang Yang |
Int. J. Softw. Eng. Knowl. Eng. | 6 |
| 2024 | Special Session: Estimation and Optimization of DNNs for Embedded PlatformsabstractSeveral state of the art estimation and optimization techniques for CNNs and LLMs on embedded devices are summarized. For LLMs an Activation-aware Weight Quantization and on-the-fly dequantization techniques is presented. For CNNs various pruning algorithms and an integrated optimization and implementation flow is discussed. To estimate inference latency of CNNs on specific hardware platforms, three different techniques are reviewed: A mixed analytic-stochastic model, an analytic model based on step-wise linear functions, and a method that uses a detailed architecture description of the hardware. Axel Jantsch, Song Han 0003, Lin Meng 0001, Oliver Bringmann 0001, Haotian Tang, Shang Yang, Matthias Wess, Martin Lechner |
CODES+ISSS | 6 |
| 2024 | Sparse Refinement for Efficient High-Resolution Semantic Segmentation
Zhuoyang Zhang, Samir Khaki, Shang Yang, Haotian Tang, Chenfeng Xu, Kurt Keutzer, Song Han 0003 |
ECCV (67) | 4 |
| 2024 | Reliability allocation method based on minimizing implementation risk
Axita, Chuanhai Chen, Jinyan Guo, Shang Yang |
Expert Syst. Appl. | 5 |
| 2023 | FlatFormer: Flattened Window Attention for Efficient Point Cloud TransformerabstractTransformer, as an alternative to CNN, has been proven effective in many modalities (e.g., texts and images). For 3D point cloud transformers, existing efforts focus primarily on pushing their accuracy to the state-of-the-art level. However, their latency lags behind sparse convolution-based models (3 × slower), hindering their usage in resource-constrained, latency-sensitive applications (such as autonomous driving). This inefficiency comes from point clouds' sparse and irregular nature, whereas transformers are designed for dense, regular workloads. This paper presents FlatFormer to close this latency gap by trading spatial proximity for better computational regularity. We first flatten the point cloud with window-based sorting and partition points into groups of equal sizes rather than windows of equal shapes. This effectively avoids expensive structuring and padding overheads. We then apply self-attention within groups to extract local features, alternate sorting axis to gather features from different directions, and shift windows to exchange features across groups. FlatFormer delivers state-of-the-art accuracy on Waymo Open Dataset with 4.6× speedup over (transformer-based) SST and 1.4× speedup over (sparse convolutional) CenterPoint. This is the first point cloud transformer that achieves real-time performance on edge GPUs and is faster than sparse convolutional methods while achieving on-par or even superior accuracy on large-scale benchmarks. Xinyu Yang 0002, Haotian Tang, Shang Yang, Song Han 0003 |
CVPR | 4 |
| 2023 | CLAP: Locality Aware and Parallel Triangle Counting with Content Addressable MemoryabstractTriangle counting (TC) is one of the most fundamental graph analysis tools with a wide range of applications. Modern triangle counting algorithms traverse the graph and perform set intersections of neighbor sets to find triangles. However, existing triangle counting approaches suffer from the heavy off-chip memory access and set intersection overhead. Thus, we propose CLAP, the first content addressable memory (CAM) based triangle counting architecture with the software and hardware co-optimizations. To reduce off-chip memory access and the number of set intersections, we propose the first force-based node index reorder method. It simultaneously optimizes both data locality and the computation amount. Compared with random node indices, the reorder method reduces the off-chip memory access and the set intersections by 61% and 64%, respectively, while providing$\mathbf{2.19}\times$end-to-end speedup. To improve the set intersection parallelism, we propose the first CAM-based triangle counting architecture under chip area constraints. We enable the high parallel set intersection by translating it into content search on CAM with full parallelism. Thus, the time complexity of the set intersection reduces from$O(m+n)$or$O(n\log m)$to$O(n)$. Extensive experiments on real-world graphs show that CLAP achieves$\mathbf{39}\times, \mathbf{27}\times$, and$\mathbf{78}\times$speedup over state-of-the-art CPU, GPU, and processing-in-memory baselines, respectively. The software code is available at: https://github.com/thu-nics/CLAP-triangle-counting Tianyu Fu 0004, Chiyue Wei, Zhenhua Zhu 0002, Shang Yang, Zhongming Yu, Guohao Dai 0001, Huazhong Yang, Yu Wang 0002 |
DATE | 4 |
| 2023 | K-Fold Cross-Valuation for Machine Learning Using Shapley Value
Qiangqiang He, Mujie Zhang, Jie Zhang 0152, Shang Yang, Chong-Jun Wang |
ICANN (3) | 4 |
| 2023 | Hierarchical Vision and Language Transformer for Efficient Visual Dialog
Qiangqiang He, Mujie Zhang, Jie Zhang 0152, Shang Yang, Chong-Jun Wang |
ICANN (6) | 4 |
| 2023 | TorchSparse++: Efficient Training and Inference Framework for Sparse Convolution on GPUsabstractSparse convolution plays a pivotal role in emerging workloads, including point cloud processing in AR/VR, autonomous driving, and graph understanding in recommendation systems. Since the computation pattern is sparse and irregular, specialized high-performance kernels are required. Existing GPU libraries offer two dataflow types for sparse convolution. The gather-GEMM-scatter dataflow is easy to implement but not optimal in performance, while the dataflows with overlapped computation and memory access (e.g. implicit GEMM) are highly performant but have very high engineering costs. In this paper, we introduce TorchSparse++, a new GPU library that achieves the best of both worlds. We create a highly efficient Sparse Kernel Generator that generates performant sparse convolution kernels at less than one-tenth of the engineering cost of the current state-of-the-art system. On top of this, we design the Sparse Autotuner, which extends the design space of existing sparse convolution libraries and searches for the best dataflow configurations for training and inference workloads. Consequently, TorchSparse++ achieves 2.9 × , 3.3 × , 2.2 × and 1.7 × measured end-to-end speedup on an NVIDIA A100 GPU over state-of-the-art MinkowskiEngine, SpConv 1.2, TorchSparse and SpConv v2 in inference; and is 1.2-1.3 × faster than SpConv v2 in mixed precision training across seven representative autonomous driving benchmarks. It also seamlessly supports graph convolutions, achieving 2.6-7.6 × faster inference speed compared with state-of-the-art graph deep learning libraries. Our code is publicly released at https://github.com/mit-han-lab/torchsparse. Haotian Tang, Shang Yang, Ke Hong, Zhongming Yu, Xiuyu Li, Guohao Dai 0001, Yu Wang 0002, Song Han 0003 |
MICRO | 2 |
| 2022 | Heuristic adaptability to input dynamics for SpMM on CPUsabstractSparse Matrix-Matrix Multiplication (SpMM) has served as fundamental components in various domains. Many previous studies exploit GPUs for SpMM acceleration because GPUs provide high bandwidth and parallelism. We point out that a static design does not always improve the performance of SpMM on different input data (e.g., >85% performance loss with a single algorithm). In this paper, we consider the challenge of input dynamics from a novel auto-tuning perspective, while following issues remain to be solved: (1) Orthogonal design principles considering sparsity. Orthogonal design principles for such a sparse problem should be extracted to form different algorithms, and further used for performance tuning. (2) Nontrivial implementations in the algorithm space. Combining orthogonal design principles to create new algorithms needs to tackle with new challenges like thread race handling. (3) Heuristic adaptability to input dynamics. The heuristic adaptability is required to dynamically optimize code for input dynamics. Guohao Dai 0001, Guyue Huang, Shang Yang, Zhongming Yu, Yufei Ding 0001, Yuan Xie 0001, Huazhong Yang, Yu Wang 0002 |
DAC | 3 |
| 2021 | Robust and Efficient Mechanism Design for Heterogeneous Task Crowdsensing
Qiangqiang He, Yu Qiao 0006, Shang Yang, Chong-Jun Wang |
WASA (3) | 3 |
| 2021 | Equitable Valuation of Crowdsensing for Machine Learning via Game Theory
Qiangqiang He, Yu Qiao 0006, Shang Yang, Chong-Jun Wang |
WASA (3) | 3 |
| 2021 | Knowledge-based integrated product design framework towards sustainable low-carbon manufacturing
Shang Yang, Shanhe Lou, Yicong Gao, Yixiong Feng |
Adv. Eng. Informatics | 2 |
| 2021 | Distributed aggregation-based attributed graph summarization for summary-based approximate attributed graph queries
Shang Yang, Xiaona Chen, Jingpeng Zhao, Yinglong Ma 0001 |
Expert Syst. Appl. | 1 |
| 2014 | Convex transversals
Esther M. Arkin, Claudia Dieckmann, Christian Knauer, Joseph S. B. Mitchell, Valentin Polishchuk, Lena Schlipf, Shang Yang |
Comput. Geom. | 7 |
| 2012 | Routing multi-class traffic flows in the plane
Joondong Kim, Joseph S. B. Mitchell, Valentin Polishchuk, Shang Yang, Jingyu Zou |
Comput. Geom. | 4 |
| 2011 | Convex Transversals
Esther M. Arkin, Claudia Dieckmann, Christian Knauer, Joseph S. B. Mitchell, Valentin Polishchuk, Lena Schlipf, Shang Yang |
WADS | 7 |