EDBT 2026 Demo / reviewers in the wild / expert
Hangyu Zheng
dblp:352/6382
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2026
0009-0004-5972-2929ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Efficient and distributed learning · 60% Deep learning architectures and training · 24% Optimization for machine learning · 16% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Memory systems · 30% GPUs and heterogeneous computing · 23% Hardware accelerators and domain-specific architectures · 23% | |
| Databases, data mining, and information retrieval
2 papers |
Graph data management · 36% Query processing and optimization · 32% Indexing and storage engines · 32% |
Topics — the 14 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
model compression |
1.6 | 2 | 2025 | Sparse Learning for State Space Models on Mobile · ICLR 2025 Exploring Token Pruning in Vision State Space Models · NeurIPS 2024 |
Machine learning › Deep learning architectures and training
state space model |
1.1 | 2 | 2025 | Sparse Learning for State Space Models on Mobile · ICLR 2025 Exploring Token Pruning in Vision State Space Models · NeurIPS 2024 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
1.0 | 1 | 2026 | FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy Optimizations · ASPLOS (2) 2026 |
Memory systems › memory hierarchy
memory hierarchy optimization |
1.0 | 1 | 2026 | FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy Optimizations · ASPLOS (2) 2026 |
GPUs and heterogeneous computing › embedded GPU
mobile GPU |
1.0 | 1 | 2026 | FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy Optimizations · ASPLOS (2) 2026 |
Embedded and real-time systems › on-device inference
mobile inference |
1.0 | 1 | 2026 | FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy Optimizations · ASPLOS (2) 2026 |
Machine learning › Optimization for machine learning
sparse learning |
0.9 | 1 | 2025 | Sparse Learning for State Space Models on Mobile · ICLR 2025 |
Machine learning › Efficient and distributed learning › model compression › pruning
structured pruning |
0.9 | 1 | 2025 | Sparse Learning for State Space Models on Mobile · ICLR 2025 |
Machine learning › Efficient and distributed learning › model compression
token pruning |
0.8 | 1 | 2024 | Exploring Token Pruning in Vision State Space Models · NeurIPS 2024 |
Graph data management
graph query processing |
0.8 | 1 | 2024 | Vertex Encoding for Edge Nonexistence Determination With SIMD Acceleration · IEEE Trans. Knowl. Data Eng. 2024 |
Query processing and optimization › query optimization
graph query optimization |
0.7 | 1 | 2023 | VEND: Vertex Encoding for Edge Nonexistence Determination · ICDE 2023 |
Memory systems › memory access optimization
memory streaming |
0.3 | 1 | 2026 | FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy Optimizations · ASPLOS (2) 2026 |
Compilers and program optimization
compiler optimization |
0.3 | 1 | 2025 | Sparse Learning for State Space Models on Mobile · ICLR 2025 |
Machine learning › Deep learning architectures and training › state space model
vision state space model |
0.2 | 1 | 2024 | Exploring Token Pruning in Vision State Space Models · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
pruning · 1.7kernel sparsity · 1.7compiler optimization · 1.7static scheduling · 1.02.5d texture memory · 1.0token importance evaluation · 0.8pruning-aware hidden state alignment · 0.8compression · 0.8SIMD · 0.8vertex encoding · 0.7maintenance algorithms · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy OptimizationsabstractThe increasing size and complexity of modern deep neural networks (DNNs) pose significant challenges for on-device inference on mobile GPUs, with limited memory and computational resources. Existing DNN acceleration frameworks primarily deploy a weight preloading strategy, where all model parameters are loaded into memory before execution on mobile GPUs. We posit that this approach is not adequate for modern DNN workloads that comprise very large model(s) and possibly execution of several distinct models in succession. In this work, we introduce FlashMem, a memory streaming framework designed to efficiently execute large-scale modern DNNs and multi-DNN workloads while minimizing memory consumption and reducing inference latency. Instead of fully preloading weights, FlashMem statically determines model loading schedules and dynamically streams them on demand, leveraging 2.5D texture memory to minimize data transformations and improve execution efficiency. Experimental results on 11 models demonstrate that FlashMem achieves 2.0× to 8.4× memory reduction and 1.7× to 75.0× speedup compared to existing frameworks, enabling efficient execution of large-scale models and multi-DNN support on resource-constrained mobile GPUs. Zhihao Shu, Md. Musfiqur Rahman Sanim, Hangyu Zheng, Kunxiong Zhu, Miao Yin, Gagan Agrawal, Wei Niu 0002 |
ASPLOS (2) | 3 |
| 2025 | Sparse Learning for State Space Models on MobileabstractTransformer models have been widely investigated in different domains by providing long-range dependency handling and global contextual awareness, driving the development of popular AI applications such as ChatGPT, Gemini, and Alexa.
State Space Models (SSMs) have emerged as strong contenders in the field of sequential modeling, challenging the dominance of Transformers. SSMs incorporate a selective mechanism that allows for dynamic parameter adjustment based on input data, enhancing their performance.
However, this mechanism also comes with increasing computational complexity and bandwidth demands, posing challenges for deployment on resource-constraint mobile devices.
To address these challenges without sacrificing the accuracy of the selective mechanism, we propose a sparse learning framework that integrates architecture-aware compiler optimizations. We introduce an end-to-end solution--$\mathbf{C}_4^n$ kernel sparsity, which prunes $n$ elements from every four contiguous weights, and develop a compiler-based acceleration solution to ensure execution efficiency for this sparsity on mobile devices.
Based on the kernel sparsity, our framework generates optimized sparse models targeting specific sparsity or latency requirements for various model sizes. We further leverage pruned weights to compensate for the remaining weights, enhancing downstream task performance.
For practical hardware acceleration, we propose $\mathbf{C}_4^n$-specific optimizations combined with a layout transformation elimination strategy.
This approach mitigates inefficiencies arising from fine-grained pruning in linear layers and improves performance across other operations.
Experimental results demonstrate that our method achieves superior task performance compared to other semi-structured pruning methods and achieves up-to 7$\times$ speedup compared to llama.cpp framework on mobile devices. Xuan Shen, Hangyu Zheng, Yifan Gong 0004, Zhenglun Kong, Changdi Yang, Zheng Zhan 0001, Yushu Wu, Xue Lin 0001, Yanzhi Wang 0001, Pu Zhao 0001, Wei Niu 0002 |
ICLR | 2 |
| 2024 | Exploring Token Pruning in Vision State Space ModelsabstractState Space Models (SSMs) have the advantage of keeping linear computational complexity compared to attention modules in transformers, and have been applied to vision tasks as a new type of powerful vision foundation model. Inspired by the observations that the final prediction in vision transformers (ViTs) is only based on a subset of most informative tokens, we take the novel step of enhancing the efficiency of SSM-based vision models through token-based pruning. However, direct applications of existing token pruning techniques designed for ViTs fail to deliver good performance, even with extensive fine-tuning. To address this issue, we revisit the unique computational characteristics of SSMs and discover that naive application disrupts the sequential token positions. This insight motivates us to design a novel and general token pruning method specifically for SSM-based vision models. We first introduce a pruning-aware hidden state alignment method to stabilize the neighborhood of remaining tokens for performance enhancement. Besides, based on our detailed analysis, we propose a token importance evaluation method adapted for SSM models, to guide the token pruning. With efficient implementation and practical acceleration methods, our method brings actual speedup. Extensive experiments demonstrate that our approach can achieve significant computation reduction with minimal impact on performance across different tasks. Notably, we achieve 81.7\% accuracy on ImageNet with a 41.6\% reduction in the FLOPs for pruned PlainMamba-L3. Furthermore, our work provides deeper insights into understanding the behavior of SSM-based vision models for future research. Zheng Zhan 0001, Zhenglun Kong, Yifan Gong 0004, Yushu Wu, Zichong Meng, Hangyu Zheng, Xuan Shen, Stratis Ioannidis, Wei Niu 0002, Pu Zhao 0001, Yanzhi Wang 0001 |
NeurIPS | 6 |
| 2024 | Vertex Encoding for Edge Nonexistence Determination With SIMD AccelerationabstractWe propose to design vertex encoding for determinations of no-result edge queries that should not be executed. Edge query is one of the core operations in mainstream graph databases, which is to retrieve edges connecting two given vertices. Real-world graphs may be too large to be stored in memory and frequently accessing edge data on disk usually incurs much overhead. The average degree of real-world graph tends to be much less than the vertex number, and edges may not exist in most pairs of vertices. Efficiently avoiding no-result edge query executions will certainly improve the performance of graph database. In this paper, we propose a new and important problem for determining no-result edge queries: vertex encoding for edge nonexistence determination (VEND, for short). We build a low dimensional vertex encoding for all vertices, and we can efficiently determine most vertex pairs that are connected by no edges just with their corresponding codes. The encoding can be efficiently adjusted when data updates happen. With VEND, we can utilize in-memory efficient operations to filter no-result disk accesses for edge query. We also design SIMD-oriented compression optimizations to further improve performance. Extensive experiments on real-world datasets confirm the effectiveness of our solution. Hangyu Zheng, Youhuan Li, Fang Xiong, Xiaosen Li, Lei Zou 0001, Peifan Shi, Zheng Qin 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2023 | VEND: Vertex Encoding for Edge Nonexistence DeterminationabstractWe propose to design vertex encoding for determinations of no-result edge queries that should not be executed. Edge query is one of the core operations in mainstream graph databases, which is to retrieve the corresponding edges connecting two given vertices. Real-world graphs may be too large to be stored in memory and frequently accessing edge data on disk usually incurs much overhead. Average degree of real-world graph tends to be much less than the vertex number, and edges may not exist in most pairs of vertices. Efficiently avoiding no-result edge query executions will certainly improve performance of graph database. In this paper, we propose a new and important problem for determining no-result edge queries: vertex encoding for edge nonexistence determination (VEND, for short). We build a low dimensional vertex encoding for all vertices, and we can efficiently determine most vertex pairs that are connected by no edges just with their corresponding codes. With VEND, we can utilize in-memory efficient operations to filter no-result disk accesses for edge query. We also design maintenance algorithms for the proposed solution when data updates happen. Extensive experiments on many real-world datasets confirm the ability of our solution on determining a quite high proportion of non-edge vertex pairs, as well as the acceleration for edge queries. Youhuan Li, Hangyu Zheng, Lei Zou 0001, Xiaosen Li, Ziming Li 0004, Pin Xiao, Yangyu Tao, Zheng Qin 0001 |
ICDE | 2 |