VLDB 2026 Research / reviewers in the wild / expert
Minghui Yu
dblp:160/2206
· DBLP profile ↗
5ranked-venue papers
1as first author
3since 2021 · last 2026
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Efficient and distributed learning · 40% Language models and text generation · 40% 3D vision · 10% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Hardware accelerators and domain-specific architectures · 69% Performance modeling and evaluation · 21% Cloud and datacenter computing · 10% |
Topics — the 10 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures › machine learning accelerator › transformer accelerator
LLM inference accelerator |
1.0 | 1 | 2026 | Characterizing Cloud-Native LLM Inference at Bytedance and Exposing Optimization Challenges and Opportunities for Future AI Accelerators · HPCA 2026 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
1.0 | 1 | 2026 | Characterizing Cloud-Native LLM Inference at Bytedance and Exposing Optimization Challenges and Opportunities for Future AI Accelerators · HPCA 2026 |
Natural language and speech › Language models and text generation › decoding
efficient decoding |
0.9 | 1 | 2025 | HShare: Fast LLM Decoding by Hierarchical Key-Value Sharing · ICLR 2025 |
Machine learning › Efficient and distributed learning
inference efficiency |
0.9 | 1 | 2025 | HShare: Fast LLM Decoding by Hierarchical Key-Value Sharing · ICLR 2025 |
Machine learning › Efficient and distributed learning
KV cache management |
0.9 | 1 | 2025 | HShare: Fast LLM Decoding by Hierarchical Key-Value Sharing · ICLR 2025 |
Natural language and speech › Language models and text generation › large language model inference
KV cache sharing |
0.9 | 1 | 2025 | HShare: Fast LLM Decoding by Hierarchical Key-Value Sharing · ICLR 2025 |
Computer vision › 3D vision
point cloud segmentation |
0.4 | 1 | 2020 | Self-Prediction for Joint Instance and Semantic Segmentation of Point Clouds · ECCV (22) 2020 |
Computer vision › Segmentation and scene understanding › image segmentation
semantic and instance segmentation |
0.4 | 1 | 2020 | Self-Prediction for Joint Instance and Semantic Segmentation of Point Clouds · ECCV (22) 2020 |
Performance modeling and evaluation › benchmarking › computer architecture benchmarking
accelerator benchmarking |
0.3 | 1 | 2026 | Characterizing Cloud-Native LLM Inference at Bytedance and Exposing Optimization Challenges and Opportunities for Future AI Accelerators · HPCA 2026 |
Performance modeling and evaluation
benchmarking |
0.3 | 1 | 2026 | Characterizing Cloud-Native LLM Inference at Bytedance and Exposing Optimization Challenges and Opportunities for Future AI Accelerators · HPCA 2026 |
Methods — techniques the papers use, named apart from their topics
workload characterization · 1.0design space analysis · 1.0hierarchical KV sharing · 0.9greedy algorithm · 0.9self-prediction · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Characterizing Cloud-Native LLM Inference at Bytedance and Exposing Optimization Challenges and Opportunities for Future AI AcceleratorsabstractAs a major provider of LLM inference services, ByteDance has continuously explored diverse accelerator options to meet the rapidly growing inference demands of various heterogeneous LLM scenarios with higher cost-effectiveness, thereby enabling LLMs to serve more people worldwide. However, during this process, we have found that the complexity and opacity of cloud scenarios and corresponding cloud accelerators make it difficult for academia and many innovative chip startups to fully understand the real demands and challenges of these scenarios, which in turn severely restricts innovation and application potential in this field. To bridge this gap, we first present and analyze the data and characteristics of the ByteDance Doubao LLM app across multiple dimensions, helping the community understand real-world cloud scenarios, and detail the challenges and opportunities we have identified. Second, we propose and plan to open-source our multi-level evaluation framework, XPU-Perf, which includes benchmarks spanning instructions, operators, and models. This framework improves interpretability and trustworthiness, and helps promising new accelerator architectures gain wider adoption and development. Finally, we present comparative results of four typical accelerators, summarize their shortcomings and challenges, conduct in-depth analysis, and highlight numerous architectural and scheduling innovation opportunities we have observed. Jingwei Cai, Dehao Kong, Hantao Huang, Zishan Jiang, Zixuan Ma, Qingyu Guo, Guiming Shi, Mingyu Gao 0001, Kaisheng Ma, Minghui Yu |
HPCA | 11 |
| 2025 | HShare: Fast LLM Decoding by Hierarchical Key-Value SharingabstractThe frequent retrieval of Key-Value (KV) cache data has emerged as a significant factor contributing to the inefficiency of the inference process in large language models. Previous research has demonstrated that a small subset of critical KV cache tokens largely influences attention outcomes, leading to methods that either employ fixed sparsity patterns or dynamically select critical tokens based on the query. While dynamic sparse patterns have proven to be more effective, they introduce significant computational overhead, as critical tokens must be reselected for each self-attention computation. In this paper, we reveal substantial similarities in KV cache token criticality across neighboring queries, layers, and heads. Motivated by this insight, we propose HShare, a hierarchical KV sharing framework. HShare facilitates the sharing of critical KV cache token indices across layers, heads, and queries, which significantly reduces the computational overhead associated with query-aware dynamic token sparsity. In addition, we introduce a greedy algorithm that dynamically determines the optimal layer-level and head-level sharing configuration for the decoding phase. We evaluate the effectiveness and efficiency of HShare across various tasks using three models: LLaMA2-7b, LLaMA3-70b, and Mistral-7b. Experimental results demonstrate that HShare achieves competitive accuracy with different sharing ratios, while delivering up to an $8.6\times$ speedup in self-attention operations and a $2.7\times$ improvement in end-to-end throughput compared with FlashAttention2 and GPT-fast respectively. The source code is publicly available at ~\url{https://github.com/wuhuaijin/HShare}. Huaijin Wu, Lianqiang Li, Hantao Huang, Tu Yi, Jihang Zhang, Minghui Yu, Junchi Yan |
ICLR | 6 |
| 2025 | SALS: Sparse Attention in Latent Space for KV Cache CompressionabstractLarge Language Models (LLMs) capable of handling extended contexts are in high demand, yet their inference remains challenging due to substantial Key-Value (KV) cache size and high memory bandwidth requirements. Previous research has demonstrated that KV cache exhibits low-rank characteristics within the hidden dimension, suggesting the potential for effective compression. However, due to the widely adopted Rotary Position Embedding (RoPE) mechanism in modern LLMs, naive low‑-rank compression suffers severe accuracy degradation or creates a new speed bottleneck, as the low-rank cache must first be reconstructed in order to apply RoPE. In this paper, we introduce two key insights: first, the application of RoPE to the key vectors increases their variance, which in turn results in a higher rank; second, after the key vectors are transformed into the latent space, they largely maintain their representation across most layers. Based on these insights, we propose the Sparse Attention in Latent Space (SALS) framework. SALS projects the KV cache into a compact latent space via low-rank projection, and performs sparse token selection using RoPE-free query--key interactions in this space. By reconstructing only a small subset of important tokens, it avoids the overhead of full KV cache reconstruction. We comprehensively evaluate SALS on various tasks using two large-scale models: LLaMA2-7b-chat and Mistral-7b, and additionally verify its scalability on the RULER-128k benchmark with LLaMA3.1-8B-Instruct. Experimental results demonstrate that SALS achieves SOTA performance by maintaining competitive accuracy. Under different settings, SALS achieves 6.4-fold KV cache compression and 5.7-fold speed-up in the attention operator compared to FlashAttention2 on the 4K sequence. For the end-to-end throughput performance, we achieves 1.4-fold and 4.5-fold improvement compared to GPT-fast on 4k and 32K sequences, respectively. The source code will be publicly available in the future. Junlin Mu, Hantao Huang, Jihang Zhang, Minghui Yu, Tao Wang 0011, Yidong Li |
NeurIPS | 4 |
| 2020 | Two-Stage Relation Constraint for Semantic Segmentation of Point Clouds
Minghui Yu, Jinxian Liu, Bingbing Ni, Caiyuan Li |
3DV | 1 |
| 2020 | Self-Prediction for Joint Instance and Semantic Segmentation of Point Clouds
Jinxian Liu, Minghui Yu, Bingbing Ni, Ye Chen 0006 |
ECCV (22) | 2 |