EDBT 2026 Demo / reviewers in the wild / expert
Tianchen Huang
dblp:278/7729
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Deep learning architectures and training · 93% Language models and text generation · 7% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
attention mechanism |
1.0 | 1 | 2026 | Making Every Head Count: Sparse Attention Without the Speed-Performance Trade-off · AAAI 2026 |
Machine learning › Deep learning architectures and training › attention mechanism
multi-head attention |
1.0 | 1 | 2026 | Making Every Head Count: Sparse Attention Without the Speed-Performance Trade-off · AAAI 2026 |
Machine learning › Deep learning architectures and training › attention mechanism
sparse attention |
1.0 | 1 | 2026 | Making Every Head Count: Sparse Attention Without the Speed-Performance Trade-off · AAAI 2026 |
Machine learning › Deep learning architectures and training
transformer |
1.0 | 1 | 2026 | Making Every Head Count: Sparse Attention Without the Speed-Performance Trade-off · AAAI 2026 |
Natural language and speech › Language models and text generation › efficient language model
large language model efficiency |
0.3 | 1 | 2026 | Making Every Head Count: Sparse Attention Without the Speed-Performance Trade-off · AAAI 2026 |
Methods — techniques the papers use, named apart from their topics
structural sparsity · 1.0head specialization · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Making Every Head Count: Sparse Attention Without the Speed-Performance Trade-offabstractThe design of Large Language Models (LLMs) has long been hampered by a fundamental conflict within their core attention mechanism: its remarkable expressivity is built upon a computational complexity of O(H·N²) that grows quadratically with the context size (N) and linearly with the number of heads (H). This standard implementation harbors significant computational redundancy, as all heads independently compute attention over the same sequence space. Existing sparse methods, meanwhile, often trade information integrity for computational efficiency. To resolve this efficiency-performance trade-off, we propose SPAttention, whose core contribution is the introduction of a new paradigm we term Principled Structural Sparsity. SPAttention does not merely drop connections but instead reorganizes the computational task by partitioning the total attention workload into balanced, non-overlapping distance bands, assigning each head a unique segment. This approach transforms the multi-head attention mechanism from H independent O(N²) computations into a single, collaborative O(N²) computation, fundamentally reducing complexity by a factor of H. The structured inductive bias compels functional specialization among heads, enabling a more efficient allocation of computational resources from redundant modeling to distinct dependencies across the entire sequence span. Extensive empirical validation on the OLMoE-1B-7B and 0.25B-1.75B model series demonstrates that while delivering an approximately two-fold increase in training throughput, its performance is on par with standard dense attention, even surpassing it on select key metrics, while consistently outperforming representative sparse attention methods including Longformer, Reformer, and BigBird across all evaluation metrics. Our work demonstrates that thoughtfully designed structural sparsity can serve as an effective inductive bias that simultaneously improves both computational efficiency and model performance, opening a new avenue for the architectural design of next-generation, high-performance LLMs. Mingkuan Zhao, Jiayin Wang 0002, Xin Lai 0003, Tianchen Huang, Yuheng Min, Xiaoyan Zhu 0003 |
AAAI | 5 |