EDBT 2026 Demo / reviewers in the wild / expert
Ameya Mahabaleshwarkar
dblp:308/8911 · also Ameya Sunil Mahabaleshwarkar
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Efficient and distributed learning · 50% Deep learning architectures and training · 35% Language models and text generation · 15% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
attention mechanism |
0.9 | 1 | 2025 | Hymba: A Hybrid-head Architecture for Small Language Models · ICLR 2025 |
Machine learning › Deep learning architectures and training › attention mechanism
hybrid attention |
0.9 | 1 | 2025 | Hymba: A Hybrid-head Architecture for Small Language Models · ICLR 2025 |
Machine learning › Efficient and distributed learning
model compression |
0.9 | 1 | 2025 | Efficient Hybrid Language Model Compression through Group-Aware SSM Pruning · NeurIPS 2025 |
Machine learning › Efficient and distributed learning › model compression
pruning |
0.9 | 1 | 2025 | Efficient Hybrid Language Model Compression through Group-Aware SSM Pruning · NeurIPS 2025 |
Machine learning › Efficient and distributed learning › model compression › lightweight neural network
small language models |
0.9 | 1 | 2025 | Hymba: A Hybrid-head Architecture for Small Language Models · ICLR 2025 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.3 | 1 | 2025 | Efficient Hybrid Language Model Compression through Group-Aware SSM Pruning · NeurIPS 2025 |
Machine learning › Deep learning architectures and training
state space model |
0.3 | 1 | 2025 | Hymba: A Hybrid-head Architecture for Small Language Models · ICLR 2025 |
Methods — techniques the papers use, named apart from their topics
state space model pruning · 0.9meta tokens · 0.9knowledge distillation · 0.9group-aware pruning · 0.9cross-layer key-value sharing · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Hymba: A Hybrid-head Architecture for Small Language ModelsabstractWe propose Hymba, a family of small language models featuring a hybrid-head parallel architecture that integrates attention mechanisms and state space models (SSMs) within the same layer, offering parallel and complementary processing of the same inputs. In this hybrid-head module, attention heads provide high-resolution recall, while SSM heads facilitate efficient context summarization. Additionally, we introduce learnable meta tokens, which are prepended to prompts to store critical meta information, guiding subsequent tokens and alleviating the “forced-to-attend” burden associated with attention mechanisms. Thanks to the global context summarized by SSMs, the attention heads in our model can be further optimized through cross-layer key-value (KV) sharing and a mix of global and local attention, resulting in a compact cache size without compromising accuracy. Notably, Hymba achieves state-of-the-art performance among small LMs: Our Hymba-1.5B-Base model surpasses all sub-2B public models and even outperforms Llama-3.2-3B, achieving 1.32\% higher average accuracy, an 11.67$\times$ reduction in cache size, and 3.49$\times$ higher throughput. Xin Dong 0009, Yonggan Fu, Shizhe Diao, Wonmin Byeon, Zijia Chen, Ameya Mahabaleshwarkar, Shih-Yang Liu, Matthijs Van Keirsbilck, Min-Hung Chen, Yoshi Suhara, Yingyan (Celine) Lin, Jan Kautz, Pavlo Molchanov 0001 |
ICLR | 6 |
| 2025 | When2Call: When (not) to Call ToolsabstractHayley Ross, Ameya Sunil Mahabaleshwarkar, Yoshi Suhara. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Hayley Ross, Ameya Mahabaleshwarkar, Yoshi Suhara |
NAACL (Long Papers) | 2 |
| 2025 | Efficient Hybrid Language Model Compression through Group-Aware SSM PruningabstractHybrid language models that combine Attention and State Space Models (SSMs) have been shown to achieve state-of-the-art accuracy and runtime performance. Recent work has also demonstrated that applying pruning and distillation to Attention-only models yields smaller, more accurate models at a fraction of the training cost. In this work, we explore the effectiveness of compressing Hybrid architectures. To this end, we introduce a novel group-aware pruning method for Mamba layers that preserves the structural integrity of SSM blocks and their sequence modeling capabilities. We combine this method with FFN, embedding dimension, and layer pruning, along with knowledge distillation-based retraining to obtain a unified compression recipe for hybrid models. Using this recipe, we compress the Nemotron-H 8B Hybrid model down to 4B parameters with up to $40\times$ fewer training tokens compared to similarly-sized models. The resulting model surpasses the accuracy of similarly-sized models while achieving $\sim2\times$ faster inference throughput, significantly advancing the Pareto frontier. Ali Taghibakhshi, Sharath Turuvekere Sreenivas, Saurav Muralidharan, Marcin Chochowski, Yashaswi Karnati, Raviraj Joshi, Ameya Mahabaleshwarkar, Zijia Chen, Yoshi Suhara, Oluwatobi Olabiyi, Daniel Korzekwa, Mostofa Patwary, Mohammad Shoeybi, Jan Kautz, Bryan Catanzaro, Ashwath Aithal, Nima Tajbakhsh, Pavlo Molchanov 0001 |
NeurIPS | 7 |