EDBT 2026 Demo / reviewers in the wild / expert
Dwaraknath Gnaneshwar
dblp:266/2830
· DBLP profile ↗
5ranked-venue papers
0as first author
3since 2021 · last 2025
0000-0002-3771-8433ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 3 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Deep learning architectures and training · 54% Language models and text generation · 21% Knowledge representation and reasoning · 9% |
Topics — the 12 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
attention mechanism |
0.9 | 1 | 2025 | Rope to Nope and Back Again: A New Hybrid Attention Strategy · NeurIPS 2025 |
Machine learning › Deep learning architectures and training › attention mechanism
hybrid attention |
0.9 | 1 | 2025 | Rope to Nope and Back Again: A New Hybrid Attention Strategy · NeurIPS 2025 |
Natural language and speech › Language models and text generation
large language model reasoning |
0.9 | 1 | 2025 | Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models · ICLR 2025 |
Machine learning › Deep learning architectures and training
positional encoding |
0.9 | 1 | 2025 | Rope to Nope and Back Again: A New Hybrid Attention Strategy · NeurIPS 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge structures
procedural knowledge |
0.9 | 1 | 2025 | Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models · ICLR 2025 |
Machine learning › Deep learning architectures and training › positional encoding
rotary position embedding |
0.9 | 1 | 2025 | Rope to Nope and Back Again: A New Hybrid Attention Strategy · NeurIPS 2025 |
Machine learning › Deep learning architectures and training
transformer |
0.9 | 1 | 2025 | Rope to Nope and Back Again: A New Hybrid Attention Strategy · NeurIPS 2025 |
Machine learning › Deep learning architectures and training
mixture of experts |
0.8 | 1 | 2024 | BAM! Just Like That: Simple and Efficient Parameter Upcycling for Mixture of Experts · NeurIPS 2024 |
Natural language and speech › Information extraction and text analysis › text classification
sentence classification |
0.4 | 1 | 2020 | Leveraging BERT with Mixup for Sentence Classification (Student Abstract) · AAAI 2020 |
Natural language and speech › Language models and text generation › language modeling › long-context language modeling › context utilization
long-context modeling |
0.3 | 1 | 2025 | Rope to Nope and Back Again: A New Hybrid Attention Strategy · NeurIPS 2025 |
Machine learning › Learning theory
generalization |
0.1 | 1 | 2020 | Leveraging BERT with Mixup for Sentence Classification (Student Abstract) · AAAI 2020 |
Machine learning › Trustworthy machine learning
robustness |
0.1 | 1 | 2020 | Leveraging BERT with Mixup for Sentence Classification (Student Abstract) · AAAI 2020 |
Methods — techniques the papers use, named apart from their topics
query-key normalization · 0.9influence functions · 0.9NoPE · 0.9parameter upcycling · 0.8mixture of attention · 0.8manifold mixup · 0.4ablation study · 0.4BERT · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Procedural Knowledge in Pretraining Drives Reasoning in Large Language ModelsabstractThe capabilities and limitations of Large Language Models (LLMs) have been sketched out in great detail in recent years, providing an intriguing yet conflicting picture. On the one hand, LLMs demonstrate a general ability to solve problems. On the other hand, they show surprising reasoning gaps when compared to humans, casting doubt on the robustness of their generalisation strategies. The sheer volume of data used in the design of LLMs has precluded us from applying the method traditionally used to measure generalisation: train-test set separation. To overcome this, we study what kind of generalisation strategies LLMs employ when performing reasoning tasks by investigating the pretraining data they rely on. For two models of different sizes (7B and 35B) and 2.5B of their pretraining tokens, we identify what documents influence the model outputs for three simple mathematical reasoning tasks and contrast this to the data that are influential for answering factual questions. We find that, while the models rely on mostly distinct sets of data for each factual question, a document often has a similar influence across different reasoning questions within the same task, indicating the presence of procedural knowledge. We further find that the answers to factual questions often show up in the most influential data. However, for reasoning questions the answers usually do not show up as highly influential, nor do the answers to the intermediate reasoning steps. When we characterise the top ranked documents for the reasoning questions qualitatively, we confirm that the influential documents often contain procedural knowledge, like demonstrating how to obtain a solution using formulae or code. Our findings indicate that the approach to reasoning the models use is unlike retrieval, and more like a generalisable strategy that synthesises procedural knowledge from documents doing a similar form of reasoning. Laura Ruis, Maximilian Mozes, Juhan Bae, Siddhartha Rao Kamalakara, Dwaraknath Gnaneshwar, Acyr Locatelli, Robert Kirk, Tim Rocktäschel, Edward Grefenstette, Max Bartolo |
ICLR | 5 |
| 2025 | Rope to Nope and Back Again: A New Hybrid Attention StrategyabstractLong-context large language models (LLMs) have achieved remarkable advancements, driven by techniques like Rotary Position Embedding (RoPE) (Su et al., 2023) and its extensions (Chen et al., 2023; Liu et al., 2024c; Peng et al., 2023). By adjusting RoPE parameters and incorporating training data with extended contexts, we can train performant models with considerably longer input sequences. However, existing RoPE-based methods exhibit performance limitations when applied to extended context lengths. This paper presents a comprehensive analysis of various attention mechanisms, including RoPE, No Positional Embedding (NoPE), and Query-Key Normalization (QK-Norm), identifying their strengths and shortcomings in long-context modeling. Our investigation identifies distinctive attention patterns in these methods and highlights their impact on long-context performance, providing valuable insights for architectural design. on long context performance, providing valuable insights for architectural design. Building on these findings, we propose a novel architecture featuring a hybrid attention mechanism that integrates global and local attention spans. This design not only surpasses conventional RoPE-based transformer models with full attention in both long and short context tasks but also delivers substantial efficiency gains during training and inference. Bharat Venkitesh, Dwaraknath Gnaneshwar, David Cairuz, Phil Blunsom, Acyr Locatelli |
NeurIPS | 3 |
| 2024 | BAM! Just Like That: Simple and Efficient Parameter Upcycling for Mixture of ExpertsabstractMixture of Experts (MoE) framework has become a popular architecture for large language models due to its superior performance compared to dense models. However, training MoEs from scratch in a large-scale regime is prohibitively expensive. Previous work addresses this challenge by independently training multiple dense expert models and using them to initialize an MoE. In particular, state-of-the-art approaches initialize MoE layers using experts' feed-forward parameters while merging all other parameters, limiting the advantages of the specialized dense models when upcycling them as MoEs. We propose BAM (Branch-Attend-Mix), a simple yet effective improvement to MoE training. BAM makes full use of specialized dense models by not only using their feed-forward network (FFN) to initialize the MoE layers but also leveraging experts' attention weights fully by leveraging them as mixture-of-attention (MoA) layers. We explore two methods for upcycling MoA layers: 1) initializing separate attention experts from dense models including key, value, and query matrices; and 2) initializing only Q projections while sharing key-value pairs across all experts to facilitate efficient inference. Our experiments using seed models ranging from 590 million to 2 billion parameters show that our approach outperforms state-of-the-art approaches under the same data and compute budget in both perplexity and downstream tasks evaluations, confirming the effectiveness of BAM. Qizhen Zhang 0002, Nikolas Gritsch, Dwaraknath Gnaneshwar, Simon Guo 0003, David Cairuz, Bharat Venkitesh, Jakob N. Foerster, Phil Blunsom, Sebastian Ruder, Ahmet Üstün, Acyr Locatelli |
NeurIPS | 3 |
| 2020 | Leveraging BERT with Mixup for Sentence Classification (Student Abstract)abstractGood generalization capability is an important quality of well-trained and robust neural networks. However, networks usually struggle when faced with samples outside the training distribution. Mixup is a technique that improves generalization, reduces memorization, and increases adversarial robustness. We apply a variant of Mixup called Manifold Mixup to the sentence classification problem, and present the results along with an ablation study. Our methodology outperforms CNN, LSTM, and vanilla BERT models in generalization. Amit Jindal, Dwaraknath Gnaneshwar, Ramit Sawhney, Rajiv Ratn Shah |
AAAI | 2 |
| 2020 | NABU - Multilingual Graph-Based Neural RDF Verbalizer
Diego Moussallem, Dwaraknath Gnaneshwar, Thiago Castro Ferreira, Axel-Cyrille Ngonga Ngomo |
ISWC (1) | 2 |