Dwaraknath Gnaneshwar

dblp:266/2830 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
3since 2021 · last 2025
0000-0002-3771-8433ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Deep learning architectures and training · 54% Language models and text generation · 21% Knowledge representation and reasoning · 9%

Topics — the 12 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
attention mechanism
0.912025
Rope to Nope and Back Again: A New Hybrid Attention Strategy · NeurIPS 2025
Machine learning › Deep learning architectures and training › attention mechanism
hybrid attention
0.912025
Rope to Nope and Back Again: A New Hybrid Attention Strategy · NeurIPS 2025
Natural language and speech › Language models and text generation
large language model reasoning
0.912025
Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models · ICLR 2025
Machine learning › Deep learning architectures and training
positional encoding
0.912025
Rope to Nope and Back Again: A New Hybrid Attention Strategy · NeurIPS 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge structures
procedural knowledge
0.912025
Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models · ICLR 2025
Machine learning › Deep learning architectures and training › positional encoding
rotary position embedding
0.912025
Rope to Nope and Back Again: A New Hybrid Attention Strategy · NeurIPS 2025
Machine learning › Deep learning architectures and training
transformer
0.912025
Rope to Nope and Back Again: A New Hybrid Attention Strategy · NeurIPS 2025
Machine learning › Deep learning architectures and training
mixture of experts
0.812024
BAM! Just Like That: Simple and Efficient Parameter Upcycling for Mixture of Experts · NeurIPS 2024
Natural language and speech › Information extraction and text analysis › text classification
sentence classification
0.412020
Leveraging BERT with Mixup for Sentence Classification (Student Abstract) · AAAI 2020
Natural language and speech › Language models and text generation › language modeling › long-context language modeling › context utilization
long-context modeling
0.312025
Rope to Nope and Back Again: A New Hybrid Attention Strategy · NeurIPS 2025
Machine learning › Learning theory
generalization
0.112020
Leveraging BERT with Mixup for Sentence Classification (Student Abstract) · AAAI 2020
Machine learning › Trustworthy machine learning
robustness
0.112020
Leveraging BERT with Mixup for Sentence Classification (Student Abstract) · AAAI 2020

Methods — techniques the papers use, named apart from their topics

query-key normalization · 0.9influence functions · 0.9NoPE · 0.9parameter upcycling · 0.8mixture of attention · 0.8manifold mixup · 0.4ablation study · 0.4BERT · 0.4
YearPublicationVenuePosition
2025 Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models
abstract
The capabilities and limitations of Large Language Models (LLMs) have been sketched out in great detail in recent years, providing an intriguing yet conflicting picture. On the one hand, LLMs demonstrate a general ability to solve problems. On the other hand, they show surprising reasoning gaps when compared to humans, casting doubt on the robustness of their generalisation strategies. The sheer volume of data used in the design of LLMs has precluded us from applying the method traditionally used to measure generalisation: train-test set separation. To overcome this, we study what kind of generalisation strategies LLMs employ when performing reasoning tasks by investigating the pretraining data they rely on. For two models of different sizes (7B and 35B) and 2.5B of their pretraining tokens, we identify what documents influence the model outputs for three simple mathematical reasoning tasks and contrast this to the data that are influential for answering factual questions. We find that, while the models rely on mostly distinct sets of data for each factual question, a document often has a similar influence across different reasoning questions within the same task, indicating the presence of procedural knowledge. We further find that the answers to factual questions often show up in the most influential data. However, for reasoning questions the answers usually do not show up as highly influential, nor do the answers to the intermediate reasoning steps. When we characterise the top ranked documents for the reasoning questions qualitatively, we confirm that the influential documents often contain procedural knowledge, like demonstrating how to obtain a solution using formulae or code. Our findings indicate that the approach to reasoning the models use is unlike retrieval, and more like a generalisable strategy that synthesises procedural knowledge from documents doing a similar form of reasoning.
Laura Ruis, Maximilian Mozes, Juhan Bae, Siddhartha Rao Kamalakara, Dwaraknath Gnaneshwar, Acyr Locatelli, Robert Kirk, Tim Rocktäschel, Edward Grefenstette, Max Bartolo
ICLR5
2025 Rope to Nope and Back Again: A New Hybrid Attention Strategy
abstract
Long-context large language models (LLMs) have achieved remarkable advancements, driven by techniques like Rotary Position Embedding (RoPE) (Su et al., 2023) and its extensions (Chen et al., 2023; Liu et al., 2024c; Peng et al., 2023). By adjusting RoPE parameters and incorporating training data with extended contexts, we can train performant models with considerably longer input sequences. However, existing RoPE-based methods exhibit performance limitations when applied to extended context lengths. This paper presents a comprehensive analysis of various attention mechanisms, including RoPE, No Positional Embedding (NoPE), and Query-Key Normalization (QK-Norm), identifying their strengths and shortcomings in long-context modeling. Our investigation identifies distinctive attention patterns in these methods and highlights their impact on long-context performance, providing valuable insights for architectural design. on long context performance, providing valuable insights for architectural design. Building on these findings, we propose a novel architecture featuring a hybrid attention mechanism that integrates global and local attention spans. This design not only surpasses conventional RoPE-based transformer models with full attention in both long and short context tasks but also delivers substantial efficiency gains during training and inference.
Bharat Venkitesh, Dwaraknath Gnaneshwar, David Cairuz, Phil Blunsom, Acyr Locatelli
NeurIPS3
2024 BAM! Just Like That: Simple and Efficient Parameter Upcycling for Mixture of Experts
abstract
Mixture of Experts (MoE) framework has become a popular architecture for large language models due to its superior performance compared to dense models. However, training MoEs from scratch in a large-scale regime is prohibitively expensive. Previous work addresses this challenge by independently training multiple dense expert models and using them to initialize an MoE. In particular, state-of-the-art approaches initialize MoE layers using experts' feed-forward parameters while merging all other parameters, limiting the advantages of the specialized dense models when upcycling them as MoEs. We propose BAM (Branch-Attend-Mix), a simple yet effective improvement to MoE training. BAM makes full use of specialized dense models by not only using their feed-forward network (FFN) to initialize the MoE layers but also leveraging experts' attention weights fully by leveraging them as mixture-of-attention (MoA) layers. We explore two methods for upcycling MoA layers: 1) initializing separate attention experts from dense models including key, value, and query matrices; and 2) initializing only Q projections while sharing key-value pairs across all experts to facilitate efficient inference. Our experiments using seed models ranging from 590 million to 2 billion parameters show that our approach outperforms state-of-the-art approaches under the same data and compute budget in both perplexity and downstream tasks evaluations, confirming the effectiveness of BAM.
Qizhen Zhang 0002, Nikolas Gritsch, Dwaraknath Gnaneshwar, Simon Guo 0003, David Cairuz, Bharat Venkitesh, Jakob N. Foerster, Phil Blunsom, Sebastian Ruder, Ahmet Üstün, Acyr Locatelli
NeurIPS3
2020 Leveraging BERT with Mixup for Sentence Classification (Student Abstract)
abstract
Good generalization capability is an important quality of well-trained and robust neural networks. However, networks usually struggle when faced with samples outside the training distribution. Mixup is a technique that improves generalization, reduces memorization, and increases adversarial robustness. We apply a variant of Mixup called Manifold Mixup to the sentence classification problem, and present the results along with an ablation study. Our methodology outperforms CNN, LSTM, and vanilla BERT models in generalization.
Amit Jindal, Dwaraknath Gnaneshwar, Ramit Sawhney, Rajiv Ratn Shah
AAAI2
2020 NABU - Multilingual Graph-Based Neural RDF Verbalizer
Diego Moussallem, Dwaraknath Gnaneshwar, Thiago Castro Ferreira, Axel-Cyrille Ngonga Ngomo
ISWC (1)2