VLDB 2026 Research / reviewers in the wild / expert
Chanakya Ajit Ekbote
dblp:258/1037 · also Chanakya Ekbote
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Deep learning architectures and training · 24% Representation and self-supervised learning · 20% Graph learning · 13% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Medical and health informatics · 60% Bioinformatics and computational biology · 40% |
Topics — the 15 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
transformer |
1.6 | 2 | 2025 | What One Cannot, Two Can: Two-Layer Transformers Provably Represent Induction Heads on Any-Order Markov Chains · NeurIPS 2025 Local to Global: Learning Dynamics and Effect of Initialization for Transformers · NeurIPS 2024 |
Machine learning › Representation and self-supervised learning › representation matching
feature alignment |
0.9 | 1 | 2025 | Understanding the Emergence of Multimodal Representation Alignment · ICML 2025 |
Natural language and speech › Language models and text generation
in-context learning |
0.9 | 1 | 2025 | What One Cannot, Two Can: Two-Layer Transformers Provably Represent Induction Heads on Any-Order Markov Chains · NeurIPS 2025 |
Machine learning › Representation and self-supervised learning
multimodal representation learning |
0.9 | 1 | 2025 | Understanding the Emergence of Multimodal Representation Alignment · ICML 2025 |
Machine learning › Reinforcement learning
policy optimization |
0.9 | 1 | 2025 | QoQ-Med: Building Multimodal Clinical Foundation Models with Domain-Aware GRPO Training · NeurIPS 2025 |
Machine learning › Learning theory
learning dynamics |
0.8 | 1 | 2024 | Local to Global: Learning Dynamics and Effect of Initialization for Transformers · NeurIPS 2024 |
Machine learning › Optimization for machine learning
optimization landscape |
0.8 | 1 | 2024 | Local to Global: Learning Dynamics and Effect of Initialization for Transformers · NeurIPS 2024 |
Machine learning › Deep learning architectures and training
weight initialization |
0.8 | 1 | 2024 | Local to Global: Learning Dynamics and Effect of Initialization for Transformers · NeurIPS 2024 |
Machine learning › Graph learning › graph representation learning
node representation learning |
0.7 | 1 | 2023 | FiGURe: Simple and Efficient Unsupervised Node Representations with Filter Augmentations · NeurIPS 2023 |
Machine learning › Graph learning › graph representation learning
unsupervised graph representation learning |
0.7 | 1 | 2023 | FiGURe: Simple and Efficient Unsupervised Node Representations with Filter Augmentations · NeurIPS 2023 |
Machine learning › Generative modeling
generative flow networks |
0.6 | 1 | 2022 | Biological Sequence Design with GFlowNets · ICML 2022 |
Bioinformatics and computational biology › synthetic biology
biological sequence design |
0.6 | 1 | 2022 | Biological Sequence Design with GFlowNets · ICML 2022 |
Natural language and speech › Language models and text generation
multimodal language model |
0.3 | 1 | 2025 | QoQ-Med: Building Multimodal Clinical Foundation Models with Domain-Aware GRPO Training · NeurIPS 2025 |
Algorithms and data structures
markov chains |
0.3 | 1 | 2025 | What One Cannot, Two Can: Two-Layer Transformers Provably Represent Induction Heads on Any-Order Markov Chains · NeurIPS 2025 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.2 | 1 | 2023 | FiGURe: Simple and Efficient Unsupervised Node Representations with Filter Augmentations · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 1.7instruction tuning · 1.7induction head analysis · 1.7group relative policy optimization · 1.7circuit construction · 1.7empirical investigation · 0.9markov chain · 0.8gradient descent · 0.8filter augmentation · 0.7contrastive learning · 0.7epistemic uncertainty estimation · 0.6active learning · 0.6GFlowNets · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Understanding the Emergence of Multimodal Representation AlignmentabstractMultimodal representation learning is fundamentally about transforming incomparable modalities into comparable representations. While prior research has primarily focused on explicitly aligning these representations through targeted learning objectives and model architectures, a recent line of work has found that independently trained unimodal models of increasing scale and performance can become implicitly aligned with each other. These findings raise fundamental questions regarding the emergence of aligned representations in multimodal learning. Specifically: (1) when and why does alignment emerge implicitly? and (2) is alignment a reliable indicator of performance? Through a comprehensive empirical investigation, we demonstrate that both the emergence of alignment and its relationship with task performance depend on several critical data characteristics. These include, but are not necessarily limited to, the degree of similarity between the modalities and the balance between redundant and unique information they provide for the task. Our findings suggest that alignment may not be universally beneficial; rather, its impact on performance varies depending on the dataset and task. These insights can help practitioners determine whether increasing alignment between modalities is advantageous or, in some cases, detrimental to achieving optimal performance. Megan Tjandrasuwita, Chanakya Ajit Ekbote, Liu Ziyin 0001, Paul Pu Liang |
ICML | 2 |
| 2025 | QoQ-Med: Building Multimodal Clinical Foundation Models with Domain-Aware GRPO TrainingabstractClinical decision‑making routinely demands reasoning over heterogeneous data, yet existing multimodal language models (MLLMs) remain largely vision‑centric and fail to generalize across clinical specialties. To bridge this gap, we introduce QoQ-Med-7B/32B, the first open generalist clinical foundation model that jointly reasons across medical images, time‑series signals, and text reports. QoQ-Med is trained with Domain‑aware Relative Policy Optimization (DRPO), a novel reinforcement‑learning objective that hierarchically scales normalized rewards according to domain rarity and modality difficulty, mitigating performance imbalance caused by skewed clinical data distributions. Trained on 2.61 million instruction tuning pairs spanning 9 clinical domains, we show that DRPO training boosts diagnostic performance by 43% in macro‑F1 on average across all visual domains as compared to other critic-free training methods like GRPO. Furthermore, with QoQ-Med trained on intensive segmentation data, it is able to highlight salient regions related to the diagnosis, with an IoU 10x higher than open models while reaching the performance of OpenAI o4-mini. To foster reproducibility and downstream research, we release (i) the full model weights, (ii) the modular training pipeline, and (iii) all intermediate reasoning traces. David Dai, Chanakya Ajit Ekbote, Paul Pu Liang |
NeurIPS | 3 |
| 2025 | What One Cannot, Two Can: Two-Layer Transformers Provably Represent Induction Heads on Any-Order Markov ChainsabstractIn-context learning (ICL) is a hallmark capability of transformers, through which trained models learn to adapt to new tasks by leveraging information from the input context. Prior work has shown that ICL emerges in transformers due to the presence of special circuits called induction heads. Given the equivalence between induction heads and conditional $k$-grams, a recent line of work modeling sequential inputs as Markov processes has revealed the fundamental impact of model depth on its ICL capabilities: while a two-layer transformer can efficiently represent a conditional $1$-gram model, its single-layer counterpart cannot solve the task unless it is exponentially large. However, for higher order Markov sources, the best known constructions require at least three layers (each with a single attention head) - leaving open the question: *can a two-layer single-head transformer represent any $k^{\text{th}}$-order Markov process?* In this paper, we precisely address this and theoretically show that a two-layer transformer with one head per layer can indeed represent any conditional $k$-gram. Thus, our result provides the tightest known characterization of the interplay between transformer depth and Markov order for ICL. Building on this, we further analyze the learning dynamics of our two-layer construction, focusing on a simplified variant for first-order Markov chains, illustrating how effective in-context representations emerge during training. Together, these results deepen our current understanding of transformer-based ICL and illustrate how even shallow architectures can surprisingly exhibit strong ICL capabilities on structured sequence modeling tasks. Chanakya Ajit Ekbote, Ashok Vardhan Makkuva, Marco Bondaschi, Nived Rajaraman, Michael Gastpar, Jason D. Lee, Paul Pu Liang |
NeurIPS | 1 |
| 2024 | Local to Global: Learning Dynamics and Effect of Initialization for TransformersabstractIn recent years, transformer-based models have revolutionized deep learning, particularly in sequence modeling. To better understand this phenomenon, there is a growing interest in using Markov input processes to study transformers. However, our current understanding in this regard remains limited with many fundamental questions about how transformers learn Markov chains still unanswered. In this paper, we address this by focusing on first-order Markov chains and single-layer transformers, providing a comprehensive characterization of the learning dynamics in this context. Specifically, we prove that transformer parameters trained on next-token prediction loss can either converge to global or local minima, contingent on the initialization and the Markovian data properties, and we characterize the precise conditions under which this occurs. To the best of our knowledge, this is the first result of its kind highlighting the role of initialization. We further demonstrate that our theoretical findings are corroborated by empirical evidence. Based on these insights, we provide guidelines for the initialization of single-layer transformers and demonstrate their effectiveness. Finally, we outline several open problems in this arena. Code is available at: \url{https://github.com/Bond1995/Markov}. Ashok Vardhan Makkuva, Marco Bondaschi, Adway Girish, Alliot Nagle, Hyeji Kim, Michael Gastpar, Chanakya Ajit Ekbote |
NeurIPS | 7 |
| 2023 | FiGURe: Simple and Efficient Unsupervised Node Representations with Filter AugmentationsabstractUnsupervised node representations learnt using contrastive learning-based methods have shown good performance on downstream tasks. However, these methods rely on augmentations that mimic low-pass filters, limiting their performance on tasks requiring different eigen-spectrum parts. This paper presents a simple filter-based augmentation method to capture different parts of the eigen-spectrum. We show significant improvements using these augmentations. Further, we show that sharing the same weights across these different filter augmentations is possible, reducing the computational load. In addition, previous works have shown that good performance on downstream tasks requires high dimensional representations. Working with high dimensions increases the computations, especially when multiple augmentations are involved. We mitigate this problem and recover good performance through lower dimensional embeddings using simple random Fourier feature projections. Our method, FiGURe, achieves an average gain of up to 4.4\%, compared to the state-of-the-art unsupervised models, across all datasets in consideration, both homophilic and heterophilic. Our code can be found at: https://github.com/Microsoft/figure. Chanakya Ajit Ekbote, Ajinkya Pankaj Deshpande, Arun Iyer, Sundararajan Sellamanickam, Ramakrishna Bairi |
NeurIPS | 1 |
| 2022 | Biological Sequence Design with GFlowNetsabstractDesign of de novo biological sequences with desired properties, like protein and DNA sequences, often involves an active loop with several rounds of molecule ideation and expensive wet-lab evaluations. These experiments can consist of multiple stages, with increasing levels of precision and cost of evaluation, where candidates are filtered. This makes the diversity of proposed candidates a key consideration in the ideation phase. In this work, we propose an active learning algorithm leveraging epistemic uncertainty estimation and the recently proposed GFlowNets as a generator of diverse candidate solutions, with the objective to obtain a diverse batch of useful (as defined by some utility function, for example, the predicted anti-microbial activity of a peptide) and informative candidates after each round. We also propose a scheme to incorporate existing labeled datasets of candidates, in addition to a reward function, to speed up learning in GFlowNets. We present empirical results on several biological sequence design tasks, and we find that our method generates more diverse and novel batches with high scoring candidates compared to existing approaches. Moksh Jain, Emmanuel Bengio, Alex Hernández-García, Jarrid Rector-Brooks, Bonaventure F. P. Dossou, Chanakya Ajit Ekbote, Jie Fu 0001, Michael Kilgour, Dinghuai Zhang, Lena Simine, Yoshua Bengio |
ICML | 6 |
| 2022 | A Piece-Wise Polynomial Filtering Approach for Graph Neural Networks
Vijay Lingam, Manan Sharma, Chanakya Ajit Ekbote, Rahul Ragesh, Arun Iyer, Sundararajan Sellamanickam |
ECML/PKDD (2) | 3 |