Chanakya Ajit Ekbote

dblp:258/1037 · also Chanakya Ekbote · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Deep learning architectures and training · 24% Representation and self-supervised learning · 20% Graph learning · 13%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Medical and health informatics · 60% Bioinformatics and computational biology · 40%

Topics — the 15 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
transformer
1.622025
What One Cannot, Two Can: Two-Layer Transformers Provably Represent Induction Heads on Any-Order Markov Chains · NeurIPS 2025
Local to Global: Learning Dynamics and Effect of Initialization for Transformers · NeurIPS 2024
Machine learning › Representation and self-supervised learning › representation matching
feature alignment
0.912025
Understanding the Emergence of Multimodal Representation Alignment · ICML 2025
Natural language and speech › Language models and text generation
in-context learning
0.912025
What One Cannot, Two Can: Two-Layer Transformers Provably Represent Induction Heads on Any-Order Markov Chains · NeurIPS 2025
Machine learning › Representation and self-supervised learning
multimodal representation learning
0.912025
Understanding the Emergence of Multimodal Representation Alignment · ICML 2025
Machine learning › Reinforcement learning
policy optimization
0.912025
QoQ-Med: Building Multimodal Clinical Foundation Models with Domain-Aware GRPO Training · NeurIPS 2025
Machine learning › Learning theory
learning dynamics
0.812024
Local to Global: Learning Dynamics and Effect of Initialization for Transformers · NeurIPS 2024
Machine learning › Optimization for machine learning
optimization landscape
0.812024
Local to Global: Learning Dynamics and Effect of Initialization for Transformers · NeurIPS 2024
Machine learning › Deep learning architectures and training
weight initialization
0.812024
Local to Global: Learning Dynamics and Effect of Initialization for Transformers · NeurIPS 2024
Machine learning › Graph learning › graph representation learning
node representation learning
0.712023
FiGURe: Simple and Efficient Unsupervised Node Representations with Filter Augmentations · NeurIPS 2023
Machine learning › Graph learning › graph representation learning
unsupervised graph representation learning
0.712023
FiGURe: Simple and Efficient Unsupervised Node Representations with Filter Augmentations · NeurIPS 2023
Machine learning › Generative modeling
generative flow networks
0.612022
Biological Sequence Design with GFlowNets · ICML 2022
Bioinformatics and computational biology › synthetic biology
biological sequence design
0.612022
Biological Sequence Design with GFlowNets · ICML 2022
Natural language and speech › Language models and text generation
multimodal language model
0.312025
QoQ-Med: Building Multimodal Clinical Foundation Models with Domain-Aware GRPO Training · NeurIPS 2025
Algorithms and data structures
markov chains
0.312025
What One Cannot, Two Can: Two-Layer Transformers Provably Represent Induction Heads on Any-Order Markov Chains · NeurIPS 2025
Machine learning › Representation and self-supervised learning
contrastive learning
0.212023
FiGURe: Simple and Efficient Unsupervised Node Representations with Filter Augmentations · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 1.7instruction tuning · 1.7induction head analysis · 1.7group relative policy optimization · 1.7circuit construction · 1.7empirical investigation · 0.9markov chain · 0.8gradient descent · 0.8filter augmentation · 0.7contrastive learning · 0.7epistemic uncertainty estimation · 0.6active learning · 0.6GFlowNets · 0.6
YearPublicationVenuePosition
2025 Understanding the Emergence of Multimodal Representation Alignment
abstract
Multimodal representation learning is fundamentally about transforming incomparable modalities into comparable representations. While prior research has primarily focused on explicitly aligning these representations through targeted learning objectives and model architectures, a recent line of work has found that independently trained unimodal models of increasing scale and performance can become implicitly aligned with each other. These findings raise fundamental questions regarding the emergence of aligned representations in multimodal learning. Specifically: (1) when and why does alignment emerge implicitly? and (2) is alignment a reliable indicator of performance? Through a comprehensive empirical investigation, we demonstrate that both the emergence of alignment and its relationship with task performance depend on several critical data characteristics. These include, but are not necessarily limited to, the degree of similarity between the modalities and the balance between redundant and unique information they provide for the task. Our findings suggest that alignment may not be universally beneficial; rather, its impact on performance varies depending on the dataset and task. These insights can help practitioners determine whether increasing alignment between modalities is advantageous or, in some cases, detrimental to achieving optimal performance.
Megan Tjandrasuwita, Chanakya Ajit Ekbote, Liu Ziyin 0001, Paul Pu Liang
ICML2
2025 QoQ-Med: Building Multimodal Clinical Foundation Models with Domain-Aware GRPO Training
abstract
Clinical decision‑making routinely demands reasoning over heterogeneous data, yet existing multimodal language models (MLLMs) remain largely vision‑centric and fail to generalize across clinical specialties. To bridge this gap, we introduce QoQ-Med-7B/32B, the first open generalist clinical foundation model that jointly reasons across medical images, time‑series signals, and text reports. QoQ-Med is trained with Domain‑aware Relative Policy Optimization (DRPO), a novel reinforcement‑learning objective that hierarchically scales normalized rewards according to domain rarity and modality difficulty, mitigating performance imbalance caused by skewed clinical data distributions. Trained on 2.61 million instruction tuning pairs spanning 9 clinical domains, we show that DRPO training boosts diagnostic performance by 43% in macro‑F1 on average across all visual domains as compared to other critic-free training methods like GRPO. Furthermore, with QoQ-Med trained on intensive segmentation data, it is able to highlight salient regions related to the diagnosis, with an IoU 10x higher than open models while reaching the performance of OpenAI o4-mini. To foster reproducibility and downstream research, we release (i) the full model weights, (ii) the modular training pipeline, and (iii) all intermediate reasoning traces.
David Dai, Chanakya Ajit Ekbote, Paul Pu Liang
NeurIPS3
2025 What One Cannot, Two Can: Two-Layer Transformers Provably Represent Induction Heads on Any-Order Markov Chains
abstract
In-context learning (ICL) is a hallmark capability of transformers, through which trained models learn to adapt to new tasks by leveraging information from the input context. Prior work has shown that ICL emerges in transformers due to the presence of special circuits called induction heads. Given the equivalence between induction heads and conditional $k$-grams, a recent line of work modeling sequential inputs as Markov processes has revealed the fundamental impact of model depth on its ICL capabilities: while a two-layer transformer can efficiently represent a conditional $1$-gram model, its single-layer counterpart cannot solve the task unless it is exponentially large. However, for higher order Markov sources, the best known constructions require at least three layers (each with a single attention head) - leaving open the question: *can a two-layer single-head transformer represent any $k^{\text{th}}$-order Markov process?* In this paper, we precisely address this and theoretically show that a two-layer transformer with one head per layer can indeed represent any conditional $k$-gram. Thus, our result provides the tightest known characterization of the interplay between transformer depth and Markov order for ICL. Building on this, we further analyze the learning dynamics of our two-layer construction, focusing on a simplified variant for first-order Markov chains, illustrating how effective in-context representations emerge during training. Together, these results deepen our current understanding of transformer-based ICL and illustrate how even shallow architectures can surprisingly exhibit strong ICL capabilities on structured sequence modeling tasks.
Chanakya Ajit Ekbote, Ashok Vardhan Makkuva, Marco Bondaschi, Nived Rajaraman, Michael Gastpar, Jason D. Lee, Paul Pu Liang
NeurIPS1
2024 Local to Global: Learning Dynamics and Effect of Initialization for Transformers
abstract
In recent years, transformer-based models have revolutionized deep learning, particularly in sequence modeling. To better understand this phenomenon, there is a growing interest in using Markov input processes to study transformers. However, our current understanding in this regard remains limited with many fundamental questions about how transformers learn Markov chains still unanswered. In this paper, we address this by focusing on first-order Markov chains and single-layer transformers, providing a comprehensive characterization of the learning dynamics in this context. Specifically, we prove that transformer parameters trained on next-token prediction loss can either converge to global or local minima, contingent on the initialization and the Markovian data properties, and we characterize the precise conditions under which this occurs. To the best of our knowledge, this is the first result of its kind highlighting the role of initialization. We further demonstrate that our theoretical findings are corroborated by empirical evidence. Based on these insights, we provide guidelines for the initialization of single-layer transformers and demonstrate their effectiveness. Finally, we outline several open problems in this arena. Code is available at: \url{https://github.com/Bond1995/Markov}.
Ashok Vardhan Makkuva, Marco Bondaschi, Adway Girish, Alliot Nagle, Hyeji Kim, Michael Gastpar, Chanakya Ajit Ekbote
NeurIPS7
2023 FiGURe: Simple and Efficient Unsupervised Node Representations with Filter Augmentations
abstract
Unsupervised node representations learnt using contrastive learning-based methods have shown good performance on downstream tasks. However, these methods rely on augmentations that mimic low-pass filters, limiting their performance on tasks requiring different eigen-spectrum parts. This paper presents a simple filter-based augmentation method to capture different parts of the eigen-spectrum. We show significant improvements using these augmentations. Further, we show that sharing the same weights across these different filter augmentations is possible, reducing the computational load. In addition, previous works have shown that good performance on downstream tasks requires high dimensional representations. Working with high dimensions increases the computations, especially when multiple augmentations are involved. We mitigate this problem and recover good performance through lower dimensional embeddings using simple random Fourier feature projections. Our method, FiGURe, achieves an average gain of up to 4.4\%, compared to the state-of-the-art unsupervised models, across all datasets in consideration, both homophilic and heterophilic. Our code can be found at: https://github.com/Microsoft/figure.
Chanakya Ajit Ekbote, Ajinkya Pankaj Deshpande, Arun Iyer, Sundararajan Sellamanickam, Ramakrishna Bairi
NeurIPS1
2022 Biological Sequence Design with GFlowNets
abstract
Design of de novo biological sequences with desired properties, like protein and DNA sequences, often involves an active loop with several rounds of molecule ideation and expensive wet-lab evaluations. These experiments can consist of multiple stages, with increasing levels of precision and cost of evaluation, where candidates are filtered. This makes the diversity of proposed candidates a key consideration in the ideation phase. In this work, we propose an active learning algorithm leveraging epistemic uncertainty estimation and the recently proposed GFlowNets as a generator of diverse candidate solutions, with the objective to obtain a diverse batch of useful (as defined by some utility function, for example, the predicted anti-microbial activity of a peptide) and informative candidates after each round. We also propose a scheme to incorporate existing labeled datasets of candidates, in addition to a reward function, to speed up learning in GFlowNets. We present empirical results on several biological sequence design tasks, and we find that our method generates more diverse and novel batches with high scoring candidates compared to existing approaches.
Moksh Jain, Emmanuel Bengio, Alex Hernández-García, Jarrid Rector-Brooks, Bonaventure F. P. Dossou, Chanakya Ajit Ekbote, Jie Fu 0001, Michael Kilgour, Dinghuai Zhang, Lena Simine, Yoshua Bengio
ICML6
2022 A Piece-Wise Polynomial Filtering Approach for Graph Neural Networks
Vijay Lingam, Manan Sharma, Chanakya Ajit Ekbote, Rahul Ragesh, Arun Iyer, Sundararajan Sellamanickam
ECML/PKDD (2)3