Devaansh Gupta

dblp:351/9786 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2025
0009-0002-3007-0109ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Reinforcement learning · 32% Generative modeling · 16% Efficient and distributed learning · 16%
Databases, data mining, and information retrieval
3 papers
Data mining · 96% Information retrieval · 4%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining › predictive modeling › classification › multi-label classification
extreme multi-label classification
2.332025
UniDEC : Unified Dual Encoder and Classifier Training for Extreme Multi-Label Classification · WWW 2025
Gandalf: Learning Label-label Correlations in Extreme Multi-label Classification via Label Features · KDD 2024
InceptionXML: A Lightweight Framework with Synchronized Negative Sampling for Short Text Extreme Classification · SIGIR 2023
Machine learning › Generative modeling › diffusion model › discrete diffusion model
diffusion language model
0.912025
d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement Learning · NeurIPS 2025
Machine learning › Reinforcement learning › policy optimization
group relative policy optimization
0.912025
d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement Learning · NeurIPS 2025
Machine learning › Reinforcement learning › policy optimization
policy gradient
0.912025
d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement Learning · NeurIPS 2025
Data mining › predictive modeling
classification
0.812024
Gandalf: Learning Label-label Correlations in Extreme Multi-label Classification via Label Features · KDD 2024
Data mining › predictive modeling › classification › multi-label classification
label correlation learning
0.812024
Gandalf: Learning Label-label Correlations in Extreme Multi-label Classification via Label Features · KDD 2024
Data mining › predictive modeling › classification
multi-label classification
0.812024
Gandalf: Learning Label-label Correlations in Extreme Multi-label Classification via Label Features · KDD 2024
Natural language and speech › Machine translation
multimodal machine translation
0.712023
CLIPTrans: Transferring Visual Knowledge with Pre-trained Models for Multimodal Machine Translation · ICCV 2023
Machine learning › Representation and self-supervised learning › pre-training
multimodal pretraining
0.712023
CLIPTrans: Transferring Visual Knowledge with Pre-trained Models for Multimodal Machine Translation · ICCV 2023
Machine learning › Transfer learning and domain adaptation › cross-modal transfer
vision-language model transfer
0.712023
CLIPTrans: Transferring Visual Knowledge with Pre-trained Models for Multimodal Machine Translation · ICCV 2023
Information retrieval › retrieval models › neural retrieval
neural ranking model
0.212023
InceptionXML: A Lightweight Framework with Synchronized Negative Sampling for Short Text Extreme Classification · SIGIR 2023

Methods — techniques the papers use, named apart from their topics

dual encoder · 1.7supervised fine-tuning · 0.9policy gradient · 0.9one-vs-all classifiers · 0.9one-vs-all classifier · 0.9masked diffusion · 0.9label features · 0.8two-stage training · 0.7prefix sequence mapping · 0.7mBART · 0.7hard negative mining · 0.7convolutional neural network · 0.7CLIP · 0.7
YearPublicationVenuePosition
2025 d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement Learning
abstract
Recent large language models (LLMs) have demonstrated strong reasoning capabilities that benefits from online reinforcement learning (RL). These capabilities have primarily been demonstrated within the left-to-right autoregressive (AR) generation paradigm. In contrast, non-autoregressive paradigms based on diffusion generate text in a coarse-to-fine manner. Although recent diffusion-based large language models (dLLMs) have achieved competitive language modeling performance compared to their AR counterparts, it remains unclear if dLLMs can also leverage recent advances in LLM reasoning. To this end, we propose, a framework to adapt pre-trained masked dLLMs into reasoning models via a combination of supervised finetuning (SFT) and RL. Specifically, we develop and extend techniques to improve reasoning in pretrained dLLMs: (a) we utilize a masked SFT technique to distill knowledge and instill self-improvement behavior directly from existing datasets, and (b) we introduce a novel critic-free, policy-gradient based RL algorithm called diffu-GRPO, the first integration of policy gradient methods to masked dLLMs. Through empirical studies, we investigate the performance of different post-training recipes on multiple mathematical and planning benchmarks. We find that d1 yields the best performance and significantly improves performance of a state-of-the-art dLLM.
Siyan Zhao, Devaansh Gupta, Qinqing Zheng, Aditya Grover
NeurIPS2
2025 UniDEC : Unified Dual Encoder and Classifier Training for Extreme Multi-Label Classification
abstract
Extreme Multi-label Classification (XMC) involves predicting a subset of relevant labels from an extremely large label space, given an input query and labels with textual features.Models developed for this problem have conventionally made use of dual encoder (DE) to embed the queries and label texts and one-vs-all (OvA) classifiers to rerank the shortlisted labels by the DE.While such methods have shown empirical success, a major drawback is their computational cost, often requiring upto 16 GPUs to train on the largest public dataset.Such a high cost is a consequence of calculating the loss over the entire label space.While shortlisting strategies have been proposed for classifiers, we aim to study such methods for the DE framework.In this work, we develop UniDEC, a loss-independent, end-to-end trainable framework which trains the DE and classifier together in a unified manner with a multi-class loss, while reducing the computational cost by 4 -16×.This is done via the proposed pick-some-label (PSL) reduction, which aims to compute the loss on only a subset of positive and negative labels.These labels are carefully chosen in-batch so as to maximise their supervisory signals.Not only does the proposed framework achieve state-of-the-art results on datasets with labels in the order of millions, it is also computationally and resource efficient in achieving this performance on a single GPU.Code is made available at https://github.com/the-catalyst/UniDEC.
Siddhant Kharbanda, Devaansh Gupta, Gururaj K, Pankaj Malhotra, Amit Singh 0003, Cho-Jui Hsieh, Rohit Babbar
WWW2
2024 Gandalf: Learning Label-label Correlations in Extreme Multi-label Classification via Label Features
abstract
Publisher Copyright: © 2024 Copyright held by the owner/author(s).
Siddhant Kharbanda, Devaansh Gupta, Erik Schultheis, Atmadeep Banerjee, Cho-Jui Hsieh, Rohit Babbar
KDD2
2023 CLIPTrans: Transferring Visual Knowledge with Pre-trained Models for Multimodal Machine Translation
abstract
There has been a growing interest in developing multimodal machine translation (MMT) systems that enhance neural machine translation (NMT) with visual knowledge. This problem setup involves using images as auxiliary information during training, and more recently, eliminating their use during inference. Towards this end, previous works face a challenge in training powerful MMT models from scratch due to the scarcity of annotated multilingual vision-language data, especially for low-resource languages. Simultaneously, there has been an influx of multilingual pretrained models for NMT and multimodal pre-trained models for vision-language tasks, primarily in English, which have shown exceptional generalisation ability. However, these are not directly applicable to MMT since they do not provide aligned multimodal multilingual features for generative tasks. To alleviate this issue, instead of designing complex modules for MMT, we propose CLIPTrans, which simply adapts the independently pre-trained multimodal M-CLIP and the multilingual mBART. In order to align their embedding spaces, mBART is conditioned on the M-CLIP features by a prefix sequence generated through a lightweight mapping network. We train this in a two-stage pipeline which warms up the model with image captioning before the actual translation task. Through experiments, we demonstrate the merits of this framework and consequently push forward the state-of-the-art across standard benchmarks by an average of +2.67 BLEU. The code can be found at www.github.com/devaansh100/CLIPTrans.
Devaansh Gupta, Siddhant Kharbanda, Wanhua Li 0001, Hanspeter Pfister, Donglai Wei 0001
ICCV1
2023 InceptionXML: A Lightweight Framework with Synchronized Negative Sampling for Short Text Extreme Classification
abstract
Automatic annotation of short-text data to a large number of target labels, referred to as Short Text Extreme Classification, has found numerous applications including prediction of related searches and product recommendation. In this paper, we propose a convolutional architecture InceptionXML which is light-weight, yet powerful, and robust to the inherent lack of word-order in short-text queries encountered in search and recommendation. We demonstrate the efficacy of applying convolutions by recasting the operation along the embedding dimension instead of the word dimension as applied in conventional CNNs for text classification. Towards scaling our model to datasets with millions of labels, we also propose SyncXML pipeline which improves upon the shortcomings of the recently proposed dynamic hard-negative mining technique for label shortlisting by synchronizing the label-shortlister and extreme classifier. SyncXML not only reduces the inference time to half but is also an order of magnitude smaller than state-of-the-art Astec in terms of model size. Through a comprehensive empirical comparison, we show that not only can InceptionXML outperform existing approaches on benchmark datasets but also the transformer baselines requiring only 2% FLOPs. The code for InceptionXML is available at https://github.com/xmc-aalto.
Siddhant Kharbanda, Atmadeep Banerjee, Devaansh Gupta, Akash Palrecha, Rohit Babbar
SIGIR3