EDBT 2026 Demo / reviewers in the wild / expert
Akshat Gupta
dblp:98/8230
· DBLP profile ↗
17ranked-venue papers
4as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OrthoEdit: Principled and Stable Knowledge Editing via Orthogonal Subspace ProjectionabstractAbstract Large language models (LLMs) encode extensive factual knowledge through pretraining, yet often require targeted updates to correct errors, incorporate new information, or revise outdated facts. Recent approaches to knowledge editing, such as projection-based constraints, parameter pruning, and regularization, have proven effective in improving editing accuracy and stability. However, these methods often fail to maintain a clear separation between new edits and existing knowledge, leading to interference and degradation over time. We propose OrthoEdit, a principled framework for stable and scalable knowledge editing that ensures each parameter update is orthogonal to both pre-existing and previously edited knowledge, remains strictly non-interfering and preserving the integrity of prior edits. OrthoEdit enables exact subspace control through three coordinated steps: progressive null space refinement, principal subspace extraction, and orthogonal projection. This yields compact and well-aligned updates that systematically satisfy all accumulated constraints. Comprehensive experiments across diverse models and benchmarks demonstrate that OrthoEdit consistently enhances editing accuracy and robustness while preserving general capabilities—even through extended sequences of batched edits. Our code is available at https://github.com/JoveReCode/OrthoEdit.git. Shanbao Qiao, Xuebing Liu, Akshat Gupta, Seung-Hoon Na |
Trans. Assoc. Comput. Linguistics | 3 |
| 2025 | PokerBench: Training Large Language Models to Become Professional Poker PlayersabstractWe introduce PokerBench - a benchmark for evaluating the poker-playing abilities of large language models (LLMs). As LLMs excel in traditional NLP tasks, their application to complex, strategic games like poker poses a new challenge. Poker, an incomplete information game, demands a multitude of skills such as mathematics, reasoning, planning, strategy, and a deep understanding of game theory and human psychology. This makes Poker the ideal next frontier for large language models. PokerBench consists of a comprehensive compilation of 11,000 most important scenarios, split between pre-flop and post-flop play, developed in collaboration with trained poker players. We evaluate prominent models including GPT-4, ChatGPT 3.5, and various Llama and Gemma series models, finding that all state-of-the-art LLMs underperform in playing optimal poker. However, after fine-tuning, these models show marked improvements. We validate PokerBench by having models with different scores compete with each other, demonstrating that higher scores on PokerBench leads to higher win rates in actual poker games. Through gameplay between our fine-tuned model and GPT-4, we also identify limitations of simple supervised fine-tuning for learning optimal playing strategy, suggesting the need for more advanced methodologies for effectively training language models to excel in games. PokerBench thus presents a unique benchmark for a quick and reliable evaluation of the poker-playing ability of LLMs as well as a comprehensive benchmark to study the progress of LLMs in complex game-playing scenarios. Richard Zhuang, Akshat Gupta, Richard Yang, Aniket Rahane, Gopala Krishna Anumanchipalli |
AAAI | 2 |
| 2025 | InterAct: Advancing Large-Scale Versatile 3D Human-Object Interaction GenerationabstractWhile large-scale human motion capture datasets have advanced human motion generation, modeling and generating dynamic 3D human-object interactions (HOIs) remain challenging due to dataset limitations. Existing datasets often lack extensive, high-quality motion and annotation and exhibit artifacts such as contact penetration, floating, and incorrect hand motions. To address these issues, we introduce InterAct, a large-scale 3D HOI benchmark featuring dataset and methodological advancements. First, we consolidate and standardize 21.81 hours of HOI data from diverse sources, enriching it with detailed textual annotations. Second, we propose a unified optimization framework to enhance data quality by reducing artifacts and correcting hand motions. Leveraging the principle of contact invariance, we maintain human-object relationships while introducing motion variations, expanding the dataset to 30.70 hours. Third, we define six benchmarking tasks and develop a unified HOI generative modeling perspective, achieving state-of-the-art performance. Extensive experiments validate the utility of our dataset as a foundational resource for advancing 3D human-object interaction generation. The dataset will be publicly accessible to support further research in the field. Sirui Xu 0002, Dongting Li 0001, Xiyan Xu, Qi Long, Ziyin Wang, Yunzhi Lu, Shuchang Dong, Hezi Jiang, Akshat Gupta, Yu-Xiong Wang, Liangyan Gui |
CVPR | 10 |
| 2025 | CardioRiskNet: Attention-based CVAE-enabled GCN for Risk Prediction in STEMIabstractCardiovascular diseases (CVDs) are a major cause of death worldwide, taking almost 18 million lives each year. ST Elevation Myocardial Infarction (STEMI) is one of the highest contributors to the same. The immediate 30-day period post-STEMI is critical in judging long-term patient outcomes. Thus, there is a need for an accurate risk predictor to guide clinical interventions immediately after STEMI. In this paper, we propose CardioRiskNet, a post-STEMI 30-day mortality predictor based on Graph Convolutional Networks, designed to adapt to different populations with the relational nature of graph-based models. To address class imbalance, we propose a data synthesis method for CVD data by introducing a self-attention mechanism in a Conditional Variational Autoencoder. To demonstrate robustness, the model has been tested on three datasets including two publicly available datasets. CardioRiskNet shows better performance compared to the state-of-the-art methods. Posthoc interpretability analysis also suggests that CardioRiskNet offers promising advancements in data-driven risk assessment, providing clinicians with a precise tool for patient management. Akshat Gupta, Anubha Gupta, Manu Kumar Shetty, Dixit Goyal, Girish M. P, Mohit D. Gupta |
ICASSP | 1 |
| 2025 | Sylber: Syllabic Embedding Representation of Speech from Raw AudioabstractSyllables are compositional units of spoken language that efficiently structure human speech perception and production. However, current neural speech representations lack such structure, resulting in dense token sequences that are costly to process. To bridge this gap, we propose a new model, Sylber, that produces speech representations with clean and robust syllabic structure. Specifically, we propose a self-supervised learning (SSL) framework that bootstraps syllabic embeddings by distilling from its own initial unsupervised syllabic segmentation. This results in a highly structured representation of speech features, offering three key benefits: 1) a fast, linear-time syllable segmentation algorithm, 2) efficient syllabic tokenization with an average of 4.27 tokens per second, and 3) novel phonological units suited for efficient spoken language modeling. Our proposed segmentation method is highly robust and generalizes to out-of-domain data and unseen languages without any tuning. By training token-to-speech generative models, fully intelligible speech can be reconstructed from Sylber tokens with a significantly lower bitrate than baseline SSL tokens. This suggests that our model effectively compresses speech into a compact sequence of tokens with minimal information loss. Lastly, we demonstrate that categorical perception—a linguistic phenomenon in speech perception—emerges naturally in Sylber, making the embedding space more categorical and sparse than previous speech features and thus supporting the high efficiency of our tokenization. Together, we present a novel SSL approach for representing speech as syllables, with significant potential for efficient speech tokenization and spoken language modeling. Cheol Jun Cho, Nicholas Lee, Akshat Gupta, Dhruv Agarwal 0005, Ethan Chen, Alan W. Black, Gopala Krishna Anumanchipalli |
ICLR | 3 |
| 2024 | Rebuilding ROME : Resolving Model Collapse during Sequential Model EditingabstractRecent work using Rank-One Model Editing (ROME), a popular model editing method, has shown that there are certain facts that the algorithm is unable to edit without breaking the model.Such edits have previously been called disabling edits (Gupta et al., 2024a).These disabling edits cause immediate model collapse and limits the use of ROME for sequential editing.In this paper, we show that disabling edits are an artifact of irregularities in the implementation of ROME.With this paper, we provide a more stable implementation ROME, which we call r-ROME and show that model collapse is no longer observed when making large scale sequential edits with r-ROME, while further improving generalization and locality of model editing compared to the original implementation of ROME. Akshat Gupta, Sidharth Baskaran, Gopala Krishna Anumanchipalli |
EMNLP | 1 |
| 2024 | Sampling Rate Adaptive Speaker Verification from Raw Waveforms
Vinayak Abrol, Anshul Thakur, Akshat Gupta, Xiaomo Liu, Sameena Shah |
ICPR (28) | 3 |
| 2024 | ActNeRF: Uncertainty-aware Active Learning of NeRF-based Object Models for Robot Manipulators using Visual and Re-orientation ActionsabstractManipulating unseen objects is challenging without a 3D representation, as objects generally have occluded surfaces. This requires physical interaction with objects to build their internal representations. This paper presents an approach that enables a robot to rapidly learn the complete 3D model of a given object for manipulation in unfamiliar orientations. We use an ensemble of partially constructed NeRF models to quantify model uncertainty to determine the next action (a visual or re-orientation action) by optimizing informativeness and feasibility. Further, our approach determines when and how to grasp and re-orient an object given its partial NeRF model and re-estimates the object pose to rectify misalignments introduced during the interaction. Experiments with a simulated Franka Emika Robot Manipulator operating in a tabletop environment with benchmark objects demonstrate an improvement of (i) 14% in visual reconstruction quality (PSNR), (ii) 20% in the geometric/depth reconstruction of the object surface (F-score) and (iii) 71% in the task success rate of manipulating objects a-priori unseen orientations/stable configurations in the scene; over current methods. The project page can be found at https://actnerf.github.io/ Saptarshi Dasgupta, Akshat Gupta, Shreshth Tuli, Rohan Paul |
IROS | 2 |
| 2023 | REFinD: Relation Extraction Financial DatasetabstractA number of datasets for Relation Extraction (RE) have been created to aide downstream tasks such as information retrieval, semantic search, question answering and textual entailment. However, these datasets fail to capture financial-domain specific challenges since most of these datasets are compiled using general knowledge sources such as Wikipedia, web-based text and news articles, hindering real-life progress and adoption within the financial world. To address this limitation, we propose REFinD, the first large-scale annotated dataset of relations, with ~29K instances and 22 relations amongst 8 types of entity pairs, generated entirely over financial documents. We also provide an empirical evaluation with various state-of-the-art models as benchmarks for the RE task and highlight the challenges posed by our dataset. We observed that various state-of-the-art deep learning models struggle with numeric inference, relational and directional ambiguity. To encourage further research in this direction, REFinD is available at https://www.jpmorgan.com/technology/artificial-intelligence/initiatives/refind-dataset/problem-motivation-outcome. Simerjot Kaur, Charese Smiley, Akshat Gupta, Joy Prakash Sain, Dongsheng Wang 0005, Suchetha Siddagangappa, Toyin Aguda, Sameena Shah |
SIGIR | 3 |
| 2023 | Quantifying and Leveraging User Fatigue for Interventions in Recommender SystemsabstractPredicting churn and designing intervention strategies are crucial for online platforms to maintain user engagement. We hypothesize that predicting churn, i.e. users leaving from the system without further return, is often a delayed act, and it might get too late for the system to intervene. We propose detecting early signs of users losing interest, allowing time for intervention, and introduce a new formulation ofuser fatigue as short-term dissatisfaction, providing early signals to predict long-term churn. We identify behavioral signals predicting fatigue and develop models for fatigue prediction. Furthermore, we leverage the predicted fatigue estimates to develop fatigue-aware ad-load balancing intervention strategy that reduces churn, improving short- and long-term user retention. Results from deployed recommendation system and multiple live A/B tests across over 80 million users generating over 200 million sessions highlight gains for user engagement and platform strategic metrics. Hitesh Sagtani, Madan Gopal Jhawar, Akshat Gupta, Rishabh Mehrotra |
SIGIR | 3 |
| 2023 | Knowledge Discovery from Unstructured Data in Financial Services (KDF) WorkshopabstractKnowledge discovery from unstructured data, including business documents, web content, and news articles, has been a key AI challenge for the financial services industry. Comprehending these corpora and discovering knowledge from them, which could be textual, tabular, or graphic, are the cornerstone of supporting business decisions in the financial services domain, where information retrieval and content analysis techniques are of fundamental importance. We propose a workshop on knowledge discovery from unstructured data in financial services at SIGIR 2023 to highlight the current and emerging opportunities, invite original research, and prompt success sharing between researchers. Sameena Shah, Xiaodan Zhu 0001, Wenhu Chen, Manling Li, Armineh Nourbakhsh, Xiaomo Liu, Charese Smiley, Yulong Pei, Akshat Gupta |
SIGIR | 10 |
| 2022 | Intent classification using pre-trained language agnostic embeddings for low resource languagesabstractBuilding Spoken Language Understanding (SLU) systems that do not rely on language specific Automatic Speech Recognition (ASR) is an important yet less explored problem in language processing.In this paper, we present a comparative study aimed at employing a pre-trained language agnostic acoustic model to perform SLU in low resource scenarios.Specifically, we use three different embedding settings extracted using Allosaurus, a pre-trained universal phone decoder: (1) Phonelabels (2) Panphone, and (3) Allo embeddings (proposed by us).These embeddings are then used in identifying the spoken intent.We perform experiments across three different languages: English, Sinhala, and Tamil each with different data sizes to simulate high, medium, and low resource scenarios.Our system improves on the state-of-the-art (SOTA) intent classification accuracy by absolute 2.11% for Sinhala and 7.00% for Tamil and achieves competitive results in English.Furthermore, we also present a quantitative analysis to show how the performance scales with the number of training examples. Hemant Yadav, Akshat Gupta, Sai Krishna Rallabandi, Alan W. Black, Rajiv Ratn Shah |
INTERSPEECH | 2 |
| 2022 | Multilingual Speech Emotion Recognition with Multi-Gating Mechanism and Neural Architecture SearchabstractSpeech emotion recognition (SER) classifies audio into emotion categories such as Happy, Angry, Fear, Disgust and Neutral. While Speech Emotion Recognition (SER) is a common application for popular languages, it continues to be a problem for low-resourced languages, i.e., languages with no pre-trained speech-to-text recognition models. This paper firstly proposes a language-specific model that extract emotional information from multiple pre-trained speech models, and then designs a multi-domain model that simultaneously performs SER for various languages. Our multi-domain model employs a multi-gating mechanism to generate unique weighted feature combination for each language, and also searches for specific neural network structure for each language through a neural architecture search module. In addition, we introduce a contrastive auxiliary loss to build more separable rep-resentations for audio data. Our experiments show that our model raises the state-of-the-art accuracy by 3% for German and 14.3% for French. Zihan Wang 0006, HaiFeng Lan, KeHao Guo, Akshat Gupta |
SLT | 6 |
| 2021 | Intent Recognition and Unsupervised Slot Identification for Low-Resourced Spoken Dialog SystemsabstractIntent Recognition and Slot Identification are crucial components in spoken language understanding (SLU) systems. In this paper, we present a novel approach towards both these tasks in the context of low-resourced and unwritten languages. We use an acoustic based SLU system that converts speech to its phonetic transcription using a universal phone recognition system. We build a word-free natural language understanding module that does intent recognition and slot identification from these phonetic transcription. Our proposed SLU system performs competitively for resource rich scenarios and significantly outperforms existing approaches as the amount of available data reduces. We train both recurrent and transformer based neural networks and test our system on five natural speech datasets in five different languages. We observe more than 10% improvement for intent classification in Tamil and more than 5% improvement for intent classification in Sinhala. Additionally, we present a novel approach towards unsupervised slot identification using normalized attention scores. This approach can be used for unsupervised slot labelling, data augmentation and to generate data for a new slot in a one-shot way with only one speech recording. Akshat Gupta, Olivia Deng, Akruti Kushwaha, Saloni Mittal, William Zeng, Sai Krishna Rallabandi, Alan W. Black |
ASRU | 1 |
| 2021 | Acoustics Based Intent Recognition Using Discovered Phonetic Units for Low Resource LanguagesabstractWith recent advancements in language technologies, humans are now speaking to devices. Increasing the reach of spoken language technologies requires building systems in local languages. A major bottleneck here are the underlying data-intensive parts that make up such systems, including automatic speech recognition (ASR) systems that require large amounts of labelled data. With the aim of aiding development of spoken dialog systems in low resourced languages, we propose a novel acoustics based intent recognition system that uses discovered phonetic units for intent classification. The system is made up of two blocks - the first block is a universal phone recognition system that generates a transcript of discovered phonetic units for the input audio, and the second block performs intent classification from the generated phonetic transcripts. We propose a CNN+LSTM based architecture and present results for two languages families - Indic languages and Romance languages, for two different intent recognition tasks. We also perform multilingual training of our intent classifier and show improved cross-lingual transfer and zero-shot performance on an unknown language within the same language family. Akshat Gupta, Sai Krishna Rallabandi, Alan W. Black |
ICASSP | 1 |
| 2021 | Cuttlefish: library for achieving energy efficiency in multicore parallel programsabstractA low-cap power budget is challenging for exascale computing. Dynamic Voltage and Frequency Scaling (DVFS) and Uncore Frequency Scaling (UFS) are the two widely used techniques for limiting the HPC application's energy footprint. However, existing approaches fail to provide a unified solution that can work with different types of parallel programming models and applications. Sunil Kumar 0001, Akshat Gupta, Vivek Kumar 0001, Sridutt Bhalachandra |
SC | 2 |
| 2020 | Multimodal Word Sense Disambiguation in Creative PracticeabstractLanguage is ambiguous; many terms and expressions can convey the same idea. This is especially true in creative practice, where ideas and design intents are highly subjective. We present a dataset-Ambiguous Descriptions of Art Images (ADARI)-of contemporary workpieces, which aims to provide a foundational resource for subjective image description and multimodal word disambiguation in the context of creative practice. The dataset contains a total of 240k images labeled with 260k descriptive sentences. It is additionally organized into sub-domains of architecture, art, design, fashion, furniture, product design and technology. In subjective image description, labels do not necessarily correspond to well-defined entities i.e. cars, quantitative attributes such as the color red, or actions like playing. For example, the ambiguous label dynamic is a qualitative attribute of an extensive amount of objects and thus, the data's variance is high. To understand this complexity, we analyze the ambiguity and relevance of text with respect to images using the state-of-the-art pre-trained BERT model for sentence classification. We provide a baseline for multi-label classification tasks and demonstrate the potential of multimodal approaches for understanding ambiguity in design intentions. We hope that ADARI dataset and baselines constitute a first step towards subjective label classification. Manuel Ladron de Guevara, Christopher George, Akshat Gupta, Daragh Byrne, Ramesh Krishnamurti |
ICMLA | 3 |