Akshara Prabhakar

dblp:307/3028 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Language models and text generation · 25% Information extraction and text analysis · 15% Question answering and dialogue systems · 13%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 100%

Topics — the 12 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › domain model learning
action model learning
0.912025
ActionStudio: A Lightweight Framework for Data and Training of Large Action Models · EMNLP 2025
Natural language and speech › Language models and text generation
agentic language model
0.912025
APIGen-MT: Agentic Pipeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay · NeurIPS 2025
Natural language and speech › Language models and text generation
large language model training
0.912025
APIGen-MT: Agentic Pipeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay · NeurIPS 2025
Machine learning › Generative modeling
synthetic training data
0.912025
APIGen-MT: Agentic Pipeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay · NeurIPS 2025
Natural language and speech › Question answering and dialogue systems › domain-specific question answering
science question answering
0.812024
Language Models as Science Tutors · ICML 2024
Machine learning › Reinforcement learning › reinforcement learning environment
environment design
0.712023
InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback · NeurIPS 2023
Program synthesis and code generation › interactive program synthesis
interactive code generation
0.712023
InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback · NeurIPS 2023
Natural language and speech › Machine translation
annotation projection
0.612022
CL-NERIL: A Cross-Lingual Model for NER in Indian Languages (Student Abstract) · AAAI 2022
Natural language and speech › Information extraction and text analysis › named entity recognition
cross-lingual named entity recognition
0.612022
CL-NERIL: A Cross-Lingual Model for NER in Indian Languages (Student Abstract) · AAAI 2022
Machine learning › Transfer learning and domain adaptation › domain adaptation
low-resource domain adaptation
0.612022
CL-NERIL: A Cross-Lingual Model for NER in Indian Languages (Student Abstract) · AAAI 2022
Natural language and speech › Information extraction and text analysis
named entity recognition
0.612022
CL-NERIL: A Cross-Lingual Model for NER in Indian Languages (Student Abstract) · AAAI 2022
Natural language and speech › Language models and text generation
code generation
0.212023
InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

long-context language models · 1.5fine-tuning · 1.5reinforcement learning · 1.3prompting · 1.3training pipeline · 0.9simulated agent-human interplay · 0.9data curation · 0.9LLM-based data generation · 0.9weak supervision · 0.6teacher-student model · 0.6
YearPublicationVenuePosition
2025 ActionStudio: A Lightweight Framework for Data and Training of Large Action Models
abstract
Jianguo Zhang, Thai Quoc Hoang, Ming Zhu, Zuxin Liu, Shiyu Wang, Tulika Manoj Awalgaonkar, Akshara Prabhakar, Haolin Chen, Weiran Yao, Zhiwei Liu, Juntao Tan, Juan Carlos Niebles, Shelby Heinecke, Huan Wang, Silvio Savarese, Caiming Xiong. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Jianguo Zhang 0005, Thai Hoang, Zuxin Liu, Tulika Manoj Awalgaonkar, Akshara Prabhakar, Weiran Yao, Zhiwei Liu 0001, Juntao Tan, Juan Carlos Niebles, Shelby Heinecke, Huan Wang 0016, Silvio Savarese, Caiming Xiong
EMNLP7
2025 CRMArena: Understanding the Capacity of LLM Agents to Perform Professional CRM Tasks in Realistic Environments
abstract
Kung-Hsiang Huang, Akshara Prabhakar, Sidharth Dhawan, Yixin Mao, Huan Wang, Silvio Savarese, Caiming Xiong, Philippe Laban, Chien-Sheng Wu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Kung-Hsiang Huang, Akshara Prabhakar, Sidharth Dhawan, Yixin Mao, Huan Wang 0016, Silvio Savarese, Caiming Xiong, Philippe Laban, Chien-Sheng Wu
NAACL (Long Papers)2
2025 xLAM: A Family of Large Action Models to Empower AI Agent Systems
abstract
Jianguo Zhang, Tian Lan, Ming Zhu, Zuxin Liu, Thai Quoc Hoang, Shirley Kokane, Weiran Yao, Juntao Tan, Akshara Prabhakar, Haolin Chen, Zhiwei Liu, Yihao Feng, Tulika Manoj Awalgaonkar, Rithesh R N, Zeyuan Chen, Ran Xu, Juan Carlos Niebles, Shelby Heinecke, Huan Wang, Silvio Savarese, Caiming Xiong. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Jianguo Zhang 0005, Tian Lan 0006, Zuxin Liu, Thai Hoang, Shirley Kokane, Weiran Yao, Juntao Tan, Akshara Prabhakar, Zhiwei Liu 0001, Yihao Feng, Tulika Manoj Awalgaonkar, Rithesh R. N., Zeyuan Chen 0001, Ran Xu 0001, Juan Carlos Niebles, Shelby Heinecke, Huan Wang 0016, Silvio Savarese, Caiming Xiong
NAACL (Long Papers)9
2025 APIGen-MT: Agentic Pipeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay
abstract
Training effective AI agents for multi-turn interactions requires high-quality data that captures realistic human-agent dynamics, yet such data is scarce and expensive to collect manually. We introduce APIGen-MT, a two-phase framework that generates verifiable and diverse multi-turn agent data. In the first phase, our agentic pipeline produces detailed task blueprints with ground-truth actions, leveraging a committee of LLM reviewers and iterative feedback loops. These blueprints are then transformed into complete interaction trajectories through simulated human-agent interplay. We train a family of models---the xLAM-2-fc-r series with sizes ranging from 1B to 70B parameters. Our models outperform frontier models such as GPT-4o and Claude 3.5 on $\tau$-bench and BFCL benchmarks, with the smaller models surpassing their larger counterparts, particularly in multi-turn settings, while maintaining superior consistency across multiple trials. Comprehensive experiments demonstrate that our verified blueprint-to-details approach yields high-quality training data, enabling the development of more reliable, efficient, and capable agents. We open-source both the synthetic data collected and the trained xLAM-2-fc-r models to advance research in AI agents.Dataset: https://huggingface.co/datasets/Salesforce/APIGen-MT-5k & Models: https://huggingface.co/collections/Salesforce/xlam-2-67ef5be12949d8dcdae354c4
Akshara Prabhakar, Zuxin Liu, Tulika Manoj Awalgaonkar, Zhiwei Liu 0001, Thai Hoang, Juan Carlos Niebles, Shelby Heinecke, Weiran Yao, Huan Wang 0016, Silvio Savarese, Caiming Xiong
NeurIPS1
2024 Language Models as Science Tutors
abstract
NLP has recently made exciting progress toward training language models (LMs) with strong scientific problem-solving skills. However, model development has not focused on real-life use-cases of LMs for science, including applications in education that require processing long scientific documents. To address this, we introduce TutorEval and TutorChat. TutorEval is a diverse question-answering benchmark consisting of questions about long chapters from STEM textbooks, written by experts. TutorEval helps measure real-life usability of LMs as scientific assistants, and it is the first benchmark combining long contexts, free-form generation, and multi-disciplinary scientific knowledge. Moreover, we show that fine-tuning base models with existing dialogue datasets leads to poor performance on TutorEval. Therefore, we create TutorChat, a dataset of 80,000 long synthetic dialogues about textbooks. We use TutorChat to fine-tune Llemma models with 7B and 34B parameters. These LM tutors specialized in math have a 32K-token context window, and they excel at TutorEval while performing strongly on GSM8K and MATH. Our datasets build on open-source materials, and we release our models, data, and evaluations publicly.
Alexis Chevalier, Jiayi Geng, Alexander Wettig, Howard Chen 0003, Sebastian Mizera, Toni Annala, Max Jameson Aragon, Arturo Rodríguez Fanlo, Simon Frieder, Simon Machado, Akshara Prabhakar, Ellie Thieu, Jiachen T. Wang, Xindi Wu, Mengzhou Xia, Wenhan Xia, Jiatong Yu, Zhiyong Jason Ren, Sanjeev Arora, Danqi Chen 0001
ICML11
2023 InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback
abstract
Humans write code in a fundamentally interactive manner and rely on constant execution feedback to correct errors, resolve ambiguities, and decompose tasks. While LLMs have recently exhibited promising coding capabilities, current coding benchmarks mostly consider a static instruction-to-code sequence transduction process, which has the potential for error propagation and a disconnect between the generated code and its final execution environment. To address this gap, we introduce InterCode, a lightweight, flexible, and easy-to-use framework of interactive coding as a standard reinforcement learning (RL) environment, with code as actions and execution feedback as observations. Our framework is language and platform agnostic, uses self-contained Docker environments to provide safe and reproducible execution, and is compatible out-of-the-box with traditional seq2seq coding methods, while enabling the development of new methods for interactive code generation. We use InterCode to create three interactive code environments with Bash, SQL, and Python as action spaces, leveraging data from the static NL2Bash, Spider, and MBPP datasets. We demonstrate InterCode’s viability as a testbed by evaluating multiple state-of-the-art LLMs configured with different prompting strategies such as ReAct and Plan & Solve. Our results showcase the benefits of interactive code generation and demonstrate that InterCode can serve as a challenging benchmark for advancing code understanding and generation capabilities. InterCode is designed to be easily extensible and can even be used to create new tasks such as Capture the Flag, a popular coding puzzle that is inherently multi-step and involves multiple programming languages.
John Yang 0002, Akshara Prabhakar, Karthik Narasimhan, Shunyu Yao 0006
NeurIPS2
2022 CL-NERIL: A Cross-Lingual Model for NER in Indian Languages (Student Abstract)
abstract
Developing Named Entity Recognition (NER) systems for Indian languages has been a long-standing challenge, mainly owing to the requirement of a large amount of annotated clean training instances. This paper proposes an end-to-end framework for NER for Indian languages in a low-resource setting by exploiting parallel corpora of English and Indian languages and an English NER dataset. The proposed framework includes an annotation projection method that combines word alignment score and NER tag prediction confidence score on source language (English) data to generate weakly labeled data in a target Indian language. We employ a variant of the Teacher-Student model and optimize it jointly on the pseudo labels of the Teacher model and predictions on the generated weakly labeled data. We also present manually annotated test sets for three Indian languages: Hindi, Bengali, and Gujarati. We evaluate the performance of the proposed framework on the test sets of the three Indian languages. Empirical results show a minimum 10% performance improvement compared to the zero-shot transfer learning model on all languages. This indicates that weakly labeled data generated using the proposed annotation projection method in target Indian languages can complement well-annotated source language data to enhance performance. Our code is publicly available at https://github.com/aksh555/CL-NERIL.
Akshara Prabhakar, Gouri Sankar Majumder, Ashish Anand
AAAI1
2022 Commonsense and Named Entity Aware Knowledge Grounded Dialogue Generation
abstract
Grounding dialogue on external knowledge and interpreting linguistic patterns in dialogue history context, such as ellipsis, anaphora, and co-references is critical for dialogue comprehension and generation.In this paper, we present a novel open-domain dialogue generation model which effectively utilizes the large-scale commonsense and named entity based knowledge in addition to the unstructured topic-specific knowledge associated with each utterance.We enhance the commonsense knowledge with named entity-aware structures using co-references.Our proposed model utilizes a multi-hop attention layer to preserve the most accurate and critical parts of the dialogue history and the associated knowledge.In addition, we employ a Commonsense and Named Entity Enhanced Attention Module, which starts with the extracted triples from various sources and gradually finds the relevant supporting set of triples using multi-hop attention with the query vector obtained from the interactive dialogue-knowledge module.Empirical results on two benchmark dataset demonstrate that our model significantly outperforms the state-of-the-art methods in terms of both automatic evaluation metrics and human judgment.Our code is publicly available at https://github.com/deekshaVarshney/CNTF;https://www.iitp.ac.in/-ai-nlp-ml/resources/ codes/CNTF.zip.
Deeksha Varshney, Akshara Prabhakar, Asif Ekbal
NAACL-HLT2