Pritika Ramu

dblp:348/5232 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
7since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Information extraction and text analysis · 50% Multi-agent systems · 14% Language models and text generation · 12%
Computer graphics and multimedia
1 paper
Visualization and visual analytics · 100%

Topics — the 5 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis
document understanding
1.222026
Doc2Chart: Intent-Driven Zero-Shot Chart Generation from Documents · EMNLP 2025
Decisive: Guiding User Decisions with Optimal Preference Elicitation from Unstructured Documents · ACL (1) 2026
Knowledge, reasoning and agents › Multi-agent systems › social choice › computational social choice
preference elicitation
1.012026
Decisive: Guiding User Decisions with Optimal Preference Elicitation from Unstructured Documents · ACL (1) 2026
Visualization and visual analytics › visualization generation › automated visualization generation
chart generation
0.912025
Doc2Chart: Intent-Driven Zero-Shot Chart Generation from Documents · EMNLP 2025
Natural language and speech › Information extraction and text analysis
fact decomposition
0.812024
Enhancing Post-Hoc Attributions in Long Document Comprehension via Coarse Grained Answer Decomposition · EMNLP 2024
Natural language and speech › Information extraction and text analysis › structure prediction
text-to-table generation
0.812024
Is This a Bad Table? A Closer Look at the Evaluation of Table Generation from Text · EMNLP 2024

Methods — techniques the papers use, named apart from their topics

large language model · 2.6retrieval-augmented generation · 1.7unstructured documents · 1.0preference elicitation · 1.0template-based decomposition · 0.8negative sampling · 0.8in-context learning · 0.8entailment · 0.8atomic statement decomposition · 0.8
YearPublicationVenuePosition
2026 Decisive: Guiding User Decisions with Optimal Preference Elicitation from Unstructured Documents
abstract
Akriti Jain, Anish Mulay, Divyansh Verma, Aishani Pandey, Pritika Ramu, Aparna Garimella. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Akriti Jain 0001, Anish Mulay, Divyansh Verma, Aishani Pandey, Pritika Ramu, Aparna Garimella
ACL (1)5
2025 Infogen: Generating Complex Statistical Infographics from Documents
abstract
Akash Ghosh, Aparna Garimella, Pritika Ramu, Sambaran Bandyopadhyay, Sriparna Saha. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Akash Ghosh, Aparna Garimella, Pritika Ramu, Sambaran Bandyopadhyay, Sriparna Saha 0001
ACL (1)3
2025 Doc2Chart: Intent-Driven Zero-Shot Chart Generation from Documents
abstract
Large Language Models (LLMs) have demonstrated strong capabilities in transforming text descriptions or tables to data visualizations via instruction-tuning methods.However, it is not straightforward to apply these methods directly for a more real-world use case of visualizing data from long documents based on user-given intents, as opposed to the user pre-selecting the relevant content manually.We introduce the task of intent-based chart generation from documents: given a user-specified intent and document(s), the goal is to generate a chart adhering to the intent and grounded on the document(s) in a zero-shot setting.We propose an unsupervised, two-staged framework in which an LLM first extracts relevant information from the document(s) by decomposing the intent and iteratively validates and refines this data.Next, a heuristic-guided module selects an appropriate chart type before final code generation.To assess the data accuracy of the generated charts, we propose an attribution-based metric that uses a structured textual representation of charts, instead of relying on visual decoding metrics that often fail to capture the chart data effectively.To validate our approach, we curate a dataset comprising of 1,242 tuples from two domains, finance and scientific, in contrast to the existing datasets that are largely limited to parallel text descriptions/ tables and their corresponding charts.We compare our approach with baselines using single-shot chart generation using LLMs and query-based retrieval methods; our method outperforms by upto 9 points and 17 points in terms of chart data accuracy and chart type respectively over the best baselines.
Akriti Jain 0001, Pritika Ramu, Aparna Garimella, Apoorv Saxena
EMNLP2
2024 Is This a Bad Table? A Closer Look at the Evaluation of Table Generation from Text
abstract
Understanding whether a generated table is of good quality is important to be able to use it in creating or editing documents using automatic methods.In this work, we underline that existing measures for table quality evaluation fail to capture the overall semantics of the tables, and sometimes unfairly penalize good tables and reward bad ones.We propose TABEVAL, a novel table evaluation strategy that captures table semantics by first breaking down a table into a list of natural language atomic statements and then compares them with ground truth statements using entailment-based measures.To validate our approach, we curate a dataset comprising of text descriptions for 1,250 diverse Wikipedia tables, covering a range of topics and structures, in contrast to the limited scope of existing datasets.We compare TABEVAL with existing metrics using unsupervised and supervised textto-table generation methods, demonstrating its stronger correlation with human judgments of table quality across four datasets.
Pritika Ramu, Aparna Garimella, Sambaran Bandyopadhyay
EMNLP1
2024 Enhancing Post-Hoc Attributions in Long Document Comprehension via Coarse Grained Answer Decomposition
abstract
Accurately attributing answer text to its source document is crucial for developing a reliable question-answering system.However, attribution for long documents remains largely unexplored.Post-hoc attribution systems are designed to map answer text back to the source document, yet the granularity of this mapping has not been addressed.Furthermore, a critical question arises: What exactly should be attributed?This involves identifying the specific information units within an answer that require grounding.In this paper, we propose and investigate a novel approach to the factual decomposition of generated answers for attribution, employing template-based in-context learning.To accomplish this, we utilize the question and integrate negative sampling during few-shot in-context learning for decomposition.This approach enhances the semantic understanding of both abstractive and extractive answers.We examine the impact of answer decomposition by providing a thorough examination of various attribution approaches, ranging from retrieval-based techniques to LLM-based attributors.
Pritika Ramu, Koustava Goswami, Apoorv Saxena, Balaji Vasan Srinivasan
EMNLP1
2024 Zooming in on Zero-Shot Intent-Guided and Grounded Document Generation using LLMs
abstract
Repurposing existing content on-the-fly to suit author's goals for creating initial drafts is crucial for document creation.We introduce the task of intent-guided and grounded document generation: given a user-specified intent (e.g., section title) and a few reference documents, the goal is to generate section-level multimodal documents spanning text and images, grounded on the given references, in a zero-shot setting.We present a data curation strategy to obtain general-domain samples from Wikipedia, and collect 1,000 Wikipedia sections consisting of textual and image content along with appropriate intent specifications and references.We propose a simple yet effective planningbased prompting strategy Multimodal Plan-And-Write (MM-PAW), to prompt LLMs to generate an intermediate plan with text and image descriptions, to guide the subsequent generation.We compare the performances of MM-PAW and a text-only variant of it with those of zero-shot Chain-of-Thought (CoT) using recent close and open-domain LLMs.Both of them lead to significantly better performances in terms of content relevance, structure, and groundedness to the references, more so in the smaller models (upto 12.5 points ↑ in Rouge 1-F1) than in the larger ones (upto 4 points ↑ R1-F1).They are particularly effective in improving relatively smaller models' performances, to be on par or higher than those of their larger counterparts for this task.
Pritika Ramu, Pranshu Gaur, Rishita Emandi, Himanshu Maheshwari, Danish Javed, Aparna Garimella
INLG1
2024 RE²: Region-Aware Relation Extraction from Visually Rich Documents
abstract
Pritika Ramu, Sijia Wang, Lalla Mouatadid, Joy Rimchala, Lifu Huang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Pritika Ramu, Lalla Mouatadid, Joy Rimchala, Lifu Huang
NAACL-HLT1