Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Varre Chaitanya

dblp:378/5664 · also Varre Suman Chaitanya · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 50% Visual content generation and editing · 50%
Artificial intelligence
1 paper
Language models and text generation · 100%

Topics — the 1 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › text generation
paraphrase generation
0.912025
Deep Submodular Optimization and LLM for Multimodal Content Extraction and Automatic Poster Generation from Long Document · AAAI 2025

Methods — techniques the papers use, named apart from their topics

large language model · 1.7deep submodular optimization · 1.7
YearPublicationVenuePosition
2025 Deep Submodular Optimization and LLM for Multimodal Content Extraction and Automatic Poster Generation from Long Document
abstract
A poster from a long input document can be considered as a one-page easy-to-read multimodal (text and images) summary presented on a nice template with good design elements. Automatic transformation of a long document into a poster is a very less studied but challenging task. It involves content summarization of the input document followed by template generation and harmonization. In this work, we propose a novel deep submodular function which can be trained on ground truth summaries to extract multimodal content from the document and explicitly ensures good coverage, diversity and alignment of text and images. Then, we use an LLM based paraphraser and propose to generate a template with various design aspects conditioned on the input content. We show the merits of our approach through extensive automated and human evaluations.
Vijay Jaisankar, Sambaran Bandyopadhyay, Kalp Vyas, Varre Chaitanya, Shwetha Somasundaram
AAAI4