Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Abhay Puri

dblp:383/3753 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
6since 2021 · last 2025
0009-0003-9370-0814ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Vision and language · 57% Generative modeling · 13% Language models and text generation · 13%
Computer graphics and multimedia
3 papers
Visual content generation and editing · 100%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 100%

Topics — the 11 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Visual content generation and editing
vector graphics generation
2.632025
Rendering-Aware Reinforcement Learning for Vector Graphics Generation · NeurIPS 2025
StarVector: Generating Scalable Vector Graphics Code from Images and Text · CVPR 2025
StarVector: Generating Scalable Vector Graphics Code from Images and Text · AAAI 2025
Computer vision › Vision and language › multimodal understanding
multimodal document understanding
1.722025
AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding · NeurIPS 2025
BigDocs: An Open Dataset for Training Multimodal Models on Document and Code Tasks · ICLR 2025
Visual content generation and editing › vector graphics generation
SVG generation
1.722025
Rendering-Aware Reinforcement Learning for Vector Graphics Generation · NeurIPS 2025
StarVector: Generating Scalable Vector Graphics Code from Images and Text · CVPR 2025
Natural language and speech › Language models and text generation
LLM agents
0.912025
InsightBench: Evaluating Business Analytics Agents Through Multi-Step Insight Generation · ICLR 2025
Machine learning › Generative modeling
multimodal generation
0.912025
StarVector: Generating Scalable Vector Graphics Code from Images and Text · CVPR 2025
Computer vision › Vision and language › vision-language model
vision-language model alignment
0.912025
AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding · NeurIPS 2025
Information retrieval › evaluation
benchmark
0.912025
InsightBench: Evaluating Business Analytics Agents Through Multi-Step Insight Generation · ICLR 2025
Information retrieval
evaluation
0.912025
InsightBench: Evaluating Business Analytics Agents Through Multi-Step Insight Generation · ICLR 2025
Program synthesis and code generation › code generation with language models
image-to-code generation
0.912025
BigDocs: An Open Dataset for Training Multimodal Models on Document and Code Tasks · ICLR 2025
Computer vision › Vision and language
vision-language model
0.312025
Rendering-Aware Reinforcement Learning for Vector Graphics Generation · NeurIPS 2025
Visual content generation and editing
diagram generation
0.312025
StarVector: Generating Scalable Vector Graphics Code from Images and Text · CVPR 2025

Methods — techniques the papers use, named apart from their topics

vision-language transformer · 1.7multimodal large language model · 1.7dataset curation · 1.7benchmark construction · 1.7autoregressive decoding · 1.7LLaMA-3 · 1.7LLM-based evaluation · 1.7vision-language model · 0.9vision encoder · 0.9reinforcement learning · 0.9differentiable rendering · 0.9MLP connector · 0.9LLM text embeddings · 0.9
YearPublicationVenuePosition
2025 StarVector: Generating Scalable Vector Graphics Code from Images and Text
abstract
Scalable Vector Graphics (SVG) have become integral to modern image rendering applications due to their infinite scalability and versatility, especially in graphic design and web development. SVGs are essentially long strings of code that adhere to a structured syntax with validity constraints. With the rise of large language models, which excel at generating code in various languages, we aim to generate SVG code in a similar way. Our findings show that a vision-language model can be conditioned to produce valid SVG code that closely resembles input images, effectively enabling vectorization. Additionally, we harness the rich SVG syntax, encompassing all possible primitives—such as lines, paths, polygons, text, and effects like color gradients—that previous methods often missed. We briefly explain how the StarVector model operates, primarily leveraging a vision-language transformer architecture to generate SVG code. We also detail our training and inference procedures. Finally, we provide an interactive demo that allows users to input an image and generate its SVG code autoregressively, featuring real-time rendering that visually demonstrates the SVG generation process.
Juan A. Rodríguez, Abhay Puri, Issam H. Laradji, Sai Rajeswar, David Vázquez 0001, Christopher Joseph Pal, Marco Pedersoli
AAAI2
2025 StarVector: Generating Scalable Vector Graphics Code from Images and Text
abstract
Scalable Vector Graphics (SVGs) are vital for modern image rendering due to their scalability and versatility. Previous SVG generation methods have focused on curve-based vectorization, lacking semantic understanding, often producing artifacts, and struggling with SVG primitives beyond path curves. To address these issues, we introduce StarVector, a multimodal large language model for SVG generation. It performs image vectorization by understanding image semantics and using SVG primitives for compact, precise outputs. Unlike traditional methods, StarVector works directly in the SVG code space, leveraging visual understanding to apply accurate SVG primitives. To train StarVector, we create SVG-Stack, a diverse dataset of 2M samples that enables generalization across vectorization tasks and precise use of primitives like ellipses, polygons, and text. We address challenges in SVG evaluation, showing that pixel-based metrics like MSE fail to capture the unique qualities of vector graphics. We introduce SVG-Bench, a benchmark across 10 datasets, and 3 tasks: Image-to-SVG, Text-to-SVG generation, and diagram generation. Using this setup, StarVector achieves state-of-the-art performance, producing more compact and semantically rich SVGs.
Juan A. Rodríguez, Abhay Puri, Issam H. Laradji, Pau Rodríguez, Sai Rajeswar, David Vázquez 0001, Christopher Joseph Pal, Marco Pedersoli
CVPR2
2025 BigDocs: An Open Dataset for Training Multimodal Models on Document and Code Tasks
abstract
Multimodal AI has the potential to significantly enhance document-understanding tasks, such as processing receipts, understanding workflows, extracting data from documents, and summarizing reports. Code generation tasks that require long-structured outputs can also be enhanced by multimodality. Despite this, their use in commercial applications is often limited due to limited access to relevant training data and restrictive licensing, which hinders open access. To address these limitations, we introduce BigDocs-7.5M, a high-quality, open-access dataset comprising 7.5 million multimodal documents across 30 tasks. We use an efficient data curation process to ensure that our data is high quality and license-permissive. Our process emphasizes accountability, responsibility, and transparency through filtering rules, traceable metadata, and careful content analysis. Additionally, we introduce BigDocs-Bench,, a benchmark suite with 10 novel tasks where we carefully create datasets that reflect real-world use cases involving reasoning over Graphical User Interfaces (GUI) and code generation from images. Our experiments show that training with BigDocs-Bench, improves average performance up to 25.8% over closed-source GPT-4o in document reasoning and structured output tasks such as Screenshot2HTML or Image2Latex generation. Finally, human evaluations revealed that participants preferred the outputs from models trained with BigDocs over those from GPT-4o. This suggests that BigDocs can help both academics and the open-source community utilize and improve AI tools to enhance multimodal capabilities and document reasoning.
Juan A. Rodríguez, Xiangru Jian, Siba Smarak Panigrahi, Aarash Feizi, Abhay Puri, Akshay Kalkunte Suresh, François Savard, Ahmed Masry, Shravan Nayak, Rabiul Awal, Mahsa Massoud, Amirhossein Abaskohi, Suyuchen Wang, Pierre-André Noël, Mats Leon Richter, Saverio Vadacchino, Sanket Biswas
ICLR6
2025 InsightBench: Evaluating Business Analytics Agents Through Multi-Step Insight Generation
abstract
Data analytics is essential for extracting valuable insights from data that can assist organizations in making effective decisions. We introduce InsightBench, a benchmark dataset with three key features. First, it consists of 100 datasets representing diverse business use cases such as finance and incident management, each accompanied by a carefully curated set of insights planted in the datasets. Second, unlike existing benchmarks focusing on answering single queries, InsightBench evaluates agents based on their ability to perform end-to-end data analytics, including formulating questions, interpreting answers, and generating a summary of insights and actionable steps. Third, we conducted comprehensive quality assurance to ensure that each dataset in the benchmark had clear goals and included relevant and meaningful questions and analysis. Furthermore, we implement a two-way evaluation mechanism using LLaMA-3 as an effective, open-source evaluator to assess agents’ ability to extract insights. We also propose AgentPoirot, our baseline data analysis agent capable of performing end-to-end data analytics. Our evaluation on InsightBench shows that AgentPoirot outperforms existing approaches (such as Pandas Agent) that focus on resolving single queries. We also compare the performance of open- and closed-source LLMs and various evaluation strategies. Overall, this benchmark serves as a testbed to motivate further development in comprehensive automated data analytics and can be accessed here: https://github.com/ServiceNow/insight-bench.
Gaurav Sahu, Abhay Puri, Juan A. Rodríguez, Amirhossein Abaskohi, Mohammad Chegini, Alexandre Drouin, Perouz Taslakian, Valentina Zantedeschi, Alexandre Lacoste, David Vázquez 0001, Nicolas Chapados, Christopher Joseph Pal, Sai Rajeswar, Issam H. Laradji
ICLR2
2025 AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding
abstract
Aligning visual features with language embeddings is a key challenge in vision-language models (VLMs). The performance of such models hinges on having a good connector that maps visual features generated by a vision encoder to a shared embedding space with the LLM while preserving semantic similarity. Existing connectors, such as multilayer perceptrons (MLPs), lack inductive bias to constrain visual features within the linguistic structure of the LLM’s embedding space, making them data-hungry and prone to cross-modal misalignment. In this work, we propose a novel vision-text alignment method, AlignVLM, that maps visual features to a weighted average of LLM text embeddings. Our approach leverages the linguistic priors encoded by the LLM to ensure that visual features are mapped to regions of the space that the LLM can effectively interpret. AlignVLM is particularly effective for document understanding tasks, where visual and textual modalities are highly correlated. Our extensive experiments show that AlignVLM achieves state-of-the-art performance compared to prior alignment methods, with larger gains on document understanding and under low-resource setups. We provide further analysis demonstrating its efficiency and robustness to noise.
Ahmed Masry, Juan A. Rodríguez, Suyuchen Wang, Aarash Feizi, Akshay Kalkunte Suresh, Abhay Puri, Xiangru Jian, Pierre-André Noël, Sathwik Tejaswi Madhusudhan, Marco Pedersoli, Bang Liu 0003, Nicolas Chapados, Yoshua Bengio, Enamul Hoque Prince, Christopher Joseph Pal, Issam H. Laradji, David Vázquez 0001, Perouz Taslakian, Spandana Gella, Sai Rajeswar
NeurIPS8
2025 Rendering-Aware Reinforcement Learning for Vector Graphics Generation
abstract
Scalable Vector Graphics (SVG) offer a powerful format for representing visual designs as interpretable code. Recent advances in vision-language models (VLMs) have enabled high-quality SVG generation by framing the problem as a code generation task and leveraging large-scale pretraining. VLMs are particularly suitable for this task as they capture both global semantics and fine-grained visual patterns, while transferring knowledge across vision, natural language, and code domains. However, existing VLM approaches often struggle to produce faithful and efficient SVGs because they never observe the rendered images during training. Although differentiable rendering for autoregressive SVG code generation remains unavailable, rendered outputs can still be compared to original inputs, enabling evaluative feedback suitable for reinforcement learning (RL). We introduce Reinforcement Learning from Rendering Feedback, an RL method that enhances SVG generation in autoregressive VLMs by leveraging feedback from rendered SVG outputs. Given an input image, the model generates SVG roll-outs that are rendered and compared to the original image to compute a reward. This visual fidelity feedback guides the model toward producing more accurate, efficient, and semantically coherent SVGs. \method significantly outperforms supervised fine-tuning, addressing common failure modes and enabling precise, high-quality SVG generation with strong structural understanding and generalization.
Juan A. Rodríguez, Abhay Puri, Rishav Pramanik, Aarash Feizi, Pascal Wichmann, Arnab Kumar Mondal, Mohammad Reza Samsami, Rabiul Awal, Perouz Taslakian, Spandana Gella, Sai Rajeswar, David Vázquez 0001, Christopher Joseph Pal, Marco Pedersoli
NeurIPS3