Aarash Feizi

dblp:275/3823 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Vision and language · 73% Reinforcement learning · 17% Representation and self-supervised learning · 5%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%
Software engineering, system software, and programming languages
2 papers
Compilers and program optimization · 50% Program synthesis and code generation · 50%
Databases, data mining, and information retrieval
1 paper
Knowledge graphs · 100%

Topics — the 11 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › multimodal understanding
multimodal document understanding
1.722025
AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding · NeurIPS 2025
BigDocs: An Open Dataset for Training Multimodal Models on Document and Code Tasks · ICLR 2025
Computer vision › Vision and language › multimodal understanding
multimodal web understanding
0.912025
WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation · EMNLP 2025
Computer vision › Vision and language › vision-language model
vision-language model alignment
0.912025
AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding · NeurIPS 2025
Visual content generation and editing › vector graphics generation
SVG generation
0.912025
Rendering-Aware Reinforcement Learning for Vector Graphics Generation · NeurIPS 2025
Visual content generation and editing
vector graphics generation
0.912025
Rendering-Aware Reinforcement Learning for Vector Graphics Generation · NeurIPS 2025
Compilers and program optimization
code generation
0.912025
WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation · EMNLP 2025
Program synthesis and code generation › code generation with language models
image-to-code generation
0.912025
BigDocs: An Open Dataset for Training Multimodal Models on Document and Code Tasks · ICLR 2025
Knowledge graphs
knowledge graph embedding
0.412020
Structure Aware Negative Sampling in Knowledge Graphs · EMNLP (1) 2020
Computer vision › Vision and language
vision-language model
0.312025
Rendering-Aware Reinforcement Learning for Vector Graphics Generation · NeurIPS 2025
Machine learning › Representation and self-supervised learning
contrastive learning
0.112020
Structure Aware Negative Sampling in Knowledge Graphs · EMNLP (1) 2020
Machine learning › Representation and self-supervised learning › contrastive learning
negative sampling
0.112020
Structure Aware Negative Sampling in Knowledge Graphs · EMNLP (1) 2020

Methods — techniques the papers use, named apart from their topics

vision-language model · 1.7reinforcement learning · 1.7multimodal large language model · 1.7differentiable rendering · 1.7dataset curation · 1.7benchmark construction · 1.7vision encoder · 0.9MLP connector · 0.9LLM text embeddings · 0.9contrastive estimation · 0.9k-hop neighborhood sampling · 0.4
YearPublicationVenuePosition
2025 WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation
abstract
Rabiul Awal, Mahsa Massoud, Aarash Feizi, Zichao Li, Suyuchen Wang, Christopher Pal, Aishwarya Agrawal, David Vazquez, Siva Reddy, Juan A. Rodriguez, Perouz Taslakian, Spandana Gella, Sai Rajeswar. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Rabiul Awal, Mahsa Massoud, Aarash Feizi, Suyuchen Wang, Christopher Joseph Pal, Aishwarya Agrawal, David Vázquez 0001, Siva Reddy, Juan A. Rodríguez, Perouz Taslakian, Spandana Gella, Sai Rajeswar
EMNLP3
2025 BigDocs: An Open Dataset for Training Multimodal Models on Document and Code Tasks
abstract
Multimodal AI has the potential to significantly enhance document-understanding tasks, such as processing receipts, understanding workflows, extracting data from documents, and summarizing reports. Code generation tasks that require long-structured outputs can also be enhanced by multimodality. Despite this, their use in commercial applications is often limited due to limited access to relevant training data and restrictive licensing, which hinders open access. To address these limitations, we introduce BigDocs-7.5M, a high-quality, open-access dataset comprising 7.5 million multimodal documents across 30 tasks. We use an efficient data curation process to ensure that our data is high quality and license-permissive. Our process emphasizes accountability, responsibility, and transparency through filtering rules, traceable metadata, and careful content analysis. Additionally, we introduce BigDocs-Bench,, a benchmark suite with 10 novel tasks where we carefully create datasets that reflect real-world use cases involving reasoning over Graphical User Interfaces (GUI) and code generation from images. Our experiments show that training with BigDocs-Bench, improves average performance up to 25.8% over closed-source GPT-4o in document reasoning and structured output tasks such as Screenshot2HTML or Image2Latex generation. Finally, human evaluations revealed that participants preferred the outputs from models trained with BigDocs over those from GPT-4o. This suggests that BigDocs can help both academics and the open-source community utilize and improve AI tools to enhance multimodal capabilities and document reasoning.
Juan A. Rodríguez, Xiangru Jian, Siba Smarak Panigrahi, Aarash Feizi, Abhay Puri, Akshay Kalkunte Suresh, François Savard, Ahmed Masry, Shravan Nayak, Rabiul Awal, Mahsa Massoud, Amirhossein Abaskohi, Suyuchen Wang, Pierre-André Noël, Mats Leon Richter, Saverio Vadacchino, Sanket Biswas
ICLR5
2025 AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding
abstract
Aligning visual features with language embeddings is a key challenge in vision-language models (VLMs). The performance of such models hinges on having a good connector that maps visual features generated by a vision encoder to a shared embedding space with the LLM while preserving semantic similarity. Existing connectors, such as multilayer perceptrons (MLPs), lack inductive bias to constrain visual features within the linguistic structure of the LLM’s embedding space, making them data-hungry and prone to cross-modal misalignment. In this work, we propose a novel vision-text alignment method, AlignVLM, that maps visual features to a weighted average of LLM text embeddings. Our approach leverages the linguistic priors encoded by the LLM to ensure that visual features are mapped to regions of the space that the LLM can effectively interpret. AlignVLM is particularly effective for document understanding tasks, where visual and textual modalities are highly correlated. Our extensive experiments show that AlignVLM achieves state-of-the-art performance compared to prior alignment methods, with larger gains on document understanding and under low-resource setups. We provide further analysis demonstrating its efficiency and robustness to noise.
Ahmed Masry, Juan A. Rodríguez, Suyuchen Wang, Aarash Feizi, Akshay Kalkunte Suresh, Abhay Puri, Xiangru Jian, Pierre-André Noël, Sathwik Tejaswi Madhusudhan, Marco Pedersoli, Bang Liu 0003, Nicolas Chapados, Yoshua Bengio, Enamul Hoque Prince, Christopher Joseph Pal, Issam H. Laradji, David Vázquez 0001, Perouz Taslakian, Spandana Gella, Sai Rajeswar
NeurIPS6
2025 Rendering-Aware Reinforcement Learning for Vector Graphics Generation
abstract
Scalable Vector Graphics (SVG) offer a powerful format for representing visual designs as interpretable code. Recent advances in vision-language models (VLMs) have enabled high-quality SVG generation by framing the problem as a code generation task and leveraging large-scale pretraining. VLMs are particularly suitable for this task as they capture both global semantics and fine-grained visual patterns, while transferring knowledge across vision, natural language, and code domains. However, existing VLM approaches often struggle to produce faithful and efficient SVGs because they never observe the rendered images during training. Although differentiable rendering for autoregressive SVG code generation remains unavailable, rendered outputs can still be compared to original inputs, enabling evaluative feedback suitable for reinforcement learning (RL). We introduce Reinforcement Learning from Rendering Feedback, an RL method that enhances SVG generation in autoregressive VLMs by leveraging feedback from rendered SVG outputs. Given an input image, the model generates SVG roll-outs that are rendered and compared to the original image to compute a reward. This visual fidelity feedback guides the model toward producing more accurate, efficient, and semantically coherent SVGs. \method significantly outperforms supervised fine-tuning, addressing common failure modes and enabling precise, high-quality SVG generation with strong structural understanding and generalization.
Juan A. Rodríguez, Abhay Puri, Rishav Pramanik, Aarash Feizi, Pascal Wichmann, Arnab Kumar Mondal, Mohammad Reza Samsami, Rabiul Awal, Perouz Taslakian, Spandana Gella, Sai Rajeswar, David Vázquez 0001, Christopher Joseph Pal, Marco Pedersoli
NeurIPS5
2024 Party Prediction for Twitter
abstract
A large number of studies on social media compare the behaviour of users from different political parties. As a basic step, they employ a predictive model for inferring their political affiliation. The accuracy of this model can change the conclusions of a downstream analysis significantly, yet the choice between different models seems to be made arbitrarily. In this paper, we provide a comprehensive survey and an empirical comparison of the current party prediction practices and propose several new approaches which are competitive with or outperform state-of-the-art methods, yet require less computational resources. Party prediction models rely on the content generated by the users (e.g., tweet texts), the relations they have (e.g., who they follow), or their activities and interactions (e.g., which tweets they like). We examine all of these and compare their signal strength for the party prediction task. This paper lets the practitioner select from a wide range of data types that all give strong performance. Finally, we conduct extensive experiments on different aspects of these methods, such as data collection speed and transfer capabilities, which can provide further insights for both applied and methodological research.
Kellin Pelrine, Anne Imouza, Zachary Yang, Jacob-Junqi Tian, Sacha Levy, Gabrielle Desrosiers-Brisebois, Aarash Feizi, Cécile Amadoro, André Blais, Jean-François Godbout, Reihaneh Rabbany
ICWSM7
2020 Structure Aware Negative Sampling in Knowledge Graphs
abstract
Learning low-dimensional representations for entities and relations in knowledge graphs using contrastive estimation represents a scalable and effective method for inferring connectivity patterns.A crucial aspect of contrastive learning approaches is the choice of corruption distribution that generates hard negative samples, which force the embedding model to learn discriminative representations and find critical characteristics of observed data.While earlier methods either employ too simple corruption distributions, i.e. uniform, yielding easy uninformative negatives or sophisticated adversarial distributions with challenging optimization schemes, they do not explicitly incorporate known graph structure resulting in suboptimal negatives.In this paper, we propose Structure Aware Negative Sampling (SANS), an inexpensive negative sampling strategy that utilizes the rich graph structure by selecting negative samples from a node's k-hop neighborhood.Empirically, we demonstrate that SANS finds semantically meaningful negatives and is competitive with SOTA approaches while requires no additional parameters nor difficult adversarial optimization.
Kian Ahrabian, Aarash Feizi, Yasmin Salehi, William L. Hamilton, Joey Bose
EMNLP (1)2