Sina Malakouti

dblp:357/5308 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Generative modeling · 25% Trustworthy machine learning · 25% Language models and text generation · 24%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › fairness › fairness in generative models
bias in generative models
0.912025
Role Bias in Diffusion Models: Diagnosing and Mitigating through Intermediate Decomposition · NeurIPS 2025
Natural language and speech › Language models and text generation
compositional generalization
0.912025
Role Bias in Diffusion Models: Diagnosing and Mitigating through Intermediate Decomposition · NeurIPS 2025
Machine learning › Generative modeling
diffusion model
0.912025
Role Bias in Diffusion Models: Diagnosing and Mitigating through Intermediate Decomposition · NeurIPS 2025
Machine learning › Trustworthy machine learning
fairness
0.912025
Role Bias in Diffusion Models: Diagnosing and Mitigating through Intermediate Decomposition · NeurIPS 2025
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.912025
Role Bias in Diffusion Models: Diagnosing and Mitigating through Intermediate Decomposition · NeurIPS 2025
Computer vision › Image recognition and object detection
object recognition
0.812024
Incorporating Geo-Diverse Knowledge into Prompting for Increased Geographical Robustness in Object Recognition · CVPR 2024
Natural language and speech › Language models and text generation › prompt tuning
soft prompt tuning
0.812024
Incorporating Geo-Diverse Knowledge into Prompting for Increased Geographical Robustness in Object Recognition · CVPR 2024
Computer vision › Vision and language › vision-language model › prompt learning
vision-language model prompting
0.812024
Incorporating Geo-Diverse Knowledge into Prompting for Increased Geographical Robustness in Object Recognition · CVPR 2024
Machine learning › Transfer learning and domain adaptation
domain shift
0.212024
Incorporating Geo-Diverse Knowledge into Prompting for Increased Geographical Robustness in Object Recognition · CVPR 2024

Methods — techniques the papers use, named apart from their topics

fine-tuning · 0.9large language model prompting · 0.8knowledge regularization · 0.8CLIP · 0.8
YearPublicationVenuePosition
2025 Role Bias in Diffusion Models: Diagnosing and Mitigating through Intermediate Decomposition
abstract
Text-to-image (T2I) diffusion models exhibit impressive photorealistic image generation capabilities, yet they struggle in compositional image generation. In this work, we introduce RoleBench, a benchmark focused on evaluating compositional generalization in action-based relations (e.g., "mouse chasing cat"). We show that state-of-the-art T2I models and compositional generation methods consistently default to frequent reversed relations (i.e., "cat chasing mouse"), a phenomenon we call role collapse. Related works attribute this to the model’s architectural limitation or underrepresentation in the data. Our key insight reveals that while models fail on rare compositions when their inversions are common, they can successfully generate similar intermediate compositions (e.g., "mouse chasing boy"), suggesting that this limitation is also due to the presence of frequent counterparts rather than just the absence of rare compositions. Motivated by this, we hypothesize that directional decomposition can gradually mitigate role collapse. We test this via ReBind, a lightweight framework that teaches role bindings using carefully selected active/passive intermediate compositions. Experiments suggest that intermediate compositions through simple fine-tuning can significantly reduce role collapse, with humans preferring ReBind more than 78% compared to state-of-the-art methods. Our findings highlight the role of distributional asymmetries in compositional failures and offer a simple, effective path for improving generalization.
Sina Malakouti, Adriana Kovashka
NeurIPS1
2025 Benchmarking VLMs' Reasoning About Persuasive Atypical Images
abstract
Vision-language models (VLMs) have shown strong zero-shot generalization across various tasks, especially when integrated with large language models (LLMs). However, their ability to comprehend rhetorical and persuasive visual media, such as advertisements, remains understudied. Ads often employ atypical imagery, using surprising object juxtapositions to convey shared properties. For example, Fig. 1(e) shows a beer with a feather-like texture. This requires advanced reasoning to deduce that this atypical representation signifies the beer's lightness. We introduce three novel tasks, Multi-label Atypicality Classification, Atypicality Statement Retrieval, and Atypical Object Recognition, to benchmark VLMs' understanding of atypicality in persuasive images. We evaluate how well VLMs use atypicality to infer an ad's message and test their reasoning abilities by employing semantically challenging negatives. Finally, we pioneer atypicality-aware verbalization by extracting comprehensive image descriptions sensitive to atypical elements. Findings reveal that: (1) VLMs lack advanced reasoning capabilities compared to LLMs; (2) simple, effective strategies can extract atypicality-aware information, leading to comprehensive image verbalization; (3) atypicality aids persuasive ad understanding. Code and data is available at aysanaghazadeh.github.io/PersuasiveAdVLMBenchmark/
Sina Malakouti, Aysan Aghazadeh, Ashmit Khandelwal, Adriana Kovashka
WACV1
2024 Incorporating Geo-Diverse Knowledge into Prompting for Increased Geographical Robustness in Object Recognition
abstract
Existing object recognition models have been shown to lack robustness in diverse geographical scenarios due to domain shifts in design and context. Class representations need to be adapted to more accurately reflect an object concept under these shifts. In the absence of training data from target geographies, we hypothesize that geographically diverse descriptive knowledge of categories can enhance robustness. For this purpose, we explore the feasibility of probing a large language model for geography-based object knowledge, and we examine the effects of integrating knowledge into zero-shot and learnable soft prompting with CLIP. Within this exploration, we propose geog-raphy knowledge regularization to ensure that soft prompts trained on a source set of geographies generalize to an un-seen target set. Accuracy gains over prompting baselines on DollarStreet while training only on Europe data are up to +2.8/1.2/1.6 on target data from Africa/Asia/Americas, and +4.6 overall on the hardest classes. Competitive performance is shown vs. few-shot target training, and analysis is provided to direct future study of geographical robustness.
Kyle Buettner, Sina Malakouti, Xiang Li 0069, Adriana Kovashka
CVPR2
2023 Semi-Supervised Domain Generalization for Object Detection via Language-Guided Feature Alignment
Sina Malakouti, Adriana Kovashka
BMVC1