Afsaneh Fazly

dblp:84/1969 · DBLP profile ↗
← Back
40ranked-venue papers
4as first author
13since 2021 · last 2025
0000-0002-4479-4901ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 35 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 8Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Augmenting Perceptual Super-Resolution via Image Quality Predictors
abstract
Super-resolution (SR), a classical inverse problem in computer vision, is inherently ill-posed, inducing a distribution of plausible solutions for every input. However, the desired result is not simply the expectation of this distribution, which is the blurry image obtained by minimizing pixelwise error, but rather the sample with the highest image quality. A variety of techniques, from perceptual metrics to adversarial losses, are employed to this end. In this work, we explore an alternative: utilizing powerful non-reference image quality assessment (NR-IQA) models in the SR context. We begin with a comprehensive analysis of NR-IQA metrics on human-derived SR data, identifying both the accuracy (human alignment) and complementarity of different metrics. Then, we explore two methods of applying NR-IQA models to SR learning: (i) altering data sampling, by building on an existing multi-ground-truth SR framework, and (ii) directly optimizing a differentiable quality score. Our results demonstrate a more human-centric perception-distortion tradeoff, focusing less on non-perceptual pixelwise distortion, instead improving the balance between perceptual fidelity and human-tuned NR-IQA measures.
Fengjia Zhang, Samrudhdhi B. Rangrej, Tristan Aumentado-Armstrong, Afsaneh Fazly, Alex Levinshtein
CVPR4
2024 End-to-end Parsing of Procedural Text into Flow Graphs
abstract
We focus on the problem of parsing procedural text into fine-grained flow graphs that encode actions and entities, as well as their interactions. Specifically, we focus on parsing cooking recipes, and address a few limitations of existing parsers. Unlike SOTA approaches to flow graph parsing that work in two separate stages identifying actions and entities (tagging) and encoding their interactions via connecting edges (graph generation). we propose an end-to-end multi-task framework that simultaneously performs tagging and graph generation. In addition, due to the end-to-end nature of our proposed model, we can unify the input representation, and moreover can use compact encoders, resulting in small models with significantly fewer parameters than SOTA models. Another key challenge in training flow graph parsers is the lack of sufficient annotated data, due to the costly nature of the fine-grained annotations. We address this problem by taking advantage of the abundant unlabelled recipes, and show that pre-training on automatically-generated noisy silver annotations (from unlabelled recipes) results in a large improvement in flow graph parsing.
Dhaivat Bhatt, Seyed Ahmad Abdollahpouri Hosseini, Federico Fancellu, Afsaneh Fazly
LREC/COLING4
2024 Graph Guided Question Answer Generation for Procedural Question-Answering
abstract
Hai Pham, Isma Hadji, Xinnuo Xu, Ziedune Degutyte, Jay Rainey, Evangelos Kazakos, Afsaneh Fazly, Georgios Tzimiropoulos, Brais Martinez. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Hai X. Pham, Isma Hadji, Xinnuo Xu, Ziedune Degutyte, Jay Rainey, Evangelos Kazakos, Afsaneh Fazly, Georgios Tzimiropoulos, Brais Martínez
EACL (1)7
2024 CIC-BART-SSA: Controllable Image Captioning with Structured Semantic Augmentation
Kalliopi Basioti, Mohamed Ashraf Abdelsalam, Federico Fancellu, Vladimir Pavlovic 0001, Afsaneh Fazly
ECCV (66)5
2024 Realizing Efficient On-Device Language-based Image Retrieval
abstract
Advances in deep learning have enabled accurate language-based search and retrieval (e.g., over user photos) in the cloud. Many users prefer to store their photos in the home due to privacy concerns. As such, a need arises for models that can perform cross-modal search on resource-limited devices. State-of-the-art (SOTA) cross-modal retrieval models achieve high accuracy through learning entangled representations that enable fine-grained similarity calculation between a language query and an image, but at the expense of having a prohibitively high retrieval latency. Alternatively, there is a new class of methods that exhibits good performance with low latency but requires a lot more computational resources and an order of magnitude more training data (i.e., large web-scraped datasets consisting of millions of image–caption pairs), making them infeasible to use in a commercial context. From a pragmatic perspective, none of the existing methods are suitable for developing commercial applications for low-latency cross-modal retrieval on low-resource devices. We propose CrispSearch, a cascaded approach that greatly reduces the retrieval latency with minimal loss in ranking accuracy for on-device language-based image retrieval. The idea behind our approach is to combine a light-weight and runtime-efficient coarse model with a fine re-ranking stage. Given a language query, the coarse model effectively filters out many of the irrelevant image candidates. After this filtering, only a handful of strong candidates will be selected and sent to a fine model for re-ranking. Extensive experimental results with two SOTA models for the fine re-ranking stage on standard benchmark datasets show that CrispSearch results in a speedup of up to 38 times over the SOTA fine methods with negligible performance degradation. Moreover, our method does not require millions of training instances, making it a pragmatic solution to on-device search and retrieval.
Zhiming Hu 0005, Mete Kemertas, Caleb Phillips, Iqbal Mohomed, Afsaneh Fazly
ACM Trans. Multim. Comput. Commun. Appl.6
2023 Identifying Multimodal Context Awareness Requirements for Supporting User Interaction with Procedural Videos
abstract
Following along how-to videos requires alternating focus between understanding procedural video instructions and performing them. Examining how to support these continuous context switches for the user has been largely unexplored. In this paper, we describe a user study with thirty participants who performed an hour-long cooking task while interacting with a wizard-of-oz hands-free interactive system that is aware of both their cooking progress and environment contexts. Through analysis of the session scripts, we identify a dichotomy between participant query differences and workflow alignment similarities, under-studied interactions that require AI functionality beyond video navigation alone, and queries that call for multimodal sensing of a user’s environment. By understanding the assistant experience through the participants’ interactions, we identify design implications for a smart assistant that can discern a user’s task completion flow and personal characteristics, accommodate requests within and external to the task domain, and support nonvoice-based queries.
Georgianna Lin, Jin Yi Li, Afsaneh Fazly, Vladimir Pavlovic 0001, Khai N. Truong
CHI3
2023 GePSAn: Generative Procedure Step Anticipation in Cooking Videos
abstract
We study the problem of future step anticipation in procedural videos. Given a video of an ongoing procedural activity, we predict a plausible next procedure step described in rich natural language. While most previous work focuses on the problem of data scarcity in procedural video datasets, another core challenge of future anticipation is how to account for multiple plausible future realizations in natural settings. This problem has been largely overlooked in previous work. To address this challenge, we frame future step prediction as modelling the distribution of all possible candidates for the next step. Specifically, we design a generative model that takes a series of video clips as input, and generates multiple plausible and diverse candidates (in natural language) for the next step. Following previous work, we side-step the video annotation scarcity by pretraining our model on a large text-based corpus of procedural activities, and then transfer the model to the video domain. Our experiments, both in textual and video domains, show that our model captures diversity in the next step prediction and generates multiple plausible future predictions. Moreover, our model establishes new state-of-the-art results on YouCookII, where it outperforms existing baselines on the next step anticipation. Finally, we also show that our model can successfully transfer from text to the video domain zero-shot, i.e., without fine-tuning or adaptation, and produces good-quality future step predictions from video.
Mohamed Ashraf Abdelsalam, Samrudhdhi B. Rangrej, Isma Hadji, Nikita Dvornik, Konstantinos G. Derpanis, Afsaneh Fazly
ICCV6
2022 SAGE: Saliency-Guided Mixup with Optimal Rearrangements
Avery Ma, Nikita Dvornik, Leila Pishdad, Konstantinos G. Derpanis, Afsaneh Fazly
BMVC6
2022 Visual Semantic Parsing: From Images to Abstract Meaning Representation
abstract
Mohamed Ashraf Abdelsalam, Zhan Shi, Federico Fancellu, Kalliopi Basioti, Dhaivat Bhatt, Vladimir Pavlovic, Afsaneh Fazly. Proceedings of the 26th Conference on Computational Natural Language Learning (CoNLL). 2022.
Mohamed Ashraf Abdelsalam, Federico Fancellu, Kalliopi Basioti, Dhaivat Bhatt, Vladimir Pavlovic 0001, Afsaneh Fazly
CoNLL7
2022 Flow Graph to Video Grounding for Weakly-Supervised Multi-step Localization
Nikita Dvornik, Isma Hadji, Hai X. Pham, Dhaivat Bhatt, Brais Martínez, Afsaneh Fazly, Allan Douglas Jepson
ECCV (35)6
2022 CrispSearch: low-latency on-device language-based image retrieval
abstract
Advances in deep learning have enabled accurate language-based search and retrieval, e.g., over user photos, in the cloud. Many users prefer to store their photos in the home due to privacy concerns. As such, a need arises for models that can perform cross-modal search on resource-limited devices. State-of-the-art cross-modal retrieval models achieve high accuracy through learning entangled representations that enable fine-grained similarity calculation between a language query and an image, but at the expense of having a prohibitively high retrieval latency. Alternatively, there is a new class of methods that exhibits good performance with low latency, but requires a lot more computational resources, and an order of magnitude more training data (i.e. large web-scraped datasets consisting of millions of image-caption pairs) making them infeasible to use in a commercial context. From a pragmatic perspective, none of the existing methods are suitable for developing commercial applications for low-latency cross-modal retrieval on low-resource devices. We propose CrispSearch, a cascaded approach that greatly reduces the retrieval latency with minimal loss in ranking accuracy for on-device language-based image retrieval. The idea behind our approach is to combine a light-weight and runtime-efficient coarse model with a fine re-ranking stage. Given a language query, the coarse model effectively filters out many of the irrelevant image candidates. After this filtering, only a handful of strong candidates will be selected and sent to a fine model for re-ranking. Extensive experimental results with two SOTA models for the fine re-ranking stage, on two standard benchmark datasets (namely, MSCOCO and Flickr30k) show that CrispSearch results in a speedup of up to 38 times over the SOTA fine methods with negligible performance degradation. Moreover, our method does not require millions of training instances, making it a pragmatic solution to on-device search and retrieval.
Zhiming Hu 0005, Mete Kemertas, Caleb Phillips, Iqbal Mohomed, Afsaneh Fazly
MMSys6
2021 Dependency parsing with structure preserving embeddings
abstract
Ákos Kádár, Lan Xiao, Mete Kemertas, Federico Fancellu, Allan Jepson, Afsaneh Fazly. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Ákos Kádár, Mete Kemertas, Federico Fancellu, Allan Douglas Jepson, Afsaneh Fazly
EACL6
2021 Personalized Multi-modal Video Retrieval on Mobile Devices
abstract
Current video retrieval systems on mobile devices cannot process complex natural language queries, especially if they contain personalized concepts, such as proper names. To address these shortcomings, we propose an efficient and privacy-preserving video retrieval system that works well with personalized queries containing proper names, without re-training using personalized labelled data from users. Our system first computes an initial ranking of a video collection by using a generic attention-based video-text matching model (i.e., a model designed for non-personalized queries), and then uses a face detector to conduct personalized adjustments to these initial rankings. These adjustments are done by reasoning over the face information from the detector and the attention information provided by the generic model. We show that our system significantly outperforms existing keyword-based retrieval systems, and achieves comparable performance to the generic matching model fine-tuned on plenty of labelled data. Our results suggest that the proposed system can effectively capture both semantic context and personalized information in queries.
Allan Douglas Jepson, Iqbal Mohomed, Konstantinos G. Derpanis, Afsaneh Fazly
ACM Multimedia6
2020 How coherent are neural models of coherence?
abstract
Despite the recent advances in coherence modelling, most such models including state-of-the-art neural ones, are evaluated on either contrived proxy tasks such as the standard order discrimination benchmark, or tasks that require special expert annotation.Moreover, most evaluations are conducted on small newswire corpora.To address these shortcomings, in this paper we propose four generic evaluation tasks that draw on different aspects of coherence at both the lexical and document levels, and can be applied to any corpora.In designing these tasks, we aim at capturing coherence-specific properties, such as the correct use of discourse connectives, lexical cohesion, as well as the overall temporal and causal consistency among events and participants in a story.Importantly, our proposed tasks either rely on automatically-generated data, or data annotated for other purposes, hence alleviating the need for annotation specifically targeted to the task of coherence modelling.We perform experiments with several existing state-of-the-art neural models of coherence on these tasks, across large corpora from different domains, including newswire, dialogue, as well as narrative and instructional text.Our findings point to a strong need for revisiting the common practices in the development and evaluation of coherence models.
Leila Pishdad, Federico Fancellu, Afsaneh Fazly
COLING4
2020 RankMI: A Mutual Information Maximizing Ranking Loss
abstract
We introduce an information-theoretic loss function, RankMI, and an associated training algorithm for deep representation learning for image retrieval. Our proposed framework consists of alternating updates to a network that estimates the divergence between distance distributions of matching and non-matching pairs of learned embeddings, and an embedding network that maximizes this estimate via sampled negatives. In addition, under this information-theoretic lens we draw connections between RankMI and commonly-used ranking losses, e.g., triplet loss. We extensively evaluate RankMI on several standard image retrieval datasets, namely, CUB-200-2011, CARS-196, and Stanford Online Products. Our method achieves competitive results or significant improvements over previous reported results on all datasets.
Mete Kemertas, Leila Pishdad, Konstantinos G. Derpanis, Afsaneh Fazly
CVPR4
2020 VASTA: a vision and language-assisted smartphone task automation system
abstract
We present VASTA, a novel vision and language-assisted Programming By Demonstration (PBD) system for smartphone task automation. Development of a robust PBD automation system requires overcoming three key challenges: first, how to make a particular demonstration robust to positional and visual changes in the user interface (UI) elements; secondly, how to recognize changes in the automation parameters to make the demonstration as generalizable as possible; and thirdly, how to recognize from the user utterance what automation the user wishes to carry out. To address the first challenge, VASTA leverages state-of-the-art computer vision techniques, including object detection and optical character recognition, to accurately label interactions demonstrated by a user, without relying on the underlying UI structures. To address the second and third challenges, VASTA takes advantage of advanced natural language understanding algorithms for analyzing the user utterance to trigger the VASTA automation scripts, and to determine the automation parameters for generalization. We run an initial user study that demonstrates the effectiveness of VASTA at clustering user utterances, understanding changes in the automation parameters, detecting desired UI elements, and, most importantly, automating various tasks. A demo video of the system is available here: http://y2u.be/kr2xE-FixjI.
Alborz Rezazadeh Sereshkeh, Gary Leung, Krish Perumal, Caleb Phillips, Minfan Zhang, Afsaneh Fazly, Iqbal Mohomed
IUI6
2016 Bilingual Contexts from Comparable Corpora to Mine for Translations of Collocations
Shiva Taslimipoor, Ruslan Mitkov, Gloria Corpas Pastor, Afsaneh Fazly
CICLing (2)4
2016 Classifying Out-of-vocabulary Terms in a Domain-Specific Social Media Corpus
SoHyun Park, Afsaneh Fazly, Annie Lee, Brandon Seibel, Wenjie Zi, Paul Cook
LREC2
2014 Gradual Acquisition of Mental State Meaning: A Computational Investigation
Libby Barak, Afsaneh Fazly, Suzanne Stevenson
CogSci2
2014 Learning Meaning without Primitives: Typology Predicts Developmental Patterns
Barend Beekhuizen, Afsaneh Fazly, Suzanne Stevenson
CogSci2
2014 Structural Differences in the Semantic Networks of Simulated Word Learners
Aida Nematzadeh, Afsaneh Fazly, Suzanne Stevenson
CogSci2
2014 A Cognitive Model of Semantic Network Learning
abstract
Child semantic development includes learning the meaning of words as well as the semantic relations among words.A presumed outcome of semantic development is the formation of a semantic network that reflects this knowledge.We present an algorithm for simultaneously learning word meanings and gradually growing a semantic network, which adheres to the cognitive plausibility requirements of incrementality and limited computations.We demonstrate that the semantic connections among words in addition to their context is necessary in forming a semantic network that resembles an adult's semantic knowledge.
Aida Nematzadeh, Afsaneh Fazly, Suzanne Stevenson
EMNLP2
2013 Modeling the Emergence of an Exemplar Verb in Construction Learning
Libby Barak, Afsaneh Fazly, Suzanne Stevenson
CogSci2
2013 Word Learning in the Wild: What Natural Data Can Tell Us
Barend Beekhuizen, Afsaneh Fazly, Aida Nematzadeh, Suzanne Stevenson
CogSci2
2013 Desirable Difficulty in Learning: A Computational Investigation
Aida Nematzadeh, Afsaneh Fazly, Suzanne Stevenson
CogSci2
2013 Acquisition of Desires before Beliefs: A Computional Investigation
Libby Barak, Afsaneh Fazly, Suzanne Stevenson
CoNLL2
2012 Automatic Identification of Persian Light Verb Constructions
Bahar Salehi, Narjes Askarian, Afsaneh Fazly
CICLing (1)3
2012 Interaction of Word Learning and Semantic Category Formation in Late Talking
Aida Nematzadeh, Afsaneh Fazly, Suzanne Stevenson
CogSci2
2012 Using Noun Similarity to Adapt an Acceptability Measure for Persian Light Verb Constructions
Shiva Taslimipoor, Afsaneh Fazly, Ali Hamzeh
LREC2
2012 Discovering hierarchical object models from captioned images
Michael Jamieson, Yulia Eskin, Afsaneh Fazly, Suzanne Stevenson, Sven J. Dickinson
Comput. Vis. Image Underst.3
2011 A Computational Study of Late Talking in Word-Meaning Acquisition
Aida Nematzadeh, Afsaneh Fazly, Suzanne Stevenson
CogSci2
2010 Discovering Multipart Appearance Models from Captioned Images
Michael Jamieson, Yulia Eskin, Afsaneh Fazly, Suzanne Stevenson, Sven J. Dickinson
ECCV (5)3
2010 Using Language to Learn Structured Appearance Models for Image Annotation
abstract
Given an unstructured collection of captioned images of cluttered scenes featuring a variety of objects, our goal is to simultaneously learn the names and appearances of the objects. Only a small fraction of local features within any given image are associated with a particular caption word, and captions may contain irrelevant words not associated with any image object. We propose a novel algorithm that uses the repetition of feature neighborhoods across training images and a measure of correspondence with caption words to learn meaningful feature configurations (representing named objects). We also introduce a graph-based appearance model that captures some of the structure of an object by encoding the spatial relationships among the local visual features. In an iterative procedure, we use language (the words) to drive a perceptual grouping process that assembles an appearance model for a named object. Results of applying our method to three data sets in a variety of conditions demonstrate that, from complex, cluttered, real-world scenes with noisy captions, we can learn both the names and appearances of objects, resulting in a set of models invariant to translation, scale, orientation, occlusion, and minor changes in viewpoint or articulation. These named models, in turn, are used to automatically annotate new, uncaptioned images, thereby facilitating keyword-based image retrieval.
Michael Jamieson, Afsaneh Fazly, Suzanne Stevenson, Sven J. Dickinson, Sven Wachsmuth
IEEE Trans. Pattern Anal. Mach. Intell.2
2009 Unsupervised Type and Token Identification of Idiomatic Expressions
abstract
Idiomatic expressions are plentiful in everyday language, yet they remain mysterious, as it is not clear exactly how people learn and understand them. They are of special interest to linguists, psycholinguists, and lexicographers, mainly because of their syntactic and semantic idiosyncrasies as well as their unclear lexical status. Despite a great deal of research on the properties of idioms in the linguistics literature, there is not much agreement on which properties are characteristic of these expressions. Because of their peculiarities, idiomatic expressions have mostly been overlooked by researchers in computational linguistics. In this article, we look into the usefulness of some of the identified linguistic properties of idioms for their automatic recognition. Specifically, we develop statistical measures that each model a specific property of idiomatic expressions by looking at their actual usage patterns in text. We use these statistical measures in a type-based classification task where we automatically separate idiomatic expressions (expressions with a possible idiomatic interpretation) from similar-on-the-surface literal phrases (for which no idiomatic interpretation is possible). In addition, we use some of the measures in a token identification task where we distinguish idiomatic and literal usages of potentially idiomatic expressions in context.
Afsaneh Fazly, Paul Cook, Suzanne Stevenson
Comput. Linguistics1
2008 Fast Mapping in Word Learning: What Probabilities Tell Us
Afra Alishahi, Afsaneh Fazly, Suzanne Stevenson
CoNLL2
2008 An Incremental Bayesian Model for Learning Syntactic Categories
Christopher Parisien, Afsaneh Fazly, Suzanne Stevenson
CoNLL2
2007 Learning Structured Appearance Models from Captioned Images of Cluttered Scenes
abstract
Given an unstructured collection of captioned images of cluttered scenes featuring a variety of objects, our goal is to learn both the names and appearances of the objects. Only a small number of local features within any given image are associated with a particular caption word. We describe a connected graph appearance model where vertices represent local features and edges encode spatial relationships. We use the repetition of feature neighborhoods across training images and a measure of correspondence with caption words to guide the search for meaningful feature configurations. We demonstrate improved results on a dataset to which an unstructured object model was previously applied. We also apply the new method to a more challenging collection of captioned images from the Web, detecting and annotating objects within highly cluttered realistic scenes.
Michael Jamieson, Afsaneh Fazly, Sven J. Dickinson, Suzanne Stevenson, Sven Wachsmuth
ICCV2
2006 Automatically Determining Allowable Combinations of a Class of Flexible Multiword Expressions
Afsaneh Fazly, Ryan North, Suzanne Stevenson
CICLing1
2006 Automatically Constructing a Lexicon of Verb Phrase Idiomatic Combinations
Afsaneh Fazly, Suzanne Stevenson
EACL1
2005 Automatic Acquisition of Knowledge About Multiword Predicates
Afsaneh Fazly, Suzanne Stevenson
PACLIC1