Danna Gurari

dblp:118/9699 · DBLP profile ↗
← Back
46ranked-venue papers
10as first author
21since 2021 · last 2026
0000-0003-1306-0283ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 27 · 6 first-author · 15 since 2021Artificial intelligence and machine learning · 18 · 6 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 17 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-authorSecurity and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Hierarchical Instance Tracking to Balance Privacy Preservation with Accessible Information
abstract
We propose a novel task, hierarchical instance tracking, which entails tracking all instances of predefined categories of objects and parts, while maintaining their hierarchical relationships. We introduce the first benchmark dataset supporting this task, consisting of 2,765 unique entities that are tracked in 552 videos and belong to 40 categories (across objects and parts). Evaluation of seven variants of four models tailored to our novel task reveals the new dataset is challenging. Our dataset is available at https://vizwiz.org/tasks-and-datasets/hierarchical-instance-tracking/
Neelima Prasad, Jarek Reynolds, Neel Karsanbhai, Tanusree Sharma, Lotus Hanzi Zhang, Abigale Stangl, Yang Wang 0005, Leah Findlater, Danna Gurari
WACV9
2025 The Accessibility, Security, and Privacy Nexus: Trends and Opportunities
abstract
Insights into the unique security and privacy practices, risks, and solutions for people with disabilities are currently fragmented across disciplines.In this work, we present a literature review of 33 papers published at leading human-computer interaction, accessibility, and usable security and privacy venues.We categorize the contributions of these papers-ranging from interventions to empirical studies of risks and behaviors-and identify key themes and implications.Papers in this corpus highlight 1) the opportunities and risks of the data collected by assistive technologies and security and privacy tools, 2) the inaccessibility or low usability of security and privacy solutions for people with disabilities, and 3) the utility of customized, contextual security and privacy solutions.We conclude with best practices for collecting data from disabled communities and implications for the design of assistive technologies and security/privacy tools.
Kelly Mack, Yu-Jie Chen, Lotus Hanzi Zhang, Danna Gurari, Tanusree Sharma, Yang Wang 0005, Leah Findlater
ASSETS4
2025 "Before, I Asked My Mom, Now I Ask ChatGPT": Visual Privacy Management with Generative AI for Blind and Low-Vision People
abstract
Blind and low vision (BLV) individuals use Generative AI (GenAI) tools to interpret and manage visual content in their daily lives.While such tools can enhance the accessibility of visual content and enable greater user independence, they also introduce complex challenges around visual privacy.In this paper, we investigate the current practices and future design preferences of blind and low vision individuals through an interview study with 21 participants.Our findings reveal a range of current practices with GenAI that balance privacy, efficiency, and emotional agency, with users accounting for privacy risks across six key scenarios: selfpresentation, indoor spatial privacy, outdoor spatial privacy, social media sharing, sharing with employer or professional setup, and handling professional content as employers.Our findings reveal design preferences, including on-device processing, zero-retention guarantees, sensitive content redaction, privacy-aware appearance indicators, and multimodal tactile mirrored interaction methods.We conclude with actionable design recommendations to support user-centered visual privacy through GenAI, expanding the notion of privacy and responsible handling of others' information.
Tanusree Sharma, Yu-Yun Tseng, Lotus Hanzi Zhang, Ayae Ide, Kelly Mack, Leah Findlater, Danna Gurari, Yang Wang 0005
ASSETS7
2025 Acknowledging Focus Ambiguity in Visual Questions
Chongyan Chen, Yu-Yun Tseng, Zhuoheng Li, Anush Venkatesh, Danna Gurari
ICCV5
2025 BIV-Priv-Seg: Locating Private Content in Images Taken by People With Visual Impairments
abstract
Individuals who are blind or have low vision (BLV) are at a heightened risk of sharing private information if they share photographs they have taken. To facilitate developing technologies that can help them preserve privacy, we introduce BIV-Priv-Seg, the first localization dataset originating from people with visual impairments that shows private content. It contains 1,028 images with segmentation annotations for 16 private object categories. We first characterize BIV-Priv-Seg and then evaluate modern models' performance for locating private content in the dataset. We find modern models struggle most with locating private objects that are not salient, small, and lack text as well as recognizing when private content is absent from an image. We facilitate future extensions by sharing our new dataset with the evaluation server at https://vizwiz.org/tasks-and-datasets/object-localization/
Yu-Yun Tseng, Tanusree Sharma, Lotus Hanzi Zhang, Abigale Stangl, Leah Findlater, Yang Wang 0005, Danna Gurari
WACV7
2024 Designing Accessible Obfuscation Support for Blind Individuals' Visual Privacy Management
abstract
Blind individuals commonly share photos in everyday life. Despite substantial interest from the blind community in being able to independently obfuscate private information in photos, existing tools are designed without their inputs. In this study, we prototyped a preliminary screen reader-accessible obfuscation interface to probe for feedback and design insights. We implemented a version of the prototype through off-the-shelf AI models (e.g., SAM, BLIP2, ChatGPT) and a Wizard-of-Oz version that provides human-authored guidance. Through a user study with 12 blind participants who obfuscated diverse private photos using the prototype, we uncovered how they understood and approached visual private content manipulation, how they reacted to frictions such as inaccuracy with existing AI models and cognitive load, and how they envisioned such tools to be better designed to support their needs (e.g., guidelines for describing visual obfuscation effects, co-creative interaction design that respects blind users’ agency).
Lotus Hanzi Zhang, Abigale Stangl, Tanusree Sharma, Yu-Yun Tseng, Inan Xu, Danna Gurari, Yang Wang 0005, Leah Findlater
CHI6
2024 Fully Authentic Visual Question Answering Dataset from Online Communities
Chongyan Chen, Mengchen Liu, Noel Codella, Yunsheng Li, Lu Yuan 0001, Danna Gurari
ECCV (48)6
2024 SPIN: Hierarchical Segmentation with Subpart Granularity in Natural Images
Josh Myers-Dean, Jarek Reynolds, Brian L. Price, Danna Gurari
ECCV (24)5
2024 Interactive Segmentation for Diverse Gesture Types Without Context
abstract
Interactive segmentation entails a human marking an image to guide how a model either creates or edits a segmentation. Our work addresses limitations of existing methods: they either only support one gesture type for marking an image (e.g., either clicks or scribbles) or require knowledge of the gesture type being employed, and require specifying whether marked regions should be included versus excluded in the final segmentation. We instead propose a simplified interactive segmentation task where a user only must mark an image, where the input can be of any gesture type without specifying the gesture type. We support this new task by introducing the first interactive segmentation dataset with multiple gesture types as well as a new evaluation metric capable of holistically evaluating interactive segmentation algorithms. We then analyze numerous interactive segmentation algorithms, including ones adapted for our novel task. While we observe promising performance overall, we also highlight areas for future improvement. To facilitate further extensions of this work, we publicly share our new dataset at https://github.com/joshmyersdean/dig.
Josh Myers-Dean, Brian L. Price, Wilson Chan, Danna Gurari
WACV5
2024 Salient Object Detection for Images Taken by People With Vision Impairments
abstract
Salient object detection is the task of producing a binary mask for an image that deciphers which pixels belong to the foreground object versus background. We introduce a new salient object detection dataset using images taken by people who are visually impaired who were seeking to better understand their surroundings, which we call VizWiz-SalientObject. Compared to seven existing datasets, VizWiz-SalientObject is the largest (i.e., 32,000 human-annotated images) and contains unique characteristics including a higher prevalence of text in the salient objects (i.e., in 68% of images) and salient objects that occupy a larger ratio of the images (i.e., on average, ∼50% coverage). We benchmarked ten modern models on our dataset. One method achieves nearly human performance while the rest struggle, mostly for images with salient objects that are large, have less complex boundaries, and lack text as well as for lower quality images. To facilitate future extensions, we share the dataset at https://vizwiz.org/tasks-anddatasets/salient-object-detection.
Jarek Reynolds, Chandra Kanth Nagesh, Danna Gurari
WACV3
2023 Disability-First Design and Creation of A Dataset Showing Private Visual Information Collected With People Who Are Blind
abstract
We present the design and creation of a disability-first dataset, “BIV-Priv,” which contains 728 images and 728 videos of 14 private categories captured by 26 blind participants to support downstream development of artificial intelligence (AI) models. While best practices in dataset creation typically attempt to eliminate private content, some applications require such content for model development. We describe our approach in creating this dataset with private content in an ethical way, including using props rather than participants’ own private objects and balancing multi-disciplinary perspectives (e.g., accessibility, privacy, computer vision) to meet the tangible metrics (e.g., diversity, category, amount of content) to support AI innovations. We observed challenges that our participants encountered during the data collection, including accessibility issues (e.g., understanding foreground vs. background object placement) and issues due to the sensitive nature of the content (e.g., discomfort in capturing some props such as condoms around family members).
Tanusree Sharma, Abigale Stangl, Lotus Hanzi Zhang, Yu-Yun Tseng, Inan Xu, Leah Findlater, Danna Gurari, Yang Wang 0005
CHI7
2023 A New Dataset Based on Images Taken by Blind People for Testing the Robustness of Image Classification Models Trained for ImageNet Categories
abstract
Our goal is to improve upon the status quo for designing image classification models trained in one domain that perform well on images from another domain. Complementing existing work in robustness testing, we introduce the first dataset for this purpose which comes from an authentic use case where photographers wanted to learn about the content in their images. We built a new test set using 8,900 images taken by people who are blind for which we collected metadata to indicate the presence versus absence of 200 ImageNet object categories. We call this dataset VizWiz-Classification. We characterize this dataset and how it compares to the mainstream datasets for evaluating how well ImageNet-trained classification models generalize. Finally, we analyze the performance of 100 ImageNet classification models on our new test dataset. Our fine-grained analysis demonstrates that these models struggle on images with quality issues. To enable future extensions to this work, we share our new dataset with evaluation server at: https://VizWiz.org/tasks-and-datasets/image-classification.
Reza Akbarian Bafghi, Danna Gurari
CVPR2
2023 VQA Therapy: Exploring Answer Differences by Visually Grounding Answers
abstract
Visual question answering is a task of predicting the answer to a question about an image. Given that different people can provide different answers to a visual question, we aim to better understand why with answer groundings. We introduce the first dataset that visually grounds each unique answer to each visual question, which we call VQA-AnswerTherapy. We then propose two novel problems of predicting whether a visual question has a single answer grounding and localizing all answer groundings. We benchmark modern algorithms for these novel problems to show where they succeed and struggle. The dataset and evaluation server can be found publicly at https://vizwiz.org/tasks-and-datasets/vqa-answer-therapy/.
Chongyan Chen, Samreen Anjum, Danna Gurari
ICCV3
2023 ImageAlly: A Human-AI Hybrid Approach to Support Blind People in Detecting and Redacting Private Image Content
Zhuohao (Jerry) Zhang, Smirity Kaushik, Jooyoung Seo, Haolin Yuan, Sauvik Das, Leah Findlater, Danna Gurari, Abigale Stangl, Yang Wang 0005
SOUPS7
2023 Line Search-Based Feature Transformation for Fast, Stable, and Tunable Content-Style Control in Photorealistic Style Transfer
abstract
Photorealistic style transfer is the task of synthesizing a realistic-looking image when adapting the content from one image to appear in the style of another image. Modern models commonly embed a transformation that fuses features describing the content image and style image and then decodes the resulting feature into a stylized image. We introduce a general-purpose transformation that enables controlling the balance between how much content is preserved and the strength of the infused style. We offer the first experiments that demonstrate the performance of existing transformations across different style transfer models, and demonstrate how our transformation performs better in its ability to simultaneously run fast, produce consistently reasonable results, and control the balance between content and style in different models. To support reproducing our method and models, we share the code at https://github.com/chiutaiyin/LS-FT.
Tai-Yin Chiu, Danna Gurari
WACV2
2023 Helping Visually Impaired People Take Better Quality Pictures
abstract
Perception-based image analysis technologies can be used to help visually impaired people take better quality pictures by providing automated guidance, thereby empowering them to interact more confidently on social media. The photographs taken by visually impaired users often suffer from one or both of two kinds of quality issues: technical quality (distortions), and semantic quality, such as framing and aesthetic composition. Here we develop tools to help them minimize occurrences of common technical distortions, such as blur, poor exposure, and noise. We do not address the complementary problems of semantic quality, leaving that aspect for future work. The problem of assessing, and providing actionable feedback on the technical quality of pictures captured by visually impaired users is hard enough, owing to the severe, commingled distortions that often occur. To advance progress on the problem of analyzing and measuring the technical quality of visually impaired user-generated content (VI-UGC), we built a very large and unique subjective image quality and distortion dataset. This new perceptual resource, which we call the LIVE-Meta VI-UGC Database, contains 40K real-world distorted VI-UGC images and 40K patches, on which we recorded 2.7M human perceptual quality judgments and 2.7M distortion labels. Using this psychometric resource we also created an automatic limited vision picture quality and distortion predictor that learns local-to-global spatial quality relationships, achieving state-of-the-art prediction performance on VI-UGC pictures, significantly outperforming existing picture quality models on this unique class of distorted picture data. We also created a prototype feedback system that helps to guide users to mitigate quality issues and take better quality pictures, by creating a multi-task learning framework. The dataset and models can be accessed at: https://github.com/mandal-cv/visimpaired.
Maniratnam Mandal, Deepti Ghadiyaram, Danna Gurari, Alan C. Bovik
IEEE Trans. Image Process.3
2022 Grounding Answers for Visual Questions Asked by Visually Impaired People
abstract
Visual question answering is the task of answering questions about images. We introduce the VizWiz-VQA-Grounding dataset, the first dataset that visually grounds answers to visual questions asked by people with visual impairments. We analyze our dataset and compare it with five VQA-Grounding datasets to demonstrate what makes it similar and different. We then evaluate the SOTA VQA and VQA-Grounding models and demonstrate that current SOTA algorithms often fail to identify the correct visual evidence where the answer is located. These models regularly struggle when the visual evidence occupies a small fraction of the image, for images that are higher quality, as well as for visual questions that require skills in text recognition. The dataset, evaluation server, and leader-board all can be found at the following link: https://vizwiz.org/tasks-and-datasets/answer-grounding-for-vqa/.
Chongyan Chen, Samreen Anjum, Danna Gurari
CVPR3
2022 PCA-Based Knowledge Distillation Towards Lightweight and Content-Style Balanced Photorealistic Style Transfer Models
abstract
Photorealistic style transfer entails transferring the style of a reference image to another image so the result seems like a plausible photo. Our work is inspired by the ob-servation that existing models are slow due to their large sizes. We introduce PCA-based knowledge distillation to distill lightweight models and show it is motivated by the-ory. To our knowledge, this is the first knowledge dis-tillation method for photorealistic style transfer. Our ex-periments demonstrate its versatility for use with differ-ent backbone architectures, VGG and MobileNet, across six image resolutions. Compared to existing models, our top-performing model runs at speeds 5-20x faster using at most 1% of the parameters. Additionally, our dis-tilled models achieve a better balance between stylization strength and content preservation than existing models. To support reproducing our method and models, we share the code at https://github.com/chiutaiyin/PCA-Knowledge-Distillation.
Tai-Yin Chiu, Danna Gurari
CVPR2
2022 VizWiz-FewShot: Locating Objects in Images Taken by People with Visual Impairments
Yu-Yun Tseng, Alexander Bell, Danna Gurari
ECCV (8)3
2022 PhotoWCT2: Compact Autoencoder for Photorealistic Style Transfer Resulting from Blockwise Training and Skip Connections of High-Frequency Residuals
abstract
Photorealistic style transfer is an image editing task with the goal to modify an image to match the style of another image while ensuring the result looks like a real photograph. A limitation of existing models is that they have many parameters, which in turn prevents their use for larger image resolutions and leads to slower run-times. We introduce two mechanisms that enable our design of a more compact model that we call PhotoWCT2, which preserves state-of-art stylization strength and photorealism. First, we introduce blockwise training to perform coarse-to-fine feature transformations that enable state-of-art stylization strength in a single autoencoder in place of the inefficient cascade of four autoencoders used in PhotoWCT. Second, we introduce skip connections of high-frequency residuals in order to preserve image quality when applying the sequential coarse-to-fine feature transformations. Our PhotoWCT2model requires fewer parameters (e.g., 30.3% fewer) while supporting higher resolution images (e.g., 4K) and achieving faster stylization than existing models.
Tai-Yin Chiu, Danna Gurari
WACV2
2021 Going Beyond One-Size-Fits-All Image Descriptions to Satisfy the Information Wants of People Who are Blind or Have Low Vision
abstract
Image descriptions are how people who are blind or have low vision (BLV) access information depicted within images. To our knowledge, no prior work has examined how a description for an image should be designed for different scenarios in which users encounter images. Scenarios consist of the information goal the person has when seeking information from or about an image, paired with the source where the image is found. To address this gap, we interviewed 28 people who are BLV to learn how the scenario impacts what image content (information) should go into an image description. We offer our findings as a foundation for considering how to design next-generation image description technologies that can both (A) support a departure from one-size-fits-all image descriptions to context-aware descriptions, and (B) reveal what content to include in minimum viable descriptions for a large range of scenarios.
Abigale Stangl, Nitin Verma, Kenneth R. Fleischmann, Meredith Ringel Morris, Danna Gurari
ASSETS5
2020 Visual Content Considered Private by People Who are Blind
abstract
We present an empirical study into the visual content people who are blind consider to be private. We conduct a two-stage interview with 18 participants that identifies what they deem private in general and with respect to their use of services that describe their visual surroundings based on camera feeds from their personal devices. We then describe a taxonomy of private visual content that is reflective of our participants’ privacy-related concerns and values. We discuss how this taxonomy can benefit services that collect and sell visual data containing private information so such services are better aligned with their users.
Abigale Stangl, Kristina Shiroma, Bo Xie 0001, Kenneth R. Fleischmann, Danna Gurari
ASSETS5
2020 "Person, Shoes, Tree. Is the Person Naked?" What People with Vision Impairments Want in Image Descriptions
abstract
Access to digital images is important to people who are blind or have low vision (BLV). Many contemporary image description efforts do not take into account this population's nuanced image description preferences. In this paper, we present a qualitative study that provides insight into 28 BLV people's experiences with descriptions of digital images from news websites, social networking sites/platforms, eCommerce websites, employment websites, online dating websites/platforms, productivity applications, and e-publications. Our findings reveal how image description preferences vary based on the source where digital images are encountered and the surrounding context. We provide recommendations for the development of next-generation image description technologies inspired by our empirical analysis.
Abigale Stangl, Meredith Ringel Morris, Danna Gurari
CHI3
2020 Assessing Image Quality Issues for Real-World Problems
abstract
We introduce a new large-scale dataset that links the assessment of image quality issues to two practical vision tasks: image captioning and visual question answering. First, we identify for 39,181 images taken by people who are blind whether each is sufficient quality to recognize the content as well as what quality flaws are observed from six options. These labels serve as a critical foundation for us to make the following contributions: (1) a new problem and algorithms for deciding whether an image is insufficient quality to recognize the content and so not captionable, (2) a new problem and algorithms for deciding which of six quality flaws an image contains, (3) a new problem and algorithms for deciding whether a visual question is unanswerable due to unrecognizable content versus the content of interest being missing from the field of view, and (4) a novel application of more efficiently creating a large-scale image captioning dataset by automatically deciding whether an image is insufficient quality and so should not be captioned. We publicly-share our datasets and code to facilitate future extensions of this work: https://vizwiz.org.
Tai-Yin Chiu, Danna Gurari
CVPR3
2020 Iterative Feature Transformation for Fast and Versatile Universal Style Transfer
Tai-Yin Chiu, Danna Gurari
ECCV (19)2
2020 Captioning Images Taken by People Who Are Blind
Danna Gurari, Nilavra Bhattacharya
ECCV (17)1
2020 CrowdMOT: Crowdsourcing Strategies for Tracking Multiple Objects in Videos
abstract
Crowdsourcing is a valuable approach for tracking objects in videos in a more scalable manner than possible with domain experts. However, existing frameworks do not produce high quality results with non-expert crowdworkers, especially for scenarios where objects split. To address this shortcoming, we introduce a crowdsourcing platform called CrowdMOT, and investigate two micro-task design decisions: (1) whether to decompose the task so that each worker is in charge of annotating all objects in a sub-segment of the video versus annotating a single object across the entire video, and (2) whether to show annotations from previous workers to the next individuals working on the task. We conduct experiments on a diversity of videos which show both familiar objects (aka - people) and unfamiliar objects (aka - cells). Our results highlight strategies for efficiently collecting higher quality annotations than observed when using strategies employed by today's state-of-art crowdsourcing system.
Samreen Anjum, Chi Lin 0001, Danna Gurari
Proc. ACM Hum. Comput. Interact.3
2020 "I Hope This Is Helpful": Understanding Crowdworkers' Challenges and Motivations for an Image Description Task
abstract
AI image captioning challenges encourage broad participation in designing algorithms that automatically create captions for a variety of images and users. To create large datasets necessary for these challenges, researchers typically employ a shared crowdsourcing task design for image captioning. This paper discusses findings from our thematic analysis of 1,064 comments left by Amazon Mechanical Turk workers using this task design to create captions for images taken by people who are blind. Workers discussed difficulties in understanding how to complete this task, provided suggestions of how to improve the task, gave explanations or clarifications about their work, and described why they found this particular task rewarding or interesting. Our analysis provides insights both into this particular genre of task as well as broader considerations for how to employ crowdsourcing to generate large datasets for developing AI algorithms.
Rachel N. Simons, Danna Gurari, Kenneth R. Fleischmann
Proc. ACM Hum. Comput. Interact.2
2020 Vision Skills Needed to Answer Visual Questions
abstract
The task of answering questions about images has garnered attention as a practical service for assisting populations with visual impairments as well as a visual Turing test for the artificial intelligence community. Our first aim is to identify the common vision skills needed for both scenarios. To do so, we analyze the need for four vision skills--object recognition, text recognition, color recognition, and counting--on over 27,000 visual questions from two datasets representing both scenarios. We next quantify the difficulty of these skills for both humans and computers on both datasets. Finally, we propose a novel task of predicting what vision skills are needed to answer a question about an image. Our results reveal (mis)matches between aims of real users of such services and the focus of the AI community. We conclude with a discussion about future directions for addressing the visual question answering task.
Xiaoyu Zeng, Yanan Wang 0008, Tai-Yin Chiu, Nilavra Bhattacharya, Danna Gurari
Proc. ACM Hum. Comput. Interact.5
2019 VizWiz-Priv: A Dataset for Recognizing the Presence and Purpose of Private Visual Information in Images Taken by Blind People
abstract
We introduce the first visual privacy dataset originating from people who are blind in order to better understand their privacy disclosures and to encourage the development of algorithms that can assist in preventing their unintended disclosures. It includes 8,862 regions showing private content across 5,537 images taken by blind people. Of these, 1,403 are paired with questions and 62\% of those directly ask about the private content. Experiments demonstrate the utility of this data for predicting whether an image shows private information and whether a question asks about the private content in an image. The dataset is publicly-shared at http://vizwiz.org/data/.
Danna Gurari, Qing Li 0003, Chi Lin 0001, Anhong Guo, Abigale Stangl, Jeffrey P. Bigham
CVPR1
2019 Why Does a Visual Question Have Different Answers?
abstract
Visual question answering is the task of returning the answer to a question about an image. A challenge is that different people often provide different answers to the same visual question. To our knowledge, this is the first work that aims to understand why. We propose a taxonomy of nine plausible reasons, and create two labelled datasets consisting of ~45,000 visual questions indicating which reasons led to answer differences. We then propose a novel problem of predicting directly from a visual question which reasons will cause answer differences as well as a novel algorithm for this purpose. Experiments demonstrate the advantage of our approach over several related baselines on two diverse datasets. We publicly share the datasets and code at https://vizwiz.org.
Nilavra Bhattacharya, Qing Li 0003, Danna Gurari
ICCV3
2019 Unconstrained Foreground Object Search
abstract
Many people search for foreground objects to use when editing images. While existing methods can retrieve candidates to aid in this, they are constrained to returning objects that belong to a pre-specified semantic class. We instead propose a novel problem of unconstrained foreground object (UFO) search and introduce a solution that supports efficient search by encoding the background image in the same latent space as the candidate foreground objects. A key contribution of our work is a cost-free, scalable approach for creating a large-scale training dataset with a variety of foreground objects of differing semantic categories per image location. Quantitative and human-perception experiments with two diverse datasets demonstrate the advantage of our UFO search solution over related baselines.
Brian L. Price, Scott Cohen, Danna Gurari
ICCV4
2019 Guided Image Inpainting: Replacing an Image Region by Pulling Content From Another Image
abstract
Deep generative models have shown success in automatically synthesizing missing image regions using surrounding context. However, users cannot directly decide what content to synthesize with such approaches.We propose an end-to-end network for image inpainting that uses a different image to guide the synthesis of new content to fill the hole. A key challenge addressed by our approach is synthesizing new content in regions where the guidance image and the context of the original image are inconsistent. We conduct four studies that demonstrate our method yields more realistic image inpainting results over seven baselines.
Brian L. Price, Scott Cohen, Danna Gurari
WACV4
2019 Predicting How to Distribute Work Between Algorithms and Humans to Segment an Image Batch
Danna Gurari, Suyog Dutt Jain, Margrit Betke, Kristen Grauman
Int. J. Comput. Vis.1
2018 BrowseWithMe: An Online Clothes Shopping Assistant for People with Visual Impairments
abstract
Our interviews with people who have visual impairments show clothes shopping is an important activity in their lives. Unfortunately, clothes shopping web sites remain largely inaccessible. We propose design recommendations to address online accessibility issues reported by visually impaired study participants and an implementation, which we call BrowseWithMe, to address these issues. BrowseWithMe employs artificial intelligence to automatically convert a product web page into a structured representation that enables a user to interactively ask the BrowseWithMe system what the user wants to learn about a product (e.g., What is the price? Can I see a magnified image of the pants?). This enables people to be active solicitors of the specific information they are seeking rather than passive listeners of unparsed information. Experiments demonstrate BrowseWithMe can make online clothes shopping more accessible and produce accurate image descriptions.
Abigale Stangl, Esha Kothari, Suyog Dutt Jain, Tom Yeh, Kristen Grauman, Danna Gurari
ASSETS6
2018 VizWiz Grand Challenge: Answering Visual Questions From Blind People
abstract
The study of algorithms to automatically answer visual questions currently is motivated by visual question answering (VQA) datasets constructed in artificial VQA settings. We propose VizWiz, the first goal-oriented VQA dataset arising from a natural VQA setting. VizWiz consists of over 31,000 visual questions originating from blind people who each took a picture using a mobile phone and recorded a spoken question about it, together with 10 crowdsourced answers per visual question. VizWiz differs from the many existing VQA datasets because (1) images are captured by blind photographers and so are often poor quality, (2) questions are spoken and so are more conversational, and (3) often visual questions cannot be answered. Evaluation of modern algorithms for answering visual questions and deciding if a visual question is answerable reveals that VizWiz is a challenging dataset. We introduce this dataset to encourage a larger community to develop more generalized algorithms that can assist blind people.
Danna Gurari, Qing Li 0003, Abigale Stangl, Anhong Guo, Chi Lin 0001, Kristen Grauman, Jiebo Luo 0001, Jeffrey P. Bigham
CVPR1
2018 Visual Question Answer Diversity
abstract
Visual questions (VQs) can lead multiple people to respond with different answers rather than a single, agreed upon response. Moreover, the answers from a crowd can include different numbers of unique answers that arise with different relative frequencies. Such answer diversity arises for a variety of reasons including that VQs are subjective, difficult, or ambiguous. We propose a new problem of predicting the answer distribution that would be observed from a crowd for any given VQ; i.e., the number of unique answers and their relative frequencies. Our experiments confirm that the answer distribution can be predicted accurately for VQs asked by both blind and sighted people. We then propose a novel crowd-powered VQA system that uses the answer distribution predictions to reason about how many answers are needed to capture the diversity of possible human responses. Experiments demonstrate this proposed system accelerates capturing the diversity of answers with considerably less human effort than is required with a state-of-art system.
Chun-Ju Yang, Kristen Grauman, Danna Gurari
HCOMP3
2018 Predicting Foreground Object Ambiguity and Efficiently Crowdsourcing the Segmentation(s)
Danna Gurari, Kun He 0003, Jianming Zhang 0001, Mehrnoosh Sameki, Suyog Dutt Jain, Stan Sclaroff, Margrit Betke, Kristen Grauman
Int. J. Comput. Vis.1
2017 CrowdVerge: Predicting If People Will Agree on the Answer to a Visual Question
abstract
Visual question answering systems empower users to ask any question about any image and receive a valid answer. However, existing systems do not yet account for the fact that a visual question can lead to a single answer or multiple different answers. While a crowd often agrees, disagreements do arise for many reasons including that visual questions are ambiguous, subjective, or difficult. We propose a model, CrowdVerge, for automatically predicting from a visual question whether a crowd would agree on one answer. We then propose how to exploit these predictions in a novel application to efficiently collect all valid answers to visual questions. Specifically, we solicit fewer human responses when answer agreement is expected and more human responses otherwise. Experiments on 121,811 visual questions asked by sighted and blind people show that, compared to existing crowdsourcing systems, our system captures the same answer diversity with typically 14-23% less crowd involvement.
Danna Gurari, Kristen Grauman
CHI1
2017 Crowd-O-Meter: Predicting if a Person Is Vulnerable to Believe Political Claims
abstract
Social media platforms have been criticized for promoting false information during the 2016 U.S. presidential election campaign. Our work is motivated by the idea that a platform could reduce the circulation of false information if it could estimate whether its users are vulnerable to believing political claims. We here explore whether such a vulnerability could be measured in a crowdsourcing setting. We propose Crowd-O-Meter, a framework that automatically predicts if a crowd worker will be consistent in his/her beliefs about political claims; i.e., consistently believes the claims are true or consistently believes the claims are not true. Crowd-O-Meter is a user-centered approach which interprets a combination of cues characterizing the user's implicit and explicit opinion bias. Experiments on 580 quotes from PolitiFact's fact checking corpus of 2016 U.S. presidential candidates show that Crowd-O-Meter is precise and accurate for two news modalities: text and video. Our analysis also reveals which are the most informative cues of a person's vulnerability.
Mehrnoosh Sameki, Linli Ding, Margrit Betke, Danna Gurari
HCOMP5
2016 Pull the Plug? Predicting If Computers or Humans Should Segment Images
abstract
Foreground object segmentation is a critical step for many image analysis tasks. While automated methods can produce high-quality results, their failures disappoint users in need of practical solutions. We propose a resource allocation framework for predicting how best to allocate a fixed budget of human annotation effort in order to collect higher quality segmentations for a given batch of images and automated methods. The framework is based on a proposed prediction module that estimates the quality of given algorithm-drawn segmentations. We demonstrate the value of the framework for two novel tasks related to "pulling the plug" on computer and human annotators. Specifically, we implement two systems that automatically decide, for a batch of images, when to replace 1) humans with computers to create coarse segmentations required to initialize segmentation tools and 2) computers with humans to create final, fine-grained segmentations. Experiments demonstrate the advantage of relying on a mix of human and computer efforts over relying on either resource alone for segmenting objects in three diverse datasets representing visible, phase contrast microscopy, and fluorescence microscopy images.
Danna Gurari, Suyog Dutt Jain, Margrit Betke, Kristen Grauman
CVPR1
2016 Investigating the Influence of Data Familiarity to Improve the Design of a Crowdsourcing Image Annotation System
abstract
Crowdsourced demarcations of object boundaries in images (segmentations) are important for many vision-based applications. A commonly reported challenge is that a large percentage of crowd results are discarded due to concerns about quality. We conducted three studies to examine (1) how does the quality of crowdsourced segmentations differ for familiar everyday images versus unfamiliar biomedical images?, (2) how does making familiar images less recognizable (rotating images upside down) influence crowd work with respect to the quality of results, segmentation time, and segmentation detail?, and (3) how does crowd workers’ judgments of the ambiguity of the segmentation task, collected by voting, differ for familiar everyday images and unfamiliar biomedical images? We analyzed a total of 2,525 segmentations collected from 121 crowd workers and 1,850 votes from 55 crowd workers. Our results illustrate the potential benefit of explicitly accounting for human familiarity with the data when designing computer interfaces for human interaction.
Danna Gurari, Mehrnoosh Sameki, Margrit Betke
HCOMP1
2015 Predicting Quality of Crowdsourced Image Segmentations from Crowd Behavior
abstract
Quality control (QC) is an integral part of many crowd- sourcing systems. However, popular QC methods, such as aggregating multiple annotations, filtering workers, or verifying the quality of crowd work, introduce additional costs and delays. We propose a complementary paradigm to these QC methods based on predicting the quality of submitted crowd work. In particular, we pro- pose to predict the quality of a given crowd drawing directly from a crowd worker’s drawing time, number of user clicks, and average time per user click. We focus on the task of drawing the boundary of a single object in an image. To train and test our prediction models, we collected a total of 2,025 crowd-drawn segmentations for 405 familiar everyday images and unfamiliar biomedical images from 90 unique crowd workers. We first evaluated five prediction models learned using different combinations of the three worker behavior cues for all images. Experiments revealed that time per number of user clicks was the most effective cue for predicting segmentation quality. We next inspected the predictive power of models learned using crowd annotations collected for familiar and unfamiliar data independently. Prediction models were significantly more effective for estimating the segmentation quality from crowd worker behavior for familiar image content than unfamiliar image content.
Mehrnoosh Sameki, Danna Gurari, Margrit Betke
HCOMP2
2015 How to Collect Segmentations for Biomedical Images? A Benchmark Evaluating the Performance of Experts, Crowdsourced Non-experts, and Algorithms
abstract
Analyses of biomedical images often rely on demarcating the boundaries of biological structures (segmentation). While numerous approaches are adopted to address the segmentation problem including collecting annotations from domain-experts and automated algorithms, the lack of comparative benchmarking makes it challenging to determine the current state-of-art, recognize limitations of existing approaches, and identify relevant future research directions. To provide practical guidance, we evaluated and compared the performance of trained experts, crowd sourced non-experts, and algorithms for annotating 305 objects coming from six datasets that include phase contrast, fluorescence, and magnetic resonance images. Compared to the gold standard established by expert consensus, we found the best annotators were experts, followed by non-experts, and then algorithms. This analysis revealed that online paid crowd sourced workers without domain-specific backgrounds are reliable annotators to use as part of the laboratory protocol for segmenting biomedical images. We also found that fusing the segmentations created by crowd sourced internet workers and algorithms yielded improved segmentation results over segmentations created by single crowd sourced or algorithm annotations respectively. We invite extensions of our work by sharing our data sets and associated segmentation annotations (http://www.cs.bu.edu/~betke/Biomedical Image Segmentation).
Danna Gurari, Diane H. Theriault, Mehrnoosh Sameki, Brett Isenberg, Tuan A. Pham 0002, Alberto Purwada, Patricia Solski, Matthew L. Walker, Chentian Zhang, Joyce Y. Wong, Margrit Betke
WACV1
2013 SAGE: An approach and implementation empowering quick and reliable quantitative analysis of segmentation quality
abstract
Finding the outline of an object in an image is a fundamental step in many vision-based applications. It is important to demonstrate that the segmentation found accurately represents the contour of the object in the image. The discrepancy measure model for segmentation analysis focuses on selecting an appropriate discrepancy measure to compute a score that indicates how similar a query segmentation is to a gold standard segmentation. Observing that the score depends on the gold standard segmentation, we propose a framework that expands this approach by introducing the consideration of how to establish the gold standard segmentation. The framework shows how to obtain project-specific performance indicators in a principled way that links annotation tools, fusion methods, and evaluation algorithms into a unified model we call SAGE. We also describe a freely available implementation of SAGE that enables quick segmentation validation against either a single annotation or a fused annotation. Finally, three studies are presented to highlight the impact of annotation tools, an-notators, and fusion methods on establishing trusted gold standard segmentations for cell and artery images.
Danna Gurari, Suele Ki Kim, Brett Isenberg, Tuan A. Pham 0002, Alberto Purwada, Patricia Solski, Matthew L. Walker, Joyce Y. Wong, Margrit Betke
WACV1
2012 Hierarchical Partial Matching and Segmentation of Interacting Cells
Zheng Wu 0003, Danna Gurari, Joyce Y. Wong, Margrit Betke
MICCAI (1)2