VLDB 2026 Research / reviewers in the wild / expert
David Crandall
dblp:c/DavidCrandall · also David J. Crandall
· DBLP profile ↗
141ranked-venue papers
11as first author
60since 2021 · last 2026
0000-0002-5827-5344ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 89 · 10 first-author · 43 since 2021Graphics, computer vision, multimedia, augmented reality and games · 55 · 6 first-author · 18 since 2021Human-computer interaction and ubiquitous computing · 24 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 8 · 3 first-authorSystems, architecture and hardware · 7 · 3 since 2021Security and privacy · 4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BehaviorKit: A Multi-modal Real-Time Behavior Analysis Library for RobotsabstractWe present BehaviorKit, an open-source plug-and-play module that enables the addition of real-time behavior analysis using deep learning to any robot with internet access. The library bundles together gaze tracking, real-time transcription, textual sentiment analysis, facial emotion (valence and arousal) estimation, and face and pose landmark localization. The library is designed to be run on a GPU-powered computer, either on the robot or externally, to process the incoming visual and auditory inputs that the robot receives. The library connects via WebSocket to the robot, which receives the processed outputs from all of the models. The WebSocket client can easily connect to a ROS-powered robot, or a custom client can be written to adapt to any robot; alternatively, the library can also be run between two laptop computers. A small dataset is provided to test the framework. By packaging together and optimizing many commonly used models, we hope to enable easier access to high-performing behavior models for the HRI community. The code is available at https://github.com/IUB-RHouse/BehaviorKit. William Valentine, Selma Sabanovic, David Crandall, Weslie Khoo |
HRI | 3 |
| 2026 | Combining Deep Learning and Large Language Models for Retrieval-Based Image Classification
Zachary Wilkerson, David B. Leake, David Crandall |
ICCBR | 3 |
| 2026 | TimeRefine: Temporal Grounding with Time Refining Video LLMabstractVideo temporal grounding aims to localize relevant temporal boundaries in a video given a textual prompt. Recent work has focused on enabling Video LLMs to perform video temporal grounding via next-token prediction of temporal timestamps. However, accurately localizing timestamps in videos remains challenging for Video LLMs when relying solely on temporal token prediction. Our proposed TimeRefine addresses this challenge in two ways. First, instead of directly predicting the start and end timestamps, we reformulate temporal grounding as a temporal refinement task: the model first makes rough predictions and then refines them by predicting offsets to the target segment. This refining process is repeated multiple times, through which the model progressively improves its own temporal localization accuracy. Second, to enhance the model’s temporal perception capabilities, we incorporate an auxiliary prediction head that applies a larger penalty as a predicted segment deviates further from the ground truth, encouraging more precise temporal localizations. Our plug-and-play method can be integrated into most LLM-based temporal grounding approaches. The experimental results demonstrate that TimeRefine achieves 3.6% and 5.0% mIoU improvements on the ActivityNet and Charades-STA datasets, respectively. Code and pretrained models are available at https://github.com/SJTUwxz/TimeRefine_code. Md Mohaiminul Islam, Lorenzo Torresani, Mohit Bansal, Gedas Bertasius, David Crandall |
WACV | 9 |
| 2026 | GateFusion: Hierarchical Gated Cross-Modal Fusion for Active Speaker DetectionabstractActive Speaker Detection (ASD) aims to identify who is currently speaking in each frame of a video. Most state-of-the-art approaches rely on late fusion to combine visual and audio features, but late fusion often fails to capture fine-grained cross-modal interactions, which can be critical for robust performance in unconstrained scenarios. In this paper, we introduce GateFusion, a novel architecture that combines strong pretrained unimodal encoders with a Hierarchical Gated Fusion Decoder (HiGate). HiGate enables progressive, multi-depth fusion by adaptively injecting contextual features from one modality into the other at multiple layers of the Transformer backbone, guided by learnable, bimodally-conditioned gates. To further strengthen multimodal learning, we propose two auxiliary objectives: Masked Alignment Loss (MAL) to align unimodal outputs with multimodal predictions, and Over-Positive Penalty (OPP) to suppress spurious video-only activations. GateFusion establishes new state-of-the-art results on several challenging ASD benchmarks, achieving 77.8% mAP (+9.4%), 86.1% mAP (+2.9%), and 96.1% mAP (+0.5%) on Ego4D-ASD, UniTalk, and WASD benchmarks, respectively, and delivering competitive performance on AVA-ActiveSpeaker. Out-of-domain experiments demonstrate the generalization of our model, while comprehensive ablations show the complementary benefits of each component. Yu Wang 0164, Juhyung Ha, Frangil Ramirez, David Crandall |
WACV | 5 |
| 2025 | Research as Care: A Reflection on Incorporating the Ethics of Care in Design Research with People Living with Dementia
Long-Jing Hsu, Janice K. Bays, Manasi Swaminathan, Weslie Khoo, Hiroki Sato 0002, Kyrie Jig Amon, Sathvika Dobbala, Min Min Thant, Alex Foster, Katherine M. Tsui, Philip B. Stafford, David Crandall, Selma Sabanovic |
Conference on Designing Interactive Systems | 12 |
| 2025 | Bittersweet Snapshots of Life: Designing to Address Complex Emotions in a Reminiscence Interaction between Older Adults and a Robot
Long-Jing Hsu, Manasi Swaminathan, Weslie Khoo, Kyrie Jig Amon, Hiroki Sato 0002, Sathvika Dobbala, Katherine M. Tsui, David Crandall, Selma Sabanovic |
CHI | 8 |
| 2025 | Million Eyes on the "Robot Umps": The Case for Studying Sports in HRI Through BaseballabstractIn this position paper, we argue that baseball-and sports more broadly-provide a unique and under-explored opportunity for researchers to study human-robot interaction (HRI) in real-world settings. Using the rise of robot umpires in baseball as a primary example, we examine emerging themes such as power dynamics among players and umpires, labor implications, and technical challenges. We emphasize the affordances and benefits of studying sports within HRI, including the integration of interdisciplinary perspectives, the large-scale deployment of robots, and the examination of their role in deeply rooted cultural practices. Waki Kamino, Andrea W. Wen-Yi, Dhruv Agarwal 0001, Sil Hamilton, Eun Jeong Kang, Keigo Kusumegi, Pegah Moradi, Daniel Mwesigwa, Yan Tao, I-Ting Tsai, Ethan Yang, Shengqi Zhu 0002, Shu-Jung Han, Chi-Jung Lee, Michael J. Sack, Tianhong Catherine Yu, Weslie Khoo, Andy Elliot Ricci, Yoyo Tsung-Yu Hou, Selma Sabanovic, David Crandall, Karen Levy, Malte F. Jung |
HRI | 23 |
| 2025 | Learning Case Features with Proxy-Guided Deep Neural Networks
Vibhas Vats, Zachary Wilkerson, Hiroki Sato 0002, David B. Leake, David Crandall |
ICCBR | 5 |
| 2025 | Extracting Features with Deep Learning for Ensemble-Driven Case-Based Classification
Zachary Wilkerson, David B. Leake, David Crandall, Benjamin Wilkerson |
ICCBR | 3 |
| 2025 | HVPUNet: Hybrid-Voxel Point-Cloud Upsampling Network
Juhyung Ha, Vibhas Vats, Soon-Heung Jung, Md. Alimoor Reza, David Crandall |
ICCV | 5 |
| 2025 | What Changed and What Could Have Changed? State-Change Counterfactuals for Procedure-Aware Video Representation LearningabstractUnderstanding a procedural activity requires modeling both how action steps transform the scene, and how evolving scene transformations can influence the sequence of action steps, even those that are accidental or erroneous. Existing work has studied procedure-aware video representations by modeling the temporal order of actions, but has not explicitly learned the state changes (scene transformations). In this work, we study procedure-aware video representation learning by incorporating state-change descriptions generated by Large Language Models (LLMs) as supervision signals for video encoders. Moreover, we generate state-change counterfactuals that simulate hypothesized failure outcomes, allowing models to learn by imagining unseen "What if" scenarios. This counterfactual reasoning facilitates the model's ability to understand the cause and effect of each step in an activity. We conduct extensive experiments on procedure-aware tasks, including temporal action segmentation, error detection, action phase classification, frame retrieval, multi-instance retrieval, and action recognition. Our results demonstrate the effectiveness of the proposed state-change descriptions and their counterfactuals, and achieve significant improvements on multiple tasks. Chi-Hsi Kung, Frangil Ramirez, Juhyung Ha, Yi-Ting Chen 0001, David Crandall, Yi-Hsuan Tsai |
ICCV | 5 |
| 2025 | Run Like a Neural Network, Explain Like k-Nearest NeighborabstractDeep neural networks have achieved remarkable performance across a variety of applications. However, their decision-making processes are opaque. In contrast, k-nearest neighbor (k-NN) provides interpretable predictions by relying on similar cases, but it lacks important capabilities of neural networks. The neural network k-nearest neighbor (NN-kNN) model is designed to bridge this gap, combining the benefits of neural networks with the instance-based interpretability of k-NN. However, the initial formulation of NN-kNN had limitations including scalability issues, reliance on surface-level features, and an excessive number of parameters. This paper improves NN-kNN by enhancing its scalability, parameter efficiency, ease of integration with feature extractors, and training simplicity. An evaluation of the revised architecture for image and language classification tasks illustrates its promise as a flexible and interpretable method. Xiaomeng Ye, David B. Leake, Yu Wang 0164, David Crandall |
IJCAI | 4 |
| 2025 | Multi-Resolution Guided 3D GANs for Medical Image TranslationabstractMedical image translation is the process of converting from one imaging modality to another, in order to reduce the need for multiple image acquisitions from the same patient. This can enhance the efficiency of treatment by reducing the time, equipment, and labor needed. In this paper, we introduce a multi-resolution guided Generative Adversarial Network (GAN)-based framework for 3D medical image translation. Our framework uses a 3D multi-resolution DenseAttention UNet (3D-mDAUNet) as the generator and a 3D multi-resolution UNet as the discriminator, optimized with a unique combination of loss functions including voxel-wise GAN loss and 2.5D perception loss. Our approach yields promising results in volumetric image quality assessment (IQA) across a variety of imaging modalities, body regions, and age groups, demonstrating its robustness. Furthermore, we propose a synthetic-to-real applicability assessment as an additional evaluation to assess the effectiveness of synthetic data in downstream applications such as segmentation. This comprehensive evaluation shows that our method produces synthetic medical images not only of high-quality but also potentially useful in clinical applications. Our code is available at github.com/juhha/3D-mADUNet. Juhyung Ha, Jong Sung Park, David Crandall, Eleftherios Garyfallidis, Xuhong Zhang 0001 |
WACV | 3 |
| 2025 | Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person PerspectivesabstractWe present Ego-Exo4D, a diverse, large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric video of skilled human activities (e.g., sports, music, dance, bike repair). 740 participants from 13 cities worldwide performed these activities in 123 different natural scene contexts, yielding long-form captures from 1 to 42 minutes each and 1,286 hours of video combined. The multimodal nature of the dataset is unprecedented: the video is accompanied by multichannel audio, eye gaze, 3D point clouds, camera poses, IMU, and multiple paired language descriptions—including a novel “expert commentary” done by coaches and teachers and tailored to the skilled-activity domain. To push the frontier of first-person video understanding of skilled human activity, we also present a suite of benchmark tasks and their annotations, including fine-grained activity understanding, proficiency estimation, cross-view translation, and 3D hand/body pose. All resources are open sourced to fuel new research in the community. https://ego-exo4d-data.org/ Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Makoto Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, Eugene Byrne, Zachary Chavis, Joya Chen, Fu-Jen Chu, Sean Crane, Avijit Dasgupta, Jing Dong 0002, María Escobar, Cristhian Forigua, Abrham Gebreselasie, Sanjay Haresh, Jing Huang 0020, Md Mohaiminul Islam, Suyog Dutt Jain, Rawal Khirodkar, Devansh Kukreja, Kevin J. Liang, Jia-Wei Liu, Sagnik Majumder, Yongsen Mao, Effrosyni Mavroudi, Tushar Nagarajan, Francesco Ragusa, Santhosh K. Ramakrishnan, Luigi Seminara, Arjun Somayazulu, Yale Song, Shan Su, Zihui Xue, Jinxu Zhang, Angela Castillo, Changan Chen, Xinzhu Fu, Ryosuke Furuta, Cristina González, Prince Gupta, Jiabo Hu, Yifei Huang 0002, Yiming Huang 0011, Weslie Khoo, Anush Kumar, Robert Kuo, Sach Lakhavani, Miao Liu 0007, Mi Luo, Zhengyi Luo 0002, Brighid Meredith, Austin Miller, Oluwatumininu Oguntola, Xiaqing Pan, Penny Peng, Shraman Pramanick, Merey Ramazanova, Fiona Ryan, Kiran K. Somasundaram, Chenan Song, Audrey Southerland, Masatoshi Tateno, Takuma Yagi, Mingfei Yan, Xitong Yang, Zecheng Yu, Shengxin Cindy Zha, Chen Zhao 0002, Ziwei Zhao 0003, Zhifan Zhu 0001, Jeff Zhuo, Pablo Andrés Arbeláez, Gedas Bertasius, David Crandall, Dima Damen, Jakob J. Engel, Giovanni Maria Farinella, Antonino Furnari, Bernard Ghanem, Judy Hoffman, C. V. Jawahar, Richard A. Newcombe, Hyun Soo Park, James M. Rehg, Yoichi Sato 0001, Manolis Savva, Jianbo Shi, Mike Zheng Shout, Michael Wray |
Int. J. Comput. Vis. | 86 |
| 2025 | Transformer for Object Re-identification: A Survey
Mang Ye, Shuoyi Chen, Chenyue Li, Wei-Shi Zheng 0001, David Crandall, Bo Du 0001 |
Int. J. Comput. Vis. | 5 |
| 2025 | Blending 3D geometry and machine learning for multi-view stereopsis
Vibhas Vats, Md. Alimoor Reza, David Crandall, Soon-Heung Jung |
Neurocomputing | 3 |
| 2025 | Ego4D: Around the World in 3,600 Hours of Egocentric VideoabstractWe introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countries. The approach to collection is designed to uphold rigorous privacy and ethics standards, with consenting participants and robust de-identification procedures where relevant. Ego4D dramatically expands the volume of diverse egocentric video footage publicly available to the research community. Portions of the video are accompanied by audio, 3D meshes of the environment, eye gaze, stereo, and/or synchronized videos from multiple egocentric cameras at the same event. Furthermore, we present a host of new benchmark challenges centered around understanding the first-person visual experience in the past (querying an episodic memory), present (analyzing hand-object manipulation, audio-visual conversation, and social interactions), and future (forecasting activities). By publicly sharing this massive annotated dataset and benchmark suite, we aim to push the frontier of first-person perception. Kristen Grauman, Andrew Westbury, Eugene Byrne, Vincent Cartillier, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang 0007, Devansh Kukreja, Miao Liu 0007, Xingyu Liu 0001, Tushar Nagarajan, Ilija Radosavovic, Santhosh K. Ramakrishnan, Fiona Ryan, Jayant Sharma 0002, Michael Wray, Mengmeng Xu 0006, Eric Zhongcong Xu, Chen Zhao 0002, Siddhant Bansal, Dhruv Batra, Sean Crane, Tien Do, Morrie Doulaty, Akshay Erapalli, Christoph Feichtenhofer, Adriano Fragomeni, Qichen Fu, Abrham Gebreselasie, Cristina González, James Hillis, Xuhua Huang, Yifei Huang 0002, Wenqi Jia 0001, Weslie Khoo, Jáchym Kolár, Satwik Kottur, Anurag Kumar 0003, Federico Landini, Yanghao Li, Zhenqiang Li 0002, Karttikeya Mangalam, Raghava Modhugu, Jonathan Munro, Tullie Murrell, Takumi Nishiyasu, Will Price, Paola Ruiz Puentes, Merey Ramazanova, Leda Sari, Kiran K. Somasundaram, Audrey Southerland, Yusuke Sugano, Ruijie Tao, Minh Vo, Xindi Wu, Takuma Yagi, Ziwei Zhao 0003, Yunyi Zhu, Pablo Andrés Arbeláez, David Crandall, Dima Damen, Giovanni Maria Farinella, Christian Fügen, Bernard Ghanem, Vamsi K. Ithapu, C. V. Jawahar, Hanbyul Joo, Kris Makoto Kitani, Haizhou Li 0001, Richard A. Newcombe, Aude Oliva, Hyun Soo Park, James M. Rehg, Yoichi Sato 0001, Jianbo Shi, Zheng Shou 0001, Antonio Torralba 0001, Lorenzo Torresani, Mingfei Yan, Jitendra Malik |
IEEE Trans. Pattern Anal. Mach. Intell. | 66 |
| 2025 | Reconstructing High Quality Raw Video Using Temporal Affinity and Diffusion PriorabstractDue to the rich information and original data distribution, RAW data are widely used in many computer vision applications. However, the use of RAW video remains limited because of the high storage costs associated with data collection. Previous works have attempted to reconstruct RAW frames from sRGB data using small sampled metadata from the original RAW frames. Yet, these algorithms struggle with RAW video reconstruction due to the high computational cost of sampling metadata on cameras. To address these issues, we propose a new RAW video reconstruction pipeline that de-renders high-quality RAW videos from sRGB data using only one initial RAW frame as a reference. Specifically, we introduce three new models to achieve this goal. First, we present the Temporal-Affinity Guided De-rendering Network. This network leverages the temporal affinity between adjacent frames to construct a reference RAW image from previous RAW pixels. The corresponding RAW pixels in the previous frame provide valuable information about the original RAW data distribution, aiding in the precise reconstruction of the current frame. Second, to recover the missing RAW pixels caused by camera and foreground movement, we fully exploit the rich prior information from a pre-trained diffusion model and propose the RAW In-painting Model. This model can accurately fill in hollow regions in a RAW image based on the corresponding sRGB image and the surrounding RAW context. Lastly, we present a lightweight content-aware video clipper that automatically adjusts the clip length used for RAW video reconstruction, thereby balancing storage requirements with reconstruction quality. To better evaluate the performance of the proposed framework across different devices, we introduce the first RAW video reconstruction benchmark that comprises RAW videos from six types of camera devices with challenging scenarios. Experimental results demonstrate that our algorithm can accurately reconstruct RAW videos across all the scenarios. Wencheng Han, Jianbing Shen, David Crandall, Cheng-Zhong Xu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Guest Editorial Introduction to the Special Issue on Segment Anything for Videos and Beyond
Wenguan Wang, Hengshuang Zhao, Xinggang Wang, Fisher Yu 0001, David Crandall |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | "Give it Time: " Longitudinal Panels Scaffold Older Adults' Learning and Robot Co-DesignabstractParticipatory robot design projects with older adults often use multiple sessions to encourage design feedback and active participation from users. Prior projects have, however, not analyzed the learning outcomes for older adults across co-design sessions and how they support constructive design feedback and meaningful participation. To bridge this gap, we examined the learning outcomes within a "longitudinal panel." This panel comprised seven co-design sessions with 11 older adults of varying cognitive abilities over six months, aimed at designing a robot to guide a photograph-based conversational activity. Using Nelson and Stolterman's framework of the hierarchy of design-learning, we demonstrate how older adult panelists achieved multiple design-learning outcomes- capacity, confidence, capability, competence, courage, and connection- which allowed them to provide actionable design suggestions. We provide guidelines for conducting longitudinal panels that can enhance user design-learning and participation in robot design. Long-Jing Hsu, Philip B. Stafford, Weslie Khoo, Manasi Swaminathan, Kyrie Jig Amon, Hiroki Sato 0002, Katherine M. Tsui, David Crandall, Selma Sabanovic |
HRI | 8 |
| 2024 | Extracting Indexing Features for CBR from Deep Neural Networks: A Transfer Learning Approach
Zachary Wilkerson, David B. Leake, Vibhas Vats, David Crandall |
ICCBR | 4 |
| 2024 | Towards Network Implementation of CBR: Case Study of a Neural Network K-NN Algorithm
Xiaomeng Ye, David B. Leake, Yu Wang 0164, Ziwei Zhao 0003, David Crandall |
ICCBR | 5 |
| 2024 | SePaint: Semantic Map Inpainting via Multinomial DiffusionabstractPrediction beyond partial observations is crucial for robots to navigate in unknown environments because it can provide extra information regarding the surroundings beyond the current sensing range or resolution. In this work, we consider the inpainting of semantic Bird’s-Eye-View maps. We propose SePaint, an inpainting model for semantic data based on generative multinomial diffusion. To maintain semantic consistency, we need to condition the prediction for the missing regions on the known regions. We propose a novel and efficient condition strategy, Look-Back Condition (LB-Con), which performs one-step look-back operations during the reverse diffusion process. By doing so, we are able to strengthen the harmonization between unknown and known parts, leading to better completion performance. We have conducted extensive experiments on different datasets, showing our proposed model outperforms commonly used interpolation methods in various robotic applications. Zheng Chen 0016, Deepak Duggirala, David Crandall, Lei Jiang 0001, Lantao Liu |
IROS | 3 |
| 2024 | Let's Talk About You: Development and Evaluation of an Autonomous Robot to Support Ikigai Reflection in Older AdultsabstractThe sources of a person’s ikigai—their sense of meaning and purpose in life—often change as they age. Reflecting on past and new sources of ikigai may help people renew their sense of meaning as their life circumstances shift. Building on insights from an initial Wizard-of-Oz robot prototype [1], we describe the design of an autonomous robot that uses a semi-structured conversation format to help older adults reflect on what gives their life meaning and purpose. The robot uses both pre-determined (scripted) and Large Language Model (LLM) generated questions to personalize conversations with older adults around themes of social interaction, planning, accomplishments, goal setting, and the recent past. We evaluated the autonomous robot with 19 older adult participants in a lab setting and at two eldercare facilities. Analysis of the older adults’ conversations with the robot and their responses to an evaluative survey allowed us to identify several design considerations for an autonomous robot that can support ikigai reflection. Interweaving simple yet detailed predetermined questions with LLM-generated follow-up questions yielded enjoyable, in-depth conversations with older adults. We also recognized the need for the robot to be able to offer relevant suggestions when participants cannot recall events and people they find meaningful. These findings aim to further refine the design of an interactive robot that can support users in their exploration of life’s purpose. Long-Jing Hsu, Weslie Khoo, Manasi Swaminathan, Kyrie Jig Amon, Rasika Muralidharan, Hiroki Sato 0002, Min Min Thant, Anna S. Kim, Katherine M. Tsui, David Crandall, Selma Sabanovic |
RO-MAN | 10 |
| 2024 | GC-MVSNet: Multi-View, Multi-Scale, Geometrically-Consistent Multi-View StereoabstractTraditional multi-view stereo (MVS) methods rely heavily on photometric and geometric consistency constraints, but newer machine learning-based MVS methods check geometric consistency across multiple source views only as a post-processing step. In this paper, we present a novel approach that explicitly encourages geometric consistency of reference view depth maps across multiple source views at different scales during learning (see Fig. 1). We find that adding this geometric consistency loss significantly accelerates learning by explicitly penalizing geometrically inconsistent pixels, reducing the training iteration requirements to nearly half that of other MVS methods. Our extensive experiments show that our approach achieves a new state-of-the-art on the DTU and BlendedMVS datasets, and competitive results on the Tanks and Temples benchmark. To the best of our knowledge, GC-MVSNet is the first attempt to enforce multi-view, multi-scale geometric consistency during learning. Vibhas Vats, Sripad Joshi, David Crandall, Md. Alimoor Reza, Soon-Heung Jung |
WACV | 3 |
| 2024 | Attention is All They Need: Exploring the Media Archaeology of the Computer Vision Research PaperabstractResearch papers, in addition to textual documents, are a designed interface through which researchers communicate. Recently, rapid growth has transformed that interface in many fields of computing. In this work, we examine the effects of this growth from a media archaeology perspective, through the changes to figures and tables in research papers. Specifically, we study these changes in computer vision over the past decade, as the deep learning revolution has driven unprecedented growth in the discipline. We ground our investigation through interviews with veteran researchers spanning computer vision, graphics, and visualization. Our analysis focuses on the research attention economy: how research paper elements contribute towards advertising, measuring, and disseminating an increasingly commodified "contribution." Through this work, we seek to motivate future discussion surrounding the design of both the research paper itself as well as the larger sociotechnical research publishing system, including tools for finding, reading, and writing research papers. Samuel Goree, Gabriel Appleby, David Crandall, Norman Makoto Su |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2024 | Asymmetric Convolution: An Efficient and Generalized Method to Fuse Feature Maps in Multiple Vision TasksabstractFusing features from different sources is a critical aspect of many computer vision tasks. Existing approaches can be roughly categorized as parameter-free or learnable operations. However, parameter-free modules are limited in their ability to benefit from offline learning, leading to poor performance in some challenging situations. Learnable fusing methods are often space-consuming and time-consuming, particularly when fusing features with different shapes. To address these shortcomings, we conducted an in-depth analysis of the limitations associated with both fusion methods. Based on our findings, we propose a generalized module named Asymmetric Convolution Module (ACM). This module can learn to encode effective priors during offline training and efficiently fuse feature maps with different shapes in specific tasks. Specifically, we propose a mathematically equivalent method for replacing costly convolutions on concatenated features. This method can be widely applied to fuse feature maps across different shapes. Furthermore, distinguished from parameter-free operations that can only fuse two features of the same type, our ACM is general, flexible, and can fuse multiple features of different types. To demonstrate the generality and efficiency of ACM, we integrate it into several state-of-the-art models on three representative vision tasks. Extensive experimental results on three tasks and several datasets demonstrate that our new module can bring significant improvements and noteworthy efficiency. Wencheng Han, Xingping Dong, David Crandall, Cheng-Zhong Xu 0001, Jianbing Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Reinforcing Generated Images via Meta-Learning for One-Shot Fine-Grained Visual RecognitionabstractOne-shot fine-grained visual recognition often suffers from the problem of having few training examples for new fine-grained classes. To alleviate this problem, off-the-shelf image generation techniques based on Generative Adversarial Networks (GANs) can potentially create additional training images. However, these GAN-generated images are often not helpful for actually improving the accuracy of one-shot fine-grained recognition. In this paper, we propose a meta-learning framework to combine generated images with original images, so that the resulting "hybrid" training images improve one-shot learning. Specifically, the generic image generator is updated by a few training instances of novel classes, and a Meta Image Reinforcing Network (MetaIRNet) is proposed to conduct one-shot fine-grained recognition as well as image reinforcement. Our experiments demonstrate consistent improvement over baselines on one-shot fine-grained image classification benchmarks. Furthermore, our analysis shows that the reinforced images have more diversity compared to the original and GAN-generated images. Satoshi Tsutsui, Yanwei Fu 0001, David Crandall |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Correct for Whom? Subjectivity and the Evaluation of Personalized Image Aesthetics Assessment ModelsabstractThe problem of image aesthetic quality assessment is surprisingly difficult to define precisely. Most early work attempted to estimate the average aesthetic rating of a group of observers, while some recent work has shifted to an approach based on few-shot personalization. In this paper, we connect few-shot personalization, via Immanuel Kant's concept of disinterested judgment, to an argument from feminist aesthetics about the biased tendencies of objective standards for subjective pleasures. To empirically investigate this philosophical debate, we introduce PR-AADB, a relabeling of the existing AADB dataset with labels for pairs of images, and measure how well the existing groundtruth predicts our new pairwise labels. We find, consistent with the feminist critique, that both the existing groundtruth and few-shot personalized predictions represent some users' preferences significantly better than others, but that it is difficult to predict when and for whom the existing groundtruth will be correct. We thus advise against using benchmark datasets to evaluate models for personalized IAQA, and recommend caution when attempting to account for subjective difference using machine learning more generally. Samuel Goree, Weslie Khoo, David Crandall |
AAAI | 3 |
| 2023 | Using manual actions to create visual saliency: an outside-in solution to sustained attention and joint attention
Jane Yang, Linda B. Smith, David Crandall, Chen Yu 0001 |
CogSci | 3 |
| 2023 | VindLU: A Recipe for Effective Video-and-Language PretrainingabstractThe last several years have witnessed remarkable progress in video-and-language (VidL) understanding. However, most modern VidL approaches use complex and specialized model architectures and sophisticated pretraining protocols, making the reproducibility, analysis and comparisons of these frameworks difficult. Hence, instead of proposing yet another new VidL model, this paper conducts a thorough empirical study demystifying the most important factors in the VidL model design. Among the factors that we investigate are (i) the spatiotemporal architecture design, (ii) the multimodal fusion schemes, (iii) the pretraining objectives, (iv) the choice of pretraining data, (v) pretraining and finetuning protocols, and (vi) dataset and model scaling. Our empirical study reveals that the most important design factors include: temporal modeling, video-to-text multimodal fusion, masked modeling objectives, and joint training on images and videos. Using these empirical insights, we then develop a step-by-step recipe, dubbed VindLU, for effective VidL pretraining. Our final model trained using our recipe achieves comparable or better than state-of-the-art results on several VidL tasks without relying on external CLIP pretraining. In particular, on the text-to-video retrieval task, our approach obtains 61.2% on DiDeMo, and 55.0% on ActivityNet, outperforming current SOTA by 7.8% and 6.1% respectively. Furthermore, our model also obtains state-of-the-art video question-answering results on ActivityNet-QA, MSRVTT-QA, MSRVTT-MC and TVQA. Our code and pretrained models are publicly available at: https://github.com/klauscc/VindLU. Jie Lei 0003, David Crandall, Mohit Bansal, Gedas Bertasius |
CVPR | 4 |
| 2023 | Examining the Impact of Network Architecture on Extracted Feature Quality for CBR
David B. Leake, Zachary Wilkerson, Vibhas Vats, Karan Acharya, David Crandall |
ICCBR | 5 |
| 2023 | Few-Shot Segmentation and Semantic Segmentation for Underwater ImageryabstractThis paper tackles image segmentation problems for underwater environments. First, we introduce a novel under-water animal-centric dataset with dense pixel-level annotations containing diverse fine-grained animal categories to mitigate the lack of diverse categories in the existing benchmarks. Then, we solve two image segmentation tasks using underwater images in this dataset: (i) few-shot segmentation, and (ii) semantic segmentation. For the segmentation task in a few-shot learning framework, we propose a novel attention-guided deep neural network architecture by infusing attention modules in various stages of our proposed network. We systematically explore how the learned attention maps can improve few-shot segmentation performance for underwater imagery. Finally, we assess the semantic segmentation problem on our proposed dataset by benchmarking it with two state-of-the-art semantic segmentation methods. We believe our new problem setup, i.e., few-shot segmentation for underwater environments, will be a valuable addition to the existing underwater semantic segmentation task. We believe our novel dataset will pave the way for developing better algorithms and exploring new research directions for marine robotics and underwater image understanding. We publicly release our dataset and the code to advance image understanding research in underwater environments: https://github.com/Imran220S/uwsnet. Imran Kabir, Shubham Shaurya, Vijayalaxmi Maigur, Nikhil Thakurdesai, Mahesh Latnekar, Mayank Raunak, David Crandall, Md. Alimoor Reza |
IROS | 7 |
| 2023 | Finding its Voice: The Influence of Robot Voice on Fit, Social Attributes, and Willingness to Use Among Older Adults in the U.S. and JapanabstractRobots may be able to significantly assist older adults through making activity recommendations. Prior research suggests that gender and age of a robot’s voice may affect how people respond to such recommendations, but few studies have explored how a robot’s voice is perceived by older adults, and whether their perceptions differ across cultures. We conducted a survey study with older adult participants (aged 65+) in the U.S. (N=225) and Japan (N=466), asking them to evaluate a humanoid robot speaking with three different voices (male, female, child). After seeing a video of a robot making recommendations, participants rated the fit of the voice to the robot, its sociality (via the Robotic Social Attributes Scale - RoSAS), and their willingness to use the robot in various contexts. We discovered that robot’s social attributes and participants’ culture impacted willingness to use the robot in both countries. Having positive social attributes and lower negative attributes increases willingness to use the robot. The U.S. older adults preferred the adult robot voices, had more positive social attributes, less negative social attributes, and were more likely to accept lifestyle recommendations than Japanese older adults. This study contributes to our understanding of older adults’ perceptions of robot voice and provides design implications for robots that make recommendations to older adults. Long-Jing Hsu, Weslie Khoo, Natasha Randall, Waki Kamino, Swapna Joshi, Hiroki Sato 0002, David Crandall, Katherine M. Tsui, Selma Sabanovic |
RO-MAN | 7 |
| 2023 | Editorial: Special Section on Egocentric PerceptionabstractThe papers in this special issue focus on egocentric perception. It gathers recent advances in this field that brings together multiple communities including computer vision, machine learning, and multimedia. Antonino Furnari, David Crandall, Dima Damen, Kristen Grauman, Giovanni Maria Farinella |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | DoTA: Unsupervised Detection of Traffic Anomaly in Driving VideosabstractVideo anomaly detection (VAD) has been extensively studied for static cameras but is much more challenging in egocentric driving videos where the scenes are extremely dynamic. This paper proposes an unsupervised method for traffic VAD based on future object localization. The idea is to predict future locations of traffic participants over a short horizon, and then monitor the accuracy and consistency of these predictions as evidence of an anomaly. Inconsistent predictions tend to indicate an anomaly has occurred or is about to occur. To evaluate our method, we introduce a new large-scale benchmark dataset called Detection of Traffic Anomaly (DoTA)containing 4,677 videos with temporal, spatial, and categorical annotations. We also propose a new VAD evaluation metric, called spatial-temporal area under curve (STAUC), and show that it captures how well a model detects both temporal and spatial locations of anomalies unlike existing metrics that focus only on temporal localization. Experimental results show our method outperforms state-of-the-art methods on DoTA in terms of both metrics. We offer rich categorical annotations in DoTA to benchmark video action detection and online action detection methods. The DoTA dataset has been made available at: https://github.com/MoonBlvd/Detection-of-Traffic-Anomaly. Yu Yao 0006, Zelin Pu, Ella M. Atkins, David Crandall |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2023 | Graph Neural Network and Spatiotemporal Transformer Attention for 3D Video Object Detection From Point CloudsabstractPrevious works for LiDAR-based 3D object detection mainly focus on the single-frame paradigm. In this paper, we propose to detect 3D objects by exploiting temporal information in multiple frames, i.e., point cloud videos. We empirically categorize the temporal information into short-term and long-term patterns. To encode the short-term data, we present a Grid Message Passing Network (GMPNet), which considers each grid (i.e., the grouped points) as a node and constructs a k-NN graph with the neighbor grids. To update features for a grid, GMPNet iteratively collects information from its neighbors, thus mining the motion cues in grids from nearby frames. To further aggregate long-term frames, we propose an Attentive Spatiotemporal Transformer GRU (AST-GRU), which contains a Spatial Transformer Attention (STA) module and a Temporal Transformer Attention (TTA) module. STA and TTA enhance the vanilla GRU to focus on small objects and better align moving objects. Our overall framework supports both online and offline video object detection in point clouds. We implement our algorithm based on prevalent anchor-based and anchor-free detectors. Evaluation results on the challenging nuScenes benchmark show superior performance of our method, achieving first on the leaderboard (at the time of paper submission) without any "bells and whistles." Our source code is available at https://github.com/shenjianbing/GMP3D. Junbo Yin, Jianbing Shen, Xin Gao 0001, David Crandall, Ruigang Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | A Survey on Deep Learning Technique for Video SegmentationabstractVideo segmentation-partitioning video frames into multiple segments or objects-plays a critical role in a broad range of practical applications, from enhancing visual effects in movie, to understanding scenes in autonomous driving, to creating virtual background in video conferencing. Recently, with the renaissance of connectionism in computer vision, there has been an influx of deep learning based approaches for video segmentation that have delivered compelling performance. In this survey, we comprehensively review two basic lines of research - generic object segmentation (of unknown categories) in videos, and video semantic segmentation - by introducing their respective task settings, background concepts, perceived need, development history, and main challenges. We also offer a detailed overview of representative literature on both methods and datasets. We further benchmark the reviewed methods on several well-known datasets. Finally, we point out open issues in this field, and suggest opportunities for further research. We also provide a public website to continuously track developments in this fast advancing field: https://github.com/tfzhou/VS-Survey. Tianfei Zhou, Fatih Porikli, David Crandall, Luc Van Gool, Wenguan Wang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | "It Was Really All About Books:" Speech-like Techno-Masculinity in the Rhetoric of Dot-Com Era Web Design BooksabstractThe future of Human-computer interaction (HCI) communication requires researchers to develop a strong understanding of the factors that influence design practitioners. As a step towards building that understanding, based on interviews conducted with veteran web designers, we analyze a corpus of popular web design books published during and shortly after the dot-com boom. Using a combination of ethnographic methods and discourse analysis, we identify the rhetorical strategies in these books and why they were successful in shaping our participants’ ideas about web design. We find that the books exhibit a particular style of technical writing defined by a speech-like techno-masculinity . Despite their short shelf-lives, the books and their writing style contributed to the disciplinary identity of web design which exists today. Studying the history of best practice books is an important opportunity to reflect on the genre of best practices in design, and how we should frame them in the future. Samuel Goree, David Crandall, Norman Makoto Su |
ACM Trans. Comput. Hum. Interact. | 2 |
| 2022 | Grounding Action Verbs in Egocentric Visual Perception
Yayun Zhang, Ellis Cain, David Crandall, Chen Yu 0001 |
CogSci | 3 |
| 2022 | Ego4D: Around the World in 3, 000 Hours of Egocentric VideoabstractWe introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of dailylife activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countries. The approach to collection is designed to uphold rigorous privacy and ethics standards, with consenting participants and robust de-identification procedures where relevant. Ego4D dramatically expands the volume of diverse egocentric video footage publicly available to the research community. Portions of the video are accompanied by audio, 3D meshes of the environment, eye gaze, stereo, and/or synchronized videos from multiple egocentric cameras at the same event. Furthermore, we present a host of new benchmark challenges centered around understanding the first-person visual experience in the past (querying an episodic memory), present (analyzing hand-object manipulation, audio-visual conversation, and social interactions), and future (forecasting activities). By publicly sharing this massive annotated dataset and benchmark suite, we aim to push the frontier of first-person perception. Project page: https://ego4d-data.org/ Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang 0007, Miao Liu 0007, Xingyu Liu 0001, Tushar Nagarajan, Ilija Radosavovic, Santhosh K. Ramakrishnan, Fiona Ryan, Jayant Sharma 0002, Michael Wray, Mengmeng Xu 0006, Eric Zhongcong Xu, Chen Zhao 0002, Siddhant Bansal, Dhruv Batra, Vincent Cartillier, Sean Crane, Tien Do, Morrie Doulaty, Akshay Erapalli, Christoph Feichtenhofer, Adriano Fragomeni, Qichen Fu, Abrham Gebreselasie, Cristina González, James Hillis, Xuhua Huang, Yifei Huang 0002, Wenqi Jia 0001, Weslie Khoo, Jáchym Kolár, Satwik Kottur, Anurag Kumar 0003, Federico Landini, Yanghao Li, Zhenqiang Li 0002, Karttikeya Mangalam, Raghava Modhugu, Jonathan Munro, Tullie Murrell, Takumi Nishiyasu, Will Price, Paola Ruiz Puentes, Merey Ramazanova, Leda Sari, Kiran K. Somasundaram, Audrey Southerland, Yusuke Sugano, Ruijie Tao, Minh Vo, Xindi Wu, Takuma Yagi, Ziwei Zhao 0003, Yunyi Zhu, Pablo Andrés Arbeláez, David Crandall, Dima Damen, Giovanni Maria Farinella, Christian Fügen, Bernard Ghanem, Vamsi K. Ithapu, C. V. Jawahar, Hanbyul Joo, Kris Makoto Kitani, Haizhou Li 0001, Richard A. Newcombe, Aude Oliva, Hyun Soo Park, James M. Rehg, Yoichi Sato 0001, Jianbo Shi, Zheng Shou 0001, Antonio Torralba 0001, Lorenzo Torresani, Mingfei Yan, Jitendra Malik |
CVPR | 65 |
| 2022 | Can Gaze Inform Egocentric Action Recognition?abstractWe investigate the hypothesis that gaze-signal can improve egocentric action recognition on the standard benchmark, EGTEA Gaze++ dataset. In contrast to prior work where gaze-signal was only used during training, we formulate a novel neural fusion approach, Cross-modality Attention Blocks (CMA), to leverage gaze-signal for action recognition during inference as well. CMA combines information from different modalities at different levels of abstraction to achieve state-of-the-art performance for egocentric action recognition. Specifically, fusing the video-stream with optical-flow with CMA outperforms the current state-of-the-art by 3%. However, when CMA is employed to fuse gaze-signal with video-stream data, no improvements are observed. Further investigation of this counter-intuitive finding indicates that small spatial overlap between the network’s attention-map and gaze ground-truth renders the gaze-signal uninformative for this benchmark. Based on our empirical findings, we recommend improvements to the current benchmark to develop practical systems for egocentric video understanding with gaze-signal. David Crandall, Michael J. Proulx, Sachin S. Talathi |
ETRA | 2 |
| 2022 | Extracting Case Indices from Convolutional Neural Networks: A Comparative Study
David B. Leake, Zachary Wilkerson, David Crandall |
ICCBR | 3 |
| 2022 | Case Adaptation with Neural Networks: Capabilities and Limitations
Xiaomeng Ye, David B. Leake, David Crandall |
ICCBR | 3 |
| 2022 | Generation and Evaluation of Creative Images from Limited Data: A Class-to-Class VAE Approach
Xiaomeng Ye, Ziwei Zhao 0003, David B. Leake, David Crandall |
ICCC | 4 |
| 2022 | Semantically Stealthy Adversarial Attacks against Segmentation ModelsabstractSegmentation models have been found to be vulnerable to targeted and non-targeted adversarial attacks. However, the resulting segmentation outputs are often so damaged that it is easy to spot an attack. In this paper, we propose semantically stealthy adversarial attacks which can manipulate targeted labels while preserving non-targeted labels at the same time. One challenge is making semantically meaningful manipulations across datasets and models. Another challenge is avoiding damaging non-targeted labels. To solve these challenges, we consider each input image as prior knowledge to generate perturbations. We also design a special regularizer to help extract features. To evaluate our model’s performance, we design three basic attack types, namely ‘vanishing into the context,’ ‘embedding fake labels,’ and ‘displacing target objects.’ Our experiments show that our stealthy adversarial model can attack segmentation models with a relatively high success rate on Cityscapes, Mapillary, and BDD100K. Our framework shows good empirical generalization across datasets and models. Zhenhua Chen 0003, Chuhua Wang, David Crandall |
WACV | 3 |
| 2022 | Hierarchically Decoupled Spatial-Temporal Contrast for Self-supervised Video Representation LearningabstractWe present a novel technique for self-supervised video representation learning by: (a) decoupling the learning objective into two contrastive subtasks respectively emphasizing spatial and temporal features, and (b) performing it hierarchically to encourage multi-scale understanding. Motivated by their effectiveness in supervised learning, we first introduce spatial-temporal feature learning decoupling and hierarchical learning to the context of unsupervised video learning. We show by experiments that augmentations can be manipulated as regularization to guide the network to learn desired semantics in contrastive learning, and we propose a way for the model to separately capture spatial and temporal features at multiple scales. We also introduce an approach to overcome the problem of divergent levels of instance invariance at different hierarchies by modeling the invariance as loss weights for objective re-weighting. Experiments on downstream action recognition benchmarks on UCF101 and HMDB51 show that our proposed Hierarchically Decoupled Spatial-Temporal Contrast (HDC) makes substantial improvements over directly learning spatial-temporal features as a whole and achieves competitive performance when compared with other state-of-the-art unsupervised methods. Code will be made available. David Crandall |
WACV | 2 |
| 2022 | Segmenting Objects From Relational Visual DataabstractIn this article, we model a set of pixelwise object segmentation tasks - automatic video segmentation (AVS), image co-segmentation (ICS) and few-shot semantic segmentation (FSS) - in a unified view of segmenting objects from relational visual data. To this end, we propose an attentive graph neural network (AGNN) that addresses these tasks in a holistic fashion, by formulating them as a process of iterative information fusion over data graphs. It builds a fully-connected graph to efficiently represent visual data as nodes and relations between data instances as edges. The underlying relations are described by a differentiable attention mechanism, which thoroughly examines fine-grained semantic similarities between all the possible location pairs in two data instances. Through parametric message passing, AGNN is able to capture knowledge from the relational visual data, enabling more accurate object discovery and segmentation. Experiments show that AGNN can automatically highlight primary foreground objects from video sequences (i.e., automatic video segmentation), and extract common objects from noisy collections of semantically related images (i.e., image co-segmentation). AGNN can even generalize segment new categories with little annotated data (i.e., few-shot semantic segmentation). Taken together, our results demonstrate that AGNN provides a powerful tool that is applicable to a wide range of pixel-wise object pattern understanding tasks with relational visual data. Our algorithm implementations have been made publicly available at https://github.com/carrierlxk/AGNN. Xiankai Lu, Wenguan Wang, Jianbing Shen, David Crandall, Luc Van Gool |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Zero-Shot Video Object Segmentation With Co-Attention Siamese NetworksabstractWe introduce a novel network, called CO-attention siamese network (COSNet), to address the zero-shot video object segmentation task in a holistic fashion. We exploit the inherent correlation among video frames and incorporate a global co-attention mechanism to further improve the state-of-the-art deep learning based solutions that primarily focus on learning discriminative foreground representations over appearance and motion in short-term temporal segments. The co-attention layers in COSNet provide efficient and competent stages for capturing global correlations and scene context by jointly computing and appending co-attention responses into a joint feature space. COSNet is a unified and end-to-end trainable framework where different co-attention variants can be derived for capturing diverse properties of the learned joint feature space. We train COSNet with pairs (or groups) of video frames, and this naturally augments training data and allows increased learning capacity. During the segmentation stage, the co-attention model encodes useful information by processing multiple reference frames together, which is leveraged to infer the frequently reappearing and salient foreground objects better. Our extensive experiments over three large benchmarks demonstrate that COSNet outperforms the current alternatives by a large margin. Our implementations are available at https://github.com/carrierlxk/COSNet. Xiankai Lu, Wenguan Wang, Jianbing Shen, David Crandall, Jiebo Luo 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | Investigating the Homogenization of Web Design: A Mixed-Methods ApproachabstractVisual design provides the backdrop to most of our interactions over the Internet, but has not received as much analytical attention as textual content. Combining computational with qualitative approaches, we investigate the growing concern that visual design of the World Wide Web has homogenized over the past decade. By applying computer vision techniques to a large dataset of representative websites images from 2003–2019, we show that designs have become significantly more similar since 2007, especially for page layouts where the average distance between sites decreased by over 30%. Synthesizing interviews from 11 experienced web design professionals with our computational analyses, we discuss causes of this homogenization including overlap in source code and libraries, color scheme standardization, and support for mobile devices. Our results seek to motivate future discussion of the factors that influence designers and their implications on the future trajectory of web design. Samuel Goree, Bardia Doosti, David Crandall, Norman Makoto Su |
CHI | 3 |
| 2021 | In-the-Moment Visual Information from the Infant's Egocentric View Determines the Success of Infant Word Learning: A Computational Study
Andrei Amatuni, Sara E. Schroer, Yayun Zhang, Ryan E. Peters, Md. Alimoor Reza, David Crandall, Chen Yu 0001 |
CogSci | 6 |
| 2021 | Modeling joint attention from egocentric vision
Ryan E. Peters, Andrei Amatuni, Sara E. Schroer, Shujon Naha, David Crandall, Chen Yu 0001 |
CogSci | 5 |
| 2021 | Human Learners Integrate Visual and Linguistic Information Cross-Situational Verb Learning
Yayun Zhang, Andrei Amatuni, Ellis Cain, David Crandall, Chen Yu 0001 |
CogSci | 5 |
| 2021 | The Affective Growth of Computer VisionabstractThe success of deep learning has led to intense growth and interest in computer vision, along with concerns about its potential impact on society. Yet we know little about how these changes have affected the people that research and practice computer vision: we as a community spend so much effort trying to replicate the abilities of humans, but so little time considering the impact of this work on ourselves. In this paper, we report on a study in which we asked computer vision researchers and practitioners to write stories about emotionally-salient events that happened to them. Our analysis of over 50 responses found tremendous affective (emotional) strain in the computer vision community. While many describe excitement and success, we found strikingly frequent feelings of isolation, cynicism, apathy, and exasperation over the state of the field. This is especially true among people who do not share the unbridled enthusiasm for normative standards for computer vision research and who do not see themselves as part of the "incrowd." Our findings suggest that these feelings are closely tied to the kinds of research and professional practices now expected in computer vision. We argue that as a community with significant stature, we need to work towards an inclusive culture that makes transparent and addresses the real emotional toil of its members. Norman Makoto Su, David Crandall |
CVPR | 2 |
| 2021 | On Combining Knowledge-Engineered and Network-Extracted Features for Retrieval
Zachary Wilkerson, David B. Leake, David Crandall |
ICCBR | 3 |
| 2021 | Learning Adaptations for Case-Based Classification: A Neural Network Approach
Xiaomeng Ye, David B. Leake, Vahid Jalali, David Crandall |
ICCBR | 4 |
| 2021 | Deep Tiered Image Segmentation for Detecting Internal ICE Layers in Radar ImageryabstractUnderstanding the structure of Earth’s polar ice sheets is important for modeling how global warming will impact polar ice and, in turn, the Earth’s climate. Ground-penetrating radar is able to collect observations of the internal structure of snow and ice, but the process of manually labeling these observations is slow and laborious. Recent work has developed automatic techniques for finding the boundaries between the ice and the bedrock, but finding internal layers – the subtle boundaries that indicate where one year’s ice accumulation ended and the next began – is much more challenging because the number of layers varies and the boundaries often merge and split. In this paper, we propose a novel deep neural network for solving a general class of tiered segmentation problems. We then apply it to detecting internal layers in polar ice, evaluating on a large-scale dataset of polar ice radar data with human-labeled annotations as ground truth. John Paden, Lora Koenig, Geoffrey C. Fox, David Crandall |
ICME | 6 |
| 2021 | Error Diagnosis of Deep Monocular Depth Estimation ModelsabstractEstimating depth from a monocular image is an ill-posed problem: when the camera projects a 3D scene onto a 2D plane, depth information is inherently and permanently lost. Nevertheless, recent work has shown impressive results in estimating 3D structure from 2D images using deep learning. In this paper, we put on an introspective hat and analyze state-of-the-art monocular depth estimation models in indoor scenes to understand these models’ limitations and error patterns. To address errors in depth estimation, we introduce a novel Depth Error Detection Network (DEDN) that spatially identifies erroneous depth predictions in the monocular depth estimation models. By experimenting with multiple state-of-the-art monocular indoor depth estimation models on multiple datasets, we show that our proposed depth error detection network can identify a significant number of errors in the predicted depth maps. Our module is flexible and can be readily plugged into any monocular depth prediction network to help diagnose its results. Additionally, we propose a simple yet effective Depth Error Correction Network (DECN) that iteratively corrects errors based on our initial error diagnosis. Jagpreet Chawla, Nikhil Thakurdesai, Anuj Godase, Md. Alimoor Reza, David Crandall, Soon-Heung Jung |
IROS | 5 |
| 2021 | Part Segmentation of Unseen Objects using Keypoint GuidanceabstractWhile object part segmentation is useful for many applications, typical approaches require a large amount of labeled data to train a model for good performance. To reduce the labeling effort, weak supervision cues such as object keypoints have been used to generate pseudo-part annotations which can subsequently be used to train larger models. However, previous weakly-supervised part segmentation methods require the same object classes during both training and testing. We propose a new model to use key-point guidance for segmenting parts of novel object classes given that they have similar structures as seen objects - different types of four-legged animals, for example. We show that a non-parametric template matching approach is more effective than pixel classification for part segmentation, especially for small or less frequent parts. To evaluate the generalizability of our approach, we introduce two new datasets that contain 200 quadrupeds in total with both key-point and part segmentation annotations. We show that our approach can outperform existing models by a large margin on the novel object part segmentation task using limited part segmentation labels during training. Shujon Naha, Qingyang Xiao, Prianka Banik, Md. Alimoor Reza, David Crandall |
WACV | 5 |
| 2021 | Whose hand is this? Person Identification from Egocentric Hand GesturesabstractRecognizing people by faces and other biometrics has been extensively studied in computer vision. But these techniques do not work for identifying the wearer of an egocentric (first-person) camera because that person rarely (if ever) appears in their own first-person view. But while one's own face is not frequently visible, their hands are: in fact, hands are among the most common objects in one's own field of view. It is thus natural to ask whether the appearance and motion patterns of people's hands are distinctive enough to recognize them. In this paper, we systematically study the possibility of Egocentric Hand Identification (EHI) with unconstrained egocentric hand gestures. We explore several different visual cues, including color, shape, skin texture, and depth maps to identify users' hands. Extensive ablation experiments are conducted to analyze the properties of hands that are most distinctive. Finally, we show that EHI can improve generalization of other tasks, such as gesture recognition, by training adversarially to encourage these models to ignore differences between users. Satoshi Tsutsui, Yanwei Fu 0001, David Crandall |
WACV | 3 |
| 2020 | Localizing Novel Attended Objects in Egocentric Views
Shujon Naha, Md. Alimoor Reza, Chen Yu 0001, David Crandall |
BMVC | 4 |
| 2020 | A Computational Model of Early Word Learning from the Infant's Point of View
Satoshi Tsutsui, Arjun Chandrasekaran, Md. Alimoor Reza, David Crandall, Chen Yu 0001 |
CogSci | 4 |
| 2020 | HOPE-Net: A Graph-Based Model for Hand-Object Pose EstimationabstractHand-object pose estimation (HOPE) aims to jointly detect the poses of both a hand and of a held object. In this paper, we propose a lightweight model called HOPE-Net which jointly estimates hand and object pose in 2D and 3D in real-time. Our network uses a cascade of two adaptive graph convolutional neural networks, one to estimate 2D coordinates of the hand joints and object corners, followed by another to convert 2D coordinates to 3D. Our experiments show that through end-to-end training of the full network, we achieve better accuracy for both the 2D and 3D coordinate estimation problems. The proposed 2D to 3D graph convolution-based model could be applied to other 3D landmark detection problems, where it is possible to first predict the 2D keypoints and then transform them to 3D. Bardia Doosti, Shujon Naha, Majid Mirbagheri, David Crandall |
CVPR | 4 |
| 2020 | Learning Video Object Segmentation From Unlabeled VideosabstractWe propose a new method for video object segmentation (VOS) that addresses object pattern learning from unlabeled videos, unlike most existing methods which rely heavily on extensive annotated data. We introduce a unified unsupervised/weakly supervised learning framework, called MuG, that comprehensively captures intrinsic properties of VOS at multiple granularities. Our approach can help advance understanding of visual patterns in VOS and significantly reduce annotation burden. With a carefully-designed architecture and strong representation learning ability, our learned model can be applied to diverse VOS settings, including object-level zero-shot VOS, instance-level zero-shot VOS, and one-shot VOS. Experiments demonstrate promising performance in these settings, as well as the potential of MuG in leveraging unlabeled data to further improve the segmentation accuracy. Xiankai Lu, Wenguan Wang, Jianbing Shen, Yu-Wing Tai, David Crandall, Steven C. H. Hoi |
CVPR | 5 |
| 2020 | Dynamic Dual-Attentive Aggregation Learning for Visible-Infrared Person Re-identification
Mang Ye, Jianbing Shen, David Crandall, Ling Shao 0001, Jiebo Luo 0001 |
ECCV (17) | 3 |
| 2020 | On Bringing Case-Based Reasoning Methodology to Deep Learning
David B. Leake, David Crandall |
ICCBR | 2 |
| 2020 | Interaction Graphs for Object Importance Estimation in On-road Driving VideosabstractA vehicle driving along the road is surrounded by many objects, but only a small subset of them influence the driver's decisions and actions. Learning to estimate the importance of each object on the driver's real-time decision-making may help better understand human driving behavior and lead to more reliable autonomous driving systems. Solving this problem requires models that understand the interactions between the ego-vehicle and the surrounding objects. However, interactions among other objects in the scene can potentially also be very helpful, e.g., a pedestrian beginning to cross the road between the ego-vehicle and the car in front will make the car in front less important. We propose a novel framework for object importance estimation using an interaction graph, in which the features of each object node are updated by interacting with others through graph convolution. Experiments show that our model outperforms state-of-the-art baselines with much less input and pre-processing. Ashish Tawari, Sujitha Martin, David Crandall |
ICRA | 4 |
| 2020 | Snow Radar Layer Tracking Using Iterative Neural Network ApproachabstractThis paper presents preliminary results using a fully connected neural network (NN) to automatically track the internal layers of snow radar echograms using an iterative “row-block-column” approach. Snow radar images, when accurately tracked, provide relevant information for estimating snow accumulation rates in polar regions which is a key measurement needed to understand and predict the impact of climate warming in Greenland and Antarctica. A multiclass NN was designed and trained with a training set of 121,408 columns of simulated snow radar data and learns to automatically track the internal layers with an accuracy of 92.8%, a RMSE of 0.24 pixels, and with 98% of pixel errors less than or equal to 1 pixel. Ibikunle Oluwanisola, John Paden, Maryam Rahnemoonfar, David Crandall, Masoud Yari |
IGARSS | 4 |
| 2020 | Automatically Detecting Bystanders in Photos to Reduce Privacy RisksabstractPhotographs taken in public places often contain bystanders - people who are not the main subject of a photo. These photos, when shared online, can reach a large number of viewers and potentially undermine the bystanders' privacy. Furthermore, recent developments in computer vision and machine learning can be used by online platforms to identify and track individuals. To combat this problem, researchers have proposed technical solutions that require bystanders to be proactive and use specific devices or applications to broadcast their privacy policy and identifying information to locate them in an image.We explore the prospect of a different approach - identifying bystanders solely based on the visual information present in an image. Through an online user study, we catalog the rationale humans use to classify subjects and bystanders in an image, and systematically validate a set of intuitive concepts (such as intentionally posing for a photo) that can be used to automatically identify bystanders. Using image data, we infer those concepts and then use them to train several classifier models. We extensively evaluate the models and compare them with human raters. On our initial dataset, with a 10-fold cross validation, our best model achieves a mean detection accuracy of 93% for images when human raters have 100% agreement on the class label and 80% when the agreement is only 67%. We validate this model on a completely different dataset and achieve similar results, demonstrating that our model generalizes well. Rakibul Hasan 0001, David Crandall, Mario Fritz, Apu Kapadia |
SP | 2 |
| 2020 | Privacy Norms and Preferences for Photos Posted OnlineabstractWe are surrounded by digital images of personal lives posted online. Changes in information and communications technology have enabled widespread sharing of personal photos, increasing access to aspects of private life previously less observable. Most studies of privacy online explore differences in individual privacy preferences. Here we examine privacy perceptions of online photos considering both social norms, collectively—shared expectations of privacy and individual preferences. We conducted an online factorial vignette study on Amazon’s Mechanical Turk ( n = 279). Our findings show that people share common expectations about the privacy of online images, and these privacy norms are socially contingent and multidimensional. Use of digital technologies to share personal photos is influenced by social context as well as individual preferences, while such sharing can affect the social meaning of privacy. Roberto Hoyle, Luke Stark, Qatrunnada Ismail, David Crandall, Apu Kapadia, Denise L. Anthony |
ACM Trans. Comput. Hum. Interact. | 4 |
| 2019 | Can Privacy Be Satisfying?: On Improving Viewer Satisfaction for Privacy-Enhanced Photos Using Aesthetic TransformsabstractPervasive photo sharing in online social media platforms can cause unintended privacy violations when elements of an image reveal sensitive information. Prior studies have identified image obfuscation methods (e.g., blurring) to enhance privacy, but many of these methods adversely affect viewers' satisfaction with the photo, which may cause people to avoid using them. In this paper, we study the novel hypothesis that it may be possible to restore viewers' satisfaction by 'boosting' or enhancing the aesthetics of an obscured image, thereby compensating for the negative effects of a privacy transform. Using a between-subjects online experiment, we studied the effects of three artistic transformations on images that had objects obscured using three popular obfuscation methods validated by prior research. Our findings suggest that using artistic transformations can mitigate some negative effects of obfuscation methods, but more exploration is needed to retain viewer satisfaction. Rakibul Hasan 0001, Yifang Li, Eman T. Hassan, Kelly Caine, David Crandall, Roberto Hoyle, Apu Kapadia |
CHI | 5 |
| 2019 | How do infants start learning object names in a sea of clutter?
Hadar Karmazyn Raz, Drew H. Abney, David Crandall, Chen Yu 0001, Linda B. Smith |
CogSci | 3 |
| 2019 | Semantic structure of infant first-person scenes changes with development
Linda B. Smith, David Crandall |
CogSci | 3 |
| 2019 | Zero-Shot Video Object Segmentation via Attentive Graph Neural NetworksabstractThis work proposes a novel attentive graph neural network (AGNN) for zero-shot video object segmentation (ZVOS). The suggested AGNN recasts this task as a process of iterative information fusion over video graphs. Specifically, AGNN builds a fully connected graph to efficiently represent frames as nodes, and relations between arbitrary frame pairs as edges. The underlying pair-wise relations are described by a differentiable attention mechanism. Through parametric message passing, AGNN is able to efficiently capture and mine much richer and higher-order relations between video frames, thus enabling a more complete understanding of video content and more accurate foreground estimation. Experimental results on three video segmentation datasets show that AGNN sets a new state-of-the-art in each case. To further demonstrate the generalizability of our framework, we extend AGNN to an additional task: image object co-segmentation (IOCS). We perform experiments on two famous IOCS datasets and observe again the superiority of our AGNN model. The extensive experiments verify that AGNN is able to learn the underlying semantic/appearance relationships among video frames or related images, and discover the common objects. Wenguan Wang, Xiankai Lu, Jianbing Shen, David Crandall, Ling Shao 0001 |
ICCV | 4 |
| 2019 | Temporal Recurrent Networks for Online Action DetectionabstractMost work on temporal action detection is formulated as an offline problem, in which the start and end times of actions are determined after the entire video is fully observed. However, important real-time applications including surveillance and driver assistance systems require identifying actions as soon as each video frame arrives, based only on current and historical observations. In this paper, we propose a novel framework, the Temporal Recurrent Network (TRN), to model greater temporal context of each frame by simultaneously performing online action detection and anticipation of the immediate future. At each moment in time, our approach makes use of both accumulated historical evidence and predicted future information to better recognize the action that is currently occurring, and integrates both of these into a unified end-to-end architecture. We evaluate our approach on two popular online action detection datasets, HDD and TVSeries, as well as another widely used dataset, THUMOS'14. The results show that TRN significantly outperforms the state-of-the-art. Mingfei Gao, Yi-Ting Chen 0001, Larry Davis 0001, David Crandall |
ICCV | 5 |
| 2019 | Embodied Amodal Recognition: Learning to Move to Perceive ObjectsabstractPassive visual systems typically fail to recognize objects in the amodal setting where they are heavily occluded. In contrast, humans and other embodied agents have the ability to move in the environment and actively control the viewing angle to better understand object shapes and semantics. In this work, we introduce the task of Embodied Amodel Recognition (EAR): an agent is instantiated in a 3D environment close to an occluded target object, and is free to move in the environment to perform object classification, amodal object localization, and amodal object segmentation. To address this problem, we develop a new model called Embodied Mask R-CNN for agents to learn to move strategically to improve their visual recognition abilities. We conduct experiments using a simulator for indoor environments. Experimental results show that: 1) agents with embodiment (movement) achieve better visual recognition performance than passive ones and 2) in order to improve visual recognition abilities, agents can learn strategic paths that are different from shortest paths. Zhile Ren, Xinlei Chen, David Crandall, Devi Parikh, Dhruv Batra |
ICCV | 5 |
| 2019 | Egocentric Vision-based Future Vehicle Localization for Intelligent Driving Assistance SystemsabstractPredicting the future location of vehicles is essential for safety-critical applications such as advanced driver assistance systems (ADAS) and autonomous driving. This paper introduces a novel approach to simultaneously predict both the location and scale of target vehicles in the first-person (egocentric) view of an ego-vehicle. We present a multi-stream recurrent neural network (RNN) encoder-decoder model that separately captures both object location and scale and pixel-level observations for future vehicle localization. We show that incorporating dense optical flow improves prediction results significantly since it captures information about motion as well as appearance change. We also find that explicitly modeling future motion of the ego-vehicle improves the prediction accuracy, which could be especially beneficial in intelligent and automated vehicles that have motion planning capability. To evaluate the performance of our approach, we present a new dataset of first-person videos collected from a variety of scenarios at road intersections, which are particularly challenging moments for prediction because vehicle trajectories are diverse and dynamic. Code and dataset have been made available at: https://usa.honda-ri.com/hevi. Yu Yao 0006, Chiho Choi, David Crandall, Ella M. Atkins, Behzad Dariush |
ICRA | 4 |
| 2019 | Automatic Annotation for Semantic Segmentation in Indoor ScenesabstractDomestic robots could eventually transform our lives, but safely operating in home environments requires a rich understanding of indoor scenes. Learning-based techniques for scene segmentation require large-scale, pixel-level annotations, which are laborious and expensive to collect. We propose an automatic method for pixel-wise semantic annotation of video sequences, that gathers cues from object detectors and indoor 3D room-layout estimation and then annotates all the image pixels in an energy minimization framework. Extensive experiments on a publicly available video dataset (SUN3D) evaluate the approach and demonstrate its effectiveness. Md. Alimoor Reza, Akshay U. Naik, David Crandall |
IROS | 4 |
| 2019 | Unsupervised Traffic Accident Detection in First-Person VideosabstractRecognizing abnormal events such as traffic violations and accidents in natural driving scenes is essential for successful autonomous driving and advanced driver assistance systems. However, most work on video anomaly detection suffers from two crucial drawbacks. First, they assume cameras are fixed and videos have static backgrounds, which is reasonable for surveillance applications but not for vehicle-mounted cameras. Second, they pose the problem as one-class classification, relying on arduously hand-labeled training datasets that limit recognition to anomaly categories that have been explicitly trained. This paper proposes an unsupervised approach for traffic accident detection in first-person (dashboard-mounted camera) videos. Our major novelty is to detect anomalies by predicting the future locations of traffic participants and then monitoring the prediction accuracy and consistency metrics with three different strategies. We evaluate our approach using a new dataset of diverse traffic accidents, AnAn Accident Detection (A3D), as well as another publicly-available dataset. Experimental results show that our approach outperforms the state-of-the-art. Code and the dataset developed in this work are available at: https:llgithub.comlMoonBtvdltad-IROS2019. Yu Yao 0006, David Crandall, Ella M. Atkins |
IROS | 4 |
| 2019 | Meta-Reinforced Synthetic Data for One-Shot Fine-Grained Visual RecognitionabstractThis paper studies the task of one-shot fine-grained recognition, which suffers from the problem of data scarcity of novel fine-grained classes. To alleviate this problem, a off-the-shelf image generator can be applied to synthesize additional images to help one-shot learning. However, such synthesized images may not be helpful in one-shot fine-grained recognition, due to a large domain discrepancy between synthesized and original images. To this end, this paper proposes a meta-learning framework to reinforce the generated images by original images so that these images can facilitate one-shot learning. Specifically, the generic image generator is updated by few training instances of novel classes; and a Meta Image Reinforcing Network (MetaIRNet) is proposed to conduct one-shot fine-grained recognition as well as image reinforcement. The model is trained in an end-to-end manner, and our experiments demonstrate consistent improvement over baseline on one-shot fine-grained image classification benchmarks. Satoshi Tsutsui, Yanwei Fu 0001, David Crandall |
NeurIPS | 3 |
| 2019 | A Self Validation Network for Object-Level Human Attention EstimationabstractDue to the foveated nature of the human vision system, people can focus their visual attention on a small region of their visual field at a time, which usually contains only a single object. Estimating this object of attention in first-person (egocentric) videos is useful for many human-centered real-world applications such as augmented reality applications and driver assistance systems. A straightforward solution for this problem is to pick the object whose bounding box is hit by the gaze, where eye gaze point estimation is obtained from a traditional eye gaze estimator and object candidates are generated from an off-the-shelf object detector. However, such an approach can fail because it addresses the where and the what problems separately, despite that they are highly related, chicken-and-egg problems. In this paper, we propose a novel unified model that incorporates both spatial and temporal evidence in identifying as well as locating the attended object in firstperson videos. It introduces a novel Self Validation Module that enforces and leverages consistency of the where and the what concepts. We evaluate on two public datasets, demonstrating that Self Validation Module significantly benefits both training and testing and that our model outperforms the state-of-the-art. Chen Yu 0001, David Crandall |
NeurIPS | 3 |
| 2019 | Hello Research! Developing an Intensive Research Experience for Undergraduate WomenabstractThis paper describes the design and implementation of a three-day intensive research experience (IRE) workshop for undergraduate women in Computer Science. Expanding on a model pioneered at Carnegie Mellon University, we developed and piloted a regional variant called HelloResearch at Indiana University. Participants were actively recruited from our own and neighboring states. Industry partners provided travel scholarships for low-income and first-generation college students, people with disabilities, and students at Historically Black Colleges and Universities (HBCUs) across the country. The primary goal of HelloResearch was to encourage the pursuit of research careers, enabling participants to reach the highest levels of leadership in their fields. In this paper, we report on the demographics of our 92 participants, outline best practices to ensure an authentic short-term research experience for the students, describe our assessment plans, and share our survey instruments to assist others in jump-starting their own regional workshops. Suzanne Menzel, Katie A. Siek, David Crandall |
SIGCSE | 3 |
| 2019 | Observing Pianist Accuracy and Form with Computer VisionabstractWe present a first step towards developing an interactive piano tutoring system that can observe a student playing the piano and give feedback about hand movements and musical accuracy. In particular, we have two primary aims: 1) to determine which notes on a piano are being played at any moment in time, 2) to identify which finger is pressing each note. We introduce a novel two-stream convolutional neural network that takes video and audio inputs together for detecting pressed notes and finger presses. We formulate our two problems in terms of multi-task learning and extend a state-of-the-art object detection model to incorporate both audio and visual features. In addition, we introduce a novel finger identification solution based on pressed piano note information. We experimentally confirm that our approach is able to detect pressed piano keys and the piano player's fingers with a high accuracy. Jangwon Lee 0002, Bardia Doosti, Yupeng Gu, David Cartledge, David Crandall, Christopher Raphael |
WACV | 5 |
| 2018 | Diverse Beam Search for Improved Description of Complex ScenesabstractA single image captures the appearance and position of multiple entities in a scene as well as their complex interactions. As a consequence, natural language grounded in visual contexts tends to be diverse---with utterances differing as focus shifts to specific objects, interactions, or levels of detail. Recently, neural sequence models such as RNNs and LSTMs have been employed to produce visually-grounded language. Beam Search, the standard work-horse for decoding sequences from these models, is an approximate inference algorithm that decodes the top-B sequences in a greedy left-to-right fashion. In practice, the resulting sequences are often minor rewordings of a common utterance, failing to capture the multimodal nature of source images. To address this shortcoming, we propose Diverse Beam Search (DBS), a diversity promoting alternative to BS for approximate inference. DBS produces sequences that are significantly different from each other by incorporating diversity constraints within groups of candidate sequences during decoding; moreover, it achieves this with minimal computational or memory overhead. We demonstrate that our method improves both diversity and quality of decoded sequences over existing techniques on two visually-grounded language generation tasks---image captioning and visual question generation---particularly on complex scenes containing diverse visual content. We also show similar improvements at language-only machine translation tasks, highlighting the generality of our approach. Ashwin K. Vijayakumar, Michael Cogswell, Ramprasaath R. Selvaraju, Qing Sun 0001, Stefan Lee, David Crandall, Dhruv Batra |
AAAI | 6 |
| 2018 | From Coarse Attention to Fine-Grained Gaze: A Two-stage 3D Fully Convolutional Network for Predicting Eye Gaze in First Person Video
David Crandall, Chen Yu 0001, Sven Bambach |
BMVC | 2 |
| 2018 | Viewer Experience of Obscuring Scene Elements in Photos to Enhance PrivacyabstractWith the rise of digital photography and social networking, people are sharing personal photos online at an unprecedented rate. In addition to their main subject matter, photographs often capture various incidental information that could harm people's privacy. While blurring and other image filters may help obscure private content, they also often affect the utility and aesthetics of the photos, which is important since images shared in social media are mainly for human consumption. Existing studies of privacy-enhancing image filters either primarily focus on obscuring faces, or do not systematically study how filters affect image utility. To understand the trade-offs when obscuring various sensitive aspects of images, we study eleven filters applied to obfuscate twenty different objects and attributes, and evaluate how effectively they protect privacy and preserve image quality for human viewers. Rakibul Hasan 0001, Eman T. Hassan, Yifang Li, Kelly Caine, David Crandall, Roberto Hoyle, Apu Kapadia |
CHI | 5 |
| 2018 | Joint Person Segmentation and Identification in Synchronized First- and Third-Person Videos
Chenyou Fan, Michael S. Ryoo, David Crandall |
ECCV (1) | 5 |
| 2018 | Estimating Head Motion from Egocentric VisionabstractThe recent availability of lightweight, wearable cameras allows for collecting video data from a "first-person' perspective, capturing the visual world of the wearer in everyday interactive contexts. In this paper, we investigate how to exploit egocentric vision to infer multimodal behaviors from people wearing head-mounted cameras. More specifically, we estimate head (camera) motion from egocentric video, which can be further used to infer non-verbal behaviors such as head turns and nodding in multimodal interactions. We propose several approaches based on Convolutional Neural Networks (CNNs) that combine raw images and optical flow fields to learn to distinguish regions with optical flow caused by global ego-motion from those caused by other motion in a scene. Our results suggest that CNNs do not directly learn useful visual features with end-to-end training from raw images alone; instead, a better approach is to first extract optical flow explicitly and then train CNNs to integrate optical flow and visual information. Satoshi Tsutsui, Sven Bambach, David Crandall, Chen Yu 0001 |
ICMI | 3 |
| 2018 | Automated Tracking of 2D and 3D Ice Radar Imagery Using Viterbi and TRW-SabstractWe present improvements to existing implementations of the Viterbi and TRW-S algorithms applied to ice-bottom layer tracking on 2D and 3D radar imagery, respectively. Along with an explanation of our modifications and the reasoning behind them, we present a comparison between our results, the results obtained with the original implementations, and those obtained with other proposed methods of performing ice-bottom layer tracking. Victor Berger, Shane Chu, David Crandall, John Paden, Geoffrey C. Fox |
IGARSS | 4 |
| 2018 | Toddler-Inspired Visual Object LearningabstractReal-world learning systems have practical limitations on the quality and quantity of the training datasets that they can collect and consider. How should a system go about choosing a subset of the possible training examples that still allows for learning accurate, generalizable models? To help address this question, we draw inspiration from a highly efficient practical learning system: the human child. Using head-mounted cameras, eye gaze trackers, and a model of foveated vision, we collected first-person (egocentric) images that represents a highly accurate approximation of the "training data" that toddlers' visual systems collect in everyday, naturalistic learning contexts. We used state-of-the-art computer vision learning models (convolutional neural networks) to help characterize the structure of these data, and found that child data produce significantly better object models than egocentric data experienced by adults in exactly the same environment. By using the CNNs as a modeling tool to investigate the properties of the child data that may enable this rapid learning, we found that child data exhibit a unique combination of quality and diversity, with not only many similar large, high-quality object views but also a greater number and diversity of rare views. This novel methodology of analyzing the visual "training data" used by children may not only reveal insights to improve machine learning, but also may suggest new experimental tools to better understand infant learning in developmental psychology. Sven Bambach, David Crandall, Linda B. Smith, Chen Yu 0001 |
NeurIPS | 2 |
| 2018 | Multi-task Spatiotemporal Neural Networks for Structured Surface ReconstructionabstractDeep learning methods have surpassed the performance of traditional techniques on a wide range of problems in computer vision, but nearly all of this work has studied consumer photos, where precisely correct output is often not critical. It is less clear how well these techniques may apply on structured prediction problems where fine-grained output with high precision is required, such as in scientific imaging domains. Here we consider the problem of segmenting echogram radar data collected from the polar ice sheets, which is challenging because segmentation boundaries are often very weak and there is a high degree of noise. We propose a multi-task spatiotemporal neural network that combines 3D ConvNets and Recurrent Neural Networks (RNNs) to estimate ice surface boundaries from sequences of tomographic radar images. We show that our model outperforms the state-of-the-art on this problem by (1) avoiding the need for hand-tuned parameters, (2) extracting multiple surfaces (ice-air and ice-bed) simultaneously, (3) requiring less non-visual metadata, and (4) being about 6 times faster. Chenyou Fan, John Paden, Geoffrey C. Fox, David Crandall |
WACV | 5 |
| 2018 | Fully-Coupled Two-Stream Spatiotemporal Networks for Extremely Low Resolution Action RecognitionabstractA major emerging challenge is how to protect people's privacy as cameras and computer vision are increasingly integrated into our daily lives, including in smart devices inside homes. A potential solution is to capture and record just the minimum amount of information needed to perform a task of interest. In this paper, we propose a fully-coupled two-stream spatiotemporal architecture for reliable human action recognition on extremely low resolution (e.g., 1216 pixel) videos. We provide an efficient method to extract spatial and temporal features and to aggregate them into a robust feature representation for an entire action video sequence. We also consider how to incorporate high resolution videos during training in order to build better low resolution action recognition models. We evaluate on two publicly-available datasets, showing significant improvements over the state-of-the-art. Aidean Sharghi, Xin Chen 0071, David Crandall |
WACV | 4 |
| 2018 | Introduction to the special issue: Egocentric Vision and Lifelogging
Mariella Dimiccoli, Cathal Gurrin, David Crandall, Xavier Giró-i-Nieto, Petia Radeva |
J. Vis. Commun. Image Represent. | 3 |
| 2018 | Deepdiary: Lifelogging image captioning and summarization
Chenyou Fan, David Crandall |
J. Vis. Commun. Image Represent. | 3 |
| 2017 | Understanding the Aesthetic Evolution of Websites: Towards a Notion of Design PeriodsabstractIn art and music, time periods like "classical" and "impressionist" are powerful means for academics and practitioners to compare and contrast artifacts that share aesthetics or philosophies. While web designs have undergone changes for 25 years, we lack theories to describe or explain these changes. In this paper, we take a first step towards identifying and understanding the design periods of websites. Drawing from humanistic HCI methods, we asked subject experts of web design to critically analyze a dataset of prominent websites whose lifetimes span over a decade. These informed judgments reveal a set of key markers that signal shifts in design periods. For instance, advances in display technologies and changes in company strategies help explain how design periods demarcated by particular layout templates and navigation models arise. We suggest that designers and marketers can draw inspiration from website designs curated into design periods. Future work should examine the utility of applying design periods to any computationally embedded artifact that is an interaction design. David Crandall, Norman Makoto Su |
CHI | 2 |
| 2017 | Identifying First-Person Camera Wearers in Third-Person VideosabstractWe consider scenarios in which we wish to perform joint scene understanding, object tracking, activity recognition, and other tasks in scenarios in which multiple people are wearing body-worn cameras while a third-person static camera also captures the scene. To do this, we need to establish person-level correspondences across first-and third-person videos, which is challenging because the camera wearer is not visible from his/her own egocentric video, preventing the use of direct feature matching. In this paper, we propose a new semi-Siamese Convolutional Neural Network architecture to address this novel challenge. We formulate the problem as learning a joint embedding space for first-and third-person videos that considers both spatial-and motion-domain cues. A new triplet loss function is designed to minimize the distance between correct first-and third-person matches while maximizing the distance between incorrect ones. This end-to-end approach performs significantly better than several baselines, in part by learning the first-and third-person features optimized for matching jointly with the distance measure itself. Chenyou Fan, Jangwon Lee 0002, Krishna Kumar Singh, Yong Jae Lee, David Crandall, Michael S. Ryoo |
CVPR | 6 |
| 2017 | A Unified Model for Near and Remote SensingabstractWe propose a novel convolutional neural network architecture for estimating geospatial functions such as population density, land cover, or land use. In our approach, we combine overhead and ground-level images in an end-toend trainable neural network, which uses kernel regression and density estimation to convert features extracted from the ground-level images into a dense feature map. The output of this network is a dense estimate of the geospatial function in the form of a pixel-level labeling of the overhead image. To evaluate our approach, we created a large dataset of overhead and ground-level images from a major urban area with three sets of labels: land use, building function, and building age. We find that our approach is more accurate for all tasks, in some cases dramatically so. Scott Workman, Menghua Zhai, David Crandall, Nathan Jacobs |
ICCV | 3 |
| 2017 | A Data Driven Approach for Compound Figure Separation Using Convolutional Neural NetworksabstractA key problem in automatic analysis and understanding of scientific papers is to extract semantic information from non-textual paper components like figures, diagrams, tables, etc. Much of this work requires a very first preprocessing step: decomposing compound multi-part figures into individual sub-figures. Previous work in compound figure separation has been based on manually designed features and separation rules, which often fail for less common figure types and layouts. Moreover, few implementations for compound figure decomposition are publicly available. This paper proposes a data driven approach to separate compound figures using modern deep Convolutional Neural Networks (CNNs) to train the separator in an end-to-end manner. CNNs eliminate the need for manually designing features and separation rules, but require a large amount of annotated training data. We overcome this challenge using transfer learning as well as automatically synthesizing training exemplars. We evaluate our technique on the ImageCLEF Medical dataset, achieving 85.9% accuracy and outperforming previous techniques. We have released our implementation as an easy-to-use Python library, aiming to promote further research in scientific figure mining. Satoshi Tsutsui, David Crandall |
ICDAR | 2 |
| 2017 | Automatic estimation of ice bottom surfaces from radar imageryabstractGround-penetrating radar on planes and satellites now makes it practical to collect 3D observations of the subsurface structure of the polar ice sheets, providing crucial data for understanding and tracking global climate change. But converting these noisy readings into useful observations is generally done by hand, which is impractical at a continental scale. In this paper, we propose a computer vision-based technique for extracting 3D ice-bottom surfaces by viewing the task as an inference problem on a probabilistic graphical model. We first generate a seed surface subject to a set of constraints, and then incorporate additional sources of evidence to refine it via discrete energy minimization. We evaluate the performance of the tracking algorithm on 7 topographic sequences (each with over 3000 radar images) collected from the Canadian Arctic Archipelago with respect to human-labeled ground truth. David Crandall, Geoffrey C. Fox, John Paden |
ICIP | 2 |
| 2017 | DEM extraction of the basal topography of the Canadian archipelago ICE caps via 2D automated layer-trackerabstractThe basal topography of most of the glaciers that drain the ice caps of the Canadian Arctic Archipelago is largely unknown. To measure the basal topography, NASA Operation IceBridge flew a radar depth sounder in a wide swath mode with three transmit beams to image the glacier beds during three flights over the archipelago in 2014. We describe the measurement setup of the radar system, the algorithms used to process the data to produce a 3D image of the glacier bed, show digital elevation model (DEM) results of the beds, and provide a basic assessment of the tracking algorithm used to extract the DEM. Mohanad Al-Ibadi, Jordan Sprick, Sravya Athinarapu, Theresa Stumpf, John Paden, Carlton J. Leuschen, Fernando Rodriguez-Morales, David Crandall, Geoffrey C. Fox, David Burgess, Martin Sharp, Luke Copland, Wesley Van Wychen |
IGARSS | 9 |
| 2016 | Enhancing Lifelogging Privacy by Detecting ScreensabstractLow-cost, lightweight wearable cameras let us record (or 'lifelog') our lives from a 'first-person' perspective for purposes ranging from fun to therapy. But they also capture private information that people may not want to be recorded, especially if images are stored in the cloud or visible to other people. For example, recent studies suggest that computer screens may be lifeloggers' single greatest privacy concern, because many people spend a considerable amount of time in front of devices that display private information. In this paper, we investigate using computer vision to automatically detect computer screens in photo lifelogs. We evaluate our approach on an existing in-situ dataset of 36 people who wore cameras for a week, and show that our technique could help manage privacy in the upcoming era of wearable cameras. Mohammed Korayem, Robert Templeman, Dennis Chen, David Crandall, Apu Kapadia |
CHI | 4 |
| 2016 | Active Viewing in Toddlers Facilitates Visual Object Learning: An Egocentric Vision Approach
Sven Bambach, David Crandall, Linda B. Smith, Chen Yu 0001 |
CogSci | 2 |
| 2016 | Tracking Natural Events through Social Media and Computer VisionabstractAccurate, efficient, global observation of natural events is important for ecologists, meteorologists, governments, and the public. Satellites are effective but limited by their perspective and by atmospheric conditions. Public images on photo-sharing websites could provide crowd-sourced ground data to complement satellites, since photos contain evidence of the state of the natural world. In this work, we test the ability of computer vision to observe natural events in millions of geo-tagged Flickr photos, over nine years and an entire continent. We use satellites as (noisy) ground truth to train two types of classifiers, one that estimates if a Flickr photo has evidence of an event, and one that aggregates these estimates to produce an observation for given times and places. We present a web tool for visualizing the satellite and photo observations, allowing scientists to explore this novel combination of data sources. Mohammed Korayem, Saúl A. Blanco, David Crandall |
ACM Multimedia | 4 |
| 2016 | Stochastic Multiple Choice Learning for Training Diverse Deep EnsemblesabstractMany practical perception systems exist within larger processes which often include interactions with users or additional components that are capable of evaluating the quality of predicted solutions. In these contexts, it is beneficial to provide these oracle mechanisms with multiple highly likely hypotheses rather than a single prediction. In this work, we pose the task of producing multiple outputs as a learning problem over an ensemble of deep networks -- introducing a novel stochastic gradient descent based approach to minimize the loss with respect to an oracle. Our method is simple to implement, agnostic to both architecture and loss function, and parameter-free. Our approach achieves lower oracle error compared to existing methods on a wide range of tasks and deep architectures. We also show qualitatively that solutions produced from our approach often provide interpretable representations of task ambiguity. Stefan Lee, Senthil Purushwalkam, Michael Cogswell, Viresh Ranjan, David Crandall, Dhruv Batra |
NIPS | 5 |
| 2016 | Addressing Physical Safety, Security, and Privacy for People with Visual Impairments
Tousif Ahmed, Patrick Shaffer, Kay Connelly, David Crandall, Apu Kapadia |
SOUPS | 4 |
| 2015 | Privacy Concerns and Behaviors of People with Visual ImpairmentsabstractVarious technologies have been developed to help make the world more accessible to visually impaired people, and recent advances in low-cost wearable and mobile computing are likely to drive even moreadvances. However, the unique privacy and security needs of visually impaired people remain largely unaddressed. We conducted an exploratory user study with 14 visually impaired participants to understand the techniques they currently use for protecting privacy, their remaining privacy concerns,and how new technologies may be able to help. The interviews explored privacy not only in the physical world (e.g., bystanders overhearing private conversations) and the online world (e.g., determining if a URL is legitimate), but also in the interface between the two (e.g. bystanders `shoulder-surfing' data from screens). The study revealed serious concerns that are not adequately solved by current technology, and suggested new directions for improving the privacy of this significant fraction of the population. Tousif Ahmed, Roberto Hoyle, Kay Connelly, David Crandall, Apu Kapadia |
CHI | 4 |
| 2015 | Sensitive Lifelogs: A Privacy Analysis of Photos from Wearable CamerasabstractWhile media reports about wearable cameras have focused on the privacy concerns of bystanders, the perspectives of the `lifeloggers' themselves have not been adequately studied. We report on additional analysis of our previous in-situ lifelogging study in which 36 participants wore a camera for a week and then reviewed the images to specify privacy and sharing preferences. In this Note, we analyze the photos themselves, seeking to understand what makes a photo private, what participants said about their images, and what we can learn about privacy in this new and very different context where photos are captured automatically by one's wearable camera. We find that these devices record many moments that may not be captured by traditional (deliberate) photography, with camera owners concerned about impression management and protecting private information of both themselves and bystanders. Roberto Hoyle, Robert Templeman, Denise L. Anthony, David Crandall, Apu Kapadia |
CHI | 4 |
| 2015 | Linking Past to Present: Discovering Style in Two Centuries of ArchitectureabstractWith vast quantities of imagery now available online, researchers have begun to explore whether visual patterns can be discovered automatically. Here we consider the particular domain of architecture, using huge collections of street-level imagery to find visual patterns that correspond to semantic-level architectural elements distinctive to particular time periods. We use this analysis both to date buildings, as well as to discover how functionally-similar architectural elements (e.g. windows, doors, balconies, etc.) have changed over time due to evolving styles. We validate the methods by combining a large dataset of nearly 150,000 Google Street View images from Paris with a cadastre map to infer approximate construction date for each facade. Not only could our analysis be used for dating or geo- localizing buildings based on architectural features, but it also could give architects and historians new tools for confirming known theories or even discovering new ones. Stefan Lee, Nicolas Maisonneuve, David Crandall, Alexei A. Efros, Josef Sivic |
ICCP | 3 |
| 2015 | Lending A Hand: Detecting Hands and Recognizing Activities in Complex Egocentric InteractionsabstractHands appear very often in egocentric video, and their appearance and pose give important cues about what people are doing and what they are paying attention to. But existing work in hand detection has made strong assumptions that work well in only simple scenarios, such as with limited interaction with other people or in lab settings. We develop methods to locate and distinguish between hands in egocentric video using strong appearance models with Convolutional Neural Networks, and introduce a simple candidate region generation approach that outperforms existing techniques at a fraction of the computational cost. We show how these high-quality bounding boxes can be used to create accurate pixelwise hand regions, and as an application, we investigate the extent to which hand segmentation alone can distinguish between different activities. We evaluate these techniques on a new dataset of 48 first-person videos of people interacting in realistic environments, with pixel-level ground truth for over 15,000 hand instances. Sven Bambach, Stefan Lee, David Crandall, Chen Yu 0001 |
ICCV | 3 |
| 2015 | Viewpoint Integration for Hand-Based Recognition of Social Interactions from a First-Person ViewabstractWearable devices are becoming part of everyday life, from first-person cameras (GoPro, Google Glass), to smart watches (Apple Watch), to activity trackers (FitBit). These devices are often equipped with advanced sensors that gather data about the wearer and the environment. These sensors enable new ways of recognizing and analyzing the wearer's everyday personal activities, which could be used for intelligent human-computer interfaces and other applications. We explore one possible application by investigating how egocentric video data collected from head-mounted cameras can be used to recognize social activities between two interacting partners (e.g. playing chess or cards). In particular, we demonstrate that just the positions and poses of hands within the first-person view are highly informative for activity recognition, and present a computer vision approach that detects hands to automatically estimate activities. While hand pose detection is imperfect, we show that combining evidence across first-person views from the two social partners significantly improves activity recognition accuracy. This result highlights how integrating weak but complimentary sources of evidence from social partners engaged in the same task can help to recognize the nature of their interaction. Sven Bambach, David Crandall, Chen Yu 0001 |
ICMI | 2 |
| 2015 | Predicting Geo-informative Attributes in Large-Scale Image Collections Using Convolutional Neural NetworksabstractGeographic location is a powerful property for organizing large-scale photo collections, but only a small fraction of online photos are geo-tagged. Most work in automatically estimating geo-tags from image content is based on comparison against models of buildings or landmarks, or on matching to large reference collections of geotagged images. These approaches work well for frequently photographed places like major cities and tourist destinations, but fail for photos taken in sparsely photographed places where few reference photos exist. Here we consider how to recognize general geo-informative attributes of a photo, e.g. the elevation gradient, population density, demographics, etc. of where it was taken, instead of trying to estimate a precise geo-tag. We learn models for these attributes using a large (noisy) set of geo-tagged images from Flickr by training deep convolutional neural networks (CNNs). We evaluate on over a dozen attributes, showing that while automatically recognizing some attributes is very difficult, others can be automatically estimated with about the same accuracy as a human. Stefan Lee, Haipeng Zhang 0004, David Crandall |
WACV | 3 |
| 2015 | Human pose estimation via multi-layer composite models
Kun Duan, Dhruv Batra, David Crandall |
Signal Process. | 3 |
| 2014 | Detecting Hands in Children's Egocentric Views to Understand Embodied Attention during Social Interaction
Sven Bambach, John M. Franchak, David Crandall, Chen Yu 0001 |
CogSci | 3 |
| 2014 | Multimodal Learning in Loosely-Organized Web ImagesabstractPhoto-sharing websites have become very popular in the last few years, leading to huge collections of online images. In addition to image data, these websites collect a variety of multimodal metadata about photos including text tags, captions, GPS coordinates, camera metadata, user profiles, etc. However, this metadata is not well constrained and is often noisy, sparse, or missing altogether. In this paper, we propose a framework to model these "loosely organized" multimodal datasets, and show how to perform loosely-supervised learning using a novel latent Conditional Random Field framework. We learn parameters of the LCRF automatically from a small set of validation data, using Information Theoretic Metric Learning (ITML) to learn distance functions and a structural SVM formulation to learn the potential functions. We apply our framework on four datasets of images from Flickr, evaluating both qualitatively and quantitatively against several baselines. Kun Duan, David Crandall, Dhruv Batra |
CVPR | 2 |
| 2014 | Privacy behaviors of lifeloggers using wearable camerasabstractA number of wearable 'lifelogging' camera devices have been released recently, allowing consumers to capture images and other sensor data continuously from a first-person perspective. Unlike traditional cameras that are used deliberately and sporadically, lifelogging devices are always 'on' and automatically capturing images. Such features may challenge users' (and bystanders') expectations about privacy and control of image gathering and dissemination. While lifelogging cameras are growing in popularity, little is known about privacy perceptions of these devices or what kinds of privacy challenges they are likely to create. Roberto Hoyle, Robert Templeman, Steven Armes, Denise L. Anthony, David Crandall, Apu Kapadia |
UbiComp | 5 |
| 2014 | Estimating bedrock and surface layer boundaries and confidence intervals in ice sheet radar imagery using MCMCabstractClimate models that predict polar ice sheet behavior require accurate measurements of the bedrock-ice and ice-air boundaries in ground-penetrating radar imagery. Identifying these features is typically performed by hand, which can be tedious and error prone. We propose an approach for automatically estimating layer boundaries by viewing this task as a probabilistic inference problem. Our solution uses Markov-Chain Monte Carlo to sample from the joint distribution over all possible layers conditioned on an image. Layer boundaries can then be estimated from the expectation over this distribution, and confidence intervals can be estimated from the variance of the samples. We evaluate the method on 560 echograms collected in Antarctica, and compare to a state-of-the-art technique with respect to hand-labeled images. These experiments show an approximately 50% reduction in error for tracing both bedrock and surface layers. Stefan Lee, Jerome E. Mitchell, David Crandall, Geoffrey C. Fox |
ICIP | 3 |
| 2014 | PlaceAvoider: Steering First-Person Cameras away from Sensitive Spaces
Robert Templeman, Mohammed Korayem, David Crandall, Apu Kapadia |
NDSS | 3 |
| 2014 | Attribute-based vehicle recognition using viewpoint-aware multiple instance SVMsabstractVehicle recognition is a challenging task with many useful applications. State-of-the-art methods usually learn discriminative classifiers for different vehicle categories or different viewpoint angles, but little work has explored vehicle recognition using semantic visual attributes. In this paper, we propose a novel iterative multiple instance learning method to model local attributes and viewpoint angles together in the same framework. We expand the standard MISVM formulation to incorporate pairwise constraints based on viewpoint relations within positive exemplars. We show that our method is able to generate discriminative and semantic local attributes for vehicle categories. We also show that we can estimate viewpoint labels more accurately than baselines when these annotations are not available in the training set. We test the technique on the Stanford cars and INRIA vehicles datasets, and compare with other methods. Kun Duan, Luca Marchesotti, David Crandall |
WACV | 3 |
| 2013 | De-Anonymizing Users Across Heterogeneous Social Computing Platforms
Mohammed Korayem, David Crandall |
ICWSM | 2 |
| 2013 | A semi-automatic approach for estimating near surface internal layers from snow radar imageryabstractThe near surface layer signatures in polar firn are preserved from the glaciological behaviors of past climate and are important to understanding the rapidly changing polar ice sheets. Identifying and tracing near surface internal layers in snow radar echograms can be used to produce high-resolution accumulation maps. This process is typically performed manually, which requires time-consuming, dense hand-selection and interpolation between sections, for each echogram. We have developed an approach for semi-automatically estimating near surface internal layers and have applied it to snow radar echograms acquired from Antarctica. Our solution utilizes an active contour (“snakes”) model to find high-intensity edges likely to correspond to layer boundaries, while simultaneously imposing constraints on smoothness of layer depth and parallelism among layers. Jerome E. Mitchell, David Crandall, Geoffrey C. Fox, John Paden |
IGARSS | 2 |
| 2013 | PlaceRaider: Virtual Theft in Physical Spaces with Smartphones
Robert Templeman, Zahid Rahman, David Crandall, Apu Kapadia |
NDSS | 3 |
| 2013 | SfM with MRFs: Discrete-Continuous Optimization for Large-Scale Structure from MotionabstractRecent work in structure from motion (SfM) has built 3D models from large collections of images downloaded from the Internet. Many approaches to this problem use incremental algorithms that solve progressively larger bundle adjustment problems. These incremental techniques scale poorly as the image collection grows, and can suffer from drift or local minima. We present an alternative framework for SfM based on finding a coarse initial solution using hybrid discrete-continuous optimization and then improving that solution using bundle adjustment. The initial optimization step uses a discrete Markov random field (MRF) formulation, coupled with a continuous Levenberg-Marquardt refinement. The formulation naturally incorporates various sources of information about both the cameras and points, including noisy geotags and vanishing point (VP) estimates. We test our method on several large-scale photo collections, including one with measured camera positions, and show that it produces models that are similar to or better than those produced by incremental bundle adjustment, but more robustly and in a fraction of the time. David Crandall, Andrew Owens, Noah Snavely, Daniel P. Huttenlocher |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2012 | A Multi-layer Composite Model for Human Pose EstimationabstractWe introduce a new approach for part-based human pose estimation using multi-layer composite models, in which each layer is a tree-structured pictorial structure that models pose at a different scale and with a different graphical structure. At the highest level, the submodel acts as a person detector, while at the lowest level, the body is decomposed into a collection of many local parts. Edges between adjacent layers of the composite model encode cross-model constraints. This multi-layer composite model is able to relax the independence assumptions of traditional tree-structured pictorial-structure models while permitting efficient inference using dual-decomposition. We propose an optimization procedure for joint learning of the entire composite model. Our approach outperforms the state-of-the-art on the challenging Parse and UIUC Sport datasets. Kun Duan, Dhruv Batra, David Crandall |
BMVC | 3 |
| 2012 | Discovering localized attributes for fine-grained recognitionabstractAttributes are visual concepts that can be detected by machines, understood by humans, and shared across categories. They are particularly useful for fine-grained domains where categories are closely related to one other (e.g. bird species recognition). In such scenarios, relevant attributes are often local (e.g. “white belly”), but the question of how to choose these local attributes remains largely unexplored. In this paper, we propose an interactive approach that discovers local attributes that are both discriminative and semantically meaningful from image datasets annotated only with fine-grained category labels and object bounding boxes. Our approach uses a latent conditional random field model to discover candidate attributes that are detectable and discriminative, and then employs a recommender system that selects attributes likely to be semantically meaningful. Human interaction is used to provide semantic names for the discovered attributes. We demonstrate our method on two challenging datasets, Caltech-UCSD Birds-200-2011 and Leeds Butterflies, and find that our discovered attributes outperform those generated by traditional approaches. Kun Duan, Devi Parikh, David Crandall, Kristen Grauman |
CVPR | 3 |
| 2012 | Learning Visual Features for the Avatar Captcha Recognition ChallengeabstractCaptchas are frequently used on the modern world wide web to differentiate human users from automated bots by giving tests that are easy for humans to answer but difficult or impossible for algorithms. As artificial intelligence algorithms have improved, new types of Captchas have had to be developed. Recent work has proposed a new system called Avatar Captcha, in which a user is asked to distinguish between facial images of real humans and those of avatars generated by computer graphics. This novel system has been proposed on the assumption that this Captcha is very difficult for computers to break. In this paper we test a variety of modern visual features and learning algorithms on this avatar recognition task. We find that relatively simple techniques can perform very well on this task, and in some cases can even surpass human performance. Mohammed Korayem, Abdallah A. Mohamed, David Crandall, Roman V. Yampolskiy |
ICMLA (2) | 3 |
| 2012 | Layer-finding in radar echograms using probabilistic graphical models
David Crandall, Geoffrey C. Fox, John Paden |
ICPR | 1 |
| 2012 | Beyond co-occurrence: discovering and visualizing tag relationships from geo-spatial and temporal similaritiesabstractStudying relationships between keyword tags on social sharing websites has become a popular topic of research, both to improve tag suggestion systems and to discover connections between the concepts that the tags represent. Existing approaches have largely relied on tag co-occurrences. In this paper, we show how to find connections between tags by comparing their distributions over time and space, discovering tags with similar geographic and temporal patterns of use. Geo-spatial, temporal and geo-temporal distributions of tags are extracted and represented as vectors which can then be compared and clustered. Using a dataset of tens of millions of geo-tagged Flickr photos, we show that we can cluster Flickr photo tags based on their geographic and temporal patterns, and we evaluate the results both qualitatively and quantitatively using a panel of human judges. We also develop visualizations of temporal and geographic tag distributions, and show that they help humans recognize semantic relationships between tags. This approach to finding and visualizing similar tags is potentially useful for exploring any data having geographic and temporal annotations. Haipeng Zhang 0004, Mohammed Korayem, Erkang You, David Crandall |
WSDM | 4 |
| 2012 | Mining photo-sharing websites to study ecological phenomenaabstractThe popularity of social media websites like Flickr and Twitter has created enormous collections of user-generated content online. Latent in these content collections are observations of the world: each photo is a visual snapshot of what the world looked like at a particular point in time and space, for example, while each tweet is a textual expression of the state of a person and his or her environment. Aggregating these observations across millions of social sharing users could lead to new techniques for large-scale monitoring of the state of the world and how it is changing over time. In this paper we step towards that goal, showing that by analyzing the tags and image features of geo-tagged, time-stamped photos we can measure and quantify the occurrence of ecological phenomena including ground snow cover, snow fall and vegetation density. We compare several techniques for dealing with the large degree of noise in the dataset, and show how machine learning can be used to reduce errors caused by misleading tags and ambiguous visual content. We evaluate the accuracy of these techniques by comparing to ground truth data collected both by surface stations and by Earth-observing satellites. Besides the immediate application to ecology, our study gives insight into how to accurately crowd-source other types of information from large, noisy social sharing datasets. Haipeng Zhang 0004, Mohammed Korayem, David Crandall, Gretchen LeBuhn |
WWW | 3 |
| 2011 | Discrete-continuous optimization for large-scale structure from motionabstractRecent work in structure from motion (SfM) has successfully built 3D models from large unstructured collections of images downloaded from the Internet. Most approaches use incremental algorithms that solve progressively larger bundle adjustment problems. These incremental techniques scale poorly as the number of images grows, and can drift or fall into bad local minima. We present an alternative formulation for SfM based on finding a coarse initial solution using a hybrid discrete-continuous optimization, and then improving that solution using bundle adjustment. The initial optimization step uses a discrete Markov random field (MRF) formulation, coupled with a continuous Levenberg-Marquardt refinement. The formulation naturally incorporates various sources of information about both the cameras and the points, including noisy geotags and vanishing point estimates. We test our method on several large-scale photo collections, including one with measured camera positions, and show that it can produce models that are similar to or better than those produced with incremental bundle adjustment, but more robustly and in a fraction of the time. David Crandall, Andrew Owens, Noah Snavely, Daniel P. Huttenlocher |
CVPR | 1 |
| 2009 | Landmark classification in large-scale image collectionsabstractWith the rise of photo-sharing websites such as Facebook and Flickr has come dramatic growth in the number of photographs online. Recent research in object recognition has used such sites as a source of image data, but the test images have been selected and labeled by hand, yielding relatively small validation sets. In this paper we study image classification on a much larger dataset of 30 million images, including nearly 2 million of which have been labeled into one of 500 categories. The dataset and categories are formed automatically from geotagged photos from Flickr, by looking for peaks in the spatial geotag distribution corresponding to frequently-photographed landmarks. We learn models for these landmarks with a multiclass support vector machine, using vector-quantized interest point descriptors as features. We also explore the non-visual information available on modern photo-sharing sites, showing that using textual tags and temporal constraints leads to significant improvements in classification rate. We find that in some cases image features alone yield comparable classification accuracy to using text tags as well as to the performance of human observers. Yunpeng Li 0002, David Crandall, Daniel P. Huttenlocher |
ICCV | 2 |
| 2009 | Mapping the world's photosabstractWe investigate how to organize a large collection of geotagged photos, working with a dataset of about 35 million images collected from Flickr. Our approach combines content analysis based on text tags and image data with structural analysis based on geospatial data. We use the spatial distribution of where people take photos to define a relational structure between the photos that are taken at popular places. We then study the interplay between this structure and the content, using classification methods for predicting such locations from visual, textual and temporal features of the photos. We find that visual and temporal features improve the ability to estimate the location of a photo, compared to using just textual features. We illustrate using these techniques to organize a large photo collection, while also revealing various interesting properties about popular cities and landmarks at a global scale. David Crandall, Lars Backstrom, Daniel P. Huttenlocher, Jon M. Kleinberg |
WWW | 1 |
| 2008 | Feedback effects between similarity and social influence in online communitiesabstractA fundamental open question in the analysis of social networks is to understand the interplay between similarity and social ties. People are similar to their neighbors in a social network for two distinct reasons: first, they grow to resemble their current friends due to social influence; and second, they tend to form new links to others who are already like them, a process often termed selection by sociologists. While both factors are present in everyday social processes, they are in tension: social influence can push systems toward uniformity of behavior, while selection can lead to fragmentation. As such, it is important to understand the relative effects of these forces, and this has been a challenge due to the difficulty of isolating and quantifying them in real settings. David Crandall, Dan Cosley, Daniel P. Huttenlocher, Jon M. Kleinberg, Siddharth Suri |
KDD | 1 |
| 2007 | Composite Models of Objects and Scenes for Category RecognitionabstractThis paper presents a method of learning and recognizing generic object categories using part-based spatial models. The models are multiscale, with a scene component that specifies relationships between the object and surrounding scene context, and an object component that specifies relationships between parts of the object. The underlying graphical model forms a tree structure, with a star topology for both the contextual and object components. A partially supervised paradigm is used for learning the models, where each training image is labeled with bounding boxes indicating the overall location of object instances, but parts or regions of the objects and scene are not specified. The parts, regions and spatial relationships are learned automatically. We demonstrate the method on the detection task on the PASCAL 2006 Visual Object Classes Challenge dataset, where objects must be correctly localized. Our results demonstrate better overall performance than those of previously reported techniques, in terms of the average precision measure used in the PASCAL detection evaluation. Our results also show that incorporating scene context into the models improves performance in comparison with not using such contextual information. David Crandall, Daniel P. Huttenlocher |
CVPR | 1 |
| 2006 | Weakly Supervised Learning of Part-Based Spatial Models for Visual Object Recognition
David Crandall, Daniel P. Huttenlocher |
ECCV (1) | 1 |
| 2006 | Color object detection using spatial-color joint probability functionsabstractObject detection in unconstrained images is an important image understanding problem with many potential applications. There has been little success in creating a single algorithm that can detect arbitrary objects in unconstrained images; instead, algorithms typically must be customized for each specific object. Consequently, it typically requires a large number of exemplars (for rigid objects) or a large amount of human intuition (for nonrigid objects) to develop a robust algorithm. We present a robust algorithm designed to detect a class of compound color objects given a single model image. A compound color object is defined as having a set of multiple, particular colors arranged spatially in a particular way, including flags, logos, cartoon characters, people in uniforms, etc. Our approach is based on a particular type of spatial-color joint probability function called the color edge co-occurrence histogram. In addition, our algorithm employs perceptual color naming to handle color variation, and prescreening to limit the search scope (i.e., size and location) for the object. Experimental results demonstrated that the proposed algorithm is insensitive to object rotation, scaling, partial occlusion, and folding, outperforming a closely related algorithm based on color co-occurrence histograms by a decisive margin. Jiebo Luo 0001, David Crandall |
IEEE Trans. Image Process. | 2 |
| 2005 | Spatial Priors for Part-Based Recognition Using Statistical ModelsabstractWe present a class of statistical models for part-based object recognition that are explicitly parameterized according to the degree of spatial structure they can represent. These models provide a way of relating different spatial priors that have been used for recognizing generic classes of objects, including joint Gaussian models and tree-structured models. By providing explicit control over the degree of spatial structure, our models make it possible to study the extent to which additional spatial constraints among parts are actually helpful in detection and localization, and to consider the tradeoff in representational power and computational cost. We consider these questions for object classes that have substantial geometric structure, such as airplanes, faces and motorbikes, using datasets employed by other researchers to facilitate evaluation. We find that for these classes of objects, a relatively small amount of spatial structure in the model can provide statistically indistinguishable recognition performance from more powerful models, and at a substantially lower computational cost. David Crandall, Pedro F. Felzenszwalb, Daniel P. Huttenlocher |
CVPR (1) | 1 |
| 2004 | Robust Color Object Detection Using Spatial-Color Joint Probability Functions
David Crandall, Jiebo Luo 0001 |
CVPR (1) | 1 |
| 2003 | Extraction of special effects caption text events from digital video
David Crandall, Sameer K. Antani, Rangachar Kasturi |
Int. J. Document Anal. Recognit. | 1 |
| 2001 | Robust Detection of Stylized Text Events in Digital VideoabstractAutomatic content-based video indexing is an important research problem. One approach is to extract text appearing in video as an indication of a scene's semantic content. Most work so far has focused only on detecting the spatial extent of text instances in individual video frames. But text occurring in video usually persists for several seconds. This constitutes a text event that should be entered only once in the video index. Therefore it is necessary to determine the temporal extent of text events by combining the results of text detection on individual frames, over time. This is a nontrivial problem because a text event may move, rotate, grow, shrink, or otherwise change throughout its lifetime. Such text effects are common in television programs and commercials to attract viewer attention, but have so far been ignored in the literature. We present a method for detecting and tracking moving, changing caption text events in MPEG-1 compressed video. David Crandall, Rangachar Kasturi |
ICDAR | 1 |
| 2000 | Robust Extraction of Text in VideoabstractDespite advances in the archiving of digital video, we are still unable to efficiently search and retrieve the portions that interest us. Video indexing by shot segmentation has been a proposed solution and several research efforts are seen in the literature. Shot segmentation alone cannot solve the problem of content based access to video. Recognition of text in video has been proposed as an additional feature. Several research efforts are found in the literature for text extraction from complex images and video with applications for video indexing. We present an update of our system for detection and extraction of an unconstrained variety of text from general purpose video. The text detection results from a variety of methods are fused and each single text instance is segmented to enable it for OCR. Problems in segmenting text from video are similar to those faced in detection and localization phases. Video has low resolution and the text often has poor contrast with a changing background. The proposed system applies a variety of methods and takes advantage of the temporal redundancy in video resulting in good text segmentation. Sameer K. Antani, David Crandall, Rangachar Kasturi |
ICPR | 2 |
| 1999 | A System for Automatic Text Detection in VideoabstractVideo indexing is an important problem that has occupied recent research efforts. The text appearing in video can provide semantic information about the scene content. Detecting and recognizing text events can provide indices into the video for content based querying. We describe a system for detecting, tracking, and extracting artificial and scene text in MPEG-1 video. Preliminary results are presented. Ullas Gargi, David Crandall, Sameer K. Antani, Tarak Gandhi, Ryan Keener, Rangachar Kasturi |
ICDAR | 2 |