VLDB 2026 Research / reviewers in the wild / expert
Laurent Itti
dblp:31/3256
· DBLP profile ↗
108ranked-venue papers
10as first author
26since 2021 · last 2025
0000-0002-0168-2977ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 82 · 7 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 45 · 4 first-author · 14 since 2021Systems, architecture and hardware · 18 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 9Human-computer interaction and ubiquitous computing · 5 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | NaviSense: A Multimodal Assistive Mobile application for Object Retrieval by Persons with Visual ImpairmentabstractFigure 1: System architecture of NaviSense.Gradient-filled boxes indicate on-device processing components (e.g., speech recognition and synthesis); dashed-border boxes denote cloud-based processing (LLM and VLM); thick-border boxes represent iterative operations.The system uses speech input, real-time camera and LiDAR feeds, and provides multimodal feedback (audio and haptic) to guide blind or visually impaired users to the requested object. Ajay Narayanan Sridhar, Fuli Qiao, Nelson Daniel Troncoso Aldas, Yanpei Shi, Mehrdad Mahdavi, Laurent Itti, Narayanan Vijaykrishnan |
ASSETS | 6 |
| 2025 | Riemannian-Geometric Fingerprints of Generative ModelsabstractRecent breakthroughs and rapid integration of generative models (GMs) have sparked interest in the problem of model attribution and their fingerprints. For instance, service providers need reliable methods of authenticating their models to protect their IP, while users and law enforcement seek to verify the source of generated content for accountability and trust. In addition, a growing threat of model collapse is arising, as more model-generated data are being fed back into sources (e.g., YouTube) that are often harvested for training ("regurgitative training"), heightening the need to differentiate synthetic from human data. Yet, a gap still exists in understanding generative models' fingerprints, we believe, stemming from the lack of a formal framework that can define, represent, and analyze the fingerprints in a principled way. To address this gap, we take a geometric approach and propose a new definition of artifact and fingerprint of GMs using Riemannian geometry, which allows us to leverage the rich theory of differential geometry. Our new definition generalizes previous work (Song et al., 2024) to non-Euclidean manifolds by learning Riemannian metrics from data and replacing the Euclidean distances and nearest-neighbor search with geodesic distances and kNN-based Riemannian center of mass. We apply our theory to a new gradient-based algorithm for computing the fingerprints in practice. Results show that it is more effective in distinguishing a large array of GMs, spanning across 4 different datasets in 2 different resolutions (64 by 64, 256 by 256), 27 model architectures, and 2 modalities (Vision, Vision-Language). Using our proposed definition significantly improves the performance on model attribution, as well as a generalization to unseen datasets, model types, and modalities, suggesting its practical efficacy. Hae Jin Song, Laurent Itti |
ICCV | 2 |
| 2025 | Bayesian Surprise for Small and Sub-Pixel Moving Target DetectionabstractDetecting very small moving targets (sub-pixel to a few pixels) has not benefited much from recent advances in deep neural networks, as these typically rely on texture features that are absent in very small targets. We present a new algorithm using Bayesian surprise. It computes how an artificial observer’s beliefs about the contents of a video sequence evolve over space and time. Surprise builds a distribution of beliefs for each pixel in the video. On every new video frame, Bayes’ theorem updates the prior belief distributions into posterior distributions. Surprise arises when/where these updates are significant, as measured by the Kullback-Leibler divergence. Surprise outperforms 27 baselines on downscaled versions of the public Neovision2-Tower dataset (50 videos, 45,000 frames total, 5 target types). Notably, surprise is more robust to wide-area perturbations (e.g., foliage fluttering in the wind, clouds passing by) than other pixel-level models. Laurent Itti, Daben Liu, Corinne Teeter, Srideep Musuvathy |
ICIP | 1 |
| 2025 | DreamDistribution: Learning Prompt Distribution for Diverse In-distribution GenerationabstractThe popularization of Text-to-Image (T2I) diffusion models enables the generation of high-quality images from text descriptions. However, generating diverse customized images with reference visual attributes remains challenging. This work focuses on personalizing T2I diffusion models at a more abstract concept or category level, adapting commonalities from a set of reference images while creating new instances with sufficient variations. We introduce a solution that allows a pretrained T2I diffusion model to learn a set of soft prompts, enabling the generation of novel images by sampling prompts from the learned distribution. These prompts offer text-guided editing capabilities and additional flexibility in controlling variation and mixing between multiple distributions. We also show the adaptability of the learned prompt distribution to other tasks, such as text-to-3D. Finally we demonstrate effectiveness of our approach through quantitative analysis including automatic evaluation and human assessment. Brian Nlong Zhao, Xinyang Jiang, Yifan Yang 0004, Dongsheng Li 0002, Laurent Itti, Vibhav Vineet, Yunhao Ge |
ICLR | 7 |
| 2024 | BEHAVIOR Vision Suite: Customizable Dataset Generation via SimulationabstractThe systematic evaluation and understanding of computer vision models under varying conditions require large amounts of data with comprehensive and customized labels, which real-world vision datasets rarely satisfy. While current synthetic data generators offer a promising alternative, particularly for embodied AI tasks, they often fall short for computer vision tasks due to low asset and rendering quality, limited diversity, and unrealistic physical properties. We introduce the BEHAVIOR Vision Suite (BVS), a set of tools and assets to generate fully customized synthetic data for systematic evaluation of computer vision models, based on the newly developed embodied AI benchmark, BEHAVIOR-1 K. BVS supports a large number of adjustable parameters at the scene level (e.g., lighting, object placement), the object level (e.g., joint configuration, attributes such as “filled” and “folded”), and the camera level (e.g., field of view, focal length). Researchers can arbitrarily vary these parameters during data generation to perform controlled experiments. We showcase three example application scenarios: systematically evaluating the robustness of models across different continuous axes of domain shift, evaluating scene understanding models on the same set of images, and training and evaluating simulation-to-real transfer for a novel vision task: unary and binary state prediction. Project website: https://behavior-vision-suite.github.io/ Yunhao Ge, Yihe Tang, Cem Gökmen, Chengshu Li 0001, Wensi Ai, Benjamin Jose Martinez, Arman Aydin, Mona Anvari, Ayush K. Chakravarthy, Hong-Xing Yu, Josiah Wong, Sanjana Srivastava, Sharon Lee, Shengxin Zha, Laurent Itti, Yunzhu Li, Roberto Martin Martin, Miao Liu 0007, Pengchuan Zhang, Li Fei-Fei 0001, Jiajun Wu 0001 |
CVPR | 16 |
| 2024 | Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment CollaborationabstractLarge, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for many applications. Can such a consolidation happen in robotics? Conventionally, robotic learning methods train a separate model for every application, every robot, and even every environment. Can we instead train "generalist" X-robot policy that can be adapted efficiently to new robots, tasks, and environments? In this paper, we provide datasets in standardized data formats and models to make it possible to explore this possibility in the context of robotic manipulation, alongside experimental results that provide an example of effective X-robot policies. We assemble a dataset from 22 different robots collected through a collaboration between 21 institutions, demonstrating 527 skills (160266 tasks). We show that a high-capacity model trained on this data, which we call RT-X, exhibits positive transfer and improves the capabilities of multiple robots by leveraging experience from other platforms. The project website is robotics-transformer-x.github.io. Abigail O'Neill, Abhiram Maddukuri, Abhishek Gupta 0004, Abhishek Padalkar, Abraham Lee, Acorn Pooley, Agrim Gupta, Ajay Mandlekar, Ajinkya Jain, Albert Tung, Alex Bewley, Alex Irpan, Alexander Khazatsky, Anant Rai, Anchit Gupta, Andrew E. Wang, Anikait Singh, Animesh Garg, Aniruddha Kembhavi, Annie Xie, Anthony Brohan, Antonin Raffin, Archit Sharma, Arefeh Yavary, Arhan Jain, Ashwin Balakrishna, Ayzaan Wahid, Ben Burgess-Limerick, Bernhard Schölkopf, Blake Wulfe, Brian Ichter, Cewu Lu, Charles Xu 0003, Charlotte Le, Chelsea Finn, Chen Wang 0053, Chenfeng Xu, Cheng Chi 0001, Chenguang Huang, Christine Chan, Christopher Agia, Chuer Pan, Chuyuan Fu, Coline Devin, Danfei Xu, Daniel Morton, Danny Drieß, Daphne Chen, Deepak Pathak, Dhruv Shah, Dieter Büchler, Dinesh Jayaraman, Dmitry Kalashnikov, Dorsa Sadigh, Edward Johns, Ethan Paul Foster, Fangchen Liu, Federico Ceola, Fei Xia 0002, Feiyu Zhao, Freek Stulp, Gaoyue Zhou, Gaurav S. Sukhatme, Gautam Salhotra, Gilbert Feng, Giulio Schiavi, Glen Berseth, Gregory Kahn, Guanzhi Wang, Hao Su 0001, Haoshu Fang, Henghui Bao, Heni Ben Amor, Henrik I. Christensen, Hiroki Furuta, Homer Walke, Hongjie Fang, Huy Ha, Igor Mordatch, Ilija Radosavovic, Isabel Leal, Jacky Liang, Jad Abou-Chakra, Jaehyung Kim 0001, Jaimyn Drake, Jan Peters 0001, Jan Schneider 0007, Jasmine Hsu, Jeannette Bohg, Jeffrey T. Bingham, Jensen Gao, Jiaheng Hu, Jiajun Wu 0001, Jiankai Sun, Jianlan Luo, Jiayuan Gu, Jie Tan 0001, Jihoon Oh, Jimmy Wu, Jingpei Lu, Jitendra Malik, João Silvério, Joey Hejna, Jonathan Booher, Jonathan Tompson, Jonathan Yang, Jordi Salvador, Joseph J. Lim, Junhyek Han, Kanishka Rao, Karl Pertsch, Karol Hausman, Keegan Go, Keerthana Gopalakrishnan, Kenneth Y. Goldberg, Kendra Byrne, Kenneth Oslund, Kento Kawaharazuka, Kevin Black, Kevin Zhang 0002, Kiana Ehsani, Kiran Lekkala, Kirsty Ellis, Krishan Rana, Krishnan Srinivasan, Kuan Fang, Kunal Pratap Singh, Kuo-Hao Zeng, Kyle Hatch, Kyle Hsu, Laurent Itti, Yunliang Chen 0001, Lerrel Pinto, Li Fei-Fei 0001, Liam Tan, Linxi Fan, Lionel Ott, Lisa Lee, Luca Weihs, Magnum Chen, Marion Lepert, Marius Memmel, Masayoshi Tomizuka, Masha Itkina, Mateo Guaman Castro, Max Spero, Maximilian Du, Michael Ahn, Michael C. Yip, Mingtong Zhang 0003, Mingyu Ding, Minho Heo, Mohan Kumar Srirama, Mohit Sharma 0001, Moo Jin Kim, Naoaki Kanazawa, Nicklas Hansen 0001, Nicolas Heess, Nikhil J. Joshi, Niko Sünderhauf, Norman Di Palo, Nur Muhammad Shafiullah, Oier Mees, Oliver Kroemer, Osbert Bastani, Pannag R. Sanketi, Patrick Tree Miller, Patrick Yin, Paul Wohlhart, Peng Xu 0010, Peter David Fagan, Peter Mitrano, Pierre Sermanet, Pieter Abbeel, Priya Sundaresan, Qiuyu Chen, Rafael Rafailov, Ria Doshi, Roberto Martin Martin, Rohan Baijal, Rosario Scalise, Rose Hendrix, Roy Lin, Runjia Qian, Russell Mendonca, Rutav Shah, Ryan Hoque, Ryan Julian, Samuel Bustamante-Gomez, Sean Kirmani, Sergey Levine, Sherry Moore, Shikhar Bahl, Shivin Dass, Shubham D. Sonawani, Shuran Song, Sichun Xu, Siddhant Haldar, Siddharth Karamcheti, Simeon Adebola, Simon Guist, Soroush Nasiriany, Stefan Schaal, Stefan Welker, Stephen Tian, Subramanian Ramamoorthy, Sudeep Dasari, Suneel Belkhale, Sungjae Park, Suraj Nair 0003, Suvir Mirchandani, Takayuki Osa, Tanmay Gupta, Tatsuya Harada, Tatsuya Matsushima, Ted Xiao, Thomas Kollar, Tianhe Yu, Tianli Ding, Todor Davchev, Tony Z. Zhao, Travis Armstrong, Trevor Darrell, Trinity Chung, Vidhi Jain, Vincent Vanhoucke, Wolfram Burgard, Xiaolong Wang 0004, Xinghao Zhu, Xinyang Geng, Liangwei Xu, Yecheng Jason Ma 0001, Yejin Kim 0003, Yevgen Chebotar, Yilin Wu 0003, Yonatan Bisk, Yoonyoung Cho, Youngwoon Lee, Yuchen Cui, Yueh-Hua Wu, Yujin Tang, Yuke Zhu, Yunchu Zhang, Yunfan Jiang 0001, Yunshuang Li, Yunzhu Li, Yusuke Iwasawa, Yutaka Matsuo, Zehan Ma, Zichen Jeff Cui, Zichen Zhang 0016, Zipeng Lin |
ICRA | 141 |
| 2024 | USCILab3D: A Large-scale, Long-term, Semantically Annotated Outdoor DatasetabstractIn this paper, we introduce the \textbf{USCILab3D dataset}, a large-scale, annotated outdoor dataset designed for versatile applications across multiple domains, including computer vision, robotics, and machine learning. The dataset was acquired using a mobile robot equipped with 5 cameras and a 32-beam, $360^{\circ}$ scanning LIDAR. The robot was teleoperated, over the course of a year and under a variety of weather and lighting conditions, through a rich variety of paths within the USC campus (229 acres = $\sim 92.7$ hectares). The raw data was annotated using state-of-the-art large foundation models, and processed to provide multi-view imagery, 3D reconstructions, semantically-annotated images and point clouds (267 semantic categories), and text descriptions of images and objects within. The dataset also offers a diverse array of complex analyses using pose-stamping and trajectory data. In sum, the dataset offers 1.4M point clouds and 10M images ($\sim 6$TB of data). Despite covering a narrower geographical scope compared to a whole-city dataset, our dataset prioritizes intricate intersections along with denser multi-view scene images and semantic point clouds, enabling more precise 3D labelling and facilitating a broader spectrum of 3D vision tasks. For data, code and more details, please visit our website. Kiran Lekkala, Henghui Bao, Peixu Cai, Wei Lim, Laurent Itti |
NeurIPS | 6 |
| 2023 | Reproducibility Requires Consolidated ArtifactsabstractMachine learning is facing a ‘reproducibility crisis’ where a significant number of works report failures when attempting to reproduce previously published results. We evaluate the sources of reproducibility failures using a meta-analysis of 142 replication studies from ReScience C and 204 code repositories. We find that missing experiment details such as hyperparameters are potential causes of unreproducibility. We experimentally show the bias of different hyperparameter selection strategies and conclude that consolidated artifacts with a unified framework can help support reproducibility. Iordanis Fostiropoulos, Bowman Brown, Laurent Itti |
CAIN | 3 |
| 2023 | Batch Model Consolidation: A Multi-Task Model Consolidation FrameworkabstractIn Continual Learning (CL), a model is required to learn a stream of tasks sequentially without significant performance degradation on previously learned tasks. Current approaches fail for a long sequence of tasks from diverse domains and difficulties. Many of the existing CL approaches are difficult to apply in practice due to excessive memory cost or training time, or are tightly coupled to a single device. With the intuition derived from the widely applied mini-batch training, we propose Batch Model Consolidation (BMC) to support more realistic CL under conditions where multiple agents are exposed to a range of tasks. During a regularization phase, BMC trains multiple expert models in parallel on a set of disjoint tasks. Each expert maintains weight similarity to a base model through a stability loss, and constructs a buffer from a fraction of the task's data. During the consolidation phase, we combine the learned knowledge on 'batches' of expert models using a batched consolidation loss in memory data that aggregates all buffers. We thoroughly evaluate each component of our method in an ablation study and demonstrate the effectiveness on standardized benchmark datasets Split-CIFAR-100, Tiny-ImageNet, and the Stream dataset composed of 71 image classification tasks from diverse domains and difficulties. Our method outperforms the next best CL approach by 70% and is the only approach that can maintain performance at the end of 71 tasks. Iordanis Fostiropoulos, Jiaye Zhu, Laurent Itti |
CVPR | 3 |
| 2023 | Improving Zero-shot Generalization and Robustness of Multi-Modal ModelsabstractMulti-modal- image-text models such as CLIP and LiT have demonstrated impressive performance on image classification benchmarks and their zero-shot generalization ability is particularly exciting. While the top-5 zero-shot accuracies of these models are very high, the top-1 accuracies are much lower (over 25% gap in some cases). We investigate the reasons for this performance gap and find that many of the failure cases are caused by ambiguity in the text prompts. First, we develop a simple and efficient zero-shot post-hoc method to identify images whose top-1 prediction is likely to be incorrect, by measuring consistency of the predictions w.r.t. multiple prompts and image transformations. We show that our procedure better predicts mistakes, outperforming the popular max logit baseline on selective prediction tasks. Next, we propose a simple and efficient way to improve accuracy on such uncertain images by making use of the WordNet hierarchy; specifically we augment the original class by incorporating its parent and children from the semantic label hierarchy, and plug the augmentation into text prompts. We conduct experiments on both CLIP and LiT models with five different ImageNet-based datasets. For CLIP, our method improves the top-1 accuracy by 17.13% on the uncertain subset and 3.6% on the entire ImageNet validation set. We also show that our method improves across ImageNet shifted datasets, four other datasets, and other model architectures such as LiT. The proposed method11Work carried out mainly at Google is hyperparameter-free, requires no additional model training and can be easily scaled to other large multi-modal architectures. Code is available at https://github.com/gyhandy/Hierarchy-CLIP. Yunhao Ge, Jie Ren 0006, Andrew Gallagher, Yuxiao Wang 0001, Ming-Hsuan Yang 0001, Hartwig Adam, Laurent Itti, Balaji Lakshminarayanan, Jiaping Zhao |
CVPR | 7 |
| 2023 | CLR: Channel-wise Lightweight Reprogramming for Continual LearningabstractContinual learning aims to emulate the human ability to continually accumulate knowledge over sequential tasks. The main challenge is to maintain performance on previously learned tasks after learning new tasks, i.e., to avoid catastrophic forgetting. We propose a Channel-wise Lightweight Reprogramming (CLR) approach that helps convolutional neural networks (CNNs) overcome catastrophic forgetting during continual learning. We show that a CNN model trained on an old task (or self-supervised proxy task) could be "reprogrammed" to solve a new task by using our proposed lightweight (very cheap) reprogramming parameter. With the help of CLR, we have a better stability-plasticity trade-off to solve continual learning problems: To maintain stability and retain previous task ability, we use a common task-agnostic immutable part as the shared "anchor" parameter set. We then add task-specific lightweight reprogramming parameters to reinterpret the outputs of the immutable parts, to enable plasticity and integrate new knowledge. To learn sequential tasks, we only train the lightweight reprogramming parameters to learn each new task. Reprogramming parameters are task-specific and exclusive to each task, which makes our method immune to catastrophic forgetting. To minimize the parameter requirement of reprogramming to learn new tasks, we make reprogramming lightweight by only adjusting essential kernels and learning channel-wise linear mappings from anchor parameters to task-specific domain knowledge. We show that, for general CNNs, the CLR parameter increase is less than 0.6% for any new task. Our method outperforms 13 state-of-the-art continual learning baselines on a new challenging sequence of 53 image classification datasets. Code and data are in the top link. Yunhao Ge, Yuecheng Li, Shuo Ni, Jiaping Zhao, Ming-Hsuan Yang 0001, Laurent Itti |
ICCV | 6 |
| 2023 | 3D Copy-Paste: Physically Plausible Object Insertion for Monocular 3D DetectionabstractA major challenge in monocular 3D object detection is the limited diversity and quantity of objects in real datasets. While augmenting real scenes with virtual objects holds promise to improve both the diversity and quantity of the objects, it remains elusive due to the lack of an effective 3D object insertion method in complex real captured scenes. In this work, we study augmenting complex real indoor scenes with virtual objects for monocular 3D object detection. The main challenge is to automatically identify plausible physical properties for virtual assets (e.g., locations, appearances, sizes, etc.) in cluttered real scenes. To address this challenge, we propose a physically plausible indoor 3D object insertion approach to automatically copy virtual objects and paste them into real scenes. The resulting objects in scenes have 3D bounding boxes with plausible physical locations and appearances. In particular, our method first identifies physically feasible locations and poses for the inserted objects to prevent collisions with the existing room layout. Subsequently, it estimates spatially-varying illumination for the insertion location, enabling the immersive blending of the virtual objects into the original scene with plausible appearances and cast shadows. We show that our augmentation method significantly improves existing monocular 3D object models and achieves state-of-the-art performance. For the first time, we demonstrate that a physically plausible 3D object insertion, serving as a generative data augmentation technique, can lead to significant improvements for discriminative downstream tasks such as monocular 3D object detection. Project website: https://gyhandy.github.io/3D-Copy-Paste/. Yunhao Ge, Hong-Xing Yu, Cheng Zhao 0002, Yuliang Guo, Xinyu Huang 0001, Liu Ren 0001, Laurent Itti, Jiajun Wu 0001 |
NeurIPS | 7 |
| 2023 | RoboCLIP: One Demonstration is Enough to Learn Robot PoliciesabstractReward specification is a notoriously difficult problem in reinforcement learning, requiring extensive expert supervision to design robust reward functions. Imitation learning (IL) methods attempt to circumvent these problems by utilizing expert demonstrations instead of using an extrinsic reward function but typically require a large number of in-domain expert demonstrations. Inspired by advances in the field of Video-and-Language Models (VLMs), we present RoboCLIP, an online imitation learning method that uses a single demonstration (overcoming the large data requirement) in the form of a video demonstration or a textual description of the task to generate rewards without manual reward function design. Additionally, RoboCLIP can also utilize out-of-domain demonstrations, like videos of humans solving the task for reward generation, circumventing the need to have the same demonstration and deployment domains.
RoboCLIP utilizes pretrained VLMs without any finetuning for reward generation. Reinforcement learning agents trained with RoboCLIP rewards demonstrate 2-3 times higher zero-shot performance than competing imitation learning methods on downstream robot manipulation tasks, doing so using only one video/text demonstration. Visit our website at https://sites.google.com/view/roboclip/home for experiment videos. Sumedh A. Sontakke, Jesse Zhang, Sébastien M. R. Arnold, Karl Pertsch, Erdem Biyik, Dorsa Sadigh, Chelsea Finn, Laurent Itti |
NeurIPS | 8 |
| 2023 | Encouraging Disentangled and Convex Representation with Controllable Interpolation RegularizationabstractWe focus on controllable disentangled representation learning (C-Dis-RL), where users can control the partition of the disentangled latent space to factorize dataset attributes (concepts) for downstream tasks. Two general problems remain under-explored in current methods: (1) They lack comprehensive disentanglement constraints, especially missing the minimization of mutual information between different attributes across latent and observation domains. (2) They lack convexity constraints, which is important for meaningfully manipulating specific attributes for downstream tasks. To encourage both comprehensive C-Dis-RL and convexity simultaneously, we propose a simple yet efficient method: Controllable Interpolation Regularization (CIR), which creates a positive loop where disentanglement and convexity can help each other. Specifically, we conduct controlled interpolation in latent space during training, and we reuse the encoder to help form a ’perfect disentanglement’ regularization. In that case, (a) disentanglement loss implicitly enlarges the potential understandable distribution to encourage convexity; (b) convexity can in turn improve robust and precise disentanglement. CIR is a general module and we merge CIR with three different algorithms: ELEGANT, I2I-Dis, and GZS-Net to show the compatibility and effectiveness. Qualitative and quantitative experiments show improvement in C-Dis-RL and latent convexity by CIR. This further improves downstream tasks: controllable image synthesis, cross-modality image translation and zero-shot synthesis. Yunhao Ge, Zhi Xu 0013, Gan Xin, Yunkui Pang, Laurent Itti |
WACV | 6 |
| 2023 | HOOT: Heavy Occlusions in Object Tracking BenchmarkabstractIn this paper, we present HOOT, the Heavy Occlusions in Object Tracking Benchmark, a new visual object tracking dataset aimed towards handling high occlusion scenarios for single-object tracking tasks. The dataset consists of 581 high-quality videos, which have 436K frames densely annotated with rotated bounding boxes for targets spanning 74 object classes. The dataset is geared for development, evaluation and analysis of visual tracking algorithms that are robust to occlusions. It is comprised of videos with high occlusion levels, where the median percentage of occluded frames per-video is 68%. It also provides critical attributes on occlusions, which include defining a taxonomy for occluders, providing occlusion masks for every bounding box, per-frame partial/full occlusion labels and more. HOOT has been compiled to encourage development of new methods targeting occlusion handling in visual tracking, by providing training and test splits with high occlusion levels. This makes HOOT the first densely-annotated, large dataset designed for single-object tracking under severe occlusion. We evaluate 15 state-of-the-art trackers on this new dataset to act as a baseline for future work focusing on occlusions. Gozde Sahin, Laurent Itti |
WACV | 2 |
| 2022 | GalilAI: Out-of-Task Distribution Detection using Causal Active Experimentation for Safe Transfer RLabstractOut-of-distribution (OOD) detection is a well-studied topic in supervised learning. Extending the successes in supervised learning methods to the reinforcement learning (RL) setting, however, is difficult due to the data generating process - RL agents actively query their environment for data and this data is a function of the policy followed by the agent. Thus, an agent could neglect a shift in the environment if its policy did not lead it to explore the aspect of the environment that shifted. Therefore, to achieve safe and robust generalization in RL, there exists an unmet need for OOD detection through active experimentation. Here, we attempt to bridge this lacuna by first - defining a causal framework for OOD scenarios or environments encountered by RL agents in the wild. Then, we propose a novel task - that of Out-of-Task Distribution (OOTD) detection. We introduce an RL agent which actively experiments in a test environment and subsequently concludes whether it is OOTD or not. We name our method GalilAI, in honor of Galileo Galilei, as it also discovers, among other causal processes, that gravitational acceleration is independent of the mass of a body. Finally, we propose a simple probabilistic neural network baseline for comparison, which extends extant Model-Based RL. We find that our method outperforms the baseline significantly. Sumedh A. Sontakke, Stephen Iota, Zizhao Hu, Arash Mehrjou, Laurent Itti, Bernhard Schölkopf |
AISTATS | 5 |
| 2022 | Neural-Sim: Learning to Generate Training Data with NeRF
Yunhao Ge, Harkirat S. Behl, Suriya Gunasekar, Neel Joshi, Yale Song, Xin Wang 0066, Laurent Itti, Vibhav Vineet |
ECCV (23) | 8 |
| 2022 | Contributions of Shape, Texture, and Color in Visual Recognition
Yunhao Ge, Zhi Xu 0013, Xingrui Wang, Laurent Itti |
ECCV (12) | 5 |
| 2022 | incDFM: Incremental Deep Feature Modeling for Continual Novelty Detection
Amanda Rios, Nilesh A. Ahuja, Ibrahima J. Ndiour, Ergin Utku Genc, Laurent Itti, Omesh Tickoo |
ECCV (25) | 5 |
| 2022 | SHERLock: Self-Supervised Hierarchical Event Representation LearningabstractTemporal event representations are an essential aspect of learning among humans. They allow for succinct encoding of the experiences we have through a variety of sensory inputs. Also, they are believed to be arranged hierarchically, allowing for an efficient representation of complex long-horizon experiences. Additionally, these representations are acquired in a self-supervised manner. Analogously, here we propose a model that learns temporal representations from long-horizon visual demonstration data and associated textual descriptions, without explicit temporal supervision. Our method produces a hierarchy of representations that align more closely with ground-truth human-annotated events (+15.3%) than state-of-the-art unsupervised baselines. Our results are comparable to heavily-supervised baselines in complex visual domains such as Chess Openings, YouCook2 and TutorialVQA datasets. Finally, we perform ablation studies illustrating the robustness of our approach. We release our code and demo visualizations in the Supplementary Material. S. Roychowdhury, Sumedh A. Sontakke, Laurent Itti, M. Sarkar, Milan Aggarwal, Pinkesh Badjatiya, Nikaash Puri, Balaji Krishnamurthy |
ICPR | 3 |
| 2022 | Beneficial Perturbation Network for Designing General Adaptive Artificial Intelligence SystemsabstractThe human brain is the gold standard of adaptive learning. It not only can learn and benefit from experience, but also can adapt to new situations. In contrast, deep neural networks only learn one sophisticated but fixed mapping from inputs to outputs. This limits their applicability to more dynamic situations, where the input to output mapping may change with different contexts. A salient example is continual learning-learning new independent tasks sequentially without forgetting previous tasks. Continual learning of multiple tasks in artificial neural networks using gradient descent leads to catastrophic forgetting, whereby a previously learned mapping of an old task is erased when learning new mappings for new tasks. Herein, we propose a new biologically plausible type of deep neural network with extra, out-of-network, task-dependent biasing units to accommodate these dynamic situations. This allows, for the first time, a single network to learn potentially unlimited parallel input to output mappings, and to switch on the fly between them at runtime. Biasing units are programed by leveraging beneficial perturbations (opposite to well-known adversarial perturbations) for each task. Beneficial perturbations for a given task bias the network toward that task, essentially switching the network into a different mode to process that task. This largely eliminates catastrophic interference between tasks. Our approach is memory-efficient and parameter-efficient, can accommodate many tasks, and achieves the state-of-the-art performance across different tasks and domains. Shixian Wen, Amanda Rios, Yunhao Ge, Laurent Itti |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | A Peek Into the Reasoning of Neural Networks: Interpreting With Structural Visual ConceptsabstractDespite substantial progress in applying neural networks (NN) to a wide variety of areas, they still largely suffer from a lack of transparency and interpretability. While recent developments in explainable artificial intelligence attempt to bridge this gap (e.g., by visualizing the correlation between input pixels and final outputs), these approaches are limited to explaining low-level relationships, and crucially, do not provide insights on error correction. In this work, we propose a framework (VRX) to interpret classification NNs with intuitive structural visual concepts. Given a trained classification model, the proposed VRX extracts relevant class-specific visual concepts and organizes them using structural concept graphs (SCG) based on pairwise concept relationships. By means of knowledge distillation, we show VRX can take a step towards mimicking the reasoning process of NNs and provide logical, concept-level explanations for final model decisions. With extensive experiments, we empirically show VRX can meaningfully answer "why" and "why not" questions about the prediction, providing easy-to-understand insights about the reasoning process. We also show that these insights can potentially provide guidance on improving NN’s performance. Yunhao Ge, Zhi Xu 0013, Meng Zheng 0002, Srikrishna Karanam, Terrence Chen, Laurent Itti, Ziyan Wu 0001 |
CVPR | 7 |
| 2021 | Multi-Task Occlusion Learning for Real-Time Visual Object TrackingabstractOcclusion handling is one of the important challenges in the field of visual tracking, especially for real-time applications, where further processing for occlusion reasoning may not always be possible. In this paper, an occlusion-aware real-time object tracker is proposed, which enhances the baseline SiamRPN model with an additional branch that directly predicts the occlusion level of the object. Experimental results on GOT-10k and VOT benchmarks show that learning to predict occlusion levels end-to-end in this multi-task learning framework helps improve tracking accuracy, especially on frames that contain occlusions. Up to 7% improvement on EAO scores can be observed for occluded frames, which are only 11% of the data. The performance results over all frames also indicate the model does favorably compared to the other trackers. Gozde Sahin, Laurent Itti |
ICIP | 2 |
| 2021 | Zero-shot Synthesis with Group-Supervised Learning
Yunhao Ge, Sami Abu-El-Haija, Gan Xin, Laurent Itti |
ICLR | 4 |
| 2021 | Causal Curiosity: RL Agents Discovering Self-supervised Experiments for Causal Representation LearningabstractHumans show an innate ability to learn the regularities of the world through interaction. By performing experiments in our environment, we are able to discern the causal factors of variation and infer how they affect the dynamics of our world. Analogously, here we attempt to equip reinforcement learning agents with the ability to perform experiments that facilitate a categorization of the rolled-out trajectories, and to subsequently infer the causal factors of the environment in a hierarchical manner. We introduce a novel intrinsic reward, called causal curiosity, and show that it allows our agents to learn optimal sequences of actions, and to discover causal factors in the dynamics. The learned behavior allows the agent to infer a binary quantized representation for the ground-truth causal factors in every environment. Additionally, we find that these experimental behaviors are semantically meaningful (e.g., to differentiate between heavy and light blocks, our agents learn to lift them), and are learnt in a self-supervised manner with approximately 2.5 times less data than conventional supervised planners. We show that these behaviors can be re-purposed and fine-tuned (e.g., from lifting to pushing or other downstream tasks). Finally, we show that the knowledge of causal factor representations aids zero-shot learning for more complex tasks. Sumedh A. Sontakke, Arash Mehrjou, Laurent Itti, Bernhard Schölkopf |
ICML | 3 |
| 2021 | Shaped Policy Search for Evolutionary Strategies using Waypoints*
Kiran Lekkala, Laurent Itti |
ICRA | 2 |
| 2020 | Pose Augmentation: Class-Agnostic Object Pose Transformation for Object Recognition
Yunhao Ge, Jiaping Zhao, Laurent Itti |
ECCV (28) | 3 |
| 2020 | Lifelong Learning Without a Task OracleabstractSupervised deep neural networks are known to undergo a sharp decline in the accuracy of older tasks when new tasks are learned, termed “catastrophic forgetting”. Many state-of-the-art solutions to continual learning rely on biasing and/or partitioning a model to accommodate successive tasks incrementally. However, these methods largely depend on the availability of a task-oracle to confer task identities to each test sample, without which the models are entirely unable to perform. To address this shortcoming, we propose and compare several candidate task-assigning mappers which require very little memory overhead: (1) Incremental unsupervised prototype assignment using either nearest means, Gaussian Mixture Models or fuzzy ART backbones; (2) Supervised incremental prototype assignment with fast fuzzy ARTMAP; (3) Shallow perceptron trained via a dynamic coreset. Our proposed model variants are trained either from pre-trained feature extractors or task-dependent feature embeddings of the main classifier network. We apply these pipeline variants to continual learning benchmarks, comprised of either sequences of several datasets or within one single dataset. Overall, these methods, despite their simplicity and compactness, perform very close to a ground truth oracle, especially in experiments of inter-dataset task assignment. Moreover, best-performing variants only impose an average cost of 1.7% parameter memory increase. Amanda Rios, Laurent Itti |
ICTAI | 2 |
| 2020 | Learning visual variation for object recognition
Jatuporn Toy Leksut, Jiaping Zhao, Laurent Itti |
Image Vis. Comput. | 3 |
| 2019 | Closed-Loop Memory GAN for Continual LearningabstractSequential learning of tasks using gradient descent leads to an unremitting decline in the accuracy of tasks for which training data is no longer available, termed catastrophic forgetting. Generative models have been explored as a means to approximate the distribution of old tasks and bypass storage of real data. Here we propose a cumulative closed-loop memory replay GAN (CloGAN) provided with external regularization by a small memory unit selected for maximum sample diversity. We evaluate incremental class learning using a notoriously hard paradigm, single-headed learning, in which each task is a disjoint subset of classes in the overall dataset, and performance is evaluated on all previous classes. First, we show that when constructing a dynamic memory unit to preserve sample heterogeneity, model performance asymptotically approaches training on the full dataset. We then show that using a stochastic generator to continuously output fresh new images during training increases performance significantly further meanwhile generating quality images. We compare our approach to several baselines including fine-tuning by gradient descent (FGD), Elastic Weight Consolidation (EWC), Deep Generative Replay (DGR) and Memory Replay GAN (MeRGAN). Our method has very low long-term memory cost, the memory unit, as well as negligible intermediate memory storage. Amanda Rios, Laurent Itti |
IJCAI | 2 |
| 2019 | Inertial-based Motion Capturing and Smart Training SystemabstractSmart coaching platforms are emerging which combine BodySensor-Networks with AI-based training software to monitor and analyze body motions of athletes, workers, or medical patients. This allows for new opportunities to explore algorithms to interpret body sensor data and provide analytical feedback for learning a physical task, refining body motions, or to protect from work-related injuries. This paper presents a solution to non-invasively equip a person with sensors of a Smart Training System (STS) to improve training efficiency during sport activities. Our system calculates the significance of each body part during physical activities and provides targeted feedback on which body locations are under-performing. In experiments, the system collected data from 13 inertial sensors attached to the entire body of inexperienced golf learners. Using an indoor golf training net with a central target with 3 concentric zones, 1,080 real-world golf swings of 11 participants were analyzed. During the first 30 swings of each participant, the system learned distributions of motions from each sensor, conditioned on swing performance reported by users from their hitting location on the target. In the later 70 swings, feedback was provided to a subgroup of 8 participants, by computing, for an optimal set of features determined during training, the largest discrepancy. The remaining 3 (control) participants received no feedback. From only 100 golf swings for each participant, our system led to significantly improved scores by on average 3.7x (t-test, p <; 0.0001) over the latter 70 swings. Our results suggest that the combination of motion sensors and processing developed here was able to yield significantly improved golf swing training. Jens Windau, Laurent Itti |
IROS | 2 |
| 2019 | Surprise! Predicting Infant Visual Attention in a Socially Assistive Robot Contingent Learning ParadigmabstractEarly intervention to address developmental disability in infants has the potential to promote improved outcomes in neurodevelopmental structure and function [1]. Researchers are starting to explore Socially Assistive Robotics (SAR) as a tool for delivering early interventions that are synergistic with and enhance human-administered therapy. For SAR to be effective, the robot must be able to consistently attract the attention of the infant in order to engage the infant in a desired activity. This work presents the analysis of eye gaze tracking data from five 6-8 month old infants interacting with a Nao robot that kicked its leg as a contingent reward for infant leg movement. We evaluate a Bayesian model of low-level surprise on video data from the infants' head-mounted camera and on the timing of robot behaviors as a predictor of infant visual attention. The results demonstrate that over 67% of infant gaze locations were in areas the model evaluated to be more surprising than average. We also present an initial exploration using surprise to predict the extent to which the robot attracts infant visual attention during specific intervals in the study. This work is the first to validate the surprise model on infants; our results indicate the potential for using surprise to inform robot behaviors that attract infant attention during SAR interactions. Lauren Klein, Laurent Itti, Beth A. Smith, Marcelo R. Rosales, Stefanos Nikolaidis, Maja J. Mataric |
RO-MAN | 2 |
| 2019 | Learning Invariant Features in Modulatory Networks through Conflict and AmbiguityabstractThis work lays the foundation for a framework of cortical learning based on the idea of a competitive column, which is inspired by the functional organization of neurons in the cortex. A column describes a prototypical organization for neurons that gives rise to an ability to learn scale, rotation, and translation-invariant features. This is empowered by a recently developed learning rule, conflict learning, which enables the network to learn over both driving and modulatory feedforward, feedback, and lateral inputs. The framework is further supported by introducing both a notion of neural ambiguity and an adaptive threshold scheme. Ambiguity, which captures the idea that too many decisions lead to indecision, gives the network a dynamic way to resolve locally ambiguous decisions. The adaptive threshold operates over multiple timescales to regulate neural activity under the varied arrival timings of input in a highly interconnected multilayer network with feedforward and feedback. The competitive column architecture is demonstrated on a large-scale (54,000 neurons and 18 million synapses), invariant model of border ownership. The model is trained on four simple, fixed-scale shapes: two squares, one rectangle, and one symmetric L-shape. Tested on 1899 synthetic shapes of varying scale and complexity, the model correctly assigned border ownership with 74% accuracy. The model's abilities were also illustrated on contours of objects taken from natural images. Combined with conflict learning, the competitive column and ambiguity give a better intuitive understanding of how feedback, modulation, and inhibition may interact in the brain to influence activation and learning. W. Shane Grant, Laurent Itti |
Neural Comput. | 2 |
| 2018 | Born-Again Neural NetworksabstractKnowledge Distillation (KD) consists of transferring “knowledge” from one machine learning model (the teacher) to another (the student). Commonly, the teacher is a high-capacity model with formidable performance, while the student is more compact. By transferring knowledge, one hopes to benefit from the student’s compactness, without sacrificing too much performance. We study KD from a new perspective: rather than compressing models, we train students parameterized identically to their teachers. Surprisingly, these Born-Again Networks (BANs), outperform their teachers significantly, both on computer vision and language modeling tasks. Our experiments with BANs based on DenseNets demonstrate state-of-the-art performance on the CIFAR-10 (3.5%) and CIFAR-100 (15.5%) datasets, by validation error. Additional experiments explore two distillation objectives: (i) Confidence-Weighted by Teacher Max (CWTM) and (ii) Dark Knowledge with Permuted Predictions (DKPP). Both methods elucidate the essential components of KD, demonstrating the effect of the teacher outputs on both predicted and non-predicted classes. Tommaso Furlanello, Zachary C. Lipton, Michael Tschannen, Laurent Itti, Anima Anandkumar |
ICML | 4 |
| 2018 | DeepVP: Deep Learning for Vanishing Point Detection on 1 Million Street View ImagesabstractWe propose a novel approach to detect vanishing points in images using a convolutional neural network (CNN) trained on a newly collected Google street-view image dataset. By utilizing the camera parameters and road direction data from Google street view, we collected a total of 1,053,425 images with inferred ground-truth vanishing points, along 23 worldwide routes totaling 125,165 kilometers. We then formulate vanishing point detection as a CNN classification problem using an output layer with 225 discrete possible vanishing point locations. Experimental results show that our deep vanishing point system outperforms the state-of-the-art algorithmic vanishing point detector. We achieved 99% accuracy in recovering the horizon line and 92% in locating the vanishing point within a ±5-degree range. Chin-Kai Chang, Jiaping Zhao, Laurent Itti |
ICRA | 3 |
| 2018 | Inertial Machine Monitoring System for Automated Failure DetectionabstractSmart manufacturing technologies are emerging which combine industrial equipment with Internet-of-Things (IoT) sensors to monitor and improve productivity of manufacturing. This allows for new opportunities to explore algorithms for predicting machine failures from attached sensor data. This paper presents a solution to non-invasively upgrade an existing machine with an Inertial Machine Monitoring System (IMMS) to detect and classify equipment failure or degraded state. We also provide a strategy to optimize the amount, placement locations, and efficiency of the sensors. In experiments, the system collected data from 36 inertial sensors placed at multiple locations on a 3D printer. Normal operation vs. 10 types of realworld abnormal equipment behavior (loose belt, failures of machine components) were detected and classified by Support Vector Machines and Neural Networks. Using under 1 minute of recording while running a test print, a recursively discovered best subset of 4 to 9 sensors yielded 11-way classification accuracy over 99%. Our results suggest that even a small sensor network and short test program can yield effective detection of machine degraded state and can facilitate early remediation. Jens Windau, Laurent Itti |
ICRA | 2 |
| 2018 | Salient object detection via a local and global method based on deep residual network
Dandan Zhu 0001, Ye Luo 0004, Xuan Shao, Qiangqiang Zhou, Laurent Itti |
J. Vis. Commun. Image Represent. | 6 |
| 2018 | shapeDTW: Shape Dynamic Time Warping
Jiaping Zhao, Laurent Itti |
Pattern Recognit. | 2 |
| 2017 | Saliency prediction based on new deep multi-layer convolution neural networkabstractRecent advances in saliency detection have utilized deep learning to obtain high-level features to detect salient regions. These advances have demonstrated superior results over previous works that utilize hand-crafted low-level features for saliency detection. In this paper, we propose a new multilayer Convolutional Neural Network (CNN) model to learn high-level features for saliency detection. Compared to other methods, our method presents two merits. First, when performing features extraction, apart from the convolution and pooling step in our method, we add Restricted Boltzmann Machine (RBM) into the CNN framework to obtain more accurate features in intermediate step. Second, in order to deal with case of non-linear classification, we add the Deep Belief Network (DBN) classifier at the end of this model to classify the salient and non-salient regions. Quantitative and qualitative experiments on three benchmark datasets demonstrate that our method performs favorably against the state-of-the-art methods. Dandan Zhu 0001, Ye Luo 0004, Xuan Shao, Laurent Itti |
ICIP | 4 |
| 2017 | Scanpath Prediction Based on High-Level Features and Memory Bias
Xuan Shao, Ye Luo 0004, Dandan Zhu 0001, Laurent Itti |
ICONIP (3) | 5 |
| 2017 | Deep Salient Object Detection via Hierarchical Network Learning
Dandan Zhu 0001, Ye Luo 0004, Xuan Shao, Laurent Itti |
ICONIP (3) | 5 |
| 2017 | Learning to Recognize Objects by Retaining Other Factors of VariationabstractMost ConvNets formulate object recognition from natural images as a single task classification problem, and attempt to learn features useful for object categories, but invariant to other factors of variation such as pose and illumination. They do not explicitly learn these other factors, instead, they usually discard them by pooling and normalization. Here, we take the opposite approach: we train ConvNets for object recognition by retaining other factors (pose in our case) and learning them jointly with object category. We design a new multi-task leaning (MTL) ConvNet, named disentangling CNN (disCNN), which explicitly enforces the disentangled representations of object identity and pose, and is trained to predict object categories and pose transformations. disCNN achieves significantly better object recognition accuracies than the baseline CNN trained solely to predict object categories on the iLab-20M dataset, a large-scale turntable dataset with detailed pose and lighting information. We further show that the pretrained features on iLab-20M generalize to both Washington RGB-D and ImageNet datasets, and the pretrained dis-CNN features are significantly better than the pretrained baseline CNN features for fine-tuning on ImageNet. Jiaping Zhao, Chin-Kai Chang, Laurent Itti |
WACV | 3 |
| 2017 | Improved Deep Learning of Object Category Using Pose InformationabstractDespite significant recent progress, the best available computer vision algorithms still lag far behind human capabilities, even for recognizing individual discrete objects under various poses, illuminations, and backgrounds. Here we present a new approach to using object pose information to improve deep network learning. While existing large-scale datasets, e.g. ImageNet, do not have pose information, we leverage the newly published turntable dataset, iLab-20M, which has 22M images of 704 object instances shot under different lightings, camera viewpoints and turntable rotations, to do more controlled object recognition experiments. We introduce a new convolutional neural network architecture, what/where CNN (2W-CNN), built on a linear-chain feedforward CNN (e.g., AlexNet), augmented by hierarchical layers regularized by object poses. Pose information is only used as feedback signal during training, in addition to category information, but is not needed during test. To validate the approach, we train both 2W-CNN and AlexNet using a fraction of the dataset, and 2W-CNN achieves 6% performance improvement in category prediction. We show mathematically that 2W-CNN has inherent advantages over AlexNet under the stochastic gradient descent (SGD) optimization procedure. Furthermore, we fine-tune object recognition on ImageNet by using the pretrained 2W-CNN and AlexNet features on iLab-20M, results show significant improvement compared with training AlexNet from scratch. Moreover, fine-tuning 2W-CNN features performs even better than fine-tuning the pretrained AlexNet features. These results show that pretrained features on iLab-20M generalize well to natural image datasets, and 2W-CNN learns better features for object recognition than AlexNet. Jiaping Zhao, Laurent Itti |
WACV | 2 |
| 2017 | Biologically plausible learning in neural networks with modulatory feedbackabstractAlthough Hebbian learning has long been a key component in understanding neural plasticity, it has not yet been successful in modeling modulatory feedback connections, which make up a significant portion of connections in the brain. We develop a new learning rule designed around the complications of learning modulatory feedback and composed of three simple concepts grounded in physiologically plausible evidence. Using border ownership as a prototypical example, we show that a Hebbian learning rule fails to properly learn modulatory connections, while our proposed rule correctly learns a stimulus-driven model. To the authors' knowledge, this is the first time a border ownership network has been learned. Additionally, we show that the rule can be used as a drop-in replacement for a Hebbian learning rule to learn a biologically consistent model of orientation selectivity, a network which lacks any modulatory connections. Our results predict that the mechanisms we use are integral for learning modulatory connections in the brain and furthermore that modulatory connections have a strong dependence on inhibition. W. Shane Grant, James Tanner, Laurent Itti |
Neural Networks | 3 |
| 2016 | iLab-20M: A Large-Scale Controlled Object Dataset to Investigate Deep LearningabstractTolerance to image variations (e.g., translation, scale, pose, illumination, background) is an important desired property of any object recognition system, be it human or machine. Moving towards increasingly bigger datasets has been trending in computer vision especially with the emergence of highly popular deep learning models. While being very useful for learning invariance to object inter-and intra-class shape variability, these large-scale wild datasets are not very useful for learning invariance to other parameters urging researchers to resort to other tricks for training models. In this work, we introduce a large-scale synthetic dataset, which is freely and publicly available, and use it to answer several fundamental questions regarding selectivity and invariance properties of convolutional neural networks. Our dataset contains two parts: a) objects shot on a turntable: 15 categories, 8 rotation angles, 11 cameras on a semi-circular arch, 5 lighting conditions, 3 focus levels, variety of backgrounds (23.4 per instance) generating 1320 images per instance (about 22 million images in total), and b) scenes: in which a robotic arm takes pictures of objects on a 1:160 scale scene. We study: 1) invariance and selectivity of different CNN layers, 2) knowledge transfer from one object category to another, 3) systematic or random sampling of images to build a train set, 4) domain adaptation from synthetic to natural scenes, and 5) order of knowledge delivery to CNNs. We also discuss how our analyses can lead the field to develop more efficient deep learning methods. Ali Borji, Saeed Izadi, Laurent Itti |
CVPR | 3 |
| 2016 | Walking compass with head-mounted IMU sensorabstractEmerging wearable technologies offer new sensor placement options on the human body. Particularly, head-mounted glass-wear opens up new data capturing possibilities directly from the human head. This allows exploring new cyber-robotics algorithms (robotics sensors and human motor plant). Glass-wear systems, however, require additional compensation for head motions that will affect the captured sensor data. Particularly, pedestrian dead-reckoning (PDR), activity recognition, and other applications are limited or restricted when head-mounted sensors are used, because of possible confusion between head and body movements. Thus, previous PDR approaches typically required to keep the head pointing direction aligned with the walking direction to avoid positional errors. This paper presents a head-mounted orientation system (HOS) that identifies and filters out interfering head motions in 3 steps. Step 1 transforms inertial sensor data into a stable normalized coordinate system (roll/pitch motion compensated). Step 2 compares walking patterns before and after a rotating motion. Step 3 eliminates interfering head motions from sensor data by dynamically adjusting the noise parameters of the extended Kalman filter. HOS has been implemented on a Google Glass platform and achieved high accuracy in tracking a person's path even in the presence of head movements (within 2.5% of traveled distance) when tested in multiple real-world scenarios. By eliminating head motions, HOS not only enables accurate PDR, but also facilitates the task for downstream activity recognition algorithms. Jens Windau, Laurent Itti |
ICRA | 2 |
| 2016 | Decomposing time series with application to temporal segmentationabstractWe propose a novel univariate time series decomposition algorithm to partition temporal sequences into homogeneous segments. Unlike most existing temporal segmentation approaches, which generally build statistical models of temporal observations and then detect change points using inference or hypothesis testing techniques, our algorithm requires no domain knowledge, is insensitive to the choice of design parameters and has low time complexity. Our algorithm first symbolizes the time series into a string, and then decomposes the string recursively, similar to the construction process of a decision-tree classifier. We extend this univariate decomposition algorithm to multivariate cases by decomposing each dimension as an univariate time series and then searching for temporal transition points in a coarse-to-fine manner. We evaluate and compare our algorithm to two state-of-the-art approaches on synthetic data, CMU motion capture data, and action videos. Experimental results demonstrate the effectiveness of our approach, which yields both significantly higher precision and recall of temporal transition points. Jiaping Zhao, Laurent Itti |
WACV | 2 |
| 2016 | Learning a Combined Model of Visual Saliency for Fixation PredictionabstractA large number of saliency models, each based on a different hypothesis, have been proposed over the past 20 years. In practice, while subscribing to one hypothesis or computational principle makes a model that performs well on some types of images, it hinders the general performance of a model on arbitrary images and large-scale data sets. One natural approach to improve overall saliency detection accuracy would then be fusing different types of models. In this paper, inspired by the success of late-fusion strategies in semantic analysis and multi-modal biometrics, we propose to fuse the state-of-the-art saliency models at the score level in a para-boosting learning fashion. First, saliency maps generated by several models are used as confidence scores. Then, these scores are fed into our para-boosting learner (i.e., support vector machine, adaptive boosting, or probability density estimator) to generate the final saliency map. In order to explore the strength of para-boosting learners, traditional transformation-based fusion strategies, such as Sum, Min, and Max, are also explored and compared in this paper. To further reduce the computation cost of fusing too many models, only a few of them are considered in the next step. Experimental results show that score-level fusion outperforms each individual model and can further reduce the performance gap between the current models and the human inter-observer model. Ali Borji, C.-C. Jay Kuo, Laurent Itti |
IEEE Trans. Image Process. | 4 |
| 2016 | Classifying Time Series Using Local Descriptors with Hybrid SamplingabstractTime series classification (TSC) arises in many fields and has a wide range of applications. Here, we adopt the bag-of-words (BoW) framework to classify time series. Our algorithm first samples local subsequences from time series at feature-point locations when available. It then builds local descriptors, and models their distribution by Gaussian mixture models (GMM), and at last it computes a Fisher Vector (FV) to encode each time series. The encoded FV representations of time series are readily used by existing classifiers, e.g., SVM, for training and prediction. In our work, we focus on detecting better feature points and crafting better local representations, while using existing techniques to learn codebook and encode time series. Specifically, we develop an efficient and effective peak and valley detection algorithm from real-case time series data. Subsequences are sampled from these peaks and valleys, instead of sampled randomly or uniformly as was done previously. Then, two local descriptors, Histogram of Oriented Gradients (HOG-1D) and Dynamic time warping-Multidimensional scaling (DTW-MDS), are designed to represent sampled subsequences. Both descriptors complement each other, and their fused representation is shown to be more descriptive than individual ones. We test our approach extensively on 43 UCR time series datasets, and obtain significantly improved classification accuracies over existing approaches, including NNDTW and shapelet transform. Jiaping Zhao, Laurent Itti |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2015 | Fixation bank: Learning to reweight fixation candidatesabstractPredicting where humans will fixate in a scene has many practical applications. Biologically-inspired saliency models decompose visual stimuli into feature maps across multiple scales, and then integrate different feature channels, e.g., in a linear, MAX, or MAP. However, to date there is no universally accepted feature integration mechanism. Here, we propose a new a data-driven solution: We first build a “fixation bank” by mining training samples, which maintains the association between local patterns of activation, in 4 feature channels (color, intensity, orientation, motion) around a given location, and corresponding human fixation density at that location. During testing, we decompose feature maps into blobs, extract local activation patterns around each blob, match those patterns against the fixation bank by group lasso, and determine weights of blobs based on reconstruction errors. Our final saliency map is the weighted sum of all blobs. Our system thus incorporates some amount of spatial and featural context information into the location-dependent weighting mechanism. Tested on two standard data sets (DIEM for training and test, and CRCNS for test only; total of 23,670 training and 15,793 + 4,505 test frames), our model slightly but significantly outperforms 7 state-of-the-art saliency models. Jiaping Zhao, Christian Siagian, Laurent Itti |
CVPR | 3 |
| 2015 | BIK-BUS: Biologically Motivated 3D Keypoint Based on Bottom-Up SaliencyabstractOne of the major problems found when developing a 3D recognition system involves the choice of keypoint detector and descriptor. To help solve this problem, we present a new method for the detection of 3D keypoints on point clouds and we perform benchmarking between each pair of 3D keypoint detector and 3D descriptor to evaluate their performance on object and category recognition. These evaluations are done in a public database of real 3D objects. Our keypoint detector is inspired by the behavior and neural architecture of the primate visual system. The 3D keypoints are extracted based on a bottom-up 3D saliency map, that is, a map that encodes the saliency of objects in the visual environment. The saliency map is determined by computing conspicuity maps (a combination across different modalities) of the orientation, intensity, and color information in a bottom-up and in a purely stimulus-driven manner. These three conspicuity maps are fused into a 3D saliency map and, finally, the focus of attention (or keypoint location) is sequentially directed to the most salient points in this map. Inhibiting this location automatically allows the system to attend to the next most salient location. The main conclusions are: with a similar average number of keypoints, our 3D keypoint detector outperforms the other eight 3D keypoint detectors evaluated by achieving the best result in 32 of the evaluated metrics in the category and object recognition experiments, when the second best detector only obtained the best result in eight of these metrics. The unique drawback is the computational time, since biologically inspired 3D keypoint based on bottom-up saliency is slower than the other detectors. Given that there are big differences in terms of recognition performance, size and time requirements, the selection of the keypoint detector and descriptor has to be matched to the desired task and we give some directions to facilitate this choice. Sílvio Filipe, Laurent Itti, Luís A. Alexandre |
IEEE Trans. Image Process. | 2 |
| 2014 | Human vs. Computer in Scene and Object RecognitionabstractSeveral decades of research in computer and primate vision have resulted in many models (some specialized for one problem, others more general) and invaluable experimental data. Here, to help focus research efforts onto the hardest unsolved problems, and bridge computer and human vision, we define a battery of 5 tests that measure the gap between human and machine performances in several dimensions (generalization across scene categories, generalization from images to edge maps and line drawings, invariance to rotation and scaling, local/global information with jumbled images, and object recognition performance). We measure model accuracy and the correlation between model and human error patterns. Experimenting over 7 datasets, where human data is available, and gauging 14 well-established models, we find that none fully resembles humans in all aspects, and we learn from each test which models and features are more promising in approaching humans in the tested dimension. Across all tests, we find that models based on local edge histograms consistently resemble humans more, while several scene statistics or "gist" models do perform well with both scenes and objects. While computer vision has long been inspired by human vision, we believe systematic efforts, such as this, will help better identify shortcomings of models and find new paths forward. Ali Borji, Laurent Itti |
CVPR | 2 |
| 2014 | Integrating human context and occlusion reasoning to improve handheld object trackingabstractTracking an unknown number of various objects involving occlusion and multiple entry and exit points automatically is a challenging problem. Here we integrate spatial knowledge of human-object interactions into a high performing tracker to show that human context can further improve both detection and tracking. We use the DARPA Mind's Eye Action Recognition Dataset, which is comprised of street level scenes with humans interacting with handheld objects, to show this improvement. We find that human context can greatly reduce the number of false positive detections at the expense of increasing false negatives over a large test set (>230k frames). To minimize this, we add occlusion reasoning, where object detections are hallucinated when a human detection overlaps an object detection. These components together result in an average F1improvement of 107% per object category and a 69% reduction in track latency. Daniel F. Parks, Laurent Itti |
ICIP | 2 |
| 2014 | Performance Evaluation of Neuromorphic-Vision Object Recognition AlgorithmsabstractThe U.S. Defense Advanced Research Projects Agency's (DARPA) Neovision2 program aims to develop artificial vision systems based on the design principles employed by mammalian vision systems. Three such algorithms are briefly described in this paper. These neuromorphic-vision systems' performance in detecting objects in video was measured using a set of annotated clips. This paper describes the results of these evaluations including the data domains, metrics, methodologies, performance over a range of operating points and a comparison with computer vision based baseline algorithms. Rangachar Kasturi, Dmitry B. Goldgof, Ekambaram Rajmadhan, Gill A. Pratt, Eric Krotkov, Douglas Hackett, Yang Ran, Qinfen Zheng, Rajeev Sharma, Mark Peot, Mario Aguilar, Deepak Khosla, Kyungnam Kim, Lior Elazary, Randolph Voorhies, Daniel F. Parks, Laurent Itti |
ICPR | 19 |
| 2014 | What/Where to Look Next? Modeling Top-Down Visual Attention in Complex Interactive EnvironmentsabstractSeveral visual attention models have been proposed for describing eye movements over simple stimuli and tasks such as free viewing or visual search. Yet, to date, there exists no computational framework that can reliably mimic human gaze behavior in more complex environments and tasks such as urban driving. In addition, benchmark datasets, scoring techniques, and top-down model architectures are not yet well understood. In this paper, we describe new task-dependent approaches for modeling top-down overt visual attention based on graphical models for probabilistic inference and reasoning. We describe a dynamic Bayesian network that infers probability distributions over attended objects and spatial locations directly from observed data. Probabilistic inference in our model is performed over object-related functions that are fed from manual annotations of objects in video scenes or by state-of-the-art object detection/recognition algorithms. Evaluating over approximately 3 h (approximately 315 000 eye fixations and 12 000 saccades) of observers playing three video games (time-scheduling, driving, and flight combat), we show that our approach is significantly more predictive of eye fixations compared to: 1) simpler classifier-based models also developed here that map a signature of a scene (multimodal information from gist, bottom-up saliency, physical actions, and events) to eye positions; 2) 14 state-of-the-art bottom-up saliency models; and 3) brute-force algorithms such as mean eye position. Our results show that the proposed model is more effective in employing and reasoning over spatio-temporal visual data compared with the state-of-the-art. Ali Borji, Dicky N. Sihite, Laurent Itti |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2013 | Schema-Driven, Space-Supported Random Accessible Memory Systems for Manipulation of Symbolic Working Memory
Nader Noori, Laurent Itti |
CogSci | 2 |
| 2013 | Where What You Count is What Really Counts
Nader Noori, Laurent Itti |
CogSci | 2 |
| 2013 | Traces of Intellectual Working Memory Tasks on Visuospatial Short-Term Memory
Nader Noori, Laurent Itti |
CogSci | 2 |
| 2013 | Analysis of Scores, Datasets, and Models in Visual Saliency PredictionabstractSignificant recent progress has been made in developing high-quality saliency models. However, less effort has been undertaken on fair assessment of these models, over large standardized datasets and correctly addressing confounding factors. In this study, we pursue a critical and quantitative look at challenges (e.g., center-bias, map smoothing) in saliency modeling and the way they affect model accuracy. We quantitatively compare 32 state-of-the-art models (using the shuffled AUC score to discount center-bias) on 4 benchmark eye movement datasets, for prediction of human fixation locations and scan path sequence. We also account for the role of map smoothing. We find that, although model rankings vary, some (e.g., AWS, LG, AIM, and HouNIPS) consistently outperform other models over all datasets. Some models work well for prediction of both fixation locations and scan path sequence (e.g., Judd, GBVS). Our results show low prediction accuracy for models over emotional stimuli from the NUSEF dataset. Our last benchmark, for the first time, gauges the ability of models to decode the stimulus category from statistics of fixations, saccades, and model saliency values at fixated locations. In this test, ITTI and AIM models win over other models. Our benchmark provides a comprehensive high-level picture of the strengths and weaknesses of many popular models, and suggests future research directions in saliency modeling. Ali Borji, Hamed Rezazadegan Tavakoli, Dicky N. Sihite, Laurent Itti |
ICCV | 4 |
| 2013 | Mobile robot navigation system in outdoor pedestrian environment using vision-based road recognitionabstractWe present a mobile robot navigation system guided by a novel vision-based road recognition approach. The system represents the road as a set of lines extrapolated from the detected image contour segments. These lines enable the robot to maintain its heading by centering the vanishing point in its field of view, and to correct the long term drift from its original lateral position. We integrate odometry and our visual road recognition system into a grid-based local map that estimates the robot pose as well as its surroundings to generate a movement path. Our road recognition system is able to estimate the road center on a standard dataset with 25,076 images to within 11.42 cm (with respect to roads at least 3 m wide). It outperforms three other state-of-the-art systems. In addition, we extensively test our navigation system in four busy college campus environments using a wheeled robot. Our tests cover more than 5 km of autonomous driving without failure. This demonstrates robustness of the proposed approach against challenges that include occlusion by pedestrians, non-standard complex road markings and shapes, shadows, and miscellaneous obstacle objects. Christian Siagian, Chin-Kai Chang, Laurent Itti |
ICRA | 3 |
| 2013 | Deep Learning on Natural Viewing Behaviors to Differentiate Children with Fetal Alcohol Spectrum Disorder
Po-He Tseng, Angelina Paolozza, Douglas P. Munoz, James N. Reynolds, Laurent Itti |
IDEAL | 5 |
| 2013 | Beobot 2.0: Autonomous mobile robot localization and navigation in outdoor pedestrian environmentabstractWe present Beobot 2.0 [1], an autonomous mobile robot designed to operate in unconstrained urban environments. The goal of the project is to create service robots that can be deployed for various tasks that require long range travel. Over the past two years, Beobot has successfully traversed various paths across the USC campus, demonstrating its robustness in recognizing and following different types of roads, avoiding obstacles such as pedestrians and service vehicles, and finding its way to the goal. Chin-Kai Chang, Christian Siagian, Laurent Itti |
IROS | 3 |
| 2013 | Finding planes in LiDAR point clouds for real-time registrationabstractWe present a robust plane finding algorithm that when combined with plane-based frame-to-frame registration gives accurate real-time pose estimation. Our plane extraction is capable of handling large and sparse datasets such as those generated from spinning multi-laser sensors such as the Velodyne HDL-32E LiDAR. We test our algorithm on frame-to-frame registration in a closed-loop indoor path comprising 827 successive 3D laser scans (over 57 million points), using no additional information (e.g., odometry, IMU). Our algorithm outperforms, in both accuracy and time, three state-of-the-art methods, based on iterative closest point (ICP), plane-based randomized Hough transform, and planar region growing. W. Shane Grant, Randolph Voorhies, Laurent Itti |
IROS | 3 |
| 2013 | Situation awareness via sensor-equipped eyeglassesabstractNew smartphone technologies are emerging which combine head-mounted displays (HMD) with standard functions such as receiving phone calls, emails, and helping with navigation. This opens new opportunities to explore cyber robotics algorithms (robotics sensors and human motor plant). To make these devices more adaptive to the environmental conditions, user behavior, and user preferences, it is important to allow the sensor-equipped devices to efficiently adapt and respond to user activities (e.g., disable incoming phone calls in an elevator, activate video recording while car driving). This paper hence presents a situation awareness system (SAS) for head-mounted smartphones. After collecting data from inertial sensors (accelerometers, gyroscopes), and video data (camera), SAS performs activity classification in three steps. Step 1 transforms inertial sensor data into a head orientation-independent and stable normalized coordinate system. Step 2 extracts critical features (statistical, physical, GIST). Step 3 classifies activities (Naive Bayes classifier), distinguishes between environments (Support Vector Machine), and finally combines both results (Hidden Markov Model) for further improvement. SAS has been implemented on a sensor-equipped eyeglasses prototype and achieved high accuracy (81.5%) when distinguishing between 20 real-world activities. Jens Windau, Laurent Itti |
IROS | 2 |
| 2013 | Bayesian optimization explains human active searchabstractMany real-world problems have complicated objective functions. To optimize such functions, humans utilize sophisticated sequential decision-making strategies. Many optimization algorithms have also been developed for this same purpose, but how do they compare to humans in terms of both performance and behavior? We try to unravel the general underlying algorithm people may be using while searching for the maximum of an invisible 1D function. Subjects click on a blank screen and are shown the ordinate of the function at each clicked abscissa location. Their task is to find the function’s maximum in as few clicks as possible. Subjects win if they get close enough to the maximum location. Analysis over 23 non-maths undergraduates, optimizing 25 functions from different families, shows that humans outperform 24 well-known optimization algorithms. Bayesian Optimization based on Gaussian Processes, which exploit all the x values tried and all the f(x) values obtained so far to pick the next x, predicts human performance and searched locations better. In 6 follow-up controlled experiments over 76 subjects, covering interpolation, extrapolation, and optimization tasks, we further confirm that Gaussian Processes provide a general and unified theoretical account to explain passive and active function learning and search in humans. Ali Borji, Laurent Itti |
NIPS | 2 |
| 2013 | State-of-the-Art in Visual Attention ModelingabstractModeling visual attention--particularly stimulus-driven, saliency-based attention--has been a very active research area over the past 25 years. Many different models of attention are now available which, aside from lending theoretical contributions to other fields, have demonstrated successful applications in computer vision, mobile robotics, and cognitive systems. Here we review, from a computational perspective, the basic concepts of attention implemented in these models. We present a taxonomy of nearly 65 models, which provides a critical comparison of approaches, their capabilities, and shortcomings. In particular, 13 criteria derived from behavioral and computational studies are formulated for qualitative comparison of attention models. Furthermore, we address several challenging issues with models, including biological plausibility of the computations, correlation with eye movement datasets, bottom-up and top-down dissociation, and constructing meaningful performance measures. Finally, we highlight current research trends in attention modeling and provide insights for future. Ali Borji, Laurent Itti |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2013 | Quantitative Analysis of Human-Model Agreement in Visual Saliency Modeling: A Comparative StudyabstractVisual attention is a process that enables biological and machine vision systems to select the most relevant regions from a scene. Relevance is determined by two components: 1) top-down factors driven by task and 2) bottom-up factors that highlight image regions that are different from their surroundings. The latter are often referred to as "visual saliency." Modeling bottom-up visual saliency has been the subject of numerous research efforts during the past 20 years, with many successful applications in computer vision and robotics. Available models have been tested with different datasets (e.g., synthetic psychological search arrays, natural images or videos) using different evaluation scores (e.g., search slopes, comparison to human eye tracking) and parameter settings. This has made direct comparison of models difficult. Here, we perform an exhaustive comparison of 35 state-of-the-art saliency models over 54 challenging synthetic patterns, three natural image datasets, and two video datasets, using three evaluation scores. We find that although model rankings vary, some models consistently perform better. Analysis of datasets reveals that existing datasets are highly center-biased, which influences some of the evaluation scores. Computational complexity analysis shows that some models are very fast, yet yield competitive eye movement prediction accuracy. Different models often have common easy/difficult stimuli. Furthermore, several concerns in visual saliency modeling, eye movement datasets, and evaluation scores are discussed and insights for future work are provided. Our study allows one to assess the state-of-the-art, helps to organizing this rapidly growing field, and sets a unified comparison framework for gauging future efforts, similar to the PASCAL VOC challenge in the object recognition and detection domains. Ali Borji, Dicky N. Sihite, Laurent Itti |
IEEE Trans. Image Process. | 3 |
| 2012 | An Object-Based Bayesian Framework for Top-Down Visual AttentionabstractWe introduce a new task-independent framework to model top-down overt visual attention based on graph-ical models for probabilistic inference and reasoning. We describe a Dynamic Bayesian Network (DBN) that infers probability distributions over attended objects and spatial locations directly from observed data. Probabilistic inference in our model is performed over object-related functions which are fed from manual annotations of objects in video scenes or by state-of-the-art object detection models. Evaluating over ∼3 hours (appx. 315,000 eye fixations and 12,600 saccades) of observers playing 3 video games (time-scheduling, driving, and flight combat), we show that our approach is significantly more predictive of eye fixations compared to: 1) simpler classifier-based models also developed here that map a signature of a scene (multi-modal information from gist, bottom-up saliency, physical actions, and events) to eye positions, 2) 14 state-of-the-art bottom-up saliency models, and 3) brute-force algorithms such as mean eye position. Our results show that the proposed model is more effective in employing and reasoning over spatio-temporal visual data. Ali Borji, Dicky N. Sihite, Laurent Itti |
AAAI | 3 |
| 2012 | Neuromorphic Bayesian Surprise for Far-Range Event DetectionabstractIn this paper we address the problem of detecting small, rare events in very high resolution, far-field video streams. Rather than learning color distributions for individual pixels, our method utilizes a uniquely structured network of Bayesian learning units which compute a combined measure of "surprise" across multiple spatial and temporal scales on various visual features. The features used, as well as the learning rules for these units are derived from recent work in computational neuroscience. We test the system extensively on both real and virtual data, and show that it out-performs a standard foreground/background segmentation approach as well as a standard visual saliency algorithm. Randolph Voorhies, Lior Elazary, Laurent Itti |
AVSS | 3 |
| 2012 | Exploiting local and global patch rarities for saliency detectionabstractWe introduce a saliency model based on two key ideas. The first one is considering local and global image patch rarities as two complementary processes. The second one is based on our observation that for different images, one of the RGB and Lab color spaces outperforms the other in saliency detection. We propose a framework that measures patch rarities in each color space and combines them in a final map. For each color channel, first, the input image is partitioned into non-overlapping patches and then each patch is represented by a vector of coefficients that linearly reconstruct it from a learned dictionary of patches from natural scenes. Next, two measures of saliency (Local and Global) are calculated and fused to indicate saliency of each patch. Local saliency is distinctiveness of a patch from its surrounding patches. Global saliency is the inverse of a patch's probability of happening over the entire image. The final saliency map is built by normalizing and fusing local and global saliency maps of all channels from both color systems. Extensive evaluation over four benchmark eye-tracking datasets shows the significant advantage of our approach over 10 state-of-the-art saliency models. Ali Borji, Laurent Itti |
CVPR | 2 |
| 2012 | Probabilistic learning of task-specific visual attentionabstractDespite a considerable amount of previous work on bottom-up saliency modeling for predicting human fixations over static and dynamic stimuli, few studies have thus far attempted to model top-down and task-driven influences of visual attention. Here, taking advantage of the sequential nature of real-world tasks, we propose a unified Bayesian approach for modeling task-driven visual attention. Several sources of information, including global context of a scene, previous attended locations, and previous motor actions, are integrated over time to predict the next attended location. Recording eye movements while subjects engage in 5 contemporary 2D and 3D video games, as modest counterparts of everyday tasks, we show that our approach is able to predict human attention and gaze better than the state-of-the-art, with a large margin (about 15% increase in prediction accuracy). The advantage of our approach is that it is automatic and applicable to arbitrary visual tasks. Ali Borji, Dicky N. Sihite, Laurent Itti |
CVPR | 3 |
| 2012 | Salient Object Detection: A Benchmark
Ali Borji, Dicky N. Sihite, Laurent Itti |
ECCV (2) | 3 |
| 2012 | Saliency mapping enhanced by symmetry from local phaseabstractWe describe a method of generating saliency maps that combines symmetry with traditional contrast information. Log Gabor filtering is used to obtain local frequency information that can be used to both calculate local symmetry and edge responses. The symmetry information is combined with center-surround responses from color (either in RGB or LAB), orientation, and intensity features. The algorithm is evaluated on the Kootstra dataset and shown to significantly outperform three state-of-the-art models that utilize either contrast only or symmetry only based responses. A subjective evaluation by 13 human observers of image regions selected by 3 variants of our algorithm shows the significant benefit of including symmetry information in addition to contrast-based features. W. Shane Grant, Laurent Itti |
ICIP | 2 |
| 2012 | Modeling the influence of action on spatial attention in visual interactive environmentsabstractA large number of studies have been reported on top-down influences of visual attention. However, less progress have been made in understanding and modeling its mechanisms in real-world tasks. In this paper, we propose an approach for learning spatial attention taking into account influences of physical actions on top-down attention. For this purpose, we focus on interactive visual environments (video games) which are modest real-world simulations, where a player has to attend to certain aspects of visual stimuli and perform actions to achieve a goal. The basic idea is to learn a mapping from current mental state of the game player, represented by past actions and observations, to its gaze fixation. A data-driven approach is followed where we train a model from the data of some players and test it over a new subject. In particular, two contributions this paper makes are: 1) employing multi-modal information including mean eye position, gist of a scene, physical actions, bottom-up saliency, and tagged events for state representation and 2) analysis of different methods of combining bottom-up and top-down influences. Comparing with other top-down task-driven and bottom-up spatio-temporal models, our approach shows higher NSS scores in predicting eye positions. Ali Borji, Dicky N. Sihite, Laurent Itti |
ICRA | 3 |
| 2012 | Mobile robot monocular vision navigation based on road region and boundary estimationabstractWe present a monocular vision-based navigation system that incorporates two contrasting approaches: region segmentation that computes the road appearance, and road boundary detection that estimates the road shape. The former approach segments the image into multiple regions, then selects and tracks the most likely road appearance. On the other hand, the latter detects the vanishing point and road boundaries to estimate the shape of the road. Our algorithm operates in urban road settings and requires no training or camera calibration to maximize its adaptability to many environments. We tested our system in 1 indoor and 3 outdoor urban environments using our ground-based robot, Beobot 2.0, for real-time autonomous visual navigation. In 20 trial runs the robot was able to travel autonomously for 98.19% of the total route length of 316.60m. Chin-Kai Chang, Christian Siagian, Laurent Itti |
IROS | 3 |
| 2011 | Computational Modeling of Top-down Visual Attention in Interactive EnvironmentsabstractModeling how visual saliency guides the deployment of attention over visual scenes has attracted much interest recently — among both computer vision and experimental/computational researchers — since visual attention is a key function of both machine and biological vision systems. Research efforts in computer vision have mostly been focused on modeling bottom-up saliency. Strong influences o n attention and eye movements, however, come from instantaneous task demands. Here, we propose models of top-down visual guidance considering task influences. The n ew models estimate the state of a human subject performing a task (here, playing video games), and map that state to an eye position. Factors influencing state come from scene gi st, physical actions, events, and bottom-up saliency. Proposed models fall into two categories. In the first category, we use classical discriminative classifiers, including Reg ression, kNN and SVM. In the second category, we use Bayesian Networks to combine all the multi-modal factors in a unified framework. Our approaches significantly outperfor m 15 competing bottom-up and top-down attention models in predicting future eye fixat ions on 18,000 and 75,00 video frames and eye movement samples from a driving and a flig ht combat video game, respectively. We further test and validate our approaches on 1.4M video frames and 11M fixations samples and in all cases obtain higher prediction s cores that reference models. Ali Borji, Dicky N. Sihite, Laurent Itti |
BMVC | 3 |
| 2011 | Modeling forward and backward serial recall using a spatial registry assumption
Nader Noori, Laurent Itti |
CogSci | 2 |
| 2011 | Spatial Registry Model : Towards a Grounded Account for Executive Attention
Nader Noori, Laurent Itti |
CogSci | 2 |
| 2011 | Scene classification with a sparse set of salient regionsabstractThis work proposes an approach for scene classification by extracting and matching visual features only at the focuses of visual attention instead of the entire scene. Analysis over a database of natural scenes demonstrates that regions proposed by the saliency-based model of visual attention are robust to image transformations. Using a nearest neighbor classifier and a distance measure defined over the salient regions, we obtained 97.35% and 78.28% classification rates with SIFT and C2 features from the HMAX model at 5 salient regions covering at most 31% of the image. Classification with features extracted from the entire image results in 99.3% and 82.32% using SIFT and C2 features, respectively. Comparing attentional and adhoc approaches shows that classification rate of the first approach is 0.95 of the second. Overall, our results prove that efficient scene classification, in terms of reducing the complexity of feature extraction is possible without a significant drop in performance. Ali Borji, Laurent Itti |
ICRA | 2 |
| 2011 | Multilayer real-time video image stabilizationabstractIn many camera-based robotics applications, stabilizing video images in real-time is often critical for successful performance. In particular vision-based navigation, localization and tracking tasks cannot be performed reliably when landmarks are blurry, poorly focused or disappear from the camera view due to strong vibrations. Thus a reliable video image stabilization system would be invaluable for these applications. This paper presents a real-time video image stabilization system (VISS) primarily developed for aerial robots. Its unique architecture combines four independent stabilization layers. Layer 1 detects vibrations via an inertial measurement unit (IMU) and performs external counter-movements with a motorized gimbal. Layer 2 damps vibrations by using mechanical devices. The internal optical image stabilization of the camera represents Layer 3, while Layer 4 filters remaining vibrations using software. VISS is low-cost and robust. It has been implemented on a ¿Photoship One¿ gimbal, using GUMBOT hardware for processing Sparkfun-IMU data (Layer 1). Lord Mount vibration isolators damp vibrations (Layer 2). Video images of Panasonic's Lumix DMCTZ5 camera are optically stabilized with Panasonic's ¿Mega O.I.S.¿ technique (Layer 3) and digitally stabilized with ¿Deshaker¿ software (Layer 4). VISS significantly improved the stability of shaky video images in a series of experiments. Jens Windau, Laurent Itti |
IROS | 2 |
| 2011 | Visual attention guided bit allocation in video compression
Zhicheng Li 0001, Shiyin Qin, Laurent Itti |
Image Vis. Comput. | 3 |
| 2011 | Saliency and Gist Features for Target Detection in Satellite ImagesabstractReliably detecting objects in broad-area overhead or satellite images has become an increasingly pressing need, as the capabilities for image acquisition are growing rapidly. The problem is particularly difficult in the presence of large intraclass variability, e.g., finding "boats" or "buildings," where model-based approaches tend to fail because no good model or template can be defined for the highly variable targets. This paper explores an automatic approach to detect and classify targets in high-resolution broad-area satellite images, which relies on detecting statistical signatures of targets, in terms of a set of biologically-inspired low-level visual features. Broad-area images are cut into small image chips, analyzed in two complementary ways: "attention/saliency" analysis exploits local features and their interactions across space, while "gist" analysis focuses on global nonspatial features and their statistics. Both feature sets are used to classify each chip as containing target(s) or not, using a support vector machine. Four experiments were performed to find "boats" (Experiments 1 and 2), "buildings" (Experiment 3) and "airplanes" (Experiment 4). In experiment 1, 14 416 image chips were randomly divided into training (300 boat, 300 nonboat) and test sets (13 816), and classification was performed on the test set (ROC area: 0.977 ± 0.003). In experiment 2, classification was performed on another test set of 11 385 chips from another broad-area image, keeping the same training set as in experiment 1 (ROC area: 0.952 ± 0.006). In experiment 3, 600 training chips (300 for each type) were randomly selected from 108 885 chips, and classification was conducted (ROC area: 0.922 ± 0.005). In experiment 4, 20 training chips (10 for each type) were randomly selected to classify the remaining 2581 chips (ROC area: 0.976 ± 0.003). The proposed algorithm outperformed the state-of-the-art SIFT, HMAX, and hidden-scale salient structure methods, and previous gist-only features in all four experiments. This study shows that the proposed target search method can reliably and effectively detect highly variable target objects in large image datasets. Zhicheng Li 0001, Laurent Itti |
IEEE Trans. Image Process. | 2 |
| 2010 | Mobile robot vision navigation & localization using Gist and SaliencyabstractWe present a vision-based navigation and localization system using two biologically-inspired scene understanding models which are studied from human visual capabilities: (1) Gist model which captures the holistic characteristics and layout of an image and (2) Saliency model which emulates the visual attention of primates to identify conspicuous regions in the image. Here the localization system utilizes the gist features and salient regions to accurately localize the robot, while the navigation system uses the salient regions to perform visual feedback control to direct its heading and go to a user-provided goal location. We tested the system on our robot, Beobot2.0, in an indoor and outdoor environment with a route length of 36.67m (10,890 video frames) and 138.27m (28,971 frames), respectively. On average, the robot is able to drive within 3.68cm and 8.78cm (respectively) of the center of the lane. Chin-Kai Chang, Christian Siagian, Laurent Itti |
IROS | 3 |
| 2010 | Of bits and wows: A Bayesian theory of surprise with applications to attention
Pierre Baldi, Laurent Itti |
Neural Networks | 2 |
| 2009 | Centralized server environment for educational roboticsabstractOne of the main challenges when creating an undergraduate introduction to robotics course is connecting the theory taught in the lectures with the current practices of research. The primary cause of this difficulty is an inability to find a hardware solution that is powerful enough to run complex cutting-edge algorithms yet inexpensive enough to be purchased by an undergraduate class budget. An ideal system needs to have a gentle learning curve to allow students with minimal background in the field to get a robot up and running. Lastly, a fleet of classroom robots needs to be easy to administrate and maintain given the limited time of a Teaching Assistant. Our approach is to implement a centralized server system. In this system individual robots are inexpensive yet capable of establishing a WiFi link to a main server so that all the compilation and system administration, as well as much of the computationally intensive processing, are done on that server. We find that this solution saves both time and money and provides an effective teaching tool. This paper describes the hardware and software architecture of the system, and example applications implemented by undergraduate students. Randolph Voorhies, Christian Siagian, Lior Elazary, Laurent Itti |
IROS | 4 |
| 2009 | Biologically Inspired Mobile Robot Vision LocalizationabstractWe present a robot localization system using biologically inspired vision. Our system models two extensively studied human visual capabilities: (1) extracting the ldquogistrdquo of a scene to produce a coarse localization hypothesis and (2) refining it by locating salient landmark points in the scene. Gist is computed here as a holistic statistical signature of the image, thereby yielding abstract scene classification and layout. Saliency is computed as a measure of interest at every image location, which efficiently directs the time-consuming landmark-identification process toward the most likely candidate locations in the image. The gist features and salient regions are then further processed using a Monte Carlo localization algorithm to allow the robot to generate its position. We test the system in three different outdoor environments-building complex (38.4 m times 54.86 m area, 13 966 testing images), vegetation-filled park (82.3 m times 109.73 m area, 26 397 testing images), and open-field park (137.16 m times 178.31 m area, 34 711 testing images)-each with its own challenges. The system is able to localize, on average, within 0.98, 2.63, and 3.46 m, respectively, even with multiple kidnapped-robot instances. Christian Siagian, Laurent Itti |
IEEE Trans. Robotics | 2 |
| 2008 | Storing and recalling information for vision localizationabstractIn implementing a vision localization system, a crucial issue to consider is how to efficiently store and recall the necessary information so that the robot is not only able to accurately localize itself, but does so in a timely manner. In the presented system, we discuss a strategy to minimize the amount of stored data by analyzing the strengths and weaknesses of several cooperating recognition modules, and by using them through a prioritization scheme, which orders the data entries from the most likely to match to the least. We validate the system is a series of experiments at three large scale outdoor environments: a building complex (126 times 180 ft. area, 3583 testing images), a vegetation-filled park (270 times 360 ft. area, 6006 testing images), and an open-field area (450 times 585 ft. area, 8823 testing images) - each with its own set of challenges. Not only is the system able to localize in these environments (on average 3.46 ft., 6.55 ft. 12.96 ft. of error, respectively), it does so while searching through only 7.35%, 3.50%, and 6.12% of all the stored information, respectively. Christian Siagian, Laurent Itti |
ICRA | 2 |
| 2008 | Applying computational tools to predict gaze direction in interactive visual environmentsabstractFuture interactive virtual environments will be “attention-aware,” capable of predicting, reacting to, and ultimately influencing the visual attention of their human operators. Before such environments can be realized, it is necessary to operationalize our understanding of the relevant aspects of visual perception, in the form of fully automated computational heuristics that can efficiently identify locations that would attract human gaze in complex dynamic environments. One promising approach to designing such heuristics draws on ideas from computational neuroscience. We compared several neurobiologically inspired heuristics with eye-movement recordings from five observers playing video games, and found that human gaze was better predicted by heuristics that detect outliers from the global distribution of visual features than by purely local heuristics. Heuristics sensitive to dynamic events performed best overall. Further, heuristic prediction power differed more between games than between different human observers. While other factors clearly also influence eye position, our findings suggest that simple neurally inspired algorithmic methods can account for a significant portion of human gaze behavior in a naturalistic, interactive setting. These algorithms may be useful in the implementation of interactive virtual environments, both to predict the cognitive state of human operators, as well as to effectively endow virtual agents in the system with humanlike visual behavior. Robert J. Peters, Laurent Itti |
ACM Trans. Appl. Percept. | 2 |
| 2007 | Beyond bottom-up: Incorporating task-dependent influences into a computational model of spatial attentionabstractA critical function in both machine vision and biological vision systems is attentional selection of scene regions worthy of further analysis by higher-level processes such as object recognition. Here we present the first model of spatial attention that (1) can be applied to arbitrary static and dynamic image sequences with interactive tasks and (2) combines a general computational implementation of both bottom-up (BU) saliency and dynamic top-down (TD) task relevance; the claimed novelty lies in the combination of these elements and in the fully computational nature of the model. The BU component computes a saliency map from 12 low-level multi-scale visual features. The TD component computes a low-level signature of the entire image, and learns to associate different classes of signatures with the different gaze patterns recorded from human subjects performing a task of interest. We measured the ability of this model to predict the eye movements of people playing contemporary video games. We found that the TD model alone predicts where humans look about twice as well as does the BU model alone; in addition, a combined BU*TD model performs significantly better than either individual component. Qualitatively, the combined model predicts some easy-to-describe but hard-to-compute aspects of attentional selection, such as shifting attention leftward when approaching a left turn along a racing track. Thus, our study demonstrates the advantages of integrating BU factors derived from a saliency map and TD factors learned from image and task contexts in predicting where humans look while performing complex visually-guided behavior. Robert J. Peters, Laurent Itti |
CVPR | 2 |
| 2007 | Biologically-inspired robotics vision monte-carlo localization in the outdoor environmentabstractWe present a robot localization system using biologically-inspired vision. Our system models two extensively studied human visual capabilities: (1) extracting the "gist" of a scene to produce a coarse localization hypothesis, and (2) refining it by locating salient landmark regions in the scene. Gist is computed here as a holistic statistical signature of the image, yielding abstract scene classification and layout. Saliency is computed as a measure of interest at every image location, efficiently directing the time-consuming landmark identification process towards the most likely candidate locations in the image. The gist and salient landmark features are then further processed using a Monte-Carlo localization algorithm to allow the robot to generate its position. We test the system in three different outdoor environments - building complex (126times180 ft. area, 3794 testing images), vegetation-filled park (270times360 ft. area, 7196 testing images), and open-field park (450times585 ft. area, 8287 testing images) - each with its own challenges. The system is able to localize, on average, within 6.0, 10.73, and 32.24 ft., respectively, even with multiple kidnapped-robot instances. Christian Siagian, Laurent Itti |
IROS | 2 |
| 2007 | Congruence between model and human attention reveals unique signatures of critical visual eventsabstractCurrent computational models of bottom-up and top-down components of atten- tion are predictive of eye movements across a range of stimuli and of simple, fixed visual tasks (such as visual search for a target among distractors). How- ever, to date there exists no computational framework which can reliably mimic human gaze behavior in more complex environments and tasks, such as driving a vehicle through traffic. Here, we develop a hybrid computational/behavioral framework, combining simple models for bottom-up salience and top-down rel- evance, and looking for changes in the predictive power of these components at different critical event times during 4.7 hours (500,000 video frames) of observers playing car racing and flight combat video games. This approach is motivated by our observation that the predictive strengths of the salience and relevance mod- els exhibit reliable temporal signatures during critical event windows in the task sequence—for example, when the game player directly engages an enemy plane in a flight combat game, the predictive strength of the salience model increases significantly, while that of the relevance model decreases significantly. Our new framework combines these temporal signatures to implement several event detec- tors. Critically, we find that an event detector based on fused behavioral and stim- ulus information (in the form of the model’s predictive strength) is much stronger than detectors based on behavioral information alone (eye position) or image in- formation alone (model prediction maps). This approach to event detection, based on eye tracking combined with computational models applied to the visual input, may have useful applications as a less-invasive alternative to other event detection approaches based on neural signatures derived from EEG or fMRI recordings. Robert J. Peters, Laurent Itti |
NIPS | 2 |
| 2007 | Rapid Biologically-Inspired Scene Classification Using Features Shared with Visual AttentionabstractWe describe and validate a simple context-based scene recognition algorithm for mobile robotics applications. The system can differentiate outdoor scenes from various sites on a college campus using a multiscale set of early-visual features, which capture the "gist" of the scene into a low-dimensional signature vector. Distinct from previous approaches, the algorithm presents the advantage of being biologically plausible and of having low-computational complexity, sharing its low-level features with a model for visual attention that may operate concurrently on a robot. We compare classification accuracy using scenes filmed at three outdoor sites on campus (13,965 to 34,711 frames per site). Dividing each site into nine segments, we obtain segment classification rates between 84.21 percent and 88.62 percent. Combining scenes from all sites (75,073 frames in total) yields 86.45 percent correct classification, demonstrating the generalization and scalability of the approach. Christian Siagian, Laurent Itti |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2006 | An Integrated Model of Top-Down and Bottom-Up Attention for Optimizing Detection SpeedabstractIntegration of goal-driven, top-down attention and image-driven, bottom-up attention is crucial for visual search. Yet, previous research has mostly focused on models that are purely top-down or bottom-up. Here, we propose a new model that combines both. The bottom-up component computes the visual salience of scene locations in different feature maps extracted at multiple spatial scales. The topdown component uses accumulated statistical knowledge of the visual features of the desired search target and background clutter, to optimally tune the bottom-up maps such that target detection speed is maximized. Testing on 750 artificial and natural scenes shows that the model’s predictions are consistent with a large body of available literature on human psychophysics of visual search. These results suggest that our model may provide good approximation of how humans combine bottom-up and top-down cues such as to optimize target detection speed. Vidhya Navalpakkam, Laurent Itti |
CVPR (2) | 2 |
| 2006 | Causal saliency effects during natural visionabstractSalient stimuli, such as color or motion contrasts, attract human attention, thus providing a fast heuristic for focusing limited neural resources on behaviorally relevant sensory inputs. Here we address the following questions: What types of saliency attract attention and how do they compare to each other during natural vision? We asked human participants to inspect scene-shuffled video clips, tracked their instantaneous eye-position, and quantified how well a battery of computational saliency models predicted overt attentional selections (saccades). Saliency effects were measured as a function of total viewing time, proximity to abrupt scene transitions (jump cuts), and inter-participant consistency. All saliency models predicted overall attentional selection well above chance, with dynamic models being equally predictive to each other, and up to 3.6 times more predictive than static models. The prediction accuracy of all dynamic models was twice higher than their average for saccades that were initiated immediately after jump cuts, and led to maximal inter-participant consistency. Static models showed mixed results in these circumstances, with some models having weaker prediction accuracy than their average. These results demonstrate that dynamic visual cues play a dominant causal role in attracting attention, while static visual cues correlate with attentional selection mostly due to top-down causes. Ran Carmi, Laurent Itti |
ETRA | 2 |
| 2006 | Computational mechanisms for gaze direction in interactive visual environmentsabstractNext-generation immersive virtual environments and video games will require virtual agents with human-like visual attention and gaze behaviors. A critical step is to devise efficient visual processing heuristics to select locations that would attract human gaze in complex dynamic environments. One promising approach to designing such heuristics draws on ideas from computational neuroscience. We compared several such heuristics with eye movement recordings from five observers playing video games, and found that heuristics which detect outliers from the global distribution of visual features were better predictors of human gaze than were purely local heuristics. Heuristics sensitive to dynamic events performed best overall. Further, heuristic prediction power differed more between games than between different human observers. Our findings suggest simple neurally-inspired algorithmic methods to predict where humans look while playing video games. Robert J. Peters, Laurent Itti |
ETRA | 2 |
| 2006 | Photorealistic Attention-Based Gaze AnimationabstractWe apply a neurobiological model of visual attention and gaze control to the automatic animation of a photorealistic virtual human head. The attention model simulates biological visual processing along the occipito-parietal pathway of the primate brain. The gaze control model is derived from motion capture of human subjects, using high-speed video-based eye and head tracking apparatus. Given an arbitrary video clip, the model predicts visual locations most likely to attract an observer's attention, and simulates the dynamics of eye and head movements towards these locations. Tested on 85 video clips including synthetic stimuli, video games, TV news, sports, and outdoor scenes, the model demonstrates a strong ability at saccading towards and tracking salient targets. The resulting autonomous virtual human animation is of photorealistic quality Laurent Itti, Nitin Dhavale, Frédéric H. Pighin |
ICME | 1 |
| 2006 | The use of attention and spatial information for rapid facial recognition in video
James Bonaiuto, Laurent Itti |
Image Vis. Comput. | 2 |
| 2005 | A Principled Approach to Detecting Surprising Events in VideoabstractPrimates demonstrate unparalleled ability at rapidly orienting towards important events in complex dynamic environments. During rapid guidance of attention and gaze towards potential objects of interest or threats, often there is no time for detailed visual analysis. Thus, heuristic computations are necessary to locate the most interesting events in quasi real-time. We present a new theory of sensory surprise, which provides a principled and computable shortcut to important information. We develop a model that computes instantaneous low-level surprise at every location in video streams. The algorithm significantly correlates with eye movements of two humans watching complex video clips, including television programs (17,936 frames, 2,152 saccadic gaze shifts). The system allows more sophisticated and time-consuming image analysis to be efficiently focused onto the most surprising subsets of the incoming data. Laurent Itti, Pierre Baldi |
CVPR (1) | 1 |
| 2005 | Bayesian Surprise Attracts Human AttentionabstractThe concept of surprise is central to sensory processing, adaptation, learning, and attention. Yet, no widely-accepted mathematical theory currently exists to quantitatively characterize surprise elicited by a stimulus or event, for observers that range from single neurons to complex natural or engineered systems. We describe a formal Bayesian definition of surprise that is the only consistent formulation under minimal axiomatic assumptions. Surprise quantifies how data affects a natural or artificial observer, by measuring the difference between posterior and prior beliefs of the observer. Using this framework we measure the extent to which humans direct their gaze towards surprising items while watching television and video games. We find that subjects are strongly attracted towards surprising locations, with 72% of all human gaze shifts directed towards locations more surprising than the average, a figure which rises to 84% when considering only gaze targets simultaneously selected by all subjects. The resulting theory of surprise is applicable across different spatio-temporal scales, modalities, and levels of abstraction. Life is full of surprises, ranging from a great christmas gift or a new magic trick, to wardrobe malfunctions, reckless drivers, terrorist attacks, and tsunami waves. Key to survival is our ability to rapidly attend to, identify, and learn from surprising events, to decide on present and future courses of action [1]. Yet, little theoretical and computational understanding exists of the very essence of surprise, as evidenced by the absence from our everyday vocabulary of a quantitative unit of surprise: Qualities such as the "wow factor" have remained vague and elusive to mathematical analysis. Informal correlates of surprise exist at nearly all stages of neural processing. In sensory neuroscience, it has been suggested that only the unexpected at one stage is transmitted to the next stage [2]. Hence, sensory cortex may have evolved to adapt to, to predict, and to quiet down the expected statistical regularities of the world [3, 4, 5, 6], focusing instead on events that are unpredictable or surprising. Electrophysiological evidence for this early sensory emphasis onto surprising stimuli exists from studies of adaptation in visual [7, 8, 4, 9], olfactory [10, 11], and auditory cortices [12], subcortical structures like the LGN [13], and even retinal ganglion cells [14, 15] and cochlear hair cells [16]: neural response greatly attenuates with repeated or prolonged exposure to an initially novel stimulus. Surprise and novelty are also central to learning and memory formation [1], to the point that surprise is believed to be a necessary trigger for associative learning [17, 18], as supported by mounting evidence for a role of the hippocampus as a novelty detector [19, 20, 21]. Finally, seeking novelty is a well-identified human character trait, with possible association with the dopamine D4 receptor gene [22, 23, 24]. In the Bayesian framework, we develop the only consistent theory of surprise, in terms of the difference between the posterior and prior distributions of beliefs of an observer over the available class of models or hypotheses about the world. We show that this definition derived from first principles presents key advantages over more ad-hoc formulations, typically relying on detecting outlier stimuli. Armed with this new framework, we provide direct experimental evidence that surprise best characterizes what attracts human gaze in large amounts of natural video stimuli. We here extend a recent pilot study [25], adding more comprehensive theory, large-scale human data collection, and additional analysis. Laurent Itti, Pierre Baldi |
NIPS | 1 |
| 2005 | Optimal cue selection strategyabstractSurvival in the natural world demands the selection of relevant visual cues to rapidly and reliably guide attention towards prey an d predators in cluttered environments. We investigate whether our visu al system selects cues that guide search in an optimal manner. We formall y obtain the optimal cue selection strategy by maximizing the signal to noise ratio (S N R) between a search target and surrounding distractors. This optimal strategy successfully accounts for several phenom ena in visual search behavior, including the effect of target-distracto r discriminability, uncertainty in target's features, distractor heterogenei ty, and linear separability. Furthermore, the theory generates a new predict ion, which we verify through psychophysical experiments with human subj ects. Our results provide direct experimental evidence that humans sel ect visual cues so as to maximize S N R between the targets and surrounding clutter. Vidhya Navalpakkam, Laurent Itti |
NIPS | 2 |
| 2005 | Robot steering with spectral image informationabstractWe introduce a method for rapidly classifying visual scenes globally along a small number of navigationally relevant dimensions: depth of scene, presence of obstacles, path versus nonpath, and orientation of path. We show that the algorithm reliably classifies scenes in terms of these high-level features, based on global or coarsely localized spectral analysis analogous to early-stage biological vision. We use this analysis to implement a real-time visual navigational system on a mobile robot, trained online by a human operator. We demonstrate successful training and subsequent autonomous path following for two different outdoor environments, a running track and a concrete trail. Our success with this technique suggests a general applicability to autonomous robot navigation in a variety of environments. Christopher Ackerman, Laurent Itti |
IEEE Trans. Robotics | 2 |
| 2004 | Automatic foveation for video compression using a neurobiological model of visual attentionabstractWe evaluate the applicability of a biologically-motivated algorithm to select visually-salient regions of interest in video streams for multiply-foveated video compression. Regions are selected based on a nonlinear integration of low-level visual cues, mimicking processing in primate occipital, and posterior parietal cortex. A dynamic foveation filter then blurs every frame, increasingly with distance from salient locations. Sixty-three variants of the algorithm (varying number and shape of virtual foveas, maximum blur, and saliency competition) are evaluated against an outdoor video scene, using MPEG-1 and constant-quality MPEG-4 (DivX) encoding. Additional compression radios of 1.1 to 8.5 are achieved by foveation. Two variants of the algorithm are validated against eye fixations recorded from four to six human observers on a heterogeneous collection of 50 video clips (over 45 000 frames in total). Significantly higher overlap than expected by chance is found between human and algorithmic foveations. With both variants, foveated clips are, on average, approximately half the size of unfoveated clips, for both MPEG-1 and MPEG-4. These results suggest a general-purpose usefulness of the algorithm in improving compression ratios of unconstrained video. Laurent Itti |
IEEE Trans. Image Process. | 1 |
| 2003 | CINNIC, a new computational algorithm for the modeling of early visual contour integration in humans
T. Nathan Mundhenk, Laurent Itti |
Neurocomputing | 2 |
| 2001 | Modeling the Modulatory Effect of Attention on Human Spatial VisionabstractWe present new simulation results, in which a computational model of interacting visual neurons simultaneously predicts the modula(cid:173) tion of spatial vision thresholds by focal visual attention, for five dual-task human psychophysics experiments. This new study com(cid:173) plements our previous findings that attention activates a winner(cid:173) take-all competition among early visual neurons within one cortical hypercolumn. This "intensified competition" hypothesis assumed that attention equally affects all neurons, and yielded two single(cid:173) unit predictions: an increase in gain and a sharpening of tuning with attention. While both effects have been separately observed in electrophysiology, no single-unit study has yet shown them si(cid:173) multaneously. Hence, we here explore whether our model could still predict our data if attention might only modulate neuronal gain, but do so non-uniformly across neurons and tasks. Specifically, we investigate whether modulating the gain of only the neurons that are loudest, best-tuned, or most informative about the stimulus, or of all neurons equally but in a task-dependent manner, may ac(cid:173) count for the data. We find that none of these hypotheses yields predictions as plausible as the intensified competition hypothesis, hence providing additional support for our original findings. Laurent Itti, Jochen Braun, Christof Koch |
NIPS | 1 |
| 1999 | A quantitative model relating visual neuronal activity to psychophysical thresholds
Laurent Itti, Christof Koch, Jochen Braun |
Neurocomputing | 1 |
| 1998 | Attentional Modulation of Human Pattern Discrimination Psychophysics Reproduced by a Quantitative Model
Laurent Itti, Jochen Braun, Dale K. Lee, Christof Koch |
NIPS | 1 |
| 1998 | A Model of Saliency-Based Visual Attention for Rapid Scene AnalysisabstractA visual attention system, inspired by the behavior and the neuronal architecture of the early primate visual system, is presented. Multiscale image features are combined into a single topographical saliency map. A dynamical neural network then selects attended locations in order of decreasing saliency. The system breaks down the complex problem of scene understanding by rapidly selecting, in a computationally efficient manner, conspicuous locations to be analyzed in detail. Laurent Itti, Christof Koch, Ernst Niebur |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1997 | A Model of Early Visual Processing
Laurent Itti, Jochen Braun, Dale K. Lee, Christof Koch |
NIPS | 1 |