Camilo Fosco

dblp:256/5485 · also Camilo Luciano Fosco · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
5since 2021 · last 2026
0009-0001-9491-8175ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Video understanding and tracking · 43% Efficient and distributed learning · 25% Representation and self-supervised learning · 17%
Computer graphics and multimedia
4 papers
Multimedia analysis and retrieval · 68% Image and video processing · 16% Multimedia systems and quality of experience · 11%
Human-computer interaction and pervasive computing
2 papers
Interaction techniques and input · 50% User interface design and tools · 50%

Topics — the 14 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › brain decoding
video reconstruction from brain signals
0.812024
Brain Netflix: Scaling Data to Reconstruct Videos from Brain Signals · ECCV (26) 2024
Computer vision › Video understanding and tracking
action anticipation
0.712023
Leveraging Temporal Context in Low Representational Power Regimes · CVPR 2023
Computer vision › Video understanding and tracking
action recognition
0.712023
Leveraging Temporal Context in Low Representational Power Regimes · CVPR 2023
Machine learning › Representation and self-supervised learning › representation learning
modular representation learning
0.712023
Modular Memorability: Tiered Representations for Video Memorability Prediction · CVPR 2023
Machine learning › Efficient and distributed learning
adaptive computation
0.512021
VA-RED2: Video Adaptive Redundancy Reduction · ICLR 2021
Machine learning › Efficient and distributed learning
model compression
0.512021
VA-RED2: Video Adaptive Redundancy Reduction · ICLR 2021
Machine learning › Representation and self-supervised learning
redundancy reduction
0.512021
VA-RED2: Video Adaptive Redundancy Reduction · ICLR 2021
Computer vision › Video understanding and tracking
video classification
0.512021
VA-RED2: Video Adaptive Redundancy Reduction · ICLR 2021
Image and video processing
saliency detection
0.412020
How Much Time Do You Have? Modeling Multi-Duration Saliency · CVPR 2020
Multimedia analysis and retrieval
video understanding
0.412020
Multimodal Memorability: Modeling Effects of Semantics and Decay on Video Memorability · ECCV (16) 2020
User interface design and tools
design tools
0.412020
Predicting Visual Importance Across Graphic Design Types · UIST 2020
Medical and health informatics
brain-computer interface
0.212024
Brain Netflix: Scaling Data to Reconstruct Videos from Brain Signals · ECCV (26) 2024
Computer vision › Image recognition and object detection
saliency prediction
0.112020
Predicting Visual Importance Across Graphic Design Types · UIST 2020
Visualization and visual analytics › information visualization
attention visualization
0.112020
TurkEyes: A Web-Based Toolbox for Crowdsourcing Attention Data · CHI 2020

Methods — techniques the papers use, named apart from their topics

deep learning · 2.4brain signal decoding · 1.5literature review · 1.0benchmark analysis · 1.0saliency modeling · 0.9eye tracking · 0.9crowdsourcing · 0.9temporal context modeling · 0.7neural network · 0.7event transition matrix · 0.7ablation study · 0.7reinforcement learning · 0.5adaptive inference · 0.5multimodal modeling · 0.4LSTM · 0.4
YearPublicationVenuePosition
2026 A Review of Computational Memorability: A Benchmark Framework
abstract
Abstract One of the powers of visual media lies in its ability to create a lasting impression on the viewer’s memory. In this digital age, where media is abundant and attention spans are fleeting, the task of predicting which content will stick in the viewer’s mind has become a critical challenge in computer vision. Computational memorability seeks to address this by developing models that estimate how memorable a piece of media is likely to be. In this review we focus on the MediaEval Predicting Video Memorability benchmark, a recurring evaluation task that has run annually since 2018. This benchmark provides a unique and consistent framework for researchers to compare and refine their memorability prediction techniques using standardised datasets and metrics. Its reproducible framework has proven invaluable for tracking progress and fostering innovation in this rapidly evolving domain. We analyse the evolution of the benchmark across its 2018–2023 editions, discussing the challenges that still remain, such as the need for more interpretability in models and the difficulty of predicting subjective and context-dependent memorability. By analysing and synthesising the collective insights gained from this task, we endeavour to inspire new avenues of inquiry and drive progress towards a more comprehensive understanding of this topic.
Mihai Gabriel Constantin, Claire-Hélène Demarty, Camilo Fosco, Sebastian Halder 0001, Graham Healy, Bogdan Ionescu, Stefan Valentin Luncanu, Iván Martín-Fernández, Ana Matran-Fernandez, Rukiye Savran Kiziltepe, Alan F. Smeaton, Liviu-Daniel Stefan, Lorin Sweeney, Alba Garcia Seco de Herrera
Int. J. Comput. Vis.3
2024 Brain Netflix: Scaling Data to Reconstruct Videos from Brain Signals
Camilo Fosco, Benjamin Lahner, Bowen Pan, Alex Andonian, Emilie Josephs, Alex Lascelles, Aude Oliva
ECCV (26)1
2023 Modular Memorability: Tiered Representations for Video Memorability Prediction
abstract
The question of how to best estimate the memorability of visual content is currently a source of debate in the memorability community. In this paper, we propose to explore how different key properties of images and videos affect their consolidation into memory. We analyze the impact of several features and develop a model that emulates the most important parts of a proposed “pathway to memory”: a simple but effective way of representing the different hurdles that new visual content needs to surpass to stay in memory. This framework leads to the construction of our M3-S model, a novel memorability network that processes input videos in a modular fashion. Each module of the network emulates one of the four key steps of the pathway to memory: raw encoding, scene understanding, event understanding and memory consolidation. We find that the different representations learned by our modules are non-trivial and substantially different from each other. Additionally, we observe that certain representations tend to perform better at the task of memorability prediction than others, and we introduce an in-depth ablation study to support our results. Our proposed approach surpasses the state of the art on the two largest video memorability datasets and opens the door to new applications in the field. Our code is available at https://github.com/tekal-ai/modular-memorability.
Théo Dumont, Juan Segundo Hevia, Camilo Fosco
CVPR3
2023 Leveraging Temporal Context in Low Representational Power Regimes
abstract
Computer vision models are excellent at identifying and exploiting regularities in the world. However, it is computationally costly to learn these regularities from scratch. This presents a challenge for low-parameter models, like those running on edge devices (e.g. smartphones). Can the performance of models with low representational power be improved by supplementing training with additional information about these statistical regularities? We explore this in the domains of action recognition and action anticipation, leveraging the fact that actions are typically embedded in stereotypical sequences. We introduce the Event Transition Matrix (ETM), computed from action labels in an untrimmed video dataset, which captures the temporal context of a given action, operationalized as the likelihood that it was preceded or followed by each other action in the set. We show that including information from the ETM during training improves action recognition and anticipation performance on various egocentric video datasets. Through ablation and control studies, we show that the coherent sequence of information captured by our ETM is key to this effect, and we find that the benefit of this explicit representation of temporal context is most pronounced for smaller models. Code, matrices and models are available in our project page: https://camilofosco.com/etm_website.
Camilo Fosco, SouYoung Jin, Emilie Josephs, Aude Oliva
CVPR1
2021 VA-RED2: Video Adaptive Redundancy Reduction
Bowen Pan, Rameswar Panda, Camilo Fosco, Chung-Ching Lin, Alex Andonian, Kate Saenko, Aude Oliva, Rogério Feris
ICLR3
2020 TurkEyes: A Web-Based Toolbox for Crowdsourcing Attention Data
abstract
Eye movements provide insight into what parts of an image a viewer finds most salient, interesting, or relevant to the task at hand. Unfortunately, eye tracking data, a commonly-used proxy for attention, is cumbersome to collect. Here we explore an alternative: a comprehensive web-based toolbox for crowdsourcing visual attention. We draw from four main classes of attention-capturing methodologies in the literature. ZoomMaps is a novel zoom-based interface that captures viewing on a mobile phone. CodeCharts is a self-reporting methodology that records points of interest at precise viewing durations. ImportAnnots is an "annotation" tool for selecting important image regions, and cursor-based BubbleView lets viewers click to deblur a small area. We compare these methodologies using a common analysis framework in order to develop appropriate use cases for each interface. This toolbox and our analyses provide a blueprint for how to gather attention data at scale without an eye tracker.
Anelise Newman, Barry A. McNamara, Camilo Fosco, Yun Bin Zhang, Pat Sukhum, Matthew Tancik, Zoya Bylinskii
CHI3
2020 How Much Time Do You Have? Modeling Multi-Duration Saliency
abstract
What jumps out in a single glance of an image is different than what you might notice after closer inspection. Yet conventional models of visual saliency produce predictions at an arbitrary, fixed viewing duration, offering a limited view of the rich interactions between image content and gaze location. In this paper we propose to capture gaze as a series of snapshots, by generating population-level saliency heatmaps for multiple viewing durations. We collect the CodeCharts1K dataset, which contains multiple distinct heatmaps per image corresponding to 0.5, 3, and 5 seconds of free-viewing. We develop an LSTM-based model of saliency that simultaneously trains on data from multiple viewing durations. Our Multi-Duration Saliency Excited Model (MD-SEM) achieves competitive performance on the LSUN 2017 Challenge with 57% fewer parameters than comparable architectures. It is the first model that produces heatmaps at multiple viewing durations, enabling applications where multi-duration saliency can be used to prioritize visual content to keep, transmit, and render.
Camilo Fosco, Anelise Newman, Pat Sukhum, Yun Bin Zhang, Nanxuan Zhao, Aude Oliva, Zoya Bylinskii
CVPR1
2020 We Have So Much in Common: Modeling Semantic Relational Set Abstractions in Videos
Alex Andonian, Camilo Fosco, Mathew Monfort, Allen Lee, Rogério Feris, Carl Vondrick, Aude Oliva
ECCV (18)2
2020 Multimodal Memorability: Modeling Effects of Semantics and Decay on Video Memorability
Anelise Newman, Camilo Fosco, Vincent Casser, Allen Lee, Barry A. McNamara, Aude Oliva
ECCV (16)2
2020 Predicting Visual Importance Across Graphic Design Types
abstract
This paper introduces a Unified Model of Saliency and Importance (UMSI), which learns to predict visual importance in input graphic designs, and saliency in natural images, along with a new dataset and applications. Previous methods for predicting saliency or visual importance are trained individually on specialized datasets, making them limited in application and leading to poor generalization on novel image classes, while requiring a user to know which model to apply to which input. UMSI is a deep learning-based model simultaneously trained on images from different design classes, including posters, infographics, mobile UIs, as well as natural images, and includes an automatic classification module to classify the input. This allows the model to work more effectively without requiring a user to label the input. We also introduce Imp1k, a new dataset of designs annotated with importance information. We demonstrate two new design interfaces that use importance prediction, including a tool for adjusting the relative importance of design elements, and a tool for reflowing designs to new aspect ratios while preserving visual importance.
Camilo Fosco, Vincent Casser, Amish Kumar Bedi, Peter O'Donovan, Aaron Hertzmann, Zoya Bylinskii
UIST1