Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Susanna Ricco

dblp:66/812 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
2since 2021 · last 2023
0009-0005-1505-7055ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 4 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-authorSystems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Video understanding and tracking · 49% 3D vision · 20% Trustworthy machine learning · 19%
Human-computer interaction and pervasive computing
1 paper
Collaborative and social computing · 100%

Topics — the 19 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
fairness
0.712023
Consensus and Subjectivity of Skin Tone Annotation for ML Fairness · NeurIPS 2023
Computer vision › Video understanding and tracking
action recognition
0.312018
AVA: A Video Dataset of Spatio-Temporally Localized Atomic Visual Actions · CVPR 2018
Computer vision › Video understanding and tracking › action detection
spatio-temporal action localization
0.312018
AVA: A Video Dataset of Spatio-Temporally Localized Atomic Visual Actions · CVPR 2018
Computer vision › 3D vision
motion estimation
0.322013
Video Motion for Every Visible Point · ICCV 2013
Dense Lagrangian motion estimation with occlusions · CVPR 2012
Computer vision › Video understanding and tracking
motion segmentation
0.212016
Discovering the Physical Parts of an Articulated Object Class from Multiple Videos · CVPR 2016
Computer vision › Video understanding and tracking › motion analysis
motion pattern discovery
0.212015
Articulated motion discovery using pairs of trajectories · CVPR 2015
Collaborative and social computing
crowdsourcing
0.212023
Consensus and Subjectivity of Skin Tone Annotation for ML Fairness · NeurIPS 2023
Computer vision › 3D vision › motion estimation
optical flow
0.212013
Video Motion for Every Visible Point · ICCV 2013
Computer vision › Video understanding and tracking › feature tracking
dense point tracking
0.112012
Dense Lagrangian motion estimation with occlusions · CVPR 2012
Computer vision › Video understanding and tracking › object tracking
occlusion handling
0.112012
Dense Lagrangian motion estimation with occlusions · CVPR 2012
Computer vision › 3D vision
structure from motion
0.112012
Simultaneous Compaction and Factorization of Sparse Image Motion Matrices · ECCV (6) 2012
Robotics › Robot navigation and mapping
localization
0.112011
Textured occupancy grids for monocular localization without features · ICRA 2011
Robotics › Robot navigation and mapping › localization › vision-based localization
monocular localization
0.112011
Textured occupancy grids for monocular localization without features · ICRA 2011
Computer vision › Image recognition and object detection
spatial alignment
0.112017
Behavior Discovery and Alignment of Articulated Object Classes from Unstructured Video · Int. J. Comput. Vis. 2017
Computer vision › 3D vision › structure from motion
non-rigid structure from motion
0.112016
Discovering the Physical Parts of an Articulated Object Class from Multiple Videos · CVPR 2016
Machine learning › Learning paradigms
unsupervised learning
0.112015
Articulated motion discovery using pairs of trajectories · CVPR 2015
Algorithms and data structures › numerical linear algebra
matrix factorization
0.012012
Simultaneous Compaction and Factorization of Sparse Image Motion Matrices · ECCV (6) 2012
Robotics › Robot navigation and mapping › robot mapping
map representation
0.012011
Textured occupancy grids for monocular localization without features · ICRA 2011
Embedded and real-time systems
cyber-physical system platforms
0.012005
Scavenging with a Laptop Robot · AAAI 2005

Methods — techniques the papers use, named apart from their topics

crowdsourced labeling · 1.3annotation experiments · 1.3action localization · 0.3trajectory displacement · 0.3thin-plate spline · 0.3location model · 0.2energy minimization · 0.2trajectory pairs descriptor · 0.2clustering · 0.2path optimization · 0.2sparse matrix compaction · 0.1matrix factorization · 0.1robotics · 0.1
YearPublicationVenuePosition
2023 Consensus and Subjectivity of Skin Tone Annotation for ML Fairness
abstract
Understanding different human attributes and how they affect model behavior may become a standard need for all model creation and usage, from traditional computer vision tasks to the newest multimodal generative AI systems. In computer vision specifically, we have relied on datasets augmented with perceived attribute signals (eg, gender presentation, skin tone, and age) and benchmarks enabled by these datasets. Typically labels for these tasks come from human annotators. However, annotating attribute signals, especially skin tone, is a difficult and subjective task. Perceived skin tone is affected by technical factors, like lighting conditions, and social factors that shape an annotator's lived experience.This paper examines the subjectivity of skin tone annotation through a series of annotation experiments using the Monk Skin Tone (MST) scale~\cite{Monk2022Monk}, a small pool of professional photographers, and a much larger pool of trained crowdsourced annotators. Along with this study we release the Monk Skin Tone Examples (MST-E) dataset, containing 1515 images and 31 videos spread across the full MST scale. MST-E is designed to help train human annotators to annotate MST effectively.Our study shows that annotators can reliably annotate skin tone in a way that aligns with an expert in the MST scale, even under challenging environmental conditions. We also find evidence that annotators from different geographic regions rely on different mental models of MST categories resulting in annotations that systematically vary across regions. Given this, we advise practitioners to use a diverse set of annotators and a higher replication count for each image when annotating skin tone for fairness research.
Candice Schumann, Femi Olanubi, Auriel Wright, Ellis Monk Jr., Courtney Heldreth, Susanna Ricco
NeurIPS6
2021 A Step Toward More Inclusive People Annotations for Fairness
abstract
The Open Images Dataset contains approximately 9 million images and is a widely accepted dataset for computer vision research. As is common practice for large datasets, the annotations are not exhaustive, with bounding boxes and attribute labels for only a subset of the classes in each image. In this paper, we present a new set of annotations on a subset of the Open Images dataset called the MIAP (More Inclusive Annotations for People) subset, containing bounding boxes and attributes for all of the people visible in those images. The attributes and labeling methodology for the MIAP subset were designed to enable research into model fairness. In addition, we analyze the original annotation methodology for the person class and its subclasses, discussing the resulting patterns in order to inform future annotation efforts. By considering both the original and exhaustive annotation sets, researchers can also now study how systematic patterns in training annotations affect modeling.
Candice Schumann, Susanna Ricco, Utsav Prabhu, Vittorio Ferrari, Caroline Pantofaru
AIES2
2018 AVA: A Video Dataset of Spatio-Temporally Localized Atomic Visual Actions
abstract
This paper introduces a video dataset of spatio-temporally localized Atomic Visual Actions (AVA). The AVA dataset densely annotates 80 atomic visual actions in 437 15-minute video clips, where actions are localized in space and time, resulting in 1.59M action labels with multiple labels per person occurring frequently. The key characteristics of our dataset are: (1) the definition of atomic visual actions, rather than composite actions; (2) precise spatio-temporal annotations with possibly multiple annotations for each person; (3) exhaustive annotation of these atomic actions over 15-minute video clips; (4) people temporally linked across consecutive segments; and (5) using movies to gather a varied set of action representations. This departs from existing datasets for spatio-temporal action recognition, which typically provide sparse annotations for composite actions in short video clips. AVA, with its realistic scene and action complexity, exposes the intrinsic difficulty of action recognition. To benchmark this, we present a novel approach for action localization that builds upon the current state-of-the-art methods, and demonstrates better performance on JHMDB and UCF101-24 categories. While setting a new state of the art on existing datasets, the overall results on AVA are low at 15.8% mAP, underscoring the need for developing new approaches for video understanding.
Chunhui Gu, Chen Sun 0002, David A. Ross, Carl Vondrick, Caroline Pantofaru, Yeqing Li, Sudheendra Vijayanarasimhan, George Toderici, Susanna Ricco, Rahul Sukthankar, Cordelia Schmid, Jitendra Malik
CVPR9
2017 Behavior Discovery and Alignment of Articulated Object Classes from Unstructured Video
abstract
We propose an automatic system for organizing the content of a collection of unstructured videos of an articulated object class (e.g., tiger, horse). By exploiting the recurring motion patterns of the class across videos, our system: (1) identifies its characteristic behaviors, and (2) recovers pixel-to-pixel alignments across different instances. Our system can be useful for organizing video collections for indexing and retrieval. Moreover, it can be a platform for learning the appearance or behaviors of object classes from Internet video. Traditional supervised techniques cannot exploit this wealth of data directly, as they require a large amount of time-consuming manual annotations. The behavior discovery stage generates temporal video intervals, each automatically trimmed to one instance of the discovered behavior, clustered by type. It relies on our novel motion representation for articulated motion based on the displacement of ordered pairs of trajectories. The alignment stage aligns hundreds of instances of the class to a great accuracy despite considerable appearance variations (e.g., an adult tiger and a cub). It uses a flexible thin plate spline deformation model that can vary through time. We carefully evaluate each step of our system on a new, fully annotated dataset. On behavior discovery, we outperform the state-of-the-art improved dense trajectory feature descriptor. On spatial alignment, we outperform the popular SIFT Flow algorithm.
Luca Del Pero, Susanna Ricco, Rahul Sukthankar, Vittorio Ferrari
Int. J. Comput. Vis.2
2016 Discovering the Physical Parts of an Articulated Object Class from Multiple Videos
abstract
We propose a motion-based method to discover the physical parts of an articulated object class (e.g. head/torso/leg of a horse) from multiple videos. The key is to find object regions that exhibit consistent motion relative to the rest of the object, across multiple videos. We can then learn a location model for the parts and segment them accurately in the individual videos using an energy function that also enforces temporal and spatial consistency in part motion. Unlike our approach, traditional methods for motion segmentation or non-rigid structure from motion operate on one video at a time. Hence they cannot discover a part unless it displays independent motion in that particular video. We evaluate our method on a new dataset of 32 videos of tigers and horses, where we significantly outperform a recent motion segmentation method on the task of part discovery (obtaining roughly twice the accuracy).
Luca Del Pero, Susanna Ricco, Rahul Sukthankar, Vittorio Ferrari
CVPR2
2015 Articulated motion discovery using pairs of trajectories
abstract
We propose an unsupervised approach for discovering characteristic motion patterns in videos of highly articulated objects performing natural, unscripted behaviors, such as tigers in the wild. We discover consistent patterns in a bottom-up manner by analyzing the relative displacements of large numbers of ordered trajectory pairs through time, such that each trajectory is attached to a different moving part on the object. The pairs of trajectories descriptor relies entirely on motion and is more discriminative than state-of-the-art features that employ single trajectories. Our method generates temporal video intervals, each automatically trimmed to one instance of the discovered behavior, and clusters them by type (e.g., running, turning head, drinking water). We present experiments on two datasets: dogs from YouTube-Objects and a new dataset of National Geographic tiger videos. Results confirm that our proposed descriptor outperforms existing appearance- and trajectory-based descriptors (e.g., HOG and DTFs) on both datasets and enables us to segment unconstrained animal video into intervals containing single behaviors.
Luca Del Pero, Susanna Ricco, Rahul Sukthankar, Vittorio Ferrari
CVPR2
2013 Video Motion for Every Visible Point
abstract
Dense motion of image points over many video frames can provide important information about the world. However, occlusions and drift make it impossible to compute long motion paths by merely concatenating optical flow vectors between consecutive frames. Instead, we solve for entire paths directly, and flag the frames in which each is visible. As in previous work, we anchor each path to a unique pixel which guarantees an even spatial distribution of paths. Unlike earlier methods, we allow paths to be anchored in any frame. By explicitly requiring that at least one visible path passes within a small neighborhood of every pixel, we guarantee complete coverage of all visible points in all frames. We achieve state-of-the-art results on real sequences including both rigid and non-rigid motions with significant occlusions.
Susanna Ricco, Carlo Tomasi
ICCV1
2012 Dense Lagrangian motion estimation with occlusions
abstract
We couple occlusion modeling and multi-frame motion estimation to compute dense, temporally extended point trajectories in video with significant occlusions. Our approach combines robust spatial regularization with spatially and temporally global occlusion labeling in a variational, Lagrangian framework with subspace constraints. We track points even through ephemeral occlusions. Experiments demonstrate accuracy superior to the state of the art while tracking more points through more frames.
Susanna Ricco, Carlo Tomasi
CVPR1
2012 Simultaneous Compaction and Factorization of Sparse Image Motion Matrices
Susanna Ricco, Carlo Tomasi
ECCV (6)1
2011 Textured occupancy grids for monocular localization without features
abstract
A textured occupancy grid map is an extremely versatile data structure. It can be used to render human readable views and for laser rangefinder localization algorithms. For camera-based localization, landmark or feature based maps tend to be favored in current research. This may be because of a tacit assumption that working with a textured occupancy grid with a camera would be impractical. We demonstrate that a textured occupancy grid can be combined with an extremely simple monocular localization algorithm to produce a viable localization solution. Our approach is simple, efficient, and produces localization results comparable to laser localization results. A consequence of this result is that a single map representation, the textured occupancy grid, can now be used for humans, robots with laser rangefinders, and robots with just a single camera.
Julian Mason, Susanna Ricco, Ronald Parr
ICRA2
2009 Fingerspelling Recognition through Classification of Letter-to-Letter Transitions
Susanna Ricco, Carlo Tomasi
ACCV (3)1
2009 Correcting Motion Artifacts in Retinal Spectral Domain Optical Coherence Tomography via Image Registration
Susanna Ricco, Hiroshi Ishikawa 0005, Gadi Wollstein, Joel S. Schuman
MICCAI (1)1
2005 Scavenging with a Laptop Robot
Alan Davidson, Mac Mason, Susanna Ricco, Ben Tribelhorn, Zachary Dodds
AAAI3