EDBT 2026 Demo / reviewers in the wild / expert
Susanna Ricco
dblp:66/812
· DBLP profile ↗
13ranked-venue papers
5as first author
2since 2021 · last 2023
0009-0005-1505-7055ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 4 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-authorSystems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
Video understanding and tracking · 49% 3D vision · 20% Trustworthy machine learning · 19% | |
| Human-computer interaction and pervasive computing
1 paper |
Collaborative and social computing · 100% |
Topics — the 19 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
fairness |
0.7 | 1 | 2023 | Consensus and Subjectivity of Skin Tone Annotation for ML Fairness · NeurIPS 2023 |
Computer vision › Video understanding and tracking
action recognition |
0.3 | 1 | 2018 | AVA: A Video Dataset of Spatio-Temporally Localized Atomic Visual Actions · CVPR 2018 |
Computer vision › Video understanding and tracking › action detection
spatio-temporal action localization |
0.3 | 1 | 2018 | AVA: A Video Dataset of Spatio-Temporally Localized Atomic Visual Actions · CVPR 2018 |
Computer vision › 3D vision
motion estimation |
0.3 | 2 | 2013 | Video Motion for Every Visible Point · ICCV 2013 Dense Lagrangian motion estimation with occlusions · CVPR 2012 |
Computer vision › Video understanding and tracking
motion segmentation |
0.2 | 1 | 2016 | Discovering the Physical Parts of an Articulated Object Class from Multiple Videos · CVPR 2016 |
Computer vision › Video understanding and tracking › motion analysis
motion pattern discovery |
0.2 | 1 | 2015 | Articulated motion discovery using pairs of trajectories · CVPR 2015 |
Collaborative and social computing
crowdsourcing |
0.2 | 1 | 2023 | Consensus and Subjectivity of Skin Tone Annotation for ML Fairness · NeurIPS 2023 |
Computer vision › 3D vision › motion estimation
optical flow |
0.2 | 1 | 2013 | Video Motion for Every Visible Point · ICCV 2013 |
Computer vision › Video understanding and tracking › feature tracking
dense point tracking |
0.1 | 1 | 2012 | Dense Lagrangian motion estimation with occlusions · CVPR 2012 |
Computer vision › Video understanding and tracking › object tracking
occlusion handling |
0.1 | 1 | 2012 | Dense Lagrangian motion estimation with occlusions · CVPR 2012 |
Computer vision › 3D vision
structure from motion |
0.1 | 1 | 2012 | Simultaneous Compaction and Factorization of Sparse Image Motion Matrices · ECCV (6) 2012 |
Robotics › Robot navigation and mapping
localization |
0.1 | 1 | 2011 | Textured occupancy grids for monocular localization without features · ICRA 2011 |
Robotics › Robot navigation and mapping › localization › vision-based localization
monocular localization |
0.1 | 1 | 2011 | Textured occupancy grids for monocular localization without features · ICRA 2011 |
Computer vision › Image recognition and object detection
spatial alignment |
0.1 | 1 | 2017 | Behavior Discovery and Alignment of Articulated Object Classes from Unstructured Video · Int. J. Comput. Vis. 2017 |
Computer vision › 3D vision › structure from motion
non-rigid structure from motion |
0.1 | 1 | 2016 | Discovering the Physical Parts of an Articulated Object Class from Multiple Videos · CVPR 2016 |
Machine learning › Learning paradigms
unsupervised learning |
0.1 | 1 | 2015 | Articulated motion discovery using pairs of trajectories · CVPR 2015 |
Algorithms and data structures › numerical linear algebra
matrix factorization |
0.0 | 1 | 2012 | Simultaneous Compaction and Factorization of Sparse Image Motion Matrices · ECCV (6) 2012 |
Robotics › Robot navigation and mapping › robot mapping
map representation |
0.0 | 1 | 2011 | Textured occupancy grids for monocular localization without features · ICRA 2011 |
Embedded and real-time systems
cyber-physical system platforms |
0.0 | 1 | 2005 | Scavenging with a Laptop Robot · AAAI 2005 |
Methods — techniques the papers use, named apart from their topics
crowdsourced labeling · 1.3annotation experiments · 1.3action localization · 0.3trajectory displacement · 0.3thin-plate spline · 0.3location model · 0.2energy minimization · 0.2trajectory pairs descriptor · 0.2clustering · 0.2path optimization · 0.2sparse matrix compaction · 0.1matrix factorization · 0.1robotics · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Consensus and Subjectivity of Skin Tone Annotation for ML FairnessabstractUnderstanding different human attributes and how they affect model behavior may become a standard need for all model creation and usage, from traditional computer vision tasks to the newest multimodal generative AI systems. In computer vision specifically, we have relied on datasets augmented with perceived attribute signals (eg, gender presentation, skin tone, and age) and benchmarks enabled by these datasets. Typically labels for these tasks come from human annotators. However, annotating attribute signals, especially skin tone, is a difficult and subjective task. Perceived skin tone is affected by technical factors, like lighting conditions, and social factors that shape an annotator's lived experience.This paper examines the subjectivity of skin tone annotation through a series of annotation experiments using the Monk Skin Tone (MST) scale~\cite{Monk2022Monk}, a small pool of professional photographers, and a much larger pool of trained crowdsourced annotators. Along with this study we release the Monk Skin Tone Examples (MST-E) dataset, containing 1515 images and 31 videos spread across the full MST scale. MST-E is designed to help train human annotators to annotate MST effectively.Our study shows that annotators can reliably annotate skin tone in a way that aligns with an expert in the MST scale, even under challenging environmental conditions. We also find evidence that annotators from different geographic regions rely on different mental models of MST categories resulting in annotations that systematically vary across regions. Given this, we advise practitioners to use a diverse set of annotators and a higher replication count for each image when annotating skin tone for fairness research. Candice Schumann, Femi Olanubi, Auriel Wright, Ellis Monk Jr., Courtney Heldreth, Susanna Ricco |
NeurIPS | 6 |
| 2021 | A Step Toward More Inclusive People Annotations for FairnessabstractThe Open Images Dataset contains approximately 9 million images and is a widely accepted dataset for computer vision research. As is common practice for large datasets, the annotations are not exhaustive, with bounding boxes and attribute labels for only a subset of the classes in each image. In this paper, we present a new set of annotations on a subset of the Open Images dataset called the MIAP (More Inclusive Annotations for People) subset, containing bounding boxes and attributes for all of the people visible in those images. The attributes and labeling methodology for the MIAP subset were designed to enable research into model fairness. In addition, we analyze the original annotation methodology for the person class and its subclasses, discussing the resulting patterns in order to inform future annotation efforts. By considering both the original and exhaustive annotation sets, researchers can also now study how systematic patterns in training annotations affect modeling. Candice Schumann, Susanna Ricco, Utsav Prabhu, Vittorio Ferrari, Caroline Pantofaru |
AIES | 2 |
| 2018 | AVA: A Video Dataset of Spatio-Temporally Localized Atomic Visual ActionsabstractThis paper introduces a video dataset of spatio-temporally localized Atomic Visual Actions (AVA). The AVA dataset densely annotates 80 atomic visual actions in 437 15-minute video clips, where actions are localized in space and time, resulting in 1.59M action labels with multiple labels per person occurring frequently. The key characteristics of our dataset are: (1) the definition of atomic visual actions, rather than composite actions; (2) precise spatio-temporal annotations with possibly multiple annotations for each person; (3) exhaustive annotation of these atomic actions over 15-minute video clips; (4) people temporally linked across consecutive segments; and (5) using movies to gather a varied set of action representations. This departs from existing datasets for spatio-temporal action recognition, which typically provide sparse annotations for composite actions in short video clips. AVA, with its realistic scene and action complexity, exposes the intrinsic difficulty of action recognition. To benchmark this, we present a novel approach for action localization that builds upon the current state-of-the-art methods, and demonstrates better performance on JHMDB and UCF101-24 categories. While setting a new state of the art on existing datasets, the overall results on AVA are low at 15.8% mAP, underscoring the need for developing new approaches for video understanding. Chunhui Gu, Chen Sun 0002, David A. Ross, Carl Vondrick, Caroline Pantofaru, Yeqing Li, Sudheendra Vijayanarasimhan, George Toderici, Susanna Ricco, Rahul Sukthankar, Cordelia Schmid, Jitendra Malik |
CVPR | 9 |
| 2017 | Behavior Discovery and Alignment of Articulated Object Classes from Unstructured VideoabstractWe propose an automatic system for organizing the content of a collection of unstructured videos of an articulated object class (e.g., tiger, horse). By exploiting the recurring motion patterns of the class across videos, our system: (1) identifies its characteristic behaviors, and (2) recovers pixel-to-pixel alignments across different instances. Our system can be useful for organizing video collections for indexing and retrieval. Moreover, it can be a platform for learning the appearance or behaviors of object classes from Internet video. Traditional supervised techniques cannot exploit this wealth of data directly, as they require a large amount of time-consuming manual annotations. The behavior discovery stage generates temporal video intervals, each automatically trimmed to one instance of the discovered behavior, clustered by type. It relies on our novel motion representation for articulated motion based on the displacement of ordered pairs of trajectories. The alignment stage aligns hundreds of instances of the class to a great accuracy despite considerable appearance variations (e.g., an adult tiger and a cub). It uses a flexible thin plate spline deformation model that can vary through time. We carefully evaluate each step of our system on a new, fully annotated dataset. On behavior discovery, we outperform the state-of-the-art improved dense trajectory feature descriptor. On spatial alignment, we outperform the popular SIFT Flow algorithm. Luca Del Pero, Susanna Ricco, Rahul Sukthankar, Vittorio Ferrari |
Int. J. Comput. Vis. | 2 |
| 2016 | Discovering the Physical Parts of an Articulated Object Class from Multiple VideosabstractWe propose a motion-based method to discover the physical parts of an articulated object class (e.g. head/torso/leg of a horse) from multiple videos. The key is to find object regions that exhibit consistent motion relative to the rest of the object, across multiple videos. We can then learn a location model for the parts and segment them accurately in the individual videos using an energy function that also enforces temporal and spatial consistency in part motion. Unlike our approach, traditional methods for motion segmentation or non-rigid structure from motion operate on one video at a time. Hence they cannot discover a part unless it displays independent motion in that particular video. We evaluate our method on a new dataset of 32 videos of tigers and horses, where we significantly outperform a recent motion segmentation method on the task of part discovery (obtaining roughly twice the accuracy). Luca Del Pero, Susanna Ricco, Rahul Sukthankar, Vittorio Ferrari |
CVPR | 2 |
| 2015 | Articulated motion discovery using pairs of trajectoriesabstractWe propose an unsupervised approach for discovering characteristic motion patterns in videos of highly articulated objects performing natural, unscripted behaviors, such as tigers in the wild. We discover consistent patterns in a bottom-up manner by analyzing the relative displacements of large numbers of ordered trajectory pairs through time, such that each trajectory is attached to a different moving part on the object. The pairs of trajectories descriptor relies entirely on motion and is more discriminative than state-of-the-art features that employ single trajectories. Our method generates temporal video intervals, each automatically trimmed to one instance of the discovered behavior, and clusters them by type (e.g., running, turning head, drinking water). We present experiments on two datasets: dogs from YouTube-Objects and a new dataset of National Geographic tiger videos. Results confirm that our proposed descriptor outperforms existing appearance- and trajectory-based descriptors (e.g., HOG and DTFs) on both datasets and enables us to segment unconstrained animal video into intervals containing single behaviors. Luca Del Pero, Susanna Ricco, Rahul Sukthankar, Vittorio Ferrari |
CVPR | 2 |
| 2013 | Video Motion for Every Visible PointabstractDense motion of image points over many video frames can provide important information about the world. However, occlusions and drift make it impossible to compute long motion paths by merely concatenating optical flow vectors between consecutive frames. Instead, we solve for entire paths directly, and flag the frames in which each is visible. As in previous work, we anchor each path to a unique pixel which guarantees an even spatial distribution of paths. Unlike earlier methods, we allow paths to be anchored in any frame. By explicitly requiring that at least one visible path passes within a small neighborhood of every pixel, we guarantee complete coverage of all visible points in all frames. We achieve state-of-the-art results on real sequences including both rigid and non-rigid motions with significant occlusions. Susanna Ricco, Carlo Tomasi |
ICCV | 1 |
| 2012 | Dense Lagrangian motion estimation with occlusionsabstractWe couple occlusion modeling and multi-frame motion estimation to compute dense, temporally extended point trajectories in video with significant occlusions. Our approach combines robust spatial regularization with spatially and temporally global occlusion labeling in a variational, Lagrangian framework with subspace constraints. We track points even through ephemeral occlusions. Experiments demonstrate accuracy superior to the state of the art while tracking more points through more frames. Susanna Ricco, Carlo Tomasi |
CVPR | 1 |
| 2012 | Simultaneous Compaction and Factorization of Sparse Image Motion Matrices
Susanna Ricco, Carlo Tomasi |
ECCV (6) | 1 |
| 2011 | Textured occupancy grids for monocular localization without featuresabstractA textured occupancy grid map is an extremely versatile data structure. It can be used to render human readable views and for laser rangefinder localization algorithms. For camera-based localization, landmark or feature based maps tend to be favored in current research. This may be because of a tacit assumption that working with a textured occupancy grid with a camera would be impractical. We demonstrate that a textured occupancy grid can be combined with an extremely simple monocular localization algorithm to produce a viable localization solution. Our approach is simple, efficient, and produces localization results comparable to laser localization results. A consequence of this result is that a single map representation, the textured occupancy grid, can now be used for humans, robots with laser rangefinders, and robots with just a single camera. Julian Mason, Susanna Ricco, Ronald Parr |
ICRA | 2 |
| 2009 | Fingerspelling Recognition through Classification of Letter-to-Letter Transitions
Susanna Ricco, Carlo Tomasi |
ACCV (3) | 1 |
| 2009 | Correcting Motion Artifacts in Retinal Spectral Domain Optical Coherence Tomography via Image Registration
Susanna Ricco, Hiroshi Ishikawa 0005, Gadi Wollstein, Joel S. Schuman |
MICCAI (1) | 1 |
| 2005 | Scavenging with a Laptop Robot
Alan Davidson, Mac Mason, Susanna Ricco, Ben Tribelhorn, Zachary Dodds |
AAAI | 3 |