EDBT 2026 Demo / reviewers in the wild / expert
Juana Valeria Hurtado
dblp:263/2033
· DBLP profile ↗
4ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0002-7878-4269ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Video understanding and tracking · 40% 3D vision · 12% Efficient and distributed learning · 11% |
Topics — the 15 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
depth estimation |
0.9 | 1 | 2025 | Panoptic-Depth Forecasting · ICRA 2025 |
Computer vision › Segmentation and scene understanding
panoptic segmentation |
0.9 | 1 | 2025 | Panoptic-Depth Forecasting · ICRA 2025 |
Computer vision › Video understanding and tracking › video prediction
panoptic segmentation forecasting |
0.9 | 1 | 2025 | Panoptic-Depth Forecasting · ICRA 2025 |
Computer vision › Video understanding and tracking › spatio-temporal modeling
spatiotemporal representation |
0.9 | 1 | 2025 | Panoptic-Depth Forecasting · ICRA 2025 |
Computer vision › Video understanding and tracking
video prediction |
0.9 | 1 | 2025 | Panoptic-Depth Forecasting · ICRA 2025 |
Machine learning › Trustworthy machine learning
fairness |
0.8 | 1 | 2024 | Fairness and Bias in Robot Learning · Proc. IEEE 2024 |
Robotics › Motion planning and robot control
robot learning |
0.8 | 1 | 2024 | Fairness and Bias in Robot Learning · Proc. IEEE 2024 |
Computer vision › Video understanding and tracking › object tracking › multi-modal tracking
audio-visual tracking |
0.5 | 1 | 2021 | There Is More Than Meets the Eye: Self-Supervised Multi-Object Detection and Tracking With Sound by Distilling Multimodal Knowledge · CVPR 2021 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.5 | 1 | 2021 | There Is More Than Meets the Eye: Self-Supervised Multi-Object Detection and Tracking With Sound by Distilling Multimodal Knowledge · CVPR 2021 |
Machine learning › Efficient and distributed learning › model compression › knowledge distillation › cross-modal distillation
multimodal distillation |
0.5 | 1 | 2021 | There Is More Than Meets the Eye: Self-Supervised Multi-Object Detection and Tracking With Sound by Distilling Multimodal Knowledge · CVPR 2021 |
Computer vision › Image recognition and object detection › object detection
multi-object detection |
0.5 | 1 | 2021 | There Is More Than Meets the Eye: Self-Supervised Multi-Object Detection and Tracking With Sound by Distilling Multimodal Knowledge · CVPR 2021 |
Computer vision › Video understanding and tracking
multi-object tracking |
0.5 | 1 | 2021 | There Is More Than Meets the Eye: Self-Supervised Multi-Object Detection and Tracking With Sound by Distilling Multimodal Knowledge · CVPR 2021 |
Robotics › Autonomous driving
perception |
0.3 | 1 | 2025 | Panoptic-Depth Forecasting · ICRA 2025 |
Computer vision › 3D vision › 3d scene understanding › dynamic scene understanding
scene forecasting |
0.3 | 1 | 2025 | Panoptic-Depth Forecasting · ICRA 2025 |
Machine learning › Trustworthy machine learning › fairness
bias mitigation |
0.2 | 1 | 2024 | Fairness and Bias in Robot Learning · Proc. IEEE 2024 |
Methods — techniques the papers use, named apart from their topics
transformer · 0.9encoder-decoder architecture · 0.9self-supervised pretext task · 0.5multimodal knowledge distillation · 0.5MTA loss · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Panoptic-Depth ForecastingabstractForecasting the semantics and 3D structure of scenes is essential for robots to navigate and plan actions safely. Recent methods have explored semantic and panoptic scene forecasting; however, they do not consider the geometry of the scene. In this work, we propose the panoptic-depth forecasting task for jointly predicting the panoptic segmentation and depth maps of unobserved future frames, from monocular camera images. To facilitate this work, we extend the popular KITTI-360 and Cityscapes benchmarks by computing depth maps from LiDAR point clouds and leveraging sequential labeled data. We also introduce a suitable evaluation metric that quantifies both the panoptic quality and depth estimation accuracy of forecasts in a coherent manner. Furthermore, we present two baselines and propose the novel PDcast architecture that learns rich spatio-temporal representations by incorporating a transformer-based encoder, a forecasting module, and task-specific decoders to predict future panoptic-depth outputs. Extensive evaluations demonstrate the effectiveness of PDcast across two datasets and three forecasting tasks, consistently addressing the primary challenges. We make the code publicly available at https://pdcast.cs.uni-freiburg.de Juana Valeria Hurtado, Riya Mohan, Abhinav Valada |
ICRA | 1 |
| 2025 | Learning Appearance and Motion Cues for Panoptic TrackingabstractPanoptic tracking enables pixel-level scene interpretation of videos by integrating instance tracking in panoptic segmentation. This provides robots with a spatio-temporal understanding of the environment, an essential attribute for their operation in dynamic environments. In this paper, we propose a novel approach for panoptic tracking that simultaneously captures general semantic information and instance-specific appearance and motion features. Unlike existing methods that overlook dynamic scene attributes, our approach leverages both appearance and motion cues through dedicated network heads. These interconnected heads employ multi-scale deformable convolutions that reason about scene motion offsets with semantic context and motion-enhanced appearance features to learn tracking embeddings. Furthermore, we introduce a novel two-step fusion module that integrates the outputs from both heads by first matching instances from the current time step with propagated instances from previous time steps and subsequently refines associations using motion-enhanced appearance embeddings, improving robustness in challenging scenarios. Extensive evaluations of our proposed MAPT model on two benchmark datasets demonstrate that it achieves state-of-the-art performance in panoptic tracking accuracy, surpassing prior methods in maintaining object identities over time. To facilitate future research, we make the code available at http://panoptictracking.cs.uni-freiburg.de. Juana Valeria Hurtado, Sajad Marvi, Rohit Mohan, Abhinav Valada |
IROS | 1 |
| 2024 | Fairness and Bias in Robot LearningabstractMachine learning (ML) has significantly enhanced the abilities of robots, enabling them to perform a wide range of tasks in human environments and adapt to our uncertain real world. Recent works in various ML domains have highlighted the importance of accounting for fairness to ensure that these algorithms do not reproduce human biases and consequently lead to discriminatory outcomes. With robot learning systems increasingly performing more and more tasks in our everyday lives, it is crucial to understand the influence of such biases to prevent unintended behavior toward certain groups of people. In this work, we present the first survey on fairness in robot learning from an interdisciplinary perspective spanning technical, ethical, and legal challenges. We propose a taxonomy for sources of bias and the resulting types of discrimination due to them. Using examples from different robot learning domains, we examine scenarios of unfair outcomes and strategies to mitigate them. We present early advances in the field by covering different fairness definitions, ethical and legal considerations, and methods for fair robot learning. With this work, we aim to pave the road for groundbreaking developments in fair robot learning. Laura Londoño, Juana Valeria Hurtado, Nora Hertz, Philipp Kellmeyer, Silja Voeneky, Abhinav Valada |
Proc. IEEE | 2 |
| 2021 | There Is More Than Meets the Eye: Self-Supervised Multi-Object Detection and Tracking With Sound by Distilling Multimodal KnowledgeabstractAttributes of sound inherent to objects can provide valuable cues to learn rich representations for object detection and tracking. Furthermore, the co-occurrence of audiovisual events in videos can be exploited to localize objects over the image field by solely monitoring the sound in the environment. Thus far, this has only been feasible in scenarios where the camera is static and for single object detection. Moreover, the robustness of these methods has been limited as they primarily rely on RGB images which are highly susceptible to illumination and weather changes. In this work, we present the novel self-supervised MM-DistillNet framework consisting of multiple teachers that leverage diverse modalities including RGB, depth and thermal images, to simultaneously exploit complementary cues and distill knowledge into a single audio student network. We propose the new MTA loss function that facilitates the distillation of information from multimodal teachers in a self-supervised manner. Additionally, we propose a novel self-supervised pretext task for the audio student that enables us to not rely on labor-intensive manual annotations. We introduce a large-scale multimodal dataset with over 113,000 time-synchronized frames of RGB, depth, thermal, and audio modalities. Extensive experiments demonstrate that our approach outperforms state-of-the-art methods while being able to detect multiple objects using only sound during inference and even while moving. Francisco Rivera Valverde, Juana Valeria Hurtado, Abhinav Valada |
CVPR | 2 |