Julie Stephany Berrio

dblp:227/3179 · also Julie Stephany Berrio Perez, Stephany Berrio Perez · DBLP profile ↗
← Back
18ranked-venue papers
5as first author
14since 2021 · last 2026
0000-0003-3126-7042ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 3 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 3 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Blinking Beyond EAR: A Stable Eyelid Angle Metric for Driver Drowsiness Detection and Data Augmentation
Mathis Wolter, Julie Stephany Berrio, Mao Shan
IV2
2026 Been There, Scanned That: Nostalgia-Driven Point Cloud Compression for Self-Driving Cars
abstract
An autonomous vehicle generates several terabytes of sensor data per day. A significant portion of this data consists of 3D point clouds produced by depth sensors such as LiDAR. This data is transferred to cloud storage, where it is utilized for training machine learning models or conducting analyses, e.g., forensic investigations in the event of an accident. To reduce network and storage costs, this paper introduces DejaView that searches for and uses redundancies on larger temporal scales (days and months) for more effective compression. We designed DejaView with the insight that the operating area of autonomous vehicles is limited and that vehicles mostly traverse the same routes daily. Consequently, the daily collected 3D data is likely similar to the data they’ve captured in the past. To capture this, the core of DejaView is a diff operation that compactly represents point clouds as delta w.r.t. 3D data from the past. Using two months of LiDAR data, DejaView can compress point clouds by a factor of 210 at a reconstruction error of only 15 cm.
Ali Khalid, Jaiaid Mobin, Sumanth Rao Appala, Avinash Maurya, Julie Stephany Berrio, M. Mustafa Rafique, Fawad Ahmad 0002
SenSys5
2026 What demands attention in urban street scenes? From scene understanding towards road safety: A survey of vision-driven datasets and studies
abstract
Advances in vision-based sensors and computer vision algorithms have significantly improved the analysis and understanding of traffic scenarios. To facilitate the use of these improvements for road safety, this survey systematically categorizes the critical elements that demand attention in traffic scenarios and comprehensively analyzes available vision-driven tasks and datasets. Compared to existing surveys that focus on isolated domains, our taxonomy categorizes attention-worthy traffic entities into two main groups, namely anomalies (abnormal entities) and pertinent entities (normal but critical elements), integrating eleven categories and twenty-three subclasses. It establishes connections between inherently related fields and provides a unified analytical framework. Based on the proposed taxonomy, our survey highlights the analysis of 40 vision-driven tasks and the comprehensive examinations and visualizations of 78 available datasets, including their basic characteristics, sensor settings, label design, visualization practices, annotation schemas, and the resulting implications. The cross-domain investigation reveals substantial variations in benchmark quality across tasks, with recurring limitations including uneven task coverage, imbalanced distributions, inconsistent or insufficient annotations, and limited multimodal and cross-task support. Our article further outlines promising solutions from the perspectives of task formulation, benchmark evaluations, dataset adoption and future dataset construction. The integrated taxonomy, comprehensive analysis, and recapitulatory tables provide researchers with a holistic overview of this rapidly evolving field, guiding strategic resource selection, and highlighting critical yet underexplored areas.
Yaoqi Huang, Julie Stephany Berrio, Mao Shan, Stewart Worrall 0002
Eng. Appl. Artif. Intell.2
2025 Animal Interaction with Autonomous Mobility Systems: Designing for Multi-Species Coexistence
abstract
Autonomous mobility systems increasingly operate in environments shared with animals, from urban pets to wildlife.However, their design has largely focused on human interaction, with limited understanding of how non-human species perceive, respond to, or are affected by these systems.Motivated by research in Animal-Computer Interaction (ACI) and more-than-human design, this study investigates animal interactions with autonomous mobility through a multi-method approach combining a scoping review (45 articles), online ethnography (39 YouTube videos and 11 Reddit discussions), and expert interviews (8 participants).Our analysis surfaces five key areas of concern: Physical Impact (e.g., collisions, failures to detect), Behavioural Effects (e.g., avoidance, stress), Accessibility Concerns (particularly for service animals), Ethics and Regulations, and Urban Disturbance.We conclude with design and policy directions aimed at supporting multispecies coexistence in the age of autonomous systems.This work underscores the importance of incorporating non-human perspectives to ensure safer, more inclusive futures for all species.
Tram Thi Minh Tran, Xinyan Yu 0005, Marius Hoggenmüller, Callum Parker, Paul Schmitt, Julie Stephany Berrio, Stewart Worrall 0002, Martin Tomitsch
AutomotiveUI6
2025 Spotting the Unexpected (STU): A 3D LiDAR Dataset for Anomaly Segmentation in Autonomous Driving
abstract
To operate safely, autonomous vehicles (AVs) need to detect and handle unexpected objects or anomalies on the road. While significant research exists for anomaly detection and segmentation in 2D, research progress in 3D is underexplored. Existing datasets lack high-quality multimodal data that are typically found in AVs. This paper presents a novel dataset for anomaly segmentation in driving scenarios. To the best of our knowledge, it is the first publicly available dataset focused on road anomaly segmentation with dense 3D semantic labeling, incorporating both LiDAR and camera data, as well as sequential information to enable anomaly detection across various ranges. This capability is critical for the safe navigation of autonomous vehicles. We adapted and evaluated several baseline models for 3D segmentation, highlighting the challenges of 3D anomaly detection in driving environments. Our dataset and evaluation code will be openly available, facilitating the testing and performance comparison of different approaches.
Alexey Nekrasov 0001, Malcolm Burdorf, Stewart Worrall 0002, Bastian Leibe, Julie Stephany Berrio
CVPR5
2025 Mixed Signals: A Diverse Point Cloud Dataset for Heterogeneous LiDAR V2X Collaboration
abstract
Vehicle-to-everything (V2X) collaborative perception has emerged as a promising solution to address the limitations of single-vehicle perception systems. However, existing V2X datasets are limited in scope, diversity, and quality. To address these gaps, we present Mixed Signals, a comprehensive V2X dataset featuring 45.1k point clouds and 240.6k bounding boxes collected from three connected autonomous vehicles (CAVs) equipped with two different configurations of LiDAR sensors, plus a roadside unit with dual LiDARs. Our dataset provides point clouds and bounding box annotations across 10 classes, ensuring reliable data for perception training. We provide detailed statistical analysis on the quality of our dataset and extensively benchmark existing V2X methods on it. The Mixed Signals dataset is ready-to-use, with precise alignment and consistent annotations across time and viewpoints. Dataset website is available at https://mixedsignalsdataset.cs.cornell.edu/.
Katie Luo, Minh-Quan Dao, Mark E. Campbell, Wei-Lun Chao, Kilian Q. Weinberger, Ezio Malis, Vincent Frémont, Bharath Hariharan, Mao Shan, Stewart Worrall 0002, Julie Stephany Berrio
ICCV12
2024 InverseMatrixVT3D: An Efficient Projection Matrix-Based Approach for 3D Occupancy Prediction
abstract
This paper introduces InverseMatrixVT3D, an efficient method for transforming multi-view image features into 3D feature volumes for 3D semantic occupancy prediction. Existing methods for constructing 3D volumes often rely on depth estimation, device-specific operators, or transformer queries, which hinders the widespread adoption of 3D occupancy models. In contrast, our approach leverages two projection matrices to store the static mapping relationships and matrix multiplications to efficiently generate global Bird’s Eye View (BEV) features and local 3D feature volumes. Specifically, we achieve this by performing matrix multiplications between multi-view image feature maps and two sparse projection matrices. We introduce a sparse matrix handling technique for the projection matrices to optimize GPU memory usage. Moreover, a global-local attention fusion module is proposed to integrate the global BEV features with the local 3D feature volumes to obtain the final 3D volume. We also employ a multi-scale supervision mechanism to enhance performance further. Extensive experiments performed on the nuScenes and SemanticKITTI datasets reveal that our approach not only stands out for its simplicity and effectiveness but also achieves the top performance in detecting vulnerable road users (VRU), crucial for autonomous driving and road safety. The code has been made available at: https://github.com/DanielMing123/InverseMatrixVT3D
Zhenxing Ming, Julie Stephany Berrio, Mao Shan, Stewart Worrall 0002
IROS2
2024 Practical Collaborative Perception: A Framework for Asynchronous and Multi-Agent 3D Object Detection
abstract
Occlusion is a major challenge for LiDAR-based object detection methods as it renders regions of interest unobservable to the ego vehicle. A proposed solution to this problem comes from collaborative perception via Vehicle-to-Everything (V2X) communication, which leverages a diverse perspective thanks to the presence of connected agents (vehicles and intelligent roadside units) at multiple locations to form a complete scene representation. The major challenge of V2X collaboration is the performance-bandwidth tradeoff which presents two questions (i) which information should be exchanged over the V2X network, and (ii) how the exchanged information is fused. The current state-of-the-art resolves to the mid-collaboration approach where Birds-Eye View (BEV) images of point clouds are communicated to enable a deep interaction among connected agents while reducing bandwidth consumption. While achieving strong performance, the real-world deployment of most mid-collaboration approaches are hindered by their overly complicated architectures and unrealistic assumptions about inter-agent synchronization. In this work, we devise a simple yet effective collaboration method based on exchanging the outputs from each agent that achieves a better bandwidth-performance tradeoff while minimising the required changes to the single-vehicle detection models. Moreover, we relax the assumptions used in existing state-of-the-art approaches about inter-agent synchronization to only require a common time reference among connected agents, which can be achieved in practice using GPS time. Experiments on the V2X-Sim dataset show that our collaboration method reaches 76.72 mean average precision which is 99% the performance of the early collaboration method while consuming as much bandwidth as the late collaboration (0.01 MB on average). The code is released in https://github.com/quan-dao/practical-collab-perception.
Minh-Quan Dao, Julie Stephany Berrio, Vincent Frémont, Mao Shan, Elwan Héry, Stewart Worrall 0002
IV2
2024 Label-Efficient 3D Object Detection For Road-Side Units
abstract
Occlusion presents a significant challenge for safety-critical applications such as autonomous driving. Collaborative perception has recently attracted a large research interest thanks to the ability to enhance the perception of autonomous vehicles via deep information fusion with intelligent roadside units (RSU), thus minimizing the impact of occlusion. While significant advancement has been made, the data-hungry nature of these methods creates a major hurdle for their realworld deployment, particularly due to the need for annotated RSU data. Manually annotating the vast amount of RSU data required for training is prohibitively expensive, given the sheer number of intersections and the effort involved in annotating point clouds. We address this challenge by devising a label-efficient object detection method for RSU based on unsupervised object discovery. Our paper introduces two new modules: one for object discovery based on a spatial temporal aggregation of point clouds, and another for refinement. Furthermore, we demonstrate that fine-tuning on a small portion of annotated data allows our object discovery models to narrow the performance gap with, or even surpass, fully supervised models. Extensive experiments are carried out in simulated and real-world datasets to evaluate our method†.
Minh-Quan Dao, Holger Caesar, Julie Stephany Berrio, Mao Shan, Stewart Worrall 0002, Vincent Frémont, Ezio Malis
IV3
2024 Safety Driver Attention on Autonomous Vehicle Operation Based on Head Pose and Vehicle Perception
abstract
Despite the continual advances in Advanced Driver Assistance Systems (ADAS) and the development of high-level autonomous vehicles (AV), there is a consensus that for the short to medium term, there is a requirement for a human supervisor to handle the edge cases that inevitably arise. Given this requirement, the state of the autonomous vehicle operator (referred to as the safety driver) must be monitored to ensure their contribution to the vehicle's safe operation. This paper introduces a dual-source approach integrating data from an infrared camera facing the safety driver and vehicle perception systems to produce a metric for safety driver alertness to promote and ensure safe operator behaviour. The infrared camera detects the safety driver’s head, enabling the calculation of head orientation, which is relevant as the head typically moves according to the individual's focus of attention. By incorporating environmental data from the perception system, it becomes possible to determine whether the safety driver observes objects in the surroundings. Experiments were conducted using data collected in Sydney, Australia, simulating AV operations in an urban environment. Our results demonstrate that the proposed system effectively determines a metric for the attention levels of the safety driver, enabling interventions such as warnings or reducing autonomous functionality as appropriate. The results indicate reduced awareness on subsequent laps during the study, demonstrating the "automation complacency" phenomenon. This comprehensive solution shows promise in contributing to ADAS and AVs’ overall safety and efficiency in a real-world setting.
Santiago Gerling Konrad, Julie Stephany Berrio, Mao Shan, Favio R. Masson, Eduardo M. Nebot, Stewart Worrall 0002
IV2
2024 Practical Collaborative Perception: A Framework for Asynchronous and Multi-Agent 3D Object Detection
abstract
Occlusion is a major challenge for LiDAR-based object detection methods as it renders regions of interest unobservable to the ego vehicle. A proposed solution to this problem comes from collaborative perception via Vehicle-to-Everything (V2X) communication, which leverages a diverse perspective thanks to the presence of connected agents (vehicles and intelligent roadside units) at multiple locations to form a complete scene representation. The major challenge of V2X collaboration is the performance-bandwidth tradeoff which presents two questions 1) which information should be exchanged over the V2X network and 2) how the exchanged information is fused. The current state-of-the-art resolves to the mid-collaboration approach where Birds-Eye View (BEV) images of point clouds are communicated to enable a deep interaction among connected agents while reducing bandwidth consumption. While achieving strong performance, the real-world deployment of most mid-collaboration approaches are hindered by their overly complicated architectures and unrealistic assumptions about inter-agent synchronization. In this work, we devise a simple yet effective collaboration method based on exchanging the outputs from each agent that achieves a better bandwidth-performance tradeoff while minimising the required changes to the single-vehicle detection models. Moreover, we relax the assumptions used in existing state-of-the-art approaches about inter-agent synchronization to only require a common time reference among connected agents, which can be achieved in practice using GPS time. Experiments on the V2X-Sim dataset show that our collaboration method reaches 76.72 mean average precision which is 99% the performance of the early collaboration method while consuming as much bandwidth as the late collaboration (0.01 MB on average). The code will be released in https://github.com/quan-dao/practical-collab-perception.
Minh-Quan Dao, Julie Stephany Berrio, Vincent Frémont, Mao Shan, Elwan Héry, Stewart Worrall 0002
IEEE Trans. Intell. Transp. Syst.2
2023 Viewer-Centred Surface Completion for Unsupervised Domain Adaptation in 3D Object Detection
abstract
Every autonomous driving dataset has a different configuration of sensors, originating from distinct geographic regions and covering various scenarios. As a result, 3D detectors tend to overfit the datasets they are trained on. This causes a drastic decrease in accuracy when the detectors are trained on one dataset and tested on another. We observe that lidar scan pattern differences form a large component of this reduction in performance. We address this in our approach, SEE-VCN, by designing a novel viewer-centred surface completion network (VCN) to complete the surfaces of objects of interest within an unsupervised domain adaptation framework, SEE [1]. With SEE-VCN, we obtain a unified representation of objects across datasets, allowing the network to focus on learning geometry, rather than overfitting on scan patterns. By adopting a domain-invariant representation, SEE-VCN can be classed as a multi-target domain adaptation approach where no annotations or re-training is required to obtain 3D detections for new scan patterns. Through extensive experiments, we show that our approach outperforms previous domain adaptation methods in multiple domain adaptation settings. Our code and data are available at https://github.com/darrenjkt/SEE-VCN.
Darren Tsai, Julie Stephany Berrio, Mao Shan, Eduardo M. Nebot, Stewart Worrall 0002
ICRA2
2022 Camera-LIDAR Integration: Probabilistic Sensor Fusion for Semantic Mapping
abstract
An automated vehicle operating in an urban environment must be able to perceive and recognise objects and obstacles in a three-dimensional world for navigation and path planning. In order to plan and execute accurate and sophisticated driving maneuvers, a high-level contextual understanding of the surroundings is essential. Due to the recent progress in image processing, it is now possible to obtain high definition semantic information in 2D from monocular cameras, though cameras cannot reliably provide the high accuracy 3D information provided by lasers. The fusion of these two sensor modalities can overcome the shortcomings of each individual sensor, though there are a number of important challenges that need to be addressed in a probabilistic manner. In this paper we address the common, yet challenging, LIDAR/camera/semantic fusion problems which are seldom approached in a wholly probabilistic manner. Our approach is capable of using a multi-sensor platform to build a three-dimensional semantic voxelized map that considers the uncertainty of all of the processes involved. We present a probabilistic pipeline that incorporates uncertainty from the sensor readings (cameras, LIDAR, IMU and wheel encoders), compensation for the motion of the vehicle, and heuristic label probabilities for the semantic images depicted inFig. 1. We also present a novel and efficient viewpoint validation algorithm to check for occlusions within the camera frame. A probabilistic projection is performed from the camera images to the LIDAR point cloud. Each labelled LIDAR scan then feeds into an octree map-building algorithm that updates the class probabilities of the map voxels every time a new observation is available. We validate our approach using a set of qualitative and quantitative experiments using the USyd Campus Dataset.1These tests demonstrate the usefulness of a probabilistic sensor fusion approach by evaluating the performance of the perception system in a typical autonomous vehicle application.
Julie Stephany Berrio, Mao Shan, Stewart Worrall 0002, Eduardo M. Nebot
IEEE Trans. Intell. Transp. Syst.1
2022 Long-Term Map Maintenance Pipeline for Autonomous Vehicles
abstract
For autonomous vehicles to operate persistently in a typical urban environment, it is essential to have high accuracy position information. This requires a mapping and localisation system that can adapt to changes over time. A localisation approach based on a single-survey map will not be suitable for long-term operation as it does not incorporate variations in the environment. In this paper, we present new algorithms to maintain a featured-based map. A map maintenance pipeline is proposed that can continuously update a map with the most relevant features taking advantage of the changes in the surroundings. Our pipeline detects and removes transient features based on their geometrical relationships with the vehicle’s pose. Newly identified features became part of a new feature map and are assessed by the pipeline as candidates for the localisation map. By purging out-of-date features and adding newly detected features, we continually update the prior map to more accurately represent the most recent environment. We have validated our approach using the USyd Campus Dataset, which includes more than 18 months of data. The results presented demonstrate that our maintenance pipeline produces a resilient map which can provide sustained localisation performance over time.
Julie Stephany Berrio, Stewart Worrall 0002, Mao Shan, Eduardo M. Nebot
IEEE Trans. Intell. Transp. Syst.1
2020 Automated Evaluation of Semantic Segmentation Robustness for Autonomous Driving
abstract
One of the fundamental challenges in the design of perception systems for autonomous vehicles is validating the performance of each algorithm under a comprehensive variety of operating conditions. In the case of vision-based semantic segmentation, there are known issues when encountering new scenarios that are sufficiently different to the training data. In addition, even small variations in environmental conditions, such as illumination and precipitation, can affect the classification performance of the segmentation model. Given the reliance on visual information, these effects often translate into poor semantic pixel classification which can potentially lead to catastrophic consequences when driving autonomously. This paper presents a novel method for analyzing the robustness of semantic segmentation models and provides a number of metrics to evaluate the classification performance over a variety of environmental conditions. The process incorporates an additional sensor (lidar) to automate the process and improve the system integrity, eliminating the need for labor-intensive hand labeling of validation data. The experimental results are presented based on multiple datasets collected at different times of the year with different environmental conditions. We extract the ”Road” class using the lidar to demonstrate the concepts, but this could be extended to other classes with different feature detection algorithms. These results show that the semantic segmentation performance varies depending on the weather, camera parameters, and existence of shadows. The results also demonstrate how the metrics can be used to compare and validate the performance after making improvements to a model, and compare the performance of different networks.
Wei Zhou 0025, Julie Stephany Berrio, Stewart Worrall 0002, Eduardo M. Nebot
IEEE Trans. Intell. Transp. Syst.2
2019 Identifying robust landmarks in feature-based maps
abstract
To operate in an urban environment, an automated vehicle must be capable of accurately estimating its position within a global map reference frame. This is necessary for optimal path planning and safe navigation. To accomplish this over an extended period of time, the global map requires long term maintenance. This includes the addition of newly observable features and the removal of transient features belonging to dynamic objects. The latter is especially important for the long-term use of the map as matching against a map with features that no longer exist can result in incorrect data associations, and consequently erroneous localisation. This paper addresses the problem of removing features from the map that correspond to objects that are no longer observable/present in the environment. This is achieved by assigning a single score which depends on the geometric distribution and characteristics when the features are re-detected (or not) on different occasions. Our approach not only eliminates ephemeral features, but can also be used as a reduction algorithm for highly dense maps. We tested our approach using half a year of weekly drives over the same 500 metre section of road in an urban environment. The results presented demonstrate the validity of the long term approach to map maintenance.
Julie Stephany Berrio, James R. Ward, Stewart Worrall 0002, Eduardo M. Nebot
IV1
2019 Updating the visibility of a feature-based map for long-term maintenance
abstract
Mobile vehicles operating in urban navigation applications can achieve high integrity localisation with high accuracy by using maps of the surroundings. To accomplish this, the map should always have an accurate representation of the environment. Thus, it is necessary to detect and remove the map components that no longer exist in the current environment. This maintains the map compactness and dependability while simplifying the data association problem. This paper addresses the problem of deletion of transient map components by taking advantage of the geometric connection between the map and agent poses in order to establish and update the visibility of each feature. Once the map is created an initial visibility vector is associated with every map element and updated over time. The visibility of a map element which no longer exists is reduced and ultimately removed from the map. We demonstrate our approach in a 2D feature-based map composed of poles and corners extracted from information provided by a Iidar sensor. The experimental results show the map update using a seven-month data set collected in the University of Sydney campus.
Julie Stephany Berrio, James R. Ward, Stewart Worrall 0002, Eduardo M. Nebot
IV1
2018 Octree map based on sparse point cloud and heuristic probability distribution for labeled images
abstract
To navigate through urban roads, an automated vehicle must be able to perceive and recognize objects in a three-dimensional environment. A high level contextual understanding of the surroundings is necessary to execute accurate driving maneuvers. This paper presents a novel approach to build three dimensional semantic octree maps from lidar scans and the output of a convolutional neural network (CNN) to obtain the labels of the environment. We present a heuristic method to associate uncertainties to the labels from the images based on a combination of the labels themselves, score maps retrieved by the CNN and the raw images. These uncertainties and the camera-lidar calibration parameters for multiple cameras are considered in the projection of the labels and their uncertainties into the point cloud. Every labeled lidar scan works as an input to an octree map building algorithm that calculates and updates the label probabilities of the voxels in the map. This paper also presents a qualitative and quantitative evaluation of accuracy, analyzing projection in single lidar scans and complete maps built with our probabilistic octree framework.
Julie Stephany Berrio, Wei Zhou 0025, James R. Ward, Stewart Worrall 0002, Eduardo M. Nebot
IROS1