Simon Stent

dblp:146/2461 · DBLP profile ↗
← Back
26ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0002-2623-6383ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 3 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 2 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 See Something, Say Something: Layered Driver Situational Awareness from Language
Pranay Gupta, Simon Stent, John Gideon, Kayli Battel, Laporsha Dees, Patricio Reyes Gomez, Megan Applegate-Kenton, Emily S. Sumner, Guy Rosman
IV2
2025 Beyond Breathalyzers: Towards Pre-Driving Sobriety Testing with a Driver Monitoring Camera
abstract
Field sobriety tests and breathalyzers are commonly used to prevent alcohol-impaired driving, but are expensive and time-consuming to administer. We propose a set of sobriety tests which, in contrast, can feasibly be automated and deployed to modern vehicles equipped with a driver monitoring camera. Our tests are inspired by research on the physiological effects of alcohol, with particular focus on eye movements and gaze behavior. We run an exploratory in-lab study with N=50 subjects (20 alcohol-impaired, 30 control), and train a variety of models to detect alcohol impairment. We find that, using only 10 seconds of observations of the driver, one of the four proposed tests performs comparably to existing non-breathalyzer field sobriety tests. We make our code and data available to support further research efforts to combat alcohol-impaired driving: https://toyotaresearchinstitute.github.io/IV25-beyond-breathalysers/.
Simon Stent, John Gideon, Kimimasa Tamura, Avinash Balachandran, Guy Rosman
IV1
2025 Cognitive Distraction Detection Using Gaze and Pupil with an Interpretable Approach
abstract
Cognitive distraction (CD) is one of the major causes of traffic accidents, but there remains room to improve its detection. Most prior research on CD detection has commonly used basic statistical measures (e.g., mean, standard deviation) of driver-facing camera signals such as gaze and pupil size. However, these signals often exhibit subtle and complex patterns that conventional approaches cannot fully capture. In this paper, we evaluate a wide range of machine learning models and feature extraction methods using data from 52 participants in a driving simulator under two cognitive distraction inducing tasks (n-back and statement tasks). Our results demonstrate that combining gaze, pupil, and features derived from physiological signals (e.g., fixation saccade ratio and gaze entropy) and comprehensive time-series feature extraction boosts detection performance. While deep neural networks (Transformers) excel at modeling intricate relationships, our results show that tree-based ensemble methods (e.g., CatBoost) achieve comparable or higher detection performance while maintaining their advantage of better interpretability. Cross-task experiments further show that models trained on one type of task can generalize to another task. Feature analyses (via SHAP and Sobol) reveal that nonlinearity in vertical gaze movements, baseline pupil size, and greater minimum gaze distance are related to CD. These findings suggest that integrating multiple modalities, sophisticated feature engineering, and employing models capable of capturing nonlinear interactions are effective strategies for detecting CD. To support future research in this field, we release our code and preprocessed data: https://toyotaresearchinstitute.github.io/IV25-cognitive-distraction/.
Kimimasa Tamura, Simon Stent, John Gideon, Kohei Shintani, Guy Rosman
IV2
2024 Seeing Faces in Things: A Model and Dataset for Pareidolia
Mark Hamilton, Simon Stent, Vasha DuTell, Anne Harrington, Jennifer Corbett, Ruth Rosenholtz, William T. Freeman
ECCV (65)2
2024 COCO-Periph: Bridging the Gap Between Human and Machine Perception in the Periphery
abstract
Evaluating deep neural networks (DNNs) as models of human perception has given rich insights into both human visual processing and representational properties of DNNs. We extend this work by analyzing how well DNNs perform compared to humans when constrained by peripheral vision -- which limits human performance on a variety of tasks, but also benefits the visual system significantly. We evaluate this by (1) modifying the Texture Tiling Model (TTM), a well tested model of peripheral vision to be more flexibly used with DNNs, (2) generating a large dataset which we call COCO-Periph that contains images transformed to capture the information available in human peripheral vision, and (3) comparing DNNs to humans at peripheral object detection using a psychophysics experiment. Our results show that common DNNs underperform at object detection compared to humans when simulating peripheral vision with TTM. Training on COCO-Periph begins to reduce the gap between human and DNN performance and leads to small increases in corruption robustness, but DNNs still struggle to capture human-like sensitivity to peripheral clutter. Our work brings us closer to accurately modeling human vision, and paves the way for DNNs to mimic and sometimes benefit from properties of human visual processing.
Anne Harrington, Vasha DuTell, Mark Hamilton, Ayush Tewari, Simon Stent, William T. Freeman, Ruth Rosenholtz
ICLR5
2024 Can Pupillometry be used to Detect Driver Hazard Awareness?
abstract
Modern Advanced Driver-Assistance Systems (ADAS) increasingly rely on interactions between vehicle and human driver. To inform these interactions, it is helpful for a vehicle system to have a good understanding of a driver's situational awareness. In this work we explore a relatively under-exploited, passively measurable signal which might provide insight into a driver's awareness: the constriction and dilation of their pupils over time, or pupillometry. We ask whether pupillometry might be practically useful to detect if and when a driver becomes aware of a road hazard. Using a dataset of driver responses to both hazardous and routine scenarios during simulated semi-automated driving, we compare models trained on pupillometric data to a model trained on facial responses, and demonstrate how their performances differ in terms of accuracy and latency. While a driver's facial expressions are, as expected, a useful cue to determine awareness (0.82 AUC on held-out test stimuli), we find that pupillometric data alone can provide an even more meaningful signal (0.93 AUC). In addition, we find that the pupillometric model performance degrades more gracefully than the face model when tested on unseen subjects, while fusing models yields further accuracy and latency improvements given sufficient training data. We characterize the shape of the performance vs. latency curve for all models and make our code available for reproducibility.
Kimimasa Tamura, John Gideon, Simon Stent, Guy Rosman
SMC3
2023 Tracking Through Containers and Occluders in the Wild
abstract
Tracking objects with persistence in cluttered and dynamic environments remains a difficult challenge for computer vision systems. In this paper, we introduce TCOW, a new benchmark and model for visual tracking through heavy occlusion and containment. We set up a task where the goal is to, given a video sequence, segment both the projected extent of the target object, as well as the surrounding container or occluder whenever one exists. To study this task, we create a mixture of synthetic and annotated real datasets to support both supervised learning and structured evaluation of model performance under various forms of task variation, such as moving or nested containment. We evaluate two recent transformer-based video models and find that while they can be surprisingly capable of tracking targets under certain settings of task variation, there remains a considerable performance gap before we can claim a tracking model to have acquired a true notion of object permanence.
Basile Van Hoorick, Pavel Tokmakov, Simon Stent, Carl Vondrick
CVPR3
2023 What You Can Reconstruct from a Shadow
abstract
3D reconstruction is a fundamental problem in computer vision, and the task is especially challenging when the object to reconstruct is partially or fully occluded. We introduce a method that uses the shadows cast by an unobserved object in order to infer the possible 3D volumes under occlusion. We create a differentiable image formation model that allows us to jointly infer the 3D shape of an object, its pose, and the position of a light source. Since the approach is end-to-end differentiable, we are able to integrate learned priors of object geometry in order to generate realistic 3D shapes of different object categories. Experiments and visualizations show that the method is able to generate multiple possible solutions that are consistent with the observation of the shadow. Our approach works even when the position of the light source and object pose are both unknown. Our approach is also robust to real-world images where ground-truth shadow mask is unknown.
Ruoshi Liu, Sachit Menon, Chengzhi Mao, Dennis Park, Simon Stent, Carl Vondrick
CVPR5
2023 Exploring perceptual straightness in learned visual representations
Anne Harrington, Vasha DuTell, Ayush Tewari, Mark Hamilton, Simon Stent, Ruth Rosenholtz, William T. Freeman
ICLR5
2022 Revealing Occlusions with 4D Neural Fields
abstract
For computer vision systems to operate in dynamic situations, they need to be able to represent and reason about object permanence. We introduce a framework for learning to estimate 4D visual representations from monocular RGB-D video, which is able to persist objects, even once they become obstructed by occlusions. Unlike traditional video representations, we encode point clouds into a continuous representation, which permits the model to attend across the spatiotemporal context to resolve occlusions. On two large video datasets that we release along with this paper, our experiments show that the representation is able to successfully reveal occlusions for several tasks, without any architectural changes. Visualizations show that the attention mechanism automatically learns to follow occluded objects. Since our approach can be trained end-to-end and is easily adaptable, we believe it will be useful for handling occlusions in many video understanding tasks. Data, code, and models are available at occ1usions. cs. co1umbia. edu.
Basile Van Hoorick, Purva Tendulkar, Didac Suris, Dennis Park, Simon Stent, Carl Vondrick
CVPR5
2022 Look Both Ways: Self-supervising Driver Gaze Estimation and Road Scene Saliency
Isaac Kasahara, Simon Stent, Hyun Soo Park
ECCV (13)2
2022 Fine-Grained Egocentric Hand-Object Segmentation: Dataset, Model, and Applications
Lingzhi Zhang, Shenghao Zhou, Simon Stent, Jianbo Shi
ECCV (29)3
2022 Boosting Supervised Learning in Small Data Regimes with Conditional GAN Augmentation
abstract
In many applied computer vision tasks, training data is a scarce resource. This can result in poor performance of deep neural networks trained via supervised learning. We propose an approach to help boost performance in small training data regimes. Our method, called "CONGA", uses a Conditional GAN to Augment training data. Unlike previous work it can be added to any target discriminative model and allows the trade-off of computational cost for improved training sample efficiency. To further improve the quality of generated images and our method’s performance, and distinct from normal conditional GANs, we also propose to supervise the generator’s output via the target model. We compare our approach to similarly motivated methods on various image classification datasets (CIFAR-10, CIFAR-100, and Street View House Numbers), showing significant quantitative improvements.
Tetsuya Ishikawa, Simon Stent
ICIP2
2021 The Way to my Heart is through Contrastive Learning: Remote Photoplethysmography from Unlabelled Video
abstract
The ability to reliably estimate physiological signals from video is a powerful tool in low-cost, pre-clinical health monitoring. In this work we propose a new approach to remote photoplethysmography (rPPG)–the measurement of blood volume changes from observations of a person’s face or skin. Similar to current state-of-the-art methods for rPPG, we apply neural networks to learn deep representations with invariance to nuisance image variation. In contrast to such methods, we employ a fully self-supervised training approach, which has no reliance on expensive ground truth physiological training data. Our proposed method uses contrastive learning with a weak prior over the frequency and temporal smoothness of the target signal of interest. We evaluate our approach on four rPPG datasets, showing that comparable or better results can be achieved compared to recent supervised deep learning methods but without using any annotation. In addition, we incorporate a learned saliency resampling module into both our unsupervised approach and supervised baseline. We show that by allowing the model to learn where to sample the input image, we can reduce the need for hand-engineered features while providing some interpretability into the model’s behavior and possible failure modes. We release code for our complete training and evaluation pipeline to encourage re-producible progress in this exciting new direction.1
John Gideon, Simon Stent
ICCV2
2021 LocTex: Learning Data-Efficient Visual Representations from Localized Textual Supervision
abstract
Computer vision tasks such as object detection and semantic/instance segmentation rely on the painstaking annotation of large training datasets. In this paper, we propose LocTex that takes advantage of the low-cost localized textual annotations (i.e., captions and synchronized mouse-over gestures) to reduce the annotation effort. We introduce a contrastive pre-training framework between images and captions, and propose to supervise the cross-modal attention map with rendered mouse traces to provide coarse localization signals. Our learned visual features capture rich semantics (from free-form captions) and accurate localization (from mouse traces), which are very effective when transferred to various downstream vision tasks. Compared with ImageNet supervised pre-training, LocTex can reduce the size of the pre-training dataset by 10× or the target dataset by 2× while achieving comparable or even improved performance on COCO instance segmentation. When provided with the same amount of annotations, LocTex achieves around 4% higher accuracy than the previous state-of-the-art "vision+language" pre-training approach on the task of PASCAL VOC image classification.
Simon Stent, John Gideon, Song Han 0003
ICCV2
2019 Is Now A Good Time?: An Empirical Study of Vehicle-Driver Communication Timing
abstract
Advances in automotive sensing systems and speech interfaces provide new opportunities for smarter driving assistants or infotainment systems. For both safety and consumer satisfaction reasons, any new system which interacts with drivers must do so at appropriate times. We asked 63 drivers, ''Is now a good time?'' to receive non-driving information during a 50-minute drive. We analyzed 2,734 responses and synchronized automotive and video data, and show that while the chances of choosing a good time can be determined with better success using easily accessible automotive data, certain nuances in the problem require a richer understanding of the driver and environment states in order to achieve higher performance. We illustrate several of these nuances with quantitative and qualitative analyses to contribute to the understanding of how to design a system that might simultaneously minimize the risk of interacting at a bad time while maximizing the window of allowable interruption.
Rob Semmens, Nikolas Martelaro, Pushyami Kaveti, Simon Stent, Wendy Ju
CHI4
2019 Gaze360: Physically Unconstrained Gaze Estimation in the Wild
abstract
Understanding where people are looking is an informative social cue. In this work, we present Gaze360, a large-scale remote gaze-tracking dataset and method for robust 3D gaze estimation in unconstrained images. Our dataset consists of 238 subjects in indoor and outdoor environments with labelled 3D gaze across a wide range of head poses and distances. It is the largest publicly available dataset of its kind by both subject and variety, made possible by a simple and efficient collection method. Our proposed 3D gaze model extends existing models to include temporal information and to directly output an estimate of gaze uncertainty. We demonstrate the benefits of our model via an ablation study, and show its generalization performance via a cross-dataset evaluation against other recent gaze benchmark datasets. We furthermore propose a simple self-supervised approach to improve cross-dataset domain adaptation. Finally, we demonstrate an application of our model for estimating customer attention in a supermarket setting. Our dataset and models will be made available at http://gaze360.csail.mit.edu.
Petr Kellnhofer, Adrià Recasens, Simon Stent, Wojciech Matusik, Antonio Torralba 0001
ICCV3
2019 Spatial Focal Loss for Pedestrian Detection in Fisheye Imagery
abstract
Objects in the periphery of fisheye images can become extremely distorted. This distortion can cause false positives and missed detections for automated object detection systems. This is problematic, not only for systems which have been trained on perspective images, but also for those that have been explicitly trained on fisheye data. In this paper we propose a new cost function for training object detectors on fisheye images. We model fisheye image distortion as an imbalanced domain problem and develop a domain association loss function to approach it with deep learning. We define separate domains based on the level of distortion within the image plane and propose a new objective function, inspired by the recently introduced focal loss for object detection, which we call a spatial focal loss. Our proposed loss incorporates a domain-modulating term which re-weights samples from different domains to encourage the learning of domain-invariant features. We implement spatial focal loss function in the YOLOv2 architecture and evaluate it on the task of pedestrian detection in a fisheye dataset captured by a 360 camera system mounted on a moving vehicle and labeled with over 11,000 pedestrian instances. Our experiments demonstrate that spatial focal loss can improve model performance in the highly distorted image periphery versus existing loss functions, including focal loss, without sacrificing performance in the less distorted image center, with no adaptations to network architectures required. By analyzing the locations of missed detections, we show further evidence that our loss function can improve the learning of domain-invariant features.
Xishuai Peng, Yi Lu Murphey, Simon Stent
WACV3
2018 Learning to Zoom: A Saliency-Based Sampling Layer for Neural Networks
Adrià Recasens, Petr Kellnhofer, Simon Stent, Wojciech Matusik, Antonio Torralba 0001
ECCV (9)3
2018 A Multi-Camera Deep Neural Network for Detecting Elevated Alertness in Drivers
abstract
We present a system for the detection of elevated levels of driver alertness in driver-facing video captured from multiple viewpoints. This problem is important in automotive safety as a helpful feedback signal to determine driver engagement and as a means of automatically flagging anomalous driving events. We generated a dataset of videos from 25 participants overseeing an hour each of driving sequences in a simulator consisting of a mixture of normal and near-miss driving events. Our proposed system consists of a deep neural network which fuses information from three driver-facing cameras to estimate moments of elevated driver alertness. A novel aspect of the system is that it learns to actively re-weight the importance of camera inputs depending on their content. We demonstrate that this approach is not only resilient to dropped or occluded frames, but also has significantly improved performance compared to a system trained on any single stream.
John Gideon, Simon Stent, Luke Fletcher
ICASSP2
2018 Driving Maneuver Detection via Sequence Learning from Vehicle Signals and Video Images
abstract
Driving maneuver detection is one of the most challenging tasks in Advanced Driver Assistance Systems (ADAS). Research has shown that the early notification of improper driving maneuvers is helpful to avoid fatalities and serious accidents. In this paper, we introduce a driver maneuvering detection (DMD) system. The DMD system contains three major computational components, distance based representation of driving context, combined features of vehicle trajectory and VGG-19 network features extracted from the video images of vehicle front view, and a Long Short-Term Memory (LSTM)-based neural network model to learn sequence knowledge in driving maneuvering events. We show through experiments that the DMD system is capable of learning the latent features of five different classes of driving maneuvers and achieving significantly better performance than traditional classification methods on real-world driving trips.
Xishuai Peng, Yi Lu Murphey, Simon Stent
ICPR4
2016 Understanding RealWorld Indoor Scenes with Synthetic Data
abstract
Scene understanding is a prerequisite to many high level tasks for any automated intelligent machine operating in real world environments. Recent attempts with supervised learning have shown promise in this direction but also highlighted the need for enormous quantity of supervised data- performance increases in proportion to the amount of data used. However, this quickly becomes prohibitive when considering the manual labour needed to collect such data. In this work, we focus our attention on depth based semantic per-pixel labelling as a scene understanding problem and show the potential of computer graphics to generate virtually unlimited labelled data from synthetic 3D scenes. By carefully synthesizing training data with appropriate noise models we show comparable performance to state-of-the-art RGBD systems on NYUv2 dataset despite using only depth data as input and set a benchmark on depth-based segmentation on SUN RGB-D dataset.
Ankur Handa, Viorica Patraucean, Vijay Badrinarayanan, Simon Stent, Roberto Cipolla
CVPR4
2016 SceneNet: An annotated model generator for indoor scene understanding
abstract
We introduce SceneNet, a framework for generating high-quality annotated 3D scenes to aid indoor scene understanding. SceneNet leverages manually-annotated datasets of real world scenes such as NYUv2 to learn statistics about object co-occurrences and their spatial relationships. Using a hierarchical simulated annealing optimisation, these statistics are exploited to generate a potentially unlimited number of new annotated scenes, by sampling objects from various existing databases of 3D objects such as ModelNet, and textures such as OpenSurfaces and ArchiveTextures. Depending on the task, SceneNet can be used directly in the form of annotated 3D models for supervised training and 3D reconstruction benchmarking, or in the form of rendered annotated sequences of RGB-D frames or videos.
Ankur Handa, Viorica Patraucean, Simon Stent, Roberto Cipolla
ICRA3
2016 Precise deterministic change detection for smooth surfaces
abstract
We introduce a precise deterministic approach for pixel-wise change detection in images taken of a scene of interest over time. Our motivation is for applications such as artefact condition monitoring and structural inspection, where a common problem is the need to efficiently and accurately identify subtle signs of damage and deterioration. The approach we describe is designed to compensate for the three most common sources of nuisance variation encountered when tackling the problem of change detection, namely: viewpoint variation due to camera motion between images, photometric variation due to lighting differences, and changes in image resolution/focal settings. To tackle viewpoint variation, particularly in areas of low texture, we propose the use of the generalised PatchMatch (PM) correspondence algorithm to compute a dense flow field. The flow field is regularized using a Thin Plate Spline (TPS) model which assumes a smooth underlying geometry and allows registration to be interpolated precisely through areas of low texture or uncertain flow. To compensate for low-frequency lighting variation, we fit a second TPS model to the photometric differences between registered images. Finally, to account for changes in focal settings, we estimate and apply a blurring kernel via optimisation over image differences. We provide a thorough evaluation of the performance of our method on an illustrative toy dataset and on two recent, real-world inspection datasets. Our approach performs favourably versus state-of-the-art baselines in both cases, while remaining relatively transparent to understand and simple to compute.
Simon Stent, Riccardo Gherardi, Björn Stenger, Roberto Cipolla
WACV1
2016 Visual change detection on tunnel linings
Simon Stent, Riccardo Gherardi, Björn Stenger, Kenichi Soga, Roberto Cipolla
Mach. Vis. Appl.1
2015 Detecting Change for Multi-View, Long-Term Surface Inspection
abstract
We describe a system for the detection of changes in multiple views of a tunnel surface. From data gathered by a robotic inspection rig, we use a structure-from-motion pipeline to build panoramas of the surface and register images from different time instances. Reliably detecting changes such as hairline cracks, water ingress and other surface damage between the registered images is a challenging problem: achieving the best possible performance for a given set of data requires sub-pixel precision and careful modelling of the noise sources. The task is further complicated by factors such as unavoidable registration error and changes in image sensors, capture settings and lighting. Our contribution is a novel approach to change detection using a two-channel convolutional neural network. The network accepts pairs of approximately registered image patches taken at different times and classifies them to detect anomalous changes. To train the network, we take advantage of synthetically generated training examples and the homogeneity of the tunnel surfaces to eliminate most of the manual labelling effort. We evaluate our method on field data gathered from a live tunnel over several months, demonstrating it to outperform existing approaches from recent literature and industrial practice.
Simon Stent, Riccardo Gherardi, Björn Stenger, Roberto Cipolla
BMVC1