Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Steven Lovegrove

dblp:03/8535 · also Steven J. Lovegrove · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
3since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
3D vision · 88% Speech recognition and synthesis · 10% Robot navigation and mapping · 2%
Computer graphics and multimedia
7 papers
Geometric modeling and processing · 36% Multimedia analysis and retrieval · 23% Computer animation and physical simulation · 20%

Topics — the 27 heaviest of 30, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
neural radiance field
1.122022
Neural 3D Video Synthesis from Multi-view Video · CVPR 2022
STaR: Self-Supervised Tracking and Reconstruction of Rigid Objects in Motion With Neural Rendering · CVPR 2021
Computer vision › 3D vision › 3d shape representation › implicit surface representation
signed distance function
0.822020
Deep Local Shapes: Learning Local SDF Priors for Detailed 3D Reconstruction · ECCV (29) 2020
DeepSDF: Learning Continuous Signed Distance Functions for Shape Representation · CVPR 2019
Computer vision › 3D vision
3d reconstruction
0.822020
Deep Local Shapes: Learning Local SDF Priors for Detailed 3D Reconstruction · ECCV (29) 2020
Reconstructing scenes with mirror and glass surfaces · ACM Trans. Graph. 2018
Multimedia analysis and retrieval › multimedia analysis
multimodal conversation analysis
0.712023
EgoCom: A Multi-Person Multi-Modal Egocentric Communications Dataset · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Computer vision › 3D vision › neural radiance field
dynamic neural radiance field
0.612022
Neural 3D Video Synthesis from Multi-view Video · CVPR 2022
Computer vision › 3D vision › novel view synthesis
multi-view video generation
0.612022
Neural 3D Video Synthesis from Multi-view Video · CVPR 2022
Computer vision › 3D vision
novel view synthesis
0.612022
Neural 3D Video Synthesis from Multi-view Video · CVPR 2022
Computer animation and physical simulation › motion synthesis
motion interpolation
0.612022
Neural 3D Video Synthesis from Multi-view Video · CVPR 2022
Computer vision › 3D vision › 3d scene reconstruction
dynamic scene reconstruction
0.512021
STaR: Self-Supervised Tracking and Reconstruction of Rigid Objects in Motion With Neural Rendering · CVPR 2021
Computer vision › 3D vision › object pose estimation
object pose tracking
0.512021
STaR: Self-Supervised Tracking and Reconstruction of Rigid Objects in Motion With Neural Rendering · CVPR 2021
Computer vision › 3D vision › motion estimation
rigid motion estimation
0.512021
STaR: Self-Supervised Tracking and Reconstruction of Rigid Objects in Motion With Neural Rendering · CVPR 2021
Computer vision › 3D vision › pose estimation
shape and pose estimation
0.412020
FroDO: From Detections to 3D Objects · CVPR 2020
Geometric modeling and processing
3d reconstruction
0.412020
FroDO: From Detections to 3D Objects · CVPR 2020
Computer vision › 3D vision
3d shape representation
0.412019
DeepSDF: Learning Continuous Signed Distance Functions for Shape Representation · CVPR 2019
Geometric modeling and processing › surface reconstruction
shape reconstruction
0.412019
DeepSDF: Learning Continuous Signed Distance Functions for Shape Representation · CVPR 2019
Computer vision › 3D vision › 3d reconstruction › surface reconstruction
planar surface reconstruction
0.312018
Reconstructing scenes with mirror and glass surfaces · ACM Trans. Graph. 2018
Computer vision › 3D vision › 3d reconstruction › non-lambertian surface reconstruction
reflective surface reconstruction
0.312018
Reconstructing scenes with mirror and glass surfaces · ACM Trans. Graph. 2018
Computational photography and imaging › image signal processing
rolling shutter correction
0.212015
A Spline-Based Trajectory Representation for Sensor Fusion and Rolling Shutter Cameras · Int. J. Comput. Vis. 2015
Geometric modeling and processing
trajectory representation
0.212015
A Spline-Based Trajectory Representation for Sensor Fusion and Rolling Shutter Cameras · Int. J. Comput. Vis. 2015
Natural language and speech › Speech recognition and synthesis › spoken dialogue model
turn-taking prediction
0.212023
EgoCom: A Multi-Person Multi-Modal Egocentric Communications Dataset · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Rendering
novel view synthesis
0.112021
STaR: Self-Supervised Tracking and Reconstruction of Rigid Objects in Motion With Neural Rendering · CVPR 2021
Robotics › Robot navigation and mapping › SLAM › visual SLAM › dense visual SLAM
dense tracking and mapping
0.112011
DTAM: Dense tracking and mapping in real-time · ICCV 2011
Computer vision › 3D vision › 3d reconstruction
real-time reconstruction
0.112011
DTAM: Dense tracking and mapping in real-time · ICCV 2011
Computational photography and imaging
image stitching
0.112010
Real-Time Spherical Mosaicing Using Whole Image Alignment · ECCV (3) 2010
Computational photography and imaging › image stitching
spherical mosaicing
0.112010
Real-Time Spherical Mosaicing Using Whole Image Alignment · ECCV (3) 2010
Robotics › Robot navigation and mapping
sensor fusion
0.112015
A Spline-Based Trajectory Representation for Sensor Fusion and Rolling Shutter Cameras · Int. J. Comput. Vis. 2015
Computer vision › 3D vision
image registration
0.012010
Real-Time Spherical Mosaicing Using Whole Image Alignment · ECCV (3) 2010

Methods — techniques the papers use, named apart from their topics

speech-to-text · 1.3bayesian modeling · 1.3time-conditioned neural radiance field · 1.1ray importance sampling · 1.1hierarchical training · 1.1volume rendering · 1.0self-supervised learning · 1.0joint optimization · 1.0encoder network · 0.9deep signed distance function · 0.9
YearPublicationVenuePosition
2023 EgoCom: A Multi-Person Multi-Modal Egocentric Communications Dataset
abstract
Multi-modal datasets in artificial intelligence (AI) often capture a third-person perspective, but our embodied human intelligence evolved with sensory input from the egocentric, first-person perspective. Towards embodied AI, we introduce the Egocentric Communications (EgoCom) dataset to advance the state-of-the-art in conversational AI, natural language, audio speech analysis, computer vision, and machine learning. EgoCom is a first-of-its-kind natural conversations dataset containing multi-modal human communication data captured simultaneously from the participants' egocentric perspectives. EgoCom includes 38.5 hours of synchronized embodied stereo audio, egocentric video with 240,000 ground-truth, time-stamped word-level transcriptions and speaker labels from 34 diverse speakers. We study baseline performance on two novel applications that benefit from embodied data: (1) predicting turn-taking in conversations and (2) multi-speaker transcription. For (1), we investigate Bayesian baselines to predict turn-taking within 5 percent of human performance. For (2), we use simultaneous egocentric capture to combine Google speech-to-text outputs, improving global transcription by 79 percent relative to a single perspective. Both applications exploit EgoCom's synchronous multi-perspective data to augment performance of embodied AI tasks.
Curtis G. Northcutt, Shengxin Zha, Steven Lovegrove, Richard A. Newcombe
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Neural 3D Video Synthesis from Multi-view Video
abstract
We propose a novel approach for 3D video synthesis that is able to represent multi-view video recordings of a dynamic real-world scene in a compact, yet expressive representation that enables high-quality view synthesis and motion interpolation. Our approach takes the high quality and compactness of static neural radiance fields in a new direction: to a model-free, dynamic setting. At the core of our approach is a novel time-conditioned neural radiance field that represents scene dynamics using a set of compact latent codes. We are able to significantly boost the training speed and perceptual quality of the generated imagery by a novel hierarchical training scheme in combination with ray importance sampling. Our learned representation is highly compact and able to represent a 10 second 30 FPS multi-view video recording by 18 cameras with a model size of only 28MB. We demonstrate that our method can render high-fidelity wide-angle novel views at over 1K resolution, even for complex and dynamic scenes. We perform an extensive qualitative and quantitative evaluation that shows that our approach outperforms the state of the art. Project website: https://neural-3d-video.github.io/.
Tianye Li, Mira Slavcheva, Michael Zollhöfer, Simon Green, Christoph Lassner, Changil Kim 0001, Tanner Schmidt, Steven Lovegrove, Michael Goesele, Richard A. Newcombe, Zhaoyang Lv
CVPR8
2021 STaR: Self-Supervised Tracking and Reconstruction of Rigid Objects in Motion With Neural Rendering
abstract
We present STaR, a novel method that performs Self-supervised Tracking and Reconstruction of dynamic scenes with rigid motion from multi-view RGB videos without any manual annotation. Recent work has shown that neural networks are surprisingly effective at the task of compressing many views of a scene into a learned function which maps from a viewing ray to an observed radiance value via volume rendering. Unfortunately, these methods lose all their predictive power once any object in the scene has moved. In this work, we explicitly model rigid motion of objects in the context of neural representations of radiance fields. We show that without any additional human specified supervision, we can reconstruct a dynamic scene with a single rigid object in motion by simultaneously decomposing it into its two constituent parts and encoding each with its own neural representation. We achieve this by jointly optimizing the parameters of two neural radiance fields and a set of rigid poses which align the two fields at each frame. On both synthetic and real world datasets, we demonstrate that our method can render photorealistic novel views, where novelty is measured on both spatial and temporal axes. Our factored representation furthermore enables animation of unseen object motion.
Zhaoyang Lv, Tanner Schmidt, Steven Lovegrove
CVPR4
2020 FroDO: From Detections to 3D Objects
abstract
Object-oriented maps are important for scene understanding since they jointly capture geometry and semantics, allow individual instantiation and meaningful reasoning about objects. We introduce FroDO, a method for accurate 3D reconstruction of object instances from RGB video that infers their location, pose and shape in a coarse to fine manner. Key to FroDO is to embed object shapes in a novel learnt shape space that allows seamless switching between sparse point cloud and dense DeepSDF decoding. Given an input sequence of localized RGB frames, FroDO first aggregates 2D detections to instantiate a 3D bounding box per object. A shape code is regressed using an encoder network before optimizing shape and pose further under the learnt shape priors using sparse or dense shape representations. The optimization uses multi-view geometric, photometric and silhouette losses. We evaluate on real-world datasets, including Pix3D, Redwood-OS, and ScanNet, for single-view, multi-view, and multi-object reconstruction.
Martin Rünz, Kejie Li, Meng Tang 0001, Lingni Ma, Chen Kong, Tanner Schmidt, Ian D. Reid 0001, Lourdes Agapito, Julian Straub, Steven Lovegrove, Richard A. Newcombe
CVPR10
2020 Deep Local Shapes: Learning Local SDF Priors for Detailed 3D Reconstruction
Rohan Chabra, Jan Eric Lenssen, Eddy Ilg, Tanner Schmidt, Julian Straub, Steven Lovegrove, Richard A. Newcombe
ECCV (29)6
2020 Raycast Calibration for Augmented Reality HMDs with Off-Axis Reflective Combiners
abstract
Augmented reality overlays virtual objects on the real world. To do so, the head mounted display (HMD) needs to be calibrated to establish a mapping between 3D points in the real world with 2D pixels on display panels. This distortion is a high-dimensional function that also depends on pupil position and varifocal settings. We present Raycast calibration, an efficient approach to geometrically calibrate AR displays with off-axis reflective combiners. Our approach requires a small amount of data to estimate a compact, physics-based, and ray-traceable model of the HMD optics. We apply this technique to automatically calibrate an AR prototype with display, SLAM and eye-tracker, without user in the loop.
Qi Guo 0009, Huixuan Tang, Aaron Schmitz, Yang Lou, Alexander Fix, Steven Lovegrove, Hauke Strasdat
ICCP7
2019 DeepSDF: Learning Continuous Signed Distance Functions for Shape Representation
abstract
Computer graphics, 3D computer vision and robotics communities have produced multiple approaches to representing 3D geometry for rendering and reconstruction. These provide trade-offs across fidelity, efficiency and compression capabilities. In this work, we introduce DeepSDF, a learned continuous Signed Distance Function (SDF) representation of a class of shapes that enables high quality shape representation, interpolation and completion from partial and noisy 3D input data. DeepSDF, like its classical counterpart, represents a shape's surface by a continuous volumetric field: the magnitude of a point in the field represents the distance to the surface boundary and the sign indicates whether the region is inside (-) or outside (+) of the shape, hence our representation implicitly encodes a shape's boundary as the zero-level-set of the learned function while explicitly representing the classification of space as being part of the shapes interior or not. While classical SDF's both in analytical or discretized voxel form typically represent the surface of a single shape, DeepSDF can represent an entire class of shapes. Furthermore, we show state-of-the-art performance for learned 3D shape representation and completion while reducing the model size by an order of magnitude compared with previous work.
Jeong Joon Park, Peter R. Florence, Julian Straub, Richard A. Newcombe, Steven Lovegrove
CVPR5
2018 Reconstructing scenes with mirror and glass surfaces
abstract
Planar reflective surfaces such as glass and mirrors are notoriously hard to reconstruct for most current 3D scanning techniques. When treated naïvely, they introduce duplicate scene structures, effectively destroying the reconstruction altogether. Our key insight is that an easy to identify structure attached to the scanner---in our case an AprilTag---can yield reliable information about the existence and the geometry of glass and mirror surfaces in a scene. We introduce a fully automatic pipeline that allows us to reconstruct the geometry and extent of planar glass and mirror surfaces while being able to distinguish between the two. Furthermore, our system can automatically segment observations of multiple reflective surfaces in a scene based on their estimated planes and locations. In the proposed setup, minimal additional hardware is needed to create high-quality results. We demonstrate this using reconstructions of several scenes with a variety of real mirrors and glass.
Thomas Whelan, Michael Goesele, Steven Lovegrove, Julian Straub, Simon Green, Richard Szeliski, Steven Butterfield, Shobhit Verma, Richard A. Newcombe
ACM Trans. Graph.3
2015 A Spline-Based Trajectory Representation for Sensor Fusion and Rolling Shutter Cameras
Alonso Patron-Perez, Steven Lovegrove, Gabe Sibley
Int. J. Comput. Vis.2
2013 Spline Fusion: A continuous-time representation for visual-inertial fusion with application to rolling shutter cameras
abstract
This paper describes a general continuous-time framework for visual-inertial simultaneous localization and mapping and calibration. We show how to use a spline parameterization that closely matches the torque-minimal motion of the sensor. Compared to traditional discrete-time solutions, the continuous-time formulation is particularly useful for solving problems with high-frame rate sensors and multiple unsynchronized devices. We demonstrate the applicability of the method for multi-sensor visual-inertial SLAM and calibration by accurately establishing the relative pose and internal parameters of multiple unsynchronized devices. We also show the advantages of the approach through evaluation and uniform treatment of both global and rolling shutter cameras within visual and visual-inertial SLAM systems.
Steven Lovegrove, Alonso Patron-Perez, Gabe Sibley
BMVC1
2011 DTAM: Dense tracking and mapping in real-time
abstract
DTAM is a system for real-time camera tracking and reconstruction which relies not on feature extraction but dense, every pixel methods. As a single hand-held RGB camera flies over a static scene, we estimate detailed textured depth maps at selected keyframes to produce a surface patchwork with millions of vertices. We use the hundreds of images available in a video stream to improve the quality of a simple photometric data term, and minimise a global spatially regularised energy functional in a novel non-convex optimisation framework. Interleaved, we track the camera's 6DOF motion precisely by frame-rate whole image alignment against the entire dense model. Our algorithms are highly parallelisable throughout and DTAM achieves real-time performance using current commodity GPU hardware. We demonstrate that a dense model permits superior tracking performance under rapid motion compared to a state of the art method using features; and also show the additional usefulness of the dense model for real-time scene interaction in a physics-enhanced augmented reality application.
Richard A. Newcombe, Steven Lovegrove, Andrew J. Davison
ICCV2
2011 Accurate visual odometry from a rear parking camera
abstract
As an increasing number of automatic safety and navigation features are added to modern vehicles, the crucial job of providing real-time localisation is predominantly performed by a single sensor, GPS, despite its well-known failings, particularly in urban environments. Various attempts have been made to supplement GPS to improve localisation performance, but these usually require additional specialised and expensive sensors. Offering increased value to vehicle OEMs, we show that it is possible to use just the video stream from a rear parking camera to produce smooth and locally accurate visual odometry in real-time. We use an efficient whole image alignment approach based on ESM, taking account of both the difficulties and advantages of the fact that a parking camera views only the road surface directly behind a vehicle. Visual odometry is complementary to GPS in offering localisation information at 30 Hz which is smooth and highly accurate locally whilst GPS is course but offers absolute measurements. We demonstrate our system in a large scale experiment covering real urban driving. We also present real-time fusion of our visual estimation with automotive GPS to generate a commodity-cost localisation solution which is smooth, accurate and drift free in global coordinates.
Steven Lovegrove, Andrew J. Davison, Javier Ibañez-Guzmán
Intelligent Vehicles Symposium1
2010 Real-Time Spherical Mosaicing Using Whole Image Alignment
Steven Lovegrove, Andrew J. Davison
ECCV (3)1