VLDB 2026 Research / reviewers in the wild / expert
Steven Lovegrove
dblp:03/8535 · also Steven J. Lovegrove
· DBLP profile ↗
13ranked-venue papers
3as first author
3since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
3D vision · 88% Speech recognition and synthesis · 10% Robot navigation and mapping · 2% | |
| Computer graphics and multimedia
7 papers |
Geometric modeling and processing · 36% Multimedia analysis and retrieval · 23% Computer animation and physical simulation · 20% |
Topics — the 27 heaviest of 30, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
neural radiance field |
1.1 | 2 | 2022 | Neural 3D Video Synthesis from Multi-view Video · CVPR 2022 STaR: Self-Supervised Tracking and Reconstruction of Rigid Objects in Motion With Neural Rendering · CVPR 2021 |
Computer vision › 3D vision › 3d shape representation › implicit surface representation
signed distance function |
0.8 | 2 | 2020 | Deep Local Shapes: Learning Local SDF Priors for Detailed 3D Reconstruction · ECCV (29) 2020 DeepSDF: Learning Continuous Signed Distance Functions for Shape Representation · CVPR 2019 |
Computer vision › 3D vision
3d reconstruction |
0.8 | 2 | 2020 | Deep Local Shapes: Learning Local SDF Priors for Detailed 3D Reconstruction · ECCV (29) 2020 Reconstructing scenes with mirror and glass surfaces · ACM Trans. Graph. 2018 |
Multimedia analysis and retrieval › multimedia analysis
multimodal conversation analysis |
0.7 | 1 | 2023 | EgoCom: A Multi-Person Multi-Modal Egocentric Communications Dataset · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Computer vision › 3D vision › neural radiance field
dynamic neural radiance field |
0.6 | 1 | 2022 | Neural 3D Video Synthesis from Multi-view Video · CVPR 2022 |
Computer vision › 3D vision › novel view synthesis
multi-view video generation |
0.6 | 1 | 2022 | Neural 3D Video Synthesis from Multi-view Video · CVPR 2022 |
Computer vision › 3D vision
novel view synthesis |
0.6 | 1 | 2022 | Neural 3D Video Synthesis from Multi-view Video · CVPR 2022 |
Computer animation and physical simulation › motion synthesis
motion interpolation |
0.6 | 1 | 2022 | Neural 3D Video Synthesis from Multi-view Video · CVPR 2022 |
Computer vision › 3D vision › 3d scene reconstruction
dynamic scene reconstruction |
0.5 | 1 | 2021 | STaR: Self-Supervised Tracking and Reconstruction of Rigid Objects in Motion With Neural Rendering · CVPR 2021 |
Computer vision › 3D vision › object pose estimation
object pose tracking |
0.5 | 1 | 2021 | STaR: Self-Supervised Tracking and Reconstruction of Rigid Objects in Motion With Neural Rendering · CVPR 2021 |
Computer vision › 3D vision › motion estimation
rigid motion estimation |
0.5 | 1 | 2021 | STaR: Self-Supervised Tracking and Reconstruction of Rigid Objects in Motion With Neural Rendering · CVPR 2021 |
Computer vision › 3D vision › pose estimation
shape and pose estimation |
0.4 | 1 | 2020 | FroDO: From Detections to 3D Objects · CVPR 2020 |
Geometric modeling and processing
3d reconstruction |
0.4 | 1 | 2020 | FroDO: From Detections to 3D Objects · CVPR 2020 |
Computer vision › 3D vision
3d shape representation |
0.4 | 1 | 2019 | DeepSDF: Learning Continuous Signed Distance Functions for Shape Representation · CVPR 2019 |
Geometric modeling and processing › surface reconstruction
shape reconstruction |
0.4 | 1 | 2019 | DeepSDF: Learning Continuous Signed Distance Functions for Shape Representation · CVPR 2019 |
Computer vision › 3D vision › 3d reconstruction › surface reconstruction
planar surface reconstruction |
0.3 | 1 | 2018 | Reconstructing scenes with mirror and glass surfaces · ACM Trans. Graph. 2018 |
Computer vision › 3D vision › 3d reconstruction › non-lambertian surface reconstruction
reflective surface reconstruction |
0.3 | 1 | 2018 | Reconstructing scenes with mirror and glass surfaces · ACM Trans. Graph. 2018 |
Computational photography and imaging › image signal processing
rolling shutter correction |
0.2 | 1 | 2015 | A Spline-Based Trajectory Representation for Sensor Fusion and Rolling Shutter Cameras · Int. J. Comput. Vis. 2015 |
Geometric modeling and processing
trajectory representation |
0.2 | 1 | 2015 | A Spline-Based Trajectory Representation for Sensor Fusion and Rolling Shutter Cameras · Int. J. Comput. Vis. 2015 |
Natural language and speech › Speech recognition and synthesis › spoken dialogue model
turn-taking prediction |
0.2 | 1 | 2023 | EgoCom: A Multi-Person Multi-Modal Egocentric Communications Dataset · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Rendering
novel view synthesis |
0.1 | 1 | 2021 | STaR: Self-Supervised Tracking and Reconstruction of Rigid Objects in Motion With Neural Rendering · CVPR 2021 |
Robotics › Robot navigation and mapping › SLAM › visual SLAM › dense visual SLAM
dense tracking and mapping |
0.1 | 1 | 2011 | DTAM: Dense tracking and mapping in real-time · ICCV 2011 |
Computer vision › 3D vision › 3d reconstruction
real-time reconstruction |
0.1 | 1 | 2011 | DTAM: Dense tracking and mapping in real-time · ICCV 2011 |
Computational photography and imaging
image stitching |
0.1 | 1 | 2010 | Real-Time Spherical Mosaicing Using Whole Image Alignment · ECCV (3) 2010 |
Computational photography and imaging › image stitching
spherical mosaicing |
0.1 | 1 | 2010 | Real-Time Spherical Mosaicing Using Whole Image Alignment · ECCV (3) 2010 |
Robotics › Robot navigation and mapping
sensor fusion |
0.1 | 1 | 2015 | A Spline-Based Trajectory Representation for Sensor Fusion and Rolling Shutter Cameras · Int. J. Comput. Vis. 2015 |
Computer vision › 3D vision
image registration |
0.0 | 1 | 2010 | Real-Time Spherical Mosaicing Using Whole Image Alignment · ECCV (3) 2010 |
Methods — techniques the papers use, named apart from their topics
speech-to-text · 1.3bayesian modeling · 1.3time-conditioned neural radiance field · 1.1ray importance sampling · 1.1hierarchical training · 1.1volume rendering · 1.0self-supervised learning · 1.0joint optimization · 1.0encoder network · 0.9deep signed distance function · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | EgoCom: A Multi-Person Multi-Modal Egocentric Communications DatasetabstractMulti-modal datasets in artificial intelligence (AI) often capture a third-person perspective, but our embodied human intelligence evolved with sensory input from the egocentric, first-person perspective. Towards embodied AI, we introduce the Egocentric Communications (EgoCom) dataset to advance the state-of-the-art in conversational AI, natural language, audio speech analysis, computer vision, and machine learning. EgoCom is a first-of-its-kind natural conversations dataset containing multi-modal human communication data captured simultaneously from the participants' egocentric perspectives. EgoCom includes 38.5 hours of synchronized embodied stereo audio, egocentric video with 240,000 ground-truth, time-stamped word-level transcriptions and speaker labels from 34 diverse speakers. We study baseline performance on two novel applications that benefit from embodied data: (1) predicting turn-taking in conversations and (2) multi-speaker transcription. For (1), we investigate Bayesian baselines to predict turn-taking within 5 percent of human performance. For (2), we use simultaneous egocentric capture to combine Google speech-to-text outputs, improving global transcription by 79 percent relative to a single perspective. Both applications exploit EgoCom's synchronous multi-perspective data to augment performance of embodied AI tasks. Curtis G. Northcutt, Shengxin Zha, Steven Lovegrove, Richard A. Newcombe |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Neural 3D Video Synthesis from Multi-view VideoabstractWe propose a novel approach for 3D video synthesis that is able to represent multi-view video recordings of a dynamic real-world scene in a compact, yet expressive representation that enables high-quality view synthesis and motion interpolation. Our approach takes the high quality and compactness of static neural radiance fields in a new direction: to a model-free, dynamic setting. At the core of our approach is a novel time-conditioned neural radiance field that represents scene dynamics using a set of compact latent codes. We are able to significantly boost the training speed and perceptual quality of the generated imagery by a novel hierarchical training scheme in combination with ray importance sampling. Our learned representation is highly compact and able to represent a 10 second 30 FPS multi-view video recording by 18 cameras with a model size of only 28MB. We demonstrate that our method can render high-fidelity wide-angle novel views at over 1K resolution, even for complex and dynamic scenes. We perform an extensive qualitative and quantitative evaluation that shows that our approach outperforms the state of the art. Project website: https://neural-3d-video.github.io/. Tianye Li, Mira Slavcheva, Michael Zollhöfer, Simon Green, Christoph Lassner, Changil Kim 0001, Tanner Schmidt, Steven Lovegrove, Michael Goesele, Richard A. Newcombe, Zhaoyang Lv |
CVPR | 8 |
| 2021 | STaR: Self-Supervised Tracking and Reconstruction of Rigid Objects in Motion With Neural RenderingabstractWe present STaR, a novel method that performs Self-supervised Tracking and Reconstruction of dynamic scenes with rigid motion from multi-view RGB videos without any manual annotation. Recent work has shown that neural networks are surprisingly effective at the task of compressing many views of a scene into a learned function which maps from a viewing ray to an observed radiance value via volume rendering. Unfortunately, these methods lose all their predictive power once any object in the scene has moved. In this work, we explicitly model rigid motion of objects in the context of neural representations of radiance fields. We show that without any additional human specified supervision, we can reconstruct a dynamic scene with a single rigid object in motion by simultaneously decomposing it into its two constituent parts and encoding each with its own neural representation. We achieve this by jointly optimizing the parameters of two neural radiance fields and a set of rigid poses which align the two fields at each frame. On both synthetic and real world datasets, we demonstrate that our method can render photorealistic novel views, where novelty is measured on both spatial and temporal axes. Our factored representation furthermore enables animation of unseen object motion. Zhaoyang Lv, Tanner Schmidt, Steven Lovegrove |
CVPR | 4 |
| 2020 | FroDO: From Detections to 3D ObjectsabstractObject-oriented maps are important for scene understanding since they jointly capture geometry and semantics, allow individual instantiation and meaningful reasoning about objects. We introduce FroDO, a method for accurate 3D reconstruction of object instances from RGB video that infers their location, pose and shape in a coarse to fine manner. Key to FroDO is to embed object shapes in a novel learnt shape space that allows seamless switching between sparse point cloud and dense DeepSDF decoding. Given an input sequence of localized RGB frames, FroDO first aggregates 2D detections to instantiate a 3D bounding box per object. A shape code is regressed using an encoder network before optimizing shape and pose further under the learnt shape priors using sparse or dense shape representations. The optimization uses multi-view geometric, photometric and silhouette losses. We evaluate on real-world datasets, including Pix3D, Redwood-OS, and ScanNet, for single-view, multi-view, and multi-object reconstruction. Martin Rünz, Kejie Li, Meng Tang 0001, Lingni Ma, Chen Kong, Tanner Schmidt, Ian D. Reid 0001, Lourdes Agapito, Julian Straub, Steven Lovegrove, Richard A. Newcombe |
CVPR | 10 |
| 2020 | Deep Local Shapes: Learning Local SDF Priors for Detailed 3D Reconstruction
Rohan Chabra, Jan Eric Lenssen, Eddy Ilg, Tanner Schmidt, Julian Straub, Steven Lovegrove, Richard A. Newcombe |
ECCV (29) | 6 |
| 2020 | Raycast Calibration for Augmented Reality HMDs with Off-Axis Reflective CombinersabstractAugmented reality overlays virtual objects on the real world. To do so, the head mounted display (HMD) needs to be calibrated to establish a mapping between 3D points in the real world with 2D pixels on display panels. This distortion is a high-dimensional function that also depends on pupil position and varifocal settings. We present Raycast calibration, an efficient approach to geometrically calibrate AR displays with off-axis reflective combiners. Our approach requires a small amount of data to estimate a compact, physics-based, and ray-traceable model of the HMD optics. We apply this technique to automatically calibrate an AR prototype with display, SLAM and eye-tracker, without user in the loop. Qi Guo 0009, Huixuan Tang, Aaron Schmitz, Yang Lou, Alexander Fix, Steven Lovegrove, Hauke Strasdat |
ICCP | 7 |
| 2019 | DeepSDF: Learning Continuous Signed Distance Functions for Shape RepresentationabstractComputer graphics, 3D computer vision and robotics communities have produced multiple approaches to representing 3D geometry for rendering and reconstruction. These provide trade-offs across fidelity, efficiency and compression capabilities. In this work, we introduce DeepSDF, a learned continuous Signed Distance Function (SDF) representation of a class of shapes that enables high quality shape representation, interpolation and completion from partial and noisy 3D input data. DeepSDF, like its classical counterpart, represents a shape's surface by a continuous volumetric field: the magnitude of a point in the field represents the distance to the surface boundary and the sign indicates whether the region is inside (-) or outside (+) of the shape, hence our representation implicitly encodes a shape's boundary as the zero-level-set of the learned function while explicitly representing the classification of space as being part of the shapes interior or not. While classical SDF's both in analytical or discretized voxel form typically represent the surface of a single shape, DeepSDF can represent an entire class of shapes. Furthermore, we show state-of-the-art performance for learned 3D shape representation and completion while reducing the model size by an order of magnitude compared with previous work. Jeong Joon Park, Peter R. Florence, Julian Straub, Richard A. Newcombe, Steven Lovegrove |
CVPR | 5 |
| 2018 | Reconstructing scenes with mirror and glass surfacesabstractPlanar reflective surfaces such as glass and mirrors are notoriously hard to reconstruct for most current 3D scanning techniques. When treated naïvely, they introduce duplicate scene structures, effectively destroying the reconstruction altogether. Our key insight is that an easy to identify structure attached to the scanner---in our case an AprilTag---can yield reliable information about the existence and the geometry of glass and mirror surfaces in a scene. We introduce a fully automatic pipeline that allows us to reconstruct the geometry and extent of planar glass and mirror surfaces while being able to distinguish between the two. Furthermore, our system can automatically segment observations of multiple reflective surfaces in a scene based on their estimated planes and locations. In the proposed setup, minimal additional hardware is needed to create high-quality results. We demonstrate this using reconstructions of several scenes with a variety of real mirrors and glass. Thomas Whelan, Michael Goesele, Steven Lovegrove, Julian Straub, Simon Green, Richard Szeliski, Steven Butterfield, Shobhit Verma, Richard A. Newcombe |
ACM Trans. Graph. | 3 |
| 2015 | A Spline-Based Trajectory Representation for Sensor Fusion and Rolling Shutter Cameras
Alonso Patron-Perez, Steven Lovegrove, Gabe Sibley |
Int. J. Comput. Vis. | 2 |
| 2013 | Spline Fusion: A continuous-time representation for visual-inertial fusion with application to rolling shutter camerasabstractThis paper describes a general continuous-time framework for visual-inertial simultaneous localization and mapping and calibration. We show how to use a spline parameterization that closely matches the torque-minimal motion of the sensor. Compared to traditional discrete-time solutions, the continuous-time formulation is particularly useful for solving problems with high-frame rate sensors and multiple unsynchronized devices. We demonstrate the applicability of the method for multi-sensor visual-inertial SLAM and calibration by accurately establishing the relative pose and internal parameters of multiple unsynchronized devices. We also show the advantages of the approach through evaluation and uniform treatment of both global and rolling shutter cameras within visual and visual-inertial SLAM systems. Steven Lovegrove, Alonso Patron-Perez, Gabe Sibley |
BMVC | 1 |
| 2011 | DTAM: Dense tracking and mapping in real-timeabstractDTAM is a system for real-time camera tracking and reconstruction which relies not on feature extraction but dense, every pixel methods. As a single hand-held RGB camera flies over a static scene, we estimate detailed textured depth maps at selected keyframes to produce a surface patchwork with millions of vertices. We use the hundreds of images available in a video stream to improve the quality of a simple photometric data term, and minimise a global spatially regularised energy functional in a novel non-convex optimisation framework. Interleaved, we track the camera's 6DOF motion precisely by frame-rate whole image alignment against the entire dense model. Our algorithms are highly parallelisable throughout and DTAM achieves real-time performance using current commodity GPU hardware. We demonstrate that a dense model permits superior tracking performance under rapid motion compared to a state of the art method using features; and also show the additional usefulness of the dense model for real-time scene interaction in a physics-enhanced augmented reality application. Richard A. Newcombe, Steven Lovegrove, Andrew J. Davison |
ICCV | 2 |
| 2011 | Accurate visual odometry from a rear parking cameraabstractAs an increasing number of automatic safety and navigation features are added to modern vehicles, the crucial job of providing real-time localisation is predominantly performed by a single sensor, GPS, despite its well-known failings, particularly in urban environments. Various attempts have been made to supplement GPS to improve localisation performance, but these usually require additional specialised and expensive sensors. Offering increased value to vehicle OEMs, we show that it is possible to use just the video stream from a rear parking camera to produce smooth and locally accurate visual odometry in real-time. We use an efficient whole image alignment approach based on ESM, taking account of both the difficulties and advantages of the fact that a parking camera views only the road surface directly behind a vehicle. Visual odometry is complementary to GPS in offering localisation information at 30 Hz which is smooth and highly accurate locally whilst GPS is course but offers absolute measurements. We demonstrate our system in a large scale experiment covering real urban driving. We also present real-time fusion of our visual estimation with automotive GPS to generate a commodity-cost localisation solution which is smooth, accurate and drift free in global coordinates. Steven Lovegrove, Andrew J. Davison, Javier Ibañez-Guzmán |
Intelligent Vehicles Symposium | 1 |
| 2010 | Real-Time Spherical Mosaicing Using Whole Image Alignment
Steven Lovegrove, Andrew J. Davison |
ECCV (3) | 1 |