Carl Yuheng Ren

dblp:119/1485 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
3since 2021 · last 2025
0009-0000-0942-2929ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
3D vision · 56% Video understanding and tracking · 25% Robot manipulation · 9%
Human-computer interaction and pervasive computing
1 paper
Wearable and physiological sensing · 100%
Computer graphics and multimedia
4 papers
Virtual and augmented reality · 72% Image and video processing · 19% Geometric modeling and processing · 9%

Topics — the 24 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d shape reconstruction
0.912025
Digital Twin Catalog: A Large-Scale Photorealistic 3D Object Digital Twin Dataset · CVPR 2025
Computer vision › Video understanding and tracking
activity recognition
0.912025
Reading Recognition in the Wild · NeurIPS 2025
Robotics › Robot manipulation
digital twin generation
0.912025
Digital Twin Catalog: A Large-Scale Photorealistic 3D Object Digital Twin Dataset · CVPR 2025
Computer vision › 3D vision
egocentric vision
0.912025
Reading Recognition in the Wild · NeurIPS 2025
Wearable and physiological sensing › wearable camera › egocentric vision
egocentric sensing
0.912025
Reading Recognition in the Wild · NeurIPS 2025
Computer vision › 3D vision › 3d object detection
3d object detection and tracking
0.712023
Aria Digital Twin: A New Benchmark Dataset for Egocentric 3D Machine Perception · ICCV 2023
Computer vision › 3D vision
3d scene reconstruction
0.712023
Aria Digital Twin: A New Benchmark Dataset for Egocentric 3D Machine Perception · ICCV 2023
Computer vision › 3D vision › 3d scene understanding
egocentric 3d perception
0.712023
Aria Digital Twin: A New Benchmark Dataset for Egocentric 3D Machine Perception · ICCV 2023
Computer vision › Video understanding and tracking › multi-object tracking
object detection and tracking
0.712023
Aria Digital Twin: A New Benchmark Dataset for Egocentric 3D Machine Perception · ICCV 2023
Computer vision › Segmentation and scene understanding
scene understanding
0.712023
Aria Digital Twin: A New Benchmark Dataset for Egocentric 3D Machine Perception · ICCV 2023
Computer vision › 3D vision
3d reconstruction
0.532015
Very High Frame Rate Volumetric Integration of Depth Images on Mobile Devices · IEEE Trans. Vis. Comput. Graph. 2015
STAR3D: Simultaneous Tracking and Reconstruction of 3D Objects Using RGB-D Data · ICCV 2013
Dense Reconstruction Using 3D Object Shape Priors · CVPR 2013
Computer vision › Video understanding and tracking
object tracking
0.532017
Real-Time Tracking of Single and Multiple Objects from Depth-Colour Imagery Using 3D Signed Distance Functions · Int. J. Comput. Vis. 2017
STAR3D: Simultaneous Tracking and Reconstruction of 3D Objects Using RGB-D Data · ICCV 2013
Regressing Local to Global Shape Properties for Online Segmentation and Tracking · Int. J. Comput. Vis. 2014
Computer vision › Video understanding and tracking
multi-object tracking
0.312017
Real-Time Tracking of Single and Multiple Objects from Depth-Colour Imagery Using 3D Signed Distance Functions · Int. J. Comput. Vis. 2017
Virtual and augmented reality › augmented reality display
AR glasses
0.312025
Digital Twin Catalog: A Large-Scale Photorealistic 3D Object Digital Twin Dataset · CVPR 2025
Virtual and augmented reality › tracking
egocentric motion capture
0.312025
Digital Twin Catalog: A Large-Scale Photorealistic 3D Object Digital Twin Dataset · CVPR 2025
Wearable and physiological sensing › wearable display
smart glasses
0.312025
Reading Recognition in the Wild · NeurIPS 2025
Image and video processing
image segmentation
0.212014
Regressing Local to Global Shape Properties for Online Segmentation and Tracking · Int. J. Comput. Vis. 2014
Computer vision › Video understanding and tracking › object tracking
3d object tracking
0.212013
STAR3D: Simultaneous Tracking and Reconstruction of 3D Objects Using RGB-D Data · ICCV 2013
Robotics › Robot navigation and mapping › SLAM
dense SLAM
0.212013
Dense Reconstruction Using 3D Object Shape Priors · CVPR 2013
Computer vision › 3D vision
object pose estimation
0.212013
Dense Reconstruction Using 3D Object Shape Priors · CVPR 2013
Computer vision › 3D vision › 3d reconstruction › multimodal 3d reconstruction
RGB-D reconstruction
0.212013
STAR3D: Simultaneous Tracking and Reconstruction of 3D Objects Using RGB-D Data · ICCV 2013
Robotics › Robot navigation and mapping
SLAM
0.212013
Dense Reconstruction Using 3D Object Shape Priors · CVPR 2013
Geometric modeling and processing › shape representation › implicit representation
signed distance function
0.112017
Real-Time Tracking of Single and Multiple Objects from Depth-Colour Imagery Using 3D Signed Distance Functions · Int. J. Comput. Vis. 2017
GPUs and heterogeneous computing › embedded GPU
mobile GPU
0.112015
Very High Frame Rate Volumetric Integration of Depth Images on Mobile Devices · IEEE Trans. Vis. Comput. Graph. 2015

Methods — techniques the papers use, named apart from their topics

transformer · 1.7neural reconstruction · 1.7inverse rendering · 1.7head pose · 1.7eye gaze · 1.7sim-to-real learning · 1.3image translation · 1.3probabilistic generative model · 0.6maximum a posteriori · 0.6IMU integration · 0.4ray casting · 0.2shape regression · 0.2
YearPublicationVenuePosition
2025 Digital Twin Catalog: A Large-Scale Photorealistic 3D Object Digital Twin Dataset
abstract
We introduce Digital Twin Catalog (DTC), a new large-scale photorealistic 3D object digital twin dataset. A digital twin of a 3D object is a highly detailed, virtually indistinguishable representation of a physical object, accurately capturing its shape, appearance, physical properties, and other attributes. Recent advances in neural-based 3D reconstruction and inverse rendering have significantly improved the quality of 3D object reconstruction. Despite these advancements, there remains a lack of a large-scale, digital twin quality real-world dataset and benchmark that can quantitatively assess and compare the performance of different reconstruction methods, as well as improve reconstruction quality through training or fine-tuning. Moreover, to democratize 3D digital twin creation, it is essential to integrate creation techniques with next-generation egocentric computing platforms, such as AR glasses. Currently, there is no dataset available to evaluate 3D object reconstruction using egocentric captured images. To address these gaps, the DTC dataset features 2,000 scanned digital twin-quality 3D objects, along with image sequences captured under different lighting conditions using DSLR cameras and egocentric AR glasses. This dataset establishes the first comprehensive real-world evaluation benchmark for 3D digital twin creation tasks, offering a robust foundation for comparing and improving existing reconstruction methods. The DTC dataset is already released at https://www.projectaria.com/datasets/dtc/ and we will also make the baseline evaluations open-source.
Zhao Dong 0001, Ka Chen, Zhaoyang Lv, Hong-Xing Yu, Yufeng Zhu, Stephen Tian, Zhengqin Li, Geordie Moffatt, Sean Christofferson, James Fort, Xiaqing Pan, Mingfei Yan, Jiajun Wu 0001, Carl Yuheng Ren, Richard A. Newcombe
CVPR16
2025 Reading Recognition in the Wild
abstract
To enable egocentric contextual AI in always-on smart glasses, it is crucial to be able to keep a record of the user's interactions with the world, including during reading. In this paper, we introduce a new task of reading recognition to determine when the user is reading. We first introduce the first-of-its-kind large-scale multimodal Reading in the Wild dataset, containing 100 hours of reading and non-reading videos in diverse and realistic scenarios. We then identify three modalities (egocentric RGB, eye gaze, head pose) that can be used to solve the task, and present a flexible transformer model that performs the task using these modalities, either individually or combined. We show that these modalities are relevant and complementary to the task, and investigate how to efficiently and effectively encode each modality. Additionally, we show the usefulness of this dataset towards classifying types of reading, extending current reading understanding studies conducted in constrained settings to larger scale, diversity and realism. Code, model, and data will be public.
Charig Yang, Samiul Alam, Shakhrul Iman Siam, Michael J. Proulx, Lambert Mathias, Kiran K. Somasundaram, Luis Pesqueira, James Fort, Sheroze Sheriffdeen, Omkar M. Parkhi, Carl Yuheng Ren, Mi Zhang 0002, Yuning Chai, Richard A. Newcombe, Hyo Jin Kim 0004
NeurIPS11
2023 Aria Digital Twin: A New Benchmark Dataset for Egocentric 3D Machine Perception
abstract
We introduce the Aria Digital Twin (ADT)1- an egocentric dataset captured using Aria glasses with extensive object, environment, and human level ground truth. This ADT release contains 200 sequences of real-world activities conducted by Aria wearers in two real indoor scenes with 398 object instances (344 stationary and 74 dynamic). Each sequence consists of: a) raw data of two monochrome camera streams, one RGB camera stream, two IMU streams; b) complete sensor calibration; c) ground truth data including continuous 6-degree-of-freedom (6DoF) poses of the Aria devices, object 6DoF poses, 3D eye gaze vectors, 3D human poses, 2D image segmentations, image depth maps; and d) photo-realistic synthetic renderings. To the best of our knowledge, there is no existing egocentric dataset with a level of accuracy, photo-realism and comprehensiveness comparable to ADT. By contributing ADT to the research community, our mission is to set a new standard for evaluation in the egocentric machine perception domain, which includes very challenging research problems such as 3D object detection and tracking, scene reconstruction and understanding, sim-to-real learning, human pose prediction - while also inspiring new machine perception tasks for augmented reality (AR) applications. To kick start exploration of the ADT research use cases, we evaluated several existing state-of-the-art methods for object detection, segmentation and image translation tasks that demonstrate the usefulness of ADT as a benchmarking dataset.
Xiaqing Pan, Nicholas Charron, Yongqian Yang, Scott Peters, Thomas Whelan, Chen Kong, Omkar M. Parkhi, Richard A. Newcombe, Carl Yuheng Ren
ICCV9
2017 Real-Time Tracking of Single and Multiple Objects from Depth-Colour Imagery Using 3D Signed Distance Functions
abstract
We describe a novel probabilistic framework for real-time tracking of multiple objects from combined depth-colour imagery. Object shape is represented implicitly using 3D signed distance functions. Probabilistic generative models based on these functions are developed to account for the observed RGB-D imagery, and tracking is posed as a maximum a posteriori problem. We present first a method suited to tracking a single rigid 3D object, and then generalise this to multiple objects by combining distance functions into a shape union in the frame of the camera. This second model accounts for similarity and proximity between objects, and leads to robust real-time tracking without recourse to bolt-on or ad-hoc collision detection.
Carl Yuheng Ren, Victor Adrian Prisacariu, Olaf Kähler, Ian D. Reid 0001, David William Murray 0001
Int. J. Comput. Vis.1
2015 Very High Frame Rate Volumetric Integration of Depth Images on Mobile Devices
abstract
Volumetric methods provide efficient, flexible and simple ways of integrating multiple depth images into a full 3D model. They provide dense and photorealistic 3D reconstructions, and parallelised implementations on GPUs achieve real-time performance on modern graphics hardware. To run such methods on mobile devices, providing users with freedom of movement and instantaneous reconstruction feedback, remains challenging however. In this paper we present a range of modifications to existing volumetric integration methods based on voxel block hashing, considerably improving their performance and making them applicable to tablet computer applications. We present (i) optimisations for the basic data structure, and its allocation and integration; (ii) a highly optimised raycasting pipeline; and (iii) extensions to the camera tracker to incorporate IMU data. In total, our system thus achieves frame rates up 47 Hz on a Nvidia Shield Tablet and 910 Hz on a Nvidia GTX Titan XGPU, or even beyond 1.1 kHz without visualisation.
Olaf Kähler, Victor Adrian Prisacariu, Carl Yuheng Ren, Xin Sun 0010, Philip Torr 0001, David William Murray 0001
IEEE Trans. Vis. Comput. Graph.3
2014 3D Tracking of Multiple Objects with Identical Appearance Using RGB-D Input
abstract
Most current approaches for 3D object tracking rely on distinctive object appearances. While several such trackers can be instantiated to track multiple objects independently, this not only neglects that objects should not occupy the same space in 3D, but also fails when objects have highly similar or identical appearances. In this paper we develop a probabilistic graphical model that accounts for similarity and proximity and leads to robust real-time tracking of multiple objects from RGB-D data, without recourse to bolton collision detection.
Carl Yuheng Ren, Victor Adrian Prisacariu, Olaf Kähler, Ian D. Reid 0001, David William Murray 0001
3DV1
2014 Regressing Local to Global Shape Properties for Online Segmentation and Tracking
Carl Yuheng Ren, Victor Adrian Prisacariu, Ian D. Reid 0001
Int. J. Comput. Vis.1
2013 Dense Reconstruction Using 3D Object Shape Priors
abstract
We propose a formulation of monocular SLAM which combines live dense reconstruction with shape priors-based 3D tracking and reconstruction. Current live dense SLAM approaches are limited to the reconstruction of visible surfaces. Moreover, most of them are based on the minimisation of a photo-consistency error, which usually makes them sensitive to specularities. In the 3D pose recovery literature, problems caused by imperfect and ambiguous image information have been dealt with by using prior shape knowledge. At the same time, the success of depth sensors has shown that combining joint image and depth information drastically increases the robustness of the classical monocular 3D tracking and 3D reconstruction approaches. In this work we link dense SLAM to 3D object pose and shape recovery. More specifically, we automatically augment our SLAM system with object specific identity, together with 6D pose and additional shape degrees of freedom for the object(s) of known class in the scene, combining image data and depth information for the pose and shape recovery. This leads to a system that allows for full scaled 3D reconstruction with the known object(s) segmented from the scene. The segmentation enhances the clarity, accuracy and completeness of the maps built by the dense SLAM system, while the dense 3D data aids the segmentation process, yielding faster and more reliable convergence than when using 2D image data alone.
Amaury Dame, Victor Adrian Prisacariu, Carl Yuheng Ren, Ian D. Reid 0001
CVPR3
2013 STAR3D: Simultaneous Tracking and Reconstruction of 3D Objects Using RGB-D Data
abstract
We introduce a probabilistic framework for simultaneous tracking and reconstruction of 3D rigid objects using an RGB-D camera. The tracking problem is handled using a bag-of-pixels representation and a back-projection scheme. Surface and background appearance models are learned online, leading to robust tracking in the presence of heavy occlusion and outliers. In both our tracking and reconstruction modules, the 3D object is implicitly embedded using a 3D level-set function. The framework is initialized with a simple shape primitive model (e.g. a sphere or a cube), and the real 3D object shape is tracked and reconstructed online. Unlike existing depth-based 3D reconstruction works, which either rely on calibrated/fixed camera set up or use the observed world map to track the depth camera, our framework can simultaneously track and reconstruct small moving objects. We use both qualitative and quantitative results to demonstrate the superior performance of both tracking and reconstruction of our method.
Carl Yuheng Ren, Victor Adrian Prisacariu, David William Murray 0001, Ian D. Reid 0001
ICCV1
2011 Regressing Local to Global Shape Properties for Online Segmentation and Tracking
abstract
We propose a regression based learning framework that learns a set of shapes online, which can then be used to recover occluded object shapes. We represent shapes using their 2D discrete cosine transforms, and the key insight we propose is to regress low frequency harmonics, which represent the global properties of the shape, from high frequency harmonics, that encode the details of the object's shape. We learn the regression model using Locally Weighted Projection Regression (LWPR) which expedites online, incremental learning. After sufficient observation of a set of unoccluded shapes, the learned model can detect occlusion and recover the full shapes from the occluded ones. We demonstrate the ideas using a level-set based tracking system that provides shape and pose, however, the framework could be embedded in any segmentation-based tracking system. Our experiments demonstrate the efficacy of the method on a variety of objects using both real data and artificial data.
Carl Yuheng Ren, Victor Adrian Prisacariu, Ian D. Reid 0001
BMVC1