Lingni Ma

dblp:126/4514 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
12since 2021 · last 2025
0009-0002-3698-0787ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 2 first-author · 12 since 2021Artificial intelligence and machine learning · 14 · 3 first-author · 9 since 2021Systems, architecture and hardware · 3 · 2 first-author
YearPublicationVenuePosition
2025 HMD2: Environment-Aware Motion Generation from Single Egocentric Head-Mounted Device
abstract
This paper investigates the generation of realistic full-body human motion using a single head-mounted device with an outward-facing color camera and the ability to perform visual SLAM. To address the ambiguity of this setup, we present HMD2, a novel system that balances motion reconstruction and generation. From a reconstruction stand-point, it aims to maximally utilize the camera streams to produce both analytical and learned features, including head motion, SLAM point cloud, and image embeddings. On the generative front, HMD2 employs a multi-modal conditional motion diffusion model with a Transformer back-bone to maintain temporal coherence of generated motions, and utilizes autoregressive inpainting to facilitate online motion inference with minimal latency (0.17 seconds). We show that our system provides an effective and robust so-lution that scales to a diverse dataset of over 200 hours of motion in complex indoor and outdoor environments.
Vladimir Guzov, Yifeng Jiang 0002, Fangzhou Hong, Gerard Pons-Moll, Richard A. Newcombe, C. Karen Liu, Yuting Ye, Lingni Ma
3DV8
2025 EgoLM: Multi-Modal Language Model of Egocentric Motions
abstract
As wearable devices become more prevalent, understanding the user’s motion is crucial for improving contextual AI systems. We introduce EgoLM, a versatile framework designed for egocentric motion understanding using multi-modal data. EgoLM integrates the rich contextual information from egocentric videos and motion sensors afforded by wearable devices. It also combines dense supervision signals from motion and language, leveraging the vast knowledge encoded in pre-trained large language models (LLMs). EgoLM models the joint distribution of egocentric motions and natural language using LLMs, conditioned on observations from egocentric videos and motion sensors. It unifies a range of motion understanding tasks, including motion narration from video or motion data, as well as motion generation from text or sparse sensor data. Unique to wearable devices, it also enables a novel task to generate text descriptions from sparse sensors. Through extensive experiments, we validate the effectiveness of EgoLM in addressing the challenges of under-constrained egocentric motion learning, and demonstrate its capability as a generalist model through a variety of applications. Project page: https://hongfz16.github.io/projects/EgoLM.
Fangzhou Hong, Vladimir Guzov, Hyo Jin Kim 0004, Yuting Ye, Richard A. Newcombe, Ziwei Liu 0002, Lingni Ma
CVPR7
2024 Nymeria: A Massive Collection of Multimodal Egocentric Daily Motion in the Wild
Lingni Ma, Yuting Ye, Fangzhou Hong, Vladimir Guzov, Yifeng Jiang 0002, Rowan Postyeni, Luis Pesqueira, Alexander Gamino, Vijay Baiyya, Hyo Jin Kim 0004, Kevin Bailey, David Soriano Fosas, C. Karen Liu, Ziwei Liu 0002, Jakob J. Engel, Renzo De Nardi, Richard A. Newcombe
ECCV (24)1
2024 FoundPose: Unseen Object Pose Estimation with Foundation Features
Evin Pinar Örnek, Yann Labbé, Bugra Tekin, Lingni Ma, Cem Keskin, Christian Forster, Tomas Hodan
ECCV (26)4
2024 DivaTrack: Diverse Bodies and Motions from Acceleration-Enhanced Three-Point Trackers
abstract
Abstract Full‐body avatar presence is important for immersive social and environmental interactions in digital reality. However, current devices only provide three six degrees of freedom (DOF) poses from the headset and two controllers (i.e. three‐point trackers). Because it is a highly under‐constrained problem, inferring full‐body pose from these inputs is challenging, especially when supporting the full range of body proportions and use cases represented by the general population. In this paper, we propose a deep learning framework, DivaTrack, which outperforms existing methods when applied to diverse body sizes and activities. We augment the sparse three‐point inputs with linear accelerations from Inertial Measurement Units (IMU) to improve foot contact prediction. We then condition the otherwise ambiguous lower‐body pose with the predictions of foot contact and upper‐body pose in a two‐stage model. We further stabilize the inferred full‐body pose in a wide range of configurations by learning to blend predictions that are computed in two reference frames, each of which is designed for different types of motions. We demonstrate the effectiveness of our design on a large dataset that captures 22 subjects performing challenging locomotion for three‐point tracking, including lunges, hula‐hooping, and sitting. As shown in a live demo using the Meta VR headset and Xsens IMUs, our method runs in real‐time while accurately tracking a user's motion when they perform a diverse set of movements.
Dongseok Yang, Jiho Kang, Lingni Ma, Joseph D. Greer, Yuting Ye, Sung-Hee Lee
Comput. Graph. Forum3
2023 In-Hand 3D Object Scanning from an RGB Sequence
abstract
We propose a method for in-hand 3D scanning of an unknown object with a monocular camera. Our method relies on a neural implicit surface representation that captures both the geometry and the appearance of the object, however, by contrast with most NeRF-based methods, we do not assume that the camera-object relative poses are known. Instead, we simultaneously optimize both the object shape and the pose trajectory. As direct optimization over all shape and pose parameters is prone to fail without coarse-level initialization, we propose an incremental approach that starts by splitting the sequence into carefully selected overlapping segments within which the optimization is likely to succeed. We reconstruct the object shape and track its poses independently within each segment, then merge all the segments before performing a global optimization. We show that our method is able to reconstruct the shape and color of both textured and challenging texture-less objects, outperforms classical methods that rely only on appearance features, and that its performance is close to recent methods that assume known camera poses.
Shreyas Hampali, Tomas Hodan, Luan Tran, Lingni Ma, Cem Keskin, Vincent Lepetit
CVPR4
2023 EgoHumans: An Egocentric 3D Multi-Human Benchmark
abstract
We present EgoHumans, a new multi-view multi-human video benchmark to advance the state-of-the-art of egocentric human 3D pose estimation and tracking. Existing ego-centric benchmarks either capture single subject or indoor-only scenarios, which limit the generalization of computer vision algorithms for real-world applications. We propose a novel 3D capture setup to construct a comprehensive ego-centric multi-human benchmark in the wild with annotations to support diverse tasks such as human detection, tracking, 2D/3D pose estimation, and mesh recovery. We leverage consumer-grade wearable camera-equipped glasses for the egocentric view, which enables us to capture dynamic activities like playing tennis, fencing, volleyball, etc. Furthermore, our multi-view setup generates accurate 3D ground truth even under severe or complete occlusion. The dataset consists of more than 125k egocentric images, spanning diverse scenes with a particular focus on challenging and unchoreographed multi-human activities and fast-moving egocentric views. We rigorously evaluate existing state-of-the-art methods and highlight their limitations in the egocentric scenario, specifically on multi-human tracking. To address such limitations, we propose EgoFormer, a novel approach with a multi-stream transformer architecture and explicit 3D spatial reasoning to estimate and track the human pose. EgoFormer significantly outperforms prior art by 13.6% IDF1 on the EgoHumans dataset.
Rawal Khirodkar, Aayush Bansal, Lingni Ma, Richard A. Newcombe, Minh Vo, Kris Makoto Kitani
ICCV3
2023 Snipper: A Spatiotemporal Transformer for Simultaneous Multi-Person 3D Pose Estimation Tracking and Forecasting on a Video Snippet
abstract
Multi-person pose understanding from RGB videos involves three complex tasks: pose estimation, tracking and motion forecasting. Intuitively, accurate multi-person pose estimation facilitates robust tracking, and robust tracking builds crucial history for correct motion forecasting. Most existing works either focus on a single task or employ multi-stage approaches to solving multiple tasks separately, which tends to make sub-optimal decision at each stage and also fail to exploit correlations among the three tasks. In this paper, we propose Snipper, a unified framework to perform multi-person 3D pose estimation, tracking, and motion forecasting simultaneously in a single stage. We propose an efficient yet powerful deformable attention mechanism to aggregate spatiotemporal information from the video snippet. Building upon this deformable attention, a video transformer is learned to encode the spatiotemporal features from the multi-frame snippet and to decode informative pose features for multi-person pose queries. Finally, these pose queries are regressed to predict multi-person pose trajectories and future motions in a single shot. In the experiments, we show the effectiveness of Snipper on three challenging public datasets where our generic model rivals specialized state-of-art baselines for pose estimation, tracking, and forecasting. Code is available athttps://github.com/JimmyZou/Snipper.
Shihao Zou, Yuanlu Xu, Chao Li 0021, Lingni Ma, Li Cheng 0001, Minh Vo
IEEE Trans. Circuits Syst. Video Technol.4
2022 LISA: Learning Implicit Shape and Appearance of Hands
abstract
This paper proposes a do-it-all neural model of human hands, named LISA. The model can capture accurate hand shape and appearance, generalize to arbitrary hand sub-jects, provide dense surface correspondences, be reconstructed from images in the wild, and can be easily an-imated. We train LISA by minimizing the shape and appearance losses on a large set of multi-view RGB image se-quences annotated with coarse 3D poses of the hand skele-ton. For a 3D point in the local hand coordinates, our model predicts the color and the signed distance with respect to each hand bone independently, and then combines the per-bone predictions using the predicted skinning weights. The shape, color, and pose representations are disentangled by design, enabling fine control of the selected hand param-eters. We experimentally demonstrate that LISA can ac-curately reconstruct a dynamic hand from monocular or multi-view sequences, achieving a noticeably higher qual-ity of reconstructed hand shapes compared to baseline approaches. Project page: https://www.iri.upc.edu/people/ecorona/lisa/.
Enric Corona, Tomas Hodan, Minh Vo, Francesc Moreno-Noguer, Chris Sweeney, Richard A. Newcombe, Lingni Ma
CVPR7
2022 Self-supervised Neural Articulated Shape and Appearance Models
abstract
Learning geometry, motion, and appearance priors of object classes is important for the solution of a large variety of computer vision problems. While the majority of approaches has focused on static objects, dynamic objects, especially with controllable articulation, are less explored. We propose a novel approach for learning a representation of the geometry, appearance, and motion of a class of articulated objects given only a set of color images as input. In a self-supervised manner, our novel representation learns shape, appearance, and articulation codes that enable independent control of these semantic dimensions. Our model is trained end-to-end without requiring any articulation annotations. Experiments show that our approach performs well for different joint types, such as revolute and prismatic joints, as well as different combinations of these joints. Compared to state of the art that uses direct 3D supervision and does not output appearance, we recover more faithful geometry and appearance from 2D observations only. In addition, our representation enables a large variety of applications, such as few-shot reconstruction, the generation of novel articulations, and novel view-synthesis. Project page: https://weify627.github.io/nasam/.
Fangyin Wei, Rohan Chabra, Lingni Ma, Christoph Lassner, Michael Zollhöfer, Szymon Rusinkiewicz, Chris Sweeney, Richard A. Newcombe, Mira Slavcheva
CVPR3
2022 Neural Correspondence Field for Object Pose Estimation
Lin Huang 0004, Tomas Hodan, Lingni Ma, Linguang Zhang, Luan Tran, Christopher D. Twigg, Po-Chen Wu, Junsong Yuan 0001, Cem Keskin, Robert Wang 0002
ECCV (10)3
2022 Egocentric Activity Recognition and Localization on a 3D Map
Miao Liu 0007, Lingni Ma, Kiran K. Somasundaram, Yin Li 0003, Kristen Grauman, James M. Rehg
ECCV (13)2
2020 FroDO: From Detections to 3D Objects
abstract
Object-oriented maps are important for scene understanding since they jointly capture geometry and semantics, allow individual instantiation and meaningful reasoning about objects. We introduce FroDO, a method for accurate 3D reconstruction of object instances from RGB video that infers their location, pose and shape in a coarse to fine manner. Key to FroDO is to embed object shapes in a novel learnt shape space that allows seamless switching between sparse point cloud and dense DeepSDF decoding. Given an input sequence of localized RGB frames, FroDO first aggregates 2D detections to instantiate a 3D bounding box per object. A shape code is regressed using an encoder network before optimizing shape and pose further under the learnt shape priors using sparse or dense shape representations. The optimization uses multi-view geometric, photometric and silhouette losses. We evaluate on real-world datasets, including Pix3D, Redwood-OS, and ScanNet, for single-view, multi-view, and multi-object reconstruction.
Martin Rünz, Kejie Li, Meng Tang 0001, Lingni Ma, Chen Kong, Tanner Schmidt, Ian D. Reid 0001, Lourdes Agapito, Julian Straub, Steven Lovegrove, Richard A. Newcombe
CVPR4
2017 De-noising, stabilizing and completing 3D reconstructions on-the-go using plane priors
abstract
Creating 3D maps on robots and other mobile devices has become a reality in recent years. Online 3D reconstruction enables many exciting applications in robotics and AR/VR gaming. However, the reconstructions are noisy and generally incomplete. Moreover, during online reconstruction, the surface changes with every newly integrated depth image which poses a significant challenge for physics engines and path planning algorithms. This paper presents a novel, fast and robust method for obtaining and using information about planar surfaces, such as walls, floors, and ceilings as a stage in 3D reconstruction based on Signed Distance Fields (SDFs). Our algorithm recovers clean and accurate surfaces, reduces the movement of individual mesh vertices caused by noise during online reconstruction and fills in the occluded and unobserved regions. We implemented and evaluated two different strategies to generate plane candidates and two strategies for merging them. Our implementation is optimized to run in real-time on mobile devices such as the Tango tablet. In an extensive set of experiments, we validated that our approach works well in a large number of natural environments despite the presence of significant amount of occlusion, clutter and noise, which occur frequently. We further show that plane fitting enables in many cases a meaningful semantic segmentation of real-world scenes.
Maksym Dzitsiuk, Jürgen Sturm, Robert Maier 0001, Lingni Ma, Daniel Cremers
ICRA4
2017 Multi-view deep learning for consistent semantic mapping with RGB-D cameras
abstract
Visual scene understanding is an important capability that enables robots to purposefully act in their environment. In this paper, we propose a novel deep neural network approach to predict semantic segmentation from RGB-D sequences. The key innovation is to train our network to predict multi-view consistent semantics in a self-supervised way. At test time, its semantics predictions can be fused more consistently in semantic keyframe maps than predictions of a network trained on individual views. We base our network architecture on a recent single-view deep learning approach to RGB and depth fusion for semantic object-class segmentation and enhance it with multi-scale loss minimization. We obtain the camera trajectory using RGB-D SLAM and warp the predictions of RGB-D images into ground-truth annotated frames in order to enforce multi-view consistency during training. At test time, predictions from multiple views are fused into keyframes. We propose and analyze several methods for enforcing multi-view consistency during training and testing. We evaluate the benefit of multi-view consistency training and demonstrate that pooling of deep features and fusion over multiple views outperforms single-view baselines on the NYUDv2 benchmark for semantic segmentation. Our end-to-end trained network achieves state-of-the-art performance on the NYUDv2 dataset in single-view segmentation as well as multi-view semantic fusion.
Lingni Ma, Jörg Stückler, Christian Kerl, Daniel Cremers
IROS1
2016 FuseNet: Incorporating Depth into Semantic Segmentation via Fusion-Based CNN Architecture
Caner Hazirbas, Lingni Ma, Csaba Domokos, Daniel Cremers
ACCV (1)2
2016 CPA-SLAM: Consistent plane-model alignment for direct RGB-D SLAM
abstract
Planes are predominant features of man-made environments which have been exploited in many mapping approaches. In this paper, we propose a real-time capable RGB-D SLAM system that consistently integrates frame-to-keyframe and frame-to-plane alignment. Our method models the environment with a global plane model and - besides direct image alignment - it uses the planes for tracking and global graph optimization. This way, our method makes use of the dense image information available in keyframes for accurate short-term tracking. At the same time it uses a global model to reduce drift. Both components are integrated consistently in an expectation-maximization framework. In experiments, we demonstrate the benefits our approach and its state-of-the-art accuracy on challenging benchmarks.
Lingni Ma, Christian Kerl, Jörg Stückler, Daniel Cremers
ICRA1
2013 On photo-realistic 3D reconstruction of large-scale and arbitrary-shaped environments
abstract
This paper presents a system architecture for reconstructing photorealistic and accurate 3D models of indoor environments. The system specifically targets large-scale and arbitrary-shaped environments and enables processing of data obtained with an arbitrary-chosen capturing path. The system extends the baseline Kinect Fusion algorithm with a buffering algorithm to remove scene-size capturing limitations. Beside this, the paper presents the complete chain of advanced algorithms for point cloud segmentation/decimation, camera pose correction and texture mapping with post-processing filters. The presented architecture features memory- and processor-efficient processing, such that it can be executed on a conventional PC with a mainstream GPU card at the consumer premises.
Egor Bondarev, Francisco Heredia, Rafael Favier, Lingni Ma, Peter H. N. de With
CCNC4
2013 Plane segmentation and decimation of point clouds for 3D environment reconstruction
abstract
Three-dimensional (3D) models of environments are a promising technique for serious gaming and professional engineering applications. In this paper, we introduce a fast and memory-efficient system for the reconstruction of large-scale environments based on point clouds. Our main contribution is the emphasis on the data processing of large planes, for which two algorithms have been designed to improve the overall performance of the 3D reconstruction. First, a flatness-based segmentation algorithm is presented for plane detection in point clouds. Second, a quadtree-based algorithm is proposed for decimating the point cloud involved with the segmented plane and consequently improving the efficiency of triangulation. Our experimental results have shown that the proposed system and algorithms have a high efficiency in speed and memory for environment reconstruction. Depending on the amount of planes in the scene, the obtained efficiency gain varies between 20% and 50%.
Lingni Ma, Raphael Favier, Luat Do, Egor Bondarev, Peter H. N. de With
CCNC1
2012 Depth-guided inpainting algorithm for Free-Viewpoint Video
abstract
Free-Viewpoint Video (FVV) is a novel technique which creates virtual images of multiple direction by view synthesis. In this paper, an exemplar-based depth-guided inpainting algorithm is proposed to fill disocclusions due to uncovered areas after projection. We develop an improved priority function which uses the depth information to impose a desirable inpainting order. We also propose an efficient background-foreground separation technique to enhance the accuracy of hole filling. Furthermore, a gradient-based searching approach is developed to reduce the computational cost and the location distance is incorporated into patch matching criteria to improve the accuracy. The experimental results have shown that the gradient-based search in our algorithm requires a much lower computational cost (factor of 6 compared to global search), while producing significantly improved visual results.
Lingni Ma, Luat Do, Peter H. N. de With
ICIP1