EDBT 2026 Demo / reviewers in the wild / expert
Lahav Lipson
dblp:302/0769
· DBLP profile ↗
10ranked-venue papers
4as first author
10since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Mesh Extraction for Unbounded Scenes Using Camera-Aware OctreesabstractMesh extraction from occupancy functions is a useful tool in creating synthetic datasets for computer vision. However, existing mesh extraction methods have artifacts or performance profiles that limit their use. We propose OcMesher, a mesh extractor that efficiently handles high-detail unbounded scenes with perfect view consistency, with easy export to downstream real-time engines. The main novelty is an algorithm to construct an octree based on a given occupancy function and multiple camera views. We performed extensive experiments, and demonstrate OcMesher's usefulness for synthetic training & benchmark datasets, generating real-time environments for embodied AI and mesh extraction from depthmaps or novel view synthesis methods. Zeyu Ma 0004, Alexander Raistrick, Lahav Lipson, Jia Deng 0001 |
3DV | 3 |
| 2025 | CoMotion: Concurrent Multi-person 3D MotionabstractWe introduce an approach for detecting and tracking detailed 3D poses of multiple people from a single monocular camera stream. Our system maintains temporally coherent predictions in crowded scenes filled with difficult poses and occlusions. Our model performs both strong per-frame detection and a learned pose update to track people from frame to frame. Rather than match detections across time, poses are updated directly from a new input image, which enables online tracking through occlusion. We train on numerous image and video datasets leveraging pseudo-labeled annotations to produce a model that matches state-of-the-art systems in 3D pose estimation accuracy while being faster and more accurate in tracking multiple people through time. Alejandro Newell, Peiyun Hu, Lahav Lipson, Stephan R. Richter, Vladlen Koltun |
ICLR | 3 |
| 2024 | Multi-Session SLAM with Differentiable Wide-Baseline Pose OptimizationabstractWe introduce a new system for Multi-Session SLAM, which tracks camera motion across multiple disjoint videos under a single global reference. Our approach couples the prediction of optical flow with solver layers to estimate camera pose. The backbone is trained end-to-end using a novel differentiable solver for wide-baseline two-view pose. The full system can connect disjoint sequences, perform visualodometry, and global optimization. Compared to existing approaches, our design is accurate and robust to catas-trophic failures. Code is available at h t tps://github. com/princeton-v1/MultiSlam_DiffPose Lahav Lipson, Jia Deng 0001 |
CVPR | 1 |
| 2024 | Infinigen Indoors: Photorealistic Indoor Scenes using Procedural GenerationabstractWe introduce Infinigen Indoors, a Blender-based procedural generator of photorealistic indoor scenes. It builds upon the existing Infinigen system, which focuses on natural scenes, but expands its coverage to indoor scenes by introducing a diverse library of procedural indoor assets, including furniture, architecture elements, appliances, and other day-to-day objects. It also introduces a constraint-based arrangement system, which consists of a domain-specific language for expressing diverse constraints on scene composition, and a solver that generates scene compositions that maximally satisfy the constraints. We provide an export tool that allows the generated 3D objects and scenes to be directly used for training embodied agents in real-time simulators such as Omniverse and Unreal. Infinigen Indoors is open-sourced under the BSD license. Please visit infinigen.org for code and videos. Alexander Raistrick, Lingjie Mei, Karhan Kayan, David Yan, Yiming Zuo 0001, Beining Han, Hongyu Wen, Meenal Parakh, Stamatis Alexandropoulos, Lahav Lipson, Zeyu Ma 0004, Jia Deng 0001 |
CVPR | 10 |
| 2024 | Deep Patch Visual SLAM
Lahav Lipson, Zachary Teed, Jia Deng 0001 |
ECCV (2) | 1 |
| 2024 | SEA-RAFT: Simple, Efficient, Accurate RAFT for Optical Flow
Lahav Lipson, Jia Deng 0001 |
ECCV (7) | 2 |
| 2023 | Infinite Photorealistic Worlds Using Procedural GenerationabstractWe introduce Infinigen, a procedural generator of photorealistic 3D scenes of the natural world. Infinigen is entirely procedural: every asset, from shape to texture, is generated from scratch via randomized mathematical rules, using no external source and allowing infinite variation and composition. Infinigen offers broad coverage of objects and scenes in the natural world including plants, animals, terrains, and natural phenomena such as fire, cloud, rain, and snow. Infinigen can be used to generate unlimited, diverse training data for a wide range of computer vision tasks including object detection, semantic segmentation, optical flow, and 3D reconstruction. We expect Infinigen to be a useful resource for computer vision research and beyond. Please visit infinigen.org for videos, code and pre-generated data. Alexander Raistrick, Lahav Lipson, Zeyu Ma 0004, Lingjie Mei, Yiming Zuo 0001, Karhan Kayan, Hongyu Wen, Beining Han, Alejandro Newell, Hei Law, Ankit Goyal 0001, Kaiyu Yang, Jia Deng 0001 |
CVPR | 2 |
| 2023 | Deep Patch Visual OdometryabstractWe propose Deep Patch Visual Odometry (DPVO), a new deep learning system for monocular Visual Odometry (VO). DPVO uses a novel recurrent network architecture designed for tracking image patches across time. Recent approaches to VO have significantly improved the state-of-the-art accuracy by using deep networks to predict dense flow between video frames. However, using dense flow incurs a large computational cost, making these previous methods impractical for many use cases. Despite this, it has been assumed that dense flow is important as it provides additional redundancy against incorrect matches. DPVO disproves this assumption, showing that it is possible to get the best accuracy and efficiency by exploiting the advantages of sparse patch-based matching over dense flow. DPVO introduces a novel recurrent update operator for patch based correspondence coupled with differentiable bundle adjustment. On Standard benchmarks, DPVO outperforms all prior work, including the learning-based state-of-the-art VO-system (DROID) using a third of the memory while running 3x faster on average. Code is available at https://github.com/princeton-vl/DPVO Zachary Teed, Lahav Lipson, Jia Deng 0001 |
NeurIPS | 2 |
| 2022 | Coupled Iterative Refinement for 6D Multi-Object Pose EstimationabstractWe address the task of 6D multi-object pose: given a set of known 3D objects and an RGB or RGB-D input image, we detect and estimate the 6D pose of each object. We propose a new approach to 6D object pose estimation which consists of an end-to-end differentiable architecture that makes use of geometric knowledge. Our approach iteratively refines both pose and correspondence in a tightly coupled manner, allowing us to dynamically remove outliers to improve accuracy. We use a novel differentiable layer to perform pose refinement by solving an optimization problem we refer to as Bidirectional Depth-Augmented Perspective-N-Point (BD-PnP). Our method achieves state-of-the-art accuracy on standard 6D Object Pose benchmarks. Code is available at https://github.com/princeton-vl/Coupled-Iterative-Refinement. Lahav Lipson, Zachary Teed, Ankit Goyal 0001, Jia Deng 0001 |
CVPR | 1 |
| 2021 | RAFT-Stereo: Multilevel Recurrent Field Transforms for Stereo MatchingabstractWe introduce RAFT-Stereo, a new deep architecture for rectified stereo based on the optical flow network RAFT [35]. We introduce multi-level convolutional GRUs, which more efficiently propagate information across the image. A modified version of RAFT-Stereo can perform accurate real-time inference. RAFT-stereo ranks first on the Middlebury leaderboard, outperforming the next best method on 1px error by 29% and outperforms all published work on the ETH3D two-view stereo benchmark. Code is available at https://github.com/princeton-vl/RAFT-Stereo. Lahav Lipson, Zachary Teed, Jia Deng 0001 |
3DV | 1 |