Yixing Lao

dblp:213/7784 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
10since 2021 · last 2025
0000-0001-8338-3577ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
3D vision · 77% Representation and self-supervised learning · 12% Autonomous driving · 5%
Computer graphics and multimedia
3 papers
Rendering · 87% Geometric modeling and processing · 13%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
GPUs and heterogeneous computing · 100%

Topics — the 26 heaviest of 28, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
novel view synthesis
1.522024
LiDAR-NeRF: Novel LiDAR View Synthesis via Neural Radiance Fields · ACM Multimedia 2024
AlignMiF: Geometry-Aligned Multimodal Implicit Field for LiDAR-Camera Joint Synthesis · CVPR 2024
Rendering
neural radiance fields
1.422024
LiDAR-NeRF: Novel LiDAR View Synthesis via Neural Radiance Fields · ACM Multimedia 2024
CorresNeRF: Image Correspondence Priors for Neural Radiance Fields · NeurIPS 2023
Computer vision › 3D vision › geometric deep learning
3d representation learning
0.912025
Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations · NeurIPS 2025
Computer vision › 3D vision
3d scene understanding
0.912025
Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations · NeurIPS 2025
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › multimodal self-supervised learning
cross-modal self-supervised learning
0.912025
Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations · NeurIPS 2025
Computer vision › 3D vision › depth estimation
monocular depth estimation
0.912025
StableDepth: Scene-Consistent and Scale-Invariant Monocular Depth · ICCV 2025
Computer vision › 3D vision › 3d scene understanding
point cloud scene understanding
0.912025
Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations · NeurIPS 2025
Machine learning › Representation and self-supervised learning
spatial representation learning
0.912025
Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations · NeurIPS 2025
Computer vision › 3D vision › 3d shape representation
implicit function
0.812024
AlignMiF: Geometry-Aligned Multimodal Implicit Field for LiDAR-Camera Joint Synthesis · CVPR 2024
Robotics › Autonomous driving › perception
LiDAR perception
0.812024
LiT: Unifying LiDAR "Languages" with LiDAR Translator · NeurIPS 2024
Computer vision › 3D vision › novel view synthesis
LiDAR view synthesis
0.812024
LiDAR-NeRF: Novel LiDAR View Synthesis via Neural Radiance Fields · ACM Multimedia 2024
Rendering › gaussian splatting
3d gaussian splatting
0.812024
Pixel-GS: Density Control with Pixel-Aware Gradient for 3D Gaussian Splatting · ECCV (19) 2024
Rendering
neural rendering
0.812024
LiDAR-NeRF: Novel LiDAR View Synthesis via Neural Radiance Fields · ACM Multimedia 2024
Rendering
point-based rendering
0.812024
Pixel-GS: Density Control with Pixel-Aware Gradient for 3D Gaussian Splatting · ECCV (19) 2024
Computer vision › 3D vision
point cloud registration
0.712023
ASH: A Modern Framework for Parallel Spatial Hashing in 3D Perception · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Computer vision › 3D vision › 3d reconstruction
volumetric reconstruction
0.712023
ASH: A Modern Framework for Parallel Spatial Hashing in 3D Perception · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Rendering
novel view synthesis
0.712023
CorresNeRF: Image Correspondence Priors for Neural Radiance Fields · NeurIPS 2023
Geometric modeling and processing
surface reconstruction
0.712023
CorresNeRF: Image Correspondence Priors for Neural Radiance Fields · NeurIPS 2023
GPUs and heterogeneous computing
GPU computing
0.712023
ASH: A Modern Framework for Parallel Spatial Hashing in 3D Perception · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Computer vision › 3D vision
point cloud analysis
0.612022
Point Transformer V2: Grouped Vector Attention and Partition-based Pooling · NeurIPS 2022
Computer vision › 3D vision › point cloud analysis
point cloud classification
0.612022
Point Transformer V2: Grouped Vector Attention and Partition-based Pooling · NeurIPS 2022
Computer vision › 3D vision
point cloud segmentation
0.612022
Point Transformer V2: Grouped Vector Attention and Partition-based Pooling · NeurIPS 2022
Computer vision › 3D vision › point cloud processing
point transformer
0.612022
Point Transformer V2: Grouped Vector Attention and Partition-based Pooling · NeurIPS 2022
Machine learning › Deep learning architectures and training
transformer
0.612022
Point Transformer V2: Grouped Vector Attention and Partition-based Pooling · NeurIPS 2022
Machine learning › Transfer learning and domain adaptation
domain adaptation
0.212024
LiT: Unifying LiDAR "Languages" with LiDAR Translator · NeurIPS 2024
Computer vision › 3D vision
correspondence estimation
0.212023
CorresNeRF: Image Correspondence Priors for Neural Radiance Fields · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

neural radiance field · 1.5adaptive augmentation and filtering · 1.3self-distillation · 0.9contrastive learning · 0.9CLIP · 0.9shared geometry initialization · 0.8scene reconstruction · 0.8ray casting · 0.8pixel-aware gradient · 0.8geometry-aware alignment · 0.8adaptive density control · 0.8LiDAR ray simulation · 0.8tensor interface · 0.7spatial hashing · 0.7image correspondence priors · 0.7depth loss · 0.7correspondence pixel reprojection loss · 0.7
YearPublicationVenuePosition
2025 StableDepth: Scene-Consistent and Scale-Invariant Monocular Depth
Lihe Yang, Tianyu Yang 0003, Chaohui Yu, Yixing Lao, Hengshuang Zhao
ICCV6
2025 Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations
abstract
Humans learn abstract concepts through multisensory synergy, and once formed, such representations can often be recalled from a single modality. Inspired by this principle, we introduce Concerto, a minimalist simulation of human concept learning for spatial cognition, combining 3D intra-modal self-distillation with 2D-3D cross-modal joint embedding. Despite its simplicity, Concerto learns more coherent and informative spatial features, as demonstrated by zero-shot visualizations. It outperforms both standalone SOTA 2D and 3D self-supervised models by 14.2\% and 4.8\%, respectively, as well as their feature concatenation, in linear probing for 3D scene perception. With full fine-tuning, Concerto sets new SOTA results across multiple scene understanding benchmarks (e.g., 80.7\% mIoU on ScanNet). We further present a variant of Concerto tailored for video-lifted point cloud spatial understanding, and a translator that linearly projects Concerto representations into CLIP’s language space, enabling open-world perception. These results highlight that Concerto emerges spatial representations with superior fine-grained geometric and semantic consistency.
Yujia Zhang 0003, Xiaoyang Wu 0002, Yixing Lao, Chengyao Wang, Zhuotao Tian, Naiyan Wang, Hengshuang Zhao
NeurIPS3
2024 Objects With Lighting: A Real-World Dataset for Evaluating Reconstruction and Rendering for Object Relighting
abstract
Reconstructing an object from photos and placing it virtually in a new environment goes beyond the standard novel view synthesis task as the appearance of the object has to not only adapt to the novel viewpoint but also to the new lighting conditions and yet evaluations of inverse rendering methods rely on novel view synthesis data or simplistic synthetic datasets for quantitative analysis. This work presents a real-world dataset for measuring the reconstruction and rendering of objects for relighting. To this end, we capture the environment lighting and ground truth images of the same objects in multiple environments allowing to reconstruct the objects from images taken in one environment and quantify the quality of the rendered views for the unseen lighting environments. Further, we introduce a simple baseline composed of off-the-shelf methods and test several state-of-the-art methods on the relighting task and show that novel view synthesis is not a reliable proxy to measure performance. Code and dataset are available at https://github.com/isl-org/objects-with-lighting.
Benjamin Ummenhofer, Sanskar Agrawal, Rene Sepúlveda, Yixing Lao, Tianhang Cheng, Stephan R. Richter, Shenlong Wang, Germán Ros 0001
3DV4
2024 AlignMiF: Geometry-Aligned Multimodal Implicit Field for LiDAR-Camera Joint Synthesis
abstract
Neural implicit fields have been a de facto standard in novel view synthesis. Recently, there exist some methods exploring fusing multiple modalities within a single field, aiming to share implicit features from different modalities to enhance reconstruction performance. However, these modalities often exhibit misaligned behaviors: optimizing for one modality, such as LiDAR, can adversely affect another, like camera performance, and vice versa. In this work, we conduct comprehensive analyses on the multimodal implicit field of LiDAR-camera joint synthesis, revealing the underlying issue lies in the misalignment of different sensors. Furthermore, we introduce AlignMiF, a geometrically aligned multimodal implicit field with two proposed modules: Geometry-Aware Alignment (GAA) and Shared Geometry Initialization (SGI). These modules effectively align the coarse geometry across different modalities, significantly enhancing the fusion process between LiDAR and camera data. Through extensive experiments across various datasets and scenes, we demonstrate the effectiveness of our approach in facilitating better interaction between LiDAR and camera modalities within a unified neural field. Specifically, our proposed AlignMiF, achieves remarkable improvement over recent implicit fusion methods (+2.01 and +3.11 image PSNR on the KITTI-360 and Waymo datasets) and consistently surpasses single modality performance (13.8% and 14.2% reduction in LiDAR Cham-fer Distance on the respective datasets). Code release: https://github.com/tangtaogo/alignmif.
Tang Tao, Guangrun Wang, Yixing Lao, Peng Chen 0054, Liang Lin 0004, Kaicheng Yu, Xiaodan Liang
CVPR3
2024 Pixel-GS: Density Control with Pixel-Aware Gradient for 3D Gaussian Splatting
Wenbo Hu 0002, Yixing Lao, Tong He 0001, Hengshuang Zhao
ECCV (19)3
2024 LiDAR-NeRF: Novel LiDAR View Synthesis via Neural Radiance Fields
Tang Tao, Longfei Gao, Guangrun Wang, Yixing Lao, Peng Chen 0054, Hengshuang Zhao, Dayang Hao, Xiaodan Liang, Mathieu Salzmann, Kaicheng Yu
ACM Multimedia4
2024 LiT: Unifying LiDAR "Languages" with LiDAR Translator
abstract
LiDAR data exhibits significant domain gaps due to variations in sensors, vehicles, and driving environments, creating “language barriers” that limit the effective use of data across domains and the scalability of LiDAR perception models. To address these challenges, we introduce the LiDAR Translator (LiT), a framework that directly translates LiDAR data across domains, enabling both cross-domain adaptation and multi-domain joint learning. LiT integrates three key components: a scene modeling module for precise foreground and background reconstruction, a LiDAR modeling module that models LiDAR rays statistically and simulates ray-drop, and a fast, hardware-accelerated ray casting engine. LiT enables state-of-the-art zero-shot and unified domain detection across diverse LiDAR datasets, marking a step toward data-driven domain unification for autonomous driving systems. Source code and demos are available at: https://yxlao.github.io/lit.
Yixing Lao, Xiaoyang Wu 0002, Peng Chen 0054, Kaicheng Yu, Hengshuang Zhao
NeurIPS1
2023 CorresNeRF: Image Correspondence Priors for Neural Radiance Fields
abstract
Neural Radiance Fields (NeRFs) have achieved impressive results in novel view synthesis and surface reconstruction tasks. However, their performance suffers under challenging scenarios with sparse input views. We present CorresNeRF, a novel method that leverages image correspondence priors computed by off-the-shelf methods to supervise NeRF training. We design adaptive processes for augmentation and filtering to generate dense and high-quality correspondences. The correspondences are then used to regularize NeRF training via the correspondence pixel reprojection and depth loss terms. We evaluate our methods on novel view synthesis and surface reconstruction tasks with density-based and SDF-based NeRF models on different datasets. Our method outperforms previous methods in both photometric and geometric metrics. We show that this simple yet effective technique of using correspondence priors can be applied as a plug-and-play module across different NeRF variants. The project page is at https://yxlao.github.io/corres-nerf/.
Yixing Lao, Xiaogang Xu 0002, Xihui Liu, Hengshuang Zhao
NeurIPS1
2023 ASH: A Modern Framework for Parallel Spatial Hashing in 3D Perception
abstract
We present ASH, a modern and high-performance framework for parallel spatial hashing on GPU. Compared to existing GPU hash map implementations, ASH achieves higher performance, supports richer functionality, and requires fewer lines of code (LoC) when used for implementing spatially varying operations from volumetric geometry reconstruction to differentiable appearance reconstruction. Unlike existing GPU hash maps, the ASH framework provides a versatile tensor interface, hiding low-level details from the users. In addition, by decoupling the internal hashing data structures and key-value data in buffers, we offer direct access to spatially varying data via indices, enabling seamless integration to modern libraries such as PyTorch. To achieve this, we 1) detach stored key-value data from the low-level hash map implementation; 2) bridge the pointer-first low level data structures to index-first high-level tensor interfaces via an index heap; 3) adapt both generic and non-generic integer-only hash map implementations as backends to operate on multi-dimensional keys. We first profile our hash map against state-of-the-art hash maps on synthetic data to show the performance gain from this architecture. We then show that ASH can consistently achieve higher performance on various large-scale 3D perception tasks with fewer LoC by showcasing several applications, including 1) point cloud voxelization, 2) retargetable volumetric scene reconstruction, 3) non-rigid point cloud registration and volumetric deformation, and 4) spatially varying geometry and appearance refinement. ASH and its example applications are open sourced in Open3D (http://www.open3d.org).
Yixing Lao, Michael Kaess, Vladlen Koltun
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Point Transformer V2: Grouped Vector Attention and Partition-based Pooling
abstract
As a pioneering work exploring transformer architecture for 3D point cloud understanding, Point Transformer achieves impressive results on multiple highly competitive benchmarks. In this work, we analyze the limitations of the Point Transformer and propose our powerful and efficient Point Transformer V2 model with novel designs that overcome the limitations of previous work. In particular, we first propose group vector attention, which is more effective than the previous version of vector attention. Inheriting the advantages of both learnable weight encoding and multi-head attention, we present a highly effective implementation of grouped vector attention with a novel grouped weight encoding layer. We also strengthen the position information for attention by an additional position encoding multiplier. Furthermore, we design novel and lightweight partition-based pooling methods which enable better spatial alignment and more efficient sampling. Extensive experiments show that our model achieves better performance than its predecessor and achieves state-of-the-art on several challenging 3D point cloud understanding benchmarks, including 3D point cloud segmentation on ScanNet v2 and S3DIS and 3D point cloud classification on ModelNet40. Our code will be available at https://github.com/Gofinge/PointTransformerV2.
Xiaoyang Wu 0002, Yixing Lao, Li Jiang 0009, Xihui Liu, Hengshuang Zhao
NeurIPS2
2019 nGraph-HE: a graph compiler for deep learning on homomorphically encrypted data
abstract
Homomorphic encryption (HE)---the ability to perform computation on encrypted data---is an attractive remedy to increasing concerns about data privacy in deep learning (DL). However, building DL models that operate on ciphertext is currently labor-intensive and requires simultaneous expertise in DL, cryptography, and software engineering. DL frameworks and recent advances in graph compilers have greatly accelerated the training and deployment of DL models to various computing platforms. We introduce nGraph-HE, an extension of nGraph, Intel's DL graph compiler, which enables deployment of trained models with popular frameworks such as TensorFlow while simply treating HE as another hardware target. Our graph-compiler approach enables HE-aware optimizations- implemented at compile-time, such as constant folding and HE-SIMD packing, and at run-time, such as special value plaintext bypass. Furthermore, nGraph-HE integrates with DL frameworks such as TensorFlow, enabling data scientists to benchmark DL models with minimal overhead.
Fabian Boemer, Yixing Lao, Rosario Cammarota, Casimir Wierzynski
CF2