Beichen Li 0005

dblp:201/8574-5 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0002-9271-0055ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Robot navigation and mapping · 30% Vision and language · 15% 3D vision · 15%
Computer graphics and multimedia
5 papers
Rendering · 56% Visual content generation and editing · 21% Geometric modeling and processing · 20%

Topics — the 18 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Rendering
material appearance
2.032025
VLMaterial: Procedural Material Generation with Large Vision-Language Models · ICLR 2025
End-to-end Procedural Material Capture with Proxy-Free Mixed-Integer Optimization · ACM Trans. Graph. 2023
Match: differentiable material graphs for procedural material capture · ACM Trans. Graph. 2020
Computer vision › Vision and language
vision-language model
0.912025
VLMaterial: Procedural Material Generation with Large Vision-Language Models · ICLR 2025
Machine learning › Generative modeling
molecular generation
0.612022
Data-Efficient Graph Grammar Learning for Molecular Generation · ICLR 2022
Machine learning › Representation and self-supervised learning
structured representation
0.612022
Data-Efficient Graph Grammar Learning for Molecular Generation · ICLR 2022
Computer vision › 3D vision
3d reconstruction
0.412020
HeteroFusion: Dense Scene Reconstruction Integrating Multi-Sensors · IEEE Trans. Vis. Comput. Graph. 2020
Computer vision › 3D vision › 3d scene reconstruction
dense scene reconstruction
0.412020
HeteroFusion: Dense Scene Reconstruction Integrating Multi-Sensors · IEEE Trans. Vis. Comput. Graph. 2020
Robotics › Robot navigation and mapping › SLAM
multi-sensor SLAM
0.412020
HeteroFusion: Dense Scene Reconstruction Integrating Multi-Sensors · IEEE Trans. Vis. Comput. Graph. 2020
Robotics › Robot navigation and mapping
sensor fusion
0.412020
HeteroFusion: Dense Scene Reconstruction Integrating Multi-Sensors · IEEE Trans. Vis. Comput. Graph. 2020
Robotics › Robot navigation and mapping › sensor fusion
sensor fusion for pose estimation
0.412020
HeteroFusion: Dense Scene Reconstruction Integrating Multi-Sensors · IEEE Trans. Vis. Comput. Graph. 2020
Robotics › Robot navigation and mapping
SLAM
0.412020
HeteroFusion: Dense Scene Reconstruction Integrating Multi-Sensors · IEEE Trans. Vis. Comput. Graph. 2020
Geometric modeling and processing
3d reconstruction
0.412020
Noise-Resilient Reconstruction of Panoramas and 3D Scenes Using Robot-Mounted Unsynchronized Commodity RGB-D Cameras · ACM Trans. Graph. 2020
Geometric modeling and processing › 3d reconstruction
3d scene reconstruction
0.412020
Noise-Resilient Reconstruction of Panoramas and 3D Scenes Using Robot-Mounted Unsynchronized Commodity RGB-D Cameras · ACM Trans. Graph. 2020
Rendering
differentiable rendering
0.412020
Match: differentiable material graphs for procedural material capture · ACM Trans. Graph. 2020
Robotics › Legged, aerial and field robots
aerial robots
0.412019
Learning to fly: computational controller design for hybrid UAVs with reinforcement learning · ACM Trans. Graph. 2019
Robotics › Motion planning and robot control › robot control
learning control
0.412019
Learning to fly: computational controller design for hybrid UAVs with reinforcement learning · ACM Trans. Graph. 2019
Robotics › Motion planning and robot control
robot control
0.412019
Learning to fly: computational controller design for hybrid UAVs with reinforcement learning · ACM Trans. Graph. 2019
Computational photography and imaging › depth sensing
RGB-D imaging
0.112020
Noise-Resilient Reconstruction of Panoramas and 3D Scenes Using Robot-Mounted Unsynchronized Commodity RGB-D Cameras · ACM Trans. Graph. 2020
Visual content generation and editing
style transfer
0.112020
Match: differentiable material graphs for procedural material capture · ACM Trans. Graph. 2020

Methods — techniques the papers use, named apart from their topics

program-level augmentation · 1.7large vision-language model · 1.7large language model · 1.7fine-tuning · 1.7gradient-based optimization · 1.1transformer · 0.8reinforcement learning · 0.8mixed-integer optimization · 0.7gradient-free parameter search · 0.7graph grammar induction · 0.6data-efficient learning · 0.6loop-closure optimization · 0.4deep neural feature-based graph selection · 0.4classifier for pose validation · 0.4TSDF · 0.4IMU fusion · 0.4
YearPublicationVenuePosition
2025 VLMaterial: Procedural Material Generation with Large Vision-Language Models
abstract
Procedural materials, represented as functional node graphs, are ubiquitous in computer graphics for photorealistic material appearance design. They allow users to perform intuitive and precise editing to achieve desired visual appearances. However, creating a procedural material given an input image requires professional knowledge and significant effort. In this work, we leverage the ability to convert procedural materials into standard Python programs and fine-tune a large pre-trained vision-language model (VLM) to generate such programs from input images. To enable effective fine-tuning, we also contribute an open-source procedural material dataset and propose to perform program-level augmentation by prompting another pre-trained large language model (LLM). Through extensive evaluation, we show that our method outperforms previous methods on both synthetic and real-world examples.
Beichen Li 0005, Rundi Wu, Armando Solar-Lezama, Changxi Zheng, Liang Shi 0003, Bernd Bickel, Wojciech Matusik
ICLR1
2024 Procedural Material Generation with Reinforcement Learning
abstract
Modern 3D content creation heavily relies on procedural assets. In particular, procedural materials are ubiquitous in the industry, but their manipulation remains challenging. Previous work [Hu et al. 2023] conditionally generates procedural graphs that match a given input image. However, the parameter generation step limits how accurately the generated graph matches the input image, due to a reliance on supervision with scarcely available procedural data. We propose to improve parameter prediction accuracy for image-conditioned procedural material generation by leveraging reinforcement learning (RL) and present the first RL approach for procedural materials. RL circumvents the limited availability of procedural data, the domain gap between real and synthetic materials, and the need for end-to-end differentiable loss functions. Given a target image, we retrieve a procedural material and use an RL-trained transformer model to predict a set of parameters that reconstruct the target image as closely as possible. We show that using RL significantly improves parameter prediction to match a given target image compared to supervised methods on both synthetic and real target images.
Beichen Li 0005, Paul Guerrero 0001, Milos Hasan, Liang Shi 0003, Valentin Deschaintre, Wojciech Matusik
ACM Trans. Graph.1
2023 End-to-end Procedural Material Capture with Proxy-Free Mixed-Integer Optimization
abstract
Node-graph-based procedural materials are vital to 3D content creation within the computer graphics industry. Leveraging the expressive representation of procedural materials, artists can effortlessly generate diverse appearances by altering the graph structure or node parameters. However, manually reproducing a specific appearance is a challenging task that demands extensive domain knowledge and labor. Previous research has sought to automate this process by converting artist-created material graphs into differentiable programs and optimizing node parameters against a photographed material appearance using gradient descent. These methods involve implementing differentiable filter nodes [Shi et al. 2020] and training differentiable neural proxies for generator nodes to optimize continuous and discrete node parameters [Hu et al. 2022a] jointly. Nevertheless, Neural Proxies exhibits critical limitations, such as long training times, inaccuracies, fixed resolutions, and confined parameter ranges, which hinder their scalability towards the broad spectrum of production-grade material graphs. These constraints fundamentally stem from the absence of faithful and efficient implementations of generic noise and pattern generator nodes, both differentiable and non-differentiable. Such deficiency prevents the direct optimization of continuous and discrete generator node parameters without relying on surrogate models. We present Diffmat v2 , an improved differentiable procedural material library, along with a fully-automated, end-to-end procedural material capture framework that combines gradient-based optimization and gradient-free parameter search to match existing production-grade procedural materials against user-taken flash photos. Diffmat v2 expands the range of differentiable material graph nodes in Diffmat [Shi et al. 2020] by adding generic noise/pattern generator nodes and user-customizable per-pixel filter nodes. This allows for the complete translation and optimization of procedural materials across various categories without the need for external proprietary tools or pre-cached noise patterns. Consequently, our method can capture a considerably broader array of materials, encompassing those with highly regular or stochastic geometries. We demonstrate that our end-to-end approach yields a closer match to the target than MATch [Shi et al. 2020] and Neural Proxies [Hu et al. 2022a] when starting from initially unmatched continuous and discrete parameters.
Beichen Li 0005, Liang Shi 0003, Wojciech Matusik
ACM Trans. Graph.1
2022 Data-Efficient Graph Grammar Learning for Molecular Generation
Veronika Thost, Beichen Li 0005, Jie Chen 0007, Wojciech Matusik
ICLR3
2020 Match: differentiable material graphs for procedural material capture
abstract
We present MATch , a method to automatically convert photographs of material samples into production-grade procedural material models. At the core of MATch is a new library DiffMat that provides differentiable building blocks for constructing procedural materials, and automatic translation of large-scale procedural models, with hundreds to thousands of node parameters, into differentiable node graphs. Combining these translated node graphs with a rendering layer yields an end-to-end differentiable pipeline that maps node graph parameters to rendered images. This facilitates the use of gradient-based optimization to estimate the parameters such that the resulting material, when rendered, matches the target image appearance, as quantified by a style transfer loss. In addition, we propose a deep neural feature-based graph selection and parameter initialization method that efficiently scales to a large number of procedural graphs. We evaluate our method on both rendered synthetic materials and real materials captured as flash photographs. We demonstrate that MATch can reconstruct more accurate, general, and complex procedural materials compared to the state-of-the-art. Moreover, by producing a procedural output, we unlock capabilities such as constructing arbitrary-resolution material maps and parametrically editing the material appearance.
Liang Shi 0003, Beichen Li 0005, Milos Hasan, Kalyan Sunkavalli, Tamy Boubekeur, Radomír Mech, Wojciech Matusik
ACM Trans. Graph.2
2020 Noise-Resilient Reconstruction of Panoramas and 3D Scenes Using Robot-Mounted Unsynchronized Commodity RGB-D Cameras
abstract
We present a two-stage approach to first constructing 3D panoramas and then stitching them for noise-resilient reconstruction of large-scale indoor scenes. Our approach requires multiple unsynchronized RGB-D cameras, mounted on a robot platform, which can perform in-place rotations at different locations in a scene. Such cameras rotate on a common (but unknown) axis, which provides a novel perspective for coping with unsynchronized cameras, without requiring sufficient overlap of their Field-of-View (FoV). Based on this key observation, we propose novel algorithms to track these cameras simultaneously. Furthermore, during the integration of raw frames onto an equirectangular panorama, we derive uncertainty estimates from multiple measurements assigned to the same pixels. This enables us to appropriately model the sensing noise and consider its influence, so as to achieve better noise resilience, and improve the geometric quality of each panorama and the accuracy of global inter-panorama registration. We evaluate and demonstrate the performance of our proposed method for enhancing the geometric quality of scene reconstruction from both real-world and synthetic scans.
Sheng Yang 0007, Beichen Li 0005, Yan-Pei Cao 0001, Hongbo Fu 0001, Yukun Lai, Leif Kobbelt, Shi-Min Hu 0001
ACM Trans. Graph.2
2020 HeteroFusion: Dense Scene Reconstruction Integrating Multi-Sensors
abstract
We present a novel approach to integrate data from multiple sensor types for dense 3D reconstruction of indoor scenes in realtime. Existing algorithms are mainly based on a single RGBD camera and thus require continuous scanning of areas with sufficient geometric features. Otherwise, tracking may fail due to unreliable frame registration. Inspired by the fact that the fusion of multiple sensors can combine their strengths towards a more robust and accurate self-localization, we incorporate multiple types of sensors which are prevalent in modern robot systems, including a 2D range sensor, an inertial measurement unit (IMU), and wheel encoders. We fuse their measurements to reinforce the tracking process and to eventually obtain better 3D reconstructions. Specifically, we develop a 2D truncated signed distance field (TSDF) volume representation for the integration and ray-casting of laser frames, leading to a unified cost function in the pose estimation stage. For validation of the estimated poses in the loop-closure optimization process, we train a classifier for the features extracted from heterogeneous sensors during the registration progress. To evaluate our method on challenging use case scenarios, we assembled a scanning platform prototype to acquire real-world scans. We further simulated synthetic scans based on high-fidelity synthetic scenes for quantitative evaluation. Extensive experimental evaluation on these two types of scans demonstrate that our system is capable of robustly acquiring dense 3D reconstructions and outperforms state-of-the-art RGBD and LiDAR systems.
Sheng Yang 0007, Beichen Li 0005, Minghua Liu, Yukun Lai, Leif Kobbelt, Shi-Min Hu 0001
IEEE Trans. Vis. Comput. Graph.2
2019 Learning to fly: computational controller design for hybrid UAVs with reinforcement learning
abstract
Hybrid unmanned aerial vehicles (UAV) combine advantages of multicopters and fixed-wing planes: vertical take-off, landing, and low energy use. However, hybrid UAVs are rarely used because controller design is challenging due to its complex, mixed dynamics. In this paper, we propose a method to automate this design process by training a mode-free, model-agnostic neural network controller for hybrid UAVs. We present a neural network controller design with a novel error convolution input trained by reinforcement learning. Our controller exhibits two key features: First, it does not distinguish among flying modes, and the same controller structure can be used for copters with various dynamics. Second, our controller works for real models without any additional parameter tuning process, closing the gap between virtual simulation and real fabrication. We demonstrate the efficacy of the proposed controller both in simulation and in our custom-built hybrid UAVs (Figure 1, 8). The experiments show that the controller is robust to exploit the complex dynamics when both rotors and wings are active in flight tests.
Jie Xu 0028, Tao Du 0001, Michael Foshey, Beichen Li 0005, Bo Zhu 0002, Adriana Schulz, Wojciech Matusik
ACM Trans. Graph.4
2018 Student cluster competition 2017, team Tsinghua University: Reproducing vectorization of the tersoff multi-body potential on the Intel Skylake and NVIDIA Volta architectures
Ka Cheong Jason Lau, Qian Xie 0005, Beichen Li 0005, Guanyu Feng, Jiping Yu, Xinjian Yu, Jidong Zhai
Parallel Comput.5