EDBT 2026 Demo / reviewers in the wild / expert
Xinhang Liu
dblp:291/3884
· DBLP profile ↗
11ranked-venue papers
5as first author
11since 2021 · last 2025
0009-0003-0494-4877ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
3D vision · 42% Segmentation and scene understanding · 17% Generative modeling · 15% | |
| Computer graphics and multimedia
7 papers |
Rendering · 62% Computer animation and physical simulation · 20% Visual content generation and editing · 12% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 30 heaviest of 32, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Rendering › temporal rendering
dynamic scene rendering |
1.3 | 2 | 2024 | Gear-NeRF: Free-Viewpoint Rendering and Tracking with Motion-Aware Spatio-Temporal Sampling · CVPR 2024 Fourier PlenOctrees for Dynamic Radiance Field Rendering in Real-time · CVPR 2022 |
Rendering
neural radiance fields |
1.3 | 2 | 2024 | Gear-NeRF: Free-Viewpoint Rendering and Tracking with Motion-Aware Spatio-Temporal Sampling · CVPR 2024 Editable free-viewpoint video using a layered neural representation · ACM Trans. Graph. 2021 |
Computer vision › 3D vision
3d reconstruction |
1.1 | 2 | 2025 | Martian World Model: Controllable Video Synthesis with Physically Accurate 3D Reconstructions · NeurIPS 2025 CryoFastAR: Fast Cryo-EM AB Initio Reconstruction Made Easy · ICCV 2025 |
Machine learning › Generative modeling › video generation
controllable video generation |
0.9 | 1 | 2025 | Martian World Model: Controllable Video Synthesis with Physically Accurate 3D Reconstructions · NeurIPS 2025 |
Natural language and speech › Language models and text generation
LLM agents |
0.9 | 1 | 2025 | Motion-Agent: A Conversational Framework for Human Motion Generation with LLMs · ICLR 2025 |
Computer vision › 3D vision › object pose estimation › 6d object pose estimation
novel object pose estimation |
0.9 | 1 | 2025 | MixRI: Mixing Features of Reference Images for Novel Object Pose Estimation · ICCV 2025 |
Computer vision › 3D vision
object pose estimation |
0.9 | 1 | 2025 | MixRI: Mixing Features of Reference Images for Novel Object Pose Estimation · ICCV 2025 |
Computer vision › 3D vision
pose estimation |
0.9 | 1 | 2025 | CryoFastAR: Fast Cryo-EM AB Initio Reconstruction Made Easy · ICCV 2025 |
Machine learning › Generative modeling
video generation |
0.9 | 1 | 2025 | Martian World Model: Controllable Video Synthesis with Physically Accurate 3D Reconstructions · NeurIPS 2025 |
Bioinformatics and computational biology › structural biology
cryo-electron microscopy |
0.9 | 1 | 2025 | CryoFastAR: Fast Cryo-EM AB Initio Reconstruction Made Easy · ICCV 2025 |
Computer animation and physical simulation › motion synthesis
human motion synthesis |
0.9 | 1 | 2025 | Motion-Agent: A Conversational Framework for Human Motion Generation with LLMs · ICLR 2025 |
Computer animation and physical simulation › motion synthesis › human motion synthesis
text-to-motion generation |
0.9 | 1 | 2025 | Motion-Agent: A Conversational Framework for Human Motion Generation with LLMs · ICLR 2025 |
Natural language and speech › Question answering and dialogue systems
conversational agents |
0.8 | 1 | 2024 | ChatCam: Empowering Camera Control through Conversational AI · NeurIPS 2024 |
Computer vision › Video understanding and tracking
object tracking |
0.8 | 1 | 2024 | Gear-NeRF: Free-Viewpoint Rendering and Tracking with Motion-Aware Spatio-Temporal Sampling · CVPR 2024 |
Visual content generation and editing
camera control |
0.8 | 1 | 2024 | ChatCam: Empowering Camera Control through Conversational AI · NeurIPS 2024 |
Rendering › novel view synthesis
free-viewpoint rendering |
0.8 | 1 | 2024 | Gear-NeRF: Free-Viewpoint Rendering and Tracking with Motion-Aware Spatio-Temporal Sampling · CVPR 2024 |
Rendering
novel view synthesis |
0.8 | 1 | 2024 | Deceptive-NeRF/3DGS: Diffusion-Generated Pseudo-observations for High-Quality Sparse-View Reconstruction · ECCV (16) 2024 |
Computer vision › Segmentation and scene understanding › image segmentation
multi-view segmentation |
0.6 | 1 | 2022 | Unsupervised Multi-View Object Segmentation Using Radiance Field Propagation · NeurIPS 2022 |
Computer vision › 3D vision
neural radiance field |
0.6 | 1 | 2022 | Unsupervised Multi-View Object Segmentation Using Radiance Field Propagation · NeurIPS 2022 |
Computer vision › Segmentation and scene understanding
object segmentation |
0.6 | 1 | 2022 | Unsupervised Multi-View Object Segmentation Using Radiance Field Propagation · NeurIPS 2022 |
Computer vision › Segmentation and scene understanding › image segmentation
unsupervised segmentation |
0.6 | 1 | 2022 | Unsupervised Multi-View Object Segmentation Using Radiance Field Propagation · NeurIPS 2022 |
Rendering › neural radiance fields
neural radiance field rendering |
0.6 | 1 | 2022 | Fourier PlenOctrees for Dynamic Radiance Field Rendering in Real-time · CVPR 2022 |
Rendering
real-time rendering |
0.6 | 1 | 2022 | Fourier PlenOctrees for Dynamic Radiance Field Rendering in Real-time · CVPR 2022 |
Virtual and augmented reality › immersive video
free-viewpoint video |
0.5 | 1 | 2021 | Editable free-viewpoint video using a layered neural representation · ACM Trans. Graph. 2021 |
Computer vision › Vision and language › motion-language model
text-motion alignment |
0.3 | 1 | 2025 | Motion-Agent: A Conversational Framework for Human Motion Generation with LLMs · ICLR 2025 |
Visual content generation and editing
video generation |
0.3 | 1 | 2025 | Martian World Model: Controllable Video Synthesis with Physically Accurate 3D Reconstructions · NeurIPS 2025 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.2 | 1 | 2024 | Gear-NeRF: Free-Viewpoint Rendering and Tracking with Motion-Aware Spatio-Temporal Sampling · CVPR 2024 |
Computer vision › 3D vision › 3d reconstruction › multi-view reconstruction
sparse-view reconstruction |
0.2 | 1 | 2024 | Deceptive-NeRF/3DGS: Diffusion-Generated Pseudo-observations for High-Quality Sparse-View Reconstruction · ECCV (16) 2024 |
Computer vision › 3D vision › novel view synthesis
free-viewpoint video |
0.2 | 1 | 2022 | Fourier PlenOctrees for Dynamic Radiance Field Rendering in Real-time · CVPR 2022 |
Rendering
dynamic scene representation |
0.1 | 1 | 2021 | Editable free-viewpoint video using a layered neural representation · ACM Trans. Graph. 2021 |
Methods — techniques the papers use, named apart from their topics
neural radiance field · 3.5stereo navigation image processing · 1.7progressive training · 1.7multimodal conditioning · 1.7multi-view feature integration · 1.7motion tokenization · 1.7large language model · 1.7contrast transfer function modeling · 1.7adapter fine-tuning · 1.7feature mixing · 0.9spatio-temporal sampling · 0.8semantic embedding · 0.8diffusion model · 0.83d gaussian splatting · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MixRI: Mixing Features of Reference Images for Novel Object Pose Estimation
Xinhang Liu, Zheng Dang, Yuchao Dai |
ICCV | 1 |
| 2025 | CryoFastAR: Fast Cryo-EM AB Initio Reconstruction Made EasyabstractPose estimation from unordered images is fundamental for 3D reconstruction, robotics, and scientific imaging. Recent geometric foundation models, such as DUSt3R, enable end-to-end dense 3D reconstruction but remain underexplored in scientific imaging fields like cryo-electron microscopy (cryo-EM) for near-atomic protein reconstruction. In cryo-EM, pose estimation and 3D reconstruction from unordered particle images still depend on time-consuming iterative optimization, primarily due to challenges such as low signal-to-noise ratios (SNR) and distortions from the contrast transfer function (CTF). We introduce CryoFastAR, the first geometric foundation model that can directly predict poses from Cryo-EM noisy images for Fast ab initio Reconstruction. By integrating multi-view features and training on large-scale simulated cryo-EM data with realistic noise and CTF modulations, CryoFastAR enhances pose estimation accuracy and generalization. To enhance training stability, we propose a progressive training strategy that first allows the model to extract essential features under simpler conditions before gradually increasing difficulty to improve robustness. Experiments show that CryoFastAR achieves comparable quality while significantly accelerating inference over traditional iterative approaches on both synthetic and real datasets. Jiakai Zhang, Shouchen Zhou, Haizhao Dai, Xinhang Liu, Peihao Wang, Zhiwen Fan, Yuan Pei, Jingyi Yu 0001 |
ICCV | 4 |
| 2025 | Motion-Agent: A Conversational Framework for Human Motion Generation with LLMsabstractWhile previous approaches to 3D human motion generation have achieved notable success, they often rely on extensive training and are limited to specific tasks. To address these challenges, we introduce **Motion-Agent**, an efficient conversational framework designed for general human motion generation, editing, and understanding.
Motion-Agent employs an open-source pre-trained language model to develop a generative agent, **MotionLLM**, that bridges the gap between motion and text. This is accomplished by encoding and quantizing motions into discrete tokens that align with the language model's vocabulary. With only 1-3% of the model's parameters fine-tuned using adapters, MotionLLM delivers performance on par with diffusion models and other transformer-based methods trained from scratch. By integrating MotionLLM with GPT-4 without additional training, Motion-Agent is able to generate highly complex motion sequences through multi-turn conversations, a capability that previous models have struggled to achieve.
Motion-Agent supports a wide range of motion-language tasks, offering versatile capabilities for generating and customizing human motion through interactive conversational exchanges. Xinhang Liu, Yu-Wing Tai, Chi-Keung Tang |
ICLR | 4 |
| 2025 | Martian World Model: Controllable Video Synthesis with Physically Accurate 3D ReconstructionsabstractThe synthesis of realistic Martian landscape videos, essential for mission rehearsal and robotic simulation, presents unique challenges. These primarily stem from the scarcity of high-quality Martian data and the significant domain gap relative to terrestrial imagery.To address these challenges, we introduce a holistic solution comprising two main components: 1) a data curation framework, Multimodal Mars Synthesis (M3arsSynth), which processes stereo navigation images to render high-fidelity 3D video sequences. 2) a video-based Martian terrain generator (MarsGen), that utilizes multimodal conditioning data to accurately synthesize novel, 3D-consistent frames. Our data are sourced from NASA’s Planetary Data System (PDS), covering diverse Martian terrains and dates, enabling the production of physics-accurate 3D surface models at metric-scale resolution. During inference, MarsGen is conditioned on an initial image frame and can be guided by specified camera trajectories or textual prompts to generate new environments.Experimental results demonstrate that our solution surpasses video synthesis approaches trained on terrestrial data, achieving superior visual quality and 3D structural consistency. Zhiwen Fan, Wenyan Cong, Xinhang Liu, Yuyang Yin, Matthew Foutter, Panwang Pan, Chenyu You, Yue Wang 0041, Zhangyang Wang, Yao Zhao 0001, Marco Pavone 0001, Yunchao Wei |
NeurIPS | 4 |
| 2024 | Gear-NeRF: Free-Viewpoint Rendering and Tracking with Motion-Aware Spatio-Temporal SamplingabstractExtensions of Neural Radiance Fields (NeRFs) to model dynamic scenes have enabled their near photo-realistic, free-viewpoint rendering. Although these methods have shown some potential in creating immersive experiences, two drawbacks limit their ubiquity: ( i) a significant reduction in reconstruction quality when the computing budget is limited, and (ii) a lack of semantic understanding of the underlying scenes. To address these issues, we introduce Gear-NeRF, which leverages semantic information from powerful image segmentation models. Our approach presents a principled way for learning a spatio-temporal (4D) semantic embedding, based on which we introduce the concept of gears to allow for stratified modeling of dynamic regions of the scene based on the extent of their motion. Such differentiation allows us to adjust the spatio- temporal sampling resolution for each region in proportion to its motion scale, achieving more photo-realistic dynamic novel view synthesis. At the same time, almost for free, our approach enables free-viewpoint tracking of objects of interest - a functionality not yet achieved by existing NeRF-based methods. Empirical studies validate the effectiveness of our method, where we achieve state-of-the-art rendering and tracking performance on multiple challenging datasets. The project page is available at: https://merl.com/research/highlights/gear-nerf Xinhang Liu, Yu-Wing Tai, Chi-Keung Tang, Pedro Miraldo, Suhas Lohit, Moitreya Chatterjee |
CVPR | 1 |
| 2024 | Deceptive-NeRF/3DGS: Diffusion-Generated Pseudo-observations for High-Quality Sparse-View Reconstruction
Xinhang Liu, Jiaben Chen, Shiu-Hong Kao, Yu-Wing Tai, Chi-Keung Tang |
ECCV (16) | 1 |
| 2024 | ChatCam: Empowering Camera Control through Conversational AIabstractCinematographers adeptly capture the essence of the world, crafting compelling visual narratives through intricate camera movements. Witnessing the strides made by large language models in perceiving and interacting with the 3D world, this study explores their capability to control cameras with human language guidance. We introduce ChatCam, a system that navigates camera movements through conversations with users, mimicking a professional cinematographer's workflow. To achieve this, we propose CineGPT, a GPT-based autoregressive model for text-conditioned camera trajectory generation. We also develop an Anchor Determinator to ensure precise camera trajectory placement. ChatCam understands user requests and employs our proposed tools to generate trajectories, which can be used to render high-quality video footage on radiance field representations. Our experiments, including comparisons to state-of-the-art approaches and user studies, demonstrate our approach's ability to interpret and execute complex instructions for camera operation, showing promising applications in real-world production settings. Project page: https://xinhangliu.com/chatcam. Xinhang Liu, Yu-Wing Tai, Chi-Keung Tang |
NeurIPS | 1 |
| 2023 | Revisiting Event-Based Video Frame InterpolationabstractDynamic vision sensors or event cameras provide rich complementary information for video frame interpolation. Existing state-of-the-art methods follow the paradigm of combining both synthesis-based and warping networks. However, few of those methods fully respect the intrinsic characteristics of events streams. Given that event cameras only encode intensity changes and polarity rather than color intensities, estimating optical flow from events is arguably more difficult than from RGB information. We therefore propose to incorporate RGB information in an event-guided optical flow refinement strategy. Moreover, in light of the quasi-continuous nature of the time signals provided by event cameras, we propose a divide-and-conquer strategy in which event-based intermediate frame synthesis happens incrementally in multiple simplified stages rather than in a single, long stage. Extensive experiments on both synthetic and real-world datasets show that these modifications lead to more reliable and realistic intermediate frame results than previous video frame interpolation methods. Our findings underline that a careful consideration of event characteristics such as high temporal density and elevated noise benefits interpolation accuracy. Jiaben Chen, Dongze Lian, Yifu Wang, Renrui Zhang, Xinhang Liu, Shenhan Qian, Laurent Kneip, Shenghua Gao |
IROS | 7 |
| 2022 | Fourier PlenOctrees for Dynamic Radiance Field Rendering in Real-timeabstractImplicit neural representations such as Neural Radiance Field (NeRF) have focused mainly on modeling static objects captured under multi-view settings where real-time rendering can be achieved with smart data structures, e.g., PlenOctree. In this paper, we present a novel Fourier PlenOctree (FPO) technique to tackle efficient neural mod-eling and real-time rendering of dynamic scenes captured under the free-view video (FVV) setting. The key idea in our FPO is a novel combination of generalized NeRF, PlenOctree representation, volumetric fusion and Fourier transform. To accelerate FPO construction, we present a novel coarse-to-fine fusion scheme that leverages the gen-eralizable NeRF technique to generate the tree via spatial blending. To tackle dynamic scenes, we tailor the implicit network to model the Fourier coefficients of time-varying density and color attributes. Finally, we construct the FPO and train the Fourier coefficients directly on the leaves of a union PlenOctree structure of the dynamic sequence. We show that the resulting FPO enables compact memory overload to handle dynamic objects and supports efficient fine-tuning. Extensive experiments show that the proposed method is 3000 times faster than the original NeRF and achieves over an order of magnitude acceleration over SOTA while preserving high visual quality for the free-viewpoint rendering of unseen dynamic scenes. Jiakai Zhang, Xinhang Liu, Fuqiang Zhao, Yanshun Zhang, Yingliang Zhang, Minye Wu, Jingyi Yu 0001, Lan Xu 0003 |
CVPR | 3 |
| 2022 | Unsupervised Multi-View Object Segmentation Using Radiance Field PropagationabstractWe present radiance field propagation (RFP), a novel approach to segmenting objects in 3D during reconstruction given only unlabeled multi-view images of a scene. RFP is derived from emerging neural radiance field-based techniques, which jointly encodes semantics with appearance and geometry. The core of our method is a novel propagation strategy for individual objects' radiance fields with a bidirectional photometric loss, enabling an unsupervised partitioning of a scene into salient or meaningful regions corresponding to different object instances. To better handle complex scenes with multiple objects and occlusions, we further propose an iterative expectation-maximization algorithm to refine object masks. To the best of our knowledge, RFP is the first unsupervised approach for tackling 3D scene object segmentation for neural radiance field (NeRF) without any supervision, annotations, or other cues such as 3D bounding boxes and prior knowledge of object class. Experiments demonstrate that RFP achieves feasible segmentation results that are more accurate than previous unsupervised image/scene segmentation approaches, and are comparable to existing supervised NeRF-based methods. The segmented object representations enable individual 3D object editing operations. Codes and datasets will be made publicly available. Xinhang Liu, Jiaben Chen, Huai Yu, Yu-Wing Tai, Chi-Keung Tang |
NeurIPS | 1 |
| 2021 | Editable free-viewpoint video using a layered neural representationabstractGenerating free-viewpoint videos is critical for immersive VR/AR experience, but recent neural advances still lack the editing ability to manipulate the visual perception for large dynamic scenes. To fill this gap, in this paper, we propose the first approach for editable free-viewpoint video generation for large-scale view-dependent dynamic scenes using only 16 cameras. The core of our approach is a new layered neural representation, where each dynamic entity, including the environment itself, is formulated into a spatio-temporal coherent neural layered radiance representation called ST-NeRF. Such a layered representation supports manipulations of the dynamic scene while still supporting a wide free viewing experience. In our ST-NeRF, we represent the dynamic entity/layer as a continuous function, which achieves the disentanglement of location, deformation as well as the appearance of the dynamic entity in a continuous and self-supervised manner. We propose a scene parsing 4D label map tracking to disentangle the spatial information explicitly and a continuous deform module to disentangle the temporal motion implicitly. An object-aware volume rendering scheme is further introduced for the re-assembling of all the neural layers. We adopt a novel layered loss and motion-aware ray sampling strategy to enable efficient training for a large dynamic scene with multiple performers, Our framework further enables a variety of editing functions, i.e., manipulating the scale and location, duplicating or retiming individual neural layers to create numerous visual effects while preserving high realism. Extensive experiments demonstrate the effectiveness of our approach to achieve high-quality, photo-realistic, and editable free-viewpoint video generation for dynamic scenes. Jiakai Zhang, Xinhang Liu, Fuqiang Zhao, Yanshun Zhang, Minye Wu, Yingliang Zhang, Lan Xu 0003, Jingyi Yu 0001 |
ACM Trans. Graph. | 2 |