VLDB 2026 Research / reviewers in the wild / expert
Xuangeng Chu
dblp:261/3590
· DBLP profile ↗
9ranked-venue papers
4as first author
8since 2021 · last 2025
0009-0008-1925-7566ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
6 papers |
Rendering · 64% Computer animation and physical simulation · 13% Visual content generation and editing · 8% | |
| Artificial intelligence
4 papers |
3D vision · 66% Image recognition and object detection · 30% Video understanding and tracking · 4% |
Topics — the 23 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Rendering
novel view synthesis |
1.7 | 2 | 2025 | Real-Time High-Resolution View Synthesis of Complex Scenes With Explicit 3D Visibility Reasoning · IEEE Trans. Vis. Comput. Graph. 2025 Luminance-GS: Adapting 3D Gaussian Splatting to Challenging Lighting Conditions with View-Adaptive Curve Adjustment · CVPR 2025 |
Rendering
volume rendering |
1.7 | 2 | 2025 | Real-Time High-Resolution View Synthesis of Complex Scenes With Explicit 3D Visibility Reasoning · IEEE Trans. Vis. Comput. Graph. 2025 I2-NeRF: Learning Neural Radiance Fields Under Physically-Grounded Media Interactions · NeurIPS 2025 |
Rendering
neural rendering |
1.6 | 2 | 2025 | Real-Time High-Resolution View Synthesis of Complex Scenes With Explicit 3D Visibility Reasoning · IEEE Trans. Vis. Comput. Graph. 2025 GPAvatar: Generalizable and Precise Head Avatar from Image(s) · ICLR 2024 |
Computer vision › 3D vision
3d reconstruction |
1.0 | 2 | 2025 | GPAvatar: Generalizable and Precise Head Avatar from Image(s) · ICLR 2024 Real-Time High-Resolution View Synthesis of Complex Scenes With Explicit 3D Visibility Reasoning · IEEE Trans. Vis. Comput. Graph. 2025 |
Rendering › volume rendering
differentiable volume rendering |
0.9 | 1 | 2025 | Real-Time High-Resolution View Synthesis of Complex Scenes With Explicit 3D Visibility Reasoning · IEEE Trans. Vis. Comput. Graph. 2025 |
Image and video processing › image enhancement
low-light image enhancement |
0.9 | 1 | 2025 | Luminance-GS: Adapting 3D Gaussian Splatting to Challenging Lighting Conditions with View-Adaptive Curve Adjustment · CVPR 2025 |
Rendering
neural radiance fields |
0.9 | 1 | 2025 | I2-NeRF: Learning Neural Radiance Fields Under Physically-Grounded Media Interactions · NeurIPS 2025 |
Computer animation and physical simulation › facial animation
speech-driven facial animation |
0.9 | 1 | 2025 | ARTalk: Speech-Driven 3D Head Animation via Autoregressive Model · SIGGRAPH Asia 2025 |
Computer vision › 3D vision › 3d reconstruction › object reconstruction
head avatar reconstruction |
0.8 | 1 | 2024 | GPAvatar: Generalizable and Precise Head Avatar from Image(s) · ICLR 2024 |
Visual content generation and editing
3d content creation |
0.8 | 1 | 2024 | GPAvatar: Generalizable and Precise Head Avatar from Image(s) · ICLR 2024 |
Rendering › novel view synthesis
free-viewpoint rendering |
0.8 | 1 | 2024 | GPAvatar: Generalizable and Precise Head Avatar from Image(s) · ICLR 2024 |
Rendering
gaussian splatting |
0.8 | 1 | 2024 | Generalizable and Animatable Gaussian Head Avatar · NeurIPS 2024 |
Geometric modeling and processing › 3d reconstruction › avatar reconstruction
head avatar reconstruction |
0.8 | 1 | 2024 | Generalizable and Animatable Gaussian Head Avatar · NeurIPS 2024 |
Computer vision › 3D vision
3d face reconstruction |
0.7 | 1 | 2023 | Accurate 3D Face Reconstruction with Facial Component Tokens · ICCV 2023 |
Computer vision › 3D vision › 3d face reconstruction
single-image 3d face reconstruction |
0.7 | 1 | 2023 | Accurate 3D Face Reconstruction with Facial Component Tokens · ICCV 2023 |
Computer vision › Image recognition and object detection › object detection › multi-object detection
crowded scene detection |
0.4 | 1 | 2020 | Detection in Crowded Scenes: One Proposal, Multiple Predictions · CVPR 2020 |
Computer vision › Image recognition and object detection › object detection › object detection post-processing
non-maximum suppression |
0.4 | 1 | 2020 | Detection in Crowded Scenes: One Proposal, Multiple Predictions · CVPR 2020 |
Computer vision › Image recognition and object detection
object detection |
0.4 | 1 | 2020 | Detection in Crowded Scenes: One Proposal, Multiple Predictions · CVPR 2020 |
Computer vision › Image recognition and object detection › object detection
proposal-based detection |
0.4 | 1 | 2020 | Detection in Crowded Scenes: One Proposal, Multiple Predictions · CVPR 2020 |
Computer vision › 3D vision › geometric deep learning › set learning
set prediction |
0.4 | 1 | 2020 | Detection in Crowded Scenes: One Proposal, Multiple Predictions · CVPR 2020 |
Computer vision › 3D vision › 3d reconstruction
geometric reconstruction |
0.3 | 1 | 2025 | Real-Time High-Resolution View Synthesis of Complex Scenes With Explicit 3D Visibility Reasoning · IEEE Trans. Vis. Comput. Graph. 2025 |
Computational photography and imaging
underwater imaging |
0.3 | 1 | 2025 | I2-NeRF: Learning Neural Radiance Fields Under Physically-Grounded Media Interactions · NeurIPS 2025 |
Visual content generation and editing › face editing
face reenactment |
0.2 | 1 | 2024 | Generalizable and Animatable Gaussian Head Avatar · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
multi-view supervision · 1.7explicit 3d visibility reasoning · 1.7attention · 1.5reverse-stratified upsampling · 0.9per-view color matrix mapping · 0.9motion codebook · 0.9curve adjustment · 0.9beer-lambert attenuation · 0.9autoregressive model · 0.93d gaussian splatting · 0.9tri-plane · 0.8point cloud · 0.8transformer · 0.7token-based representation · 0.7temporal transformer · 0.7resnet · 0.4feature pyramid network · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Luminance-GS: Adapting 3D Gaussian Splatting to Challenging Lighting Conditions with View-Adaptive Curve AdjustmentabstractCapturing high-quality photographs under diverse real-world lighting conditions is challenging, as both natural lighting (e.g., low-light) and camera exposure settings (e.g., exposure time) significantly impact image quality. This challenge becomes more pronounced in multi-view scenarios, where variations in lighting and image signal processor (ISP) settings across viewpoints introduce photometric inconsistencies. Such lighting degradations and view-dependent variations pose substantial challenges to novel view synthesis (NVS) frameworks based on Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS).To address this, we introduce Luminance-GS, a novel approach to achieving high-quality novel view synthesis results under diverse challenging lighting conditions using 3DGS. By adopting per-view color matrix mapping and view adaptive curve adjustments, Luminance-GS achieves state-of-the-art (SOTA) results across various lighting conditions—including low-light, overexposure, and varying exposure—while not altering the original 3DGS explicit representation. Compared to previous NeRF- and 3DGS-based baselines, Luminance-GS provides real-time rendering speed with improved reconstruction quality. The source code is available at1. Ziteng Cui, Xuangeng Chu, Tatsuya Harada |
CVPR | 2 |
| 2025 | I2-NeRF: Learning Neural Radiance Fields Under Physically-Grounded Media InteractionsabstractParticipating in efforts to endow generative AI with the 3D physical world perception, we propose I2-NeRF, a novel neural radiance field framework that enhances isometric and isotropic metric perception under media degradation. While existing NeRF models predominantly rely on object-centric sampling, I2-NeRF introduces a reverse-stratified upsampling strategy to achieve near-uniform sampling across 3D space, thereby preserving isometry. We further present a general radiative formulation for media degradation that unifies emission, absorption, and scattering into a particle model governed by the Beer–Lambert attenuation law. By matting direct and media-induced in-scatter radiance, this formulation extends naturally to complex media environments such as underwater, haze, and even low-light scenes. By treating light propagation uniformly in both vertical and horizontal directions, I2-NeRF enables isotropic metric perception and can even estimate medium properties such as water depth. Experiments on real-world datasets demonstrate that our method significantly improves both reconstruction fidelity and physical plausibility compared to existing approaches. The source code is available at https://github.com/ShuhongLL/I2-NeRF. Shuhong Liu, Lin Gu 0003, Ziteng Cui, Xuangeng Chu, Tatsuya Harada |
NeurIPS | 4 |
| 2025 | Intend to Move: A Multimodal Dataset for Intention-Aware Human Motion UnderstandingabstractHuman motion is inherently intentional, yet most motion modeling paradigms focus on low-level kinematics, overlooking the semantic and causal factors that drive behavior. Existing datasets further limit progress: they capture short, decontextualized actions in static scenes, providing little grounding for embodied reasoning. To address these limitations, we introduce $\textit{Intend to Move (I2M)}$, a large-scale, multimodal dataset for intention-grounded motion modeling. I2M contains 10.1 hours of two-person 3D motion sequences recorded in dynamic realistic home environments, accompanied by multi-view RGB-D video, 3D scene geometry, and language annotations of each participant’s evolving intentions. Benchmark experiments reveal a fundamental gap in current motion models: they fail to translate high-level goals into physically and socially coherent motion. I2M thus serves not only as a dataset but as a benchmark for embodied intelligence, enabling research on models that can reason about, predict, and act upon the ``why'' behind human motion. Ryo Umagami, Liu Yue, Xuangeng Chu, Ryuto Fukushima, Tetsuya Narita, Yusuke Mukuta, Tomoyuki Takahata, Tatsuya Harada |
NeurIPS | 3 |
| 2025 | ARTalk: Speech-Driven 3D Head Animation via Autoregressive ModelabstractSpeech-driven 3D facial animation aims to generate realistic lip movements and facial expressions for 3D head models from arbitrary audio clips. Although existing diffusion-based methods are capable of producing natural motions, their slow generation speed limits their application potential. In this paper, we introduce a novel autoregressive model that achieves real-time generation of highly synchronized lip movements and realistic head poses and eye blinks by learning a mapping from speech to a multi-scale motion codebook. Furthermore, our model can adapt to unseen speaking styles, enabling the creation of 3D talking avatars with unique personal styles beyond the identities seen during training. Extensive evaluations and user studies demonstrate that our method outperforms existing approaches in lip synchronization accuracy and perceived quality. Demos and codes are available at https://xg-chu.site/project_artalk/. Xuangeng Chu, Nabarun Goswami, Ziteng Cui, Hanqin Wang, Tatsuya Harada |
SIGGRAPH Asia | 1 |
| 2025 | Real-Time High-Resolution View Synthesis of Complex Scenes With Explicit 3D Visibility ReasoningabstractRendering photo-realistic novel-view images of complex scenes has been a long-standing challenge in computer graphics. In recent years, great research progress has been made in enhancing rendering quality and accelerating rendering speed in the realm of view synthesis. However, when rendering complex dynamic scenes with sparse views, the rendering quality remains limited due to occlusion problems. Besides, for rendering high-resolution images on dynamic scenes, the rendering speed is still far from real-time. In this work, we propose a generalizable view synthesis method that can render high-resolution novel-view images of complex static and dynamic scenes in real-time from sparse views. To address the occlusion problems arising from the sparsity of input views and the complexity of captured scenes, we introduce an explicit 3D visibility reasoning approach that can efficiently estimate the visibility of sampled 3D points to the input views. The proposed visibility reasoning approach is fully differentiable and can gracefully fit inside the volume rendering pipeline, allowing us to train our networks with only multi-view images as supervision while refining geometry and texture simultaneously. Besides, each module in our pipeline is carefully designed to bypass the time-consuming MLP querying process and enhance the rendering quality of high-resolution images, enabling us to render high-resolution novel-view images in real-time. Experimental results show that our method outperforms previous view synthesis methods in both rendering quality and speed, particularly when dealing with complex dynamic scenes with sparse views. Tiansong Zhou, Yu Li 0003, Xuangeng Chu, Chengkun Cao, Changyin Zhou, F. Richard Yu, Yebin Liu |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | GPAvatar: Generalizable and Precise Head Avatar from Image(s)abstractHead avatar reconstruction, crucial for applications in virtual reality, online meetings, gaming, and film industries, has garnered substantial attention within the computer vision community. The fundamental objective of this field is to faithfully recreate the head avatar and precisely control expressions and postures. Existing methods, categorized into 2D-based warping, mesh-based, and neural rendering approaches, present challenges in maintaining multi-view consistency, incorporating non-facial information, and generalizing to new identities. In this paper, we propose a framework named GPAvatar that reconstructs 3D head avatars from one or several images in a single forward pass. The key idea of this work is to introduce a dynamic point-based expression field driven by a point cloud to precisely and effectively capture expressions. Furthermore, we use a Multi Tri-planes Attention (MTA) fusion module in tri-planes canonical field to leverage information from multiple input images. The proposed method achieves faithful identity reconstruction, precise expression control, and multi-view consistency, demonstrating promising results for free-viewpoint rendering and novel view synthesis. Xuangeng Chu, Ailing Zeng, Lijian Lin, Tatsuya Harada |
ICLR | 1 |
| 2024 | Generalizable and Animatable Gaussian Head AvatarabstractIn this paper, we propose Generalizable and Animatable Gaussian head Avatar (GAGA) for one-shot animatable head avatar reconstruction.
Existing methods rely on neural radiance fields, leading to heavy rendering consumption and low reenactment speeds.
To address these limitations, we generate the parameters of 3D Gaussians from a single image in a single forward pass.
The key innovation of our work is the proposed dual-lifting method, which produces high-fidelity 3D Gaussians that capture identity and facial details.
Additionally, we leverage global image features and the 3D morphable model to construct 3D Gaussians for controlling expressions.
After training, our model can reconstruct unseen identities without specific optimizations and perform reenactment rendering at real-time speeds.
Experiments show that our method exhibits superior performance compared to previous methods in terms of reconstruction quality and expression accuracy.
We believe our method can establish new benchmarks for future research and advance applications of digital avatars. Xuangeng Chu, Tatsuya Harada |
NeurIPS | 1 |
| 2023 | Accurate 3D Face Reconstruction with Facial Component TokensabstractAccurately reconstructing 3D faces from monocular images and videos is crucial for various applications, such as digital avatar creation. However, the current deep learning-based methods face significant challenges in achieving accurate reconstruction with disentangled facial parameters and ensuring temporal stability in single-frame methods for 3D face tracking on video data. In this paper, we propose TokenFace, a transformer-based monocular 3D face reconstruction model. TokenFace uses separate tokens for different facial components to capture information about different facial parameters and employs temporal transformers to capture temporal information from video data. This design can naturally disentangle different facial components and is flexible to both 2D and 3D training data. Trained on hybrid 2D and 3D data, our model shows its power in accurately reconstructing faces from images and producing stable results for video data. Experimental results on popular benchmarks NoWand Stirling demonstrate that TokenFace achieves state-of-the-art performance, outperforming existing methods on all metrics by a large margin. Tianke Zhang, Xuangeng Chu, Yunfei Liu 0001, Lijian Lin, Zhendong Yang, Zhengzhuo Xu, Chengkun Cao, F. Richard Yu, Changyin Zhou, Chun Yuan 0003, Yu Li 0003 |
ICCV | 2 |
| 2020 | Detection in Crowded Scenes: One Proposal, Multiple PredictionsabstractWe propose a simple yet effective proposal-based object detector, aiming at detecting highly-overlapped instances in crowded scenes. The key of our approach is to let each proposal predict a set of correlated instances rather than a single one in previous proposal-based frameworks. Equipped with new techniques such as EMD Loss and Set NMS, our detector can effectively handle the difficulty of detecting highly overlapped objects. On a FPN-Res50 baseline, our detector can obtain 4.9\% AP gains on challenging CrowdHuman dataset and 1.0\% $\text{MR}^{-2}$ improvements on CityPersons dataset, without bells and whistles. Moreover, on less crowed datasets like COCO, our approach can still achieve moderate improvement, suggesting the proposed method is robust to crowdedness. Xuangeng Chu, Anlin Zheng, Xiangyu Zhang 0005, Jian Sun 0001 |
CVPR | 1 |