EDBT 2026 Demo / reviewers in the wild / expert
Haiwei Chen
dblp:194/4223
· DBLP profile ↗
13ranked-venue papers
6as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bringing Diversity from Diffusion Models to Semantic-Guided Face Asset GenerationabstractHigh-quality 3D face asset creation remains costly due to reliance on controlled capture setups and manual processing, limiting scalability and diversity. We introduce a fully automated, semantically controllable framework for generating PBR-ready 3D facial assets without requiring dedicated scans. Our pipeline begins with a diffusion-based data synthesis stage, where 2D portrait samples from a pre-trained diffusion model are converted into 44K textured 3D face reconstructions via our proposed geometry recovery and texture normalization algorithm, which aligns arbitrarily shaded outputs into clean albedo space. Using this dataset, we train a disentangled adversarial generator that maps semantic attributes (age, gender, ethnicity) to UV-space geometry and albedo, enabling both direct sampling and continuous latent editing while preserving identity. A refinement stage further produces PBR materials and secondary assets (eyeballs, teeth, gums). The resulting system supports controllable face generation and post-editing in real time and exports directly to standard rendering and animation pipelines. We evaluate each component extensively and provide a web-based interactive interface to showcase practical deployment. Yunxuan Cai, Sitao Xiang, Zongjian Li, Haiwei Chen |
ACM Trans. Graph. | 4 |
| 2025 | Geometry-Aware Feature Matching for Large-Scale Structure from MotionabstractEstablishing consistent and dense correspondences across multiple images is crucial for Structure from Motion (SfM) systems. Significant view changes, such as air-to-ground with very sparse view overlap, pose an even greater challenge to the correspondence solvers. We present a novel optimization-based approach that significantly enhances existing feature matching methods by introducing geometry cues in addition to color cues. This helps fill gaps when there is less overlap in large-scale scenarios. Our method formulates geometric verification as an optimization problem, guiding feature matching within detector-free methods and using sparse correspondences from detector-based methods as anchor points. By enforcing geometric constraints via the Sampson Distance, our approach ensures that the denser correspondences from detector-free methods are geometrically consistent and more accurate. This hybrid strategy significantly improves correspondence density and accuracy, mitigates multi-view inconsistencies, and leads to notable advancements in camera pose accuracy and point cloud density. It outperforms state-of-the-art feature matching methods on benchmark datasets and enables feature matching in challenging extreme large-scale settings. Project page: https://xtcpete.github.io/geo-website/. Gonglin Chen, Jinsen Wu, Haiwei Chen, Wenbin Teng, Andrew Feng, Rongjun Qin |
3DV | 3 |
| 2025 | RDD: Robust Feature Detector and Descriptor using Deformable TransformerabstractAs a core step in structure-from-motion and SLAM, robust feature detection and description under challenging scenarios such as significant viewpoint changes remain unresolved despite their ubiquity. While recent works have identified the importance of local features in modeling geometric transformations, these methods fail to learn the visual cues present in long-range relationships. We present Robust Deformable Detector (RDD), a novel and robust keypoint detector/descriptor leveraging the deformable transformer, which captures global context and geometric invariance through deformable self-attention mechanisms. Specifically, we observed that deformable attention focuses on key locations, effectively reducing the search space complexity and modeling the geometric invariance. Furthermore, we collected an Air-to-Ground dataset for training in addition to the standard MegaDepth dataset. Our proposed method outperforms all state-of-the-art keypoint detection/description methods in sparse matching tasks and is also capable of semi-dense matching. To ensure comprehensive evaluation, we introduce two challenging benchmarks: one emphasizing large viewpoint and scale variations, and the other being an Air-to-Ground benchmark — an evaluation setting that has recently gaining popularity for 3D reconstruction across different altitudes. Project page: https://xtcpete.github.io/rdd/. Gonglin Chen, Tianwen Fu, Haiwei Chen, Wenbin Teng, Hanyuan Xiao |
CVPR | 3 |
| 2025 | FVGen: Accelerating Novel-View Synthesis with Adversarial Video Diffusion DistillationabstractRecent progress in 3D reconstruction has enabled realistic 3D models from dense image captures, yet challenges persist with sparse views, often leading to artifacts in unseen areas. Recent works leverage Video Diffusion Models (VDMs) to generate dense observations, filling the gaps when only sparse views are available for 3D reconstruction tasks. A significant limitation of these methods is their slow sampling speed when using VDMs. In this paper, we present FVGen, a novel framework that addresses this challenge by enabling fast novel view synthesis using VDMs in as few as four sampling steps. We propose a novel video diffusion model distillation method that distills a multi-step denoising teacher model into a few-step denoising student model using Generative Adversarial Networks (GANs) and softened reverse KL-divergence minimization. Extensive experiments on real-world datasets show that, compared to previous works, our framework generates the same number of novel views with similar (or even better) visual quality while reducing sampling time by more than 90%. FVGen significantly improves time efficiency for downstream reconstruction tasks, particularly when working with sparse input views (more than 2) where pre-trained VDMs need to be run multiple times to achieve better spatial coverage. Wenbin Teng, Gonglin Chen, Haiwei Chen |
ICCV | 3 |
| 2024 | Don't Look into the Dark: Latent Codes for Pluralistic Image InpaintingabstractWe present a method for large-mask pluralistic image in-painting based on the generative framework of discrete latent codes. Our method learns latent priors, discretized as tokens, by only performing computations at the visible locations of the image. This is realized by a restrictive partial encoder that predicts the token label for each visible block, a bidirectional transformer that infers the missing labels by only looking at these tokens, and a dedicated synthesis network that couples the tokens with the partial image priors to generate coherent and pluralistic complete image even under extreme mask settings. Experiments on public benchmarks validate our design choices as the proposed method outperforms strong baselines in both visual quality and diversity metrics. Haiwei Chen |
CVPR | 1 |
| 2023 | AutoBCS: Block-Based Image Compressive Sensing With Data-Driven Acquisition and Noniterative ReconstructionabstractBlock compressive sensing (CS) is a well-known signal acquisition and reconstruction paradigm with widespread application prospects in science, engineering, and cybernetic systems. However, state-of-the-art block-based image CS (BCS) methods generally suffer from two issues. The sparsifying domain and the sensing matrices widely used for image acquisition are not data driven and, thus, both the features of the image and the relationships among subblock images are ignored. Moreover, it requires to address a high-dimensional optimization problem with extensive computational complexity for image reconstruction. In this article, we provide a deep learning (DL) strategy for BCS, called AutoBCS, which automatically takes the prior knowledge of images into account in the acquisition step and establishes a reconstruction model for performing fast image reconstruction. More precisely, we present a learning-based sensing matrix to accomplish image acquisition, thereby capturing and preserving more image characteristics than those captured by the existing methods. In addition, we build a noniterative reconstruction network, which provides an end-to-end BCS reconstruction framework to maximize image reconstruction efficiency. Furthermore, we investigate comprehensive comparison studies with both traditional BCS approaches and newly developed DL methods. Compared with these approaches, our proposed AutoBCS can not only provide superior performance in terms of image quality metrics (SSIM and PSNR) and visual perception but also automatically benefit reconstruction speed. Hongping Gan, Yang Gao 0030, Chunyi Liu, Haiwei Chen, Tao Zhang 0027, Feng Liu 0005 |
IEEE Trans. Cybern. | 4 |
| 2022 | Exemplar-based Pattern Synthesis with Implicit Periodic Field NetworkabstractSynthesis of ergodic, stationary visual patterns is widely applicable in texturing, shape modeling, and digital content creation. The wide applicability of this technique thus requires the pattern synthesis approaches to be scalable, diverse, and authentic. In this paper, we propose an exemplar-based visual pattern synthesis framework that aims to model the inner statistics of visual patterns and generate new, versatile patterns that meet the aforementioned requirements. To this end, we propose an implicit network based on generative adversarial network (GAN) and periodic encoding, thus calling our network the Implicit Periodic Field Network (IPFN). The design of IPFN ensures scalability: the implicit formulation directly maps the input coordinates to features, which enables synthesis of arbitrary size and is computationally efficient for 3D shape synthesis. Learning with a periodic encoding scheme encourages diversity: the network is constrained to model the inner statistics of the exemplar based on spatial latent codes in a periodic field. Coupled with continuously designed GAN training procedures, IPFN is shown to synthesize tileable patterns with smooth transitions and local variations. Last but not least, thanks to both the adversarial training technique and the encoded Fourier features, IPFN learns high-frequency functions that produce authentic, high-quality results. To validate our approach, we present novel experimental results on various applications in 2D texture synthesis and 3D shape synthesis. Haiwei Chen, Weikai Chen 0001, Shichen Liu |
CVPR | 1 |
| 2022 | Rapid Face Asset Acquisition with Recurrent Feature AlignmentabstractWe present Re current F eature A lignment (ReFA), an end-to-end neural network for the very rapid creation of production-grade face assets from multi-view images. ReFA is on par with the industrial pipelines in quality for producing accurate, complete, registered, and textured assets directly applicable to physically-based rendering, but produces the asset end-to-end, fully automatically at a significantly faster speed at 4.5 FPS, which is unprecedented among neural-based techniques. Our method represents face geometry as a position map in the UV space. The network first extracts per-pixel features in both the multi-view image space and the UV space. A recurrent module then iteratively optimizes the geometry by projecting the image-space features to the UV space and comparing them with a reference UV-space feature. The optimized geometry then provides pixel-aligned signals for the inference of high-resolution textures. Experiments have validated that ReFA achieves a median error of 0.603 mm in geometry reconstruction, is robust to extreme pose and expression, and excels in sparse-view settings. We believe that the progress achieved by our network enables lightweight, fast face assets acquisition that significantly boosts the downstream applications, such as avatar creation and facial performance capture. It will also enable massive database capturing for deep learning purposes. Shichen Liu, Yunxuan Cai, Haiwei Chen, Yichao Zhou 0003 |
ACM Trans. Graph. | 3 |
| 2021 | Equivariant Point Network for 3D Point Cloud AnalysisabstractFeatures that are equivariant to a larger group of symmetries have been shown to be more discriminative and powerful in recent studies [4], [40], [5]. However, higher-order equivariant features often come with an exponentially-growing computational cost. Furthermore, it remains relatively less explored how rotation-equivariant features can be leveraged to tackle 3D shape alignment tasks. While many past approaches have been based on either non-equivariant or invariant descriptors to align 3D shapes, we argue that such tasks may benefit greatly from an equivariant framework. In this paper, we propose an effective and practical SE(3) (3D translation and rotation) equivariant network for point cloud analysis that addresses both problems. First, we present SE(3) separable point convolution, a novel framework that breaks down the 6D convolution into two separable convolutional operators alternatively performed in the 3D Euclidean and SO(3) spaces respectively. This significantly reduces the computational cost without compromising the performance. Second, we introduce an attention layer to effectively harness the expressiveness of the equivariant features. While jointly trained with the network, the attention layer implicitly derives the intrinsic local frame in the feature space and generates attention vectors that can be integrated with different alignment tasks. We evaluate our approach through extensive studies and visual interpretations. The empirical results demonstrate that our proposed model outperforms strong baselines in a variety of benchmarks. Code is available at https://github.com/nintendops/EPN_PointCloud. Haiwei Chen, Shichen Liu, Weikai Chen 0001, Hao Li 0015, Randall W. Hill Jr. |
CVPR | 1 |
| 2018 | Redirected Walking in Irregularly Shaped Physical Environments with Dynamic ObstaclesabstractRedirected walking (RDW) is a virtual reality (VR) locomotion technique that enables the exploration of a large virtual environment (VE) within a small physical space via real walking. Thus far, the physical environment has generally been assumed to be rectangular, static, and free of obstacles. However, it is unlikely that real-world locations that may be used for VR fulfill these constraints. In addition, accounting for dynamic obstacles such as people helps increase user safety when the view of the physical world is occluded by a head-mounted display. In this work, we present the design and initial implementation of a RDW planning algorithm that can redirect the user in an irregularly shaped physical environment with dynamically moving obstacles. This represents an important step towards the use of RDW in more dynamic, real-world environments. Haiwei Chen, Samantha Chen 0003, Evan A. Suma |
VR | 1 |
| 2017 | Supporting free walking in a large virtual environment: imperceptible redirected walking with an immersive distractorabstractRedirected walking, a technique in which the user's orientation in the physical space is constantly and imperceptibly changed from their orientation in the virtual world, has been shown to be an effective technique when only a limited physical space is available. Unfortunately, previous efforts have restricted redirected walking applications to operate under the constant supervision of researchers to prevent the users from leaving the tracked area. In addition, Virtual Environments (VE) used in these applications were often limited to narrow hallways, mazes or predefined waypoints, while the performance of redirected walking in a large, open VE is not well explored. In this paper, we introduce the idea and implementation for an imperceptible redirected walking system that supports the illusion of free walking in a large, open virtual environment with minimal amount of physical interventions, by integrating the distractor into the user's main immersive activity in the VE. We demonstrate this new approach with two user studies of an immersive interactive game. Our study indicates that for the majority of the subjects, the illusion is maintained of unconstrained walking in a very large area (a full-size basketball court, 50 feet × 95 feet), even while they were limited to a physical area of a mere 6% of the size of the basketball court (16 feet × 16 feet tracked area). Our result demonstrates that the illusion of free walking is created, since a majority of them was not interrupted by the researchers and did not realize they were redirected, and the 34 subjects took vastly different routes to reach the distant goal (See Figure 2). We believe that this technique demonstrates a more immersive way of designing redirected walking application and shows possibility of bringing redirected walking applications out of the monitored lab environments. Our result may provide insights for the designers of immersive experiences to create other redirected walking applications for consumer VR systems with room-size tracking. Haiwei Chen, Henry Fuchs |
CGI | 1 |
| 2017 | Towards imperceptible redirected walking: integrating a distractor into the immersive experienceabstractPhysically walking in a virtual world has been repeatedly demonstrated to be superior to navigation with game controllers and the like. The problem is that virtual worlds are often much larger than the available physical space. Redirected walking, a technique in which the users orientation in the physical space is constantly and imperceptibly changed from their orientation in the virtual world, has been shown to be an effective technique for walking in a limited physical space [Razzaque et al. 2001], but it needs frequent and rapid head rotation to keep the reorientation changes imperceptible. A rapidly moving object in the environment, a distractor, has been shown to be effective at inspiring the users to rapidly rotate their head, but such a distractor can be distracting to the user's main activity in the virtual world. We introduce the notion of the distractor being integrated into the user's main activity in the virtual world - in our example, a fire-breathing dragon into an immersive adventure game. With such an integrated character, the users do not need any special instructions about redirected walking; they are simply performing their intended activities. We report the results of a small (N=24) user study which indicates that for the majority of subjects (17 of 24) the illusion is maintained of unconstrained walking in a very large area (a full-sized basketball court, 45 × 90 feet) even while they were limited to a small, 16 × 16 feet region. We speculate that this technique may extend to many other applications in which the distractor can be integrated into the major immersive activity and thus enable the illusion of natural, unconstrained walking in large virtual worlds even when only a small physical space is available. Haiwei Chen, Henry Fuchs |
I3D | 1 |
| 2016 | The Design of a Patient-Centered Personal Health Record with Patients as Co-Designers
Arlene Chung, Haiwei Chen, Grace Shin, Ketan K. Mane, Hye-Chung Kum |
AMIA | 2 |