EDBT 2026 Demo / reviewers in the wild / expert
Di Chang
dblp:122/2664
· DBLP profile ↗
15ranked-venue papers
5as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning-Based Multi-View Stereo: A Surveyabstract3D reconstruction aims to recover the dense 3D structure of a scene. It plays an essential role in various applications such as Augmented/Virtual Reality (AR/VR), autonomous driving and robotics. Leveraging multiple views of a scene captured from different viewpoints, Multi-View Stereo (MVS) algorithms synthesize a comprehensive 3D representation, enabling precise reconstruction in complex environments. Due to its efficiency and effectiveness, MVS has become a pivotal method for image-based 3D reconstruction. Recently, with the success of deep learning, many learning-based MVS methods have been proposed, achieving impressive performance against traditional methods. We categorize these learning-based methods as: depth map-based, voxel-based, NeRF-based, 3D Gaussian Splatting-based, and large feed-forward methods. Among these, we focus significantly on depth map-based methods, which are the main family of MVS due to their conciseness, flexibility and scalability. In this survey, we provide a comprehensive review of the literature at the time of this writing. We investigate these learning-based methods, summarize their performances on popular benchmarks, and discuss promising future research directions in this area. Fangjinhua Wang, Qingtian Zhu, Di Chang, Quankai Gao, Junlin Han, Tong Zhang 0023, Richard I. Hartley, Marc Pollefeys |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | X-Dyna: Expressive Dynamic Human Image AnimationabstractWe introduce X-Dyna, a novel zero-shot, diffusion-based pipeline for animating a single human image using facial expressions and body movements derived from a driving video, that generates realistic, context-aware dynamics for both the subject and the surrounding environment. Building on prior approaches centered on human pose control, X-Dyna addresses key shortcomings causing the loss of dynamic details, enhancing the lifelike qualities of human video animations. At the core of our approach is the Dynamics-Adapter, a lightweight module that effectively integrates reference appearance context into the spatial attentions of the diffusion backbone while preserving the capacity of motion modules in synthesizing fluid and intricate dynamic details. Beyond body pose control, we connect a local control module with our model to capture identity-disentangled facial expressions, facilitating accurate expression transfer for enhanced realism in animated scenes. Together, these components form a unified framework capable of learning physical human motion and natural scene dynamics from a diverse blend of human and scene videos. Comprehensive qualitative and quantitative evaluations demonstrate that X-Dyna outperforms state-of- the-art methods, creating highly lifelike and expressive animations. The code is available at https://github.com/bytedance/X-Dyna Di Chang, You Xie, Yipeng Gao, Zhengfei Kuang, Shengqu Cai, Guoxian Song, Chao Wang 0088, Yichun Shi, Shijie Zhou 0003, Linjie Luo, Gordon Wetzstein, Mohammad Soleymani 0001 |
CVPR | 1 |
| 2025 | X-Dancer: Expressive Music to Human Dance Video GenerationabstractWe present X-Dancer, a novel zero-shot music-driven image animation pipeline that creates diverse and long-range lifelike human dance videos from a single static image. As its core, we introduce a unified transformer-diffusion framework, featuring an autoregressive transformer model that synthesize extended and music-synchronized token sequences for 2D body, head and hands poses, which then guide a diffusion model to produce coherent and realistic dance video frames. Unlike traditional methods that primarily generate human motion in 3D, X-Dancer addresses data limitations and enhances scalability by modeling a wide spectrum of 2D dance motions, capturing their nuanced alignment with musical beats through readily available monocular videos. To achieve this, we first build a spatially compositional token representation from 2D human pose labels associated with keypoint confidences, encoding both large articulated body movements (e.g., upper and lower body) and fine-grained motions (e.g., head and hands). We then design a music-to-motion transformer model that autoregressively generates music-aligned dance pose token sequences, incorporating global attention to both musical style and prior motion context. Finally we leverage a diffusion backbone to animate the reference image with these synthesized pose tokens through AdaIN, forming a fully differentiable end-to-end framework. Experimental results demonstrate that X-Dancer is able to produce both diverse and characterized dance videos, substantially outperforming state-of-the-art methods in term of diversity, expressiveness and realism. Code and model will be available for research purposes. Guoxian Song, You Xie, Xin Chen 0040, Chao Wang 0088, Di Chang, Linjie Luo |
ICCV | 8 |
| 2025 | Ditailistener: Controllable High Fidelity Listener Video Generation with DiffusionabstractGenerating naturalistic and nuanced listener motions for extended interactions remains an open problem. Existing methods often rely on low-dimensional motion codes for facial behavior generation followed by photorealistic rendering, limiting both visual fidelity and expressive richness. To address these challenges, we introduce DiTaiListener, powered by a video diffusion model with multimodal conditions. Our approach first generates short segments of listener responses conditioned on the speaker's speech and facial motions with DiTaiListener-Gen. It then refines the transitional frames via DiTaiListener-Edit for a seamless transition. Specifically, DiTaiListener-Gen adapts a Diffusion Transformer (DiT) for the task of listener head portrait generation by introducing a Causal Temporal Multimodal Adapter (CTM-Adapter) to process speakers' auditory and visual cues. CTM-Adapter integrates speakers' input in a causal manner into the video generation process to ensure temporally coherent listener responses. For long-form video generation, we introduce DiTaiListener-Edit, a transition refinement video-to-video diffusion model. The model fuses video segments into smooth and continuous videos, ensuring temporal consistency in facial expressions and image quality when merging short video segments produced by DiTaiListener-Gen. Quantitatively, DiTaiListener achieves the state-of-the-art performance on benchmark datasets in both photorealism (+73.8% in FID on RealTalk) and motion representation (+6.1% in FD metric on VICO) spaces. User studies confirm the superior performance of DiTaiListener, with the model being the clear preference in terms of feedback, diversity, and smoothness, outperforming competitors by a significant margin. Maksim Siniukov, Di Chang, Minh Tran 0004, Hongkun Gong, Ashutosh Chaubey, Mohammad Soleymani 0001 |
ICCV | 2 |
| 2025 | VLM4D: Towards Spatiotemporal Awareness in Vision Language ModelsabstractVision language models (VLMs) have shown remarkable capabilities in integrating linguistic and visual reasoning but remain fundamentally limited in understanding dynamic spatiotemporal interactions. Humans effortlessly track and reason about object movements, rotations, and perspective shifts-abilities essential for robust dynamic real-world understanding yet notably lacking in current VLMs. In this paper, we introduce VLM4D, the first benchmark specifically designed to evaluate the spatiotemporal reasoning capabilities of VLMs. Our benchmark comprises diverse real-world and synthetic videos accompanied by carefully curated question-answer pairs emphasizing translational and rotational motions, perspective awareness, and motion continuity. Through comprehensive evaluations of state-of-the-art open and closed-source VLMs, we identify significant performance gaps compared to human baselines, highlighting fundamental deficiencies in existing models. Extensive analysis reveals that VLMs struggle particularly with integrating multiple visual cues and maintaining temporal coherence. We further explore promising directions, such as leveraging 4D feature field reconstruction and targeted spatiotemporal supervised fine-tuning, demonstrating their effectiveness in enhancing spatiotemporal comprehension. Our work aims to encourage deeper exploration into improving VLMs' spatial and temporal grounding, paving the way towards more capable and reliable visual intelligence for dynamic environments. Shijie Zhou 0003, Alexander Vilesov, Xuehai He, Ziyu Wan, Shuwang Zhang, Aditya Nagachandra, Di Chang, Xin Wang 0061, Achuta Kadambi |
ICCV | 7 |
| 2024 | DiffPortrait3D: Controllable Diffusion for Zero-Shot Portrait View SynthesisabstractWe present DiffPortrait3D, a conditional diffusion model that is capable of synthesizing 3D-consistent photo-realistic novel views from as few as a single in-the-wild portrait. Specifically, given a single RGB input, we aim to synthesize plausible but consistent facial details rendered from novel camera views with retained both identity and facial expression. In lieu of time-consuming optimization and fine-tuning, our zero-shot method generalizes well to arbitrary face portraits with unposed camera views, extreme facial expressions, and diverse artistic depictions. At its core, we leverage the generative prior of 2D diffusion models pre-trained on large-scale image datasets as our rendering backbone, while the denoising is guided with disentangled attentive control of appearance and camera pose. To achieve this, we first inject the appearance context from the reference image into the self-attention layers of the frozen UNets. The rendering view is then manipulated with a novel conditional control module that interprets the camera pose by watching a condition image of a crossed subject from the same view. Furthermore, we insert a trainable cross-view attention module to enhance view consistency, which is further strengthened with a novel 3D-aware noise generation process during inference. We demonstrate state-of-the-art results both qualitatively and quantitatively on our challenging in-the-wild and multi-view benchmarks. Yuming Gu, You Xie, Guoxian Song, Yichun Shi, Di Chang, Linjie Luo |
CVPR | 6 |
| 2024 | DIM: Dyadic Interaction Modeling for Social Behavior Generation
Minh Tran 0004, Di Chang, Maksim Siniukov, Mohammad Soleymani 0001 |
ECCV (37) | 2 |
| 2024 | MagicPose: Realistic Human Poses and Facial Expressions Retargeting with Identity-aware DiffusionabstractIn this work, we propose MagicPose, a diffusion-based model for 2D human pose and facial expression retargeting. Specifically, given a reference image, we aim to generate a person’s new images by controlling the poses and facial expressions while keeping the identity unchanged. To this end, we propose a two-stage training strategy to disentangle human motions and appearance (e.g., facial expressions, skin tone, and dressing), consisting of (1) the pre-training of an appearance-control block and (2) learning appearance-disentangled pose control. Our novel design enables robust appearance control over generated human images, including body, facial attributes, and even background. By leveraging the prior knowledge of image diffusion models, MagicPose generalizes well to unseen human identities and complex poses without the need for additional fine-tuning. Moreover, the proposed model is easy to use and can be considered as a plug-in module/extension to Stable Diffusion. The project website is here. The code is available here. Di Chang, Yichun Shi, Quankai Gao, Jessica Fu, Guoxian Song, Yizhe Zhu, Mohammad Soleymani 0001 |
ICML | 1 |
| 2024 | FG-Net: Facial Action Unit Detection with Generalizable Pyramidal FeaturesabstractAutomatic detection of facial Action Units (AUs) allows for objective facial expression analysis. Due to the high cost of AU labeling and the limited size of existing benchmarks, previous AU detection methods tend to overfit the dataset, resulting in a significant performance loss when evaluated across corpora. To address this problem, we propose FG-Net for generalizable facial action unit detection. Specifically, FG-Net extracts feature maps from a Style-GAN2 model pre-trained on a large and diverse face image dataset. Then, these features are used to detect AUs with a Pyramid CNN Interpreter, making the training efficient and capturing essential local features. The proposed FG-Net achieves a strong generalization ability for heatmap-based AU detection thanks to the generalizable and semantic-rich features extracted from the pre-trained generative model. Extensive experiments are conducted to evaluate within- and cross-corpus AU detection with the widely-used DISFA and BP4D datasets. Compared with the state-of-the-art, the proposed method achieves superior cross-domain performance while maintaining competitive within-domain performance. In addition, FG-Net is dataefficient and achieves competitive performance even when trained on 1000 samples. Our code will be released at https://github.com/ihp-lab/FG-Net Yufeng Yin 0002, Di Chang, Guoxian Song, Shen Sang, Tiancheng Zhi, Linjie Luo, Mohammad Soleymani 0001 |
WACV | 2 |
| 2024 | LibreFace: An Open-Source Toolkit for Deep Facial Expression AnalysisabstractFacial expression analysis is an important tool for human-computer interaction. In this paper, we introduce LibreFace, an open-source toolkit for facial expression analysis. This open-source toolbox offers real-time and offline analysis of facial behavior through deep learning models, including facial action unit (AU) detection, AU intensity estimation, and facial expression recognition. To accomplish this, we employ several techniques, including the utilization of a large-scale pre-trained network, feature-wise knowledge distillation, and task-specific fine-tuning. These approaches are designed to effectively and accurately analyze facial expressions by leveraging visual information, thereby facilitating the implementation of real-time interactive applications. In terms of Action Unit (AU) intensity estimation, we achieve a Pearson Correlation Coefficient (PCC) of 0.63 on DISFA, which is 7% higher than the performance of OpenFace 2.0 [4] while maintaining highly-efficient inference that runs two times faster than OpenFace 2.0 [4]. Despite being compact, our model also demonstrates competitive performance to state-of-the-art facial expression analysis methods on AffecNet, FFHQ, and RAF-DB. Our code will be released at https://github.com/ihp-lab/LibreFace Di Chang, Yufeng Yin 0002, Zongjian Li, Minh Tran 0004, Mohammad Soleymani 0001 |
WACV | 1 |
| 2022 | Generalized Binary Search Network for Highly-Efficient Multi-View StereoabstractMulti-view Stereo (MVS) with known camera parameters is essentially a 1D search problem within a valid depth range. Recent deep learning-based MVS methods typically densely sample depth hypotheses in the depth range, and then construct prohibitively memory-consuming 3D cost volumes for depth prediction. Although coarse-to-fine sampling strategies alleviate this overhead issue to a certain extent, the efficiency of MVS is still an open challenge. In this work, we propose a novel method for highly efficient MVS that remarkably decreases the memory footprint, meanwhile clearly advancing state-of-the-art depth prediction performance. We investigate what a search strategy can be reasonably optimal for MVS taking into account of both efficiency and effectiveness. We first formulate MVS as a binary search problem, and accordingly propose a generalized binary search network for MVS. Specifically, in each step, the depth range is split into 2 bins with extra 1 error tolerance bin on both sides. A classification is performed to identify which bin contains the true depth. We also design three mechanisms to respectively handle classification errors, deal with out-of-range samples and decrease the training memory. The new formulation makes our method only sample a very small number of depth hypotheses in each step, which is highly memory efficient, and also greatly facilitates quick training convergence. Experiments on competitive benchmarks show that our method achieves state-of-the-art accuracy with much less memory. Particularly, our method obtains an overall score of 0.289 on DTU dataset and tops the first place on challenging Tanks and Temples advanced dataset among all the learning-based methods. Our code will be released at https://github.com/MiZhenxing/GBi-Net. Zhenxing Mi, Di Chang, Dan Xu 0002 |
CVPR | 2 |
| 2022 | RC-MVSNet: Unsupervised Multi-View Stereo with Neural Rendering
Di Chang, Aljaz Bozic, Tong Zhang 0023, Qingsong Yan, Ying-Cong Chen, Sabine Süsstrunk, Matthias Nießner |
ECCV (31) | 1 |
| 2021 | Fused-like angles: replacement for roll-pitch-yaw angles for a six-degree-of-freedom grating interferometerabstractRepresentation of orientation is important in a six-degree-of-freedom grating interferometer but only a few studies have focused on this topic. Roll-pitch-yaw angles, widely used in aviation, navigation, and robotics, are now being brought to the field of multi-degree-of-freedom interferometric measurement. However, the roll-pitch-yaw angles are not the exact definitions the metrologists expected in interferometry, because they require a certain sequential order of rotations and may cause errors in describing complicated rotations. The errors increase as the tip and tilt angles of the grating increase. Therefore, a replacement based on fused angles in robotics is proposed and named “fused-like angles.” The fused-like angles are error-free, so they are more in line with the definitions in grating interferometry and more suitable for six-degree-of-freedom measurements. Fused-like angles have already been used in research on the kinematic model and decoupling algorithm of the six-degree-of-freedom grating interferometer. Di Chang, Pengcheng Hu 0003, Jiubin Tan |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2020 | scHaplotyper: haplotype construction and visualization for genetic diagnosis using single cell DNA sequencing dataabstractBACKGROUND: Haplotyping reveals chromosome blocks inherited from parents to in vitro fertilized (IVF) embryos in preimplantation genetic diagnosis (PGD), enabling the observation of the transmission of disease alleles between generations. However, the methods of haplotyping that are suitable for single cells are limited because a whole genome amplification (WGA) process is performed before sequencing or genotyping in PGD, and true haplotype profiles of embryos need to be constructed based on genotypes that can contain many WGA artifacts. RESULTS: Here, we offer scHaplotyper as a genetic diagnosis tool that reconstructs and visualizes the haplotype profiles of single cells based on the Hidden Markov Model (HMM). scHaplotyper can trace the origin of each haplotype block in the embryo, enabling the detection of carrier status of disease alleles in each embryo. We applied this method in PGD in two families affected with genetic disorders, and the result was the healthy live births of two children in the two families, demonstrating the clinical application of this method. CONCLUSION: Next generation sequencing (NGS) of preimplantation embryos enable genetic screening for families with genetic disorders, avoiding the birth of affected babies. With the validation and successful clinical application, we showed that scHaplotyper is a convenient and accurate method to screen out embryos. More patients with genetic disorder will benefit from the genetic diagnosis of embryos. The source code of scHaplotyper is available at GitHub repository: https://github.com/yzqheart/scHaplotyper. Yanli Nie, Shuo Guan, Ying Kuo, Di Chang, Jie Qiao, Liying Yan |
BMC Bioinform. | 7 |
| 2019 | Displacement measuring grating interferometer: a reviewabstractA grating interferometer, called the “optical encoder,” is a commonly used tool for precise displacement measurements. In contrast to a laser interferometer, a grating interferometer is insensitive to the air refractive index and can be easily applied to multi-degree-of-freedom measurements, which has made it an extensively researched and widely used device. Classified based on the measuring principle and optical configuration, a grating interferometer experiences three distinct stages of development: homodyne, heterodyne, and spatially separated heterodyne. Compared with the former two, the spatially separated heterodyne grating interferometer could achieve a better resolution with a feature of eliminating periodic nonlinear errors. Meanwhile, numerous structures of grating interferometers with a high optical fold factor, a large measurement range, good usability, and multi-degree-of-freedom measurements have been investigated. The development of incremental displacement measuring grating interferometers achieved in recent years is summarized in detail, and studies on error analysis of a grating interferometer are briefly introduced. Pengcheng Hu 0003, Di Chang, Jiubin Tan, Ruitao Yang, Hongxing Yang, Haijin Fu |
Frontiers Inf. Technol. Electron. Eng. | 2 |