VLDB 2026 Research / reviewers in the wild / expert
Duomin Wang
dblp:285/9743
· DBLP profile ↗
6ranked-venue papers
1as first author
6since 2021 · last 2025
0009-0004-9507-6741ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Taming Teacher Forcing for Masked Autoregressive Video GenerationabstractWe introduce MAGI, a hybrid video generation framework that combines masked modeling for intra-frame generation with causal modeling for next-frame generation. Our key innovation, Complete Teacher Forcing (CTF), conditions masked frames on complete observation frames rather than masked ones (namely Masked Teacher Forcing, MTF), enabling a smooth transition from token-level (patch-level) to frame-level autoregressive generation. CTF significantly outperforms MTF, achieving a +23% improvement in FVD scores on first-frame conditioned video prediction. To address issues like exposure bias, we employ targeted training strategies, setting a new benchmark in autoregressive video generation. Experiments show that MAGI can generate long, coherent video sequences exceeding 100 frames, even when trained on as few as 16 frames, highlighting its potential for scalable, high-quality video generation. Yuang Peng, Kun Yan 0004, Runpei Dong, Duomin Wang, Zheng Ge, Nan Duan 0001, Xiangyu Zhang 0005 |
CVPR | 6 |
| 2024 | Portrait4D: Learning One-Shot 4D Head Avatar Synthesis using Synthetic DataabstractExisting one-shot 4D head synthesis methods usually learn from monocular videos with the aid of 3DMM re-construction, yet the latter is evenly challenging which re-stricts them from reasonable 4D head synthesis. We present a method to learn one-shot 4D head synthesis via large-scale synthetic data. The key is to first learn a part-wise 4D generative model from monocular images via adver-sarial learning, to synthesize multi-view images of diverse identities and full motions as training data; then leverage a transformer-based animatable triplane reconstructor to learn 4D head reconstruction using the synthetic data. A novel learning strategy is enforced to enhance the general-izability to real images by disentangling the learning pro-cess of 3D reconstruction and reenactment. Experiments demonstrate our superiority over the prior art. Duomin Wang, Xiaohang Ren, Baoyuan Wang |
CVPR | 2 |
| 2024 | PICTURE: PhotorealistIC Virtual Try-on from UnconstRained dEsignsabstractIn this paper, we propose a novel virtual try-on from unconstrained designs (ucVTON) task to enable photorealistic synthesis of personalized composite clothing on input human images. Unlike prior arts constrained by specific input types, our method allows flexible specification of style (text or image) and texture (full garment, cropped sections, or texture patches) conditions. To address the entanglement challenge when using full garment images as conditions, we develop a two-stage pipeline with explicit disentanglement of style and texture. In the first stage, we generate a human parsing map reflecting the desired style conditioned on the input. In the second stage, we composite textures onto the parsing map areas based on the texture input. To represent complex and non-stationary textures that have never been achieved in previous fashion editing works, we first propose extracting hierarchical and balanced CLIP features and applying position encoding in VTON. Experiments demonstrate superior synthesis quality and personalization enabled by our method. The flexible control over style and texture mixing brings virtual try-on to a new level of user experience for online shopping and fashion design. Shuliang Ning, Duomin Wang, Yipeng Qin, Zirong Jin, Baoyuan Wang, Xiaoguang Han 0001 |
CVPR | 2 |
| 2024 | Portrait4D-V2: Pseudo Multi-view Data Creates Better 4D Head Synthesizer
Yu Deng 0006, Duomin Wang, Baoyuan Wang |
ECCV (17) | 2 |
| 2023 | Progressive Disentangled Representation Learning for Fine-Grained Controllable Talking Head SynthesisabstractWe present a novel one-shot talking head synthesis method that achieves disentangled and fine-grained control over lip motion, eye gaze&blink, head pose, and emotional expression. We represent different motions via disentangled latent representations and leverage an image generator to synthesize talking heads from them. To effectively disentangle each motion factor, we propose a progressive disentangled representation learning strategy by separating the factors in a coarse-to-fine manner, where we first extract unified motion feature from the driving signal, and then isolate each fine-grained motion from the unified feature. We leverage motion-specific contrastive learning and regressing for non-emotional motions, and introduce feature-level decorrelation and self-reconstruction for emotional expression, to fully utilize the inherent properties of each motion factor in unstructured video data to achieve disentanglement. Experiments show that our method provides high quality speech&lip-motion synchronization along with precise and disentangled control over multiple extra facial motions, which can hardly be achieved by previous methods. Duomin Wang, Zixin Yin, Harry Shum, Baoyuan Wang |
CVPR | 1 |
| 2023 | Talking Head Generation with Probabilistic Audio-to-Visual Diffusion PriorsabstractWe introduce a novel framework for one-shot audio-driven talking head generation. Unlike prior works that require additional driving sources for controlled synthesis in a deterministic manner, we instead sample all holistic lip-irrelevant facial motions (i.e. pose, expression, blink, gaze, etc.) to semantically match the input audio while still maintaining both the photo-realism of audio-lip synchronization and overall naturalness. This is achieved by our newly proposed audio-to-visual diffusion prior trained on top of the mapping between audio and non-lip representations. Thanks to the probabilistic nature of the diffusion prior, one big advantage of our framework is it can synthesize diverse facial motion sequences given the same audio clip, which is quite user-friendly for many real applications. Through comprehensive evaluations of public benchmarks, we conclude that (1) our diffusion prior outperforms auto-regressive prior significantly on all the concerned metrics; (2) our overall system is competitive with prior works in terms of audio-lip synchronization but can effectively sample rich and natural-looking lip-irrelevant facial motions while still semantically harmonized with the audio input. Zhentao Yu, Zixin Yin, Duomin Wang, Finn Wong, Baoyuan Wang |
ICCV | 4 |