VLDB 2026 Research / reviewers in the wild / expert
Pengfei Wan 0001
dblp:119/0306-1
· DBLP profile ↗
85ranked-venue papers
8as first author
67since 2021 · last 2026
0000-0001-7225-565XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 61 · 8 first-author · 44 since 2021Artificial intelligence and machine learning · 51 · 51 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Boosting Resolution Generalization of Diffusion Transformers with Randomized Positional EncodingsabstractResolution generalization in image generation tasks enables the production of higher-resolution images with lower training resolution overhead. However, a key obstacle for diffusion transformers in addressing this problem is the mismatch between positional encodings seen at inference and those used during training. Existing strategies such as positional encodings interpolation, extrapolation, or hybrids, do not fully resolve this mismatch. In this paper, we propose a novel two-dimensional randomized positional encodings, namely RPE-2D, that prioritizes the order of image patches rather than their absolute distances, enabling seamless high- and low-resolution generation without training on multiple resolutions. Concretely, RPE-2D independently samples positions along the horizontal and vertical axes over an expanded range during training, ensuring that the encodings used at inference lie within the training distribution and thereby improving resolution generalization. We further introduce a simple random resize-and-crop augmentation to strengthen order modeling and add micro-conditioning to indicate the applied cropping pattern. On the ImageNet dataset, RPE-2D achieves state-of-the-art resolution generalization performance, outperforming competitive methods when trained at 256^2 and evaluated at 384^2 and 512^2, and when trained at 512^2 and evaluated at 768^2 and 1024^2. RPE-2D also exhibits outstanding capabilities in low-resolution image generation, multi-stage training acceleration, and multi-resolution inheritance. Mingwu Zheng, Xin Tao 0001, Pengfei Wan 0001, Di Zhang 0026, Kun Gai |
AAAI | 5 |
| 2026 | Bridging Cognitive Gap: Hierarchical Description Learning for Artistic Image Aesthetics AssessmentabstractThe aesthetic quality assessment task is crucial for developing a human-aligned quantitative evaluation system for AIGC. However, its inherently complex nature—spanning visual perception, cognition, and emotion—poses fundamental challenges. Although aesthetic descriptions offer a viable representation of this complexity, two critical challenges persist: (1) data scarcity and imbalance: existing dataset overly focuses on visual perception and neglects deeper dimensions due to the expensive manual annotation; and (2) model fragmentation: current visual networks isolate aesthetic attributes with multi-branch encoder, while multimodal methods represented by contrastive learning struggle to effectively process long-form textual descriptions. To resolve challenge (1), we first present the Refined Aesthetic Description (RAD) dataset, a large-scale (70k), multi-dimensional structured dataset, generated via an iterative pipeline without heavy annotation costs and easy to scale. To address challenge (2), we propose ArtQuant, an aesthetics assessment framework for artistic image which not only couple isolated aesthetic dimensions through joint description generation, but also better model long-text semantics with the help of LLM decoders. Besides, theoretical analysis confirms this symbiosis: RAD's semantic adequacy (data) and generation paradigm (model) collectively minimize prediction entropy, providing mathematical grounding for the framework. Our approach achieves state-of-the-art performance on several datasets while requiring only 33% of conventional training epochs, narrowing the cognitive gap between artistic image and aesthetic judgment. We will release both code and dataset to support future research. Henglin Liu, Nisha Huang, Chang Liu 0071, Jiangpeng Yan, Huijuan Huang 0001, Jixuan Ying, Tong-Yee Lee, Pengfei Wan 0001, Xiangyang Ji |
AAAI | 8 |
| 2026 | FilmWeaver: Weaving Consistent Multi-Shot Videos with Cache-Guided Autoregressive DiffusionabstractCurrent video generation models perform well at single-shot synthesis but struggle with multi-shot videos, facing critical challenges in maintaining character and background consistency across shots and flexibly generating videos of arbitrary length and shot count. To address these limitations, we introduce \textbf{FilmWeaver}, a novel framework designed to generate consistent, multi-shot videos of arbitrary length. First, it employs an autoregressive diffusion paradigm to achieve arbitrary-length video generation. To address the challenge of consistency, our key insight is to decouple the problem into inter-shot consistency and intra-shot coherence. We achieve this through a dual-level cache mechanism: a shot memory caches keyframes from preceding shots to maintain character and scene identity, while a temporal memory retains a history of frames from the current shot to ensure smooth, continuous motion. The proposed framework allows for flexible, multi-round user interaction to create multi-shot videos. Furthermore, due to this decoupled design, our method demonstrates high versatility by supporting downstream tasks such as multi-concept injection and video extension. To facilitate the training of our consistency-aware method, we also developed a comprehensive pipeline to construct a high-quality multi-shot video dataset. Extensive experimental results demonstrate that our method surpasses existing approaches on metrics for both consistency and aesthetic quality, opening up new possibilities for creating more consistent, controllable, and narrative-driven video content. Xiaokun Liu, Wenyu Qin, Meng Wang 0001, Pengfei Wan 0001, Di Zhang 0026, Kun Gai, Shao-Lun Huang |
AAAI | 7 |
| 2026 | MultiPaint: A Unified Framework for Multi-Task, Multi-Object, and Multi-Condition Video InpaintingabstractVideo inpainting modifies local regions in video while ensuring spatial and temporal coherence. However, existing methods-both traditional and recent diffusion-based ones-face key limitations: they lack unified support for both insertion and completion, and are restricted to single-object inpainting, making it difficult to handle multi-object scenarios involving grounding and interaction. In this article, we propose MultiPaint, a unified framework for multi-task, multi-object, and multi-condition video inpainting. First, we introduce dual-branch adapters to unify the insertion and completion tasks within a single model. Moreover, we propose a test-time scheduled feature composition strategy that enables multi-object inpainting with user-specified locations while better preserving interactions among objects, a setting that has been insufficiently addressed in prior work. Additionally, we introduce a multi-condition inpainting scheme that integrates text-guided, image-guided, and keyframe-guided modes via dynamic frame masking, providing more controllability in appearance customization. Extensive experiments show that MultiPaint achieves state-of-the-art performance on object insertion and scene completion among the recent works. We further demonstrate its versatility in downstream tasks including grounded video generation, object editing, object removal, image-guided inpainting, and long video inpainting. Zheng Gu 0001, Xin Tao 0001, Pengfei Wan 0001, Xiaodong Chen 0009, Jing Liao 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2026 | ExGes: Expressive Human Motion Retrieval and Modulation for Audio-Driven Gesture SynthesisabstractAudio-driven human gesture synthesis is a crucial task with broad applications in virtual avatars, human-computer interaction, and creative content generation. Despite notable progress, existing methods often produce coarse gestures, lack expressiveness, and fail to fully align with audio semantics. To address these challenges, we propose ExGes, a novel retrieval-enhanced diffusion framework with three key designs: (1) a Motion Base Construction, which builds a gesture library from the training dataset; (2) a Motion Retrieval Module, employing contrastive learning and momentum distillation for retrieving fine-grained reference poses; and (3) a Precise Control Module, integrating partial masking and stochastic masking to enable flexible and fine-grained control. Experimental evaluations on BEAT2 demonstrate that ExGes reduces Fréchet Gesture Distance by 4.55%and improves motion diversity by 5.3% over EMAGE, with user studies revealing a 71.3% preference for its naturalness and semantic relevance. Xukun Zhou, Fengxin Li, Yan Zhou 0003, Pengfei Wan 0001, Yeying Jin, Hongyuan Zhang 0001, Hongyan Liu 0002, Zhaoxin Fan, Jun He 0008, Xuelong Li 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | PatchVSR: Breaking Video Diffusion Resolution Limits with Patch-wise Video Super-ResolutionabstractPre-trained video generation models hold great potential for generative video super-resolution (VSR). However, adapting them for full-size VSR, as most existing methods do, suffers from unnecessary intensive full-attention computation and fixed output resolution. To overcome these limitations, we make the first exploration into utilizing video diffusion priors for patch-wise VSR. This is non-trivial because pre-trained video diffusion models are not native for patch-level detail generation. To mitigate this challenge, we propose an innovative approach, called PatchVSR, which integrates a dual-stream adapter for conditional guidance. The patch branch extracts features from input patches to maintain content fidelity while the global branch extracts context features from the resized full video to bridge the generation gap caused by incomplete semantics of patches. Particularly, we also inject the patch’s location information into the model to better contextualize patch synthesis within the global video frame. Experiments demonstrate that our method can synthesize high-fidelity, high-resolution details at the patch level. A tailor-made multi-patch joint modulation is proposed to ensure visual consistency across individually enhanced patches. Due to the flexibility of our patch-based paradigm, we can achieve highly competitive 4K VSR based on a 512×512 resolution base model, with extremely high efficiency. Shian Du, Menghan Xia, Chang Liu 0071, Xintao Wang 0002, Jing Wang 0021, Pengfei Wan 0001, Di Zhang 0026, Xiangyang Ji |
CVPR | 6 |
| 2025 | GPAvatar: High-fidelity Head Avatars by Learning Efficient Gaussian ProjectionsabstractExisting radiance field-based head avatar methods have mostly relied on pre-computed explicit priors (e.g., mesh, point) or neural implicit representations, making it challenging to achieve high fidelity with both computational efficiency and low memory consumption. To overcome this, we present GPAvatar, a novel and efficient Gaussian splatting-based method for reconstructing high-fidelity dynamic 3D head avatars from monocular videos. We extend Gaussians in 3D space to a high-dimensional embedding space encompassing Gaussian’s spatial position and avatar expression, enabling the representation of the head avatar with arbitrary pose and expression. To enable splatting-based rasterization, a linear transformation is learned to project each high-dimensional Gaussian back to the 3D space, which is sufficient to capture expression variations instead of using complex neural networks. Furthermore, we propose an adaptive densification strategy that dynamically allocates Gaussians to regions with high expression variance, improving the facial detail representation. Experimental results on three datasets show that our method outperforms existing state-of-the-art methods in rendering quality and speed while reducing memory usage in training and rendering. Wei-Qi Feng, Ze-Kang Zhou, Shunkai Li, Pengfei Wan 0001, Di Zhang 0026, Miao Wang 0004 |
CVPR | 6 |
| 2025 | SketchVideo: Sketch-based Video Generation and EditingabstractVideo generation and editing conditioned on text prompts or images have undergone significant advancements. However, challenges remain in accurately controlling global layout and geometry details solely by texts, and supporting motion control and local modification through images. In this paper, we aim to achieve sketch-based spatial and motion control for video generation and support fine-grained editing of real or synthetic videos. Based on the DiT video generation model, we propose a memory-efficient control structure with sketch control blocks that predict residual features of skipped DiT blocks. Sketches are drawn on one or two keyframes (at arbitrary time points) for easy interaction. To propagate such temporally sparse sketch conditions across all frames, we propose an inter-frame attention mechanism to analyze the relationship between the keyframes and each video frame. For sketch-based video editing, we design an additional video insertion module that maintains consistency between the newly edited content and the original video’s spatial feature and dynamic motion. During inference, we use latent fusion for the accurate preservation of unedited regions. Extensive experiments demonstrate that our SketchVideo achieves superior performance in controllable video generation and editing. Feng-Lin Liu, Hongbo Fu 0001, Xintao Wang 0004, Weicai Ye, Pengfei Wan 0001, Di Zhang 0026, Lin Gao 0004 |
CVPR | 5 |
| 2025 | Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene SimulationabstractRealistic simulation of dynamic scenes requires accurately capturing diverse material properties and modeling complex object interactions grounded in physical principles. However, existing methods are constrained to basic material types with limited predictable parameters, making them insufficient to represent the complexity of real-world materials. We introduce PhysFlow, a novel approach that leverages multi-modal foundation models and video diffusion to achieve enhanced 4D dynamic scene simulation. Our method utilizes multi-modal models to identify material types and initialize material parameters through image queries, while simultaneously inferring 3D Gaussian splats for detailed scene representation. We further refine these material parameters using video diffusion with a differentiable Material Point Method (MPM) and optical flow guidance rather than render loss or Score Distillation Sampling (SDS) loss. This integrated framework enables accurate prediction and realistic simulation of dynamic interactions in real-world scenarios, advancing both accuracy and flexibility in physics-based simulations. Our code and data are available at https://zhuomanliu.github.io/PhysFlow Zhuoman Liu, Weicai Ye, Yan Luximon, Pengfei Wan 0001, Di Zhang 0026 |
CVPR | 4 |
| 2025 | Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video ContentabstractWith the continuous progress of visual generation technologies, the scale of video datasets has grown exponentially. The quality of these datasets plays a pivotal role in the performance of video generation models. We assert that temporal splitting, detailed captions, and video quality filtering are three crucial determinants of dataset quality. However, existing datasets exhibit various limitations in these areas. To address these challenges, we introduce Koala-36M, a large-scale, high-quality video dataset featuring accurate temporal splitting, detailed captions, and superior video quality. The essence of our approach lies in improving the consistency between fine-grained conditions and video content. Specifically, we employ a linear classifier on probability distributions to enhance the accuracy of transition detection, ensuring better temporal consistency. We then provide structured captions for the splitted videos, with an average length of 200 words, to improve text-video alignment. Additionally, we develop a Video Training Suitability Score (VTSS) that integrates multiple sub-metrics, allowing us to filter high-quality videos from the original corpus. Finally, we incorporate several metrics into the training process of the generation model, further refining the fine-grained conditions. Our experiments demonstrate the effectiveness of our data processing pipeline and the quality of the proposed Koala-36M dataset. Our dataset and code have been released at https://koala36m.github.io/. Qiuheng Wang, Yukai Shi, Jiarong Ou, Boyuan Jiang, Mingwu Zheng, Xin Tao 0001, Pengfei Wan 0001, Di Zhang 0026 |
CVPR | 12 |
| 2025 | StyleMaster: Stylize Your Video with Artistic Generation and TranslationabstractStyle control has been popular in video generation models. Existing methods often generate videos far from the given style, cause content leakage, and struggle to transfer one video to the desired style. Our first observation is that the style extraction stage matters, whereas existing methods emphasize global style but ignore local textures. In order to bring texture features while preventing content leakage, we filter content-related patches while retaining style ones based on prompt-patch similarity; for global style extraction, we generate a paired style dataset through model illusion to facilitate contrastive learning, which greatly enhances the absolute style consistency. Moreover, to fill in the image-to-video gap, we train a lightweight motion adapter on still videos, which implicitly enhances stylization extent, and enables our image-trained model to be seamlessly applied to videos. Benefited from these efforts, our approach, StyleMaster, not only achieves significant improvement in both style resemblance and temporal coherence, but also can easily generalize to video style transfer with a gray tile ControlNet. Extensive experiments and visualizations demonstrate that StyleMaster significantly outperforms competitors, effectively generating high-quality stylized videos that align with textual content and closely resemble the style of reference images. Zixuan Ye, Huijuan Huang 0001, Xintao Wang 0004, Pengfei Wan 0001, Di Zhang 0026, Wenhan Luo |
CVPR | 4 |
| 2025 | Towards Precise Scaling Laws for Video Diffusion TransformersabstractAchieving optimal performance of video diffusion transformers within given data and compute budget is crucial due to their high training costs. This necessitates precisely determining the optimal model size and training hyperparameters before large-scale training. While scaling laws are employed in language models to predict performance, their existence and accurate derivation in visual generation models remain underexplored. In this paper, we systematically analyze scaling laws for video diffusion transformers and confirm their presence. Moreover, we discover that, unlike language models, video diffusion models are more sensitive to learning rate and batch size—two hyperparameters often not precisely modeled. To address this, we propose a new scaling law that predicts optimal hyperparameters for any model size and compute budget. Under these optimal settings, we achieve comparable performance and reduce inference costs by 40.1% compared to conventional scaling methods, within a compute budget of 1e10 TFlops. Furthermore, we establish a more generalized and precise relationship among validation loss, any model size, and compute budget. This enables performance prediction for non-optimal model sizes, which may also be appealed under practical inference cost constraints, achieving a better trade-off. Yuanyang Yin, Mingwu Zheng, Jiarong Ou, Victor Shea-Jay Huang, Xin Tao 0001, Pengfei Wan 0001, Di Zhang 0026, Baoqun Yin, Wentao Zhang 0001, Kun Gai |
CVPR | 10 |
| 2025 | RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual ReconstructionabstractYuchi Wang, Yishuo Cai, Shuhuai Ren, Sihan Yang, Linli Yao, Yuanxin Liu, Yuanxing Zhang, Pengfei Wan, Xu Sun. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yuchi Wang, Yishuo Cai, Shuhuai Ren, Linli Yao, Yuanxin Liu, Yuanxing Zhang, Pengfei Wan 0001, Xu Sun 0001 |
EMNLP | 8 |
| 2025 | SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMsabstractYuanyang Yin, Yaqi Zhao, Yajie Zhang, Yuanxing Zhang, Ke Lin, Jiahao Wang, Xin Tao, Pengfei Wan, Wentao Zhang, Feng Zhao. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yuanyang Yin, Yuanxing Zhang, Xin Tao 0001, Pengfei Wan 0001, Wentao Zhang 0001 |
EMNLP | 8 |
| 2025 | Recammaster: Camera-Controlled Generative Rendering From a Single VideoabstractCamera control has been actively studied in text or image conditioned video generation tasks. However, altering camera trajectories of a given video remains under-explored, despite its importance in the field of video creation. It is non-trivial due to the extra constraints of maintaining multiple-frame appearance and dynamic synchronization. To address this, we present ReCamMaster, a camera-controlled generative video re-rendering framework that reproduces the dynamic scene of an input video at novel camera trajectories. The core innovation lies in harnessing the generative capabilities of pre-trained text-to-video models through a simple yet powerful video conditioning mechanism--its capability is often overlooked in current research. To overcome the scarcity of qualified training data, we construct a comprehensive multi-camera synchronized video dataset using Unreal Engine 5, which is carefully curated to follow real-world filming characteristics, covering diverse scenes and camera movements. It helps the model generalize to in-the-wild videos. Lastly, we further improve the robustness to diverse inputs through a meticulously designed training strategy. Extensive experiments show that our method substantially outperforms existing state-of-the-art approaches. Our method also finds promising applications in video stabilization, super-resolution, and outpainting. Our code and dataset are publicly available at: https://github.com/KwaiVGI/ReCamMaster. Jianhong Bai, Menghan Xia, Xintao Wang 0002, Lianrui Mu, Jinwen Cao, Zuozhu Liu, Haoji Hu, Xiang Bai, Pengfei Wan 0001, Di Zhang 0026 |
ICCV | 10 |
| 2025 | How Far are AI-Generated Videos from Simulating the 3D Visual World: A Learned 3D Evaluation Approach
Chirui Chang, Jiahui Liu 0012, Zhengzhe Liu, Xiaoyang Lyu, Xin Tao 0001, Pengfei Wan 0001, Di Zhang 0026, Xiaojuan Qi 0001 |
ICCV | 7 |
| 2025 | GGTalker: Talking Head Systhesis with Generalizable Gaussian Priors and Identity-Specific AdaptationabstractCreating high-quality, generalizable speech-driven 3D talking heads remains a persistent challenge. Previous methods achieve satisfactory results for fixed viewpoints and small-scale audio variations, but they struggle with large head rotations and out-of-distribution (OOD) audio. Moreover, they are constrained by the need for time-consuming, identity-specific training. We believe the core issue lies in the lack of sufficient 3D priors, which limits the extrapolation capabilities of synthesized talking heads. To address this, we propose GGTalker, which synthesizes talking heads through a combination of generalizable priors and identity-specific adaptation. We introduce a two-stage Prior-Adaptation training strategy to learn Gaussian head priors and adapt to individual characteristics. We train Audio-Expression and Expression-Visual priors to capture the universal patterns of lip movements and the general distribution of head textures. During the Customized Adaptation, individual speaking styles and texture details are precisely modeled. Additionally, we introduce a color MLP to generate fine-grained, motion-aligned textures and a Body Inpainter to blend rendered results with the background, producing indistinguishable, photorealistic video frames. Comprehensive experiments show that GGTalker achieves state-of-the-art performance in rendering quality, 3D consistency, lip-sync accuracy, and training efficiency. Shunkai Li, Ziqiao Peng, Haoxian Zhang, Pengfei Wan 0001, Di Zhang 0026 |
ICCV | 7 |
| 2025 | FullDiT: Video Generative Foundation Models with Multimodal Control via Full Attention
Xuan Ju, Weicai Ye, Quande Liu, Qiulin Wang, Xintao Wang 0002, Pengfei Wan 0001, Di Zhang 0026, Kun Gai, Qiang Xu 0001 |
ICCV | 6 |
| 2025 | Scene Graph Guided Generation: Enable Accurate Relations Generation in Text-to-Image Models via Textural Rectification
Guibao Shen, Luozhou Wang, Jiantao Lin, Wenhang Ge, Chaozhe Zhang, Xin Tao 0001, Di Zhang 0026, Pengfei Wan 0001, Guangyong Chen, Yijun Li 0001, Ying-Cong Chen |
ICCV | 8 |
| 2025 | Imbalance in Balance: Online Concept Balancing in Generation ModelsabstractIn visual generation tasks, the responses and combinations of complex concepts often lack stability and are error-prone, which remains an under-explored area. In this paper, we attempt to explore the causal factors for poor concept responses through elaborately designed experiments. We also design a concept-wise equalization loss function (IMBA loss) to address this issue. Our proposed method is online, eliminating the need for offline dataset processing, and requires minimal code changes. In our newly proposed complex concept benchmark Inert-CompBench and two other public test sets, our method significantly enhances the concept response capability of baseline models and yields highly competitive results with only a few codes released at https://github.com/KwaiVGI/IMBA-Loss. Yukai Shi, Jiarong Ou, Xin Tao 0001, Pengfei Wan 0001, Di Zhang 0026, Kun Gai |
ICCV | 7 |
| 2025 | BadVideo: Stealthy Backdoor Attack Against Text-to-Video GenerationabstractText-to-video (T2V) generative models have rapidly advanced and found widespread applications across fields like entertainment, education, and marketing. However, the adversarial vulnerabilities of these models remain rarely explored. We observe that in T2V generation tasks, the generated videos often contain substantial redundant information not explicitly specified in the text prompts, such as environmental elements, secondary objects, and additional details, providing opportunities for malicious attackers to embed hidden harmful content. Exploiting this inherent redundancy, we introduce BadVideo, the first backdoor attack framework tailored for T2V generation. Our attack focuses on designing target adversarial outputs through two key strategies: (1) Spatio-Temporal Composition, which combines different spatiotemporal features to encode malicious information; (2) Dynamic Element Transformation, which introduces transformations in redundant elements over time to convey malicious information. Based on these strategies, the attacker's malicious target seamlessly integrates with the user's textual instructions, providing high stealthiness. Moreover, by exploiting the temporal dimension of videos, our attack successfully evades traditional content moderation systems that primarily analyze spatial information within individual frames. Extensive experiments demonstrate that BadVideo achieves high attack success rates while preserving original semantics and maintaining excellent performance on clean inputs. Overall, our work reveals the adversarial vulnerability of T2V models, calling attention to potential risks and misuse. Our project page is at https://wrt2000.github.io/BadVideo2025/. Ruotong Wang 0008, Mingli Zhu, Jiarong Ou, Xin Tao 0001, Pengfei Wan 0001, Baoyuan Wu |
ICCV | 6 |
| 2025 | GameFactorly: Creating New Games with Generative Interactive Videos
Jiwen Yu, Yiran Qin, Xintao Wang 0002, Pengfei Wan 0001, Di Zhang 0026, Xihui Liu |
ICCV | 4 |
| 2025 | SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse ViewpointsabstractRecent advancements in video diffusion models demonstrate remarkable capabilities in simulating real-world dynamics and 3D consistency. This progress motivates us to explore the potential of these models to maintain dynamic consistency across diverse viewpoints, a feature highly sought after in applications like virtual filming. Unlike existing methods focused on multi-view generation of single objects for 4D reconstruction, our interest lies in generating open-world videos from arbitrary viewpoints, incorporating six degrees of freedom (6 DoF) camera poses.
To achieve this, we propose a plug-and-play module that enhances a pre-trained text-to-video model for multi-camera video generation, ensuring consistent content across different viewpoints. Specifically, we introduce a multi-view synchronization module designed to maintain appearance and geometry consistency across these viewpoints. Given the scarcity of high-quality training data, we also propose a progressive training scheme that leverages multi-camera images and monocular videos as a supplement to Unreal Engine-rendered multi-camera videos. This comprehensive approach significantly benefits our model.
Experimental results demonstrate the superiority of our proposed method over existing competitors and several baselines. Furthermore, our method enables intriguing extensions, such as re-rendering a video from multiple novel viewpoints. Project webpage: https://jianhongbai.github.io/SynCamMaster/ Jianhong Bai, Menghan Xia, Xintao Wang 0002, Ziyang Yuan, Zuozhu Liu, Haoji Hu, Pengfei Wan 0001, Di Zhang 0026 |
ICLR | 7 |
| 2025 | Cafe-Talk: Generating 3D Talking Face Animation with Multimodal Coarse- and Fine-grained ControlabstractSpeech-driven 3D talking face method should offer both accurate lip synchronization and controllable expressions. Previous methods solely adopt discrete emotion labels to globally control expressions throughout sequences while limiting flexible fine-grained facial control within the spatiotemporal domain. We propose a diffusion-transformer-based 3D talking face generation model, Cafe-Talk, which simultaneously incorporates coarse- and fine-grained multimodal control conditions. Nevertheless, the entanglement of multiple conditions challenges achieving satisfying performance. To disentangle speech audio and fine-grained conditions, we employ a two-stage training pipeline. Specifically, Cafe-Talk is initially trained using only speech audio and coarse-grained conditions. Then, a proposed fine-grained control adapter gradually adds fine-grained instructions represented by action units (AUs), preventing unfavorable speech-lip synchronization. To disentangle coarse- and fine-grained conditions, we design a swap-label training mechanism, which enables the dominance of the fine-grained conditions. We also devise a mask-based CFG technique to regulate the occurrence and intensity of fine-grained control. In addition, a text-based detector is introduced with text-AU alignment to enable natural language user input and further support multimodal control. Extensive experimental results prove that Cafe-Talk achieves state-of-the-art lip synchronization and expressiveness performance and receives wide acceptance in fine-grained control in user studies. Hejia Chen, Haoxian Zhang, Shoulong Zhang, Sisi Zhuang, Yuan Zhang 0020, Pengfei Wan 0001, Di Zhang 0026, Shuai Li 0001 |
ICLR | 7 |
| 2025 | Stable Segment Anything ModelabstractThe Segment Anything Model (SAM) achieves remarkable promptable segmentation given high-quality prompts which, however, often require good skills to specify. To make SAM robust to casual prompts, this paper presents the first comprehensive analysis on SAM’s segmentation stability across a diverse spectrum of prompt qualities, notably imprecise bounding boxes and insufficient points. Our key finding reveals that given such low-quality prompts, SAM’s mask decoder tends to activate image features that are biased towards the background or confined to specific object parts. To mitigate this issue, our key idea consists of calibrating solely SAM’s mask attention by adjusting the sampling locations and amplitudes of image features, while the original SAM model architecture and weights remain unchanged. Consequently, our deformable sampling plugin (DSP) enables SAM to adaptively shift attention to the prompted target regions in a data-driven manner. During inference, dynamic routing plugin (DRP) is proposed that toggles SAM between the deformable and regular grid sampling modes, conditioned on the input prompt quality. Thus, our solution, termed Stable-SAM, offers several advantages: 1) improved SAM’s segmentation stability across a wide range of prompt qualities, while 2) retaining SAM’s powerful promptable segmentation efficiency and generality, with 3) minimal learnable parameters (0.08 M) and fast adaptation. Extensive experiments validate the effectiveness and advantages of our approach, underscoring Stable-SAM as a more robust solution for segmenting anything. Codes are at https://github.com/fanq15/Stable-SAM. Xin Tao 0001, Lei Ke, Mingqiao Ye, Di Zhang 0026, Pengfei Wan 0001, Yu-Wing Tai, Chi-Keung Tang |
ICLR | 6 |
| 2025 | 3DTrajMaster: Mastering 3D Trajectory for Multi-Entity Motion in Video GenerationabstractThis paper aims to manipulate multi-entity 3D motions in video generation. Previous methods on controllable video generation primarily leverage 2D control signals to manipulate object motions and have achieved remarkable synthesis results. However, 2D control signals are inherently limited in expressing the 3D nature of object motions. To overcome this problem, we introduce 3DTrajMaster, a robust controller that regulates multi-entity dynamics in 3D space, given user-desired 6DoF pose (location and rotation) sequences of entities. At the core of our approach is a plug-and-play 3D-motion grounded object injector that fuses multiple input entities with their respective 3D trajectories through a gated self-attention mechanism. In addition, we exploit an injector architecture to preserve the video diffusion prior, which is crucial for generalization ability. To mitigate video quality degradation, we introduce a domain adaptor during training and employ an annealed sampling strategy during inference. To address the lack of suitable training data, we construct a 360-Motion Dataset, which first correlates collected 3D human and animal assets with GPT-generated trajectory and then captures their motion with 12 evenly-surround cameras on diverse 3D UE platforms. Extensive experiments show that 3DTrajMaster sets a new state-of-the-art in both accuracy and generalization for controlling multi-entity 3D motions. Project page: http://fuxiao0719.github.io/projects/3dtrajmaster Xintao Wang 0002, Sida Peng, Menghan Xia, Xiaoyu Shi 0002, Ziyang Yuan, Pengfei Wan 0001, Di Zhang 0026, Dahua Lin |
ICLR | 8 |
| 2025 | MODA: MOdular Duplex Attention for Multimodal Perception, Cognition, and Emotion UnderstandingabstractMultimodal large language models (MLLMs) recently showed strong capacity in integrating data among multiple modalities, empowered by generalizable attention architecture. Advanced methods predominantly focus on language-centric tuning while less exploring multimodal tokens mixed through attention, posing challenges in high-level tasks that require fine-grained cognition and emotion understanding. In this work, we identify the attention deficit disorder problem in multimodal learning, caused by inconsistent cross-modal attention and layer-by-layer decayed attention activation. To address this, we propose a novel attention mechanism, termed MOdular Duplex Attention (MODA), simultaneously conducting the inner-modal refinement and inter-modal interaction. MODA employs a correct-after-align strategy to effectively decouple modality alignment from cross-layer token mixing. In the alignment phase, tokens are mapped to duplex modality spaces based on the basis vectors, enabling the interaction between visual and language modality. Further, the correctness of attention scores is ensured through adaptive masked attention, which enhances the model's flexibility by allowing customizable masking patterns for different modalities. Extensive experiments on 21 benchmark datasets verify the effectiveness of MODA in perception, cognition, and emotion tasks. Wuyou Xia, Chenxi Zhao 0002, Zhou Yan, Yongjie Zhu, Wenyu Qin, Pengfei Wan 0001, Di Zhang 0026, Jufeng Yang |
ICML | 8 |
| 2025 | EditWorld: Simulating World Dynamics for Instruction-Following Image EditingabstractDiffusion models have significantly improved the performance of image editing. Existing methods realize various approaches to achieve high-quality image editing, including but not limited to text control, dragging operation, and mask-and-inpainting. Among these, instruction-based editing stands out for its convenience and effectiveness in following human instructions across diverse scenarios. However, it still focuses on simple editing operations like adding, replacing, or deleting, and falls short of understanding aspects of world dynamics that convey the realistic dynamic nature in the physical world. Therefore, this work EditWorld introduces a new editing task, namely world-instructed image editing, which defines and categorizes the instructions grounded by various world scenarios. We curate a new image editing dataset with world instructions using a set of large pretrained models (e.g., GPT, Video-LLava and SDXL). To enable sufficient simulation of world dynamics for image editing, our EditWorld trains model in the curated dataset, and improves instruction-following ability with designed post-edit strategy. Extensive experiments demonstrate our method significantly outperforms existing editing methods in this new task. https://github.com/YangLing0818/EditWorld Bohan Zeng, Ling Yang 0006, Yuanxing Zhang, Pengfei Wan 0001, Wentao Zhang 0001, Shuicheng Yan |
ACM Multimedia | 6 |
| 2025 | Flow-GRPO: Training Flow Matching Models via Online RLabstractWe propose Flow-GRPO, the first method to integrate online policy gradient reinforcement learning (RL) into flow matching models. Our approach uses two key strategies: (1) an ODE-to-SDE conversion that transforms a deterministic Ordinary Differential Equation (ODE) into an equivalent Stochastic Differential Equation (SDE) that matches the original model's marginal distribution at all timesteps, enabling statistical sampling for RL exploration; and (2) a Denoising Reduction strategy that reduces training denoising steps while retaining the original number of inference steps, significantly improving sampling efficiency without sacrificing performance. Empirically, Flow-GRPO is effective across multiple text-to-image tasks. For compositional generation, RL-tuned SD3.5-M generates nearly perfect object counts, spatial relations, and fine-grained attributes, increasing GenEval accuracy from $63\%$ to $95\%$. In visual text rendering, accuracy improves from $59\%$ to $92\%$, greatly enhancing text generation. Flow-GRPO also achieves substantial gains in human preference alignment. Notably, very little reward hacking occurred, meaning rewards did not increase at the cost of appreciable image quality or diversity degradation. Jie Liu 0047, Gongye Liu, Jiajun Liang, Yangguang Li 0001, Xintao Wang 0002, Pengfei Wan 0001, Di Zhang 0026, Wanli Ouyang |
NeurIPS | 7 |
| 2025 | Improving Video Generation with Human FeedbackabstractVideo generation has achieved significant advances through rectified flow techniques, but issues like unsmooth motion and misalignment between videos and prompts persist. In this work, we develop a systematic pipeline that harnesses human feedback to mitigate these problems and refine the video generation model. Specifically, we begin by constructing a large-scale human preference dataset focused on modern video generation models, incorporating pairwise annotations across multi-dimensions. We then introduce VideoReward, a multi-dimensional video reward model, and examine how annotations and various design choices impact its rewarding efficacy. From a unified reinforcement learning perspective aimed at maximizing reward with KL regularization, we introduce three alignment algorithms for flow-based models. These include two training-time strategies: direct preference optimization for flow (Flow-DPO) and reward weighted regression for flow (Flow-RWR), and an inference-time technique, Flow-NRG, which applies reward guidance directly to noisy videos. Experimental results indicate that VideoReward significantly outperforms existing reward models, and Flow-DPO demonstrates superior performance compared to both Flow-RWR and supervised fine-tuning methods. Additionally, Flow-NRG lets users assign custom weights to multiple objectives during inference, meeting personalized video quality needs. Jie Liu 0047, Gongye Liu, Jiajun Liang, Ziyang Yuan, Xiaokun Liu, Mingwu Zheng, Xiele Wu, Qiulin Wang, Menghan Xia, Xintao Wang 0002, Xiaohong Liu 0001, Pengfei Wan 0001, Di Zhang 0026, Kun Gai, Yujiu Yang 0001, Wanli Ouyang |
NeurIPS | 13 |
| 2025 | OmniSync: Towards Universal Lip Synchronization via Diffusion TransformersabstractLip synchronization is the task of aligning a speaker’s lip movements in video with corresponding speech audio, and it is essential for creating realistic, expressive video content.
However, existing methods often rely on reference frames and masked-frame inpainting, which limit their robustness to identity consistency, pose variations, facial occlusions, and stylized content. In addition, since audio signals provide weaker conditioning than visual cues, lip shape leakage from the original video will affect lip sync quality.
In this paper, we present OmniSync, a universal lip synchronization framework for diverse visual scenarios. Our approach introduces a mask-free training paradigm using Diffusion Transformer models for direct frame editing without explicit masks, enabling unlimited-duration inference while maintaining natural facial dynamics and preserving character identity.
During inference, we propose a flow-matching-based progressive noise initialization to ensure pose and identity consistency, while allowing precise mouth-region editing. To address the weak conditioning signal of audio, we develop a Dynamic Spatiotemporal Classifier-Free Guidance (DS-CFG) mechanism that adaptively adjusts guidance strength over time and space.
We also establish the AIGC-LipSync Benchmark, the first evaluation suite for lip synchronization in diverse AI-generated videos. Extensive experiments demonstrate that OmniSync significantly outperforms prior methods in both visual quality and lip sync accuracy, achieving superior results in both real-world and AI-generated videos. Ziqiao Peng, Jiwen Liu, Haoxian Zhang, Songlin Tang, Pengfei Wan 0001, Di Zhang 0026, Hongyan Liu 0002, Jun He 0008 |
NeurIPS | 6 |
| 2025 | MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMsabstractThe advent of Multimodal Large Language Models (MLLMs) has expanded AI capabilities to visual modalities, yet existing evaluation benchmarks remain limited to single-video understanding, overlooking the critical need for multi-video understanding in real-world scenarios (e.g., sports analytics and autonomous driving). To address this significant gap, we introduce MVU-Eval, the first comprehensive benchmark for evaluating Multi-Video Understanding for MLLMs. Specifically, our MVU-Eval mainly assesses eight core competencies through 1,824 meticulously curated question-answer pairs spanning 4,959 videos from diverse domains, addressing both fundamental perception tasks and high-order reasoning tasks. These capabilities are rigorously aligned with real-world applications such as multi-sensor synthesis in autonomous systems and cross-angle sports analytics. Through extensive evaluation of state-of-the-art open-source and closed-source models, we reveal significant performance discrepancies and limitations in current MLLMs' ability to perform understanding across multiple videos.The benchmark will be made publicly available to foster future research. Yuanxing Zhang, Noah Wang, Ge Zhang 0009, Jian Yang 0037, Yanghai Wang, Xintao Wang 0002, Houyi Li, Wei Ji 0011, Pengfei Wan 0001, Wenhao Huang 0001, Zhaoxiang Zhang 0001 |
NeurIPS | 13 |
| 2025 | MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video ScenariosabstractMultimodal Large Language Models (MLLMs) have achieved considerable accuracy in Optical Character Recognition (OCR) from static images. However, their efficacy in video OCR is significantly diminished due to factors such as motion blur, temporal variations, and visual effects inherent in video content. To provide clearer guidance for training practical MLLMs, we introduce MME-VideoOCR benchmark, which encompasses a comprehensive range of video OCR application scenarios. MME-VideoOCR features 10 task categories comprising 25 individual tasks and spans 44 diverse scenarios. These tasks extend beyond text recognition to incorporate deeper comprehension and reasoning of textual content within videos. The benchmark consists of 1,464 videos with varying resolutions, aspect ratios, and durations, along with 2,000 meticulously curated, manually annotated question-answer pairs. We evaluate 18 state-of-the-art MLLMs on MME-VideoOCR, revealing that even the best-performing model (Gemini-2.5 Pro) achieves only an accuracy of 73.7%. Fine-grained analysis indicates that while existing MLLMs demonstrate strong performance on tasks where relevant texts are contained within a single or few frames, they exhibit limited capability in effectively handling tasks that demand holistic video comprehension. These limitations are especially evident in scenarios that require spatio-temporal reasoning, cross-frame information integration, or resistance to language prior bias. Our findings also highlight the importance of high-resolution visual input and sufficient temporal coverage for reliable OCR in dynamic video scenarios. Yang Shi 0009, Huanqian Wang, Wulin Xie, Huanyao Zhang, Lijie Zhao, Yifan Zhang 0004, Xinfeng Li, Chaoyou Fu, Zhuoer Wen, Zhuoran Zhang 0003, Xinlong Chen, Bohan Zeng, Yushuo Guan, Zhang Zhang 0001, Liang Wang 0001, Haoxuan Li 0001, Zhouchen Lin, Yuanxing Zhang, Pengfei Wan 0001, Haotian Wang 0001, Wenjing Yang 0002 |
NeurIPS | 21 |
| 2025 | VidEmo: Affective-Tree Reasoning for Emotion-Centric Video Foundation ModelsabstractUnderstanding and predicting emotions from videos has gathered significant attention in recent studies, driven by advancements in video large language models (VideoLLMs). While advanced methods have made progress in video emotion analysis, the intrinsic nature of emotions—characterized by their open-set, dynamic, and context-dependent properties—poses challenge in understanding complex and evolving emotional states with reasonable rationale. To tackle these challenges, we propose a novel affective cues-guided reasoning framework that unifies fundamental attribute perception, expression analysis, and high-level emotional understanding in a stage-wise manner. At the core of our approach is a family of video emotion foundation models (VidEmo), specifically designed for emotion reasoning and instruction-following. These models undergo a two-stage tuning process: first, curriculum emotion learning for injecting emotion knowledge, followed by affective-tree reinforcement learning for emotion reasoning. Moreover, we establish a foundational data infrastructure and introduce a emotion-centric fine-grained dataset (Emo-CFG) consisting of 2.1M diverse instruction-based samples. Emo-CFG includes explainable emotional question-answering, fine-grained captions, and associated rationales, providing essential resources for advancing emotion understanding tasks. Experimental results demonstrate that our approach achieves competitive performance, setting a new milestone across 15 face perception tasks. Yongjie Zhu, Wenyu Qin, Pengfei Wan 0001, Di Zhang 0026, Jufeng Yang |
NeurIPS | 5 |
| 2025 | Training-Free Efficient Video Generation via Dynamic Token CarvingabstractDespite the remarkable generation quality of video Diffusion Transformer (DiT) models, their practical deployment is severely hindered by extensive computational requirements. This inefficiency stems from two key challenges: the quadratic complexity of self-attention with respect to token length and the multi-step nature of diffusion models. To address these limitations, we present Jenga, a novel inference pipeline that combines dynamic attention carving with progressive resolution generation. Our approach leverages two key insights: (1) early denoising steps do not require high-resolution latents, and (2) later steps do not require dense attention. Jenga introduces a block-wise attention mechanism that dynamically selects relevant token interactions using 3D space-filling curves, alongside a progressive resolution strategy that gradually increases latent resolution during generation. Experimental results demonstrate that Jenga achieves substantial speedups across multiple state-of-the-art video diffusion models while maintaining comparable generation quality (8.83$\times$ speedup with 0.01\% performance drop on VBench). As a plug-and-play solution, Jenga enables practical, high-quality video generation on modern hardware by reducing inference time from minutes to seconds---without requiring model retraining. Yuechen Zhang, Jinbo Xing, Bin Xia 0014, Shaoteng Liu, Bohao Peng, Xin Tao 0001, Pengfei Wan 0001, Eric Lo 0001, Jiaya Jia |
NeurIPS | 7 |
| 2025 | VFRTok: Variable Frame Rates Video Tokenizer with Duration-Proportional Information AssumptionabstractModern video generation frameworks based on Latent Diffusion Models suffer from inefficiencies in tokenization due to the Frame-Proportional Information Assumption.
Existing tokenizers provide fixed temporal compression rates, causing the computational cost of the diffusion model to scale linearly with the frame rate.
The paper proposes the Duration-Proportional Information Assumption: the upper bound on the information capacity of a video is proportional to the duration rather than the number of frames.
Based on this insight, the paper introduces VFRTok, a Transformer-based video tokenizer, that enables variable frame rate encoding and decoding through asymmetric frame rate training between the encoder and decoder.
Furthermore, the paper proposes Partial Rotary Position Embeddings (RoPE) to decouple position and content modeling, which groups correlated patches into unified tokens.
The Partial RoPE effectively improves content-awareness, enhancing the video generation capability.
Benefiting from the compact and continuous spatio-temporal representation, VFRTok achieves competitive reconstruction quality and state-of-the-art generation fidelity while using only $1/8$ tokens compared to existing tokenizers. Tianxiong Zhong, Xingye Tian, Boyuan Jiang, Xuebo Wang, Xin Tao 0001, Pengfei Wan 0001 |
NeurIPS | 6 |
| 2025 | CamCloneMaster: Enabling Reference-based Camera Control for Video GenerationabstractCamera control is crucial for generating expressive and cinematic videos. Existing methods rely on explicit sequences of camera parameters as control conditions, which can be cumbersome for users to construct, particularly for intricate camera movements. To provide a more intuitive camera control method, we propose CamCloneMaster, a framework that enables users to replicate camera movements from reference videos without requiring camera parameters or test-time fine-tuning. CamCloneMaster seamlessly supports reference-based camera control for both Image-to-Video and Video-to-Video tasks within a unified framework. Furthermore, we present the Camera Clone Dataset, a large-scale synthetic dataset designed for camera clone learning, encompassing diverse scenes, subjects, and camera movements. Extensive experiments and user studies demonstrate that CamCloneMaster outperforms existing methods in terms of both camera controllability and visual quality. Dataset and Code can be found at https://camclonemaster.github.io/. Yawen Luo, Xiaoyu Shi 0002, Jianhong Bai, Menghan Xia, Tianfan Xue, Xintao Wang 0002, Pengfei Wan 0001, Di Zhang 0026, Kun Gai |
SIGGRAPH Asia | 7 |
| 2025 | Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory RetrievalabstractRecent advances in interactive video generation have shown promising results, yet existing approaches struggle with scene-consistent memory capabilities in long video generation due to limited use of historical context. In this work, we propose Context-as-Memory, which utilizes historical context as memory for video generation. It includes two simple yet effective designs: (1) storing context in frame format without additional post-processing; (2) conditioning by concatenating context and frames to be predicted along the frame dimension at the input, requiring no external control modules. Furthermore, considering the enormous computational overhead of incorporating all historical context, we propose the Memory Retrieval module to select truly relevant context frames by determining FOV (Field of View) overlap between camera poses, which significantly reduces the number of candidate frames without substantial information loss. Experiments demonstrate that Context-as-Memory achieves superior memory capabilities in interactive long video generation compared to SOTAs, even generalizing effectively to open-domain scenarios not seen during training. Our project page are publicly available at https://context-as-memory.github.io/. Jiwen Yu, Jianhong Bai, Yiran Qin, Quande Liu, Xintao Wang 0002, Pengfei Wan 0001, Di Zhang 0026, Xihui Liu |
SIGGRAPH Asia | 6 |
| 2025 | DVIS++: Improved Decoupled Framework for Universal Video SegmentationabstractWe present the Decoupled VIdeo Segmentation (DVIS) framework, a novel approach for the challenging task of universal video segmentation, including video instance segmentation (VIS), video semantic segmentation (VSS), and video panoptic segmentation (VPS). Unlike previous methods that model video segmentation in an end-to-end manner, our approach decouples video segmentation into three cascaded sub-tasks: segmentation, tracking, and refinement. This decoupling design allows for simpler and more effective modeling of the spatio-temporal representations of objects, especially in complex scenes and long videos. Accordingly, we introduce two novel components: the referring tracker and the temporal refiner. These components track objects frame by frame and model spatio-temporal representations based on pre-aligned features. To improve the tracking capability of DVIS, we propose a denoising training strategy and introduce contrastive learning, resulting in a more robust framework named DVIS++. The proposed decoupled framework efficiently handles universal and open-vocabulary object representations, allowing DVIS++ to conduct universal and open-vocabulary video segmentation. We conduct extensive experiments on six mainstream benchmarks, including the VIS, VSS, and VPS datasets. Using a unified architecture, DVIS++ significantly outperforms state-of-the-art specialized methods on these benchmarks in closed- and open-vocabulary settings. Tao Zhang 0042, Xingye Tian, Yikang Zhou, Shunping Ji, Xuebo Wang, Xin Tao 0001, Yuan Zhang 0020, Pengfei Wan 0001, Zhongyuan Wang 0006, Yu Wu 0011 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2025 | A-SDM: Accelerating Stable Diffusion Through Model Assembly and Feature Inheritance StrategiesabstractThe stable diffusion model (SDM) is a prevalent and effective model for text-to-image (T2I) and image-to-image (I2I) generation. Despite various attempts at sampler optimization, model distillation, and network quantification, these approaches typically maintain the original network architecture. The extensive parameter scale and substantial computational demands have limited research into adjusting the model architecture. This study focuses on reducing redundant computation in SDM and optimizes the model through both tuning and tuning-free methods: 1) for the tuning method, we design a model assembly strategy to reconstruct a lightweight model while preserving performance and ensuring semantic stability through distillation and 2) for the tuning-free method, we propose a feature inheritance strategy to accelerate inference by skipping local computations at the block, layer, or unit level within the network structure. We also examine multiple sampling modes for feature inheritance at the time-step level. Experiments demonstrate that both the proposed tuning and the tuning-free methods can improve the speed and performance of the SDM. The lightweight model reconstructed by the model assembly strategy increases generation speed by 22.4%, while the feature inheritance strategy enhances the SDM generation speed by 40.0%. Jinchao Zhu, Siyuan Pan, Pengfei Wan 0001, Di Zhang 0026, Gao Huang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | NeRFFaceShop: Learning a Photo-Realistic 3D-Aware Generative Model of Animatable and Relightable Heads From Large-Scale in-the-Wild VideosabstractAnimatable and relightable 3D facial generation has fundamental applications in computer vision and graphics. Although animation and relighting are highly correlated, previous methods usually address them separately. Effectively combining animation methods and relighting methods is nontrivial. In terms of explicit shading models, animatable methods cannot be easily extended to achieve realistic relighting results, such as shadow effects, due to prohibitive computational training costs. Regarding implicit lighting representations, current animatable methods cannot be incorporated due to their inharmonious animation representations, i.e., deforming spatial points. This paper, armed with a lightweight but effective lighting representation, presents a compatible animation representation to achieve a disentangled generative model of 3D animatable and relightable heads. Our represented animation allows for updating and control of realistic lighting effects. Due to the disentangled nature of our representations, we learn the animation and relighting from large-scale, in-the-wild videos instead of relying on a morphable model. We show that our method can synthesize geometrically consistent and detailed motion along with the disentangled control of lighting conditions. We further show that our method is still compatible with morphable models for driving generated avatars. Our method can also be extended to domains without video data by domain transfer to achieve a broader range of animatable and relightable head synthesis. We will release the code for reproducibility and facilitating future research. Feng-Lin Liu, Pengfei Wan 0001, Yuan Zhang 0020, Yukun Lai, Hongbo Fu 0001, Lin Gao 0004 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | DiffCap: Diffusion-Based Real-Time Human Motion Capture Using Sparse IMUs and a Monocular CameraabstractCombining sparse IMUs and a monocular camera is a new promising setting to perform real-time human motion capture. This paper proposes a diffusion-based solution to learn human motion priors and fuse the two modalities of signals together seamlessly in a unified framework. By delicately considering the characteristics of the two signals, the sequential visual information is considered as a whole and transformed into a condition embedding, while the inertial measurement is concatenated with the noisy body pose frame by frame to construct a sequential input for the diffusion model. Firstly, we observe that the visual information may be unavailable in some frames due to occlusions or subjects moving out of the camera view. Thus incorporating the sequential visual features as a whole to get a single feature embedding is robust to the occasional degenerations of visual information in those frames. On the other hand, the IMU measurements are robust to occlusions and always stable when signal transmission has no problem. So incorporating them frame-wisely could better explore the temporal information for the system. Experiments have demonstrated the effectiveness of the system design and its state-of-the-art performance in pose estimation compared with the previous works. The code will be released. Shaohua Pan 0002, Xinyu Yi, Yan Zhou 0003, Weihua Jian, Yuan Zhang 0020, Pengfei Wan 0001, Feng Xu 0005 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | MotionCrafter: Plug-and-Play Motion Guidance for Diffusion ModelsabstractThe essence of a video lies in the dynamic motions. While text-to-video generative diffusion models have made significant strides in creating diverse content, effectively controlling specific motions through text prompts remains a challenge. By utilizing user-specified reference videos, the more precise guidance for character actions, object movements, and camera movements can be achieved. This gives rise to the task of motion customization, where the primary challenge lies in effectively decoupling the appearance and motion within a video clip. To address this challenge, we introduce MotionCrafter, a novel one-shot instance-guided motion customization method that is suitable for both pre-trained text-to-video and text-to-image diffusion models. MotionCrafter employs a parallel spatial-temporal architecture that integrates the reference motion into the temporal component of the base model, while independently adjusting the spatial module for character or style control. To enhance the disentanglement of motion and appearance, we propose an innovative dual-branch motion disentanglement approach, which includes a motion disentanglement loss and an appearance prior enhancement strategy. To facilitate more efficient learning of motions, we further propose a novel timestep-layered tuning strategy that directs the diffusion model to focus on motion-level information. Through comprehensive quantitative and qualitative experiments, along with user preference tests, we demonstrate that MotionCrafter can successfully integrate dynamic motions while maintaining the coherence and quality of the base model, providing a wide range of appearance generation capabilities. MotionCrafter can be applied to various personalized backbones in the community to generate videos with a variety of artistic styles. Yuxin Zhang 0006, Weiming Dong, Fan Tang, Nisha Huang, Chongyang Ma, Pengfei Wan 0001, Tong-Yee Lee, Changsheng Xu |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2024 | Agent Attention: On the Integration of Softmax and Linear Attention
Dongchen Han, Tianzhu Ye, Yizeng Han, Zhuofan Xia, Siyuan Pan, Pengfei Wan 0001, Shiji Song, Gao Huang 0001 |
ECCV (50) | 6 |
| 2024 | PlacidDreamer: Advancing Harmony in Text-to-3D GenerationabstractRecently, text-to-3D generation has attracted significant attention, resulting in notable performance enhancements. Previous methods utilize end-to-end 3D generation models to initialize 3D Gaussians, multi-view diffusion models to enforce multi-view consistency, and text-to-image diffusion models to refine details with score distillation algorithms. However, these methods exhibit two limitations. Firstly, they encounter conflicts in generation directions since different models aim to produce diverse 3D assets. Secondly, the issue of over-saturation in score distillation has not been thoroughly investigated and solved. To address these limitations, we propose PlacidDreamer, a text-to-3D framework that harmonizes initialization, multi-view generation, and text-conditioned generation with a single multi-view diffusion model, while simultaneously employing a novel score distillation algorithm to achieve balanced saturation. To unify the generation direction, we introduce the Latent-Plane module, a training-friendly plug-in extension that enables multi-view diffusion models to provide fast geometry reconstruction for initialization and enhanced multi-view images to personalize the text-to-image diffusion model. To address the over-saturation problem, we propose to view score distillation as a multi-objective optimization problem and introduce the Balanced Score Distillation algorithm, which offers a Pareto Optimal solution that achieves both rich details and balanced saturation. Extensive experiments validate the outstanding capabilities of our PlacidDreamer. The code is available at https://github.com/HansenHuang0823/PlacidDreamer. Shuo Huang 0005, Shikun Sun, Zixuan Wang 0026, Xiaoyu Qin 0001, Yanmin Xiong, Yuan Zhang 0020, Pengfei Wan 0001, Di Zhang 0026, Jia Jia 0001 |
ACM Multimedia | 7 |
| 2024 | VideoTetris: Towards Compositional Text-to-Video GenerationabstractDiffusion models have demonstrated great success in text-to-video (T2V) generation. However, existing methods may face challenges when handling complex (long) video generation scenarios that involve multiple objects or dynamic changes in object numbers. To address these limitations, we propose VideoTetris, a novel framework that enables compositional T2V generation. Specifically, we propose spatio-temporal compositional diffusion to precisely follow complex textual semantics by manipulating and composing the attention maps of denoising networks spatially and temporally. Moreover, we propose a new dynamic-aware data processing pipeline and a consistency regularization method to enhance the consistency of auto-regressive video generation. Extensive experiments demonstrate that our VideoTetris achieves impressive qualitative and quantitative results in compositional T2V generation. Code is available at: https://github.com/YangLing0818/VideoTetris Ling Yang 0006, Yuan Gao 0015, Yufan Deng, Xintao Wang 0002, Zhaochen Yu, Xin Tao 0001, Pengfei Wan 0001, Di Zhang 0026, Bin Cui 0001 |
NeurIPS | 9 |
| 2024 | Towards Unified 3D Hair Reconstruction from Single-View Portraits
Yujian Zheng, Yuda Qiu, Leyang Jin 0001, Chongyang Ma, Di Zhang 0026, Pengfei Wan 0001, Xiaoguang Han 0001 |
SIGGRAPH Asia | 7 |
| 2023 | FEditNet: Few-Shot Editing of Latent Semantics in GAN SpacesabstractGenerative Adversarial networks (GANs) have demonstrated their powerful capability of synthesizing high-resolution images, and great efforts have been made to interpret the semantics in the latent spaces of GANs. However, existing works still have the following limitations: (1) the majority of works rely on either pretrained attribute predictors or large-scale labeled datasets, which are difficult to collect in most cases, and (2) some other methods are only suitable for restricted cases, such as focusing on interpretation of human facial images using prior facial semantics. In this paper, we propose a GAN-based method called FEditNet, aiming to discover latent semantics using very few labeled data without any pretrained predictors or prior knowledge. Specifically, we reuse the knowledge from the pretrained GANs, and by doing so, avoid overfitting during the few-shot training of FEditNet. Moreover, our layer-wise objectives which take content consistency into account also ensure the disentanglement between attributes. Qualitative and quantitative results demonstrate that our method outperforms the state-of-the-art methods on various datasets. The code is available at https://github.com/THU-LYJ-Lab/FEditNet. Mengfei Xia, Yezhi Shu, Yuji Wang, Yukun Lai, Qiang Li 0024, Pengfei Wan 0001, Zhongyuan Wang 0006, Yong-Jin Liu 0001 |
AAAI | 6 |
| 2023 | DVIS: Decoupled Video Instance Segmentation FrameworkabstractVideo instance segmentation (VIS) is a critical task with diverse applications, including autonomous driving and video editing. Existing methods often underperform on complex and long videos in real world, primarily due to two factors. Firstly, offline methods are limited by the tightly-coupled modeling paradigm, which treats all frames equally and disregards the interdependencies between adjacent frames. Consequently, this leads to the introduction of excessive noise during long-term temporal alignment. Secondly, online methods suffer from inadequate utilization of temporal information. To tackle these challenges, we propose a decoupling strategy for VIS by dividing it into three independent sub-tasks: segmentation, tracking, and refinement. The efficacy of the decoupling strategy relies on two crucial elements: 1) attaining precise long-term alignment outcomes via frame-by-frame association during tracking, and 2) the effective utilization of temporal information predicated on the aforementioned accurate alignment outcomes during refinement. We introduce a novel referring tracker and temporal refiner to construct the Decoupled VIS framework (DVIS). DVIS achieves new SOTA performance in both VIS and VPS, surpassing the current SOTA methods by 7.3 AP and 9.6 VPQ on the OVIS and VIPSeg datasets, which are the most challenging and realistic benchmarks. Moreover, thanks to the decoupling strategy, the referring tracker and temporal refiner are super light-weight (only 1.69% of the segmenter FLOPs), allowing for efficient training and inference on a single GPU with 11G memory. The code is available at https://github.com/zhang-tao-whu/DVIS. Tao Zhang 0042, Xingye Tian, Yu Wu 0011, Shunping Ji, Xuebo Wang, Yuan Zhang 0020, Pengfei Wan 0001 |
ICCV | 7 |
| 2023 | Automatic Human Scene Interaction through Contact Estimation and Motion AdaptationabstractHuman scene interaction (HSI) aims to understand and accommodate the various ways humans interact with their environment. However, existing works typically struggle to produce high-precision contact estimation for understanding these interactions and lack sufficient fine-grained semantics to adapt to the complexities of diverse characters and scenes, resulting in unnatural and inaccurate interacting motions. In this paper, we present a novel approach to automatic human scene interaction that effectively recovers the human body mesh and high-precision contact information, subsequently enabling adaptation to different environments. Our main contributions include the proposal of a contact estimation framework that leverages semantic features from 2D images and 3D model recovered with inverse kinematics to guide the learning of vertex-level human scene contact (HSC) estimation. For motion adaptation, we propose an enhanced Laplacian semantics descriptor combined with kinematic constraints, enabling precise retargeting between variously sized human models and distinct 3D scenes. Through extensive experiments, we demonstrate our method's superiority against state-of-the-art approaches and showcase the results of the automatic human scene interaction process. Yan Zhou 0003, Li Chen 0031, Weihua Jian, Pengfei Wan 0001 |
ACM Multimedia | 6 |
| 2023 | Augmentation-Aware Self-Supervision for Data-Efficient GAN TrainingabstractTraining generative adversarial networks (GANs) with limited data is challenging because the discriminator is prone to overfitting. Previously proposed differentiable augmentation demonstrates improved data efficiency of training GANs. However, the augmentation implicitly introduces undesired invariance to augmentation for the discriminator since it ignores the change of semantics in the label space caused by data transformation, which may limit the representation learning ability of the discriminator and ultimately affect the generative modeling performance of the generator. To mitigate the negative impact of invariance while inheriting the benefits of data augmentation, we propose a novel augmentation-aware self-supervised discriminator that predicts the augmentation parameter of the augmented data. Particularly, the prediction targets of real data and generated data are required to be distinguished since they are different during training. We further encourage the generator to adversarially learn from the self-supervised discriminator by generating augmentation-predictable real and not fake data. This formulation connects the learning objective of the generator and the arithmetic $-$ harmonic mean divergence under certain assumptions. We compare our method with state-of-the-art (SOTA) methods using the class-conditional BigGAN and unconditional StyleGAN2 architectures on data-limited CIFAR-10, CIFAR-100, FFHQ, LSUN-Cat, and five low-shot datasets. Experimental results demonstrate significant improvements of our method over SOTA methods in training data-efficient GANs. Qi Cao 0005, Yige Yuan, Songtao Zhao, Chongyang Ma, Siyuan Pan, Pengfei Wan 0001, Huawei Shen, Xueqi Cheng 0001 |
NeurIPS | 7 |
| 2023 | Towards Practical Capture of High-Fidelity Relightable AvatarsabstractIn this paper, we propose a novel framework, Tracking-free Relightable Avatar (TRAvatar), for capturing and reconstructing high-fidelity 3D avatars. Compared to previous methods, TRAvatar works in a more practical and efficient setting. Specifically, TRAvatar is trained with dynamic image sequences captured in a Light Stage under varying lighting conditions, enabling realistic relighting and real-time animation for avatars in diverse scenes. Additionally, TRAvatar allows for tracking-free avatar capture and obviates the need for accurate surface tracking under varying illumination conditions. Our contributions are two-fold: First, we propose a novel network architecture that explicitly builds on and ensures the satisfaction of the linear nature of lighting. Trained on simple group light captures, TRAvatar can predict the appearance in real-time with a single forward pass, achieving high-quality relighting effects under illuminations of arbitrary environment maps. Second, we jointly optimize the facial geometry and relightable appearance from scratch based on image sequences, where the tracking is implicitly learned. This tracking-free approach brings robustness for establishing temporal correspondences between frames under different lighting conditions. Extensive qualitative and quantitative experiments demonstrate that our framework achieves superior performance for photorealistic avatar animation and relighting. Mingwu Zheng, Wanquan Feng, Yukun Lai, Pengfei Wan 0001, Zhongyuan Wang 0006, Chongyang Ma |
SIGGRAPH Asia | 6 |
| 2023 | Multi-Modal Face Stylization with a Generative PriorabstractAbstract In this work, we introduce a new approach for face stylization. Despite existing methods achieving impressive results in this task, there is still room for improvement in generating high‐quality artistic faces with diverse styles and accurate facial reconstruction. Our proposed framework, MMFS, supports multi‐modal face stylization by leveraging the strengths of StyleGAN and integrates it into an encoder‐decoder architecture. Specifically, we use the mid‐resolution and high‐resolution layers of StyleGAN as the decoder to generate high‐quality faces, while aligning its low‐resolution layer with the encoder to extract and preserve input facial details. We also introduce a two‐stage training strategy, where we train the encoder in the first stage to align the feature maps with StyleGAN and enable a faithful reconstruction of input faces. In the second stage, the entire network is fine‐tuned with artistic data for stylized face generation. To enable the fine‐tuned model to be applied in zero‐shot and one‐shot stylization tasks, we train an additional mapping network from the large‐scale Contrastive‐Language‐Image‐Pre‐training (CLIP) space to a latent w+ space of fine‐tuned StyleGAN. Qualitative and quantitative experiments show that our framework achieves superior performance in both one‐shot and zero‐shot face stylization tasks, outperforming state‐of‐the‐art methods by a large margin. Mengtian Li 0003, Minxuan Lin, Pengfei Wan 0001, Chongyang Ma |
Comput. Graph. Forum | 5 |
| 2023 | PMP-Net++: Point Cloud Completion by Transformer-Enhanced Multi-Step Point Moving PathsabstractPoint cloud completion concerns to predict missing part for incomplete 3D shapes. A common strategy is to generate complete shape according to incomplete input. However, unordered nature of point clouds will degrade generation of high-quality 3D shapes, as detailed topology and structure of unordered points are hard to be captured during the generative process using an extracted latent code. We address this problem by formulating completion as point cloud deformation process. Specifically, we design a novel neural network, named PMP-Net++, to mimic behavior of an earth mover. It moves each point of incomplete input to obtain a complete point cloud, where total distance of point moving paths (PMPs) should be the shortest. Therefore, PMP-Net++ predicts unique PMP for each point according to constraint of point moving distances. The network learns a strict and unique correspondence on point-level, and thus improves quality of predicted complete shape. Moreover, since moving points heavily relies on per-point features learned by network, we further introduce a transformer-enhanced representation learning network, which significantly improves completion performance of PMP-Net++. We conduct comprehensive experiments in shape completion, and further explore application on point cloud up-sampling, which demonstrate non-trivial improvement of PMP-Net++ over state-of-the-art point cloud completion/up-sampling methods. Xin Wen 0003, Peng Xiang 0002, Zhizhong Han, Yan-Pei Cao 0001, Pengfei Wan 0001, Yu-Shen Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Snowflake Point Deconvolution for Point Cloud Completion and Generation With Skip-TransformerabstractMost existing point cloud completion methods suffer from the discrete nature of point clouds and the unstructured prediction of points in local regions, which makes it difficult to reveal fine local geometric details. To resolve this issue, we propose SnowflakeNet with snowflake point deconvolution (SPD) to generate complete point clouds. SPD models the generation of point clouds as the snowflake-like growth of points, where child points are generated progressively by splitting their parent points after each SPD. Our insight into the detailed geometry is to introduce a skip-transformer in the SPD to learn the point splitting patterns that can best fit the local regions. The skip-transformer leverages attention mechanism to summarize the splitting patterns used in the previous SPD layer to produce the splitting in the current layer. The locally compact and structured point clouds generated by SPD precisely reveal the structural characteristics of the 3D shape in local patches, which enables us to predict highly detailed geometries. Moreover, since SPD is a general operation that is not limited to completion, we explore its applications in other generative tasks, including point cloud auto-encoding, generation, single image reconstruction, and upsampling. Our experimental results outperform state-of-the-art methods under widely used benchmarks. Peng Xiang 0002, Xin Wen 0003, Yu-Shen Liu, Yan-Pei Cao 0001, Pengfei Wan 0001, Zhizhong Han |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Predicting Personalized Head Movement From Short Video and Speech SignalabstractAudio-driven talking face video generation has attracted much attention recently. However, few existing works pay attention to machine learning of talking head movement, especially based on the phonetic study. Observing that real-world talking faces often accompany natural head movement, in this paper, we model the relation between speech signal and talking head movement, which is a typical one-to-many mapping problem. To solve this problem, we propose a novel two-step mapping strategy: (1) in the first step, we train an encoder that predicts a head motion behavior pattern (modeled as a feature vector) from the head motion sequence of a short video of 10–15 seconds, and (2) in the second step, we train a decoder that predict a unique head motion sequence from both the motion behavior pattern and the auditory features of an arbitrary speech signal. Based on the proposed mapping strategy, we build a deep neural network model that takes a speech signal of a source person and a short video of a target person as input, and outputs a synthesized high-fidelity talking face video with personalized head pose. Extensive experiments and a user study show that our method can generate high-quality personalized head movement in synthesized talking face videos, and meanwhile, has comparable facial animation quality (e.g., lip synchronization and expression) with the state-of-the-art methods. Ran Yi 0002, Zipeng Ye, Zhiyao Sun, Juyong Zhang, Guo-Xin Zhang, Pengfei Wan 0001, Hujun Bao, Yong-Jin Liu 0001 |
IEEE Trans. Multim. | 6 |
| 2022 | Assessing a Single Image in Reference-Guided Image SynthesisabstractAssessing the performance of Generative Adversarial Networks (GANs) has been an important topic due to its practical significance. Although several evaluation metrics have been proposed, they generally assess the quality of the whole generated image distribution. For Reference-guided Image Synthesis (RIS) tasks, i.e., rendering a source image in the style of another reference image, where assessing the quality of a single generated image is crucial, these metrics are not applicable. In this paper, we propose a general learning-based framework, Reference-guided Image Synthesis Assessment (RISA) to quantitatively evaluate the quality of a single generated image. Notably, the training of RISA does not require human annotations. In specific, the training data for RISA are acquired by the intermediate models from the training procedure in RIS, and weakly annotated by the number of models' iterations, based on the positive correlation between image quality and iterations. As this annotation is too coarse as a supervision signal, we introduce two techniques: 1) a pixel-wise interpolation scheme to refine the coarse labels, and 2) multiple binary classifiers to replace a naïve regressor. In addition, an unsupervised contrastive loss is introduced to effectively capture the style similarity between a generated image and its reference image. Empirical results on various datasets demonstrate that RISA is highly consistent with human preference and transfers well across models. Chaoqun Du, Jiangshan Wang, Huijuan Huang 0001, Pengfei Wan 0001, Gao Huang 0001 |
AAAI | 5 |
| 2022 | Exploring Set Similarity for Dense Self-supervised Representation LearningabstractBy considering the spatial correspondence, dense self-supervised representation learning has achieved superior performance on various dense prediction tasks. However, the pixel-level correspondence tends to be noisy because of many similar misleading pixels, e.g., backgrounds. To address this issue, in this paper, we propose to explore set similarity (SetSim) for dense self-supervised representation learning. We generalize pixel-wise similarity learning to set-wise one to improve the robustness because sets contain more semantic and structure information. Specifically, by resorting to attentional features of views, we establish the corresponding set, thus filtering out noisy backgrounds that may cause incorrect correspondences. Meanwhile, these at-tentional features can keep the coherence of the same image across different views to alleviate semantic inconsistency. We further search the cross-view nearest neighbours of sets and employ the structured neighbourhood information to enhance the robustness. Empirical evaluations demonstrate that SetSim surpasses or is on par with state-of-the-art meth-ods on object detection, keypoint detection, instance segmen-tation, and semantic segmentation. Zhaoqing Wang, Qiang Li 0024, Pengfei Wan 0001, Nannan Wang 0001, Mingming Gong, Tongliang Liu |
CVPR | 4 |
| 2022 | Wavelet Knowledge Distillation: Towards Efficient Image-to-Image TranslationabstractRemarkable achievements have been attained with Generative Adversarial Networks (GANs) in image-to-image translation. However, due to a tremendous amount of parameters, state-of-the-art GANs usually suffer from low efficiency and bulky memory usage. To tackle this challenge, firstly, this paper investigates GANs performance from a frequency perspective. The results show that GANs, especially small GANs lack the ability to generate high-quality high frequency information. To address this problem, we propose a novel knowledge distillation method referred to as wavelet knowledge distillation. Instead of directly distilling the generated images of teachers, wavelet knowledge distillation first decomposes the images into different frequency bands with discrete wavelet transformation and then only distills the high frequency bands. As a result, the student GAN can pay more attention to its learning on high frequency bands. Experiments demonstrate that our method leads to 7.08× compression and 6.80× acceleration on CycleGAN with almost no performance drop. Additionally, we have studied the relation between discriminators and generators which shows that the compression of discriminators can promote the performance of compressed generators. Linfeng Zhang 0001, Xin Chen 0071, Xiaobing Tu, Pengfei Wan 0001, Kaisheng Ma |
CVPR | 4 |
| 2022 | Learning an Inference-accelerated Network from a Pre-trained Model with Frequency-enhanced Feature DistillationabstractConvolution neural networks (CNNs) have achieved great success in various computer vision tasks, but they are still suffering from the heavy computation costs, which are mainly resulted from the substantial redundancy of the feature maps. In order to reduce these redundancy, we proposed a simple but effective frequency-enhanced feature distillation strategy to train an inference-accelerated network with a pre-trained model. Traditionally, one CNN can be regarded as a hierarchical structure, which can generate the low-level, middle-level and high-level feature maps from different convolution layers. In order to accelerate the inference time of CNNs, in this paper, we propose to resize the low-level and middle-level feature maps to smaller scales to reduce the spatial computation costs of CNNs. A frequency-enhanced feature distillation training strategy with a pre-trained model is then used to help the inference-accelerated network to maintain the core information after resizing the feature maps. To be specific, the original pre-trained network and the inference-accelerated network with resized feature maps are regarded as the teacher network and student network respectively. Considering that the low-frequency domain of the feature maps contribute the most parts to the final classification, we then transform the feature maps of different levels into a frequency-enhanced feature space, which highlights the low-frequency features for both the teacher and student networks. The frequency-enhanced features are used to transfer the knowledge from the teacher network to the student network. At the same time, knowledge for the final classification, i.e., the classification feature and predicted probabilities, are also used for distillation. Experiments on multiple databases based on various network structure types, e.g., ResNet, Res2Net, MobileNetV2, and ConvNeXt, have shown that with the proposed frequency-enhanced feature distillation training strategy, our method could get an inference-accelerated network with comparable performance and much less computation cost. Xuesong Niu, Jili Gu, Pengfei Wan 0001, Zhongyuan Wang 0006 |
ACM Multimedia | 4 |
| 2022 | Debiased Self-Training for Semi-Supervised LearningabstractDeep neural networks achieve remarkable performances on a wide range of tasks with the aid of large-scale labeled datasets. Yet these datasets are time-consuming and labor-exhaustive to obtain on realistic tasks. To mitigate the requirement for labeled data, self-training is widely used in semi-supervised learning by iteratively assigning pseudo labels to unlabeled samples. Despite its popularity, self-training is well-believed to be unreliable and often leads to training instability. Our experimental studies further reveal that the bias in semi-supervised learning arises from both the problem itself and the inappropriate training with potentially incorrect pseudo labels, which accumulates the error in the iterative self-training process. To reduce the above bias, we propose Debiased Self-Training (DST). First, the generation and utilization of pseudo labels are decoupled by two parameter-independent classifier heads to avoid direct error accumulation. Second, we estimate the worst case of self-training bias, where the pseudo labeling function is accurate on labeled samples, yet makes as many mistakes as possible on unlabeled samples. We then adversarially optimize the representations to improve the quality of pseudo labels by avoiding the worst case. Extensive experiments justify that DST achieves an average improvement of 6.3% against state-of-the-art methods on standard semi-supervised learning benchmark datasets and 18.9% against FixMatch on 13 diverse tasks. Furthermore, DST can be seamlessly adapted to other self-training methods and help stabilize their training and balance performance across classes in both cases of training from scratch and finetuning from pre-trained models. Baixu Chen, Junguang Jiang, Ximei Wang, Pengfei Wan 0001, Jianmin Wang 0001, Mingsheng Long |
NeurIPS | 4 |
| 2021 | Camera-Space Hand Mesh Recovery via Semantic Aggregation and Adaptive 2D-1D RegistrationabstractRecent years have witnessed significant progress in 3D hand mesh recovery. Nevertheless, because of the intrinsic 2D-to-3D ambiguity, recovering camera-space 3D information from a single RGB image remains challenging. To tackle this problem, we divide camera-space mesh recovery into two sub-tasks, i.e., root-relative mesh recovery and root recovery. First, joint landmarks and silhouette are extracted from a single input image to provide 2D cues for the 3D tasks. In the root-relative mesh recovery task, we exploit semantic relations among joints to generate a 3D mesh from the extracted 2D cues. Such generated 3D mesh coordinates are expressed relative to a root position, i.e., wrist of the hand. In the root recovery task, the root position is registered to the camera space by aligning the generated 3D mesh back to 2D cues, thereby completing camera-space 3D mesh recovery. Our pipeline is novel in that (1) it explicitly makes use of known semantic relations among joints and (2) it exploits 1D projections of the silhouette and mesh to achieve robust registration. Extensive experiments on popular datasets such as FreiHAND, RHD, and Human3.6M demonstrate that our approach achieves state-of-the-art performance on both root-relative mesh recovery and root recovery. Our code is publicly available at https://github.com/SeanChenxy/HandMesh. Chongyang Ma, Jianlong Chang, Huayan Wang, Pengfei Wan 0001 |
CVPR | 8 |
| 2021 | Cycle4Completion: Unpaired Point Cloud Completion Using Cycle Transformation With Missing Region CodingabstractIn this paper, we present a novel unpaired point cloud completion network, named Cycle4Completion, to infer the complete geometries from a partial 3D object. Previous unpaired completion methods merely focus on the learning of geometric correspondence from incomplete shapes to complete shapes, and ignore the learning in the reverse direction, which makes them suffer from low completion accuracy due to the limited 3D shape understanding ability. To address this problem, we propose two simultaneous cycle transformations between the latent spaces of complete shapes and incomplete ones. Specifically, the first cycle transforms shapes from incomplete domain to complete domain, and then projects them back to the incomplete domain. This process learns the geometric characteristic of complete shapes, and maintains the shape consistency between the complete prediction and the incomplete input. Similarly, the inverse cycle transformation starts from complete domain to incomplete domain, and goes back to complete domain to learn the characteristic of incomplete shapes. We experimentally show that our model with the learned bidirectional geometry correspondence outperforms state-of-the-art unpaired completion methods. Code will be available at https://github.com/diviswen/Cycle4Completion. Xin Wen 0003, Zhizhong Han, Yan-Pei Cao 0001, Pengfei Wan 0001, Yu-Shen Liu |
CVPR | 4 |
| 2021 | PMP-Net: Point Cloud Completion by Learning Multi-Step Point Moving PathsabstractThe task of point cloud completion aims to predict the missing part for an incomplete 3D shape. A widely used strategy is to generate a complete point cloud from the incomplete one. However, the unordered nature of point clouds will degrade the generation of high-quality 3D shapes, as the detailed topology and structure of discrete points are hard to be captured by the generative process only using a latent code. In this paper, we address the above problem by reconsidering the completion task from a new perspective, where we formulate the prediction as a point cloud deformation process. Specifically, we design a novel neural network, named PMP-Net, to mimic the behavior of an earth mover. It moves move each point of the incomplete input to complete the point cloud, where the total distance of point moving paths (PMP) should be shortest. Therefore, PMP-Net predicts a unique point moving path for each point according to the constraint of total point moving distances. As a result, the network learns a strict and unique correspondence on point-level, and thus improves the quality of the predicted complete shape. We conduct comprehensive experiments on Completion3D and PCN datasets, which demonstrate our advantages over the state-of-the-art point cloud completion methods. Code will be available at https://github.com/diviswen/PMP-Net. Xin Wen 0003, Peng Xiang 0002, Zhizhong Han, Yan-Pei Cao 0001, Pengfei Wan 0001, Yu-Shen Liu |
CVPR | 5 |
| 2021 | SnowflakeNet: Point Cloud Completion by Snowflake Point Deconvolution with Skip-TransformerabstractPoint cloud completion aims to predict a complete shape in high accuracy from its partial observation. However, previous methods usually suffered from discrete nature of point cloud and unstructured prediction of points in local regions, which makes it hard to reveal fine local geometric details on the complete shape. To resolve this issue, we propose SnowflakeNet with Snowflake Point Deconvolution (SPD) to generate the complete point clouds. The SnowflakeNet models the generation of complete point clouds as the snowflake-like growth of points in 3D space, where the child points are progressively generated by splitting their parent points after each SPD. Our insight of revealing detailed geometry is to introduce skip-transformer in SPD to learn point splitting patterns which can fit local regions the best. Skip-transformer leverages attention mechanism to summarize the splitting patterns used in the previous SPD layer to produce the splitting in the current SPD layer. The locally compact and structured point cloud generated by SPD is able to precisely capture the structure characteristic of 3D shape in local patches, which enables the network to predict highly detailed geometries, such as smooth regions, sharp edges and corners. Our experimental results outperform the state-of-the-art point cloud completion methods under widely used benchmarks. Code will be available at https://github.com/AllenXiangX/SnowflakeNet. Peng Xiang 0002, Xin Wen 0003, Yu-Shen Liu, Yan-Pei Cao 0001, Pengfei Wan 0001, Zhizhong Han |
ICCV | 5 |
| 2021 | BlendGAN: Implicitly GAN Blending for Arbitrary Stylized Face GenerationabstractGenerative Adversarial Networks (GANs) have made a dramatic leap in high-fidelity image synthesis and stylized face generation. Recently, a layer-swapping mechanism has been developed to improve the stylization performance. However, this method is incapable of fitting arbitrary styles in a single model and requires hundreds of style-consistent training images for each style. To address the above issues, we propose BlendGAN for arbitrary stylized face generation by leveraging a flexible blending strategy and a generic artistic dataset. Specifically, we first train a self-supervised style encoder on the generic artistic dataset to extract the representations of arbitrary styles. In addition, a weighted blending module (WBM) is proposed to blend face and style representations implicitly and control the arbitrary stylization effect. By doing so, BlendGAN can gracefully fit arbitrary styles in a unified model while avoiding case-by-case preparation of style-consistent training images. To this end, we also present a novel large-scale artistic face dataset AAHQ. Extensive experiments demonstrate that BlendGAN outperforms state-of-the-art methods in terms of visual quality and style diversity for both latent-guided and reference-guided stylized face synthesis. Mingcong Liu, Qiang Li 0024, Zekui Qin, Pengfei Wan 0001 |
NeurIPS | 5 |
| 2021 | Write-An-Animation: High-level Text-based Animation Editing with Character-Scene InteractionabstractAbstract 3D animation production for storytelling requires essential manual processes of virtual scene composition, character creation, and motion editing, etc. Although professional artists can favorably create 3D animations using software, it remains a complex and challenging task for novice users to handle and learn such tools for content creation. In this paper, we present Write‐An‐Animation, a 3D animation system that allows novice users to create, edit, preview, and render animations, all through text editing. Based on the input texts describing virtual scenes and human motions in natural languages, our system first parses the texts as semantic scene graphs, then retrieves 3D object models for virtual scene composition and motion clips for character animation. Character motion is synthesized with the combination of generative locomotions using neural state machine as well as template action motions retrieved from the dataset. Moreover, to make the virtual scene layout compatible with character motion, we propose an iterative scene layout and character motion optimization algorithm that jointly considers character‐object collision and interaction. We demonstrate the effectiveness of our system with customized texts and public film scripts. Experimental results indicate that our system can generate satisfactory animations from texts. Jia-Qi Zhang, Zhi-Meng Shen, Zehuan Huang, Yan-Pei Cao 0001, Pengfei Wan 0001, Miao Wang 0004 |
Comput. Graph. Forum | 7 |
| 2020 | Advancing Image Understanding in Poor Visibility Environments: A Collective Benchmark StudyabstractExisting enhancement methods are empirically expected to help the high-level end computer vision task: however, that is observed to not always be the case in practice. We focus on object or face detection in poor visibility enhancements caused by bad weathers (haze, rain) and low light conditions. To provide a more thorough examination and fair comparison, we introduce three benchmark sets collected in real-world hazy, rainy, and low-light conditions, respectively, with annotated objects/faces. We launched the UG2+ challenge Track 2 competition in IEEE CVPR 2019, aiming to evoke a comprehensive discussion and exploration about whether and how low-level vision techniques can benefit the high-level automatic visual recognition in various scenarios. To our best knowledge, this is the first and currently largest effort of its kind. Baseline results by cascading existing enhancement and detection models are reported, indicating the highly challenging nature of our new data as well as the large room for further technical innovations. Thanks to a large participation from the research community, we are able to analyze representative team solutions, striving to better identify the strengths and limitations of existing mindsets as well as the future directions. Wenhan Yang, Ye Yuan 0012, Wenqi Ren, Jiaying Liu 0001, Walter J. Scheirer, Zhangyang Wang, Taiheng Zhang, Qiaoyong Zhong, Di Xie, Shiliang Pu, Yuqiang Zheng, Yanyun Qu, Yuhong Xie, Hao Jiang 0014, Siyuan Yang 0001, Yan Liu 0041, Xiaochao Qu, Pengfei Wan 0001, Shuai Zheng 0005, Minhui Zhong, Taiyi Su, Lingzhi He, Yandong Guo, Yao Zhao 0001, Zhenfeng Zhu, Jinxiu Liang, Jingwen Wang 0003, Yuhui Quan, Yong Xu 0007, Bo Liu 0112, Xin Liu 0012, Tingyu Lin 0003, Xiaochuan Li 0001, Feng Lu 0005, Lin Gu 0003, Shengdi Zhou, Cong Cao 0005, Cheng Chi 0003, Chubin Zhuang, Zhen Lei 0001, Stan Z. Li, Shizheng Wang, Ruizhe Liu, Dong Yi, Zheming Zuo, Jianning Chi, Huan Wang 0014, Kai Wang 0036, Yixiu Liu, Xingyu Gao 0001, Zhenyu Chen 0003, Yongzhou Li, Huicai Zhong, Jing Huang 0017, Heng Guo 0003, Jianfei Yang 0001, Wenjuan Liao, Jiangang Yang, Liguo Zhou, Mingyue Feng, Likun Qin |
IEEE Trans. Image Process. | 21 |
| 2016 | Image Bit-Depth Enhancement via Maximum A Posteriori Estimation of AC SignalabstractWhen images at low bit-depth are rendered at high bit-depth displays, missing least significant bits needs to be estimated. We study the image bit-depth enhancement problem: estimating an original image from its quantized version from a minimum mean squared error (MMSE) perspective. We first argue that a graph-signal smoothness prior-one defined on a graph embedding the image structure-is an appropriate prior for the bit-depth enhancement problem. We next show that directly solving for the MMSE solution is, in general, too computationally expensive to be practical. We then propose an efficient approximation strategy. In particular, we first estimate the ac component of the desired signal in a maximum a posteriori formulation, efficiently computed via convex programming. We then compute the dc component with an MMSE criterion in a closed form given the computed ac component. Experiments show that our proposed two-step approach has improved performance over the conventional bit-depth enhancement schemes in both objective and subjective comparisons. Pengfei Wan 0001, Gene Cheung, Dinei A. F. Florêncio, Cha Zhang, Oscar C. Au |
IEEE Trans. Image Process. | 1 |
| 2015 | Motion vector fields based video codingabstractMotion vector fields (MVFs) are able to produce a more accurate prediction image than conventional block based motion compensation. However, MVFs are not used in conventional video coding standards due to the difficulty of efficient estimation and compression. In this work, we propose an MVF based video coding framework. We formulate the estimation of the MVF as a discrete optimization problem by both optimizing the residual energy and MVF smoothness, which can be efficiently solved by a graph cut algorithm with initialized motion vectors for each pixel. We then propose a modified rate distortion optimization approach for the MVF compression. Experimental results show that the proposed method has comparable performance in terms of object quality compared to the state-of-art of HEVC, while it has a better subjective performance by overcoming the block artifacts problem. Amin Zheng, Yuan Yuan 0002, Hong Zhang 0024, Haitao Yang 0001, Pengfei Wan 0001, Oscar C. Au |
ICIP | 5 |
| 2015 | Precision Enhancement of 3-D Surfaces from Compressed Multiview Depth MapsabstractTransmitting depth maps captured from multiple viewpoints of a 3-D scene enables a wide range of receiver-side 3-D applications, including virtual view synthesis via depth-image-based rendering (DIBR). Observing that compressed depth maps from different viewpoints constitute multiple descriptions (MD) of the same signal, we propose to reconstruct 3-D surfaces of the scene by considering multiple compressed depth maps jointly. Specifically, we propose an alternating projection algorithm, inspired by the theory of projection onto convex sets (POCS), which at convergence returns a 3-D surface that satisfies three sets of conditions: spatial smoothness prior, quantization bin constraints in the block transform domain, and inter-view consistency. We present a theoretical proof that shows convergence of our algorithm under benign conditions. Compared to existing multiview depth map denoising schemes and single image de-quantization schemes, our proposed solution achieves higher objective quality for both reconstructed depth maps and synthesized virtual views. Pengfei Wan 0001, Gene Cheung, Philip A. Chou, Dinei A. F. Florêncio, Cha Zhang, Oscar C. Au |
IEEE Signal Process. Lett. | 1 |
| 2014 | SSIM-based rate-distortion optimization in H.264abstractIn the current video coding standards, rate-distortion optimization (RDO) plays an important role in achieving best tradeoff between the perceived distortion and transmission rate. It is widely used in all kinds of encoder decisions, including block mode decision, motion vector selection and so on. Generally, the sum of absolute difference (SAD) or the sum of square difference (SSD) is used as the distortion measurement. However, it is well known that both of them cannot always reflect the perceptual quality of the encoded video. In this paper, an objective quality measurement structural similarity (SSIM) index is proposed as the distortion measurement in the RDO framework for video coding standards. By fully exploiting the relationship between SSIM and mean square error (MSE), the SSIM-based RDO framework can be approximated by the original SSD-based RDO framework with only a scaling of the Lagrange multiplier. Experimental results show that the proposed method outperforms the latest H.264 codec and also the state-of-the-art SSIM-based RDO video codec. Wei Dai 0002, Oscar C. Au, Pengfei Wan 0001, Wei Hu 0003, Jiantao Zhou 0001 |
ICASSP | 4 |
| 2014 | Palette-based compound image compression in HEVC by exploiting non-local spatial correlationabstractNon-camera captured images (also known as compound image) contain a mixture of camera-captured natural images and computer-generated graphics and texts. Nowadays, there are more and more applications calling for non-camera captured image/video compression scheme. However, current video coding standards, which are designed for natural video, treat non-camera captured video less carefully. For example, the state-of-the-art video coding standard High Efficiency Video Coding (HEVC) may blur or even remove edges in text/graphic region. A lot of schemes are proposed to preserve direction property of texts and graphics, such as palette-based intra coding. In this paper, a novel palette coding scheme is proposed for palette-based intra coding in HEVC. The palette in a block is predicted from an adaptive palette template, which records the statistical non-local spatial correlation of an image. Every block chooses its own palette using the palette template as the prediction in a rate-distortion optimized manner. Experimental results show that the proposed scheme can achieve up to 5.2% bit-rate saving compared to the state-of-the-art palette-based coding scheme in HEVC. Oscar C. Au, Wei Dai 0002, Haitao Yang 0001, Luheng Jia, Jin Zeng 0004, Pengfei Wan 0001 |
ICASSP | 8 |
| 2014 | DCT coefficients generation model for film grain noise and its application in super-resolutionabstractFilm grain noise (FGN) is generated by the procedure of capturing pictures using photographic film. Images with FGN are subjectively pleasing. However, FGN is difficult to compress and its pleasant features are difficult to preserve when the images are resized. So in literature, FGN is extracted first, then regenerated for the processed noise-free images. In this paper a new method is proposed to generate FGN. In contrast to some other models which generate FGN in spatial domain, our method captures the statistic feature of FGN in frequency domain. FGN is signal dependent and the signal independent scaled noise image (SNI) is obtained by scaling FGN by corresponding noise-free image. We model each discrete cosine transform (DCT) coefficient of SNI as a Gaussian random variable. The Gaussian model parameters can be estimated from the stack of all the blocks in SNI based on stationary assumption. Experimental results show that proposed model recovers FGN with similar visual and spectrum properties to the original FGN. We also apply proposed model in superresolution and the quality of resultant images are improved. Ting Sun 0001, Luhong Liang, King Hung Chiu, Pengfei Wan 0001, Oscar C. Au |
ICIP | 4 |
| 2014 | High bit-precision image acquisition and reconstruction by planned sensor distortionabstractWe present a novel framework for high bit-precision image acquisition and reconstruction. This framework is designed based on the inherent Markov property of image signals. In acquisition stage, we add planned sensor distortion (PSD) to the analog image signal before feeding it to A/D converters (or quantizers) in camera sensor. In reconstruction stage, the acquired quantized pixel values are jointly combined to get the reconstructed signal with reduced uncertainty range. Advantages of proposed PSD framework include 1) simplicity: it does not require any change to the core hardware of existing A/D converters; 2) effectiveness: experiment results demonstrate significant PSNR gain over traditional methods (up to 10 dB when quantizer bit-depth is relatively low); and 3) generality: this framework can also be applied for acquisition of other analog signals, including audio, video, etc. Pengfei Wan 0001, Oscar C. Au, Jiahao Pang, Ketan Tang |
ICIP | 1 |
| 2014 | Image bit-depth enhancement via maximum-a-posteriori estimation of graph AC componentabstractWhile modern displays offer high dynamic range (HDR) with large bit-depth for each rendered pixel, the bulk of legacy image and video contents were captured using cameras with shallower bit-depth. In this paper, we study the bit-depth enhancement problem for images, so that a high bit-depth (HBD) image can be reconstructed from an input low bit-depth (LBD) image. The key idea is to apply appropriate smoothing given the constraints that reconstructed signal must lie within the per-pixel quantization bins. Specifically, we first define smoothness via a signal-dependent graph Laplacian, so that natural image gradients can nonetheless be interpreted as low frequencies. Given defined smoothness prior and observed LBD image, we then demonstrate that computing the most probable signal via maximum a posteriori (MAP) estimation can lead to large expected distortion. However, we argue that MAP can still be used to efficiently estimate the AC component of the desired HBD signal, which along with a distortion-minimizing DC component, can result in a good approximate solution that minimizes the expected distortion. Experimental results show that our proposed method outperforms existing bit-depth enhancement methods in terms of reconstruction error. Pengfei Wan 0001, Gene Cheung, Dinei A. F. Florêncio, Cha Zhang, Oscar C. Au |
ICIP | 1 |
| 2014 | A fast intermode decision algorithm based on analysis of inter prediction residualabstractRate-distortion-optimized (RDO) intermode decision is one of the most effective tools that greatly improves the coding performance of modern coding standards, for example, H.264/AVC and HEVC. However, RDO intermode decision also leads to extremely intense computation. To reduce the complexity, a fast intermode decision algorithm is presented in this paper. Mathematical analysis of inter prediction residual is performed which explicitly shows the impact of edge information and motion characteristics to the prediction accuracy. Moreover, It is shown that with fixed quantization step that minimizing of R-D costs over different partition types is equivalent to minimizing the variance of transformed residual which can be expressed by motion vector and edge gradient components. In consequence, the complex calculation of rate and distortion is replaced by simple pre-analysis of video content. The repetition of motion estimation (ME) and entropy coding process are avoided. Inspired by the theoretical analysis, a fast inter mode decision algorithm is proposed. Experimental results show that the fast method achieves considerable complexity reduction with negligible coding performance degradation. Luheng Jia, Oscar C. Au, Chi-Ying Tsui, Wei Dai 0002, Pengfei Wan 0001 |
MMSP | 5 |
| 2014 | Solving dense stereo matching via quadratic programmingabstractWe study the problem of formulating the discrete dense stereo matching using continuous convex optimization. One of the previous work derived a relaxed convex formulation by establishing the relationship between the disparity vector and a warping matrix. However it suffers from high computational complexity. In this paper, the previous convex formulation is translated into an equivalent quadratic programming (QP). Then redundant variables and constraints are eliminated by exploiting the internal sparse property of the warping matrix. The resulting QP can be efficiently tackled using interior point solvers. Moreover, enhanced smoothness term and effective post-processing procedures are also incorporated to further improve the disparity accuracy. Experimental results show that the proposed method is much faster and better than the previous convex formulation, and provides competitive results against existing convex approaches. Oscar C. Au, Pengfei Wan 0001, Wenxiu Sun, Lingfeng Xu, Luheng Jia |
VCIP | 3 |
| 2013 | 3D motion in visual saliency modelingabstractVisual saliency is a probabilistic estimate of how likely a given spatial area in an image or video is to attract human visual attention relative to other areas. Bottom-up saliency models aggregate low-level image features like luminance and color contrast, flicker, 2D motion, etc. to construct a plausible saliency map. In this paper, we introduce 3D motion (object movements towards or away from the observer) into bottom-up video saliency modeling. Given availability of per-pixel depth maps, we first propose a novel algorithm to estimate 3D motion vectors (3DMVs) for arbitrarily shaped sub-blocks in texture-plus-depth videos. We then derive two feature channels from 3DMVs to be incorporated into a widely accepted bottom-up saliency model. Experiments on subjective quality of Region-of-Interest (ROI) based video coding show that our enriched saliency model with 3DMV channels is more accurate in estimating human visual attention. Pengfei Wan 0001, Yunlong Feng, Gene Cheung, Ivan V. Bajic, Oscar C. Au, Yusheng Ji |
ICASSP | 1 |
| 2013 | A robust interpolation-free approach for sub-pixel accuracy motion estimationabstractMotion estimation (ME) is one of the key elements in video coding standard which eliminates the temporal redundancy by using a motion vector (MV) to indicate the best match between the current frame and reference frame. A coarse to fine process is taken to find the best MV. First of all, integer-pixel ME finds a coarse MV and followed by the sub-pixel ME around the best integer-pixel point. The sub-pixel ME plays an important role in improving the coding efficiency. However, the computational complexity of searching one sub-pixel point is much higher than the integer-pixel point searching because of the interpolation and Hadamard transform operation. In this paper, an accurate optimal sub-pixel position prediction algorithm is presented. With the information of the 8 neighboring integer-pixel points, the optimal sub-pixel position is predicted directly without explicitly solving model parameters. Moreover, an outlier rejection scheme is applied to improve the robustness of the proposed algorithm. Experimental results show that the proposed algorithm outperforms the state of the art interpolation-freesub-pixel ME algorithms. Wei Dai 0002, Oscar C. Au, Wei Hu 0003, Pengfei Wan 0001 |
ICIP | 5 |
| 2013 | Optimal dependent bit allocation for AVS intra-frame coding via successive convex approximationabstractWe consider the optimal dependent bit allocation strategy for AVS intra-frame coding. Due to the block-based predictive coding, the rate-distortion (R-D) characteristics of neighboring blocks are dependent with each other. However, the interblock coding dependency is neglected in most of the existing bit allocation methods. Different from the conventional methods, the proposed method fully exploit the interblock coding dependency and carefully leverage it in the problem formulation. Then successive convex optimization techniques are employed to convert the original nonconvex optimization problem into a series of convex optimization problems which can be solved efficiently and optimally. Experimental results have proved the superiority of the proposed method in terms of significant R-D performance improvement. Oscar C. Au, Feng Zou 0006, Wei Hu 0003, Pengfei Wan 0001 |
ICIP | 6 |
| 2013 | Personal photo album compression and managementabstractThe advance in multimedia technologies have resulted in an explosive growth of pictures in personal computers and in cloud. Typically many pictures taken in the same occasion are similar. The cost to store and transmit them can be very significant. Thus it is important to find an efficient method to store these pictures. This paper proposed a compression scheme for similar images. Our approach is to arrange all the similar images into tree structure then apply video coding technique along each branch. To maximize the inter-image correlation between adjacent photos, we consider the minimum spanning tree (MST) subjecting to a maximum depth limit to ensure fast access to all images. This structure is encoded by the latest video coding technique High Efficiency Video Coding (HEVC), which is reported to has advantage in high definition video/image compression. It also supports deleting, adding and modifying images. Experiments show that the proposed method saved 75% space comparing to JPEG format. Ruobing Zou, Oscar C. Au, Guyue Zhou, Wei Dai 0002, Wei Hu 0003, Pengfei Wan 0001 |
ISCAS | 6 |
| 2013 | 3-D Motion Estimation for Visual Saliency ModelingabstractVisual saliency is a probabilistic estimate of how likely a spatial area in an image or video frame is to attract human visual attention relative to other areas. When existing bottom-up saliency models aggregate low-level features to construct a plausible saliency map, only 2-D motion cues are used as motion features, even though videos typically capture dynamic 3-D scenes. In this paper, we introduce 3-D motion into bottom-up saliency modeling for texture-plus-depth videos. We first propose an efficient 3-D motion estimation algorithm, which computes a 3-D motion vector (3DMV) for each sub-block in the frame. Using the computed 3DMVs, we then derive several saliency channels (called 3DMV channels), which are incorporated into a bottom-up saliency model to obtain enhanced saliency maps. Experiments tracking human gaze show that incorporating our 3DMV channels into bottom-up saliency model significantly improves the accuracy of derived saliency maps. Pengfei Wan 0001, Yunlong Feng, Gene Cheung, Ivan V. Bajic, Oscar C. Au |
IEEE Signal Process. Lett. | 1 |
| 2012 | Image de-quantization via spatially varying sparsity priorabstractWe address the problem of image de-quantization, which is also known as bit-depth expansion if the reconstructed 2D signal is re-quantized into higher bit-precision. In this paper, a novel image de-quantization method based on convex optimization theory is proposed, which exploits the spatially varying characteristics of image surface. We test our method on image bit-depth expansion problems, and the experimental results show that proposed method can achieve superior PSNR and SSIM performance. Pengfei Wan 0001, Oscar C. Au, Ketan Tang, Yuanfang Guo |
ICIP | 1 |
| 2012 | From 2D Extrapolation to 1D Interpolation: Content Adaptive Image Bit-Depth ExpansionabstractIn this paper, we address the problem of image bit-depth expansion and present a novel method to generate high bit-depth (HBD) images from a single low bit-depth (LBD) image. We expand image bit-depth by reconstructing the least significant bits (LSBs) for the LBD image after it is rescaled to high bit-depth. For image regions whose intensities are neither locally maximum nor minimum, neighborhood flooding is applied to convert 2D interpolation problem into 1D interpolation, for local maxima/minima (LMM) regions where interpolation is not applicable, a virtual skeleton marking algorithm is proposed to convert problematic 2D extrapolation problem into 1D interpolation. At last, a content-adaptive reconstruction model is proposed to obtain the output HBD image. The experimental results show that proposed method significantly outperforms existing methods in PSNR and SSIM without contouring artifacts. Pengfei Wan 0001, Oscar C. Au, Ketan Tang, Yuanfang Guo, Lu Fang 0001 |
ICME | 1 |