VLDB 2026 Research / reviewers in the wild / expert
Menghan Xia
dblp:169/4908
· DBLP profile ↗
44ranked-venue papers
7as first author
34since 2021 · last 2026
0000-0001-9664-4967ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 32 · 6 first-author · 25 since 2021Artificial intelligence and machine learning · 22 · 3 first-author · 18 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FreeTraj: Tuning-Free Trajectory Control via Noise Guided Video Diffusion
Haonan Qiu, Zhaoxi Chen 0009, Zhouxia Wang, Yingqing He, Menghan Xia, Ziwei Liu 0002 |
Int. J. Comput. Vis. | 5 |
| 2025 | PatchVSR: Breaking Video Diffusion Resolution Limits with Patch-wise Video Super-ResolutionabstractPre-trained video generation models hold great potential for generative video super-resolution (VSR). However, adapting them for full-size VSR, as most existing methods do, suffers from unnecessary intensive full-attention computation and fixed output resolution. To overcome these limitations, we make the first exploration into utilizing video diffusion priors for patch-wise VSR. This is non-trivial because pre-trained video diffusion models are not native for patch-level detail generation. To mitigate this challenge, we propose an innovative approach, called PatchVSR, which integrates a dual-stream adapter for conditional guidance. The patch branch extracts features from input patches to maintain content fidelity while the global branch extracts context features from the resized full video to bridge the generation gap caused by incomplete semantics of patches. Particularly, we also inject the patch’s location information into the model to better contextualize patch synthesis within the global video frame. Experiments demonstrate that our method can synthesize high-fidelity, high-resolution details at the patch level. A tailor-made multi-patch joint modulation is proposed to ensure visual consistency across individually enhanced patches. Due to the flexibility of our patch-based paradigm, we can achieve highly competitive 4K VSR based on a 512×512 resolution base model, with extremely high efficiency. Shian Du, Menghan Xia, Chang Liu 0071, Xintao Wang 0002, Jing Wang 0021, Pengfei Wan 0001, Di Zhang 0026, Xiangyang Ji |
CVPR | 2 |
| 2025 | Recammaster: Camera-Controlled Generative Rendering From a Single VideoabstractCamera control has been actively studied in text or image conditioned video generation tasks. However, altering camera trajectories of a given video remains under-explored, despite its importance in the field of video creation. It is non-trivial due to the extra constraints of maintaining multiple-frame appearance and dynamic synchronization. To address this, we present ReCamMaster, a camera-controlled generative video re-rendering framework that reproduces the dynamic scene of an input video at novel camera trajectories. The core innovation lies in harnessing the generative capabilities of pre-trained text-to-video models through a simple yet powerful video conditioning mechanism--its capability is often overlooked in current research. To overcome the scarcity of qualified training data, we construct a comprehensive multi-camera synchronized video dataset using Unreal Engine 5, which is carefully curated to follow real-world filming characteristics, covering diverse scenes and camera movements. It helps the model generalize to in-the-wild videos. Lastly, we further improve the robustness to diverse inputs through a meticulously designed training strategy. Extensive experiments show that our method substantially outperforms existing state-of-the-art approaches. Our method also finds promising applications in video stabilization, super-resolution, and outpainting. Our code and dataset are publicly available at: https://github.com/KwaiVGI/ReCamMaster. Jianhong Bai, Menghan Xia, Xintao Wang 0002, Lianrui Mu, Jinwen Cao, Zuozhu Liu, Haoji Hu, Xiang Bai, Pengfei Wan 0001, Di Zhang 0026 |
ICCV | 2 |
| 2025 | SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse ViewpointsabstractRecent advancements in video diffusion models demonstrate remarkable capabilities in simulating real-world dynamics and 3D consistency. This progress motivates us to explore the potential of these models to maintain dynamic consistency across diverse viewpoints, a feature highly sought after in applications like virtual filming. Unlike existing methods focused on multi-view generation of single objects for 4D reconstruction, our interest lies in generating open-world videos from arbitrary viewpoints, incorporating six degrees of freedom (6 DoF) camera poses.
To achieve this, we propose a plug-and-play module that enhances a pre-trained text-to-video model for multi-camera video generation, ensuring consistent content across different viewpoints. Specifically, we introduce a multi-view synchronization module designed to maintain appearance and geometry consistency across these viewpoints. Given the scarcity of high-quality training data, we also propose a progressive training scheme that leverages multi-camera images and monocular videos as a supplement to Unreal Engine-rendered multi-camera videos. This comprehensive approach significantly benefits our model.
Experimental results demonstrate the superiority of our proposed method over existing competitors and several baselines. Furthermore, our method enables intriguing extensions, such as re-rendering a video from multiple novel viewpoints. Project webpage: https://jianhongbai.github.io/SynCamMaster/ Jianhong Bai, Menghan Xia, Xintao Wang 0002, Ziyang Yuan, Zuozhu Liu, Haoji Hu, Pengfei Wan 0001, Di Zhang 0026 |
ICLR | 2 |
| 2025 | 3DTrajMaster: Mastering 3D Trajectory for Multi-Entity Motion in Video GenerationabstractThis paper aims to manipulate multi-entity 3D motions in video generation. Previous methods on controllable video generation primarily leverage 2D control signals to manipulate object motions and have achieved remarkable synthesis results. However, 2D control signals are inherently limited in expressing the 3D nature of object motions. To overcome this problem, we introduce 3DTrajMaster, a robust controller that regulates multi-entity dynamics in 3D space, given user-desired 6DoF pose (location and rotation) sequences of entities. At the core of our approach is a plug-and-play 3D-motion grounded object injector that fuses multiple input entities with their respective 3D trajectories through a gated self-attention mechanism. In addition, we exploit an injector architecture to preserve the video diffusion prior, which is crucial for generalization ability. To mitigate video quality degradation, we introduce a domain adaptor during training and employ an annealed sampling strategy during inference. To address the lack of suitable training data, we construct a 360-Motion Dataset, which first correlates collected 3D human and animal assets with GPT-generated trajectory and then captures their motion with 12 evenly-surround cameras on diverse 3D UE platforms. Extensive experiments show that 3DTrajMaster sets a new state-of-the-art in both accuracy and generalization for controlling multi-entity 3D motions. Project page: http://fuxiao0719.github.io/projects/3dtrajmaster Xintao Wang 0002, Sida Peng, Menghan Xia, Xiaoyu Shi 0002, Ziyang Yuan, Pengfei Wan 0001, Di Zhang 0026, Dahua Lin |
ICLR | 5 |
| 2025 | Improving Video Generation with Human FeedbackabstractVideo generation has achieved significant advances through rectified flow techniques, but issues like unsmooth motion and misalignment between videos and prompts persist. In this work, we develop a systematic pipeline that harnesses human feedback to mitigate these problems and refine the video generation model. Specifically, we begin by constructing a large-scale human preference dataset focused on modern video generation models, incorporating pairwise annotations across multi-dimensions. We then introduce VideoReward, a multi-dimensional video reward model, and examine how annotations and various design choices impact its rewarding efficacy. From a unified reinforcement learning perspective aimed at maximizing reward with KL regularization, we introduce three alignment algorithms for flow-based models. These include two training-time strategies: direct preference optimization for flow (Flow-DPO) and reward weighted regression for flow (Flow-RWR), and an inference-time technique, Flow-NRG, which applies reward guidance directly to noisy videos. Experimental results indicate that VideoReward significantly outperforms existing reward models, and Flow-DPO demonstrates superior performance compared to both Flow-RWR and supervised fine-tuning methods. Additionally, Flow-NRG lets users assign custom weights to multiple objectives during inference, meeting personalized video quality needs. Jie Liu 0047, Gongye Liu, Jiajun Liang, Ziyang Yuan, Xiaokun Liu, Mingwu Zheng, Xiele Wu, Qiulin Wang, Menghan Xia, Xintao Wang 0002, Xiaohong Liu 0001, Pengfei Wan 0001, Di Zhang 0026, Kun Gai, Yujiu Yang 0001, Wanli Ouyang |
NeurIPS | 9 |
| 2025 | CamCloneMaster: Enabling Reference-based Camera Control for Video GenerationabstractCamera control is crucial for generating expressive and cinematic videos. Existing methods rely on explicit sequences of camera parameters as control conditions, which can be cumbersome for users to construct, particularly for intricate camera movements. To provide a more intuitive camera control method, we propose CamCloneMaster, a framework that enables users to replicate camera movements from reference videos without requiring camera parameters or test-time fine-tuning. CamCloneMaster seamlessly supports reference-based camera control for both Image-to-Video and Video-to-Video tasks within a unified framework. Furthermore, we present the Camera Clone Dataset, a large-scale synthetic dataset designed for camera clone learning, encompassing diverse scenes, subjects, and camera movements. Extensive experiments and user studies demonstrate that CamCloneMaster outperforms existing methods in terms of both camera controllability and visual quality. Dataset and Code can be found at https://camclonemaster.github.io/. Yawen Luo, Xiaoyu Shi 0002, Jianhong Bai, Menghan Xia, Tianfan Xue, Xintao Wang 0002, Pengfei Wan 0001, Di Zhang 0026, Kun Gai |
SIGGRAPH Asia | 4 |
| 2025 | Screentone-Preserved Manga RetargetingabstractAbstract As a popular comic style, manga offers a unique impression by utilizing a rich set ofbitonal patterns, or screentones, for illustration. However, screentones can easily be degraded when manga is resized in terms of aspect ratio and resolution for manga re‐layout and e‐manga migration applications. To tackle this problem, we propose the first automatic manga retargeting method that synthesizes a retargeted manga image while preserving the prominent structure and fine screentone intended by the manga artist. While modern natural photo retargeting methods can achieve prominent structure preservation, preserving screentones within arbitrarily shaped regions is very challenging due to two properties of manga: (i) pattern constancy under translation, and (ii) non‐compatibility with interpolation. To circumvent this barrier, we propose learning a quantized representation of screentones that is translation‐invariant and pointwisely representable through a tailored manga reconstruction network with a screentone‐anchored codebook. Thanks to these merits, we can perform the re‐synthesis operation using existing photo retargeting methods and achieve the desired manga retargeting results. We conducted extensive qualitative and quantitative experiments to validate the effectiveness of our method, and we achieved notably compelling results compared to alternative methods. Minshan Xie, Menghan Xia, Chengze Li, Xueting Liu 0001, Tien-Tsin Wong |
Comput. Graph. Forum | 2 |
| 2025 | T2EA: Target-Aware Taylor Expansion Approximation Network for Infrared and Visible Image FusionabstractIn the image fusion mission, the crucial task is to generate high-quality images for highlighting the key objects while enhancing the scenes to be understood. To complete this task and provide a powerful interpretability as well as a strong generalization ability in producing enjoyable fusion results which are comfortable for vision tasks (such as objects detection and their segmentation), we present a novel interpretable decomposition scheme and develop a target-aware Taylor expansion approximation (T2EA) network for infrared and visible image fusion, where our T2EA includes the following key procedures: Firstly, visible and infrared images are both decomposed into feature maps through a designed Taylor expansion approximation (TEA) network. Then, the Taylor feature maps are hierarchically fused by a dual-branch feature fusion (DBFF) network. Next, the fused map of each layer is contributed to synthesize an enjoyable fusion result by the inverse Taylor expansion. Finally, a segmentation network is jointed to refine the fusion network parameters which can promote the pleasing fusion results to be more suitable for segmenting the objects. To validate the effectiveness of our reported T2EA network, we first discuss the selection of Taylor expansion layers and fusion strategies. Then, both quantitatively and qualitatively experimental results generated by the selected SOTA approaches on three datasets (MSRS, TNO, andLLVIP) are compared in testing, generalization, and target detection and segmentation, demonstrating that our T2EA can produce more competitive fusion results for vision tasks and is more powerful for image adaption. The code will be available at https://github.com/MysterYxby/T2EA. Zhenghua Huang, Biyun Xu, Menghan Xia, Qian Li 0019, Yansheng Li 0001, Nong Sang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | CodePhys: Robust Video-Based Remote Physiological Measurement Through Latent Codebook QueryingabstractRemote photoplethysmography (rPPG) aims to measure non-contact physiological signals from facial videos, which has shown great potential in many applications. Most existing methods directly extract video-based rPPG features by designing neural networks for heart rate estimation. Although they can achieve acceptable results, the recovery of rPPG signal faces intractable challenges when interference from real-world scenarios takes place on facial video. Specifically, facial videos are inevitably affected by non-physiological factors (e.g., camera device noise, defocus, and motion blur), leading to the distortion of extracted rPPG signals. Recent rPPG extraction methods are easily affected by interference and degradation, resulting in noisy rPPG signals. In this paper, we propose a novel method named CodePhys, which innovatively treats rPPG measurement as a code query task in a noise-free proxy space (i.e., codebook) constructed by ground-truth PPG signals. We consider noisy rPPG features as queries and generate high-fidelity rPPG features by matching them with noise-free PPG features from the codebook. Our approach also incorporates a spatial-aware encoder network with a spatial attention mechanism to highlight physiologically active areas and uses a distillation loss to reduce the influence of non-periodic visual interference. Experimental results on four benchmark datasets demonstrate that CodePhys outperforms state-of-the-art methods in both intra-dataset and cross-dataset settings. Shuyang Chu, Menghan Xia, Mengyao Yuan, Xin Liu 0012, Tapio Seppänen, Guoying Zhao 0001, Jingang Shi |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | Make-Your-Video: Customized Video Generation Using Textual and Structural GuidanceabstractCreating a vivid video from the event or scenario in our imagination is a truly fascinating experience. Recent advancements in text-to-video synthesis have unveiled the potential to achieve this with prompts only. While text is convenient in conveying the overall scene context, it may be insufficient to control precisely. In this paper, we explore customized video generation by utilizing text as context description and motion structure (e.g., frame-wise depth) as concrete guidance. Our method, dubbed Make-Your-Video, involves joint-conditional video generation using a Latent Diffusion Model that is pre-trained for still image synthesis and then promoted for video generation with the introduction of temporal modules. This two-stage learning scheme not only reduces the computing resources required, but also improves the performance by transferring the rich concepts available in image datasets solely into video generation. Moreover, we use a simple yet effective causal attention mask strategy to enable longer video synthesis, which mitigates the potential quality degradation effectively. Experimental results show the superiority of our method over existing baselines, particularly in terms of temporal coherence and fidelity to users' guidance. In addition, our model enables several intriguing applications that demonstrate potential for practical usage. Jinbo Xing, Menghan Xia, Yuechen Zhang, Yong Zhang 0034, Yingqing He, Hanyuan Liu, Haoxin Chen, Xiaodong Cun, Xintao Wang 0002, Ying Shan, Tien-Tsin Wong |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion ModelsabstractText-to-video generation aims to produce a video based on a given prompt. Recently, several commercial video models have been able to generate plausible videos with mini-mal noise, excellent details, and high aesthetic scores. However, these models rely on large-scale, well-filtered, high-quality videos that are not accessible to the community. Many existing research works, which train models using the low-quality WebVid-10M dataset, struggle to generate high-quality videos because the models are optimized to fit WebVid-10M. In this work, we explore the training scheme of video models extended from Stable Diffusion and investigate the feasibility of leveraging low-quality videos and synthesized high-quality images to obtain a high-quality video model. We first analyze the connection between the spatial and temporal modules of video models and the distribution shift to low-quality videos. We observe that full training of all modules results in a stronger coupling between spatial and temporal modules than only training temporal modules. Based on this stronger coupling, we shift the distribution to higher quality without motion degradation by finetuning spatial modules with high-quality images, resulting in a generic high-quality video model. Evaluations are conducted to demonstrate the superiority of the proposed method, particularly in picture quality, motion, and concept composition. Haoxin Chen, Yong Zhang 0034, Xiaodong Cun, Menghan Xia, Xintao Wang 0002, Chao Weng, Ying Shan |
CVPR | 4 |
| 2024 | Make a Cheap Scaling: A Self-Cascade Diffusion Model for Higher-Resolution Adaptation
Lanqing Guo, Yingqing He, Haoxin Chen, Menghan Xia, Xiaodong Cun, Yufei Wang 0006, Siyu Huang, Yong Zhang 0034, Xintao Wang 0002, Qifeng Chen 0001, Ying Shan, Bihan Wen |
ECCV (36) | 4 |
| 2024 | DynamiCrafter: Animating Open-Domain Images with Video Diffusion Priors
Jinbo Xing, Menghan Xia, Yong Zhang 0034, Hao Chen 0011, Wangbo Yu, Hanyuan Liu, Gongye Liu, Xintao Wang 0002, Ying Shan, Tien-Tsin Wong |
ECCV (46) | 2 |
| 2024 | Noise Calibration: Plug-and-Play Content-Preserving Video Enhancement Using Pre-trained Video Diffusion Models
Qinyu Yang, Hao Chen 0011, Yong Zhang 0034, Menghan Xia, Xiaodong Cun, Zhixun Su, Ying Shan |
ECCV (36) | 4 |
| 2024 | ScaleCrafter: Tuning-free Higher-Resolution Visual Generation with Diffusion ModelsabstractIn this work, we investigate the capability of generating images from pre-trained diffusion models at much higher resolutions than the training image sizes. In addition, the generated images should have arbitrary image aspect ratios. When generating images directly at a higher resolution, 1024 x 1024, with the pre-trained Stable Diffusion using training images of resolution 512 x 512, we observe persistent problems of object repetition and unreasonable object structures. Existing works for higher-resolution generation, such as attention-based and joint-diffusion approaches, cannot well address these issues. As a new perspective, we examine the structural components of the U-Net in diffusion models and identify the crucial cause as the limited perception field of convolutional kernels. Based on this key observation, we propose a simple yet effective re-dilation that can dynamically adjust the convolutional perception field during inference. We further propose the dispersed convolution and noise-damped classifier-free guidance, which can enable ultra-high-resolution image generation (e.g., 4096 x 4096). Notably, our approach does not require any training or optimization. Extensive experiments demonstrate that our approach can address the repetition issue well and achieve state-of-the-art performance on higher-resolution image synthesis, especially in texture details. Our work also suggests that a pre-trained diffusion model trained on low-resolution images can be directly used for high-resolution visual generation without further tuning, which may provide insights for future research on ultra-high-resolution image and video synthesis. More results are available at the anonymous website: https://scalecrafter.github.io/ScaleCrafter/ Yingqing He, Shaoshu Yang, Haoxin Chen, Xiaodong Cun, Menghan Xia, Yong Zhang 0034, Xintao Wang 0002, Ran He 0001, Qifeng Chen 0001, Ying Shan |
ICLR | 5 |
| 2024 | FreeNoise: Tuning-Free Longer Video Diffusion via Noise ReschedulingabstractWith the availability of large-scale video datasets and the advances of diffusion models, text-driven video generation has achieved substantial progress. However, existing video generation models are typically trained on a limited number of frames, resulting in the inability to generate high-fidelity long videos during inference. Furthermore, these models only support single-text conditions, whereas real-life scenarios often require multi-text conditions as the video content changes over time. To tackle these challenges, this study explores the potential of extending the text-driven capability to generate longer videos conditioned on multiple texts. 1) We first analyze the impact of initial noise in video diffusion models. Then building upon the observation of noise, we propose FreeNoise, a tuning-free and time-efficient paradigm to enhance the generative capabilities of pretrained video diffusion models while preserving content consistency. Specifically, instead of initializing noises for all frames, we reschedule a sequence of noises for long-range correlation and perform temporal attention over them by window-based fusion. 2) Additionally, we design a novel motion injection method to support the generation of videos conditioned on multiple text prompts. Extensive experiments validate the superiority of our paradigm in extending the generative capabilities of video diffusion models. It is noteworthy that compared with the previous best-performing method which brought about 255% extra time cost, our method incurs only negligible time cost of approximately 17%. Generated video samples are available at our website: http://haonanqiu.com/projects/FreeNoise.html. Haonan Qiu, Menghan Xia, Yong Zhang 0034, Yingqing He, Xintao Wang 0002, Ying Shan, Ziwei Liu 0002 |
ICLR | 2 |
| 2024 | Sketch Video SynthesisabstractAbstract Understanding semantic intricacies and high‐level concepts is essential in image sketch generation, and this challenge becomes even more formidable when applied to the domain of videos. To address this, we propose a novel optimization‐based framework for sketching videos represented by the frame‐wise Bézier Curves. In detail, we first propose a cross‐frame stroke initialization approach to warm up the location and the width of each curve. Then, we optimize the locations of these curves by utilizing a semantic loss based on CLIP features and a newly designed consistency loss using the self‐decomposed 2D atlas network. Built upon these design elements, the resulting sketch video showcases notable visual abstraction and temporal coherence. Furthermore, by transforming a video into vector lines through the sketching process, our method unlocks applications in sketch‐based video editing and video doodling, enabled through video composition. Yudian Zheng, Xiaodong Cun, Menghan Xia, Chi-Man Pun |
Comput. Graph. Forum | 3 |
| 2024 | Measurement Guidance in Diffusion Models: Insight from Medical Image SynthesisabstractIn the field of healthcare, the acquisition of sample is usually restricted by multiple considerations, including cost, labor- intensive annotation, privacy concerns, and radiation hazards, therefore, synthesizing images-of-interest is an important tool to data augmentation. Diffusion models have recently attained state-of-the-art results in various synthesis tasks, and embedding energy functions has been proved that can effectively guide the pre-trained model to synthesize target samples. However, we notice that current method development and validation are still limited to improving indicators, such as Fréchet Inception Distance score (FID) and Inception Score (IS), and have not provided deeper investigations on downstream tasks, like disease grading and diagnosis. Moreover, existing classifier guidance which can be regarded as a special case of energy function can only has a singular effect on altering the distribution of the synthetic dataset. This may contribute to in-distribution synthetic sample that has limited help to downstream model optimization. All these limitations remind that we still have a long way to go to achieve controllable generation. In this work, we first conducted an analysis on previous guidance as well as its contributions on further applications from the perspective of data distribution. To synthesize samples which can help downstream applications, we then introduce uncertainty guidance in each sampling step and design an uncertainty-guided diffusion models. Extensive experiments on four medical datasets, with ten classic networks trained on the augmented sample sets provided a comprehensive evaluation on the practical contributions of our methodology. Furthermore, we provide a theoretical guarantee for general gradient guidance in diffusion models, which would benefit future research on investigating other forms of measurement guidance for specific generative tasks. Yimin Luo, Qinyu Yang, Haikun Qi, Menghan Xia |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | StyleCrafter: Taming Artistic Video Diffusion with Reference-Augmented Adapter LearningabstractText-to-video (T2V) models have shown remarkable capabilities in generating diverse videos. However, they struggle to produce user-desired artistic videos due to (i) text's inherent clumsiness in expressing specific styles and (ii) the generally degraded style fidelity. To address these challenges, we introduce StyleCrafter, a generic method that enhances pretrained T2V models with a style control adapter, allowing video generation in any style by feeding a reference image. Considering the scarcity of artistic video data, we propose to first train a style control adapter using style-rich image datasets, then transfer the learned stylization ability to video generation through a tailor-made finetuning paradigm. To promote content-style disentanglement, we employ carefully designed data augmentation strategies to enhance decoupled learning. Additionally, we propose a scale-adaptive fusion module to balance the influences of text-based content features and image-based style features, which helps generalization across various text and style combinations. StyleCrafter efficiently generates high-quality stylized videos that align with the content of the texts and resemble the style of the reference images. Experiments demonstrate that our approach is more flexible and efficient than existing competitors. Project page: https://gongyeliu.github.io/StyleCrafter.github.io/ Gongye Liu, Menghan Xia, Yong Zhang 0034, Haoxin Chen, Jinbo Xing, Yibo Wang 0039, Xintao Wang 0002, Ying Shan, Yujiu Yang 0001 |
ACM Trans. Graph. | 2 |
| 2024 | ToonCrafter: Generative Cartoon InterpolationabstractWe introduce ToonCrafter, a novel approach that transcends traditional correspondence-based cartoon video interpolation, paving the way for generative interpolation. Traditional methods, that implicitly assume linear motion and the absence of complicated phenomena like dis-occlusion, often struggle with the exaggerated non-linear and large motions with occlusion commonly found in cartoons, resulting in implausible or even failed interpolation results. To overcome these limitations, we explore the potential of adapting live-action video priors to better suit cartoon interpolation within a generative framework. ToonCrafter effectively addresses the challenges faced when applying live-action video motion priors to generative cartoon interpolation. First, we design a toon rectification learning strategy that seamlessly adapts live-action video priors to the cartoon domain, resolving the domain gap and content leakage issues. Next, we introduce a dual-reference-based 3D decoder to compensate for lost details due to the highly compressed latent prior spaces, ensuring the preservation of fine details in interpolation results. Finally, we design a flexible sketch encoder that empowers users with interactive control over the interpolation results. Experimental results demonstrate that our proposed method not only produces visually convincing and more natural dynamics, but also effectively handles dis-occlusion. The comparative evaluation demonstrates the notable superiority of our approach over existing competitors. Code and model weights are available at https://doubiiu.github.io/projects/ToonCrafter Jinbo Xing, Hanyuan Liu, Menghan Xia, Yong Zhang 0034, Xintao Wang 0002, Ying Shan, Tien-Tsin Wong |
ACM Trans. Graph. | 3 |
| 2024 | Taming Reversible Halftoning Via Predictive LuminanceabstractTraditional halftoning usually drops colors when dithering images with binary dots, which makes it difficult to recover the original color information. We proposed a novel halftoning technique that converts a color image into a binary halftone with full restorability to its original version. Our novel base halftoning technique consists of two convolutional neural networks (CNNs) to produce the reversible halftone patterns, and a noise incentive block (NIB) to mitigate the flatness degradation issue of CNNs. Furthermore, to tackle the conflicts between the blue-noise quality and restoration accuracy in our novel base method, we proposed a predictor-embedded approach to offload predictable information from the network, which in our case is the luminance information resembling from the halftone pattern. Such an approach allows the network to gain more flexibility to produce halftones with better blue-noise quality without compromising the restoration quality. Detailed studies on the multiple-stage training method and loss weightings have been conducted. We have compared our predictor-embedded method and our novel method regarding spectrum analysis on halftone, halftone accuracy, restoration accuracy, and the data embedding studies. Our entropy evaluation evidences our halftone contains less encoding information than our novel base method. The experiments show our predictor-embedded method gains more flexibility to improve the blue-noise quality of halftones and maintains a comparable restoration quality with a higher tolerance for disturbances. Cheuk-Kit Lau, Menghan Xia, Tien-Tsin Wong |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | LF2MV: Learning an Editable Meta-View Towards Light Field RepresentationabstractLight fields are 4D scene representations that are typically structured as arrays of views or several directional samples per pixel in a single view. However, this highly correlated structure is not very efficient to transmit and manipulate, especially for editing. To tackle this issue, we propose a novel representation learning framework that can encode the light field into a single meta-view that is both compact and editable. Specifically, the meta-view composes of three visual channels and a complementary meta channel that is embedded with geometric and residual appearance information. The visual channels can be edited using existing 2D image editing tools, before reconstructing the whole edited light field. To facilitate edit propagation against occlusion, we design a special editing-aware decoding network that consistently propagates the visual edits to the whole light field upon reconstruction. Extensive experiments show that our proposed method achieves competitive representation accuracy and meanwhile enables consistent edit propagation. Menghan Xia, Jose Echevarria, Minshan Xie, Tien-Tsin Wong |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2023 | CoordFill: Efficient High-Resolution Image Inpainting via Parameterized Coordinate QueryingabstractImage inpainting aims to fill the missing hole of the input. It is hard to solve this task efficiently when facing high-resolution images due to two reasons: (1) Large reception field needs to be handled for high-resolution image inpainting. (2) The general encoder and decoder network synthesizes many background pixels synchronously due to the form of the image matrix. In this paper, we try to break the above limitations for the first time thanks to the recent development of continuous implicit representation. In detail, we down-sample and encode the degraded image to produce the spatial-adaptive parameters for each spatial patch via an attentional Fast Fourier Convolution (FFC)-based parameter generation network. Then, we take these parameters as the weights and biases of a series of multi-layer perceptron (MLP), where the input is the encoded continuous coordinates and the output is the synthesized color value. Thanks to the proposed structure, we only encode the high-resolution image in a relatively low resolution for larger reception field capturing. Then, the continuous position encoding will be helpful to synthesize the photo-realistic high-frequency textures by re-sampling the coordinate in a higher resolution. Also, our framework enables us to query the coordinates of missing pixels only in parallel, yielding a more efficient solution than the previous methods. Experiments show that the proposed method achieves real-time performance on the 2048X2048 images using a single GTX 2080 Ti GPU and can handle 4096X4096 images, with much better performance than existing state-of-the-art methods visually and numerically. The code is available at: https://github.com/NiFangBaAGe/CoordFill. Weihuang Liu, Xiaodong Cun, Chi-Man Pun, Menghan Xia, Yong Zhang 0034, Jue Wang 0001 |
AAAI | 4 |
| 2023 | CodeTalker: Speech-Driven 3D Facial Animation with Discrete Motion PriorabstractSpeech-driven 3D facial animation has been widely studied, yet there is still a gap to achieving realism and vividness due to the highly ill-posed nature and scarcity of audio-visual data. Existing works typically formulate the cross-modal mapping into a regression task, which suffers from the regression-to-mean problem leading to over-smoothed facial motions. In this paper, we propose to cast speech-driven facial animation as a code query task in a finite proxy space of the learned codebook, which effectively promotes the vividness of the generated motions by reducing the cross-modal mapping uncertainty. The codebook is learned by self-reconstruction over real facial motions and thus embedded with realistic facial motion priors. Over the discrete motion space, a temporal autoregressive model is employed to sequentially synthesize facial motions from the input speech signal, which guarantees lip-sync as well as plausible facial expressions. We demonstrate that our approach outperforms current state-of-the-art methods both qualitatively and quantitatively. Also, a user study further justifies our superiority in perceptual quality. Code and video demo are available at https://doubiiu.github.io/projects/codetalker. Jinbo Xing, Menghan Xia, Yuechen Zhang, Xiaodong Cun, Jue Wang 0001, Tien-Tsin Wong |
CVPR | 2 |
| 2023 | Interactive Story Visualization with Multiple CharactersabstractAccurate Story visualization requires several necessary elements, such as identity consistency across frames, the alignment between plain text and visual content, and a reasonable layout of objects in images. Most previous works endeavor to meet these requirements by fitting a text-to-image (T2I) model on a set of videos in the same style and with the same characters, e.g., the FlintstonesSV dataset. However, the learned T2I models typically struggle to adapt to new characters, scenes, and styles, and often lack the flexibility to revise the layout of the synthesized images. This paper proposes a system for generic interactive story visualization, capable of handling multiple novel characters and supporting the editing of layout and local structure. It is developed by leveraging the prior knowledge of large language and T2I models, trained on massive corpora. The system comprises four interconnected components: story-to-prompt generation (S2P), text-to-layout generation (T2L), controllable text-to-image generation (C-T2I), and image-to-video animation (I2V). First, the S2P module converts concise story information into detailed prompts required for subsequent stages. Next, T2L generates diverse and reasonable layouts based on the prompts, offering users the ability to adjust and refine the layout to their preferences. The core component, C-T2I, enables the creation of images guided by layouts, sketches, and actor-specific identifiers to maintain consistency and detail across visualizations. Finally, I2V enriches the visualization process by animating the generated images. Extensive experiments and a user study are conducted to validate the effectiveness and flexibility of interactive editing of the proposed system. Yuan Gong 0002, Youxin Pang, Xiaodong Cun, Menghan Xia, Yingqing He, Haoxin Chen, Longyue Wang, Yong Zhang 0034, Xintao Wang 0002, Ying Shan, Yujiu Yang 0001 |
SIGGRAPH Asia | 4 |
| 2023 | Scale-Arbitrary Invertible Image DownscalingabstractConventional social media platforms usually downscale high-resolution (HR) images to restrict their resolution to a specific size for saving transmission/storage cost, which makes those visual details inaccessible to other users. To bypass this obstacle, recent invertible image downscaling methods jointly model the downscaling/upscaling problems and achieve impressive performance. However, they only consider fixed integer scale factors and may be inapplicable to generic downscaling tasks towards resolution restriction as posed by social media platforms. In this paper, we propose an effective and universal Scale-Arbitrary Invertible Image Downscaling Network (AIDN), to downscale HR images with arbitrary scale factors in an invertible manner. Particularly, the HR information is embedded in the downscaled low-resolution (LR) counterparts in a nearly imperceptible form such that our AIDN can further restore the original HR images solely from the LR images. The key to supporting arbitrary scale factors is our proposed Conditional Resampling Module (CRM) that conditions the downscaling/upscaling kernels and sampling locations on both scale factors and image content. Extensive experimental results demonstrate that our AIDN achieves top performance for invertible downscaling with both arbitrary integer and non-integer scale factors. Also, both quantitative and qualitative evaluations show our AIDN is robust to the lossy image compression standard. The source code and trained models are publicly available at https://github.com/Doubiiu/AIDN. Jinbo Xing, Wenbo Hu 0002, Menghan Xia, Tien-Tsin Wong |
IEEE Trans. Image Process. | 3 |
| 2022 | PalGAN: Image Colorization with Palette Generative Adversarial Networks
Yi Wang 0074, Menghan Xia, Lu Qi 0001, Yu Qiao 0001 |
ECCV (15) | 2 |
| 2022 | VideoReTalking: Audio-based Lip Synchronization for Talking Head Video Editing In the WildabstractWe present VideoReTalking, a new system to edit the faces of a real-world talking head video according to input audio, producing a high-quality and lip-syncing output video even with a different emotion. Our system disentangles this objective into three sequential tasks: (1) face video generation with a canonical expression; (2) audio-driven lip-sync; and (3) face enhancement for improving photo-realism. Given a talking-head video, we first modify the expression of each frame according to the same expression template using the expression editing network, resulting in a video with the canonical expression. This video, together with the given audio, is then fed into the lip-sync network to generate a lip-syncing video. Finally, we improve the photo-realism of the synthesized faces through an identity-aware face enhancement network and post-processing. We use learning-based approaches for all three steps and all our modules can be tackled in a sequential pipeline without any user intervention. Furthermore, our system is a generic approach that does not need to be retrained to a specific person. Evaluations on two widely-used datasets and in-the-wild examples demonstrate the superiority of our framework over other state-of-the-art methods in terms of lip-sync accuracy and visual quality. Xiaodong Cun, Yong Zhang 0034, Menghan Xia, Mingrui Zhu, Xuan Wang 0009, Jue Wang 0001, Nannan Wang 0001 |
SIGGRAPH Asia | 4 |
| 2022 | Disentangled Image Colorization via Global AnchorsabstractColorization is multimodal by nature and challenges existing frameworks to achieve colorful and structurally consistent results. Even the sophisticated autoregressive model struggles to maintain long-distance color consistency due to the fragility of sequential dependence. To overcome this challenge, we propose a novel colorization framework that disentangles color multimodality and structure consistency through global color anchors, so that both aspects could be learned effectively. Our key insight is that several carefully located anchors could approximately represent the color distribution of an image, and conditioned on the anchor colors, we can predict the image color in a deterministic manner by utilizing internal correlation. To this end, we construct a colorization model with dual branches, where the color modeler predicts the color distribution for anchor color representation, and the color generator predicts the pixel colors by referring the sampled anchor colors. Importantly, the anchors are located under two principles: color independence and global coverage, which is realized with clustering analysis on the deep color features. To simplify the computation, we creatively adopt soft superpixel segmentation to reduce the image primitives, which still nicely reserves the reversibility to pixel-wise representation. Extensive experiments show that our method achieves notable superiority over various mainstream frameworks in perceptual quality. Thanks to anchor-based color representation, our model has the flexibility to support diverse and controllable colorization as well. Menghan Xia, Wenbo Hu 0005, Tien-Tsin Wong, Jue Wang 0001 |
ACM Trans. Graph. | 1 |
| 2021 | Exploiting Aliasing for Manga RestorationabstractAs a popular entertainment art form, manga enriches the line drawings details with bitonal screentones. However, manga resources over the Internet usually show screen-tone artifacts because of inappropriate scanning/rescaling resolution. In this paper, we propose an innovative two-stage method to restore quality bitonal manga from de-graded ones. Our key observation is that the aliasing induced by downsampling bitonal screentones can be utilized as informative clues to infer the original resolution and screentones. First, we predict the target resolution from the degraded manga via the Scale Estimation Network (SE-Net) with spatial voting scheme. Then, at the target resolution, we restore the region-wise bitonal screentones via the Manga Restoration Network (MR-Net) discriminatively, depending on the degradation degree. Specifically, the original screentones are directly restored in pattern-identifiable regions, and visually plausible screentones are synthesized in pattern-agnostic regions. Quantitative evaluation on synthetic data and visual assessment on real-world cases illustrate the effectiveness of our method. Minshan Xie, Menghan Xia, Tien-Tsin Wong |
CVPR | 2 |
| 2021 | Deep Halftoning with Reversible Binary PatternabstractExisting halftoning algorithms usually drop colors and fine details when dithering color images with binary dot patterns, which makes it extremely difficult to recover the original information. To dispense the recovery trouble in future, we propose a novel halftoning technique that converts a color image into binary halftone with full restorability to the original version. The key idea is to implicitly embed those previously dropped information into the halftone patterns. So, the halftone pattern not only serves to reproduce the image tone, maintain the blue-noise randomness, but also represents the color information and fine details. To this end, we exploit two collaborative convolutional neural networks (CNNs) to learn the dithering scheme, under a nontrivial self-supervision formulation. To tackle the flatness degradation issue of CNNs, we propose a novel noise incentive block (NIB) that can serve as a generic CNN plug-in for performance promotion. At last, we tailor a guiding-aware training scheme that secures the convergence direction as regulated. We evaluate the invertible halftones in multiple aspects, which evidences the effectiveness of our method. Menghan Xia, Wenbo Hu 0002, Xueting Liu 0001, Tien-Tsin Wong |
ICCV | 1 |
| 2021 | Grid Model-Based Global Color Correction for Multiple Image MosaickingabstractColor consistency optimization for multiple images is a challenging problem in image mosaicking. To facilitate the global color optimization, existing approaches mainly use less flexible models, e.g., linear or gamma function, to eliminate the color differences between multiple images. However, these models often struggle to eliminate the color differences that existed in the local areas and preserve the image gradient information. To solve this problem, we creatively propose a novel color-correction model, which comprised a series of local grid linear models. This model is simple, but it is flexible enough to approximate a variety of complicated local color variations. To obtain the optimal model parameters for each image globally, a specific cost function that considers both color consistency and gradient preservation is designed and solved. The aim of our approach is to generate a composite image with visually consistent color. The original color information may be destroyed. Thus, this approach is unsuitable for the quantitative remote sensing applications. The experimental results on several challenging data sets show that the proposed approach outperforms state-of-the-art approaches in both visual quality and quantitative metrics. Li Li 0047, Yunmeng Li, Menghan Xia, Yinxuan Li, Jian Yao 0002, Bin Wang 0100 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2021 | Seamless manga inpainting with semantics awarenessabstractManga inpainting fills up the disoccluded pixels due to the removal of dialogue balloons or "sound effect" text. This process is long needed by the industry for the language localization and the conversion to animated manga. It is mostly done manually, as existing methods (mostly for natural image inpainting) cannot produce satisfying results. Manga inpainting is more tricky than natural image inpainting because its highly abstract illustration using structural lines and screentone patterns, which confuses the semantic interpretation and visual content synthesis. In this paper, we present the first manga inpainting method, a deep learning model, that generates high-quality results. Instead of direct inpainting, we propose to separate the complicated inpainting into two major phases, semantic inpainting and appearance synthesis. This separation eases both the feature understanding and hence the training of the learning model. A key idea is to disentangle the structural line and screentone, that helps the network to better distinguish the structural line and the screentone features for semantic interpretation. Both the visual comparison and the quantitative experiments evidence the effectiveness of our method and justify its superiority over existing state-of-the-art methods in the application of manga inpainting. Minshan Xie, Menghan Xia, Xueting Liu 0001, Chengze Li, Tien-Tsin Wong |
ACM Trans. Graph. | 2 |
| 2020 | Mononizing binocular videosabstractThis paper presents the idea of mono-nizing binocular videos and a framework to effectively realize it. Mono-nize means we purposely convert a binocular video into a regular monocular video with the stereo information implicitly encoded in a visual but nearly-imperceptible form. Hence, we can impartially distribute and show the mononized video as an ordinary monocular video. Unlike ordinary monocular videos, we can restore from it the original binocular video and show it on a stereoscopic display. To start, we formulate an encoding-and-decoding framework with the pyramidal deformable fusion module to exploit long-range correspondences between the left and right views, a quantization layer to suppress the restoring artifacts, and the compression noise simulation module to resist the compression noise introduced by modern video codecs. Our framework is self-supervised, as we articulate our objective function with loss terms defined on the input: a monocular term for creating the mononized video, an invertibility term for restoring the original video, and a temporal term for frame-to-frame coherence. Further, we conducted extensive experiments to evaluate our generated mononized videos and restored binocular videos for diverse types of images and 3D movies. Quantitative results on both standard metrics and user perception studies show the effectiveness of our method. Wenbo Hu 0002, Menghan Xia, Chi-Wing Fu, Tien-Tsin Wong |
ACM Trans. Graph. | 2 |
| 2019 | RoadNet: Learning to Comprehensively Analyze Road Networks in Complex Urban Scenes From High-Resolution Remotely Sensed ImagesabstractIt is a classical task to automatically extract road networks from very high-resolution (VHR) images in remote sensing. This paper presents a novel method for extracting road networks from VHR remotely sensed images in complex urban scenes. Inspired by image segmentation, edge detection, and object skeleton extraction, we develop a multitask convolutional neural network (CNN), called RoadNet, to simultaneously predict road surfaces, edges, and centerlines, which is the first work in such field. The RoadNet solves seven important issues in this vision problem: 1) automatically learning multiscale and multilevel features [gained by the deeply supervised nets (DSN) providing integrated direct supervision] to cope with the roads in various scenes and scales; 2) holistically training the mentioned tasks in a cascaded end-to-end CNN model; 3) correlating the predictions of road surfaces, edges, and centerlines in a network model to improve the multitask prediction; 4) designing elaborate architecture and loss function, by which the well-trained model produces approximately single-pixel width road edges/centerlines without nonmaximum suppression postprocessing; 5) cropping and bilinear blending to deal with the large VHR images with finite-computing resources; 6) introducing rough and simple user interaction to obtain desired predictions in the challenging regions; and 7) establishing a benchmark data set which consists of a series of VHR remote sensing images with pixelwise annotation. Different from the previous works, we pay more attention to the challenging situations, in which there are lots of shadows and occlusions along the road regions. Experimental results on two benchmark data sets show the superiority of our proposed approaches. Jian Yao 0002, Xiaohu Lu, Menghan Xia, Xingbo Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2019 | Colorblind-shareable videos by synthesizing temporal-coherent polynomial coefficientsabstractTo share the same visual content between color vision deficiencies (CVD) and normal-vision people, attempts have been made to allocate the two visual experiences of a binocular display (wearing and not wearing glasses) to CVD and normal-vision audiences. However, existing approaches only work for still images. Although state-of-the-art temporal filtering techniques can be applied to smooth the per-frame generated content, they may fail to maintain the multiple binocular constraints needed in our applications, and even worse, sometimes introduce color inconsistency (same color regions map to different colors). In this paper, we propose to train a neural network to predict the temporal coherent polynomial coefficients in the domain of global color decomposition. This indirect formulation solves the color inconsistency problem. Our key challenge is to design a neural network to predict the temporal coherent coefficients, while maintaining all required binocular constraints. Our method is evaluated on various videos and all metrics confirm that it outperforms all existing solutions. Xinghong Hu, Xueting Liu 0001, Zhuming Zhang, Menghan Xia, Chengze Li, Tien-Tsin Wong |
ACM Trans. Graph. | 4 |
| 2018 | Deep Inverse Halftoning via Progressively Residual Learning
Menghan Xia, Tien-Tsin Wong |
ACCV (6) | 1 |
| 2018 | Invertible grayscaleabstractOnce a color image is converted to grayscale, it is a common belief that the original color cannot be fully restored, even with the state-of-the-art colorization methods. In this paper, we propose an innovative method to synthesize invertible grayscale. It is a grayscale image that can fully restore its original color. The key idea here is to encode the original color information into the synthesized grayscale, in a way that users cannot recognize any anomalies. We propose to learn and embed the color-encoding scheme via a convolutional neural network (CNN). It consists of an encoding network to convert a color image to grayscale, and a decoding network to invert the grayscale to color. We then design a loss function to ensure the trained network possesses three required properties: (a) color invertibility, (b) grayscale conformity, and (c) resistance to quantization error. We have conducted intensive quantitative experiments and user studies over a large amount of color images to validate the proposed method. Regardless of the genre and content of the color input, convincing results are obtained in all cases. Menghan Xia, Xueting Liu 0001, Tien-Tsin Wong |
ACM Trans. Graph. | 1 |
| 2017 | Optimal seamline detection in dynamic scenes via graph cuts for image mosaicking
Li Li 0047, Jian Yao 0002, Haoang Li, Menghan Xia, Wei Zhang 0021 |
Mach. Vis. Appl. | 4 |
| 2017 | Globally consistent alignment for planar mosaicking via topology analysis
Menghan Xia, Jian Yao 0002, Renping Xie, Li Li 0047, Wei Zhang 0021 |
Pattern Recognit. | 1 |
| 2016 | Joint point and line segment matching on wide-baseline stereo imagesabstractThis paper presents an method that matches points and line segments jointly on wide-baseline stereo images. In both two images to be matched, line segments are extracted and those spatially adjacent ones are intersected to generate V-junctions. To match V-junctions from the two images, we extract for each of them an affine and scale invariant local region and describe it with SIFT. The putative V-junction matches obtained from evaluating their description vectors are refined subsequently by the epipolar line constraint and topological distribution constraint among neighbor V-junctions. Since once a pair of V-junctions are matched, the two pairs of line segments forming them are matched accordingly. A part of line segments from the two images are therefore matched along with V-junction matches. To get more line segment matches, we further match those left unmatched line segments by the local homographies estimated from their adjacent V-junction matches. Experiments verify the robustness of the proposed method and its superiority to both some famous point and line segment matching methods on wide-baseline stereo images. In addition, we also show the proposed method can make it easier for 3D line segment reconstruction. Kai Li 0015, Jian Yao 0002, Menghan Xia, Li Li 0047 |
WACV | 3 |
| 2015 | Line-based Multi-Label Energy Optimization for fisheye image rectification and calibrationabstractFisheye image rectification and estimation of intrinsic parameters for real scenes have been addressed in the literature by using line information on the distorted images. In this paper, we propose an easily implemented fisheye image rectification algorithm with line constrains in the undistorted perspective image plane. A novel Multi-Label Energy Optimization (MLEO) method is adopted to merge short circular arcs sharing the same or the approximately same circular parameters and select long circular arcs for camera rectification. Further we propose an efficient method to estimate intrinsic parameters of the fisheye camera by automatically selecting three properly arranged long circular arcs from previously obtained circular arcs in the calibration procedure. Experimental results on a number of real images and simulated data show that the proposed method can achieve good results and outperforms the existing approaches and the commercial software in most cases. Mi Zhang 0004, Jian Yao 0002, Menghan Xia, Kai Li 0015 |
CVPR | 3 |
| 2015 | Globally consistent alignment for mosaicking aerial imagesabstractIn this paper, we present a robust method to efficiently create a globally consistent and seamless mosaic from aerial images. Firstly, a globally consistent registration strategy is proposed to align the aerial images in a common coordinate system, which combines the affine model with the homographic model effectively. To suppress the accumulation of perspective distortions induced by a sequential set of aerial images taken from a wide-range region, we proposed to initially align each image by an affine model and then perform a homographic refinement in groups to increase the global consistency. Secondly, to efficiently conceal the parallax between aligned images in overlap regions with large depth differences where it is impossible to recover a highly accurate consistent image registration, a novel optimized seamline detection algorithm in the graph cuts energy minimization framework is proposed to find optimal seamlines within overlap regions for image mosaicking through rounding visually obvious foreground objects. Finally, experimental results on several representative image sets illustrate the superiority of our proposed approaches. Menghan Xia, Man Yao, Li Li 0047, Xiaohu Lu |
ICIP | 1 |