VLDB 2026 Research / reviewers in the wild / expert
Shuai Yang 0001
dblp:72/7503-1
· DBLP profile ↗
79ranked-venue papers
26as first author
51since 2021 · last 2026
0000-0002-5576-8629ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 59 · 20 first-author · 34 since 2021Artificial intelligence and machine learning · 48 · 16 first-author · 39 since 2021Systems, architecture and hardware · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action ModelingabstractRecent vision-language-action (VLA) models built on pretrained vision-language models (VLMs) have demonstrated strong performance in robotic manipulation. However, these models remain constrained by the single-frame image paradigm and fail to fully leverage the temporal information offered by multi-frame histories, as directly feeding multiple frames into VLM backbones incurs substantial computational overhead and inference latency. We propose CronusVLA, a unified framework that extends single-frame VLA models to the multi-frame paradigm. CronusVLA follows a two-stage process: (1) Single-frame pretraining on large-scale embodied datasets with autoregressive prediction of action tokens, establishing an effective embodied vision-language foundation; (2) Multi-frame post-training, which adapts the prediction of the vision-language backbone from discrete tokens to learnable features, and aggregates historical information via feature chunking. CronusVLA effectively addresses the existing challenges of multi-frame modeling while enhancing performance. To evaluate the robustness under temporal and spatial disturbances, we introduce SimplerEnv-OR, a novel benchmark featuring 24 types of observational disturbances and 120 severity levels. Experiments across three embodiments in simulated and real-world environments demonstrate that CronusVLA achieves leading performance and superior robustness, with a 70.9% success rate on SimplerEnv, a 26.8% improvement over OpenVLA on LIBERO, and the highest robustness score on SimplerEnv-OR, showing the promise of efficient multi-frame adaptation for real-world VLA deployment. Hao Li 0069, Shuai Yang 0001, Xiaoda Yang, Dahua Lin, Feng Zhao 0004, Jiangmiao Pang |
AAAI | 2 |
| 2026 | Self-Supervised Skeleton-Based Action Representation Learning: A Benchmark and Beyond
Jiahang Zhang 0001, Lilang Lin, Shuai Yang 0001, Jiaying Liu 0001 |
Int. J. Comput. Vis. | 3 |
| 2026 | LN3Diff++: Scalable Latent Neural Fields Diffusion for Speedy 3D GenerationabstractThe field of neural rendering has seen remarkable progress, driven by advancements in generative models and differentiable rendering techniques. While 2D diffusion has achieved notable success, the development of a unified 3D diffusion pipeline remains an open challenge. This paper presents a novel framework, LN3Diff++, designed to bridge this gap and facilitate fast, high-quality, and versatile conditional 3D generation. Our method leverages a 3D-aware architecture and a variational autoencoder (VAE) to encode input image(s) into a structured, compact 3D latent space. The latent representation is then decoded by a transformer-based decoder into a high-capacity 3D neural field. By training a diffusion model on this 3D-aware latent space, our method achieves superior performance for category-specific 3D generation on ShapeNet and FFHQ, as well as category-free image/text-conditioned 3D generation over Objaverse. Moreover, it surpasses existing 3D diffusion methods in inference speed, requiring no per-instance optimization. Yushi Lan, Fangzhou Hong, Shangchen Zhou, Shuai Yang 0001, Xuyi Meng, Yongwei Chen, Zhaoyang Lyu, Bo Dai 0002, Xingang Pan, Chen Change Loy |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Alias-Free Latent Diffusion Models: Improving Fractional Shift Equivariance of Diffusion Latent SpaceabstractLatent Diffusion Models (LDMs) are known to have an unstable generation process, where even small perturbations or shifts in the input noise can lead to significantly different outputs. This hinders their applicability in applications requiring consistent results. In this work, we redesign LDMs to enhance consistency by making them shift-equivariant. While introducing anti-aliasing operations can partially improve shift-equivariance, significant aliasing and inconsistency persist due to the unique challenges in LDMs, including 1) aliasing amplification during VAE training and multiple U-Net inferences, and 2) selfattention modules that inherently lack shift-equivariance. To address these issues, we redesign the attention modules to be shift-equivariant and propose an equivariance loss that effectively suppresses the frequency bandwidth of the features in the continuous domain. The resulting alias-free LDM (AF-LDM) achieves strong shift-equivariance and is also robust to irregular warping. Extensive experiments demonstrate that AF-LDM produces significantly more consistent results than vanilla LDM across various applications, including video editing and image-to-image translation. Code is available at: https://github.com/SingleZombie/AFLDM Yifan Zhou 0001, Zeqi Xiao, Shuai Yang 0001, Xingang Pan |
CVPR | 3 |
| 2025 | MEAT: Multiview Diffusion Model for Human Generation on Megapixels with Mesh AttentionabstractMultiview diffusion models have shown considerable success in image-to-3D generation for general objects. However, when applied to human data, existing methods have yet to deliver promising results, largely due to the challenges of scaling multiview attention to higher resolutions. In this paper, we explore human multiview diffusion models at the megapixel level and introduce a solution called mesh attention to enable training at 10242resolution. Using a clothed human mesh as a central coarse geometric representation, the proposed mesh attention leverages rasterization and projection to establish direct cross-view coordinate correspondences. This approach significantly reduces the complexity of multiview attention while maintaining cross-view consistency. Building on this foundation, we devise a mesh attention block and combine it with keypoint conditioning to create our human-specific multiview diffusion model, MEAT. In addition, we present valuable insights into applying multiview human motion videos for diffusion training, addressing the longstanding issue of data scarcity. Extensive experiments show that MEAT effectively generates dense, consistent multiview human images at the megapixel level, outperforming existing multiview diffusion methods. Code is available at https://johann.wang/MEAT/. Yuhan Wang 0002, Fangzhou Hong, Shuai Yang 0001, Liming Jiang 0001, Wayne Wu, Chen Change Loy |
CVPR | 3 |
| 2025 | GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following ManipulationabstractRobotic manipulation in real-world settings remains challenging, especially regarding robust generalization. Existing simulation platforms lack sufficient support for exploring how policies adapt to varied instructions and scenarios. Thus, they lag behind the growing interest in instruction-following foundation models like LLMs, whose adaptability is crucial yet remains underexplored in fair comparisons. To bridge this gap, we introduce GenManip, a realistic tabletop simulation platform tailored for policy generalization studies. It features an automatic pipeline via LLM-driven task-oriented scene graph to synthesize large-scale, diverse tasks using 10K annotated 3D object assets. To systematically assess generalization, we present GenManip-Bench, a benchmark of 200 scenarios refined via human-in-the-loop corrections. We evaluate two policy types: (1) modular manipulation systems integrating foundation models for perception, reasoning, and planning, and (2) end-to-end policies trained through scalable data collection. Results show that while data scaling benefits end-to-end methods, modular systems enhanced with foundation models generalize more effectively across diverse scenarios. We anticipate this platform to facilitate critical insights for advancing policy generalization in realistic conditions. All code will be available at project page. Shuai Yang 0001, Hao Li 0009, Haifeng Huang 0001, Jiangmiao Pang |
CVPR | 3 |
| 2025 | PTDiffusion: Free Lunch for Generating Optical Illusion Hidden Pictures with Phase-Transferred Diffusion ModelabstractOptical illusion hidden picture is an interesting visual perceptual phenomenon where an image is cleverly integrated into another picture. Established on the off-the-shelf text-to-image (T2I) diffusion model, we propose a novel text-guided image-to-image (I2I) translation framework dubbed as Phase-Transferred Diffusion Model (PTDiffusion) for hidden art syntheses, which harmoniously embeds an input reference image into arbitrary scenes described by the text prompts. At the heart of our method is a plug-and-play phase transfer mechanism that dynamically and progressively transplants diffusion features’ phase spectrum from the denoising process to reconstruct the reference image into the one to sample the generated illusion image, realizing deep fusion of the reference structural information and the textual semantic information. Furthermore, we propose asynchronous phase transfer to enable flexible control over the degree of hidden content discernability. Our method bypasses any model training and fine-tuning process, all while substantially outperforming related methods in image quality, text fidelity, visual discernibility, and contextual naturalness for illusion picture synthesis, as demonstrated by extensive qualitative and quantitative experiments. Our project is publically available at this web page. Xiang Gao 0014, Shuai Yang 0001, Jiaying Liu 0001 |
CVPR | 2 |
| 2025 | AnyPortal: Zero-Shot Consistent Video Background ReplacementabstractDespite the rapid advancements in video generation technology, creating high-quality videos that precisely align with user intentions remains a significant challenge. Existing methods often fail to achieve fine-grained control over video details, limiting their practical applicability. We introduce ANYPORTAL, a novel zero-shot framework for video background replacement that leverages pre-trained diffusion models. Our framework collaboratively integrates the temporal prior of video diffusion models with the relighting capabilities of image diffusion models in a zero-shot setting. To address the critical challenge of foreground consistency, we propose a Refinement Projection Algorithm, which enables pixel-level detail manipulation to ensure precise foreground preservation. ANYPORTAL is training-free and overcomes the challenges of achieving foreground consistency and temporally coherent relighting. Experimental results demonstrate that ANYPORTAL achieves high-quality results on consumer-grade GPUs, offering a practical and efficient solution for video content creation and editing. Wenshuo Gao, Xicheng Lan, Shuai Yang 0001 |
ICCV | 3 |
| 2025 | Balanced Image Stylization with Style Matching Score
Liming Jiang 0001, Shuai Yang 0001, Jia-Wei Liu, Ivor W. Tsang, Zheng Shou 0001 |
ICCV | 3 |
| 2025 | TokensGen: Harnessing Condensed Tokens for Long Video GenerationabstractGenerating consistent long videos is a complex challenge: while diffusion-based generative models generate visually impressive short clips, extending them to longer durations often leads to memory bottlenecks and long-term inconsistency. In this paper, we propose TokensGen, a novel two-stage framework that leverages condensed tokens to address these issues. Our method decomposes long video generation into three core tasks: (1) inner-clip semantic control, (2) long-term consistency control, and (3) inter-clip smooth transition. First, we train To2V (Token-to-Video), a short video diffusion model guided by text and video tokens, with a Video Tokenizer that condenses short clips into semantically rich tokens. Second, we introduce T2To (Text-to-Token), a video token diffusion transformer that generates all tokens at once, ensuring global consistency across clips. Finally, during inference, an adaptive FIFO-Diffusion strategy seamlessly connects adjacent clips, reducing boundary artifacts and enhancing smooth transitions. Experimental results demonstrate that our approach significantly enhances long-term temporal and content coherence without incurring prohibitive computational overhead. By leveraging condensed tokens and pre-trained short video models, our method provides a scalable, modular solution for long video generation, opening new possibilities for storytelling, cinematic production, and immersive simulations. Please see our project page at https://vicky0522.github.io/tokensgen-webpage/ . Wenqi Ouyang, Zeqi Xiao, Danni Yang, Yifan Zhou 0001, Shuai Yang 0001, Lei Yang 0045, Jianlou Si, Xingang Pan |
ICCV | 5 |
| 2025 | GaussianAnything: Interactive Point Cloud Flow Matching for 3D GenerationabstractRecent advancements in diffusion models and large-scale datasets have revolutionized image and video generation, with increasing focus on 3D content generation. While existing methods show promise, they face challenges in input formats, latent space structures, and output representations. This paper introduces a novel 3D generation framework that addresses these issues, enabling scalable and high-quality 3D generation with an interactive Point Cloud-structured Latent space. Our approach utilizes a VAE with multi-view posed RGB-D-N renderings as input, features a unique latent space design that preserves 3D shape information, and incorporates a cascaded latent flow-based model for improved shape-texture disentanglement. The proposed method, GaussianAnything, supports multi-modal conditional 3D generation, allowing for point cloud, caption, and single-view image inputs. Experimental results demonstrate superior performance on various datasets, advancing the state-of-the-art in 3D content generation. Yushi Lan, Shangchen Zhou, Zhaoyang Lyu, Fangzhou Hong, Shuai Yang 0001, Bo Dai 0002, Xingang Pan, Chen Change Loy |
ICLR | 5 |
| 2025 | Trajectory attention for fine-grained video motion controlabstractRecent advancements in video generation have been greatly driven by video diffusion models, with camera motion control emerging as a crucial challenge in creating view-customized visual content. This paper introduces trajectory attention, a novel approach that performs attention along available pixel trajectories for fine-grained camera motion control. Unlike existing methods that often yield imprecise outputs or neglect temporal correlations, our approach possesses a stronger inductive bias that seamlessly injects trajectory information into the video generation process. Importantly, our approach models trajectory attention as an auxiliary branch alongside traditional temporal attention. This design enables the original temporal attention and the trajectory attention to work in synergy, ensuring both
precise motion control and new content generation capability, which is critical when the trajectory is only partially available. Experiments on camera motion control for images and videos demonstrate significant improvements in precision and long-range consistency while maintaining high-quality generation. Furthermore, we show that our approach can be extended to other video motion control tasks, such as first-frame-guided video editing, where it excels in maintaining content consistency over large spatial and temporal ranges. Zeqi Xiao, Wenqi Ouyang, Yifan Zhou 0001, Shuai Yang 0001, Lei Yang 0045, Jianlou Si, Xingang Pan |
ICLR | 4 |
| 2025 | Splice, Focus and Relife: High-Resolution Periodic Pattern GenerationabstractThe printing and dyeing industry requires periodic and high-resolution patterns to ensure seamless designs on large fabric sections and high-quality final products. However, current manual approaches to pattern creation are time-consuming and labor-intensive. Leveraging powerful image generative models, such as Latent Diffusion Models (LDMs), offers a promising alternative, but challenges persist in generating strictly periodic and high-resolution patterns due to the inherent randomness and high computational demands of LDMs. In this paper, we propose a novel text-driven framework for generating periodic and high-resolution patterns. We introduce a new training-free Splice-and-Focus Mechanism, which enhances the model by constraining latent features and modifying the attention mechanism to produce natural and strictly periodic patterns. Additionally, we present a ReLife Pipeline, which integrates super-resolution and guided image synthesis to enhance pattern resolution while eliminating artifacts and distortions. Experimental results demonstrate that our framework produces patterns of superior quality. Xicheng Lan, Wenshuo Gao, Luyao Zhang 0007, Jiaying Liu 0001, Shuai Yang 0001 |
ISCAS | 5 |
| 2025 | Imagine360: Immersive 360 Video Generation from Perspective Anchorabstract$360^\circ$ videos offer a hyper-immersive experience that allows the viewers to explore a dynamic scene from full 360 degrees.
To achieve more accessible and personalized content creation in $360^\circ$ video format, we seek to lift standard perspective videos into $360^\circ$ equirectangular videos. To this end, we introduce **Imagine360**, the first perspective-to-$360^\circ$ video generation framework that creates high-quality $360^\circ$ videos with rich and diverse motion patterns from video anchors.
Imagine360 learns fine-grained spherical visual and motion patterns from limited $360^\circ$ video data with several key designs.
**1)** Firstly we adopt the dual-branch design, including a perspective and a panorama video denoising branch to provide local and global constraints for $360^\circ$ video generation, with motion module and spatial LoRA layers fine-tuned on $360^\circ$ videos.
**2)** Additionally, an antipodal mask is devised to capture long-range motion dependencies, enhancing the reversed camera motion between antipodal pixels across hemispheres.
**3)** To handle diverse perspective video inputs, we propose rotation-aware designs that adapt to varying video masking due to changing camera poses across frames.
**4)** Lastly, we introduce a new 360 video dataset featuring 10K high-quality, trimmed 360 video clips with structured motion to facilitate training.
Extensive experiments show Imagine360 achieves superior graphics quality and motion coherence with our curated dataset among state-of-the-art $360^\circ$ video generation methods. We believe Imagine360 holds promise for advancing personalized, immersive $360^\circ$ video creation. Jing Tan 0002, Shuai Yang 0001, Jingwen He, Yuwei Guo 0002, Ziwei Liu 0002, Dahua Lin |
NeurIPS | 2 |
| 2025 | Video World Models with Long-term Spatial MemoryabstractEmerging world models autoregressively generate video frames in response to actions, such as camera movements and text prompts, among other control signals. Due to limited temporal context window sizes, these models often struggle to maintain scene consistency during revisits, leading to severe forgetting of previously generated environments. Inspired by the mechanisms of human memory, we introduce a novel framework to enhancing long-term consistency of video world models through a geometry-grounded long-term spatial memory. Our framework includes mechanisms to store and retrieve information from the long-term spatial memory and we curate custom datasets to train and evaluate world models with explicitly stored 3D memory mechanisms. Our evaluations show improved quality, consistency, and context length compared to relevant baselines, paving the way towards long-term consistent world generation. Shuai Yang 0001, Ryan Po, Yinghao Xu 0001, Ziwei Liu 0002, Dahua Lin, Gordon Wetzstein |
NeurIPS | 2 |
| 2025 | WorldMem: Long-term Consistent World Simulation with MemoryabstractWorld simulation has gained increasing popularity due to its ability to model virtual environments and predict the consequences of actions. However, the limited temporal context window often leads to failures in maintaining long-term consistency, particularly in preserving 3D spatial consistency. In this work, we present WorldMem, a framework that enhances scene generation with a memory bank consisting of memory units that store memory frames and states (e.g., poses and timestamps). By employing state-aware memory attention that effectively extracts relevant information from these memory frames based on their states, our method is capable of accurately reconstructing previously observed scenes, even under significant viewpoint or temporal gaps. Furthermore, by incorporating timestamps into the states, our framework not only models a static world but also captures its dynamic evolution over time, enabling both perception and interaction within the simulated world. Extensive experiments in both virtual and real scenarios validate the effectiveness of our approach. Zeqi Xiao, Yushi Lan, Yifan Zhou 0001, Wenqi Ouyang, Shuai Yang 0001, Yanhong Zeng, Xingang Pan |
NeurIPS | 5 |
| 2025 | SGAR: Structural Generative Augmentation for 3D Human Motion Retrievalabstract3D human motion-text retrieval is essential for accurate motion understanding, targeted at cross-modal alignment learning. Existing methods typically align the global motion-text concepts directly, suffering from sub-optimal generalization due to the uncertainty of correspondence learning between multiple motion concepts coupled in a single motion/text sequence. Therefore, we study the explicit fine-grained concept decomposition for alignment learning and present a novel framework, Structural Generative Augmentation for 3D Human Motion Retrieval (SGAR), to enable generation-augmented retrieval. Specifically, relying on the strong priors of existing large language model (LLM) assets, we effectively decompose human motions structurally into subtler semantic units, \ie, body parts, for fine-grained motion modeling. Based on this, we develop part-mixture learning to better decouple the local motion concept learning, boosting part-level alignment. Moreover, a directional relation alignment strategy exploiting the correspondence between full-body and part motions is incorporated to regularize feature manifold for better consistency. Extensive experiments on three benchmarks, including motion-text retrieval as well as recognition and generation applications, demonstrate the superior performance and promising transferability of our method. Jiahang Zhang 0001, Lilang Lin, Shuai Yang 0001, Jiaying Liu 0001 |
NeurIPS | 3 |
| 2025 | E3DGE: Self-Supervised Geometry-Aware Encoder for Style-Based 3D GAN Inversion
Yushi Lan, Xuyi Meng, Shuai Yang 0001, Chen Change Loy, Bo Dai 0002 |
Int. J. Comput. Vis. | 3 |
| 2024 | VideoBooth: Diffusion-based Video Generation with Image PromptsabstractText-driven video generation witnesses rapid progress. However, merely using text prompts is not enough to depict the desired subject appearance that accurately aligns with users' intents, especially for customized content creation. In this paper, we study the task of video generation with image prompts, which provide more accurate and direct content control beyond the text prompts. Specifically, we propose a feed-forward framework VideoBooth, with two dedicated designs: 1) We propose to embed image prompts in a coarse-to-fine manner. Coarse visual embeddings from image encoder provide high-level encodings of image prompts, while fine visual embeddings from the proposed attention injection module provide multi-scale and detailed encoding of image prompts. These two complementary embeddings can faithfully capture the desired appearance. 2) In the attention injection module at fine level, multi-scale image prompts are fed into different cross-frame attention layers as additional keys and values. This extra spatial in-formation refines the details in the first frame and then it is propagated to the remaining frames, which maintains temporal consistency. Extensive experiments demonstrate that Video Booth achieves state-of-the-art performance in gener-ating customized high-quality videos with subjects specified in image prompts. Notably, VideoBooth is a generalizable framework where a single model works for a wide range of image prompts with only feed-forward passes. Yuming Jiang 0003, Tianxing Wu 0002, Shuai Yang 0001, Chenyang Si, Dahua Lin, Yu Qiao 0001, Chen Change Loy, Ziwei Liu 0002 |
CVPR | 3 |
| 2024 | Low-Rank Approximation for Sparse Attention in Multi-Modal LLMsabstractThis paper focuses on the high computational complexity in Large Language Models (LLMs), a significant challenge in both natural language processing (NLP) and multi-modal tasks. We propose Low-Rank Approximation for Sparse Attention (LoRA -Sparse), an innovative approach that strategically reduces this complexity. LoRA -Sparse introduces low-rank linear projection layers for sparse attention approximation. It utilizes an order-mimic training methodology, which is crucial for efficiently approximating the self-attention mechanism in LLMs. We empirically show that sparse attention not only reduces computational demands, but also enhances model performance in both NLP and multi-modal tasks. This surprisingly shows that redundant attention in LLMs might be non-beneficial. We extensively validate LoRA -Sparse through rigorous empirical studies in both (NLP) and multi-modal tasks, demonstrating its effectiveness and general applicability. Based on LLaMA and LLaVA models, our methods can reduce more than half of the self-attention computation with even better performance than full-attention baselines. Lin Song 0002, Yukang Chen, Shuai Yang 0001, Xiaohan Ding, Yixiao Ge, Ying-Cong Chen, Ying Shan |
CVPR | 3 |
| 2024 | Fresco: Spatial-Temporal Correspondence for Zero-Shot Video TranslationabstractThe remarkable efficacy of text-to-image diffusion models has motivated extensive exploration of their potential application in video domains. Zero-shot methods seek to extend image diffusion models to videos without necessitating model training. Recent methods mainly focus on incorporating inter-frame correspondence into attention mechanisms. However, the soft constraint imposed on determining where to attend to valid features can sometimes be insufficient, resulting in temporal inconsistency. In this paper, we introduce FRESCO, intra-frame correspondence alongside inter-frame correspondence to establish a more robust spatial-temporal constraint. This enhancement ensures a more consistent transformation of semantically similar content across frames. Beyond mere attention guidance, our approach involves an explicit update of features to achieve high spatial-temporal consistency with the input video, significantly improving the visual coherence of the resulting translated videos. Extensive experiments demonstrate the effectiveness of our proposed framework in producing high-quality, coherent videos, marking a notable improvement over existing zero-shot methods. Shuai Yang 0001, Yifan Zhou 0001, Ziwei Liu 0002, Chen Change Loy |
CVPR | 1 |
| 2024 | GroupDiff: Diffusion-Based Group Portrait Editing
Yuming Jiang 0003, Nanxuan Zhao, Qing Liu 0017, Krishna Kumar Singh, Shuai Yang 0001, Chen Change Loy, Ziwei Liu 0002 |
ECCV (34) | 5 |
| 2024 | LN3Diff: Scalable Latent Neural Fields Diffusion for Speedy 3D Generation
Yushi Lan, Fangzhou Hong, Shuai Yang 0001, Shangchen Zhou, Xuyi Meng, Bo Dai 0002, Xingang Pan, Chen Change Loy |
ECCV (4) | 3 |
| 2024 | Defect Spectrum: A Granular Look of Large-Scale Defect Datasets with Rich Semantics
Shuai Yang 0001, Zhifei Chen, Pengguang Chen, Yixun Liang, Shu Liu 0005, Ying-Cong Chen |
ECCV (7) | 1 |
| 2024 | Denoising Diffusion Step-aware ModelsabstractDenoising Diffusion Probabilistic Models (DDPMs) have garnered popularity for data generation across various domains. However, a significant bottleneck is the necessity for whole-network computation during every step of the generative process, leading to high computational overheads. This paper presents a novel framework, Denoising Diffusion Step-aware Models (DDSM), to address this challenge. Unlike conventional approaches, DDSM employs a spectrum of neural networks whose sizes are adapted according to the importance of each generative step, as determined through evolutionary search. This step-wise network variation effectively circumvents redundant computational efforts, particularly in less critical steps, thereby enhancing the efficiency of the diffusion model. Furthermore, the step-aware design can be seamlessly integrated with other efficiency-geared diffusion models such as DDIMs and latent diffusion, thus broadening the scope of computational savings. Empirical evaluations demonstrate that DDSM achieves computational savings of 49% for CIFAR-10, 61% for CelebA-HQ, 59% for LSUN-bedroom, 71% for AFHQ, and 76% for ImageNet, all without compromising the generation quality. Our code and models are available at https://github.com/EnVision-Research/DDSM. Shuai Yang 0001, Yukang Chen, Luozhou Wang, Shu Liu 0005, Ying-Cong Chen |
ICLR | 1 |
| 2024 | COCO-LC: Colorfulness Controllable Language-based ColorizationabstractLanguage-based image colorization aims to convert grayscale images to plausible and visually pleasing color images with language guidance, enjoying wide applications in historical photo restoration and film industry. Existing methods mainly leverage large language models and diffusion models to incorporate language guidance into the colorization process. However, it is still a great challenge to build accurate correspondence between the gray image and the semantic instructions, leading to mismatched, overflowing and under-saturated colors. In this paper, we introduce a novel coarse-to-fine framework, COlorfulness COntrollable Language-based Colorization (COCO-LC), that effectively reinforces the image-text correspondence with a coarsely colorized results. In addition, a multi-level condition that leverages both low-level and high-level cues of the gray image is introduced to realize accurate semantic-aware colorization without color overflows. Furthermore, we condition COCO-LC with a scale factor to determine the colorfulness of the output, flexibly meeting the different needs of users. We validate the superiority of COCO-LC over state-of-the-art image colorization methods in accurate, realistic and controllable colorization through extensive experiments. The code and demo will be released at https://lyf1212.github.io/COCO-LC. Shuai Yang 0001, Jiaying Liu 0001 |
ACM Multimedia | 3 |
| 2024 | CoolColor: Text-guided COherent OLd film COLORization
Zichuan Huang, Shuai Yang 0001, Jiaying Liu 0001 |
MMAsia | 3 |
| 2024 | MMScan: A Multi-Modal 3D Scene Dataset with Hierarchical Grounded Language AnnotationsabstractWith the emergence of LLMs and their integration with other data modalities, multi-modal 3D perception attracts more attention due to its connectivity to the physical world and makes rapid progress. However, limited by existing datasets, previous works mainly focus on understanding object properties or inter-object spatial relationships in a 3D scene. To tackle this problem, this paper builds the first largest ever multi-modal 3D scene dataset and benchmark with hierarchical grounded language annotations, MMScan. It is constructed based on a top-down logic, from region to object level, from a single target to inter-target relationships, covering holistic aspects of spatial and attribute understanding. The overall pipeline incorporates powerful VLMs via carefully designed prompts to initialize the annotations efficiently and further involve humans' correction in the loop to ensure the annotations are natural, correct, and comprehensive. Built upon existing 3D scanning data, the resulting multi-modal 3D dataset encompasses 1.4M meta-annotated captions on 109k objects and 7.7k regions as well as over 3.04M diverse samples for 3D visual grounding and question-answering benchmarks. We evaluate representative baselines on our benchmarks, analyze their capabilities in different aspects, and showcase the key problems to be addressed in the future. Furthermore, we use this high-quality dataset to train state-of-the-art 3D visual grounding and LLMs and obtain remarkable performance improvement both on existing benchmarks and in-the-wild evaluation. Ruiyuan Lyu, Jingli Lin, Shuai Yang 0001, Xiaohan Mao, Runsen Xu, Haifeng Huang 0001, Chenming Zhu, Dahua Lin, Jiangmiao Pang |
NeurIPS | 4 |
| 2024 | Video Diffusion Models are Training-free Motion Interpreter and ControllerabstractVideo generation primarily aims to model authentic and customized motion across frames, making understanding and controlling the motion a crucial topic. Most diffusion-based studies on video motion focus on motion customization with training-based paradigms, which, however, demands substantial training resources and necessitates retraining for diverse models. Crucially, these approaches do not explore how video diffusion models encode cross-frame motion information in their features, lacking interpretability and transparency in their effectiveness. To answer this question, this paper introduces a novel perspective to understand, localize, and manipulate motion-aware features in video diffusion models. Through analysis using Principal Component Analysis (PCA), our work discloses that robust motion-aware feature already exists in video diffusion models. We present a new MOtion FeaTure (MOFT) by eliminating content correlation information and filtering motion channels. MOFT provides a distinct set of benefits, including the ability to encode comprehensive motion information with clear interpretability, extraction without the need for training, and generalizability across diverse architectures. Leveraging MOFT, we propose a novel training-free video motion control framework. Our method demonstrates competitive performance in generating natural and faithful motion, providing architecture-agnostic insights and applicability in a variety of downstream tasks. Zeqi Xiao, Yifan Zhou 0001, Shuai Yang 0001, Xingang Pan |
NeurIPS | 3 |
| 2023 | Self-Supervised Geometry-Aware Encoder for Style-Based 3D GAN InversionabstractStyleGAN has achieved great progress in 2D face reconstruction and semantic editing via image inversion and latent editing. While studies over extending 2D StyleGAN to 3D faces have emerged, a corresponding generic 3D GAN inversion framework is still missing, limiting the applications of 3D face reconstruction and semantic editing. In this paper, we study the challenging problem of 3D GAN inversion where a latent code is predicted given a single face image to faithfully recover its 3D shapes and detailed textures. The problem is ill-posed: innumerable compositions of shape and texture could be rendered to the current image. Furthermore, with the limited capacity of a global latent code, 2D inversion methods cannot preserve faithful shape and texture at the same time when applied to 3D models. To solve this problem, we devise an effective self-training scheme to constrain the learning of inversion. The learning is done efficiently without any real-world 2D-3D training pairs but proxy samples generated from a 3D GAN. In addition, apart from a global latent code that captures the coarse shape and texture information, we augment the generation network with a local branch, where pixel-aligned features are added to faithfully reconstruct face details. We further consider a new pipeline to perform 3D view-consistent editing. Extensive experiments show that our method outperforms state-of-the-art inversion methods in both shape and texture reconstruction quality. Yushi Lan, Xuyi Meng, Shuai Yang 0001, Chen Change Loy, Bo Dai 0002 |
CVPR | 3 |
| 2023 | Text2Performer: Text-Driven Human Video GenerationabstractText-driven content creation has evolved to be a transformative technique that revolutionizes creativity. Here we study the task of text-driven human video generation, where a video sequence is synthesized from texts describing the appearance and motions of a target performer. Compared to general text-driven video generation, human-centric video generation requires maintaining the appearance of synthesized human while performing complex motions. In this work, we present Text2Performer to generate vivid human videos with articulated motions from texts. Text2Performer has two novel designs: 1) decomposed human representation and 2) diffusion-based motion sampler. First, we decompose the VQVAE latent space into human appearance and pose representation in an unsupervised manner by utilizing the nature of human videos. In this way, the appearance is well maintained along the generated frames. Then, we propose continuous VQ-diffuser to sample a sequence of pose embeddings. Unlike existing VQ-based methods that operate in the discrete space, continuous VQ-diffuser directly outputs the continuous pose embeddings for better motion modeling. Finally, motion-aware masking strategy is designed to mask the pose embeddings spatial-temporally to enhance the temporal coherence. Moreover, to facilitate the task of text-driven human video generation, we contribute a Fashion-Text2Video dataset with manually annotated action labels and text descriptions. Extensive experiments demonstrate that Text2Performer generates high-quality human videos (up to 512 × 256 resolution) with diverse appearances and flexible motions. Our project page is https://yumingj.github.io/projects/Text2Performer.html Yuming Jiang 0003, Shuai Yang 0001, Tong Liang Koh, Wayne Wu, Chen Change Loy, Ziwei Liu 0002 |
ICCV | 2 |
| 2023 | Scenimefy: Learning to Craft Anime Scene via Semi-Supervised Image-to-Image TranslationabstractAutomatic high-quality rendering of anime scenes from complex real-world images is of significant practical value. The challenges of this task lie in the complexity of the scenes, the unique features of anime style, and the lack of high-quality datasets to bridge the domain gap. Despite promising attempts, previous efforts are still incompetent in achieving satisfactory results with consistent semantic preservation, evident stylization, and fine details. In this study, we propose Scenimefy, a novel semi-supervised image-to-image translation framework that addresses these challenges. Our approach guides the learning with structure-consistent pseudo paired data, simplifying the pure unsupervised setting. The pseudo data are derived uniquely from a semantic-constrained StyleGAN leveraging rich model priors like CLIP. We further apply segmentation-guided data selection to obtain high-quality pseudo supervision. A patch-wise contrastive style loss is introduced to improve stylization and fine details. Besides, we contribute a high-resolution anime scene dataset to facilitate future research. Our extensive experiments demonstrate the superiority of our method over state-of-the-art baselines in terms of both perceptual quality and quantitative performance. Project page: https://yuxinn-j.github.io/projects/Scenimefy.html. Liming Jiang 0001, Shuai Yang 0001, Chen Change Loy |
ICCV | 3 |
| 2023 | Not All Steps are Created Equal: Selective Diffusion Distillation for Image ManipulationabstractConditional diffusion models have demonstrated impressive performance in image manipulation tasks. The general pipeline involves adding noise to the image and then denoising it. However, this method faces a trade-off problem: adding too much noise affects the fidelity of the image while adding too little affects its editability. This largely limits their practical applicability. In this paper, we propose a novel framework, Selective Diffusion Distillation (SDD), that ensures both the fidelity and editability of images. Instead of directly editing images with a diffusion model, we train a feedforward image manipulation network under the guidance of the diffusion model. Besides, we propose an effective indicator to select the semantic-related timestep to obtain the correct semantic guidance from the diffusion model. This approach successfully avoids the dilemma caused by the diffusion process. Our extensive experiments demonstrate the advantages of our framework. Code is released at https://github.com/AndysonYs/Selective-Diffusion-Distillation. Luozhou Wang, Shuai Yang 0001, Shu Liu 0005, Ying-Cong Chen |
ICCV | 2 |
| 2023 | StyleGANEX: StyleGAN-Based Manipulation Beyond Cropped Aligned FacesabstractRecent advances in face manipulation using StyleGAN have produced impressive results. However, StyleGAN is inherently limited to cropped aligned faces at a fixed image resolution it is pre-trained on. In this paper, we propose a simple and effective solution to this limitation by using dilated convolutions to rescale the receptive fields of shallow layers in StyleGAN, without altering any model parameters. This allows fixed-size small features at shallow layers to be extended into larger ones that can accommodate variable resolutions, making them more robust in characterizing unaligned faces. To enable real face inversion and manipulation, we introduce a corresponding encoder that provides the first-layer feature of the extended StyleGAN in addition to the latent style code. We validate the effectiveness of our method using unaligned face inputs of various resolutions in a diverse set of face manipulation tasks, including facial attribute editing, super-resolution, sketch/mask-to-face translation, and face toonification. Project page https://www.mmlab-ntu.com/project/styleganex Shuai Yang 0001, Liming Jiang 0001, Ziwei Liu 0002, Chen Change Loy |
ICCV | 1 |
| 2023 | DeformToon3d: Deformable Neural Radiance Fields for 3D ToonificationabstractIn this paper, we address the challenging problem of 3D toonification, which involves transferring the style of an artistic domain onto a target 3D face with stylized geometry and texture. Although fine-tuning a pre-trained 3D GAN on the artistic domain can produce reasonable performance, this strategy has limitations in the 3D domain. In particular, fine-tuning can deteriorate the original GAN latent space, which affects subsequent semantic editing, and requires independent optimization and storage for each new style, limiting flexibility and efficient deployment. To overcome these challenges, we propose DeformToon3d, an effective toonification framework tailored for hierarchical 3D GAN. Our approach decomposes 3D toonification into subproblems of geometry and texture stylization to better preserve the original latent space. Specifically, we devise a novel StyleField that predicts conditional 3D deformation to align a real-space NeRF to the style space for geometry stylization. Thanks to the StyleField formulation, which already handles geometry stylization well, texture stylization can be achieved conveniently via adaptive style mixing that injects information of the artistic domain into the decoder of the pre-trained 3D GAN. Due to the unique design, our method enables flexible style degree control and shape-texture-specific style swap. Furthermore, we achieve efficient training without any real-world 2D-3D training pairs but proxy samples synthesized from off-the-shelf 2D toonification models. Code is released at https://github.com/junzhezhang/DeformToon3D. Junzhe Zhang 0002, Yushi Lan, Shuai Yang 0001, Fangzhou Hong, Chai Kiat Yeo, Ziwei Liu 0002, Chen Change Loy |
ICCV | 3 |
| 2023 | HyperDreamer: Hyper-Realistic 3D Content Generation and Editing from a Single Imageabstract3D content creation from a single image is a long-standing yet highly desirable task. Recent advances introduce 2D diffusion priors, yielding reasonable results. However, existing methods are not hyper-realistic enough for post-generation usage, as users cannot view, render and edit the resulting 3D content from a full range. To address these challenges, we introduce HyperDreamer with several key designs and appealing properties: 1) Full-range viewable: 360° mesh modeling with high-resolution textures enables the creation of visually compelling 3D models from a full range of observation points. 2) Full-range renderable: Fine-grained semantic segmentation and data-driven priors are incorporated as guidance to learn reasonable albedo, roughness, and specular properties of the materials, enabling semantic-aware arbitrary material estimation. 3) Full-range editable: For a generated model or their own data, users can interactively select any region via a few clicks and efficiently edit the texture with text-based guidance. Extensive experiments demonstrate the effectiveness of HyperDreamer in modeling region-aware materials with high-resolution textures and enabling user-friendly editing. We believe that HyperDreamer holds promise for advancing 3D content creation and finding applications in various domains. Zhibing Li, Shuai Yang 0001, Pan Zhang 0001, Xingang Pan, Jiaqi Wang 0003, Dahua Lin, Ziwei Liu 0002 |
SIGGRAPH Asia | 3 |
| 2023 | Rerender A Video: Zero-Shot Text-Guided Video-to-Video TranslationabstractLarge text-to-image diffusion models have exhibited impressive proficiency in generating high-quality images. However, when applying these models to video domain, ensuring temporal consistency across video frames remains a formidable challenge. This paper proposes a novel zero-shot text-guided video-to-video translation framework to adapt image models to videos. The framework includes two parts: key frame translation and full video translation. The first part uses an adapted diffusion model to generate key frames, with hierarchical cross-frame constraints applied to enforce coherence in shapes, textures and colors. The second part propagates the key frames to other frames with temporal-aware patch matching and frame blending. Our framework achieves global style and local texture temporal consistency at a low cost (without re-training or optimization). The adaptation is compatible with existing image diffusion techniques, allowing our framework to take advantage of them, such as customizing a specific subject with LoRA, and introducing extra spatial guidance with ControlNet. Extensive experimental results demonstrate the effectiveness of our proposed framework over existing methods in rendering high-quality and temporally-coherent videos. Code is available at our project page: https://www.mmlab-ntu.com/project/rerender/ Shuai Yang 0001, Yifan Zhou 0001, Ziwei Liu 0002, Chen Change Loy |
SIGGRAPH Asia | 1 |
| 2023 | GP-UNIT: Generative Prior for Versatile Unsupervised Image-to-Image TranslationabstractRecent advances in deep learning have witnessed many successful unsupervised image-to-image translation models that learn correspondences between two visual domains without paired data. However, it is still a great challenge to build robust mappings between various domains especially for those with drastic visual discrepancies. In this paper, we introduce a novel versatile framework, Generative Prior-guided UNsupervised Image-to-image Translation (GP-UNIT), that improves the quality, applicability and controllability of the existing translation models. The key idea of GP-UNIT is to distill the generative prior from pre-trained class-conditional GANs to build coarse-level cross-domain correspondences, and to apply the learned prior to adversarial translations to excavate fine-level correspondences. With the learned multi-level content correspondences, GP-UNIT is able to perform valid translations between both close domains and distant domains. For close domains, GP-UNIT can be conditioned on a parameter to determine the intensity of the content correspondences during translation, allowing users to balance between content and style consistency. For distant domains, semi-supervised learning is explored to guide GP-UNIT to discover accurate semantic correspondences that are hard to learn solely from the appearance. We validate the superiority of GP-UNIT over state-of-the-art translation models in robust, high-quality and diversified translations between various domains through extensive experiments. Shuai Yang 0001, Liming Jiang 0001, Ziwei Liu 0002, Chen Change Loy |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Intelligent Typography: Artistic Text Style Transfer for Complex Texture and StructureabstractText style transfer is an important task to render artistic texts from a reference image or style, and is widely desired in many visual creations. Previous works have brought some efficient methods for text style transfer, which facilitate users to design various artistic texts automatically. However, these works mainly focus on relatively simple text effects, and do not perform well on complex reference styles. In this paper, we propose a coarse-to-fine framework to generate exquisite texts with complex texture and structure in an unsupervised way, achieving real-time control of style scales (i.e., text stylistic degree or deformation degree). The key idea is to decouple the overall task into two steps, prototype generation and detail refinement, and explore delicate networks for each step to imitate the features at different levels. Based on this idea, in the first step, we present a novel pro-gen GAN to generate prototypes of artistic texts using the reference style, and develop a deformable module to empower the pro-gen GAN to continuously characterize the multi-scale shape features without network retraining. Furthermore, we propose a mix-attention training scheme for text style transfer, which can avoid artifacts and retain a clear text background. In the second step, we introduce two optimized networks for detail refinements. Experimental results show that the proposed method can synthesize exquisite stylized texts with complex reference styles, and surpass the state of the arts in texture reconstruction, contour imitation, and text image quality drastically. Wendong Mao, Shuai Yang 0001, Huihong Shi, Jiaying Liu 0001, Zhongfeng Wang 0001 |
IEEE Trans. Multim. | 2 |
| 2022 | Pastiche Master: Exemplar-Based High-Resolution Portrait Style TransferabstractRecent studies on StyleGAN show high performance on artistic portrait generation by transfer learning with limited data. In this paper, we explore more challenging exemplar-based high-resolution portrait style transfer by introducing a novel DualStyleGAN with flexible control of dual styles of the original face domain and the extended artistic portrait domain. Different from StyleGAN, DualStyleGAN provides a natural way of style transfer by characterizing the content and style of a portrait with an intrinsic style path and a new extrinsic style path, respectively. The del-icately designed extrinsic style path enables our model to modulate both the color and complex structural styles hierarchically to precisely pastiche the style example. Furthermore, a novel progressive fine-tuning scheme is introduced to smoothly transform the generative space of the model to the target domain, even with the above modifications on the network architecture. Experiments demonstrate the superiority of DualStyleGAN over state-of-the-art methods in high-quality portrait style transfer and flexible stylecontrol. Code is available at https://github.com/williamyang1991/DualStyleGAN. Shuai Yang 0001, Liming Jiang 0001, Ziwei Liu 0002, Chen Change Loy |
CVPR | 1 |
| 2022 | Unsupervised Image-to-Image Translation with Generative PriorabstractUnsupervised image-to-image translation aims to learn the translation between two visual domains without paired data. Despite the recent progress in image translation models, it remains challenging to build mappings between complex domains with drastic visual discrepancies. In this work, we present a novel framework, Generative Priorguided UNsupervised Image-to-image Translation (GP-UNIT), to improve the overall quality and applicability of the translation algorithm. Our key insight is to leverage the generative prior from pre-trained class-conditional GANs (e.g., BigGAN) to learn rich content correspondences across various domains. We propose a novel coarse-to-fine scheme: we first distill the generative prior to capture a robust coarse-level content representation that can link objects at an abstract semantic level, based on which finelevel content features are adaptively learned for more accurate multi-level content correspondences. Extensive experiments demonstrate the superiority of our versatile framework over state-of-the-art methods in robust, high-quality and diversified translations, even for challenging and distant domains. Code is available at https://github.com/williamyang1991/GP-UNIT. Shuai Yang 0001, Liming Jiang 0001, Ziwei Liu 0002, Chen Change Loy |
CVPR | 1 |
| 2022 | Shape-Matching GAN++: Scale Controllable Dynamic Artistic Text Style TransferabstractDynamic artistic text style transfer aims to migrate the style in terms of both the appearance and motion patterns from a reference style video to the target text to create artistic text animation. Recent researches have improved the usability of transfer models by introducing texture control. However, it remains an important open challenge to investigate the control of the stylistic degree with respect to shape deformation. In this paper, we explore a new problem of dynamic artistic text style transfer with glyph stylistic degree control. The key idea is to build multi-scale glyph-style shape mappings through a novel bidirectional shape matching framework. Following this idea, we first introduce a scale-ware Shape-Matching GAN to learn such mappings to simultaneously model the style shape features at multiple scales and transfer them onto the target glyph. Furthermore, an advanced Shape-Matching GAN++ is proposed to animate a static text image based on the reference style video. Our Shape-Matching GAN++ characterizes the short-term consistency of motion patterns via shape matchings within consecutive frames, which are propagated to achieve effective long-term consistency. Experiments show that the proposed method outperforms previous state-of-the-arts both qualitatively and quantitatively, and generate high-quality and controllable artistic text. Shuai Yang 0001, Zhangyang Wang, Jiaying Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | CLAST: Contrastive Learning for Arbitrary Style TransferabstractArbitrary style transfer aims at migrating the style of a reference style painting to a target content image. Existing methods find it challenging to achieve good content fidelity and style migration at the same time. Moreover, they all rely on manually defined content and style, which is of limited universality and robustness. In this paper, we propose to introduce contrastive learning into style transfer, instructing the network to automatically learn to model the structural content and artistic style based on natural contrastive relationships in style transfer. Compared with existing methods, our learned modeling of content and style is more robust and universal. In addition, we further propose instance-wise contrastive style losses and a patch-wise contrastive content loss to guide style transfer. Combining the proposed contrastive losses and two self-reconstruction strategies, we develop a new style transfer framework, which is pluggable and can be flexibly applied to various style transfer modules. Experimental results demonstrate that our method has strong flexibility and synthesizes stylized images with higher quality. Wenjing Wang 0001, Shuai Yang 0001, Jiaying Liu 0001 |
IEEE Trans. Image Process. | 3 |
| 2022 | Text2Human: text-driven controllable human image generationabstractGenerating high-quality and diverse human images is an important yet challenging task in vision and graphics. However, existing generative models often fall short under the high diversity of clothing shapes and textures. Furthermore, the generation process is even desired to be intuitively controllable for layman users. In this work, we present a text-driven controllable framework, Text2Human, for a high-quality and diverse human generation. We synthesize full-body human images starting from a given human pose with two dedicated steps. 1) With some texts describing the shapes of clothes, the given human pose is first translated to a human parsing map. 2) The final human image is then generated by providing the system with more attributes about the textures of clothes. Specifically, to model the diversity of clothing textures, we build a hierarchical texture-aware codebook that stores multi-scale neural representations for each type of texture. The codebook at the coarse level includes the structural representations of textures, while the codebook at the fine level focuses on the details of textures. To make use of the learned hierarchical codebook to synthesize desired images, a diffusion-based transformer sampler with mixture of experts is firstly employed to sample indices from the coarsest level of the codebook, which then is used to predict the indices of the codebook at finer levels. The predicted indices at different levels are translated to human images by the decoder learned accompanied with hierarchical codebooks. The use of mixture-of-experts allows for the generated image conditioned on the fine-grained text input. The prediction for finer level indices refines the quality of clothing textures. Extensive quantitative and qualitative evaluations demonstrate that our proposed Text2Human framework can generate more diverse and realistic human images compared to state-of-the-art methods. Our project page is https://yumingj.github.io/projects/Text2Human.html. Code and pretrained models are available at https://github.com/yumingj/Text2Human. Yuming Jiang 0003, Shuai Yang 0001, Haonan Qiu, Wayne Wu, Chen Change Loy, Ziwei Liu 0002 |
ACM Trans. Graph. | 2 |
| 2022 | VToonify: Controllable High-Resolution Portrait Video Style TransferabstractGenerating high-quality artistic portrait videos is an important and desirable task in computer graphics and vision. Although a series of successful portrait image toonification models built upon the powerful StyleGAN have been proposed, these image-oriented methods have obvious limitations when applied to videos, such as the fixed frame size, the requirement of face alignment, missing non-facial details and temporal inconsistency. In this work, we investigate the challenging controllable high-resolution portrait video style transfer by introducing a novel VToonify framework. Specifically, VToonify leverages the mid- and high-resolution layers of StyleGAN to render high-quality artistic portraits based on the multi-scale content features extracted by an encoder to better preserve the frame details. The resulting fully convolutional architecture accepts non-aligned faces in videos of variable size as input, contributing to complete face regions with natural motions in the output. Our framework is compatible with existing StyleGAN-based image toonification models to extend them to video toonification, and inherits appealing features of these models for flexible style control on color and intensity. This work presents two instantiations of VToonify built upon Toonify and DualStyleGAN for collection-based and exemplar-based portrait video style transfer, respectively. Extensive experimental results demonstrate the effectiveness of our proposed VToonify framework over existing methods in generating high-quality and temporally-coherent artistic portrait videos with flexible style controls. Code and pretrained models are available at our project page: www.mmlab-ntu.com/project/vtoonify/. Shuai Yang 0001, Liming Jiang 0001, Ziwei Liu 0002, Chen Change Loy |
ACM Trans. Graph. | 1 |
| 2021 | Instance-Aware Coherent Video Style Transfer for Chinese Ink Wash PaintingabstractRecent researches have made remarkable achievements in fast video style transfer based on western paintings. However, due to the inherent different drawing techniques and aesthetic expressions of Chinese ink wash painting, existing methods either achieve poor temporal consistency or fail to transfer the key freehand brushstroke characteristics of Chinese ink wash painting. In this paper, we present a novel video style transfer framework for Chinese ink wash paintings. The two key ideas are a multi-frame fusion for temporal coherence and an instance-aware style transfer. The frame reordering and stylization based on reference frame fusion are proposed to improve temporal consistency. Meanwhile, the proposed method is able to adaptively leave the white spaces in the background and to select proper scales to extract features and depict the foreground subject by leveraging instance segmentation. Experimental results demonstrate the superiority of the proposed method over state-of-the-art style transfer methods in terms of both temporal coherence and visual quality. Our project website is available at https://oblivioussy.github.io/InkVideo/. Shuai Yang 0001, Wenjing Wang 0001, Jiaying Liu 0001 |
IJCAI | 2 |
| 2021 | Edit Like A Designer: Modeling Design Workflows for Unaligned Fashion EditingabstractFashion editing has drawn increasing research interest with its extensive application prospect. Instead of directly manipulating the real fashion item image, it is more intuitive for designers to modify it via the design draft. In this paper, we model design workflows for a novel task of unaligned fashion editing, allowing the user to edit a fashion item through manipulating its corresponding design draft. The challenge lies in the large misalignment between the real fashion item and the design draft, which could severely degrade the quality of editing results. To address this issue, we propose an Unaligned Fashion Editing Network (UFE-Net). A coarsely rendered fashion item is firstly generated from the edited design draft via a translation module. With this as guidance, we align and manipulate the original unedited fashion item via a novel alignment-driven fashion editing module, and then optimize the details and shape via a reference-guided refinement module. Furthermore, a joint training strategy is introduced to exploit the synergy between the alignment and editing tasks. Our UFE-Net enables the edited fashion item to have semantically consistent geometric shape and realistic details to the edited draft in the edited region, as well as to keep the unedited region intact. Experiments demonstrate our superiority over the competing methods on unaligned fashion editing. Qiyu Dai, Shuai Yang 0001, Wenjing Wang 0001, Jiaying Liu 0001 |
ACM Multimedia | 2 |
| 2021 | Mask-guided GAN for robust text editing in the scene
Boxi Yu, Yong Xu 0007, Shuai Yang 0001, Jiaying Liu 0001 |
Neurocomputing | 4 |
| 2021 | TE141K: Artistic Text Benchmark for Text Effect TransferabstractText effects are combinations of visual elements such as outlines, colors and textures of text, which can dramatically improve its artistry. Although text effects are extensively utilized in the design industry, they are usually created by human experts due to their extreme complexity; this is laborious and not practical for normal users. In recent years, some efforts have been made toward automatic text effect transfer; however, the lack of data limits the capabilities of transfer models. To address this problem, we introduce a new text effects dataset, TE141K,11.Project page:https://daooshee.github.io/TE141K/.with 141,081 text effect/glyph pairs in total. Our dataset consists of 152 professionally designed text effects rendered on glyphs, including English letters, Chinese characters, and Arabic numerals. To the best of our knowledge, this is the largest dataset for text effect transfer to date. Based on this dataset, we propose a baseline approach called text effect transfer GAN (TET-GAN), which supports the transfer of all 152 styles in one model and can efficiently extend to new styles. Finally, we conduct a comprehensive comparison in which 14 style transfer models are benchmarked. Experimental results demonstrate the superiority of TET-GAN both qualitatively and quantitatively and indicate that our dataset is effective and challenging. Shuai Yang 0001, Wenjing Wang 0001, Jiaying Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Controllable Sketch-to-Image Translation for Robust Face SynthesisabstractIn this paper, we propose a novel controllable sketch-to-image translation framework that allows users to interactively and robustly synthesize and edit face images with hand-drawn sketches. Inspired by the coarse-to-fine painting process of human artists, we propose a novel dilation-based sketch refinement method to refine sketches at varied coarse levels without the need for real sketch training data. We further investigate multi-level refinement that enables users to flexibly define how "reliable" the input sketch should be considered for the final output through a refinement level control parameter, which helps balance between the realism of the output and its structural consistency with the input sketch. It is realized by leveraging scale-aware style transfer to model and adjust the style features of sketches at different coarse levels. Moreover, advanced user controllability in terms of the editing region control, facial attribute editing, and spatially non-uniform refinement is further explored for fine-grained and semantic editing. We demonstrate the effectiveness of the proposed method in terms of visual quality and user controllability through extensive experiments including qualitative and quantitative comparison with state-of-the-art methods, ablation studies and various applications. Shuai Yang 0001, Zhangyang Wang, Jiaying Liu 0001, Zongming Guo |
IEEE Trans. Image Process. | 1 |
| 2021 | Towards Coding for Human and Machine Vision: Scalable Face Image CodingabstractThe past decades have witnessed the rapid development of image and video coding techniques in the era of big data. However, the signal fidelity-driven coding pipeline design limits the capability of the existing image/video coding frameworks to fulfill the needs of both machine and human vision. In this paper, we come up with a novel face image coding framework by leveraging both the compressive and the generative models, to support machine vision and human perception tasks jointly. Given an input image, the feature analysis is first applied, and then the generative model is employed to reconstruct image with compact structure and color features, where sparse edges are extracted to connect both kinds of vision and a key reference pixel selection method is proposed to determine the priorities of the reference color pixels for scalable coding. The compact edge map serves as the basic layer for machine vision tasks, and the reference pixels act as an enhanced layer to guarantee signal fidelity for human vision. By introducing advanced generative models, we train a decoding network to reconstruct images from compact structure and color representations, which is flexible to accept inputs in a scalable way and to control the imagery effect of the outputs between signal fidelity and visual realism. Experimental results and comprehensive performance analysis over the face image dataset demonstrate the superiority of our framework in both human vision tasks and machine vision tasks, which provide useful evidence on the emerging standardization efforts on MPEG VCM (Video Coding for Machine). Shuai Yang 0001, Yueyu Hu, Wenhan Yang, Ling-Yu Duan, Jiaying Liu 0001 |
IEEE Trans. Multim. | 1 |
| 2020 | Deep Plastic Surgery: Robust and Controllable Image Editing with Human-Drawn Sketches
Shuai Yang 0001, Zhangyang Wang, Jiaying Liu 0001, Zongming Guo |
ECCV (15) | 1 |
| 2020 | Towards Coding For Human And Machine Vision: A Scalable Image Coding ApproachabstractThe past decades have witnessed the rapid development of image and video coding techniques in the era of big data. However, the signal fidelity-driven coding pipeline design limits the capability of the existing image/video coding frameworks to fulfill the needs of both machine and human vision. In this paper, we come up with a novel image coding framework by leveraging both the compressive and the generative models, to support machine vision and human perception tasks jointly. Given an input image, the feature analysis is first applied, and then the generative model is employed to perform image reconstruction with features and additional reference pixels, in which compact edge maps are extracted in this work to connect both kinds of vision in a scalable way. The compact edge map serves as the basic layer for machine vision tasks, and the reference pixels act as a sort of enhanced layer to guarantee signal fidelity for human vision. By introducing advanced generative models, we train a flexible network to reconstruct images from compact feature representations and the reference pixels. Experimental results demonstrate the superiority of our framework in both human visual quality and facial landmark detection, which provide useful evidence on the emerging standardization efforts on MPEG VCM (Video Coding for Machine). Our project website is available at https://williamyang1991.github.io/projects/VCM-Face/. Yueyu Hu, Shuai Yang 0001, Wenhan Yang, Ling-Yu Duan, Jiaying Liu 0001 |
ICME | 2 |
| 2020 | Multitask Attentive Network For Text Effects Quality AssessmentabstractAlong with the fast development of image style transfer, large amounts of style transfer algorithms were proposed. However, not enough attention has been paid to assess the quality of stylized images, which is of great value in allowing users to efficiently search for high quality images as well as guiding the designing of style transfer algorithms. In this paper, we focus on artistic text stylization and build a novel deep neural network equipped with multitask learning and attention mechanism for text effects quality assessment. We first select stylized images from TE141K [1] dataset and then collect the corresponding visual scores from users. Then through multitask learning, the network learns to extract features related to both style and content information. Furthermore, we employ an attention module to simulate the process of human high-level visual judgement. Experimental results demonstrate the superiority of our network in achieving a high judgement accuracy over the state-of-the-art methods. Our project website is available at https://ykq98.github.io/projects/TEA/. Keqiang Yan, Shuai Yang 0001, Wenjing Wang 0001, Jiaying Liu 0001 |
ICME | 2 |
| 2020 | From Design Draft to Real Attire: Unaligned Fashion Image TranslationabstractFashion manipulation has attracted growing interest due to its great application value, which inspires many researches towards fashion images. However, little attention has been paid to fashion design draft. In this paper, we study a new unaligned translation problem between design drafts and real fashion items, whose main challenge lies in the huge misalignment between the two modalities. We first collect paired design drafts and real fashion item images without pixel-wise alignment. To solve the misalignment problem, our main idea is to train a sampling network to adaptively adjust the input to an intermediate state with structure alignment to the output. Moreover, built upon the sampling network, we present design draft to real fashion item translation network (D2RNet), where two separate translation streams that focus on texture and shape, respectively, are combined tactfully to get both benefits. D2RNet is able to generate realistic garments with both texture and shape consistency to their design drafts. We show that this idea can be effectively applied to the reverse translation problem and present R2DNet accordingly. Extensive experiments on unaligned fashion design translation demonstrate the superiority of our method over state-of-the-art methods. Our project website is available at: https://victoriahy.github.io/MM2020/. Yu Han 0008, Shuai Yang 0001, Wenjing Wang 0001, Jiaying Liu 0001 |
ACM Multimedia | 2 |
| 2020 | Consistent Video Style Transfer via Relaxation and RegularizationabstractIn recent years, neural style transfer has attracted more and more attention, especially for image style transfer. However, temporally consistent style transfer for videos is still a challenging problem. Existing methods, either relying on a significant amount of video data with optical flows or using singleframe regularizers, fail to handle strong motions or complex variations, therefore have limited performance on real videos. In this paper, we address the problem by jointly considering the intrinsic properties of stylization and temporal consistency. We first identify the cause of the conflict between style transfer and temporal consistency, and propose to reconcile this contradiction by relaxing the objective function, so as to make the stylization loss term more robust to motions. Through relaxation, style transfer is more robust to inter-frame variation without degrading the subjective effect. Then, we provide a novel formulation and understanding of temporal consistency. Based on the formulation, we analyze the drawbacks of existing training strategies and derive a new regularization. We show by experiments that the proposed regularization can better balance the spatial and temporal performance. Based on relaxation and regularization, we design a zero-shot video style transfer framework. Moreover, for better feature migration, we introduce a new module to dynamically adjust inter-channel distributions. Quantitative and qualitative results demonstrate the superiority of our method over state-of-the-art style transfer methods. Wenjing Wang 0001, Shuai Yang 0001, Jizheng Xu, Jiaying Liu 0001 |
IEEE Trans. Image Process. | 2 |
| 2019 | TET-GAN: Text Effects Transfer via Stylization and DestylizationabstractText effects transfer technology automatically makes the text dramatically more impressive. However, previous style transfer methods either study the model for general style, which cannot handle the highly-structured text effects along the glyph, or require manual design of subtle matching criteria for text effects. In this paper, we focus on the use of the powerful representation abilities of deep neural features for text effects transfer. For this purpose, we propose a novel Texture Effects Transfer GAN (TET-GAN), which consists of a stylization subnetwork and a destylization subnetwork. The key idea is to train our network to accomplish both the objective of style transfer and style removal, so that it can learn to disentangle and recombine the content and style features of text effects images. To support the training of our network, we propose a new text effects dataset with as much as 64 professionally designed styles on 837 characters. We show that the disentangled feature representations enable us to transfer or remove all these styles on arbitrary glyphs using one network. Furthermore, the flexible network design empowers TET-GAN to efficiently extend to a new text style via oneshot learning where only one example is required. We demonstrate the superiority of the proposed method in generating high-quality stylized text over the state-of-the-art methods. Shuai Yang 0001, Jiaying Liu 0001, Wenjing Wang 0001, Zongming Guo |
AAAI | 1 |
| 2019 | Typography With Decor: Intelligent Text Style TransferabstractText effects transfer can dramatically make the text visually pleasing. In this paper, we present a novel framework to stylize the text with exquisite decor, which is ignored by the previous text stylization methods. Decorative elements pose a challenge to spontaneously handle basal text effects and decor, which are two different styles. To address this issue, our key idea is to learn to separate, transfer and recombine the decors and the basal text effect. A novel text effect transfer network is proposed to infer the styled version of the target text. The stylized text is finally embellished with decor where the placement of the decor is carefully determined by a novel structure-aware strategy. Furthermore, we propose a domain adaptation strategy for decor detection and a one-shot training strategy for text effects transfer, which greatly enhance the robustness of our network to new styles. We base our experiments on our collected topography dataset including 59,000 professionally styled text and demonstrate the superiority of our method over other state-of-the-art style transfer methods. Wenjing Wang 0001, Jiaying Liu 0001, Shuai Yang 0001, Zongming Guo |
CVPR | 3 |
| 2019 | Controllable Artistic Text Style Transfer via Shape-Matching GANabstractArtistic text style transfer is the task of migrating the style from a source image to the target text to create artistic typography. Recent style transfer methods have considered texture control to enhance usability. However, controlling the stylistic degree in terms of shape deformation remains an important open challenge. In this paper, we present the first text style transfer network that allows for real-time control of the crucial stylistic degree of the glyph through an adjustable parameter. Our key contribution is a novel bidirectional shape matching framework to establish an effective glyph-style mapping at various deformation levels without paired ground truth. Based on this idea, we propose a scale-controllable module to empower a single network to continuously characterize the multi-scale shape features of the style image and transfer these features to the target text. The proposed method demonstrates its superiority over previous state-of-the-arts in generating diverse, controllable and high-quality stylized text. Shuai Yang 0001, Zhangyang Wang, Ning Xu 0007, Jiaying Liu 0001, Zongming Guo |
ICCV | 1 |
| 2019 | Artistic Text Stylization for Visual-Textual Presentation SynthesisabstractIn this research, we study a specific task of visual-textual presentation synthesis, where artistic text is generated and embedded in a background photo. The art form of visual-textual presentation is widely used in graphic design such as posters, billboards and trademarks, and therefore is of high application value. We propose a new framework to complete this task. First, the shape of the target text is adjusted and the textures are rendered to match the reference style image to generate artistic text. By considering both aesthetics and seamlessness, the layout where the artistic text is placed is determined. Finally the artistic text is blended with the background photo to obtain the visual-textual presentations. The experimental results demonstrate the effectiveness of the proposed framework in creating professionally designed visual-textual presentations. Shuai Yang 0001 |
MMAsia | 1 |
| 2019 | Selfie retoucher: subject-oriented self-portrait enhancement
Sifeng Xia, Shuai Yang 0001, Jiaying Liu 0001 |
Multim. Tools Appl. | 2 |
| 2019 | D3R-Net: Dynamic Routing Residue Recurrent Network for Video Rain RemovalabstractIn this paper, we address the problem of video rain removal by considering rain occlusion regions, i.e., very low light transmittance for rain streaks. Different from additive rain streaks, in such occlusion regions, the details of backgrounds are completely lost. Therefore, we propose a hybrid rain model to depict both rain streaks and occlusions. Integrating the hybrid model and useful motion segmentation context information, we present a Dynamic Routing Residue Recurrent Network (D3R-Net). D3R-Net first extracts the spatial features by a residual network. Then, the spatial features are aggregated by recurrent units along the temporal axis. In the temporal fusion, the context information is embedded into the network in a "dynamic routing" way. A heap of recurrent units takes responsibility for handling the temporal fusion in given contexts, e.g., rain or non-rain regions. In the certain forward and backward processes, one of these recurrent units is mainly activated. Then, a context selection gate is employed to detect the context and select one of these temporally fused features generated by these recurrent units as the final fused feature. Finally, this last feature plays a role of "residual feature." It is combined with the spatial feature and then used to reconstruct the negative rain streaks. In such a D3R-Net, we incorporate motion segmentation, which denotes whether a pixel belongs to fast moving edges or not, and rain type indicator, indicating whether a pixel belongs to rain streaks, rain occlusions, and non-rain regions, as the context variables. Extensive experiments on a series of synthetic and real videos with rain streaks verify not only the superiority of the proposed method over state of the art but also the effectiveness of our network design and its each component. Jiaying Liu 0001, Wenhan Yang, Shuai Yang 0001, Zongming Guo |
IEEE Trans. Image Process. | 3 |
| 2019 | Context-Aware Text-Based Binary Image Stylization and SynthesisabstractIn this work, we present a new framework for the stylization of text-based binary images. First, our method stylizes the stroke-based geometric shape like text, symbols and icons in the target binary image based on an input style image. Second, the composition of the stylized geometric shape and a background image is explored. To accomplish the task, we propose legibilitypreserving structure and texture transfer algorithms, which progressively narrow the visual differences between the binary image and the style image. The stylization is then followed by a contextaware layout design algorithm, where cues for both seamlessness and aesthetics are employed to determine the optimal layout of the shape in the background. Given the layout, the binary image is seamlessly embedded into the background by texture synthesis under a context-aware boundary constraint. According to the contents of binary images, our method can be applied to many fields.We show that the proposed method is capable of addressing the unsupervised text stylization problem and is superior to stateof- the-art style transfer methods in automatic artistic typography creation. Besides, extensive experiments on various tasks, such as visual-textual presentation synthesis, icon/symbol rendering and structure-guided image inpainting, demonstrate the effectiveness of the proposed method. Shuai Yang 0001, Jiaying Liu 0001, Wenhan Yang, Zongming Guo |
IEEE Trans. Image Process. | 1 |
| 2019 | Scale-Free Single Image Deraining Via Visibility-Enhanced Recurrent Wavelet LearningabstractIn this paper, we address a rain removal problem from a single image, even in the presence of large rain streaks and rain streak accumulation (where individual streaks cannot be seen, and thus visually similar to mist or fog). For rain streak removal, the mismatch problem between different streak sizes in training and testing phases leads to a poor performance, especially when there are large streaks. To mitigate this problem, we embed a hierarchical representation of wavelet transform into a recurrent rain removal process: 1) rain removal on the low-frequency component; 2) recurrent detail recovery on highfrequency components under the guidance of the recovered lowfrequency component. Benefiting from the recurrent multi-scale modeling of wavelet transform-like design, the proposed network trained on streaks with one size can adapt to those with larger sizes, which significantly favors real rain streak removal. The dilated residual dense network is used as the basic model of the recurrent recovery process. The network includes multiple paths with different receptive fields, thus can make full use of multi-scale redundancy and utilize context information in large regions. Furthermore, to handle heavy rain cases where rain streak accumulation is presented, we construct a detail appearing rain accumulation removal to not only improve the visibility but also enhance the details in dark regions. The evaluation on both synthetic and real images, particularly on those containing large rain streaks and heavy accumulation, shows the effectiveness of our novel models, which significantly outperforms the state-ofthe- art methods. Wenhan Yang, Jiaying Liu 0001, Shuai Yang 0001, Zongming Guo |
IEEE Trans. Image Process. | 3 |
| 2018 | Erase or Fill? Deep Joint Recurrent Rain Removal and Reconstruction in VideosabstractIn this paper, we address the problem of video rain removal by constructing deep recurrent convolutional networks. We visit the rain removal case by considering rain occlusion regions, i.e. the light transmittance of rain streaks is low. Different from additive rain streaks, in such rain occlusion regions, the details of background images are completely lost. Therefore, we propose a hybrid rain model to depict both rain streaks and occlusions. With the wealth of temporal redundancy, we build a Joint Recurrent Rain Removal and Reconstruction Network (J4R-Net) that seamlessly integrates rain degradation classification, spatial texture appearances based rain removal and temporal coherence based background details reconstruction. The rain degradation classification provides a binary map that reveals whether a location is degraded by linear additive streaks or occlusions. With this side information, the gate of the recurrent unit learns to make a trade-off between rain streak removal and background details reconstruction. Extensive experiments on a series of synthetic and real videos with rain streaks verify the superiority of the proposed method over previous state-of-the-art methods. Jiaying Liu 0001, Wenhan Yang, Shuai Yang 0001, Zongming Guo |
CVPR | 3 |
| 2018 | Soft Decoding of Light Field Images Using Pocs and Fast Graph Spectrayl FiltersabstractLight field data captured by a lenslet-based image sensor is typically demosaicked, aligned and rearranged into a series of sub-aperture (viewpoint) images, before a disparity-compensated coding scheme is employed for compression. In this paper, we focus on the problem of soft decoding of block-based compressed sub-aperture images at the decoder: given quantization bin indices of DCT coefficients of non-overlapping code blocks, we select appropriate coefficient values that are low-pass filtered using graph spectral filters and view-consistent across sub-aperture images via projection on convex sets (POCS). Specifically, after an initial pixel estimate, we low-pass filter each pixel block using accelerated graph filters based on the Lanczos method. We then map filtered pixels to a neighborhood of sub-aperture images based on estimated disparity to enforce indexed quantization bin constraints of multiple images. Experimental results show that our algorithm achieves PSNR gain of 2.34dB over JPEG hard decoding. Shuai Yang 0001, Gene Cheung, Jiaying Liu 0001, Zongming Guo |
ICASSP | 1 |
| 2018 | Context-Aware Unsupervised Text StylizationabstractIn this work, we present a novel algorithm to stylize the text without supervision, which provides a flexible and convenient way to invoke fantastic text expressions. Rather than employing the fixed pair of target text and source style images, our unsupervised framework establishes an implicit mapping for them by using an abstract imagery of the style image as bridges. Based on the mapping, we progressively narrow the visual discrepancy between text and style images by the proposed legibility-preserving structure transfer and texture transfer algorithms, which effectively balance the text legibility and style consistency. Furthermore, we explore a seamless composition of the stylized text and a background image, in which the optimal text layout is determined by a context-aware layout design algorithm utilizing cues for both seamlessness and aesthetics. Given the layout, the text can be seamlessly embedded into the background by texture synthesis under a context-aware boundary constraint. Experimental results demonstrate the effectiveness of the proposed method in automatic artistic typography creation and visual-textual presentation synthesis. Shuai Yang 0001, Jiaying Liu 0001, Wenhan Yang, Zongming Guo |
ACM Multimedia | 1 |
| 2018 | Text effects transfer via distribution-aware texture synthesis
Shuai Yang 0001, Jiaying Liu 0001, Zhouhui Lian, Zongming Guo |
Comput. Vis. Image Underst. | 1 |
| 2018 | Automatic portrait oil painter: joint domain stylization for portrait images
Saboya Yang, Shuai Yang 0001, Wenhan Yang, Jiaying Liu 0001 |
Multim. Tools Appl. | 2 |
| 2018 | Joint-Feature Guided Depth Map Super-Resolution With Face PriorsabstractIn this paper, we present a novel method to super-resolve and recover the facial depth map nicely. The key idea is to exploit the exemplar-based method to obtain the reliable face priors from high-quality facial depth map to improve the depth image. Specifically, a new neighbor embedding (NE) framework is designed for face prior learning and depth map reconstruction. First, face components are decomposed to form specialized dictionaries and then reconstructed, respectively. Joint features, i.e., low-level depth, intensity cues and high-level position cues, are put forward for robust patch similarity measurement. The NE results are used to obtain the face priors of facial structures and smooth maps, which are then combined in an uniform optimization framework to recover high-quality facial depth maps. Finally, an edge enhancement process is implemented to estimate the final high resolution depth map. Experimental results demonstrate the superiority of our method compared to state-of-the-art depth map super-resolution techniques on both synthetic data and real-world data from Kinect. Shuai Yang 0001, Jiaying Liu 0001, Yuming Fang 0001, Zongming Guo |
IEEE Trans. Cybern. | 1 |
| 2018 | Structure-Guided Image Inpainting Using Homography TransformationabstractIn this paper, we present a novel structure-guided framework for exemplar-based image inpainting to maintain the neighborhood consistence and structure coherence of an inpainted region. The proposed method consists of a data term for pixel validity and boundary continuity, a smoothness term to depict the compatibility of neighboring pixels for contextual continuity, and a coherence term to investigate image inherent regularities to ensure image self-similarity. To better reconstruct image structures, the method utilizes image regularity statistics to extract dominant linear structures of the target image. Guided by these structures, homography transformations are estimated and combined to globally repair the missing region using the Markov random field model. To reduce computational complexity, a hierarchical process is implemented to utilize the regularity effectively. The experimental results demonstrate that our method yields better results for various real-world scenes than existing state-of-the-art image inpainting techniques. Jiaying Liu 0001, Shuai Yang 0001, Yuming Fang 0001, Zongming Guo |
IEEE Trans. Multim. | 2 |
| 2017 | Awesome Typography: Statistics-Based Text Effects TransferabstractIn this work, we explore the problem of generating fantastic special-effects for the typography. It is quite challenging due to the model diversities to illustrate varied text effects for different characters. To address this issue, our key idea is to exploit the analytics on the high regularity of the spatial distribution for text effects to guide the synthesis process. Specifically, we characterize the stylized patches by their normalized positions and the optimal scales to depict their style elements. Our method first estimates these two features and derives their correlation statistically. They are then converted into soft constraints for texture transfer to accomplish adaptive multi-scale texture synthesis and to make style element distribution uniform. It allows our algorithm to produce artistic typography that fits for both local texture patterns and the global spatial distribution in the example. Experimental results demonstrate the superiority of our method for various text effects over conventional style transfer methods. In addition, we validate the effectiveness of our algorithm with extensive artistic typography library generation. Shuai Yang 0001, Jiaying Liu 0001, Zhouhui Lian, Zongming Guo |
CVPR | 1 |
| 2017 | 1+N fusion: Cascaded self-portrait enhancementabstractIn this paper, we present a novel cascaded framework to solve a self-portrait enhancement problem we call “1+N” problem, in which a self-portrait is enhanced with the help of N supporting photos that share the same scene and similar shooting time. The key idea is to exploit the extra information of these N photos to expand the field of view of the self-portrait and improve its lighting style. We achieve this by alternatingly optimizing two complementary tasks, namely illumination unification and photo registration. Based on the correspondences extracted in the input 1+N photos, our method estimates and updates the illumination and registration coefficients in a cascaded manner. Then a Markov Random Field formulation is proposed to globally fuse the aligned photos. Experimental results demonstrate the proposed method achieves high-quality results in this novel application scenario. Shuai Yang 0001, Jiaying Liu 0001, Sifeng Xia, Zongming Guo |
ICASSP | 1 |
| 2017 | Soft segmentation-guided bipartite graph image stylizationabstractIn this paper, we propose a photo stylistic brush, an automatic robust style transfer approach based on soft segmentation-guided bipartite graph. A two-step bipartite graph algorithm with different granularity levels is employed to aggregate pixels into superpixel and find their correspondences. In the first step, with the extracted hierarchical features, a bipartite graph is constructed to describe the content similarity for pixel partition to produce superpixels. In the second step, superpixels in the input/reference image are rematched to form a new soft segmentation-guided bipartite graph, and superpixel-level correspondences are generated by a bipartite matching. Finally, the refined correspondence guides our approach to perform the transfer in a decorrelated color space. Extensive experimental results demonstrate the effectiveness and robustness of the proposed method for transferring various styles of exemplar images, even for some challenging cases, such as night images. Saboya Yang, Jiaying Liu 0001, Wenhan Yang, Shuai Yang 0001, Chunpeng Li |
ICIP | 4 |
| 2017 | Joint-domain unsupervised stylization for portraitsabstractPeople wish to own a portrait painting of themselves by Da Vinci. Unfortunately, it is impossible to make this dream come true; nevertheless, it may give us an opportunity by transferring some artistic features from one single reference painting. To address this issue, we propose a joint-domain image stylization approach, particularly for portrait oil paintings. From the view of artistic appreciation, we analyze an amount of oil painting artworks and summarize three critical factors to depict the figure, i.e. color, structure and texture. First, the tone of the input image is recolored based on semantic regions corresponding to the reference. Those semantic regions are segmented automatically via the color swatch, by considering the constraints of colors and positions. Then, we exploit sparse representation to reconstruct the layout by acquiring the structure from the reference. The paired training set for sparse dictionary learning is built with the guidance of edge features. Third, considering that texture is usually locally stochastic but regularly repetitive in global, a coarse-to-fine texture synthesis is used to enhance the detail pattern. Subjective results demonstrate the proposed method achieves desirable results compared with state-of-art methods while keeping consistent with artist's style. Saboya Yang, Jiaying Liu 0001, Shuai Yang 0001, Wenhan Yang, Zongming Guo |
ISCAS | 3 |
| 2016 | Structure-guided image completion via regularity statisticsabstractIn this paper, we propose a novel hierarchical image completion approach using regularity statistics, considering structure features. Guided by dominant structures, the target image is used to generate reference images in a self-reproductive way by image data enhancement. The structure-guided image data enhancement allows us to expand the search space for samples. A Markov Random Field model is used to guide the enhanced image data combination to globally reconstruct the target image. For lower computational complexity and more accurate structure estimation, a hierarchical process is implemented. Experiments demonstrate the effectiveness of our method comparing to several state-of-the-art image completion techniques. Shuai Yang 0001, Jiaying Liu 0001, Sijie Song, Mading Li, Zongming Quo |
ICASSP | 1 |
| 2016 | Facial depth map enhancement via neighbor embeddingabstractThe simple yet subtle structures of faces make it difficult to capture the fine differences between different facial regions in the depth map, especially for consumer devices like Kinect. To address this issue, we present a novel method to super-solve and recover the facial depth map nicely. The key idea of our approach is to exploit the learning-based method to obtain the reliable face priors from high quality facial depth map to further improve the depth image. Specifically, we utilize the neighbor embedding framework. First, face components are decomposed to train specialized dictionaries and reconstructed, respectively. Joint features, i.e. color, depth and position cues, are put forward for robust patch similarity measurement. The neighbor embedding results form high frequency cues of facial depth details and gradients. Finally, an optimization function is defined to combine these high frequency information to yield depth maps that fit the actual face structures better. Experimental results demonstrate the superiority of our method compared to state-of-the-art techniques in recovering both synthetic data and real world data from Kinect. Shuai Yang 0001, Sijie Song, Qikun Guo, Xiaoqing Lu, Jiaying Liu 0001 |
ICPR | 1 |
| 2015 | Novel autoregressive model based on adaptive window-extension and patch-geodesic distance for image interpolationabstractIn this paper, we propose a novel autoregressive (AR) model based on the adaptive window and the patch-geodesic distance for the image interpolation. The model combines the information of inner/inter-patch correlation. To model the inner-patch correlation, we introduce a patch-geodesic distance similarity metric. The proposed metric shows the desirable capacity to depict the piecewise-stationarity of natural images. For the inter-patch correlation, we introduce the inter-patch structure variation and propose an adaptive window-extension AR model. The model extends the interpolation window according to the local structural variation, increasing the adaptation without violating the consistency. Comprehensive experiments demonstrate that the proposed method is better than or competitive with state-of-the-art interpolation methods in both objective and subjective quality evaluations. Wenhan Yang, Jiaying Liu 0001, Shuai Yang 0001, Zongming Guo |
ICASSP | 3 |
| 2015 | Hierarchical oil painting stylization with limited reference via sparse representationabstractTraditional image stylization is enforced by learning the mappings with an external paired training set. But in practice, people usually encounter a specific stylish image and want to transfer its style to their own pictures without the external dataset. Thus, we propose a hierarchical stylization model with limited reference particularly for oil paintings. First, the edge patch based dictionary is trained to build connections between images and limited reference, then reconstruct the structure layer. Due to the highly structured property of saliency regions, the saliency mask is extracted to integrate the structure layer and the texture layer with different weights. Hence, the advantages of both sparse representation based methods and example based methods are integrated. Moreover, the color layer and the surface layer are considered to make the output more consistent with the artist's individual oil painting style. Subjective results demonstrate the proposed method produces desirable results with state-of-art methods while keeping consistent with the artist's oil painting style. Saboya Yang, Jiaying Liu 0001, Shuai Yang 0001, Sifeng Xia, Zongming Guo |
MMSP | 3 |