EDBT 2026 Demo / reviewers in the wild / expert
Xin Tao 0001
dblp:98/7443-1
· DBLP profile ↗
41ranked-venue papers
3as first author
23since 2021 · last 2026
0000-0001-9126-4746ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 35 · 3 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 32 · 3 first-author · 15 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Boosting Resolution Generalization of Diffusion Transformers with Randomized Positional EncodingsabstractResolution generalization in image generation tasks enables the production of higher-resolution images with lower training resolution overhead. However, a key obstacle for diffusion transformers in addressing this problem is the mismatch between positional encodings seen at inference and those used during training. Existing strategies such as positional encodings interpolation, extrapolation, or hybrids, do not fully resolve this mismatch. In this paper, we propose a novel two-dimensional randomized positional encodings, namely RPE-2D, that prioritizes the order of image patches rather than their absolute distances, enabling seamless high- and low-resolution generation without training on multiple resolutions. Concretely, RPE-2D independently samples positions along the horizontal and vertical axes over an expanded range during training, ensuring that the encodings used at inference lie within the training distribution and thereby improving resolution generalization. We further introduce a simple random resize-and-crop augmentation to strengthen order modeling and add micro-conditioning to indicate the applied cropping pattern. On the ImageNet dataset, RPE-2D achieves state-of-the-art resolution generalization performance, outperforming competitive methods when trained at 256^2 and evaluated at 384^2 and 512^2, and when trained at 512^2 and evaluated at 768^2 and 1024^2. RPE-2D also exhibits outstanding capabilities in low-resolution image generation, multi-stage training acceleration, and multi-resolution inheritance. Mingwu Zheng, Xin Tao 0001, Pengfei Wan 0001, Di Zhang 0026, Kun Gai |
AAAI | 4 |
| 2026 | Less is More: Improving LLM Reasoning with Minimal Test-Time InterventionabstractZhen Yang, Mingyang Zhang, Feng Chen, Ganggui Ding, Liang Hou, Xin Tao, Ying-Cong Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Ganggui Ding, Xin Tao 0001, Ying-Cong Chen |
ACL (1) | 6 |
| 2026 | MultiPaint: A Unified Framework for Multi-Task, Multi-Object, and Multi-Condition Video InpaintingabstractVideo inpainting modifies local regions in video while ensuring spatial and temporal coherence. However, existing methods-both traditional and recent diffusion-based ones-face key limitations: they lack unified support for both insertion and completion, and are restricted to single-object inpainting, making it difficult to handle multi-object scenarios involving grounding and interaction. In this article, we propose MultiPaint, a unified framework for multi-task, multi-object, and multi-condition video inpainting. First, we introduce dual-branch adapters to unify the insertion and completion tasks within a single model. Moreover, we propose a test-time scheduled feature composition strategy that enables multi-object inpainting with user-specified locations while better preserving interactions among objects, a setting that has been insufficiently addressed in prior work. Additionally, we introduce a multi-condition inpainting scheme that integrates text-guided, image-guided, and keyframe-guided modes via dynamic frame masking, providing more controllability in appearance customization. Extensive experiments show that MultiPaint achieves state-of-the-art performance on object insertion and scene completion among the recent works. We further demonstrate its versatility in downstream tasks including grounded video generation, object editing, object removal, image-guided inpainting, and long video inpainting. Zheng Gu 0001, Xin Tao 0001, Pengfei Wan 0001, Xiaodong Chen 0009, Jing Liao 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video ContentabstractWith the continuous progress of visual generation technologies, the scale of video datasets has grown exponentially. The quality of these datasets plays a pivotal role in the performance of video generation models. We assert that temporal splitting, detailed captions, and video quality filtering are three crucial determinants of dataset quality. However, existing datasets exhibit various limitations in these areas. To address these challenges, we introduce Koala-36M, a large-scale, high-quality video dataset featuring accurate temporal splitting, detailed captions, and superior video quality. The essence of our approach lies in improving the consistency between fine-grained conditions and video content. Specifically, we employ a linear classifier on probability distributions to enhance the accuracy of transition detection, ensuring better temporal consistency. We then provide structured captions for the splitted videos, with an average length of 200 words, to improve text-video alignment. Additionally, we develop a Video Training Suitability Score (VTSS) that integrates multiple sub-metrics, allowing us to filter high-quality videos from the original corpus. Finally, we incorporate several metrics into the training process of the generation model, further refining the fine-grained conditions. Our experiments demonstrate the effectiveness of our data processing pipeline and the quality of the proposed Koala-36M dataset. Our dataset and code have been released at https://koala36m.github.io/. Qiuheng Wang, Yukai Shi, Jiarong Ou, Boyuan Jiang, Mingwu Zheng, Xin Tao 0001, Pengfei Wan 0001, Di Zhang 0026 |
CVPR | 10 |
| 2025 | Towards Precise Scaling Laws for Video Diffusion TransformersabstractAchieving optimal performance of video diffusion transformers within given data and compute budget is crucial due to their high training costs. This necessitates precisely determining the optimal model size and training hyperparameters before large-scale training. While scaling laws are employed in language models to predict performance, their existence and accurate derivation in visual generation models remain underexplored. In this paper, we systematically analyze scaling laws for video diffusion transformers and confirm their presence. Moreover, we discover that, unlike language models, video diffusion models are more sensitive to learning rate and batch size—two hyperparameters often not precisely modeled. To address this, we propose a new scaling law that predicts optimal hyperparameters for any model size and compute budget. Under these optimal settings, we achieve comparable performance and reduce inference costs by 40.1% compared to conventional scaling methods, within a compute budget of 1e10 TFlops. Furthermore, we establish a more generalized and precise relationship among validation loss, any model size, and compute budget. This enables performance prediction for non-optimal model sizes, which may also be appealed under practical inference cost constraints, achieving a better trade-off. Yuanyang Yin, Mingwu Zheng, Jiarong Ou, Victor Shea-Jay Huang, Xin Tao 0001, Pengfei Wan 0001, Di Zhang 0026, Baoqun Yin, Wentao Zhang 0001, Kun Gai |
CVPR | 9 |
| 2025 | SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMsabstractYuanyang Yin, Yaqi Zhao, Yajie Zhang, Yuanxing Zhang, Ke Lin, Jiahao Wang, Xin Tao, Pengfei Wan, Wentao Zhang, Feng Zhao. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yuanyang Yin, Yuanxing Zhang, Xin Tao 0001, Pengfei Wan 0001, Wentao Zhang 0001 |
EMNLP | 7 |
| 2025 | How Far are AI-Generated Videos from Simulating the 3D Visual World: A Learned 3D Evaluation Approach
Chirui Chang, Jiahui Liu 0012, Zhengzhe Liu, Xiaoyang Lyu, Xin Tao 0001, Pengfei Wan 0001, Di Zhang 0026, Xiaojuan Qi 0001 |
ICCV | 6 |
| 2025 | Scene Graph Guided Generation: Enable Accurate Relations Generation in Text-to-Image Models via Textural Rectification
Guibao Shen, Luozhou Wang, Jiantao Lin, Wenhang Ge, Chaozhe Zhang, Xin Tao 0001, Di Zhang 0026, Pengfei Wan 0001, Guangyong Chen, Yijun Li 0001, Ying-Cong Chen |
ICCV | 6 |
| 2025 | Imbalance in Balance: Online Concept Balancing in Generation ModelsabstractIn visual generation tasks, the responses and combinations of complex concepts often lack stability and are error-prone, which remains an under-explored area. In this paper, we attempt to explore the causal factors for poor concept responses through elaborately designed experiments. We also design a concept-wise equalization loss function (IMBA loss) to address this issue. Our proposed method is online, eliminating the need for offline dataset processing, and requires minimal code changes. In our newly proposed complex concept benchmark Inert-CompBench and two other public test sets, our method significantly enhances the concept response capability of baseline models and yields highly competitive results with only a few codes released at https://github.com/KwaiVGI/IMBA-Loss. Yukai Shi, Jiarong Ou, Xin Tao 0001, Pengfei Wan 0001, Di Zhang 0026, Kun Gai |
ICCV | 6 |
| 2025 | BadVideo: Stealthy Backdoor Attack Against Text-to-Video GenerationabstractText-to-video (T2V) generative models have rapidly advanced and found widespread applications across fields like entertainment, education, and marketing. However, the adversarial vulnerabilities of these models remain rarely explored. We observe that in T2V generation tasks, the generated videos often contain substantial redundant information not explicitly specified in the text prompts, such as environmental elements, secondary objects, and additional details, providing opportunities for malicious attackers to embed hidden harmful content. Exploiting this inherent redundancy, we introduce BadVideo, the first backdoor attack framework tailored for T2V generation. Our attack focuses on designing target adversarial outputs through two key strategies: (1) Spatio-Temporal Composition, which combines different spatiotemporal features to encode malicious information; (2) Dynamic Element Transformation, which introduces transformations in redundant elements over time to convey malicious information. Based on these strategies, the attacker's malicious target seamlessly integrates with the user's textual instructions, providing high stealthiness. Moreover, by exploiting the temporal dimension of videos, our attack successfully evades traditional content moderation systems that primarily analyze spatial information within individual frames. Extensive experiments demonstrate that BadVideo achieves high attack success rates while preserving original semantics and maintaining excellent performance on clean inputs. Overall, our work reveals the adversarial vulnerability of T2V models, calling attention to potential risks and misuse. Our project page is at https://wrt2000.github.io/BadVideo2025/. Ruotong Wang 0008, Mingli Zhu, Jiarong Ou, Xin Tao 0001, Pengfei Wan 0001, Baoyuan Wu |
ICCV | 5 |
| 2025 | Stable Segment Anything ModelabstractThe Segment Anything Model (SAM) achieves remarkable promptable segmentation given high-quality prompts which, however, often require good skills to specify. To make SAM robust to casual prompts, this paper presents the first comprehensive analysis on SAM’s segmentation stability across a diverse spectrum of prompt qualities, notably imprecise bounding boxes and insufficient points. Our key finding reveals that given such low-quality prompts, SAM’s mask decoder tends to activate image features that are biased towards the background or confined to specific object parts. To mitigate this issue, our key idea consists of calibrating solely SAM’s mask attention by adjusting the sampling locations and amplitudes of image features, while the original SAM model architecture and weights remain unchanged. Consequently, our deformable sampling plugin (DSP) enables SAM to adaptively shift attention to the prompted target regions in a data-driven manner. During inference, dynamic routing plugin (DRP) is proposed that toggles SAM between the deformable and regular grid sampling modes, conditioned on the input prompt quality. Thus, our solution, termed Stable-SAM, offers several advantages: 1) improved SAM’s segmentation stability across a wide range of prompt qualities, while 2) retaining SAM’s powerful promptable segmentation efficiency and generality, with 3) minimal learnable parameters (0.08 M) and fast adaptation. Extensive experiments validate the effectiveness and advantages of our approach, underscoring Stable-SAM as a more robust solution for segmenting anything. Codes are at https://github.com/fanq15/Stable-SAM. Xin Tao 0001, Lei Ke, Mingqiao Ye, Di Zhang 0026, Pengfei Wan 0001, Yu-Wing Tai, Chi-Keung Tang |
ICLR | 2 |
| 2025 | Training-Free Efficient Video Generation via Dynamic Token CarvingabstractDespite the remarkable generation quality of video Diffusion Transformer (DiT) models, their practical deployment is severely hindered by extensive computational requirements. This inefficiency stems from two key challenges: the quadratic complexity of self-attention with respect to token length and the multi-step nature of diffusion models. To address these limitations, we present Jenga, a novel inference pipeline that combines dynamic attention carving with progressive resolution generation. Our approach leverages two key insights: (1) early denoising steps do not require high-resolution latents, and (2) later steps do not require dense attention. Jenga introduces a block-wise attention mechanism that dynamically selects relevant token interactions using 3D space-filling curves, alongside a progressive resolution strategy that gradually increases latent resolution during generation. Experimental results demonstrate that Jenga achieves substantial speedups across multiple state-of-the-art video diffusion models while maintaining comparable generation quality (8.83$\times$ speedup with 0.01\% performance drop on VBench). As a plug-and-play solution, Jenga enables practical, high-quality video generation on modern hardware by reducing inference time from minutes to seconds---without requiring model retraining. Yuechen Zhang, Jinbo Xing, Bin Xia 0014, Shaoteng Liu, Bohao Peng, Xin Tao 0001, Pengfei Wan 0001, Eric Lo 0001, Jiaya Jia |
NeurIPS | 6 |
| 2025 | VFRTok: Variable Frame Rates Video Tokenizer with Duration-Proportional Information AssumptionabstractModern video generation frameworks based on Latent Diffusion Models suffer from inefficiencies in tokenization due to the Frame-Proportional Information Assumption.
Existing tokenizers provide fixed temporal compression rates, causing the computational cost of the diffusion model to scale linearly with the frame rate.
The paper proposes the Duration-Proportional Information Assumption: the upper bound on the information capacity of a video is proportional to the duration rather than the number of frames.
Based on this insight, the paper introduces VFRTok, a Transformer-based video tokenizer, that enables variable frame rate encoding and decoding through asymmetric frame rate training between the encoder and decoder.
Furthermore, the paper proposes Partial Rotary Position Embeddings (RoPE) to decouple position and content modeling, which groups correlated patches into unified tokens.
The Partial RoPE effectively improves content-awareness, enhancing the video generation capability.
Benefiting from the compact and continuous spatio-temporal representation, VFRTok achieves competitive reconstruction quality and state-of-the-art generation fidelity while using only $1/8$ tokens compared to existing tokenizers. Tianxiong Zhong, Xingye Tian, Boyuan Jiang, Xuebo Wang, Xin Tao 0001, Pengfei Wan 0001 |
NeurIPS | 5 |
| 2025 | DVIS++: Improved Decoupled Framework for Universal Video SegmentationabstractWe present the Decoupled VIdeo Segmentation (DVIS) framework, a novel approach for the challenging task of universal video segmentation, including video instance segmentation (VIS), video semantic segmentation (VSS), and video panoptic segmentation (VPS). Unlike previous methods that model video segmentation in an end-to-end manner, our approach decouples video segmentation into three cascaded sub-tasks: segmentation, tracking, and refinement. This decoupling design allows for simpler and more effective modeling of the spatio-temporal representations of objects, especially in complex scenes and long videos. Accordingly, we introduce two novel components: the referring tracker and the temporal refiner. These components track objects frame by frame and model spatio-temporal representations based on pre-aligned features. To improve the tracking capability of DVIS, we propose a denoising training strategy and introduce contrastive learning, resulting in a more robust framework named DVIS++. The proposed decoupled framework efficiently handles universal and open-vocabulary object representations, allowing DVIS++ to conduct universal and open-vocabulary video segmentation. We conduct extensive experiments on six mainstream benchmarks, including the VIS, VSS, and VPS datasets. Using a unified architecture, DVIS++ significantly outperforms state-of-the-art specialized methods on these benchmarks in closed- and open-vocabulary settings. Tao Zhang 0042, Xingye Tian, Yikang Zhou, Shunping Ji, Xuebo Wang, Xin Tao 0001, Yuan Zhang 0020, Pengfei Wan 0001, Zhongyuan Wang 0006, Yu Wu 0011 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | Perception-Oriented Video Frame Interpolation via Asymmetric BlendingabstractPrevious methods for Video Frame Interpolation (VFI) have encountered challenges, notably the manifestation of blur and ghosting effects. These issues can be traced back to two pivotal factors: unavoidable motion errors and misalignment in supervision. In practice, motion estimates often prove to be error-prone, resulting in misaligned features. Furthermore, the reconstruction loss tends to bring blurry results, particularly in misaligned regions. To mitigate these challenges, we propose a new paradigm called PerVFI (Perception-oriented Video Frame Interpolation). Our approach incorporates an Asymmetric Synergistic Blending module (ASB) that utilizes features from both sides to synergistically blend intermediate features. One reference frame emphasizes primary content, while the other contributes complementary information. To impose a stringent constraint on the blending process, we introduce a self-learned sparse quasi-binary mask which effectively mitigates ghosting and blur artifacts in the output. Additionally, we employ a normalizing flow-based generator and utilize the negative log-likelihood loss to learn the conditional distribution of the output, which further facilitates the generation of clear and fine details. Experimental results validate the superiority of PerVFI, demonstrating significant improvements in perceptual quality compared to existing methods. Codes are available at https://github.com/mulns/PerVFI Guangyang Wu, Xin Tao 0001, Wenyi Wang 0005, Xiaohong Liu 0001, Qingqing Zheng |
CVPR | 2 |
| 2024 | VideoTetris: Towards Compositional Text-to-Video GenerationabstractDiffusion models have demonstrated great success in text-to-video (T2V) generation. However, existing methods may face challenges when handling complex (long) video generation scenarios that involve multiple objects or dynamic changes in object numbers. To address these limitations, we propose VideoTetris, a novel framework that enables compositional T2V generation. Specifically, we propose spatio-temporal compositional diffusion to precisely follow complex textual semantics by manipulating and composing the attention maps of denoising networks spatially and temporally. Moreover, we propose a new dynamic-aware data processing pipeline and a consistency regularization method to enhance the consistency of auto-regressive video generation. Extensive experiments demonstrate that our VideoTetris achieves impressive qualitative and quantitative results in compositional T2V generation. Code is available at: https://github.com/YangLing0818/VideoTetris Ling Yang 0006, Yuan Gao 0015, Yufan Deng, Xintao Wang 0002, Zhaochen Yu, Xin Tao 0001, Pengfei Wan 0001, Di Zhang 0026, Bin Cui 0001 |
NeurIPS | 8 |
| 2023 | Compression-Aware Video Super-ResolutionabstractVideos stored on mobile devices or delivered on the Internet are usually in compressed format and are of various unknown compression parameters, but most video super-resolution (VSR) methods often assume ideal inputs resulting in large performance gap between experimental settings and real-world applications. In spite of a few pioneering works being proposed recently to super-resolve the compressed videos, they are not specially designed to deal with videos of various levels of compression. In this paper, we propose a novel and practical compression-aware video super-resolution model, which could adapt its video enhancement process to the estimated compression level. A compression encoder is designed to model compression levels of input frames, and a base VSR model is then conditioned on the implicitly computed representation by inserting compression-aware modules. In addition, we propose to further strengthen the VSR model by taking full advantage of meta data that is embedded naturally in compressed video streams in the procedure of information fusion. Extensive experiments are conducted to demonstrate the effectiveness and efficiency of the proposed method on compressed VSR benchmarks. The codes will be available at https://github.com/aprBlue/CAVSR Takashi Isobe, Xu Jia 0012, Xin Tao 0001, Huchuan Lu, Yu-Wing Tai |
CVPR | 4 |
| 2023 | Scene-Generalizable Interactive Segmentation of Radiance FieldsabstractExisting methods for interactive segmentation in radiance fields entail scene-specific optimization and thus cannot generalize across different scenes, which greatly limits their applicability. In this work we make the first attempt at Scene-Generalizable Interactive Segmentation in Radiance Fields (SGISRF) and propose a novel SGISRF method, which can perform 3D object segmentation for novel (unseen) scenes represented by radiance fields, guided by only a few interactive user clicks in a given set of multi-view 2D images. In particular, the proposed SGISRF focuses on addressing three crucial challenges with three specially designed techniques. First, we devise the Cross-Dimension Guidance Propagation to encode the scarce 2D user clicks into informative 3D guidance representations. Second, the Uncertainty-Eliminated 3D Segmentation module is designed to achieve efficient yet effective 3D segmentation. Third, Concealment-Revealed Supervised Learning scheme is proposed to reveal and correct the concealed 3D segmentation errors resulted from the supervision in 2D space with only 2D mask annotations. Extensive experiments on two real-world challenging benchmarks covering diverse scenes demonstrate 1) effectiveness and scene-generalizability of the proposed method, 2) favorable performance compared to classical method requiring scene-specific optimization. Songlin Tang, Wenjie Pei, Xin Tao 0001, Tanghui Jia, Guangming Lu 0002, Yu-Wing Tai |
ACM Multimedia | 3 |
| 2023 | Feature Decoupling-Recycling Network for Fast Interactive SegmentationabstractRecent interactive segmentation methods iteratively take source image, user guidance and previously predicted mask as the input without considering the invariant nature of the source image. As a result, the process of extracting features from the source image is repeated in each interaction, resulting in substantial computational redundancy. In this work, we propose the Feature Decoupling-Recycling Network (FDRN), which decouples the modeling components based on their intrinsic discrepancies and then recycles components that can be reused for each user interaction. Thus, the efficiency of the whole interactive process can be significantly improved. To be specific, we apply the Decoupling-Recycling strategy from three perspectives to address three types of discrepancies, respectively. First, our model decouples the learning of source image semantics from the encoding of user guidance to process two types of input domains separately. Second, FDRN decouples high-level and low-level features from stratified semantic representations to enhance feature learning. Third, during the encoding of user guidance, current user guidance is decoupled from historical guidance to highlight the effect of current user guidance. We conduct extensive experiments on 6 datasets from different domains and modalities, which demonstrate the following merits of our model: 1) superior efficiency than other methods, particularly advantageous in the challenging scenarios requiring long-term interactions (up to 4.25x faster), while achieving favorable segmentation performance; 2) strong applicability to various methods serving as a universal enhancement technique; 3) well cross-task generalizability, e.g., to medical image segmentation, and robustness against misleading user guidance. Weinong Wang, Xin Tao 0001, Zhiwei Xiong, Yu-Wing Tai, Wenjie Pei |
ACM Multimedia | 3 |
| 2022 | Look Back and Forth: Video Super-Resolution with Explicit Temporal Difference ModelingabstractTemporal modeling is crucial for video super-resolution. Most of the video super-resolution methods adopt the optical flow or deformable convolution for explicitly motion compensation. However, such temporal modeling techniques increase the model complexity and might fail in case of occlusion or complex motion, resulting in serious distortion and artifacts. In this paper, we propose to explore the role of explicit temporal difference modeling in both LR and HR space. Instead of directly feeding consecutive frames into a VSR model, we propose to compute the temporal difference between frames and divide those pixels into two subsets according to the level of difference. They are separately processed with two branches of different receptive fields in order to better extract complementary information. To further enhance the super-resolution result, not only spatial residual features are extracted, but the difference between consecutive frames in high-frequency domain is also computed. It allows the model to exploit intermediate SR results in both future and past to refine the current SR output. The difference at different time steps could be cached such that information from further distance in time could be propagated to the current frame for refinement. Experiments on several video super-resolution benchmark datasets demonstrate the effectiveness of the proposed method and its favorable performance against state-of-the-art methods. Takashi Isobe, Xu Jia 0012, Xin Tao 0001, Ruihuang Li, Yongjie Shi, Huchuan Lu, Yu-Wing Tai |
CVPR | 3 |
| 2022 | DeViT: Deformed Vision Transformers in Video InpaintingabstractThis paper presents a novel video inpainting architecture named Deformed Vision Transformers (DeViT). We make three significant contributions to this task: First, we extended previous Transformers with patch alignment by introducing Deformed Patch-based Homography Estimator (DePtH), which enriches the patch-level feature alignments in key and query with additional offsets learned from patch pairs without additional supervision. DePtH enables our method to handle challenging scenes or agile motion with in-plane or out-of-plane deformation, which previous methods usually fail. Second, we introduce the Mask Pruning-based Patch Attention (MPPA) to improve the standard patch-wised feature matching by pruning out less essential features and considering the saliency map. MPPA enhances the matching accuracy between warped tokens with invalid pixels. Third, we introduce the Spatial-Temporal weighting Adaptor (STA) module to assign more accurate attention to spatial-temporal tokens under the guidance of the Deformation Factor learned from DePtH, especially for videos with agile motions. Experimental results demonstrate that our method outperforms previous state-of-the-art methods in quality and quantity and achieves a new state-of-the-art for video inpainting. Jiayin Cai, Xin Tao 0001, Chun Yuan 0003, Yu-Wing Tai |
ACM Multimedia | 3 |
| 2022 | Text-Guided Human Image Manipulation via Image-Text Shared SpaceabstractText is a new way to guide human image manipulation. Albeit natural and flexible, text usually suffers from inaccuracy in spatial description, ambiguity in the description of appearance, and incompleteness. We in this paper address these issues. To overcome inaccuracy, we use structured information (e.g., poses) to help identify correct location to manipulate, by disentangling the control of appearance and spatial structure. Moreover, we learn the image-text shared space with derived disentanglement to improve accuracy and quality of manipulation, by separating relevant and irrelevant editing directions for the textual instructions in this space. Our model generates a series of manipulation results by moving source images in this space with different degrees of editing strength. Thus, to reduce the ambiguity in text, our model generates sequential output for manual selection. In addition, we propose an efficient pseudo-label loss to enhance editing performance when the text is incomplete. We evaluate our method on various datasets and show its precision and interactiveness to manipulate human images. Xiaogang Xu 0002, Ying-Cong Chen, Xin Tao 0001, Jiaya Jia |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | MASA-SR: Matching Acceleration and Spatial Adaptation for Reference-Based Image Super-ResolutionabstractReference-based image super-resolution (RefSR) has shown promising success in recovering high-frequency details by utilizing an external reference image (Ref). In this task, texture details are transferred from the Ref image to the low-resolution (LR) image according to their point- or patch-wise correspondence. Therefore, high-quality correspondence matching is critical. It is also desired to be computationally efficient. Besides, existing RefSR methods tend to ignore the potential large disparity in distributions between the LR and Ref images, which hurts the effectiveness of the information utilization. In this paper, we propose the MASA network for RefSR, where two novel modules are designed to address these problems. The proposed Match & Extraction Module significantly reduces the computational cost by a coarse-to-fine correspondence matching scheme. The Spatial Adaptation Module learns the difference of distribution between the LR and Ref images, and remaps the distribution of Ref features to that of LR features in a spatially adaptive way. This scheme makes the network robust to handle different reference images. Extensive quantitative and qualitative experiments validate the effectiveness of our proposed model. Liying Lu, Wenbo Li 0002, Xin Tao 0001, Jiangbo Lu, Jiaya Jia |
CVPR | 3 |
| 2020 | MuCAN: Multi-correspondence Aggregation Network for Video Super-Resolution
Wenbo Li 0002, Xin Tao 0001, Taian Guo, Lu Qi 0001, Jiangbo Lu, Jiaya Jia |
ECCV (10) | 2 |
| 2020 | VCNet: A Robust Approach to Blind Image Inpainting
Yi Wang 0074, Ying-Cong Chen, Xin Tao 0001, Jiaya Jia |
ECCV (25) | 3 |
| 2020 | Particularity Beyond Commonality: Unpaired Identity Transfer with Multiple References
Ruizheng Wu, Xin Tao 0001, Ying-Cong Chen, Xiaoyong Shen, Jiaya Jia |
ECCV (4) | 2 |
| 2019 | Dynamic Scene Deblurring With Parameter Selective Sharing and Nested Skip ConnectionsabstractDynamic Scene deblurring is a challenging low-level vision task where spatially variant blur is caused by many factors, e.g., camera shake and object motion. Recent study has made significant progress. Compared with the parameter independence scheme [19] and parameter sharing scheme [33], we develop the general principle for constraining the deblurring network structure by proposing the generic and effective selective sharing scheme. Inside the subnetwork of each scale, we propose a nested skip connection structure for the nonlinear transformation modules to replace stacked convolution layers or residual blocks. Besides, we build a new large dataset of blurred/sharp image pairs towards better restoration quality. Comprehensive experimental results show that our parameter selective sharing scheme, nested skip connection structure, and the new dataset are all significant to set a new state-of-the-art in dynamic scene deblurring. Hongyun Gao 0001, Xin Tao 0001, Xiaoyong Shen, Jiaya Jia |
CVPR | 2 |
| 2019 | Wide-Context Semantic Image ExtrapolationabstractThis paper studies the fundamental problem of extrapolating visual context using deep generative models, i.e., extending image borders with plausible structure and details. This seemingly easy task actually faces many crucial technical challenges and has its unique properties. The two major issues are size expansion and one-side constraints. We propose a semantic regeneration network with several special contributions and use multiple spatial related losses to address these issues. Our results contain consistent structures and high-quality textures. Extensive experiments are conducted on various possible alternatives and related methods. We also explore the potential of our method for various interesting applications that can benefit research in a variety of fields. Yi Wang 0074, Xin Tao 0001, Xiaoyong Shen, Jiaya Jia |
CVPR | 2 |
| 2019 | Attribute-Driven Spontaneous Motion in Unpaired Image TranslationabstractCurrent image translation methods, albeit effective to produce high-quality results in various applications, still do not consider much geometric transform. We in this paper propose the spontaneous motion estimation module, along with a refinement part, to learn attribute-driven deformation between source and target domains. Extensive experiments and visualization demonstrate effectiveness of these modules. We achieve promising results in unpaired-image translation tasks, and enable interesting applications based on spontaneous motion. Ruizheng Wu, Xin Tao 0001, Xiaoyong Shen, Jiaya Jia |
ICCV | 2 |
| 2018 | Facelet-Bank for Fast Portrait ManipulationabstractDigital face manipulation has become a popular and fascinating way to touch images with the prevalence of smart phones and social networks. With a wide variety of user preferences, facial expressions, and accessories, a general and flexible model is necessary to accommodate different types of facial editing. In this paper, we propose a model to achieve this goal based on an end-to-end convolutional neural network that supports fast inference, edit-effect control, and quick partial-model update. In addition, this model learns from unpaired image sets with different attributes. Experimental results show that our framework can handle a wide range of expressions, accessories, and makeup effects. It produces high-resolution and high-quality results in fast speed. Ying-Cong Chen, Huaijia Lin, Michelle Shu, Ruiyu Li, Xin Tao 0001, Xiaoyong Shen, Yangang Ye, Jiaya Jia |
CVPR | 5 |
| 2018 | Scale-Recurrent Network for Deep Image DeblurringabstractIn single image deblurring, the "coarse-to-fine" scheme, i.e. gradually restoring the sharp image on different resolutions in a pyramid, is very successful in both traditional optimization-based methods and recent neural-network-based approaches. In this paper, we investigate this strategy and propose a Scale-recurrent Network (SRN-DeblurNet) for this deblurring task. Compared with the many recent learning-based approaches in [25], it has a simpler network structure, a smaller number of parameters and is easier to train. We evaluate our method on large-scale deblurring datasets with complex motion. Results show that our method can produce better quality results than state-of-the-arts, both quantitatively and qualitatively. Xin Tao 0001, Hongyun Gao 0001, Xiaoyong Shen, Jue Wang 0001, Jiaya Jia |
CVPR | 1 |
| 2018 | Image Inpainting via Generative Multi-column Convolutional Neural NetworksabstractIn this paper, we propose a generative multi-column network for image inpainting. This network synthesizes different image components in a parallel manner within one stage. To better characterize global structures, we design a confidence-driven reconstruction loss while an implicit diversified MRF regularization is adopted to enhance local details. The multi-column network combined with the reconstruction and MRF loss propagates local and global information derived from context to the target inpainting regions. Extensive experiments on challenging street view, face, natural objects and scenes manifest that our method produces visual compelling results even without previously common post-processing. Yi Wang 0074, Xin Tao 0001, Xiaojuan Qi 0001, Xiaoyong Shen, Jiaya Jia |
NeurIPS | 2 |
| 2017 | High-Quality Correspondence and Segmentation Estimation for Dual-Lens Smart-Phone PortraitsabstractEstimating correspondence between two images and extracting the foreground object are two challenges in computer vision. With dual-lens smart phones, such as iPhone 7+ and Huawei P9, coming into the market, two images of slightly different views provide us new information to unify the two topics. We propose a joint method to tackle them simultaneously via a joint fully connected conditional random field (CRF) framework. The regional correspondence is used to handle textureless regions in matching and make our CRF system computationally efficient. Our method is evaluated over 2,000 new image pairs, and produces promising results on challenging portrait images. Xiaoyong Shen, Hongyun Gao 0001, Xin Tao 0001, Chao Zhou 0001, Jiaya Jia |
ICCV | 3 |
| 2017 | Detail-Revealing Deep Video Super-ResolutionabstractPrevious CNN-based video super-resolution approaches need to align multiple frames to the reference. In this paper, we show that proper frame alignment and motion compensation is crucial for achieving high quality results. We accordingly propose a “sub-pixel motion compensation” (SPMC) layer in a CNN framework. Analysis and experiments show the suitability of this layer in video SR. The final end-to-end, scalable CNN framework effectively incorporates the SPMC layer and fuses multiple frames to reveal image details. Our implementation can generate visually and quantitatively high-quality results, superior to current state-of-the-arts, without the need of parameter tuning. Xin Tao 0001, Hongyun Gao 0001, Renjie Liao 0001, Jue Wang 0001, Jiaya Jia |
ICCV | 1 |
| 2017 | Zero-Order Reverse FilteringabstractIn this paper, we study an unconventional but practically meaningful reversibility problem of commonly used image filters. We broadly define filters as operations to smooth images or to produce layers via global or local algorithms. And we raise the intriguingly problem if they are reservable to the status before filtering. To answer it, we present a novel strategy to understand general filter via contraction mappings on a metric space. A very simple yet effective zero-order algorithm is proposed. It is able to practically reverse most filters with low computational cost. We present quite a few experiments in the paper and supplementary file to thoroughly verify its performance. This method can also be generalized to solve other inverse problems and enables new applications. Xin Tao 0001, Chao Zhou 0001, Xiaoyong Shen, Jue Wang 0001, Jiaya Jia |
ICCV | 1 |
| 2016 | Deep Automatic Portrait Matting
Xiaoyong Shen, Xin Tao 0001, Hongyun Gao 0001, Chao Zhou 0001, Jiaya Jia |
ECCV (1) | 2 |
| 2016 | Regional foremost matching for internet scene imagesabstractWe analyze the dense matching problem for Internet scene images based on the fact that commonly only part of images can be matched due to the variation of view angle, motion, objects, etc. We thus proposeregional foremost matchingto reject outlier matching points while still producing dense high-quality correspondence in the remaining foremost regions. Our system initializes sparse correspondence, propagates matching with model fitting and optimization, and detects foremost regions robustly. We apply our method to several applications, including time-lapse sequence generation, Internet photo composition, automatic image morphing, and automatic rephotography. Xiaoyong Shen, Xin Tao 0001, Chao Zhou 0001, Hongyun Gao 0001, Jiaya Jia |
ACM Trans. Graph. | 2 |
| 2015 | Handling motion blur in multi-frame super-resolutionabstractUbiquitous motion blur easily fails multi-frame super-resolution (MFSR). Our method proposed in this paper tackles this issue by optimally searching least blurred pixels in MFSR. An EM framework is proposed to guide residual blur estimation and high-resolution image reconstruction. To suppress noise, we employ a family of sparse penalties as natural image priors, along with an effective solver. Theoretical analysis is performed on how and when our method works. The relationship between estimation errors of motion blur and the quality of input images is discussed. Our method produces sharp and higher-resolution results given input of challenging low-resolution noisy and blurred sequences. Ziyang Ma 0002, Renjie Liao 0001, Xin Tao 0001, Li Xu 0001, Jiaya Jia, Enhua Wu |
CVPR | 3 |
| 2015 | Video Super-Resolution via Deep Draft-Ensemble LearningabstractWe propose a new direction for fast video super-resolution (VideoSR) via a SR draft ensemble, which is defined as the set of high-resolution patch candidates before final image deconvolution. Our method contains two main components -- i.e., SR draft ensemble generation and its optimal reconstruction. The first component is to renovate traditional feedforward reconstruction pipeline and greatly enhance its ability to compute different super resolution results considering large motion variation and possible errors arising in this process. Then we combine SR drafts through the nonlinear process in a deep convolutional neural network (CNN). We analyze why this framework is proposed and explain its unique advantages compared to previous iterative methods to update different modules in passes. Promising experimental results are shown on natural video sequences. Renjie Liao 0001, Xin Tao 0001, Ruiyu Li, Ziyang Ma 0002, Jiaya Jia |
ICCV | 2 |
| 2015 | Break Ames room illusion: depth from general single imagesabstractPhotos compress 3D visual data to 2D. However, it is still possible to infer depth information even without sophisticated object learning. We propose a solution based on small-scale defocus blur inherent in optical lens and tackle the estimation problem by proposing a non-parametric matching scheme for natural images. It incorporates a matching prior with our newly constructed edgelet dataset using a non-local scheme, and includes semantic depth order cues for physically based inference. Several applications are enabled on natural images, including geometry based rendering and editing. Jianping Shi, Xin Tao 0001, Li Xu 0001, Jiaya Jia |
ACM Trans. Graph. | 2 |
| 2014 | Inverse Kernels for Fast Spatial Deconvolution
Li Xu 0001, Xin Tao 0001, Jiaya Jia |
ECCV (5) | 2 |