VLDB 2026 Research / reviewers in the wild / expert
Ying Shan
dblp:68/5910
· DBLP profile ↗
230ranked-venue papers
16as first author
198since 2021 · last 2026
0000-0001-7673-8325ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 183 · 15 first-author · 156 since 2021Graphics, computer vision, multimedia, augmented reality and games · 170 · 8 first-author · 152 since 2021Human-computer interaction and ubiquitous computing · 7 · 6 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MMhops-R1: Multimodal Multi-hop ReasoningabstractThe ability to perform multi-modal multi-hop reasoning by iteratively integrating information across various modalities and external knowledge is critical for addressing complex real-world challenges. However, existing Multi-modal Large Language Models (MLLMs) are predominantly limited to single-step reasoning, as existing benchmarks lack the complexity needed to evaluate and drive multi-hop abilities. To bridge this gap, we introduce MMhops, a novel, large-scale benchmark designed to systematically evaluate and foster multi-modal multi-hop reasoning. MMhops dataset comprises two challenging task formats, Bridging and Comparison, which necessitate that models dynamically construct complex reasoning chains by integrating external knowledge. To tackle the challenges posed by MMhops, we propose MMhops-R1, a novel multi-modal Retrieval-Augmented Generation (mRAG) framework for dynamic reasoning. Our framework utilizes reinforcement learning to optimize the model for autonomously planning reasoning paths, formulating targeted queries, and synthesizing multi-level information. Comprehensive experiments demonstrate that MMhops-R1 significantly outperforms strong baselines on MMhops, highlighting that dynamic planning and multi-modal knowledge integration are crucial for complex reasoning. Moreover, MMhops-R1 demonstrates strong generalization to tasks requiring fixed-hop reasoning, underscoring the robustness of our dynamic planning approach. Ziqi Zhang 0010, Zongyang Ma, Bing Li 0001, Chunfeng Yuan, Guangting Wang, Fengyun Rao, Ying Shan, Weiming Hu 0004 |
AAAI | 9 |
| 2026 | MuMu-LLaMA: Multi-modal music understanding and generation via large language models
Shansong Liu, Atin Sakkeer Hussain, Ying Shan, Chenshuo Sun |
Expert Syst. Appl. | 4 |
| 2026 | EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level PlanningabstractAbstract The pursuit of artificial general intelligence (AGI) has been accelerated by Multimodal Large Language Models (MLLMs), which exhibit superior reasoning, generalization capabilities, and proficiency in processing multimodal inputs. A crucial milestone in the evolution of AGI is the attainment of human-level planning, a fundamental ability for making informed decisions in complex environments, and solving a wide range of real-world problems. Despite the impressive advancements in MLLMs, a question remains: How far are current MLLMs from achieving human-level planning? To shed light on this question, we introduce EgoPlan-Bench, a comprehensive benchmark to evaluate the planning abilities of MLLMs in real-world scenarios from an egocentric perspective, mirroring human perception. EgoPlan-Bench emphasizes the evaluation of planning capabilities of MLLMs, featuring realistic tasks, diverse action plans, and intricate visual observations. Our rigorous evaluation of a wide range of MLLMs reveals that EgoPlan-Bench poses significant challenges, highlighting a substantial scope for improvement in MLLMs to achieve human-level task planning. To facilitate this advancement, we further present EgoPlan-IT, a specialized instruction-tuning dataset that effectively enhances model performance on EgoPlan-Bench. We have made all the codes, data, and a maintained benchmark leaderboard available at https://chenyi99.github.io/ego_plan/ to advance future research. Yi Chen 0019, Yuying Ge, Yixiao Ge, Mingyu Ding, Bohao Li 0002, Rui Wang 0092, Ruifeng Xu 0001, Ying Shan, Xihui Liu |
Int. J. Comput. Vis. | 8 |
| 2026 | EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World ScenariosabstractAbstract Multimodal Large Language Models (MLLMs) has recently demonstrated superior multimodal comprehension abilities, heralding a new era for artificial general intelligence (AGI). However, achieving AGI necessitates more than just comprehension. A crucial capability required is effective planning in diverse scenarios, which involves making reasonable decisions based on complex environments to solve real-world problems. Despite its importance, the planning abilities of current MLLMs in varied scenarios remain underexplored, leaving a significant gap in our understanding of their full potential. In this paper, we introduce EgoPlan-Bench2, a rigorous and comprehensive benchmark designed to assess the planning capabilities of MLLMs across a wide range of real-world scenarios . EgoPlan-Bench2 encompasses everyday tasks spanning 4 major domains and 24 detailed scenarios, closely aligned with human daily life. It is constructed through a semi-automatic process utilizing egocentric videos, complemented by manual verification. Grounded in a first-person perspective, it mirrors the way humans approach problem-solving in everyday life. We evaluate 25 competitive MLLMs and provide an in-depth analysis of their limitations, revealing that they face significant challenges in real-world planning. To diagnose the underlying bottlenecks, we investigate the effectiveness of various prompts via a training-free multimodal prompting method. We find that MLLMs’ planning performance on EgoPlan-Bench2 is critically dependent on temporally structured action sequences in historical task progress and interactions between objects and humans in current observation state. This dependency also underscores the necessity for strong reasoning abilities to integrate diverse multimodal cues and analysis before making final decision. Building on this insight, we demonstrate that EgoPlan-Bench2 is also an effective video reasoning benchmark. Experiments with Gemini-2.5-Flash and a post-trained Qwen-2.5-VL confirm its ability to distinguish between models with and without explicit deliberate reasoning mechanisms, showcasing the tangible impact of DeepSeek-R1 paradigm reasoning on planning tasks. We have made data and code available at https://qiulu66.github.io/egoplanbench2/ . Lu Qiu, Yi Chen 0019, Yuying Ge, Yixiao Ge, Ying Shan, Xihui Liu |
Int. J. Comput. Vis. | 5 |
| 2025 | Image Conductor: Precision Control for Interactive Video SynthesisabstractFilmmaking and animation production often require sophisticated techniques for coordinating camera transitions and object movements, typically involving labor-intensive real-world capturing. Despite advancements in generative AI for video creation, achieving precise control over motion for interactive video asset generation remains challenging. To this end, we propose Image Conductor, a method for precise control of camera transitions and object movements to generate video assets from a single image. An well-cultivated training strategy is proposed to separate distinct camera and object motion by camera LoRA weights and object LoRA weights. To further eliminate motion ambiguity from ill-posed trajectories, we introduce a camera-free guidance technique during inference process, enhancing object movements while eliminating camera transitions. Additionally, we develop a trajectory-oriented video motion data curation pipeline for training. Quantitative and qualitative experiments demonstrate our method's precision and fine-grained control in generating motion-controllable videos from images, advancing the practical application of interactive video synthesis. Yaowei Li 0001, Xintao Wang 0002, Zhaoyang Zhang 0004, Zhouxia Wang, Ziyang Yuan, Liangbin Xie, Ying Shan, Yuexian Zou |
AAAI | 7 |
| 2025 | CustomCrafter: Customized Video Generation with Preserving Motion and Concept Composition AbilitiesabstractCustomized video generation aims to generate high-quality videos guided by text prompts and subject's reference images. However, since it is only trained on static images, the fine-tuning process of subject learning disrupts abilities of video diffusion models (VDMs) to combine concepts and generate motions. To restore these abilities, some methods use additional video similar to the prompt to fine-tune or guide the model. This requires frequent changes of guiding videos and even re-tuning of the model when generating different motions, which is very inconvenient for users. In this paper, we propose CustomCrafter, a novel framework that preserves the model's motion generation and conceptual combination abilities without additional video and fine-tuning to recovery. For preserving conceptual combination ability, we design a plug-and-play module to update few parameters in VDMs, enhancing the model's ability to capture the appearance details and the ability of concept combinations for new subjects. For motion generation, we observed that VDMs tend to restore the motion of video in the early stage of denoising, while focusing on the recovery of subject details in the later stage. Therefore, we propose Dynamic Weighted Video Sampling Strategy. Using the pluggability of our subject learning modules, we reduce the impact of this module on motion generation in the early stage of denoising, preserving the ability to generate motion of VDMs. In the later stage of denoising, we restore this module to repair the appearance details of the specified subject, thereby ensuring the fidelity of the subject's appearance. Experimental results show that our method has a significant improvement compared to previous methods. Yong Zhang 0034, Xintao Wang 0002, Xianpan Zhou, Guangcong Zheng, Zhongang Qi, Ying Shan, Xi Li 0001 |
AAAI | 7 |
| 2025 | High-Accuracy Fractured Object Reassembly Under Arbitrary Poses
Qun-Ce Xu, Yan-Pei Cao 0001, Weihao Cheng 0002, Tai-Jiang Mu, Ying Shan, Yongliang Yang 0002, Shi-Min Hu 0001 |
CVM (2) | 5 |
| 2025 | DepthCrafter: Generating Consistent Long Depth Sequences for Open-world VideosabstractEstimating video depth in open-world scenarios is challenging due to the diversity of videos in appearance, content motion, camera movement, and length. We present DepthCrafter, an innovative method for generating temporally consistent long depth sequences with intricate details for open-world videos, without requiring any supplementary information such as camera poses or optical flow. The generalization ability to open-world videos is achieved by training the video-to-depth model from a pretrained image-to-video diffusion model, through our meticulously designed three-stage training strategy. Our training approach enables the model to generate depth sequences with variable lengths at one time, up to 110 frames, and harvest both precise depth details and rich content diversity from realistic and synthetic datasets. We also propose an inference strategy that can process extremely long videos through segment-wise estimation and seamless stitching. Comprehensive evaluations on multiple datasets reveal that DepthCrafter achieves state-of-the-art performance in open-world video depth estimation under zero-shot settings. Furthermore, DepthCrafter facilitates various downstream applications, including depth-based visual effects and conditional video generation. Wenbo Hu 0002, Xiangjun Gao, Xiaoyu Li 0002, Sijie Zhao, Xiaodong Cun, Yong Zhang 0034, Long Quan, Ying Shan |
CVPR | 8 |
| 2025 | DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video GenerationabstractSora-like video generation models have achieved remarkable progress with a Multi-Modal Diffusion Transformer (MM-DiT) architecture. However, the current video generation models predominantly focus on single-prompt, struggling to generate coherent scenes with multiple sequential prompts that better reflect real-world dynamic scenarios. While some pioneering works have explored multi-prompt video generation, they face significant challenges including strict training data requirements, weak prompt following, and unnatural transitions. To address these problems, we propose DiTCtrl, a training-free multi-prompt video generation method under MM-DiT architectures for the first time. Our key idea is to take the multi-prompt video generation task as temporal video editing with smooth transitions. To achieve this goal, we first analyze MM-DiT’s attention mechanism, finding that the 3D full attention behaves similarly to that of the cross/self-attention blocks in the UNet-like diffusion models, enabling mask-guided precise semantic control across different prompts with attention sharing for multi-prompt video generation. Based on our careful design, the video generated by DiTCtrl achieves smooth transitions and consistent object motion given multiple sequential prompts without additional training. Besides, we also present MPVBench, a new benchmark specially designed for multi-prompt video generation to evaluate the performance of multi-prompt generation. Extensive experiments demonstrate that our method achieves state-of-the-art performance without additional training. Code is available at https://github.com/TencentARC/DiTCtrl. Minghong Cai, Xiaodong Cun, Xiaoyu Li 0002, Wenze Liu, Zhaoyang Zhang 0004, Yong Zhang 0034, Ying Shan, Xiangyu Yue 0001 |
CVPR | 7 |
| 2025 | Mani-GS: Gaussian Splatting Manipulation with Triangular MeshabstractNeural 3D representations, such as Neural Radiation Fields (NeRF), excel at producing photorealistic rendering results but lack the flexibility for manipulation and editing which is crucial for content creation. However, manipulating NeRF is not highly controllable and requires a long training and inference time. With the emergence of 3D Gaussian Splatting (3DGS), extremely high-fidelity novel view synthesis can be achieved using an explicit point-based 3D representation with much faster training and rendering speed. However, there is still a lack of effective means to manipulate 3DGS freely while maintaining rendering quality. In this work, we aim to tackle the challenge of achieving manipulable photo-realistic rendering. We propose to utilize a triangular mesh to manipulate 3DGS directly with self-adaptation. This approach reduces the need to design various algorithms for different types of 3DGS manipulation. By utilizing a triangle shape-aware Gaussian binding and adapting method, we can achieve 3DGS manipulation and preserve high-fidelity rendering. In addition, our method is also effective with inaccurate meshes extracted from 3DGS. Experiments demonstrate our method’s effectiveness and superiority over baseline approaches. Xiangjun Gao, Xiaoyu Li 0002, Yiyu Zhuang, Qi Zhang 0029, Wenbo Hu 0002, Chaopeng Zhang, Yao Yao 0008, Ying Shan, Long Quan |
CVPR | 8 |
| 2025 | Divot: Diffusion Powers Video Tokenizer for Comprehension and GenerationabstractIn recent years, there has been a significant surge of interest in unifying image comprehension and generation within Large Language Models (LLMs). This growing interest has prompted us to explore extending this unification to videos. The core challenge lies in developing a versatile video tokenizer that captures both the spatial characteristics and temporal dynamics of videos to obtain representations for LLMs, and the representations can be further decoded into realistic video clips to enable video generation. In this work, we introduce Divot, a Diffusion-Powered Video Tokenizer, which leverages the diffusion process for self-supervised video representation learning. We posit that if a video diffusion model can effectively de-noise video clips by taking the features of a video tokenizer as the condition, then the tokenizer has successfully captured robust spatial and temporal information. Additionally, the video diffusion model inherently functions as a de-tokenizer, decoding videos from their representations. Building upon the Divot tokenizer, we present Divot-LLM through video-to-text auto-regression and text-to-video generation by modeling the distributions of continuous-valued Divot features with a Gaussian Mixture Model. Experimental results demonstrate that our diffusion-based video tokenizer, when integrated with a pre-trained LLM, achieves competitive performance across various video comprehension and generation benchmarks. The instruction tuned Divot-LLM also excels in video storytelling, generating interleaved narratives and corresponding videos. Models and codes are available at https://github.com/TencentARC/Divot. Yuying Ge, Yizhuo Li 0001, Yixiao Ge, Ying Shan |
CVPR | 4 |
| 2025 | NVComposer: Boosting Generative Novel View Synthesis with Multiple Sparse and Unposed ImagesabstractRecent advancements in generative models have significantly improved novel view synthesis (NVS) from multi-view data. However, existing methods depend on external multiview alignment processes, such as explicit pose estimation or pre-reconstruction, which limits their flexibility and accessibility, especially when alignment is unstable due to insufficient overlap or occlusions between views. In this paper, we propose NVComposer, a novel approach that eliminates the need for explicit external alignment. NVComposer enables the generative model to implicitly infer spatial and geometric relationships between multiple conditional views by introducing two key components: 1) an image-pose dualstream diffusion model that simultaneously generates target novel views and condition camera poses, and 2) a geometry-aware feature alignment module that distills geometric priors from dense stereo models during training. Extensive experiments demonstrate that NVComposer achieves state-of-the-art performance in generative multi-view NVS tasks, removing the reliance on external alignment and thus improving model accessibility. Our approach shows substantial improvements in synthesis quality as the number of unposed input views increases, highlighting its potential for more flexible and accessible generative NVS systems. Lingen Li, Zhaoyang Zhang 0004, Yaowei Li 0001, Wenbo Hu 0002, Xiaoyu Li 0002, Weihao Cheng 0002, Jinwei Gu, Tianfan Xue, Ying Shan |
CVPR | 10 |
| 2025 | Mono2Stereo: A Benchmark and Empirical Study for Stereo ConversionabstractWith the rapid proliferation of 3D devices and the shortage of 3D content, stereo conversion is attracting increasing attention. Recent works introduce pretrained Diffusion Models (DMs) into this task. However, due to the scarcity of large-scale training data and comprehensive benchmarks, the optimal methodologies for employing DMs in stereo conversion and the accurate evaluation of stereo effects remain largely unexplored. In this work, we introduce the Mono2Stereo dataset, providing high-quality training data and benchmark to support in-depth exploration of stereo conversion. With this dataset, we conduct an empirical study that yields two primary findings. 1) The differences between the left and right views are subtle, yet existing metrics consider overall pixels, failing to concentrate on regions critical to stereo effects. 2) Mainstream methods adopt either one-stage left-to-right generation or warp-and-inpaint pipeline, facing challenges of degraded stereo effect and image distortion respectively. Based on these findings, we introduce a new evaluation metric, Stereo Intersection-over-Union, which prioritizes disparity and achieves a high correlation with human judgments on stereo effect. Moreover, we propose a strong baseline model, harmonizing the stereo effect and image quality simultaneously, and notably surpassing current mainstream methods. Our code and data will be open-sourced to promote further research in stereo conversion. Our models are available at mono2stereo-bench.github.io. Songsong Yu, Zhongang Qi, Zeke Xie, Yifan Wang 0004, Lijun Wang 0001, Ying Shan, Huchuan Lu |
CVPR | 7 |
| 2025 | DI-PCG: Diffusion-based Efficient Inverse Procedural Content Generation for High-quality 3D Asset CreationabstractProcedural Content Generation (PCG) is powerful in creating high-quality 3D contents, yet controlling it to produce desired shapes is difficult and often requires extensive parameter tuning. Inverse Procedural Content Generation aims to automatically find the best parameters under the input condition. However, existing sampling-based and neural network-based methods still suffer from numerous sample iterations or limited controllability. In this work, we present DI-PCG, a novel and efficient method for Inverse PCG from general image conditions. At its core is a lightweight diffusion transformer model, where PCG parameters are directly treated as the denoising target and the observed images as conditions to control parameter generation. DI-PCG is efficient and effective. With only 7.6M network parameters and 30 GPU hours to train, it demonstrates superior performance in recovering parameters accurately, and generalizing well to in-the-wild images. Quantitative and qualitative experiment results validate the effectiveness of DI-PCG in inverse PCG and image-to-3D generation tasks. DI-PCG offers a promising approach for efficient inverse PCG and represents a valuable exploration step towards a 3D generation path that models how to construct a 3D asset using parametric models. Wang Zhao 0001, Yan-Pei Cao 0001, Yuejiang Dong, Ying Shan |
CVPR | 5 |
| 2025 | Editing Music with Melody and Text: Using ControlNet for Diffusion TransformerabstractDespite the significant progress in controllable music generation and editing, challenges remain in the quality and length of generated music due to the use of Mel-spectrogram representations and UNet-based model structures. To address these limitations, we propose a novel approach using a Diffusion Transformer (DiT) augmented with an additional control branch using ControlNet. This allows for long-form and variable-length music generation and editing controlled by text and melody prompts. For more precise and fine-grained melody control, we introduce a novel top-k constant-Q Transform representation as the melody prompt, reducing ambiguity compared to previous representations (e.g., chroma), particularly for music with multiple tracks or a wide range of pitch values. To effectively balance the control signals from text and melody prompts, we adopt a curriculum learning strategy that progressively masks the melody prompt, resulting in a more stable training process. Experiments have been performed on text-to-music generation and music-style transfer tasks using open-source instrumental recording data. The results demonstrate that by extending StableAudio, a pre-trained text-controlled DiT model, our approach enables superior melody-controlled editing while retaining good text-to-music generation performance. These results outperform a strong MusicGen baseline in terms of both text-based generation and melody preservation for editing. Audio examples can be found at https://stable-audio-control.github.io. Siyuan Hou, Shansong Liu, Ruibin Yuan, Wei Xue 0002, Ying Shan, Mangsuo Zhao, Chao Zhang 0031 |
ICASSP | 5 |
| 2025 | Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from Videos
Yi Chen 0019, Yuying Ge, Weiliang Tang, Yizhuo Li 0001, Yixiao Ge, Mingyu Ding, Ying Shan, Xihui Liu |
ICCV | 7 |
| 2025 | AnimeGamer: Infinite Anime Life Simulation with Next Game State PredictionabstractRecent advancements in image and video synthesis have opened up new promise in generative games. One particularly intriguing application is transforming characters from anime films into interactive, playable entities. This allows players to immerse themselves in the dynamic anime world as their favorite characters for life simulation through language instructions. Such games are defined as infinite game since they eliminate predetermined boundaries and fixed gameplay rules, where players can interact with the game world through open-ended language and experience ever-evolving storylines and environments. Recently, a pioneering approach for infinite anime life simulation employs large language models (LLMs) to translate multi-turn text dialogues into language instructions for image generation. However, it neglects historical visual context, leading to inconsistent gameplay. Furthermore, it only generates static images, failing to incorporate the dynamics necessary for an engaging gaming experience. In this work, we propose AnimeGamer, which is built upon Multimodal Large Language Models (MLLMs) to generate each game state, including dynamic animation shots that depict character movements and updates to character states, as illustrated in Figure 1. We introduce novel action-aware multimodal representations to represent animation shots, which can be decoded into high-quality video clips using a video diffusion model. By taking historical animation shot representations as context and predicting subsequent representations, AnimeGamer can generate games with contextual consistency and satisfactory dynamics. Extensive evaluations using both automated metrics and human evaluations demonstrate that AnimeGamer outperforms existing methods in various aspects of the gaming experience. Codes and checkpoints are available at https://github.com/TencentARC/AnimeGamer. Yuying Ge, Yixiao Ge, Jing Liao 0001, Ying Shan |
ICCV | 5 |
| 2025 | DepthSync: Diffusion Guidance-Based Depth Synchronization for Scale- and Geometry-Consistent Video Depth EstimationabstractDiffusion-based video depth estimation methods have achieved remarkable success with strong generalization ability. However, predicting depth for long videos remains challenging. Existing methods typically split videos into overlapping sliding windows, leading to accumulated scale discrepancies across different windows, particularly as the number of windows increases. Additionally, these methods rely solely on 2D diffusion priors, overlooking the inherent 3D geometric structure of video depths, which results in geometrically inconsistent predictions. In this paper, we propose DepthSync, a novel, training-free framework using diffusion guidance to achieve scale- and geometry-consistent depth predictions for long videos. Specifically, we introduce scale guidance to synchronize the depth scale across windows and geometry guidance to enforce geometric alignment within windows based on the inherent 3D constraints in video depths. These two terms work synergistically, steering the denoising process toward consistent depth predictions. Experiments on various datasets validate the effectiveness of our method in producing depth estimates with improved scale and geometry consistency, particularly for long videos. Yuejiang Dong, Wang Zhao 0001, Ying Shan, Song-Hai Zhang |
ICCV | 4 |
| 2025 | VisionMath: Vision-Form Mathematical Problem-Solving
Zongyang Ma, Ziqi Zhang 0010, Zhongang Oi, Chunfeng Yuan, Shaojie Zhu, Chengxiang Zhuo, Bing Li 0001, Ye Liu 0002, Zang Li, Ying Shan, Weiming Hu 0004 |
ICCV | 11 |
| 2025 | GenHancer: Imperfect Generative Models are Secretly Strong Vision-Centric EnhancersabstractThe synergy between generative and discriminative models receives growing attention. While discriminative Contrastive Language-Image Pre-Training (CLIP) excels in high-level semantics, it struggles with perceiving fine-grained visual details. Generally, to enhance representations, generative models take CLIP's visual features as conditions for reconstruction. However, the underlying principle remains underexplored. In this work, we empirically found that visually perfect generations are not always optimal for representation enhancement. The essence lies in effectively extracting fine-grained knowledge from generative models while mitigating irrelevant information. To explore critical factors, we delve into three aspects: (1) Conditioning mechanisms: We found that even a small number of local tokens can drastically reduce the difficulty of reconstruction, leading to collapsed training. We thus conclude that utilizing only global visual tokens as conditions is the most effective strategy. (2) Denoising configurations: We observed that end-to-end training introduces extraneous information. To address this, we propose a two-stage training strategy to prioritize learning useful visual knowledge. Additionally, we demonstrate that lightweight denoisers can yield remarkable improvements. (3) Generation paradigms: We explore both continuous and discrete denoisers with desirable outcomes, validating the versatility of our method. Through our in-depth explorations, we have finally arrived at an effective method, namely GenHancer, which consistently outperforms prior arts on the MMVP-VLM benchmark, e.g., 6.0% on OpenAICLIP. The enhanced CLIP can be further plugged into multimodal large language models for better vision-centric performance. All the models and codes are made publicly available. Shijie Ma, Yuying Ge, Teng Wang 0007, Yixiao Ge, Ying Shan |
ICCV | 6 |
| 2025 | Scalable Image Tokenization with Index Backpropagation Quantization
Fengyuan Shi 0001, Zhuoyan Luo, Yixiao Ge, Yujiu Yang 0001, Ying Shan, Limin Wang 0002 |
ICCV | 5 |
| 2025 | Mamba-3VL: Taming State Space Model for 3D Vision Language Learning
Zhongang Qi, Jile Jiao, Xuetao Feng, Yujia Liang, Ying Shan |
ICCV | 8 |
| 2025 | Geometrycrafter: Consistent Geometry Estimation for Open-World Videos With Diffusion PriorsabstractDespite remarkable advancements in video depth estimation, existing methods exhibit inherent limitations in achieving geometric fidelity through the affine-invariant predictions, limiting their applicability in reconstruction and other metrically grounded downstream tasks. We propose GeometryCrafter, a novel framework that recovers high-fidelity point map sequences with temporal coherence from open-world videos, enabling accurate 3D/4D reconstruction, camera parameter estimation, and other depth-based applications. At the core of our approach lies a point map Variational Autoencoder (VAE) that learns a latent space agnostic to video latent distributions for effective point map encoding and decoding. Leveraging the VAE, we train a video diffusion model to model the distribution of point map sequences conditioned on the input videos. Extensive evaluations on diverse datasets demonstrate that GeometryCrafter achieves state-of-the-art 3D accuracy, temporal consistency, and generalization capability. Tian-Xing Xu, Xiangjun Gao, Wenbo Hu 0002, Xiaoyu Li 0002, Song-Hai Zhang, Ying Shan |
ICCV | 6 |
| 2025 | FreeSplatter: Pose-free Gaussian Splatting for Sparse-view 3D ReconstructionabstractSparse-view reconstruction models typically require precise camera poses, yet obtaining these parameters from sparse-view images remains challenging. We introduce FreeSplatter, a scalable feed-forward framework that generates high-quality 3D Gaussians from uncalibrated sparse-view images while estimating camera parameters within seconds. Our approach employs a streamlined transformer architecture where self-attention blocks facilitate information exchange among multi-view image tokens, decoding them into pixel-aligned 3D Gaussian primitives within a unified reference frame. This representation enables both high-fidelity 3D modeling and efficient camera parameter estimation using off-the-shelf solvers. We develop two specialized variants--for object-centric and scene-level reconstruction--trained on comprehensive datasets. Remarkably, FreeSplatter outperforms several pose-dependent Large Reconstruction Models (LRMs) by a notable margin while achieving comparable or even better pose estimation accuracy compared to state-of-the-art pose-free reconstruction approach MASt3R in challenging benchmarks. Beyond technical benchmarks, FreeSplatter streamlines text/image-to-3D content creation pipelines, eliminating the complexity of camera pose management while delivering exceptional visual fidelity. Shenghua Gao, Ying Shan |
ICCV | 3 |
| 2025 | TrajectoryCrafter: Redirecting Camera Trajectory for Monocular Videos via Diffusion Models
Mark Yu, Wenbo Hu 0002, Jinbo Xing, Ying Shan |
ICCV | 4 |
| 2025 | HaploVL: A Single-Transformer Baseline for Multi-Modal UnderstandingabstractRecent advancements in large language models (LLMs) have significantly propelled the development of large multi-modal models (LMMs), highlighting the potential for general and intelligent assistants. However, most LMMs model visual and textual modalities separately, leading to recent efforts to develop native LMMs using a single transformer. Despite the promise, these native models are resource-intensive and often exhibit performance gaps compared to their compositional counterparts. To alleviate this issue, we propose a simple yet efficient method to construct a baseline for the native and end-to-end large multi-modal model in a single transformer. First, we propose a new early-fusion LMM that can fuse multi-modal inputs in the early stage and respond to visual instructions in an auto-regressive manner. Second, we devise an efficient training recipe for the proposed model, which harnesses the prior knowledge of the pre-trained models, addressing both the performance limitations and the challenge of resource consumption. The proposed model demonstrates superior performance compared to other LMMs using one transformer and significantly narrows the performance gap with compositional LMMs. Rui Yang 0041, Lin Song 0002, Yicheng Xiao, Runhui Huang, Yixiao Ge, Ying Shan, Hengshuang Zhao |
ICML | 6 |
| 2025 | Taming Rectified Flow for Inversion and EditingabstractRectified-flow-based diffusion transformers like FLUX and OpenSora have demonstrated outstanding performance in the field of image and video generation. Despite their robust generative capabilities, these models often struggle with inversion inaccuracies, which could further limit their effectiveness in downstream tasks such as image and video editing. To address this issue, we propose RF-Solver, a novel training-free sampler that effectively enhances inversion precision by mitigating the errors in the ODE-solving process of rectified flow. Specifically, we derive the exact formulation of the rectified flow ODE and apply the high-order Taylor expansion to estimate its nonlinear components, significantly enhancing the precision of ODE solutions at each timestep. Building upon RF-Solver, we further propose RF-Edit, a general feature-sharing-based framework for image and video editing. By incorporating self-attention features from the inversion process into the editing process, RF-Edit effectively preserves the structural information of the source image or video while achieving high-quality editing results. Our approach is compatible with any pre-trained rectified-flow-based models for image and video tasks, requiring no additional training or optimization. Extensive experiments across generation, inversion, and editing tasks in both image and video modalities demonstrate the superiority and versatility of our method. The source code is available at https://github.com/wangjiangshan0725/RF-Solver-Edit. Jiangshan Wang, Junfu Pu, Zhongang Qi, Yue Ma 0016, Nisha Huang, Xiu Li 0001, Ying Shan |
ICML | 9 |
| 2025 | LoRA-Gen: Specializing Large Language Model via Online LoRA GenerationabstractRecent advances have highlighted the benefits of scaling language models to enhance performance across a wide range of NLP tasks. However, these approaches still face limitations in effectiveness and efficiency when applied to domain-specific tasks, particularly for small edge-side models. We propose the LoRA-Gen framework, which utilizes a large cloud-side model to generate LoRA parameters for edge-side models based on task descriptions. By employing the reparameterization technique, we merge the LoRA parameters into the edge-side model to achieve flexible specialization. Our method facilitates knowledge transfer between models while significantly improving the inference efficiency of the specialized model by reducing the input context length. Without specialized training, LoRA-Gen outperforms conventional LoRA fine-tuning, which achieves competitive accuracy and a 2.1x speedup with TinyLLaMA-1.1B in reasoning tasks. Besides, our method delivers a compress ratio of 10.1x with Gemma-2B on intelligent agent tasks. Yicheng Xiao, Lin Song 0002, Rui Yan 0001, Yixiao Ge, Xiu Li 0001, Ying Shan |
ICML | 7 |
| 2025 | UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual ReasoningabstractRecent advances in Large Multi-modal Models (LMMs) have demonstrated their remarkable success as general-purpose multi-modal assistants, with particular focuses on holistic image- and video-language understanding. Conversely, less attention has been given to scaling fine-grained pixel-level understanding capabilities, where the models are expected to realize pixel-level alignment between visual signals and language semantics. Some previous studies have applied LMMs to related tasks such as region-level captioning and referring expression segmentation. However, these models are limited to performing either referring or segmentation tasks independently and fail to integrate these fine-grained perception capabilities into visual reasoning. To bridge this gap, we propose UniPixel, a large multi-modal model capable of flexibly comprehending visual prompt inputs and generating mask-grounded responses. Our model distinguishes itself by seamlessly integrating pixel-level perception with general visual understanding capabilities. Specifically, UniPixel processes visual prompts and generates relevant masks on demand, and performs subsequent reasoning conditioning on these intermediate pointers during inference, thereby enabling fine-grained pixel-level reasoning. The effectiveness of our approach has been verified on 10 benchmarks across a diverse set of tasks, including pixel-level referring/segmentation and object-centric understanding in images/videos. A novel PixelQA task that jointly requires referring, segmentation, and question answering is also designed to verify the flexibility of our method. Ye Liu 0002, Zongyang Ma, Junfu Pu, Zhongang Qi, Yang Wu 0001, Ying Shan, Chang Wen Chen |
NeurIPS | 6 |
| 2025 | MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPOabstractRecent text-to-image systems face limitations in handling multimodal inputs and complex reasoning tasks.
We introduce MindOmni, a unified multimodal large language model that addresses these challenges by incorporating reasoning generation through reinforcement learning. MindOmni leverages a three-phase training strategy: i) design of a unified vision language model with a decoder-only diffusion module, ii) supervised fine-tuning with Chain-of-Thought (CoT) instruction data, and iii) our proposed Reasoning Generation Policy Optimization (RGPO) algorithm, utilizing multimodal feedback to effectively guide policy updates.
Experimental results demonstrate that MindOmni outperforms existing models, achieving impressive performance on both understanding and generation benchmarks, meanwhile showcasing advanced fine-grained reasoning generation capabilities, especially with mathematical reasoning instruction. All codes will be made public. Yicheng Xiao, Lin Song 0002, Yukang Chen, Yingmin Luo, Yukang Gan, Wei Huang 0042, Xiu Li 0001, Xiaojuan Qi 0001, Ying Shan |
NeurIPS | 10 |
| 2025 | BlobCtrl: Taming Controllable Blob for Element-level Image EditingabstractAs user expectations for image editing continue to rise, the demand for flexible, fine-grained manipulation of specific visual elements presents a challenge for current diffusion-based methods. In this work, we present BlobCtrl, a framework for element-level image editing based on a probabilistic blob-based representation. Treating blobs as visual primitives, BlobCtrl disentangles layout from appearance, affording fine-grained, controllable object-level elements manipulation. Our key contributions are twofold: 1) an in-context dual-branch diffusion model that separates foreground and background processing, incorporating blob representations to explicitly decouple layout and appearance; and 2) a self-supervised disentangle-then-reconstruct training paradigm with an identity-preserving loss function, along with tailored strategies to efficiently leverage blob-image pairs. To foster further research, we introduce BlobData for large-scale training, and BlobBench, a benchmark for systematic evaluation. Experimental results demonstrate that BlobCtrl achieves state-of-the-art performance in a variety of element-level editing tasks—such as object addition, removal, scaling, and replacement—while maintaining computational efficiency. Yaowei Li 0001, Lingen Li, Zhaoyang Zhang 0004, Xiaoyu Li 0002, Guangzhi Wang, Hongxiang Li 0004, Xiaodong Cun, Ying Shan, Yuexian Zou |
SIGGRAPH Asia | 8 |
| 2025 | Assembler: Scalable 3D Part Assembly via Anchor Point DiffusionabstractWe present Assembler, a scalable and generalizable framework for 3D part assembly that reconstructs complete objects from input part meshes and a reference image. Unlike prior approaches that mostly rely on deterministic part pose prediction and category-specific training, Assembler is designed to handle diverse, in-the-wild objects with varying part counts, geometries, and structures. It addresses the core challenges of scaling to general 3D part assembly through innovations in task formulation, representation, and data. First, Assembler casts part assembly as a generative problem and employs diffusion models to sample plausible configurations, effectively capturing ambiguities arising from symmetry, repeated parts, and multiple valid assemblies. Second, we introduce a novel shape-centric representation based on sparse anchor point clouds, enabling scalable generation in Euclidean space and avoiding the limitations of abstract SE(3) pose prediction. Third, we construct a large-scale dataset of over 320K diverse part-object assemblies using a synthesis and filtering pipeline built on existing 3D shape repositories. Assembler achieves state-of-the-art performance on PartNet and is the first to demonstrate high-quality assembly for complex, real-world objects. Based on Assembler, we further introduce an interesting part-aware 3D modeling system that generates high-resolution, editable objects from images, demonstrating potential for interactive and compositional design. Project page: https://assembler3d.github.io/ Wang Zhao 0001, Yan-Pei Cao 0001, Yuejiang Dong, Ying Shan |
SIGGRAPH Asia | 5 |
| 2025 | StyleAdapter: A Unified Stylized Image Generation Model
Zhouxia Wang, Xintao Wang 0002, Liangbin Xie, Zhongang Qi, Ying Shan, Wenping Wang 0001, Ping Luo 0002 |
Int. J. Comput. Vis. | 5 |
| 2025 | UV Gaussians: Joint learning of mesh deformation and Gaussian textures for human avatar modeling
Yujiao Jiang, Qingmin Liao, Xiaoyu Li 0002, Qi Zhang 0029, Chaopeng Zhang, Zongqing Lu 0001, Ying Shan |
Knowl. Based Syst. | 8 |
| 2025 | Weakly-Supervised Temporal Action Localization by Progressive Complementary LearningabstractWeakly-Supervised Temporal Action Localization (WSTAL) aims to localize and classify action instances in long untrimmed videos with only video-level category labels as supervision. A critical challenge of WSTAL is the large gap between video-level supervision and unavailable snippet-level supervision. Prevailing methods typically assign pseudo labels to snippets, but these methods suffer from significant noise caused by the pseudo snippet-level labels. In this work, we address the WSTAL from a novel category exclusion perspective, which gradually enhances the snippet-level supervision to bridge the gap. Our proposed Progressive Complementary Learning (ProCL) is inspired by the fact that, video-level labels precisely indicate the categories that all snippets surely do not belong to, which is ignored by previous works. Accordingly, we first exclude these surely non-existent categories by the deterministic complementary learning. And then, we introduce the entropy-based pseudo complementary learning that is able to exclude more categories for snippets of less ambiguity. Furthermore, for the remaining ambiguous snippets, we attempt to reduce the ambiguity by distinguishing foreground actions from the background. Extensive experimental results show that our method achieves new state-of-the-art performance on THUMOS14, ActivityNet1.3, and MultiTHUMOS benchmarks. Jia-Run Du, Jia-Chang Feng, Kun-Yu Lin, Fa-Ting Hong, Zhongang Qi, Ying Shan, Jianfang Hu, Wei-Shi Zheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Learning Multi-Task Target-Specific Correlation Filters for Robust TrackingabstractIn recent years, correlation filter based trackers have shown great potentials in visual tracking because of their high computational efficiency and low memory consumption. However, their increasing tracking performance typically comes at the cost of sacrificing the computational speed and memory usage. Furthermore, training high-dimensional correlation filters with a large number of parameters usually introduces the risk of over-fitting. In this paper, we propose Multi-Task Target-Specific Correlation Filters (MTSCF) to tackle the above issues. First, we construct a novel regression formulation for multi-task filter learning to promote both competition and collaboration among correlation filters to select discriminative features for robust tracking. This significantly reduces redundancies among features at both spatial level and channel level, which produces sparse correlation filters. Then, we develop an effective filter importance evaluation criterion according to the expansion of designed regression formulation to choose a set of target-specific features for efficient tracking. This significantly reduces the number of filter parameters, which further results in compact correlation filters. Moreover, we propose to efficiently optimize the proposed MTSCF via an Alternating Direction Method of Multipliers (ADMM) algorithm. Evaluation results on six challenging benchmark datasets (i.e., OTB2013, OTB2015, VOT2016, VOT2018, UAV20L and LaSOT) show the proposed method performs favorably against existing state-of-the-art DCF based trackers, and it retains a high speed of 40 FPS on a CPU when evaluated with only hand-crafted features. Luo Xiong, Ying Shan, Hanzi Wang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | HumanRef-GS: Image-to-3D Human Generation With Reference-Guided Diffusion and 3D Gaussian SplattingabstractGenerating a 3D human model from a single reference image is a challenging task as it involves inferring textures and geometries in unseen views while maintaining consistency with the reference image. Existing methods that rely on 3D generative models are limited by the availability of 3D training data. Optimization-based approaches that distill text-to-image diffusion models into 3D models often struggle to preserve the intricate texture details of the reference image, resulting in inconsistent appearances across different views. In this paper, we propose HumanRef-GS, a novel method for single image-to-3D clothed human generation based on 3D Gaussian Splatting (3DGS). To ensure the generated 3D model is both photorealistic and consistent with the input image, HumanRef-GS employs a unique technique called reference-guided score distillation sampling (Ref-SDS). This method effectively incorporates image guidance into the generation process, enhancing the quality of the results. Additionally, we introduce region-aware attention to Ref-SDS, which ensures accurate correspondence between different body regions. To mitigate the impact of view dependence in 3DGS and enhance the view-consistency of the generated results, we substitute the anisotropic Gaussians in the vanilla representation with isotropic Gaussians. By utilizing the 3D Gaussian representation, our method significantly enhances the generation efficiency and rendering speed of 3D clothed human models. This improvement allows for faster and more efficient generation of high-quality results. Experimental results demonstrate that HumanRef-GS surpasses state-of-the-art methods in generating 3D clothed humans with fine geometry, photorealistic textures, and view-consistent appearances. We are committed to making our code and model available upon acceptance for further research and exploration. Jingbo Zhang 0002, Xiaoyu Li 0002, Hongliang Zhong, Qi Zhang 0029, Yan-Pei Cao 0001, Ying Shan, Jing Liao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Make-Your-Video: Customized Video Generation Using Textual and Structural GuidanceabstractCreating a vivid video from the event or scenario in our imagination is a truly fascinating experience. Recent advancements in text-to-video synthesis have unveiled the potential to achieve this with prompts only. While text is convenient in conveying the overall scene context, it may be insufficient to control precisely. In this paper, we explore customized video generation by utilizing text as context description and motion structure (e.g., frame-wise depth) as concrete guidance. Our method, dubbed Make-Your-Video, involves joint-conditional video generation using a Latent Diffusion Model that is pre-trained for still image synthesis and then promoted for video generation with the introduction of temporal modules. This two-stage learning scheme not only reduces the computing resources required, but also improves the performance by transferring the rich concepts available in image datasets solely into video generation. Moreover, we use a simple yet effective causal attention mask strategy to enable longer video synthesis, which mitigates the potential quality degradation effectively. Experimental results show the superiority of our method over existing baselines, particularly in terms of temporal coherence and fidelity to users' guidance. In addition, our model enables several intriguing applications that demonstrate potential for practical usage. Jinbo Xing, Menghan Xia, Yuechen Zhang, Yong Zhang 0034, Yingqing He, Hanyuan Liu, Haoxin Chen, Xiaodong Cun, Xintao Wang 0002, Ying Shan, Tien-Tsin Wong |
IEEE Trans. Vis. Comput. Graph. | 11 |
| 2025 | GP-Recon: Online Monocular Neural 3D Reconstruction With Geometric PriorabstractHigh-fidelity online 3D scene reconstruction from monocular videos continues to be challenging, especially for coherent and fine-grained geometry reconstruction. The previous learning-based online 3D reconstruction approaches with neural implicit representations have shown a promising ability for coherent scene reconstruction, but often fail to consistently reconstruct fine-grained geometric details during online reconstruction. This paper presents a new on-the-fly monocular 3D reconstruction approach, named GP-Recon, to perform high-fidelity online neural 3D reconstruction with fine-grained geometric details. We incorporate geometric prior (GP) into a scene's neural geometry learning to better capture its geometric details and, more importantly, propose an online volume rendering optimization to reconstruct and maintain geometric details during the online reconstruction task. The extensive comparisons with state-of-the-art approaches show that our GP-Recon consistently generates more accurate and complete reconstruction results with much better fine-grained details, both quantitatively and qualitatively. Zixin Zou, Shi-Sheng Huang, Yan-Pei Cao 0001, Tai-Jiang Mu, Ying Shan, Hongbo Fu 0001, Song-Hai Zhang |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2024 | SparseGNV: Generating Novel Views of Indoor Scenes with Sparse RGB-D ImagesabstractWe study to generate novel views of indoor scenes given sparse input views. The challenge is to achieve both photorealism and view consistency. We present SparseGNV: a learning framework that incorporates 3D structures and image generative models to generate novel views with three modules. The first module builds a neural point cloud as underlying geometry, providing scene context and guidance for the target novel view. The second module utilizes a transformer-based network to map the scene context and the guidance into a shared latent space and autoregressively decodes the target view in the form of discrete image tokens. The third module reconstructs the tokens back to the image of the target view. SparseGNV is trained across a large-scale indoor scene dataset to learn generalizable priors. Once trained, it can efficiently generate novel views of an unseen indoor scene in a feed-forward manner. We evaluate SparseGNV on real-world indoor scenes and demonstrate that it outperforms state-of-the-art methods based on either neural radiance fields or conditional image generation. Weihao Cheng 0002, Yan-Pei Cao 0001, Ying Shan |
AAAI | 3 |
| 2024 | SC-NeuS: Consistent Neural Surface Reconstruction from Sparse and Noisy ViewsabstractThe recent neural surface reconstruction approaches using volume rendering have made much progress by achieving impressive surface reconstruction quality, but are still limited to dense and highly accurate posed views. To overcome such drawbacks, this paper pays special attention on the consistent surface reconstruction from sparse views with noisy camera poses. Unlike previous approaches, the key difference of this paper is to exploit the multi-view constraints directly from the explicit geometry of the neural surface, which can be used as effective regularization to jointly learn the neural surface and refine the camera poses. To build effective multi-view constraints, we introduce a fast differentiable on-surface intersection to generate on-surface points, and propose view-consistent losses on such differentiable points to regularize the neural surface learning. Based on this point, we propose a joint learning strategy, named SC-NeuS, to perform geometry-consistent surface reconstruction in an end-to-end manner. With extensive evaluation on public datasets, our SC-NeuS can achieve consistently better surface reconstruction results with fine-grained details than previous approaches, especially from sparse and noisy camera views. The source code is available at https://github.com/zouzx/sc-neus.git. Shi-Sheng Huang, Zixin Zou, Yan-Pei Cao 0001, Ying Shan |
AAAI | 5 |
| 2024 | T2I-Adapter: Learning Adapters to Dig Out More Controllable Ability for Text-to-Image Diffusion ModelsabstractThe incredible generative ability of large-scale text-to-image (T2I) models has demonstrated strong power of learning complex structures and meaningful semantics. However, relying solely on text prompts cannot fully take advantage of the knowledge learned by the model, especially when flexible and accurate controlling (e.g., structure and color) is needed. In this paper, we aim to ``dig out" the capabilities that T2I models have implicitly learned, and then explicitly use them to control the generation more granularly. Specifically, we propose to learn low-cost T2I-Adapters to align internal knowledge in T2I models with external control signals, while freezing the original large T2I models. In this way, we can train various adapters according to different conditions, achieving rich control and editing effects in the color and structure of the generation results. Further, the proposed T2I-Adapters have attractive properties of practical value, such as composability and generalization ability. Extensive experiments demonstrate that our T2I-Adapter has promising generation quality and a wide range of applications. Our code is available at https://github.com/TencentARC/T2I-Adapter. Chong Mou, Xintao Wang 0002, Liangbin Xie, Yanze Wu, Jian Zhang 0018, Zhongang Qi, Ying Shan |
AAAI | 7 |
| 2024 | SphereDiffusion: Spherical Geometry-Aware Distortion Resilient Diffusion ModelabstractControllable spherical panoramic image generation holds substantial applicative potential across a variety of domains. However, it remains a challenging task due to the inherent spherical distortion and geometry characteristics, resulting in low-quality content generation. In this paper, we introduce a novel framework of SphereDiffusion to address these unique challenges, for better generating high-quality and precisely controllable spherical panoramic images. For the spherical distortion characteristic, we embed the semantics of the distorted object with text encoding, then explicitly construct the relationship with text-object correspondence to better use the pre-trained knowledge of the planar images. Meanwhile, we employ a deformable technique to mitigate the semantic deviation in latent space caused by spherical distortion. For the spherical geometry characteristic, in virtue of spherical rotation invariance, we improve the data diversity and optimization objectives in the training process, enabling the model to better learn the spherical geometry characteristic. Furthermore, we enhance the denoising process of the diffusion model, enabling it to effectively use the learned geometric characteristic to ensure the boundary continuity of the generated images. With these specific techniques, experiments on Structured3D dataset show that SphereDiffusion significantly improves the quality of controllable spherical image generation and relatively reduces around 35% FID on average. Xuewei Li 0003, Zhongang Qi, Xintao Wang 0002, Ying Shan, Xi Li 0001 |
AAAI | 6 |
| 2024 | A Pre-convolved Representation for Plug-and-Play Neural Illumination FieldsabstractRecent advances in implicit neural representation have demonstrated the ability to recover detailed geometry and material from multi-view images. However, the use of simplified lighting models such as environment maps to represent non-distant illumination, or using a network to fit indirect light modeling without a solid basis, can lead to an undesirable decomposition between lighting and material. To address this, we propose a fully differentiable framework named Neural Illumination Fields (NeIF) that uses radiance fields as a lighting model to handle complex lighting in a physically based way. Together with integral lobe encoding for roughness-adaptive specular lobe and leveraging the pre-convolved background for accurate decomposition, the proposed method represents a significant step towards integrating physically based rendering into the NeRF representation. The experiments demonstrate the superior performance of novel-view rendering compared to previous works, and the capability to re-render objects under arbitrary NeRF-style environments opens up exciting possibilities for bridging the gap between virtual and real-world scenes. Yiyu Zhuang, Qi Zhang 0029, Xuan Wang 0009, Hao Zhu 0004, Xiaoyu Li 0002, Ying Shan, Xun Cao |
AAAI | 7 |
| 2024 | Sparse3D: Distilling Multiview-Consistent Diffusion for Object Reconstruction from Sparse ViewsabstractReconstructing 3D objects from extremely sparse views is a long-standing and challenging problem. While recent techniques employ image diffusion models for generating plausible images at novel viewpoints or for distilling pre-trained diffusion priors into 3D representations using score distillation sampling (SDS), these methods often struggle to simultaneously achieve high-quality, consistent, and detailed results for both novel-view synthesis (NVS) and geometry. In this work, we present Sparse3D, a novel 3D reconstruction method tailored for sparse view inputs. Our approach distills robust priors from a multiview-consistent diffusion model to refine a neural radiance field. Specifically, we employ a controller that harnesses epipolar features from input views, guiding a pre-trained diffusion model, such as Stable Diffusion, to produce novel-view images that maintain 3D consistency with the input. By tapping into 2D priors from powerful image diffusion models, our integrated model consistently delivers high-quality results, even when faced with open-world objects. To address the blurriness introduced by conventional SDS, we introduce the category-score distillation sampling (C-SDS) to enhance detail. We conduct experiments on CO3DV2 which is a multi-view dataset of real-world objects. Both quantitative and qualitative evaluations demonstrate that our approach outperforms previous state-of-the-art works on the metrics regarding NVS and geometry reconstruction. Zixin Zou, Weihao Cheng 0002, Yan-Pei Cao 0001, Shi-Sheng Huang, Ying Shan, Song-Hai Zhang |
AAAI | 5 |
| 2024 | LLaMA Pro: Progressive LLaMA with Block ExpansionabstractChengyue Wu, Yukang Gan, Yixiao Ge, Zeyu Lu, Jiahao Wang, Ye Feng, Ying Shan, Ping Luo. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Chengyue Wu, Yukang Gan, Yixiao Ge, Jiahao Wang 0005, Ying Shan, Ping Luo 0002 |
ACL (1) | 7 |
| 2024 | DreamAvatar: Text-and-Shape Guided 3D Human Avatar Generation via Diffusion ModelsabstractWe present DreamAvatar, a text-and-shape guided framework for generating high-quality 3D human avatars with controllable poses. While encouraging results have been reported by recent methods on text-guided 3D common object generation, generating high-quality human avatars remains an open challenge due to the complexity of the human body's shape, pose, and appearance. We propose DreamAvatar to tackle this challenge, which utilizes a train-able NeRF for predicting density and color for 3D points and pretrained text-to-image diffusion models for providing 2D self-supervision. Specifically, we leverage the SMPL model to provide shape and pose guidance for the generation. We introduce a dual-observation-space design that involves the joint optimization of a canonical space and a posed space that are related by a learnable deformation field. This facilitates the generation of more complete textures and geometry faithful to the target pose. We also jointly optimize the losses computed from the full body and from the zoomed-in 3D head to alleviate the common multi-face “Janus” problem and improve facial details in the generated avatars. Extensive evaluations demonstrate that DreamAvatar significantly outperforms existing meth-ods, establishing a new state-of-the-art for text-and-shape guided 3D human avatar generation. Yan-Pei Cao 0001, Kai Han 0001, Ying Shan, Kwan-Yee Kenneth Wong |
CVPR | 4 |
| 2024 | How to Make Cross Encoder a Good Teacher for Efficient Image-Text Retrieval?abstractDominant dual-encoder models enable efficient image-text retrieval but suffer from limited accuracy, while the cross-encoder models offer higher accuracy at the expense of efficiency. Distilling cross-modality matching knowledge from cross-encoder to dual-encoder provides a natural approach to harness their strengths. Thus, we investigate the following valuable question: how to make cross-encoder a good teacher for dual-encoder? Our findings are threefold: (1) Cross-modal similarity score distribution of cross-encoder is more concentrated, while the result of dual-encoder is nearly normal, making vanilla logit distillation less effective. However, ranking distillation remains practical, as it is not affected by the score distribution. (2) Only the relative order between hard negatives conveys valid knowledge, while the order information between easy negatives has little significance. (3) Maintaining the coordination between distillation loss and dual-encoder training loss is beneficial for knowledge transfer. Based on these findings, we propose a novel Contrastive Partial Ranking Distillation (CPRD) method, which implements the objective of mimicking relative order between hard negative samples with contrastive learning. This approach coordinates with the training of the dual-encoder, effectively transferring valid knowledge from the cross-encoder to the dual-encoder. Extensive experiments on image-text retrieval and ranking tasks show that our method surpasses other distillation methods and significantly improves the accuracy of dual-encoder. Zongyang Ma, Ziqi Zhang 0010, Zhongang Qi, Chunfeng Yuan, Bing Li 0001, Junfu Pu, Ying Shan, Xiaojuan Qi 0001, Weiming Hu 0004 |
CVPR | 8 |
| 2024 | VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion ModelsabstractText-to-video generation aims to produce a video based on a given prompt. Recently, several commercial video models have been able to generate plausible videos with mini-mal noise, excellent details, and high aesthetic scores. However, these models rely on large-scale, well-filtered, high-quality videos that are not accessible to the community. Many existing research works, which train models using the low-quality WebVid-10M dataset, struggle to generate high-quality videos because the models are optimized to fit WebVid-10M. In this work, we explore the training scheme of video models extended from Stable Diffusion and investigate the feasibility of leveraging low-quality videos and synthesized high-quality images to obtain a high-quality video model. We first analyze the connection between the spatial and temporal modules of video models and the distribution shift to low-quality videos. We observe that full training of all modules results in a stronger coupling between spatial and temporal modules than only training temporal modules. Based on this stronger coupling, we shift the distribution to higher quality without motion degradation by finetuning spatial modules with high-quality images, resulting in a generic high-quality video model. Evaluations are conducted to demonstrate the superiority of the proposed method, particularly in picture quality, motion, and concept composition. Haoxin Chen, Yong Zhang 0034, Xiaodong Cun, Menghan Xia, Xintao Wang 0002, Chao Weng, Ying Shan |
CVPR | 7 |
| 2024 | YOLO-World: Real-Time Open-Vocabulary Object DetectionabstractThe You Only Look Once (YOLO) series of detectors have established themselves as efficient and practical tools. However, their reliance on predefined and trained object categories limits their applicability in open scenarios. Addressing this limitation, we introduce YOLO-World, an innovative approach that enhances YOLO with open-vocabulary detection capabilities through vision-language modeling and pre-training on large-scale datasets. Specifically, we propose a new Re-parameterizable Vision-Language Path Aggregation Network (RepVL-PAN) and region-text contrastive loss to facilitate the interaction between visual and linguistic information. Our method excels in detecting a wide range of objects in a zero-shot manner with high efficiency. On the challenging LVIS dataset, YOLO-World achieves 35.4 AP with 52.0 FPS on V100, which outperforms many state-of-the-art methods in terms of both accuracy and speed. Furthermore, the finetuned YOLO-World achieves remarkable performance on several downstream tasks, including object detection and open-vocabulary instance segmentation. Code and models are available at: https://github.com/AILab-eve/YOLO-World. Tianheng Cheng, Lin Song 0002, Yixiao Ge, Wenyu Liu 0001, Xinggang Wang, Ying Shan |
CVPR | 6 |
| 2024 | UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image RecognitionabstractLarge-kernel convolutional neural networks (ConvNets) have recently received extensive research attention, but two unresolved and critical issues demand further investigation. 1) The architectures of existing large-kernel ConvNets largely follow the design principles of conventional ConvNets or transformers, while the architectural design for large-kernel ConvNets remains under-addressed. 2) As transformers have dominated multiple modalities, it re-mains to be investigated whether ConvNets also have a strong universal perception ability in domains beyond vision. In this paper, we contribute from two aspects. 1) We propose four architectural guidelines for designing large- kernel ConvNets, the core of which is to exploit the essential characteristics of large kernels that distinguish them from small kernels - they can see wide without going deep. Fol-lowing such guidelines, our proposed large-kernel ConvNet shows leading performance in image recognition (ImageNet accuracy of 88.0%, ADE20K mIoU of 55.6%, and COCO box AP of 56.4%), demonstrating better performance and higher speed than the recent powerful competitors. 2) We discover large kernels are the key to unlocking the exceptional performance of ConvNets in domains where they were originally not proficient. With certain modality-related pre-processing approaches, the proposed model achieves state- of-the-art performance on time-series forecasting and audio recognition tasks even without modality-specific customization to the architecture. All the code and models are publicly available on GitHub and Huggingface. Xiaohan Ding, Yixiao Ge, Sijie Zhao, Lin Song 0002, Xiangyu Yue 0001, Ying Shan |
CVPR | 7 |
| 2024 | ConTex-Human: Free-View Rendering of Human from a Single Image with Texture-Consistent SynthesisabstractIn this work, we propose a method to address the chal-lenge of rendering a 3D human from a single image in a free-view manner. Some existing approaches could achieve this by using generalizable pixel-aligned implicit fields to reconstruct a textured mesh of a human or by employing a 2D diffusion model as guidance with the Score Distillation Sampling (SDS) method, to lift the 2D image into 3D space. However, a generalizable implicit field often results in an over-smooth texture field, while the SDS method tends to lead to a texture-inconsistent novel view with the input image. In this paper, we introduce a texture-consistent back view synthesis module that could transfer the reference im-age content to the back view through depth and text-guided attention injection. Moreover, to alleviate the color distortion that occurs in the side region, we propose a visibility-aware patch consistency regularization for texture mapping and refinement combined with the synthesized back view texture. With the above techniques, we can achieve high-fidelity and texture-consistent human rendering from a single image. Experiments conducted on both real and synthetic data demonstrate the effectiveness of our method and show that our approach outperforms previous baseline methods. Xiangjun Gao, Xiaoyu Li 0002, Chaopeng Zhang, Qi Zhang 0029, Yan-Pei Cao 0001, Ying Shan, Long Quan |
CVPR | 6 |
| 2024 | Rethinking the Objectives of Vector-Quantized Tokenizers for Image SynthesisabstractVector-Quantized (VQ-based) generative models usually consist of two basic components, i.e., VQ tokenizers and generative transformers. Prior research focuses on improving the reconstruction fidelity of VQ tokenizers but rarely examines how the improvement in reconstruction affects the generation ability of generative transformers. In this paper, we surprisingly find that improving the reconstruction fidelity of VQ tokenizers does not necessarily improve the generation. Instead, learning to compress semantic features within VQ tokenizers significantly improves generative transformers' ability to capture textures and structures. We thus highlight two competing objectives of VQ tokeniz-ers for image synthesis: semantic compression and details preservation. Different from previous work that pri-oritizes better details preservation, we propose Semantic-Quantized GAN (SeQ-GAN) with two learning phases to balance the two objectives. In the first phase, we propose a semantic-enhanced perceptual loss for better semantic compression. In the second phase, we fix the encoder and codebook, but enhance and finetune the decoder to achieve better details preservation. Our proposed SeQ-GAN significantly improves VQ-based generative models for both un-conditional and conditional image generation. Specifically, SeQ-GAN achieves a Fré chet Inception Distance (FID) of 6.25 and Inception Score (IS) of 140.9 on 256×256 Ima-geNet generation, which is a remarkable improvement over VIT-VQGAN (714M), which obtains 11.2 FID and 97.2 IS. Yuchao Gu, Xintao Wang 0002, Yixiao Ge, Ying Shan, Zheng Shou 0001 |
CVPR | 4 |
| 2024 | SmartEdit: Exploring Complex Instruction-Based Image Editing with Multimodal Large Language ModelsabstractCurrent instruction-based image editing methods, such as InstructPix2Pix, often fail to produce satisfactory results in complex scenarios due to their dependence on the simple CLIP text encoder in diffusion models. To rectify this, this paper introduces SmartEdit, a novel approach of instruction-based image editing that leverages Multimodal Large Language Models (MLLMs) to enhance its understanding and reasoning capabilities. However, direct integration of these elements still faces challenges in situations requiring complex reasoning. To mitigate this, we propose a Bidirectional Interaction Module (BIM) that enables comprehensive bidirectional information interactions between the input image and the MLLM output. During training, we initially incorporate perception data to boost the perception and understanding capabilities of diffusion models. Subsequently, we demonstrate that a small amount of complex instruction editing data can effectively stimulate SmartEdit’ s editing capabilities for more complex instructions. We further construct a new evaluation dataset, Reason-Edit, specifically tailored for complex instruction-based image editing. Both quantitative and qualitative results on this evaluation dataset indicate that our SmartEdit surpasses previous methods, paving the way for the practical application of complex instruction-based image editing. Yuzhou Huang, Liangbin Xie, Xintao Wang 0002, Ziyang Yuan, Xiaodong Cun, Yixiao Ge, Jiantao Zhou 0001, Chao Dong 0005, Ruimao Zhang, Ying Shan |
CVPR | 11 |
| 2024 | VIT-LENS: Towards Omni-modal RepresentationsabstractAiming to advance AI agents, large foundation models significantly improve reasoning and instruction execution, yet the current focus on vision and language neglects the potential of perceiving diverse modalities in open-world environments. However, the success of data-driven vision and language models is costly or even infeasible to be reproduced for rare modalities. In this paper, we present Vit-lens that facilitates efficient omni-modal representation learning by perceiving novel modalities with a pretrained- ViT and aligning them to a pre-defined space. Specifically, the modality-specific lens is tuned to project any-modal signals to an intermediate embedding space, which are then processed by a strong ViT with pre-trained visual knowledge. The encoded representations are optimized toward aligning with the modal-independent space, pre-defined by off-the-shelf foundation models. Vit-lensprovides a unified solution for representation learning of increasing modalities with two appealing advantages: (i) Unlocking the great potential of pretrained- ViTs to novel modalities effectively with efficient parameters and data regime; (ii) Enabling emergent down- stream capabilities through modality alignment and shared ViT parameters. We tailor Vit-lensto learn representations for 3D point cloud, depth, audio, tactile and EEG, and set new state-of-the-art results across various understanding tasks, such as zero-shot classification. By seamlessly integrating Vit-lensinto Multimodal Foundation Models, we enable Any-modality to Text and Image Generation in a zero-shot manner. Code and models are available at https://github.com/TencentARC/ViT-Lens. Weixian Lei, Yixiao Ge, Difei Gao, Dylan Sun 0001, Yuying Ge, Ying Shan, Zheng Shou 0001 |
CVPR | 8 |
| 2024 | PhotoMaker: Customizing Realistic Human Photos via Stacked ID EmbeddingabstractRecent advances in text-to-image generation have made remarkable progress in synthesizing realistic human photos conditioned on given text prompts. However, existing per-sonalized generation methods cannot simultaneously sat-isfy the requirements of high efficiency, promising identity (ID) fidelity, and flexible text controllability. In this work, we introduce PhotoMaker, an efficient personalized text-to-image generation method, which mainly encodes an arbitrary number of input ID images into a stack ID embed-ding for preserving ID information. Such an embedding, serving as a unified ID representation, can not only encap-sulate the characteristics of the same input ID comprehen-sively, but also accommodate the characteristics of differ-ent IDs for subsequent integration. This paves the way for more intriguing and practically valuable applications. Be-sides, to drive the training of our PhotoMaker, we propose an ID-oriented data construction pipeline to assemble the training data. Under the nourishment of the dataset constructed through the proposed pipeline, our PhotoMaker demonstrates better ID preservation ability than test-time fine-tuning based methods, yet provides significant speed improvements, high-quality generation results, strong gen-eralization capabilities, and a wide range of applications. Zhen Li 0031, Mingdeng Cao, Xintao Wang 0002, Zhongang Qi, Ming-Ming Cheng, Ying Shan |
CVPR | 6 |
| 2024 | SEED-Bench: Benchmarking Multimodal Large Language ModelsabstractMultimodal large language models (MLLMs), building upon the foundation of powerful large language models (LLMs), have recently demonstrated exceptional capabilities in generating not only texts but also images given in-terleaved multimodal inputs (acting like a combination of GPT-4V and DALL-E 3). However, existing MLLM benchmarks remain limited to assessing only models' comprehension ability of single image-text inputs, failing to keep up with the strides made in MLLMs. A comprehensive benchmark is imperative for investigating the progress and uncovering the limitations of current MLLMs. In this work, we categorize the capabilities of MLLMs into hierarchical levels from L0to L4based on the modalities they can ac-cept and generate, and propose SEED-Bench, a comprehensive benchmark that evaluates the hierarchical capa-bilities of MLLMs. Specifically, SEED-Bench comprises 24K multiple-choice questions with accurate human annotations, which span 27 dimensions, including the evaluation of both text and image generation. Multiple-choice questions with ground truth options derived from human annotation enable an objective and efficient assessment of model performance, eliminating the need for human or GPT intervention during evaluation. We further evaluate the performance of 22 prominent open-source MLLMs and summarize valuable observations. By revealing the limitations of existing MLLMs through extensive evaluations, we aim for SEED-Bench to provide insights that will mo-tivate future research toward the goal of General Artificial Intelligence. Dataset and evaluation code are available at https://github.com/AILab-CVC/SEED-Bench. Bohao Li 0002, Yuying Ge, Yixiao Ge, Guangzhi Wang, Rui Wang 0092, Ruimao Zhang, Ying Shan |
CVPR | 7 |
| 2024 | GS-IR: 3D Gaussian Splatting for Inverse RenderingabstractWe propose GS-IR, a novel inverse rendering approach based on 3D Gaussian Splatting (3DGS) that leverages forward mapping volume rendering to achieve photorealistic novel view synthesis and relighting results. Unlike previous works that use implicit neural representations and volume rendering (e.g. NeRF), which suffer from low expressive power and high computational complexity, we extend 3DGS, a top-performance representation for novel view synthesis, to estimate scene geometry, surface material, and environment illumination from multi-view images captured under unknown lighting conditions. There are two main problems when introducing 3DGS to inverse rendering: 1) 3DGS does not support producing plausible normal natively; 2) forward mapping (e.g. rasterization and splatting) cannot trace the occlusion like backward mapping (e.g. ray tracing). To address these challenges, our GS-IR proposes an efficient optimization scheme incorporating a depth-derivation-based regularization for normal estimation and a baking-based occlusion to model indirect lighting. The flexible and expressive 3DGS representation allows us to achieve fast and compact geometry reconstruction, photore-alistic novel view synthesis, and effective physically-based rendering. We demonstrate the superiority of our method over baseline methods through qualitative and quantitative evaluations of various challenging scenes. The source code is available at https://github.com/lzhnb/GS-IR. Zhihao Liang 0002, Qi Zhang 0029, Ying Shan, Kui Jia |
CVPR | 4 |
| 2024 | EvalCrafter: Benchmarking and Evaluating Large Video Generation ModelsabstractThe vision and language generative models have been overgrown in recent years. For video generation, various open-sourced models and public-available services have been developed to generate high-quality videos. However, these methods often use a few metrics, e.g., FVD [56] or IS [45], to evaluate the performance. We argue that it is hard to judge the large conditional generative models from the simple metrics since these models are often trained on very large datasets with multi-aspect abilities. Thus, we propose a novel framework and pipeline for exhaustively evaluating the performance of the generated videos. Our approach involves generating a diverse and comprehensive list of 700 prompts for text-to-video generation, which is based on an analysis of real-world user data and generated with the assistance of a large language model. Then, we evaluate the state-of-the-art video generative models on our carefully designed benchmark, in terms of visual qualities, content qualities, motion qualities, and text-video alignment with 17 well-selected objective metrics. To obtain the finalleaderboard of the models, we further fit a series of coefficients to align the objective metrics to the users' opinions. Based on the proposed human alignment method, our final score shows a higher correlation than simply averaging the metrics, showing the effectiveness of the proposed evaluation method. Yaofang Liu, Xiaodong Cun, Xuebo Liu 0002, Xintao Wang 0002, Yong Zhang 0034, Haoxin Chen, Yang Liu 0005, Tieyong Zeng, Raymond Chan 0001, Ying Shan |
CVPR | 10 |
| 2024 | DynVideo-E: Harnessing Dynamic NeRF for Large-Scale Motion- and View-Change Human-Centric Video EditingabstractDespite recent progress in diffusion-based video editing, existing methods are limited to short-length videos due to the contradiction between long-range consistency and frame-wise editing. Prior attempts to address this challenge by introducing video-2D representations encounter significant difficulties with large motion- and view-change videos, especially in human-centric scenarios. To overcome this, we propose to introduce the dynamic Neural Radiance Fields (NeRF) as the innovative video representation, where the editing can be performed in the 3D spaces and propagated to the entire video via the deformation field. To provide consistent and controllable editing, we propose the image-based video-NeRF editing pipeline with a set of innovative designs, including multi-view multi-pose Score Distillation Sampling (SDS) from both the 2D personalized diffusion prior and 3D diffusion prior, reconstruction losses, text-guided local parts super-resolution, and style transfer. Extensive experiments demonstrate that our method dubbed as DynVideo-E, significantly outperforms SOTA approaches on two challenging datasets by a large margin of 50% ~ 95% for human preference. Code will be released at https://showlab.github.io/DynVideo-E/. Jia-Wei Liu, Yan-Pei Cao 0001, Jay Zhangjie Wu, Weijia Mao, Yuchao Gu, Rui Zhao 0001, Jussi Keppo, Ying Shan, Zheng Shou 0001 |
CVPR | 8 |
| 2024 | BT-Adapter: Video Conversation is Feasible Without Video Instruction TuningabstractThe recent progress in Large Language Models (LLM) has spurred various advancements in image-language con-versation agents, while how to build a proficient video-based dialogue system is still under exploration. Consid-ering the extensive scale of LLM and visual backbone, min-imal GPU memory is left for facilitating effective temporal modeling, which is crucial for comprehending and providing feedback on videos. To this end, we propose Branching Temporal Adapter (BT-Adapter), a novel method for ex-tending image-language pretrained models into the video domain. Specifically, BT-Adapter serves as a plug-and-use temporal modeling branch alongside the pretrained vi-sual encoder, which is tuned while keeping the backbone frozen. Just pretrained once, BT-Adapter can be seamlessly integrated into all image conversation models using this version of CLIP, enabling video conversations without the need for video instructions. Besides, we develop a unique asymmetric token masking strategy inside the branch with tailor-made training tasks for BT-Adapter, facilitating faster convergence and better results. Thanks to BT-Adapter, we are able to empower existing multimodal dialogue models with strong video understanding capabilities without incur-ring excessive GPU costs. Without bells and whistles, BT-Adapter achieves (1) state-of-the-art zero-shot results on various video tasks using thousands of fewer GPU hours. (2) better performance than current video chatbots without any video instruction tuning. (3) state-of-the-art results of video chatting using video instruction tuning, outperforming previous SOTAs by a large margin. The code has been available at https://github.com/farewellthreeIBT-Adapter. Ruyang Liu, Chen Li 0046, Yixiao Ge, Thomas H. Li, Ying Shan, Ge Li 0002 |
CVPR | 5 |
| 2024 | Programmable Motion Generation for Open-Set Motion Control TasksabstractCharacter animation in real-world scenarios necessitates a variety of constraints, such as trajectories, keyframes, interactions, etc. Existing methodologies typically treat single or a finite set of these constraint(s) as separate control tasks. These methods are often specialized, and the tasks they address are rarely extendable or customizable. We categorize these as solutions to the close-set motion control problem. In response to the complexity of practical motion control, we propose and attempt to solve the open-set motion control problem. This problem is characterized by an open and fully customizable set of motion control tasks. To address this, we introduce a new paradigm, programmable motion generation. In this paradigm, any given motion control task is broken down into a combination of atomic constraints. These constraints are then programmed into an error function that quantifies the degree to which a motion sequence adheres to them. We utilize a pretrained motion generation model and optimize its latent code to minimize the error function of the generated motion. Consequently, the generated motion not only inherits the prior of the generative model but also satisfies the requirements of the compounded constraints. Our experiments demonstrate that our approach can generate high-quality motions when addressing a wide range of unseen tasks. These tasks encompass motion control by motion dynamics, geometric constraints, physical laws, interactions with scenes, objects or the character's own body parts, etc. All of these are achieved in a unified approach, without the need for ad-hoc paired training data collection or specialized network designs. During the programming of novel tasks, we observed the emergence of new skills beyond those of the prior model. With the assistance of large language models, we also achieved automatic programming. We hope that this work will pave the way for the motion control of general AI agents. Xiaohang Zhan, Shaoli Huang, Tai-Jiang Mu, Ying Shan |
CVPR | 5 |
| 2024 | HumanGaussian: Text-Driven 3D Human Generation with Gaussian SplattingabstractRealistic 3D human generationfrom text prompts is a de-sirable yet challenging task. Existing methods optimize 3D representations like mesh or neural fields via score distil-lation sampling (SDS), which suffers from inadequate fine details or excessive training time. In this paper, we pro-pose an efficient yet effective framework, HumanGaussian, that generates high-quality 3D humans with fine-grained geometry and realistic appearance. Our key insight is that 3D Gaussian Splatting is an efficient renderer with peri-odic Gaussian shrinkage or growing, where such adaptive density control can be naturally guided by intrinsic human structures. Specifically, 1) we first propose a Structure-Aware SDS that simultaneously optimizes human appear-ance and geometry. The multi-modal score function from both RGB and depth space is leveraged to distill the Gaus-sian densification and pruning process. 2) Moreover, we devise an Annealed Negative Prompt Guidance by decom-posing SDS into a noisier generative score and a cleaner classifier score, which well addresses the over-saturation issue. The floating artifacts are further eliminated based on Gaussian size in a prune-only phase to enhance generation smoothness. Extensive experiments demonstrate the supe-rior efficiency and competitive quality of our framework, rendering vivid 3D humans under diverse scenarios. Xiaohang Zhan, Jiaxiang Tang, Ying Shan, Dahua Lin, Xihui Liu, Ziwei Liu 0002 |
CVPR | 4 |
| 2024 | DiffEditor: Boosting Accuracy and Flexibility on Diffusion-Based Image EditingabstractLarge-scale Text-to-Image (T2I) diffusion models have revolutionized image generation over the last few years. Although owning diverse and high-quality generation capabilities, translating these abilities to fine-grained Image editing remains challenging. In this paper, we propose DiffEditor to rectify two weaknesses in existing diffusion-based image editing: (1) in complex scenarios, editing results often lack editing accuracy and exhibit unexpected artifacts; (2) lack of flexibility to harmonize editing operations, e.g., imagine new content. In our solution, we introduce image prompts in fine-grained image editing, cooperating with the text prompt to better describe the editing content. To increase the flexibility while maintaining content consistency, we locally combine stochastic differential equation (SDE) into the ordinary differential equation (ODE) sampling. In addition, we incorporate regional score-based gradient guidance and a time travel strategy into the diffusion sampling, further improving the editing quality. Extensive experiments demonstrate that our method can efficiently achieve state-of-the-art performance on various fine-grained image editing tasks, including editing within a single image (e.g., object moving, resizing, and content dragging) and across images (e.g., appearance replacing and object pasting). Our source code is released at https://github.com/MC-E/DragonDiffusion. Chong Mou, Xintao Wang 0002, Jiechong Song, Ying Shan, Jian Zhang 0018 |
CVPR | 4 |
| 2024 | Low-Rank Approximation for Sparse Attention in Multi-Modal LLMsabstractThis paper focuses on the high computational complexity in Large Language Models (LLMs), a significant challenge in both natural language processing (NLP) and multi-modal tasks. We propose Low-Rank Approximation for Sparse Attention (LoRA -Sparse), an innovative approach that strategically reduces this complexity. LoRA -Sparse introduces low-rank linear projection layers for sparse attention approximation. It utilizes an order-mimic training methodology, which is crucial for efficiently approximating the self-attention mechanism in LLMs. We empirically show that sparse attention not only reduces computational demands, but also enhances model performance in both NLP and multi-modal tasks. This surprisingly shows that redundant attention in LLMs might be non-beneficial. We extensively validate LoRA -Sparse through rigorous empirical studies in both (NLP) and multi-modal tasks, demonstrating its effectiveness and general applicability. Based on LLaMA and LLaVA models, our methods can reduce more than half of the self-attention computation with even better performance than full-attention baselines. Lin Song 0002, Yukang Chen, Shuai Yang 0001, Xiaohan Ding, Yixiao Ge, Ying-Cong Chen, Ying Shan |
CVPR | 7 |
| 2024 | Multimodal Pathway: Improve Transformers with Irrelevant Data from Other ModalitiesabstractWe propose to improve transformers of a specific modality with irrelevant data from other modalities, e.g., improve an ImageNet model with audio or point cloud datasets. We would like to highlight that the data samples of the target modality are irrelevant to the other modalities, which distinguishes our method from other works utilizing paired (e.g., CLIP) or interleaved data of different modalities. We propose a methodology named Multimodal Pathway - given a target modality and a transformer designed for it, we use an auxiliary transformer trained with data of another modality and construct pathways to connect components of the two models so that data of the target modality can be processed by both models. In this way, we utilize the universal sequence-to-sequence modeling abilities of transformers obtained from two modalities. As a concrete implementation, we use a modality-specific tokenizer and task-specific head as usual but utilize the transformer blocks of the auxiliary model via a proposed method named Cross-Modal Re-parameterization, which exploits the auxiliary weights without any inference costs. On the image, point cloud, video, and audio recognition tasks, we observe significant and consistent performance improvements with irrelevant data from other modalities. The code and models are available at https://github.com/AILab-CVC/M2PT. Xiaohan Ding, Kaixiong Gong, Yixiao Ge, Ying Shan, Xiangyu Yue 0001 |
CVPR | 5 |
| 2024 | HumanRef: Single Image to 3D Human Generation via Reference-Guided DiffusionabstractGenerating a 3D human model from a single reference image is challenging because it requires inferring textures and geometries in invisible views while maintaining consistency with the reference image. Previous methods utilizing 3D generative models are limited by the availability of 3D training data. Optimization-based methods that lift text-to-image diffusion models to 3D generation often fail to preserve the texture details of the reference image, resulting in inconsistent appearances in different views. In this paper, we propose HumanRef, a 3D human generation framework from a single-view input. To ensure the generated 3D model is photorealistic and consistent with the input image, HumanRef introduces a novel method called reference-guided score distillation sampling (Ref-SDS), which effectively incorporates image guidance into the generation process. Furthermore, we introduce region-aware attention to Ref-SDS, ensuring accurate correspondence between different body regions. Experimental results demonstrate that HumanRef outper-forms state-of-the-art methods in generating 3D clothed humans with fine geometry, photorealistic textures, and view-consistent appearances. Code and model are available at https./reckcrtrhang.github.io/HumanRef.github.io/. Jingbo Zhang 0002, Xiaoyu Li 0002, Qi Zhang 0029, Yan-Pei Cao 0001, Ying Shan, Jing Liao 0001 |
CVPR | 5 |
| 2024 | DreamDiffusion: High-Quality EEG-to-Image Generation with Temporal Masked Signal Modeling and CLIP Alignment
Yunpeng Bai, Xintao Wang 0002, Yan-Pei Cao 0001, Yixiao Ge, Chun Yuan 0003, Ying Shan |
ECCV (31) | 6 |
| 2024 | Make a Cheap Scaling: A Self-Cascade Diffusion Model for Higher-Resolution Adaptation
Lanqing Guo, Yingqing He, Haoxin Chen, Menghan Xia, Xiaodong Cun, Yufei Wang 0006, Siyu Huang, Yong Zhang 0034, Xintao Wang 0002, Qifeng Chen 0001, Ying Shan, Bihan Wen |
ECCV (36) | 11 |
| 2024 | BrushNet: A Plug-and-Play Image Inpainting Model with Decomposed Dual-Branch Diffusion
Xuan Ju, Xintao Wang 0002, Yuxuan Bian, Ying Shan, Qiang Xu 0001 |
ECCV (20) | 5 |
| 2024 | ST-LLM: Large Language Models Are Effective Temporal Learners
Ruyang Liu, Chen Li 0046, Yixiao Ge, Ying Shan, Ge Li 0002 |
ECCV (57) | 5 |
| 2024 | EA-VTR: Event-Aware Video-Text Retrieval
Zongyang Ma, Ziqi Zhang 0010, Zhongang Qi, Chunfeng Yuan, Bing Li 0001, Yingmin Luo, Xu Li 0015, Xiaojuan Qi 0001, Ying Shan, Weiming Hu 0004 |
ECCV (52) | 10 |
| 2024 | MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model
Muyao Niu, Xiaodong Cun, Xintao Wang 0002, Yong Zhang 0034, Ying Shan, Yinqiang Zheng |
ECCV (19) | 5 |
| 2024 | DynamiCrafter: Animating Open-Domain Images with Video Diffusion Priors
Jinbo Xing, Menghan Xia, Yong Zhang 0034, Hao Chen 0011, Wangbo Yu, Hanyuan Liu, Gongye Liu, Xintao Wang 0002, Ying Shan, Tien-Tsin Wong |
ECCV (46) | 9 |
| 2024 | Texture-GS: Disentangling the Geometry and Texture for 3D Gaussian Splatting Editing
Tian-Xing Xu, Wenbo Hu 0002, Yukun Lai, Ying Shan, Song-Hai Zhang |
ECCV (25) | 4 |
| 2024 | Noise Calibration: Plug-and-Play Content-Preserving Video Enhancement Using Pre-trained Video Diffusion Models
Qinyu Yang, Hao Chen 0011, Yong Zhang 0034, Menghan Xia, Xiaodong Cun, Zhixun Su, Ying Shan |
ECCV (36) | 7 |
| 2024 | DMiT: Deformable Mipmapped Tri-Plane Representation for Dynamic Scenes
Jing-Wen Yang 0002, Jia-Mu Sun, Yongliang Yang 0002, Jie Yang 0038, Ying Shan, Yan-Pei Cao 0001, Lin Gao 0004 |
ECCV (55) | 5 |
| 2024 | HiFi-123: Towards High-Fidelity One Image to 3D Content Generation
Wangbo Yu, Li Yuan 0007, Yan-Pei Cao 0001, Xiangjun Gao, Xiaoyu Li 0002, Wenbo Hu 0002, Ying Shan, Yonghong Tian 0001 |
ECCV (73) | 8 |
| 2024 | Music Understanding LLaMA: Advancing Text-to-Music Generation with Question Answering and CaptioningabstractText-to-music generation (T2M-Gen) faces a major obstacle due to the scarcity of large-scale publicly available music datasets with natural language captions. To address this, we propose the Music Understanding LLaMA (MU-LLaMA), capable of answering music-related questions and generating captions for music files. Our model utilizes audio representations from a pretrained MERT model to extract music features. However, obtaining a suitable dataset for training the MU-LLaMA model remains challenging, as existing publicly accessible audio question answering datasets lack the necessary depth for open-ended music question answering. To fill this gap, we present a methodology for generating question-answer pairs from existing audio captioning datasets and introduce the MusicQA Dataset designed for answering open-ended music-related questions. The experiments demonstrate that the proposed MU-LLaMA model, trained on our designed MusicQA dataset, achieves outstanding performance in both music question answering and music caption generation across various metrics, outperforming current state-of-the-art (SOTA) models in both fields and offering a promising advancement in the T2M-Gen research field. Shansong Liu, Atin Sakkeer Hussain, Chenshuo Sun, Ying Shan |
ICASSP | 4 |
| 2024 | Humtrans: A Novel Open-Source Dataset for Humming Melody Transcription and BeyondabstractThis paper introduces the HumTrans dataset, which is publicly available and primarily designed for humming melody transcription. The dataset can also serve as a foundation for downstream tasks such as humming melody based music generation. It consists of 500 musical compositions of different genres and languages, with each composition divided into multiple segments. In total, the dataset comprises 1000 music segments. To collect this humming dataset, we employed 10 college students, all of whom are either music majors or proficient in playing at least one musical instrument. Each of them hummed every segment twice using the web recording interface provided by our designed website1. The humming recordings were sampled at a frequency of 44,100 Hz. During the humming session, the main interface provides a musical score for students to reference, with the melody audio playing simultaneously to aid in capturing both melody and rhythm. The dataset encompasses approximately 56.22 hours of audio, making it the largest known humming dataset to date. The dataset will be released on Hugging Face2, and we will provide a GitHub repository containing baseline results and evaluation codes3. Shansong Liu, Xu Li 0015, Ying Shan |
ICASSP | 4 |
| 2024 | Unified Pretraining Target Based Video-Music Retrieval with Music Rhythm and Video Optical Flow InformationabstractBackground music (BGM) can enhance the video’s emotion. However, selecting an appropriate BGM often requires domain knowledge. This has led to the development of video-music retrieval techniques. Most existing approaches utilize pretrained video/music feature extractors trained with different target sets to obtain average video/music-level embeddings. The drawbacks are two-fold. One is that different target sets for video/music pretraining may cause the generated embeddings difficult to match. The second is that the underlying temporal correlation between video and music is ignored. In this paper, our proposed approach leverages a unified target set to perform video/music pretraining and produces clip-level embeddings to preserve temporal information. The downstream crossmodal matching is based on the clip-level features with embedded music rhythm and optical flow information. Experiments demonstrate that our proposed method can achieve superior performance over the state-of-the-art methods by a significant margin. Tianjun Mao, Shansong Liu, Yunxuan Zhang, Ying Shan |
ICASSP | 5 |
| 2024 | Neural Concatenative Singing Voice Conversion: Rethinking Concatenation-Based Approach for One-Shot Singing Voice ConversionabstractAny-to-any singing voice conversion (SVC) is confronted with the challenge of "timbre leakage" issue caused by inadequate disentanglement between the content and the speaker timbre. To address this issue, this study introduces NeuCoSVC, a novel neural concatenative SVC framework. It consists of a self-supervised learning (SSL) representation extractor, a neural harmonic signal generator, and a waveform synthesizer. The SSL extractor condenses audio into fixed-dimensional SSL features, while the harmonic signal generator leverages linear time-varying filters to produce both raw and filtered harmonic signals for pitch information. The synthesizer reconstructs waveforms using SSL features, harmonic signals, and loudness information. During inference, voice conversion is performed by substituting source SSL features with their nearest counterparts from a matching pool which comprises SSL features extracted from the reference audio, while preserving raw harmonic signals and loudness from the source audio. By directly utilizing SSL features from the reference audio, the proposed framework effectively resolves the "timbre leakage" issue caused by previous disentanglement-based approaches. Experimental results demonstrate that the proposed NeuCoSVC system outperforms the disentanglement-based speaker embedding approach in one-shot SVC across intra-language, cross-language, and cross-domain evaluations. Binzhu Sha, Xu Li 0015, Zhiyong Wu 0001, Ying Shan, Helen M. Meng |
ICASSP | 4 |
| 2024 | Making LLaMA SEE and Draw with SEED TokenizerabstractThe great success of Large Language Models (LLMs) has expanded the potential of multimodality, contributing to the gradual evolution of General Artificial Intelligence (AGI). A true AGI agent should not only possess the capability to perform predefined multi-tasks but also exhibit emergent abilities in an open-world context. However, despite the considerable advancements made by recent multimodal LLMs, they still fall short in effectively unifying comprehension and generation tasks, let alone open-world emergent abilities. We contend that the key to overcoming the present impasse lies in enabling text and images to be represented and processed interchangeably within a unified autoregressive Transformer. To this end, we introduce $\textbf{SEED}$, an elaborate image tokenizer that empowers LLMs with the ability to $\textbf{SEE}$ and $\textbf{D}$raw at the same time. We identify two crucial design principles: (1) Image tokens should be independent of 2D physical patch positions and instead be produced with a $\textit{1D causal dependency}$, exhibiting intrinsic interdependence that aligns with the left-to-right autoregressive prediction mechanism in LLMs. (2) Image tokens should capture $\textit{high-level semantics}$ consistent with the degree of semantic abstraction in words, and be optimized for both discriminativeness and reconstruction during the tokenizer training phase. With SEED tokens, LLM is able to perform scalable multimodal autoregression under its original training recipe, i.e., next-word prediction. SEED-LLaMA is therefore produced by large-scale pretraining and instruction tuning on the interleaved textual and visual data, demonstrating impressive performance on a broad range of multimodal comprehension and generation tasks. More importantly, SEED-LLaMA has exhibited compositional emergent abilities such as multi-turn in-context multimodal generation, acting like your AI assistant. The code (training and inference) and models are released in https://github.com/AILab-CVC/SEED. Yuying Ge, Sijie Zhao, Ziyun Zeng, Yixiao Ge, Chen Li 0046, Xintao Wang 0002, Ying Shan |
ICLR | 7 |
| 2024 | ScaleCrafter: Tuning-free Higher-Resolution Visual Generation with Diffusion ModelsabstractIn this work, we investigate the capability of generating images from pre-trained diffusion models at much higher resolutions than the training image sizes. In addition, the generated images should have arbitrary image aspect ratios. When generating images directly at a higher resolution, 1024 x 1024, with the pre-trained Stable Diffusion using training images of resolution 512 x 512, we observe persistent problems of object repetition and unreasonable object structures. Existing works for higher-resolution generation, such as attention-based and joint-diffusion approaches, cannot well address these issues. As a new perspective, we examine the structural components of the U-Net in diffusion models and identify the crucial cause as the limited perception field of convolutional kernels. Based on this key observation, we propose a simple yet effective re-dilation that can dynamically adjust the convolutional perception field during inference. We further propose the dispersed convolution and noise-damped classifier-free guidance, which can enable ultra-high-resolution image generation (e.g., 4096 x 4096). Notably, our approach does not require any training or optimization. Extensive experiments demonstrate that our approach can address the repetition issue well and achieve state-of-the-art performance on higher-resolution image synthesis, especially in texture details. Our work also suggests that a pre-trained diffusion model trained on low-resolution images can be directly used for high-resolution visual generation without further tuning, which may provide insights for future research on ultra-high-resolution image and video synthesis. More results are available at the anonymous website: https://scalecrafter.github.io/ScaleCrafter/ Yingqing He, Shaoshu Yang, Haoxin Chen, Xiaodong Cun, Menghan Xia, Yong Zhang 0034, Xintao Wang 0002, Ran He 0001, Qifeng Chen 0001, Ying Shan |
ICLR | 10 |
| 2024 | DragonDiffusion: Enabling Drag-style Manipulation on Diffusion ModelsabstractDespite the ability of text-to-image (T2I) diffusion models to generate high-quality images, transferring this ability to accurate image editing remains a challenge. In this paper, we propose a novel image editing method, DragonDiffusion, enabling Drag-style manipulation on Diffusion models. Specifically, we treat image editing as the change of feature correspondence in a pre-trained diffusion model. By leveraging feature correspondence, we develop energy functions that align with the editing target, transforming image editing operations into gradient guidance. Based on this guidance approach, we also construct multi-scale guidance that considers both semantic and geometric alignment. Furthermore, we incorporate a visual cross-attention strategy based on a memory bank design to ensure consistency between the edited result and original image. Benefiting from these efficient designs, all content editing and consistency operations come from the feature correspondence without extra model fine-tuning. Extensive experiments demonstrate that our method has promising performance on various image editing tasks, including within a single image (e.g., object moving, resizing, and content dragging) or across images (e.g., appearance replacing and object pasting). Code is available at https://github.com/MC-E/DragonDiffusion. Chong Mou, Xintao Wang 0002, Jiechong Song, Ying Shan, Jian Zhang 0018 |
ICLR | 4 |
| 2024 | FreeNoise: Tuning-Free Longer Video Diffusion via Noise ReschedulingabstractWith the availability of large-scale video datasets and the advances of diffusion models, text-driven video generation has achieved substantial progress. However, existing video generation models are typically trained on a limited number of frames, resulting in the inability to generate high-fidelity long videos during inference. Furthermore, these models only support single-text conditions, whereas real-life scenarios often require multi-text conditions as the video content changes over time. To tackle these challenges, this study explores the potential of extending the text-driven capability to generate longer videos conditioned on multiple texts. 1) We first analyze the impact of initial noise in video diffusion models. Then building upon the observation of noise, we propose FreeNoise, a tuning-free and time-efficient paradigm to enhance the generative capabilities of pretrained video diffusion models while preserving content consistency. Specifically, instead of initializing noises for all frames, we reschedule a sequence of noises for long-range correlation and perform temporal attention over them by window-based fusion. 2) Additionally, we design a novel motion injection method to support the generation of videos conditioned on multiple text prompts. Extensive experiments validate the superiority of our paradigm in extending the generative capabilities of video diffusion models. It is noteworthy that compared with the previous best-performing method which brought about 255% extra time cost, our method incurs only negligible time cost of approximately 17%. Generated video samples are available at our website: http://haonanqiu.com/projects/FreeNoise.html. Haonan Qiu, Menghan Xia, Yong Zhang 0034, Yingqing He, Xintao Wang 0002, Ying Shan, Ziwei Liu 0002 |
ICLR | 6 |
| 2024 | TapMo: Shape-aware Motion Generation of Skeleton-free CharactersabstractPrevious motion generation methods are limited to the pre-rigged 3D human model, hindering their applications in the animation of various non-rigged characters. In this work, we present TapMo, a Text-driven Animation PIpeline for synthesizing Motion in a broad spectrum of skeleton-free 3D characters. The pivotal innovation in TapMo is its use of shape deformation-aware features as a condition to guide the diffusion model, thereby enabling the generation of mesh-specific motions for various characters. Specifically, TapMo comprises two main components - Mesh Handle Predictor and Shape-aware Diffusion Module. Mesh Handle Predictor predicts the skinning weights and clusters mesh vertices into adaptive handles for deformation control, which eliminates the need for traditional skeletal rigging. Shape-aware Motion Diffusion synthesizes motion with mesh-specific adaptations. This module employs text-guided motions and mesh features extracted during the first stage, preserving the geometric integrity of the animations by accounting for the character's shape and deformation. Trained in a weakly-supervised manner, TapMo can accommodate a multitude of non-human meshes, both with and without associated text motions. We demonstrate the effectiveness and generalizability of TapMo through rigorous qualitative and quantitative experiments. Our results reveal that TapMo consistently outperforms existing auto-animation methods, delivering superior-quality animations for both seen or unseen heterogeneous 3D characters. Shaoli Huang, Zhigang Tu 0001, Xin Chen 0040, Xiaohang Zhan, Gang Yu 0002, Ying Shan |
ICLR | 7 |
| 2024 | AFL-Net: Integrating Audio, Facial, and Lip Modalities with a Two-step Cross-attention for Robust Speaker Diarization in the Wild
Yongkang Yin, Xu Li 0015, Ying Shan, Yuexian Zou |
INTERSPEECH | 3 |
| 2024 | SynopGround: A Large-Scale Dataset for Multi-Paragraph Video Grounding from TV Dramas and Synopses
Chaolei Tan, Zihang Lin, Junfu Pu, Zhongang Qi, Wei-Yi Pei, Yexin Wang, Ying Shan, Wei-Shi Zheng 0001, Jianfang Hu |
ACM Multimedia | 8 |
| 2024 | CustomNet: Object Customization with Variable-Viewpoints in Text-to-Image Diffusion ModelsabstractIncorporating a customized object into image generation presents an attractive feature in text-to-image (T2I) generation. Some methods finetune T2I models for each object individually at test-time, which tend to be overfitted and time-consuming. Others train an extra encoder to extract object visual information for customization efficiently but struggle to preserve the object's identity. To address these limitations, we present CustomNet, a unified encoder-based object customization framework that explicitly incorporates 3D novel view synthesis capabilities into the customization process. This integration facilitates the adjustment of spatial positions and viewpoints, producing diverse outputs while effectively preserving the object's identity. To train our model effectively, we propose a dataset construction pipeline to better handle real-world objects and complex backgrounds. Additionally, we introduce delicate designs that enable location control and flexible background control through textual descriptions or user-defined backgrounds. Our method allows for object customization without the need of test-time optimization, providing simultaneous control over viewpoints, location, and text. Experimental results show that our method outperforms other customization methods regarding identity preservation, diversity, and harmony. Codes are available at https://github.com/TencentARC/CustomNet. Ziyang Yuan, Mingdeng Cao, Xintao Wang 0002, Zhongang Qi, Chun Yuan 0003, Ying Shan |
ACM Multimedia | 6 |
| 2024 | MiraData: A Large-Scale Video Dataset with Long Durations and Structured CaptionsabstractSora's high-motion intensity and long consistent videos have significantly impacted the field of video generation, attracting unprecedented attention. However, existing publicly available datasets are inadequate for generating Sora-like videos, as they mainly contain short videos with low motion intensity and brief captions. To address these issues, we propose MiraData, a high-quality video dataset that surpasses previous ones in video duration, caption detail, motion strength, and visual quality. We curate MiraData from diverse, manually selected sources and meticulously process the data to obtain semantically consistent clips. GPT-4V is employed to annotate structured captions, providing detailed descriptions from four different perspectives along with a summarized dense caption. To better assess temporal consistency and motion intensity in video generation, we introduce MiraBench, which enhances existing benchmarks by adding 3D consistency and tracking-based motion strength metrics. MiraBench includes 150 evaluation prompts and 17 metrics covering temporal consistency, motion strength, 3D consistency, visual quality, text-video alignment, and distribution similarity. To demonstrate the utility and effectiveness of MiraData, we conduct experiments using our DiT-based video generation model, MiraDiT. The experimental results on MiraBench demonstrate the superiority of MiraData, especially in motion strength. Xuan Ju, Yiming Gao 0007, Zhaoyang Zhang 0004, Ziyang Yuan, Xintao Wang 0002, Ailing Zeng, Qiang Xu 0001, Ying Shan |
NeurIPS | 9 |
| 2024 | E.T. Bench: Towards Open-Ended Event-Level Video-Language UnderstandingabstractRecent advances in Video Large Language Models (Video-LLMs) have demonstrated their great potential in general-purpose video understanding. To verify the significance of these models, a number of benchmarks have been proposed to diagnose their capabilities in different scenarios. However, existing benchmarks merely evaluate models through video-level question-answering, lacking fine-grained event-level assessment and task diversity. To fill this gap, we introduce E.T. Bench (Event-Level & Time-Sensitive Video Understanding Benchmark), a large-scale and high-quality benchmark for open-ended event-level video understanding. Categorized within a 3-level task taxonomy, E.T. Bench encompasses 7.3K samples under 12 tasks with 7K videos (251.4h total length) under 8 domains, providing comprehensive evaluations. We extensively evaluated 8 Image-LLMs and 12 Video-LLMs on our benchmark, and the results reveal that state-of-the-art models for coarse-level (video-level) understanding struggle to solve our fine-grained tasks, e.g., grounding event-of-interests within videos, largely due to the short video context length, improper time representations, and lack of multi-event training data. Focusing on these issues, we further propose a strong baseline model, E.T. Chat, together with an instruction-tuning dataset E.T. Instruct 164K tailored for fine-grained event-level understanding. Our simple but effective solution demonstrates superior performance in multiple scenarios. Ye Liu 0002, Zongyang Ma, Zhongang Qi, Yang Wu 0001, Ying Shan, Chang Wen Chen |
NeurIPS | 5 |
| 2024 | ReVideo: Remake a Video with Motion and Content ControlabstractDespite significant advancements in video generation and editing using diffusion models, achieving accurate and localized video editing remains a substantial challenge. Additionally, most existing video editing methods primarily focus on altering visual content, with limited research dedicated to motion editing. In this paper, we present a novel attempt to Remake a Video (ReVideo) which stands out from existing methods by allowing precise video editing in specific areas through the specification of both content and motion. Content editing is facilitated by modifying the first frame, while the trajectory-based motion control offers an intuitive user interaction experience. ReVideo addresses a new task involving the coupling and training imbalance between content and motion control. To tackle this, we develop a three-stage training strategy that progressively decouples these two aspects from coarse to fine. Furthermore, we propose a spatiotemporal adaptive fusion module to integrate content and motion control across various sampling steps and spatial locations. Extensive experiments demonstrate that our ReVideo has promising performance on several accurate video editing applications, i.e., (1) locally changing video content while keeping the motion constant, (2) keeping content unchanged and customizing new motion trajectories, (3) modifying both content and motion trajectories. Our method can also seamlessly extend these applications to multi-area editing without specific training, demonstrating its flexibility and robustness. Chong Mou, Mingdeng Cao, Xintao Wang 0002, Zhaoyang Zhang 0004, Ying Shan, Jian Zhang 0018 |
NeurIPS | 5 |
| 2024 | MambaTree: Tree Topology is All You Need in State Space ModelabstractThe state space models, employing recursively propagated features, demonstrate strong representation capabilities comparable to Transformer models and superior efficiency.
However, constrained by the inherent geometric constraints of sequences, it still falls short in modeling long-range dependencies.
To address this issue, we propose the MambaTree network, which first dynamically generates a tree topology based on spatial relationships and input features.
Then, feature propagation is performed based on this graph, thereby breaking the original sequence constraints to achieve stronger representation capabilities.
Additionally, we introduce a linear complexity dynamic programming algorithm to enhance long-range interactions without increasing computational cost.
MambaTree is a versatile multimodal framework that can be applied to both visual and textual tasks.
Extensive experiments demonstrate that our method significantly outperforms existing structured state space models on image classification, object detection and segmentation.
Besides, by fine-tuning large language models, our approach achieves consistent improvements in multiple textual tasks at minor training cost. Yicheng Xiao, Lin Song 0002, Shaoli Huang, Jiangshan Wang, Siyu Song, Yixiao Ge, Xiu Li 0001, Ying Shan |
NeurIPS | 8 |
| 2024 | CV-VAE: A Compatible Video VAE for Latent Generative Video ModelsabstractSpatio-temporal compression of videos, utilizing networks such as Variational Autoencoders (VAE), plays a crucial role in OpenAI's SORA and numerous other video generative models. For instance, many LLM-like video models learn the distribution of discrete tokens derived from 3D VAEs within the VQVAE framework, while most diffusion-based video models capture the distribution of continuous latent extracted by 2D VAEs without quantization. The temporal compression is simply realized by uniform frame sampling which results in unsmooth motion between consecutive frames. Currently, there lacks of a commonly used continuous video (3D) VAE for latent diffusion-based video models in the research community. Moreover, since current diffusion-based approaches are often implemented using pre-trained text-to-image (T2I) models, directly training a video VAE without considering the compatibility with existing T2I models will result in a latent space gap between them, which will take huge computational resources for training to bridge the gap even with the T2I models as initialization. To address this issue, we propose a method for training a video VAE of latent video models, namely CV-VAE, whose latent space is compatible with that of a given image VAE, e.g., image VAE of Stable Diffusion (SD). The compatibility is achieved by the proposed novel latent space regularization, which involves formulating a regularization loss using the image VAE. Benefiting from the latent space compatibility, video models can be trained seamlessly from pre-trained T2I or video models in a truly spatio-temporally compressed latent space, rather than simply sampling video frames at equal intervals. To improve the training efficiency, we also design a novel architecture for the video VAE. With our CV-VAE, existing video models can generate four times more frames with minimal finetuning. Extensive experiments are conducted to demonstrate the effectiveness of the proposed video VAE. Sijie Zhao, Yong Zhang 0034, Xiaodong Cun, Shaoshu Yang, Muyao Niu, Xiaoyu Li 0002, Wenbo Hu 0002, Ying Shan |
NeurIPS | 8 |
| 2024 | RecDCL: Dual Contrastive Learning for RecommendationabstractSelf-supervised learning (SSL) has recently achieved great success in mining the user-item interactions for collaborative filtering. As a major paradigm, contrastive learning (CL) based SSL helps address data sparsity in Web platforms by contrasting the embeddings between raw and augmented data. However, existing CL-based methods mostly focus on contrasting in a batch-wise way, failing to exploit potential regularity in the feature dimension. This leads to redundant solutions during the representation learning of users and items. In this work, we investigate how to employ both batch-wise CL (BCL) and feature-wise CL (FCL) for recommendation. We theoretically analyze the relation between BCL and FCL, and find that combining BCL and FCL helps eliminate redundant solutions but never misses an optimal solution. We propose a dual contrastive learning recommendation framework---RecDCL. In RecDCL, the FCL objective is designed to eliminate redundant solutions on user-item positive pairs and to optimize the uniform distributions within users and items using a polynomial kernel for driving the representations to be orthogonal; The BCL objective is utilized to generate contrastive embeddings on output vectors for enhancing the robustness of the representations. Extensive experiments on four widely-used benchmarks and one industry dataset demonstrate that RecDCL can consistently outperform the state-of-the-art GNNs-based and SSL-based models (with an improvement of up to 5.65% in terms of Recall@20). The source code is publicly available https://github.com/THUDM/RecDCL Zhongang Qi, Zhiyu Chen 0014, Xing Tang 0007, Ying Shan, Yuxiao Dong, Jie Tang 0001 |
WWW | 7 |
| 2024 | Learning layout generation for virtual worldsabstractThe emergence of the metaverse has led to the rapidly increasing demand for the generation of extensive 3D worlds. We consider that an engaging world is built upon a rational layout of multiple landuse areas (e.g., forest, meadow, and farmland). To this end, we propose a generative model of landuse distribution that learns from geographic data. The model is based on a transformer architecture that generates a 2D map of the land-use layout, which can be conditioned on spatial and semantic controls, depending on whether either one or both are provided. This model enables diverse layout generation with user control and layout expansion by extending borders with partial inputs. To generate high-quality and satisfactory layouts, we devise a geometric objective function that supervises the model to perceive layout shapes and regularize generations using geometric priors. Additionally, we devise a planning objective function that supervises the model to perceive progressive composition demands and suppress generations deviating from controls. To evaluate the spatial distribution of the generations, we train an autoencoder to embed land-use layouts into vectors to enable comparison between the real and generated data using the Wasserstein metric, which is inspired by the Frechet inception distance. Weihao Cheng 0002, Ying Shan |
Comput. Vis. Media | 2 |
| 2024 | NeRF-Texture: Synthesizing Neural Radiance Field TexturesabstractTexture synthesis is a fundamental problem in computer graphics that would benefit various applications. Existing methods are effective in handling 2D image textures. In contrast, many real-world textures contain meso-structure in the 3D geometry space, such as grass, leaves, and fabrics, which cannot be effectively modeled using only 2D image textures. We propose a novel texture synthesis method with Neural Radiance Fields (NeRF) to capture and synthesize textures from given multi-view images. In the proposed NeRF texture representation, a scene with fine geometric details is disentangled into the meso-structure textures and the underlying base shape. This allows textures with meso-structure to be effectively learned as latent features situated on the base shape, which are fed into a NeRF decoder trained simultaneously to represent the rich view-dependent appearance. Using this implicit representation, we can synthesize NeRF-based textures through patch matching of latent features. However, inconsistencies between the metrics of the reconstructed content space and the latent feature space may compromise the synthesis quality. To enhance matching performance, we further regularize the distribution of latent features by incorporating a clustering constraint. In addition to generating NeRF textures over a planar domain, our method can also synthesize NeRF textures over curved surfaces, which are practically useful. Experimental results and evaluations demonstrate the effectiveness of our approach. Yihua Huang 0002, Yan-Pei Cao 0001, Yukun Lai, Ying Shan, Lin Gao 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Empowering Real-World Image Super-Resolution With Flexible Interactive ModulationabstractInteractive image restoration aims to construct an interactive pathway between users and restoration networks, which empowers users to modulate the restoration results according to their own demands. However, existing methods are primarily limited to training their networks with predefined and simplistic synthetic degradations. Consequently, these methods often encounter significant performance degradation when confronted with real-world degradations that deviate from their assumptions. Furthermore, existing interactive image restoration approaches solely support global modulation, wherein a single modulation factor governs the reconstruction process for the entire image. In this paper, we propose a novel method to perform real-world and intricate image super-resolution in an interactive manner. Specifically, we propose a metric-learning-based degradation estimation strategy to estimate not only the overall degradation level of the entire image but also the finer-grained, pixel-wise degradation within real-world scenarios. This enables local control over the restoration results by selectively modulating the corresponding regions based on the densely-estimated degradation map. Additionally, a new metric-argumented loss is proposed to further enhance the performance of real-world image super-resolution. Through extensive experimentation, we demonstrate the efficacy of our method in achieving exceptional modulation and restoration performance in real-world image super-resolution tasks, all while maintaining an appealing model complexity. Chong Mou, Xintao Wang 0002, Yanze Wu, Ying Shan, Jian Zhang 0018 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Chinese Title Generation for Short Videos: Dataset, Metric and AlgorithmabstractPrevious work for video captioning aims to objectively describe the video content but the captions lack human interest and attractiveness, limiting its practical application scenarios. The intention of video title generation (video titling) is to produce attractive titles, but there is a lack of benchmarks. This work offers CREATE, the first large-scale Chinese shoRt vidEo retrievAl and Title gEneration dataset, to assist research and applications in video titling, video captioning, and video retrieval in Chinese. CREATE comprises a high-quality labeled 210 K dataset and two web-scale 3 M and 10 M pre-training datasets, covering 51 categories, 50K+ tags, 537K+ manually annotated titles and captions, and 10M+ short videos with original video information. This work presents ACTEr, a unique Attractiveness-Consensus-based Title Evaluation, to objectively evaluate the quality of video title generation. This metric measures the semantic correlation between the candidate (model-generated title) and references (manual-labeled titles) and introduces attractive consensus weights to assess the attractiveness and relevance of the video title. Accordingly, this work proposes a novel multi-modal ALignment WIth Generation model, ALWIG, as one strong baseline to aid future model development. With the help of a tag-driven video-text alignment module and a GPT-based generation module, this model achieves video titling, captioning, and retrieval simultaneously. We believe that the release of the CREATE dataset, ACTEr metric, and ALWIG model will encourage in-depth research on the analysis and creation of Chinese short videos. Ziqi Zhang 0010, Zongyang Ma, Chunfeng Yuan, Peijin Wang, Zhongang Qi, Chenglei Hao, Bing Li 0001, Ying Shan, Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2024 | DARTScore: DuAl-Reconstruction Transformer for Video Captioning EvaluationabstractVideo captioning evaluation aims at assessing the semantic consistency between video and candidate text, which should include measurement from two aspects: faithfulness (whether the information conveyed by candidate is correct w.r.t. video) and comprehensiveness (whether the main video content is covered by candidate). However, previous approaches have difficulty in evaluating faithfulness and comprehensiveness due to heavy reliance on references or heterogeneous of visual and textual data. In this paper, we propose a vision-involved evaluation metric based on a novel DuAl-Reconstruction Transformer, named DARTScore. DARTScore formulates the caption evaluation task as a dual-reconstruction problem to evaluate both faithfulness and comprehensiveness explicitly. Since the word in a candidate is usually related to several frames, DARTScore adaptively collects relevant frames to reconstruct the word and computes the reconstruction accuracy as faithfulness to inherently reflect whether the word information is contained in the video. In the inversive way, DARTScore reconstructs each frame with relevant words to evaluate comprehensiveness. By integrating fine-grained bidirectional reconstruction accuracies, DARTScore drills into each word in candidate and each frame in video to fully evaluate the semantic consistency. Furthermore, we collect and annotate two Chinese datasets with a large domain gap, named CRAETE-EVAL and VATEX-ZH-EVAL, to systematically evaluate existing metrics and fill the blank of Chinese video captioning evaluation. Experimental results show that DARTScore achieves higher correlation with human judgments, has lower reference reliance, and generalizes well to data from different domains. Ziqi Zhang 0010, Zhongang Qi, Chunfeng Yuan, Ying Shan, Bing Li 0001, Weiming Hu 0004, Xiaohu Qie |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Analysis and Benchmarking of Extending Blind Face Image Restoration to VideosabstractRecent progress in blind face restoration has resulted in producing high-quality restored results for static images. However, efforts to extend these advancements to video scenarios have been minimal, partly because of the absence of benchmarks that allow for a comprehensive and fair comparison. In this work, we first present a fair evaluation benchmark, in which we first introduce a Real-world Low-Quality Face Video benchmark (RFV-LQ), evaluate several leading image-based face restoration algorithms, and conduct a thorough systematical analysis of the benefits and challenges associated with extending blind face image restoration algorithms to degraded face videos. Our analysis identifies several key issues, primarily categorized into two aspects: significant jitters in facial components and noise-shape flickering between frames. To address these issues, we propose a Temporal Consistency Network (TCN) cooperated with alignment smoothing to reduce jitters and flickers in restored videos. TCN is a flexible component that can be seamlessly plugged into the most advanced face image restoration algorithms, ensuring the quality of image-based restoration is maintained as closely as possible. Extensive experiments have been conducted to evaluate the effectiveness and efficiency of our proposed TCN and alignment smoothing operation. Zhouxia Wang, Jiawei Zhang 0002, Xintao Wang 0002, Tianshui Chen, Ying Shan, Wenping Wang 0001, Ping Luo 0002 |
IEEE Trans. Image Process. | 5 |
| 2024 | DropConn: Dropout Connection Based Random GNNs for Molecular Property PredictionabstractRecently, molecular data mining has attracted a lot of attention owing to its great application potential in material and drug discovery. However, this mining task faces a challenge posed by the scarcity of labeled molecular graphs. To overcome this challenge, we introduce a novel data augmentation and a semi-supervised confidence-aware consistency regularization training framework for molecular property prediction. The core of our framework is a data augmentation strategy on molecular graphs, named DropConn (Dropout Connection). DropConn generates pseudo molecular graphs by softening the hard connections of chemical bonds (as edges), where the soft weights are calculated from edge features so that the adaptive interactions between different atoms can be incorporated. Besides, to enhance the model's generalization ability, a consistency regularization training strategy is proposed to take full advantage of massive unlabeled data. Furthermore, DropConn can serve as a plugin that can be seamlessly added to many existing models. Extensive experiments under both non-pre-training setting and fine-tuning setting demonstrate that DropConn can obtain superior performance (up to 8.22%) over state-of-the-art methods on molecular property prediction tasks. The code is available athttps://github.com/THUDM/DropConn. Wenzheng Feng, Yuandong Wang 0002, Zhongang Qi, Ying Shan, Jie Tang 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2024 | Learning Music-Dance Representations Through Explicit-Implicit Rhythm SynchronizationabstractAlthough audio-visual representation has been proven to be applicable in many downstream tasks, the representation of dancing videos, which is more specific and always accompanied by music with complex auditory contents, remains challenging and uninvestigated. Considering the intrinsic alignment between the cadent movement of the dancer and music rhythm, we introduceMuDaR, a novelMusic-DanceRepresentation learning framework to perform the synchronization of music and dance rhythms both in explicit and implicit ways. Specifically, we derive the dance rhythms based on visual appearance and motion cues inspired by the music rhythm analysis. Then the visual rhythms are temporally aligned with the music counterparts, which are extracted by the amplitude of sound intensity. Meanwhile, we exploit the implicit coherence of rhythms implied in audio and visual streams by contrastive learning. The model learns the joint embedding by predicting the temporal consistency between audio-visual pairs. The music-dance representation, together with the capability of detecting audio and visual rhythms, can further be applied to three downstream tasks: (a) dance classification, (b) music-dance retrieval, and (c) music-dance retargeting. Extensive experiments demonstrate that our proposed framework outperforms other self-supervised methods by a large margin. Jiashuo Yu, Junfu Pu, Ying Cheng 0005, Rui Feng 0001, Ying Shan |
IEEE Trans. Multim. | 5 |
| 2024 | StyleCrafter: Taming Artistic Video Diffusion with Reference-Augmented Adapter LearningabstractText-to-video (T2V) models have shown remarkable capabilities in generating diverse videos. However, they struggle to produce user-desired artistic videos due to (i) text's inherent clumsiness in expressing specific styles and (ii) the generally degraded style fidelity. To address these challenges, we introduce StyleCrafter, a generic method that enhances pretrained T2V models with a style control adapter, allowing video generation in any style by feeding a reference image. Considering the scarcity of artistic video data, we propose to first train a style control adapter using style-rich image datasets, then transfer the learned stylization ability to video generation through a tailor-made finetuning paradigm. To promote content-style disentanglement, we employ carefully designed data augmentation strategies to enhance decoupled learning. Additionally, we propose a scale-adaptive fusion module to balance the influences of text-based content features and image-based style features, which helps generalization across various text and style combinations. StyleCrafter efficiently generates high-quality stylized videos that align with the content of the texts and resemble the style of the reference images. Experiments demonstrate that our approach is more flexible and efficient than existing competitors. Project page: https://gongyeliu.github.io/StyleCrafter.github.io/ Gongye Liu, Menghan Xia, Yong Zhang 0034, Haoxin Chen, Jinbo Xing, Yibo Wang 0039, Xintao Wang 0002, Ying Shan, Yujiu Yang 0001 |
ACM Trans. Graph. | 8 |
| 2024 | ToonCrafter: Generative Cartoon InterpolationabstractWe introduce ToonCrafter, a novel approach that transcends traditional correspondence-based cartoon video interpolation, paving the way for generative interpolation. Traditional methods, that implicitly assume linear motion and the absence of complicated phenomena like dis-occlusion, often struggle with the exaggerated non-linear and large motions with occlusion commonly found in cartoons, resulting in implausible or even failed interpolation results. To overcome these limitations, we explore the potential of adapting live-action video priors to better suit cartoon interpolation within a generative framework. ToonCrafter effectively addresses the challenges faced when applying live-action video motion priors to generative cartoon interpolation. First, we design a toon rectification learning strategy that seamlessly adapts live-action video priors to the cartoon domain, resolving the domain gap and content leakage issues. Next, we introduce a dual-reference-based 3D decoder to compensate for lost details due to the highly compressed latent prior spaces, ensuring the preservation of fine details in interpolation results. Finally, we design a flexible sketch encoder that empowers users with interactive control over the interpolation results. Experimental results demonstrate that our proposed method not only produces visually convincing and more natural dynamics, but also effectively handles dis-occlusion. The comparative evaluation demonstrates the notable superiority of our approach over existing competitors. Code and model weights are available at https://doubiiu.github.io/projects/ToonCrafter Jinbo Xing, Hanyuan Liu, Menghan Xia, Yong Zhang 0034, Xintao Wang 0002, Ying Shan, Tien-Tsin Wong |
ACM Trans. Graph. | 6 |
| 2024 | TIP-Editor: An Accurate 3D Editor Following Both Text-Prompts And Image-PromptsabstractText-driven 3D scene editing has gained significant attention owing to its convenience and user-friendliness. However, existing methods still lack accurate control of the specified appearance and location of the editing result due to the inherent limitations of the text description. To this end, we propose a 3D scene editing framework, TIP-Editor, that accepts both text and image prompts and a 3D bounding box to specify the editing region. With the image prompt, users can conveniently specify the detailed appearance/style of the target content in complement to the text description, enabling accurate control of the appearance. Specifically, TIP-Editor employs a stepwise 2D personalization strategy to better learn the representation of the existing scene and the reference image, in which a localization loss is proposed to encourage correct object placement as specified by the bounding box. Additionally, TIP-Editor utilizes explicit and flexible 3D Gaussian splatting (GS) as the 3D representation to facilitate local editing while keeping the background unchanged. Extensive experiments have demonstrated that TIP-Editor conducts accurate editing following the text and image prompts in the specified bounding box region, consistently outperforming the baselines in editing quality, and the alignment to the prompts, qualitatively and quantitatively. Jingyu Zhuang, Yan-Pei Cao 0001, Guanbin Li, Liang Lin 0004, Ying Shan |
ACM Trans. Graph. | 6 |
| 2023 | Tagging before Alignment: Integrating Multi-Modal Tags for Video-Text RetrievalabstractVision-language alignment learning for video-text retrieval arouses a lot of attention in recent years. Most of the existing methods either transfer the knowledge of image-text pretraining model to video-text retrieval task without fully exploring the multi-modal information of videos, or simply fuse multi-modal features in a brute force manner without explicit guidance. In this paper, we integrate multi-modal information in an explicit manner by tagging, and use the tags as the anchors for better video-text alignment. Various pretrained experts are utilized for extracting the information of multiple modalities, including object, person, motion, audio, etc. To take full advantage of these information, we propose the TABLE (TAgging Before aLignmEnt) network, which consists of a visual encoder, a tag encoder, a text encoder, and a tag-guiding cross-modal encoder for jointly encoding multi-frame visual features and multi-modal tags information. Furthermore, to strengthen the interaction between video and text, we build a joint cross-modal encoder with the triplet input of [vision, tag, text] and perform two additional supervised tasks, Video Text Matching (VTM) and Masked Language Modeling (MLM). Extensive experimental results demonstrate that the TABLE model is capable of achieving State-Of-The-Art (SOTA) performance on various video-text retrieval benchmarks, including MSR-VTT, MSVD, LSMDC and DiDeMo. Yizhen Chen, Lijian Lin, Zhongang Qi, Jin Ma 0003, Ying Shan |
AAAI | 6 |
| 2023 | Accelerating the Training of Video Super-resolution ModelsabstractDespite that convolution neural networks (CNN) have recently demonstrated high-quality reconstruction for video super-resolution (VSR), efficiently training competitive VSR models remains a challenging problem. It usually takes an order of magnitude more time than training their counterpart image models, leading to long research cycles. Existing VSR methods typically train models with fixed spatial and temporal sizes from beginning to end. The fixed sizes are usually set to large values for good performance, resulting to slow training. However, is such a rigid training strategy necessary for VSR? In this work, we show that it is possible to gradually train video models from small to large spatial/temporal sizes, \ie, in an easy-to-hard manner. In particular, the whole training is divided into several stages and the earlier stage has smaller training spatial shape. Inside each stage, the temporal size also varies from short to long while the spatial size remains unchanged. Training is accelerated by such a multigrid training strategy, as most of computation is performed on smaller spatial and shorter temporal shapes. For further acceleration with GPU parallelization, we also investigate the large minibatch training without the loss in accuracy. Extensive experiments demonstrate that our method is capable of largely speeding up training (up to $6.2\times$ speedup in wall-clock training time) without performance drop for various VSR models. Lijian Lin, Xintao Wang 0002, Zhongang Qi, Ying Shan |
AAAI | 4 |
| 2023 | Mitigating Artifacts in Real-World Video Super-resolution ModelsabstractThe recurrent structure is a prevalent framework for the task of video super-resolution, which models the temporal dependency between frames via hidden states. When applied to real-world scenarios with unknown and complex degradations, hidden states tend to contain unpleasant artifacts and propagate them to restored frames. In this circumstance, our analyses show that such artifacts can be largely alleviated when the hidden state is replaced with a cleaner counterpart. Based on the observations, we propose a Hidden State Attention (HSA) module to mitigate artifacts in real-world video super-resolution. Specifically, we first adopt various cheap filters to produce a hidden state pool. For example, Gaussian blur filters are for smoothing artifacts while sharpening filters are for enhancing details. To aggregate a new hidden state that contains fewer artifacts from the hidden state pool, we devise a Selective Cross Attention (SCA) module, in which the attention between input features and each hidden state is calculated. Equipped with HSA, our proposed method, namely FastRealVSR, is able to achieve 2x speedup while obtaining better performance than Real-BasicVSR. Codes will be available at https://github.com/TencentARC/FastRealVSR. Liangbin Xie, Xintao Wang 0002, Shuwei Shi, Jinjin Gu, Chao Dong 0005, Ying Shan |
AAAI | 6 |
| 2023 | What Does Your Face Sound Like? 3D Face Shape towards VoiceabstractFace-based speech synthesis provides a practical solution to generate voices from human faces. However, directly using 2D face images leads to the problems of uninterpretability and entanglement. In this paper, to address the issues, we introduce 3D face shape which (1) has an anatomical relationship between voice characteristics, partaking in the "bone conduction" of human timbre production, and (2) is naturally independent of irrelevant factors by excluding the blending process. We devise a three-stage framework to generate speech from 3D face shapes. Fully considering timbre production in anatomical and acquired terms, our framework incorporates three additional relevant attributes including face texture, facial features, and demographics. Experiments and subjective tests demonstrate our method can generate utterances matching faces well, with good audio quality and voice diversity. We also explore and visualize how the voice changes with the face. Case studies show that our method upgrades the face-voice inference to personalized custom-made voice creating, revealing a promising prospect in virtual human and dubbing applications. Zhiyong Wu 0001, Ying Shan, Jia Jia 0001 |
AAAI | 3 |
| 2023 | Darwinian Model Upgrades: Model Evolving with Selective CompatibilityabstractThe traditional model upgrading paradigm for retrieval requires recomputing all gallery embeddings before deploying the new model (dubbed as "backfilling"), which is quite expensive and time-consuming considering billions of instances in industrial applications. BCT presents the first step towards backward-compatible model upgrades to get rid of backfilling. It is workable but leaves the new model in a dilemma between new feature discriminativeness and new-to-old compatibility due to the undifferentiated compatibility constraints. In this work, we propose Darwinian Model Upgrades (DMU), which disentangle the inheritance and variation in the model evolving with selective backward compatibility and forward adaptation, respectively. The old-to-new heritable knowledge is measured by old feature discriminativeness, and the gallery features, especially those of poor quality, are evolved in a lightweight manner to become more adaptive in the new latent space. We demonstrate the superiority of DMU through comprehensive experiments on large-scale landmark retrieval and face recognition benchmarks. DMU effectively alleviates the new-to-new degradation at the same time improving new-to-old compatibility, rendering a more proper model upgrading paradigm in large-scale retrieval systems.Code: https://github.com/TencentARC/OpenCompatible. Binjie Zhang, Shupeng Su, Yixiao Ge, Xuyuan Xu, Yexin Wang, Chun Yuan 0003, Zheng Shou 0001, Ying Shan |
AAAI | 8 |
| 2023 | DSRM: Boost Textual Adversarial Training with Distribution Shift Risk MinimizationabstractSongYang Gao, Shihan Dou, Yan Liu, Xiao Wang, Qi Zhang, Zhongyu Wei, Jin Ma, Ying Shan. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Songyang Gao, Shihan Dou, Yan Liu 0002, Xiao Wang 0001, Qi Zhang 0001, Zhongyu Wei, Jin Ma 0003, Ying Shan |
ACL (1) | 8 |
| 2023 | SurfelNeRF: Neural Surfel Radiance Fields for Online Photorealistic Reconstruction of Indoor ScenesabstractOnline reconstructing and rendering of large-scale indoor scenes is a long-standing challenge. SLAM-based methods can reconstruct 3D scene geometry progressively in real time but can not render photorealistic results. While NeRF-based methods produce promising novel view synthesis results, their long offline optimization time and lack of geometric constraints pose challenges to efficiently handling online input. Inspired by the complementary advantages of classical 3D reconstruction and NeRF, we thus investigate marrying explicit geometric representation with NeRF rendering to achieve efficient online reconstruction and high-quality rendering. We introduce SurfelNeRF, a variant of neural radiance field which employs a flexible and scalable neural surfel representation to store geometric attributes and extracted appearance features from input images. We further extend the conventional surfel-based fusion scheme to progressively integrate incoming input frames into the reconstructed global neural scene representation. In addition, we propose a highly-efficient differentiable rasterization scheme for rendering neural surfel radiance fields, which helps SurfelNeRF achieve 10× speedups in training and inference time, respectively. Experimental results show that our method achieves the state-of-the-art 23.82 PSNR and 29.58 PSNR on ScanNet in feedforward inference and perscene optimization settings, respectively.11Project website: https://gymat.github.io/SurfelNeRF-web Yiming Gao 0007, Yan-Pei Cao 0001, Ying Shan |
CVPR | 3 |
| 2023 | HRDFuse: Monocular 360° Depth Estimation by Collaboratively Learning Holistic-with-Regional Depth DistributionsabstractDepth estimation from a monocular 360° image is a burgeoning problem owing to its holistic sensing of a scene. Recently, some methods, e.g., OmniFusion, have applied the tangent projection (TP) to represent a 360° image and predicted depth values via patch-wise regressions, which are merged to get a depth map with equirectangular projection (ERP) format. However, these methods suffer from 1) non-trivial process of merging plenty of patches; 2) capturing less holistic-with-regional contextual information by directly regressing the depth value of each pixel. In this paper, we propose a novel framework, HRDFuse, that subtly combines the potential of convolutional neural networks (CNNs) and transformers by collaboratively learning the holistic contextual information from the ERP and the regional structural information from the TP. Firstly, we propose a spatial feature alignment (SFA) module that learns feature similarities between the TP and ERP to aggregate the TP features into a complete ERP feature map in a pixelwise manner. Secondly, we propose a collaborative depth distribution classification (CDDC) module that learns the holistic-with-regional histograms capturing the ERP and TP depth distributions. As such, the final depth values can be predicted as a linear combination of histogram bin centers. Lastly, we adaptively combine the depth predictions from ERP and TP to obtain the final depth map. Extensive experiments show that our method predicts more smooth and accurate depth results while achieving favorably better results than the SOTA methods. Hao Ai, Zidong Cao, Yan-Pei Cao 0001, Ying Shan, Lin Wang 0025 |
CVPR | 4 |
| 2023 | High-fidelity Facial Avatar Reconstruction from Monocular Video with Generative PriorsabstractHigh-fidelity facial avatar reconstruction from a monocular video is a significant research problem in computer graphics and computer vision. Recently, Neural Radiance Field (NeRF) has shown impressive novel view rendering results and has been considered for facial avatar reconstruction. However, the complex facial dynamics and missing 3D information in monocular videos raise significant challenges for faithful facial reconstruction. In this work, we propose a new method for NeRF-based facial avatar reconstruction that utilizes 3D-aware generative prior. Different from existing works that depend on a conditional deformation field for dynamic modeling, we propose to learn a personalized generative prior, which is formulated as a local and low dimensional subspace in the latent space of 3D-GAN. We propose an efficient method to construct the personalized generative prior based on a small set of facial images of a given individual. After learning, it allows for photo-realistic rendering with novel views, and the face reenactment can be realized by performing navigation in the latent space. Our proposed method is applicable for different driven signals, including RGB images, 3DMM coefficients, and audio. Compared with existing works, we obtain superior novel view synthesis results and faithfully face reenactment performance. The code is available here https://github.com/bbaaii/HFA-GP. Yunpeng Bai, Yanbo Fan, Xuan Wang 0009, Yong Zhang 0034, Jingxiang Sun, Chun Yuan 0003, Ying Shan |
CVPR | 7 |
| 2023 | Local-to-Global Registration for Bundle-Adjusting Neural Radiance FieldsabstractNeural Radiance Fields (NeRF) have achieved photorealistic novel views synthesis; however, the requirement of accurate camera poses limits its application. Despite analysis-by-synthesis extensions for jointly learning neural3D representations and registering camera frames exist, they are susceptible to suboptimal solutions if poorly initialized. We propose L2G-NeRF, a Local-to-Global registration method for bundle-adjusting Neural Radiance Fields: first, a pixel-wise flexible alignment, followed by a framewise constrained parametric alignment. Pixel-wise local alignment is learned in an unsupervised way via a deep network which optimizes photometric reconstruction errors. framewise global alignment is performed using differentiable parameter estimation solvers on the pixel-wise correspondences to find a global transformation. Experiments on synthetic and real-world data show that our method outperforms the current state-of-the-art in terms of high-fidelity reconstruction and resolving large camera pose misalignment. Our module is an easy-to-use plugin that can be applied to NeRF variants and other neural field applications. The Code and supplementary materials are available at https://rover-xingyu.github.io/L2G-NeRF/. Xuan Wang 0009, Qi Zhang 0029, Yu Guo 0006, Ying Shan, Fei Wang 0008 |
CVPR | 6 |
| 2023 | ViLEM: Visual-Language Error Modeling for Image-Text RetrievalabstractDominant pre-training works for image-text retrieval adopt “dual-encoder” architecture to enable high efficiency, where two encoders are used to extract image and text representations and contrastive learning is employed for global alignment. However, coarse-grained global alignment ignores detailed semantic associations between image and text. In this work, we propose a novel proxy task, named Visual-Language Error Modeling (ViLEM), to inject detailed image-text association into “dual-encoder” model by “proofreading” each word in the text against the corresponding image. Specifically, we first edit the image-paired text to automatically generate diverse plausible negative texts with pre-trained language models. ViLEM then enforces the model to discriminate the correctness of each word in the plausible negative texts and further correct the wrong words via resorting to image information. Further-more, we propose a multi-granularity interaction framework to perform ViLEM via interacting text features with both global and local image features, which associates local text semantics with both high-level visual context and multi-level local visual information. Our method surpasses state-of-the-art “dual-encoder” methods by a large margin on the image-text retrieval task and significantly improves discriminativeness to local textual semantics. Our model can also generalize well to video-text retrieval. Zongyang Ma, Ziqi Zhang 0010, Zhongang Qi, Chunfeng Yuan, Ying Shan, Bing Li 0001, Weiming Hu 0004, Xiaohu Qie |
CVPR | 6 |
| 2023 | Improved Test-Time Adaptation for Domain GeneralizationabstractThe main challenge in domain generalization (DG) is to handle the distribution shift problem that lies between the training and test data. Recent studies suggest that test-time training (TTT), which adapts the learned model with test data, might be a promising solution to the problem. Generally, a TTT strategy hinges its performance on two main factors: selecting an appropriate auxiliary TTT task for updating and identifying reliable parameters to update during the test phase. Both previous arts and our experiments indicate that TTT may not improve but be detrimental to the learned model if those two factors are not properly considered. This work addresses those two factors by proposing an Improved Test-Time Adaptation (ITTA) method. First, instead of heuristically defining an auxiliary objective, we propose a learnable consistency loss for the TTT task, which contains learnable parameters that can be adjusted toward better alignment between our TTT task and the main prediction task. Second, we introduce additional adaptive parameters for the trained model, and we suggest only updating the adaptive parameters during the test phase. Through extensive experiments, we show that the proposed two strategies are beneficial for the learned model (see Figure 1), and ITTA could achieve superior performance to the current state-of-the-art methods on several DG benchmarks. Code is available at https://github.com/liangchen527/ITTA. Liang Chen 0030, Yong Zhang 0034, Yibing Song, Ying Shan, Lingqiao Liu |
CVPR | 4 |
| 2023 | DPE: Disentanglement of Pose and Expression for General Video Portrait EditingabstractOne-shot video-driven talking face generation aims at producing a synthetic talking video by transferring the facial motion from a video to an arbitrary portrait image. Head pose and facial expression are always entangled in facial motion and transferred simultaneously. However, the entanglement sets up a barrier for these methods to be used in video portrait editing directly, where it may require to modify the expression only while maintaining the pose unchanged. One challenge of decoupling pose and expression is the lack of paired data, such as the same pose but different expressions. Only a few methods attempt to tackle this challenge with the feat of 3D Morphable Models (3DMMs) for explicit disentanglement. But 3DMMs are not accurate enough to capture facial details due to the limited number of Blend-shapes, which has side effects on motion transfer. In this paper, we introduce a novel self-supervised disentanglement framework to decouple pose and expression without 3DMMs and paired data, which consists of a motion editing module, a pose generator, and an expression generator. The editing module projects faces into a latent space where pose motion and expression motion can be disentangled, and the pose or expression transfer can be performed in the latent space conveniently via addition. The two generators render the modified latent codes to images, respectively. Moreover, to guarantee the disentanglement, we propose a bidirectional cyclic training strategy with well-designed constraints. Evaluations demonstrate our method can control pose or expression independently and be used for general video editing. Code: https://github.com/Carlyx/DPE Youxin Pang, Yong Zhang 0034, Weize Quan, Yanbo Fan, Xiaodong Cun, Ying Shan, Dong-Ming Yan 0001 |
CVPR | 6 |
| 2023 | All in One: Exploring Unified Video-Language Pre-TrainingabstractMainstream Video-Language Pre-training (VLP) models [10, 26, 64] consist of three parts, a video encoder, a text encoder, and a video-text fusion Transformer. They pursue better performance via utilizing heavier unimodal encoders or multimodal fusion Transformers, resulting in increased parameters with lower efficiency in downstream tasks. In this work, we for the first time introduce an end-to-end VLP model, namely all-in-one Transformer, that embeds raw video and textual signals into joint representations using a unified backbone architecture. We argue that the unique temporal information of video data turns out to be a key barrier hindering the design of a modality-agnostic Transformer. To overcome the challenge, we introduce a novel and effective token rolling operation to encode temporal representations from video clips in a non-parametric manner. The careful design enables the representation learning of both video-text multimodal inputs and unimodal inputs using a unified model. Our pretrained ali-in-one Transformer is transferred to various downstream video-text tasks after fine-tuning, including text-video retrieval, video-question answering, multiple choice and video captioning. State-of-the-art performances with the minimal model FLOPs on ten datasets demonstrate the superiority of our method compared to the competitive counterparts. The code and pretrained models are available at https://github.com/showlab/all-in-one. Jinpeng Wang 0001, Yixiao Ge, Rui Yan 0001, Yuying Ge, Qinghong Lin, Satoshi Tsutsui, Xudong Lin 0003, Guanyu Cai, Ying Shan, Xiaohu Qie, Zheng Shou 0001 |
CVPR | 10 |
| 2023 | Accelerating Vision-Language Pretraining with Free Language ModelingabstractThe state of the arts in vision-language pretraining (VLP) achieves exemplary performance but suffers from high training costs resulting from slow convergence and long training time, especially on large-scale web datasets. An essential obstacle to training efficiency lies in the entangled prediction rate (percentage of tokens for reconstruction) and corruption rate (percentage of corrupted tokens) in masked language modeling (MLM), that is, a proper corruption rate is achieved at the cost of a large portion of output tokens being excluded from prediction loss. To accelerate the convergence of VLP, we propose a new pretraining task, namely, free language modeling (FLM), that enables a 100% prediction rate with arbitrary corruption rates. FLM successfully frees the prediction rate from the tie-up with the corruption rate while allowing the corruption spans to be customized for each token to be predicted. FLM-trained models are encouraged to learn better and faster given the same GPU time by exploiting bidirectional contexts more flexibly. Extensive experiments show FLM could achieve an impressive 2.5 × pretraining time reduction in comparison to the MLM-based methods, while keeping competitive performance on both vision-language understanding and generation tasks. Code will be public at https://github.com/TencentARC/FLM. Teng Wang 0007, Yixiao Ge, Feng Zheng 0001, Ran Cheng 0004, Ying Shan, Xiaohu Qie, Ping Luo 0002 |
CVPR | 5 |
| 2023 | DropMAE: Masked Autoencoders with Spatial-Attention Dropout for Tracking TasksabstractIn this paper, we study masked autoencoder (MAE) pretraining on videos for matching-based downstream tasks, including visual object tracking (VOT) and video object segmentation (VOS). A simple extension of MAE is to randomly mask out frame patches in videos and reconstruct the frame pixels. However, we find that this simple baseline heavily relies on spatial cues while ignoring temporal relations for frame reconstruction, thus leading to sub-optimal temporal matching representations for VOT and VOS. To alleviate this problem, we propose DropMAE, which adaptively performs spatial-attention dropout in the frame reconstruction to facilitate temporal correspondence learning in videos. We show that our DropMAE is a strong and efficient temporal matching learner, which achieves better finetuning results on matching-based tasks than the ImageNet-based MAE with$2\times$faster pre-training speed. Moreover, we also find that motion diversity in pre-training videos is more important than scene diversity for improving the performance on VOT and VOS. Our pre-trained DropMAE model can be directly loaded in existing ViT-based trackers for fine-tuning without further modifications. Notably, DropMAE sets new state-of-the-art performance on 8 out of 9 highly competitive video tracking and segmentation datasets. Our code and pre-trained models are available at https://github.com/jimmy-dq/DropMAE.git. Qiangqiang Wu, Tianyu Yang 0003, Ziquan Liu, Baoyuan Wu, Ying Shan, Antoni B. Chan |
CVPR | 5 |
| 2023 | Dream3D: Zero-Shot Text-to-3D Synthesis Using 3D Shape Prior and Text-to-Image Diffusion ModelsabstractRecent CLIP-guided 3D optimization methods, such as DreamFields [19] and PureCLIPNeRF [24], have achieved impressive results in zero-shot text-to-3D synthesis. However, due to scratch training and random initialization without prior knowledge, these methods often fail to generate accurate and faithful 3D structures that conform to the input text. In this paper, we make the first attempt to introduce explicit 3D shape priors into the CLIP-guided 3D optimization process. Specifically, we first generate a high-quality 3D shape from the input text in the text-to-shape stage as a 3D shape prior. We then use it as the initialization of a neural radiance field and optimize it with the full prompt. To address the challenging text-to-shape generation task, we present a simple yet effective approach that directly bridges the text and image modalities with a powerful text-to-image diffusion model. To narrow the style domain gap between the images synthesized by the text-to-image diffusion model and shape renderings used to train the image-to-shape generator, we further propose to jointly optimize a learnable text prompt and fine-tune the text-to-image diffusion model for rendering-style image generation. Our method, Dream3D, is capable of generating imaginative 3D content with superior visual quality and shape accuracy compared to state-of-the-art methods. Our project page is at https://bluestyle97.github.io/dream3d/. Xintao Wang 0002, Weihao Cheng 0002, Yan-Pei Cao 0001, Ying Shan, Xiaohu Qie, Shenghua Gao |
CVPR | 5 |
| 2023 | RILS: Masked Visual Reconstruction in Language Semantic SpaceabstractBoth masked image modeling (MIM) and natural language supervision have facilitated the progress of transferable visual pre-training. In this work, we seek the synergy between two paradigms and study the emerging properties when MIM meets natural language supervision. To this end, we present a novel masked visual Reconstruction In Language semantic Space (RILS) pre-training framework, in which sentence representations, encoded by the text encoder, serve as prototypes to transform the vision-only signals into patch-sentence probabilities as semantically meaningful MIM reconstruction targets. The vision models can therefore capture useful components with structured information by predicting proper semantic of masked tokens. Better visual representations could, in turn, improve the text encoder via the image-text alignment objective, which is essential for the effective MIM target transformation. Extensive experimental results demonstrate that our method not only enjoys the best of previous MIM and CLIP but also achieves further improvements on various tasks due to their mutual benefits. RILS exhibits advanced transferability on downstream classification, detection, and segmentation, especially for low-shot regimes. Code is available at https://github.com/hustvl/RILS. Shusheng Yang, Yixiao Ge, Ying Shan, Xiaohu Qie, Xinggang Wang |
CVPR | 5 |
| 2023 | 3D GAN Inversion with Facial Symmetry PriorabstractRecently, a surge of high-quality 3D-aware GANs have been proposed, which leverage the generative power of neural rendering. It is natural to associate 3D GANs with GAN inversion methods to project a real image into the generator's latent space, allowing free-view consistent synthesis and editing, referred as 3D GAN inversion. Although with the facial prior preserved in pre-trained 3D GANs, reconstructing a 3D portrait with only one monocular image is still an ill-pose problem. The straightforward application of 2D GAN inversion methods focuses on texture similarity only while ignoring the correctness of 3D geometry shapes. It may raise geometry collapse effects, especially when reconstructing a side face under an extreme pose. Besides, the synthetic results in novel views are prone to be blurry. In this work, we propose a novel method to promote 3D GAN inversion by introducing facial symmetry prior. We design a pipeline and constraints to make full use of the pseudo auxiliary view obtained via image flipping, which helps obtain a view-consistent and well-structured geometry shape during the inversion process. To enhance texture fidelity in unobserved viewpoints, pseudo labels from depth-guided 3D warping can provide extra supervision. We design constraints to filter out conflict areas for optimization in asymmetric situations. Comprehensive quantitative and qualitative evaluations on image reconstruction and editing demonstrate the superiority of our method. Yong Zhang 0034, Xuan Wang 0009, Tengfei Wang 0002, Xiaoyu Li 0002, Yuan Gong 0002, Yanbo Fan, Xiaodong Cun, Ying Shan, A. Cengiz Öztireli, Yujiu Yang 0001 |
CVPR | 9 |
| 2023 | OSRT: Omnidirectional Image Super-Resolution with Distortion-aware TransformerabstractOmnidirectional images (ODIs) have obtained lots of research interest for immersive experiences. Although ODIs require extremely high resolution to capture details of the entire scene, the resolutions of most ODIs are insufficient. Previous methods attempt to solve this issue by image super-resolution (SR) on equirectangular projection (ERP) images. However, they omit geometric properties of ERP in the degradation process, and their models can hardly generalize to real ERP images. In this paper, we propose Fisheye downsampling, which mimics the real-world imaging process and synthesizes more realistic low-resolution samples. Then we design a distortion-aware Transformer (OSRT) to modulate ERP distortions continuously and self-adaptively. Without a cumbersome process, OSRT outperforms previous methods by about 0.2dB on PSNR. Moreover, we propose a convenient data augmentation strategy, which synthesizes pseudo ERP images from plain images. This simple strategy can alleviate the over-fitting problem of large networks and significantly boost the performance of ODISR. Extensive experiments have demonstrated the state-of-the-art performance of our OSRT. Fanghua Yu, Xintao Wang 0002, Mingdeng Cao, Gen Li 0011, Ying Shan, Chao Dong 0005 |
CVPR | 5 |
| 2023 | SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face AnimationabstractGenerating talking head videos through a face image and a piece of speech audio still contains many challenges. i.e., unnatural head movement, distorted expression, and identity modification. We argue that these issues are mainly caused by learning from the coupled 2D motion fields. On the other hand, explicitly using 3D information also suffers problems of stiff expression and incoherent video. We present SadTalker, which generates 3D motion coefficients (head pose, expression) of the 3DMM from audio and implicitly modulates a novel 3D-aware face render for talking head generation. To learn the realistic motion coefficients, we explicitly model the connections between audio and different types of motion coefficients individually. Precisely, we present ExpNet to learn the accurate facial expression from audio by distilling both coefficients and 3D-rendered faces. As for the head pose, we design PoseVAE via a conditional VAE to synthesize head motion in different styles. Finally, the generated 3D motion coefficients are mapped to the unsupervised 3D keypoints space of the proposed face render to synthesize the final video. We conducted extensive experiments to demonstrate the superiority of our method in terms of motion and video quality.11The code and demo videos are available at https://sadtalker.github.io. Xiaodong Cun, Xuan Wang 0009, Yong Zhang 0034, Xi Shen 0001, Yu Guo 0006, Ying Shan, Fei Wang 0008 |
CVPR | 7 |
| 2023 | Skinned Motion Retargeting with Residual Perception of Motion Semantics & GeometryabstractA good motion retargeting cannot be reached without reasonable consideration of source-target differences on both the skeleton and shape geometry levels. In this work, we propose a novel Residual RETargeting network (R2ET) structure, which relies on two neural modification modules, to adjust the source motions to fit the target skeletons and shapes progressively. In particular, a skeleton-aware module is introduced to preserve the source motion semantics. A shape-aware module is designed to perceive the geometries of target characters to reduce interpenetration and contact-missing. Driven by our explored distance-based losses that explicitly model the motion semantics and geometry, these two modules can learn residual motion modifications on the source motion to generate plausible retargeted motion in a single inference without postprocessing. To balance these two modifications, we further present a balancing gate to conduct linear interpolation between them. Extensive experiments on the public dataset Mixamo demonstrate that our R2ET achieves the state-of-the-art performance, and provides a good balance between the preservation of motion semantics as well as the attenuation of interpenetration and contact-missing. Code is available at https://github.com/Kebii/R2ET. Junwu Weng, Fang Zhao 0006, Shaoli Huang, Xuefei Zhe, Linchao Bao, Ying Shan, Jue Wang 0001, Zhigang Tu 0001 |
CVPR | 8 |
| 2023 | Learning Anchor Transformations for 3D Garment AnimationabstractThis paper proposes an anchor-based deformation model, namely AnchorDEF, to predict 3D garment animation from a body motion sequence. It deforms a garment mesh template by a mixture of rigid transformations with extra nonlinear displacements. A set of anchors around the mesh surface is introduced to guide the learning of rigid transformation matrices. Once the anchor transformations are found, per-vertex nonlinear displacements of the garment template can be regressed in a canonical space, which reduces the complexity of deformation space learning. By explicitly constraining the transformed anchors to satisfy the consistencies of position, normal and direction, the physical meaning of learned anchor transformations in space is guaranteed for better generalization. Furthermore, an adaptive anchor updating is proposed to optimize the anchor position by being aware of local mesh topology for learning representative anchor transformations. Qualitative and quantitative experiments on different types of garments demonstrate that AnchorDEF achieves the state-of-the-art performance on 3D garment deformation prediction in motion, especially for loose-fitting garments. Fang Zhao 0006, Zekun Li 0002, Shaoli Huang, Junwu Weng, Tianfei Zhou, Guosen Xie, Jue Wang 0001, Ying Shan |
CVPR | 8 |
| 2023 | LayoutDiffusion: Controllable Diffusion Model for Layout-to-Image GenerationabstractRecently, diffusion models have achieved great success in image synthesis. However, when it comes to the layout-to-image generation where an image often has a complex scene of multiple objects, how to make strong control over both the global layout map and each detailed object remains a challenging task. In this paper, we propose a diffusion model named LayoutDiffusion that can obtain higher generation quality and greater controllability than the previous works. To overcome the difficult multimodal fusion of image and layout, we propose to construct a structural image patch with region information and transform the patched image into a special layout to fuse with the normal layout in a unified form. Moreover, Layout Fusion Module (LFM) and Object-aware Cross Attention (OaCA) are proposed to model the relationship among multiple objects and designed to be object-aware and position-sensitive, allowing for precisely controlling the spatial related information. Extensive experiments show that our LayoutDiffusion out-performs the previous SOTA methods on FID, CAS by relatively 46.35%,26.70% on COCO-stuff and 44.29%,41.82% on VG. Code is available at https://github.com/ZGCTroy/LayoutDiffusion. Guangcong Zheng, Xianpan Zhou, Xuewei Li 0003, Zhongang Qi, Ying Shan, Xi Li 0001 |
CVPR | 5 |
| 2023 | ERBNet: An Effective Representation Based Network for Unbiased Scene Graph GenerationabstractThe scene graph generation (SGG) task has attracted increasing attention in recent years. The goal of SGG is to predict relations between pairs of objects within an image. Due to the long-tailed distribution of the dataset annotations, the performance of SGG is still far from satisfactory. To address the long-tailed problem, existing methods try various ways to conduct unbiased learning. However, we argue that the essence of the long-tailed problem in SGG is that the classifier is seriously affected by the long-tailed data. To handle this issue, we propose a novel network named ERBNet, which contains a relation feature fusion (RFF) encoder to construct effective representations of relations between objects, and a nearest class mean (NCM) classifier to conduct relation prediction based on relation feature similarities. Extensive experimental results show that the proposed ERBNet outperforms several state-of-the-art methods on the challenging Visual Genome dataset. Wenxi Ma, Tianxiang Hou, Qianji Di, Zhongang Qi, Ying Shan, Hanzi Wang |
ICASSP | 5 |
| 2023 | Enhancing the Vocal Range of Single-Speaker Singing Voice Synthesis with Melody-Unsupervised Pre-TrainingabstractThe single-speaker singing voice synthesis (SVS) usually underperforms at pitch values that are out of the singer's vocal range or associated with limited training samples. Based on our previous work, this work proposes a melody-unsupervised multi-speaker pretraining method conducted on a multi-singer dataset to enhance the vocal range of the single-speaker, while not degrading the timbre similarity. This pre-training method can be deployed to a large-scale multi-singer dataset, which only contains audio-and-lyrics pairs without phonemic timing information and pitch annotation. Specifically, in the pre-training step, we design a phoneme predictor to produce the frame-level phoneme probability vectors as the phonemic timing information and a speaker encoder to model the timbre variations of different singers, and directly estimate the frame-level f0 values from the audio to provide the pitch information. These pre-trained model parameters are delivered into the fine-tuning step as prior knowledge to enhance the single speaker's vocal range. Moreover, this work also contributes to improving the sound quality and rhythm naturalness of the synthesized singing voices. It is the first to introduce a differentiable duration regulator to improve the rhythm naturalness of the synthesized voice, and a bi-directional flow model to improve the sound quality. Experimental results verify that the proposed SVS system outperforms the baseline on both sound quality and naturalness. Shaohuan Zhou, Xu Li 0015, Zhiyong Wu 0001, Ying Shan, Helen M. Meng |
ICASSP | 4 |
| 2023 | OmniZoomer: Learning to Move and Zoom in on Sphere at High-ResolutionabstractOmnidirectional images (ODIs) have become increasingly popular, as their large field-of-view (FoV) can offer viewers the chance to freely choose the view directions in immersive environments such as virtual reality. The Möbius transformation is typically employed to further provide the opportunity for movement and zoom on ODIs, but applying it to the image level often results in blurry effect and aliasing problem. In this paper, we propose a novel deep learning-based approach, called OmniZoomer, to incorporate the Möbius transformation into the network for movement and zoom on ODIs. By learning various transformed feature maps under different conditions, the network is enhanced to handle the increasing edge curvatures, which alleviates the blurry effect. Moreover, to address the aliasing problem, we propose two key components. Firstly, to compensate for the lack of pixels for describing curves, we enhance the feature maps in the high-resolution (HR) space and calculate the transformed index map with a spatial index generation module. Secondly, considering that ODIs are inherently represented in the spherical space, we propose a spherical resampling module that combines the index map and HR feature maps to transform the feature maps for better spherical correlation. The transformed feature maps are decoded to output a zoomed ODI. Experiments show that our method can produce HR and high-quality ODIs with the flexibility to move and zoom in to the object of interest. Project page is available at http: //vlislab22.github.io/OmniZoomer/. Zidong Cao, Hao Ai, Yan-Pei Cao 0001, Ying Shan, Xiaohu Qie, Lin Wang 0025 |
ICCV | 4 |
| 2023 | MasaCtrl: Tuning-Free Mutual Self-Attention Control for Consistent Image Synthesis and EditingabstractDespite the success in large-scale text-to-image generation and text-conditioned image editing, existing methods still struggle to produce consistent generation and editing results. For example, generation approaches usually fail to synthesize multiple images of the same objects/characters but with different views or poses. Meanwhile, existing editing methods either fail to achieve effective complex nonrigid editing while maintaining the overall textures and identity, or require time-consuming fine-tuning to capture the image-specific appearance. In this paper, we develop MasaCtrl, a tuning-free method to achieve consistent image generation and complex non-rigid image editing simultaneously. Specifically, MasaCtrl converts existing self-attention in diffusion models into mutual self-attention, so that it can query correlated local contents and textures from source images for consistency. To further alleviate the query confusion between foreground and background, we propose a mask-guided mutual self-attention strategy, where the mask can be easily extracted from the cross-attention maps. Extensive experiments show that the proposed MasaCtrl can produce impressive results in both consistent image generation and complex non-rigid real image editing. Mingdeng Cao, Xintao Wang 0002, Zhongang Qi, Ying Shan, Xiaohu Qie, Yinqiang Zheng |
ICCV | 4 |
| 2023 | Unleashing Vanilla Vision Transformer with Masked Image Modeling for Object DetectionabstractWe present an approach to efficiently and effectively adapt a masked image modeling (MIM) pre-trained vanilla Vision Transformer (ViT) for object detection, which is based on our two novel observations: (i) A MIM pre-trained vanilla ViT encoder can work surprisingly well in the challenging object-level recognition scenario even with randomly sampled partial observations, e.g., only 25% ~ 50% of the input embeddings. (ii) In order to construct multi-scale representations for object detection from single-scale ViT, a randomly initialized compact convolutional stem supplants the pre-trained patchify stem, and its intermediate features can naturally serve as the higher resolution inputs of a feature pyramid network without further upsampling or other manipulations. While the pre-trained ViT is only regarded as the 3rd-stage of our detector’s backbone instead of the whole feature extractor. This naturally results in a ConvNet-ViT hybrid architecture. The proposed detector, named MimDet, enables a MIM pre-trained vanilla ViT to outperform leading hierarchical architectures such as Swin Transformer, MViTv2 and ConvNeXt on COCO object detection & instance segmentation, and achieves better results compared with the previous best adapted vanilla ViT detector using a more modest fine-tuning recipe while converging 2.8× faster. Code and pre-trained models are available at https://github.com/hustvl/MIMDet. Shusheng Yang, Shijie Wang 0007, Yixiao Ge, Ying Shan, Xinggang Wang |
ICCV | 5 |
| 2023 | Exploring Model Transferability through the Lens of Potential EnergyabstractTransfer learning has become crucial in computer vision tasks due to the vast availability of pre-trained deep learning models. However, selecting the optimal pre-trained model from a diverse pool for a specific downstream task remains a challenge. Existing methods for measuring the transferability of pre-trained models rely on statistical correlations between encoded static features and task labels, but they overlook the impact of underlying representation dynamics during fine-tuning, leading to unreliable results, especially for self-supervised models. In this paper, we present an insightful physics-inspired approach named PED to address these challenges. We reframe the challenge of model selection through the lens of potential energy and directly model the interaction forces that influence fine-tuning dynamics. By capturing the motion of dynamic representations to decline the potential energy within a force-driven physical model, we can acquire an enhanced and more stable observation for estimating transferability. The experimental results on 10 downstream tasks and 12 self-supervised models demonstrate that our approach can seamlessly integrate into existing ranking techniques and enhance their performances, revealing its effectiveness for the model selection task and its potential for understanding the mechanism in transfer learning. Code is available at https://github.com/lixiaotong97/PED. Yixiao Ge, Ying Shan, Ling-Yu Duan |
ICCV | 4 |
| 2023 | HOSNeRF: Dynamic Human-Object-Scene Neural Radiance Fields from a Single VideoabstractWe introduce HOSNeRF, a novel 360° free-viewpoint rendering method that reconstructs neural radiance fields for dynamic human-object-scene from a single monocular in-the-wild video. Our method enables pausing the video at any frame and rendering all scene details (dynamic humans, objects, and backgrounds) from arbitrary viewpoints. The first challenge in this task is the complex object motions in human-object interactions, which we tackle by introducing the new object bones into the conventional human skeleton hierarchy to effectively estimate large object deformations in our dynamic human-object model. The second challenge is that humans interact with different objects at different times, for which we introduce two new learnable object state embeddings that can be used as conditions for learning our human-object representation and scene representation, respectively. Extensive experiments show that HOSNeRF significantly outperforms SOTA approaches on two challenging datasets by a large margin of 40%~50% in terms of LPIPS. The code, data, and compelling examples of 360° free-viewpoint renderings from single videos: https://showlab.github.io/HOSNeRF. Jia-Wei Liu, Yan-Pei Cao 0001, Tianyuan Yang, Zhongcong Xu, Jussi Keppo, Ying Shan, Xiaohu Qie, Zheng Shou 0001 |
ICCV | 6 |
| 2023 | Order-Prompted Tag Sequence Generation for Video TaggingabstractVideo Tagging intends to infer multiple tags spanning relevant content for a given video. Typically, video tags are freely defined and uploaded by a variety of users, so they have two characteristics: abundant in quantity and disordered intra-video. It is difficult for the existing multilabel classification and generation methods to adapt directly to this task. This paper proposes a novel generative model, Order-Prompted Tag Sequence Generation (OP-TSG), according to the above characteristics. It regards video tagging as a tag sequence generation problem guided by sample-dependent order prompts. These prompts are semantically aligned with tags and enable to decouple tag generation order, making the model focus on modeling the tag dependencies. Moreover, the word-based generation strategy enables the model to generate novel tags. To verify the effectiveness and generalization of the proposed method, a Chinese video tagging benchmark CREATE-tagging, and an English image tagging benchmark Pexel-tagging are established. Extensive results show that OP-TSG is significantly superior to other methods, especially the results on rare tags improve by 3.3% and 3% over SOTA methods on CREATE-tagging and Pexel-tagging, and novel tags generated on CREATE-tagging exhibit a tag gain of 7.04%. Zongyang Ma, Ziqi Zhang 0010, Zhongang Qi, Yingmin Luo, Zekun Li 0006, Chunfeng Yuan, Bing Li 0001, Xiaohu Qie, Ying Shan, Weiming Hu 0004 |
ICCV | 10 |
| 2023 | FateZero: Fusing Attentions for Zero-shot Text-based Video EditingabstractThe diffusion-based generative models have achieved remarkable success in text-based image generation. However, since it contains enormous randomness in generation progress, it is still challenging to apply such models for real-world visual content editing, especially in videos. In this paper, we propose FateZero, a zero-shot text-based editing method on real-world videos without per-prompt training or use-specific mask. To edit videos consistently, we propose several techniques based on the pre-trained models. Firstly, in contrast to the straightforward DDIM inversion technique, our approach captures intermediate attention maps during inversion, which effectively retain both structural and motion information. These maps are directly fused in the editing process rather than generated during denoising. To further minimize semantic leakage of the source video, we then fuse self-attentions with a blending mask obtained by cross-attention features from the source prompt. Furthermore, we have implemented a reform of the self-attention mechanism in denoising UNet by introducing spatial-temporal attention to ensure frame consistency. Yet succinct, our method is the first one to show the ability of zero-shot text-driven video style and local attribute editing from the trained text-to-image model. We also have a better zero-shot shape-aware editing ability based on the text-to-video model [52]. Extensive experiments demonstrate our superior temporal consistency and editing capability than previous works. Xiaodong Cun, Yong Zhang 0034, Chenyang Lei, Xintao Wang 0002, Ying Shan, Qifeng Chen 0001 |
ICCV | 6 |
| 2023 | Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video GenerationabstractTo replicate the success of text-to-image (T2I) generation, recent works employ large-scale video datasets to train a text-to-video (T2V) generator. Despite their promising results, such paradigm is computationally expensive. In this work, we propose a new T2V generation setting—One-Shot Video Tuning, where only one text-video pair is presented. Our model is built on state-of-the-art T2I diffusion models pre-trained on massive image data. We make two key observations: 1) T2I models can generate still images that represent verb terms; 2) extending T2I models to generate multiple images concurrently exhibits surprisingly good content consistency. To further learn continuous motion, we introduce Tune-A-Video, which involves a tailored spatio-temporal attention mechanism and an efficient one-shot tuning strategy. At inference, we employ DDIM inversion to provide structure guidance for sampling. Extensive qualitative and numerical experiments demonstrate the remarkable ability of our method across various applications. Jay Zhangjie Wu, Yixiao Ge, Xintao Wang 0002, Stan Weixian Lei, Yuchao Gu, Yufei Shi 0003, Wynne Hsu, Ying Shan, Xiaohu Qie, Zheng Shou 0001 |
ICCV | 8 |
| 2023 | Speech2Lip: High-fidelity Speech to Lip Generation by Learning from a Short VideoabstractSynthesizing realistic videos according to a given speech is still an open challenge. Previous works have been plagued by issues such as inaccurate lip shape generation and poor image quality. The key reason is that only motions and appearances on limited facial areas (e.g., lip area) are mainly driven by the input speech. Therefore, directly learning a mapping function from speech to the entire head image is prone to ambiguity, particularly when using a short video for training. We thus propose a decomposition-synthesis-composition framework named Speech to Lip (Speech2Lip) that disentangles speech-sensitive and speech-insensitive motion/appearance to facilitate effective learning from limited training data, resulting in the generation of natural-looking videos. First, given a fixed head pose (i.e., canonical space), we present a speech-driven implicit model for lip image generation which concentrates on learning speech-sensitive motion and appearance. Next, to model the major speech-insensitive motion (i.e., head movement), we introduce a geometry-aware mutual explicit mapping (GAMEM) module that establishes geometric mappings between different head poses. This allows us to paste generated lip images at the canonical space onto head images with arbitrary poses and synthesize talking videos with natural head movements. In addition, a Blend-Net and a contrastive sync loss are introduced to enhance the overall synthesis performance. Quantitative and qualitative results on three benchmarks demonstrate that our model can be trained by a video of just a few minutes in length and achieve state-of-the-art performance in both visual quality and speechvisual synchronization. Code: https://github.com/CVMILab/Speech2Lip. Xiuzhe Wu, Yang Wu 0001, Xiaoyang Lyu, Yan-Pei Cao 0001, Ying Shan, Wenming Yang, Zhongqian Sun, Xiaojuan Qi 0001 |
ICCV | 6 |
| 2023 | Masked Image Modeling with Denoising Contrast
Yixiao Ge, Shusheng Yang, Ying Shan, Xiaohu Qie |
ICLR | 7 |
| 2023 | Do We Really Need Temporal Convolutions in Action Segmentation?abstractRecognizing and segmenting actions from long videos is a challenging problem. Most existing methods focus on designing temporal convolutional models. However, these models are limited in their flexibility and ability to model long-term dependencies. Transformers have recently been used in various tasks. But the lack of inductive bias and the inefficiency of handling long video sequences limit the application of Transformers in action segmentation. In this paper, we present a pure Transformer-based model without temporal convolutions in action segmentation, called Temporal U-Transformer. The U-Transformer architecture not only reduces complexity but also introduces an inductive bias that neighboring frames are more likely to belong to the same class. Besides, we further propose a boundary-aware loss based on the distribution of similarity scores between frames from attention modules to improve the ability to recognize boundaries. Extensive experiments show the effectiveness of our method. Dazhao Du, Yu Li 0003, Zhongang Qi, Lingyu Si, Ying Shan |
ICME | 6 |
| 2023 | π-Tuning: Transferring Multimodal Foundation Models with Optimal Multi-task Interpolation
Chengyue Wu, Teng Wang 0007, Yixiao Ge, Ruisong Zhou, Ying Shan, Ping Luo 0002 |
ICML | 6 |
| 2023 | DeSRA: Detect and Delete the Artifacts of GAN-based Real-World Super-Resolution ModelsabstractImage super-resolution (SR) with generative adversarial networks (GAN) has achieved great success in restoring realistic details. However, it is notorious that GAN-based SR models will inevitably produce unpleasant and undesirable artifacts, especially in practical scenarios. Previous works typically suppress artifacts with an extra loss penalty in the training phase. They only work for in-distribution artifact types generated during training. When applied in real-world scenarios, we observe that those improved methods still generate obviously annoying artifacts during inference. In this paper, we analyze the cause and characteristics of the GAN artifacts produced in unseen test data without ground-truths. We then develop a novel method, namely, DeSRA, to Detect and then “Delete” those SR Artifacts in practice. Specifically, we propose to measure a relative local variance distance from MSE-SR results and GAN-SR results, and locate the problematic areas based on the above distance and semantic-aware thresholds. After detecting the artifact regions, we develop a finetune procedure to improve GAN-based SR models with a few samples, so that they can deal with similar types of artifacts in more unseen real data. Equipped with our DeSRA, we can successfully eliminate artifacts from inference and improve the ability of SR models to be applied in real-world scenarios. The code will be available at https://github.com/TencentARC/DeSRA. Liangbin Xie, Xintao Wang 0002, Xiangyu Chen 0006, Gen Li 0011, Ying Shan, Jiantao Zhou 0001, Chao Dong 0005 |
ICML | 5 |
| 2023 | SGAT4PASS: Spherical Geometry-Aware Transformer for PAnoramic Semantic SegmentationabstractAs an important and challenging problem in computer vision, PAnoramic Semantic Segmentation (PASS) gives complete scene perception based on an ultra-wide angle of view. Usually, prevalent PASS methods with 2D panoramic image input focus on solving image distortions but lack consideration of the 3D properties of original 360 degree data. Therefore, their performance will drop a lot when inputting panoramic images with the 3D disturbance. To be more robust to 3D disturbance, we propose our Spherical Geometry-Aware Transformer for PAnoramic Semantic Segmentation (SGAT4PASS), considering 3D spherical geometry knowledge. Specifically, a spherical geometry-aware framework is proposed for PASS. It includes three modules, i.e., spherical geometry-aware image projection, spherical deformable patch embedding, and a panorama-aware loss, which takes input images with 3D disturbance into account, adds a spherical geometry-aware constraint on the existing deformable patch embedding, and indicates the pixel density of original 360 degree data, respectively. Experimental results on Stanford2D3D Panoramic datasets show that SGAT4PASS significantly improves performance and robustness, with approximately a 2% increase in mIoU, and when small 3D disturbances occur in the data, the stability of our performance is improved by an order of magnitude. Our code and supplementary material are available at https://github.com/TencentARC/SGAT4PASS. Xuewei Li 0003, Zhongang Qi, Gaoang Wang, Ying Shan, Xi Li 0001 |
IJCAI | 5 |
| 2023 | Prosody Modeling with 3D Visual Information for Expressive Video Dubbing
Shansong Liu, Xu Li 0015, Haozhe Wu, Zhiyong Wu 0001, Ying Shan, Jia Jia 0001 |
INTERSPEECH | 6 |
| 2023 | Binary Embedding-based Retrieval at TencentabstractLarge-scale embedding-based retrieval (EBR) is the cornerstone of search-related industrial applications. Given a user query, the system of EBR aims to identify relevant information from a large corpus of documents that may be tens or hundreds of billions in size. The storage and computation turn out to be expensive and inefficient with massive documents and high concurrent queries, making it difficult to further scale up. Yukang Gan, Yixiao Ge, Chang Zhou 0008, Shupeng Su, Zhouchuan Xu, Xuyuan Xu, Quanchao Hui, Yexin Wang, Ying Shan |
KDD | 10 |
| 2023 | Toward Human Perception-Centric Video Thumbnail GenerationabstractVideo thumbnail plays an essential role in summarizing video content into a compact and concise image for users to browse efficiently. However, automatically generating attractive and informative video thumbnails remains an open problem due to the difficulty of formulating human aesthetic perception and the scarcity of paired training data. This work proposes a novel Human Perception-Centric Video Thumbnail Generation (HPCVTG) to address these challenges. Specifically, our framework first generates a set of thumbnails using a principle-based system, which conforms to established aesthetic and human perception principles, such as visual balance in the layout and avoiding overlapping elements. Then rather than designing from scratch, we ask human annotators to evaluate some of these thumbnails and select their preferred ones. A Transformer-based Variational Auto-Encoder (VAE) model is firstly pre-trained with Model-Agnostic Meta-Learning (MAML) and then fine-tuned on these human-selected thumbnails. The exploration of combining the MAML pre-training paradigm with human feedback in training can reduce human involvement and make the training process more efficient. Extensive experimental results show that our HPCVTG framework outperforms existing methods in objective and subjective evaluations, highlighting its potential to improve the user experience when browsing videos and inspire future research in human perception-centric content generation tasks. The code and dataset will be released via https://github.com/yangtao2019yt/HPCVTG. Junfan Lin, Zhongang Qi, Yang Wu 0001, Ying Shan, Chang Wen Chen |
ACM Multimedia | 7 |
| 2023 | VTLayout: A Multi-Modal Approach for Video Text LayoutabstractThe rapid explosion of video distribution is accompanied by a massive amount of video text, which encompasses rich information about the video content. While previous research has primarily focused on text extraction from videos like text detection, tracking, recognition and end to end spotting, the layout of video text has received limited attention. As different text categories convey distinct meanings, video text layout is critical for video understanding tasks such as video summarization and shooting environment comprehension. To bridge the gap between video OCR and understanding, we explore the study of video text layout in this work. We first optimize the layout annotation of the BOVText, a bilingual, open-world video text dataset, by expanding text categories and defining five clear categories: scene, subtitle, title, logo, and other. Additionally, we rectify the original unreasonable layout annotation based on these definitions. We also propose a Video-level Text Layout model (VTLayout) to address the layout problem, which fuses textual, visual, and spatial-temporal embedding of video text trajectories. To the best of our knowledge, this is the first method to tackle text layout on video level. Our method outperforms image-level layout methods across all text categories and exhibits faster inference speed. This study underscores the significance of video text layout in video understanding and offers an effective solution to this challenge. Our annotation is available at https://github.com/TencentARC/VTLayout. Jin Ma 0003, Zhongang Qi, Zehua Xie, Yu Luo 0013, Qiusheng Kang, Ying Shan |
ACM Multimedia | 7 |
| 2023 | PanoGRF: Generalizable Spherical Radiance Fields for Wide-baseline PanoramasabstractAchieving an immersive experience enabling users to explore virtual environments with six degrees of freedom (6DoF) is essential for various applications such as virtual reality (VR). Wide-baseline panoramas are commonly used in these applications to reduce network bandwidth and storage requirements. However, synthesizing novel views from these panoramas remains a key challenge. Although existing neural radiance field methods can produce photorealistic views under narrow-baseline and dense image captures, they tend to overfit the training views when dealing with wide-baseline panoramas due to the difficulty in learning accurate geometry from sparse $360^{\circ}$ views. To address this problem, we propose PanoGRF, Generalizable Spherical Radiance Fields for Wide-baseline Panoramas, which construct spherical radiance fields incorporating $360^{\circ}$ scene priors. Unlike generalizable radiance fields trained on perspective images, PanoGRF avoids the information loss from panorama-to-perspective conversion and directly aggregates geometry and appearance features of 3D sample points from each panoramic view based on spherical projection. Moreover, as some regions of the panorama are only visible from one view while invisible from others under wide baseline settings, PanoGRF incorporates $360^{\circ}$ monocular depth priors into spherical depth estimation to improve the geometry features. Experimental results on multiple panoramic datasets demonstrate that PanoGRF significantly outperforms state-of-the-art generalizable view synthesis methods for wide-baseline panoramas (e.g., OmniSyn) and perspective images (e.g., IBRNet, NeuRay). Zheng Chen 0016, Yan-Pei Cao 0001, Chen Wang 0049, Ying Shan, Song-Hai Zhang |
NeurIPS | 5 |
| 2023 | Meta-Adapter: An Online Few-shot Learner for Vision-Language ModelabstractThe contrastive vision-language pre-training, known as CLIP, demonstrates remarkable potential in perceiving open-world visual concepts, enabling effective zero-shot image recognition.
Nevertheless, few-shot learning methods based on CLIP typically require offline fine-tuning of the parameters on few-shot samples, resulting in longer inference time and the risk of overfitting in certain domains.
To tackle these challenges, we propose the Meta-Adapter, a lightweight residual-style adapter, to refine the CLIP features guided by the few-shot samples in an online manner.
With a few training samples, our method can enable effective few-shot learning capabilities and generalize to unseen data or tasks without additional fine-tuning, achieving competitive performance and high efficiency.
Without bells and whistles, our approach outperforms the state-of-the-art online few-shot learning method by an average of 3.6\% on eight image classification datasets with higher inference speed.
Furthermore, our model is simple and flexible, serving as a plug-and-play module directly applicable to downstream tasks.
Without further fine-tuning, Meta-Adapter obtains notable performance improvements in open-vocabulary object detection and segmentation tasks. Lin Song 0002, Ruoyi Xue, Hongbin Sun 0001, Yixiao Ge, Ying Shan |
NeurIPS | 7 |
| 2023 | Mix-of-Show: Decentralized Low-Rank Adaptation for Multi-Concept Customization of Diffusion ModelsabstractPublic large-scale text-to-image diffusion models, such as Stable Diffusion, have gained significant attention from the community. These models can be easily customized for new concepts using low-rank adaptations (LoRAs). However, the utilization of multiple-concept LoRAs to jointly support multiple customized concepts presents a challenge. We refer to this scenario as decentralized multi-concept customization, which involves single-client concept tuning and center-node concept fusion. In this paper, we propose a new framework called Mix-of-Show that addresses the challenges of decentralized multi-concept customization, including concept conflicts resulting from existing single-client LoRA tuning and identity loss during model fusion. Mix-of-Show adopts an embedding-decomposed LoRA (ED-LoRA) for single-client tuning and gradient fusion for the center node to preserve the in-domain essence of single concepts and support theoretically limitless concept fusion. Additionally, we introduce regionally controllable sampling, which extends spatially controllable sampling (e.g., ControlNet and T2I-Adapter) to address attribute binding and missing object problems in multi-concept sampling. Extensive experiments demonstrate that Mix-of-Show is capable of composing multiple customized concepts with high fidelity, including characters, objects, and scenes. Yuchao Gu, Xintao Wang 0002, Jay Zhangjie Wu, Yujun Shi, Yunpeng Chen, Zihan Fan, Wuyou Xiao, Rui Zhao 0001, Shuning Chang, Weijia Wu 0001, Yixiao Ge, Ying Shan, Zheng Shou 0001 |
NeurIPS | 12 |
| 2023 | CL-NeRF: Continual Learning of Neural Radiance Fields for Evolving Scene RepresentationabstractExisting methods for adapting Neural Radiance Fields (NeRFs) to scene changes require extensive data capture and model retraining, which is both time-consuming and labor-intensive. In this paper, we tackle the challenge of efficiently adapting NeRFs to real-world scene changes over time using a few new images while retaining the memory of unaltered areas, focusing on the continual learning aspect of NeRFs. To this end, we propose CL-NeRF, which consists of two key components: a lightweight expert adaptor for adapting to new changes and evolving scene representations and a conflict-aware knowledge distillation learning objective for memorizing unchanged parts. We also present a new benchmark for evaluating Continual Learning of NeRFs with comprehensive metrics. Our extensive experiments demonstrate that CL-NeRF can synthesize high-quality novel views of both changed and unchanged regions with high training efficiency, surpassing existing methods in terms of reducing forgetting and adapting to changes. Code and benchmark will be made available. Xiuzhe Wu, Peng Dai 0003, Weipeng Deng, Handi Chen, Yang Wu 0001, Yan-Pei Cao 0001, Ying Shan, Xiaojuan Qi 0001 |
NeurIPS | 7 |
| 2023 | GPT4Tools: Teaching Large Language Model to Use Tools via Self-instructionabstractThis paper aims to efficiently enable Large Language Models (LLMs) to use multi-modal tools.
The advanced proprietary LLMs, such as ChatGPT and GPT-4, have shown great potential for tool usage through sophisticated prompt engineering.
Nevertheless, these models typically rely on prohibitive computational costs and publicly inaccessible data.
To address these challenges, we propose the GPT4Tools based on self-instruct to enable open-source LLMs, such as LLaMA and OPT, to use tools.
It generates an instruction-following dataset by prompting an advanced teacher with various multi-modal contexts.
By using the Low-Rank Adaptation (LoRA) optimization, our approach facilitates the open-source LLMs to solve a range of visual problems, including visual comprehension and image generation.
Moreover, we provide a benchmark to evaluate the ability of LLMs to use tools, which is performed in both zero-shot and fine-tuning ways.
Extensive experiments demonstrate the effectiveness of our method on various language models, which not only significantly improves the accuracy of invoking seen tools, but also enables the zero-shot capacity for unseen tools. Rui Yang 0041, Lin Song 0002, Sijie Zhao, Yixiao Ge, Xiu Li 0001, Ying Shan |
NeurIPS | 7 |
| 2023 | Exploiting Contextual Objects and Relations for 3D Visual Groundingabstract3D visual grounding, the task of identifying visual objects in 3D scenes based on natural language inputs, plays a critical role in enabling machines to understand and engage with the real-world environment. However, this task is challenging due to the necessity to capture 3D contextual information to distinguish target objects from complex 3D scenes. The absence of annotations for contextual objects and relations further exacerbates the difficulties. In this paper, we propose a novel model, CORE-3DVG, to address these challenges by explicitly learning about contextual objects and relations. Our method accomplishes 3D visual grounding via three sequential modular networks, including a text-guided object detection network, a relation matching network, and a target identification network. During training, we introduce a pseudo-label self-generation strategy and a weakly-supervised method to facilitate the learning of contextual objects and relations, respectively. The proposed techniques allow the networks to focus more effectively on referred objects within 3D scenes by understanding their context better. We validate our model on the challenging Nr3D, Sr3D, and ScanRefer datasets and demonstrate state-of-the-art performance. Our code will be public at https://github.com/yangli18/CORE-3DVG. Li Yang 0014, Chunfeng Yuan, Ziqi Zhang 0010, Zhongang Qi, Wei Liu 0153, Ying Shan, Bing Li 0001, Weiping Yang, Yan Wang 0153, Weiming Hu 0004 |
NeurIPS | 7 |
| 2023 | Inserting Anybody in Diffusion Models via Celeb BasisabstractExquisite demand exists for customizing the pretrained large text-to-image model, $e.g.$ Stable Diffusion, to generate innovative concepts, such as the users themselves. However, the newly-added concept from previous customization methods often shows weaker combination abilities than the original ones even given several images during training. We thus propose a new personalization method that allows for the seamless integration of a unique individual into the pre-trained diffusion model using just $one\ facial\ photograph$ and only $1024\ learnable\ parameters$ under $3\ minutes$. So we can effortlessly generate stunning images of this person in any pose or position, interacting with anyone and doing anything imaginable from text prompts. To achieve this, we first analyze and build a well-defined celeb basis from the embedding space of the pre-trained large text encoder. Then, given one facial photo as the target identity, we generate its own embedding by optimizing the weight of this basis and locking all other parameters. Empowered by the proposed celeb basis, the new identity in our customized model showcases a better concept combination ability than previous personalization methods. Besides, our model can also learn several new identities at once and interact with each other where the previous customization model fails to. Project page is at: http://celeb-basis.github.io. Code is at: https://github.com/ygtxr1997/CelebBasis. Ge Yuan, Xiaodong Cun, Yong Zhang 0034, Maomao Li, Xintao Wang 0002, Ying Shan, Huicheng Zheng |
NeurIPS | 7 |
| 2023 | Interactive Story Visualization with Multiple CharactersabstractAccurate Story visualization requires several necessary elements, such as identity consistency across frames, the alignment between plain text and visual content, and a reasonable layout of objects in images. Most previous works endeavor to meet these requirements by fitting a text-to-image (T2I) model on a set of videos in the same style and with the same characters, e.g., the FlintstonesSV dataset. However, the learned T2I models typically struggle to adapt to new characters, scenes, and styles, and often lack the flexibility to revise the layout of the synthesized images. This paper proposes a system for generic interactive story visualization, capable of handling multiple novel characters and supporting the editing of layout and local structure. It is developed by leveraging the prior knowledge of large language and T2I models, trained on massive corpora. The system comprises four interconnected components: story-to-prompt generation (S2P), text-to-layout generation (T2L), controllable text-to-image generation (C-T2I), and image-to-video animation (I2V). First, the S2P module converts concise story information into detailed prompts required for subsequent stages. Next, T2L generates diverse and reasonable layouts based on the prompts, offering users the ability to adjust and refine the layout to their preferences. The core component, C-T2I, enables the creation of images guided by layouts, sketches, and actor-specific identifiers to maintain consistency and detail across visualizations. Finally, I2V enriches the visualization process by animating the generated images. Extensive experiments and a user study are conducted to validate the effectiveness and flexibility of interactive editing of the proposed system. Yuan Gong 0002, Youxin Pang, Xiaodong Cun, Menghan Xia, Yingqing He, Haoxin Chen, Longyue Wang, Yong Zhang 0034, Xintao Wang 0002, Ying Shan, Yujiu Yang 0001 |
SIGGRAPH Asia | 10 |
| 2023 | VMesh: Hybrid Volume-Mesh Representation for Efficient View SynthesisabstractWith the emergence of neural radiance fields (NeRFs), view synthesis quality has reached an unprecedented level. Compared to traditional mesh-based assets, this volumetric representation is more powerful in expressing scene geometry but inevitably suffers from high rendering costs and can hardly be involved in further processes like editing, posing significant difficulties in combination with the existing graphics pipeline. In this paper, we present a hybrid volume-mesh representation, VMesh, which depicts an object with a textured mesh along with an auxiliary sparse volume. VMesh retains the advantages of mesh-based assets, such as efficient rendering and compact storage, while also incorporating the ability to represent subtle geometric structures provided by the volumetric counterpart. VMesh can be obtained from multi-view images of an object and renders at 2K 60FPS on common consumer devices with high fidelity, unleashing new opportunities for real-time immersive applications. Yan-Pei Cao 0001, Chen Wang 0049, Yu He 0001, Ying Shan, Song-Hai Zhang |
SIGGRAPH Asia | 5 |
| 2023 | Neural Point-based Volumetric Avatar: Surface-guided Neural Points for Efficient and Photorealistic Volumetric Head AvatarabstractRendering photorealistic and dynamically moving human heads is crucial for ensuring a pleasant and immersive experience in AR/VR and video conferencing applications. However, existing methods often struggle to model challenging facial regions (e.g., mouth interior, eyes, and beard), resulting in unrealistic and blurry results. In this paper, we propose Neural Point-based Volumetric Avatar (NPVA), a method that adopts the neural point representation as well as the neural volume rendering process and discards the predefined connectivity and hard correspondence imposed by mesh-based approaches. Specifically, the neural points are strategically constrained around the surface of the target expression via a high-resolution UV displacement map, achieving increased modeling capacity and more accurate control. We introduce three technical innovations to improve the rendering and training efficiency: a patch-wise depth-guided (shading point) sampling strategy, a lightweight radiance decoding process, and a Grid-Error-Patch (GEP) ray sampling strategy during training. By design, our NPVA is better equipped to handle topologically changing regions and thin structures while also ensuring accurate expression control when animating avatars. Experiments conducted on three subjects from the Multiface dataset demonstrate the effectiveness of our designs, outperforming previous state-of-the-art methods, especially in handling challenging facial regions. Cong Wang 0045, Yan-Pei Cao 0001, Linchao Bao, Ying Shan, Song-Hai Zhang |
SIGGRAPH Asia | 5 |
| 2023 | Anti-Aliased Neural Implicit Surfaces with Encoding Level of DetailabstractWe present LoD-NeuS, an efficient neural representation for high-frequency geometry detail recovery and anti-aliased novel view rendering. Drawing inspiration from voxel-based representations with the level of detail (LoD), we introduce a multi-scale tri-plane-based scene representation that is capable of capturing the LoD of the signed distance function (SDF) and the space radiance. Our representation aggregates space features from a multi-convolved featurization within a conical frustum along a ray and optimizes the LoD feature volume through differentiable rendering. Additionally, we propose an error-guided sampling strategy to guide the growth of the SDF during the optimization. Both qualitative and quantitative evaluations demonstrate that our method achieves superior surface reconstruction and photorealistic view synthesis compared to state-of-the-art approaches. Yiyu Zhuang, Qi Zhang 0029, Hao Zhu 0004, Yao Yao 0008, Xiaoyu Li 0002, Yan-Pei Cao 0001, Ying Shan, Xun Cao |
SIGGRAPH Asia | 8 |
| 2023 | ACNet: Approaching-and-Centralizing Network for Zero-Shot Sketch-Based Image RetrievalabstractThe huge domain gap between sketches and photos poses huge challenges for Sketch-Based Image Retrieval (SBIR). The Zero-Shot Sketch-Based Image Retrieval (ZS-SBIR) is more generic and practical but brings an even greater challenge: the additional knowledge gap between the seen and unseen categories. In order to simultaneously mitigate both gaps, we propose an Approaching-and-Centralizing Network (termed “ACNet”) to jointly optimize sketch-to-photo synthesis and image retrieval. The retrieval module guides the synthesis module to generate large amounts of diverse photo-like images that help the sketch domain gradually approach the photo domain to eliminate the domain gap, and thus better serves retrieval. Meanwhile, the retrieval module itself centralizes the embeddings of training samples for learning a similarity measurement to eliminate the knowledge gap. Our approach is simple yet effective, which achieves state-of-the-art performance on two widely used ZS-SBIR datasets and surpasses previous methods by a large margin (eg, 8.2% improvement in terms of mAP@all on TU-Berlin Extended dataset). Hao Ren 0002, Ziqiang Zheng, Yang Wu 0001, Hong Lu 0001, Yang Yang 0002, Ying Shan, Sai-Kit Yeung |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | Task-Aware Dual-Representation Network for Few-Shot Action RecognitionabstractFew-shot action recognition has attracted increasing attention in recent years, but it remains challenging due to the intrinsic difficulty in learning transferable knowledge to generalize to novel classes by using a few labeled samples. Although some successful progress has been made, most few-shot action recognition methods commonly focus on the global characteristics of samples while ignoring the local characteristics of samples, which results in the weak generalization ability of the model. In this paper, we propose a task-aware dual-representation network (TADRNet) for few-shot action recognition, which learns how to adapt video representations to novel tasks in a meta-learning manner. It mainly includes a global relational graph subnetwork (GRG) and a fine-grained local representation subnetwork (FLR). Our method simultaneously considers both global and local characteristics of samples for few-shot action recognition. From a global perspective, we propose GRG to explore the relations across support-query sample pairs by using the relational graph neural network. To facilitate the few-shot visual learning, we propose a novel hybrid semantic attention module (HSA) for enhancing the discriminability of support and query features. From a local perspective, we utilize FLR to fully exploit the local characteristics of samples, which can improve the classification results obtained by GRG and thus guarantee high classification accuracy. Extensive experiments on four challenging benchmarks show that the proposed TADRNet significantly outperforms a variety of state-of-the-art few-shot action recognition methods. Xiao Wang 0072, Weirong Ye, Zhongang Qi, Guangge Wang, Ying Shan, Xiaohu Qie, Hanzi Wang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | Robust Human Matting via Semantic Guidance
Xiangguang Chen, Ye Zhu 0003, Yu Li 0003, Bingtao Fu, Lei Sun 0009, Ying Shan, Shan Liu 0001 |
ACCV (2) | 6 |
| 2022 | Bridging Video-text Retrieval with Multiple Choice QuestionsabstractPretraining a model to learn transferable video-text representation for retrieval has attracted a lot of attention in recent years. Previous dominant works mainly adopt two separate encoders for efficient retrieval, but ignore local associations between videos and texts. Another line of research uses a joint encoder to interact video with texts, but results in low efficiency since each text-video pair needs to be fed into the model. In this work, we enable fine-grained video-text interactions while maintaining high efficiency for retrieval via a novel pretext task, dubbed as Multiple Choice Questions (MCQ), where a parametric module BridgeFormer is trained to answer the “questions” constructed by the text features via resorting to the video features. Specifically, we exploit the rich semantics of text (i.e., nouns and verbs) to build questions, with which the video encoder can be trained to capture more regional content and temporal dynamics. In the form of questions and answers, the semantic associations between local video-text features can be properly established. BridgeFormer is able to be removed for downstream retrieval, rendering an efficient and flexible model with only two encoders. Our method outperforms state-of-the-art methods on the popular text-to-video retrieval task in five datasets with different experimental setups (i.e., zero-shot andfine-tune), including HowTo100M (one million videos). We further conduct zero-shot action recognition, which can be cast as video-to-text retrieval, and our approach also significantly surpasses its counterparts. As an additional benefit, our method achieves competitive results with much shorter pre-training videos on single-modality downstream tasks, e.g., action recognition with linear evaluation. Yuying Ge, Yixiao Ge, Xihui Liu, Ying Shan, Xiaohu Qie, Ping Luo 0002 |
CVPR | 5 |
| 2022 | UMT: Unified Multi-modal Transformers for Joint Video Moment Retrieval and Highlight DetectionabstractFinding relevant moments and highlights in videos according to natural language queries is a natural and highly valuable common need in the current video content explosion era. Nevertheless, jointly conducting moment retrieval and highlight detection is an emerging research topic, even though its component problems and some related tasks have already been studied for a while. In this paper, we present the first unified framework, named Unified Multi-modal Transformers (UMT), capable of realizing such joint optimization while can also be easily degenerated for solving individual problems. As far as we are aware, this is the first scheme to integrate multi-modal (visual-audio) learning for either joint optimization or the individual moment retrieval task, and tackles moment retrieval as a keypoint detection problem using a novel query generator and query decoder. Extensive comparisons with existing methods and ablation studies on QVHighlights, Charades-STA, YouTube Highlights, and TVSum datasets demonstrate the effectiveness, superiority, and flexibility of the proposed method under various settings. Source code and pre-trained models are available at https://github.com/TencentARC/UMT. Ye Liu 0002, Siyuan Li 0026, Yang Wu 0001, Chang Wen Chen, Ying Shan, Xiaohu Qie |
CVPR | 5 |
| 2022 | Object-aware Video-language Pre-training for RetrievalabstractRecently, by introducing large-scale dataset and strong transformer network, video-language pre-training has shown great success especially for retrieval. Yet, existing video-language transformer models do not explicitly fine-grained semantic align. In this work, we present Object-aware Transformers, an object-centric approach that extends video-language transformer to incorporate object representations. The key idea is to leverage the bounding boxes and object tags to guide the training process. We evaluate our model on three standard sub-tasks of video-text matching on four widely used benchmarks. We also provide deep analysis and detailed ablation about the proposed method. We show clear improvement in performance across all tasks and datasets considered, demonstrating the value of a model that incorporates object representations into a video-language architecture. The code has been released in https://github.com/FingerRec/OA-Transformer. Alex Jinpeng Wang, Yixiao Ge, Guanyu Cai, Rui Yan 0001, Xudong Lin 0003, Ying Shan, Xiaohu Qie, Zheng Shou 0001 |
CVPR | 6 |
| 2022 | BTS: A Bi-lingual Benchmark for Text Segmentation in the WildabstractAs a prerequisite of many text-related tasks such as text erasing and text style transfer, text segmentation arouses more and more attention recently. Current researches mainly focus on only English characters and digits, while few work studies Chinese characters due to the lack of pub-lic large-scale and high-quality Chinese datasets, which limits the practical application scenarios of text segmentation. Different from English which has a limited alphabet of letters, Chinese has much more basic characters with com-plex structures, making the problem more difficult to deal with. To better analyze this problem, we propose the Bi-lingual Text Segmentation (BTS) dataset, a benchmark that covers various common Chinese scenes including 14,250 diverse and fine-annotated text images. BTS mainly focuses on Chinese characters, and also contains English words and digits. We also introduce Prior Guided Text Segmen-tation Network (PGTSNet), the first baseline to handle bi-lingual and complex-structured text segmentation. A plug-in text region highlighting module and a text perceptual dis-criminator are proposed in PGTSNet to supervise the model with text prior, and guide for more stable and finer text seg-mentation. A variation loss is also employed for suppressing background noise under complex scene. Extensive ex-periments are conducted not only to demonstrate the neces-sity and superiority of the proposed dataset BTS, but also to show the effectiveness of the proposed PGTSNet compared with a variety of state-of-the-art text segmentation methods. Xixi Xu, Zhongang Qi, Jianqi Ma, Honglun Zhang, Ying Shan, Xiaohu Qie |
CVPR | 5 |
| 2022 | Temporally Efficient Vision Transformer for Video Instance SegmentationabstractRecently vision transformer has achieved tremendous success on image-level visual recognition tasks. To effectively and efficiently model the crucial temporal information within a video clip, we propose a Temporally Efficient Vision Transformer (TeViT) for video instance segmentation (VIS). Different from previous transformer-based VIS methods, TeViT is nearly convolution-free, which contains a transformer backbone and a query-based video instance segmentation head. In the backbone stage, we propose a nearly parameter-free messenger shift mechanism for early temporal context fusion. In the head stages, we propose a parameter-shared spatiotemporal query interaction mechanism to build the one-to-one correspondence between video instances and queries. Thus, TeViT fully utilizes both frame-level and instance-level temporal context information and obtains strong temporal modeling capacity with negligible extra computational cost. On three widely adopted VIS benchmarks, i.e., YouTube-VIS-2019, YouTube-VIS-2021, and OVIS, TeViT obtains state-of-the-art results and maintains high inference speed, e.g., 46.6 AP with 68.9 FPS on YouTube-VIS-2019. Code is available at https://github.com/hustvl/TeViT. Shusheng Yang, Xinggang Wang, Yu Li 0003, Jiemin Fang, Wenyu Liu 0001, Ying Shan |
CVPR | 8 |
| 2022 | MILES: Visual BERT Pre-training with Injected Language Semantics for Video-Text Retrieval
Yuying Ge, Yixiao Ge, Xihui Liu, Jinpeng Wang 0001, Ying Shan, Xiaohu Qie, Ping Luo 0002 |
ECCV (35) | 6 |
| 2022 | VQFR: Blind Face Restoration with Vector-Quantized Dictionary and Parallel Decoder
Yuchao Gu, Xintao Wang 0002, Liangbin Xie, Chao Dong 0005, Gen Li 0011, Ying Shan, Ming-Ming Cheng |
ECCV (18) | 6 |
| 2022 | mc-BEiT: Multi-choice Discretization for Image BERT Pre-training
Yixiao Ge, Ying Shan, Ling-Yu Duan |
ECCV (30) | 5 |
| 2022 | Metric Learning Based Interactive Modulation for Real-World Super-Resolution
Chong Mou, Yanze Wu, Xintao Wang 0002, Chao Dong 0005, Jian Zhang 0018, Ying Shan |
ECCV (17) | 6 |
| 2022 | Not All Models Are Equal: Predicting Model Transferability in a Self-challenging Fisher Space
Wenqi Shao, Yixiao Ge, Zhaoyang Zhang 0004, Lei Yang 0059, Xiaogang Wang 0001, Ying Shan, Ping Luo 0002 |
ECCV (34) | 7 |
| 2022 | Audio-To-Symbolic Arrangement Via Cross-Modal Music Representation LearningabstractCould we automatically derive the score of a piano accompaniment based on the audio of a pop song? This is the audio-to-symbolic arrangement problem we tackle in this paper. A good arrangement model should not only consider the audio content but also have prior knowledge of piano composition (so that the generation "sounds like" the audio and meanwhile maintains musicality). To this end, we contribute a cross-modal representation-learning model, which 1) extracts chord and melodic information from the audio, and 2) learns texture representation from both audio and a corrupted ground truth arrangement. We further introduce a tailored training strategy that gradually shifts the source of texture information from corrupted score to audio. In the end, the score-based texture posterior is reduced to a standard normal distribution, and only audio is needed for inference. Experiments show that our model captures major audio information and outperforms baselines in generation quality.1 Ziyu Wang 0008, Dejing Xu, Gus Xia, Ying Shan |
ICASSP | 4 |
| 2022 | Uncertainty Modeling for Out-of-Distribution Generalization
Yongxing Dai, Yixiao Ge, Jun Liu 0036, Ying Shan, Ling-Yu Duan |
ICLR | 5 |
| 2022 | Dynamic Token Normalization improves Vision Transformers
Wenqi Shao, Yixiao Ge, Zhaoyang Zhang 0004, Xuyuan Xu, Xiaogang Wang 0001, Ying Shan, Ping Luo 0002 |
ICLR | 6 |
| 2022 | Hot-Refresh Model Upgrades with Regression-Free Compatible Training in Image Retrieval
Binjie Zhang, Yixiao Ge, Yantao Shen 0003, Yu Li 0003, Chun Yuan 0003, Xuyuan Xu, Yexin Wang, Ying Shan |
ICLR | 8 |
| 2022 | Convolutional Transformer with Similarity-based Boundary Prediction for Action SegmentationabstractAction classification has made great progress, but segmenting and recognizing actions from long videos remains a challenging problem. Recently, Transformer-based models with strong sequence modeling ability have succeeded in many se-quence modeling tasks. However, the lack of inductive bias and the difficulty of handling long video sequences limit the application of the Transformer in the action segmentation task. In order to explore the potential of the Transformer in this task, we replace some specific linear layers in the vanilla Transformer with dilated temporal convolution, and a sparse attention mechanism is utilized to reduce the time and space complexities to process long video sequences. Besides, directly using frame-wise classification loss to train the model will cause that frames at boundaries of actions are treated equally with those in the middle of actions, and the learned features are not sensitive to boundaries. We propose a new local log-context attention module to predict whether each frame is at the beginning, middle, or end of an action. Since boundary frames are similar to their neighboring frames of different classes, our similarity-based boundary prediction helps learn more discriminative features. Extensive experiments on three datasets show the effectiveness of our method. Dazhao Du, Yu Li 0003, Zhongang Qi, Lingyu Si, Ying Shan |
ICTAI | 6 |
| 2022 | Towards Universal Backward-Compatible Representation LearningabstractConventional model upgrades for visual search systems require offline refresh of gallery features by feeding gallery images into new models (dubbed as “backfill”), which is time-consuming and expensive, especially in large-scale applications. The task of backward-compatible representation learning is therefore introduced to support backfill-free model upgrades, where the new query features are interoperable with the old gallery features. Despite the success, previous works only investigated a close-set training scenario (i.e., the new training set shares the same classes as the old one), and are limited by more realistic and challenging open-set scenarios. To this end, we first introduce a new problem of universal backward-compatible representation learning, covering all possible data split in model upgrades. We further propose a simple yet effective method, dubbed as Universal Backward-Compatible Training (UniBCT) with a novel structural prototype refinement algorithm, to learn compatible representations in all kinds of model upgrading benchmarks in a unified manner. Comprehensive experiments on the large-scale face recognition datasets MS1Mv3 and IJB-C fully demonstrate the effectiveness of our method. Source code is available at https://github.com/TencentARC/OpenCompatible. Binjie Zhang, Yixiao Ge, Yantao Shen 0003, Shupeng Su, Fanzi Wu, Chun Yuan 0003, Xuyuan Xu, Yexin Wang, Ying Shan |
IJCAI | 9 |
| 2022 | A Hierarchical Speaker Representation Framework for One-shot Singing Voice ConversionabstractTypically, singing voice conversion (SVC) depends on an embedding vector, extracted from either a speaker lookup table (LUT) or a speaker recognition network (SRN), to model speaker identity.However, singing contains more expressive speaker characteristics than conversational speech.It is suspected that a single embedding vector may only capture averaged and coarse-grained speaker characteristics, which is insufficient for the SVC task.To this end, this work proposes a novel hierarchical speaker representation framework for SVC, which can capture fine-grained speaker characteristics at different granularity.It consists of an up-sampling stream and three down-sampling streams.The up-sampling stream transforms the linguistic features into audio samples, while one downsampling stream of the three operates in the reverse direction.It is expected that the temporal statistics of each down-sampling block can represent speaker characteristics at different granularity, which will be engaged in the up-sampling blocks to enhance the speaker modeling.Experiment results verify that the proposed method outperforms both the LUT and SRN based SVC systems.Moreover, the proposed system supports the one-shot SVC with only a few seconds of reference audio. Xu Li 0015, Shansong Liu, Ying Shan |
INTERSPEECH | 3 |
| 2022 | PC-Dance: Posture-controllable Music-driven Dance SynthesisabstractMusic-driven dance synthesis is a task to generate high-quality dance according to the music given by the user, which has promising entertainment applications. However, most of the existing methods cannot provide an efficient and effective way for user intervention in dance generation, e.g., posture-controllable. In this work, we propose a powerful framework named PC-Dance to perform adaptive posture-controllable music-driven dance synthesis. Consisting of an music-to-dance alignment embedding network (M2D-Align) and a posture-controllable dance synthesis (PC-Syn), PC-Dance allows fine-grained control by input anchor poses efficiently without artist participation. Specifically, to relieve the cost of artist participation but ensure generating high-quality dance efficiently, a self-supervised rhythm alignment module is designed to further learn the music-to-dance alignment embedding. As for PC-Syn, we introduce an efficient scheme for adaptive motion graph construction (AMGC), which could improve the efficiency of graph-based optimization and preserve the diversity of motions. Since there is few related public dataset, we collect an MMD-ARC dataset for music-driven dance synthesis. The experimental results on MMD-ARC dataset demonstrate the effectiveness of our framework and the feasibility for dance synthesis with adaptive posture controlling. Jibin Gao, Junfu Pu, Honglun Zhang, Ying Shan, Wei-Shi Zheng 0001 |
ACM Multimedia | 4 |
| 2022 | RepSR: Training Efficient VGG-style Super-Resolution Networks with Structural Re-Parameterization and Batch NormalizationabstractThis paper explores training efficient VGG-style super-resolution (SR) networks with the structural re-parameterization technique. The general pipeline of re-parameterization is to train networks with multi-branch topology first, and then merge them into standard 3x3 convolutions for efficient inference. In this work, we revisit those primary designs and investigate essential components for re-parameterizing SR networks. First of all, we find that batch normalization (BN) is important to bring training non-linearity and improve the final performance. However, BN is typically ignored in SR, as it usually degrades the performance and introduces unpleasant artifacts. We carefully analyze the cause of BN issue and then propose a straightforward yet effective solution. In particular, we first train SR networks with mini-batch statistics as usual, and then switch to using population statistics at the later training period. While we have successfully re-introduced BN into SR, we further design a new re-parameterizable block tailored for SR, namely RepSR. It consists of a clean residual path and two expand-and-squeeze convolution paths with the modified BN. Extensive experiments demonstrate that our simple RepSR is capable of achieving superior performance to previous SR re-parameterization methods among different model sizes. In addition, our RepSR can achieve a better trade-off between performance and actual running time (throughput) than previous SR methods. Codes are available at https://github.com/TencentARC/RepSR. Xintao Wang 0002, Chao Dong 0005, Ying Shan |
ACM Multimedia | 3 |
| 2022 | DeVRF: Fast Deformable Voxel Radiance Fields for Dynamic ScenesabstractModeling dynamic scenes is important for many applications such as virtual reality and telepresence. Despite achieving unprecedented fidelity for novel view synthesis in dynamic scenes, existing methods based on Neural Radiance Fields (NeRF) suffer from slow convergence (i.e., model training time measured in days). In this paper, we present DeVRF, a novel representation to accelerate learning dynamic radiance fields. The core of DeVRF is to model both the 3D canonical space and 4D deformation field of a dynamic, non-rigid scene with explicit and discrete voxel-based representations. However, it is quite challenging to train such a representation which has a large number of model parameters, often resulting in overfitting issues. To overcome this challenge, we devise a novel static-to-dynamic learning paradigm together with a new data capture setup that is convenient to deploy in practice. This paradigm unlocks efficient learning of deformable radiance fields via utilizing the 3D volumetric canonical space learnt from multi-view static images to ease the learning of 4D voxel deformation field with only few-view dynamic sequences. To further improve the efficiency of our DeVRF and its synthesized novel view's quality, we conduct thorough explorations and identify a set of strategies. We evaluate DeVRF on both synthetic and real-world dynamic scenes with different types of deformation. Experiments demonstrate that DeVRF achieves two orders of magnitude speedup (100× faster) with on-par high-fidelity results compared to the previous state-of-the-art approaches. The code and dataset are released in https://github.com/showlab/DeVRF. Yan-Pei Cao 0001, Weijia Mao, Wenqiao Zhang, Junhao Zhang 0001, Jussi Keppo, Ying Shan, Xiaohu Qie, Zheng Shou 0001 |
NeurIPS | 7 |
| 2022 | AnimeSR: Learning Real-World Super-Resolution Models for Animation VideosabstractThis paper studies the problem of real-world video super-resolution (VSR) for animation videos, and reveals three key improvements for practical animation VSR. First, recent real-world super-resolution approaches typically rely on degradation simulation using basic operators without any learning capability, such as blur, noise, and compression. In this work, we propose to learn such basic operators from real low-quality animation videos, and incorporate the learned ones into the degradation generation pipeline. Such neural-network-based basic operators could help to better capture the distribution of real degradations. Second, a large-scale high-quality animation video dataset, AVC, is built to facilitate comprehensive training and evaluations for animation VSR. Third, we further investigate an efficient multi-scale network structure. It takes advantage of the efficiency of unidirectional recurrent networks and the effectiveness of sliding-window-based methods. Thanks to the above delicate designs, our method, AnimeSR, is capable of restoring real-world low-quality animation videos effectively and efficiently, achieving superior performance to previous state-of-the-art methods. Yanze Wu, Xintao Wang 0002, Gen Li 0011, Ying Shan |
NeurIPS | 4 |
| 2022 | Depth-Aware Shadow RemovalabstractAbstract Shadow removal from a single image is an ill‐posed problem because shadow generation is affected by the complex interactions of geometry, albedo, and illumination. Most recent deep learning‐based methods try to directly estimate the mapping between the non‐shadow and shadow image pairs to predict the shadow‐free image. However, they are not very effective for shadow images with complex shadows or messy backgrounds. In this paper, we propose a novel end‐to‐end depth‐aware shadow removal method without using depth images, which estimates depth information from RGB images and leverages the depth feature as guidance to enhance shadow removal and refinement. The proposed framework consists of three components, including depth prediction, shadow removal, and boundary refinement. First, the depth prediction module is used to predict the corresponding depth map of the input shadow image. Then, we propose a new generative adversarial network (GAN) method integrated with depth information to remove shadows in the RGB image. Finally, we propose an effective boundary refinement framework to alleviate the artifact around boundaries after shadow removal by depth cues. We conduct experiments on several public datasets and real‐world shadow images. The experimental results demonstrate the efficiency of the proposed method and superior performance against state‐of‐the‐art methods. Yanping Fu, Zhenyu Gai, Haifeng Zhao 0001, Shaojie Zhang 0002, Ying Shan, Yang Wu 0001, Jin Tang 0001 |
Comput. Graph. Forum | 5 |
| 2022 | Hybrid Warping Fusion for Video Frame Interpolation
Yu Li 0003, Ye Zhu 0003, Ruoteng Li, Xintao Wang 0002, Ying Shan |
Int. J. Comput. Vis. | 6 |
| 2021 | Distilling Audio-Visual Knowledge by Compositional Contrastive LearningabstractHaving access to multi-modal cues (e.g. vision and audio) empowers some cognitive tasks to be done faster compared to learning from a single modality. In this work, we propose to transfer knowledge across heterogeneous modalities, even though these data modalities may not be semantically correlated. Rather than directly aligning the representations of different modalities, we compose audio, image, and video representations across modalities to uncover richer multi-modal knowledge. Our main idea is to learn a compositional embedding that closes the cross-modal semantic gap and captures the task-relevant semantics, which facilitates pulling together representations across modalities by compositional contrastive learning. We establish a new, comprehensive multi-modal distillation benchmark on three video datasets: UCF101, ActivityNet, and VGGSound. Moreover, we demonstrate that our model significantly outperforms a variety of existing knowledge distillation methods in transferring audio-visual knowledge to improve video representation learning. Code is released here: https://github.com/yanbeic/CCL. Yanbei Chen, Yongqin Xian, A. Sophia Koepke, Ying Shan, Zeynep Akata |
CVPR | 4 |
| 2021 | Towards Real-World Blind Face Restoration With Generative Facial PriorabstractBlind face restoration usually relies on facial priors, such as facial geometry prior or reference prior, to restore realistic and faithful details. However, very low-quality inputs cannot offer accurate geometric prior while high-quality references are inaccessible, limiting the applicability in real-world scenarios. In this work, we propose GFP-GAN that leverages rich and diverse priors encapsulated in a pretrained face GAN for blind face restoration. This Generative Facial Prior (GFP) is incorporated into the face restoration process via spatial feature transform layers, which allow our method to achieve a good balance of realness and fidelity. Thanks to the powerful generative facial prior and delicate designs, our GFP-GAN could jointly restore facial details and enhance colors with just a single forward pass, while GAN inversion methods require image-specific optimization at inference. Extensive experiments show that our method achieves superior performance to prior art on both synthetic and real-world datasets. Xintao Wang 0002, Yu Li 0003, Honglun Zhang, Ying Shan |
CVPR | 4 |
| 2021 | Open-Book Video Captioning With Retrieve-Copy-Generate NetworkabstractIn this paper, we convert traditional video captioning task into a new paradigm, i.e., Open-book Video Captioning, which generates natural language under the prompts of video-content-relevant sentences, not limited to the video itself. To address the open-book video captioning problem, we propose a novel Retrieve-Copy-Generate network, where a pluggable video-to-text retriever is constructed to retrieve sentences as hints from the training corpus effectively, and a copy-mechanism generator is introduced to extract expressions from multi-retrieved sentences dynamically. The two modules can be trained end-to-end or separately, which is flexible and extensible. Our framework co-ordinates the conventional retrieval-based methods with orthodox encoder-decoder methods, which can not only draw on the diverse expressions in the retrieved sentences but also generate natural and accurate content of the video. Extensive experiments on several benchmark datasets show that our proposed approach surpasses the state-of-the-art performance, indicating the effectiveness and promising of the proposed paradigm in the task of video captioning. Ziqi Zhang 0010, Zhongang Qi, Chunfeng Yuan, Ying Shan, Bing Li 0001, Weiming Hu 0004 |
CVPR | 4 |
| 2021 | Instances as QueriesabstractWe present QueryInst, a new perspective for instance segmentation. QueryInst is a multi-stage end-to-end system that treats instances of interest as learnable queries, enabling query based object detectors, e.g., Sparse RCNN, to have strong instance segmentation performance. The attributes of instances such as categories, bounding boxes, instance masks, and instance association embeddings are represented by queries in a unified manner. In QueryInst, a query is shared by both detection and segmentation via dynamic convolutions and driven by parallellysupervised multi-stage learning. We conduct extensive experiments on three challenging benchmarks, i.e., COCO, CityScapes, and YouTube-VIS to evaluate the effectiveness of QueryInst in object detection, instance segmentation, and video instance segmentation tasks. For the first time, we demonstrate that a simple end-to-end query based framework can achieve the state-of-the-art performance in various instance-level recognition tasks. Code is available at https://github.com/hustvl/QueryInst. Shusheng Yang, Xinggang Wang, Yu Li 0003, Ying Shan, Bin Feng 0001, Wenyu Liu 0001 |
ICCV | 6 |
| 2021 | Towards Vivid and Diverse Image Colorization with Generative Color PriorabstractColorization has attracted increasing interest in recent years. Classic reference-based methods usually rely on external color images for plausible results. A large image database or online search engine is inevitably required for retrieving such exemplars. Recent deep-learning-based methods could automatically colorize images at a low cost. However, unsatisfactory artifacts and incoherent colors are always accompanied. In this work, we aim at recovering vivid colors by leveraging the rich and diverse color priors encapsulated in a pretrained Generative Adversarial Networks (GAN). Specifically, we first "retrieve" matched features (similar to exemplars) via a GAN encoder and then incorporate these features into the colorization process with feature modulations. Thanks to the powerful generative color prior and delicate designs, our method could produce vivid colors with a single forward pass. Moreover, it is highly convenient to obtain diverse results by modifying GAN latent codes. Our method also inherits the merit of interpretable controls of GANs and could attain controllable and smooth transitions by walking through GAN latent space. Extensive experiments and user studies demonstrate that our method achieves superior performance than previous works. Yanze Wu, Xintao Wang 0002, Yu Li 0003, Honglun Zhang, Ying Shan |
ICCV | 6 |
| 2021 | Crossover Learning for Fast Online Video Instance SegmentationabstractModeling temporal visual context across frames is critical for video instance segmentation (VIS) and other video understanding tasks. In this paper, we propose a fast on-line VIS model termed CrossVIS. For temporal information modeling in VIS, we present a novel crossover learning scheme that uses the instance feature in the current frame to pixel-wisely localize the same instance in other frames. Different from previous schemes, crossover learning does not require any additional network parameters for feature enhancement. By integrating with the instance segmentation loss, crossover learning enables efficient cross-frame instance-to-pixel relation learning and brings cost-free improvement during inference. Besides, a global balanced instance embedding branch is proposed for better and more stable online instance association. We conduct extensive experiments on three challenging VIS benchmarks, i.e., YouTube-VIS-2019, OVIS, and YouTube-VIS-2021 to evaluate our methods. CrossVIS achieves state-of-the-art online VIS performance and shows a decent trade-off between latency and accuracy. Code is available at https://github.com/hustvl/CrossVIS. Shusheng Yang, Xinggang Wang, Yu Li 0003, Ying Shan, Bin Feng 0001, Wenyu Liu 0001 |
ICCV | 6 |
| 2021 | Cross-modal Consensus Network for Weakly Supervised Temporal Action LocalizationabstractWeakly supervised temporal action localization (WS-TAL) is a challenging task that aims to localize action instances in the given video with video-level categorical supervision. Previous works use the appearance and motion features extracted from pre-trained feature encoder directly,e.g., feature concatenation or score-level fusion. In this work, we argue that the features extracted from the pre-trained extractors,e.g., I3D, which are trained for trimmed video action classification, but not specific for WS-TAL task, leading to inevitable redundancy and sub-optimization. Therefore, the feature re-calibration is needed for reducing the task-irrelevant information redundancy. Here, we propose a cross-modal consensus network(CO2-Net) to tackle this problem. In CO2-Net, we mainly introduce two identical proposed cross-modal consensus modules (CCM) that design a cross-modal attention mechanism to filter out the task-irrelevant information redundancy using the global information from the main modality and the cross-modal local information from the auxiliary modality. Moreover, we further explore inter-modality consistency, where we treat the attention weights derived from each CCM as the pseudo targets of the attention weights derived from another CCM to maintain the consistency between the predictions derived from two CCMs, forming a mutual learning manner. Finally, we conduct extensive experiments on two commonly used temporal action localization datasets, THUMOS14 and ActivityNet1.2, to verify our method, which we achieve state-of-the-art results. The experimental results show that our proposed cross-modal consensus module can produce more representative features for temporal action localization. Fa-Ting Hong, Jia-Chang Feng, Dan Xu 0002, Ying Shan, Wei-Shi Zheng 0001 |
ACM Multimedia | 4 |
| 2021 | TransFusion: Multi-Modal Fusion for Video Tag Inference via Translation-based Knowledge EmbeddingabstractTag inference is an important task in the business of video platforms with wide applications such as recommendation, interpretation, and more. Existing works are mainly based on extracting video information from multiple modalities such as frames or music, and then infer tags through classification or object detection. This, however, does not apply to inferring generic tags or taxonomy that are less relevant to video contents, such as video originality or its broader category, which are important in practice. In this paper, we claim that these generic tags can be modeled through the semantic relations between videos and tags, and can be utilized simultaneously with the multi-modal features to achieve better video tagging. We propose TransFusion, an end-to-end supervised learning framework that fuses multi-modal embeddings (e.g., vision, audio, texts, etc.) with the knowledge embedding to derive the video representation. To infer the diverse tags following heterogeneous relations, TransFusion adopts a dual attentive approach to learn both the modality importance in fusion and relation importance in inference. Besides, it is general enough and can be used with the existing translation-based knowledge embedding approaches. Extensive experiments show that TransFusion outperforms the baseline methods with lowered mean rank and at least 9.59% improvement in [email protected] on the real-world video knowledge graph. Di Jin 0003, Zhongang Qi, Yingmin Luo, Ying Shan |
ACM Multimedia | 4 |
| 2021 | Semantic-Guided Relation Propagation Network for Few-shot Action RecognitionabstractFew-shot action recognition has drawn growing attention as it can recognize novel action classes by using only a few labeled samples. In this paper, we propose a novel semantic-guided relation propagation network (SRPN), which leverages semantic information together with visual information for few-shot action recognition. Different from most previous works that neglect semantic information in the labeled data, our SRPN directly utilizes the semantic label as an additional supervisory signal to improve the generalization ability of the network. Besides, we treat the relation of each visual-semantic pair as a relational node, and we use a graph convolutional network to model and propagate such sample relations across visual-semantic pairs, including both intra-class commonality and inter-class uniqueness, to guide the relation propagation in the graph. However, since videos contain crucial sequences and ordering information, we propose a novel spatial-temporal difference module, which can facilitate the network to enhance the visual feature learning ability at both feature level and granular level for videos. Extensive experiments conducted on several challenging benchmarks demonstrate that our SRPN outperforms several state-of-the-art methods with a significant margin. Xiao Wang 0072, Weirong Ye, Zhongang Qi, Guangge Wang, Ying Shan, Hanzi Wang |
ACM Multimedia | 6 |
| 2021 | Finding Discriminative Filters for Specific Degradations in Blind Super-ResolutionabstractRecent blind super-resolution (SR) methods typically consist of two branches, one for degradation prediction and the other for conditional restoration. However, our experiments show that a one-branch network can achieve comparable performance to the two-branch scheme. Then we wonder: how can one-branch networks automatically learn to distinguish degradations? To find the answer, we propose a new diagnostic tool -- Filter Attribution method based on Integral Gradient (FAIG). Unlike previous integral gradient methods, our FAIG aims at finding the most discriminative filters instead of input pixels/features for degradation removal in blind SR networks. With the discovered filters, we further develop a simple yet effective method to predict the degradation of an input image. Based on FAIG, we show that, in one-branch blind SR networks, 1) we could find a very small number of (1%) discriminative filters for each specific degradation; 2) The weights, locations and connections of the discovered filters are all important to determine the specific network function. 3) The task of degradation prediction can be implicitly realized by these discriminative filters without explicit supervised learning. Our findings can not only help us better understand network behaviors inside one-branch blind SR networks, but also provide guidance on designing more efficient architectures and diagnosing networks for blind SR. Liangbin Xie, Xintao Wang 0002, Chao Dong 0005, Zhongang Qi, Ying Shan |
NeurIPS | 5 |
| 2020 | Fast Video Object Segmentation Using the Global Context Module
Yu Li 0003, Zhuoran Shen, Ying Shan |
ECCV (10) | 3 |
| 2020 | Feature Augmented Memory with Global Attention Network for VideoQAabstractRecently, Recurrent Neural Network (RNN) based methods and Self-Attention (SA) based methods have achieved promising performance in Video Question Answering (VideoQA). Despite the success of these works, RNN-based methods tend to forget the global semantic contents due to the inherent drawbacks of the recurrent units themselves, while SA-based methods cannot precisely capture the dependencies of the local neighborhood, leading to insufficient modeling for temporal order. To tackle these problems, we propose a novel VideoQA framework which progressively refines the representations of videos and questions from fine to coarse grain in a sequence-sensitive manner. Specifically, our model improves the feature representations via the following two steps: (1) introducing two fine-grained feature-augmented memories to strengthen the information augmentation of video and text which can improve memory capacity by memorizing more relevant and targeted information. (2) appending the self-attention and co-attention module to the memory output thus the module is able to capture global interaction between high-level semantic informations. Experimental results show that our approach achieves state-of-the-art performance on VideoQA benchmark datasets. Jiayin Cai, Chun Yuan 0003, Lei Li 0051, Yangyang Cheng, Ying Shan |
IJCAI | 6 |
| 2020 | Dual Semantic Fusion Network for Video Object DetectionabstractVideo object detection is a tough task due to the deteriorated quality of video sequences captured under complex environments. Currently, this area is dominated by a series of feature enhancement based methods, which distill beneficial semantic information from multiple frames and generate enhanced features through fusing the distilled information. However, the distillation and fusion operations are usually performed at either frame level or instance level with external guidance using additional information, such as optical flow and feature memory. In this work, we propose a dual semantic fusion network (abbreviated as DSFNet) to fully exploit both frame-level and instance-level semantics in a unified fusion framework without external guidance. Moreover, we introduce a geometric similarity measure into the fusion process to alleviate the influence of information distortion caused by noise. As a result, the proposed DSFNet can generate more robust features through the multi-granularity fusion and avoid being affected by the instability of external guidance. To evaluate the proposed DSFNet, we conduct extensive experiments on the ImageNet VID dataset. Notably, the proposed dual semantic fusion network achieves, to the best of our knowledge, the best performance of 84.1% mAP among the current state-of-the-art video object detectors with ResNet-101 and 85.4% mAP with ResNeXt-101 without using any post-processing steps. Lijian Lin, Haosheng Chen 0001, Honglun Zhang, Yu Li 0003, Ying Shan, Hanzi Wang |
ACM Multimedia | 6 |
| 2020 | Detecting Interactions from Neural Networks via Topological AnalysisabstractDetecting statistical interactions between input features is a crucial and challenging task. Recent advances demonstrate that it is possible to extract learned interactions from trained neural networks. It has also been observed that, in neural networks, any interacting features must follow a strongly weighted connection to common hidden units. Motivated by the observation, in this paper, we propose to investigate the interaction detection problem from a novel topological perspective by analyzing the connectivity in neural networks. Specially, we propose a new measure for quantifying interaction strength, based upon the well-received theory of persistent homology. Based on this measure, a Persistence Interaction Dection (PID) algorithm is developed to efficiently detect interactions. Our proposed algorithm is evaluated across a number of interaction detection tasks on several synthetic and real-world datasets with different hyperparameters. Experimental results validate that the PID algorithm outperforms the state-of-the-art baselines. Zirui Liu 0001, Qingquan Song, Kaixiong Zhou, Ting-Hsiang Wang, Ying Shan, Xia Ben Hu |
NeurIPS | 5 |
| 2019 | Overview of the NLPCC 2019 Shared Task: Open Domain Conversation Evaluation
Ying Shan, Anqi Cui, Luchen Tan, Kun Xiong |
NLPCC (2) | 1 |
| 2018 | Recurrent Binary Embedding for GPU-Enabled Exhaustive Retrieval from Billion-Scale Semantic VectorsabstractRapid advances in GPU hardware and multiple areas of Deep Learning open up a new opportunity for billion-scale information retrieval with exhaustive search. Building on top of the powerful concept of semantic learning, this paper proposes a Recurrent Binary Embedding (RBE) model that learns compact representations for real-time retrieval. The model has the unique ability to refine a base binary vector by progressively adding binary residual vectors to meet the desired accuracy. The refined vector enables efficient implementation of exhaustive similarity computation with bit-wise operations, followed by a near-lossless k-NN selection algorithm, also proposed in this paper. The proposed algorithms are integrated into an end-to-end multi-GPU system that retrieves thousands of top items from over a billion candidates in real-time. The RBE model and the retrieval system were evaluated with data from a major paid search engine. When measured against the state-of-the-art model for binary representation and the full precision model for semantic embedding, RBE significantly outperformed the former, and filled in over 80% of the AUC gap in-between. Experiments comparing with our production retrieval system also demonstrated superior performance. While the primary focus of this paper is to build RBE based on a particular class of semantic models, generalizing to other types is straightforward, as exemplified by two different models at the end of the paper. Ying Shan, Jian Jiao 0007, Jie Zhu 0007, J. C. Mao |
KDD | 1 |
| 2017 | Deep Embedding Forest: Forest-based Serving with Deep Embedding FeaturesabstractDeep Neural Networks (DNN) have demonstrated superior ability to extract high level embedding vectors from low level features. Despite the success, the serving time is still the bottleneck due to expensive run-time computation of multiple layers of dense matrices. GPGPU, FPGA, or ASIC-based serving systems require additional hardware that are not in the mainstream design of most commercial applications. In contrast, tree or forest-based models are widely adopted because of low serving cost, but heavily depend on carefully engineered features. This work proposes a Deep Embedding Forest model that benefits from the best of both worlds. The model consists of a number of embedding layers and a forest/tree layer. The former maps high dimensional (hundreds of thousands to millions) and heterogeneous low-level features to the lower dimensional (thousands) vectors, and the latter ensures fast serving. Jie Zhu 0007, Ying Shan, J. C. Mao, Holakou Rahmanian |
KDD | 2 |
| 2016 | Deep Crossing: Web-Scale Modeling without Manually Crafted Combinatorial FeaturesabstractManually crafted combinatorial features have been the "secret sauce" behind many successful models. For web-scale applications, however, the variety and volume of features make these manually crafted features expensive to create, maintain, and deploy. This paper proposes the Deep Crossing model which is a deep neural network that automatically combines features to produce superior models. The input of Deep Crossing is a set of individual features that can be either dense or sparse. The important crossing features are discovered implicitly by the networks, which are comprised of an embedding and stacking layer, as well as a cascade of Residual Units. Deep Crossing is implemented with a modeling tool called the Computational Network Tool Kit (CNTK), powered by a multi-GPU platform. It was able to build, from scratch, two web-scale models for a major paid search engine, and achieve superior results with only a sub-set of the features used in the production models. This demonstrates the potential of using Deep Crossing as a general modeling paradigm to improve existing products, as well as to speed up the development of new models with a fraction of the investment in feature engineering and acquisition of deep domain knowledge. Ying Shan, T. Ryan Hoens, Jian Jiao 0007, Haijing Wang, J. C. Mao |
KDD | 1 |
| 2010 | Internet VisionabstractThe ten papers in this special issue focus on Internet vision. The goal is to provide the reader with a general sense of the research being conducted in Internet vision, defined as the intersection of computer vision and the Internet. This issue includes coverage of a number of significant advances in this field. Shai Avidan, Simon Baker, Ying Shan |
Proc. IEEE | 3 |
| 2009 | Efficient Scale-Space Spatiotemporal Saliency Tracking for Distortion-Free Video Retargeting
Gang Hua 0001, Cha Zhang, Zicheng Liu 0001, Zhengyou Zhang, Ying Shan |
ACCV (2) | 5 |
| 2009 | Kernel PCA Regression for Missing Data Estimation in DNA Microarray AnalysisabstractThe DNA microarray may contain missing expression data. Estimation of missing values is a necessary step in microarray analysis, because data mining procedures require a complete expression as their input. In this paper, we propose a missing data estimation algorithm, named KPCAimpute, based on kernel principal component analysis. We consider a family of heavy-tailed kernel functions, which is a generalization of the famous Gaussian kernel. The performance of the proposed KPCAimpute algorithm is compared with two state-of-the-art linear regression methods, i.e., Bayesian principal component analysis imputation (BPCA) and local least squares imputation (LLSimpute). The KPCAimpute outperforms the LL-Simpute when the missing percentage increases. The performance of the KPCAimpute is similar to that of the BPCA imputation. Therefore, it is an effective and promising algorithm in estimating missing values for DNA microarray profiles. Ying Shan, Guang Deng |
ISCAS | 1 |
| 2008 | Discovering class specific composite features through discriminative sampling with Swendsen-Wang CutabstractThis paper proposes a novel approach to discover a set of class specific ldquocomposite featuresrdquo as the feature pool for the detection and classification of complex objects using AdaBoost. Each composite feature is constructed from the combination of multiple individual features. Unlike previous works that design features manually or with certain restrictions, the class specific features are selected from the space of all combinations of a set of individual features. To achieve this, we first establish an analogue between the problem of discriminative feature selection and generative image segmentation, and then draw discriminative samples from the combinatory space with a novel algorithm called discriminative generalized Swendsen-Wang cut. These samples form the initial pool of features, where AdaBoost is applied to learn a strong classifier combining the most discriminative composite features. We demonstrate the efficacy of our approach by comparing with existing detection algorithms for finding people in general pose. Feng Han 0002, Ying Shan, Harpreet Sawhney, Rakesh Kumar 0001 |
CVPR | 2 |
| 2008 | Unsupervised Learning of Discriminative Edge Measures for Vehicle Matching between Nonoverlapping CamerasabstractThis paper proposes a novel unsupervised algorithm learning discriminative features in the context of matching road vehicles between two non-overlapping cameras. The matching problem is formulated as a same-different classification problem, which aims to compute the probability of vehicle images from two distinct cameras being from the same vehicle or different vehicle(s). We employ a novel measurement vector that consists of three independent edge-based measures and their associated robust measures computed from a pair of aligned vehicle edge maps. The weight of each measure is determined by an unsupervised learning algorithm that optimally separates the same-different classes in the combined measurement space. This is achieved with a weak classification algorithm that automatically collects representative samples from same-different classes, followed by a more discriminative classifier based on Fisher' s Linear Discriminants and Gibbs Sampling. The robustness of the match measures and the use of unsupervised discriminant analysis in the classification ensures that the proposed method performs consistently in the presence of missing/false features, temporally and spatially changing illumination conditions, and systematic misalignment caused by different camera configurations. Extensive experiments based on real data of over 200 vehicles at different times of day demonstrate promising results. Ying Shan, Harpreet Sawhney, Rakesh Kumar 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2007 | PEET: Prototype Embedding and Embedding Transition for Matching Vehicles over Disparate ViewpointsabstractThis paper presents a novel framework, prototype embedding and embedding transition (PEET), for matching objects, especially vehicles, that undergo drastic pose, appearance, and even modality changes. The problem of matching objects seen under drastic variations is reduced to matching embeddings of object appearances instead of matching the object images directly. An object appearance is first embedded in the space of a representative set of model prototypes (prototype embedding (PE)). Objects captured at disparate temporal and spatial sites are embedded in the space of prototypes that are rendered with the pose of the cameras at the respective sites. Low dimensional embedding vectors are subsequently matched. A significant feature of our approach is that no mapping function is needed to compute the distance between embedding vectors extracted from objects viewed from disparate pose and appearance changes, instead, an embedding transition (ET) scheme is utilized to implicitly realize the complex and non-linear mapping with high accuracy. The heterogeneous nature of matching between high-resolution and low-resolution image objects in PEET is discussed, and an unsupervised learning scheme based on the exploitation of the heterogeneous nature is developed to improve the overall matching performance of mixed resolution objects. The proposed approach has been applied to vehicular object classification and query application, and the extensive experimental results demonstrate the efficacy and versatility of the PEET framework. Yanlin Guo, Ying Shan, Harpreet Sawhney, Rakesh Kumar 0001 |
CVPR | 2 |
| 2007 | Robust Object Matching for Persistent Tracking with Heterogeneous FeaturesabstractThis paper addresses the problem of matching vehicles across multiple sightings under variations in illumination and camera poses. Since multiple observations of a vehicle are separated in large temporal and/or spatial gaps, thus prohibiting the use of standard frame-to-frame data association, we employ features extracted over a sequence during one time interval as a vehicle fingerprint that is used to compute the likelihood that two or more sequence observations are from the same or different vehicles. Furthermore, since our domain is aerial video tracking, in order to deal with poor image quality and large resolution and quality variations, our approach employs robust alignment and match measures for different stages of vehicle matching. Most notably, we employ a heterogeneous collection of features such as lines, points, and regions in an integrated matching framework. Heterogeneous features are shown to be important. Line and point features provide accurate localization and are employed for robust alignment across disparate views. The challenges of change in pose, aspect, and appearances across two disparate observations are handled by combining a novel feature-based quasi-rigid alignment with flexible matching between two or more sequences. However, since lines and points are relatively sparse, they are not adequate to delineate the object and provide a comprehensive matching set that covers the complete object. Region features provide a high degree of coverage and are employed for continuous frames to provide a delineation of the vehicle region for subsequent generation of a match measure. Our approach reliably delineates objects by representing regions as robust blob features and matching multiple regions to multiple regions using Earth Mover's Distance (EMD). Extensive experimentation under a variety of real-world scenarios and over hundreds of thousands of Confirmatory Identification (CID) trails has demonstrated about 95 percent accuracy in vehicle reacquisition with both visible and Infrared (IR) imaging cameras. Yanlin Guo, Steven C. Hsu, Harpreet Sawhney, Rakesh Kumar 0001, Ying Shan |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2006 | Learning Exemplar-Based Categorization for the Detection of Multi-View Multi-Pose ObjectsabstractThis paper proposes a novel approach for multi-view multi-pose object detection using discriminative shapebased exemplars. The key idea underlying this method is motivated by numerous previous observations that manually clustering multi-view multi-pose training data into different categories and then combining the separately trained two-class classifiers greatly improved the detection performance. A novel computational framework is proposed to unify different processes of categorization, training individual classifier for each intra-class category, and training a strong classifier combining the individual classifiers. The individual processes employ a single objective function that is optimized using two nested AdaBoost loops. The outer AdaBoost loop is used to select discriminative exemplars and the inner AdaBoost is used to select discriminative features on the selected exemplars. The proposed approach replaces the manual time-consuming process of exemplar selection as well as addresses the problem of labeling ambiguity inherent in this process. Also, our approach fully complies with the standard AdaBoost-based object detection framework in terms of real-time implementation. Experiments on multi-view multi-pose people and vehicle data demonstrate the efficacy of the proposed approach. Ying Shan, Feng Han 0002, Harpreet Sawhney, Rakesh Kumar 0001 |
CVPR (2) | 1 |
| 2006 | Rapid Object Indexing Using Locality Sensitive Hashing and Joint 3D-Signature Space EstimationabstractWe propose a new method for rapid 3D object indexing that combines feature-based methods with coarse alignment-based matching techniques. Our approach achieves a sublinear complexity on the number of models, maintaining at the same time a high degree of performance for real 3D sensed data that is acquired in largely uncontrolled settings. The key component of our method is to first index surface descriptors computed at salient locations from the scene into the whole model database using the Locality Sensitive Hashing (LSH), a probabilistic approximate nearest neighbor method. Progressively complex geometric constraints are subsequently enforced to further prune the initial candidates and eliminate false correspondences due to inaccuracies in the surface descriptors and the errors of the LSH algorithm. The indexed models are selected based on the MAP rule using posterior probability of the models estimated in the joint 3D-signature space. Experiments with real 3D data employing a large database of vehicles, most of them very similar in shape, containing 1,000,000 features from more than 365 models demonstrate a high degree of performance in the presence of occlusion and obscuration, unmodeled vehicle interiors and part articulations, with an average processing time between 50 and 100 seconds per query. Bogdan Matei, Ying Shan, Harpreet Sawhney, Rakesh Kumar 0001, Daniel F. Huber, Martial Hebert |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2006 | Shapeme Histogram Projection and Matching for Partial Object RecognitionabstractHistograms of shape signature or prototypical shapes, called shapemes, have been used effectively in previous work for 2D/3D shape matching and recognition. We extend the idea of shapeme histogram to recognize partially observed query objects from a database of complete model objects. We propose representing each model object as a collection of shapeme histograms and match the query histogram to this representation in two steps: 1) compute a constrained projection of the query histogram onto the subspace spanned by all the shapeme histograms of the model and 2) compute a match measure between the query histogram and the projection. The first step is formulated as a constrained optimization problem that is solved by a sampling algorithm. The second step is formulated under a Bayesian framework, where an implicit feature selection process is conducted to improve the discrimination capability of shapeme histograms. Results of matching partially viewed range objects with a 243 model database demonstrate better performance than the original shapeme histogram matching algorithm and other approaches. Ying Shan, Harpreet Sawhney, Bogdan Matei, Rakesh Kumar 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2005 | Vehicle Fingerprinting for Reacquisition and Tracking in VideosabstractVisual recognition of objects through multiple observations is an important component of object tracking. We address the problem of vehicle matching when multiple observations of a vehicle are separated in time such that frames of observations are not contiguous, thus prohibiting the use of standard frame-to-frame data association. We employ features extracted over a sequence during one time interval as a vehicle fingerprint that is used to compute the likelihood that two or more sequence observations are from the same or different vehicles. The challenges of change in pose, aspect and appearances across two disparate observations are handled by combining feature-based quasi-rigid alignment with flexible matching between two or more sequences. The current work uses the domain of vehicle tracking from aerial platforms where typically both the imaging platform and the vehicles are moving and the number of pixels on the object are limited to fairly low resolutions. Extensive evaluation with respect to ground truth is reported in the paper. Yanlin Guo, Steven C. Hsu, Ying Shan, Harpreet Sawhney, Rakesh Kumar 0001 |
CVPR (2) | 3 |
| 2005 | Unsupervised Learning of Discriminative Edge Measures for Vehicle Matching between Non-Overlapping CamerasabstractThis paper proposes a method for matching road vehicles between two non-overlapping cameras. The matching problem is formulated as a same-different classification problem: probability of two observations from two distinct cameras being from the same vehicle or from different vehicles. We employ a measurement vector consists of three independent edge-based measures and their associated robust measures computed from a pair of aligned vehicle edge maps. The weight of each match measure in the final decision is determined by a unsupervised learning process so that the same-different classification can be optimally separated in the combined measurement space. The robustness of the match measures and the use of discriminant analysis in the classification ensure that the proposed method performs better than existing edge-based approaches, especially in the presence of missing/false edges caused by shadows and different illumination conditions, and systematic misalignment caused by different camera configurations. Extensive experiments based on real data of over 200 vehicles at different times of day demonstrate promising results. Ying Shan, Harpreet Sawhney, Rakesh Kumar 0001 |
CVPR (1) | 1 |
| 2005 | Vehicle Identification between Non-Overlapping Cameras without Direct Feature MatchingabstractWe propose a novel method for identifying road vehicles between two nonoverlapping cameras. The problem is formulated as a same-different classification problem: probability of two vehicle images from two distinct cameras being from the same vehicle or from different vehicles. The key idea is to compute the probability without matching the two vehicle images directly, which is a process vulnerable to drastic appearance and aspect changes. We represent each vehicle image as an embedding amongst representative exemplars of vehicles within the same camera. The embedding is computed as a vector each of whose components is a nonmetric distance for a vehicle to an exemplar. The nonmetric distances are computed using robust matching of oriented edge images. A set of truthed training examples of same-different vehicle pairings across the two cameras is used to learn a classifier that encodes the probability distributions. A pair of the embeddings representing two vehicles across two cameras is then used to compute the same-different probability. In order for the vehicle exemplars to be representative for both cameras, we also propose a method for jointly selection of corresponding exemplars using the training data. Experiments on observations of over 400 vehicles under drastically illumination and camera conditions demonstrate promising results. Ying Shan, Harpreet Sawhney, Rakesh Kumar 0001 |
ICCV | 1 |
| 2005 | Clustering multiple image sequences with a sequence-to-sequence similarity measureabstractWe propose a novel similarity measure of two image sequences based on shapeme histograms. The idea of shapeme histogram has been used for single image/texture recognition, but is used here to solve the sequence-to-sequence matching problem. We develop techniques to represent each sequence as a set of shapeme histograms, which captures different variations of the object appearances within the sequence. These shapeme histograms are computed from the set of 2D invariant features that are stable across multiple images in the sequence, and therefore minimizes the effect of both background clutter, and 2D pose variations. We define sequence similarity measure as the similarity of the most similar pair of images from both sequences. This definition maximizes the chance of matching between two sequences of the same object, because it requires only part of the sequences being similar. We also introduce a weighting scheme to conduct an implicit feature selection process during the matching of two shapeme histograms. Experiments on clustering image sequences of tracked objects demonstrate the efficacy of the proposed method. Ying Shan, Harpreet Sawhney, Art Pope |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2004 | Linear Model Hashing and Batch RANSAC for Rapid and Accurate Object Recognition
Ying Shan, Bogdan Matei, Harpreet Sawhney, Rakesh Kumar 0001, Daniel F. Huber, Martial Hebert |
CVPR (2) | 1 |
| 2004 | Partial Object Matching with Shapeme Histograms
Ying Shan, Harpreet Sawhney, Bogdan Matei, Rakesh Kumar 0001 |
ECCV (3) | 1 |
| 2004 | Robust and Rapid Generation of Animated Faces from Video Images: A Model-Based Modeling Approach
Zhengyou Zhang, Zicheng Liu 0001, Dennis Adler, Michael F. Cohen, Erik Hanson, Ying Shan |
Int. J. Comput. Vis. | 6 |
| 2003 | Incremental motion estimation through modified bundle adjustmentabstractAn incremental motion estimation scheme for long image sequence analysis is introduced. It applies to a sliding window of triplet images and maintains local motion consistency without resort to post-concatenation. This is possible thanks to our newly developed local process called three-view partial bundle adjustment. Unlike previous approaches that rely only on point matches across three or more views, our local process also takes into account all available two-view matches, leading to more accurate motion estimation. For sparse image sequences, two-view matches are very reliable and it becomes even more important to use them since the number of matches across more views decreases quickly. In this case, our incremental method can produce results very close to those obtained with a global bundle adjustment but in a fraction of time. Experiments with both synthetic and real data have been conducted to compare the proposed technique with other techniques, and have shown our technique to be clearly superior. Zhengyou Zhang, Ying Shan |
ICIP (2) | 2 |
| 2002 | New Measurements and Corner-Guidance for Curve Matching with Probabilistic Relaxation
Ying Shan, Zhengyou Zhang |
Int. J. Comput. Vis. | 1 |
| 2001 | Image-Based Surface Detail TransferabstractWe present a novel technique, called Image-Based Surface Detail Transfer, to transfer geometric details from one surface to another with simple 2D image operations. The basic observation is that, without knowing its 3D geometry, geometric details (local deformations) can be extracted from a single image of an object in a way independent of its surface reflectance, and furthermore, these geometric details can be transferred to modify the appearance of other objects directly in images. We show examples including surface detail transfer between real objects, as well as between real and synthesized objects. Ying Shan, Zicheng Liu 0001, Zhengyou Zhang |
CVPR (2) | 1 |
| 2001 | Model-Based Bundle Adjustment with Application to Face ModelingabstractWe present a new model-based bundle adjustment algorithm to recover the 3D model of a scene/object from a sequence of images with unknown motions. Instead of representing scene/object by a collection of isolated 3D features (usually points), our algorithm uses a surface controlled by a small set of parameters. Compared with previous model-based approaches, our approach has the following advantages. First instead of using the model space as a regularizer we directly use it as our search space, thus resulting in a more elegant formulation with fewer unknowns and fewer equations. Second, our algorithm automatically associates tracked points with their correct locations on the surfaces, thereby eliminating the need for a prior 2D-to-3D association. Third, regarding face modeling, we use a very small set of face metrics (meaningful deformations) to parameterize the face geometry, resulting in a smaller search space and a better posed system. Experiments with both synthetic and real data show that this new algorithm is faster, more accurate and more stable than existing ones. Ying Shan, Zicheng Liu 0001, Zhengyou Zhang |
ICCV | 1 |
| 2001 | Cloning Your Own Face with a Desktop CameraabstractWe have developed an easy and cost-effective system that constructs textured 3D animated face models from videos with minimal user interaction. Our system first takes, with an ordinary video camera, images of a face of a person sitting in front of the camera turning the head from one side to the other. After five manual clicks on two images to tell the system where the eye corners, nose top and mouth corners are, the system automatically generates a realistic looking 3D human head model and the constructed model can be animated immediately (different poses, facial expressions and talking). A user, with a PC and a video camera, can use our system to generate hisher face model in a few minutes. The face model can then be imported in hisher favorite game, and the user sees themselves and their friends take part in the game they are playing. We will demonstrate the system on a laptop computer live at the conference, and participants can try it to model their own faces. Zhengyou Zhang, Zicheng Liu 0001, Dennis Adler, Michael F. Cohen, Erik Hanson, Ying Shan |
ICCV | 6 |
| 2001 | Expressive expression mapping with ratio imagesabstractFacial expressions exhibit not only facial feature motions, but also subtle changes in illumination and appearance (e.g., facial creases and wrinkles). These details are important visual cues, but they are difficult to synthesize. Traditional expression mapping techniques consider feature motions while the details in illumination changes are ignored. In this paper, we present a novel technique for facial expression mapping. We capture the illumination change of one person's expression in what we call an expression ratio image (ERI). Together with geometric warping, we map an ERI to any other person's face image to generate more expressive facial expressions. Zicheng Liu 0001, Ying Shan, Zhengyou Zhang |
SIGGRAPH | 2 |
| 2000 | Corner Guided Curve Matching and its Application to Scene ReconstructionabstractCorners and curves are important image features in many vision-based applications. Corners are usually more stable and easier to match than curves, while curves contain richer information of scene structure. In previously work, corners are often used to recover the epipolar geometry between two views, which is then used in curve matching to reduce the search space. However, information of the scene structure contained in this set of matched corners is ignored. In this paper we present a curve matching algorithm that is guided by a set of matched corners. Within a probabilistic framework, the role of the corner guidance is explicitly defined by a set of similarity-invariant unary measurements and by a similarity function. The similarity function provides stronger capability of resolving matching ambiguity than the epipolar constraint, and is integrated into a relaxation scheme to reduce computational complexity and improve accuracy of curve matching. Experimental results clearly demonstrate the benefit of integrating corner matches into the curve matching procedure. Ying Shan, Zhengyou Zhang |
CVPR | 1 |