EDBT 2026 Demo / reviewers in the wild / expert
Ying Tai
dblp:158/1384
· DBLP profile ↗
111ranked-venue papers
8as first author
74since 2021 · last 2026
0000-0002-4665-6852ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 90 · 6 first-author · 61 since 2021Graphics, computer vision, multimedia, augmented reality and games · 88 · 4 first-author · 61 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scene Prior Filtering for Depth Super-Resolution
Zhengxue Wang, Zhiqiang Yan 0001, Ming-Hsuan Yang 0001, Jinshan Pan, Guangwei Gao, Ying Tai, Jian Yang 0003 |
Int. J. Comput. Vis. | 6 |
| 2026 | Correction: Scene Prior Filtering for Depth Super-Resolution
Zhengxue Wang, Zhiqiang Yan 0001, Ming-Hsuan Yang 0001, Jinshan Pan, Guangwei Gao, Ying Tai, Jian Yang 0003 |
Int. J. Comput. Vis. | 6 |
| 2026 | Curriculum adaptation for one-stream RGB-T tracking
Xiantao Hu, Fansheng Zeng, Bineng Zhong 0001, Zhangyong Tang, Wenxuan Fang 0001, Jun Li 0027, Ying Tai, Jian Yang 0003 |
Pattern Recognit. | 7 |
| 2026 | AddSR: Accelerating diffusion-based blind super-resolution with adversarial diffusion distillation
Ying Tai, Rui Xie 0005, Chen Zhao 0002, Kai Zhang 0008, Zhenyu Zhang 0005, Jian Yang 0003 |
Pattern Recognit. | 1 |
| 2026 | Spiking pyramid wavelet transformation for high-efficient and low-energy image restoration
Chen Zhao 0002, Xiantao Hu, Rui Xie 0005, Jian Yang 0003, Ying Tai |
Pattern Recognit. | 8 |
| 2026 | Learning multi-scale spatial-frequency features for image denoising
Xu Zhao 0001, Chen Zhao 0002, Xiantao Hu, Hongliang Zhang 0002, Ying Tai, Jian Yang 0003 |
Pattern Recognit. | 5 |
| 2026 | MambaGesture2: Co-Speech Gesture Generation via Hierarchical Fusion and Spatiotemporal AggregationabstractCo-speech gesture generation plays a vital role in producing synchronized and natural human gestures, thereby enhancing the realism of avatars in virtual environments. Although diffusion models have shown strong generative capabilities, their combination with transformer-based architectures often incurs high computational costs due to the quadratic complexity of self-attention. Moreover, as a temporal sequence modeling task, existing methods frequently struggle to effectively capture multi-scale temporal dynamics inherent in speech and gesture signals. To address these challenges, we propose MambaGesture2, a novel framework that integrates a Mamba-based denoising network, Hierarchical U-Net Gesture Mamba (HUG-Mamba), with a multimodal feature fusion module, SEAD. HUG-Mamba combines the efficient state-space modeling of Mamba blocks with the hierarchical sampling of the U-Net architecture, significantly improving temporal coherence and computational efficiency. We further introduce the Temporal-Stratified Fusion (TSF) module to capture diverse temporal scales via multi-scale learning, and the Spatial-Temporal Cascaded Aggregation (STCA) module to enhance spatial-temporal feature aggregation. Extensive experiments on the multi-modal BEAT2 and SHOW datasets demonstrate that our approach achieves state-of-the-art performance across multiple quantitative metrics, while substantially reducing model complexity and inference time. The results validate the effectiveness of our architectural innovations in generating diverse, realistic, and temporally consistent co-speech gestures. Project page:https://fcchit.github.io/mambagesture2. Chencan Fu, Yabiao Wang, Haoyang He, Chengjie Wang 0001, Ying Tai, Yong Liu 0007, Jiangning Zhang |
IEEE Trans. Multim. | 6 |
| 2026 | DreamBarbie: Text to Barbie-Style 3D AvatarsabstractTo integrate digital humans into everyday life, there is a strong demand for generating high-quality, fine-grained disentangled 3D avatars that support expressive animation and simulation capabilities, ideally from low-cost textual inputs. Although text-driven 3D avatar generation has made significant progress by leveraging 2D generative priors, existing methods still struggle to fulfill all these requirements simultaneously. To address this challenge, we propose DreamBarbie, a novel text-driven framework for generating animatable 3D avatars with separable shoes, accessories, and simulation-ready garments, truly capturing the iconic "Barbie doll" aesthetic. The core of our framework lies in an expressive 3D representation combined with appropriate modeling constraints. Unlike prior methods, we use G-Shell to uniformly model watertight components (e.g., bodies, shoes) and non-watertight garments. By reformulating boundaries as euclidean field intersections instead of manifold geodesics, we propose an SDF-based initialization and a hole regularization loss that together achieve a $100\times$100× speedup and stable open topology without image input. These disentangled 3D representations are then optimized by specialized expert diffusion models tailored to each domain, ensuring high-fidelity outputs. To mitigate geometric artifacts and texture conflicts when combining different expert models, we further propose several effective geometric losses and strategies. Extensive experiments demonstrate that DreamBarbie outperforms existing methods in both dressed human and outfit generation. Our framework further enables diverse applications, including apparel combination, editing, expressive animation, and physical simulation. Xiaokun Sun, Zhenyu Zhang 0005, Ying Tai, Hao Tang 0005, Zili Yi, Jian Yang 0003 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | Guided Real Image Dehazing Using YCbCr Color SpaceabstractImage dehazing, particularly with learning-based methods, has gained significant attention due to its importance in real-world applications. However, relying solely on the RGB color space often fall short, frequently leaving residual haze. This arises from two main issues: the difficulty in obtaining clear textural features from hazy RGB images and the complexity of acquiring real haze/clean image pairs outside controlled environments like smoke-filled scenes. To address these issues, we first propose a novel Structure Guided Dehazing Network (SGDN) that leverages the superior structural properties of YCbCr features over RGB. It comprises two key modules: Bi-Color Guidance Bridge (BGB) and Color Enhancement Module (CEM). BGB integrates a phase integration module and an interactive attention module, utilizing the rich texture features of the YCbCr space to guide the RGB space, thereby recovering clearer features in both frequency and spatial domains. To maintain tonal consistency, CEM further enhances the color perception of RGB features by aggregating YCbCr channel information. Furthermore, for effective supervised learning, we introduce a Real-World Well-Aligned Haze dataset, which includes a diverse range of scenes from various geographical regions and climate conditions. Experimental results demonstrate that our method surpasses existing state-of-the-art methods across multiple real-world smoke/haze datasets. Wenxuan Fang 0001, Junkai Fan, Yu Zheng 0036, Jiangwei Weng, Ying Tai, Jun Li 0027 |
AAAI | 5 |
| 2025 | Exploiting Multimodal Spatial-temporal Patterns for Video Object TrackingabstractMultimodal tracking has garnered widespread attention as a result of its ability to effectively address the inherent limitations of traditional RGB tracking. However, existing multimodal trackers mainly focus on the fusion and enhancement of spatial features or merely leverage the sparse temporal relationships between video frames. These approaches do not fully exploit the temporal correlations in multimodal videos, making it difficult to capture the dynamic changes and motion information of targets in complex scenarios. To alleviate this problem, we propose a unified multimodal spatial-temporal tracking approach named STTrack. In contrast to previous paradigms that solely relied on updating reference information, we introduced a temporal state generator (TSG) that continuously generates a sequence of tokens containing multimodal temporal information. These temporal information tokens are used to guide the localization of the target in the next time state, establish long-range contextual relationships between video frames, and capture the temporal trajectory of the target. Furthermore, at the spatial level, we introduced the mamba fusion and background suppression interactive (BSI) modules. These modules establish a dual-stage mechanism for coordinating information interaction and fusion between modalities. Extensive comparisons on five benchmark datasets illustrate that STTrack achieves state-of-the-art performance across various multimodal tracking scenarios. Xiantao Hu, Ying Tai, Xu Zhao 0001, Chen Zhao 0002, Zhenyu Zhang 0005, Jun Li 0027, Bineng Zhong 0001, Jian Yang 0003 |
AAAI | 2 |
| 2025 | Learning Generalized Residual Exchange-Correlation-Uncertain Functional for Density Functional TheoryabstractDensity Functional Theory (DFT) stands as a widely used and efficient approach for addressing the many-electron Schrödinger equation across various domains such as physics, chemistry, and biology. However, a core challenge that persists over the long term pertains to refining the exchange-correlation (XC) approximation. This approximation significantly influences the triumphs and shortcomings observed in DFT applications. Nonetheless, a prevalent issue among XC approximations is the presence of systematic errors, stemming from deviations from the mathematical properties of the exact XC functional. For example, although both B3LYP and DM21 (DeepMind 21) exhibit improvements over previous benchmarks, there is still potential for further refinement. In this paper, we propose a strategy for enhancing XC approximations by estimating the neural uncertainty of the XC functional, named Residual XC-Uncertain Functional. Specifically, our approach involves training a neural network to predict both the mean and variance of the XC functional, treating it as a Gaussian distribution. To ensure stability in each sampling point, we construct the mean by combining traditional XC approximations with our neural predictions, mitigating the risk of divergence or vanishing values. It is crucial to highlight that our methodology excels particularly in cases where systematic errors are pronounced. Empirical outcomes from three benchmark tests substantiate the superiority of our approach over existing state-of-the-art methods. Our approach not only surpasses related techniques but also significantly outperforms both the popular B3LYP and the recent DM21 methods, achieving average RMSE improvements of 62% and 37%, respectively, across the three benchmarks: W4-17, G21EA, and G21IP. Sizhuo Jin, Jianjun Qian, Ying Tai |
AAAI | 4 |
| 2025 | Anywhere: A Multi-Agent Framework for User-Guided, Reliable, and Diverse Foreground-Conditioned Image GenerationabstractRecent advancements in image-conditioned image generation have demonstrated substantial progress. However, foreground-conditioned image generation remains underexplored, encountering challenges such as compromised object integrity, foreground-background inconsistencies, limited diversity, and reduced control flexibility. These challenges arise from current end-to-end inpainting models, which suffer from inaccurate training masks, limited foreground semantic understanding, data distribution biases, and inherent interference between visual and textual prompts. To overcome these limitations, we present Anywhere, a multi-agent framework that departs from the traditional end-to-end approach. In this framework, each agent is specialized in a distinct aspect, such as foreground understanding, diversity enhancement, object integrity protection, and textual prompt consistency. Our framework is further enhanced with the ability to incorporate optional user textual inputs, perform automated quality assessments, and initiate re-generation as needed. Comprehensive experiments demonstrate that this modular design effectively overcomes the limitations of existing end-to-end models, resulting in higher fidelity, quality, diversity and controllability in foreground-conditioned image generation. Additionally, the Anywhere framework is extensible, allowing it to benefit from future advancements in each individual agent. Tianyidan Xie, Rui Ma 0011, Xiaoqian Ye, Feixuan Liu, Ying Tai, Zhenyu Zhang 0005, Lanjun Wang, Zili Yi |
AAAI | 6 |
| 2025 | From Words to Worth: Newborn Article Impact Prediction with LLMabstractPredicting the future impact of newly published articles is pivotal for advancing scientific discovery in an era of unprecedented scholarly expansion. This paper introduces a promising approach, leveraging the capabilities of LLMs to predict the future impact of newborn articles solely based on titles and abstracts. Breaking away from traditional methods heavily reliant on external data, we propose fine-tuning the LLM to uncover the intrinsic semantic patterns shared by highly impactful articles from a vast collection of text-score pairs. These semantic features are further utilized to predict the proposed indicator, TNCSIsp, which incorporates favorable normalization properties across value, field, and time. To facilitate parameter-efficient fine-tuning of the LLM, we have also meticulously curated a dataset containing over 12,000 entries, each annotated with titles, abstracts, and their corresponding TNCSIsp values. Experimental results reveal an MAE of 0.216 and an NDCG@20 of 0.901, setting new benchmarks in predicting the impact of newborn articles. Finally, we present a real-world application example for predicting the impact of newborn journal articles to demonstrate its noteworthy practical value. Overall, our findings challenge existing paradigms and propose a shift towards a more content-focused prediction of academic impact, offering new insights for article impact prediction. Penghai Zhao, Kairan Dou, Jinyu Tian 0006, Ying Tai, Jian Yang 0003, Ming-Ming Cheng, Xiang Li 0041 |
AAAI | 5 |
| 2025 | InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured CaptionabstractText-to-video generation has evolved rapidly in recent years, delivering remarkable results. Training typically relies on video-caption paired data, which plays a crucial role in enhancing generation performance. However, current video captions often suffer from insufficient details, hallucinations and imprecise motion depiction, affecting the fidelity and consistency of generated videos. In this work, we propose a novel instance-aware structured caption framework, termed InstanceCap, to achieve instance-level and fine-grained video caption for the first time. Based on this scheme, we design an auxiliary models cluster to convert original video into instances to enhance instance fidelity. Video instances are further used to refine dense prompts into structured phrases, achieving concise yet precise descriptions. Furthermore, a 22K InstanceVid dataset is curated for training, and an enhancement pipeline that tailored to InstanceCap structure is proposed for inference. Experimental results demonstrate that our proposed InstanceCap significantly outperform previous models, ensuring high fidelity between captions and videos while reducing hallucinations. Tiehan Fan, Kepan Nan, Rui Xie 0005, Penghao Zhou, Zhenheng Yang, Chaoyou Fu, Xiang Li 0041, Jian Yang 0003, Ying Tai |
CVPR | 9 |
| 2025 | Towards Universal Dataset Distillation via Task-Driven DiffusionabstractDataset distillation (DD) condenses key information from large-scale datasets into smaller synthetic datasets, reducing storage and computational costs for training networks. However, most recent research has primarily focused on image classification tasks, with limited exploration in detection and segmentation. Two key challenges remain: (i) Task Optimization Heterogeneity, where existing methods focus on class-level information but fail to address the diverse needs of detection and segmentation, and (ii) Inflexible Image Generation, where current generation methods rely on global updates for single-class targets and lack localized optimization for specific object regions. To address these challenges, we propose UniDD, a universal dataset distillation framework built on a task-driven diffusion model for diverse DD tasks, as shown in Fig. 1. Our approach operates in two stages: Universal Task Knowledge Mining, which captures task-relevant information through task-specific proxy model training, and Universal Task-Driven Diffusion, where these proxies guide the diffusion process to generate task-specific synthetic images. Extensive experiments across ImageNet-1K, Pascal VOC, and MS COCO demonstrate that UniDD consistently outperforms state-of-the-art methods. In particular, on ImageNet-1K with IPC-10, UniDD surpasses previous diffusion-based methods by 6.1%, while also reducing deployment costs. Ding Qi, Jian Li 0062, Junyao Gao 0002, Shuguang Dou, Ying Tai, Jianlong Hu, Bo Zhao 0015, Yabiao Wang, Chengjie Wang 0001, Cairong Zhao |
CVPR | 5 |
| 2025 | From Zero to Detail: Deconstructing Ultra-High-Definition Image Restoration from Progressive Spectral PerspectiveabstractUltra-high-definition (UHD) image restoration faces significant challenges due to its high resolution, complex content, and intricate details. To cope with these challenges, we analyze the restoration process in depth through a progressive spectral perspective, and deconstruct the complex UHD restoration problem into three progressive stages: zero-frequency enhancement, low-frequency restoration, and high-frequency refinement. Building on this insight, we propose a novel framework, ERR, which comprises three collaborative sub-networks: the zero-frequency enhancer (ZFE), the low-frequency restorer (LFR), and the high-frequency refiner (HFR). Specifically, the ZFE integrates global priors to learn global mapping, while the LFR restores low-frequency information, emphasizing reconstruction of coarse-grained content. Finally, the HFR employs our designed frequency-windowed kolmogorov-arnold networks (FW-KAN) to refine textures and details, producing high-quality image restoration. Our approach significantly outperforms previous UHD methods across various tasks, with extensive ablation studies validating the effectiveness of each component. The code is available at here. Zhizhou Chen, Yunzhe Xu, Enxuan Gu, Zili Yi, Ying Tai |
CVPR | 9 |
| 2025 | RAGD: Regional-Aware Diffusion Model for Text-to-Image Generation
Zhennan Chen, Zhibo Chen 0011, Zhengkai Jiang 0003, Ying Tai |
ICCV | 9 |
| 2025 | Describe, Don't Dictate: Semantic Image Editing with Natural Language IntentabstractDespite the progress in text-to-image generation, semantic image editing remains a challenge. Inversion-based algorithms unavoidably introduce reconstruction errors, while instruction-based models mainly suffer from limited dataset quality and scale. To address these problems, we propose a descriptive-prompt-based editing framework, named DescriptiveEdit. The core idea is to re-frame `instruction-based image editing' as `reference-image-based text-to-image generation', which preserves the generative power of well-trained Text-to-Image models without architectural modifications or inversion. Specifically, taking the reference image and a prompt as input, we introduce a Cross-Attentive UNet, which newly adds attention bridges to inject reference image features into the prompt-to-edit-image generation process. Owing to its text-to-image nature, DescriptiveEdit overcomes limitations in instruction dataset quality, integrates seamlessly with ControlNet, IP-Adapter, and other extensions, and is more scalable. Experiments on the Emu Edit benchmark show it improves editing accuracy and consistency. En Ci, Shanyan Guan, Yanhao Ge, Zhenyu Zhang 0005, Jian Yang 0003, Ying Tai |
ICCV | 8 |
| 2025 | Reverse Convolution and its Applications to Image Restoration
Xuhong Huang, Ying Tai |
ICCV | 4 |
| 2025 | StrandHead: Text to Hair-Disentangled 3D Head Avatars Using Human-Centric Priors
Xiaokun Sun, Ying Tai, Jian Yang 0003, Zhenyu Zhang 0005 |
ICCV | 3 |
| 2025 | Star: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-ResolutionabstractImage diffusion models have been adapted for real-world video super-resolution to tackle over-smoothing issues in GAN-based methods. However, these models struggle to maintain temporal consistency, as they are trained on static images, limiting their ability to capture temporal dynamics effectively. Integrating text-to-video (T2V) models into video super-resolution for improved temporal modeling is straightforward. However, two key challenges remain: artifacts introduced by complex degradations in real-world scenarios, and compromised fidelity due to the strong generative capacity of powerful T2V models (\textit{e.g.}, CogVideoX-5B). To enhance the spatio-temporal quality of restored videos, we introduce\textbf{~\name} (\textbf{S}patial-\textbf{T}emporal \textbf{A}ugmentation with T2V models for \textbf{R}eal-world video super-resolution), a novel approach that leverages T2V models for real-world video super-resolution, achieving realistic spatial details and robust temporal consistency. Specifically, we introduce a Local Information Enhancement Module (LIEM) before the global attention block to enrich local details and mitigate degradation artifacts. Moreover, we propose a Dynamic Frequency (DF) Loss to reinforce fidelity, guiding the model to focus on different frequency components across diffusion steps. Extensive experiments demonstrate\textbf{~\name}~outperforms state-of-the-art methods on both synthetic and real-world datasets. Rui Xie 0005, Yinhong Liu, Penghao Zhou, Chen Zhao 0002, Kai Zhang 0008, Zhenyu Zhang 0005, Jian Yang 0003, Zhenheng Yang, Ying Tai |
ICCV | 10 |
| 2025 | OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video GenerationabstractText-to-video (T2V) generation has recently garnered significant attention thanks to the large multi-modality model Sora. However, T2V generation still faces two important challenges: 1) Lacking a precise open sourced high-quality dataset. The previously popular video datasets, e.g.WebVid-10M and Panda-70M, overly emphasized large scale, resulting in the inclusion of many low-quality videos and
short, imprecise captions. Therefore, it is challenging but crucial to collect a precise high-quality dataset while maintaining a scale of millions for T2V generation. 2) Ignoring to fully utilize textual information. Recent T2V methods have focused on vision transformers, using a simple cross attention module for video generation, which falls short of making full use of semantic information from text tokens. To address these issues, we introduce OpenVid-1M, a precise high-quality dataset with expressive captions. This open-scenario dataset contains over 1 million text-video pairs, facilitating research on T2V generation. Furthermore, we curate 433K 1080p videos from OpenVid-1M to create OpenVidHD-0.4M, advancing high-definition video generation. Additionally, we propose a novel Multi-modal Video Diffusion Transformer (MVDiT) capable of mining both structure information from visual tokens and semantic information from text tokens. Extensive experiments and ablation studies verify the superiority of OpenVid-1M over previous datasets and the effectiveness of our MVDiT. Kepan Nan, Rui Xie 0005, Penghao Zhou, Tiehan Fan, Zhenheng Yang, Xiang Li 0041, Jian Yang 0003, Ying Tai |
ICLR | 9 |
| 2025 | EasySplat: View-Adaptive Learning makes 3D Gaussian Splatting Easyabstract3D Gaussian Splatting (3DGS) techniques have achieved satisfactory 3D scene representation. Despite their impressive performance, they confront challenges due to the limitation of structure-from-motion (SfM) methods on acquiring accurate scene initialization, or the inefficiency of densification strategy. In this paper, we introduce a novel framework EasySplat to achieve high-quality 3DGS modeling. Instead of using SfM for scene initialization, we employ a novel method to release the power of large-scale pointmap approaches. Specifically, we propose an efficient grouping strategy based on view similarity, and use robust pointmap priors to obtain high-quality point clouds and camera poses for 3D scene initialization. After obtaining a reliable scene structure, we propose a novel densification approach that adaptively splits Gaussian primitives based on the average shape of neighboring Gaussian ellipsoids, utilizing KNN scheme. In this way, the proposed method tackles the limitation on initialization and optimization, leading to an efficient and accurate 3DGS modeling. Extensive experiments demonstrate that EasySplat outperforms the current state-of-the-art (SOTA) in handling novel view synthesis. Ao Gao, Luosong Guo, Ying Tai, Jian Yang 0003, Zhenyu Zhang 0005 |
ICME | 5 |
| 2025 | Category-Aware 3D Object Composition with Disentangled Texture and Shape Multi-view Diffusion
Zeren Xiong, Zedong Zhang, Xiang Li 0041, Ying Tai, Jian Yang 0003, Jun Li 0027 |
ACM Multimedia | 5 |
| 2025 | UltraHR-100K: Enhancing UHR Image Synthesis with A Large-Scale High-Quality DatasetabstractUltra-high-resolution (UHR) text-to-image (T2I) generation has seen notable progress. However, two key challenges remain : 1) the absence of a large-scale high-quality UHR T2I dataset, and (2) the neglect of tailored training strategies for fine-grained detail synthesis in UHR scenarios. To tackle the first challenge, we introduce \textbf{UltraHR-100K}, a high-quality dataset of 100K UHR images with rich captions, offering diverse content and strong visual fidelity. Each image exceeds 3K resolution and is rigorously curated based on detail richness, content complexity, and aesthetic quality. To tackle the second challenge, we propose a frequency-aware post-training method that enhances fine-detail generation in T2I diffusion models. Specifically, we design (i) \textit{Detail-Oriented Timestep Sampling (DOTS)} to focus learning on detail-critical denoising steps, and (ii) \textit{Soft-Weighting Frequency Regularization (SWFR)}, which leverages Discrete Fourier Transform (DFT) to softly constrain frequency components, encouraging high-frequency detail preservation. Extensive experiments on our proposed UltraHR-eval4K benchmarks demonstrate that our approach significantly improves the fine-grained detail quality and overall fidelity of UHR image generation. The code is available at \href{https://github.com/NJU-PCALab/UltraHR-100k}{here}. Chen Zhao 0002, En Ci, Yunzhe Xu, Tiehan Fan, Shanyan Guan, Yanhao Ge, Jian Yang 0003, Ying Tai |
NeurIPS | 8 |
| 2025 | AGSwap: Overcoming Category Boundaries in Object Fusion via Adaptive Group SwappingabstractFusing cross-category objects to a single coherent object has gained increasing attention in text-to-image (T2I) generation due to its broad applications in virtual reality, digital media, film, and gaming. However, existing methods often produce biased, visually chaotic, or semantically inconsistent results due to overlapping artifacts and poor integration. Moreover, progress in this field has been limited by the absence of a comprehensive benchmark dataset. To address these problems, we propose Adaptive Group Swapping (AGSwap), a simple yet highly effective approach comprising two key components: (1) Group-wise Embedding Swapping, which fuses semantic attributes from different concepts through feature manipulation, and (2) Adaptive Group Updating, a dynamic optimization mechanism guided by a balance evaluation score to ensure coherent synthesis. Additionally, we introduce Cross-category Object Fusion (COF), a large-scale, hierarchically structured dataset built upon ImageNet-1K and WordNet. COF includes 95 superclasses, each with 10 subclasses, enabling 451,250 unique fusion pairs. Extensive experiments demonstrate that AGSwap outperforms state-of-the-art compositional T2I methods, including GPT-Image-1 using simple and complex prompts. Project Page Zedong Zhang, Ying Tai, Jianjun Qian, Jian Yang 0003, Jun Li 0027 |
SIGGRAPH Asia | 2 |
| 2025 | Daytime-Mixed Non-Aligned Learning for Real Nighttime Image EnhancementabstractEnhancing real nighttime images is a significant challenge due to the deterioration of visual quality caused by limited perceptibility under adverse illumination conditions, leading to loss of details and color deviation. In this paper, we propose a novel nighttime image enhancement framework using daytime-mixed non-aligned supervision. It aims to couple the information between non-aligned daytime and nighttime image pairs. Specifically, our framework consists of a simple yet effective daytime-mixed supervised learning phase and a Retinex-based reconstruction phase. In the first phase, we employ a multi-instance with adaptive information fusion (AIF) module integrated within a UNet enhancement network called MIFUNet, which is trained via a daytime-mixed supervised loss. In the second phase, the Retinex-based reconstruction employs both a light-effect estimation network and an illumination adjustment network to restore the nighttime image, guided by physical principles. To evaluate the effectiveness of our approach, we collect a real non-aligned day-night dataset named the NANE dataset, which contains 748 non-aligned image pairs and 100 nighttime images solely for testing. Extensive experiments demonstrate that our method achieves superior performance compared to state-of-the-art image enhancement methods. Jiangwei Weng, Junkai Fan, Jianjun Qian, Haiyang Zou, Ying Tai, Jian Yang 0003, Jun Li 0027 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | PortraitBooth: A Versatile Portrait Model for Fast Identity-Preserved PersonalizationabstractRecent advancements in personalized image generation using diffusion models have been noteworthy. However, existing methods suffer from inefficiencies due to the requirement for subject-specific fine-tuning. This computationally intensive process hinders efficient deployment, limiting practical usability. Moreover, these methods often grapple with identity distortion and limited expression diversity. In light of these challenges, we propose Portrait-Booth, an innovative approach designed for high efficiency, robust identity preservation, and expression-editable text-to-image generation, without the need for fine-tuning. Por-traitBooth leverages subject embeddingsfrom aface recognition model for personalized image generation without fine-tuning. It eliminates computational overhead and mitigates identity distortion. The introduced dynamic identity preservation strategy further ensures close resemblance to the original image identity. Moreover, PortraitBooth incorporates emotion-aware cross-attention control for diverse facial expressions in generated images, supporting text-driven expression editing. Its scalability enables efficient and high-quality image creation, including multi-subject generation. Extensive results demonstrate superior performance over other state-of-the-art methods in both single and multiple image generation scenarios. Our project page is at https://portraitbooth.github.io. Boyuan Jiang, Ying Tai, Donghao Luo 0001, Jiangning Zhang, Wei Lin 0004, Taisong Jin, Chengjie Wang 0001, Rongrong Ji |
CVPR | 4 |
| 2024 | FaceChain-ImagineID: Freely Crafting High-Fidelity Diverse Talking Faces from Disentangled AudioabstractIn this paper, we abstract the process of people hearing speech, extracting meaningful cues, and creating vari-ous dynamically audio-consistent talking faces, termed Lis-tening and Imagining, into the task of high-fidelity diverse talking faces generation from a single audio. Specifically, it involves two critical challenges: one is to effectively de-couple identity, content, and emotion from entangled au-dio, and the other is to maintain intra-video diversity and inter- video consistency. To tackle the issues, we first dig out the intricate relationships among facial factors and sim-plify the decoupling process, tailoring a Progressive Audio Disentanglement for accurate facial geometry and seman-tics learning, where each stage incorporates a customized training module responsible for a specific factor. Secondly, to achieve visually diverse and audio-synchronized animation solely from input audio within a single model, we intro-duce the Controllable Coherent Frame generation, which involves the flexible integration of three trainable adapters with frozen Latent Diffusion Models (LDMs) to focus on maintaining facial geometry and semantics, as well as texsture and temporal coherence between frames. In this way, we inherit high-quality diverse generation from LDMs while significantly improving their controllability at a low training cost. Extensive experiments demonstrate the flexibility and effectiveness of our method in handling this paradigm. The codes will be released at FaceChain. Chao Xu 0023, Yang Liu 0356, Jiazheng Xing, Weida Wang, Jun Dan, Tianxin Huang, Siyuan Li 0002, Zhi-Qi Cheng, Ying Tai, Baigui Sun |
CVPR | 10 |
| 2024 | HybridBooth: Hybrid Prompt Inversion for Efficient Subject-Driven Generation
Shanyan Guan, Yanhao Ge, Ying Tai, Jian Yang 0003, Mingyu You |
ECCV (9) | 3 |
| 2024 | T-Pixel2Mesh: Combining Global and Local Transformer for 3D Mesh Generation from a Single ImageabstractPixel2Mesh (P2M) is a classical approach for reconstructing 3D shapes from a single color image through coarse-to-fine mesh deformation. Although P2M is capable of generating plausible global shapes, its Graph Convolution Network (GCN) often produces overly smooth results, causing the loss of fine-grained geometry details. Moreover, P2M generates non-credible features for occluded regions and struggles with the domain gap from synthetic data to real-world images, which is a common challenge for single-view 3D reconstruction methods. To address these challenges, we propose a novel Transformer-boosted architecture, named T-Pixel2Mesh, inspired by the coarse-to-fine approach of P2M. Specifically, we use a global Transformer to control the holistic shape and a local Transformer to progressively refine the local geometry details with graph-based point upsampling. To enhance real-world reconstruction, we present the simple yet effective Linear Scale Search (LSS), which serves as prompt tuning during the input preprocessing. Our experiments on ShapeNet demonstrate state-of-the-art performance, while results on real-world data show the generalization capability. Boyan Jiang, Keke He, Ying Tai, Chengjie Wang 0001, Yinda Zhang 0001, Yanwei Fu 0001 |
ICASSP | 5 |
| 2024 | Learning to Decouple the Lights for 3D Face Texture ModelingabstractExisting research has made impressive strides in reconstructing human facial shapes and textures from images with well-illuminated faces and minimal external occlusions.
Nevertheless, it remains challenging to recover accurate facial textures from scenarios with complicated illumination affected by external occlusions, \eg a face that is partially obscured by items such as a hat.
Existing works based on the assumption of single and uniform illumination cannot correctly process these data.
In this work, we introduce a novel approach to model 3D facial textures under such unnatural illumination. Instead of assuming single illumination, our framework learns to imitate the unnatural illumination as a composition of multiple separate light conditions combined with learned neural representations, named Light Decoupling.
According to experiments on both single images and video sequences, we demonstrate the effectiveness of our approach in modeling facial textures under challenging illumination affected by occlusions. Tianxin Huang, Zhenyu Zhang 0005, Ying Tai, Gim Hee Lee |
NeurIPS | 3 |
| 2024 | MambaLLIE: Implicit Retinex-Aware Low Light Enhancement with Global-then-Local State SpaceabstractRecent advances in low light image enhancement have been dominated by Retinex-based learning framework, leveraging convolutional neural networks (CNNs) and Transformers. However, the vanilla Retinex theory primarily addresses global illumination degradation and neglects local issues such as noise and blur in dark conditions. Moreover, CNNs and Transformers struggle to capture global degradation due to their limited receptive fields. While state space models (SSMs) have shown promise in the long-sequence modeling, they face challenges in combining local invariants and global context in visual data. In this paper, we introduce MambaLLIE, an implicit Retinex-aware low light enhancer featuring a global-then-local state space design. We first propose a Local-Enhanced State Space Module (LESSM) that incorporates an augmented local bias within a 2D selective scan mechanism, enhancing the original SSMs by preserving local 2D dependency. Additionally, an Implicit Retinex-aware Selective Kernel module (IRSK) dynamically selects features using spatially-varying operations, adapting to varying inputs through an adaptive kernel selection process. Our Global-then-Local State Space Block (GLSSB) integrates LESSM and IRSK with layer normalization (LN) as its core. This design enables MambaLLIE to achieve comprehensive global long-range modeling and flexible local feature aggregation. Extensive experiments demonstrate that MambaLLIE significantly outperforms state-of-the-art CNN and Transformer-based methods. Our code is available at https://github.com/wengjiangwei/MambaLLIE. Jiangwei Weng, Zhiqiang Yan 0001, Ying Tai, Jianjun Qian, Jian Yang 0003, Jun Li 0027 |
NeurIPS | 3 |
| 2024 | Efficiently Exploiting Spatially Variant Knowledge for Video DeblurringabstractVideo deblurring is a challenging task as the blur is often spatially variant. Existing methods mainly engage in building the spatial-temporal correspondence among the frames. As one of the widely-used frameworks, the long-range temporal propagation usually suffers from the expensive computation cost and error accumulation caused by the numerous connections among temporal frames. Meanwhile, the exploration of spatial-variant information from the neighbor frames is often ignored in video deblurring. To tackle these issues, we tailor an efficient short-range multi-scale framework slimming the long-range propagation and exploiting the most relevant neighbor temporal knowledge. For capturing spatial knowledge, we further propose a spatial feature extractor, named the spatially variant adaptive block, to adaptively generate the location-wise kernel to cater to the spatially variant character of blur. For efficient temporal exploitation, a simple inter-frame shift as a motion compensation is developed to avoid expensive long temporal relevance modeling. Both quantitative and qualitative evaluation results on benchmark datasets demonstrate that the proposed algorithm performs favorably against state-of-the-art methods. Qian Xu 0014, Xiaobin Hu, Donghao Luo 0001, Ying Tai, Chengjie Wang 0001, Yuntao Qian |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | High-Resolution Iterative Feedback Network for Camouflaged Object DetectionabstractSpotting camouflaged objects that are visually assimilated into the background is tricky for both object detection algorithms and humans who are usually confused or cheated by the perfectly intrinsic similarities between the foreground objects and the background surroundings. To tackle this challenge, we aim to extract the high-resolution texture details to avoid the detail degradation that causes blurred vision in edges and boundaries. We introduce a novel HitNet to refine the low-resolution representations by high-resolution features in an iterative feedback manner, essentially a global loop-based connection among the multi-scale resolutions. To design better feedback feature flow and avoid the feature corruption caused by recurrent path, an iterative feedback strategy is proposed to impose more constraints on each feedback connection. Extensive experiments on four challenging datasets demonstrate that our HitNet breaks the performance bottleneck and achieves significant improvements compared with 29 state-of-the-art methods. In addition, to address the data scarcity in camouflaged scenarios, we provide an application example to convert the salient objects to camouflaged objects, thereby generating more camouflaged training samples from the diverse salient object datasets. Code will be made publicly available. Xiaobin Hu, Xuebin Qin, Hang Dai, Wenqi Ren, Donghao Luo 0001, Ying Tai, Ling Shao 0001 |
AAAI | 7 |
| 2023 | High-Resolution GAN Inversion for Degraded Images in Large Diverse DatasetsabstractThe last decades are marked by massive and diverse image data, which shows increasingly high resolution and quality. However, some images we obtained may be corrupted, affecting the perception and the application of downstream tasks. A generic method for generating a high-quality image from the degraded one is in demand. In this paper, we present a novel GAN inversion framework that utilizes the powerful generative ability of StyleGAN-XL for this problem. To ease the inversion challenge with StyleGAN-XL, Clustering \& Regularize Inversion (CRI) is proposed. Specifically, the latent space is firstly divided into finer-grained sub-spaces by clustering. Instead of initializing the inversion with the average latent vector, we approximate a centroid latent vector from the clusters, which generates an image close to the input image. Then, an offset with a regularization term is introduced to keep the inverted latent vector within a certain range. We validate our CRI scheme on multiple restoration tasks (i.e., inpainting, colorization, and super-resolution) of complex natural images, and show preferable quantitative and qualitative results. We further demonstrate our technique is robust in terms of data and different GAN models. To our best knowledge, we are the first to adopt StyleGAN-XL for generating high-quality natural images from diverse degraded inputs. Code is available at https://github.com/Booooooooooo/CRI. Yanbo Wang 0003, Chuming Lin, Donghao Luo 0001, Ying Tai, Zhizhong Zhang 0001, Yuan Xie 0006 |
AAAI | 4 |
| 2023 | Learning Neural Proto-Face Field for Disentangled 3D Face Modeling in the WildabstractGenerative models show good potential for recovering 3D faces beyond limited shape assumptions. While plausible details and resolutions are achieved, these models easily fail under extreme conditions of pose, shadow or appearance, due to the entangled fitting or lack of multi-view priors. To address this problem, this paper presents a novel Neural Proto-face Field (NPF) for unsupervised robust 3D face modeling. Instead of using constrained images as Neural Radiance Field (NeRF), NPF disentangles the common/specific facial cues, i.e., ID, expression and scene-specific details from in-the-wild photo collections. Specifically, NPF learns a face prototype to aggregate 3D-consistent identity via uncertainty modeling, extracting multi-image priors from a photo collection. NPF then learns to deform the prototype with the appropriate facial expressions, constrained by a loss of expression consistency and personal idiosyncrasies. Finally, NPF is optimized to fit a target image in the collection, recovering specific details of appearance and geometry. In this way, the generative model benefits from multi-image priors and meaningful facial structures. Extensive experiments on benchmarks show that NPF recovers superior or competitive facial shapes and textures, compared to state-of-the-art methods. Zhenyu Zhang 0005, Renwang Chen, Weijian Cao, Ying Tai, Chengjie Wang 0001 |
CVPR | 4 |
| 2023 | Learning to Measure the Point Cloud Reconstruction Loss in a Representation SpaceabstractFor point cloud reconstruction-related tasks, the reconstruction losses to evaluate the shape differences between reconstructed results and the ground truths are typically used to train the task networks. Most existing works measure the training loss with point-to-point distance, which may introduce extra defects as predefined matching rules may deviate from the real shape differences. Although some learning-based works have been proposed to overcome the weaknesses of manually-defined rules, they still measure the shape differences in 3D Euclidean space, which may limit their ability to capture defects in reconstructed shapes. In this work, we propose a learning-based Contrastive Adver-sarial Loss (CALoss) to measure the point cloud reconstruction loss dynamically in a non-linear representation space by combining the contrastive constraint with the adversarial strategy. Specifically, we use the contrastive constraint to help CALoss learn a representation space with shape similarity, while we introduce the adversarial strategy to help CALoss mine differences between reconstructed results and ground truths. According to experiments on reconstruction-related tasks, CALoss can help task networks improve re-construction performances and learn more representative representations. Tianxin Huang, Zhonggan Ding, Jiangning Zhang, Ying Tai, Zhenyu Zhang 0005, Mingang Chen, Chengjie Wang 0001, Yong Liu 0007 |
CVPR | 4 |
| 2023 | High-Fidelity Generalized Emotional Talking Face Generation with Multi-Modal Emotion Space LearningabstractRecently, emotional talking face generation has received considerable attention. However, existing methods only adopt one-hot coding, image, or audio as emotion conditions, thus lacking flexible control in practical applications and failing to handle unseen emotion styles due to limited semantics. They either ignore the one-shot setting or the quality of generated faces. In this paper, we propose a more flexible and generalized framework. Specifically, we supplement the emotion style in text prompts and use an Aligned Multi-modal Emotion encoder to embed the text, image, and audio emotion modality into a unified space, which inherits rich semantic prior from CLIP. Consequently, effective multi-modal emotion space learning helps our method support arbitrary emotion modality during testing and could generalize to unseen emotion styles. Besides, an Emotion-aware Audio-to-3DMM Convertor is proposed to connect the emotion condition and the audio sequence to structural representation. A followed style-based High-fidelity Emotional Face generator is designed to generate arbitrary high-resolution realistic identities. Our texture generator hierarchically learns flow fields and animated faces in a residual manner. Extensive experiments demonstrate the flexibility and generalization of our method in emotion control and the effectiveness of high-quality face synthesis. Chao Xu 0023, Jiangning Zhang, Wenqing Chu, Ying Tai, Chengjie Wang 0001, Yong Liu 0007 |
CVPR | 6 |
| 2023 | Learning Versatile 3D Shape Generation with Improved Auto-regressive ModelsabstractAuto-Regressive (AR) models have achieved impressive results in 2D image generation by modeling joint distributions in the grid space. While this approach has been extended to the 3D domain for powerful shape generation, it still has two limitations: expensive computations on volumetric grids and ambiguous auto-regressive order along grid dimensions. To overcome these limitations, we propose the Improved Auto-regressive Model (ImAM) for 3D shape generation, which applies discrete representation learning based on a latent vector instead of volumetric grids. Our approach not only reduces computational costs but also preserves essential geometric details by learning the joint distribution in a more tractable order. Moreover, thanks to the simplicity of our model architecture, we can naturally extend it from unconditional to conditional generation by concatenating various conditioning inputs, such as point clouds, categories, images, and texts. Extensive experiments demonstrate that ImAM can synthesize diverse and faithful shapes of multiple categories, achieving state-of-the-art performance. Simian Luo, Xuelin Qian, Yanwei Fu 0001, Yinda Zhang 0001, Ying Tai, Zhenyu Zhang 0005, Chengjie Wang 0001, Xiangyang Xue 0001 |
ICCV | 5 |
| 2023 | Dynamic Frame Interpolation in Wavelet DomainabstractVideo frame interpolation is an important low-level vision task, which can increase frame rate for more fluent visual experience. Existing methods have achieved great success by employing advanced motion models and synthesis networks. However, the spatial redundancy when synthesizing the target frame has not been fully explored, that can result in lots of inefficient computation. On the other hand, the computation compression degree in frame interpolation is highly dependent on both texture distribution and scene motion, which demands to understand the spatial-temporal information of each input frame pair for a better compression degree selection. In this work, we propose a novel two-stage frame interpolation framework termed WaveletVFI to address above problems. It first estimates intermediate optical flow with a lightweight motion perception network, and then a wavelet synthesis network uses flow aligned context features to predict multi-scale wavelet coefficients with sparse convolution for efficient target frame reconstruction, where the sparse valid masks that control computation in each scale are determined by a crucial threshold ratio. Instead of setting a fixed value like previous methods, we find that embedding a classifier in the motion perception network to learn a dynamic threshold for each sample can achieve more computation reduction with almost no loss of accuracy. On the common high resolution and animation frame interpolation benchmarks, proposed WaveletVFI can reduce computation up to 40% while maintaining similar accuracy, making it perform more efficiently against other state-of-the-arts. Lingtong Kong, Boyuan Jiang, Donghao Luo 0001, Wenqing Chu, Ying Tai, Chengjie Wang 0001, Jie Yang 0002 |
IEEE Trans. Image Process. | 5 |
| 2022 | LCTR: On Awakening the Local Continuity of Transformer for Weakly Supervised Object LocalizationabstractWeakly supervised object localization (WSOL) aims to learn object localizer solely by using image-level labels. The convolution neural network (CNN) based techniques often result in highlighting the most discriminative part of objects while ignoring the entire object extent. Recently, the transformer architecture has been deployed to WSOL to capture the long-range feature dependencies with self-attention mechanism and multilayer perceptron structure. Nevertheless, transformers lack the locality inductive bias inherent to CNNs and therefore may deteriorate local feature details in WSOL. In this paper, we propose a novel framework built upon the transformer, termed LCTR (Local Continuity TRansformer), which targets at enhancing the local perception capability of global features among long-range feature dependencies. To this end, we propose a relational patch-attention module (RPAM), which considers cross-patch information on a global basis. We further design a cue digging module (CDM), which utilizes local features to guide the learning trend of the model for highlighting the weak local responses. Finally, comprehensive experiments are carried out on two widely used datasets, ie, CUB-200-2011 and ILSVRC, to verify the effectiveness of our method. Changan Wang, Yabiao Wang, Guannan Jiang, Yunhang Shen, Ying Tai, Chengjie Wang 0001, Wei Zhang 0217, Liujuan Cao |
AAAI | 6 |
| 2022 | DIRL: Domain-Invariant Representation Learning for Generalizable Semantic SegmentationabstractModel generalization to the unseen scenes is crucial to real-world applications, such as autonomous driving, which requires robust vision systems. To enhance the model generalization, domain generalization through learning the domain-invariant representation has been widely studied. However, most existing works learn the shared feature space within multi-source domains but ignore the characteristic of the feature itself (e.g., the feature sensitivity to the domain-specific style). Therefore, we propose the Domain-invariant Representation Learning (DIRL) for domain generalization which utilizes the feature sensitivity as the feature prior to guide the enhancement of the model generalization capability. The guidance reflects in two folds: 1) Feature re-calibration that introduces the Prior Guided Attention Module (PGAM) to emphasize the insensitive features and suppress the sensitive features. 2): Feature whiting that proposes the Guided Feature Whiting (GFW) to remove the feature correlations which are sensitive to the domain-specific style. We construct the domain-invariant representation which suppresses the effect of the domain-specific style on the quality and correlation of the features. As a result, our method is simple yet effective, and can enhance the robustness of various backbone networks with little computational cost. Extensive experiments over multiple domains generalizable segmentation tasks show the superiority of our approach to other methods. Zhengkai Jiang 0001, Guannan Jiang, Wenqing Chu, Wenhui Han, Wei Zhang 0217, Chengjie Wang 0001, Ying Tai |
AAAI | 9 |
| 2022 | SCSNet: An Efficient Paradigm for Learning Simultaneously Image Colorization and Super-resolutionabstractIn the practical application of restoring low-resolution gray-scale images, we generally need to run three separate processes of image colorization, super-resolution, and dows-sampling operation for the target device. However, this pipeline is redundant and inefficient for the independent processes, and some inner features could have been shared. Therefore, we present an efficient paradigm to perform Simultaneously Image Colorization and Super-resolution (SCS) and propose an end-to-end SCSNet to achieve this goal. The proposed method consists of two parts: colorization branch for learning color information that employs the proposed plug-and-play Pyramid Valve Cross Attention (PVCAttn) module to aggregate feature maps between source and reference images; and super-resolution branch for integrating color and texture information to predict target images, which uses the designed Continuous Pixel Mapping (CPM) module to predict high-resolution images at continuous magnification. Furthermore, our SCSNet supports both automatic and referential modes that is more flexible for practical application. Abundant experiments demonstrate the superiority of our method for generating authentic images over state-of-the-art methods, e.g., averagely decreasing FID by 1.8 and 5.1 compared with current best scores for automatic and referential modes, respectively, while owning fewer parameters (more than x2) and faster running speed (more than x3). Jiangning Zhang, Chao Xu 0023, Jian Li 0062, Yabiao Wang, Ying Tai, Yong Liu 0007 |
AAAI | 6 |
| 2022 | Blind Face Restoration via Integrating Face Shape and Generative PriorsabstractBlind face restoration, which aims to reconstruct high-quality images from low-quality inputs, can benefit many applications. Although existing generative-based methods achieve significant progress in producing high-quality images, they often fail to restore natural face shapes and high-fidelity facial details from severely-degraded inputs. In this work, we propose to integrate shape and generative priors to guide the challenging blind face restoration. Firstly, we set up a shape restoration module to recover reason-able facial geometry with 3D reconstruction. Secondly, a pretrained facial generator is adopted as decoder to generate photo-realistic high-resolution images. To ensure high-fidelity, hierarchical spatial features extracted from the low-quality inputs and rendered 3D images are inserted into the decoder with our proposed Adaptive Feature Fusion Block (AFFB). Moreover, we introduce hybrid-level losses to Jointly train the shape and generative priors together with other network parts such that these two priors better adapt to our blind face restoration task. The proposed Shape and Generative Prior integrated Network (SGPN) can re-store high-quality images with clear face shapes and real-istic facial details. Experimental results on synthetic and real-world datasets demonstrate SGPN performs favorably against state-of-the-art blind face restoration methods. Feida Zhu 0002, Wenqing Chu, Xinyi Zhang 0005, Xiaozhong Ji, Chengjie Wang 0001, Ying Tai |
CVPR | 7 |
| 2022 | Physically-guided Disentangled Implicit Rendering for 3D Face ModelingabstractThis paper presents a novel Physically-guided Disentangled Implicit Rendering (PhyDIR) framework for highfidelity 3D face modeling. The motivation comes from two observations: Widely-used graphics renderers yield excessive approximations against photo-realistic imaging, while neural rendering methods produce superior appearances but are highly entangled to perceive 3D-aware operations. Hence, we learn to disentangle the implicit rendering via explicit physical guidance, while guaranteeing the properties of: (1) 3D-aware comprehension and (2) high-reality image formation. For the former one, PhyDIR explicitly adopts 3D shading and rasterizing modules to control the renderer, which disentangles the light, facial shape, and viewpoint from neural reasoning. Specifically, PhyDIR proposes a novel multi-image shading strategy to compensate for the monocular limitation, so that the lighting variations are accessible to the neural renderer. For the latter, PhyDIR learns the face-collection implicit texture to avoid ill-posed intrinsic factorization, then leverages a series of consistency losses to constrain the rendering robustness. With the disentangled method, we make 3D face modeling benefit from both kinds of rendering strategies. Extensive experiments on benchmarks show that PhyDIR obtains superior performance than state-of-the-art explicit/implicit methods on geometry/texture modeling. Zhenyu Zhang 0005, Yanhao Ge, Ying Tai, Weijian Cao, Renwang Chen, Kunlin Liu, Hao Tang 0005, Chengjie Wang 0001, Dongjin Huang |
CVPR | 3 |
| 2022 | Learning to Restore 3D Face from In-the-Wild Degraded ImagesabstractIn-the-wild 3D face modelling is a challenging problem as the predicted facial geometry and texture suffer from a lack of reliable clues or priors, when the input images are degraded. To address such a problem, in this paper we propose a novel Learning to Restore (L2R) 3D face framework for unsupervised high-quality face reconstruction from low-resolution images. Rather than directly refining 2D image appearance, L2R learns to recover fine-grained 3D details on the proxy against degradation via extracting generative facial priors. Concretely, L2R proposes a novel albedo restoration network to model high-quality 3D facial texture, in which the diverse guidance from the pre-trained Generative Adversarial Networks (GANs) is leveraged to complement the lack of input facial clues. With the finer details of the restored 3D texture, L2R then learns displacement maps from scratch to enhance the significant facial structure and geometry. Both of the procedures are mutually optimized with a novel 3D-aware adversarial loss, which further improves the modelling performance and suppresses the potential uncertainty. Extensive experiments on benchmarks show that L2R outperforms state-of-the-art methods under the condition of low-quality inputs, and obtains superior performances than 2D pre-processed modelling approaches with limited 3D proxy. Zhenyu Zhang 0005, Yanhao Ge, Ying Tai, Chengjie Wang 0001, Hao Tang 0005, Dongjin Huang |
CVPR | 3 |
| 2022 | IFRNet: Intermediate Feature Refine Network for Efficient Frame InterpolationabstractPrevailing video frame interpolation algorithms, that generate the intermediate frames from consecutive inputs, typically rely on complex model architectures with heavy parameters or large delay, hindering them from diverse real-time applications. In this work, we devise an efficient encoder-decoder based network, termed IFRNet, for fast in-termediate frame synthesizing. It first extracts pyramid features from given inputs, and then refines the bilateral in-termediate flow fields together with a powerful intermedi-ate feature until generating the desired output. The gradu-ally refined intermediate feature can not only facilitate in-termediate flow estimation, but also compensate for con-textual details, making IFRNet do not need additional syn-thesis or refinement module. To fully release its potential, we further propose a novel task-oriented optical flow dis-tillation loss to focus on learning the useful teacher knowl-edge towards frame synthesizing. Meanwhile, a new ge-ometry consistency regularization term is imposed on the gradually refined intermediate features to keep better structure layout. Experiments on various benchmarks demon-strate the excellent performance and fast inference speed of proposed approaches. Code is available at https://github.com/ltkong218/IFRNet. Lingtong Kong, Boyuan Jiang, Donghao Luo 0001, Wenqing Chu, Ying Tai, Chengjie Wang 0001, Jie Yang 0002 |
CVPR | 6 |
| 2022 | Learning to Memorize Feature Hallucination for One-Shot Image GenerationabstractThis paper studies the task of One-Shot image Generation (OSG), where generation network learned on base dataset should be generalizable to synthesize images of novel categories with only one available sample per novel category. Most existing methods for feature transfer in oneshot image generation only learn reusable features implicitly on pre-training tasks. Such methods would be likely to overfit pre-training tasks. In this paper, we propose a novel model to explicitly learn and memorize reusable features that can help hallucinate novel category images. To be specific, our algorithm learns to decompose image features into the Category-Related (CR) and Category-Independent(CI) features. Our model learning to memorize class-independent CI features which are further utilized by our feature hallucination component to generate target novel category images. We validate our model on several benchmarks. Extensive experiments demonstrate that our model effectively boosts the OSG performance and can generate compelling and diverse samples. Yanwei Fu 0001, Ying Tai, Yun Cao 0002, Chengjie Wang 0001 |
CVPR | 3 |
| 2022 | ColorFormer: Image Colorization via Color Memory Assisted Hybrid-Attention Transformer
Xiaozhong Ji, Boyuan Jiang, Donghao Luo 0001, Guangpin Tao, Wenqing Chu, Chengjie Wang 0001, Ying Tai |
ECCV (16) | 8 |
| 2022 | Prototypical Contrast Adaptation for Domain Adaptive Semantic Segmentation
Zhengkai Jiang 0001, Yuxi Li 0009, Ceyuan Yang, Peng Gao 0007, Yabiao Wang, Ying Tai, Chengjie Wang 0001 |
ECCV (34) | 6 |
| 2022 | StyleFace: Towards Identity-Disentangled Face Generation on Megapixels
Keke He, Wenqing Chu, Ying Tai, Chengjie Wang 0001, Junchi Yan |
ECCV (16) | 5 |
| 2022 | Designing One Unified Framework for High-Fidelity Face Reenactment and Swapping
Chao Xu 0023, Jiangning Zhang, Guanzhong Tian, Xianfang Zeng, Ying Tai, Yabiao Wang, Chengjie Wang 0001, Yong Liu 0007 |
ECCV (15) | 6 |
| 2022 | SeedFormer: Patch Seeds Based Point Cloud Completion with Upsample Transformer
Yun Cao 0002, Wenqing Chu, Tong Lu 0002, Ying Tai, Chengjie Wang 0001 |
ECCV (3) | 6 |
| 2022 | HifiHead: One-Shot High Fidelity Neural Head Synthesis with 3D ControlabstractWe propose HifiHead, a high fidelity neural talking head synthesis method, which can well preserve the source image's appearance and control the motion (e.g., pose, expression, gaze) flexibly with 3D morphable face models (3DMMs) parameters derived from a driving image or indicated by users. Existing head synthesis works mainly focus on low-resolution inputs. Instead, we exploit the powerful generative prior embedded in StyleGAN to achieve high-quality head synthesis and editing. Specifically, we first extract the source image's appearance and driving image's motion to construct 3D face descriptors, which are employed as latent style codes for the generator. Meanwhile, hierarchical representations are extracted from the source and rendered 3D images respectively to provide faithful appearance and shape guidance. Considering the appearance representations need high-resolution flow fields for spatial transform, we propose a coarse-to-fine style-based generator consisting of a series of feature alignment and refinement (FAR) blocks. Each FAR block updates the dense flow fields and refines RGB outputs simultaneously for efficiency. Extensive experiments show that our method blends source appearance and target motion more accurately along with more photo-realistic results than previous state-of-the-art approaches. Feida Zhu 0002, Wenqing Chu, Ying Tai, Chengjie Wang 0001 |
IJCAI | 4 |
| 2022 | AutoGAN-Synthesizer: Neural Architecture Search for Cross-Modality MRI Synthesis
Xiaobin Hu, Ruolin Shen, Donghao Luo 0001, Ying Tai, Chengjie Wang 0001, Bjoern Menze |
MICCAI (6) | 4 |
| 2022 | Joint Learning Content and Degradation Aware Feature for Blind Super-ResolutionabstractTo achieve promising results on blind image super-resolution (SR),some attempts leveraged the low resolution (LR) images to predict the kernel and improve the SR performance. However, these Supervised Kernel Prediction (SKP) methods are impractical due to the unavailable real-world blur kernels. Although some Unsupervised Degradation Prediction (UDP) methods are proposed to bypass this problem, the inconsistency between degradation embedding and SR feature is still challenging. By exploring the correlations between degradation embedding and SR feature, we observe that jointly learning the content and degradation aware feature is optimal. Based on this observation, a Content and Degradation aware SR Network dubbed CDSR is proposed. Specifically, CDSR contains three newly-established modules: (1) a Lightweight Patch-based Encoder (LPE) is applied to jointly extract content and degradation features; (2) a Domain Query Attention based module (DQA) is employed to adaptively reduce the inconsistency; (3) a Codebook-based Space Compress module (CSC) that can suppress the redundant information. Extensive experiments on several benchmarks demonstrate that the proposed CDSR outperforms the existing UDP models and achieves competitive performance on PSNR and SSIM even compared with the state-of-the-art SKP methods. Chuming Lin, Donghao Luo 0001, Yong Liu 0032, Ying Tai, Chengjie Wang 0001, Mingang Chen |
ACM Multimedia | 5 |
| 2022 | 3QNet: 3D Point Cloud Geometry Quantization Compression NetworkabstractSince the development of 3D applications, the point cloud, as a spatial description easily acquired by sensors, has been widely used in multiple areas such as SLAM and 3D reconstruction. Point Cloud Compression (PCC) has also attracted more attention as a primary step before point cloud transferring and saving, where the geometry compression is an important component of PCC to compress the points geometrical structures. However, existing non-learning-based geometry compression methods are often limited by manually pre-defined compression rules. Though learning-based compression methods can significantly improve the algorithm performances by learning compression rules from data, they still have some defects. Voxel-based compression networks introduce precision errors due to the voxelized operations, while point-based methods may have relatively weak robustness and are mainly designed for sparse point clouds. In this work, we propose a novel learning-based point cloud compression framework named 3D Point Cloud Geometry Quantiation Compression Network (3QNet), which overcomes the robustness limitation of existing point-based methods and can handle dense points. By learning a codebook including common structural features from simple and sparse shapes, 3QNet can efficiently deal with multiple kinds of point clouds. According to experiments on object models, indoor scenes, and outdoor scans, 3QNet can achieve better compression performances than many representative methods. Tianxin Huang, Jiangning Zhang, Jun Chen 0023, Zhonggan Ding, Ying Tai, Zhenyu Zhang 0005, Chengjie Wang 0001, Yong Liu 0007 |
ACM Trans. Graph. | 5 |
| 2021 | Generalizable Representation Learning for Mixture Domain Face Anti-SpoofingabstractFace anti-spoofing approach based on domain generalization (DG) has drawn growing attention due to its robustness for unseen scenarios. Existing DG methods assume that the domain label is known. However, in real-world applications, the collected dataset always contains mixture domains, where the domain label is unknown. In this case, most of existing methods may not work. Further, even if we can obtain the domain label as existing methods, we think this is just a sub-optimal partition. To overcome the limitation, we propose domain dynamic adjustment meta-learning (D$^2$AM) without using domain labels, which iteratively divides mixture domains via discriminative domain representation and trains a generalizable face anti-spoofing with meta-learning. Specifically, we design a domain feature based on Instance Normalization (IN) and propose a domain representation learning module (DRLM) to extract discriminative domain features for clustering. Moreover, to reduce the side effect of outliers on clustering performance, we additionally utilize maximum mean discrepancy (MMD) to align the distribution of sample features to a prior distribution, which improves the reliability of clustering. Extensive experiments show that the proposed method outperforms conventional DG-based face anti-spoofing methods, including those utilizing domain labels. Furthermore, we enhance the interpretability through visualization. Taiping Yao, Kekai Sheng, Shouhong Ding, Ying Tai, Feiyue Huang |
AAAI | 5 |
| 2021 | Frequency Consistent Adaptation for Real World Super ResolutionabstractRecent deep-learning based Super-Resolution (SR) methods have achieved remarkable performance on images with known degradation. However, these methods always fail in real-world scene, since the Low-Resolution (LR) images after the ideal degradation (e.g., bicubic down-sampling) deviate from real source domain. The domain gap between the LR images and the real-world images can be observed clearly on frequency density, which inspires us to explicitly narrow the undesired gap caused by incorrect degradation. From this point of view, we design a novel Frequency Consistent Adaptation (FCA) that ensures the frequency domain consistency when applying existing SR methods to the real scene. We estimate degradation kernels from unsupervised images and generate the corresponding LR images. To provide useful gradient information for kernel estimation, we propose Frequency Density Comparator (FDC) by distinguishing the frequency density of images on different scales. Based on the domain-consistent LR-HR pairs, we train easy-implemented Convolutional Neural Network (CNN) SR models. Extensive experiments show that the proposed FCA improves the performance of the SR model under real-world setting achieving state-of-the-art results with high fidelity and plausible perception, thus providing a novel effective framework for real-world SR application. Xiaozhong Ji, Guangpin Tao, Yun Cao 0002, Ying Tai, Tong Lu 0002, Chengjie Wang 0001, Feiyue Huang |
AAAI | 4 |
| 2021 | To Choose or to Fuse? Scale Selection for Crowd CountingabstractIn this paper, we address the large scale variation problem in crowd counting by taking full advantage of the multi-scale feature representations in a multi-level network. We implement such an idea by keeping the counting error of a patch as small as possible with a proper feature level selection strategy, since a specific feature level tends to perform better for a certain range of scales. However, without scale annotations, it is sub-optimal and error-prone to manually assign the predictions for heads of different scales to specific feature levels. Therefore, we propose a Scale-Adaptive Selection Network (SASNet), which automatically learns the internal correspondence between the scales and the feature levels. Instead of directly using the predictions from the most appropriate feature level as the final estimation, our SASNet also considers the predictions from other feature levels via weighted average, which helps to mitigate the gap between discrete feature levels and continuous scale variation. Since the heads in a local patch share roughly a same scale, we conduct the adaptive selection strategy in a patch-wise style. However, pixels within a patch contribute different counting errors due to the various difficulty degrees of learning. Thus, we further propose a Pyramid Region Awareness Loss (PRA Loss) to recursively select the most hard sub-regions within a patch until reaching the pixel level. With awareness of whether the parent patch is over-estimated or under-estimated, the fine-grained optimization with the PRA Loss for these region-aware hard pixels helps to alleviate the inconsistency problem between training target and evaluation metric. The state-of-the-art results on four datasets demonstrate the superiority of our approach. The code will be available at: https://github.com/TencentYoutuResearch/CrowdCounting-SASNet. Qingyu Song 0001, Changan Wang, Yabiao Wang, Ying Tai, Chengjie Wang 0001, Jian Wu 0001, Jiayi Ma 0001 |
AAAI | 4 |
| 2021 | Learning Comprehensive Motion Representation for Action RecognitionabstractFor action recognition learning, 2D CNN-based methods are efficient but may yield redundant features due to applying the same 2D convolution kernel to each frame. Recent efforts attempt to capture motion information by establishing inter-frame connections while still suffering the limited temporal receptive field or high latency. Moreover, the feature enhancement is often only performed by channel or space dimension in action recognition. To address these issues, we first devise a Channel-wise Motion Enhancement (CME) module to adaptively emphasize the channels related to dynamic information with a channel-wise gate vector. The channel gates generated by CME incorporate the information from all the other frames in the video. We further propose a Spatial-wise Motion Enhancement (SME) module to focus on the regions with the critical target in motion, according to the point-to-point similarity between adjacent feature maps. The intuition is that the change of background is typically slower than the motion area. Both CME and SME have clear physical meaning in capturing action clues. By integrating the two modules into the off-the-shelf 2D network, we finally obtain a Comprehensive Motion Representation (CMR) learning method for action recognition, which achieves competitive performance on Something-Something V1 & V2 and Kinetics-400. On the temporal reasoning datasets Something-Something V1 and V2, our method outperforms the current state-of-the-art by 2.3% and 1.9% when using 16 frames as input, respectively. Mingyu Wu 0002, Boyuan Jiang, Donghao Luo 0001, Junchi Yan, Yabiao Wang, Ying Tai, Chengjie Wang 0001, Feiyue Huang, Xiaokang Yang 0001 |
AAAI | 6 |
| 2021 | Learning To Aggregate and Personalize 3D Face From In-the-Wild Photo CollectionabstractNon-parametric face modeling aims to reconstruct 3D face only from images without shape assumptions. While plausible facial details are predicted, the models tend to over-depend on local color appearance and suffer from ambiguous noise. To address such problem, this paper presents a novel Learning to Aggregate and Personalize (LAP) framework for unsupervised robust 3D face modeling. Instead of using controlled environment, the proposed method implicitly disentangles ID-consistent and scene-specific face from unconstrained photo set. Specifically, to learn ID-consistent face, LAP adaptively aggregates intrinsic face factors of an identity based on a novel curriculum learning approach with relaxed consistency loss. To adapt the face for a personalized scene, we propose a novel attribute-refining network to modify ID-consistent face with target attribute and details. Based on the proposed method, we make unsupervised 3D face modeling benefit from meaningful image facial structure and possibly higher resolutions. Extensive experiments on benchmarks show LAP recovers superior or competitive face shape and texture, compared with state-of-the-art (SOTA) methods with or without prior and supervision. Zhenyu Zhang 0005, Yanhao Ge, Renwang Chen, Ying Tai, Yan Yan 0002, Jian Yang 0003, Chengjie Wang 0001, Feiyue Huang |
CVPR | 4 |
| 2021 | Learning Salient Boundary Feature for Anchor-free Temporal Action LocalizationabstractTemporal action localization is an important yet challenging task in video understanding. Typically, such a task aims at inferring both the action category and localization of the start and end frame for each action instance in a long, untrimmed video. While most current models achieve good results by using pre-defined anchors and numerous actionness, such methods could be bothered with both large number of outputs and heavy tuning of locations and sizes corresponding to different anchors. Instead, anchor-free methods is lighter, getting rid of redundant hyper-parameters, but gains few attention. In this paper, we propose the first purely anchor-free temporal localization method, which is both efficient and effective. Our model includes (i) an end-to-end trainable basic predictor, (ii) a saliency-based refinement module to gather more valuable boundary features for each proposal with a novel boundary pooling, and (iii) several consistency constraints to make sure our model can find the accurate boundary given arbitrary proposals. Extensive experiments show that our method beats all anchor-based and actionness-guided methods with a remarkable margin on THUMOS14, achieving state-of-the-art results, and comparable ones on ActivityNet v1.3. Code is available at https://github.com/TencentYoutuResearch/ActionDetection-AFSD. Chuming Lin, Chengming Xu 0001, Donghao Luo 0001, Yabiao Wang, Ying Tai, Chengjie Wang 0001, Feiyue Huang, Yanwei Fu 0001 |
CVPR | 5 |
| 2021 | Learning To Restore Hazy Video: A New Real-World Dataset and a New MethodabstractMost of the existing deep learning-based dehazing methods are trained and evaluated on the image dehazing datasets, where the dehazed images are generated by only exploiting the information from the corresponding hazy ones. On the other hand, video dehazing algorithms, which can acquire more satisfying dehazing results by exploiting the temporal redundancy from neighborhood hazy frames, receive less attention due to the absence of the video dehazing datasets. Therefore, we propose the first REal-world VIdeo DEhazing (REVIDE) dataset which can be used for the supervised learning of the video dehazing algorithms. By utilizing a well-designed video acquisition system, we can capture paired real-world hazy and haze-free videos that are perfectly aligned by recording the same scene (with or without haze) twice. Considering the challenge of exploiting temporal redundancy among the hazy frames, we also develop a Confidence Guided and Improved Deformable Network (CG-IDN) for video dehazing. The experiments demonstrate that the hazy scenes in the REVIDE dataset are more realistic than the synthetic datasets and the proposed algorithm also performs favorably against state-of-the-art dehazing methods. Xinyi Zhang 0005, Hang Dong 0001, Jinshan Pan, Chao Zhu 0007, Ying Tai, Chengjie Wang 0001, Feiyue Huang, Fei Wang 0008 |
CVPR | 5 |
| 2021 | Rethinking Counting and Localization in Crowds: A Purely Point-Based FrameworkabstractLocalizing individuals in crowds is more in accordance with the practical demands of subsequent high-level crowd analysis tasks than simply counting. However, existing localization based methods relying on intermediate representations (i.e., density maps or pseudo boxes) serving as learning targets are counter-intuitive and error-prone. In this paper, we propose a purely point-based framework for joint crowd counting and individual localization. For this framework, instead of merely reporting the absolute counting error at image level, we propose a new metric, called density Normalized Average Precision (nAP), to provide more comprehensive and more precise performance evaluation. Moreover, we design an intuitive solution under this framework, which is called Point to Point Network (P2PNet). P2PNet discards superfluous steps and directly predicts a set of point proposals to represent heads in an image, being consistent with the human annotation results. By thorough analysis, we reveal the key step towards implementing such a novel idea is to assign optimal learning targets for these proposals. Therefore, we propose to conduct this crucial association in an one-to-one matching manner using the Hungarian algorithm. The P2PNet not only significantly surpasses state-of-the-art methods on popular counting benchmarks, but also achieves promising localization accuracy. The codes will be available at: TencentYoutuResearch/CrowdCounting-P2PNet. Qingyu Song 0001, Changan Wang, Zhengkai Jiang 0001, Yabiao Wang, Ying Tai, Chengjie Wang 0001, Feiyue Huang, Yang Wu 0001 |
ICCV | 5 |
| 2021 | Uniformity in Heterogeneity: Diving Deep into Count Interval Partition for Crowd CountingabstractRecently, the problem of inaccurate learning targets in crowd counting draws increasing attention. Inspired by a few pioneering work, we solve this problem by trying to predict the indices of pre-defined interval bins of counts instead of the count values themselves. However, an inappropriate interval setting might make the count error contributions from different intervals extremely imbalanced, leading to inferior counting performance. Therefore, we propose a novel count interval partition criterion called Uniform Error Partition (UEP), which always keeps the expected counting error contributions equal for all intervals to minimize the prediction risk. Then to mitigate the inevitably introduced discretization errors in the count quantization process, we propose another criterion called Mean Count Proxies (MCP). The MCP criterion selects the best count proxy for each interval to represent its count value during inference, making the overall expected discretization error of an image nearly negligible. As far as we are aware, this work is the first to delve into such a classification task and ends up with a promising solution for count interval partition. Following the above two theoretically demonstrated criterions, we propose a simple yet effective model termed Uniform Error Partition Network (UEPNet), which achieves state-of-the-art performance on several challenging datasets. The codes will be available at: TencentYoutuResearch/CrowdCounting-UEPNet. Changan Wang, Qingyu Song 0001, Boshen Zhang, Yabiao Wang, Ying Tai, Xuyi Hu, Chengjie Wang 0001, Jiayi Ma 0001, Yang Wu 0001 |
ICCV | 5 |
| 2021 | HifiFace: 3D Shape and Semantic Prior Guided High Fidelity Face SwappingabstractIn this work, we propose a high fidelity face swapping method, called HifiFace, which can well preserve the face shape of the source face and generate photo-realistic results. Unlike other existing face swapping works that only use face recognition model to keep the identity similarity, we propose 3D shape-aware identity to control the face shape with the geometric supervision from 3DMM and 3D face reconstruction method. Meanwhile, we introduce the Semantic Facial Fusion module to optimize the combination of encoder and decoder features and make adaptive blending, which makes the results more photo-realistic. Extensive experiments on faces in the wild demonstrate that our method can preserve better identity, especially on the face shape, and can generate more photo-realistic results than previous state-of-the-art methods. Code is available at: https://johann.wang/HifiFace Yuhan Wang 0002, Xu Chen 0024, Wenqing Chu, Ying Tai, Chengjie Wang 0001, Yongjian Wu 0001, Feiyue Huang, Rongrong Ji |
IJCAI | 5 |
| 2021 | Dual Reweighting Domain Generalization for Face Presentation Attack DetectionabstractFace anti-spoofing approaches based on domain generalization (DG) have drawn growing attention due to their robustness for unseen scenarios. Previous methods treat each sample from multiple domains indiscriminately during the training process, and endeavor to extract a common feature space to improve the generalization. However, due to complex and biased data distribution, directly treating them equally will corrupt the generalization ability. To settle the issue, we propose a novel Dual Reweighting Domain Generalization (DRDG) framework which iteratively reweights the relative importance between samples to further improve the generalization. Concretely, Sample Reweighting Module is first proposed to identify samples with relatively large domain bias, and reduce their impact on the overall optimization. Afterwards, Feature Reweighting Module is introduced to focus on these samples and extract more domain-irrelevant features via a self-distilling mechanism. Combined with the domain discriminator, the iteration of the two modules promotes the extraction of generalized features. Extensive experiments and visualizations are presented to demonstrate the effectiveness and interpretability of our method against the state-of-the-art competitors. Shubao Liu, Ke-Yue Zhang, Taiping Yao, Kekai Sheng, Shouhong Ding, Ying Tai, Yuan Xie 0006, Lizhuang Ma |
IJCAI | 6 |
| 2021 | SiamRCR: Reciprocal Classification and Regression for Visual Object TrackingabstractRecently, most siamese network based trackers locate targets via object classification and bounding-box regression. Generally, they select the bounding-box with maximum classification confidence as the final prediction. This strategy may miss the right result due to the accuracy misalignment between classification and regression. In this paper, we propose a novel siamese tracking algorithm called SiamRCR, addressing this problem with a simple, light and effective solution. It builds reciprocal links between classification and regression branches, which can dynamically re-weight their losses for each positive sample. In addition, we add a localization branch to predict the localization accuracy, so that it can work as the replacement of the regression assistance link during inference. This branch makes the training and inference more consistent. Extensive experimental results demonstrate the effectiveness of SiamRCR and its superiority over the state-of-the-art competitors on GOT-10k, LaSOT, TrackingNet, OTB-2015, VOT-2018 and VOT-2019. Moreover, our SiamRCR runs at 65 FPS, far above the real-time requirement. Jinlong Peng, Zhengkai Jiang 0001, Yueyang Gu, Yang Wu 0001, Yabiao Wang, Ying Tai, Chengjie Wang 0001, Weiyao Lin |
IJCAI | 6 |
| 2021 | Context-Aware Image Inpainting with Learned Semantic PriorsabstractRecent advances in image inpainting have shown impressive results for generating plausible visual details on rather simple backgrounds. However, for complex scenes, it is still challenging to restore reasonable contents as the contextual information within the missing regions tends to be ambiguous. To tackle this problem, we introduce pretext tasks that are semantically meaningful to estimating the missing contents. In particular, we perform knowledge distillation on pretext models and adapt the features to image inpainting. The learned semantic priors ought to be partially invariant between the high-level pretext task and low-level image inpainting, which not only help to understand the global context but also provide structural guidance for the restoration of local textures. Based on the semantic priors, we further propose a context-aware image inpainting model, which adaptively integrates global semantics and local features in a unified image generator. The semantic learner and the image generator are trained in an end-to-end manner. We name the model SPL to highlight its ability to learn and leverage semantic priors. It achieves the state of the art on Places2, CelebA, and Paris StreetView datasets Wendong Zhang 0002, Ying Tai, Yunbo Wang, Wenqing Chu, Bingbing Ni, Chengjie Wang 0001, Xiaokang Yang 0001 |
IJCAI | 3 |
| 2021 | ASFD: Automatic and Scalable Face DetectorabstractAlong with current multi-scale based detectors, Feature Aggregation and Enhancement (FAE) modules have shown superior performance gains for cutting-edge object detection. However, these hand-crafted FAE modules show inconsistent improvements on face detection, which is mainly due to the significant distribution difference between its training and applying corpus, i.e. COCO vs. WIDER Face. To tackle this problem, we essentially analyse the effect of data distribution, and consequently propose to search an effective FAE architecture, termed AutoFAE by a differentiable architecture search, which outperforms all existing FAE modules in face detection with a considerable margin. Upon the found AutoFAE and existing backbones, a supernet is further built and trained, which automatically obtains a family of detectors under the different complexity constraints. Extensive experiments conducted on popular benchmarks, i.e. WIDER Face and FDDB, demonstrate the state-of-the-art performance-efficiency trade-off for the proposed automatic and scalable face detector (ASFD) family. In particular, our strong ASFD-D6 outperforms the best competitor with AP 96.7/96.2/92.1 on WIDER Face test, and the lightweight ASFD-D0 costs about 3.1 ms, i.e. more than 320 FPS, on the V100 GPU with VGA-resolution images. Jian Li 0062, Yabiao Wang, Ying Tai, Zhenyu Zhang 0005, Chengjie Wang 0001, Yili Xia |
ACM Multimedia | 4 |
| 2021 | Spectrum-to-Kernel Translation for Accurate Blind Image Super-ResolutionabstractDeep-learning based Super-Resolution (SR) methods have exhibited promising performance under non-blind setting where blur kernel is known; however, blur kernels of Low-Resolution (LR) images in different practical applications are usually unknown. It may lead to a significant performance drop when degradation process of training images deviates from that of real images. In this paper, we propose a novel blind SR framework to super-resolve LR images degraded by arbitrary blur kernel with accurate kernel estimation in frequency domain. To our best knowledge, this is the first deep learning method which conducts blur kernel estimation in frequency domain. Specifically, we first demonstrate that feature representation in frequency domain is more conducive for blur kernel reconstruction than in spatial domain. Next, we present a Spectrum-to-Kernel (S$2$K) network to estimate general blur kernels in diverse forms. We use a conditional GAN (CGAN) combined with SR-oriented optimization target to learn the end-to-end translation from degraded images' spectra to unknown kernels. Extensive experiments on both synthetic and real-world images demonstrate that our proposed method sufficiently reduces blur kernel estimation error, thus enables the off-the-shelf non-blind SR methods to work under blind setting effectively, and achieves superior performance over state-of-the-art blind SR methods, averagely by 1.39dB, 0.48dB (Gaussian kernels) and 6.15dB, 4.57dB (motion kernels) for scales $2\times$ and $4\times$ respectively. Guangpin Tao, Xiaozhong Ji, Wenzhuo Wang, Shuo Chen 0003, Chuming Lin, Yun Cao 0002, Tong Lu 0002, Donghao Luo 0001, Ying Tai |
NeurIPS | 9 |
| 2021 | Analogous to Evolutionary Algorithm: Designing a Unified Sequence ModelabstractInspired by biological evolution, we explain the rationality of Vision Transformer by analogy with the proven practical Evolutionary Algorithm (EA) and derive that both of them have consistent mathematical representation. Analogous to the dynamic local population in EA, we improve the existing transformer structure and propose a more efficient EAT model, and design task-related heads to deal with different tasks more flexibly. Moreover, we introduce the spatial-filling curve into the current vision transformer to sequence image data into a uniform sequential format. Thus we can design a unified EAT framework to address multi-modal tasks, separating the network architecture from the data format adaptation. Our approach achieves state-of-the-art results on the ImageNet classification task compared with recent vision transformer works while having smaller parameters and greater throughput. We further conduct multi-modal tasks to demonstrate the superiority of the unified EAT, \eg, Text-Based Image Retrieval, and our approach improves the rank-1 by +3.7 points over the baseline on the CSS dataset. Jiangning Zhang, Chao Xu 0023, Jian Li 0062, Wenzhou Chen, Yabiao Wang, Ying Tai, Shuo Chen 0003, Chengjie Wang 0001, Feiyue Huang, Yong Liu 0007 |
NeurIPS | 6 |
| 2020 | Fast Learning of Temporal Action Proposal via Dense Boundary GeneratorabstractGenerating temporal action proposals remains a very challenging problem, where the main issue lies in predicting precise temporal proposal boundaries and reliable action confidence in long and untrimmed real-world videos. In this paper, we propose an efficient and unified framework to generate temporal action proposals named Dense Boundary Generator (DBG), which draws inspiration from boundary-sensitive methods and implements boundary classification and action completeness regression for densely distributed proposals. In particular, the DBG consists of two modules: Temporal boundary classification (TBC) and Action-aware completeness regression (ACR). The TBC aims to provide two temporal boundary confidence maps by low-level two-stream features, while the ACR is designed to generate an action completeness score map by high-level action-aware features. Moreover, we introduce a dual stream BaseNet (DSB) to encode RGB and optical flow information, which helps to capture discriminative boundary and actionness features. Extensive experiments on popular benchmarks ActivityNet-1.3 and THUMOS14 demonstrate the superiority of DBG over the state-of-the-art proposal generator (e.g., MGG and BMN). Chuming Lin, Jian Li 0062, Yabiao Wang, Ying Tai, Donghao Luo 0001, Zhipeng Cui, Chengjie Wang 0001, Feiyue Huang, Rongrong Ji |
AAAI | 4 |
| 2020 | TEINet: Towards an Efficient Architecture for Video RecognitionabstractEfficiency is an important issue in designing video architectures for action recognition. 3D CNNs have witnessed remarkable progress in action recognition from videos. However, compared with their 2D counterparts, 3D convolutions often introduce a large amount of parameters and cause high computational cost. To relieve this problem, we propose an efficient temporal module, termed as Temporal Enhancement-and-Interaction (TEI Module), which could be plugged into the existing 2D CNNs (denoted by TEINet). The TEI module presents a different paradigm to learn temporal features by decoupling the modeling of channel correlation and temporal interaction. First, it contains a Motion Enhanced Module (MEM) which is to enhance the motion-related features while suppress irrelevant information (e.g., background). Then, it introduces a Temporal Interaction Module (TIM) which supplements the temporal contextual information in a channel-wise manner. This two-stage modeling scheme is not only able to capture temporal structure flexibly and effectively, but also efficient for model inference. We conduct extensive experiments to verify the effectiveness of TEINet on several benchmarks (e.g., Something-Something V1&V2, Kinetics, UCF101 and HMDB51). Our proposed TEINet can achieve a good recognition accuracy on these datasets but still preserve a high efficiency. Zhaoyang Liu 0001, Donghao Luo 0001, Yabiao Wang, Limin Wang 0002, Ying Tai, Chengjie Wang 0001, Feiyue Huang, Tong Lu 0002 |
AAAI | 5 |
| 2020 | FAN: Feature Adaptation Network for Surveillance Face Recognition and Normalization
Xi Yin 0001, Ying Tai, Yuge Huang, Xiaoming Liu 0002 |
ACCV (2) | 2 |
| 2020 | CurricularFace: Adaptive Curriculum Learning Loss for Deep Face RecognitionabstractAs an emerging topic in face recognition, designing margin-based loss functions can increase the feature margin between different classes for enhanced discriminability. More recently, the idea of mining-based strategies is adopted to emphasize the misclassified samples, achieving promising results. However, during the entire training process, the prior methods either do not explicitly emphasize the sample based on its importance that renders the hard samples not fully exploited; or explicitly emphasize the effects of semi-hard/hard samples even at the early training stage that may lead to convergence issue. In this work, we propose a novel Adaptive Curriculum Learning loss (CurricularFace) that embeds the idea of curriculum learning into the loss function to achieve a novel training strategy for deep face recognition, which mainly addresses easy samples in the early training stage and hard ones in the later stage. Specifically, our CurricularFace adaptively adjusts the relative importance of easy and hard samples during different training stages. In each stage, different samples are assigned with different importance according to their corresponding difficultness. Extensive experimental results on popular benchmarks demonstrate the superiority of our CurricularFace over the state-of-the-art competitors. Yuge Huang, Yuhan Wang 0002, Ying Tai, Xiaoming Liu 0002, Pengcheng Shen, Shaoxin Li 0001, Feiyue Huang |
CVPR | 3 |
| 2020 | Learning by Analogy: Reliable Supervision From Transformations for Unsupervised Optical Flow EstimationabstractUnsupervised learning of optical flow, which leverages the supervision from view synthesis, has emerged as a promising alternative to supervised methods. However, the objective of unsupervised learning is likely to be unreliable in challenging scenes. In this work, we present a framework to use more reliable supervision from transformations. It simply twists the general unsupervised learning pipeline by running another forward pass with transformed data from augmentation, along with using transformed predictions of original data as the self-supervision signal. Besides, we further introduce a lightweight network with multiple frames by a highly-shared flow decoder. Our method consistently gets a leap of performance on several benchmarks with the best accuracy among deep unsupervised methods. Also, our method achieves competitive results to recent fully supervised methods while with much fewer parameters. Liang Liu 0007, Jiangning Zhang, Ruifei He, Yong Liu 0007, Yabiao Wang, Ying Tai, Donghao Luo 0001, Chengjie Wang 0001, Feiyue Huang |
CVPR | 6 |
| 2020 | Learning Multi-Granular Hypergraphs for Video-Based Person Re-IdentificationabstractVideo-based person re-identification (re-ID) is an important research topic in computer vision. The key to tackling the challenging task is to exploit both spatial and temporal clues in video sequences. In this work, we propose a novel graph-based framework, namely Multi-Granular Hypergraph (MGH), to pursue better representational capabilities by modeling spatiotemporal dependencies in terms of multiple granularities. Specifically, hypergraphs with different spatial granularities are constructed using various levels of part-based features across the video sequence. In each hypergraph, different temporal granularities are captured by hyperedges that connect a set of graph nodes (i.e., part-based features) across different temporal ranges. Two critical issues (misalignment and occlusion) are explicitly addressed by the proposed hypergraph propagation and feature aggregation schemes. Finally, we further enhance the overall video representation by learning more diversified graph-level representations of multiple granularities based on mutual information minimization. Extensive experiments on three widely-adopted benchmarks clearly demonstrate the effectiveness of the proposed framework. Notably, 90.0% top-1 accuracy on MARS is achieved using MGH, outperforming the state-of-the-arts. Yichao Yan, Jie Qin 0004, Jiaxin Chen 0002, Li Liu 0004, Fan Zhu 0001, Ying Tai, Ling Shao 0001 |
CVPR | 6 |
| 2020 | Adversarial Semantic Data Augmentation for Human Pose Estimation
Yanrui Bin, Xuan Cao, Xinya Chen, Yanhao Ge, Ying Tai, Chengjie Wang 0001, Feiyue Huang, Changxin Gao, Nong Sang |
ECCV (19) | 5 |
| 2020 | SSCGAN: Facial Attribute Editing via Style Skip Connections
Wenqing Chu, Ying Tai, Chengjie Wang 0001, Feiyue Huang, Rongrong Ji |
ECCV (15) | 2 |
| 2020 | Improving Face Recognition from Hard Samples via Distribution Distillation Loss
Yuge Huang, Pengcheng Shen, Ying Tai, Shaoxin Li 0001, Xiaoming Liu 0002, Feiyue Huang, Rongrong Ji |
ECCV (30) | 3 |
| 2020 | Chained-Tracker: Chaining Paired Attentive Regression Results for End-to-End Joint Multiple-Object Detection and Tracking
Jinlong Peng, Changan Wang, Fangbin Wan, Yang Wu 0001, Yabiao Wang, Ying Tai, Chengjie Wang 0001, Feiyue Huang, Yanwei Fu 0001 |
ECCV (4) | 6 |
| 2020 | Temporal Distinct Representation Learning for Action Recognition
Junwu Weng, Donghao Luo 0001, Yabiao Wang, Ying Tai, Chengjie Wang 0001, Feiyue Huang, Xudong Jiang 0001, Junsong Yuan 0001 |
ECCV (7) | 4 |
| 2020 | Face Anti-Spoofing via Disentangled Representation Learning
Ke-Yue Zhang, Taiping Yao, Jian Zhang 0079, Ying Tai, Shouhong Ding, Feiyue Huang, Lizhuang Ma |
ECCV (19) | 4 |
| 2020 | Person Search by Separated Modeling and A Mask-Guided Two-Stream CNN ModelabstractIn this work, we tackle the problem of person search, which is a challenging task consisted of pedestrian detection and person re-identification (re-ID). Instead of sharing representations in a single joint model, we find that separating detector and re-ID feature extraction yields better performance. In order to extract more representative features for each identity, we segment out the foreground person from the original image patch. We propose a simple yet effective re-ID method, which models foreground person and original image patches individually, and obtains enriched representations from two separate CNN streams. We also propose a Confidence Weighted Stream Attention method which further re-adjusts the relative importance of the two streams by incorporating the detection confidence. Furthermore, we simplify the whole pipeline by incorporating semantic segmentation into the re-ID network, which is trained by bounding boxes as weakly-annotated masks and identification labels simultaneously. From the experiments on two standard person search benchmarks i.e. CUHK-SYSU and PRW, we achieve mAP of 83.3% and 32.8% respectively, surpassing the state of the art by a large margin. The extensive ablation study and model inspection further justifies our motivation. Shanshan Zhang 0001, Wanli Ouyang, Jian Yang 0003, Ying Tai |
IEEE Trans. Image Process. | 5 |
| 2019 | Data-Adaptive Metric Learning with Scale AlignmentabstractThe central problem for most existing metric learning methods is to find a suitable projection matrix on the differences of all pairs of data points. However, a single unified projection matrix can hardly characterize all data similarities accurately as the practical data are usually very complicated, and simply adopting one global projection matrix might ignore important local patterns hidden in the dataset. To address this issue, this paper proposes a novel method dubbed “Data-Adaptive Metric Learning” (DAML), which constructs a data-adaptive projection matrix for each data pair by selectively combining a set of learned candidate matrices. As a result, every data pair can obtain a specific projection matrix, enabling the proposed DAML to flexibly fit the training data and produce discriminative projection results. The model of DAML is formulated as an optimization problem which jointly learns candidate projection matrices and their sparse combination for every data pair. Nevertheless, the over-fitting problem may occur due to the large amount of parameters to be learned. To tackle this issue, we adopt the Total Variation (TV) regularizer to align the scales of data embedding produced by all candidate projection matrices, and thus the generated metrics of these learned candidates are generally comparable. Furthermore, we extend the basic linear DAML model to the kernerlized version (denoted “KDAML”) to handle the non-linear cases, and the Iterative Shrinkage-Thresholding Algorithm (ISTA) is employed to solve the optimization model. Intensive experimental results on various applications including retrieval, classification, and verification clearly demonstrate the superiority of our algorithm to other state-of-the-art metric learning methodologies. Shuo Chen 0003, Chen Gong 0002, Jian Yang 0003, Ying Tai, Le Hui, Jun Li 0027 |
AAAI | 4 |
| 2019 | Towards Highly Accurate and Stable Face Alignment for High-Resolution VideosabstractIn recent years, heatmap regression based models have shown their effectiveness in face alignment and pose estimation. However, Conventional Heatmap Regression (CHR) is not accurate nor stable when dealing with high-resolution facial videos, since it finds the maximum activated location in heatmaps which are generated from rounding coordinates, and thus leads to quantization errors when scaling back to the original high-resolution space. In this paper, we propose a Fractional Heatmap Regression (FHR) for high-resolution video-based face alignment. The proposed FHR can accurately estimate the fractional part according to the 2D Gaussian function by sampling three points in heatmaps. To further stabilize the landmarks among continuous video frames while maintaining the precise at the same time, we propose a novel stabilization loss that contains two terms to address time delay and non-smooth issues, respectively. Experiments on 300W, 300VW and Talking Face datasets clearly demonstrate that the proposed method is more accurate and stable than the state-ofthe-art models. Ying Tai, Yicong Liang, Xiaoming Liu 0002, Lei Duan, Chengjie Wang 0001, Feiyue Huang, Yu Chen 0037 |
AAAI | 1 |
| 2019 | DSFD: Dual Shot Face DetectorabstractRecently, Convolutional Neural Network (CNN) has achieved great success in face detection. However, it remains a challenging problem for the current face detection methods owing to high degree of variability in scale, pose, occlusion, expression, appearance and illumination. In this Paper, we propose a novel detection network named Dual Shot face Detector(DSFD). which inherits the architecture of SSD and introduces a Feature Enhance Module (FEM) for transferring the original feature maps to extend the single shot detector to dual shot detector. Specially, progressive anchor loss (PAL) computed by using two set of anchors is adopted to effectively facilitate the features. Additionally, we propose an improved anchor matching (IAM) method by integrating novel data augmentation techniques and anchor design strategy in our DSFD to provide better initialization for the regressor. Extensive experiments on popular benchmarks: WIDER FACE (easy: 0.966, medium: 0.957, hard: 0.904) and FDDB ( discontinuous: 0.991, continuous: 0.862 ) demonstrate the superiority of DSFD over the state-of-the-art face detection methods (e.g., PyramidBox and SRN). Code will be made available upon publication. Jian Li 0062, Yabiao Wang, Changan Wang, Ying Tai, Jianjun Qian, Jian Yang 0003, Chengjie Wang 0001, Feiyue Huang |
CVPR | 4 |
| 2018 | FSRNet: End-to-End Learning Face Super-Resolution With Facial PriorsabstractFace Super-Resolution (SR) is a domain-specific superresolution problem. The facial prior knowledge can be leveraged to better super-resolve face images. We present a novel deep end-to-end trainable Face Super-Resolution Network (FSRNet), which makes use of the geometry prior, i.e., facial landmark heatmaps and parsing maps, to super-resolve very low-resolution (LR) face images without well-aligned requirement. Specifically, we first construct a coarse SR network to recover a coarse high-resolution (HR) image. Then, the coarse HR image is sent to two branches: a fine SR encoder and a prior information estimation network, which extracts the image features, and estimates landmark heatmaps/parsing maps respectively. Both image features and prior information are sent to a fine SR decoder to recover the HR image. To generate realistic faces, we also propose the Face Super-Resolution Generative Adversarial Network (FSRGAN) to incorporate the adversarial loss into FSRNet. Further, we introduce two related tasks, face alignment and parsing, as the new evaluation metrics for face SR, which address the inconsistency of classic metrics w.r.t. visual perception. Extensive experiments show that FSRNet and FSRGAN significantly outperforms state of the arts for very LR face SR, both quantitatively and qualitatively. Yu Chen 0037, Ying Tai, Xiaoming Liu 0002, Chunhua Shen, Jian Yang 0003 |
CVPR | 2 |
| 2018 | Person Search via a Mask-Guided Two-Stream CNN Model
Shanshan Zhang 0001, Wanli Ouyang, Jian Yang 0003, Ying Tai |
ECCV (7) | 5 |
| 2018 | SESR: Single Image Super Resolution with Recursive Squeeze and Excitation NetworksabstractSingle image super resolution is a very important computer vision task, with a wide range of applications. In recent years, the depth of the super-resolution model has been constantly increasing, but with a small increase in performance, it has brought a huge amount of computation and memory consumption. In this work, in order to make the super resolution models more effective, we proposed a novel single image super resolution method via recursive squeeze and excitation networks (SESR). By introducing the squeeze and excitation module, our SESR can model the interdependencies and relationships between channels and that makes our model more efficiency. In addition, the recursive structure and progressive reconstruction method in our model minimized the layers and parameters and enabled SESR to simultaneously train multi-scale super resolution in a single model. After evaluating on four benchmark test sets, our model is proved to be above the state-of-the-art methods in terms of speed and accuracy. Xiang Li 0041, Jian Yang 0003, Ying Tai |
ICPR | 4 |
| 2018 | Deep hierarchical guidance and regularization learning for end-to-end depth estimation
Zhenyu Zhang 0005, Chunyan Xu, Jian Yang 0003, Ying Tai, Liang Chen 0003 |
Pattern Recognit. | 4 |
| 2017 | Image Super-Resolution via Deep Recursive Residual NetworkabstractRecently, Convolutional Neural Network (CNN) based models have achieved great success in Single Image Super-Resolution (SISR). Owing to the strength of deep networks, these CNN models learn an effective nonlinear mapping from the low-resolution input image to the high-resolution target image, at the cost of requiring enormous parameters. This paper proposes a very deep CNN model (up to 52 convolutional layers) named Deep Recursive Residual Network (DRRN) that strives for deep yet concise networks. Specifically, residual learning is adopted, both in global and local manners, to mitigate the difficulty of training very deep networks, recursive learning is used to control the model parameters while increasing the depth. Extensive benchmark evaluation shows that DRRN significantly outperforms state of the art in SISR, while utilizing far fewer parameters. Code is available at https://github.com/tyshiwo/DRRN_CVPR17. Ying Tai, Jian Yang 0003, Xiaoming Liu 0002 |
CVPR | 1 |
| 2017 | MemNet: A Persistent Memory Network for Image RestorationabstractRecently, very deep convolutional neural networks (CNNs) have been attracting considerable attention in image restoration. However, as the depth grows, the longterm dependency problem is rarely realized for these very deep models, which results in the prior states/layers having little influence on the subsequent ones. Motivated by the fact that human thoughts have persistency, we propose a very deep persistent memory network (MemNet) that introduces a memory block, consisting of a recursive unit and a gate unit, to explicitly mine persistent memory through an adaptive learning process. The recursive unit learns multi-level representations of the current state under different receptive fields. The representations and the outputs from the previous memory blocks are concatenated and sent to the gate unit, which adaptively controls how much of the previous states should be reserved, and decides how much of the current state should be stored. We apply MemNet to three image restoration tasks, i.e., image denosing, super-resolution and JPEG deblocking. Comprehensive experiments demonstrate the necessity of the MemNet and its unanimous superiority on all three tasks over the state of the arts. Code is available at https://github.com/tyshiwo/MemNet. Ying Tai, Jian Yang 0003, Xiaoming Liu 0002, Chunyan Xu |
ICCV | 1 |
| 2017 | Kernel orthogonal Procrustes regression for face recognition across pose
Ying Tai, Jian Yang 0003, Lei Luo 0001, Jianjun Qian |
Neurocomputing | 1 |
| 2017 | Nuclear Norm Based Matrix Regression with Applications to Face Recognition with Occlusion and Illumination ChangesabstractRecently, regression analysis has become a popular tool for face recognition. Most existing regression methods use the one-dimensional, pixel-based error model, which characterizes the representation error individually, pixel by pixel, and thus neglects the two-dimensional structure of the error image. We observe that occlusion and illumination changes generally lead, approximately, to a low-rank error image. In order to make use of this low-rank structural information, this paper presents a two-dimensional image-matrix-based error model, namely, nuclear norm based matrix regression (NMR), for face representation and classification. NMR uses the minimal nuclear norm of representation error image as a criterion, and the alternating direction method of multipliers (ADMM) to calculate the regression coefficients. We further develop a fast ADMM algorithm to solve the approximate NMR model and show it has a quadratic rate of convergence. We experiment using five popular face image databases: the Extended Yale B, AR, EURECOM, Multi-PIE and FRGC. Experimental results demonstrate the performance advantage of NMR over the state-of-the-art regression-based methods for face recognition in the presence of occlusion and illumination variations. Jian Yang 0003, Lei Luo 0001, Jianjun Qian, Ying Tai, Fanlong Zhang, Yong Xu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2017 | Low-Rank Latent Pattern Approximation With Applications to Robust Image ClassificationabstractThis paper develops a novel method to address the structural noise in samples for image classification. Recently, regression-related classification methods have shown promising results when facing the pixelwise noise. However, they become weak in coping with the structural noise due to ignoring of relationships between pixels of noise image. Meanwhile, most of them need to implement the iterative process for computing representation coefficients, which leads to the high time consumption. To overcome these problems, we exploit a latent pattern model called low-rank latent pattern approximation (LLPA) to reconstruct the test image having structural noise. The rank function is applied to characterize the structure of the reconstruction residual between test image and the corresponding latent pattern. Simultaneously, the error between the latent pattern and the reference image is constrained by Frobenius norm to prevent overfitting. LLPA involves a closed-form solution by the virtue of a singular value thresholding operator. The provided theoretic analysis demonstrates that LLPA indeed removes the structural noise during classification task. Additionally, LLPA is further extended to the form of matrix regression by connecting multiple training samples, and alternating direction of multipliers method with Gaussian back substitution algorithm is used to solve the extended LLPA. Experimental results on several popular data sets validate that the proposed methods are more robust to image classification with occlusion and illumination changes, as compared to some existing state-of-the-art reconstruction-based methods and one deep neural network-based method. Shuo Chen 0003, Jian Yang 0003, Lei Luo 0001, Yang Wei 0003, Kaihua Zhang 0001, Ying Tai |
IEEE Trans. Image Process. | 6 |
| 2017 | Robust Nuclear Norm-Based Matrix Regression With Applications to Robust Face RecognitionabstractFace recognition (FR) via regression analysis-based classification has been widely studied in the past several years. Most existing regression analysis methods characterize the pixelwise representation error via l1-norm or l2-norm, which overlook the 2D structure of the error image. Recently, the nuclear norm-based matrix regression model is proposed to characterize low-rank structure of the error image. However, the nuclear norm cannot accurately describe the low-rank structural noise when the incoherence assumptions on the singular values does not hold, since it overpenalizes several much larger singular values. To address this problem, this paper presents the robust nuclear norm to characterize the structural error image and then extends it to deal with the mixed noise. The majorization-minimization (MM) method is applied to derive a iterative scheme for minimization of the robust nuclear norm optimization problem. Then, an efficiently alternating direction method of multipliers (ADMM) method is used to solve the proposed models. We use weighted nuclear norm as classification criterion to obtain the final recognition results. Experiments on several public face databases demonstrate the effectiveness of our models in handling with variations of structural noise (occlusion, illumination, and so on) and mixed noise. Jianchun Xie, Jian Yang 0003, Jianjun Qian, Ying Tai, Hengmin Zhang |
IEEE Trans. Image Process. | 4 |
| 2017 | Robust Image Regression Based on the Extended Matrix Variate Power Exponential Distribution of Dependent NoiseabstractDealing with partial occlusion or illumination is one of the most challenging problems in image representation and classification. In this problem, the characterization of the representation error plays a crucial role. In most current approaches, the error matrix needs to be stretched into a vector and each element is assumed to be independently corrupted. This ignores the dependence between the elements of error. In this paper, it is assumed that the error image caused by partial occlusion or illumination changes is a random matrix variate and follows the extended matrix variate power exponential distribution. This has the heavy tailed regions and can be used to describe a matrix pattern of l × m dimensional observations that are not independent. This paper reveals the essence of the proposed distribution: it actually alleviates the correlations between pixels in an error matrix E and makes E approximately Gaussian. On the basis of this distribution, we derive a Schatten p-norm-based matrix regression model with Lqregularization. Alternating direction method of multipliers is applied to solve this model. To get a closed-form solution in each step of the algorithm, two singular value function thresholding operators are introduced. In addition, the extended Schatten p-norm is utilized to characterize the distance between the test samples and classes in the design of the classifier. Extensive experimental results for image reconstruction and classification with structural noise demonstrate that the proposed algorithm works much more robustly than some existing regression-based methods. Lei Luo 0001, Jian Yang 0003, Jianjun Qian, Ying Tai, Gui-Fu Lu |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2017 | Low-Complexity Transformed Encoder Architectures for Quasi-Cyclic Nonbinary LDPC Codes Over SubfieldsabstractQuasi-cyclic low-density parity-check (QC-LDPC) codes are adopted in many digital communication and storage systems. The encoding of these codes is traditionally done by multiplying the message vector with a generator matrix consisting of dense circulant submatrices. To reduce the encoder complexity, this paper introduces two schemes making use of finite Fourier transform. We focus on QC-LDPC codes whose circulant submatrices are of dimension$(2^{r}-1)\times (2^{r}-1)$and the entries are elements of GF$(2^{p})$, where$p$divides$r$, and hence, GF$(2^{p})$is a subfield of GF$(2^{r})$. These cover a broad range of codes, and binary LDPC codes are a special case. Making use of conjugacy constraints, low-complexity architectures are developed for finite Fourier and inverse transforms over subfields in this paper. In addition, composite field arithmetic is exploited to eliminate the computations associated with message mapping and reduce the complexity of Fourier transform. For a (2016, 1074) nonbinary QC-LDPC code whose generator matrix consists of circulants of dimension$63 \times 63$with GF$(2^{2})$entries, the proposed encoders achieve 22% area reduction compared with the conventional encoders without sacrificing the throughput. Xinmiao Zhang 0001, Ying Tai |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | Volume measurement based tensor completionabstractThis paper presents a new tensor completion method named minimum volume constraint tensor completion. Unlike the nuclear norm penalization based methods, our method extends the conception of the matrix volume to the tensor volume, and uses the volume measurement as the penalization to address the tensor completion problem. The alternating direction method of multipliers (ADMM) algorithm is then employed to solve the optimization problem of the proposed model. Experimental results on several popular databases show superior performance of our method compared to the nuclear norm penalization based methods in terms of the accuracy and robustness. Jianchun Xie, Jian Yang 0003, Ying Tai, Jianjun Qian |
ICIP | 3 |
| 2016 | Structural Orthogonal Procrustes Regression for Face Recognition with Pose Variations and MisalignmentabstractRegression based method is a hot topic in the face recognition community and has achieved interesting results when dealing with well-aligned frontal face images. However, most of the existing regression analysis based methods are sensitive to pose variations. In this paper, we firstly introduce the orthogonal Procrustes problem (OPP), which is simple but effective, as a model to handle pose variations in two-dimensional face images. OPP seeks an optimal transformation between two images to correct the pose from one to the other. We integrate OPP into the regression model and propose the structural orthogonal Procrustes regression (SOPR) using the nuclear norm constraint on the error term to keep image's structural information. Moreover, a subject-wise strategy is adopted to address the problem that the gallery images may span over different poses. The proposed model is solved by an efficient iteratively reweighted algorithm and experimental results on popular face databases demonstrate the effectiveness of our method. Ying Tai, Jian Yang 0003, Fanlong Zhang, Yigong Zhang, Lei Luo 0001, Jianjun Qian |
SDM | 1 |
| 2016 | Exploring deep gradient information for biometric image feature representation
Jianjun Qian, Jian Yang 0003, Ying Tai |
Neurocomputing | 3 |
| 2016 | Adaptive noise dictionary construction via IRRPCA for face recognition
Yu Chen 0037, Jian Yang 0003, Lei Luo 0001, Hengmin Zhang, Jianjun Qian, Ying Tai, Jian Zhang 0025 |
Pattern Recognit. | 6 |
| 2016 | Learning discriminative singular value decomposition representation for face recognition
Ying Tai, Jian Yang 0003, Lei Luo 0001, Fanlong Zhang, Jianjun Qian |
Pattern Recognit. | 1 |
| 2016 | Face Recognition With Pose Variations and Misalignment via Orthogonal Procrustes RegressionabstractA linear regression-based method is a hot topic in face recognition community. Recently, sparse representation and collaborative representation-based classifiers for face recognition have been proposed and attracted great attention. However, most of the existing regression analysis-based methods are sensitive to pose variations. In this paper, we introduce the orthogonal Procrustes problem (OPP) as a model to handle pose variations existed in 2D face images. OPP seeks an optimal linear transformation between two images with different poses so as to make the transformed image best fits the other one. We integrate OPP into the regression model and propose the orthogonal Procrustes regression (OPR) model. To address the problem that the linear transformation is not suitable for handling highly non-linear pose variation, we further adopt a progressive strategy and propose the stacked OPR. As a practical framework, OPR can handle face alignment, pose correction, and face representation simultaneously. We optimize the proposed model via an efficient alternating iterative algorithm, and experimental results on three popular face databases, such as CMU PIE database, CMU Multi-PIE database, and LFW database, demonstrate the effectiveness of our proposed method. Ying Tai, Jian Yang 0003, Yigong Zhang, Lei Luo 0001, Jianjun Qian, Yu Chen 0037 |
IEEE Trans. Image Process. | 1 |
| 2015 | Nuclear-L1 norm joint regression for face reconstruction and recognition with mixed noise
Lei Luo 0001, Jian Yang 0003, Jianjun Qian, Ying Tai |
Pattern Recognit. | 4 |
| 2015 | Double Nuclear Norm-Based Matrix Decomposition for Occluded Image Recovery and Background ModelingabstractRobust principal component analysis (RPCA) is a new emerging method for exact recovery of corrupted low-rank matrices. It assumes that the real data matrix has low rank and the error matrix is sparse. This paper presents a method called double nuclear norm-based matrix decomposition (DNMD) for dealing with the image data corrupted by continuous occlusion. The method uses a unified low-rank assumption to characterize the real image data and continuous occlusion. Specifically, we assume all image vectors form a low-rank matrix, and each occlusion-induced error image is a low-rank matrix as well. Compared with RPCA, the low-rank assumption of DNMD is more intuitive for describing occlusion. Moreover, DNMD is solved by alternating direction method of multipliers. Our algorithm involves only one operator: the singular value shrinkage operator. DNMD, as a transductive method, is further extended into inductive DNMD (IDNMD). Both DNMD and IDNMD use nuclear norm for measuring the continuous occlusion-induced error, while many previous methods use L1 , L2 , or other M-estimators. Extensive experiments on removing occlusion from face images and background modeling from surveillance videos demonstrate the effectiveness of the proposed methods. Fanlong Zhang, Jian Yang 0003, Ying Tai, Jinhui Tang 0001 |
IEEE Trans. Image Process. | 3 |
| 2014 | Nuclear-L1 Norm Joint Regression for Face Reconstruction and Recognition
Lei Luo 0001, Jian Yang 0003, Jianjun Qian, Ying Tai |
ACCV (2) | 4 |