Yichao Yan

dblp:185/7881 · DBLP profile ↗
← Back
75ranked-venue papers
15as first author
58since 2021 · last 2026
0000-0003-3209-8965ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 57 · 9 first-author · 43 since 2021Artificial intelligence and machine learning · 45 · 12 first-author · 35 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Linking-free online spatio temporal action detection
Ningyu Sun, Yichao Yan, Guangtao Zhai, Xiaokang Yang 0001
Image Vis. Comput.3
2026 De-biasing Facial Albedo Estimation via Visual-Textual Cues
Xingyu Ren, Jiankang Deng, Chao Ma 0004, Yichao Yan, Wenhan Zhu, Xiaokang Yang 0001
J. Artif. Intell. Res.4
2026 Topo4D++: Realistic Physically Based 4D Head Capture With Topology-Preserving Gaussian Splatting and Expression Priors
abstract
4D head capture aims to generate dynamic facial meshes in the same topology with corresponding UV maps, which requires temporal correspondence between 3D head models. Existing pipelines either involve manual processing of artists or employ constraints such as landmark tracking and optical flow, failing to achieve a trade-off between accuracy and efficiency. To enhance this process, we propose Topo4D++, a novel framework for automatic geometry and texture reconstruction that optimizes densely aligned 4D heads and 8 K BRDF maps directly from calibrated multi-view videos. Our key insight is to represent facial models as a set of dynamic 3D Gaussians with fixed topology, where the Gaussian centers are bound to the mesh vertices. This enables tracking all vertices rather than sparse vertices on the face accurately by leveraging the inverse rendering capabilities of 3D Gaussian Splatting (3DGS), while also enabling ultra-high-resolution texture generation. To maintain face structure during dynamic 3DGS optimization, we propose to optimize geometry and texture alternatively under physical and topological constraints frame-by-frame and employ blendshape-based expression priors to address extreme expressions. Then, we propose to extract dynamic facial meshes in a regular wiring arrangement and high-fidelity textures with pore-level details from the learned Gaussians. Finally, we train a diffusion-based model to generate BRDF texture maps to achieve physically based rendering. Given the absence of a universal benchmark, we construct JHead, a novel benchmark for the comprehensive evaluation of 4D head capture methods. Extensive experiments on different datasets demonstrate that our method is generalized to different capture systems, identities, and expressions, outperforming current state-of-the-art head reconstruction methods in both mesh and texture qualitatively and quantitatively.
Yuhao Cheng, Xuanchen Li, Xingyu Ren, Haozhe Jia, Di Xu 0012, Wenhan Zhu, Bingbing Ni, Yichao Yan
IEEE Trans. Pattern Anal. Mach. Intell.9
2026 AgonicDreamer: Enhancing Multi-View Consistency in Text-to-3D Generation via Rectified Score Distillation
abstract
Score Distillation Sampling and its variants have shown strong potential in text-to-3D generation by leveraging scores estimated from pretrained text-to-image diffusion models to optimize 3D representations. However, due to the view-agnostic nature of these scores, existing methods often suffer from the multi-face Janus problem, leading to inconsistencies across different views. In this work, we propose Rectified Score Distillation, which addresses this issue by incorporating view-conditioned scores as priors. Specifically, we formulate a reverse ordinary differential equation (ODE) that is additionally conditioned on camera poses. Then we rectify the standard, view-irrelevant scores to approximate the desired gradients along this ODE. Building on these rectified scores, our full framework, named AgonicDreamer, enables the generation of photorealistic and multi-view consistent 3D content with fine-grained details, as validated by extensive experimental results.
Shanyan Guan, Yanhao Ge, Wei Li 0112, Yichao Yan, Chao Ma 0004, Xiaokang Yang 0001
IEEE Trans. Image Process.6
2026 ExpDiff: Generating High-Fidelity 3D Facial Expression Meshes and BRDF Textures via Diffusion Model
abstract
3D face generation is a critical task for immersive multimedia applications, where a key challenge is the joint synthesis of expressive geometry and BRDF textures. Existing methods often struggle with geometric-textural coherence and corresponding reflectance modeling. To overcome these limitations, we present ExpDiff, a framework that generates expression meshes and corresponding BRDF textures from a single neutral-expression face. Our method employs an attention-based diffusion model to learn the semantic transition across expressions. To ensure correspondence between geometry and texture, we introduce a unified representation that explicitly models geometric-textural interaction, which is encoded into a shared latent space by models pre-trained on a vast dataset for strong generalization. To achieve semantically coherent and physically consistent generation, we propose to guide the denoising direction with specially designed textual prompts. We further construct two novel facial expression datasets, J-Reflectance, for ultra-high-quality assets, and FFHQ-BRDFExp for diverse identities, both of which are publicly released to advance the community. Extensive experiments demonstrate our method's superior performance in photo-realistic facial expression synthesis. Project page:https://cyh-sj.github.io/expdiff/.
Yuhao Cheng, Xuanchen Li, Xingyu Ren, Zhuo Chen 0060, Chenghui Ke, Xiaokang Yang 0001, Yichao Yan
IEEE Trans. Multim.7
2026 SingingHead: A Large-Scale 4D Dataset for Singing Head Animation
abstract
Singing, as a common facial movement second only to talking, can be regarded as a universal language across ethnicities and cultures, plays an important role in emotional communication, art, and entertainment. However, it is often overlooked in the field of audio-driven 3D facial animation due to the lack of singing head datasets and the domain gap between singing and talking in rhythm and amplitude. To this end, we collect a large-scale high-quality multi-modal singing head dataset,SingingHead, which consists of more than 27 hours of synchronized singing video, 3D facial motion, singing audio, and background music from 76 individuals and 8 types of music. Along with the SingingHead dataset, we benchmark existing audio-driven 3D facial animation methods and 2D talking head methods on the singing task. Existing 3D facial animation methods and 2D talking head methods fail to produce satisfactory singing results. Focusing on the 3D singing head animation, we first utilize the proposed singing-specific dataset to retrain the 3D facial animation methods, resulting in substantial performance improvements. Besides, considering the absence of background music and the slow generation speed of existing methods, we propose a simple but efficient non-autoregressive VAE-based framework with background music as an input signal to generate diverse and accurate 3D singing facial motions in real time. Extensive experiments demonstrate the significance of the SingingHead dataset in promoting the development of singing head animation. The dataset is released for research purposes at:https://wsj-sjtu.github.io/SingingHead/.
Sijing Wu, Weitian Zhang, Jun Jia, Yucheng Zhu, Yichao Yan, Guangtao Zhai, Xiaokang Yang 0001
IEEE Trans. Multim.6
2026 Relightable and Animatable Gaussian Head Avatar From Monocular Videos
abstract
In the realm of virtual avatar creation, accurate relighting capabilities are key to enhancing realism and immersion. We propose a novel pipeline for building personalized and relightable avatars from a monocular video captured under unknown lighting. This minimal input poses challenges in material entanglement and novel-view inconsistency. To tackle these, we introduce a disentangled dynamic 3D Gaussian representation that models diverse material properties and supports photorealistic rendering and animation via a parametric face model. To resolve material ambiguity under uncontrolled lighting, we train a 2D diffusion-based model to predict canonical-lighting images and physically-based material maps from casually lit portraits. These predictions serve as supervisory signals to guide the 3D disentanglement process. Additionally, we incorporate a 3D prior to enhance novel-view consistency, improving geometry and appearance in unseen views. Experiments demonstrate that our approach significantly boosts reconstruction quality and relighting fidelity, offering a practical and cost-effective solution for creating high-quality personalized avatars.
Zhuo Chen 0060, Yichao Yan, Jingnan Gao, Zhuo Su 0008, Zhaohu Li, Yuhao Cheng, Xueying Lee, Yutong Leng, Yikun Zeng, Guidong Wang, Xiaokang Yang 0001
IEEE Trans. Vis. Comput. Graph.2
2025 Towards High-fidelity 3D Talking Avatar with Personalized Dynamic Texture
abstract
Significant progress has been made for speech-driven 3D face animation, but most works focus on learning the motion of mesh/geometry, ignoring the impact of dynamic texture. In this work, we reveal that dynamic texture plays a key role in rendering high-fidelity talking avatars, and introduce a high-resolution 4D dataset TexTalk4D, consisting of 100 minutes of audio-synced scan-level meshes with detailed 8K dynamic textures from 100 subjects. Based on the dataset, we explore the inherent correlation between motion and texture, and propose a diffusion-based framework TexTalker to simultaneously generate facial motions and dynamic textures from speech. Furthermore, we propose a novel pivot-based style injection strategy to capture the complicity of different texture and motion styles, which allows disentangled control. TexTalker, as the first method to generate audio-synced facial motion with dynamic texture, not only outperforms the prior arts in synthesising facial motions, but also produces realistic textures that are consistent with the underlying facial movements. Project page: https://xuanchenli.github.io/TexTalk/.
Xuanchen Li, Yuhao Cheng, Yikun Zeng, Xingyu Ren, Wenhan Zhu, Weiming Zhao, Yichao Yan
CVPR8
2025 S^3-Face: SSS-Compliant Facial Reflectance Estimation via Diffusion Priors
abstract
Recent 3D face reconstruction methods have made remarkable advancements, yet achieving high-quality facial reflectance from monocular input remains challenging. Existing methods rely on the light-stage captured data to learn facial reflectance models. However, limited subject diversity in these datasets poses challenges in achieving good generalization and broad applicability. This motivates us to explore whether the extensive priors captured in recent generative diffusion models (e.g., Stable Diffusion) can enable more generalizable facial reflectance estimation as these models have been pre-trained on large-scale internet image collections containing rich visual patterns. In this paper, we introduce the use of Stable Diffusion as a prior for facial reflectance estimation, achieving robust results with minimal captured data for fine-tuning. We present S3-Face, a comprehensive framework capable of producing SSS-compliant skin reflectance from in-the-wild images. Our method adopts a two-stage training approach: in the first stage, DSN-Net is trained to predict diffuse albedo, specular albedo, and normal maps from in-the-wild images using a novel joint reflectance attention module. In the second stage, HM-Net is trained to generate hemoglobin and melanin maps based on the diffuse albedo predicted in the first stage, yielding SSS-compliant and detailed reflectance maps. Extensive experiments demonstrate that our method achieves strong generalization and produces high-fidelity, SSS-compliant facial reflectance estimation.
Xingyu Ren, Jiankang Deng, Yuhao Cheng, Wenhan Zhu, Yichao Yan, Xiaokang Yang 0001, Stefanos Zafeiriou, Chao Ma 0004
CVPR5
2025 Multimodal Latent Diffusion Model for Complex Sewing Pattern Generation
abstract
Generating sewing patterns in garment design is receiving increasing attention due to its CG-friendly and flexible-editing nature. Previous sewing pattern generation methods have been able to produce exquisite clothing, but struggle to design complex garments with detailed control. To address these issues, we propose SewingLDM, a multi-modal generative model that generates sewing patterns controlled by text prompts, body shapes, and garment sketches. Initially, we extend the original vector of sewing patterns into a more comprehensive representation to cover more intricate details and then compress them into a compact latent space. To learn the sewing pattern distribution in the latent space, we design a two-step training strategy to inject the multi-modal conditions, \ie, body shapes, text prompts, and garment sketches, into a diffusion model, ensuring the generated garments are body-suited and detail-controlled. Comprehensive qualitative and quantitative experiments show the effectiveness of our proposed method, significantly surpassing previous approaches in terms of complex garment design and various body adaptability. Our project page: https://shengqiliu1.github.io/SewingLDM.
Shengqi Liu, Yuhao Cheng, Zhuo Chen 0060, Xingyu Ren, Wenhan Zhu, Lincheng Li, Mengxiao Bi, Xiaokang Yang 0001, Yichao Yan
ICCV9
2025 Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions
Liang Xu 0012, Chengqun Yang, Zili Lin, Fei Xu 0008, Congsheng Xu, Yiyi Zhang 0002, Jie Qin 0004, Xingdong Sheng, Yunhui Liu 0006, Xin Jin 0014, Yichao Yan, Wenjun Zeng 0001, Xiaokang Yang 0001
ICCV12
2025 Skip-Vision: Efficient and Scalable Acceleration of Vision-Language Models Via Adaptive Token Skipping
abstract
Transformer-based models have driven significant advancements in Multimodal Large Language Models (MLLMs), yet their computational costs surge drastically when scaling resolution, training data, and model parameters. A key bottleneck stems from the proliferation of visual tokens required for fine-grained image understanding. We propose Skip-Vision, a unified framework addressing both training and inference inefficiencies in vision-language models. On top of conventional token compression approaches, our method introduces two complementary acceleration strategies. For training acceleration, we observe that Feed-Forward Network (FFN) computations on visual tokens induce marginal feature updates. This motivates our Skip-FFN strategy, which bypasses FFN layers for redundant visual tokens. For inference acceleration, we design a selective KV-cache removal mechanism that prunes the skipped key-value pairs during decoding while preserving model performance. Experimental results demonstrate that Skip-Vision reduces training time by up to 35\%, inference FLOPs by 75\%, and latency by 45\%, while achieving comparable or superior performance to existing methods. Our work provides a practical solution for scaling high-performance MLLMs with enhanced efficiency.
Kaixiang Ji, Yichao Yan
ICCV4
2025 Disentangled Clothed Avatar Generation with Layered Representation
abstract
Clothed avatar generation has wide applications in virtual and augmented reality, filmmaking, and more. Previous methods have achieved success in generating diverse digital avatars, however, generating avatars with disentangled components (\eg, body, hair, and clothes) has long been a challenge. In this paper, we propose LayerAvatar, the first feed-forward diffusion-based method for generating component-disentangled clothed avatars. To achieve this, we first propose a layered UV feature plane representation, where components are distributed in different layers of the Gaussian-based UV feature plane with corresponding semantic labels. This representation supports high-resolution and real-time rendering, as well as expressive animation including controllable gestures and facial expressions. Based on the well-designed representation, we train a single-stage diffusion model and introduce constrain terms to address the severe occlusion problem of the innermost human body layer. Extensive experiments demonstrate the impressive performances of our method in generating disentangled clothed avatars, and we further explore its applications in component transfer. The project page is available at: https://olivia23333.github.io/LayerAvatar/
Weitian Zhang, Yichao Yan, Sijing Wu, Manwen Liao, Xiaokang Yang 0001
ICCV2
2025 AniSDF: Fused-Granularity Neural Surfaces with Anisotropic Encoding for High-Fidelity 3D Reconstruction
abstract
Neural radiance fields have recently revolutionized novel-view synthesis and achieved high-fidelity renderings. However, these methods sacrifice the geometry for the rendering quality, limiting their further applications including relighting and deformation. How to synthesize photo-realistic rendering while reconstructing accurate geometry remains an unsolved problem. In this work, we present AniSDF, a novel approach that learns fused-granularity neural surfaces with physics-based encoding for high-fidelity 3D reconstruction. Different from previous neural surfaces, our fused-granularity geometry structure balances the overall structures and fine geometric details, producing accurate geometry reconstruction. To disambiguate geometry from reflective appearance, we introduce blended radiance fields to model diffuse and specularity following the anisotropic spherical Gaussian encoding, a physics-based rendering pipeline. With these designs, AniSDF can reconstruct objects with complex structures and produce high-quality renderings. Furthermore, our method is a unified model that does not require complex hyperparameter tuning for specific objects. Extensive experiments demonstrate that our method boosts the quality of SDF-based methods by a great scale in both geometry reconstruction and novel-view synthesis.
Jingnan Gao, Zhuo Chen 0060, Xiaokang Yang 0001, Yichao Yan
ICLR4
2025 PostEdit: Posterior Sampling for Efficient Zero-Shot Image Editing
abstract
In the field of image editing, three core challenges persist: controllability, background preservation, and efficiency. Inversion-based methods rely on time-consuming optimization to preserve the features of the initial images, which results in low efficiency due to the requirement for extensive network inference. Conversely, inversion-free methods lack theoretical support for background similarity, as they circumvent the issue of maintaining initial features to achieve efficiency. As a consequence, none of these methods can achieve both high efficiency and background consistency. To tackle the challenges and the aforementioned disadvantages, we introduce PostEdit, a method that incorporates a posterior scheme to govern the diffusion sampling process. Specifically, a corresponding measurement term related to both the initial features and Langevin dynamics is introduced to optimize the estimated image generated by the given target prompt. Extensive experimental results indicate that the proposed PostEdit achieves state-of-the-art editing performance while accurately preserving unedited regions. Furthermore, the method is both inversion- and training-free, necessitating approximately 1.5 seconds and 18 GB of GPU memory to generate high-quality results.
Yichao Yan, Shanyan Guan, Yanhao Ge, Xiaokang Yang 0001
ICLR3
2025 Enhancing Visual Localization with Cross-Domain Image Generation
abstract
Visual localization aims to predict the absolute camera pose for a single query image. However, predominant methods focus on single-camera images and scenes with limited appearance variations, limiting their applicability to cross-domain scenes commonly encountered in real-world applications. Furthermore, the long-tail distribution of cross-domain datasets poses additional challenges for visual localization. In this work, we propose a novel cross-domain data generation method to enhance visual localization methods. To achieve this, we first construct a cross-domain 3DGS to accurately model photometric variations and mitigate the interference of dynamic objects in large-scale scenes. We introduce a text-guided image editing model to enhance data diversity for addressing the long-tail distribution problem and design an effective fine-tuning strategy for it. Then, we develop an anchor-based method to generate high-quality datasets for visual localization. Finally, we introduce positional attention to address data ambiguities in cross-camera images. Extensive experiments show that our method achieves state-of-the-art accuracy, outperforming existing cross-domain visual localization methods by an average of 59% across all domains. Project page: https://yzwang-sjtu.github.io/CDG-Loc.
Yuanze Wang, Yichao Yan, Shiming Song 0003, Songchang Jin, Yilan Huang, Xingdong Sheng, Dian-xi Shi
ICML2
2025 DCN: Decoupled-Coupled Network for Text-based Person Search
abstract
Text-based person search aims to identify a person based on textual descriptions, by simultaneously addressing person detection and cross-modal alignment between text queries and person images. Existing approaches often struggle with conflicts in exploiting proposals across these two sub-tasks. Specifically, cross-modal alignment requires highly precise proposals, while person detection can tolerate a certain degree of proposal inaccuracy but always needs a large number of proposals. In this paper, we propose the Decoupled-Coupled Network (DCN) to tackle the above conflicts. We first attempt to resolve the above conflicts by proposing a Decoupled Proposal Selection (DPS) strategy, inspired by the divide-and-conquer principle. DPS adaptively selects the most suitable proposals for each sub-task, ensuring their distinct requirements are adequately met. We further present a Coupled Cascade Refinement (CCR) module to jointly optimize both sub-tasks in a multi-stage manner, progressively improving detection accuracy and fostering cross-modal alignment between text and person image. In addition, we introduce two types of objective functions to optimize the inherently multi-positive contrastive learning challenge. Extensive experiments conducted on two benchmarks demonstrate the effectiveness and superiority of our DCN over existing competitors.
Rong Quan, Liangxu Su, Wentong Li 0001, Yichao Yan, Jie Qin 0004
MMAsia6
2025 Correlated Low-Rank Adaptation for ConvNets
abstract
Low-Rank Adaptation (LoRA) methods have demonstrated considerable success in achieving parameter-efficient fine-tuning (PEFT) for Transformer-based foundation models. These methods typically fine-tune individual Transformer layers using independent LoRA adaptations. However, directly applying existing LoRA techniques to convolutional networks (ConvNets) yields unsatisfactory results due to the high correlation between the stacked sequential layers of ConvNets. To overcome this challenge, we introduce a novel framework called Correlated Low-Rank Adaptation (CoLoRA), which explicitly utilizes correlated low-rank matrices to model the inter-layer dependencies among convolutional layers. Additionally, to enhance tuning efficiency, we propose a parameter-free filtering method that enlarges the receptive field of LoRA, thus minimizing interference from non-informative local regions. Comprehensive experiments conducted across various mainstream vision tasks, including image classification, semantic segmentation, and object detection, illustrate that CoLoRA significantly advances the state-of-the-art PEFT approaches. Notably, our CoLoRA achieves superior performance with only 5\% of trainable parameters, surpassing full fine-tuning in the image classification task on the VTAB-1k dataset using ConvNeXt-S. Code is available at [https://github.com/VISION-SJTU/CoLoRA](https://github.com/VISION-SJTU/CoLoRA).
Wu Ran, Shuyang Pang, Jinfan Liu, Jingsheng Liu, Yichao Yan, Chao Ma 0004
NeurIPS9
2025 TGAvatar: Reconstructing 3D Gaussian Avatars With Transformer-Based Tri-Plane
abstract
We introduce TGAvatar, a novel framework for 3D head animation and reconstruction that revolutionizes the use of 3D Gaussian Splatting (3DGS). TGAvatar significantly advances rendering quality by leveraging the intricate properties of 3DGS to achieve detailed and realistic representations of human head geometries and textures. We use an innovative application of linear blending techniques to imitate 3D Morphable Model (3DMM) coefficients within 3DGS, thereby enabling precise and dynamic facial feature and expression modeling. Further enhancing TGAvatar’s capabilities, a transformer based tri-plane module is incorporated to accurately infer spherical harmonics and alpha parameters. This integration is pivotal for the method, as it allows allows us to efficiently and precisely represent the visual characteristics of gaussians, tailored specifically to the intricate details of the head’s components. Our exhaustive evaluations show that TGAvatar not only elevates the fidelity and realism of 3D head reconstructions but also sets a new standard by surpassing existing methods in rendering quality and computational efficiency. Please see our project page athttps://hrg0417.github.io/TGAvatar/
Ruigang Hu, Xuekuan Wang, Yichao Yan, Cairong Zhao
IEEE Trans. Circuits Syst. Video Technol.3
2025 GPS: Generalizable Person Search on Large-Scale User-Generated Video Content
abstract
Person search is a challenging task that involves detecting and retrieving individuals from a large set of un-cropped scene images. Existing person search models are mostly trained and deployed in the same-origin scenarios. However, collecting and annotating training samples for each scene is difficult due to the limitation of resources and labor cost. Moreover, large-scale intra-domain data for training are generally not legally available for common developers, due to the regulation of privacy and public security. Leveraging easily accessible large-scale User Generated Video Contents (i.e. UGC videos) to train person search models can fit the real-world distribution, but still suffering a performance gap from the domain difference to surveillance scenes. In this work, we explore enhancing the out-of-domain generalization capabilities of person search models, and propose a generalizable framework on both feature-level and data-level generalization to facilitate downstream tasks in arbitrary scenarios. Specifically, we focus on learning domain-invariant representations for both detection and ReID by introducing a multi-task prototype-based domain-specific batch normalization, and a channel-wise ID-relevant feature decorrelation strategy. We also identify and address typical sources of noise in UGC training frames, including inaccurate bounding boxes, the omission of identity labels, and the absence of cross-camera data. Our framework achieves promising performance on two challenging person search benchmarks without using any human annotation or samples from the target domain. The code is available athttps://github.com/caposerenity/GPS.
Guanshuo Wang, Yichao Yan, Fufu Yu, Qiong Jia 0004, Jie Qin 0004, Shouhong Ding, Xiaokang Yang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 IPAD: Industrial Process Anomaly Detection Dataset
abstract
Video anomaly detection (VAD) is a challenging task aiming to recognize anomalies in video frames, and existing large-scale VAD researches primarily focus on road traffic and human activity scenes. In industrial scenes, there are often a variety of unpredictable anomalies, and the VAD method can play a significant role in these scenarios. However, there is a lack of applicable datasets and methods specifically tailored for industrial production scenarios due to concerns regarding privacy and security. To bridge this gap, we propose a new dataset, IPAD, specifically designed for VAD in industrial scenarios. The industrial processes in our dataset are chosen through on-site factory research and discussions with engineers. This dataset covers 16 different industrial devices and contains over 6 hours of both synthetic and real-world video footage. Moreover, we annotate the key feature of the industrial process, i.e., periodicity. Based on the proposed dataset, we introduce a period memory module and a sliding window inspection mechanism to effectively investigate the periodic information in a basic reconstruction model. Our framework leverages LoRA adapter to explore the effective migration of pretrained models, which are initially trained using synthetic data, into real-world scenarios. Our proposed dataset and method will fill the gap in the field of industrial video anomaly detection and drive the process of video understanding tasks as well as smart factory deployment. Project page:https://ljf1113.github.io/IPAD_VAD.
Jinfan Liu, Yichao Yan, Weiming Zhao, Pengzhi Chu, Xingdong Sheng, Yunhui Liu 0006, Xiaokang Yang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 Revealing Directions for Text-Guided 3D Face Editing
abstract
3D face editing is a significant task in multimedia, aimed at the manipulation of 3D face models across various control signals. The success of 3D-aware GAN provides expressive 3D models learned from 2D single-view images only, encouraging researchers to discover semantic editing directions in its latent space. However, previous methods face challenges in balancing quality, efficiency, and generalization. To solve the problem, we explore the possibility of introducing the strength of diffusion model into 3D-aware GANs. In this paper, we presentFace Clan, a fast and text-general approach for generating and manipulating 3D faces based on arbitrary attribute descriptions. To achieve disentangled editing, we propose to diffuse on the latent space under a pair of opposite prompts to estimate the mask indicating the region of interest on latent codes. Based on the mask, we then apply denoising to the masked latent codes to reveal the editing direction. Our method offers a precisely controllable manipulation method, allowing users to intuitively customize regions of interest with the text description. Experiments demonstrate the effectiveness and generalization of our Face Clan for various pre-trained GANs. It offers an intuitive and wide application for text-guided face editing that contributes to the landscape of multimedia content creation. Our project page:https://windlikestone.github.io/Face_clan_website/.
Zhuo Chen 0060, Yichao Yan, Shengqi Liu, Yuhao Cheng, Weiming Zhao, Lincheng Li, Mengxiao Bi, Xiaokang Yang 0001
IEEE Trans. Multim.2
2025 Disaggregation Distillation for Person Search
abstract
Person search is a challenging task in computer vision and multimedia understanding, which aims at localizing and identifying target individuals in realistic scenes. State-of-the-art models achieve remarkable success but suffer from overloaded computation and inefficient inference, making them impractical in most real-world applications. A promising approach to tackle this dilemma is to compress person search models with knowledge distillation (KD). Previous KD-based person search methods typically distill the knowledge from the re-identification (re-id) branch, completely overlooking the useful knowledge from the detection branch. In addition, we elucidate that the imbalance between person and background regions in feature maps has a negative impact on the distillation process. To this end, we propose a novel KD-based approach, namely Disaggregation Distillation for Person Search (DDPS), which disaggregates the distillation process and feature maps, respectively. Firstly, the distillation process is disaggregated into two task-oriented sub-processes,i.e., detection distillation and re-id distillation, to help the student learn both accurate localization capability and discriminative person embeddings. Secondly, we disaggregate each feature map into person and background regions, and distill these two regions independently to alleviate the imbalance problem. More concretely, three types of distillation modules,i.e., logit distillation (LD), correlation distillation (CD), and disaggregation feature distillation (DFD), are particularly designed to transfer comprehensive information from the teacher to the student. Note that such a simple yet effective distillation scheme can be readily applied to both homogeneous and heterogeneous teacher-student combinations. We conduct extensive experiments on two person search benchmarks, where the results demonstrate that, surprisingly, our DDPS enables the student model to surpass the performance of the corresponding teacher model, even achieving comparable results with general person search models.
Rong Quan, Haiyan Chen 0001, Jiamei Liu, Yichao Yan, Song Bai 0001, Jie Qin 0004
IEEE Trans. Multim.5
2025 EvaSurf: Efficient View-Aware Implicit Textured Surface Reconstruction
abstract
Reconstructing real-world 3D objects has numerous applications in computer vision, such as virtual reality, video games, and animations. Ideally, 3D reconstruction methods should generate high-fidelity results with 3D consistency in real-time. Traditional methods match pixels between images using photo-consistency constraints or learned features, while differentiable rendering methods like Neural Radiance Fields (NeRF) use differentiable volume rendering or surface-based representation to generate high-fidelity scenes. However, these methods require excessive runtime for rendering, making them impractical for daily applications. To address these challenges, we present EvaSurf, an Efficient View-Aware implicit textured Surface reconstruction method on mobile devices. In our method, we first employ an efficient surface-based model with a multi-view supervision module to ensure accurate mesh reconstruction. To enable high-fidelity rendering, we learn an implicit texture embedded with view-aware encoding to capture view-dependent information. Furthermore, with the explicit geometry and the implicit texture, we can employ a lightweight neural shader to reduce the expense of computation and further support real-time rendering on common mobile devices. Extensive experiments demonstrate that our method can reconstruct high-quality appearance and accurate mesh on both synthetic and real-world datasets. Moreover, our method can be trained in just 1-2 hours using a single GPU and run on mobile devices at over 40 FPS (Frames Per Second), with a final package required for rendering taking up only 40-50 MB.
Jingnan Gao, Zhuo Chen 0060, Yichao Yan, Bowen Pan, Jiangjing Lyu, Xiaokang Yang 0001
IEEE Trans. Vis. Comput. Graph.3
2024 LERE: Learning-Based Low-Rank Matrix Recovery with Rank Estimation
abstract
A fundamental task in the realms of computer vision, Low-Rank Matrix Recovery (LRMR) focuses on the inherent low-rank structure precise recovery from incomplete data and/or corrupted measurements given that the rank is a known prior or accurately estimated. However, it remains challenging for existing rank estimation methods to accurately estimate the rank of an ill-conditioned matrix. Also, existing LRMR optimization methods are heavily dependent on the chosen parameters, and are therefore difficult to adapt to different situations. Addressing these issues, A novel LEarning-based low-rank matrix recovery with Rank Estimation (LERE) is proposed. More specifically, considering the characteristics of the Gerschgorin disk's center and radius, a new heuristic decision rule in the Gerschgorin Disk Theorem is significantly enhanced and the low-rank boundary can be exactly located, which leads to a marked improvement in the accuracy of rank estimation. According to the estimated rank, we select row and column sub-matrices from the observation matrix by uniformly random sampling. A 17-iteration feedforward-recurrent-mixed neural network is then adapted to learn the parameters in the sub-matrix recovery processing. Finally, by the correlation of the row sub-matrix and column sub-matrix, LERE successfully recovers the underlying low-rank matrix. Overall, LERE is more efficient and robust than existing LRMR methods. Experimental results demonstrate that LERE surpasses state-of-the-art (SOTA) methods. The code for this work is accessible at https://github.com/zhengqinxu/LERE.
Zhengqin Xu, Yulun Zhang 0001, Chao Ma 0004, Yichao Yan, Zelin Peng, Shoulie Xie, Shiqian Wu, Xiaokang Yang 0001
AAAI4
2024 3D-Aware Face Editing via Warping-Guided Latent Direction Learning
abstract
3D facial editing, a longstanding task in computer vision with broad applications, is expected to fast and intuitively manipulate any face from arbitrary viewpoints following the user's will. Existing works have limitations in terms of intuitiveness, generalization, and efficiency. To overcome these challenges, we propose FaceEdit3D, which allows users to directly manipulate 3D points to edit a 3D face, achieving natural and rapid face editing. After one or several points are manipulated by users, we propose the tri-plane warping to directly deform the view-independent 3D representation. To address the problem of distortion caused by tri-plane warping, we train a warp-aware encoder to project the warped face onto a standardized latent space. In this space, we further propose directional latent editing to mitigate the identity bias caused by the encoder and realize the disentangled editing of various attributes. Extensive experiments show that our method achieves superior results with rich facial details and nice identity preservation. Our approach also supports general applications like multi-attribute continuous editing and cat/car editing. The project website is https://cyh-sj.github.io/FaceEdit3DI.
Yuhao Cheng, Zhuo Chen 0060, Xingyu Ren, Wenhan Zhu, Zhengqin Xu, Di Xu 0012, Changpeng Yang, Yichao Yan
CVPR8
2024 Monocular Identity-Conditioned Facial Reflectance Reconstruction
abstract
Recent 3D face reconstruction methods have made re-markable advancements, yet there remain huge challenges in monocular high-quality facial reflectance reconstruction. Existing methods rely on a large amount of light-stage captured data to learn facial reflectance models. However, the lack of subject diversity poses challenges in achieving good generalization and widespread applicability. In this paper, we learn the reflectance prior in image space rather than UV space and present a framework named ID2Reflectance. Our framework can directly estimate the reflectance maps of a single image while using limited reflectance data for training. Our key insight is that reflectance data shares facial structures with RGB faces, which enables obtaining expressive facial prior from inexpensive RGB data thus re-ducing the dependency on reflectance data. We first learn a high-quality prior for facial reflectance. Specifically, we pretrain multi-domain facial feature code books and design a codebook fusion method to align the reflectance and RGB domains. Then, we propose an identity-conditioned swapping module that injects facial identity from the target image into the pre-trained autoencoder to modify the identity of the source reflectance image. Finally, we stitch multi-view swapped reflectance images to obtain renderable assets. Extensive experiments demonstrate that our method exhibits excellent generalization capability and achieves state-of-the-art facial reflectance reconstruction results for in-the-wild faces. Our project page is https://xingyuren.github.io/id2reflectance.
Xingyu Ren, Jiankang Deng, Yuhao Cheng, Jia Guo 0003, Chao Ma 0004, Yichao Yan, Wenhan Zhu, Xiaokang Yang 0001
CVPR6
2024 Inter-X: Towards Versatile Human-Human Interaction Analysis
abstract
The analysis of the ubiquitous human-human interactions is pivotal for understanding humans as social beings. Existing human-human interaction datasets typically suffer from inaccurate body motions, lack of hand gestures and fine- grained textual descriptions. To better perceive and generate human-human interactions, we propose Inter-X, a currently largest human-human interaction dataset with accurate body movements and diverse interaction patterns, together with detailed hand gestures. The dataset includes
Liang Xu 0012, Xintao Lv, Yichao Yan, Xin Jin 0014, Shuwen Wu, Congsheng Xu, Yizhou Zhou, Fengyun Rao, Xingdong Sheng, Yunhui Liu 0006, Wenjun Zeng 0001, Xiaokang Yang 0001
CVPR3
2024 ReGenNet: Towards Human Action-Reaction Synthesis
abstract
Humans constantly interact with their surrounding environments. Current human-centric generative models mainly focus on synthesizing humans plausibly interacting with static scenes and objects, while the dynamic human action-reaction synthesis for ubiquitous causal human-human interactions is less explored. Human-human interactions can be regarded as asymmetric with actors and reactors in atomic interaction periods. In this paper, we compre-hensively analyze the asymmetric, dynamic, synchronous, and detailed nature of human-human interactions and propose the first multi-setting human action-reaction synthe-sis benchmark to generate human reactions conditioned on given human actions. To begin with, we propose to an-notate the actor-reactor order of the interaction sequences for the NTU120, InterHuman, and Chi3D datasets. Based on them, a diffusion-based generative model with a Trans-former decoder architecture called ReGenNet together with an explicit distance-based interaction loss is proposed to predict human reactions in an online manner, where the future states of actors are unavailable to reactors. Quantitative and qualitative results show that our method can gener-ate instant and plausible human reactions compared to the baselines, and can generalize to unseen actor motions and viewpoint changes.
Liang Xu 0012, Yizhou Zhou, Yichao Yan, Xin Jin 0014, Wenhan Zhu, Fengyun Rao, Xiaokang Yang 0001, Wenjun Zeng 0001
CVPR3
2024 Topo4D: Topology-Preserving Gaussian Splatting for High-fidelity 4D Head Capture
Xuanchen Li, Yuhao Cheng, Xingyu Ren, Haozhe Jia, Di Xu 0012, Wenhan Zhu, Yichao Yan
ECCV (34)7
2024 HIMO: A New Benchmark for Full-Body Human Interacting with Multiple Objects
Xintao Lv, Liang Xu 0012, Yichao Yan, Xin Jin 0014, Congsheng Xu, Shuwen Wu, Lincheng Li, Mengxiao Bi, Wenjun Zeng 0001, Xiaokang Yang 0001
ECCV (4)3
2024 MTA-PS: Towards Practical Person Search in Videos
abstract
Person search (PS) aims to simultaneously localize and identify a target person from natural, uncropped images. Existing PS datasets and research works are mostly based on individual images, exhibiting limited practicability in real-world surveillance scenarios. We contend that videos, compared to static images, offer additional temporal information, making searching for the trajectory of the target person from videos more realistic and accurate. In this paper, we propose a new practical and realistic task, namely person search in videos, and a new evaluation metric specifically tailored for it. To fulfill this, we introduce a new PS dataset, namely MTAPS, based on an existing large-scale simulated video dataset. MTA-PS is the first cross-camera PS dataset in virtual videos, consisting of 6 cameras, 60 videos, 1.8K identities, 295.2K frames, 7.3M bounding boxes, and more than 20 minutes per camera, which is challenging and comprehensive, and meanwhile avoids privacy issues. To validate the effectiveness of PS in videos and make full use of the temporal information on our dataset, we also propose a novel framework by seamlessly integrating the three sub-tasks of person detection, tracking, and re-identification. Extensive experiments demonstrate that our method performs favorably over existing counterparts on the newly-introduced MTA-PS dataset. Codes and datasets are available at https://github.com/mtmyyy/MTA-PS.
Tiancheng Ying, Rong Quan, Peng Zheng 0004, Yichao Yan, Jie Qin 0004
ICIP4
2024 A Comparative Study of Perceptual Quality Metrics For Audio-Driven Talking Head Videos
abstract
The rapid advancement of Artificial Intelligence Generated Content (AIGC) technology has propelled audio-driven talking head generation, gaining considerable research attention for practical applications. However, performance evaluation research lags behind the development of talking head generation techniques. Existing literature relies on heuristic quantitative metrics without human validation, hindering accurate progress assessment. To address this gap, we collect talking head videos generated from four generative methods and conduct controlled psychophysical experiments on visual quality, lip-audio synchronization, and head movement naturalness. Our experiments validate consistency between model predictions and human annotations, identifying metrics that align better with human opinions than widely-used measures. We believe our work will facilitate performance evaluation and model development, providing insights into AIGC in a broader context. Code is available at https://github.com/zwx8981/ADTH-QA.
Weixia Zhang, Chengguang Zhu, Jingnan Gao, Yichao Yan, Guangtao Zhai, Xiaokang Yang 0001
ICIP4
2024 HQ-Avatar: Towards High-Quality 3D Avatar Generation via Point-based Representation
abstract
Despite the flourishing of 3D object generation, generating high-quality digital avatars with detailed geometry and texture that are free to animate remains a challenging task. Existing avatar generation techniques often suffer from limitations such as low-quality geometry and blurry texture. Thus, we propose HQ-Avatar, a novel method for generating animatable avatars with high-quality geometry and texture. We enhance the geometry quality by proposing an importance sampling strategy and capturing intricate details through learned normal maps. To achieve high-quality texture, we present a neural point-based avatar representation, which enables high-resolution rendering results at a resolution of 10242, allowing detailed supervision. Extensive experiments on THuman2.0 dataset demonstrate the superiority of our method over state-of-the-art techniques in generating high-quality avatars. Furthermore, we show the applicability of our method by employing it as 3D priors to simplify the human avatar reconstruction process from scans or even single images. Code is available at https://github.com/olivia23333/HQ-Avatar
Weitian Zhang, Sijing Wu, Yichao Yan, Ben Xue, Wenhan Zhu, Xiaokang Yang 0001
ICME3
2024 MMHead: Towards Fine-grained Multi-modal 3D Facial Animation
abstract
3D facial animation has attracted considerable attention due to its extensive applications in the multimedia field. Audio-driven 3D facial animation has been widely explored with promising results. However, multi-modal 3D facial animation, especially text-guided 3D facial animation is rarely explored due to the lack of multi-modal 3D facial animation dataset. To fill this gap, we first construct a large-scale multi-modal 3D facial animation dataset, MMHead, which consists of 49 hours of 3D facial motion sequences, speech audios, and rich hierarchical text annotations. Each text annotation contains abstract action and emotion descriptions, fine-grained facial and head movements (i.e., expression and head pose) descriptions, and three possible scenarios that may cause such emotion. Concretely, we integrate five public 2D portrait video datasets, and propose an automatic pipeline to 1) reconstruct 3D facial motion sequences from monocular videos; and 2) obtain hierarchical text annotations with the help of AU detection and ChatGPT. Based on the MMHead dataset, we establish benchmarks for two new tasks: text-induced 3D talking head animation and text-to-3D facial motion generation. Moreover, a simple but efficient VQ-VAE-based method named MM2Face is proposed to unify the multi-modal information and generate diverse and plausible 3D facial motions, which achieves competitive results on both benchmarks. Extensive experiments and comprehensive analysis demonstrate the significant potential of our dataset and benchmarks in promoting the development of multi-modal 3D facial animation. The dataset will be released at: https://wsj-sjtu.github.io/MMHead/.
Sijing Wu, Yichao Yan, Huiyu Duan, Ziwei Liu 0002, Guangtao Zhai
ACM Multimedia3
2024 Infusion: Preventing Customized Text-to-Image Diffusion from Overfitting
abstract
Text-to-image (T2I) customization aims to create images that embody specific visual concepts delineated in textual descriptions. However, existing works still face a main challenge, concept overfitting. To tackle this challenge, we first analyze overfitting, categorizing it into concept-agnostic overfitting, which undermines non-customized concept knowledge, and concept-specific overfitting, which is confined to customize on limited diversities, i.e, backgrounds, layouts, styles. To evaluate the overfitting degree, we further introduce two metrics, i.e, Latent Fisher divergence and Wasserstein metric to measure the distribution changes of non-customized and customized concept respectively. Drawing from the analysis, we propose Infusion, a T2I customization method that enables the learning of target concepts to avoid being constrained by limited training diversities, while preserving non-customized knowledge. Remarkably, Infusion achieves this feat with remarkable efficiency, requiring a mere 11KB of trained parameters. Extensive experiments also demonstrate that our approach outperforms state-of-the-art methods in both single and multi-concept customized generation. Project page: https://zwl666666.github.io/infusion/.
Yichao Yan, Zhuo Chen 0060, Pengzhi Chu, Weiming Zhao, Xiaokang Yang 0001
ACM Multimedia2
2024 E3Gen: Efficient, Expressive and Editable Avatars Generation
Weitian Zhang, Yichao Yan, Yunhui Liu 0006, Xingdong Sheng, Xiaokang Yang 0001
ACM Multimedia2
2024 Multi-times Monte Carlo Rendering for Inter-reflection Reconstruction
abstract
Inverse rendering methods have achieved remarkable performance in reconstructing high-fidelity 3D objects with disentangled geometries, materials, and environmental light. However, they still face huge challenges in reflective surface reconstruction. Although recent methods model the light trace to learn specularity, the ignorance of indirect illumination makes it hard to handle inter-reflections among multiple smooth objects. In this work, we propose Ref-MC2 that introduces the multi-time Monte Carlo sampling which comprehensively computes the environmental illumination and meanwhile considers the reflective light from object surfaces. To address the computation challenge as the times of Monte Carlo sampling grow, we propose a specularity-adaptive sampling strategy, significantly reducing the computational complexity. Besides the computational resource, higher geometry accuracy is also required because geometric errors accumulate multiple times. Therefore, we further introduce a reflection-aware surface model to initialize the geometry and refine it during inverse rendering. We construct a challenging dataset containing scenes with multiple objects and inter-reflections. Experiments show that our method outperforms other inverse rendering methods on various object groups. We also show downstream applications, e.g., relighting and material editing, to illustrate the disentanglement ability of our method.
Tengjie Zhu, Zhuo Chen 0060, Jingnan Gao, Yichao Yan, Xiaokang Yang 0001
NeurIPS4
2024 Directional Texture Editing for 3D Models
abstract
Abstract Texture editing is a crucial task in 3D modelling that allows users to automatically manipulate the surface materials of 3D models. However, the inherent complexity of 3D models and the ambiguous text description lead to the challenge of this task. To tackle this challenge, we propose ITEM3D, a Texture Editing Model designed for automatic 3D object editing according to the text Instructions. Leveraging the diffusion models and the differentiable rendering, ITEM3D takes the rendered images as the bridge between text and 3D representation and further optimizes the disentangled texture and environment map. Previous methods adopted the absolute editing direction, namely score distillation sampling (SDS) as the optimization objective, which unfortunately results in noisy appearances and text inconsistencies. To solve the problem caused by the ambiguous text, we introduce a relative editing direction, an optimization objective defined by the noise difference between the source and target texts, to release the semantic ambiguity between the texts and images. Additionally, we gradually adjust the direction during optimization to further address the unexpected deviation in the texture domain. Qualitative and quantitative experiments show that our ITEM3D outperforms the state‐of‐the‐art methods on various 3D objects. We also perform text‐guided relighting to show explicit control over lighting. Our project page: https://shengqiliu1.github.io/ITEM3D/ .
Shengqi Liu, Zhuo Chen 0060, Jingnan Gao, Yichao Yan, Wenhan Zhu, Jiangjing Lyu, Xiaokang Yang 0001
Comput. Graph. Forum4
2024 A Coding Framework and Benchmark Towards Low-Bitrate Video Understanding
abstract
Video compression is indispensable to most video analysis systems. Despite saving the transportation bandwidth, it also deteriorates downstream video understanding tasks, especially at low-bitrate settings. To systematically investigate this problem, we first thoroughly review the previous methods, revealing that three principles, i.e., task-decoupled, label-free, and data-emerged semantic prior, are critical to a machine-friendly coding framework but are not fully satisfied so far. In this paper, we propose a traditional-neural mixed coding framework that simultaneously fulfills all these principles, by taking advantage of both traditional codecs and neural networks (NNs). On one hand, the traditional codecs can efficiently encode the pixel signal of videos but may distort the semantic information. On the other hand, highly non-linear NNs are proficient in condensing video semantics into a compact representation. The framework is optimized by ensuring that a transportation-efficient semantic representation of the video is preserved w.r.t. the coding procedure, which is spontaneously learned from unlabeled data in a self-supervised manner. The videos collaboratively decoded from two streams (codec and NN) are of rich semantics, as well as visually photo-realistic, empirically boosting several mainstream downstream video analysis task performances without any post-adaptation procedure. Furthermore, by introducing the attention mechanism and adaptive modeling scheme, the video semantic modeling ability of our approach is further enhanced. Fianlly, we build a low-bitrate video understanding benchmark with three downstream tasks on eight datasets, demonstrating the notable superiority of our approach. All codes, data, and models will be open-sourced for facilitating future research.
Yuan Tian 0017, Guo Lu, Yichao Yan, Guangtao Zhai, Li Chen 0021
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 HyperStyle3D: Text-Guided 3D Portrait Stylization via Hypernetworks
abstract
Portrait stylization is a long-standing task enabling extensive applications. Although 2D-based methods have made great progress in recent years, real-world applications such as metaverse and games often demand 3D content. On the other hand, the requirement of 3D data, which is costly to acquire, significantly impedes the development of 3D portrait stylization methods. In this paper, inspired by the success of 3D-aware GANs that bridge 2D and 3D domains with 3D fields as the intermediate representation for rendering 2D images, we propose a novel method, dubbed HyperStyle3D, based on 3D-aware GANs for 3D portrait stylization. At the core of our method is a hyper-network learned to manipulate the parameters of the generator in a single forward pass. It not only offers a strong capacity to handle multiple styles with a single model, but also enables flexible fine-grained stylization that affects only texture, shape, or local part of the portrait. While the use of 3D-aware GANs bypasses the requirement of 3D data, we further alleviate the necessity of style images with the CLIP model being the style guidance. We conduct an extensive set of experiments across the style, attribute, and shape, and meanwhile, measure the 3D consistency. These experiments demonstrate the superior capability of our HyperStyle3D model in rendering 3D-consistent images in diverse styles, deforming the face shape, and editing various attributes.
Zhuo Chen 0060, Xudong Xu, Yichao Yan, Wenhan Zhu, Wayne Wu, Bo Dai 0002, Xiaokang Yang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2024 Head3D: Complete 3D Head Generation via Tri-plane Feature Distillation
abstract
Head generation with diverse identities is an important task in computer vision and computer graphics, widely used in multimedia applications. However, current full-head generation methods require a large number of three-dimensional (3D) scans or multi-view images to train the model, resulting in expensive data acquisition costs. To address this issue, we propose Head3D, a method to generate full 3D heads with limited multi-view images. Specifically, our approach first extracts facial priors represented by tri-planes learned in EG3D, a 3D-aware generative model, and then proposes feature distillation to deliver the 3D frontal faces within complete heads without compromising head integrity. To mitigate the domain gap between the face and head models, we present a dual-discriminator to guide the frontal and back head generation. Our model achieves cost-efficient and diverse complete head generation with photo-realistic renderings and high-quality geometry representations. Extensive experiments demonstrate the effectiveness of our proposed Head3D, both qualitatively and quantitatively.
Yuhao Cheng, Yichao Yan, Wenhan Zhu, Bowen Pan, Xiaokang Yang 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2024 StyleVR: Stylizing Character Animations With Normalizing Flows
abstract
The significance of artistry in creating animated virtual characters is widely acknowledged, and motion style is a crucial element in this process. There has been a long-standing interest in stylizing character animations with style transfer methods. However, this kind of models can only deal with short-term motions and yield deterministic outputs. To address this issue, we propose a generative model based on normalizing flows for stylizing long and aperiodic animations in the VR scene. Our approach breaks down this task into two sub-problems: motion style transfer and stylized motion generation, both formulated as the instances of conditional normalizing flows with multi-class latent space. Specifically, we encode high-frequency style features into the latent space for varied results and control the generation process with style-content labels for disentangled edits of style and content. We have developed a prototype, StyleVR, in Unity, which allows casual users to apply our method in VR. Through qualitative and quantitative comparisons, we demonstrate that our system outperforms other methods in terms of style transfer as well as stochastic stylized motion generation.
Bin Ji 0004, Yichao Yan, Ruizhao Chen, Xiaokang Yang 0001
IEEE Trans. Vis. Comput. Graph.3
2023 3D-Aware Face Swapping
abstract
Face swapping is an important research topic in computer vision with wide applications in entertainment and privacy protection. Existing methods directly learn to swap 2D facial images, taking no account of the geometric information of human faces. In the presence of large pose variance between the source and the target faces, there always exist undesirable artifacts on the swapped face. In this paper, we present a novel 3D-aware face swapping method that generates high-fidelity and multi-view-consistent swapped faces from single-view source and target images. To achieve this, we take advantage of the strong geometry and texture prior of 3D human faces, where the 2D faces are projected into the latent space of a 3D generative model. By disentangling the identity and attribute features in the latent space, we succeed in swapping faces in a 3D-aware manner, being robust to pose variations while transferring fine-grained facial details. Extensive experiments demonstrate the superiority of our 3D-aware face swapping framework in terms of visual quality, identity similarity, and multi-view consistency. Code is available at https://lyx0208.github.io/3dSwap.
Chao Ma 0004, Yichao Yan, Wenhan Zhu, Xiaokang Yang 0001
CVPR3
2023 Improving Fairness in Facial Albedo Estimation via Visual-Textual Cues
abstract
Recent 3D face reconstruction methods have made significant advances in geometry prediction, yet further cosmetic improvements are limited by lagged albedo because inferring albedo from appearance is an ill-posed problem. Although some existing methods consider prior knowledge from illumination to improve albedo estimation, they still produce a light-skin bias due to racially biased albedo models and limited light constraints. In this paper, we reconsider the relationship between albedo and face attributes and propose a ID2Albedo to directly estimate albedo without constraining illumination. Our key insight is that intrinsic semantic attributes such as race, skin color, and age can be used to constrain the albedo map. We first introduce visual-textual cues and design a semantic loss to supervise facial albedo estimation. Specifically, we pre-define text labels such as race, skin color, age, and wrinkles. Then, we employ the text-image model (CLIP) to compute the similarity between the text and the input image, and assign a pseudo-label to each facial image. We constrain generated albedos in the training phase to have the same attributes as the inputs. In addition, we train a high-quality, unbiased facial albedo generator and utilize the semantic loss to learn the mapping from illumination-robust identity features to the albedo latent codes. Finally, our ID2Albedo is trained in a self-supervised way and outperforms state-of-the-art albedo estimation methods in terms of accuracy and fidelity. It is worth mentioning that our approach has excellent generalizability and fairness, especially on in-the- wild data.
Xingyu Ren, Jiankang Deng, Chao Ma 0004, Yichao Yan, Xiaokang Yang 0001
CVPR4
2023 GANHead: Towards Generative Animatable Neural Head Avatars
abstract
To bring digital avatars into people's lives, it is highly demanded to efficiently generate complete, realistic, and animatable head avatars. This task is challenging, and it is difficult for existing methods to satisfy all the requirements at once. To achieve these goals, we propose GANHead (Generative Animatable Neural Head Avatar), a novel generative head model that takes advantages of both the fine-grained control over the explicit expression parameters and the realistic rendering results of implicit representations. Specifically, GANHead represents coarse geometry, fine-gained details and texture via three networks in canonical space to obtain the ability to generate complete and realistic head avatars. To achieve flexible animation, we define the deformation filed by standard linear blend skinning (LBS), with the learned continuous pose and expression bases and LBS weights. This allows the avatars to be directly animated by FLAME [22] parameters and generalize well to unseen poses and expressions. Compared to state-of-the-art (SOTA) methods, GANHead achieves superior performance on head avatar generation and raw scan fitting.
Sijing Wu, Yichao Yan, Yuhao Cheng, Wenhan Zhu, Ke Gao 0012, Guangtao Zhai
CVPR2
2023 Movienet-PS: A Large-Scale Person Search Dataset in the Wild
abstract
Person search (PS) aims to jointly localize and identify a query person from natural, uncropped images. Existing works unintentionally adopt pedestrians (with similar poses and unchanging clothing) as the query and restrict the application scenarios in surveillance. This is due to that most PS datasets are collected from surveillance cameras with a limited diversity of views, scenes, appearances, etc. In this paper, we study a more general and realistic task in the wild, where we aim to search target persons with a much higher degree of diversity. To this end, we introduce a new PS dataset, namely MovieNet-PS, based on an existing large-scale movie dataset. MovieNet-PS is currently the largest and most diverse PS dataset, consisting of 160K images (100K for training), 274K bounding boxes, and 3K identities. It stands out from existing counterparts from two levels of diversities, i.e., scene-level and identity-level, with 92,043 scenes and significant variations in poses, clothing, scales, etc. for the same identity. To validate the rich context information on our dataset and make full use of it, we propose a novel global-local context network, which exploits scene and group context to boost the search performance. Extensive experiments demonstrate that MovieNet-PS is more challenging and comprehensive than existing datasets, and our approach further pushes the state of the art by a large margin (relatively 34% in mAP) on this dataset. Codes, models, and the dataset are available at: https://github.com/ZhengPeng7/GLCNet.
Jie Qin 0004, Peng Zheng 0004, Yichao Yan, Rong Quan, Xiaogang Cheng, Bingbing Ni
ICASSP3
2023 ActFormer: A GAN-based Transformer towards General Action-Conditioned 3D Human Motion Generation
abstract
We present a GAN-based Transformer for general action-conditioned 3D human motion generation, including not only single-person actions but also multi-person interactive actions. Our approach consists of a powerful Action-conditioned motion TransFormer (ActFormer) under a GAN training scheme, equipped with a Gaussian Process latent prior. Such a design combines the strong spatio-temporal representation capacity of Transformer, superiority in generative modeling of GAN, and inherent temporal correlations from the latent prior. Furthermore, ActFormer can be naturally extended to multi-person motions by alternately modeling temporal correlations and human interactions with Transformer encoders. To further facilitate research on multi-person motion generation, we introduce a new synthetic dataset of complex multi-person combat behaviors. Extensive experiments on NTU-13, NTU RGB+D 120, BABEL and the proposed combat dataset show that our method can adapt to various human motion representations and achieve superior performance over the state-of-the-art methods on both single-person and multi-person motion generation tasks, demonstrating a promising step towards a general human motion generator. The project website can be found at https://liangxuy.github.io/actformer/.
Liang Xu 0012, Jing Su 0005, Zhicheng Fang, Chenjing Ding, Weihao Gan, Yichao Yan, Xin Jin 0014, Xiaokang Yang 0001, Wenjun Zeng 0001, Wei Wu 0021
ICCV8
2023 NeRF-IBVS: Visual Servo Based on NeRF for Visual Localization and Navigation
abstract
Visual localization is a fundamental task in computer vision and robotics. Training existing visual localization methods requires a large number of posed images to generalize to novel views, while state-of-the-art methods generally require dense ground truth 3D labels for supervision. However, acquiring a large number of posed images and dense 3D labels in the real world is challenging and costly. In this paper, we present a novel visual localization method that achieves accurate localization while using only a few posed images compared to other localization methods. To achieve this, we first use a few posed images with coarse pseudo-3D labels provided by NeRF to train a coordinate regression network. Then a coarse pose is estimated from the regression network with PNP. Finally, we use the image-based visual servo (IBVS) with the scene prior provided by NeRF for pose optimization. Furthermore, our method can provide effective navigation prior, which enable navigation based on IBVS without using custom markers and depth sensor. Extensive experiments on 7-Scenes and 12-Scenes datasets demonstrate that our method outperforms state-of-the-art methods under the same setting, with only 5\% to 25\% training data. Furthermore, our framework can be naturally extended to the visual navigation task based on IBVS, and its effectiveness is verified in simulation experiments.
Yuanze Wang, Yichao Yan, Dian-xi Shi, Wenhan Zhu, Jianqiang Xia, Jeff Tan, Songchang Jin, Ke Gao 0012, Xiaokang Yang 0001
NeurIPS2
2023 Efficient Person Search: An Anchor-Free Approach
Yichao Yan, Jinpeng Li 0004, Jie Qin 0004, Peng Zheng 0004, Shengcai Liao, Xiaokang Yang 0001
Int. J. Comput. Vis.1
2023 Learning Multi-Attention Context Graph for Group-Based Re-Identification
abstract
Learning to re-identify or retrieve a group of people across non-overlapped camera systems has important applications in video surveillance. However, most existing methods focus on (single) person re-identification (re-id), ignoring the fact that people often walk in groups in real scenarios. In this work, we take a step further and consider employing context information for identifying groups of people, i.e., group re-id. On the one hand, group re-id is more challenging than single person re-id, since it requires both a robust modeling of local individual person appearance (with different illumination conditions, pose/viewpoint variations, and occlusions), as well as full awareness of global group structures (with group layout and group member variations). On the other hand, we believe that person re-id can be greatly enhanced by incorporating additional visual context from neighboring group members, a task which we formulate as group-aware (single) person re-id. In this paper, we propose a novel unified framework based on graph neural networks to simultaneously address the above two group-based re-id tasks, i.e., group re-id and group-aware person re-id. Specifically, we construct a context graph with group members as its nodes to exploit dependencies among different people. A multi-level attention mechanism is developed to formulate both intra-group and inter-group context, with an additional self-attention module for robust graph-level representations by attentively aggregating node-level features. The proposed model can be directly generalized to tackle group-aware person re-id using node-level representations. Meanwhile, to facilitate the deployment of deep learning models on these tasks, we build a new group re-id dataset which contains more than 3.8K images with 1.5K annotated groups, an order of magnitude larger than existing group re-id datasets. Extensive experiments on the novel dataset as well as three existing datasets clearly demonstrate the effectiveness of the proposed framework for both group-based re-id tasks.
Yichao Yan, Jie Qin 0004, Bingbing Ni, Jiaxin Chen 0002, Li Liu 0004, Fan Zhu 0001, Wei-Shi Zheng 0001, Xiaokang Yang 0001, Ling Shao 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 CLSA: A Contrastive Learning Framework With Selective Aggregation for Video Rescaling
abstract
Video rescaling has recently drawn extensive attention for its practical applications such as video compression. Compared to video super-resolution, which focuses on upscaling bicubic-downscaled videos, video rescaling methods jointly optimize a downscaler and a upscaler. However, the inevitable loss of information during downscaling makes the upscaling procedure still ill-posed. Furthermore, the network architecture of previous methods mostly relies on convolution to aggregate information within local regions, which cannot effectively capture the relationship between distant locations. To address the above two issues, we propose a unified video rescaling framework by introducing the following designs. First, we propose to regularize the information of the downscaled videos via a contrastive learning framework, where, particularly, hard negative samples for learning are synthesized online. With this auxiliary contrastive learning objective, the downscaler tends to retain more information that benefits the upscaler. Second, we present a selective global aggregation module (SGAM) to efficiently capture long-range redundancy in high-resolution videos, where only a few representative locations are adaptively selected to participate in the computationally-heavy self-attention (SA) operations. SGAM enjoys the efficiency of the sparse modeling scheme while preserving the global modeling capability of SA. We refer to the proposed framework as Contrastive Learning framework with Selective Aggregation (CLSA) for video rescaling. Comprehensive experimental results show that CLSA outperforms video rescaling and rescaling-based video compression methods on five datasets, achieving state-of-the-art performance.
Yuan Tian 0017, Yichao Yan, Guangtao Zhai, Li Chen 0021
IEEE Trans. Image Process.2
2022 Exploring Visual Context for Weakly Supervised Person Search
abstract
Person search has recently emerged as a challenging task that jointly addresses pedestrian detection and person re-identification. Existing approaches follow a fully supervised setting where both bounding box and identity annotations are available. However, annotating identities is labor-intensive, limiting the practicability and scalability of current frameworks. This paper inventively considers weakly supervised person search with only bounding box annotations. We propose to address this novel task by investigating three levels of context clues (i.e., detection, memory and scene) in unconstrained natural images. The first two are employed to promote local and global discriminative capabilities, while the latter enhances clustering accuracy. Despite its simple design, our CGPS boosts the baseline model by 8.8% in mAP on CUHK-SYSU. Surprisingly, it even achieves comparable performance with several supervised person search models. Our code is available at https://github. com/ljpadam/CGPS.
Yichao Yan, Jinpeng Li 0004, Shengcai Liao, Jie Qin 0004, Bingbing Ni, Ke Lu 0002, Xiaokang Yang 0001
AAAI1
2022 Domain Adaptive Person Search
Yichao Yan, Guanshuo Wang, Fufu Yu, Qiong Jia 0004, Shouhong Ding
ECCV (14)2
2022 CageNeRF: Cage-based Neural Radiance Field for Generalized 3D Deformation and Animation
abstract
While implicit representations have achieved high-fidelity results in 3D rendering, it remains challenging to deforming and animating the implicit field. Existing works typically leverage data-dependent models as deformation priors, such as SMPL for human body animation. However, this dependency on category-specific priors limits them to generalize to other objects. To solve this problem, we propose a novel framework for deforming and animating the neural radiance field learned on \textit{arbitrary} objects. The key insight is that we introduce a cage-based representation as deformation prior, which is category-agnostic. Specifically, the deformation is performed based on an enclosing polygon mesh with sparsely defined vertices called \textit{cage} inside the rendering space, where each point is projected into a novel position based on the barycentric interpolation of the deformed cage vertices. In this way, we transform the cage into a generalized constraint, which is able to deform and animate arbitrary target objects while preserving geometry details. Based on extensive experiments, we demonstrate the effectiveness of our framework in the task of geometry editing, object animation and deformation transfer.
Yicong Peng, Yichao Yan, Shengqi Liu, Yuhao Cheng, Shanyan Guan, Bowen Pan, Guangtao Zhai, Xiaokang Yang 0001
NeurIPS2
2022 EAN: Event Adaptive Network for Enhanced Action Recognition
Yuan Tian 0017, Yichao Yan, Guangtao Zhai, Guodong Guo
Int. J. Comput. Vis.2
2022 Fine-Grained Video Captioning via Graph-based Multi-Granularity Interaction Learning
abstract
Learning to generate continuous linguistic descriptions for multi-subject interactive videos in great details has particular applications in team sports auto-narrative. In contrast to traditional video caption, this task is more challenging as it requires simultaneous modeling of fine-grained individual actions, uncovering of spatio-temporal dependency structures of frequent group interactions, and then accurate mapping of these complex interaction details into long and detailed commentary. To explicitly address these challenges, we propose a novel framework Graph-based Learning for Multi-Granularity Interaction Representation (GLMGIR) for fine-grained team sports auto-narrative task. A multi-granular interaction modeling module is proposed to extract among-subjects' interactive actions in a progressive way for encoding both intra- and inter-team interactions. Based on the above multi-granular representations, a multi-granular attention module is developed to consider action/event descriptions of multiple spatio-temporal resolutions. Both modules are integrated seamlessly and work in a collaborative way to generate the final narrative. In the meantime, to facilitate reproducible research, we collect a new video dataset from YouTube.com called Sports Video Narrative dataset (SVN). It is a novel direction as it contains 6K team sports videos (i.e., NBA basketball games) with 10K ground-truth narratives(e.g., sentences). Furthermore, as previous metrics such as METEOR (i.e., used in coarse-grained video caption task) DO NOT cope with fine-grained sports narrative task well, we hence develop a novel evaluation metric named Fine-grained Captioning Evaluation (FCE), which measures how accurate the generated linguistic description reflects fine-grained action details as well as the overall spatio-temporal interactional structure. Extensive experiments on our SVN dataset have demonstrated the effectiveness of the proposed framework for fine-grained team sports video auto-narrative.
Yichao Yan, Ning Zhuang, Bingbing Ni, Jian Zhang 0079, Qi Tian 0001, Yi Xu 0001, Xiaokang Yang 0001, Wenjun Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 Anchor-Free Person Search
abstract
Person search aims to simultaneously localize and identify a query person from realistic, uncropped images, which can be regarded as the unified task of pedestrian detection and person re-identification (re-id). Most existing works employ two-stage detectors like Faster-RCNN, yielding encouraging accuracy but with high computational overhead. In this work, we present the Feature-Aligned Person Search Network (AlignPS), the first anchor-free framework to efficiently tackle this challenging task. AlignPS explicitly addresses the major challenges, which we summarize as the misalignment issues in different levels (i.e., scale, region, and task), when accommodating an anchor-free detector for this task. More specifically, we propose an aligned feature aggregation module to generate more discriminative and robust feature embeddings by following a "re-id first" principle. Such a simple design directly improves the baseline anchor-free model on CUHK-SYSU by more than 20% in mAP. Moreover, AlignPS outperforms state-of-the-art two-stage methods, with a higher speed. The code is available at https://github.com/daodaofr/AlignPS.
Yichao Yan, Jinpeng Li 0004, Jie Qin 0004, Song Bai 0001, Shengcai Liao, Li Liu 0004, Fan Zhu 0001, Ling Shao 0001
CVPR1
2020 Learning Multi-Granular Hypergraphs for Video-Based Person Re-Identification
abstract
Video-based person re-identification (re-ID) is an important research topic in computer vision. The key to tackling the challenging task is to exploit both spatial and temporal clues in video sequences. In this work, we propose a novel graph-based framework, namely Multi-Granular Hypergraph (MGH), to pursue better representational capabilities by modeling spatiotemporal dependencies in terms of multiple granularities. Specifically, hypergraphs with different spatial granularities are constructed using various levels of part-based features across the video sequence. In each hypergraph, different temporal granularities are captured by hyperedges that connect a set of graph nodes (i.e., part-based features) across different temporal ranges. Two critical issues (misalignment and occlusion) are explicitly addressed by the proposed hypergraph propagation and feature aggregation schemes. Finally, we further enhance the overall video representation by learning more diversified graph-level representations of multiple granularities based on mutual information minimization. Extensive experiments on three widely-adopted benchmarks clearly demonstrate the effectiveness of the proposed framework. Notably, 90.0% top-1 accuracy on MARS is achieved using MGH, outperforming the state-of-the-arts.
Yichao Yan, Jie Qin 0004, Jiaxin Chen 0002, Li Liu 0004, Fan Zhu 0001, Ying Tai, Ling Shao 0001
CVPR1
2020 Deep Local Binary Coding for Person Re-Identification by Delving into the Details
abstract
Person re-identification (ReID) has recently received extensive research interests due to its diverse applications in multimedia analysis and computer vision. However, the majority of existing works focus on improving matching accuracy, while ignoring matching efficiency. In this work, we present a novel binary representation learning framework for efficient person ReID, namely Deep Local Binary Coding (DLBC). Different from existing deep binary ReID approaches, DLBC attempts to learn discriminative binary codes by explicitly interacting with local visual details. Specifically, DLBC first extracts a set of local features from spatially salient regions of pedestrian images. Subsequently, DLBC formulates a new binary-local semantic mutual information (BSMI) maximization term, based on which a self-lifting (SL) block is built to further exploit the semantic importance of local features. The BSMI term together with the SL block simultaneously enhances the dependency of binary codes on selected local features as well as their robustness to cross-view visual inconsistency. In addition, an efficient optimizing method is developed to train the proposed deep models with orthogonal and binary constraints. Extensive experiments reveal that DLBC significantly minimizes the accuracy gap between binary ReID methods and the state-of-the-art real-valued ones, whilst remarkably reducing query time and memory cost.
Jiaxin Chen 0002, Jie Qin 0004, Yichao Yan, Lei Huang 0015, Li Liu 0004, Fan Zhu 0001, Ling Shao 0001
ACM Multimedia3
2020 Fine-grained image analysis via progressive feature learning
Yichao Yan, Bingbing Ni, Huawei Wei, Xiaokang Yang 0001
Neurocomputing1
2019 Learning Context Graph for Person Search
abstract
Person re-identification has achieved great progress with deep convolutional neural networks. However, most previous methods focus on learning individual appearance feature embedding, and it is hard for the models to handle difficult situations with different illumination, large pose variance and occlusion. In this work, we take a step further and consider employing context information for person search. For a probe-gallery pair, we first propose a contextual instance expansion module, which employs a relative attention module to search and filter useful context information in the scene. We also build a graph learning framework to effectively employ context pairs to update target similarity. These two modules are built on top of a joint detection and instance feature learning framework, which improves the discriminativeness of the learned features. The proposed framework achieves state-of-the-art performance on two widely used person search datasets.
Yichao Yan, Bingbing Ni, Wendong Zhang 0002, Xiaokang Yang 0001
CVPR1
2019 Cross-modality motion parameterization for fine-grained video prediction
Yichao Yan, Bingbing Ni, Wendong Zhang 0002, Xiaokang Yang 0001
Comput. Vis. Image Underst.1
2019 Multi-level attention model for person re-identification
Yichao Yan, Bingbing Ni, Jinxian Liu, Xiaokang Yang 0001
Pattern Recognit. Lett.1
2019 Structure-Constrained Motion Sequence Generation
abstract
Video generation is a challenging task due to the extremely high-dimensional distribution of the solution space. Good constraints in the solution domain would thus reduce the difficulty of approximating optimal solutions. In this paper, instead of directly generating high-dimensional video data, we propose using object landmarks as explicit structure constraints to address this issue. Specifically, we propose a two-stage framework for an action-conditioned video generation task. In our framework, the first stage aims to generate landmark sequences according to predefined motion types, and a recurrent model (RNN/LSTM) is adopted for this purpose. The landmark sequence can be regarded as a low-dimensional structure embedding of high-dimensional video data, and generating landmark sequences is much easier than generating videos. The second stage is inspired by a conditional generative adversarial network (CGAN), and we take the generated landmark sequence as a structure condition to learn a landmark-to-image translation network. Such a one-to-one translation framework avoids the difficulty of generating videos and instead transfers the video generation task to image generation, which is resolvable due to the maturity of current GAN-based models. The experimental results demonstrate that our model not only achieves promising results on rigid/nonrigid motion generation tasks but also can be extended to multiobject motion situations.
Yichao Yan, Bingbing Ni, Wendong Zhang 0002, Jingwei Xu 0005, Xiaokang Yang 0001
IEEE Trans. Multim.1
2018 Video Summarization via Semantic Attended Networks
abstract
The goal of video summarization is to distill a raw video into a more compact form without losing much semantic information. However, previous methods mainly consider the diversity and representation interestingness of the obtained summary, and they seldom pay sufficient attention to semantic information of resulting frame set, especially the long temporal range semantics. To explicitly address this issue, we propose a novel technique which is able to extract the most semantically relevant video segments (i.e., valid for a long term temporal duration) and assemble them into an informative summary. To this end, we develop a semantic attended video summarization network (SASUM) which consists of a frame selector and video descriptor to select an appropriate number of video shots by minimizing the distance between the generated description sentence of the summarized video and the human annotated text of the original video. Extensive experiments show that our method achieves a superior performance gain over previous methods on two benchmark datasets.
Huawei Wei, Bingbing Ni, Yichao Yan, Huanyu Yu, Xiaokang Yang 0001
AAAI3
2018 Pose Transferrable Person Re-Identification
abstract
Person re-identification (ReID) is an important task in the field of intelligent security. A key challenge is how to capture human pose variations, while existing benchmarks (i.e., Market1501, DukeMTMC-reID, CUHK03, etc.) do NOT provide sufficient pose coverage to train a robust ReID system. To address this issue, we propose a pose-transferrable person ReID framework which utilizes pose-transferred sample augmentations (i.e., with ID supervision) to enhance ReID model training. On one hand, novel training samples with rich pose variations are generated via transferring pose instances from MARS dataset, and they are added into the target dataset to facilitate robust training. On the other hand, in addition to the conventional discriminator of GAN (i.e., to distinguish between REAL/FAKE samples), we propose a novel guider sub-network which encourages the generated sample (i.e., with novel pose) towards better satisfying the ReID loss (i.e., cross-entropy ReID loss, triplet ReID loss). In the meantime, an alternative optimization procedure is proposed to train the proposed Generator-Guider-Discriminator network. Experimental results on Market-1501, DukeMTMC-reID and CUHK03 show that our method achieves great performance improvement, and outperforms most state-of-the-art methods without elaborate designing the ReID model.
Jinxian Liu, Bingbing Ni, Yichao Yan, Peng Zhou 0010, Jianguo Hu
CVPR3
2018 Depth Structure Preserving Scene Image Generation
abstract
Key to automatically generate natural scene images is to properly arrange amongst various spatial elements, especially in the depth cue. To this end, we introduce a novel depth structure preserving scene image generation network (DSP-GAN), which favors a hierarchical architecture, for the purpose of depth structure preserving scene image generation. The main trunk of the proposed infrastructure is built upon a Hawkes point process that models high-order spatial dependency between different depth layers. Within each layer generative adversarial sub-networks are trained collaboratively to generate realistic scene components, conditioned on the layer information produced by the point process. We experiment our model on annotated natural scene images collected from SUN dataset and demonstrate that our models are capable of generating depth-realistic natural scene image.
Wendong Zhang 0002, Feng Gao 0014, Bingbing Ni, Ling-Yu Duan, Yichao Yan, Jingwei Xu 0005, Xiaokang Yang 0001
ACM Multimedia5
2017 Image Matching via Loopy RNN
abstract
Most existing matching algorithms are one-off algorithms, i.e., they usually measure the distance between the two image feature representation vectors for only one time. In contrast, human's vision system achieves this task, i.e., image matching, by recursively looking at specific/related parts of both images and then making the final judgement. Towards this end, we propose a novel loopy recurrent neural network (Loopy RNN), which is capable of aggregating relationship information of two input images in a progressive/iterative manner and outputting the consolidated matching score in the final iteration. A Loopy RNN features two uniqueness. First, built on conventional long short-term memory (LSTM) nodes, it links the output gate of the tail node to the input gate of the head node, thus it brings up symmetry property required for matching. Second, a monotonous loss designed for the proposed network guarantees increasing confidence during the recursive matching process. Extensive experiments on several image matching benchmarks demonstrate the great potential of the proposed method.
Donghao Luo 0001, Bingbing Ni, Yichao Yan, Xiaokang Yang 0001
IJCAI3
2017 Predicting Human Interaction via Relative Attention Model
abstract
Predicting human interaction is challenging as the on-going activity has to be inferred based on a partially observed video. Essentially, a good algorithm should effectively model the mutual influence between the two interacting subjects. Also, only a small region in the scene is discriminative for identifying the on-going interaction. In this work, we propose a relative attention model to explicitly address these difficulties. Built on a tri-coupled deep recurrent structure representing both interacting subjects and global interaction status, the proposed network collects spatio-temporal information from each subject, rectified with global interaction information, yielding effective interaction representation. Moreover, the proposed network also unifies an attention module to assign higher importance to the regions which are relevant to the on-going action. Extensive experiments have been conducted on two public datasets, and the results demonstrate that the proposed relative attention network successfully predicts informative regions between interacting subjects, which in turn yields superior human interaction prediction accuracy.
Yichao Yan, Bingbing Ni, Xiaokang Yang 0001
IJCAI1
2017 Deep Cross-Modality Alignment for Multi-Shot Person Re-IDentification
abstract
Multi-shot person Re-IDentification (Re-ID) has recently received more research attention as its problem setting is more realistic compared to single-shot Re-ID in terms of application. While many large-scale single-shot Re-ID human image datasets have been released, most existing multishot Re-ID video sequence datasets containonly a few (i.e., several hundreds) human instances, which hinders further improvement of multi-shot Re-ID performance. To this end, we propose a deep cross-modality alignment network, which jointly explores both human sequence pairs and image pairs to facilitate training better multi-shot human Re-ID models, i.e., via transferring knowledge from image data to sequence data. To mitigate modality-to-modality mismatch issue, the proposed network is equipped with an image-to-sequence adaption module called cross-modality alignment sub-network, which successfully maps each human image into a pseudo human sequence to facilitate knowledge transferring and joint training. Extensive experimental results on several multi-shot person Re-ID benchmarks demonstrate great performance gain brought up by the proposed network.
Zhichao Song, Bingbing Ni, Yichao Yan, Zhe Ren, Yi Xu 0001, Xiaokang Yang 0001
ACM Multimedia3
2017 Fine-Grained Recognition via Attribute-Guided Attentive Feature Aggregation
abstract
Fine-grained object recognition is challenging due to large intra-class variation and inter-class ambiguity. A good algorithm should be able to: 1) discover discriminative local details and 2) align and aggregate these local discriminative patch-level features in an effective way to facilitate object level classification. Towards this end, we propose a novel local feature discovery, discriminative alignment and aggregation framework, inspired by the recent success of deep recurrent attention model. First, we develop a novel attribute-guided attentive network to sequentially discover informative parts/regions, by seeking a good registration between attentive regions and predefined object attributes. This could be considered as a semantic guided salient region discovery and alignment network, which might be more robust than conventional attention model. Second, these discovered regions are actively and progressively fed into a recurrent neural network, to yield the object-level representation. This could be considered as a discriminant aggregation network and informative patch-level features are propagated and accumulated to the deeper nodes of the recurrent network for final classification. We extensively test our framework on two fine-grained image benchmarks and the results demonstrate the effectiveness of the proposed framework.
Yichao Yan, Bingbing Ni, Xiaokang Yang 0001
ACM Multimedia1
2017 Skeleton-Aided Articulated Motion Generation
abstract
This work makes the first attempt to generate articulated human motion sequence from a single image. On one hand, we utilize paired inputs including human skeleton information as motion embedding and a single human image as appearance reference, to generate novel motion frames based on the conditional GAN infrastructure. On the other hand, a triplet loss is employed to pursue appearance smoothness between consecutive frames. As the proposed framework is capable of jointly exploiting the image appearance space and articulated/kinematic motion space, it generates realistic articulated motion sequence, in contrast to most previous video generation methods which yield blurred motion effects. We test our model on two human action datasets including KTH and Human3.6M, and the proposed framework generates very promising results on both datasets.
Yichao Yan, Jingwei Xu 0005, Bingbing Ni, Wendong Zhang 0002, Xiaokang Yang 0001
ACM Multimedia1
2016 Person Re-identification via Recurrent Feature Aggregation
Yichao Yan, Bingbing Ni, Zhichao Song, Chao Ma 0004, Yan Yan 0002, Xiaokang Yang 0001
ECCV (6)1
2016 Exploiting neural models for no-reference image quality assessment
abstract
We propose an improved algorithm for no-reference image quality assessment (NR-IQA) using the convolutional neural network (CNN) and neural theory based saliency detection. Firstly, we extract non-overlapping patches from the input image. For each patch, we obtain the quality score by CNN network, which consists of seven layers and integrates feature learning and regression into image patch quality estimation. Considering that the patches attracting much attention take significant role in visual perception, an efficient technique based on free energy based neural model is used to detect the saliency map. This saliency map is then applied as a weighting mask to output the quality score of the whole image. Results of experiments show that our algorithm achieves state-of-the-art performance, as compared with the prevailing IQA methods.
Cenhui Pan, Yi Xu 0001, Yichao Yan, Ke Gu 0001, Xiaokang Yang 0001
VCIP3