EDBT 2026 Demo / reviewers in the wild / expert
Yuhao Cheng
dblp:242/4245
· DBLP profile ↗
25ranked-venue papers
7as first author
23since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 5 first-author · 16 since 2021Artificial intelligence and machine learning · 14 · 3 first-author · 14 since 2021Computer networks · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Topo4D++: Realistic Physically Based 4D Head Capture With Topology-Preserving Gaussian Splatting and Expression Priorsabstract4D head capture aims to generate dynamic facial meshes in the same topology with corresponding UV maps, which requires temporal correspondence between 3D head models. Existing pipelines either involve manual processing of artists or employ constraints such as landmark tracking and optical flow, failing to achieve a trade-off between accuracy and efficiency. To enhance this process, we propose Topo4D++, a novel framework for automatic geometry and texture reconstruction that optimizes densely aligned 4D heads and 8 K BRDF maps directly from calibrated multi-view videos. Our key insight is to represent facial models as a set of dynamic 3D Gaussians with fixed topology, where the Gaussian centers are bound to the mesh vertices. This enables tracking all vertices rather than sparse vertices on the face accurately by leveraging the inverse rendering capabilities of 3D Gaussian Splatting (3DGS), while also enabling ultra-high-resolution texture generation. To maintain face structure during dynamic 3DGS optimization, we propose to optimize geometry and texture alternatively under physical and topological constraints frame-by-frame and employ blendshape-based expression priors to address extreme expressions. Then, we propose to extract dynamic facial meshes in a regular wiring arrangement and high-fidelity textures with pore-level details from the learned Gaussians. Finally, we train a diffusion-based model to generate BRDF texture maps to achieve physically based rendering. Given the absence of a universal benchmark, we construct JHead, a novel benchmark for the comprehensive evaluation of 4D head capture methods. Extensive experiments on different datasets demonstrate that our method is generalized to different capture systems, identities, and expressions, outperforming current state-of-the-art head reconstruction methods in both mesh and texture qualitatively and quantitatively. Yuhao Cheng, Xuanchen Li, Xingyu Ren, Haozhe Jia, Di Xu 0012, Wenhan Zhu, Bingbing Ni, Yichao Yan |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | ConsistentID: Portrait Generation With Multimodal Fine-Grained Identity PreservingabstractDiffusion-based technologies have made significant strides, particularly in personalized and customized facial generation. However, existing methods struggle to achieve high-fidelity and detailed identity (ID) consistency. This is mainly due to two challenges: insufficient fine-grained control over specific facial areas and the absence of a comprehensive strategy for ID preservation that accounts for both intricate facial details and the overall facial structure. To address these limitations, we introduce ConsistentID, an innovative method crafted for diverse identity-preserving portrait generation under fine-grained multimodal facial prompts, utilizing only a single reference image. ConsistentID comprises two core components: a multimodal facial prompt generator and an ID-preservation network. The facial prompt generator combines localized facial features, facial feature descriptions, and overall facial descriptions to enhance the precision of facial detail reconstruction. The ID-preservation network, optimized with a facial attention localization strategy, ensures consistent identity preservation across facial regions. Together, these components leverage fine-grained multimodal identity information to improve identity preservation accuracy significantly. To drive ConsistentID's training, we propose a fine-grained portrait dataset, FGID, with over 500,000 facial images, offering greater diversity and comprehensiveness than existing public facial datasets. Experimental results substantiate that our ConsistentID achieves exceptional precision and diversity in personalized facial generation, surpassing existing methods in the MyStyle dataset. In addition, although ConsistentID introduces more multimodal ID information, it still maintains rapid inference speed during the generation process. Jiehui Huang, Wenhui Song, Zheng Chong, Zhenchao Tang, Yuhao Cheng, Long Chen 0005, Yiqiang Yan, Shengcai Liao, Xiaodan Liang |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | ExpDiff: Generating High-Fidelity 3D Facial Expression Meshes and BRDF Textures via Diffusion Modelabstract3D face generation is a critical task for immersive multimedia applications, where a key challenge is the joint synthesis of expressive geometry and BRDF textures. Existing methods often struggle with geometric-textural coherence and corresponding reflectance modeling. To overcome these limitations, we present ExpDiff, a framework that generates expression meshes and corresponding BRDF textures from a single neutral-expression face. Our method employs an attention-based diffusion model to learn the semantic transition across expressions. To ensure correspondence between geometry and texture, we introduce a unified representation that explicitly models geometric-textural interaction, which is encoded into a shared latent space by models pre-trained on a vast dataset for strong generalization. To achieve semantically coherent and physically consistent generation, we propose to guide the denoising direction with specially designed textual prompts. We further construct two novel facial expression datasets, J-Reflectance, for ultra-high-quality assets, and FFHQ-BRDFExp for diverse identities, both of which are publicly released to advance the community. Extensive experiments demonstrate our method's superior performance in photo-realistic facial expression synthesis. Project page:https://cyh-sj.github.io/expdiff/. Yuhao Cheng, Xuanchen Li, Xingyu Ren, Zhuo Chen 0060, Chenghui Ke, Xiaokang Yang 0001, Yichao Yan |
IEEE Trans. Multim. | 1 |
| 2026 | Relightable and Animatable Gaussian Head Avatar From Monocular VideosabstractIn the realm of virtual avatar creation, accurate relighting capabilities are key to enhancing realism and immersion. We propose a novel pipeline for building personalized and relightable avatars from a monocular video captured under unknown lighting. This minimal input poses challenges in material entanglement and novel-view inconsistency. To tackle these, we introduce a disentangled dynamic 3D Gaussian representation that models diverse material properties and supports photorealistic rendering and animation via a parametric face model. To resolve material ambiguity under uncontrolled lighting, we train a 2D diffusion-based model to predict canonical-lighting images and physically-based material maps from casually lit portraits. These predictions serve as supervisory signals to guide the 3D disentanglement process. Additionally, we incorporate a 3D prior to enhance novel-view consistency, improving geometry and appearance in unseen views. Experiments demonstrate that our approach significantly boosts reconstruction quality and relighting fidelity, offering a practical and cost-effective solution for creating high-quality personalized avatars. Zhuo Chen 0060, Yichao Yan, Jingnan Gao, Zhuo Su 0008, Zhaohu Li, Yuhao Cheng, Xueying Lee, Yutong Leng, Yikun Zeng, Guidong Wang, Xiaokang Yang 0001 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2025 | Towards High-fidelity 3D Talking Avatar with Personalized Dynamic TextureabstractSignificant progress has been made for speech-driven 3D face animation, but most works focus on learning the motion of mesh/geometry, ignoring the impact of dynamic texture. In this work, we reveal that dynamic texture plays a key role in rendering high-fidelity talking avatars, and introduce a high-resolution 4D dataset TexTalk4D, consisting of 100 minutes of audio-synced scan-level meshes with detailed 8K dynamic textures from 100 subjects. Based on the dataset, we explore the inherent correlation between motion and texture, and propose a diffusion-based framework TexTalker to simultaneously generate facial motions and dynamic textures from speech. Furthermore, we propose a novel pivot-based style injection strategy to capture the complicity of different texture and motion styles, which allows disentangled control. TexTalker, as the first method to generate audio-synced facial motion with dynamic texture, not only outperforms the prior arts in synthesising facial motions, but also produces realistic textures that are consistent with the underlying facial movements. Project page: https://xuanchenli.github.io/TexTalk/. Xuanchen Li, Yuhao Cheng, Yikun Zeng, Xingyu Ren, Wenhan Zhu, Weiming Zhao, Yichao Yan |
CVPR | 3 |
| 2025 | S^3-Face: SSS-Compliant Facial Reflectance Estimation via Diffusion PriorsabstractRecent 3D face reconstruction methods have made remarkable advancements, yet achieving high-quality facial reflectance from monocular input remains challenging. Existing methods rely on the light-stage captured data to learn facial reflectance models. However, limited subject diversity in these datasets poses challenges in achieving good generalization and broad applicability. This motivates us to explore whether the extensive priors captured in recent generative diffusion models (e.g., Stable Diffusion) can enable more generalizable facial reflectance estimation as these models have been pre-trained on large-scale internet image collections containing rich visual patterns. In this paper, we introduce the use of Stable Diffusion as a prior for facial reflectance estimation, achieving robust results with minimal captured data for fine-tuning. We present S3-Face, a comprehensive framework capable of producing SSS-compliant skin reflectance from in-the-wild images. Our method adopts a two-stage training approach: in the first stage, DSN-Net is trained to predict diffuse albedo, specular albedo, and normal maps from in-the-wild images using a novel joint reflectance attention module. In the second stage, HM-Net is trained to generate hemoglobin and melanin maps based on the diffuse albedo predicted in the first stage, yielding SSS-compliant and detailed reflectance maps. Extensive experiments demonstrate that our method achieves strong generalization and produces high-fidelity, SSS-compliant facial reflectance estimation. Xingyu Ren, Jiankang Deng, Yuhao Cheng, Wenhan Zhu, Yichao Yan, Xiaokang Yang 0001, Stefanos Zafeiriou, Chao Ma 0004 |
CVPR | 3 |
| 2025 | Multimodal Latent Diffusion Model for Complex Sewing Pattern GenerationabstractGenerating sewing patterns in garment design is receiving increasing attention due to its CG-friendly and flexible-editing nature. Previous sewing pattern generation methods have been able to produce exquisite clothing, but struggle to design complex garments with detailed control. To address these issues, we propose SewingLDM, a multi-modal generative model that generates sewing patterns controlled by text prompts, body shapes, and garment sketches. Initially, we extend the original vector of sewing patterns into a more comprehensive representation to cover more intricate details and then compress them into a compact latent space. To learn the sewing pattern distribution in the latent space, we design a two-step training strategy to inject the multi-modal conditions, \ie, body shapes, text prompts, and garment sketches, into a diffusion model, ensuring the generated garments are body-suited and detail-controlled. Comprehensive qualitative and quantitative experiments show the effectiveness of our proposed method, significantly surpassing previous approaches in terms of complex garment design and various body adaptability. Our project page: https://shengqiliu1.github.io/SewingLDM. Shengqi Liu, Yuhao Cheng, Zhuo Chen 0060, Xingyu Ren, Wenhan Zhu, Lincheng Li, Mengxiao Bi, Xiaokang Yang 0001, Yichao Yan |
ICCV | 2 |
| 2025 | GDrag: Towards General-Purpose Interactive Editing with Anti-ambiguity Point DiffusionabstractRecent interactive point-based image manipulation methods have gained considerable attention for being user-friendly. However, these methods still face two types of ambiguity issues that can lead to unsatisfactory outcomes, namely, intention ambiguity which misinterprets the purposes of users, and content ambiguity where target image areas are distorted by distracting elements. To address these issues and achieve general-purpose manipulations, we propose a novel task-aware, training-free framework called GDrag. Specifically, GDrag defines a taxonomy of atomic manipulations, which can be parameterized and combined unitedly to represent complex manipulations, thereby reducing intention ambiguity. Furthermore, GDrag introduces two strategies to mitigate content ambiguity, including an anti-ambiguity dense trajectory calculation method (ADT) and a self-adaptive motion supervision method (SMS). Given an atomic manipulation, ADT converts the sparse user-defined handle points into a dense point set by selecting their semantic and geometric neighbors, and calculates the trajectory of the point set. Unlike previous motion supervision methods relying on a single global scale for low-rank adaption, SMS jointly optimizes point-wise adaption scales and latent feature biases. These two methods allow us to model fine-grained target contexts and generate precise trajectories. As a result, GDrag consistently produces precise and appealing results in different editing tasks. Extensive experiments on the challenging DragBench dataset demonstrate that GDrag outperforms state-of-the-art methods significantly. The code of GDrag will be released upon acceptance. Xiaojian Lin, Yuhao Cheng, Yiqiang Yan, Xiaodan Liang |
ICLR | 3 |
| 2025 | LaVieID: Local Autoregressive Diffusion Transformers for Identity-Preserving Video CreationabstractIn this paper, we present LaVieID, a novel local a utoregressive vi deo diffusion framework designed to tackle the challenging id entity-preserving text-to-video task. The key idea of LaVieID is to mitigate the loss of identity information inherent in the stochastic global generation process of diffusion transformers (DiTs) from both spatial and temporal perspectives. Specifically, unlike the global and unstructured modeling of facial latent states in existing DiTs, LaVieID introduces a local router to explicitly represent latent states by weighted combinations of fine-grained local facial structures. This alleviates undesirable feature interference and encourages DiTs to capture distinctive facial characteristics. Furthermore, a temporal autoregressive module is integrated into LaVieID to refine denoised latent tokens before video decoding. This module divides latent tokens temporally into chunks, exploiting their long-range temporal dependencies to predict biases for rectifying tokens, thereby significantly enhancing inter-frame identity consistency. Consequently, LaVieID can generate high-fidelity personalized videos and achieve state-of-the-art performance. Our code and models are available at https://github.com/ssugarwh/LaVieID. Wenhui Song, Jiehui Huang, Panwen Hu, Yuhao Cheng, Yiqiang Yan, Xiaodan Liang |
ACM Multimedia | 5 |
| 2025 | Revealing Directions for Text-Guided 3D Face Editingabstract3D face editing is a significant task in multimedia, aimed at the manipulation of 3D face models across various control signals. The success of 3D-aware GAN provides expressive 3D models learned from 2D single-view images only, encouraging researchers to discover semantic editing directions in its latent space. However, previous methods face challenges in balancing quality, efficiency, and generalization. To solve the problem, we explore the possibility of introducing the strength of diffusion model into 3D-aware GANs. In this paper, we presentFace Clan, a fast and text-general approach for generating and manipulating 3D faces based on arbitrary attribute descriptions. To achieve disentangled editing, we propose to diffuse on the latent space under a pair of opposite prompts to estimate the mask indicating the region of interest on latent codes. Based on the mask, we then apply denoising to the masked latent codes to reveal the editing direction. Our method offers a precisely controllable manipulation method, allowing users to intuitively customize regions of interest with the text description. Experiments demonstrate the effectiveness and generalization of our Face Clan for various pre-trained GANs. It offers an intuitive and wide application for text-guided face editing that contributes to the landscape of multimedia content creation. Our project page:https://windlikestone.github.io/Face_clan_website/. Zhuo Chen 0060, Yichao Yan, Shengqi Liu, Yuhao Cheng, Weiming Zhao, Lincheng Li, Mengxiao Bi, Xiaokang Yang 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | 3D-Aware Face Editing via Warping-Guided Latent Direction Learningabstract3D facial editing, a longstanding task in computer vision with broad applications, is expected to fast and intuitively manipulate any face from arbitrary viewpoints following the user's will. Existing works have limitations in terms of intuitiveness, generalization, and efficiency. To overcome these challenges, we propose FaceEdit3D, which allows users to directly manipulate 3D points to edit a 3D face, achieving natural and rapid face editing. After one or several points are manipulated by users, we propose the tri-plane warping to directly deform the view-independent 3D representation. To address the problem of distortion caused by tri-plane warping, we train a warp-aware encoder to project the warped face onto a standardized latent space. In this space, we further propose directional latent editing to mitigate the identity bias caused by the encoder and realize the disentangled editing of various attributes. Extensive experiments show that our method achieves superior results with rich facial details and nice identity preservation. Our approach also supports general applications like multi-attribute continuous editing and cat/car editing. The project website is https://cyh-sj.github.io/FaceEdit3DI. Yuhao Cheng, Zhuo Chen 0060, Xingyu Ren, Wenhan Zhu, Zhengqin Xu, Di Xu 0012, Changpeng Yang, Yichao Yan |
CVPR | 1 |
| 2024 | Monocular Identity-Conditioned Facial Reflectance ReconstructionabstractRecent 3D face reconstruction methods have made re-markable advancements, yet there remain huge challenges in monocular high-quality facial reflectance reconstruction. Existing methods rely on a large amount of light-stage captured data to learn facial reflectance models. However, the lack of subject diversity poses challenges in achieving good generalization and widespread applicability. In this paper, we learn the reflectance prior in image space rather than UV space and present a framework named ID2Reflectance. Our framework can directly estimate the reflectance maps of a single image while using limited reflectance data for training. Our key insight is that reflectance data shares facial structures with RGB faces, which enables obtaining expressive facial prior from inexpensive RGB data thus re-ducing the dependency on reflectance data. We first learn a high-quality prior for facial reflectance. Specifically, we pretrain multi-domain facial feature code books and design a codebook fusion method to align the reflectance and RGB domains. Then, we propose an identity-conditioned swapping module that injects facial identity from the target image into the pre-trained autoencoder to modify the identity of the source reflectance image. Finally, we stitch multi-view swapped reflectance images to obtain renderable assets. Extensive experiments demonstrate that our method exhibits excellent generalization capability and achieves state-of-the-art facial reflectance reconstruction results for in-the-wild faces. Our project page is https://xingyuren.github.io/id2reflectance. Xingyu Ren, Jiankang Deng, Yuhao Cheng, Jia Guo 0003, Chao Ma 0004, Yichao Yan, Wenhan Zhu, Xiaokang Yang 0001 |
CVPR | 3 |
| 2024 | Topo4D: Topology-Preserving Gaussian Splatting for High-fidelity 4D Head Capture
Xuanchen Li, Yuhao Cheng, Xingyu Ren, Haozhe Jia, Di Xu 0012, Wenhan Zhu, Yichao Yan |
ECCV (34) | 2 |
| 2024 | GarmentAligner: Text-to-Garment Generation via Retrieval-Augmented Multi-level Corrections
Zheng Chong, Xujie Zhang, Yuhao Cheng, Yiqiang Yan, Xiaodan Liang |
ECCV (25) | 5 |
| 2024 | Learning Interaction-aware 3D Gaussian Splatting for One-shot Hand AvatarsabstractIn this paper, we propose to create animatable avatars for interacting hands with 3D Gaussian Splatting (GS) and single-image inputs. Existing GS-based methods designed for single subjects often yield unsatisfactory results due to limited input views, various hand poses, and occlusions. To address these challenges, we introduce a novel two-stage interaction-aware GS framework that exploits cross-subject hand priors and refines 3D Gaussians in interacting areas. Particularly, to handle hand variations, we disentangle the 3D presentation of hands into optimization-based identity maps and learning-based latent geometric features and neural texture maps. Learning-based features are captured by trained networks to provide reliable priors for poses, shapes, and textures, while optimization-based identity maps enable efficient one-shot fitting of out-of-distribution hands. Furthermore, we devise an interaction-aware attention module and a self-adaptive Gaussian refinement module. These modules enhance image rendering quality in areas with intra- and inter-hand interactions, overcoming the limitations of existing GS-based methods. Our proposed method is validated via extensive experiments on the large-scale InterHand2.6M dataset, and it significantly improves the state-of-the-art performance in image quality. Code and models will be released upon acceptance. Wanquan Liu, Xiaodan Liang, Yiqiang Yan, Yuhao Cheng, Chenqiang Gao |
NeurIPS | 6 |
| 2024 | Multi-label arrhythmia classification using 12-lead ECG based on lead feature guide network
Yuhao Cheng, Deyin Li, Duoduo Wang, Lirong Wang |
Eng. Appl. Artif. Intell. | 1 |
| 2024 | Esophageal tissue segmentation on OCT images with hybrid attention network
Deyin Li, Yuhao Cheng, Lirong Wang |
Multim. Tools Appl. | 2 |
| 2024 | Head3D: Complete 3D Head Generation via Tri-plane Feature DistillationabstractHead generation with diverse identities is an important task in computer vision and computer graphics, widely used in multimedia applications. However, current full-head generation methods require a large number of three-dimensional (3D) scans or multi-view images to train the model, resulting in expensive data acquisition costs. To address this issue, we propose Head3D, a method to generate full 3D heads with limited multi-view images. Specifically, our approach first extracts facial priors represented by tri-planes learned in EG3D, a 3D-aware generative model, and then proposes feature distillation to deliver the 3D frontal faces within complete heads without compromising head integrity. To mitigate the domain gap between the face and head models, we present a dual-discriminator to guide the frontal and back head generation. Our model achieves cost-efficient and diverse complete head generation with photo-realistic renderings and high-quality geometry representations. Extensive experiments demonstrate the effectiveness of our proposed Head3D, both qualitatively and quantitatively. Yuhao Cheng, Yichao Yan, Wenhan Zhu, Bowen Pan, Xiaokang Yang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | GANHead: Towards Generative Animatable Neural Head AvatarsabstractTo bring digital avatars into people's lives, it is highly demanded to efficiently generate complete, realistic, and animatable head avatars. This task is challenging, and it is difficult for existing methods to satisfy all the requirements at once. To achieve these goals, we propose GANHead (Generative Animatable Neural Head Avatar), a novel generative head model that takes advantages of both the fine-grained control over the explicit expression parameters and the realistic rendering results of implicit representations. Specifically, GANHead represents coarse geometry, fine-gained details and texture via three networks in canonical space to obtain the ability to generate complete and realistic head avatars. To achieve flexible animation, we define the deformation filed by standard linear blend skinning (LBS), with the learned continuous pose and expression bases and LBS weights. This allows the avatars to be directly animated by FLAME [22] parameters and generalize well to unseen poses and expressions. Compared to state-of-the-art (SOTA) methods, GANHead achieves superior performance on head avatar generation and raw scan fitting. Sijing Wu, Yichao Yan, Yuhao Cheng, Wenhan Zhu, Ke Gao 0012, Guangtao Zhai |
CVPR | 4 |
| 2022 | CageNeRF: Cage-based Neural Radiance Field for Generalized 3D Deformation and AnimationabstractWhile implicit representations have achieved high-fidelity results in 3D rendering, it remains challenging to deforming and animating the implicit field. Existing works typically leverage data-dependent models as deformation priors, such as SMPL for human body animation. However, this dependency on category-specific priors limits them to generalize to other objects. To solve this problem, we propose a novel framework for deforming and animating the neural radiance field learned on \textit{arbitrary} objects. The key insight is that we introduce a cage-based representation as deformation prior, which is category-agnostic. Specifically, the deformation is performed based on an enclosing polygon mesh with sparsely defined vertices called \textit{cage} inside the rendering space, where each point is projected into a novel position based on the barycentric interpolation of the deformed cage vertices. In this way, we transform the cage into a generalized constraint, which is able to deform and animate arbitrary target objects while preserving geometry details. Based on extensive experiments, we demonstrate the effectiveness of our framework in the task of geometry editing, object animation and deformation transfer. Yicong Peng, Yichao Yan, Shengqi Liu, Yuhao Cheng, Shanyan Guan, Bowen Pan, Guangtao Zhai, Xiaokang Yang 0001 |
NeurIPS | 4 |
| 2022 | Cross-modal Graph Matching Network for Image-text RetrievalabstractImage-text retrieval is a fundamental cross-modal task whose main idea is to learn image-text matching. Generally, according to whether there exist interactions during the retrieval process, existing image-text retrieval methods can be classified into independent representation matching methods and cross-interaction matching methods. The independent representation matching methods generate the embeddings of images and sentences independently and thus are convenient for retrieval with hand-crafted matching measures (e.g., cosine or Euclidean distance). As to the cross-interaction matching methods, they achieve improvement by introducing the interaction-based networks for inter-relation reasoning, yet suffer the low retrieval efficiency. This article aims to develop a method that takes the advantages of cross-modal inter-relation reasoning of cross-interaction methods while being as efficient as the independent methods. To this end, we propose a graph-based Cross-modal Graph Matching Network (CGMN) , which explores both intra- and inter-relations without introducing network interaction. In CGMN, graphs are used for both visual and textual representation to achieve intra-relation reasoning across regions and words, respectively. Furthermore, we propose a novel graph node matching loss to learn fine-grained cross-modal correspondence and to achieve inter-relation reasoning. Experiments on benchmark datasets MS-COCO, Flickr8K, and Flickr30K show that CGMN outperforms state-of-the-art methods in image retrieval. Moreover, CGMM is much more efficient than state-of-the-art methods using interactive matching. The code is available at https://github.com/cyh-sj/CGMN . Yuhao Cheng, Xiaoguang Zhu, Jiuchao Qian, Fei Wen 0005 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2021 | SDAN: Stacked Diverse Attention Network for Video Action RecognitionabstractRecently, deep learning methods have proved exceptional performance in video action recognition. However, it remains a challenging problem to extract discriminative features from videos effectively. Most existing methods mainly focus on spatial-temporal information separately. In this paper, we propose a novel Stacked Diverse Attention Network (SDAN). It uses a Multi-dimensional Attention Module to emphasize informative maps along the channel and spatial-temporal dimension. Besides, a Supervised Attention Module is designed to generate a weighted feature map in a class-supervised way. Compared with previous methods, the proposed method has the following advantages: (1) The Multi-dimensional Attention Module can exploit effective combinations and correlation of attention mechanisms along different dimensions. (2) The Supervised Attention Module improves the network capability supervised directly by action labels which reveal the class-related object and limb information. Extensive experimental evaluation demonstrates the effectiveness of the proposed approach and establishes significant results on Kinetics400, UCF101 and HMDB51 Datasets. Codes are available on https://github.com/jeff62802217/SDAN-Pytorch. Xiaoguang Zhu, Siran Huang, Wenjing Fan, Yuhao Cheng, Huaqing Shao |
ISCAS | 4 |
| 2021 | Pose-Guided Tracking-by-Detection: Robust Multi-Person Pose TrackingabstractMulti-person pose tracking task aims to estimate and track person keypoints in videos. Most of the previous methods follow the general track-by-detection strategy that ignores the consistent pose information during the whole framework. Thus, they often suffer from missing detections or inaccurate human association in challenging scenes with motion blur or person occlusion. To handle those problems, we propose a pose-guided tracking-by-detection framework that fuses pose information into both video human detection and human association procedures. In the video human detection stage, we adopt the pose-guided person location prediction exploiting the temporal information to make up missing detections. Technically, pose heatmaps are utilized to cope with the person-specific intra-class distractors. Furthermore, in the human association stage, we propose an appearance discriminative model based on the hierarchical pose-guided graph convolutional networks (PoseGCN). The PoseGCN-based model exploits human structural relations to boost person representation. Extensive experiments show the superiority of our method on the challenging pose tracking benchmark. Our proposed method ranks first on the PoseTrack leaderboard.11http://posetrack.net/leaderboard.php till the submission date (22-Aug-2019) of this paper. Our code has been publicly available at https://github.com/human-centric982/PGPT. Qian Bao, Wu Liu 0005, Yuhao Cheng, Boyan Zhou, Tao Mei 0001 |
IEEE Trans. Multim. | 3 |
| 2020 | PyAnomaly: A Pytorch-based Toolkit for Video Anomaly DetectionabstractVideo anomaly detection is an essential task in computer vision which attracts massive attention from academia and industry. The existing approaches are implemented in diverse deep learning frameworks and settings, making it difficult to reproduce the results published by the original authors. Undoubtedly, this phenomenon is detrimental to the development of Video Anomaly detection and community communication. In this paper, we present a PyTorch-based video anomaly detection toolbox, namely PyAnomaly that contains high modular and extensible components, comprehensive and impartial evaluation platforms, a friendly manageable system configuration, and the abundant engineering deployment functions. To make it easy-to-use and easy-to-extend, we implement the architecture by hooks and registers functionality. Remarkably, we have reproduced the comparable experimental results of six representative methods as those published by the original authors, and we will release these pre-trained models with more rich configurations. To our best knowledge, the PyAnomaly is the first open-source tool in video anomaly detection and is available at https://github.com/YuhaoCheng/PyAnomaly. Yuhao Cheng, Wu Liu 0005, Pengrui Duan, Jingen Liu, Tao Mei 0001 |
ACM Multimedia | 1 |
| 2019 | POINet: Pose-Guided Ovonic Insight Network for Multi-Person Pose TrackingabstractMulti-person pose tracking aims to jointly estimate and track multi-person keypoints in the unconstrained videos. The most popular solution to this task follows the tracking-by-detection strategy that relies on human detection and data association. While human detection has been boosted by deep learning, existing works mainly exploit several separated stages with hand-crafted metrics to realize data association, leading to great uncertainty and feeble adaption in complex scenes. To handle these problems, we propose an end-to-end pose-guided ovonic insight network (POINet) for the data association in multi-person pose tracking, which jointly learns feature extraction, similarity estimation, and identity assignment. Specifically, we design a pose-guided representation network to integrate pose information into hierarchical convolutional features, generating a pose-aligned person representation for person, which helps handle partial occlusions. Moreover, we propose an ovonic insight network to adaptively encode the cross-frame identity transformation, which can cope with the tough tracking cases of person leaving and entering the scene. In general, the proposed POINet provides a new insight to realize multi-person pose tracking in an end-to-end fashion. Extensive experiments conducted on the PoseTrack benchmark demonstrate that our POINet outperforms the state-of-the-art methods. Weijian Ruan, Wu Liu 0005, Qian Bao, Jun Chen 0001, Yuhao Cheng, Tao Mei 0001 |
ACM Multimedia | 5 |