VLDB 2026 Research / reviewers in the wild / expert
Yanbo Fan
dblp:181/4574
· DBLP profile ↗
48ranked-venue papers
5as first author
41since 2021 · last 2026
0000-0002-8530-485XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 39 · 3 first-author · 32 since 2021Graphics, computer vision, multimedia, augmented reality and games · 30 · 2 first-author · 24 since 2021Computer networks · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Boosting backdoor attack with a learnable poisoning sample selection strategy
Zihao Zhu 0001, Shaokui Wei, Li Shen 0008, Yanbo Fan, Baoyuan Wu |
Neurocomputing | 5 |
| 2026 | Seeking Flat Minima Over Diverse Surrogates for Improved Adversarial Transferability: A Theoretical Framework and Algorithmic InstantiationabstractThe transfer-based black-box adversarial attack setting poses the challenge of crafting an adversarial example (AE) on known surrogate models that remains effective against unseen target models. Due to the practical importance of this task, numerous methods have been proposed to address this challenge. However, most previous methods are heuristically designed and intuitively justified, lacking a theoretical foundation. To bridge this gap, we derive a novel transferability bound that offers provable guarantees for adversarial transferability. Our theoretical analysis has the advantages of (i) deepening our understanding of previous methods by building a general attack framework and (ii) providing guidance for designing an effective attack algorithm. Our theoretical results demonstrate that optimizing AEs toward flat minima over the surrogate model set, while controlling the surrogate-target model shift measured by the adversarial model discrepancy, yields a comprehensive guarantee for AE transferability. The results further lead to a general transfer-based attack framework, within which we observe that previous methods consider only partial factors contributing to the transferability. Algorithmically, inspired by our theoretical results, we first elaborately construct the surrogate model set in which models exhibit diverse adversarial vulnerabilities with respect to AEs to narrow the instantiated adversarial model discrepancy. Then, a model-Diversity-compatible Reverse Adversarial Perturbation (DRAP) is generated to effectively promote the flatness of AEs over diverse surrogate models to improve transferability. Extensive experiments on NIPS2017 and CIFAR-10 datasets against various target models demonstrate the effectiveness of our proposed attack. Meixi Zheng, Kehan Wu, Yanbo Fan, Rui Huang 0001, Baoyuan Wu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Cooperative HAP-UAV Optimization for IoRT Data Collection: A Green Transmission Strategy for Maximizing Energy EfficiencyabstractSupported by space-air-ground integrated networks (SAGIN), Internet of Remote Things (IoRT) is regarded as a cornerstone for realizing global connectivity in 6 G networks. The integration of high-altitude platforms (HAPs) and unmanned aerial vehicles (UAVs), offering both wide coverage and agile data access, becomes a promising paradigm for IoRT data collection. However, sustaining reliable and efficient transmission is challenged by the mobility and constrained onboard energy of HAPs and UAVs, as well as atmospheric fading effects. To address these issues, we propose a green and efficient HAP-UAV collaborative design for IoRT data collection, which jointly considers both transmission performance and energy consumption. Firstly, we introduce a novel metric, Overall Energy Efficiency (OEE), to quantify the balance between cooperative transmission performance and the total energy cost under dynamic trajectory planning. Secondly, we formulate a joint optimization problem that simultaneously optimizes UAV/HAP trajectories, UAV power control, HAP selection, and bandwidth allocation. Thirdly, to address the formulated non-convex fractional problem, we develop an energy efficiency maximization strategy based on the successive convex approximation technique. Extensive simulation results demonstrate that the proposed strategy achieves significant gains in OEE, achieving superior trade-offs between energy consumption and transmission performance in HAP-UAV-assisted IoRT networks. Yanbo Fan, Yuanguo Bi, Xingyu Ji, Dusit Niyato, Enchao Zhang, Liang Zhao 0004, Qiang He 0002 |
IEEE Trans. Mob. Comput. | 1 |
| 2025 | HERA: Hybrid Explicit Representation for Ultra-Realistic Head AvatarsabstractWe introduce a novel approach to creating ultra-realistic head avatars and rendering them in real time (≥ 30 fps at 2048 × 1334 resolution). First, we propose a hybrid explicit representation that combines the advantages of two primitive based efficient rendering techniques. UV-mapped 3D mesh is utilized to capture sharp and rich textures on smooth surfaces, while 3D Gaussian Splatting is employed to represent complex geometric structures. In the pipeline of modeling an avatar, after tracking parametric models based on captured multi-view RGB videos, our goal is to simultaneously optimize the texture and opacity map of mesh, as well as a set of 3D Gaussian splats localized and rigged onto the mesh facets. Specifically, we perform α-blending on the color and opacity values based on the merged and reordered z-buffer from the rasterization results of mesh and 3DGS. This process involves the mesh and 3DGS adaptively fitting the captured visual information to outline a high-fidelity digital avatar. To avoid artifacts caused by Gaussian splats crossing the mesh facets, we design a stable hybrid depth sorting strategy. Experiments illustrate that our modeled results exceed those of state-of-the-art approaches. Hongrui Cai, Xuan Wang 0009, Jiafei Li, Yanbo Fan, Shenghua Gao, Juyong Zhang |
CVPR | 6 |
| 2025 | AvatarArtist: Open-Domain 4D AvatarizationabstractThis work focuses on open-domain 4D avatarization, with the purpose of creating a 4D avatar from a portrait image in an arbitrary style. We select parametric triplanes as the intermediate 4D representation, and propose a practical training paradigm that takes advantage of both generative adversarial networks (GANs) and diffusion models. Our design stems from the observation that 4D GANs excel at bridging images and triplanes without supervision yet usually face challenges in handling diverse data distributions. A robust 2D diffusion prior emerges as the solution, assisting the GAN in transferring its expertise across various domains. The synergy between these experts permits the construction of a multi-domain image-triplane dataset, which drives the development of a general 4D avatar creator. Extensive experiments suggest that our model, termed AvatarArtist, is capable of producing high-quality 4D avatars with strong robustness to various source image domains. The code, the data, and the models will be made publicly available to facilitate future studies. Xuan Wang 0009, Ziyu Wan, Yue Ma 0016, Jingye Chen, Yanbo Fan, Yujun Shen, Yibing Song, Qifeng Chen 0001 |
CVPR | 6 |
| 2025 | DualTalk: Dual-Speaker Interaction for 3D Talking Head ConversationsabstractIn face-to-face conversations, individuals need to switch between speaking and listening roles seamlessly. Existing 3D talking head generation models focus solely on speaking or listening, neglecting the natural dynamics of interactive conversation, which leads to unnatural interactions and awkward transitions. To address this issue, we propose a new task—multi-round dual-speaker interaction for 3D talking head generation—which requires models to handle and generate both speaking and listening behaviors in continuous conversation. To solve this task, we introduce DualTalk, a novel unified framework that integrates the dynamic behaviors of speakers and listeners to simulate realistic and coherent dialogue interactions. This framework not only synthesizes lifelike talking heads when speaking but also generates continuous and vivid non-verbal feedback when listening, effectively capturing the interplay between the roles. We also create a new dataset featuring 50 hours of multi-round conversations with over 1,000 characters, where participants continuously switch between speaking and listening roles. Extensive experiments demonstrate that our method significantly enhances the naturalness and expressiveness of 3D talking heads in dual-speaker conversations. We recommend watching the supplementary video: https://ziqiaopeng.github.io/dualtalk Ziqiao Peng, Yanbo Fan, Xuan Wang 0009, Hongyan Liu 0002, Jun He 0008, Zhaoxin Fan |
CVPR | 2 |
| 2025 | Diffusion-based Realistic Listening Head Generation via Hybrid Motion ModelingabstractListening head generation aims to synthesize non-verbal responsive listening head videos that naturally react to a certain speaker, for which, both realistic head movements, expressive facial expressions, and high visual qualities are expected. Previous approaches typically follow a two-stage pipeline that first generates intermediate 3D motion signals such as 3DMM coefficients, and then synthesizes the videos by deterministic rendering, suffering from limited motion expressiveness and low visual quality (e.g. 256×256). In this work, we propose a novel listening head generation method that harnesses the generative capabilities of the diffusion model for both motion generation and high-quality rendering. Crucially, we propose an effective hybrid motion modeling module that addresses training difficulties caused by the scarcity of listening head data while preserving the intricate details that may be lost in explicit motion representations. We further develop a tailored control guidance for head pose and facial expression, by integrating their intrinsic motion characteristics. Our method enables high-fidelity video generation with 512 × 512 resolution and delivers vivid listener motion feedback. We conduct comprehensive experiments and obtain superior performance in terms of both visual quality and motion expressiveness compared with existing methods. Yanbo Fan, Xuan Wang 0009, Yu Guo 0006, Fei Wang 0008 |
CVPR | 2 |
| 2025 | 3D Gaussian Head Avatars with Expressive Dynamic Appearances by Compact Tensorial RepresentationsabstractRecent studies have combined 3D Gaussian and 3D Morphable Models (3DMM) to construct high-quality 3D head avatars. In this line of research, existing methods either fail to capture the dynamic textures or incur significant overhead in terms of runtime speed or storage space. To this end, we propose a novel method that addresses all the aforementioned demands. In specific, we introduce an expressive and compact representation that encodes texture-related attributes of the 3D Gaussians in the tensorial format. We store appearance of neutral expression in static tri-planes, and represents dynamic texture details for different expressions using lightweight 1D feature lines, which are then decoded into opacity offset relative to the neutral face. We further propose adaptive truncated opacity penalty and class-balanced sampling to improve generalization across different expressions. Experiments show this design enables accurate face dynamic details capturing while maintains real-time rendering and significantly reduces storage costs, thus broadening the applicability to more scenarios. Xuan Wang 0009, Ran Yi 0002, Yanbo Fan, Jichen Hu, Jingcheng Zhu, Lizhuang Ma |
CVPR | 4 |
| 2025 | DGTalker: Disentangled Generative Latent Space Learning for Audio-Driven Gaussian Talking Heads
Xiaoxi Liang, Yanbo Fan, Qiya Yang, Xuan Wang 0009, Wei Gao 0003, Ge Li 0002 |
ICCV | 2 |
| 2025 | Fine-Grained 3D Gaussian Head Avatars Modeling from Static Captures Via Joint Reconstruction and Registration
Yuan Sun 0003, Xuan Wang 0009, WeiLi Zhang, Yanbo Fan, Yu Guo 0006, Fei Wang 0008 |
ICCV | 5 |
| 2025 | Echo: Enhancing Conversational Behavior Generation via Hierarchical Semantic Comprehension with Large Language ModelsabstractConversational behavior generation, being a crucial capability of embodied agents, is a significant factor influencing human-computer interaction. Generating high-quality conversational motions requires not only appropriate audio-motion mapping but also interactive responses to interlocutor behaviors and comprehensive understanding of conversational semantics. Existing methods primarily rely on audio signals and interlocutor motions for main agent motion generation, lacking high-level semantic understanding of the conversational content, leading to moderate quality motions that are not appropriate for the dialogue. To address these limitations, we leverage the powerful semantic understanding capabilities of large language models, to comprehend complex conversational contexts. Inspired by human conversation processes that conversational motions are highly related to both global and local semantic factors, including the conversational context, and the intentions, emotions, and passive or active states of the participants, we propose an agentic system named Echo that analyzes such information. To achieve comprehensive conversational understanding, Echo leverages multiple prompts and test-time recipes to guide large language models in decomposing conversational structures and extracting fine-grained semantic information. Furthermore, we design a hierarchical feature fusion network that systematically integrates from frame-level audio-motion features to sentence-level semantic understanding and finally to conversation-level contextual comprehension, organically combining fine-grained semantic features from large language models with audio and motion characteristics. Experimental results demonstrate that our framework can be effectively integrated with several state-of-the-art motion generation models to enhance their performance in generating high-quality conversational behaviors. Haiwei Xue, Yanbo Fan, Xuan Wang 0009, Zhiyong Wu 0001 |
SIGGRAPH Asia | 2 |
| 2025 | Understanding adversarial robustness against on-manifold adversarial examples
Jiancong Xiao, Liusha Yang, Yanbo Fan, Jue Wang 0001, Zhi-Quan Luo |
Pattern Recognit. | 3 |
| 2025 | Trajectory Optimization and Power Allocation for Multi-UAV Wireless Networks: A Communication-Based Multi-Agent Deep Reinforcement Learning ApproachabstractUnmanned Aerial Vehicles (UAVs) play a crucial role in next-generation mobile communication systems, serving as aerial base stations to provide services when ground base stations fail to meet coverage requirements. However, trajectory planning and power allocation for collaborative UAVs as Aerial Base Stations (UAV-ABSs) face several challenges, including energy limitations, flight time constraints, high optimization complexity due to dynamic environment interactions, and insufficient decision-making information. To address these challenges, this paper proposes a multi-agent reinforcement learning algorithm, namely Communication Actor Centralized Attention Critic Algorithm (CATEN), to jointly optimize the flight trajectory and power allocation strategies of UAV-ABSs. The proposed algorithm aims to maximize the number of users meeting Quality of Service (QoS) requirements while minimizing UAV-ABSs energy consumption. To achieve this, firstly, an information sharing mechanism is designed to improve the collaboration efficiency among UAV-ABSs. It leverages distributed storage, intelligent scheduling of UAV-ABSs interaction experiences, and gating units to enhance information screening and fusion. Secondly, a multihead attention critic network is proposed to capture correlations among UAV-ABSs from different subspaces. This allows the network to prioritize value information, reduce redundancy, and strengthen UAV-ABSs collaboration and decision-making capabilities. Simulation results demonstrate that CATEN achieves better performance in terms of the number of served users and energy consumption compared to existing algorithms, exhibiting good robustness and adaptability in dynamic environments. Zimeng Yuan, Yuanguo Bi, Yanbo Fan, Lianbo Ma 0004, Liang Zhao 0004, Qiang He 0002 |
IEEE Trans. Computers | 3 |
| 2025 | GATO: Global Transmission Optimization for SAGIN-Assisted IoRT Data Collection
Yanbo Fan, Yuanguo Bi, Yufei Liu 0005, Dusit Niyato, Liang Zhao 0004, Qiang He 0002, Ammar Hawbani |
IEEE Trans. Mob. Comput. | 1 |
| 2025 | Energy Efficient AAV-Assisted Bidirectional Relaying System for Multi-Pair User DevicesabstractUnmanned aerial vehicles(UAVs), or drones, are garnering considerable focus in the realm of wireless communications research because to their notable characteristics, including exceptional mobility, versatile deployment capabilities, and robustness in maintaining line-of-sight (LoS) links. This paper studies a UAV-assisted bidirectional relaying system for multi-pair user devices (UDs), where a rotary-wing UAV is used to serve as a mobile relay for providing information transmission between UDs belonging to a pair on the ground. The UDs in a pair communicate with each other via the UAV relay employing the physical-layer network coding (PNC) technique. To trade off fair communications with system energy consumption, we jointly optimize transmission scheduling and association, UAV relay and UD transmission power, and UAV trajectory to maximize system energy efficiency during UAV relay communications. Due to the formulated problem being mixed-integer and nonconvex programming, it proves to be excessively complex to solve. For ease of solution, this problem is initially decomposed into three sub-problems. Next, by adopting the block coordinate descent (BCD) method, the successive convex approximation (SCA) method, and the Dinkelbach method, an efficient iterative algorithm is proposed that alternately solves variables of each sub-problem while fixing others. The numerical results demonstrate that our designed scheme is capable of substantially improving the system energy efficiency in comparison with other baseline schemes and benchmark schemes. Na Lin 0001, Ammar Hawbani, Cunqian Yu, Yanbo Fan, Liang Zhao 0004 |
IEEE Trans. Mob. Comput. | 5 |
| 2024 | Learning Pseudo 3D Guidance for View-Consistent Texturing with 2D Diffusion
Kehan Li 0002, Yanbo Fan, Yang Wu 0001, Zhongqian Sun, Wei Yang 0019, Xiangyang Ji, Li Yuan 0007, Jie Chen 0001 |
ECCV (86) | 2 |
| 2024 | Regional Adversarial Training for Better Robust Generalization
Chuanbiao Song, Yanbo Fan, Aoyang Zhou, Baoyuan Wu, Yiming Li 0004, Zhifeng Li 0001, Kun He 0001 |
Int. J. Comput. Vis. | 2 |
| 2024 | Generalizable Black-Box Adversarial Attack With Meta LearningabstractIn the scenario of black-box adversarial attack, the target model's parameters are unknown, and the attacker aims to find a successful adversarial perturbation based on query feedback under a query budget. Due to the limited feedback information, existing query-based black-box attack methods often require many queries for attacking each benign example. To reduce query cost, we propose to utilize the feedback information across historical attacks, dubbed example-level adversarial transferability. Specifically, by treating the attack on each benign example as one task, we develop a meta-learning framework by training a meta generator to produce perturbations conditioned on benign examples. When attacking a new benign example, the meta generator can be quickly fine-tuned based on the feedback information of the new task as well as a few historical attacks to produce effective perturbations. Moreover, since the meta-train procedure consumes many queries to learn a generalizable generator, we utilize model-level adversarial transferability to train the meta generator on a white-box surrogate model, then transfer it to help the attack against the target model. The proposed framework with the two types of adversarial transferability can be naturally combined with any off-the-shelf query-based attack methods to boost their performance, which is verified by extensive experiments. The source code is available at https://github.com/SCLBD/MCG-Blackbox. Yong Zhang 0034, Baoyuan Wu, Jingyi Zhang 0005, Yanbo Fan, Yujiu Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | High-fidelity Facial Avatar Reconstruction from Monocular Video with Generative PriorsabstractHigh-fidelity facial avatar reconstruction from a monocular video is a significant research problem in computer graphics and computer vision. Recently, Neural Radiance Field (NeRF) has shown impressive novel view rendering results and has been considered for facial avatar reconstruction. However, the complex facial dynamics and missing 3D information in monocular videos raise significant challenges for faithful facial reconstruction. In this work, we propose a new method for NeRF-based facial avatar reconstruction that utilizes 3D-aware generative prior. Different from existing works that depend on a conditional deformation field for dynamic modeling, we propose to learn a personalized generative prior, which is formulated as a local and low dimensional subspace in the latent space of 3D-GAN. We propose an efficient method to construct the personalized generative prior based on a small set of facial images of a given individual. After learning, it allows for photo-realistic rendering with novel views, and the face reenactment can be realized by performing navigation in the latent space. Our proposed method is applicable for different driven signals, including RGB images, 3DMM coefficients, and audio. Compared with existing works, we obtain superior novel view synthesis results and faithfully face reenactment performance. The code is available here https://github.com/bbaaii/HFA-GP. Yunpeng Bai, Yanbo Fan, Xuan Wang 0009, Yong Zhang 0034, Jingxiang Sun, Chun Yuan 0003, Ying Shan |
CVPR | 2 |
| 2023 | DPE: Disentanglement of Pose and Expression for General Video Portrait EditingabstractOne-shot video-driven talking face generation aims at producing a synthetic talking video by transferring the facial motion from a video to an arbitrary portrait image. Head pose and facial expression are always entangled in facial motion and transferred simultaneously. However, the entanglement sets up a barrier for these methods to be used in video portrait editing directly, where it may require to modify the expression only while maintaining the pose unchanged. One challenge of decoupling pose and expression is the lack of paired data, such as the same pose but different expressions. Only a few methods attempt to tackle this challenge with the feat of 3D Morphable Models (3DMMs) for explicit disentanglement. But 3DMMs are not accurate enough to capture facial details due to the limited number of Blend-shapes, which has side effects on motion transfer. In this paper, we introduce a novel self-supervised disentanglement framework to decouple pose and expression without 3DMMs and paired data, which consists of a motion editing module, a pose generator, and an expression generator. The editing module projects faces into a latent space where pose motion and expression motion can be disentangled, and the pose or expression transfer can be performed in the latent space conveniently via addition. The two generators render the modified latent codes to images, respectively. Moreover, to guarantee the disentanglement, we propose a bidirectional cyclic training strategy with well-designed constraints. Evaluations demonstrate our method can control pose or expression independently and be used for general video editing. Code: https://github.com/Carlyx/DPE Youxin Pang, Yong Zhang 0034, Weize Quan, Yanbo Fan, Xiaodong Cun, Ying Shan, Dong-Ming Yan 0001 |
CVPR | 4 |
| 2023 | 3D GAN Inversion with Facial Symmetry PriorabstractRecently, a surge of high-quality 3D-aware GANs have been proposed, which leverage the generative power of neural rendering. It is natural to associate 3D GANs with GAN inversion methods to project a real image into the generator's latent space, allowing free-view consistent synthesis and editing, referred as 3D GAN inversion. Although with the facial prior preserved in pre-trained 3D GANs, reconstructing a 3D portrait with only one monocular image is still an ill-pose problem. The straightforward application of 2D GAN inversion methods focuses on texture similarity only while ignoring the correctness of 3D geometry shapes. It may raise geometry collapse effects, especially when reconstructing a side face under an extreme pose. Besides, the synthetic results in novel views are prone to be blurry. In this work, we propose a novel method to promote 3D GAN inversion by introducing facial symmetry prior. We design a pipeline and constraints to make full use of the pseudo auxiliary view obtained via image flipping, which helps obtain a view-consistent and well-structured geometry shape during the inversion process. To enhance texture fidelity in unobserved viewpoints, pseudo labels from depth-guided 3D warping can provide extra supervision. We design constraints to filter out conflict areas for optimization in asymmetric situations. Comprehensive quantitative and qualitative evaluations on image reconstruction and editing demonstrate the superiority of our method. Yong Zhang 0034, Xuan Wang 0009, Tengfei Wang 0002, Xiaoyu Li 0002, Yuan Gong 0002, Yanbo Fan, Xiaodong Cun, Ying Shan, A. Cengiz Öztireli, Yujiu Yang 0001 |
CVPR | 7 |
| 2023 | UCF: Uncovering Common Features for Generalizable Deepfake DetectionabstractDeepfake detection remains a challenging task due to the difficulty of generalizing to new types of forgeries. This problem primarily stems from the overfitting of existing detection methods to forgery-irrelevant features and method-specific patterns. The latter has been rarely studied and not well addressed by previous works. This paper presents a novel approach to address the two types of overfitting issues by uncovering common forgery features. Specifically, we first propose a disentanglement framework that decomposes image information into three distinct components: forgery-irrelevant, method-specific forgery, and common forgery features. To ensure the decoupling of method-specific and common forgery features, a multi-task learning strategy is employed, including a multi-class classification that predicts the category of the forgery method and a binary classification that distinguishes the real from the fake. Additionally, a conditional decoder is designed to utilize forgery features as a condition along with forgery-irrelevant features to generate reconstructed images. Furthermore, a contrastive regularization technique is proposed to encourage the disentanglement of the common and specific forgery features. Ultimately, we only utilize the common forgery features for the purpose of generalizable deepfake detection. Extensive evaluations demonstrate that our framework can perform superior generalization than current state-of-the-art methods. Zhiyuan Yan 0002, Yong Zhang 0034, Yanbo Fan, Baoyuan Wu |
ICCV | 3 |
| 2023 | ToonTalker: Cross-Domain Face ReenactmentabstractWe target cross-domain face reenactment in this paper, i.e., driving a cartoon image with the video of a real person and vice versa. Recently, many works have focused on one-shot talking face generation to drive a portrait with a real video, i.e., within-domain reenactment. Straightforwardly applying those methods to cross-domain animation will cause inaccurate expression transfer, blur effects, and even apparent artifacts due to the domain shift between cartoon and real faces. Only a few works attempt to settle cross-domain face reenactment. The most related work AnimeCeleb [13] requires constructing a dataset with pose vector and cartoon image pairs by animating 3D characters, which makes it inapplicable anymore if no paired data is available. In this paper, we propose a novel method for cross-domain reenactment without paired data. Specifically, we propose a transformer-based framework to align the motions from different domains into a common latent space where motion transfer is conducted via latent code addition. Two domain-specific motion encoders and two learnable motion base memories are used to capture domain properties. A source query transformer and a driving one are exploited to project domain-specific motion to the canonical space. The edited motion is projected back to the domain of the source with a transformer. Moreover, since no paired data is provided, we propose a novel cross-domain training scheme using data from two domains with the designed analogy constraint. Besides, we contribute a cartoon dataset in Disney style. Extensive evaluations demonstrate the superiority of our method over competing methods. Yuan Gong 0002, Yong Zhang 0034, Xiaodong Cun, Yanbo Fan, Xuan Wang 0009, Baoyuan Wu, Yujiu Yang 0001 |
ICCV | 5 |
| 2023 | Enhancing Fine-Tuning based Backdoor Defense with Sharpness-Aware MinimizationabstractBackdoor defense, which aims to detect or mitigate the effect of malicious triggers introduced by attackers, is becoming increasingly critical for machine learning security and integrity. Fine-tuning based on benign data is a natural defense to erase the backdoor effect in a backdoored model. However, recent studies show that, given limited benign data, vanilla fine-tuning has poor defense performance. In this work, we firstly investigate the vanilla fine-tuning process for backdoor mitigation from the neuron weight perspective, and find that backdoor-related neurons are only slightly perturbed in the vanilla fine-tuning process, which explains its poor backdoor defense performance. To enhance the fine-tuning based defense, inspired by the observation that the backdoor-related neurons often have larger weight norms, we propose FT-SAM, a novel backdoor defense paradigm that aims to shrink the norms of backdoor-related neurons by incorporating sharpness-aware minimization with fine-tuning. We demonstrate the effectiveness of our method on several benchmark datasets and network architectures, where it achieves state-of-the-art defense performance, and provide extensive analysis to reveal the FT-SAM’s mechanism. Overall, our work provides a promising avenue for improving the robustness of machine learning models against backdoor attacks. Codes are available at https://github.com/SCLBD/BackdoorBench. Mingli Zhu, Shaokui Wei, Li Shen 0008, Yanbo Fan, Baoyuan Wu |
ICCV | 4 |
| 2023 | Act As You Wish: Fine-Grained Control of Motion Diffusion Model with Hierarchical Semantic GraphsabstractMost text-driven human motion generation methods employ sequential modeling approaches, e.g., transformer, to extract sentence-level text representations automatically and implicitly for human motion synthesis. However, these compact text representations may overemphasize the action names at the expense of other important properties and lack fine-grained details to guide the synthesis of subtly distinct motion. In this paper, we propose hierarchical semantic graphs for fine-grained control over motion generation. Specifically, we disentangle motion descriptions into hierarchical semantic graphs including three levels of motions, actions, and specifics. Such global-to-local structures facilitate a comprehensive understanding of motion description and fine-grained control of motion generation. Correspondingly, to leverage the coarse-to-fine topology of hierarchical semantic graphs, we decompose the text-to-motion diffusion process into three semantic levels, which correspond to capturing the overall motion, local actions, and action specifics. Extensive experiments on two benchmark human motion datasets, including HumanML3D and KIT, with superior performances, justify the efficacy of our method. More encouragingly, by modifying the edge weights of hierarchical semantic graphs, our method can continuously refine the generated motion, which may have a far-reaching impact on the community. Code and pre-trained weights are available at https://github.com/jpthu17/GraphMotion. Peng Jin 0001, Yang Wu 0001, Yanbo Fan, Zhongqian Sun, Wei Yang 0019, Li Yuan 0007 |
NeurIPS | 3 |
| 2023 | Robust Physical-World Attacks on Face Recognition
Xin Zheng 0008, Yanbo Fan, Baoyuan Wu, Yong Zhang 0034, Jue Wang 0001, Shirui Pan |
Pattern Recognit. | 2 |
| 2023 | VDTR: Video Deblurring With TransformerabstractVideo deblurring is still an unsolved problem due to the challenging spatio-temporal modeling process. While existing convolutional neural network (CNN)-based methods show a limited capacity of effective spatial and temporal modeling for video deblurring. This paper presents VDTR, an effective Transformer-based model that makes the first attempt to adapt pure Transformer for video deblurring. VDTR exploits the superior long-range and relation modeling capabilities of Transformer for both spatial and temporal modeling. However, it is challenging to design an appropriate Transformer-based model for video deblurring due to the complicated non-uniform blurs, misalignment across multiple frames and the high computational costs for high-resolution spatial modeling. To address these problems, VDTR advocates performing attention within non-overlapping windows and exploiting the hierarchical structure for long-range dependencies modeling. For frame-level spatial modeling, we propose an encoder-decoder Transformer that utilizes multi-scale features for deblurring. For multi-frame temporal modeling, we adapt Transformer to fuse multiple spatial features efficiently. Compared with CNN-based methods, the proposed method achieves highly competitive results on both synthetic and real-world video deblurring benchmarks, including DVD, GOPRO, REDS and BSD. We hope such a Transformer-based architecture can serve as a powerful alternative baseline for video deblurring and other video restoration tasks. The source code will be available athttps://github.com/ljzycmd/VDTR. Mingdeng Cao, Yanbo Fan, Yong Zhang 0034, Jue Wang 0001, Yujiu Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Fast Adversarial Training With Adaptive Step SizeabstractWhile adversarial training and its variants have shown to be the most effective algorithms to defend against adversarial attacks, their extremely slow training process makes it hard to scale to large datasets like ImageNet. The key idea of recent works to accelerate adversarial training is to substitute multi-step attacks (e.g., PGD) with single-step attacks (e.g., FGSM). However, these single-step methods suffer from catastrophic overfitting, where the accuracy against PGD attack suddenly drops to nearly 0% during training, and the network totally loses its robustness. In this work, we study the phenomenon from the perspective of training instances. We show that catastrophic overfitting is instance-dependent, and fitting instances with larger input gradient norm is more likely to cause catastrophic overfitting. Based on our findings, we propose a simple but effective method, Adversarial Training with Adaptive Step size (ATAS). ATAS learns an instance-wise adaptive step size that is inversely proportional to its gradient norm. Our theoretical analysis shows that ATAS converges faster than the commonly adopted non-adaptive counterparts. Empirically, ATAS consistently mitigates catastrophic overfitting and achieves higher robust accuracy on CIFAR10, CIFAR100, and ImageNet when evaluated on various adversarial budgets. Our code is released at https://github.com/HuangZhiChao95/ATAS. Zhichao Huang 0002, Yanbo Fan, Chen Liu 0027, Yong Zhang 0034, Mathieu Salzmann, Sabine Süsstrunk, Jue Wang 0001 |
IEEE Trans. Image Process. | 2 |
| 2023 | GREEN: A Global Energy Efficiency Maximization Strategy for Multi-UAV Enabled Communication SystemsabstractIn the scenario of limited energy supply, Unmanned Aerial Vehicles (UAVs) enabled communication systems must make efficient use of energy in order to provide long-term service. In this paper, we propose a global energy efficiency maximization (GREEN) strategy for multi-UAV enabled communication systems. In such systems, a group of UAVs communicates with their associated ground terminals (GTs) by using a UAV-enabled interference channel (UAV-IC). In particular, we optimize the UAVs' trajectory control by jointly considering both the communication throughput and the total energy consumption of the whole system. We aim to maximize the global energy efficiency (GEE) of a task for multi-UAV communications, in which the problem is challenging to optimally solve due to its non-convex nature and strongly coupled variables. To tackle this problem, first, we investigate and propose a global energy-efficient optimization problem based on the fly-hover-communicate protocol. Second, we extend our proposed solution from the single UAV-enabled system to multiple UAV-GT pairs cases. In addition, we consider the general scenario in which the UAVs also communicate while flying. Based on the successive convex approximation technique and the path discretization method, the GREEN strategy is designed for optimizing UAV trajectories in this scenario. The simulation results show that the proposed strategy can achieve significantly higher GEE than the benchmark schemes for multi-UAV enabled communications. Na Lin 0001, Yanbo Fan, Liang Zhao 0004, Xiaoming Li 0007, Mohsen Guizani |
IEEE Trans. Mob. Comput. | 2 |
| 2022 | Boosting Black-Box Attack with Partially Transferred Conditional Adversarial DistributionabstractThis work studies black-box adversarial attacks against deep neural networks (DNNs), where the attacker can only access the query feedback returned by the attacked DNN model, while other information such as model parameters or the training datasets are unknown. One promising approach to improve attack performance is utilizing the adversarial transferability between some white-box surrogate models and the target model (i.e., the attacked model). However, due to the possible differences on model architectures and training datasets between surrogate and target models, dubbed “surrogate biases”, the contribution of adversarial transferability to improving the attack performance may be weakened. To tackle this issue, we innovatively propose a black-box attack method by developing a novel mechanism of adversarial transferability, which is robust to the surrogate biases. The general idea is transferring partial parameters of the conditional adversarial distribution (CAD) of surrogate models, while learning the untransferred parameters based on queries to the target model, to keep the flexibility to adjust the CAD of the target model on any new benign sample. Extensive experiments on benchmark datasets and attacking against real-world API demonstrate the superior attack performance of the proposed method. The code will be available at https://github.com/Kira0096/CGATTACK. Baoyuan Wu, Yanbo Fan, Li Liu 0036, Zhifeng Li 0001, Shutao Xia |
CVPR | 3 |
| 2022 | High-Fidelity GAN Inversion for Image Attribute EditingabstractWe present a novel highfidelity generative adversarial network (GAN) inversion framework that enables attribute editing with image-specific details well-preserved (e.g., background, appearance, and illumination). We first analyze the challenges of highfidelity GAN inversion from the perspective of lossy data compression. With a low bitrate latent code, previous works have difficulties in preserving highfidelity details in reconstructed and edited images. Increasing the size of a latent code can improve the accuracy of GAN inversion but at the cost of inferior editability. To improve image fidelity without compromising editability, we propose a distortion consultation approach that employs a distortion map as a reference for highfidelity reconstruction. In the distortion consultation inversion (DCI), the distortion map is first projected to a high-rate latent map, which then complements the basic low-rate latent code with more details via consultation fusion. To achieve high-fidelity editing, we propose an adaptive distortion alignment (ADA) module with a self-supervised training scheme, which bridges the gap between the edited and inversion images. Extensive experiments in the face and car domains show a clear improvement in both inversion and editing quality. The project page is https://tengfei-wang.github.io/HFGI/. Tengfei Wang 0002, Yong Zhang 0034, Yanbo Fan, Jue Wang 0001, Qifeng Chen 0001 |
CVPR | 3 |
| 2022 | A Large-Scale Multiple-objective Method for Black-box Attack Against Object Detection
Siyuan Liang 0004, Longkang Li, Yanbo Fan, Xiaojun Jia, Jingzhi Li 0002, Baoyuan Wu, Xiaochun Cao |
ECCV (4) | 3 |
| 2022 | StyleHEAT: One-Shot High-Resolution Editable Talking Face Generation via Pre-trained StyleGAN
Yong Zhang 0034, Xiaodong Cun, Mingdeng Cao, Yanbo Fan, Xuan Wang 0009, Qingyan Bai, Baoyuan Wu, Jue Wang 0001, Yujiu Yang 0001 |
ECCV (17) | 5 |
| 2022 | HyP2 Loss: Beyond Hypersphere Metric Space for Multi-label Image RetrievalabstractImage retrieval has become an increasingly appealing technique with broad multimedia application prospects, where deep hashing serves as the dominant branch towards low storage and efficient retrieval. In this paper, we carried out in-depth investigations on metric learning in deep hashing for establishing a powerful metric space in multi-label scenarios, where the pair loss suffers high computational overhead and converge difficulty, while the proxy loss is theoretically incapable of expressing the profound label dependencies and exhibits conflicts in the constructed hypersphere space. To address the problems, we propose a novel metric learning framework with Hybrid Proxy-Pair Loss (HyP$^2$ Loss) that constructs an expressive metric space with efficient training complexity w.r.t. the whole dataset. The proposed HyP$^2$ Loss focuses on optimizing the hypersphere space by learnable proxies and excavating data-to-data correlations of irrelevant pairs, which integrates sufficient data correspondence of pair-based methods and high-efficiency of proxy-based methods. Extensive experiments on four standard multi-label benchmarks justify the proposed method outperforms the state-of-the-art, is robust among different hash bits and achieves significant performance gains with a faster, more stable convergence speed. Our code is available at https://github.com/JerryXu0129/HyP2-Loss. Chengyin Xu, Zenghao Chai, Zhengzhuo Xu, Chun Yuan 0003, Yanbo Fan, Jue Wang 0001 |
ACM Multimedia | 5 |
| 2022 | Boosting the Transferability of Adversarial Attacks with Reverse Adversarial PerturbationabstractDeep neural networks (DNNs) have been shown to be vulnerable to adversarial examples, which can produce erroneous predictions by injecting imperceptible perturbations. In this work, we study the transferability of adversarial examples, which is significant due to its threat to real-world applications where model architecture or parameters are usually unknown. Many existing works reveal that the adversarial examples are likely to overfit the surrogate model that they are generated from, limiting its transfer attack performance against different target models. To mitigate the overfitting of the surrogate model, we propose a novel attack method, dubbed reverse adversarial perturbation (RAP). Specifically, instead of minimizing the loss of a single adversarial point, we advocate seeking adversarial example located at a region with unified low loss value, by injecting the worst-case perturbation (the reverse adversarial perturbation) for each step of the optimization procedure. The adversarial attack with RAP is formulated as a min-max bi-level optimization problem. By integrating RAP into the iterative process for attacks, our method can find more stable adversarial examples which are less sensitive to the changes of decision boundary, mitigating the overfitting of the surrogate model. Comprehensive experimental comparisons demonstrate that RAP can significantly boost adversarial transferability. Furthermore, RAP can be naturally combined with many existing black-box attack techniques, to further boost the transferability. When attacking a real-world image recognition system, Google Cloud Vision API, we obtain 22% performance improvement of targeted attacks over the compared method. Our codes are available at https://github.com/SCLBD/TransferattackRAP. Zeyu Qin, Yanbo Fan, Li Shen 0008, Yong Zhang 0034, Jue Wang 0001, Baoyuan Wu |
NeurIPS | 2 |
| 2022 | Stability Analysis and Generalization Bounds of Adversarial TrainingabstractIn adversarial machine learning, deep neural networks can fit the adversarial examples on the training dataset but have poor generalization ability on the test set. This phenomenon is called robust overfitting, and it can be observed when adversarially training neural nets on common datasets, including SVHN, CIFAR-10, CIFAR-100, and ImageNet. In this paper, we study the robust overfitting issue of adversarial training by using tools from uniform stability. One major challenge is that the outer function (as a maximization of the inner function) is nonsmooth, so the standard technique (e.g., Hardt et al., 2016) cannot be applied. Our approach is to consider $\eta$-approximate smoothness: we show that the outer function satisfies this modified smoothness assumption with $\eta$ being a constant related to the adversarial perturbation $\epsilon$. Based on this, we derive stability-based generalization bounds for stochastic gradient descent (SGD) on the general class of $\eta$-approximate smooth functions, which covers the adversarial loss. Our results suggest that robust test accuracy decreases in $\epsilon$ when $T$ is large, with a speed between $\Omega(\epsilon\sqrt{T})$ and $\mathcal{O}(\epsilon T)$. This phenomenon is also observed in practice. Additionally, we show that a few popular techniques for adversarial training (\emph{e.g.,} early stopping, cyclic learning rate, and stochastic weight averaging) are stability-promoting in theory. Jiancong Xiao, Yanbo Fan, Ruoyu Sun 0001, Jue Wang 0001, Zhi-Quan Luo |
NeurIPS | 2 |
| 2022 | Average Top-k Aggregate Loss for Supervised LearningabstractIn this work, we introduce theaverage top-$k$k($\mathrm {AT}_k$) loss, which is the average over the$k$largest individual losses over a training data, as a new aggregate loss for supervised learning. We show that the$\mathrm {AT}_k$loss is a natural generalization of the two widely used aggregate losses, namely the average loss and the maximum loss. Yet, the$\mathrm {AT}_k$loss can better adapt to different data distributions because of the extra flexibility provided by the different choices of$k$. Furthermore, it remains a convex function over all individual losses and can be combined with different types of individual loss without significant increase in computation. We then provide interpretations of the$\mathrm {AT}_k$loss from the perspective of the modification of individual loss and robustness to training data distributions. We further study the classification calibration of the$\mathrm {AT}_k$loss and the error bounds of$\mathrm {AT}_k$-SVM model. We demonstrate the applicability of minimum average top-$k$learning for supervised learning problems including binary/multi-class classification and regression, using experiments on both synthetic and real datasets. Siwei Lyu, Yanbo Fan, Yiming Ying, Bao-Gang Hu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Semi-supervised robust training with generalized perturbed neighborhood
Yiming Li 0004, Baoyuan Wu, Yanbo Fan, Yong Jiang 0001, Zhifeng Li 0001, Shutao Xia |
Pattern Recognit. | 4 |
| 2021 | Parallel Rectangle Flip Attack: A Query-based Black-box Attack against Object DetectionabstractObject detection has been widely used in many safety- critical tasks, such as autonomous driving. However, its vulnerability to adversarial examples has not been sufficiently studied, especially under the practical scenario of black-box attacks, where the attacker can only access the query feedback of predicted bounding-boxes and top- 1 scores returned by the attacked model. Compared with black-box attack to image classification, there are two main challenges in black-box attack to detection. Firstly, even if one bounding-box is successfully attacked, another sub- optimal bounding-box may be detected near the attacked bounding-box. Secondly, there are multiple bounding- boxes, leading to very high attack cost. To address these challenges, we propose a Parallel Rectangle Flip Attack (PRFA) via random search. We explain the difference between our method with other attacks in Fig. 1. Specifically, we generate perturbations in each rectangle patch to avoid sub-optimal detection near the attacked region. Besides, utilizing the observation that adversarial perturbations mainly locate around objects’ contours and critical points under white-box attacks, the search space of attacked rectangles is reduced to improve the attack efficiency. Moreover, we develop a parallel mechanism of attacking multiple rectangles simultaneously to further accelerate the attack process. Extensive experiments demonstrate that our method can effectively and efficiently attack various popular object detectors, including anchor-based and anchor- free, and generate transferable adversarial examples. Siyuan Liang 0004, Baoyuan Wu, Yanbo Fan, Xingxing Wei 0001, Xiaochun Cao |
ICCV | 3 |
| 2021 | DAE-GAN: Dynamic Aspect-aware GAN for Text-to-Image SynthesisabstractText-to-image synthesis refers to generating an image from a given text description, the key goal of which lies in photo realism and semantic consistency. Previous methods usually generate an initial image with sentence embedding and then refine it with fine-grained word embedding. Despite the significant progress, the ‘aspect’ information (e.g., red eyes) contained in the text, referring to several words rather than a word that depicts ‘a particular part or feature of something’, is often ignored, which is highly helpful for synthesizing image details. How to make better utilization of aspect information in text-to-image synthesis still remains an unresolved challenge. To address this problem, in this paper, we propose a Dynamic Aspect-awarE GAN (DAE-GAN) that represents text information comprehensively from multiple granularities, including sentence-level, word-level, and aspect-level. Moreover, inspired by human learning behaviors, we develop a novel Aspect-aware Dynamic Re-drawer (ADR) for image refinement, in which an Attended Global Refinement (AGR) module and an Aspect-aware Local Refinement (ALR) module are alternately employed. AGR utilizes word-level embedding to globally enhance the previously generated image, while ALR dynamically employs aspect-level embedding to refine image details from a local perspective. Finally, a corresponding matching loss function is designed to ensure the text-image semantic consistency at different levels. Extensive experiments on two well-studied and publicly available datasets (i.e., CUB-200 and COCO) demonstrate the superiority and rationality of our method. Shulan Ruan, Yong Zhang 0034, Kun Zhang 0015, Yanbo Fan, Fan Tang, Qi Liu 0003, Enhong Chen |
ICCV | 4 |
| 2021 | Random Noise Defense Against Query-Based Black-Box AttacksabstractThe query-based black-box attacks have raised serious threats to machine learning models in many real applications. In this work, we study a lightweight defense method, dubbed Random Noise Defense (RND), which adds proper Gaussian noise to each query. We conduct the theoretical analysis about the effectiveness of RND against query-based black-box attacks and the corresponding adaptive attacks. Our theoretical results reveal that the defense performance of RND is determined by the magnitude ratio between the noise induced by RND and the noise added by the attackers for gradient estimation or local search. The large magnitude ratio leads to the stronger defense performance of RND, and it's also critical for mitigating adaptive attacks. Based on our analysis, we further propose to combine RND with a plausible Gaussian augmentation Fine-tuning (RND-GF). It enables RND to add larger noise to each query while maintaining the clean accuracy to obtain a better trade-off between clean accuracy and defense performance. Additionally, RND can be flexibly combined with the existing defense methods to further boost the adversarial robustness, such as adversarial training (AT). Extensive experiments on CIFAR-10 and ImageNet verify our theoretical findings and the effectiveness of RND and RND-GF. Zeyu Qin, Yanbo Fan, Hongyuan Zha, Baoyuan Wu |
NeurIPS | 2 |
| 2020 | 3D Single-Person Concurrent Activity Detection Using Stacked Relation NetworkabstractWe aim to detect real-world concurrent activities performed by a single person from a streaming 3D skeleton sequence. Different from most existing works that deal with concurrent activities performed by multiple persons that are seldom correlated, we focus on concurrent activities that are spatio-temporally or causally correlated and performed by a single person. For the sake of generalization, we propose an approach based on a decompositional design to learn a dedicated feature representation for each activity class. To address the scalability issue, we further extend the class-level decompositional design to the postural-primitive level, such that each class-wise representation does not need to be extracted by independent backbones, but through a dedicated weighted aggregation of a shared pool of postural primitives. There are multiple interdependent instances deriving from each decomposition. Thus, we propose Stacked Relation Networks (SRN), with a specialized relation network for each decomposition, so as to enhance the expressiveness of instance-wise representations via the inter-instance relationship modeling. SRN achieves state-of-the-art performance on a public dataset and a newly collected dataset. The relation weights within SRN are interpretable among the activity contexts. The new dataset and code are available at https://github.com/weiyi1991/UA_Concurrent/ Yi Wei 0006, Wenbo Li 0001, Yanbo Fan, Linghan Xu, Ming-Ching Chang, Siwei Lyu |
AAAI | 3 |
| 2020 | Sparse Adversarial Attack via Perturbation Factorization
Yanbo Fan, Baoyuan Wu, Tuanhui Li, Yong Zhang 0034, Zhifeng Li 0001, Yujiu Yang 0001 |
ECCV (22) | 1 |
| 2019 | Exact Adversarial Attack to Image Captioning via Structured Output Learning With Latent VariablesabstractIn this work, we study the robustness of a CNN+RNN based image captioning system being subjected to adversarial noises. We propose to fool an image captioning system to generate some targeted partial captions for an image polluted by adversarial noises, even the targeted captions are totally irrelevant to the image content. A partial caption indicates that the words at some locations in this caption are observed, while words at other locations are not restricted. It is the first work to study exact adversarial attacks of targeted partial captions. Due to the sequential dependencies among words in a caption, we formulate the generation of adversarial noises for targeted partial captions as a structured output learning problem with latent variables. Both the generalized expectation maximization algorithm and structural SVMs with latent variables are then adopted to optimize the problem. The proposed methods generate very successful attacks to three popular CNN+RNN based image captioning models. Furthermore, the proposed attack methods are used to understand the inner mechanism of image captioning systems, providing the guidance to further improve automatic image captioning systems towards human captioning. Yan Xu 0009, Baoyuan Wu, Fumin Shen, Yanbo Fan, Yong Zhang 0034, Heng Tao Shen, Wei Liu 0005 |
CVPR | 4 |
| 2019 | Compressing Convolutional Neural Networks via Factorized Convolutional FiltersabstractThis work studies the model compression for deep convolutional neural networks (CNNs) via filter pruning. The workflow of a traditional pruning consists of three sequential stages: pre-training the original model, selecting the pre-trained filters via ranking according to a manually designed criterion (e.g., the norm of filters), and learning the remained filters via fine-tuning. Most existing works follow this pipeline and focus on designing different ranking criteria for filter selection. However, it is difficult to control the performance due to the separation of filter selection and filter learning. In this work, we propose to conduct filter selection and filter learning simultaneously, in a unified model. To this end, we define a factorized convolutional filter (FCF), consisting of a standard real-valued convolutional filter and a binary scalar, as well as a dot-product operator between them. We train a CNN model with factorized convolutional filters (CNN-FCF) by updating the standard filter using back-propagation, while updating the binary scalar using the alternating direction method of multipliers (ADMM) based optimization method. With this trained CNN-FCF model, we only keep the standard filters corresponding to the 1-valued scalars, while all other filters and all binary scalars are discarded, to obtain a compact CNN model. Extensive experiments on CIFAR-10 and ImageNet demonstrate the superiority of the proposed method over state-of-the-art filter pruning methods. Tuanhui Li, Baoyuan Wu, Yujiu Yang 0001, Yanbo Fan, Yong Zhang 0034, Wei Liu 0005 |
CVPR | 4 |
| 2019 | Context-Aware Feature and Label Fusion for Facial Action Unit Intensity Estimation With Partially Labeled DataabstractFacial action unit (AU) intensity estimation is a fundamental task for facial behaviour analysis. Most previous methods use a whole face image as input for intensity prediction. Considering that AUs are defined according to their corresponding local appearance, a few patch-based methods utilize image features of local patches. However, fusion of local features is always performed via straightforward feature concatenation or summation. Besides, these methods require fully annotated databases for model learning, which is expensive to acquire. In this paper, we propose a novel weakly supervised patch-based deep model on basis of two types of attention mechanisms for joint intensity estimation of multiple AUs. The model consists of a feature fusion module and a label fusion module. And we augment attention mechanisms of these two modules with a learnable task-related context, as one patch may play different roles in analyzing different AUs and each AU has its own temporal evolution rule. The context-aware feature fusion module is used to capture spatial relationships among local patches while the context-aware label fusion module is used to capture the temporal dynamics of AUs. The latter enables the model to be trained on a partially annotated database. Experimental evaluations on two benchmark expression databases demonstrate the superior performance of the proposed method. Yong Zhang 0034, Haiyong Jiang, Baoyuan Wu, Yanbo Fan |
ICCV | 4 |
| 2017 | Self-Paced Learning: An Implicit Regularization PerspectiveabstractSelf-paced learning (SPL) mimics the cognitive mechanism of humans and animals that gradually learns from easy to hard samples. One key issue in SPL is to obtain better weighting strategy that is determined by the minimizer function. Existing methods usually pursue this by artificially designing the explicit form of SPL regularizer. In this paper, we study a group of new regularizer (named self-paced implicit regularizer) that is deduced from robust loss function. Based on the convex conjugacy theory, the minimizer function for self-paced implicit regularizer can be directly learned from the latent loss function, while the analytic form of the regularizer can be even unknown. A general framework (named SPL-IR) for SPL is developed accordingly. We demonstrate that the learning procedure of SPL-IR is associated with latent robust loss functions, thus can provide some theoretical insights for its working mechanism. We further analyze the relation between SPL-IR and half-quadratic optimization and provide a group of self-paced implicit regularizer. Finally, we implement SPL-IR to both supervised and unsupervised tasks, and experimental results corroborate our ideas and demonstrate the correctness and effectiveness of implicit regularizers. Yanbo Fan, Ran He 0001, Jian Liang 0001, Bao-Gang Hu |
AAAI | 1 |
| 2017 | Learning with Average Top-k LossabstractIn this work, we introduce the average top-$k$ (\atk) loss as a new ensemble loss for supervised learning. The \atk loss provides a natural generalization of the two widely used ensemble losses, namely the average loss and the maximum loss. Furthermore, the \atk loss combines the advantages of them and can alleviate their corresponding drawbacks to better adapt to different data distributions. We show that the \atk loss affords an intuitive interpretation that reduces the penalty of continuous and convex individual losses on correctly classified data. The \atk loss can lead to convex optimization problems that can be solved effectively with conventional sub-gradient based method. We further study the Statistical Learning Theory of \matk by establishing its classification calibration and statistical consistency of \matk which provide useful insights on the practical choice of the parameter $k$. We demonstrate the applicability of \matk learning combined with different individual loss functions for binary and multi-class classification and regression using synthetic and real datasets. Yanbo Fan, Siwei Lyu, Yiming Ying, Bao-Gang Hu |
NIPS | 1 |