EDBT 2026 Demo / reviewers in the wild / expert
Wei Wang 0108
dblp:35/7092-108
· DBLP profile ↗
71ranked-venue papers
14as first author
40since 2021 · last 2026
0000-0002-5477-1017ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 48 · 10 first-author · 23 since 2021Artificial intelligence and machine learning · 39 · 10 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 since 2021Computer networks · 4 · 4 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Unified Masked Jigsaw Puzzle Framework for Vision and Language ModelsabstractIn federated learning, Transformer, as a popular architecture, faces critical challenges in defending against gradient attacks and improving model performance in both Computer Vision (CV) and Natural Language Processing (NLP) tasks. It has been revealed that the gradient of Position Embeddings (PEs) in Transformer contains sufficient information, which can be used to reconstruct the input data. To mitigate this issue, we introduce a Masked Jigsaw Puzzle (MJP) framework. MJP starts with random token shuffling to break the token order, and then a learnable unknown (unk) position embedding is used to mask out the PEs of the shuffled tokens. In this manner, the local spatial information which is encoded in the position embeddings is disrupted, and the models are forced to learn feature representations that are less reliant on the local spatial information. Notably, with the careful use of MJP, we can not only improve models' robustness against gradient attacks, but also boost their performance in both vision and text application scenarios, such as classification for images (e.g., ImageNet-1 K) and sentiment analysis for text (e.g., Yelp and Amazon). Experimental results suggest that MJP is a unified framework for different Transformer-based models in both vision and language tasks. Weixin Ye, Wei Wang 0108, Yue Song 0002, Bin Ren 0005, Wei Bi, Rita Cucchiara, Nicu Sebe |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | High-Fidelity 3D Facial Avatar Synthesis With Controllable Fine-Grained ExpressionsabstractFacial expression editing methods can be mainly categorized into two types based on their architectures: 2D-based and 3D-based methods. The former lacks 3D face modeling capabilities, making it difficult to edit 3D factors effectively. The latter has demonstrated superior performance in generating high-quality and view-consistent renderings using single-view 2D face images. Although these methods have successfully used animatable models to control facial expressions, they still have limitations in achieving precise control over fine-grained expressions. To address this issue, in this paper, we propose a novel approach by simultaneously refining both the latent code of a pretrained 3D-Aware GAN model for texture editing and the expression code of the driven 3DMM model for mesh editing. Specifically, we introduce a Dual Mappers module, comprising Texture Mapper and Emotion Mapper, to learn the transformations of the given latent code for textures and the expression code for meshes, respectively. To optimize the Dual Mappers, we propose a Text-Guided Optimization method, leveraging a CLIP-based objective function with expression text prompts as targets, while integrating a SubSpace Projection mechanism to project the text embedding to the expression subspace such that we can have more precise control over fine-grained expressions. Extensive experiments and comparative analyses demonstrate the effectiveness and superiority of our proposed method. Yikang He, Jichao Zhang, Wei Wang 0108, Nicu Sebe, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Sharpness-Aware Fine-Tuning for OOD DetectionabstractThe out-of-distribution (OOD) detection task is crucial for the real-world deployment of machine learning models. In this paper, we propose to study the problem from the perspective of Sharpness-aware Minimization (SAM). Compared with traditional optimizers such as SGD, SAM can better improve the model performance and generalization ability, and this is closely related to OOD detection. Therefore, instead of using SGD, we propose to fine-tune the model with SAM, and observe that the score distributions of in-distribution (ID) data and OOD data are pushed away from each other. Besides, with our carefully designed loss, the fine-tuning process is very time-efficient. The OOD performance improvement can be observed after fine-tuning the model within 1 epoch. Moreover, our method is very flexible and can be used to improve the performance of different OOD detection methods. Extensive experiments have demonstrated that our method achieves state-of-the-art performance on various OOD benchmarks across different architectures. Moreover, comprehensive ablation studies and theoretical analyses are discussed to support the empirical results. Wei Wang 0108, Yao Zhao 0001, Nicu Sebe, Yue Song 0002 |
IEEE Trans. Image Process. | 2 |
| 2026 | IA2GNN: Imbalance-Aware Adaptive Graph Construction for Multi-Modal Image FusionabstractImage fusion aims to integrate multi-modal data, enhancing information completeness while reducing redundancy. However, image fusion faces the challenge of information imbalance. This imbalance is reflected in both intra-image variations, where different regions exhibit significant differences in information density and importance, and cross-modal inconsistencies, where corresponding regions in different modalities contribute unequally to the fused result. Most existing image fusion methods, such as models based on CNN and window attention mechanism, adopt a fixed approach that collects the same amount of neighboring information for each pixel. However, this fixed approach lacks adaptability to local variations in information density, ultimately limiting fusion performance. To this end, we propose IA2GNN, an imbalance-aware adaptive graph neural network designed for image fusion. To tackle intra- and cross-modal information imbalance, we leverage the flexibility of graph structures to dynamically adjust node connections during feature extraction and fusion. Specifically, we adjust the connectivity of nodes in the graph to allow regions with rich details to establish more connections, thereby enhancing the extraction and preservation of key information in these regions. For less informative regions, we reduce connections to prevent over-modeling. In this way, we optimize information distribution, achieving a balance between detail retention and information integration in the fusion results. Experimental results demonstrate that the proposed model outperforms baseline methods across multiple image fusion tasks, achieving up to +0.058 improvement in SSIM and +0.351 gain in MI. Yepeng Tang, Chunjie Zhang 0001, Wei Wang 0108, Xiaolong Zheng 0001, Yao Zhao 0001 |
IEEE Trans. Multim. | 4 |
| 2026 | Speech-Driven 3D Facial Animation with Natural Head MovementsabstractSpeech-driven 3D facial animation has been studied for a long time, yet many challenges still prevent achieving truly natural results. One major issue is that current methods often overlook head motion during speech. To address this, we conducted a preliminary investigation and identified several key obstacles: the scarcity of 3D facial animation datasets with head motion and the limitations of regular sequence prediction models (e.g., Transformer, LSTM, and GRU), which are designed for discrete dynamic sequences and do not align well with continuous 3D head motion sequences. To solve these issues, we propose a head motion prediction module. This module uses audio and the initial motion state of the mesh to predict head motion. By employing ordinary differential equations (ODE) to model continuous dynamic sequences, it predicts head movements that closely resemble real head motion, making the animation more realistic and natural. Additionally, recognizing that the head motion generation given audio is not a one-to-one mapping problem, we introduce a noise module to help the head motion prediction module generate varied head motions given the same audio input. We also observed that previous facial animation methods primarily focus on generating vertices for the mouth region but use a single model to generate the entire face. This approach wastes some of the model’s fitting capacity on other regions. To solve this problem, we propose a cascaded mesh generation module that uses two modules to separately generate the vertex of mouth region and other facial regions. Extensive experiments and a perceptual user study show that our approach outperforms existing methods and produces relatively natural head motion. Kuiyuan Sun, Jichao Zhang, Wei Wang 0108, Yao Zhao 0001, Nicu Sebe |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2026 | Stable-Hair V2: Real-World Hair Transfer via Multiple-View Diffusion ModelabstractWhile diffusion-based methods have shown impressive capabilities in capturing diverse and complex hairstyles, their ability to generate consistent and high-quality multi-view outputs-crucial for real-world applications such as digital humans and virtual avatars-remains underexplored. In this paper, we propose Stable-Hair v2, a novel diffusion-based multi-view hair transfer framework. To the best of our knowledge, this is the first work to leverage multiple-view diffusion models for robust, high-fidelity, and view-consistent hair transfer across multiple perspectives. We introduce a comprehensive multi-view training data generation pipeline to generate high-quality triplet data, including bald images, reference hairstyles, and view-aligned source-bald pairs. Our multi-view hair transfer model integrates polar-azimuth embeddings for pose conditioning and temporal attention layers to ensure smooth transitions between views. To optimize this model, we design a novel multi-stage training strategy consisting of Pose-Controllable Latent IdentityNet training, Hair Extractor training, and Temporal Attention training. Extensive experiments demonstrate that our method accurately transfers detailed and realistic hairstyles to source subjects while achieving seamless and consistent results across views, significantly outperforming existing methods and establishing a new benchmark in multi-view hair transfer. Kuiyuan Sun, Jichao Zhang, Wei Wang 0108, Nicu Sebe, Yao Zhao 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | Unsupervised Region-Based Image Editing of Denoising Diffusion ModelsabstractAlthough diffusion models have achieved remarkable success in the field of image generation, their latent space remains under-explored. Current methods for identifying semantics within latent space often rely on external supervision, such as textual information and segmentation masks. In this paper, we propose a method to identify semantic attributes in the latent space of pre-trained diffusion models without any further training. By projecting the Jacobian of the targeted semantic region into a low-dimensional subspace which is orthogonal to the non-masked regions, our approach facilitates precise semantic discovery and control over local masked areas, eliminating the need for annotations. We conducted extensive experiments across multiple datasets and various architectures of diffusion models, achieving state-of-the-art performance. In particular, for some specific face attributes, the performance of our proposed method even surpasses that of supervised approaches, demonstrating its superior ability in editing local image properties. Zixiang Li, Yue Song 0002, Renshuai Tao, Xiaohong Jia 0002, Yao Zhao 0001, Wei Wang 0108 |
AAAI | 6 |
| 2025 | Unlocking the Potential of Lightweight Quantized Models for Deepfake DetectionabstractDeepfake detection is increasingly crucial due to the rapid rise of AI-generated content. Existing methods achieve high performance relying on computationally intensive large models, making real-time detection on resource-constrained edge devices challenging. Given that deepfake detection is a binary classification task, there is potential for model compression and acceleration. In this paper, we propose a low-bit quantization framework for lightweight and efficient deepfake detection. The Connected Quantized Block extracts common forgery features via the quantized path and retains method-specific textures through the shortcut connections. Additionally, the Shifted Logarithmic Redistribution Quantizer mitigates information loss in near-zero domains by unfolding the unbalanced activations, enabling finer quantization granularity. Comprehensive experiments demonstrate that this new framework significantly reduces 10.8x computational costs and 12.4x storage requirements while maintaining high detection performance, even surpassing SOTA methods using less than 5% FLOPs, paving the way for efficient deepfake detection in resource-limited scenarios. Renshuai Tao, Ziheng Qin, Chuangchuang Tan, Jiakai Wang, Wei Wang 0108 |
IJCAI | 6 |
| 2025 | DCI: Dual-Conditional Inversion for Boosting Diffusion-Based Image EditingabstractDiffusion models have achieved remarkable success in image generation and editing tasks. Inversion within these models aims to recover the latent noise representation for a real or generated image, enabling reconstruction, editing, and other downstream tasks. However, to date, most inversion approaches suffer from an intrinsic trade-off between reconstruction accuracy and editing flexibility. This limitation arises from the difficulty of maintaining both semantic alignment and structural consistency during the inversion process. In this work, we introduce **Dual-Conditional Inversion (DCI)**, a novel framework that jointly conditions on the source prompt and reference image to guide the inversion process. Specifically, DCI formulates the inversion process as a dual-condition fixed-point optimization problem, minimizing both the latent noise gap and the reconstruction error under the joint guidance. This design anchors the inversion trajectory in both semantic and visual space, leading to more accurate and editable latent representations. Our novel setup brings new understanding to the inversion process. Extensive experiments demonstrate that DCI achieves state-of-the-art performance across multiple editing tasks, significantly improving both reconstruction quality and editing precision. Furthermore, we also demonstrate that our method achieves strong results in reconstruction tasks, implying a degree of robustness and generalizability approaching the ultimate goal of the inversion process. Our codes are available at: [https://github.com/Lzxhh/Dual-Conditional-Inversion](https://github.com/Lzxhh/Dual-Conditional-Inversion) Zixiang Li, Wei Wang 0108, Chuangchuang Tan, Yunchao Wei, Yao Zhao 0001 |
NeurIPS | 3 |
| 2025 | LEDNet: a multimodal foundation model for robust deepfake detection
Renshuai Tao, Shijie Tang, Haotong Qin, Wei Wang 0108, Yunchao Wei, Yao Zhao 0001 |
Sci. China Inf. Sci. | 4 |
| 2025 | Context perturbation: A Consistent alignment approach for Domain Adaptive Semantic Segmentation
Meiqin Liu 0002, Yao Zhao 0001, Wei Wang 0108, Yunchao Wei |
Comput. Vis. Image Underst. | 5 |
| 2025 | NPVForensics: Learning VA correlations in non-critical phoneme-viseme regions for deepfake detection
Yu Chen 0049, Yang Yu 0039, Haoliang Li, Wei Wang 0108, Yao Zhao 0001 |
Image Vis. Comput. | 5 |
| 2025 | Semantic-Aware Guidance for Blind Super-Resolution of Remote Sensing ImagesabstractUnlike traditional super-resolution (SR) methods that rely on fixed degradation models, blind SR (BSR) methods can capture the complex processes introduced by factors such as sensor noise and platform motion in real-world remote sensing imagery. While most BSR methods effectively remove degradation from low-resolution (LR) images, they often struggle to preserve high-frequency details, leading to reduced reconstruction accuracy. To address this issue, we propose the semantic-aware guidance BSR (SGBSR) network, which leverages semantic information to guide the entire restoration process, enabling more accurate reconstruction. Specifically, we design a semantic extractor that utilizes powerful pretrained visual models to capture rich semantic information from LR images, which is then integrated into the SR network. To further enhance the network’s ability to handle complex degradations, we introduce an implicit estimation method. Subsequently, the semantic information and degradation representations are, respectively, incorporated into the SR network through the semantic-aware block (SaB) and the degradation-aware block (DaB). Experiments on both synthetic and real-world LR images demonstrate that our method achieves superior reconstruction accuracy. Siyuan Hao, Wei Wang 0108 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2025 | RankFeat&RankWeight: Rank-1 Feature/Weight Removal for Out-of-Distribution DetectionabstractThe task of out-of-distribution (OOD) detection is crucial for deploying machine learning models in real-world settings. In this paper, we observe that the singular value distributions of the in-distribution (ID) and OOD features are quite different: the OOD feature matrix tends to have a larger dominant singular value than the ID feature, and the class predictions of OOD samples are largely determined by it. This observation motivates us to proposeRankFeat, a simple yet effectivepost hocapproach for OOD detection by removing the rank-1 matrix composed of the largest singular value and the associated singular vectors from the high-level feature.RankFeatachievesstate-of-the-artperformance and reduces the average false positive rate (FPR95) by 17.90% compared with the previous best method. The success ofRankFeatmotivates us to investigate whether a similar phenomenon would exist in the parameter matrices of neural networks. We thus proposeRankWeightwhich removes the rank-1 weight from the parameter matrices of a single deep layer. OurRankWeightis alsopost hocand only requires computing the rank-1 matrix once. As a standalone approach,RankWeighthas very competitive performance against other methods across various backbones. Moreover,RankWeightenjoys flexible compatibility with a wide range of OOD detection methods. The combination ofRankWeightandRankFeatrefreshes the newstate-of-the-artperformance, achieving the FPR95 as low as 16.13% on the ImageNet-1k benchmark. Extensive ablation studies and comprehensive theoretical analyses are presented to support the empirical results. Yue Song 0002, Wei Wang 0108, Nicu Sebe |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Towards Extensible Detection of AI-Generated Images via Content-Agnostic Adapter-Based Category-Aware Incremental LearningabstractThe rapid evolution of image generation techniques has benefited several fields, but it has also given rise to security concerns. As countermeasures, a series of AI-generated image detection methods have been developed successfully. However, existing methods exhibit an inefficiency in handling the continual emergence of new generative models. To address this issue, we formulate the detection of AI-generated images in an extensible manner using an adapter-based domain incremental learning framework. Specifically, we first investigate the global consistency property of generation artifacts and design a content-agnostic adapter equipped on a vision transformer to extract common forensic features, where a token-level shuffling strategy is constructed for the dual-stream comparison to mitigate the fitting to specific image content. Then, motivated by the compactness of real images and the diversity of fake images due to their inherent generation processes, an asymmetric category-aware domain alignment method is designed to reduce the domain shift arisen from different generators. Finally, a multi-view knowledge distillation module, considering both point-to-point and structure-to-structure forensic knowledge, is devised to alleviate catastrophic forgetting. Experiments are conducted on several protocols using various image generators, and experimental results verify the superiority of our method compared to state-of-the-art methods for extensible detection. Shuai Tang 0001, Peisong He, Haoliang Li, Wei Wang 0108, Xinghao Jiang, Yao Zhao 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | SAGNet: Decoupling Semantic-Agnostic Artifacts From Limited Training Data for Robust Generalization in Deepfake DetectionabstractDeepfake detection presents a significant challenge, particularly when the available training data is constrained to a limited set of semantic categories—a common and realistic scenario. In deepfake detection, the training labels typically indicate whether an image is real or fake, without specifying the semantic content, such as object classes. Moreover, we cannot know in advance the object categories present in an image to be detected. Ideally, a deepfake detection model should perform consistently across different semantic categories during inference, irrespective of the content. However, existing methods often exhibit significant performance bias between seen and unseen classes, struggling to generalize effectively. To address this issue, we propose Semantic-AGnostic artifact Network (SAGNet), an innovative and efficient approach designed to decouple semantic-agnostic artifacts from content-specific distributions in the training data. Our method eliminates semantic-specific biases, ensuring that the model focuses on universal artifacts related to image authenticity rather than content-dependent features. By employing this decoupling strategy, SAGNet greatly enhances the model’s generalization capacity, even when trained on limited data. Remarkably, through experiments, we demonstrate that SAGNet achieves performance comparable to models trained with 10 times more data, despite being trained on only 2 classes (comparing SAGNet trained on 2 classes in Table I with Ojha [1] trained on 20 categories in Table IV). Furthermore, through extensive experiments, we show that SAGNet’s improvements are not only evident across different semantic categories but also extend to various generative methods, including multiple GAN-based and diffusion-based models. This cross-method generalization emphasizes SAGNet’s versatility and effectiveness in diverse generative scenarios. Overall, our method represents a significant advancement in deepfake detection, particularly in realistic situations where the training data is limited. The code is released at https://github.com/rstao-bjtu/SAGNet/. Renshuai Tao, Chuangchuang Tan, Huan Liu 0030, Jiakai Wang, Haotong Qin, Yakun Chang, Wei Wang 0108, Yao Zhao 0001 |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2024 | Diffusion for Natural Image Matting
Yihan Hu 0004, Yiheng Lin 0002, Wei Wang 0108, Yao Zhao 0001, Yunchao Wei, Humphrey Shi |
ECCV (57) | 3 |
| 2024 | Multi-Modal Driven Pose-Controllable Talking Head GenerationabstractTalking head, driving a source image to generate a talking video using other modality information, has made great progress in recent years. However, there are two main issues: 1) These methods are designed to utilize a single modality of information. 2) Most methods cannot control head pose. To address these problems, we propose a novel framework that can utilize multi-modal information to generate a talking head video, while achieving arbitrary head pose control by a movement sequence. Specifically, first, to extend driving information to multiple modalities, multi-modal information is encoded to a unified semantic latent space to generate expression parameters. Secondly, to disentangle attributes, the 3D Morphable Model (3DMM) is utilized to obtain identity information from the source image, and translation and rotation information from the target image. Thirdly, to control head pose and mouth shape, the source image is warped by a motion field generated by the expression parameter, translation parameter, and angle parameter. Finally, all the above parameters are utilized to render a landmark map, and the warped source image is combined with the landmark map to generate a delicate talking head video. Our experimental results demonstrate that our proposed method is capable of achieving state-of-the-art performance in terms of visual quality, lip-audio synchronization, and head pose control. Kuiyuan Sun, Xiaolong Li 0001, Yao Zhao 0001, Wei Wang 0108 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2023 | Masked Jigsaw Puzzle: A Versatile Position Embedding for Vision TransformersabstractPosition Embeddings (PEs), an arguably indispensable component in Vision Transformers (ViTs), have been shown to improve the performance of ViTs on many vision tasks. However, PEs have a potentially high risk of privacy leakage since the spatial information of the input patches is exposed. This caveat naturally raises a series of interesting questions about the impact of PEs on accuracy, privacy, prediction consistency, etc. To tackle these issues, we propose a Masked Jigsaw Puzzle (MJP) position embedding method. In particular, MJP first shuffles the selected patches via our block-wise random jigsaw puzzle shuffle algorithm, and their corresponding PEs are occluded. Meanwhile, for the non-occluded patches, the PEs remain the original ones but their spatial relation is strengthened via our dense absolute localization regressor. The experimental results reveal that 1) PEs explicitly encode the 2D spatial relationship and lead to severe privacy leakage problems under gradient inversion attack; 2) Training ViTs with the naively shuffled patches can alleviate the problem, but it harms the accuracy; 3) Under a certain shuffle ratio, the proposed MJP not only boosts the performance and robustness on large-scale datasets (i.e., ImageNet-1K and ImageNet-C, -A/O) but also improves the privacy preservation ability under typical gradient attacks by a large margin. The source code and trained models are available at https://github.com/yhlleo/MJP. Bin Ren 0005, Yue Song 0002, Wei Bi, Rita Cucchiara, Nicu Sebe, Wei Wang 0108 |
CVPR | 7 |
| 2023 | PI-Trans: Parallel-Convmlp and Implicit-Transformation Based Gan for Cross-View Image TranslationabstractFor semantic-guided cross-view image translation, it is crucial to learn where to sample pixels from the source view image and where to reallocate them guided by the target view semantic map, especially when there is little overlap or drastic view difference between the source and target images. Hence, one not only needs to encode the long- range dependencies among pixels in both the source view image and target view semantic map but also needs to translate these learned dependencies. To this end, we propose a novel generative adversarial network, PI-Trans, which mainly consists of a novel Parallel-ConvMLP module and an Implicit Transformation module at multiple semantic levels. Extensive experimental results show that PI-Trans achieves the best qualitative and quantitative performance by a large margin compared to the state-of-the-art methods on two challenging datasets. The source code is available at https://github.com/Amazingren/PI-Trans. Bin Ren 0005, Hao Tang 0005, Yiming Wang 0002, Xia Li 0005, Wei Wang 0108, Nicu Sebe |
ICASSP | 5 |
| 2023 | Householder Projector for Unsupervised Latent Semantics DiscoveryabstractGenerative Adversarial Networks (GANs), especially the recent style-based generators (StyleGANs), have versatile semantics in the structured latent space. Latent semantics discovery methods emerge to move around the latent code such that only one factor varies during the traversal. Recently, an unsupervised method proposed a promising direction to directly use the eigenvectors of the projection matrix that maps latent codes to features as the interpretable directions. However, one overlooked fact is that the projection matrix is non-orthogonal and the number of eigenvectors is too large. The non-orthogonality would entangle semantic attributes in the top few eigenvectors, and the large dimensionality might result in meaningless variations among the directions even if the matrix is orthogonal. To avoid these issues, we propose Householder Projector, a flexible and general low-rank orthogonal matrix representation based on Householder transformations, to parameterize the projection matrix. The orthogonality guarantees that the eigenvectors correspond to disentangled interpretable semantics, while the low-rank property encourages that each identified direction has meaningful variations. We integrate our projector into pre-trained StyleGAN2/StyleGAN3 and evaluate the models on several benchmarks. Within only 1% of the original training steps for fine-tuning, our projector helps StyleGANs to discover more disentangled and precise semantic attributes without sacrificing image fidelity. Code is publicly available via https://github.com/KingJamesSong/HouseholderGAN. Yue Song 0002, Jichao Zhang, Nicu Sebe, Wei Wang 0108 |
ICCV | 4 |
| 2023 | On the Eigenvalues of Global Covariance Pooling for Fine-Grained Visual RecognitionabstractThe Fine-Grained Visual Categorization (FGVC) is challenging because the subtle inter-class variations are difficult to be captured. One notable research line uses the Global Covariance Pooling (GCP) layer to learn powerful representations with second-order statistics, which can effectively model inter-class differences. In our previous conference paper, we show that truncating small eigenvalues of the GCP covariance can attain smoother gradient and improve the performance on large-scale benchmarks. However, on fine-grained datasets, truncating the small eigenvalues would make the model fail to converge. This observation contradicts the common assumption that the small eigenvalues merely correspond to the noisy and unimportant information. Consequently, ignoring them should have little influence on the performance. To diagnose this peculiar behavior, we propose two attribution methods whose visualizations demonstrate that the seemingly unimportant small eigenvalues are crucial as they are in charge of extracting the discriminative class-specific features. Inspired by this observation, we propose a network branch dedicated to magnifying the importance of small eigenvalues. Without introducing any additional parameters, this branch simply amplifies the small eigenvalues and achieves state-of-the-art performances of GCP methods on three fine-grained benchmarks. Furthermore, the performance is also competitive against other FGVC approaches on larger datasets. Code is available at https://github.com/KingJamesSong/DifferentiableSVD. Yue Song 0002, Nicu Sebe, Wei Wang 0108 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Fast Differentiable Matrix Square Root and Inverse Square RootabstractComputing the matrix square root and its inverse in a differentiable manner is important in a variety of computer vision tasks. Previous methods either adopt the Singular Value Decomposition (SVD) to explicitly factorize the matrix or use the Newton-Schulz iteration (NS iteration) to derive the approximate solution. However, both methods are not computationally efficient enough in either the forward pass or the backward pass. In this paper, we propose two more efficient variants to compute the differentiable matrix square root and the inverse square root. For the forward propagation, one method is to use Matrix Taylor Polynomial (MTP), and the other method is to use Matrix Padé Approximants (MPA). The backward gradient is computed by iteratively solving the continuous-time Lyapunov equation using the matrix sign function. A series of numerical tests show that both methods yield considerable speed-up compared with the SVD or the NS iteration. Moreover, we validate the effectiveness of our methods in several real-world applications, including de-correlated batch normalization, second-order vision transformer, global covariance pooling for large-scale and fine-grained recognition, attentive covariance pooling for video recognition, and neural style transfer. The experiments demonstrate that our methods can also achieve competitive and even slightly better performances. Code is available at https://github.com/KingJamesSong/FastDifferentiableMatSqrt. Yue Song 0002, Nicu Sebe, Wei Wang 0108 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Orthogonal SVD Covariance Conditioning and Latent DisentanglementabstractInserting an SVD meta-layer into neural networks is prone to make the covariance ill-conditioned, which could harm the model in the training stability and generalization abilities. In this article, we systematically study how to improve the covariance conditioning by enforcing orthogonality to the Pre-SVD layer. Existing orthogonal treatments on the weights are first investigated. However, these techniques can improve the conditioning but would hurt the performance. To avoid such a side effect, we propose the Nearest Orthogonal Gradient (NOG) and Optimal Learning Rate (OLR). The effectiveness of our methods is validated in two applications: decorrelated Batch Normalization (BN) and Global Covariance Pooling (GCP). Extensive experiments on visual recognition demonstrate that our methods can simultaneously improve covariance conditioning and generalization. The combinations with orthogonal weight can further boost the performance. Moreover, we show that our orthogonality techniques can benefit generative models for better latent disentanglement through a series of experiments on various benchmarks. Code is available at: https://github.com/KingJamesSong/OrthoImproveCond. Yue Song 0002, Nicu Sebe, Wei Wang 0108 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Disentangle Saliency Detection into Cascaded Detail Modeling and Body FillingabstractSalient object detection has been long studied to identify the most visually attractive objects in images/videos. Recently, a growing amount of approaches have been proposed, all of which rely on the contour/edge information to improve detection performance. The edge labels are either put into the loss directly or used as extra supervision. The edge and body can also be learned separately and then fused afterward. Both methods either lead to high prediction errors near the edge or cannot be trained in an end-to-end manner. Another problem is that existing methods may fail to detect objects of various sizes due to the lack of efficient and effective feature fusion mechanisms. In this work, we propose to decompose the saliency detection task into two cascaded sub-tasks, i.e., detail modeling and body filling. Specifically, detail modeling focuses on capturing the object edges by supervision of explicitly decomposed detail label that consists of the pixels that are nested on the edge and near the edge. Then the body filling learns the body part that will be filled into the detail map to generate more accurate saliency map. To effectively fuse the features and handle objects at different scales, we have also proposed two novel multi-scale detail attention and body attention blocks for precise detail and body modeling. Experimental results show that our method achieves state-of-the-art performances on six public datasets. Yue Song 0002, Hao Tang 0005, Nicu Sebe, Wei Wang 0108 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Bidirectional Transformer GAN for Long-term Human Motion PredictionabstractThe mainstream motion prediction methods usually focus on short-term prediction, and their predicted long-term motions often fall into an average pose, i.e., the freezing forecasting problem [ 27 ]. To mitigate this problem, we propose a novel Bidirectional Transformer-based Generative Adversarial Network (BiTGAN) for long-term human motion prediction. The bidirectional setup leads to consistent and smooth generation in both forward and backward directions. Besides, to make full use of the history motions, we split them into two parts. The first part is fed to the Transformer encoder in our BiTGAN while the second part is used as the decoder input. This strategy can alleviate the exposure problem [ 37 ]. Additionally, to better maintain both the local (i.e., frame-level pose) and global (i.e., video-level semantic) similarities between the predicted motion sequence and the real one, the soft dynamic time warping (Soft-DTW) loss is introduced into the generator. Finally, we utilize a dual-discriminator to distinguish the predicted sequence at both frame and sequence levels. Extensive experiments on the public Human3.6M dataset demonstrate that our proposed BiTGAN achieves state-of-the-art performance on long-term (4 s ) human motion prediction, and reduces the average error of all actions by 4%. Mengyi Zhao, Hao Tang 0005, Pan Xie, Shuling Dai, Nicu Sebe, Wei Wang 0108 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2022 | Long Term Motion Prediction Using KeyposesabstractLong term human motion prediction is essential in safety-critical applications such as human-robot interaction and autonomous driving. In this paper we show that to achieve long term forecasting, predicting human pose at every time instant is unnecessary. Instead, it is more effective to predict a few keyposes and approximate intermediate ones by interpolating the keyposes. We demonstrate that our approach enables us to predict realistic motions for up to 5 seconds in the future, which is far longer than the typical 1 second encountered in the literature. Furthermore, because we model future keyposes probabilistically, we can generate multiple plausible future motions by sampling at inference time. Over this extended time period, our predictions are more realistic, more diverse and better preserve the motion dynamics than those state-of-the-art methods yield. Sena Kiciroglu, Wei Wang 0108, Mathieu Salzmann, Pascal Fua |
3DV | 2 |
| 2022 | Batch-Efficient EigenDecomposition for Small and Medium Matrices
Yue Song 0002, Nicu Sebe, Wei Wang 0108 |
ECCV (23) | 3 |
| 2022 | Improving Covariance Conditioning of the SVD Meta-layer by Orthogonality
Yue Song 0002, Nicu Sebe, Wei Wang 0108 |
ECCV (24) | 3 |
| 2022 | 3D-Aware Semantic-Guided Generative Model for Human Synthesis
Jichao Zhang, Enver Sangineto, Hao Tang 0005, Aliaksandr Siarohin, Zhun Zhong, Nicu Sebe, Wei Wang 0108 |
ECCV (15) | 7 |
| 2022 | Fast Differentiable Matrix Square Root
Yue Song 0002, Nicu Sebe, Wei Wang 0108 |
ICLR | 3 |
| 2022 | RankFeat: Rank-1 Feature Removal for Out-of-distribution DetectionabstractThe task of out-of-distribution (OOD) detection is crucial for deploying machine learning models in real-world settings. In this paper, we observe that the singular value distributions of the in-distribution (ID) and OOD features are quite different: the OOD feature matrix tends to have a larger dominant singular value than the ID feature, and the class predictions of OOD samples are largely determined by it. This observation motivates us to propose RankFeat, a simple yet effective post hoc approach for OOD detection by removing the rank-1 matrix composed of the largest singular value and the associated singular vectors from the high-level feature. RankFeat achieves state-of-the-art performance and reduces the average false positive rate (FPR95) by 17.90% compared with the previous best method. Extensive ablation studies and comprehensive theoretical analyses are presented to support the empirical results. Yue Song 0002, Nicu Sebe, Wei Wang 0108 |
NeurIPS | 3 |
| 2022 | Spectral and Spatial Feature Fusion for Hyperspectral Image ClassificationabstractCompared with traditional images, hyperspectral images (HSI) not only have spatial information, but also have rich spectral information. However, the mainstream hyperspectral image classification (HIC) methods are all based on Convolutional Neural Network (CNN), which has great advantages in extracting spatial features, but it has certain limitations in dealing with spectral continuous sequence information. Therefore Transformer which is good at processing sequences, has also been gradually applied to HIC. Besides, Since HSI are typical three-dimensional structures, we believe that the correlation of the three dimensions is also an important information. So in order to fully extract the spectral spatial information, as well as the correlation of the three dimensions. we propose a spectral and spatial feature fusion module (i.e., TransCNN) for HIC. TransCNN consists of CNNs and a Transformer. The former is in charge of mining the spatial and spectral information from different dimensions, while the latter not only undertakes the most critical fusion but also captures the deeper relationship characteristics. We transpose the data to extract features and their correlation through three CNNs branches. we believe that these feature maps still have deep spectral information. Therefore, we have embedded them into one-dimensional vectors and use Transformer’s Encoder to extract features. However, some information will be lost when embedding into one-dimensional vectors. Therefore we use Decoder, which has been ignored in the field of vision, to fuse the features before passing Encoder and the features after extracted by Encoders. Two kinds of features are fused by Decoder, and the obtained information is finally input into the classifier for classification. Experimental results on real HSIs show that the proposed architecture can achieve competitive performance compared with the state-of-the-art methods. Siyuan Hao, Yufeng Xia, Lijian Zhou, Yuanxin Ye, Wei Wang 0108 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | Robust Differentiable SVDabstractEigendecomposition of symmetric matrices is at the heart of many computer vision algorithms. However, the derivatives of the eigenvectors tend to be numerically unstable, whether using the SVD to compute them analytically or using the Power Iteration (PI) method to approximate them. This instability arises in the presence of eigenvalues that are close to each other. This makes integrating eigendecomposition into deep networks difficult and often results in poor convergence, particularly when dealing with large matrices. While this can be mitigated by partitioning the data into small arbitrary groups, doing so has no theoretical basis and makes it impossible to exploit the full power of eigendecomposition. In previous work, we mitigated this using SVD during the forward pass and PI to compute the gradients during the backward pass. However, the iterative deflation procedure required to compute multiple eigenvectors using PI tends to accumulate errors and yield inaccurate gradients. Here, we show that the Taylor expansion of the SVD gradient is theoretically equivalent to the gradient obtained using PI without relying in practice on an iterative process and thus yields more accurate gradients. We demonstrate the benefits of this increased accuracy for image classification and style transfer. Wei Wang 0108, Zheng Dang, Yinlin Hu, Pascal Fua, Mathieu Salzmann |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Quasi-Equilibrium Feature Pyramid Network for Salient Object DetectionabstractModern saliency detection models are based on the encoder-decoder framework and they use different strategies to fuse the multi-level features between the encoder and decoder to boost representation power. Motivated by recent work in implicit modelling, we propose to introduce an implicit function to simulate the equilibrium state of the feature pyramid at infinite depths. We question the existence of the ideal equilibrium and thus propose a quasi-equilibrium model by taking the first-order derivative into the black-box root solver using Taylor expansion. It models more realistic convergence states and significantly improves the network performance. We also propose a differentiable edge extractor that directly extracts edges from the saliency masks. By optimizing the extracted edges, the generated saliency masks are naturally optimized on contour constraints and the non-deterministic predictions are removed. We evaluate the proposed methodology on five public datasets and extensive experiments show that our method achieves new state-of-the-art performances on six metrics across datasets. Yue Song 0002, Hao Tang 0005, Mengyi Zhao, Nicu Sebe, Wei Wang 0108 |
IEEE Trans. Image Process. | 5 |
| 2022 | Unsupervised High-Resolution Portrait Gaze Correction and AnimationabstractThis paper proposes a gaze correction and animation method for high-resolution, unconstrained portrait images, which can be trained without the gaze angle and the head pose annotations. Common gaze-correction methods usually require annotating training data with precise gaze, and head pose information. Solving this problem using an unsupervised method remains an open problem, especially for high-resolution face images in the wild, which are not easy to annotate with gaze and head pose labels. To address this issue, we first create two new portrait datasets: CelebGaze ( 256 ×256 ) and high-resolution CelebHQGaze ( 512 ×512 ). Second, we formulate the gaze correction task as an image inpainting problem, addressed using a Gaze Correction Module (GCM) and a Gaze Animation Module (GAM). Moreover, we propose an unsupervised training strategy, i.e., Synthesis-As-Training, to learn the correlation between the eye region features and the gaze angle. As a result, we can use the learned latent space for gaze animation with semantic interpolation in this space. Moreover, to alleviate both the memory and the computational costs in the training and the inference stage, we propose a Coarse-to-Fine Module (CFM) integrated with GCM and GAM. Extensive experiments validate the effectiveness of our method for both the gaze correction and the gaze animation tasks in both low and high-resolution face datasets in the wild and demonstrate the superiority of our method with respect to the state of the art. Jichao Zhang, Hao Tang 0005, Enver Sangineto, Peng Wu 0014, Yan Yan 0002, Nicu Sebe, Wei Wang 0108 |
IEEE Trans. Image Process. | 8 |
| 2021 | Smoothing the Disentangled Latent Style Space for Unsupervised Image-to-Image TranslationabstractImage-to-Image (I2I) multi-domain translation models are usually evaluated also using the quality of their semantic interpolation results. However, state-of-the-art models frequently show abrupt changes in the image appearance during interpolation, and usually perform poorly in interpolations across domains. In this paper, we propose a new training protocol based on three specific losses which help a translation network to learn a smooth and disentangled latent style space in which: 1) Both intra- and inter-domain interpolations correspond to gradual changes in the generated images and 2) The content of the source image is better preserved during the translation. Moreover, we propose a novel evaluation metric to properly measure the smoothness of latent style space of I2I translation models. The proposed method can be plugged in existing translation approaches, and our extensive experiments on different datasets show that it can significantly boost the quality of the generated images and the graduality of the interpolations. Enver Sangineto, Linchao Bao, Haoxian Zhang, Nicu Sebe, Bruno Lepri, Wei Wang 0108, Marco De Nadai |
CVPR | 8 |
| 2021 | Why Approximate Matrix Square Root Outperforms Accurate SVD in Global Covariance Pooling?abstractGlobal Covariance Pooling (GCP) aims at exploiting the second-order statistics of the convolutional feature. Its effectiveness has been demonstrated in boosting the classification performance of Convolutional Neural Networks (CNNs). Singular Value Decomposition (SVD) is used in GCP to compute the matrix square root. However, the approximate matrix square root calculated using Newton-Schulz iteration [14] outperforms the accurate one computed via SVD [15]. We empirically analyze the reason behind the performance gap from the perspectives of data precision and gradient smoothness. Various remedies for computing smooth SVD gradients are investigated. Based on our observation and analyses, a hybrid training protocol is proposed for SVD-based GCP meta-layers such that competitive performances can be achieved against Newton-Schulz iteration. Moreover, we propose a new GCP meta-layer that uses SVD in the forward pass, and Padé approximants in the backward propagation to compute the gradients. The proposed meta-layer has been integrated into different CNN models and achieves state-of-the-art performances on both large-scale and fine-grained datasets. Yue Song 0002, Nicu Sebe, Wei Wang 0108 |
ICCV | 3 |
| 2021 | Audio-Visual Event Localization via Recursive Fusion by Joint Co-AttentionabstractThe major challenge in audio-visual event localization task lies in how to fuse information from multiple modalities effectively. Recent works have shown that the attention mechanism is beneficial to the fusion process. In this paper, we propose a novel joint attention mechanism with multi-modal fusion methods for audio-visual event localization. Particularly, we present a concise yet valid architecture that effectively learns representations from multiple modalities in a joint manner. Initially, visual features are combined with auditory features and then turned into joint representations. Next, we make use of the joint representations to attend to visual features and auditory features, respectively. With the help of this joint co-attention, new visual and auditory features are produced, and thus both features can enjoy the mutually improved benefits from each other. It is worth noting that the joint co-attention unit is recursive meaning that it can be performed multiple times for obtaining better joint representations progressively. Extensive experiments on the public AVE dataset have shown that the proposed method achieves significantly better results than the state-of-the-art methods. Bin Duan 0004, Hao Tang 0005, Wei Wang 0108, Ziliang Zong, Guowei Yang 0001, Yan Yan 0002 |
WACV | 3 |
| 2021 | Geometry-Aware Deep Recurrent Neural Networks for Hyperspectral Image ClassificationabstractVariants of deep networks have been widely used for hyperspectral image (HSI)-classification tasks. Among them, in recent years, recurrent neural networks (RNNs) have attracted considerable attention in the remote sensing community. However, complex geometries cannot be learned easily by the traditional recurrent units [e.g., long short-term memory (LSTM) and gated recurrent unit (GRU)]. In this article, we propose a geometry-aware deep recurrent neural network (Geo-DRNN) for HSI classification. We build this network upon two modules: a U-shaped network (U-Net) and RNNs. We first input the original HSI patches to the U-Net, which can be trained with very few images and obtain a preliminary classification result. We then add RNNs on the top of the U-Net so as to mimic the human brain to refine continuously the output-classification map. However, instead of using the traditional dot product in each gate of the RNNs, we introduce a Net-Gated GRU that increases the nonlinear representation power. Finally, we use a pretrained ResNet as a regularizer to improve further the ability of the proposed network to describe complex geometries. To this end, we construct a geometry-aware ResNet loss, which leverages the pretrained ResNet's knowledge about the different structures in the real world. Our experimental results on real HSIs and road topology images demonstrate that our approach outperforms the state-of-the-art classification methods and can learn complex geometries. Siyuan Hao, Wei Wang 0108, Mathieu Salzmann |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | Motion Prediction Using Temporal Inception Module
Tim Lebailly, Sena Kiciroglu, Mathieu Salzmann, Pascal Fua, Wei Wang 0108 |
ACCV (2) | 5 |
| 2020 | Single-Stage 6D Object Pose EstimationabstractMost recent 6D pose estimation frameworks first rely on a deep network to establish correspondences between 3D object keypoints and 2D image locations and then use a variant of a RANSAC-based Perspective-n-Point (PnP) algorithm. This two-stage process, however, is suboptimal: First, it is not end-to-end trainable. Second, training the deep network relies on a surrogate loss that does not directly reflect the final 6D pose estimation task. In this work, we introduce a deep architecture that directly regresses 6D poses from correspondences. It takes as input a group of candidate correspondences for each 3D keypoint and accounts for the fact that the order of the correspondences within each group is irrelevant, while the order of the groups, that is, of the 3D keypoints, is fixed. Our architecture is generic and can thus be exploited in conjunction with existing correspondence-extraction networks so as to yield single-stage 6D pose estimation frameworks. Our experiments demonstrate that these single-stage frameworks consistently outperform their two-stage counterparts in terms of both accuracy and speed. Yinlin Hu, Pascal Fua, Wei Wang 0108, Mathieu Salzmann |
CVPR | 3 |
| 2020 | Cascade Attention Guided Residue Learning GAN for Cross-Modal TranslationabstractSince we were babies, we intuitively develop the ability to correlate the input from different cognitive sensors such as vision, audio, and text. However, in machine learning, this cross-modal learning is a nontrivial task because different modalities have no homogeneous properties. Previous works discover that there should be bridges among different modalities. From a neurology and psychology perspective, humans have the capacity to link one modality with another one, e.g., associating a picture of a bird with the only hearing of its singing and vice versa. Is it possible for machine learning algorithms to recover the scene given the audio signal? In this paper, we propose a novel Cascade Attention-Guided Residue GAN (CAR-GAN), aiming at reconstructing the scenes given the corresponding audio signals. Particularly, we present a residue module to mitigate the gap between different modalities progressively. Moreover, a cascade attention guided network with a novel classification loss function is designed to tackle the cross-modal learning task. Our model keeps consistency in the high-level semantic label domain and is able to balance two different modalities. The experimental results demonstrate that our model achieves the state-of-the-art cross-modal audio-visual generation on the challenging Sub-URMP dataset. Bin Duan 0004, Wei Wang 0108, Hao Tang 0005, Hugo Latapie, Yan Yan 0002 |
ICPR | 2 |
| 2020 | Dual In-painting Model for Unsupervised Gaze Correction and Animation in the WildabstractWe address the problem of unsupervised gaze correction in the wild, presenting a solution that works without the need of precise annotations of the gaze angle and the head pose. We created a new dataset called CelebAGaze consisting of two domains X, Y, where the eyes are either staring at the camera or somewhere else. Our method consists of three novel modules: the Gaze Correction module(GCM), the Gaze Animation module(GAM), and the Pretrained Autoencoder module (PAM). Specifically, GCM and GAM separately train a dual in-painting network using data from the domain X for gaze correction and data from the domain Y for gaze animation. Additionally, a Synthesis-As-Training method is proposed when training GAM to encourage the features encoded from the eye region to be correlated with the angle information, resulting in gaze animation achieved by interpolation in the latent space. To further preserve the identity information e.g., eye shape, iris color, we propose the PAM with an Autoencoder, which is based on Self-Supervised mirror learning where the bottleneck features are angle-invariant and which works as an extra input to the dual in-painting models. Extensive experiments validate the effectiveness of the proposed method for gaze correction and gaze animation in the wild and demonstrate the superiority of our approach in producing more compelling results than state-of-the-art baselines. Our code, the pretrained models and supplementary results are available at:https://github.com/zhangqianhui/GazeAnimation. Jichao Zhang, Hao Tang 0005, Wei Wang 0108, Yan Yan 0002, Enver Sangineto, Nicu Sebe |
ACM Multimedia | 4 |
| 2020 | Learning How to Smile: Expression Video Generation With Conditional Adversarial Recurrent NetsabstractWhile several research studies have focused on analyzing human behavior and, in particular, emotional signals from visual data, the problem of synthesizing face video sequences with specific attributes (e.g. age, facial expressions) received much less attention. This paper proposes a novel deep generative model able to produce face videos from a given image of a neutral face and a label indicating a specific facial expression, e.g. spontaneous smile. Our framework consists of two main building blocks: an image generator and a frame sequence generator. The image generator is implemented as a deep neural model which combines generative adversarial networks and variational auto-encoders, while the sequence generator is a label-conditioned recurrent neural network. In the proposed framework, given as input a neural face and a label, the sequence generator outputs a set of hidden representations with smooth transitions corresponding to video frames. Then, the image generator is used to decode the hidden representations into the actual face images. To impose that the net generates videos consistent with the given label, a novel identity adversarial loss is proposed. Our experimental results demonstrate the effectiveness of the framework and the advantage of introducing an adversarial component into recurrent models for face video generation. Wei Wang 0108, Xavier Alameda-Pineda, Dan Xu 0002, Elisa Ricci 0001, Nicu Sebe |
IEEE Trans. Multim. | 1 |
| 2019 | Attribute-Guided Sketch GenerationabstractFacial attributes are important since they provide a detailed description and determine the visual appearance of human faces. In this paper, we aim at converting a face image to a sketch while simultaneously generating facial attributes. To this end, we propose a novel Attribute-Guided Sketch Generative Adversarial Network (ASGAN) which is an end-to-end framework and contains two pairs of generators and discriminators, one of which is used to generate faces with attributes while the other one is employed for image-to-sketch translation. The two generators form a W-shaped network (W-net) and they are trained jointly with a weight-sharing constraint. Additionally, we also propose two novel discriminators, the residual one focusing on attribute generation and the triplex one helping to generate realistic looking sketches. To validate our model, we have created a new large dataset with 8,804 images, named the Attribute Face Photo & Sketch (AFPS) dataset which is the first dataset containing attributes associated to face sketch images. The experimental results demonstrate that the proposed network (i) generates more photo-realistic faces with sharper facial attributes than baselines and (ii) has good generalization capability on different generative tasks. Hao Tang 0005, Xinya Chen, Wei Wang 0108, Dan Xu 0002, Jason J. Corso, Nicu Sebe, Yan Yan 0002 |
FG | 3 |
| 2019 | Recurrent U-Net for Resource-Constrained SegmentationabstractState-of-the-art segmentation methods rely on very deep networks that are not always easy to train without very large training datasets and tend to be relatively slow to run on standard GPUs. In this paper, we introduce a novel recurrent U-Net architecture that preserves the compactness of the original U-Net [33], while substantially increasing its performance to the point where it outperforms the state of the art on several benchmarks. We will demonstrate its effectiveness for several tasks, including hand segmentation, retina vessel segmentation, and road segmentation. We also introduce a large-scale dataset for hand segmentation. Wei Wang 0108, Kaicheng Yu, Joachim Hugonot, Pascal Fua, Mathieu Salzmann |
ICCV | 1 |
| 2019 | Expression Conditional Gan for Facial Expression-to-Expression TranslationabstractIn this paper, we focus on the facial expression translation task and propose a novel Expression Conditional GAN (ECGAN) which can learn the mapping from one image domain to another one based on an additional expression attribute. The proposed ECGAN is a generic framework and is applicable to different expression generation tasks where specific facial expression can be easily controlled by the conditional attribute label. Besides, we introduce a novel face mask loss to reduce the influence of background changing. Moreover, we propose an entire framework for facial expression generation and recognition in the wild, which consists of two modules, i.e., generation and recognition. Finally, we evaluate our framework on several public face datasets in which the subjects have different races, illumination, occlusion, pose, color, content and background conditions. Even though these datasets are very diverse, both the qualitative and quantitative results demonstrate that our approach is able to generate facial expressions accurately and robustly. Hao Tang 0005, Wei Wang 0108, Songsong Wu, Xinya Chen, Dan Xu 0002, Nicu Sebe, Yan Yan 0002 |
ICIP | 2 |
| 2019 | Cycle In Cycle Generative Adversarial Networks for Keypoint-Guided Image GenerationabstractIn this work, we propose a novel Cycle In Cycle Generative Adversarial Network (C2GAN) for the task of keypoint-guided image generation. The proposed C2GAN is a cross-modal framework exploring a joint exploitation of the keypoint and the image data in an interactive manner. C2GAN contains two different types of generators, i.e., keypoint-oriented generator and image-oriented generator. Both of them are mutually connected in an end-to-end learnable fashion and explicitly form three cycled sub-networks, i.e., one image generation cycle and two keypoint generation cycles. Each cycle not only aims at reconstructing the input domain, and also produces useful output involving in the generation of another cycle. By so doing, the cycles constrain each other implicitly, which provides complementary information from the two different modalities and brings extra supervision across cycles, thus facilitating more robust optimization of the whole network. Extensive experimental results on two publicly available datasets, i.e., Radboud Faces and Market-1501, demonstrate that our approach is effective to generate more photo-realistic images compared with state-of-the-art models. Hao Tang 0005, Dan Xu 0002, Gaowen Liu, Wei Wang 0108, Nicu Sebe, Yan Yan 0002 |
ACM Multimedia | 4 |
| 2019 | Backpropagation-Friendly EigendecompositionabstractEigendecomposition (ED) is widely used in deep networks. However, the backpropagation of its results tends to be numerically unstable, whether using ED directly or approximating it with the Power Iteration method, particularly when dealing with large matrices. While this can be mitigated by partitioning the data in small and arbitrary groups, doing so has no theoretical basis and makes its impossible to exploit the power of ED to the full. In this paper, we introduce a numerically stable and differentiable approach to leveraging eigenvectors in deep networks. It can handle large matrices without requiring to split them. We demonstrate the better robustness of our approach over standard ED and PI for ZCA whitening, an alternative to batch normalization, and for PCA denoising, which we introduce as a new normalization strategy for deep networks, aiming to further denoise the network's features. Wei Wang 0108, Zheng Dang, Yinlin Hu, Pascal Fua, Mathieu Salzmann |
NeurIPS | 1 |
| 2019 | Deep Micro-Dictionary Learning and Coding NetworkabstractIn this paper, we propose a novel Deep Micro-Dictionary Learning and Coding Network (DDLCN). DDLCN has most of the standard deep learning layers (pooling, fully, connected, input/output, etc.) but the main difference is that the fundamental convolutional layers are replaced by novel compound dictionary learning and coding layers. The dictionary learning layer learns an over-complete dictionary for the input training data. At the deep coding layer, a locality constraint is added to guarantee that the activated dictionary bases are close to each other. Next, the activated dictionary atoms are assembled together and passed to the next compound dictionary learning and coding layers. In this way, the activated atoms in the first layer can be represented by the deeper atoms in the second dictionary. Intuitively, the second dictionary is designed to learn the fine-grained components which are shared among the input dictionary atoms. In this way, a more informative and discriminative low-level representation of the dictionary atoms can be obtained. We empirically compare the proposed DDLCN with several dictionary learning methods and deep learning architectures. The experimental results on four popular benchmark datasets demonstrate that the proposed DDLCN achieves competitive results compared with state-of-the-art approaches. Hao Tang 0005, Heng Wei, Wei Xiao 0002, Wei Wang 0108, Dan Xu 0002, Yan Yan 0002, Nicu Sebe |
WACV | 4 |
| 2019 | Recurrent Face Aging with Hierarchical AutoRegressive MemoryabstractModeling the aging process of human faces is important for cross-age face verification and recognition. In this paper, we propose a Recurrent Face Aging (RFA) framework which takes as input a single image and automatically outputs a series of aged faces. The hidden units in the RFA are connected autoregressively allowing the framework to age the person by referring to the previous aged faces. Due to the lack of labeled face data of the same person captured in a long range of ages, traditional face aging models split the ages into discrete groups and learn a one-step face transformation for each pair of adjacent age groups. Since human face aging is a smooth progression, it is more appropriate to age the face by going through smooth transitional states. In this way, the intermediate aged faces between the age groups can be generated. Towards this target, we employ a recurrent neural network whose recurrent module is a hierarchical triple-layer gated recurrent unit which functions as an autoencoder. The bottom layer of the module encodes the input to a latent representation, and the top layer decodes the representation to a corresponding aged face. The experimental results demonstrate the effectiveness of our framework. Wei Wang 0108, Yan Yan 0002, Zhen Cui 0001, Jiashi Feng, Shuicheng Yan, Nicu Sebe |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | Dual Generator Generative Adversarial Networks for Multi-domain Image-to-Image Translation
Hao Tang 0005, Dan Xu 0002, Wei Wang 0108, Yan Yan 0002, Nicu Sebe |
ACCV (1) | 3 |
| 2018 | Structured Attention Guided Convolutional Neural Fields for Monocular Depth EstimationabstractRecent works have shown the benefit of integrating Conditional Random Fields (CRFs) models into deep architectures for improving pixel-level prediction tasks. Following this line of research, in this paper we introduce a novel approach for monocular depth estimation. Similarly to previous works, our method employs a continuous CRF to fuse multi-scale information derived from different layers of a front-end Convolutional Neural Network (CNN). Differently from past works, our approach benefits from a structured attention model which automatically regulates the amount of information transferred between corresponding features at different scales. Importantly, the proposed attention model is seamlessly integrated into the CRF, allowing end-to-end training of the entire architecture. Our extensive experimental evaluation demonstrates the effectiveness of the proposed method which is competitive with previous methods on the KITTI benchmark and outperforms the state of the art on the NYU Depth V2 dataset. Dan Xu 0002, Wei Wang 0108, Hao Tang 0005, Hong Liu 0008, Nicu Sebe, Elisa Ricci 0001 |
CVPR | 2 |
| 2018 | Every Smile Is Unique: Landmark-Guided Diverse Smile GenerationabstractEach smile is unique: one person surely smiles in different ways (e.g. closing/opening the eyes or mouth). Given one input image of a neutral face, can we generate multiple smile videos with distinctive characteristics? To tackle this one-to-many video generation problem, we propose a novel deep learning architecture named Conditional Multi-Mode Network (CMM-Net). To better encode the dynamics of facial expressions, CMM-Net explicitly exploits facial landmarks for generating smile sequences. Specifically, a variational auto-encoder is used to learn a facial landmark embedding. This single embedding is then exploited by a conditional recurrent network which generates a landmark embedding sequence conditioned on a specific expression (e.g. spontaneous smile). Next, the generated landmark embeddings are fed into a multi-mode recurrent landmark generator, producing a set of landmark sequences still associated to the given smile class but clearly distinct from each other. Finally, these landmark sequences are translated into face videos. Our experimental results demonstrate the effectiveness of our CMM-Net in generating realistic videos of multiple smile expressions. Wei Wang 0108, Xavier Alameda-Pineda, Dan Xu 0002, Pascal Fua, Elisa Ricci 0001, Nicu Sebe |
CVPR | 1 |
| 2018 | GestureGAN for Hand Gesture-to-Gesture Translation in the WildabstractHand gesture-to-gesture translation in the wild is a challenging task since hand gestures can have arbitrary poses, sizes, locations and self-occlusions. Therefore, this task requires a high-level understanding of the mapping between the input source gesture and the output target gesture. To tackle this problem, we propose a novel hand Gesture Generative Adversarial Network (GestureGAN). GestureGAN consists of a single generator G and a discriminator D, which takes as input a conditional hand image and a target hand skeleton image. GestureGAN utilizes the hand skeleton information explicitly, and learns the gesture-to-gesture mapping through two novel losses, the color loss and the cycle-consistency loss. The proposed color loss handles the issue of "channel pollution" while back-propagating the gradients. In addition, we present the Frechet ResNet Distance (FRD) to evaluate the quality of generated images. Extensive experiments on two widely used benchmark datasets demonstrate that the proposed GestureGAN achieves state-of-the-art performance on the unconstrained hand gesture-to-gesture translation task. Meanwhile, the generated images are in high-quality and are photo-realistic, allowing them to be used as data augmentation to improve the performance of a hand gesture classifier. Our model and code are available at https://github.com/Ha0Tang/GestureGAN. Hao Tang 0005, Wei Wang 0108, Dan Xu 0002, Yan Yan 0002, Nicu Sebe |
ACM Multimedia | 2 |
| 2018 | Salient object detection via robust dictionary representation
Huaxin Xiao, Weiya Ren, Wei Wang 0108, Yu Liu 0008, Maojun Zhang |
Multim. Tools Appl. | 3 |
| 2018 | Recurrent Convolutional Shape RegressionabstractThe mainstream direction in face alignment is now dominated by cascaded regression methods. These methods start from an image with an initial shape and build a set of shape increments based on features with respect to the current estimated shape. These shape increments move the initial shape to the desired location. Despite the advantages of the cascaded methods, they all share two major limitations: (i) shape increments are learned independently from each other in a cascaded manner, (ii) the use of standard generic computer vision features such SIFT, HOG, does not allow these methods to learn problem-specific features. In this work, we propose a novel Recurrent Convolutional Shape Regression (RCSR) method that overcomes these limitations. We formulate the standard cascaded alignment problem as a recurrent process and learn all shape increments jointly, by using a recurrent neural network with a gated recurrent unit. Importantly, by combining a convolutional neural network with a recurrent one we avoid hand-crafted features, widely adopted in the literature and thus we allow the model to learn task-specific features. Besides, we employ the convolutional gated recurrent unit which takes as input the feature tensors instead of flattened feature vectors. Therefore, the spatial structure of the features can be better preserved in the memory of the recurrent neural network. Moreover, both the convolutional and the recurrent neural networks are learned jointly. Experimental evaluation shows that the proposed method has better performance than the state-of-the-art methods, and further supports the importance of learning a single end-to-end model for face alignment. Wei Wang 0108, Sergey Tulyakov, Nicu Sebe |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | A Deep Network Architecture for Super-Resolution-Aided Hyperspectral Image Classification With Classwise LossabstractThe supervised deep networks have shown great potential in improving the classification performance. However, training these supervised deep networks is very challenging for hyperspectral image given the fact that usually only a small amount of labeled samples are available. In order to overcome this problem and enhance the discriminative ability of the network, in this paper, we propose a deep network architecture for a super-resolution (SR)-aided hyperspectral image classification with classwise loss (SRCL). First, a three-layer SR convolutional neural network (SRCNN) is employed to reconstruct a high-resolution image from a low-resolution image. Second, an unsupervised triplet-pipeline CNN (TCNN) with an improved classwise loss is built to encourage intraclass similarity and interclass dissimilarity. Finally, SRCNN, TCNN, and a classification module are integrated to define the SRCL, which can be fine-tuned in an end-to-end manner with a small amount of training data. Experimental results on real hyperspectral images demonstrate that the proposed SRCL approach outperforms other state-of-the-art classification methods, especially for the task in which only a small amount of training data are available. Siyuan Hao, Wei Wang 0108, Yuanxin Ye, Enyu Li, Lorenzo Bruzzone |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2018 | Two-Stream Deep Architecture for Hyperspectral Image ClassificationabstractMost traditional approaches classify hyperspectral image (HSI) pixels relying only on the spectral values of the input channels. However, the spatial context around a pixel is also very important and can enhance the classification performance. In order to effectively exploit and fuse both the spatial context and spectral structure, we propose a novel two-stream deep architecture for HSI classification. The proposed method consists of a two-stream architecture and a novel fusion scheme. In the two-stream architecture, one stream employs the stacked denoising autoencoder to encode the spectral values of each input pixel, and the other stream takes as input the corresponding image patch and deep convolutional neural networks are employed to process the image patch. In the fusion scheme, the prediction probabilities from two streams are fused by adaptive class-specific weights, which can be obtained by a fully connected layer. Finally, a weight regularizer is added to the loss function to alleviate the overfitting of the class-specific fusion weights. Experimental results on real HSIs demonstrate that the proposed two-stream deep architecture can achieve competitive performance compared with the state-of-the-art methods. Siyuan Hao, Wei Wang 0108, Yuanxin Ye, Tingyuan Nie, Lorenzo Bruzzone |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2018 | Flexible Manifold Learning With Optimal Graph for Image and Video RepresentationabstractGraph-based dimensionality reduction techniques have been widely and successfully applied to clustering and classification tasks. The basis of these algorithms is the constructed graph which dictates their performance. In general, the graph is defined by the input affinity matrix. However, the affinity matrix derived from the data is sometimes suboptimal for dimension reduction as the data used are very noisy. To address this issue, we propose the projective unsupervised flexible embedding models with optimal graph (PUFE-OG). We build an optimal graph by adjusting the affinity matrix. To tackle the out-of-sample problem, we employ a linear regression term to learn a projection matrix. The optimal graph and the projection matrix are jointly learned by integrating the manifold regularizer and regression residual into a unified model. The experimental results on the public benchmark datasets demonstrate that the proposed PUFE-OG outperforms state-of-the-art methods. Wei Wang 0108, Yan Yan 0002, Feiping Nie 0001, Shuicheng Yan, Nicu Sebe |
IEEE Trans. Image Process. | 1 |
| 2017 | FoveaNet: Perspective-Aware Urban Scene ParsingabstractParsing urban scene images benefits many applications, especially self-driving. Most of the current solutions employ generic image parsing models that treat all scales and locations in the images equally and do not consider the geometry property of car-captured urban scene images. Thus, they suffer from heterogeneous object scales caused by perspective projection of cameras on actual scenes and inevitably encounter parsing failures on distant objects as well as other boundary and recognition errors. In this work, we propose a new FoveaNet model to fully exploit the perspective geometry of scene images and address the common failures of generic parsing models. FoveaNet estimates the perspective geometry of a scene image through a convolutional network which integrates supportive evidence from contextual objects within the image. Based on the perspective geometry information, FoveaNet “undoes” the camera perspective projection - analyzing regions in the space of the actual scene, and thus provides much more reliable parsing results. Furthermore, to effectively address the recognition errors, FoveaNet introduces a new dense CRFs model that takes the perspective geometry as a prior potential. We evaluate FoveaNet on two urban scene parsing datasets, Cityspaces and CamVid, which demonstrates that FoveaNet can outperform all the well-established baselines and provide new state-of-the-art performance. Zequn Jie, Wei Wang 0108, Changsong Liu, Jimei Yang, Xiaohui Shen, Zhe Lin 0001, Qiang Chen 0007, Shuicheng Yan, Jiashi Feng |
ICCV | 3 |
| 2017 | Recurrent 3D-2D Dual Learning for Large-Pose Facial Landmark DetectionabstractDespite remarkable progress of face analysis techniques, detecting landmarks on large-pose faces is still difficult due to self-occlusion, subtle landmark difference and incomplete information. To address these challenging issues, we introduce a novel recurrent 3D-2D dual learning model that alternatively performs 2D-based 3D face model refinement and 3D-to-2D projection based 2D landmark refinement to reliably reason about self-occluded landmarks, precisely capture the subtle landmark displacement and accurately detect landmarks even in presence of extremely large poses. The proposed model presents the first loop-closed learning framework that effectively exploits the informative feedback from the 3D-2D learning and its dual 2D-3D refinement tasks in a recurrent manner. Benefiting from these two mutual-boosting steps, our proposed model demonstrates appealing robustness to large poses (up to profile pose) and outstanding ability to capture fine-scale landmark displacement compared with existing 3D models. It achieves new state-of-the-art on the challenging AFLW benchmark. Moreover, our proposed model introduces a new architectural design that economically utilizes intermediate features and achieves 4× faster speed than its deep learning based counterparts. Shengtao Xiao, Jiashi Feng, Luoqi Liu, Xuecheng Nie, Wei Wang 0108, Shuicheng Yan, Ashraf A. Kassim |
ICCV | 5 |
| 2017 | Face Aging with Contextual Generative Adversarial NetsabstractFace aging, which renders aging faces for an input face, has attracted extensive attention in the multimedia research. Recently, several conditional Generative Adversarial Nets (GANs) based methods have achieved great success. They can generate images fitting the real face distributions conditioned on each individual age group. However, these methods fail to capture the transition patterns, e.g., the gradual shape and texture changes between adjacent age groups. In this paper, we propose a novel Contextual Generative Adversarial Nets (C-GANs) to specifically take it into consideration. The C-GANs consists of a conditional transformation network and two discriminative networks. The conditional transformation network imitates the aging procedure with several specially designed residual blocks. The age discriminative network guides the synthesized face to fit the real conditional distribution. The transition pattern discriminative network is novel, aiming to distinguish the real transition patterns with the fake ones. It serves as an extra regularization term for the conditional transformation network, ensuring the generated image pairs to fit the corresponding real transition pattern distribution. Experimental results demonstrate the proposed framework produces appealing results by comparing with the state-of-the-art and ground truth. We also observe performance gain for cross-age face verification. Si Liu 0001, Yao Sun 0004, Defa Zhu, Renda Bao, Wei Wang 0108, Xiangbo Shu, Shuicheng Yan |
ACM Multimedia | 5 |
| 2017 | Class-wise dictionary learning for hyperspectral image classification
Siyuan Hao, Wei Wang 0108, Yan Yan 0002, Lorenzo Bruzzone |
Neurocomputing | 2 |
| 2016 | Sparse Code Filtering for Action Pattern Mining
Wei Wang 0108, Yan Yan 0002, Liqiang Nie, Stefan Winkler 0001, Nicu Sebe |
ACCV (2) | 1 |
| 2016 | Recurrent Convolutional Face Alignment
Wei Wang 0108, Sergey Tulyakov, Nicu Sebe |
ACCV (2) | 1 |
| 2016 | Projective Unsupervised Flexible Embedding with Optimal Graph
Wei Wang 0108, Yan Yan 0002, Feiping Nie 0001, Xavier Alameda-Pineda, Shuicheng Yan, Nicu Sebe |
BMVC | 1 |
| 2016 | Recurrent Face AgingabstractModeling the aging process of human face is important for cross-age face verification and recognition. In this paper, we introduce a recurrent face aging (RFA) framework based on a recurrent neural network which can identify the ages of people from 0 to 80. Due to the lack of labeled face data of the same person captured in a long range of ages, traditional face aging models usually split the ages into discrete groups and learn a one-step face feature transformation for each pair of adjacent age groups. However, those methods neglect the in-between evolving states between the adjacent age groups and the synthesized faces often suffer from severe ghosting artifacts. Since human face aging is a smooth progression, it is more appropriate to age the face by going through smooth transition states. In this way, the ghosting artifacts can be effectively eliminated and the intermediate aged faces between two discrete age groups can also be obtained. Towards this target, we employ a twolayer gated recurrent unit as the basic recurrent module whose bottom layer encodes a young face to a latent representation and the top layer decodes the representation to a corresponding older face. The experimental results demonstrate our proposed RFA provides better aging faces over other state-of-the-art age progression methods. Wei Wang 0108, Zhen Cui 0001, Yan Yan 0002, Jiashi Feng, Shuicheng Yan, Xiangbo Shu, Nicu Sebe |
CVPR | 1 |
| 2016 | Category Specific Dictionary Learning for Attribute Specific Feature SelectionabstractAttributes, as mid-level features, have demonstrated great potential in visual recognition tasks due to their excellent propagation capability through different categories. However, existing attribute learning methods are prone to learning the correlated attributes. To discover the genuine attribute specific features, many feature selection methods have been proposed. However, these feature selection methods are implemented at the level of raw features that might be very noisy, and these methods usually fail to consider the structural information in the feature space. To address this issue, in this paper, we propose a label constrained dictionary learning approach combined with a multilayer filter. The feature selection is implemented at dictionary level, which can better preserve the structural information. The label constrained dictionary learning suppresses the intra-class noise by encouraging the sparse representations of intra-class samples to lie close to their center. A multilayer filter is developed to discover the representative and robust attribute specific bases. The attribute specific bases are only shared among the positive samples or the negative samples. The experiments on the challenging Animals with Attributes data set and the SUN attribute data set demonstrate the effectiveness of our proposed method. Wei Wang 0108, Yan Yan 0002, Stefan Winkler 0001, Nicu Sebe |
IEEE Trans. Image Process. | 1 |
| 2015 | Attribute Guided Dictionary LearningabstractAttributes have shown great potential in visual recognition recently since they, as mid-level features, can be shared across different categories. However, existing attribute learning methods are prone to learning the correlated attributes which results in the difficulties of selecting attribute specific features. In this paper, we propose an attribute specific dictionary learning approach to address this issue. Category information is incorporated into our framework while learning the over-complete dictionary, which encourages the samples from the same category to have similar distributions over the dictionary bases. A novel scheme is developed to select the attribute specific dictionaries. The attribute specific dictionary consists of the bases which are only shared among the positive samples or the negative samples. The experiments on the Animals with Attributes (AwA) dataset show the effectiveness of our proposed method. Wei Wang 0108, Yan Yan 0002, Nicu Sebe |
ICMR | 1 |