Xiaoyu Wang 0002

dblp:58/4775-2 · DBLP profile ↗
← Back
63ranked-venue papers
6as first author
33since 2021 · last 2026
0000-0002-6431-8822ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 46 · 6 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 41 · 5 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 An Enhanced Adaptive Confidence Margin for Semi-Supervised Facial Expression Recognition
abstract
Semi-supervised learning (SSL) provides a practical framework for leveraging massive unlabeled samples, especially when labels are expensive for facial expression recognition (FER). Typical SSL methods like FixMatch select unlabeled samples with confidence scores above a fixed threshold for training. However, these methods face two primary limitations: failing to consider the varying confidence across facial expression categories and failing to utilize unlabeled facial expression samples efficiently. To address these challenges, we propose an Enhanced Adaptive Confidence Margin (EACM), consisting of dynamic thresholds for different categories, to fully learn unlabeled samples. Specifically, we employ the predictions on labeled samples at each training iteration to learn an EACM. It then partitions unlabeled samples into two subsets: (1) subset I, including samples whose confidence scores are no less than the margin; (2) subset II, including samples whose confidence scores are less than the margin. For samples in subset I, we constrain their predictions on strongly-augmented versions to match the pseudo-labels derived from the predictions on weakly-augmented versions. Meanwhile, we introduce a feature-level contrastive objective to enhance the similarity between two weakly-augmented features of a sample in subset II. We extensively evaluate EACM on image-based and video-based facial expression datasets, showing that our method achieves superior performance, significantly surpassing fully-supervised baselines in a semi-supervised manner. Additionally, our EACM is promising to leverage cross-dataset unlabeled samples for practical training to boost fully-supervised performance.
Hangyu Li 0001, Nannan Wang 0001, Xi Yang 0011, Xiaoyu Wang 0002, Xinbo Gao 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 SA-BCT: Self-Adapting Backward-Compatible Training
abstract
Backward-compatible training enables the deployment of advanced models without requiring updates to old gallery databases. However, existing methods, including old-prototype-based (i.e., those relying on prototypes from the old model) and instance-based approaches, often overlook the impact of the old model's quality. High-quality old models exhibit compact intra-class feature distributions, which facilitate effective alignment between old and new models across various methods. In contrast, low-quality old models produce dispersed features, making it difficult for old-prototype-based methods to extract sufficient information. Additionally, instance-based methods are overly restrictive, limiting the flexibility of new models. In this work, we propose SA-BCT, an extremely simple yet effective backward-compatible training method that offers a unified framework for accommodating old models of varying quality. SA-BCT employs a single loss function applied to both old and new features, self-adaptively adjusting the constraint space for new features based on the distribution of old features. Extensive experiments in diverse settings demonstrate the effectiveness of SA-BCT. Code is available athttps://github.com/yuleung/SA-BCT.
Yufeng Zhang 0001, Shiliang Zhang, Sheng Xiao, Rong Xiao 0003, Xiaoyu Wang 0002, Kenli Li 0001
IEEE Trans. Multim.6
2025 Rectified Binary Network for Single-Image Super-Resolution
abstract
Binary neural network (BNN) is an effective approach to reduce the memory usage and the computational complexity of full-precision convolutional neural networks (CNNs), which has been widely used in the field of deep learning. However, there are different properties between BNNs and real-valued models, making it difficult to draw on the experience of CNN composition to develop BNN. In this article, we study the application of binary network to the single-image super-resolution (SISR) task in which the network is trained for restoring original high-resolution (HR) images. Generally, the distribution of features in the network for SISR is more complex than those in recognition models for preserving the abundant image information, e.g., texture, color, and details. To enhance the representation ability of BNN, we explore a novel activation-rectified inference (ARI) module that achieves a more complete representation of features by combining observations from different quantitative perspectives. The activations are divided into several parts with different quantification intervals and are inferred independently. This allows the binary activations to retain more image detail and yield finer inference. In addition, we further propose an adaptive approximation estimator (AAE) for gradually learning the accurate gradient estimation interval in each layer to alleviate the optimization difficulty. Experiments conducted on several benchmarks show that our approach is able to learn a binary SISR model with superior performance over the state-of-the-art methods. The code will be released at https://github.com/jwxintt/Rectified-BSR.
Jingwei Xin, Nannan Wang 0001, Jie Li 0001, Xiaoyu Wang 0002, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.5
2025 Context-Assisted Active Learning for Weakly Supervised Person Search
abstract
Person search is a challenging task that aims to jointly detect and identify a target person from a large-scale scene image dataset. Fully supervised person search requires both bounding boxes and person identity annotations, making it hard to deploy in real-world applications. Although recent weakly supervised person search methods can alleviate annotation workloads, they often result in severe performance degradation when compared to supervised methods. To pursue better performance with a lower annotation budget, we propose to integrate active learning into weakly supervised person search, where a small number of pairwise identity annotations are actively acquired from oracles. Specifically, we propose a context-assisted active learning framework that selects informative instance pairs for labeling and refines pseudo labels for representation learning. The proposed framework consists of a split module and a merge module, which leverage two types of contextual cues for label refinement. Besides, a pairwise relationship predictor is introduced to estimate relations between instances so that annotation cost can be further reduced. Extensive experiments demonstrate that the proposed method could achieve comparable or even better performance than recent fully supervised methods at a much lower annotation cost. Notably, our method achieves 61.4% mAP on PRW dataset, which outperforms recent fully supervised methods at a much lower annotation cost.
Rinyoichi Takezoe, Hao Chen 0061, Xuefei Lv, Yaowei Wang 0001, Shiliang Zhang, Xiaoyu Wang 0002
ACM Trans. Multim. Comput. Commun. Appl.7
2024 Universal Object Detection with Large Vision Model
Feng Lin 0009, Wenze Hu, Yaowei Wang 0001, Yonghong Tian 0001, Guangming Lu 0002, Fanglin Chen 0001, Yong Xu 0007, Xiaoyu Wang 0002
Int. J. Comput. Vis.8
2024 Graph-based social relation inference with multi-level conditional attention
Xiaotian Yu, Hanling Yi, Qie Tang, Wenze Hu, Shiliang Zhang, Xiaoyu Wang 0002
Neural Networks7
2024 Unconstrained Facial Expression Recognition With No-Reference De-Elements Learning
abstract
Most unconstrained facial expression recognition (FER) methods take original facial images as inputs to learn discriminative features by well-designed loss functions, which cannot reflect important visual information in faces. Although existing methods have explored the visual information of constrained facial expressions, there is no explicit modeling of what visual information is important for unconstrained FER. To find out valuable information of unconstrained facial expressions, we pose a new problem of no-reference de-elements learning: we decompose any unconstrained facial image into the facial expression element and a neutral face without the reference of corresponding neutral faces. Importantly, the element provides visualization results to understand important facial expression information and improves the discriminative power of features. Moreover, we propose a simple yet effectiveDe-ElementsNetwork (DENet) to learn the element and introduce appropriate constraints to overcome no ground truth of corresponding neutral faces during the de-elements learning. We extensively evaluate the proposed method on in-the-wild FER datasets including RAF-DB, AffectNet, SFEW and FERPlus. The comparable results show that our method is promising to improve classification performance and achieves equivalent performance compared with state-of-the-art methods. Also, we demonstrate the strong generalization performance on realistic occlusion and pose variation datasets and the cross-dataset evaluation.
Hangyu Li 0001, Nannan Wang 0001, Xi Yang 0011, Xiaoyu Wang 0002, Xinbo Gao 0001
IEEE Trans. Affect. Comput.4
2023 Cross-Modality Person Re-identification with Memory-Based Contrastive Embedding
abstract
Visible-infrared person re-identification (VI-ReID) aims to retrieve the person images of the same identity from the RGB to infrared image space, which is very important for real-world surveillance system. In practice, VI-ReID is more challenging due to the heterogeneous modality discrepancy, which further aggravates the challenges of traditional single-modality person ReID problem, i.e., inter-class confusion and intra-class variations. In this paper, we propose an aggregated memory-based cross-modality deep metric learning framework, which benefits from the increasing number of learned modality-aware and modality-agnostic centroid proxies for cluster contrast and mutual information learning. Furthermore, to suppress the modality discrepancy, the proposed cross-modality alignment objective simultaneously utilizes both historical and up-to-date learned cluster proxies for enhanced cross-modality association. Such training mechanism helps to obtain hard positive references through increased diversity of learned cluster proxies, and finally achieves stronger ``pulling close'' effect between cross-modality image features. Extensive experiment results demonstrate the effectiveness of the proposed method, surpassing state-of-the-art works significantly by a large margin on the commonly used VI-ReID datasets.
De Cheng, Nannan Wang 0001, Zhen Wang 0037, Xiaoyu Wang 0002, Xinbo Gao 0001
AAAI5
2023 Boosting Weakly-Supervised Temporal Action Localization with Text Information
abstract
Due to the lack of temporal annotation, current Weakly-supervised Temporal Action Localization (WTAL) methods are generally stuck into over-complete or incomplete localization. In this paper, we aim to leverage the text information to boost WTAL from two aspects, i.e., (a) the discriminative objective to enlarge the inter-class difference, thus reducing the over-complete; (b) the generative objective to enhance the intra-class integrity, thus finding more complete temporal boundaries. For the discriminative objective, we propose a Text-Segment Mining (TSM) mechanism, which constructs a text description based on the action class label, and regards the text as the query to mine all class-related segments. Without the temporal annotation of actions, TSM compares the text query with the entire videos across the dataset to mine the best matching segments while ignoring irrelevant ones. Due to the shared sub-actions in different categories of videos, merely applying TSM is too strict to neglect the semantic-related segments, which results in incomplete localization. We further introduce a generative objective named Video-text Language Completion (VLC), which focuses on all semantic-related segments from videos to complete the text sentence. We achieve the state-of-the-art performance on THUMOS14 and ActivityNetl.3. Surprisingly, we also find our proposed method can be seamlessly applied to existing methods, and improve their performances with a clear margin. The code is available at https://github.com/lgzlIlIlI/Boosting-WTAL.
Guozhang Li, De Cheng, Xinpeng Ding, Nannan Wang 0001, Xiaoyu Wang 0002, Xinbo Gao 0001
CVPR5
2023 NAR-Former: Neural Architecture Representation Learning Towards Holistic Attributes Prediction
abstract
With the wide and deep adoption of deep learning models in real applications, there is an increasing need to model and learn the representations of the neural networks themselves. These models can be used to estimate attributes of different neural network architectures such as the accuracy and latency, without running the actual training or inference tasks. In this paper, we propose a neural architecture representation model that can be used to estimate these attributes holistically. Specifically, we first propose a simple and effective tokenizer to encode both the operation and topology information of a neural network into a single sequence. Then, we design a multi-stage fusion transformer to build a compact vector representation from the converted sequence. For efficient model training, we further propose an information flow consistency augmentation and correspondingly design an architecture consistency loss, which brings more benefits with less augmentation samples compared with previous random augmentation strategies. Experiment results on NAS-Bench-101, NAS-Bench-201, DARTS search space and NNLQP show that our proposed framework can be used to predict the aforementioned latency and accuracy attributes of both cell architectures and whole deep neural networks, and achieves promising performance. Code is available at https://github.com/yuny220/NAR-Former.
Yun Yi, Haokui Zhang, Wenze Hu, Nannan Wang 0001, Xiaoyu Wang 0002
CVPR5
2023 ParCNetV2: Oversized Kernel with Enhanced Attention*
abstract
Transformers have shown great potential in various computer vision tasks. By borrowing design concepts from transformers, many studies revolutionized CNNs and showed remarkable results. This paper falls in this line of studies. Specifically, we propose a new convolutional neural network, ParCNetV2, that extends the research line of ParCNetV1 by bridging the gap between CNN and ViT. It introduces two key designs: 1) Oversized Convolution (OC) with twice the size of the input, and 2) Bifurcate Gate Unit (BGU) to ensure that the model is input adaptive. Fusing OC and BGU in a unified CNN, ParCNetV2 is capable of flexibly extracting global features like ViT, while maintaining lower latency and better accuracy. Extensive experiments demonstrate the superiority of our method over other convolutional neural networks and hybrid models that combine CNNs and transformers. The code are publicly available at https://github.com/XuRuihan/ParCNetV2.
Ruihan Xu 0002, Haokui Zhang, Wenze Hu, Shiliang Zhang, Xiaoyu Wang 0002
ICCV5
2023 Fcaformer: Forward Cross Attention in Hybrid Vision Transformer
abstract
Currently, one main research line in designing a more efficient vision transformer is reducing the computational cost of self attention modules by adopting sparse attention or using local attention windows. In contrast, we propose a different approach that aims to improve the performance of transformer-based architectures by densifying the attention pattern. Specifically, we proposed forward cross attention for hybrid vision transformer (FcaFormer), where tokens from previous blocks in the same stage are secondary used. To achieve this, the FcaFormer leverages two innovative components: learnable scale factors (LSFs) and a token merge and enhancement module (TME). The LSFs enable efficient processing of cross tokens, while the TME generates representative cross tokens. By integrating these components, the proposed FcaFormer enhances the interactions of tokens across blocks with potentially different semantics, and encourages more information flows to the lower levels. Based on the forward cross attention (Fca), we have designed a series of FcaFormer models that achieve the best trade-off between model size, computational cost, memory cost, and accuracy. For example, without the need for knowledge distillation to strengthen training, our FcaFormer achieves 83.1% top-1 accuracy on Imagenet with only 16.3 million parameters and about 3.6 billion MACs. This saves almost half of the parameters and a few computational costs while achieving 0.7% higher accuracy compared to distilled EfficientFormer. Code is available at https://github.com/hkzhang-git/FcaFormer
Haokui Zhang, Wenze Hu, Xiaoyu Wang 0002
ICCV3
2023 All-to-key Attention for Arbitrary Style Transfer
abstract
Attention-based arbitrary style transfer studies have shown promising performance in synthesizing vivid local style details. They typically use the all-to-all attention mechanism—each position of content features is fully matched to all positions of style features. However, all-to-all attention tends to generate distorted style patterns and has quadratic complexity, limiting the effectiveness and efficiency of arbitrary style transfer. In this paper, we propose a novel all-to-key attention mechanism—each position of content features is matched to stable key positions of style features—that is more in line with the characteristics of style transfer. Specifically, it integrates two newly proposed attention forms: distributed and progressive attention. Distributed attention assigns attention to key style representations that depict the style distribution of local regions; Progressive attention pays attention from coarse-grained regions to fine-grained key positions. The resultant module, dubbed StyA2K, shows extraordinary performance in preserving the semantic structure and rendering consistent style patterns. Qualitative and quantitative comparisons with state-of-the-art methods demonstrate the superior performance of our approach. Codes and models are available on https://github.com/LearningHx/StyA2K.
Mingrui Zhu, Xiao He 0014, Nannan Wang 0001, Xiaoyu Wang 0002, Xinbo Gao 0001
ICCV4
2023 Controllable Face Sketch-Photo Synthesis with Flexible Generative Priors
abstract
Current face sketch-photo synthesis researches generally embrace an image-to-image (I2I) translation pipeline. However, these methods ignore the one-to-many mapping problem (i.e., multiple plausible photo results can correspond to a single input sketch) in sketch-to-photo synthesis task, resulting in significant performance degradation on diverse datasets. Besides, generating high-quality images on limited data is also a challenge for this task. To address these challenges, we propose a dual-path framework that introduces generative priors to better perform cross-domain reconstruction on limited data. The coarse path uses a layer-swapped pre-trained generator to achieve coarse cross-domain reconstruction, and the refinement path further improves the structure and texture details. To align the feature maps between the two paths, we introduce a spatial feature calibration module. Despite this, our framework still struggles to handle diverse datasets. Thanks to the flexibility of generative priors, we can extend the framework to achieve exemplar-guided I2I translation by incorporating an exemplar with style mixing and a proposed semantic-aware style refinement strategy, which addresses the one-to-many mapping problem in sketch-to-photo synthesis task. Furthermore, our framework can perform cross-domain editing by employing off-the-shelf editing methods based on the latent space, achieving fine-grained control. Extensive experiments on diverse datasets demonstrate the superiority of our framework over other state-of-the-art methods.
Mingrui Zhu, Nannan Wang 0001, Guozhang Li, Xiaoyu Wang 0002, Xinbo Gao 0001
ACM Multimedia5
2023 NAR-Former V2: Rethinking Transformer for Universal Neural Network Representation Learning
abstract
As more deep learning models are being applied in real-world applications, there is a growing need for modeling and learning the representations of neural networks themselves. An effective representation can be used to predict target attributes of networks without the need for actual training and deployment procedures, facilitating efficient network design and deployment. Recently, inspired by the success of Transformer, some Transformer-based representation learning frameworks have been proposed and achieved promising performance in handling cell-structured models. However, graph neural network (GNN) based approaches still dominate the field of learning representation for the entire network. In this paper, we revisit the Transformer and compare it with GNN to analyze their different architectural characteristics. We then propose a modified Transformer-based universal neural network representation learning model NAR-Former V2. It can learn efficient representations from both cell-structured networks and entire networks. Specifically, we first take the network as a graph and design a straightforward tokenizer to encode the network into a sequence. Then, we incorporate the inductive representation learning capability of GNN into Transformer, enabling Transformer to generalize better when encountering unseen architecture. Additionally, we introduce a series of simple yet effective modifications to enhance the ability of the Transformer in learning representation from graph structures. In encoding entire networks and then predicting the latency, our proposed method surpasses the GNN-based method NNLP by a significant margin on the NNLQP dataset. Furthermore, regarding accuracy prediction on the cell-structured NASBench101 and NASBench201 datasets, our method achieves highly comparable performance to other state-of-the-art methods. The code is available at https://github.com/yuny220/NAR-Former-V2.
Yun Yi, Haokui Zhang, Rong Xiao 0003, Nannan Wang 0001, Xiaoyu Wang 0002
NeurIPS5
2023 BiTGAN: bilateral generative adversarial networks for Chinese ink wash painting style transfer
Xiao He 0014, Mingrui Zhu, Nannan Wang 0001, Xiaoyu Wang 0002, Xinbo Gao 0001
Sci. China Inf. Sci.4
2023 Multi-proxy feature learning for robust fine-grained visual recognition
Shunan Mao, Yaowei Wang 0001, Xiaoyu Wang 0002, Shiliang Zhang
Pattern Recognit.3
2023 Adapt-Infomap: Face clustering with adaptive graph refinement in infomap
abstract
Face clustering is a critical task in computer vision due to the increasing number of applications such as augmented reality or photo album management. The primary challenge in this task arises from the imperfections in image feature representations. Given image features extracted from an existing pre-trained representation model, it remains an unresolved problem that how to leverage the inherent characteristics of similarities among unlabelled images to improve the clustering performance. In order to solve face clustering in an unsupervised manner , we develop an effective and robust framework named as Adapt-Infomap. First, we reformulate face clustering as a process of non-overlapping community detection. Specially, Adapt-Infomap achieves face clustering by minimizing the entropy of information flows (also known as the map equation) on an affinity graph of images. Since the affinity graph of images might contain noisy edges, we develop an outlier detection strategy in Adapt-Infomap to adaptively refine the affinity graph. Experiments with ablation studies demonstrate that Adapt-Infomap significantly outperforms existing methods and achieves new state-of-the-arts on three popular large-scale datasets for face clustering, e.g. , an absolute improvement of more than 10 % and 3 % comparing with prior unsupervised and supervised methods respectively in terms of average of Pairwise F-score.
Xiaotian Yu, Aibo Wang, Haokui Zhang, Hanling Yi, Guangming Lu 0002, Xiaoyu Wang 0002
Pattern Recognit.8
2023 Learning a High Fidelity Identity Representation for Face Frontalization
abstract
This paper considers the problem of face frontalization in the wild, which transforms a face image with profile views into a frontal face. Face frontalization provides an effective solution to the face recognition problem in uncontrolled scenes. However, the existing methods either focus on deep learning techniques as an end-to-end framework or combine other explicit facial prior estimation tasks, such as 3D representation, optical flow estimation and so on, where computation is highly redundant and facial identity cannot be well represented. In this paper, we focus on how to maximise the potential of the model for identity learning and representation, and propose an accurate and lightweight face frontalization approach, named identity-preserving model (IPM). IPM has a well-designed encoder-decoder architecture which restores input face to a frontal counterpart. The encoder is constructed to extract representation from the input face, where a contrastive loss function is applied that encourages representations to form compact clusters, while preserving their relationships across the corpora. Then a cross-domain rectification module is proposed to eliminate the representation differences between the recognition and reconstruction domains, thus improving the accuracy of the reconstructed face. Extensive experiments on benchmark datasets show that the proposed IPM approach not only outperforms the state-of-the-art on public datasets but also can cope with images in the uncontrolled scenes.
Jingwei Xin, Zikai Wei, Nannan Wang 0001, Jie Li 0001, Xiaoyu Wang 0002, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.5
2023 Unsupervised Across Domain Consistency- Difference Network for Hyperspectral Image Super-Resolution
abstract
Without reducing the spectral resolution, hyperspectral image super-resolution has achieved remarkable progress thanks to the success of deep neural networks. However, existing methods can not fully excavate the latent high-frequency details only in the single spatial domain. Different from existing methods that only achieves the super-resolution task in spatial domain, we optimize the amplitude spectrum and phase spectrum in frequency domain to obtain high resolution hyperspectral image (HR-HSI). We propose a new unsupervised framework to reconstruct HR-HSI using only the observed low resolution HSI and HR multispectral image. Based on triple-level modeling, the encoder-decoder learns abundant features including contextual information from multiple scales. In addition, we propose iterative across domain consistency-difference (ADCD) module, which is embedded between encoder and decoder. In ADCD module, three parallel convolution streams, (amplitude spectrum adjustment branch, phase spectrum adjustment branch and spatial domain branch) are used to explore the consistency-difference between each other, which is preserved by memory units within the module. Particularly, we embed the dilated causal convolution in the frequency domain processing branch, which is convenient to flexibly adjust the receptive field and adapt to different domains. Extensive experiments are conducted on widely-used datasets in comparison with state-of-the-art models, demonstrating the advantage of the proposed method.
Zhiling Guo, Jingwei Xin, Nannan Wang 0001, Jie Li 0001, Xiaoyu Wang 0002, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 FABNet: Frequency-Aware Binarized Network for Single Image Super-Resolution
abstract
Remarkable achievements have been obtained with binary neural networks (BNN) in real-time and energy-efficient single-image super-resolution (SISR) methods. However, existing approaches often adopt the Sign function to quantize image features while ignoring the influence of image spatial frequency. We argue that we can minimize the quantization error by considering different spatial frequency components. To achieve this, we propose a frequency-aware binarized network (FABNet) for single image super-resolution. First, we leverage the wavelet transformation to decompose the features into low-frequency and high-frequency components and then employ a "divide-and-conquer" strategy to separately process them with well-designed binary network structures. Additionally, we introduce a dynamic binarization process that incorporates learned-threshold binarization during forward propagation and dynamic approximation during backward propagation, effectively addressing the diverse spatial frequency information. Compared to existing methods, our approach is effective in reducing quantization error and recovering image textures. Extensive experiments conducted on four benchmark datasets demonstrate that the proposed methods could surpass state-of-the-art approaches in terms of PSNR and visual quality with significantly reduced computational costs. Our codes are available at https://github.com/xrjiang527/FABNet-PyTorch.
Nannan Wang 0001, Jingwei Xin, Xi Yang 0011, Jie Li 0001, Xiaoyu Wang 0002, Xinbo Gao 0001
IEEE Trans. Image Process.7
2023 An Efficient Transformer Based on Global and Local Self-Attention for Face Photo-Sketch Synthesis
abstract
Face photo-sketch synthesis tasks have been dominated by convolutional neural networks (CNNs), especially CNN-based generative adversarial networks (GANs), because of their strong texture modeling capabilities and thus their ability to generate more realistic face photos/sketches beyond traditional methods. However, due to CNNs' locality and spatial invariance properties, there have weaknesses in capturing the global and structural information which are extremely important for face images. Inspired by the recent phenomenal success of the Transformer in vision tasks, we propose replacing CNNs with Transformers that are able to model long-range dependencies to synthesize more structured and realistic face images. However, the existing vision Transformers are mainly designed for high-level vision tasks and lack the dense prediction ability to generate high resolution images due to the quadratic computational complexity of their self-attention mechanism. In addition, the original Transformer is not capable of modeling local correlations which is an important skill for image generation. To address these challenges, we propose two types of memory-friendly Transformer encoders, one for processing local correlations via local self-attention and another for modeling global information via global self-attention. By integrating the two proposed Transformer encoders, we present an efficient GL-Transformer for face photo-sketch synthesis, which can synthesize realistic face photo/sketch images from coarse to fine. Extensive experiments demonstrate that our model achieves a comparable or better performance beyond the state-of-the-art CNN-based methods both qualitatively and quantitatively.
Wangbo Yu, Mingrui Zhu, Nannan Wang 0001, Xiaoyu Wang 0002, Xinbo Gao 0001
IEEE Trans. Image Process.4
2022 Towards Semi-Supervised Deep Facial Expression Recognition with An Adaptive Confidence Margin
abstract
Only parts of unlabeled data are selected to train models for most semi-supervised learning methods, whose confidence scores are usually higher than the pre-defined threshold (i.e., the confidence margin). We argue that the recognition performance should be further improved by making full use of all unlabeled data. In this paper, we learn an Adaptive Confidence Margin (Ada-CM) to fully leverage all unlabeled data for semi-supervised deep facial expression recognition. All unlabeled samples are partitioned into two subsets by comparing their confidence scores with the adaptively learned confidence margin at each training epoch: (1) subset I including samples whose confidence scores are no lower than the margin; (2) subset II including samples whose confidence scores are lower than the margin. For samples in subset I, we constrain their predictions to match pseudo labels. Meanwhile, samples in subset II participate in the feature-level contrastive objective to learn effective facial expression features. We extensively evaluate Ada-CM on four challenging datasets, showing that our method achieves state-of-the-art performance, especially surpassing fully-supervised baselines in a semi-supervised manner. Ablation study further proves the effectiveness of our method. The source code is available at https://github.com/hangyu94/Ada-CM.
Hangyu Li 0001, Nannan Wang 0001, Xi Yang 0011, Xiaoyu Wang 0002, Xinbo Gao 0001
CVPR4
2022 ParC-Net: Position Aware Circular Convolution with Merits from ConvNets and Transformer
Haokui Zhang, Wenze Hu, Xiaoyu Wang 0002
ECCV (26)3
2022 Connecting Compression Spaces with Transformer for Approximate Nearest Neighbor Search
Haokui Zhang, Buzhou Tang, Wenze Hu, Xiaoyu Wang 0002
ECCV (14)4
2022 Improving Adversarial Robustness via Mutual Information Estimation
abstract
Deep neural networks (DNNs) are found to be vulnerable to adversarial noise. They are typically misled by adversarial samples to make wrong predictions. To alleviate this negative effect, in this paper, we investigate the dependence between outputs of the target model and input adversarial samples from the perspective of information theory, and propose an adversarial defense method. Specifically, we first measure the dependence by estimating the mutual information (MI) between outputs and the natural patterns of inputs (called natural MI) and MI between outputs and the adversarial patterns of inputs (called adversarial MI), respectively. We find that adversarial samples usually have larger adversarial MI and smaller natural MI compared with those w.r.t. natural samples. Motivated by this observation, we propose to enhance the adversarial robustness by maximizing the natural MI and minimizing the adversarial MI during the training process. In this way, the target model is expected to pay more attention to the natural pattern that contains objective semantics. Empirical evaluations demonstrate that our method could effectively improve the adversarial accuracy against multiple attacks.
Dawei Zhou 0004, Nannan Wang 0001, Xinbo Gao 0001, Bo Han 0003, Xiaoyu Wang 0002, Yibing Zhan, Tongliang Liu
ICML5
2022 Spatiotemporal consistency-enhanced network for video anomaly detection
Jie Li 0001, Nannan Wang 0001, Xiaoyu Wang 0002, Xinbo Gao 0001
Pattern Recognit.4
2021 Removing Adversarial Noise in Class Activation Feature Space
abstract
Deep neural networks (DNNs) are vulnerable to adversarial noise. Pre-processing based defenses could largely remove adversarial noise by processing inputs. However, they are typically affected by the error amplification effect, especially in the front of continuously evolving attacks. To solve this problem, in this paper, we propose to remove adversarial noise by implementing a self-supervised adversarial training mechanism in a class activation feature space. To be specific, we first maximize the disruptions to class activation features of natural examples to craft adversarial examples. Then, we train a denoising model to minimize the distances between the adversarial examples and the natural examples in the class activation feature space. Empirical evaluations demonstrate that our method could significantly enhance adversarial robustness in comparison to previous state-of-the-art approaches, especially against unseen adversarial attacks and adaptive attacks.
Dawei Zhou 0004, Nannan Wang 0001, Chunlei Peng, Xinbo Gao 0001, Xiaoyu Wang 0002, Jun Yu 0001, Tongliang Liu
ICCV5
2021 A Sketch-Transformer Network for Face Photo-Sketch Synthesis
abstract
We present a face photo-sketch synthesis model, which converts a face photo into an artistic face sketch or recover a photo-realistic facial image from a sketch portrait. Recent progress has been made by convolutional neural networks (CNNs) and generative adversarial networks (GANs), so that promising results can be obtained through real-time end-to-end architectures. However, convolutional architectures tend to focus on local information and neglect long-range spatial dependency, which limits the ability of existing approaches in keeping global structural information. In this paper, we propose a Sketch-Transformer network for face photo-sketch synthesis, which consists of three closely-related modules, including a multi-scale feature and position encoder for patch-level feature and position embedding, a self-attention module for capturing long-range spatial dependency, and a multi-scale spatially-adaptive de-normalization decoder for image reconstruction. Such a design enables the model to generate reasonable detail texture while maintaining global structural information. Extensive experiments show that the proposed method achieves significant improvements over state-of-the-art approaches on both quantitative and qualitative evaluations.
Mingrui Zhu, Changcheng Liang, Nannan Wang 0001, Xiaoyu Wang 0002, Zhifeng Li 0001, Xinbo Gao 0001
IJCAI4
2021 AIBench Training: Balanced Industry-Standard AI Training Benchmarking
abstract
Earlier-stage evaluations of a new AI architecture/system need affordable AI benchmarks. Only using a few AI component benchmarks like MLPerf alone in the other stages may lead to misleading conclusions. Moreover, the learning dynamics are not well understood, and the benchmarks' shelf-life is short. This paper proposes a balanced benchmarking methodology. We use real-world benchmarks to cover the factors space that impacts the learning dynamics to the most considerable extent. After performing an exhaustive survey on Internet service AI domains, we identify and implement nineteen representative AI tasks with state-of-the-art models. For repeatable performance ranking (RPR subset) and workload characterization (WC subset), we keep two subsets to a minimum for affordability. We contribute by far the most comprehensive AI training benchmark suite. The evaluations show: (1) AIBench Training (v1.1) outperforms MLPerf Training (v0.7) in terms of diversity and representativeness of model complexity, computational cost, convergent rate, computation, and memory access patterns, and hotspot functions; (2) Against the AIBench full benchmarks, its RPR subset shortens the benchmarking cost by 64%, while maintaining the primary workload characteristics; (3) The performance ranking shows the single-purpose AI accelerator like TPU with the optimized TensorFlow framework performs better than that of GPUs while losing the latter's general support for various AI models. The specification, source code, and performance numbers are available from the AIBench homepage https://www.benchcouncil.org/aibench-training/index.html.
Fei Tang 0003, Wanling Gao, Jianfeng Zhan, Chuanxin Lan, Lei Wang 0004, Chunjie Luo, Zheng Cao 0003, Xingwang Xiong, Zihan Jiang 0006, Tianshu Hao, Fanda Fan, Fan Zhang 0047, Yunyou Huang, Jianan Chen 0003, Mengjia Du, Chen Zheng 0001, Daoyi Zheng, Haoning Tang, Kunlin Zhan, Defei Kong, Chongkang Tan, Xinhui Tian, Yatao Li, Junchao Shao, Xiaoyu Wang 0002, Jiahui Dai, Hainan Ye
ISPASS31
2021 Deep Regionlets: Blended Representation and Deep Learning for Generic Object Detection
abstract
In this article, we propose a novel object detection algorithm named "Deep Regionlets" by integrating deep neural networks and a conventional detection schema for accurate generic object detection. Motivated by the effectiveness of regionlets for modeling object deformations and multiple aspect ratios, we incorporate regionlets into an end-to-end trainable deep learning framework. The deep regionlets framework consists of a region selection network and a deep regionlet learning module. Specifically, given a detection bounding box proposal, the region selection network provides guidance on where to select sub-regions from which features can be learned from. An object proposal typically contains three - 16 sub-regions. The regionlet learning module focuses on local feature selection and transformations to alleviate the effects of appearance variations. To this end, we first realize non-rectangular region selection within the detection framework to accommodate variations in object appearance. Moreover, we design a "gating network" within the regionlet leaning module to enable instance dependent soft feature selection and pooling. The Deep Regionlets framework is trained end-to-end without additional efforts. We present ablation studies and extensive experiments on the PASCAL VOC dataset and the Microsoft COCO dataset. The proposed method yields competitive performance over state-of-the-art algorithms, such as RetinaNet and Mask R-CNN, even without additional segmentation labels.
Hongyu Xu, Xutao Lv, Xiaoyu Wang 0002, Zhou Ren, Navaneeth Bodla, Rama Chellappa
IEEE Trans. Pattern Anal. Mach. Intell.3
2021 KFC: An Efficient Framework for Semi-Supervised Temporal Action Localization
abstract
In temporal action localization (TAL), semi-supervised learning is a promising technique to mitigate the cost of precise boundary annotations. Semi-supervised approaches employing consistency regularization (CR), encouraging models to be robust to the perturbed inputs, have achieved great success in image classification problems. The success of CR is largely depended on the perturbations, where instances are perturbed to train a robust model without altering their semantic information. However, the perturbations for image or video classification tasks are not fit to apply to TAL. Since videos in TAL are too long to train the model with raw videos in an end-to-end manner. In this paper, we devise a method named K-farthest crossover to construct perturbations based on video features and apply it to TAL. Motivated by the observation that features in the same action instance become more and more similar during the training process while those in different action instances or backgrounds become more and more divergent, we add perturbations to each feature along temporal axis and adopt CR to encourage the model to retain this observation. Specifically, for a feature, we first find the top-k dissimilar features and average them to form a perturbation. Then, similar to chromosomal crossover, we select a large part of the feature and a small part of the perturbation to recombine a perturbed feature, which preserves the feature semantics yet enough discrepancy.
Xinpeng Ding, Nannan Wang 0001, Xinbo Gao 0001, Jie Li 0001, Xiaoyu Wang 0002, Tongliang Liu
IEEE Trans. Image Process.5
2021 Progressive Learning of Low-Precision Networks for Image Classification
abstract
Recent years have witnessed a great advance of deep learning in a variety of vision tasks. Many state-of-the-art deep neural networks suffer from large size and high complexity, which makes them difficult to deploy in resource-limited platforms such as mobile devices. To this end, low-precision neural networks are widely studied that quantize weights or activations into the low-bit format. Although efficient, low-precision networks are usually difficult to train and encounter severe accuracy degradation. In this paper, we propose a new training strategy based on progressive learning for image classification. First, we equip each low-precision convolutional layer with an ancillary full-precision convolutional layer based on a low-precision network structure. Second, a decay method is introduced to reduce the output of the added full-precision convolution gradually, which keeps the resulting topology structure the same as the original low-precision convolution. Extensive experiments on SVHN, CIFAR and ILSVRC-2012 datasets reveal that the proposed method can bring faster convergence and higher accuracy for low-precision neural networks.
Zhengguang Zhou, Wengang Zhou 0001, Xutao Lv, Xiaoyu Wang 0002, Houqiang Li
IEEE Trans. Multim.5
2020 Robust Partial Matching for Person Search in the Wild
abstract
Various factors like occlusions, backgrounds, etc., would lead to misaligned detected bounding boxes , e.g., ones covering only portions of human body. This issue is common but overlooked by previous person search works. To alleviate this issue, this paper proposes an Align-to-Part Network (APNet) for person detection and re-Identification (reID). APNet refines detected bounding boxes to cover the estimated holistic body regions, from which discriminative part features can be extracted and aligned. Aligned part features naturally formulate reID as a partial feature matching procedure, where valid part features are selected for similarity computation, while part features on occluded or noisy regions are discarded. This design enhances the robustness of person search to real-world challenges with marginal computation overhead. This paper also contributes a Large-Scale dataset for Person Search in the wild (LSPS), which is by far the largest and the most challenging dataset for person search. Experiments show that APNet brings considerable performance improvement on LSPS. Meanwhile, it achieves competitive performance on existing person search benchmarks like CUHK-SYSU and PRW.
Yingji Zhong, Xiaoyu Wang 0002, Shiliang Zhang
CVPR2
2020 A Simple and Effective Framework for Pairwise Deep Metric Learning
Qi Qi 0006, Yan Yan 0006, Xiaoyu Wang 0002, Tianbao Yang
ECCV (27)4
2020 Accelerating Deep Learning with Millions of Classes
Zhuoning Yuan, Zhishuai Guo, Xiaotian Yu, Xiaoyu Wang 0002, Tianbao Yang
ECCV (23)4
2020 Stochastic Optimization for Non-convex Inf-Projection Problems
abstract
In this paper, we study a family of non-convex and possibly non-smooth inf-projection minimization problems, where the target objective function is equal to minimization of a joint function over another variable. This problem include difference of convex (DC) functions and a family of bi-convex functions as special cases. We develop stochastic algorithms and establish their first-order convergence for finding a (nearly) stationary solution of the target non-convex function under different conditions of the component functions. To the best of our knowledge, this is the first work that comprehensively studies stochastic optimization of non-convex inf-projection minimization problems with provable convergence guarantee. Our algorithms enable efficient stochastic optimization of a family of non-decomposable DC functions and a family of bi-convex functions. To demonstrate the power of the proposed algorithms we consider an important application in variance-based regularization. Experiments verify the effectiveness of our inf-projection based formulation and the proposed stochastic algorithm in comparison with previous stochastic algorithms based on the min-max formulation for achieving the same effect.
Yan Yan 0006, Yi Xu 0008, Lijun Zhang 0005, Xiaoyu Wang 0002, Tianbao Yang
ICML4
2020 Group Feedback Capsule Network
abstract
In capsule networks (CapsNets), the capsule is made up of collections of neurons. Their adjacent capsule layers are connected using routing-by-agreement mechanisms in an unsupervised way. The routing-by-agreement mechanisms have two main drawbacks: a) too many parameters and high computation complexity; b) the cluster distribution assumptions of these routing mechanisms may not hold in some complex real-world data. In this paper, we propose a novel Group Feedback Capsule Network (GF-CapsNet) which adopts a supervised routing strategy called group-routing. Compared with the previous routing strategies which globally transform each capsule, Group-routing equally splits capsules into groups where capsules locally share the same transformation weights, reducing routing parameters. To address the second drawback, we devise a distance network to directly predict capsules in a supervised way without making distribution assumptions. Our proposed group-routing captures local information of low-level capsules by group-wise transformation and supervisedly predicts high-level ones in a feedback way to address two drawbacks respectively. We conduct experiments on CIFAR-10/100 and SVHN datasets and the results show that our method can perform better against state-of-the-arts.
Xinpeng Ding, Nannan Wang 0001, Xinbo Gao 0001, Jie Li 0001, Xiaoyu Wang 0002, Tongliang Liu
IEEE Trans. Image Process.5
2020 Group-Group Loss-Based Global-Regional Feature Learning for Vehicle Re-Identification
abstract
Vehicle Re-Identification (Re-ID) is challenging because vehicles of the same model commonly show similar appearance. We tackle this challenge by proposing a Global-Regional Feature (GRF) that depicts extra local details to enhance discrimination power in addition to the global context. It is motivated by the observation that, vehicles of same color, maker, and model can be distinguished by their regional difference, e.g., the decorations on the windshields. To accelerate the GRF learning and promote its discrimination power, we propose a Group-Group Loss (GGL) to optimize the distance within and across vehicle image groups. Different from the siamese or triplet loss, GGL is directly computed on image groups rather than individual sample pairs or triplets. By avoiding traversing numerous sample combinations, GGL makes the model training easier and more efficient. Those two contributions highlight this work from previous methods on vehicle Re-ID task, which commonly learn global features with triplet loss or its variants. We evaluate our methods on two large-scale vehicle Re-ID datasets, i.e., VeRi and VehicleID. Experimental results show our methods achieve promising performance in comparison with recent works.
Shiliang Zhang, Xiaoyu Wang 0002, Richang Hong, Qi Tian 0001
IEEE Trans. Image Process.3
2019 A Robust Zero-Sum Game Framework for Pool-based Active Learning
abstract
In this paper, we present a novel robust zero- sum game framework for pool-based active learning grounded on advanced statistical learning theory. Pool-based active learning usually consists of two components, namely, learning of a classifier given labeled data and querying of unlabeled data for labeling. Most previous studies on active learning consider these as two separate tasks and propose various heuristics for selecting important unlabeled data for labeling, which may render the selection of unlabeled examples sub-optimal for minimizing the classification error. In contrast, the present work formulates active learning as a unified optimization framework for learning the classifier, i.e., the querying of labels and the learning of models are unified to minimize a common objective for statistical learning. In addition, the proposed method avoids the issues of many previous algorithms such as inefficiency, sampling bias and sensitivity to imbalanced data distribution. Besides theoretical analysis, we conduct extensive experiments on benchmark datasets and demonstrate the superior performance of the proposed active learning method compared with the state-of-the-art methods.
Dixian Zhu, Zhe Li 0008, Xiaoyu Wang 0002, Boqing Gong, Tianbao Yang
AISTATS3
2019 Group Reconstruction and Max-Pooling Residual Capsule Network
abstract
In capsule networks, the mapping of low-level capsules to high-level capsules is achieved by a routing-by-agreement algorithm. Since the capsule is made up of collections of neurons and the routing mechanism involves all the capsules instead of simply discarding some of the neurons like Max-Pooling, the capsule network has stronger representation ability than the traditional neural network. However, considering too much low-level capsules' information will cause its corresponding upper layer capsules to be interfered by other irrelevant information or noise capsules. Therefore, the original capsule network does not perform well on complex data structure. What's worse, computational complexity becomes a bottleneck in dealing with large data networks. In order to solve these shortcomings, this paper proposes a group reconstruction and max-pooling residual capsule network (GRMR-CapsNet). We build a block in which all capsules are divided into different groups and perform group reconstruction routing algorithm to obtain the corresponding high-level capsules. Between the lower and higher layers, Capsule Max-Pooling is adopted to prevent overfitting. We conduct experiments on CIFAR-10/100 and SVHN datasets and the results show that our method can perform better against state-of-the-arts.
Xinpeng Ding, Nannan Wang 0001, Xinbo Gao 0001, Jie Li 0001, Xiaoyu Wang 0002
IJCAI5
2019 Dual-alignment Feature Embedding for Cross-modality Person Re-identification
abstract
Person re-identification aims at searching pedestrians across different cameras, which is a key problem in video surveillance. With requirements in night environment, RGB-infrared person re-identification which could be regarded as a cross-modality matching problem, has gained increasing attention in recent years. Aside from cross-modality discrepancy, RGB-infrared person re-identification also suffers from human pose and view point differences. We design a dual-alignment feature embedding method to extract discriminative modality-invariant features. The concept of dual-alignment is two folds: spatial and modality alignments. We adopt the part-level features to extract fine-grained camera-invariant information. We introduce distribution loss function and correlation loss function to align the embedding features across visible and infrared modalities. Finally, we can extract modality-invariant features with robust and rich identity embeddings for cross-modality person re-identification. Experiment confirms that the proposed baseline and improvement achieves competitive results with the state-of-the-art methods on two datasets. For instance, We achieve (57.5+12.6)% rank-1 accuracy and (57.3+11.8)% mAP on the RegDB dataset.
Nannan Wang 0001, Xinbo Gao 0001, Jie Li 0001, Xiaoyu Wang 0002
ACM Multimedia5
2019 Joint Modeling of Dense and Incomplete Trajectories for Citywide Traffic Volume Inference
abstract
Real-time traffic volume inference is key to an intelligent city. It is a challenging task because accurate traffic volumes on the roads can only be measured at certain locations where sensors are installed. Moreover, the traffic evolves over time due to the influences of weather, events, holidays, etc. Existing solutions to the traffic volume inference problem often rely on dense GPS trajectories, which inevitably fail to account for the vehicles which carry no GPS devices or have them turned off. Consequently, the results are biased to taxicabs because they are almost always online for GPS tracking. In this paper, we propose a novel framework for the citywide traffic volume inference using both dense GPS trajectories and incomplete trajectories captured by camera surveillance systems. Our approach employs a high-fidelity traffic simulator and deep reinforcement learning to recover full vehicle movements from the incomplete trajectories. In order to jointly model the recovered trajectories and dense GPS trajectories, we construct spatiotemporal graphs and use multi-view graph embedding to encode the multi-hop correlations between road segments into real-valued vectors. Finally, we infer the citywide traffic volumes by propagating the traffic values of monitored road segments to the unmonitored ones through masked pairwise similarities. Extensive experiments with two big regions in a provincial capital city in China verify the effectiveness of our approach.
Xianfeng Tang, Boqing Gong, Yanwei Yu, Huaxiu Yao, Yandong Li, Haiyong Xie 0001, Xiaoyu Wang 0002
WWW7
2018 Deep Regionlets for Object Detection
Hongyu Xu, Xutao Lv, Xiaoyu Wang 0002, Zhou Ren, Navaneeth Bodla, Rama Chellappa
ECCV (11)3
2018 Adaptive Negative Curvature Descent with Applications in Non-convex Optimization
abstract
Negative curvature descent (NCD) method has been utilized to design deterministic or stochastic algorithms for non-convex optimization aiming at finding second-order stationary points or local minima. In existing studies, NCD needs to approximate the smallest eigen-value of the Hessian matrix with a sufficient precision (e.g., $\epsilon_2\ll 1$) in order to achieve a sufficiently accurate second-order stationary solution (i.e., $\lambda_{\min}(\nabla^2 f(\x))\geq -\epsilon_2)$. One issue with this approach is that the target precision $\epsilon_2$ is usually set to be very small in order to find a high quality solution, which increases the complexity for computing a negative curvature. To address this issue, we propose an adaptive NCD to allow for an adaptive error dependent on the current gradient's magnitude in approximating the smallest eigen-value of the Hessian, and to encourage competition between a noisy NCD step and gradient descent step. We consider the applications of the proposed adaptive NCD for both deterministic and stochastic non-convex optimization, and demonstrate that it can help reduce the the overall complexity in computing the negative curvatures during the course of optimization without sacrificing the iteration complexity.
Zhe Li 0008, Xiaoyu Wang 0002, Jinfeng Yi, Tianbao Yang
NeurIPS3
2018 RED-Net: A Recurrent Encoder-Decoder Network for Video-Based Face Alignment
Xi Peng 0005, Rogério Feris, Xiaoyu Wang 0002, Dimitris N. Metaxas
Int. J. Comput. Vis.3
2017 Deep Reinforcement Learning-Based Image Captioning with Embedding Reward
abstract
Image captioning is a challenging problem owing to the complexity in understanding the image content and diverse ways of describing it in natural language. Recent advances in deep neural networks have substantially improved the performance of this task. Most state-of-the-art approaches follow an encoder-decoder framework, which generates captions using a sequential recurrent prediction model. However, in this paper, we introduce a novel decision-making framework for image captioning. We utilize a policy network and a value network to collaboratively generate captions. The policy network serves as a local guidance by providing the confidence of predicting the next word according to the current state. Additionally, the value network serves as a global and lookahead guidance by evaluating all possible extensions of the current state. In essence, it adjusts the goal of predicting the correct words towards the goal of generating captions similar to the ground truth captions. We train both networks using an actor-critic reinforcement learning model, with a novel reward defined by visual-semantic embedding. Extensive experiments and analyses on the Microsoft COCO dataset show that the proposed framework outperforms state-of-the-art approaches across different evaluation metrics.
Zhou Ren, Xiaoyu Wang 0002, Xutao Lv, Li-Jia Li 0001
CVPR2
2017 Exploring Personalized Neural Conversational Models
abstract
Modeling dialog systems is currently one of the most active problems in Natural Language Processing. Recent advancement in Deep Learning has sparked an interest in the use of neural networks in modeling language, particularly for personalized conversational agents that can retain contextual information during dialog exchanges. This work carefully explores and compares several of the recently proposed neural conversation models, and carries out a detailed evaluation on the multiple factors that can significantly affect predictive performance, such as pretraining, embedding training, data cleaning, diversity reranking, evaluation setting, etc. Based on the tradeoffs of different models, we propose a new generative dialogue model conditioned on speakers as well as context history that outperforms all previous models on both retrieval and generative metrics. Our findings indicate that pretraining speaker embeddings on larger datasets, as well as bootstrapping word and speaker embeddings, can significantly improve performance (up to 3 points in perplexity), and that promoting diversity in using Mutual Information based techniques has a very strong effect in ranking metrics.
Satwik Kottur, Xiaoyu Wang 0002
IJCAI2
2016 A Recurrent Encoder-Decoder Network for Sequential Face Alignment
Xi Peng 0005, Rogério Feris, Xiaoyu Wang 0002, Dimitris N. Metaxas
ECCV (1)3
2016 Scalable Feature Matching by Dual Cascaded Scalar Quantization for Image Retrieval
abstract
In this paper, we investigate the problem of scalable visual feature matching in large-scale image search and propose a novel cascaded scalar quantization scheme in dual resolution. We formulate the visual feature matching as a range-based neighbor search problem and approach it by identifying hyper-cubes with a dual-resolution scalar quantization strategy. Specifically, for each dimension of the PCA-transformed feature, scalar quantization is performed at both coarse and fine resolutions. The scalar quantization results at the coarse resolution are cascaded over multiple dimensions to index an image database. The scalar quantization results over multiple dimensions at the fine resolution are concatenated into a binary super-vector and stored into the index list for efficient verification. The proposed cascaded scalar quantization (CSQ) method is free of the costly visual codebook training and thus is independent of any image descriptor training set. The index structure of the CSQ is flexible enough to accommodate new image features and scalable to index large-scale image database. We evaluate our approach on the public benchmark datasets for large-scale image retrieval. Experimental results demonstrate the competitive retrieval performance of the proposed method compared with several recent retrieval algorithms on feature quantization.
Wengang Zhou 0001, Ming Yang 0007, Xiaoyu Wang 0002, Houqiang Li, Yuanqing Lin, Qi Tian 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2015 Hyper-class augmented and regularized deep learning for fine-grained image classification
abstract
Deep convolutional neural networks (CNN) have seen tremendous success in large-scale generic object recognition. In comparison with generic object recognition, fine-grained image classification (FGIC) is much more challenging because (i) fine-grained labeled data is much more expensive to acquire (usually requiring domain expertise); (ii) there exists large intra-class and small inter-class variance. Most recent work exploiting deep CNN for image recognition with small training data adopts a simple strategy: pre-train a deep CNN on a large-scale external dataset (e.g., ImageNet) and fine-tune on the small-scale target data to fit the specific classification task. In this paper, beyond the fine-tuning strategy, we propose a systematic framework of learning a deep CNN that addresses the challenges from two new perspectives: (i) identifying easily annotated hyper-classes inherent in the fine-grained data and acquiring a large number of hyper-class-labeled images from readily available external sources (e.g., image search engines), and formulating the problem into multitask learning; (ii) a novel learning model by exploiting a regularization between the fine-grained recognition model and the hyper-class recognition model. We demonstrate the success of the proposed framework on two small-scale fine-grained datasets (Stanford Dogs and Stanford Cars) and on a large-scale car dataset that we collected.
Saining Xie, Tianbao Yang, Xiaoyu Wang 0002, Yuanqing Lin
CVPR3
2015 Regionlets for Generic Object Detection
abstract
Generic object detection is confronted by dealing with different degrees of variations, caused by viewpoints or deformations in distinct object classes, with tractable computations. This demands for descriptive and flexible object representations which can be efficiently evaluated in many locations. We propose to model an object class with a cascaded boosting classifier which integrates various types of features from competing local regions, each of which may consist of a group of subregions, named as regionlets. A regionlet is a base feature extraction region defined proportionally to a detection window at an arbitrary resolution (i.e., size and aspect ratio). These regionlets are organized in small groups with stable relative positions to be descriptive to delineate fine-grained spatial layouts inside objects. Their features are aggregated into a one-dimensional feature within one group so as to be flexible to tolerate deformations. The most discriminative regionlets for each object class are selected through a boosting learning procedure. Our regionlet approach achieves very competitive performance on popular multi-class detection benchmark datasets with a single method, without any context. It achieves a detection mean average precision of 41.7 percent on the PASCAL VOC 2007 dataset, and 39.7 percent on the VOC 2010 for 20 object categories. We further develop support pixel integral images to efficiently augment regionlet features with the responses learned by deep convolutional neural networks. Our regionlet based method won second place in the ImageNet Large Scale Visual Object Recognition Challenge (ILSVRC 2013).
Xiaoyu Wang 0002, Ming Yang 0007, Shenghuo Zhu, Yuanqing Lin
IEEE Trans. Pattern Anal. Mach. Intell.1
2015 Semantic-Aware Co-Indexing for Image Retrieval
abstract
In content-based image retrieval, inverted indexes allow fast access to database images and summarize all knowledge about the database. Indexing multiple clues of image contents allows retrieval algorithms search for relevant images from different perspectives, which is appealing to deliver satisfactory user experiences. However, when incorporating diverse image features during online retrieval, it is challenging to ensure retrieval efficiency and scalability. In this paper, for large-scale image retrieval, we propose a semantic-aware co-indexing algorithm to jointly embed two strong cues into the inverted indexes: 1) local invariant features that are robust to delineate low-level image contents, and 2) semantic attributes from large-scale object recognition that may reveal image semantic meanings. Specifically, for an initial set of inverted indexes of local features, we utilize semantic attributes to filter out isolated images and insert semantically similar images to this initial set. Encoding these two distinct and complementary cues together effectively enhances the discriminative capability of inverted indexes. Such co-indexing operations are totally off-line and introduce small computation overhead to online retrieval, because only local features but no semantic attributes are employed for the query. Hence, this co-indexing is different from existing image retrieval methods fusing multiple features or retrieval results. Extensive experiments and comparisons with recent retrieval methods manifest the competitive performance of our method.
Shiliang Zhang, Ming Yang 0007, Xiaoyu Wang 0002, Yuanqing Lin, Qi Tian 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2015 Cross Indexing With Grouplets
abstract
Most of the current image indexing systems for retrieval view a database as a set of individual images. It limits the flexibility of the retrieval framework to conduct sophisticated cross-image analysis, resulting in higher memory consumption and sub-optimal retrieval accuracy. To conquer this issue, we propose cross indexing with grouplets, where the core idea is to view the database images as a set of grouplets, each of which is defined as a group of highly relevant images. Because a grouplet groups similar images together, the number of grouplets is smaller than the number of images, thus naturally leading to less memory cost. Moreover, the definition of a grouplet could be based on customized relations, allowing for seamless integration of advanced image features and data mining techniques like the deep convolutional neural network (DCNN) in off-line indexing . To validate the proposed framework, we construct three different types of grouplets , which are respectively based on local similarity , regional relation, and global semantic modeling. Extensive experiments on public benchmark datasets demonstrate the efficiency and superior performance of our approach.
Shiliang Zhang, Xiaoyu Wang 0002, Yuanqing Lin, Qi Tian 0001
IEEE Trans. Multim.2
2014 Accurate Object Detection with Location Relaxation and Regionlets Re-localization
Chengjiang Long, Xiaoyu Wang 0002, Gang Hua 0001, Ming Yang 0007, Yuanqing Lin
ACCV (1)2
2014 Towards Codebook-Free: Scalable Cascaded Hashing for Mobile Image Search
abstract
State-of-the-art image retrieval algorithms using local invariant features mostly rely on a large visual codebook to accelerate the feature quantization and matching. This codebook typically contains millions of visual words, which not only demands for considerable resources to train offline but also consumes large amount of memory at the online retrieval stage. This is hardly affordable in resource limited scenarios such as mobile image search applications. To address this issue, we propose a codebook-free algorithm for large scale mobile image search. In our method, we first employ a novel scalable cascaded hashing scheme to ensure the recall rate of local feature matching. Afterwards, we enhance the matching precision by an efficient verification with the binary signatures of these local features. Consequently, our method achieves fast and accurate feature matching free of a huge visual codebook. Moreover, the quantization and binarizing functions in the proposed scheme are independent of small collections of training images and generalize well for diverse image datasets. Evaluated on two public datasets with a million distractor images, the proposed algorithm demonstrates competitive retrieval accuracy and scalability against four recent retrieval methods in literature.
Wengang Zhou 0001, Ming Yang 0007, Houqiang Li, Xiaoyu Wang 0002, Yuanqing Lin, Qi Tian 0001
IEEE Trans. Multim.4
2013 Regionlets for Generic Object Detection
abstract
Generic object detection is confronted by dealing with different degrees of variations in distinct object classes with tractable computations, which demands for descriptive and flexible object representations that are also efficient to evaluate for many locations. In view of this, we propose to model an object class by a cascaded boosting classifier which integrates various types of features from competing local regions, named as region lets. A region let is a base feature extraction region defined proportionally to a detection window at an arbitrary resolution (i.e. size and aspect ratio). These region lets are organized in small groups with stable relative positions to delineate fine grained spatial layouts inside objects. Their features are aggregated to a one-dimensional feature within one group so as to tolerate deformations. Then we evaluate the object bounding box proposal in selective search from segmentation cues, limiting the evaluation locations to thousands. Our approach significantly outperforms the state-of-the-art on popular multi-class detection benchmark datasets with a single method, without any contexts. It achieves the detection mean average precision of 41.7% on the PASCAL VOC 2007 dataset and 39.7% on the VOC 2010 for 20 object categories. It achieves 14.7% mean average precision on the Image Net dataset for 200 object categories, outperforming the latest deformable part-based model (DPM) by 4.7%.
Xiaoyu Wang 0002, Ming Yang 0007, Shenghuo Zhu, Yuanqing Lin
ICCV1
2013 Semantic-Aware Co-indexing for Image Retrieval
abstract
Inverted indexes in image retrieval not only allow fast access to database images but also summarize all knowledge about the database, so that their discriminative capacity largely determines the retrieval performance. In this paper, for vocabulary tree based image retrieval, we propose a semantic-aware co-indexing algorithm to jointly embed two strong cues into the inverted indexes: 1) local invariant features that are robust to delineate low-level image contents, and 2) semantic attributes from large-scale object recognition that may reveal image semantic meanings. For an initial set of inverted indexes of local features, we utilize 1000 semantic attributes to filter out isolated images and insert semantically similar images to the initial set. Encoding these two distinct cues together effectively enhances the discriminative capability of inverted indexes. Such co-indexing operations are totally off-line and introduce small computation overhead to online query cause only local features but no semantic attributes are used for query. Experiments and comparisons with recent retrieval methods on 3 datasets, i.e., UKbench, Holidays, Oxford5K, and 1.3 million images from Flickr as distractors, manifest the competitive performance of our method.
Shiliang Zhang, Ming Yang 0007, Xiaoyu Wang 0002, Yuanqing Lin, Qi Tian 0001
ICCV3
2012 Histogram of Oriented Normal Vectors for Object Recognition with a Depth Sensor
Xiaoyu Wang 0002, Xutao Lv, Tony X. Han, James Keller 0001, Zhihai He, Marjorie Skubic, Shihong Lao
ACCV (2)2
2012 Detection by detections: Non-parametric detector adaptation for a video
abstract
We propose an approach to improving the detection results of a generic offline trained detector on a specific video. Our method does not leverage visual tracking as most detection by tracking methods do. Instead, the proposed detection by detections approach can serve as a more confident initialization for detection by tracking methods. Different from other supervised detector adaptation methods, we constrain the task to videos and no supervised labels for the target video are required for the adaptation; we intend to fill the gap between detection by tracking and pure detection by frames. As a non-parametric detector adaptation method, confident detections are collected to re-rank and to group other detections. We focus on methods with high precision detection results since it is necessitated in real application. Extensive experiments with two state-of-the-art detectors demonstrate the efficacy of our approach.
Xiaoyu Wang 0002, Gang Hua 0001, Tony X. Han
CVPR1
2011 Contextual weighting for vocabulary tree based image retrieval
abstract
In this paper we address the problem of image retrieval from millions of database images. We improve the vocabulary tree based approach by introducing contextual weighting of local features in both descriptor and spatial domains. Specifically, we propose to incorporate efficient statistics of neighbor descriptors both on the vocabulary tree and in the image spatial domain into the retrieval. These contextual cues substantially enhance the discriminative power of individual local features with very small computational overhead. We have conducted extensive experiments on benchmark datasets, i.e., the UKbench, Holidays, and our new Mobile dataset, which show that our method reaches state-of-the-art performance with much less computation. Furthermore, the proposed method demonstrates excellent scalability in terms of both retrieval accuracy and efficiency on large-scale experiments using 1.26 million images from the ImageNet database as distractors.
Xiaoyu Wang 0002, Ming Yang 0007, Timothée Cour, Shenghuo Zhu, Kai Yu 0001, Tony X. Han
ICCV1
2010 Discriminative Tracking by Metric Learning
Xiaoyu Wang 0002, Gang Hua 0001, Tony X. Han
ECCV (3)1
2009 An HOG-LBP human detector with partial occlusion handling
abstract
By combining Histograms of Oriented Gradients (HOG) and Local Binary Pattern (LBP) as the feature set, we propose a novel human detection approach capable of handling partial occlusion. Two kinds of detectors, i.e., global detector for whole scanning windows and part detectors for local regions, are learned from the training data using linear SVM. For each ambiguous scanning window, we construct an occlusion likelihood map by using the response of each block of the HOG feature to the global detector. The occlusion likelihood map is then segmented by Mean-shift approach. The segmented portion of the window with a majority of negative response is inferred as an occluded region. If partial occlusion is indicated with high likelihood in a certain scanning window, part detectors are applied on the unoccluded regions to achieve the final classification on the current scanning window. With the help of the augmented HOG-LBP feature and the global-part occlusion handling method, we achieve a detection rate of 91.3% with FPPW= 10−6, 94.7% with FPPW= 10−5, and 97.9% with FPPW= 10−4on the INRIA dataset, which, to our best knowledge, is the best human detection performance on the INRIA dataset. The global-part occlusion handling method is further validated using synthesized occlusion data constructed from the INRIA and Pascal dataset.
Xiaoyu Wang 0002, Tony X. Han, Shuicheng Yan
ICCV1