EDBT 2026 Demo / reviewers in the wild / expert
Jiyang Xie 0001
dblp:162/0060
· DBLP profile ↗
31ranked-venue papers
5as first author
24since 2021 · last 2025
0000-0003-3659-9476ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 3 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 2 first-author · 12 since 2021Computer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Interactive triplet attention for few-shot fine-grained image classificationabstractFew-shot fine-grained classification aims to identify novel fine-grained classes from extremely few examples with ultra-high semantic similarity between classes, hence a notoriously hard task. To extract discriminative features from few samples for recognizing subtle differences between fine-grained classes , it is pivotal to exploit comprehensive interactions across all dimensions in space and channel, which, however, is unexplored yet by state-of-the-art methods in this challenging area. To address this issue, in this paper we show that a simple adjustment to the existing triplet attention module (TAM) can be highly effective for few-shot fine-grained image classification. More specifically, building on TAM which comprises three parallel branches for pairwise interactions between height, width, and channel dimensions, we introduce an additional interaction between the output of these three branches, capable of modeling the dependency across all three dimensions; the revised method is dubbed interactive triplet attention module (ITAM). ITAM is a plug-and-play module, which can be inserted into any metric-based few-shot fine-grained image classifiers for performance enhancement. Extensive experiments, on CUB-200–2011, Flowers, Stanford-Cars, and Stanford-Dogs, showcase the superiority of ITAM against state-of-the-art few-shot fine-grained image classifiers. Shaoying Xue, Jiyang Xie 0001, Zhanyu Ma, Jing-Hao Xue |
Neurocomputing | 3 |
| 2025 | CDN4: A cross-view Deep Nearest Neighbor Neural Network for fine-grained few-shot classificationabstractThe fine-grained few-shot classification is a challenging task in computer vision, aiming to classify images with subtle and detailed differences given scarce labeled samples. A promising avenue to tackle this challenge is to use spatially local features to densely measure the similarity between query and support samples. Compared with image-level global features, local features contain more low-level information that is rich and transferable across categories. However, methods based on spatially localized features have difficulty distinguishing subtle category differences due to the lack of sample diversity. To address this issue, we propose a novel method called Cross-view Deep Nearest Neighbor Neural Network (CDN4). CDN4 applies a random geometric transformation to augment a different view of support and query samples and subsequently exploits four similarities between the original and transformed views of query local features and those views of support local features. The geometric augmentation increases the diversity between samples of the same class, and the cross-view measurement encourages the model to focus more on discriminative local features for classification through the cross-measurements between the two branches. Extensive experiments validate the superiority of CDN4, which achieves new state-of-the-art results in few-shot classification across various fine-grained benchmarks. Code is available at . • A novel fine-grained FSL method for improving local feature discriminativeness. • Construct the Episodic Dual-Branch Structure to enhance sample diversity. • Exploit four cross-view metric pairs to enforce learning of discriminative features. • CDN4 achieves state-of-the-art performance on three fine-grained benchmark datasets. Jiyang Xie 0001, Zhanyu Ma, Jing-Hao Xue |
Pattern Recognit. | 3 |
| 2024 | Self-reconstruction network for fine-grained few-shot classificationabstractMetric-based methods are one of the most common methods to solve the problem of few-shot image classification. However, traditional metric-based few-shot methods suffer from overfitting and local feature misalignment. The recently proposed feature reconstruction-based approach, which reconstructs query image features from the support set features of a given class and compares the distance between the original query features and the reconstructed query features as the classification criterion, effectively solves the feature misalignment problem. However, the issue of overfitting still has not been considered. To this end, we propose a self-reconstruction metric module for diversifying query features and a restrained cross-entropy loss for avoiding over-confident predictions. By introducing them, the proposed self-reconstruction network can effectively alleviate overfitting. Extensive experiments on five benchmark fine-grained datasets demonstrate that our proposed method achieves state-of-the-art performance on both 5-way 1-shot and 5-way 5-shot classification tasks. Code is available at https://github.com/liz-lut/SRM-main. Zhen Li 0026, Jiyang Xie 0001, Jing-Hao Xue, Zhanyu Ma |
Pattern Recognit. | 3 |
| 2023 | Multi-Head Uncertainty Inference for Adversarial Attack DetectionabstractDeep neural networks (DNNs) are sensitive and susceptible to tiny perturbations by adversarial attacks which cause erroneous predictions. Various methods, including adversarial defense and uncertainty inference (UI), have been developed to overcome adversarial attacks in recent years. In this paper, we propose a multi-head uncertainty inference (MH-UI) framework for detecting adversarial attack examples. We adopt a multi-head architecture with multiple prediction heads (i.e., classifiers) to obtain predictions from different depths in the DNNs and introduce shallow information for the UI. Using independent heads at different depths, the normalized predictions are assumed to follow the same Dirichlet distribution, and we estimate the distribution parameter of it by moment matching. Cognitive uncertainty brought by the adversarial attacks will be reflected and amplified in the distribution. Experimental results show that the proposed MH-UI framework has good performance in different settings of adversarial attack detection tasks. Songyun Yang, Jiyang Xie 0001, Zhongwei Si, Ke Zhang 0005, Kongming Liang |
ICASSP | 3 |
| 2023 | Searching for Network Width With Bilaterally Coupled NetworkabstractSearching for a more compact network width recently serves as an effective way of channel pruning for the deployment of convolutional neural networks (CNNs) under hardware constraints. To fulfil the searching, a one-shot supernet is usually leveraged to efficiently evaluate the performance w.r.t. different network widths. However, current methods mainly follow a unilaterally augmented (UA) principle for the evaluation of each width, which induces the training unfairness of channels in supernet. In this article, we introduce a new supernet called Bilaterally Coupled Network (BCNet) to address this issue. In BCNet, each channel is fairly trained and responsible for the same amount of network widths, thus each network width can be evaluated more accurately. Besides, we propose to reduce the redundant search space and present the BCNetV2 as the enhanced supernet to ensure rigorous training fairness over channels. Furthermore, we leverage a stochastic complementary strategy for training the BCNet, and propose a prior initial population sampling method to boost the performance of the evolutionary search. We also propose a new open-source width search benchmark on macro structures named Channel-Bench-Macro for the better comparisons of the width search algorithms with MobileNet- and ResNet-like architectures. Extensive experiments on the benchmark datasets demonstrate that our method can achieve state-of-the-art performance. Xiu Su, Shan You, Jiyang Xie 0001, Fei Wang 0032, Chen Qian 0006, Changshui Zhang, Chang Xu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Sketch-Segformer: Transformer-Based Segmentation for Figurative and Creative SketchesabstractSketch is a well-researched topic in the vision community by now. Sketch semantic segmentation in particular, serves as a fundamental step towards finer-level sketch interpretation. Recent works use various means of extracting discriminative features from sketches and have achieved considerable improvements on segmentation accuracy. Common approaches for this include attending to the sketch-image as a whole, its stroke-level representation or the sequence information embedded in it. However, they mostly focus on only a part of such multi-facet information. In this paper, we for the first time demonstrate that there is complementary information to be explored across all these three facets of sketch data, and that segmentation performance consequently benefits as a result of such exploration of sketch-specific information. Specifically, we propose the Sketch-Segformer, a transformer-based framework for sketch semantic segmentation that inherently treats sketches as stroke sequences other than pixel-maps. In particular, Sketch-Segformer introduces two types of self-attention modules having similar structures that work with different receptive fields (i.e., whole sketch or individual stroke). The order embedding is then further synergized with spatial embeddings learned from the entire sketch as well as localized stroke-level information. Extensive experiments show that our sketch-specific design is not only able to obtain state-of-the-art performance on traditional figurative sketches (such as SPG, SketchSeg-150K datasets), but also performs well on creative sketches that do not conform to conventional object semantics (CreativeSketch dataset) thanks for our usage of multi-facet sketch information. Ablation studies, visualizations, and invariance tests further justifies our design choice and the effectiveness of Sketch-Segformer. Codes are available at https://github.com/PRIS-CV/Sketch-SF. Yixiao Zheng, Jiyang Xie 0001, Aneeshan Sain, Yi-Zhe Song, Zhanyu Ma |
IEEE Trans. Image Process. | 2 |
| 2023 | On the Comparisons of Decorrelation Approaches for Non-Gaussian Neutral Vector Variablesabstract-norm equals one. In addition, its neutral properties make it significantly different from the commonly studied vector variables (e.g., the Gaussian vector variables). Due to the aforementioned properties, the conventionally applied linear transformation approaches [e.g., principal component analysis (PCA) and independent component analysis (ICA)] are not suitable for neutral vector variables, as PCA cannot transform a neutral vector variable, which is highly negatively correlated, into a set of mutually independent scalar variables and ICA cannot preserve the bounded property after transformation. In recent work, we proposed an efficient nonlinear transformation approach, i.e., the parallel nonlinear transformation (PNT), for decorrelating neutral vector variables. In this article, we extensively compare PNT with PCA and ICA through both theoretical analysis and experimental evaluations. The results of our investigations demonstrate the superiority of PNT for decorrelating the neutral vector variables. Zhanyu Ma, Xiaoou Lu, Jiyang Xie 0001, Zhen Yang 0004, Jing-Hao Xue, Zheng-Hua Tan, Bo Xiao 0006, Jun Guo 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | ViTAS: Vision Transformer Architecture Search
Xiu Su, Shan You, Jiyang Xie 0001, Mingkai Zheng, Fei Wang 0032, Chen Qian 0006, Changshui Zhang, Xiaogang Wang 0001, Chang Xu 0002 |
ECCV (21) | 3 |
| 2022 | ScaleNet: Searching for the Model to Scale
Jiyang Xie 0001, Xiu Su, Shan You, Zhanyu Ma, Fei Wang 0032, Chen Qian 0006 |
ECCV (21) | 1 |
| 2022 | Structured Dropconnect for Uncertainty Inference in Image ClassificationabstractUncertainty inference has become an important task to prove the reliability for deep neural networks. For image classification tasks, we propose a structured DropConnect (SDC) framework to model the output of a deep neural network into a distribution. We introduce a DropConnect strategy in a structured manner in the fully connected layers, split the network into several sub-networks during testing, and choose the Dirichlet distribution to model the outputs of these sub-networks. The entropy of the parameterized Dirichlet distribution is finally utilized for uncertainty inference. In this paper, this framework is implemented on VGG16, and ResNet18 models for misclassification detection and open-set out-of-domain detection on CIFAR-10 and CIFAR-100 datasets. Experimental results show that the performance of the proposed SDC can be comparable to other uncertainty inference methods. Wenqing Zheng, Jiyang Xie 0001, Zhanyu Ma |
ICIP | 2 |
| 2022 | Cross-Layer Feature based Multi-Granularity Visual ClassificationabstractIn contrast to traditional fine-grained visual clas-sification, multi-granularity visual classification is no longer limited to identifying the different sub-classes belonging to the same super-class (e.g., bird species, cars, and aircraft models). Instead, it gives a sequence of labels from coarse to fine (e.g., Passeriformes → Corvidae → Fish Crow), which is more convenient in practice. The key to solving this task is how to use the relationships between the different levels of labels to learn feature representations that contain different levels of granularity. Interestingly, the feature pyramid structure naturally implies different granularity of feature representation, with the shallow layers representing coarse-grained features and the deep layers representing fine-grained features. Therefore, in this paper, we exploit this property of the feature pyramid structure to decouple features and obtain feature representations corre-sponding to different granularities. Specifically, we use shallow features for coarse-grained classification and deep features for fine-grained classification. In addition, to enable fine-grained features to enhance the coarse-grained classification, we propose a feature reinforcement module based on the feature pyramid structure, where deep features are first upsampled and then combined with shallow features to make decisions. Experimental results on three widely used fine-grained image classification datasets such as CUB-200-2011, Stanford Cars, and FGVC-Aircraft validate the method's effectiveness. Code available at https://github.com/PRIS-CV/CGVC. Junhan Chen, Dongliang Chang, Jiyang Xie 0001, Ruoyi Du, Zhanyu Ma |
VCIP | 3 |
| 2022 | ENDE-GNN: An Encoder-decoder GNN Framework for Sketch Semantic SegmentationabstractSketch semantic segmentation serves as an important part of sketch interpretation. Recently, some researchers have obtained significant results using graph neural networks (GNN) for this task. However, existing GNN-based methods usually neglect the drawing order of sketches thus missing out the sequence information inherent to sketches. Towards solving this problem to achieve better performance on sketch semantic segmentation, we propose an encoder-decoder GNN framework named ENDE-GNN. Working with an auxiliary decoder, our ENDE-GNN guides the GNN backbone network to not only extract the inter-stroke and intra-stroke features, but also pays attention to the drawing order of sketches. This decoder acts during training only, preventing any additional overhead during testing. The proposed ENDE-GNN obtains state-of-the-art per-formances on three public sketch semantic segmentation datasets, namely SPG, SketchSeg-150K, and CreativeSketch. We further evaluate the effectiveness of ENDE-GNN via ablation studies and visualizations. Codes are available at https://github.com/PRIS-CV/ENDE_For_SSS. Yixiao Zheng, Jiyang Xie 0001, Aneeshan Sain, Zhanyu Ma, Yi-Zhe Song, Jun Guo 0002 |
VCIP | 2 |
| 2022 | Dual-granularity feature alignment for cross-modality person re-identification
Junhui Yin, Zhanyu Ma, Jiyang Xie 0001, Shibo Nie, Kongming Liang, Jun Guo 0002 |
Neurocomputing | 3 |
| 2022 | Progressive Learning of Category-Consistent Multi-Granularity Features for Fine-Grained Visual ClassificationabstractFine-grained visual classification (FGVC) is much more challenging than traditional classification tasks due to the inherently subtle intra-class object variations. Recent works are mainly part-driven (either explicitly or implicitly), with the assumption that fine-grained information naturally rests within the parts. In this paper, we take a different stance, and show that part operations are not strictly necessary - the key lies with encouraging the network to learn at different granularities and progressively fusing multi-granularity features together. In particular, we propose: (i) a progressive training strategy that effectively fuses features from different granularities, and (ii) a consistent block convolution that encourages the network to learn the category-consistent features at specific granularities. We evaluate on several standard FGVC benchmark datasets, and demonstrate the proposed method consistently outperforms existing alternatives or delivers competitive results. Codes are available at https://github.com/PRIS-CV/PMG-V2. Ruoyi Du, Jiyang Xie 0001, Zhanyu Ma, Dongliang Chang, Yi-Zhe Song, Jun Guo 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | GPCA: A Probabilistic Framework for Gaussian Process Embedded Channel AttentionabstractChannel attention mechanisms have been commonly applied in many visual tasks for effective performance improvement. It is able to reinforce the informative channels as well as to suppress the useless channels. Recently, different channel attention modules have been proposed and implemented in various ways. Generally speaking, they are mainly based on convolution and pooling operations. In this paper, we propose Gaussian process embedded channel attention (GPCA) module and further interpret the channel attention schemes in a probabilistic way. The GPCA module intends to model the correlations among the channels, which are assumed to be captured by beta distributed variables. As the beta distribution cannot be integrated into the end-to-end training of convolutional neural networks (CNNs) with a mathematically tractable solution, we utilize an approximation of the beta distribution to solve this problem. To specify, we adapt a Sigmoid-Gaussian approximation, in which the Gaussian distributed variables are transferred into the interval [0,1]. The Gaussian process is then utilized to model the correlations among different channels. In this case, a mathematically tractable solution is derived. The GPCA module can be efficiently implemented and integrated into the end-to-end training of the CNNs. Experimental results demonstrate the promising performance of the proposed GPCA module. Codes are available at https://github.com/PRIS-CV/GPCA. Jiyang Xie 0001, Zhanyu Ma, Dongliang Chang, Guoqiang Zhang 0003, Jun Guo 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Advanced Dropout: A Model-Free Methodology for Bayesian Dropout OptimizationabstractDue to lack of data, overfitting ubiquitously exists in real-world applications of deep neural networks (DNNs). We propose advanced dropout, a model-free methodology, to mitigate overfitting and improve the performance of DNNs. The advanced dropout technique applies a model-free and easily implemented distribution with parametric prior, and adaptively adjusts dropout rate. Specifically, the distribution parameters are optimized by stochastic gradient variational Bayes in order to carry out an end-to-end training. We evaluate the effectiveness of the advanced dropout against nine dropout techniques on seven computer vision datasets (five small-scale datasets and two large-scale datasets) with various base models. The advanced dropout outperforms all the referred techniques on all the datasets. We further compare the effectiveness ratios and find that advanced dropout achieves the highest one on most cases. Next, we conduct a set of analysis of dropout rate characteristics, including convergence of the adaptive dropout rate, the learned distributions of dropout masks, and a comparison with dropout rate generation without an explicit distribution. In addition, the ability of overfitting prevention is evaluated and confirmed. Finally, we extend the application of the advanced dropout to uncertainty inference, network pruning, text classification, and regression. The proposed advanced dropout is also superior to the corresponding referred methods. Codes are available at https://github.com/PRIS-CV/AdvancedDropout. Jiyang Xie 0001, Zhanyu Ma, Jianjun Lei 0001, Guoqiang Zhang 0003, Jing-Hao Xue, Zheng-Hua Tan, Jun Guo 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | MPCCL: Multiview predictive coding with contrastive learning for person re-identification
Junhui Yin, Jiyang Xie 0001, Zhanyu Ma, Jun Guo 0002 |
Pattern Recognit. | 2 |
| 2022 | Unsupervised person re-identification via simultaneous clustering and mask prediction
Junhui Yin, Siqing Zhang 0001, Jiyang Xie 0001, Zhanyu Ma, Jun Guo 0002 |
Pattern Recognit. | 3 |
| 2022 | Learning Calibrated Class Centers for Few-Shot Classification by Pair-Wise SimilarityabstractMetric-based methods achieve promising performance on few-shot classification by learning clusters on support samples and generating shared decision boundaries for query samples. However, existing methods ignore the inaccurate class center approximation introduced by the limited number of support samples, which consequently leads to biased inference. Therefore, in this paper, we propose to reduce the approximation error by class center calibration. Specifically, we introduce the so-called Pair-wise Similarity Module (PSM) to generate calibrated class centers adapted to the query sample by capturing the semantic correlations between the support and the query samples, as well as enhancing the discriminative regions on support representation. It is worth noting that the proposed PSM is a simple plug-and-play module and can be inserted into most metric-based few-shot learning models. Through extensive experiments in metric-based models, we demonstrate that the module significantly improves the performance of conventional few-shot classification methods on four few-shot image classification benchmark datasets. Codes are available at: https://github.com/PRIS-CV/Pair-wise-Similarity-module. Yurong Guo 0001, Ruoyi Du, Jiyang Xie 0001, Zhanyu Ma |
IEEE Trans. Image Process. | 4 |
| 2022 | Dirichlet Process Mixture of Generalized Inverted Dirichlet Distributions for Positive Vector Data With Extended Variational InferenceabstractA Bayesian nonparametric approach for estimation of a Dirichlet process (DP) mixture of generalized inverted Dirichlet distributions [i.e., an infinite generalized inverted Dirichlet mixture model (InGIDMM)] has been proposed. The generalized inverted Dirichlet distribution has been proven to be efficient in modeling the vectors that contain only positive elements. Under the classical variational inference (VI) framework, the key challenge in the Bayesian estimation of InGIDMM is that the expectation of the joint distribution of data and variables cannot be explicitly calculated. Therefore, numerical methods are usually applied to simulate the optimal posterior distributions. With the recently proposed extended VI (EVI) framework, we introduce lower bound approximations to the original variational objective function in the VI framework such that an analytically tractable solution can be derived. Hence, the problem in numerical simulation has been overcome. By applying the DP mixture technique, an InGIDMM can automatically determine the number of mixture components from the observed data. Moreover, the DP mixture model with an infinite number of mixture components also avoids the problems of underfitting and overfitting. The performance of the proposed approach is demonstrated with both synthesized data and real-life data applications. Zhanyu Ma, Yuping Lai, Jiyang Xie 0001, Deyu Meng, W. Bastiaan Kleijn, Jun Guo 0002, Jingyi Yu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | Part Uncertainty Estimation Convolutional Neural Network For Person Re-IdentificationabstractDue to the large amount of noisy data in person re-identification (ReID) task, the ReID models are usually affected by the data uncertainty. Therefore, the deep uncertainty estimation method is important for improving the model robustness and matching accuracy. To this end, we propose a part-based uncertainty convolutional neural network (PUCNN), which introduces the part-based uncertainty estimation into the baseline model. On the one hand, PUCNN improves the model robustness to noisy data by distributilizing the feature embedding and constraining the part-based uncertainty. On the other hand, PUCNN improves the cumulative matching characteristics (CMC) performance of the model by filtering out low-quality training samples according to the estimated uncertainty score. The experiments on both non-video datasets, the noised Market-1501 and DukeMTMC, and video datasets, PRID2011, iLiDS-VID and MARS, demonstrate that our proposed method achieves encouraging and promising performance. Wenyu Sun, Jiyang Xie 0001, Jiayan Qiu, Zhanyu Ma |
ICIP | 2 |
| 2021 | Cross-layer Navigation Convolutional Neural Network for Fine-grained Visual ClassificationabstractFine-grained visual classification (FGVC) aims to classify sub-classes of objects in the same super-class (e.g., species of birds, models of cars). For the FGVC tasks, the essential solution is to find discriminative subtle information of the target from local regions. Traditional FGVC models preferred to use the refined features, i.e., high-level semantic information for recognition and rarely use low-level information. However, it turns out that low-level information which contains rich detail information also has effect on improving performance. Therefore, in this paper, we propose cross-layer navigation convolutional neural network for feature fusion. First, the feature maps extracted by the backbone network are fed into a convolutional long short-term memory model sequentially from high-level to low-level to perform feature aggregation. Then, attention mechanisms are used after feature fusion to extract spatial and channel information while linking the high-level semantic information and the low-level texture features, which can better locate the discriminative regions for the FGVC. In the experiments, three commonly used FGVC datasets, including CUB-200-2011, Stanford-Cars, and FGVC-Aircraft datasets, are used for evaluation and we demonstrate the superiority of the proposed method by comparing it with other referred FGVC methods to show that this method achieves superior results. https://github.com/PRIS-CV/CN-CNN.git Chenyu Guo, Jiyang Xie 0001, Kongming Liang, Zhanyu Ma |
MMAsia | 2 |
| 2021 | AP-CNN: Weakly Supervised Attention Pyramid Convolutional Neural Network for Fine-Grained Visual ClassificationabstractClassifying the sub-categories of an object from the same super-category (e.g., bird species and cars) in fine-grained visual classification (FGVC) highly relies on discriminative feature representation and accurate region localization. Existing approaches mainly focus on distilling information from high-level features. In this article, by contrast, we show that by integrating low-level information (e.g., color, edge junctions, texture patterns), performance can be improved with enhanced feature representation and accurately located discriminative regions. Our solution, named Attention Pyramid Convolutional Neural Network (AP-CNN), consists of 1) a dual pathway hierarchy structure with a top-down feature pathway and a bottom-up attention pathway, hence learning both high-level semantic and low-level detailed feature representation, and 2) an ROI-guided refinement strategy with ROI-guided dropblock and ROI-guided zoom-in operation, which refines features with discriminative local regions enhanced and background noises eliminated. The proposed AP-CNN can be trained end-to-end, without the need of any additional bounding box/part annotation. Extensive experiments on three popularly tested FGVC datasets (CUB-200-2011, Stanford Cars, and FGVC-Aircraft) demonstrate that our approach achieves state-of-the-art performance. Models and code are available at https://github.com/PRIS-CV/AP-CNN_Pytorch-master. Zhanyu Ma, Shaoguo Wen, Jiyang Xie 0001, Dongliang Chang, Zhongwei Si, Ming Wu 0001, Haibin Ling |
IEEE Trans. Image Process. | 4 |
| 2021 | DS-UI: Dual-Supervised Mixture of Gaussian Mixture Models for Uncertainty Inference in Image RecognitionabstractThis paper proposes a dual-supervised uncertainty inference (DS-UI) framework for improving Bayesian estimation-based UI in DNN-based image recognition. In the DS-UI, we combine the classifier of a DNN, i.e., the last fully-connected (FC) layer, with a mixture of Gaussian mixture models (MoGMM) to obtain an MoGMM-FC layer. Unlike existing UI methods for DNNs, which only calculate the means or modes of the DNN outputs' distributions, the proposed MoGMM-FC layer acts as a probabilistic interpreter for the features that are inputs of the classifier to directly calculate the probabilities of them for the DS-UI. In addition, we propose a dual-supervised stochastic gradient-based variational Bayes (DS-SGVB) algorithm for the MoGMM-FC layer optimization. Unlike conventional SGVB and optimization algorithms in other UI methods, the DS-SGVB not only models the samples in the specific class for each Gaussian mixture model (GMM) in the MoGMM, but also considers the negative samples from other classes for the GMM to reduce the intra-class distances and enlarge the inter-class margins simultaneously for enhancing the learning ability of the MoGMM-FC layer in the DS-UI. Experimental results show the DS-UI outperforms the state-of-the-art UI methods in misclassification detection. We further evaluate the DS-UI in open-set out-of-domain/-distribution detection and find statistically significant improvements. Visualizations of the feature spaces demonstrate the superiority of the DS-UI. Codes are available at https://github.com/PRIS-CV/DS-UI. Jiyang Xie 0001, Zhanyu Ma, Jing-Hao Xue, Guoqiang Zhang 0003, Yinhe Zheng, Jun Guo 0002 |
IEEE Trans. Image Process. | 1 |
| 2020 | Fine-Grained Visual Classification via Progressive Multi-granularity Training of Jigsaw Patches
Ruoyi Du, Dongliang Chang, Ayan Kumar Bhunia, Jiyang Xie 0001, Zhanyu Ma, Yi-Zhe Song, Jun Guo 0002 |
ECCV (20) | 4 |
| 2020 | IU-Module: Intersection and Union Module for Fine-Grained Visual ClassificationabstractA predominant viewpoint in previous works of fine-grained visual classification (FGVC) is to the localize discriminative parts by auxiliary networks and extract the part-based finegrained features for classification. In this paper, we propose a simple yet effective approach by introducing an intersection and union module (IU-Module). The IU-Module aims to capture more discriminative features by 1) dividing features into distinct groups, 2) sharing parts of interests within each group, and 3) adding a differentiation loss to reduce the similarity among those grouped feature channels. Without adding any new learnable parameters, the proposed approach imposes two straightforward operations, namely channel intersection (CI) and channel union (CU) operations, on the convolutional features and achieves competitive results compared with the state-of-the-art methods. Experimental results on three publicly available FGVC datasets show the effectiveness of the IU-Module. Ablation studies and visualizations are also provided to make further demonstrations. Yixiao Zheng, Dongliang Chang, Jiyang Xie 0001, Zhanyu Ma |
ICME | 3 |
| 2020 | Deep Neural Network-Based Impacts Analysis of Multimodal Factors on Heat Demand PredictionabstractPrediction of heat demand using artificial neural networks has attracted enormous research attention. Weather conditions, such as direct solar irradiance and wind speed, have been identified as key parameters affecting heat demand. This paper employs an Elman neural network to investigate the impacts of direct solar irradiance and wind speed on the heat demand from the perspective of the entire district heating network. Results of the overall mean absolute percentage error (MAPE) show that direct solar irradiance and wind speed have quite similar impacts. However, the involvement of direct solar irradiance can clearly reduce the maximum absolute deviation when only involving direct solar irradiance and wind speed, respectively. In addition, the simultaneous involvement of both wind speed and direct solar irradiance does not show an obvious improvement of MAPE. Moreover, the prediction accuracy can also be affected by other factors like data discontinuity and outliers. Zhanyu Ma, Jiyang Xie 0001, Qie Sun, Fredrik Wallin, Zhongwei Si, Jun Guo 0002 |
IEEE Trans. Big Data | 2 |
| 2020 | The Devil is in the Channels: Mutual-Channel Loss for Fine-Grained Image ClassificationabstractThe key to solving fine-grained image categorization is finding discriminate and local regions that correspond to subtle visual traits. Great strides have been made, with complex networks designed specifically to learn part-level discriminate feature representations. In this paper, we show that it is possible to cultivate subtle details without the need for overly complicated network designs or training mechanisms - a single loss is all it takes. The main trick lies with how we delve into individual feature channels early on, as opposed to the convention of starting from a consolidated feature map. The proposed loss function, termed as mutual-channel loss (MC-Loss), consists of two channel-specific components: a discriminality component and a diversity component. The discriminality component forces all feature channels belonging to the same class to be discriminative, through a novel channel-wise attention mechanism. The diversity component additionally constraints channels so that they become mutually exclusive across the spatial dimension. The end result is therefore a set of feature channels, each of which reflects different locally discriminative regions for a specific class. The MC-Loss can be trained end-to-end, without the need for any bounding-box/part annotations, and yields highly discriminative regions during inference. Experimental results show our MC-Loss when implemented on top of common base networks can achieve state-of-the-art performance on all four fine-grained categorization datasets (CUB-Birds, FGVC-Aircraft, Flowers-102, and Stanford Cars). Ablative studies further demonstrate the superiority of the MC-Loss when compared with other recently proposed general-purpose losses for visual classification, on two different base networks. Dongliang Chang, Jiyang Xie 0001, Ayan Kumar Bhunia, Zhanyu Ma, Ming Wu 0001, Jun Guo 0002, Yi-Zhe Song |
IEEE Trans. Image Process. | 3 |
| 2020 | Insights Into Multiple/Single Lower Bound Approximation for Extended Variational Inference in Non-Gaussian Structured Data ModelingabstractFor most of the non-Gaussian statistical models, the data being modeled represent strongly structured properties, such as scalar data with bounded support (e.g., beta distribution), vector data with unit length (e.g., Dirichlet distribution), and vector data with positive elements (e.g., generalized inverted Dirichlet distribution). In practical implementations of non-Gaussian statistical models, it is infeasible to find an analytically tractable solution to estimating the posterior distributions of the parameters. Variational inference (VI) is a widely used framework in Bayesian estimation. Recently, an improved framework, namely, the extended VI (EVI), has been introduced and applied successfully to a number of non-Gaussian statistical models. EVI derives analytically tractable solutions by introducing lower bound approximations to the variational objective function. In this paper, we compare two approximation strategies, namely, the multiple lower bounds (MLBs) approximation and the single lower bound (SLB) approximation, which can be applied to carry out the EVI. For implementation, two different conditions, the weak and the strong conditions, are discussed. Convergence of the EVI depends on the selection of the lower bound, regardless of the choice of weak or strong condition. We also discuss the convergence properties to clarify the differences between MLB and SLB. Extensive comparisons are made based on some EVI-based non-Gaussian statistical models. Theoretical analysis is conducted to demonstrate the differences between the weak and strong conditions. Experimental results based on real data show advantages of the SLB approximation over the MLB approximation. Zhanyu Ma, Jiyang Xie 0001, Yuping Lai, Jalil Taghia, Jing-Hao Xue, Jun Guo 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | FICAL: Focal Inter-Class Angular Loss for Image ClassificationabstractConvolutional Neural Networks (CNNs) have been successfully applied in various image analysis tasks and gradually become one of the most powerful machine learning approaches. In order to improve the capability of the model generalization and performance in image classification, a new trend is to learn more discriminative features via CNNs. The main contribution of this paper is to increase the angles between the categories to extract discriminative features and enlarge the inter-class variance. To this end, we propose a loss function named focal inter-class angular loss (FICAL) which introduces the confusion rate-weighted cosine distance as the similarity measurement between categories. This measurement is dynamically evaluated during each iteration to adapt the model. Compared with other loss functions, experimental results demonstrate that the proposed FICAL achieved best performance among the referred loss functions on two image classificaton datasets. Xinran Wei, Dongliang Chang, Jiyang Xie 0001, Yixiao Zheng, Chen Gong 0002, Zhanyu Ma |
VCIP | 3 |
| 2018 | A Survey on Machine Learning-Based Mobile Big Data Analysis: Challenges and ApplicationsabstractThis paper attempts to identify the requirement and the development of machine learning‐based mobile big data (MBD) analysis through discussing the insights of challenges in the mobile big data. Furthermore, it reviews the state‐of‐the‐art applications of data analysis in the area of MBD. Firstly, we introduce the development of MBD. Secondly, the frequently applied data analysis methods are reviewed. Three typical applications of MBD analysis, namely, wireless channel modeling, human online and offline behavior analysis, and speech recognition in the Internet of Vehicles, are introduced, respectively. Finally, we summarize the main challenges and future development directions of mobile big data analysis. Jiyang Xie 0001, Yanting Zhang 0001, Hong Yu 0006, Jinnan Zhan, Zhanyu Ma, Yuanyuan Qiao 0002, Jianhua Zhang 0001, Jun Guo 0002 |
Wirel. Commun. Mob. Comput. | 1 |