EDBT 2026 Demo / reviewers in the wild / expert
Masoumeh Zareapoor
dblp:154/8597 · also Mesoume Zareapoor
· DBLP profile ↗
41ranked-venue papers
10as first author
23since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 5 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 7 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Computer networks · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | BiMAC: Bidirectional Multimodal Alignment in Contrastive LearningabstractAchieving robust performance in vision-language tasks requires strong multimodal alignment, where textual and visual data interact seamlessly. Existing frameworks often combine contrastive learning with image captioning to unify visual and textual representations. However, reliance on global representations and unidirectional information flow from images to text limits their ability to reconstruct visual content accurately from textual descriptions. To address this limitation, we propose BiMAC, a novel framework that enables bidirectional interactions between images and text at both global and local levels. BiMAC employs advanced components to simultaneously reconstruct visual content from textual cues and generate textual descriptions guided by visual features. By integrating a text-region alignment mechanism, BiMAC identifies and selects relevant image patches for precise cross-modal interaction, reducing information noise and enhancing mapping accuracy. BiMAC achieves state-of-the-art performance across diverse vision-language tasks, including image-text retrieval, captioning, and classification. Masoumeh Zareapoor, Pourya Shamsolmoali, Yue Lu 0001 |
AAAI | 1 |
| 2025 | ShapeMorph: 3D Shape Completion via Blockwise Discrete Diffusion
Pourya Shamsolmoali, Yue Lu 0001, Masoumeh Zareapoor |
WACV | 4 |
| 2025 | From Missing Pieces to Masterpieces: Image Completion With Context-Adaptive DiffusionabstractImage completion is a challenging task, particularly when ensuring that generated content seamlessly integrates with existing parts of an image. While recent diffusion models have shown promise, they often struggle with maintaining coherence between known and unknown (missing) regions. This issue arises from the lack of explicit spatial and semantic alignment during the diffusion process, resulting in content that does not smoothly integrate with the original image. Additionally, diffusion models typically rely on global learned distributions rather than localized features, leading to inconsistencies between the generated and existing image parts. In this work, we propose ConFill, a novel framework that introduces a Context-Adaptive Discrepancy (CAD) model to ensure that intermediate distributions of known and unknown regions are closely aligned throughout the diffusion process. By incorporating CAD, our model progressively reduces discrepancies between generated and original images at each diffusion step, leading to contextually aligned completion. Moreover, ConFill uses a new Dynamic Sampling mechanism that adaptively increases the sampling rate in regions with high reconstruction complexity. This approach enables precise adjustments, enhancing detail and integration in restored areas. Extensive experiments demonstrate that ConFill outperforms current methods, setting a new benchmark in image completion. Pourya Shamsolmoali, Masoumeh Zareapoor, Huiyu Zhou 0001, Michael Felsberg, Dacheng Tao, Xuelong Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Hybrid Gromov-Wasserstein Embedding for Capsule LearningabstractCapsule networks (CapsNets) aim to parse images into a hierarchy of objects, parts, and their relationships using a two-step process involving part-whole transformation and hierarchical component routing. However, this hierarchical relationship modeling is computationally expensive, which has limited the wider use of CapsNet despite its potential advantages. The current state of CapsNet models primarily focuses on comparing their performance with capsule baselines, falling short of achieving the same level of proficiency as deep convolutional neural network (CNN) variants in intricate tasks. To address this limitation, we present an efficient approach for learning capsules that surpasses canonical baseline models and even demonstrates superior performance compared with high-performing convolution models. Our contribution can be outlined in two aspects: first, we introduce a group of subcapsules onto which an input vector is projected. Subsequently, we present the hybrid Gromov-Wasserstein (HGW) framework, which initially quantifies the dissimilarity between the input and the components modeled by the subcapsules, followed by determining their alignment degree through optimal transport (OT). This innovative mechanism capitalizes on new insights into defining alignment between the input and subcapsules, based on the similarity of their respective component distributions. This approach enhances CapsNets' capacity to learn from intricate, high-dimensional data while retaining their interpretability and hierarchical structure. Our proposed model offers two distinct advantages: 1) its lightweight nature facilitates the application of capsules to more intricate vision tasks, including object detection; and 2) it outperforms baseline approaches in these demanding tasks. Our empirical findings illustrate that HGW capsules (HGWCapsules) exhibit enhanced robustness against affine transformations, scale effectively to larger datasets, and surpass CNN and CapsNet models across various vision tasks. Pourya Shamsolmoali, Masoumeh Zareapoor, Swagatam Das, Eric Granger, Salvador García 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | SeTformer Is What You Need for Vision and LanguageabstractThe dot product self-attention (DPSA) is a fundamental component of transformers. However, scaling them to long sequences, like documents or high-resolution images, becomes prohibitively expensive due to the quadratic time and memory complexities arising from the softmax operation. Kernel methods are employed to simplify computations by approximating softmax but often lead to performance drops compared to softmax attention. We propose SeTformer, a novel transformer where DPSA is purely replaced by Self-optimal Transport (SeT) for achieving better performance and computational efficiency. SeT is based on two essential softmax properties: maintaining a non-negative attention matrix and using a nonlinear reweighting mechanism to emphasize important tokens in input sequences. By introducing a kernel cost function for optimal transport, SeTformer effectively satisfies these properties. In particular, with small and base-sized models, SeTformer achieves impressive top-1 accuracies of 84.7% and 86.2% on ImageNet-1K. In object detection, SeTformer-base outperforms the FocalNet counterpart by +2.2 mAP, using 38% fewer parameters and 29% fewer FLOPs. In semantic segmentation, our base-size model surpasses NAT by +3.5 mIoU with 33% fewer parameters. SeTformer also achieves state-of-the-art results in language modeling on the GLUE benchmark. These findings highlight SeTformer applicability for vision and language tasks. Pourya Shamsolmoali, Masoumeh Zareapoor, Eric Granger, Michael Felsberg |
AAAI | 2 |
| 2024 | Rethinking Fast Adversarial Training: A Splitting Technique to Overcome Catastrophic Overfitting
Masoumeh Zareapoor, Pourya Shamsolmoali |
ECCV (78) | 1 |
| 2024 | Efficient Routing in Sparse Mixture-of-ExpertsabstractSparse Mixture-of-Experts (MoE) architectures provide the distinct benefit of substantially expanding the model’s parameter space without proportionally increasing the computational load on individual input tokens or samples. However, the efficacy of these models heavily depends on the routing strategy used to assign tokens to experts. Poor routing can lead to under-trained or overly specialized experts, diminishing the overall model performance. Previous approaches have relied on the Topk router, where each token is assigned to a subset of experts. In this paper, we propose a routing mechanism that replaces the Topk router with regularized optimal transport, leveraging the Sinkhorn algorithm to optimize token-expert matching. We conducted a comprehensive evaluation comparing the pre-training efficiency of our model, using computational resources equivalent to those employed in the GShard and Switch Transformers gating mechanisms. The results demonstrate that our model expedites training convergence, achieving a speedup of over 2× compared to these baseline models. Moreover, under the same computational constraints, our model exhibits superior performance across eleven tasks from the GLUE and SuperGLUE benchmarks. We show that our model contributes to the optimization of token-expert matching in sparsely-activated MoE models, offering substantial gains in both training efficiency and task performance. Masoumeh Zareapoor, Pourya Shamsolmoali, Fateme Vesaghati |
IJCNN | 1 |
| 2024 | Fractional Correspondence Framework in Detection TransformerabstractThe Detection Transformer (DETR), by incorporating the Hungarian algorithm, has significantly simplified the matching process in object detection tasks. This algorithm facilitates optimal one-to-one matching of predicted bounding boxes to ground-truth annotations during training. While effective, this strict matching process does not inherently account for the varying densities and distributions of objects, leading to suboptimal correspondences such as failing to handle multiple detections of the same object or missing small objects. To address this, we propose the Regularized Transport Plan (RTP). RTP introduces a flexible matching strategy that captures the cost of aligning predictions with ground truths to find the most accurate correspondences between these sets. By utilizing the differentiable Sinkhorn algorithm, RTP allows for soft, fractional matching rather than strict one-to-one assignments. This approach enhances the model's capability to manage varying object densities and distributions effectively. Our extensive evaluations on the MS-COCO and VOC benchmarks demonstrate the effectiveness of our approach. RTP-DETR, surpassing the performance of the Deform-DETR and the recently introduced DINO-DETR, achieving absolute gains in mAP of +3.8% and +1.7%, respectively. Masoumeh Zareapoor, Pourya Shamsolmoali, Huiyu Zhou 0001, Yue Lu 0001, Salvador García 0001 |
ACM Multimedia | 1 |
| 2024 | Distance-based Weighted Transformer Network for image completion
Pourya Shamsolmoali, Masoumeh Zareapoor, Huiyu Zhou 0001, Xuelong Li 0001, Yue Lu 0001 |
Pattern Recognit. | 2 |
| 2023 | Image Completion Via Dual-Path Cooperative FilteringabstractGiven the recent advances with image-generating algorithms, deep image completion methods have made significant progress. However, state-of-art methods typically provide poor cross-scene generalization, and generated masked areas often contain blurry artifacts. Predictive filtering is a method for restoring images, which predicts the most effective kernels based on the input scene. Motivated by this approach, we address image completion as a filtering problem. Deep feature-level semantic filtering is introduced to fill in missing information, while preserving local structure and generating visually realistic content. In particular, a Dual-path Cooperative Filtering (DCF) model is proposed, where one path predicts dynamic kernels, and the other path extracts multi-level features by using Fast Fourier Convolution to yield semantically coherent reconstructions. Experiments on three challenging image completion datasets show that our proposed DCF outperforms state-of-art methods. Pourya Shamsolmoali, Masoumeh Zareapoor, Eric Granger |
ICASSP | 2 |
| 2023 | Entropy Transformer Networks: A Learning Approach via Tangent Bundle Data ManifoldabstractThis paper focuses on an accurate and fast interpolation approach for image transformation employed in the design of CNN architectures. Standard Spatial Transformer Networks (STNs) use bilinear or linear interpolation as their interpolation, with unrealistic assumptions about the underlying data distributions, which leads to poor performance under scale variations. Moreover, STNs do not preserve the norm of gradients in propagation due to their dependency on sparse neighboring pixels. To address this problem, a novel Entropy STN (ESTN) is proposed that interpolates on the data manifold distributions. In particular, random samples are generated for each pixel in association with the tangent space of the data manifold, and construct a linear approximation of their intensity values with an entropy regularizer to compute the transformer parameters. A simple yet effective technique is also proposed to normalize the non-zero values of the convolution operation, to fine-tune the layers for gradients' norm-regularization during training. Experiments on challenging benchmarks show that the proposed ESTN can improve predictive accuracy over a range of computer vision tasks, including image reconstruction, and classification, while reducing the computational cost. Pourya Shamsolmoali, Masoumeh Zareapoor |
IJCNN | 2 |
| 2023 | GEN: Generative Equivariant Networks for Diverse Image-to-Image TranslationabstractImage-to-image (I2I) translation has become a key asset for generative adversarial networks. Convolutional neural networks (CNNs), despite having a significant performance, are not able to capture the spatial relationships among different parts of an object and, thus, do not qualify as the ideal representative model for image translation tasks. As a remedy to this problem, capsule networks have been proposed to represent patterns for a visual object in such a way that preserves hierarchical spatial relationships. The training of capsules is constrained by learning all pairwise relationships between capsules of consecutive layers. This design would be prohibitively expensive both in time and memory. In this article, we present a new framework for capsule networks to provide a full description of the input components at various levels of semantics, which can successfully be applied to the generator-discriminator architectures without incurring computational overhead compared to the CNNs. To successfully apply the proposed capsules in the generative adversarial network, we put forth a novel Gromov-Wasserstein (GW) distance as a differentiable loss function that compares the dissimilarity between two distributions and then guides the learned distribution toward target properties, using optimal transport (OT) discrepancy. The proposed method-which is called generative equivariant network (GEN)-is an alternative architecture for GANs with equivariance capsule layers. The proposed model is evaluated through a comprehensive set of experiments on I2I translation and image generation tasks and compared with several state-of-the-art models. Results indicate that there is a principled connection between generative and capsule models that allows extracting discriminant and invariant information from image data. Pourya Shamsolmoali, Masoumeh Zareapoor, Swagatam Das, Salvador García 0001, Eric Granger, Jie Yang 0002 |
IEEE Trans. Cybern. | 2 |
| 2023 | VTAE: Variational Transformer Autoencoder With Manifolds LearningabstractDeep generative models have demonstrated successful applications in learning non-linear data distributions through a number of latent variables and these models use a non-linear function (generator) to map latent samples into the data space. On the other hand, the non-linearity of the generator implies that the latent space shows an unsatisfactory projection of the data space, which results in poor representation learning. This weak projection, however, can be addressed by a Riemannian metric, and we show that geodesics computation and accurate interpolations between data samples on the Riemannian manifold can substantially improve the performance of deep generative models. In this paper, a Variational spatial-Transformer AutoEncoder (VTAE) is proposed to minimize geodesics on a Riemannian manifold and improve representation learning. In particular, we carefully design the variational autoencoder with an encoded spatial-Transformer to explicitly expand the latent variable model to data on a Riemannian manifold, and obtain global context modelling. Moreover, to have smooth and plausible interpolations while traversing between two different objects' latent representations, we propose a geodesic interpolation network different from the existing models that use linear interpolation with inferior performance. Experiments on benchmarks show that our proposed model can improve predictive accuracy and versatility over a range of computer vision tasks, including image interpolations, and reconstructions. Pourya Shamsolmoali, Masoumeh Zareapoor, Huiyu Zhou 0001, Dacheng Tao, Xuelong Li 0001 |
IEEE Trans. Image Process. | 2 |
| 2022 | Salient Skin Lesion Segmentation via Dilated Scale-Wise Feature Fusion NetworkabstractSkin lesion detection in dermoscopic images is essential in the accurate and early diagnosis of skin cancer by a computerized apparatus. Current skin lesion segmentation approaches show poor performance in challenging circumstances such as indistinct lesion boundaries, low contrast between the lesion and the surrounding area, or heterogeneous background that causes over/under segmentation of the skin lesion. To accurately recognize the lesion from the neighboring regions, we propose a dilated scale-wise feature fusion network based on convolution factorization. Our network is designed to simultaneously extract features at different scales which are systematically fused for better detection. The proposed model has satisfactory accuracy and efficiency. Various experiments for lesion segmentation are performed along with comparisons with the state-of-the-art models. Our proposed model consistently showcases state-of-the-art results. Pourya Shamsolmoali, Masoumeh Zareapoor, Jie Yang 0002, Eric Granger, Huiyu Zhou 0001 |
ICPR | 2 |
| 2022 | Enhanced Single-Shot Detector for Small Object Detection in Remote Sensing ImagesabstractSmall-object detection is a challenging problem. In the last few years, the convolution neural networks methods have been achieved considerable progress. However, the current detectors struggle with effective features extraction for small-scale objects. To address this challenge, we propose image pyramid single-shot detector (IPSSD). In IPSSD, single-shot detector is adopted combined with an image pyramid network to extract semantically strong features for generating candidate regions. The proposed network can enhance the small-scale features from a feature pyramid network. We evaluated the performance of the proposed model on two public datasets and the results show the superior performance of our model compared to the other state-of-the-art object detectors. Pourya Shamsolmoali, Masoumeh Zareapoor, Jie Yang 0002, Eric Granger, Jocelyn Chanussot |
IGARSS | 2 |
| 2022 | Multipatch Feature Pyramid Network for Weakly Supervised Object Detection in Optical Remote Sensing ImagesabstractObject detection is a challenging task in remote sensing because objects only occupy a few pixels in the images, and the models are required to simultaneously learn object locations and detection. Even though the established approaches well perform for the objects of regular sizes, they achieve weak performance when analyzing small ones or getting stuck in the local minima (e.g. false object parts). Two possible issues stand in their way. First, the existing methods struggle to perform stably on the detection of small objects because of the complicated background. Second, most of the standard methods used hand-crafted features, and do not work well on the detection of objects parts of which are missing. We here address the above issues and propose a new architecture with a multiple patch feature pyramid network (MPFP-Net). Different from the current models that during training only pursue the most discriminative patches, in MPFPNet the patches are divided into class-affiliated subsets, in which the patches are related and based on the primary loss function, a sequence of smooth loss functions are determined for the subsets to improve the model for collecting small object parts. To enhance the feature representation for patch selection, we introduce an effective method to regularize the residual values and make the fusion transition layers strictly norm-preserving. The network contains bottom-up and crosswise connections to fuse the features of different scales to achieve better accuracy, compared to several state-of-the-art object detection models. Also, the developed architecture is more efficient than the baselines. Pourya Shamsolmoali, Jocelyn Chanussot, Masoumeh Zareapoor, Huiyu Zhou 0001, Jie Yang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Rotation Equivariant Feature Image Pyramid Network for Object Detection in Optical Remote Sensing ImageryabstractDetection of objects is extremely important in various aerial vision-based applications. Over the last few years, the methods based on convolution neural networks (CNNs) have made substantial progress. However, because of the large variety of object scales, densities, and arbitrary orientations, the current detectors struggle with the extraction of semantically strong features for small-scale objects by a predefined convolution kernel. To address this problem, we propose the rotation equivariant feature image pyramid network (REFIPN), an image pyramid network based on rotation equivariance convolution. The proposed model adopts single-shot detector in parallel with a lightweight image pyramid module (LIPM) to extract representative features and generate regions of interest in an optimization approach. The proposed network extracts feature in a wide range of scales and orientations by using novel convolution filters. These features are used to generate vector fields and determine the weight and angle of the highest-scoring orientation for all spatial locations on an image. By this approach, the performance for small-sized object detection is enhanced without sacrificing the performance for large-sized object detection. The performance of the proposed model is validated on two commonly used aerial benchmarks and the results show our proposed model can achieve state-of-the-art performance with satisfactory efficiency. Pourya Shamsolmoali, Masoumeh Zareapoor, Jocelyn Chanussot, Huiyu Zhou 0001, Jie Yang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Asymmetric Correlation Quantization Hashing for Cross-Modal RetrievalabstractIn recent years, cross-modal hashing (CMH) has attracted considerable attention due to its ability to learn across different modalities and its high efficiency for similarity retrieval applications. This procedure is computationally inexpensive when dealing with large-scale multi-modalities datasets. However, they do not form the ideal representative model to fully exploit multi-modal data’s underlying properties despite their successful performance. We identify that: (i) most CMH models in their current forms transform the real data points into discrete compact binary codes, which can limit their ability to prevent the loss of important information and thereby produce suboptimal results. (ii) the discrete-binary constraint model is hard to implement, and relaxing the binary constraints is a common property in most existing methods, which often leads to significant quantization errors. (iii) handling the CMH in a symmetry domain leads to a complex and inefficient optimization problem. This paper addresses the above challenges and proposes a novel Asymmetric Correlation Quantization Hashing (ACQH) method. ACQH learns a projection matrix for each heterogeneous modality to map the data point into a low-dimensional semantic space and constructs a compositional quantization to generate hash codes, using the pairwise semantic similarity preservation and the pointwise label regression. As a specific instantiation of our model, we use discrete iterative optimization to obtain the unified hash codes across different modalities. Extensive experiments show that ACQH outperforms state-of-the-art methods on several diverse datasets. Lu Wang 0048, Masoumeh Zareapoor, Jie Yang 0002, Zhonglong Zheng |
IEEE Trans. Multim. | 2 |
| 2021 | Imbalanced data learning by minority class augmentation using capsule adversarial networks
Pourya Shamsolmoali, Masoumeh Zareapoor, LinLin Shen, Abdul Hamid Sadka, Jie Yang 0002 |
Neurocomputing | 2 |
| 2021 | Multimodal image fusion based on point-wise mutual information
Donghao Shen, Masoumeh Zareapoor, Jie Yang 0002 |
Image Vis. Comput. | 2 |
| 2021 | Cluster-wise unsupervised hashing for cross-modal similarity search
Lu Wang 0048, Jie Yang 0002, Masoumeh Zareapoor, Zhonglong Zheng |
Pattern Recognit. | 3 |
| 2021 | Road Segmentation for Remote Sensing Images Using Adversarial Spatial Pyramid NetworksabstractRoad extraction in remote sensing images is of great importance for a wide range of applications. Because of the complex background, and high density, most of the existing methods fail to accurately extract a road network that appears correct and complete. Moreover, they suffer from either insufficient training data or high costs of manual annotation. To address these problems, we introduce a new model to apply structured domain adaption for synthetic image generation and road segmentation. We incorporate a feature pyramid (FP) network into generative adversarial networks to minimize the difference between the source and target domains. A generator is learned to produce quality synthetic images, and the discriminator attempts to distinguish them. We also propose a FP network that improves the performance of the proposed model by extracting effective features from all the layers of the network for describing different scales' objects. Indeed, a novel scale-wise architecture is introduced to learn from the multilevel feature maps and improve the semantics of the features. For optimization, the model is trained by a joint reconstruction loss function, which minimizes the difference between the fake images and the real ones. A wide range of experiments on three data sets prove the superior performance of the proposed approach in terms of accuracy and efficiency. In particular, our model achieves state-of-the-art 78.86 IOU on the Massachusetts data set with 14.89M parameters and 86.78B FLOPs, with 4× fewer FLOPs but higher accuracy (+3.47% IOU) than the top performer among state-of-the-art approaches used in the evaluation. Pourya Shamsolmoali, Masoumeh Zareapoor, Huiyu Zhou 0001, Ruili Wang 0001, Jie Yang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Equivariant Adversarial Network for Image-to-image TranslationabstractImage-to-Image translation aims to learn an image from a source domain to a target domain. However, there are three main challenges, such as lack of paired datasets, multimodality, and diversity, that are associated with these problems and need to be dealt with. Convolutional neural networks (CNNs), despite of having great performance in many computer vision tasks, they fail to detect the hierarchy of spatial relationships between different parts of an object and thus do not form the ideal representative model we look for. This article presents a new variation of generative models that aims to remedy this problem. We use a trainable transformer, which explicitly allows the spatial manipulation of data within training. This differentiable module can be augmented into the convolutional layers in the generative model, and it allows to freely alter the generated distributions for image-to-image translation. To reap the benefits of proposed module into generative model, our architecture incorporates a new loss function to facilitate an effective end-to-end generative learning for image-to-image translation. The proposed model is evaluated through comprehensive experiments on image synthesizing and image-to-image translation, along with comparisons with several state-of-the-art algorithms. Masoumeh Zareapoor, Jie Yang 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2020 | Infrared and visible image fusion via global variable consensus
Donghao Shen, Masoumeh Zareapoor, Jie Yang 0002 |
Image Vis. Comput. | 2 |
| 2020 | GAN-Poser: an improvised bidirectional GAN model for human motion prediction
Deepak Kumar Jain 0001, Masoumeh Zareapoor, Rachna Jain, Abhishek Kathuria, Shivam Bachhety |
Neural Comput. Appl. | 2 |
| 2020 | Perceptual image quality using dual generative adversarial network
Masoumeh Zareapoor, Huiyu Zhou 0001, Jie Yang 0002 |
Neural Comput. Appl. | 1 |
| 2020 | MLDNet: Multi-level dense network for multi-focus image fusion
Hafiz Tayyab Mustafa, Masoumeh Zareapoor, Jie Yang 0002 |
Signal Process. Image Commun. | 2 |
| 2020 | Deep-Learning-Based Small Surface Defect Detection via an Exaggerated Local Variation-Based Generative Adversarial NetworkabstractSurface detection of small defects plays a vital role in manufacturing and has attracted broad interest. It remains challenging primarily due to the small size of the defect relative to the large surface and the rare occurrence of defects. To address this problem, in this article we propose a novel machine vision approach for automatically identifying the tiny flaws that may appear in a single image. First, the presented defect exaggeration approach produces both the flawless image and the corresponding exaggerated version of the defect by taking the variations in the image as regularization terms. Second, a generative adversarial network (GAN) in conjunction with a convolutional neural network (CNN) is proposed to guarantee the accuracy of tiny surface defect detection by producing exaggerated defect image samples. Furthermore, the limited dataset of the training samples for defect detection is enlarged by exploiting the GAN technique with the variation exaggerated images. To evaluate the performance of our proposed method, we conduct comparison experiments between the state-of-the-art techniques with and without the proposed algorithm as well as comparison experiments between the state-of-the-art techniques and our method. The experimental results on different types of surface image samples demonstrate that the proposed method can significantly improve the performance of the state-of-the-art approaches while achieving a defect detection accuracy of 99.2%. Jian Lian, Weikuan Jia, Masoumeh Zareapoor, Yuanjie Zheng, Deepak Kumar Jain 0001, Neeraj Kumar 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2020 | AMIL: Adversarial Multi-instance Learning for Human Pose EstimationabstractHuman pose estimation has an important impact on a wide range of applications, from human-computer interface to surveillance and content-based video retrieval. For human pose estimation, joint obstructions and overlapping upon human bodies result in departed pose estimation. To address these problems, by integrating priors of the structure of human bodies, we present a novel structure-aware network to discreetly consider such priors during the training of the network. Typically, learning such constraints is a challenging task. Instead, we propose generative adversarial networks as our learning model in which we design two residual Multiple-Instance Learning (MIL) models with identical architecture—one is used as the generator, and the other one is used as the discriminator. The discriminator task is to distinguish the actual poses from the fake ones. If the pose generator generates results that the discriminator is not able to distinguish from the real ones, then the model has successfully learned the priors. In the proposed model, the discriminator differentiates the ground-truth heatmaps from the generated ones, and later the adversarial loss back-propagates to the generator. Such procedure assists the generator to learn reasonable body configurations and is proved to be advantageous to improve the pose estimation accuracy. Meanwhile, we propose a novel function for MIL. It is an adjustable structure for both instance selection and modeling to appropriately pass the information between instances in a single bag. In the proposed residual MIL neural network, the pooling action adequately updates the instance contribution to its bag. The proposed adversarial residual multi-instance neural network that is based on pooling has been validated on two datasets for the human pose estimation task and successfully outperforms the other state-of-the-art models. The code will be made available on https://github.com/pshams55/AMIL. Pourya Shamsolmoali, Masoumeh Zareapoor, Huiyu Zhou 0001, Jie Yang 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2019 | Convolutional neural network in network (CNNiN): hyperspectral image classification and dimensionality reductionabstractClassification is a principle technique in hyperspectral images (HSIs), where a label is assigned to each pixel based on its characteristics. However, due to lack of labelled training instances in HSIs and also its ultra‐high dimensionality, deep learning approaches need a special consideration for HSI classification. As one of the first works in the HSI classification, this study proposes a novel network pipeline called convolutional neural network in network (which is deeper than the existing approaches) by jointly utilising the spatial and spectral information and produces high‐level features from the original HSI. This can occur by using spatial–spectral relationships of individual pixel vector at the initial component of the proposed pipeline; the extracted features are then combined to form a joint spatial–spectral feature map. Finally, a recurrent neural network is trained on the extracted features which contain wealthy spectral and spatial properties of the HSI to predict the corresponding label of each vector. The model has been tested on two large scale hyperspectral datasets in terms of classification accuracy, training error, and computational time. Pourya Shamsolmoali, Masoumeh Zareapoor, Jie Yang 0002 |
IET Image Process. | 2 |
| 2019 | G-GANISR: Gradual generative adversarial network for image super resolution
Pourya Shamsolmoali, Masoumeh Zareapoor, Ruili Wang 0001, Deepak Kumar Jain 0001, Jie Yang 0002 |
Neurocomputing | 2 |
| 2019 | Multi-scale convolutional neural network for multi-focus image fusion
Hafiz Tayyab Mustafa, Jie Yang 0002, Masoumeh Zareapoor |
Image Vis. Comput. | 3 |
| 2019 | Image super resolution by dilated dense progressive network
Pourya Shamsolmoali, Masoumeh Zareapoor, Jie Yang 0002 |
Image Vis. Comput. | 2 |
| 2019 | High-dimensional multimedia classification using deep CNN and extended residual units
Pourya Shamsolmoali, Deepak Kumar Jain 0001, Masoumeh Zareapoor, Jie Yang 0002, Mohd. Afshar Alam |
Multim. Tools Appl. | 3 |
| 2019 | Deep convolution network for surveillance records super-resolution
Pourya Shamsolmoali, Masoumeh Zareapoor, Deepak Kumar Jain 0001, Vinay Kumar Jain, Jie Yang 0002 |
Multim. Tools Appl. | 2 |
| 2019 | Deep semantic preserving hashing for large scale image retrieval
Masoumeh Zareapoor, Jie Yang 0002, Deepak Kumar Jain 0001, Pourya Shamsolmoali, Neha Jain 0003, Surya Kant |
Multim. Tools Appl. | 1 |
| 2019 | Towards realistic image via function learning
Masoumeh Zareapoor, Jie Yang 0002 |
Multim. Tools Appl. | 1 |
| 2019 | Diverse adversarial network for image super-resolution
Masoumeh Zareapoor, M. Emre Celebi 0001, Jie Yang 0002 |
Signal Process. Image Commun. | 1 |
| 2018 | Deep Supervised Auto-encoder Hashing for Image Retrieval
Sanli Tang, Haoyuan Chi, Jie Yang 0002, Xiaolin Huang, Masoumeh Zareapoor |
PRCV (2) | 5 |
| 2018 | Hybrid deep neural networks for face emotion recognition
Neha Jain 0003, Shishir Kumar, Amit Kumar 0023, Pourya Shamsolmoali, Masoumeh Zareapoor |
Pattern Recognit. Lett. | 5 |
| 2018 | Kernelized support vector machine with deep learning: An efficient approach for extreme multiclass dataset
Masoumeh Zareapoor, Pourya Shamsolmoali, Deepak Kumar Jain 0001, Haoxiang Wang 0001, Jie Yang 0002 |
Pattern Recognit. Lett. | 1 |