EDBT 2026 Demo / reviewers in the wild / expert
Pourya Shamsolmoali
dblp:154/8477
· DBLP profile ↗
54ranked-venue papers
26as first author
39since 2021 · last 2026
0000-0002-0263-1661ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 12 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 11 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 first-author · 5 since 2021Computer networks · 2 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Patches: Mining Interpretable Part-Prototypes for Explainable AIabstractAs AI systems become more capable, it is important that their decisions are understandable and aligned with human expectations. A key challenge is the lack of interpretability in deep models. Existing methods such as GradCAM generate heatmaps but provide limited conceptual insight, while prototype-based approaches offer example-based explanations but often rely on rigid region selection and lack semantic consistency. To address these limitations, we propose PCMNet, a Part-Prototypical Concept Mining Network that learns human-comprehensible prototypes from meaningful regions without extra supervision. By clustering these into concept groups and extracting concept activation vectors, PCMNet provides structured, concept-level explanations and enhances robustness under occlusion and adversarial conditions, which are both critical for building reliable and aligned AI systems. Experiments across multiple benchmarks show that PCMNet outperforms state-of-the-art methods in interpretability, stability, and robustness. This work contributes to AI alignment by enhancing transparency, controllability, and trustworthiness in modern AI systems. Mahdi Alehdaghi, Rajarshi Bhattacharya, Pourya Shamsolmoali, Rafael M. O. Cruz, Eric Granger |
AAAI | 3 |
| 2026 | LithoMamba: High-fidelity lithography simulation with State Space ModelsabstractLithography simulation is a critical technology in modern semiconductor manufacturing, yet existing deep learning models often fail to accurately model the complex, long-range optical physics due to the inherent locality of convolution. This limitation results in insufficient simulation fidelity and poses significant challenges for optimization tasks. To overcome this challenge, we introduce LithoMamba, the first generative framework to leverage Mamba for high-fidelity lithography simulation. Our architecture uses a Mamba Generator to model global and long-range optical interactions, while a local, MLP-free Discriminator provides precise, spatial feedback to ensure fine-grained pattern fidelity. This global-local design enables our model to achieve both physical realism and exceptional detail. Our experiments show that LithoMamba outperforms existing methods, both in quantitative and qualitative results. These findings demonstrate the promise of State Space Models for improving lithography simulation and suggest new possibilities for combining physics with generative AI in chip manufacturing. Daohui Wang, Shujing Lyu, Pourya Shamsolmoali, Jiwei Shen, Yue Lu 0001 |
DATE | 4 |
| 2026 | MixER: From Cross-Modal to Mixed-Modal Visible-Infrared Re-Identification
Mahdi Alehdaghi, Rajarshi Bhattacharya, Pourya Shamsolmoali, Rafael M. O. Cruz, Eric Granger |
WACV | 3 |
| 2026 | VAMF: Variance-guided attention modulation framework for infrared and visible image fusion
Hafiz Tayyab Mustafa, Mujtaba Asad, Zhonglong Zheng, Pourya Shamsolmoali, Jie Yang 0002 |
Expert Syst. Appl. | 5 |
| 2025 | BiMAC: Bidirectional Multimodal Alignment in Contrastive LearningabstractAchieving robust performance in vision-language tasks requires strong multimodal alignment, where textual and visual data interact seamlessly. Existing frameworks often combine contrastive learning with image captioning to unify visual and textual representations. However, reliance on global representations and unidirectional information flow from images to text limits their ability to reconstruct visual content accurately from textual descriptions. To address this limitation, we propose BiMAC, a novel framework that enables bidirectional interactions between images and text at both global and local levels. BiMAC employs advanced components to simultaneously reconstruct visual content from textual cues and generate textual descriptions guided by visual features. By integrating a text-region alignment mechanism, BiMAC identifies and selects relevant image patches for precise cross-modal interaction, reducing information noise and enhancing mapping accuracy. BiMAC achieves state-of-the-art performance across diverse vision-language tasks, including image-text retrieval, captioning, and classification. Masoumeh Zareapoor, Pourya Shamsolmoali, Yue Lu 0001 |
AAAI | 2 |
| 2025 | TD-Paint: Faster Diffusion Inpainting Through Time-Aware Pixel ConditioningabstractDiffusion models have emerged as highly effective techniques for inpainting, however, they remain constrained by slow sampling rates. While recent advances have enhanced generation quality, they have also increased sampling time, thereby limiting scalability in real-world applications. We investigate the generative sampling process of diffusion-based inpainting models and observe that these models make minimal use of the input condition during the initial sampling steps. As a result, the sampling trajectory deviates from the data manifold, requiring complex synchronization mechanisms to realign the generation process. To address this, we propose Time-aware Diffusion Paint (TD-Paint), a novel approach that adapts the diffusion process by modeling variable noise levels at the pixel level. This technique allows the model to efficiently use known pixel values from the start, guiding the generation process toward the target manifold. By embedding this information early in the diffusion process, TD-Paint significantly accelerates sampling without compromising image quality. Unlike conventional diffusion-based inpainting models, which require a dedicated architecture or an expensive generation loop, TD-Paint achieves faster sampling times without architectural modifications. Experimental results across three datasets show that TD-Paint outperforms state-of-the-art diffusion models while maintaining lower complexity. Tsiry Mayet, Pourya Shamsolmoali, Simon Bernard 0001, Eric Granger, Romain Hérault, Clément Chatelain 0001 |
ICLR | 2 |
| 2025 | HalluShift: Measuring Distribution Shifts towards Hallucination Detection in LLMsabstractLarge Language Models (LLMs) have recently garnered widespread attention due to their adeptness at generating innovative responses to the given prompts across a multitude of domains. However, LLMs often suffer from the inherent limitation of hallucinations and generate incorrect information while maintaining well-structured and coherent responses. In this work, we hypothesize that hallucinations stem from the internal dynamics of LLMs. Our observations indicate that, during passage generation, LLMs tend to deviate from factual accuracy in subtle parts of responses, eventually shifting toward misinformation. This phenomenon bears a resemblance to human cognition, where individuals may hallucinate while maintaining logical coherence, embedding uncertainty within minor segments of their speech. To investigate this further, we introduce an innovative approach, HALLUSHIFT, designed to analyze the distribution shifts in the internal state space and token probabilities of the LLM-generated responses. Our method attains superior performances compared to existing baselines across various benchmark datasets. Our codebase is available at https://github.com/sharanya-dasgupta001/hallushift. Sharanya Dasgupta, Sujoy Nath, Arkaprabha Basu, Pourya Shamsolmoali, Swagatam Das |
IJCNN | 4 |
| 2025 | Force of Attraction-Based Distribution Calibration for Enhancing Minority Class RepresentationabstractImbalanced image datasets pose significant challenges for developing robust classifiers, particularly when certain classes are heavily underrepresented. To tackle this issue, we propose Density-Driven Attraction (DDA) Oversampling, a novel technique designed to improve class representation in the latent space. Our approach begins by projecting images into disentangled latent representations, ensuring clear separation between classes and precise identification of subclasses. At the core of this method is the Density-Driven Attraction Force (DDAF), a mechanism inspired by gravitational forces. DDAF quantifies the attraction between components of well-represented and underrepresented classes, adjusting the attraction based on the density of each component. This process recalibrates the distributions of underrepresented classes by leveraging their strongest attractions, effectively simulating the natural principles of mass attraction. Extensive classification experiments on six multiclass imbalanced datasets demonstrate that DDA Oversampling outperforms existing state-of-the-art methods, resulting in more accurate and balanced class distributions. Our code is made available at https://github.com/priyomondal/DDAO. Priyobrata Mondal, Faizanuddin Ansari, Swagatam Das, Pourya Shamsolmoali |
IJCNN | 4 |
| 2025 | Bidirectional Multi-Step Domain Generalization for Visible-Infrared Person Re-IdentificationabstractA key challenge in visible-infrared person re-identification (V-I ReID) is training a backbone model capable of effectively addressing the significant discrepancies across modalities. State-of-the-art methods that generate a single intermediate bridging domain are often less effective, as this generated domain may not adequately capture sufficient common discriminant information. This paper introduces Bidirectional Multi-step Domain Generalization (BMDG), a novel approach for unifying feature representations across diverse modalities. BMDG creates multiple virtual intermediate domains by learning and aligning body part features extracted from both I and V modalities. In particular, our method aims to minimize the cross-modal gap in two steps. First, BMDG aligns modalities in the feature space by learning shared and modality-invariant body part prototypes from V and I images. Then, it generalizes the feature representation by applying bidirectional multi-step learning, which progressively refines feature representations in each step and incorporates more prototypes from both modalities. Based on these prototypes, multiple bridging steps enhance the feature representation. Experiments11Our code is available at: alehdaghi.github.io/BMDG conducted on V-I ReID datasets indicate that our BMDG approach can outperform state-of-the-art part-based and intermediate generation methods, and can be integrated into other part-based methods to enhance their V-I ReID performance. Mahdi Alehdaghi, Pourya Shamsolmoali, Rafael M. O. Cruz, Eric Granger |
WACV | 2 |
| 2025 | ShapeMorph: 3D Shape Completion via Blockwise Discrete Diffusion
Pourya Shamsolmoali, Yue Lu 0001, Masoumeh Zareapoor |
WACV | 2 |
| 2025 | From Missing Pieces to Masterpieces: Image Completion With Context-Adaptive DiffusionabstractImage completion is a challenging task, particularly when ensuring that generated content seamlessly integrates with existing parts of an image. While recent diffusion models have shown promise, they often struggle with maintaining coherence between known and unknown (missing) regions. This issue arises from the lack of explicit spatial and semantic alignment during the diffusion process, resulting in content that does not smoothly integrate with the original image. Additionally, diffusion models typically rely on global learned distributions rather than localized features, leading to inconsistencies between the generated and existing image parts. In this work, we propose ConFill, a novel framework that introduces a Context-Adaptive Discrepancy (CAD) model to ensure that intermediate distributions of known and unknown regions are closely aligned throughout the diffusion process. By incorporating CAD, our model progressively reduces discrepancies between generated and original images at each diffusion step, leading to contextually aligned completion. Moreover, ConFill uses a new Dynamic Sampling mechanism that adaptively increases the sampling rate in regions with high reconstruction complexity. This approach enables precise adjustments, enhancing detail and integration in restored areas. Extensive experiments demonstrate that ConFill outperforms current methods, setting a new benchmark in image completion. Pourya Shamsolmoali, Masoumeh Zareapoor, Huiyu Zhou 0001, Michael Felsberg, Dacheng Tao, Xuelong Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Adaptive Generation of Privileged Intermediate Information for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (V-I ReID) seeks to retrieve images of the same individual captured over a distributed network of RGB and IR sensors. Several V-I ReID approaches directly integrate the V and I modalities to represent images within a shared space. However, given the significant gap in the data distributions between V and I modalities, cross-modal V-I ReID remains challenging. A solution is to involve a privileged intermediate space to bridge between modalities, but in practice, such data is not available and requires selecting or creating effective mechanisms for informative intermediate domains. This paper introduces the Adaptive Generation of Privileged Intermediate Information (AGPI2) training approach to adapt and generate a virtual domain that bridges discriminative information between the V and I modalities. AGPI2enhances the training of a deep V-I ReID backbone by generating and then leveraging bridging privileged information without modifying the model in the inference phase. This information captures shared discriminative attributes that are not easily ascertainable for the model within individual V or I modalities. Towards this goal, a non-linear generative module is trained with adversarial objectives, transforming V attributes into intermediate spaces that also contain I features. This domain exhibits less domain shift relative to the I domain compared to the V domain. Meanwhile, the embedding module within AGPI2aims to extract discriminative modality-invariant features for both modalities by leveraging modality-free descriptors from generated images, making them a bridge between the main modalities. Experiments conducted on challenging V-I ReID datasets indicate that AGPI2consistently increases matching accuracy without additional computational resources during inference. Mahdi Alehdaghi, Arthur Josi, Rafael M. O. Cruz, Pourya Shamsolmoali, Eric Granger |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | Hybrid Gromov-Wasserstein Embedding for Capsule LearningabstractCapsule networks (CapsNets) aim to parse images into a hierarchy of objects, parts, and their relationships using a two-step process involving part-whole transformation and hierarchical component routing. However, this hierarchical relationship modeling is computationally expensive, which has limited the wider use of CapsNet despite its potential advantages. The current state of CapsNet models primarily focuses on comparing their performance with capsule baselines, falling short of achieving the same level of proficiency as deep convolutional neural network (CNN) variants in intricate tasks. To address this limitation, we present an efficient approach for learning capsules that surpasses canonical baseline models and even demonstrates superior performance compared with high-performing convolution models. Our contribution can be outlined in two aspects: first, we introduce a group of subcapsules onto which an input vector is projected. Subsequently, we present the hybrid Gromov-Wasserstein (HGW) framework, which initially quantifies the dissimilarity between the input and the components modeled by the subcapsules, followed by determining their alignment degree through optimal transport (OT). This innovative mechanism capitalizes on new insights into defining alignment between the input and subcapsules, based on the similarity of their respective component distributions. This approach enhances CapsNets' capacity to learn from intricate, high-dimensional data while retaining their interpretability and hierarchical structure. Our proposed model offers two distinct advantages: 1) its lightweight nature facilitates the application of capsules to more intricate vision tasks, including object detection; and 2) it outperforms baseline approaches in these demanding tasks. Our empirical findings illustrate that HGW capsules (HGWCapsules) exhibit enhanced robustness against affine transformations, scale effectively to larger datasets, and surpass CNN and CapsNet models across various vision tasks. Pourya Shamsolmoali, Masoumeh Zareapoor, Swagatam Das, Eric Granger, Salvador García 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | SeTformer Is What You Need for Vision and LanguageabstractThe dot product self-attention (DPSA) is a fundamental component of transformers. However, scaling them to long sequences, like documents or high-resolution images, becomes prohibitively expensive due to the quadratic time and memory complexities arising from the softmax operation. Kernel methods are employed to simplify computations by approximating softmax but often lead to performance drops compared to softmax attention. We propose SeTformer, a novel transformer where DPSA is purely replaced by Self-optimal Transport (SeT) for achieving better performance and computational efficiency. SeT is based on two essential softmax properties: maintaining a non-negative attention matrix and using a nonlinear reweighting mechanism to emphasize important tokens in input sequences. By introducing a kernel cost function for optimal transport, SeTformer effectively satisfies these properties. In particular, with small and base-sized models, SeTformer achieves impressive top-1 accuracies of 84.7% and 86.2% on ImageNet-1K. In object detection, SeTformer-base outperforms the FocalNet counterpart by +2.2 mAP, using 38% fewer parameters and 29% fewer FLOPs. In semantic segmentation, our base-size model surpasses NAT by +3.5 mIoU with 33% fewer parameters. SeTformer also achieves state-of-the-art results in language modeling on the GLUE benchmark. These findings highlight SeTformer applicability for vision and language tasks. Pourya Shamsolmoali, Masoumeh Zareapoor, Eric Granger, Michael Felsberg |
AAAI | 1 |
| 2024 | Rethinking Fast Adversarial Training: A Splitting Technique to Overcome Catastrophic Overfitting
Masoumeh Zareapoor, Pourya Shamsolmoali |
ECCV (78) | 2 |
| 2024 | Autoregressive 3D Shape Completion via Sphere-Guided Disentangled RepresentationabstractThis paper introduces a novel 3D shape completion method based on sphere-guided disentangled representation. Utilizing an autoregressive transformer-based model, our approach efficiently constructs object completion distributions given incomplete point clouds. To enhance completion modeling, we propose sDVQ-DIF (sphere-guided disentangled vector quantized deep implicit function), a new approach using decoupled discrete variables to represent 3D shapes efficiently. Experimental results demonstrate our model’s superior performance in terms of completion quality and fidelity compared to state-of-the-art methods, applicable to various shape types and incomplete patterns. Pourya Shamsolmoali, Yue Lu 0001 |
ICASSP | 2 |
| 2024 | Efficient Routing in Sparse Mixture-of-ExpertsabstractSparse Mixture-of-Experts (MoE) architectures provide the distinct benefit of substantially expanding the model’s parameter space without proportionally increasing the computational load on individual input tokens or samples. However, the efficacy of these models heavily depends on the routing strategy used to assign tokens to experts. Poor routing can lead to under-trained or overly specialized experts, diminishing the overall model performance. Previous approaches have relied on the Topk router, where each token is assigned to a subset of experts. In this paper, we propose a routing mechanism that replaces the Topk router with regularized optimal transport, leveraging the Sinkhorn algorithm to optimize token-expert matching. We conducted a comprehensive evaluation comparing the pre-training efficiency of our model, using computational resources equivalent to those employed in the GShard and Switch Transformers gating mechanisms. The results demonstrate that our model expedites training convergence, achieving a speedup of over 2× compared to these baseline models. Moreover, under the same computational constraints, our model exhibits superior performance across eleven tasks from the GLUE and SuperGLUE benchmarks. We show that our model contributes to the optimization of token-expert matching in sparsely-activated MoE models, offering substantial gains in both training efficiency and task performance. Masoumeh Zareapoor, Pourya Shamsolmoali, Fateme Vesaghati |
IJCNN | 2 |
| 2024 | Fractional Correspondence Framework in Detection TransformerabstractThe Detection Transformer (DETR), by incorporating the Hungarian algorithm, has significantly simplified the matching process in object detection tasks. This algorithm facilitates optimal one-to-one matching of predicted bounding boxes to ground-truth annotations during training. While effective, this strict matching process does not inherently account for the varying densities and distributions of objects, leading to suboptimal correspondences such as failing to handle multiple detections of the same object or missing small objects. To address this, we propose the Regularized Transport Plan (RTP). RTP introduces a flexible matching strategy that captures the cost of aligning predictions with ground truths to find the most accurate correspondences between these sets. By utilizing the differentiable Sinkhorn algorithm, RTP allows for soft, fractional matching rather than strict one-to-one assignments. This approach enhances the model's capability to manage varying object densities and distributions effectively. Our extensive evaluations on the MS-COCO and VOC benchmarks demonstrate the effectiveness of our approach. RTP-DETR, surpassing the performance of the Deform-DETR and the recently introduced DINO-DETR, achieving absolute gains in mAP of +3.8% and +1.7%, respectively. Masoumeh Zareapoor, Pourya Shamsolmoali, Huiyu Zhou 0001, Yue Lu 0001, Salvador García 0001 |
ACM Multimedia | 2 |
| 2024 | TGF: Multiscale transformer graph attention network for multi-sensor image fusion
Hafiz Tayyab Mustafa, Pourya Shamsolmoali, Ikhyun Lee |
Expert Syst. Appl. | 2 |
| 2024 | Slime Mold optimization with hybrid deep learning enabled crowd-counting approach in video surveillance
Zheng Xu 0001, Deepak Kumar Jain 0001, Pourya Shamsolmoali, Alireza Goli, S. Neelakandan, Amar Jain |
Neural Comput. Appl. | 3 |
| 2024 | Distance-based Weighted Transformer Network for image completion
Pourya Shamsolmoali, Masoumeh Zareapoor, Huiyu Zhou 0001, Xuelong Li 0001, Yue Lu 0001 |
Pattern Recognit. | 1 |
| 2023 | Image Completion Via Dual-Path Cooperative FilteringabstractGiven the recent advances with image-generating algorithms, deep image completion methods have made significant progress. However, state-of-art methods typically provide poor cross-scene generalization, and generated masked areas often contain blurry artifacts. Predictive filtering is a method for restoring images, which predicts the most effective kernels based on the input scene. Motivated by this approach, we address image completion as a filtering problem. Deep feature-level semantic filtering is introduced to fill in missing information, while preserving local structure and generating visually realistic content. In particular, a Dual-path Cooperative Filtering (DCF) model is proposed, where one path predicts dynamic kernels, and the other path extracts multi-level features by using Fast Fourier Convolution to yield semantically coherent reconstructions. Experiments on three challenging image completion datasets show that our proposed DCF outperforms state-of-art methods. Pourya Shamsolmoali, Masoumeh Zareapoor, Eric Granger |
ICASSP | 1 |
| 2023 | Handling Class Imbalance by Estimating Minority Class StatisticsabstractThe problem of class imbalance arises in machine learning due to the unequal class-specific distribution of data, where most samples belong to one class, and only a few represent the others. To tackle this issue, one paradigm is to use oversampling techniques that synthesize artificial samples of the minority class using the convex combination of the minority class samples taken in some specialized way for different methods. Existing methods do not take into account any information regarding the actual distribution of the minority class, which leads to inconsistencies between the generated distribution and the actual distribution that the minority class might have. In this paper, we propose a parametrization-based method that tries to estimate the statistics of the minority class samples using the statistics of the nearby classes. Using the different hyperparameters, we can control the distribution such that it may approximate the original distribution. Experiments using synthetic and real-world benchmark datasets demonstrate the usefulness of our techniques across multiple metrics. Faizanuddin Ansari, Swagatam Das, Pourya Shamsolmoali |
IJCNN | 3 |
| 2023 | Entropy Transformer Networks: A Learning Approach via Tangent Bundle Data ManifoldabstractThis paper focuses on an accurate and fast interpolation approach for image transformation employed in the design of CNN architectures. Standard Spatial Transformer Networks (STNs) use bilinear or linear interpolation as their interpolation, with unrealistic assumptions about the underlying data distributions, which leads to poor performance under scale variations. Moreover, STNs do not preserve the norm of gradients in propagation due to their dependency on sparse neighboring pixels. To address this problem, a novel Entropy STN (ESTN) is proposed that interpolates on the data manifold distributions. In particular, random samples are generated for each pixel in association with the tangent space of the data manifold, and construct a linear approximation of their intensity values with an entropy regularizer to compute the transformer parameters. A simple yet effective technique is also proposed to normalize the non-zero values of the convolution operation, to fine-tune the layers for gradients' norm-regularization during training. Experiments on challenging benchmarks show that the proposed ESTN can improve predictive accuracy over a range of computer vision tasks, including image reconstruction, and classification, while reducing the computational cost. Pourya Shamsolmoali, Masoumeh Zareapoor |
IJCNN | 1 |
| 2023 | TransVLAD: Multi-Scale Attention-Based Global Descriptors for Visual Geo-LocalizationabstractVisual geo-localization remains a challenging task due to variations in the appearance and perspective among captured images. This paper introduces an efficient TransVLAD module, which aggregates attention-based feature maps into a discriminative and compact global descriptor. Unlike existing methods that generate feature maps using only convolutional neural networks (CNNs), we propose a sparse transformer to encode global dependencies and compute attention-based feature maps, which effectively reduces visual ambiguities that occurs in large-scale geo-localization problems. A positional embedding mechanism is used to learn the corresponding geometric configurations between query and gallery images. A grouped VLAD layer is also introduced to reduce the number of parameters, and thus construct an efficient module. Finally, rather than only learning from the global descriptors on entire images, we propose a self-supervised learning method to further encode more information from multi-scale patches between the query and positive gallery images. Extensive experiments on three challenging large-scale datasets indicate that our model outperforms state-of-the-art models, and has lower computational complexity. The code is available at: https://github.com/wacv-23/TVLAD. Yifan Xu 0034, Pourya Shamsolmoali, Eric Granger, Claire Nicodeme, Laurent Gardes, Jie Yang 0002 |
WACV | 2 |
| 2023 | GEN: Generative Equivariant Networks for Diverse Image-to-Image TranslationabstractImage-to-image (I2I) translation has become a key asset for generative adversarial networks. Convolutional neural networks (CNNs), despite having a significant performance, are not able to capture the spatial relationships among different parts of an object and, thus, do not qualify as the ideal representative model for image translation tasks. As a remedy to this problem, capsule networks have been proposed to represent patterns for a visual object in such a way that preserves hierarchical spatial relationships. The training of capsules is constrained by learning all pairwise relationships between capsules of consecutive layers. This design would be prohibitively expensive both in time and memory. In this article, we present a new framework for capsule networks to provide a full description of the input components at various levels of semantics, which can successfully be applied to the generator-discriminator architectures without incurring computational overhead compared to the CNNs. To successfully apply the proposed capsules in the generative adversarial network, we put forth a novel Gromov-Wasserstein (GW) distance as a differentiable loss function that compares the dissimilarity between two distributions and then guides the learned distribution toward target properties, using optimal transport (OT) discrepancy. The proposed method-which is called generative equivariant network (GEN)-is an alternative architecture for GANs with equivariance capsule layers. The proposed model is evaluated through a comprehensive set of experiments on I2I translation and image generation tasks and compared with several state-of-the-art models. Results indicate that there is a principled connection between generative and capsule models that allows extracting discriminant and invariant information from image data. Pourya Shamsolmoali, Masoumeh Zareapoor, Swagatam Das, Salvador García 0001, Eric Granger, Jie Yang 0002 |
IEEE Trans. Cybern. | 1 |
| 2023 | Efficient Object Detection in Optical Remote Sensing Imagery via Attention-Based Feature DistillationabstractEfficient object detection methods have recently received great attention in remote sensing. Although deep convolutional networks often have excellent detection accuracy, their deployment on resource-limited edge devices is difficult. Knowledge distillation (KD) is a strategy for addressing this issue since it makes models lightweight while maintaining accuracy. However, existing KD methods for object detection have encountered two constraints. First, they discard potentially important background information and only distill nearby foreground regions. Second, they only rely on the global context, which limits the student detector’s ability to acquire local information from the teacher detector. To address the aforementioned challenges, we propose Attention-based Feature Distillation (AFD), a new KD approach that distills both local and global information from the teacher detector. To enhance local distillation, we introduce a multi-instance attention mechanism that effectively distinguishes between background and foreground elements. This approach prompts the student detector to focus on the pertinent channels and pixels, as identified by the teacher detector. Local distillation lacks global information, thus attention global distillation is proposed to reconstruct the relationship between various pixels and pass it from teacher to student detector. The performance of AFD is evaluated on two public aerial image benchmarks, and the evaluation results demonstrate that AFD in object detection can attain the performance of other state-of-the-art models while being efficient. Pourya Shamsolmoali, Jocelyn Chanussot, Huiyu Zhou 0001, Yue Lu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | VTAE: Variational Transformer Autoencoder With Manifolds LearningabstractDeep generative models have demonstrated successful applications in learning non-linear data distributions through a number of latent variables and these models use a non-linear function (generator) to map latent samples into the data space. On the other hand, the non-linearity of the generator implies that the latent space shows an unsatisfactory projection of the data space, which results in poor representation learning. This weak projection, however, can be addressed by a Riemannian metric, and we show that geodesics computation and accurate interpolations between data samples on the Riemannian manifold can substantially improve the performance of deep generative models. In this paper, a Variational spatial-Transformer AutoEncoder (VTAE) is proposed to minimize geodesics on a Riemannian manifold and improve representation learning. In particular, we carefully design the variational autoencoder with an encoded spatial-Transformer to explicitly expand the latent variable model to data on a Riemannian manifold, and obtain global context modelling. Moreover, to have smooth and plausible interpolations while traversing between two different objects' latent representations, we propose a geodesic interpolation network different from the existing models that use linear interpolation with inferior performance. Experiments on benchmarks show that our proposed model can improve predictive accuracy and versatility over a range of computer vision tasks, including image interpolations, and reconstructions. Pourya Shamsolmoali, Masoumeh Zareapoor, Huiyu Zhou 0001, Dacheng Tao, Xuelong Li 0001 |
IEEE Trans. Image Process. | 1 |
| 2022 | Salient Skin Lesion Segmentation via Dilated Scale-Wise Feature Fusion NetworkabstractSkin lesion detection in dermoscopic images is essential in the accurate and early diagnosis of skin cancer by a computerized apparatus. Current skin lesion segmentation approaches show poor performance in challenging circumstances such as indistinct lesion boundaries, low contrast between the lesion and the surrounding area, or heterogeneous background that causes over/under segmentation of the skin lesion. To accurately recognize the lesion from the neighboring regions, we propose a dilated scale-wise feature fusion network based on convolution factorization. Our network is designed to simultaneously extract features at different scales which are systematically fused for better detection. The proposed model has satisfactory accuracy and efficiency. Various experiments for lesion segmentation are performed along with comparisons with the state-of-the-art models. Our proposed model consistently showcases state-of-the-art results. Pourya Shamsolmoali, Masoumeh Zareapoor, Jie Yang 0002, Eric Granger, Huiyu Zhou 0001 |
ICPR | 1 |
| 2022 | Weak-supervised Visual Geo-localization via Attention-based Knowledge DistillationabstractVisual geo-localization aims to estimate the geo-graphical location of a query image by identifying the best-matched reference image from a GPS-tagged database. It remains a challenging task because of image appearance changes such as lighting, scale and pose. The current approaches do not have satisfactory performance for large-scale environments owing to the lack of learning discriminative features for image matching. To address the above problem, we introduce a practical method to exploit a weak-supervised model with selective transfer for feature distillation. We propose an image matching method that uses image sub-regions to adequately analyze the potential of difficult positive images. For improving the network generations and performance, the model estimates image-to-region similarity labels at no additional parameters or manual annotations by use of soft-labeled loss. Moreover, to have optimal training we propose a novel knowledge distillation (KD) method to effectively capture and transfer knowledge of a teacher network to a student network. More specifically, our method uses an attention network to learn relative similarities within features and utilizes these similarities to enhance the distillation intensities by further exploring the potential of difficult positive images. Our model achieves significant localization performance over large variations of appearance on three challenging datasets with satisfactory efficiency. Our code is available at https://github.com/XuYifan98/WAKD. Yifan Xu 0034, Pourya Shamsolmoali, Jie Yang 0002 |
ICPR | 2 |
| 2022 | Enhanced Single-Shot Detector for Small Object Detection in Remote Sensing ImagesabstractSmall-object detection is a challenging problem. In the last few years, the convolution neural networks methods have been achieved considerable progress. However, the current detectors struggle with effective features extraction for small-scale objects. To address this challenge, we propose image pyramid single-shot detector (IPSSD). In IPSSD, single-shot detector is adopted combined with an image pyramid network to extract semantically strong features for generating candidate regions. The proposed network can enhance the small-scale features from a feature pyramid network. We evaluated the performance of the proposed model on two public datasets and the results show the superior performance of our model compared to the other state-of-the-art object detectors. Pourya Shamsolmoali, Masoumeh Zareapoor, Jie Yang 0002, Eric Granger, Jocelyn Chanussot |
IGARSS | 1 |
| 2022 | Multipatch Feature Pyramid Network for Weakly Supervised Object Detection in Optical Remote Sensing ImagesabstractObject detection is a challenging task in remote sensing because objects only occupy a few pixels in the images, and the models are required to simultaneously learn object locations and detection. Even though the established approaches well perform for the objects of regular sizes, they achieve weak performance when analyzing small ones or getting stuck in the local minima (e.g. false object parts). Two possible issues stand in their way. First, the existing methods struggle to perform stably on the detection of small objects because of the complicated background. Second, most of the standard methods used hand-crafted features, and do not work well on the detection of objects parts of which are missing. We here address the above issues and propose a new architecture with a multiple patch feature pyramid network (MPFP-Net). Different from the current models that during training only pursue the most discriminative patches, in MPFPNet the patches are divided into class-affiliated subsets, in which the patches are related and based on the primary loss function, a sequence of smooth loss functions are determined for the subsets to improve the model for collecting small object parts. To enhance the feature representation for patch selection, we introduce an effective method to regularize the residual values and make the fusion transition layers strictly norm-preserving. The network contains bottom-up and crosswise connections to fuse the features of different scales to achieve better accuracy, compared to several state-of-the-art object detection models. Also, the developed architecture is more efficient than the baselines. Pourya Shamsolmoali, Jocelyn Chanussot, Masoumeh Zareapoor, Huiyu Zhou 0001, Jie Yang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Rotation Equivariant Feature Image Pyramid Network for Object Detection in Optical Remote Sensing ImageryabstractDetection of objects is extremely important in various aerial vision-based applications. Over the last few years, the methods based on convolution neural networks (CNNs) have made substantial progress. However, because of the large variety of object scales, densities, and arbitrary orientations, the current detectors struggle with the extraction of semantically strong features for small-scale objects by a predefined convolution kernel. To address this problem, we propose the rotation equivariant feature image pyramid network (REFIPN), an image pyramid network based on rotation equivariance convolution. The proposed model adopts single-shot detector in parallel with a lightweight image pyramid module (LIPM) to extract representative features and generate regions of interest in an optimization approach. The proposed network extracts feature in a wide range of scales and orientations by using novel convolution filters. These features are used to generate vector fields and determine the weight and angle of the highest-scoring orientation for all spatial locations on an image. By this approach, the performance for small-sized object detection is enhanced without sacrificing the performance for large-sized object detection. The performance of the proposed model is validated on two commonly used aerial benchmarks and the results show our proposed model can achieve state-of-the-art performance with satisfactory efficiency. Pourya Shamsolmoali, Masoumeh Zareapoor, Jocelyn Chanussot, Huiyu Zhou 0001, Jie Yang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | An evolutionary lion optimization algorithm-based image compression technique for biomedical applicationsabstractAbstract Recently, medical image compression becomes essential to effectively handle large amounts of medical data for storage and communication purposes. Vector quantization (VQ) is a popular image compression technique, and the commonly used VQ model is Linde–Buzo–Gray (LBG) that constructs a local optimal codebook to compress images. The codebook construction was considered as an optimization problem, and a bioinspired algorithm was employed to solve it. This article proposed a VQ codebook construction approach called the L2‐LBG method utilizing the Lion optimization algorithm (LOA) and Lempel Ziv Markov chain Algorithm (LZMA). Once LOA constructed the codebook, LZMA was applied to compress the index table and further increase the compression performance of the LOA. A set of experimentation has been carried out using the benchmark medical images, and a comparative analysis was conducted with Cuckoo Search‐based LBG (CS‐LBG), Firefly‐based LBG (FF‐LBG) and JPEG2000. The compression efficiency of the presented model was validated in terms of compression ratio (CR), compression factor (CF), bit rate, and peak signal to noise ratio (PSNR). The proposed L2‐LBG method obtained a higher CR of 0.3425375 and PSNR value of 52.62459 compared to CS‐LBG, FA‐LBG, and JPEG2000 methods. The experimental values revealed that the L2‐LBG process yielded effective compression performance with a better‐quality reconstructed image. Karuppaiah Geetha, Veerasamy Anitha, Mohamed Elhoseny, K. Shankar 0002, Pourya Shamsolmoali, Mahmoud Mohamed Selim |
Expert Syst. J. Knowl. Eng. | 5 |
| 2021 | Imbalanced data learning by minority class augmentation using capsule adversarial networks
Pourya Shamsolmoali, Masoumeh Zareapoor, LinLin Shen, Abdul Hamid Sadka, Jie Yang 0002 |
Neurocomputing | 1 |
| 2021 | Advances in domain adaptation for computer vision
Pourya Shamsolmoali, Salvador García 0001, Huiyu Zhou 0001, M. Emre Celebi 0001 |
Image Vis. Comput. | 1 |
| 2021 | Road Segmentation for Remote Sensing Images Using Adversarial Spatial Pyramid NetworksabstractRoad extraction in remote sensing images is of great importance for a wide range of applications. Because of the complex background, and high density, most of the existing methods fail to accurately extract a road network that appears correct and complete. Moreover, they suffer from either insufficient training data or high costs of manual annotation. To address these problems, we introduce a new model to apply structured domain adaption for synthetic image generation and road segmentation. We incorporate a feature pyramid (FP) network into generative adversarial networks to minimize the difference between the source and target domains. A generator is learned to produce quality synthetic images, and the discriminator attempts to distinguish them. We also propose a FP network that improves the performance of the proposed model by extracting effective features from all the layers of the network for describing different scales' objects. Indeed, a novel scale-wise architecture is introduced to learn from the multilevel feature maps and improve the semantics of the features. For optimization, the model is trained by a joint reconstruction loss function, which minimizes the difference between the fake images and the real ones. A wide range of experiments on three data sets prove the superior performance of the proposed approach in terms of accuracy and efficiency. In particular, our model achieves state-of-the-art 78.86 IOU on the Massachusetts data set with 14.89M parameters and 86.78B FLOPs, with 4× fewer FLOPs but higher accuracy (+3.47% IOU) than the top performer among state-of-the-art approaches used in the evaluation. Pourya Shamsolmoali, Masoumeh Zareapoor, Huiyu Zhou 0001, Ruili Wang 0001, Jie Yang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | Introduction to the Special Issue on Advanced Approaches for Multiple Instance Learning on Multimedia ApplicationsabstractNo abstract available. Pourya Shamsolmoali, Ruili Wang 0001, Abdul Hamid Sadka |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2021 | Multi-frame feature-fusion-based model for violence detection
Mujtaba Asad, Jie Yang 0002, Pourya Shamsolmoali, Xiangjian He |
Vis. Comput. | 4 |
| 2020 | Synthetic guided domain adaptive and edge aware network for crowd counting
Zhijie Cao, Pourya Shamsolmoali, Jie Yang 0002 |
Image Vis. Comput. | 2 |
| 2020 | Deep neural learning techniques with long short-term memory for gesture recognition
Deepak Kumar Jain 0001, Aniket Mahanti, Pourya Shamsolmoali, Manikandan Ramachandran |
Neural Comput. Appl. | 3 |
| 2020 | Deep learning approaches for real-time image super-resolution
Pourya Shamsolmoali, M. Emre Celebi 0001, Ruili Wang 0001 |
Neural Comput. Appl. | 1 |
| 2020 | Advanced deep learning for image super-resolution
Pourya Shamsolmoali, Abdul Hamid Sadka, Huiyu Zhou 0001, Wankou Yang |
Signal Process. Image Commun. | 1 |
| 2020 | AMIL: Adversarial Multi-instance Learning for Human Pose EstimationabstractHuman pose estimation has an important impact on a wide range of applications, from human-computer interface to surveillance and content-based video retrieval. For human pose estimation, joint obstructions and overlapping upon human bodies result in departed pose estimation. To address these problems, by integrating priors of the structure of human bodies, we present a novel structure-aware network to discreetly consider such priors during the training of the network. Typically, learning such constraints is a challenging task. Instead, we propose generative adversarial networks as our learning model in which we design two residual Multiple-Instance Learning (MIL) models with identical architecture—one is used as the generator, and the other one is used as the discriminator. The discriminator task is to distinguish the actual poses from the fake ones. If the pose generator generates results that the discriminator is not able to distinguish from the real ones, then the model has successfully learned the priors. In the proposed model, the discriminator differentiates the ground-truth heatmaps from the generated ones, and later the adversarial loss back-propagates to the generator. Such procedure assists the generator to learn reasonable body configurations and is proved to be advantageous to improve the pose estimation accuracy. Meanwhile, we propose a novel function for MIL. It is an adjustable structure for both instance selection and modeling to appropriately pass the information between instances in a single bag. In the proposed residual MIL neural network, the pooling action adequately updates the instance contribution to its bag. The proposed adversarial residual multi-instance neural network that is based on pooling has been validated on two datasets for the human pose estimation task and successfully outperforms the other state-of-the-art models. The code will be made available on https://github.com/pshams55/AMIL. Pourya Shamsolmoali, Masoumeh Zareapoor, Huiyu Zhou 0001, Jie Yang 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2019 | Convolutional neural network in network (CNNiN): hyperspectral image classification and dimensionality reductionabstractClassification is a principle technique in hyperspectral images (HSIs), where a label is assigned to each pixel based on its characteristics. However, due to lack of labelled training instances in HSIs and also its ultra‐high dimensionality, deep learning approaches need a special consideration for HSI classification. As one of the first works in the HSI classification, this study proposes a novel network pipeline called convolutional neural network in network (which is deeper than the existing approaches) by jointly utilising the spatial and spectral information and produces high‐level features from the original HSI. This can occur by using spatial–spectral relationships of individual pixel vector at the initial component of the proposed pipeline; the extracted features are then combined to form a joint spatial–spectral feature map. Finally, a recurrent neural network is trained on the extracted features which contain wealthy spectral and spatial properties of the HSI to predict the corresponding label of each vector. The model has been tested on two large scale hyperspectral datasets in terms of classification accuracy, training error, and computational time. Pourya Shamsolmoali, Masoumeh Zareapoor, Jie Yang 0002 |
IET Image Process. | 1 |
| 2019 | G-GANISR: Gradual generative adversarial network for image super resolution
Pourya Shamsolmoali, Masoumeh Zareapoor, Ruili Wang 0001, Deepak Kumar Jain 0001, Jie Yang 0002 |
Neurocomputing | 1 |
| 2019 | Image super resolution by dilated dense progressive network
Pourya Shamsolmoali, Masoumeh Zareapoor, Jie Yang 0002 |
Image Vis. Comput. | 1 |
| 2019 | High-dimensional multimedia classification using deep CNN and extended residual units
Pourya Shamsolmoali, Deepak Kumar Jain 0001, Masoumeh Zareapoor, Jie Yang 0002, Mohd. Afshar Alam |
Multim. Tools Appl. | 1 |
| 2019 | Deep convolution network for surveillance records super-resolution
Pourya Shamsolmoali, Masoumeh Zareapoor, Deepak Kumar Jain 0001, Vinay Kumar Jain, Jie Yang 0002 |
Multim. Tools Appl. | 1 |
| 2019 | Deep semantic preserving hashing for large scale image retrieval
Masoumeh Zareapoor, Jie Yang 0002, Deepak Kumar Jain 0001, Pourya Shamsolmoali, Neha Jain 0003, Surya Kant |
Multim. Tools Appl. | 4 |
| 2019 | Extended deep neural network for facial emotion recognition
Deepak Kumar Jain 0001, Pourya Shamsolmoali, Paramjit S. Sehdev |
Pattern Recognit. Lett. | 2 |
| 2019 | Single image resolution enhancement by efficient dilated densely connected residual network
Pourya Shamsolmoali, Ruili Wang 0001 |
Signal Process. Image Commun. | 1 |
| 2018 | Hybrid deep neural networks for face emotion recognition
Neha Jain 0003, Shishir Kumar, Amit Kumar 0023, Pourya Shamsolmoali, Masoumeh Zareapoor |
Pattern Recognit. Lett. | 4 |
| 2018 | Kernelized support vector machine with deep learning: An efficient approach for extreme multiclass dataset
Masoumeh Zareapoor, Pourya Shamsolmoali, Deepak Kumar Jain 0001, Haoxiang Wang 0001, Jie Yang 0002 |
Pattern Recognit. Lett. | 2 |