Mohammad Sabokrou

dblp:163/2030 · DBLP profile ↗
← Back
36ranked-venue papers
8as first author
26since 2021 · last 2026
0000-0002-9409-2799ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 7 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 7 since 2021Computer networks · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Exploiting Layer-Specific Vulnerabilities to Backdoor Attack in Federated Learning
Mohammad Hadi Foroughi, Seyed Hamed Rastegar, Mohammad Sabokrou, Ahmad Khonsari
ICC3
2026 What are we really measuring? Rethinking dataset bias in web-scale natural image collections via unsupervised semantic clustering
Amir Hossein Saleknia, Mohammad Sabokrou
Neurocomputing2
2026 Explainability of vision transformers: a comprehensive review and new perspectives
Rojina Kashefi, Leili Barekatain, Mohammad Sabokrou, Fatemeh Aghaeipoor
Multim. Tools Appl.3
2025 Adversarially Robust Anomaly Detection through Spurious Negative Pair Mitigation
abstract
Despite significant progress in Anomaly Detection (AD), the robustness of existing detection methods against adversarial attacks remains a challenge, compromising their reliability in critical real-world applications such as autonomous driving. This issue primarily arises from the AD setup, which assumes that training data is limited to a group of unlabeled normal samples, making the detectors vulnerable to adversarial anomaly samples during testing. Additionally, implementing adversarial training as a safeguard encounters difficulties, such as formulating an effective objective function without access to labels. An ideal objective function for adversarial training in AD should promote strong perturbations both within and between the normal and anomaly groups to maximize margin between normal and anomaly distribution. To address these issues, we first propose crafting a pseudo-anomaly group derived from normal group samples. Then, we demonstrate that adversarial training with contrastive loss could serve as an ideal objective function, as it creates both inter- and intra-group perturbations. However, we notice that spurious negative pairs compromise the conventional contrastive loss for achieving robust AD. Spurious negative pairs are those that should be mapped closely but are erroneously separated. These pairs introduce noise and misguide the direction of inter-group adversarial perturbations. To overcome the effect of spurious negative pairs, we define opposite pairs and adversarially pull them apart to strengthen inter-group perturbations. Experimental results demonstrate our superior performance in both clean and adversarial scenarios, with a 26.1% improvement in robust detection across various challenging benchmark datasets.
Hossein Mirzaei, Mojtaba Nafez, Jafar Habibi, Mohammad Sabokrou, Mohammad H. Rohban
ICLR4
2025 FrameShield: Adversarially Robust Video Anomaly Detection
abstract
Weakly Supervised Video Anomaly Detection (WSVAD) has achieved notable advancements, yet existing models remain vulnerable to adversarial attacks, limiting their reliability. Due to the inherent constraints of weak supervision—where only video-level labels are provided despite the need for frame-level predictions—traditional adversarial defense mechanisms, such as adversarial training, are not effective since video-level adversarial perturbations are typically weak and inadequate. To address this limitation, pseudo-labels generated directly from the model can enable frame-level adversarial training; however, these pseudo-labels are inherently noisy, significantly degrading performance. We therefore introduce a novel Pseudo-Anomaly Generation method called Spatiotemporal Region Distortion (SRD), which creates synthetic anomalies by applying severe augmentations to localized regions in normal videos while preserving temporal consistency. Integrating these precisely annotated synthetic anomalies with the noisy pseudo-labels substantially reduces label noise, enabling effective adversarial training. Extensive experiments demonstrate that our method significantly enhances the robustness of WSVAD models against adversarial attacks, outperforming state-of-the-art methods by an average of 71.0\% in overall AUROC performance across multiple benchmarks. The implementation and code are publicly available at [FrameShield (GitHub)](https://github.com/rohban-lab/FrameShield).
Mojtaba Nafez, Mobina Poulaei, Nikan Vasei, Bardia Soltani Moakhar, Mohammad Sabokrou, Mohammad H. Rohban
NeurIPS5
2025 APML: Adaptive Probabilistic Matching Loss for Robust 3D Point Cloud Reconstruction
abstract
Training deep learning models for point cloud prediction tasks such as shape completion and generation depends critically on loss functions that measure discrepancies between predicted and ground-truth point sets. Commonly used functions such as Chamfer Distance (CD), HyperCD, InfoCD and Density-aware CD rely on nearest-neighbor assignments, which often induce many-to-one correspondences, leading to point congestion in dense regions and poor coverage in sparse regions. These losses also involve non-differentiable operations due to index selection, which may affect gradient-based optimization. Earth Mover Distance (EMD) enforces one-to-one correspondences and captures structural similarity more effectively, but its cubic computational complexity limits its practical use. We propose the Adaptive Probabilistic Matching Loss (APML), a fully differentiable approximation of one-to-one matching that leverages Sinkhorn iterations on a temperature-scaled similarity matrix derived from pairwise distances. We analytically compute the temperature to guarantee a minimum assignment probability, eliminating manual tuning. APML achieves near-quadratic runtime, comparable to Chamfer-based losses, and avoids non-differentiable operations. When integrated into state-of-the-art architectures (PoinTr, PCN, FoldingNet) on ShapeNet benchmarks and on a spatio‑temporal Transformer (CSI2PC) that generates 3‑D human point clouds from WiFi‑CSI measurements, APM loss yields faster convergence, superior spatial distribution, especially in low-density regions, and improved or on-par quantitative performance without additional hyperparameter search. The code is available at: https://github.com/apm-loss/apml.
Sasan Sharifipour, Constantino Álvarez Casado, Mohammad Sabokrou, Miguel Bordallo López
NeurIPS3
2025 A unified concept-based system for local, global, and misclassification explanations
Fatemeh Aghaeipoor, Dorsa Asgarian, Mohammad Sabokrou
Neurocomputing3
2024 Universal Novelty Detection Through Adaptive Contrastive Learning
abstract
Novelty detection is a critical task for deploying machine learning models in the open world. A crucial property of novelty detection methods is universality, which can be interpreted as generalization across various distributions of training or test data. More precisely, for novelty detection, distribution shifts may occur in the training set or the test set. Shifts in the training set refer to cases where we train a novelty detector on a new dataset and expect strong transferability. Conversely, distribution shifts in the test set indicate the methods' performance when the trained model encounters a shifted test sample. We experimentally show that existing methods falter in maintaining universality, which stems from their rigid inductive biases. Motivated by this, we aim for more generalized techniques that have more adaptable inductive biases. In this context, we leverage the fact that contrastive learning provides an efficient framework to easily switch and adapt to new inductive biases through the proper choice of augmentations in forming the negative pairs. We propose a novel probabilistic auto-negative pair generation method (AutoAugOOD), along with contrastive learning, to yield a universal novelty detector method. Our experiments demonstrate the superiority of our method under different distribution shifts in various image benchmark datasets. Notably, our method emerges universality in the lens of adaptability to different setups of novelty detection, including one-class, unlabeled multi-class, and labeled multi-class settings.
Hossein Mirzaei, Mojtaba Nafez, Mohammad Bagher Soltani, Mohammad Azizmalayeri, Jafar Habibi, Mohammad Sabokrou, Mohammad H. Rohban
CVPR7
2024 Adversarial Backdoor Attack by Naturalistic Data Poisoning on Trajectory Prediction in Autonomous Driving
abstract
In autonomous driving, behavior prediction is funda-mental for safe motion planning, hence the security and ro-bustness of prediction models against adversarial attacks are of paramount importance. We propose a novel adver-sarial backdoor attack against trajectory prediction models as a means of studying their potential vulnerabilities. Our attack affects the victim at training time via naturalistic, hence stealthy, poisoned samples crafted using a novel two-step approach. First, the triggers are crafted by perturbing the trajectory of attacking vehicle and then disguised by transforming the scene using a bi-level optimization technique. The proposed attack does not depend on a particu-lar model architecture and operates in a black-box manner, thus can be effective without any knowledge of the victim model. We conduct extensive empirical studies using state-of-the-art prediction models on two benchmark datasets using metrics customized for trajectory prediction. We show that the proposed attack is highly effective, as it can sig-nificantly hinder the performance of prediction models, un-noticeable by the victims, and efficient as it forces the vic-tim to generate malicious behavior even under constrained conditions. Via ablative studies, we analyze the impact of different attack design choices followed by an evaluation of existing defence mechanisms against the proposed attack.
Mozhgan PourKeshavarz, Mohammad Sabokrou, Amir Rasouli
CVPR2
2024 Killing It With Zero-Shot: Adversarially Robust Novelty Detection
abstract
Novelty Detection (ND) plays a crucial role in machine learning by identifying new or unseen data during model inference. This capability is especially important for the safe and reliable operation of automated systems. Despite advances in this field, existing techniques often fail to maintain their performance when subject to adversarial attacks. Our research addresses this gap by marrying the merits of nearest-neighbor algorithms with robust features obtained from models pretrained on ImageNet. We focus on enhancing the robustness and performance of ND algorithms. Experimental results demonstrate that our approach significantly outperforms current state-of-the-art methods across various benchmarks, particularly under adversarial conditions. By incorporating robust pretrained features into the k-NN algorithm, we establish a new standard for performance and robustness in the field of robust ND. This work opens up new avenues for research aimed at fortifying machine learning systems against adversarial vulnerabilities.
Hossein Mirzaei, Hamid Reza Dehbashi, Zeinab Taghavi 0001, Mohammad Sabokrou, Mohammad H. Rohban
ICASSP5
2024 Scanning Trojaned Models Using Out-of-Distribution Samples
abstract
Scanning for trojan (backdoor) in deep neural networks is crucial due to their significant real-world applications. There has been an increasing focus on developing effective general trojan scanning methods across various trojan attacks. Despite advancements, there remains a shortage of methods that perform effectively without preconceived assumptions about the backdoor attack method. Additionally, we have observed that current methods struggle to identify classifiers trojaned using adversarial training. Motivated by these challenges, our study introduces a novel scanning method named TRODO (TROjan scanning by Detection of adversarial shifts in Out-of-distribution samples). TRODO leverages the concept of "blind spots"—regions where trojaned classifiers erroneously identify out-of-distribution (OOD) samples as in-distribution (ID). We scan for these blind spots by adversarially shifting OOD samples towards in-distribution. The increased likelihood of perturbed OOD samples being classified as ID serves as a signature for trojan detection. TRODO is both trojan and label mapping agnostic, effective even against adversarially trained trojaned classifiers. It is applicable even in scenarios where training data is absent, demonstrating high accuracy and adaptability across various scenarios and datasets, highlighting its potential as a robust trojan scanning strategy.
Hossein Mirzaei, Ali Ansari 0001, Bahar Dibaei Nia, Mojtaba Nafez, Moein Madadi, Sepehr Rezaee, Zeinab Taghavi 0001, Arad Maleki, Kian Shamsaie, Mahdi Hajialilue, Jafar Habibi, Mohammad Sabokrou, Mohammad H. Rohban
NeurIPS12
2024 A survey on recent advances in Sign Language Production
Razieh Rastgoo, Kourosh Kiani, Sergio Escalera, Vassilis Athitsos, Mohammad Sabokrou
Expert Syst. Appl.5
2024 Multi-modal zero-shot dynamic hand gesture recognition
Razieh Rastgoo, Kourosh Kiani, Sergio Escalera, Mohammad Sabokrou
Expert Syst. Appl.4
2024 Projan: A probabilistic trojan attack on deep neural networks
Mehrin Saremi, Mohammad Khalooei, Razieh Rastgoo, Mohammad Sabokrou
Knowl. Based Syst.4
2024 Diverse hand gesture recognition dataset
Zahra Mohammadi, Alireza Akhavanpour, Razieh Rastgoo, Mohammad Sabokrou
Multim. Tools Appl.4
2024 STARNet: spatio-temporal aware recurrent network for efficient video object detection on embedded devices
Mohammad Hajizadeh, Mohammad Sabokrou, Adel Torkaman Rahmani
Mach. Vis. Appl.2
2023 Mitigating Bias: Enhancing Image Classification by Improving Model Explanations
Raha Ahmadi, Mohammad Javad Rajabi, Mohammad Khalooei, Mohammad Sabokrou
ACML4
2023 Fake It Until You Make It : Towards Accurate Near-Distribution Novelty Detection
Hossein Mirzaei, Mohammadreza Salehi, Sajjad Shahabi, Efstratios Gavves, Cees Snoek, Mohammad Sabokrou, Mohammad H. Rohban
ICLR6
2023 MobileDenseNet: A new approach to object detection on mobile devices
Mohammad Hajizadeh, Mohammad Sabokrou, Adel Torkaman Rahmani
Expert Syst. Appl.2
2023 Fuzzy Rule-Based Explainer Systems for Deep Neural Networks: From Local Explainability to Global Understanding
abstract
Explainability of deep neural networks has been receiving increasing attention with regard to auditability and trustworthiness purposes. Of the various post-hoc explainability approaches, rule extraction methods assist to understand the logic that underpins their functioning. Whereas the rule-based solutions are directly managed and understood by practitioners, the use of intervals or crisp values in the antecedents that rely on numerical values might not be intuitive enough. In this case, the benefits of a linguistic representation based on fuzzy sets/rules are straightforward, as these semantically meaningful components ease the model understanding. This article proposes fuzzy rule-based explainer systems for deep neural networks. The algorithm learns a compact yet accurate set of fuzzy rules based on features' importance (i.e., attribution values) distilled from the trained networks. These systems can be used for both local and global explainability purposes. The evaluation results of different applications revealed that the fuzzy explainers maintained the fidelity and accuracy of the original deep neural networks while implying lower complexity and better comprehensibility.
Fatemeh Aghaeipoor, Mohammad Sabokrou, Alberto Fernández 0001
IEEE Trans. Fuzzy Syst.2
2022 Looking Back on Learned Experiences For Class/task Incremental Learning
Mozhgan PourKeshavarz, Guoying Zhao 0001, Mohammad Sabokrou
ICLR3
2022 Deep-Disaster: Unsupervised Disaster Detection and Localization Using Visual Data
abstract
Social media plays a significant role in sharing essential information, which helps humanitarian organizations in rescue operations during and after disaster incidents. However, developing an efficient method that can provide rapid analysis of social media images in the early hours of disasters is still largely an open problem, mainly due to the lack of suitable datasets and the sheer complexity of this task. In addition, supervised methods can not generalize well to novel disaster incidents. In this paper, inspired by the success of Knowledge Distillation (KD) methods, we propose an unsupervised deep neural network to detect and localize damages in social media images. Our proposed KD architecture is a feature-based distillation approach that comprises a pre-trained teacher and a smaller student network, with both networks having similar GAN architecture containing a generator and a discriminator. The student network is trained to emulate the teacher’s behavior on training input samples, which, in turn, contain images that do not include any damaged regions. Therefore, the student network only learns the distribution of no damage data and would have different behavior from the teacher network facing damages. To detect damage, we utilize the difference between features generated by two networks using a defined score function that demonstrates the probability of damages occurring. Our experimental results on the benchmark dataset confirm that our approach outperforms state-of-the-art methods in detecting and localizing the damaged areas, especially for novel disaster types1.
Soroor Shekarizadeh, Razieh Rastgoo, Saif M. Al-Kuwari, Mohammad Sabokrou
ICPR4
2021 G2D: Generate to Detect Anomaly
abstract
In this paper, we propose a novel method for irregularity detection. Previous researches solve this problem as a One-Class Classification (OCC) task where they train a reference model on all of the available samples. Then, they consider a test sample as an anomaly if it has a diversion from the reference model. Generative Adversarial Networks (GANs) have achieved the most promising results for OCC while implementing and training such networks, especially for the OCC task, is a cumbersome and computationally expensive procedure. To cope with the mentioned challenges, we present a simple but effective method to solve the irregularity detection as a binary classification task in order to make the implementation easier along with improving the detection performance. We learn two deep neural networks (generator and discriminator) in a GAN-style setting on merely the normal samples. During training, the generator gradually becomes an expert to generate samples which are similar to the normal ones. In the training phase, when the generator fails to produce normal data (in the early stages of learning and also prior to the complete convergence), it can be considered as an irregularity generator. In this way, we simultaneously generate the irregular samples. Afterward, we train a binary classifier on the generated anomalous samples along with the normal instances in order to be capable of detecting irregularities. The proposed framework applies to different related applications of outlier and anomaly detection in images and videos, respectively. The results confirm that our proposed method is superior to the baseline and state-of-the-art solutions.
Masoud PourReza, Bahram Mohammadi, Mostafa Khaki, Samir Bouindour, Hichem Snoussi, Mohammad Sabokrou
WACV6
2021 Deep-HR: Fast heart rate estimation from face video under realistic conditions
Mohammad Sabokrou, Masoud PourReza, Mahmood Fathy, Guoying Zhao 0001
Expert Syst. Appl.1
2021 Improving the learning of self-driving vehicles based on real driving behavior using deep neural network techniques
Nayereh Zaghari, Mahmood Fathy, Seyed Mahdi Jameii, Mohammad Sabokrou, Mohammad Shahverdy
J. Supercomput.4
2021 Deep End-to-End One-Class Classifier
abstract
One-class classification (OCC) poses as an essential component in many machine learning and computer vision applications, including novelty, anomaly, and outlier detection systems. With a known definition for a target or normal set of data, one-class classifiers can determine if any given new sample spans within the distribution of the target class. Solving for this task in a general setting is particularly very challenging, due to the high diversity of samples from the target class and the absence of any supervising signal over the novelty (nontarget) concept, which makes designing end-to-end models unattainable. In this article, we propose an adversarial training approach to detect out-of-distribution samples in an end-to-end trainable deep model. To this end, we jointly train two deep neural networks, R and D . The latter plays as the discriminator while the former, during training, helps D characterize a probability distribution for the target class by creating adversarial examples and, during testing, collaborates with it to detect novelties. Using our OCC, we first test outlier detection on two image data sets, Modified National Institute of Standards and Technology (MNIST) and Caltech-256. Then, several experiments for video anomaly detection are performed on University of Minnesota (UMN) and University of California, San Diego (UCSD) data sets. Our proposed method can successfully learn the target class underlying distribution and outperforms other approaches.
Mohammad Sabokrou, Mahmood Fathy, Guoying Zhao 0001, Ehsan Adeli-Mosabbeb
IEEE Trans. Neural Networks Learn. Syst.1
2020 Driver behavior detection and classification using deep convolutional neural networks
Mohammad Shahverdy, Mahmood Fathy, Reza Berangi, Mohammad Sabokrou
Expert Syst. Appl.4
2020 Vision-based human activity recognition: a survey
abstract
Abstract Human activity recognition (HAR) systems attempt to automatically identify and analyze human activities using acquired information from various types of sensors. Although several extensive review papers have already been published in the general HAR topics, the growing technologies in the field as well as the multi-disciplinary nature of HAR prompt the need for constant updates in the field. In this respect, this paper attempts to review and summarize the progress of HAR systems from the computer vision perspective. Indeed, most computer vision applications such as human computer interaction, virtual reality, security, video surveillance and home monitoring are highly correlated to HAR tasks. This establishes new trend and milestone in the development cycle of HAR systems. Therefore, the current survey aims to provide the reader with an up to date analysis of vision-based HAR related literature and recent progress in the field. At the same time, it will highlight the main challenges and future directions.
Djamila Romaissa Beddiar, Brahim Nini, Mohammad Sabokrou, Abdenour Hadid
Multim. Tools Appl.3
2019 Unsupervised Feature Ranking and Selection Based on Autoencoders
abstract
Feature selection is one of the most important and widely-used dimension reduction techniques due to its efficiency and intractability of the results. In this paper, we propose a simple but efficient unsupervised feature ranking and selection method by exploiting the geometry of the original feature space using AutoEncoders. Average reconstruction error of training samples by ignoring features, one at time, and the contribution of feature in the latent space (bottleneck of the auto-encoder) are proposed as two useful measures for ranking the features. The proposed method is evaluated for three different tasks: (1) feature selection, (2) discovering image interest points, and (3) extracting important blocks of an images Result on standard benchmarks confirm that the performance of our method is better than state-of-the-art methods.
Sasan Sharifipour, Hossein Fayyazi, Mohammad Sabokrou, Ehsan Adeli-Mosabbeb
ICASSP3
2019 Self-Supervised Representation Learning via Neighborhood-Relational Encoding
abstract
In this paper, we propose a novel self-supervised representation learning by taking advantage of a neighborhood-relational encoding (NRE) among the training data. Conventional unsupervised learning methods only focused on training deep networks to understand the primitive characteristics of the visual data, mainly to be able to reconstruct the data from a latent space. They often neglected the relation among the samples, which can serve as an important metric for self-supervision. Different from the previous work, NRE aims at preserving the local neighborhood structure on the data manifold. Therefore, it is less sensitive to outliers. We integrate our NRE component with an encoder-decoder structure for learning to represent samples considering their local neighborhood information. Such discriminative and unsupervised representation learning scheme is adaptable to different computer vision tasks due to its independence from intense annotation requirements. We evaluate our proposed method for different tasks, including classification, detection, and segmentation based on the learned latent representations. In addition, we adopt the auto-encoding capability of our proposed method for applications like defense against adversarial example attacks and video anomaly detection. Results confirm the performance of our method is better or at least comparable with the state-of-the-art for each specific application, but with a generic and self-supervised approach.
Mohammad Sabokrou, Mohammad Khalooei, Ehsan Adeli-Mosabbeb
ICCV1
2019 End-to-End Adversarial Learning for Intrusion Detection in Computer Networks
abstract
This paper presents a simple yet efficient method for an anomaly-based Intrusion Detection System (IDS). In reality, IDSs can be defined as a one-class classification system, where the normal traffic is the target class. The high diversity of network attacks in addition to the need for generalization, motivate us to propose a semi-supervised method. Inspired by the successes of Generative Adversarial Networks (GANs) for training deep models in semi-unsupervised setting, we have proposed an end-to-end deep architecture for IDS. The proposed architecture is composed of two deep networks, each of which trained by competing with each other to understand the underlying concept of the normal traffic class. The key idea of this paper is to compensate the lack of anomalous traffic by approximately obtain them from normal flows. In this case, our method is not biased towards the available intrusions in the training set leading to more accurate detection. The proposed method has been evaluated on NSL-KDD dataset. The results confirm that our method outperforms the other state-of-the-art approaches.
Bahram Mohammadi, Mohammad Sabokrou
LCN2
2018 AVID: Adversarial Visual Irregularity Detection
Mohammad Sabokrou, Masoud PourReza, Mohsen Fayyaz, Rahim Entezari, Mahmood Fathy, Juergen Gall, Ehsan Adeli-Mosabbeb
ACCV (6)1
2018 Adversarially Learned One-Class Classifier for Novelty Detection
abstract
Novelty detection is the process of identifying the observation(s) that differ in some respect from the training observations (the target class). In reality, the novelty class is often absent during training, poorly sampled or not well defined. Therefore, one-class classifiers can efficiently model such problems. However, due to the unavailability of data from the novelty class, training an end-to-end deep network is a cumbersome task. In this paper, inspired by the success of generative adversarial networks for training deep models in unsupervised and semi-supervised settings, we propose an end-to-end architecture for one-class classification. Our architecture is composed of two deep networks, each of which trained by competing with each other while collaborating to understand the underlying concept in the target class, and then classify the testing samples. One network works as the novelty detector, while the other supports it by enhancing the inlier samples and distorting the outliers. The intuition is that the separability of the enhanced inliers and distorted outliers is much better than deciding on the original samples. The proposed framework applies to different related applications of anomaly and outlier detection in images and videos. The results on MNIST and Caltech-256 image datasets, along with the challenging UCSD Ped2 dataset for video anomaly detection illustrate that our proposed method learns the target class effectively and is superior to the baseline and state-of-the-art methods.
Mohammad Sabokrou, Mohammad Khalooei, Mahmood Fathy, Ehsan Adeli-Mosabbeb
CVPR1
2018 Deep-anomaly: Fully convolutional neural network for fast anomaly detection in crowded scenes
Mohammad Sabokrou, Mohsen Fayyaz, Mahmood Fathy, Zahra Moayed, Reinhard Klette
Comput. Vis. Image Underst.1
2017 Fast and accurate detection and localization of abnormal behavior in crowded scenes
Mohammad Sabokrou, Mahmood Fathy, Zahra Moayed, Reinhard Klette
Mach. Vis. Appl.1
2017 Deep-Cascade: Cascading 3D Deep Neural Networks for Fast Anomaly Detection and Localization in Crowded Scenes
abstract
This paper proposes a fast and reliable method for anomaly detection and localization in video data showing crowded scenes. Time-efficient anomaly localization is an ongoing challenge and subject of this paper. We propose a cubicpatch- based method, characterised by a cascade of classifiers, which makes use of an advanced feature-learning approach. Our cascade of classifiers has two main stages. First, a light but deep 3D auto-encoder is used for early identification of "many" normal cubic patches. This deep network operates on small cubic patches as being the first stage, before carefully resizing remaining candidates of interest, and evaluating those at the second stage using a more complex and deeper 3D convolutional neural network (CNN). We divide the deep autoencoder and the CNN into multiple sub-stages which operate as cascaded classifiers. Shallow layers of the cascaded deep networks (designed as Gaussian classifiers, acting as weak single-class classifiers) detect "simple" normal patches such as background patches, and more complex normal patches are detected at deeper layers. It is shown that the proposed novel technique (a cascade of two cascaded classifiers) performs comparable to current top-performing detection and localization methods on standard benchmarks, but outperforms those in general with respect to required computation time.
Mohammad Sabokrou, Mohsen Fayyaz, Mahmood Fathy, Reinhard Klette
IEEE Trans. Image Process.1