Arash Pourtaherian

dblp:142/0046 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0003-4542-1354ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 1 since 2021Systems, architecture and hardware · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 MEET: Towards Memory-Efficient Temporal Sparse Deep Neural Networks
abstract
Deep Neural Networks (DNNs) are accurate but compute-intensive, leading to substantial energy consumption during inference. Exploiting temporal redundancy through ∆-Σ convolution [26] in video processing has proven to greatly enhance computation efficiency. However, temporal ∆-Σ DNNs typically require substantial memory for storing neuron states to compute inter-frame differences, hindering their on-chip deployment. To mitigate this memory cost, directly compressing the states can disrupt the linearity of temporal ∆-Σ convolution, causing accumulated errors in long-term ∆-Σ processing. Thus, we propose MEET, an optimization framework for MEmory-Efficient Temporal ∆-Σ DNNs. MEET transfers the state compression challenge to a well-established weight compression problem by trading fewer activations for more weights and introduces a co-design of network architecture and suppression method to optimize for mixed spatial-temporal execution. Evaluations on three vision applications demonstrate a reduction of 5.1∼13.3 × in total memory compared to the most computation-efficient temporal DNNs, while preserving the computation efficiency and model accuracy in long-term ∆-Σ processing. MEET facilitates the deployment of temporal ∆-Σ DNNs within on-chip memory of embedded event-driven platforms, empowering low-power edge processing.
Zeqi Zhu, Ibrahim Batuhan Akkaya, Luc Waeijen, Egor Bondarev, Arash Pourtaherian, Orlando Moreira
CVPR5
2024 ELSE: Efficient Deep Neural Network Inference Through Line-Based Sparsity Exploration
Zeqi Zhu, Alberto García Ortiz, Luc Waeijen, Egor Bondarev, Arash Pourtaherian, Orlando Moreira
ECCV (11)5
2024 CATS: Combined Activation and Temporal Suppression for Efficient Network Inference
abstract
Brain-inspired event-driven processors execute deep neural networks (DNNs) in a sparsity-aware manner, leading to superior performance compared to conventional platforms. In the pursuit of higher event sparsity, prior studies suppress non-zero events by either eliminating the intra-frame activations (spatially) or leveraging the redundancy in the inter-frame differences for a video (temporally). However, we have empirically observed that simultaneously enhancing activation and temporal sparsity can lead to a synergistic suppression outcome. To this end, we propose an end-to-end event suppression training approach CATS −− Combined Activation and Temporal Suppression for efficient network inference. It utilizes a gradient-based method to search for the optimal temporal thresholds per layer while penalizing the presence of events in both spatial and temporal domains. Our experimental results show that CATS achieves 2 ∼ 6× higher event suppression compared to the inherent ReLU suppression across a wide range of vision applications, consistently outperforming the state-of-the-art (SOTA) methods by a significant margin at all accuracy levels. Furthermore, a case study on the commercial event-driven processor GrAI-VIP highlights that the induced event sparsity in SSD on the EgoHands dataset can be efficiently translated into a performance enhancement of 2.5× in FPS, 2.1× in latency, and 3.8× in energy consumption, while maintaining the model accuracy.
Zeqi Zhu, Arash Pourtaherian, Luc Waeijen, Ibrahim Batuhan Akkaya, Egor Bondarev, Orlando Moreira
WACV2
2023 NimbleAI: Towards Neuromorphic Sensing-Processing 3D-integrated Chips
abstract
The NimbleAI Horizon Europe project leverages key principles of energy-efficient visual sensing and processing in biological eyes and brains, and harnesses the latest advances in$\mathbf{33D}$stacked silicon integration, to create an integral sensing-processing neuromorphic architecture that efficiently and accurately runs computer vision algorithms in area-constrained endpoint chips. The rationale behind the NimbleAI architecture is: sense data only with high information value and discard data as soon as they are found not to be useful for the application (in a given context). The NimbleAI sensing-processing architecture is to be specialized after-deployment by tunning system-level trade-offs for each particular computer vision algorithm and deployment environment. The objectives of NimbleAI are: (1)$\mathbf{100x}$performance per mW gains compared to state-of-the-practice solutions (i.e., CPU/GPUs processing frame-based video); (2)$\mathbf{50x}$processing latency reduction compared to CPU/GPUs; (3) energy consumption in the order of tens of mWs; and (4) silicon area of approx. 50 mm2.
Xabier Iturbe, Nassim Abderrahmane, Jaume Abella 0001, Sergi Alcaide, Eric Beyne, Henri-Pierre Charles, Christelle Charpin-Nicolle, Lars Chittka, Angélica Dávila, Arne Erdmann, Carles Estrada, Ander Fernández, Anna Fontanelli, José Flich, Gianluca Furano, Alejandro Hernán Gloriani, Erik Isusquiza, Radu Grosu, Carles Hernández 0001, Daniele Ielmini, Maha Kooli, Nicola Lepri, Bernabé Linares-Barranco, Jean-Loup Lachese, Eric Laurent, Menno Lindwer, Frank Linsenmaier, Mikel Luján, Karel Masarík, Nele Mentens, Orlando Moreira, Chinmay Nawghane, Luca Peres, Jean-Philippe Noël, Arash Pourtaherian, Christoph Posch, Peter Priller, Zdenek Prikryl, Felix Resch, Oliver Rhodes, Todor P. Stefanov, Moritz Storring, Michele Taliercio, Rafael Tornero, Marcel D. van de Burgwal, Geert Van der Plas, Elisa Vianello, Pavel Zaykov
DATE36
2023 Synapse Compression for Event-Based Convolutional-Neural-Network Accelerators
abstract
Manufacturing-viable neuromorphic chips require novel compute architectures to achieve the massively parallel and efficient information processing the brain supports so effortlessly. The most promising architectures for that are spiking/event-based, which enables massive parallelism at low complexity. However, the large memory requirements for synaptic connectivity are a showstopper for the execution of modern convolutional neural networks (CNNs) on massively parallel, event-based architectures. The present work overcomes this roadblock by contributing a lightweight hardware scheme to compress the synaptic memory requirements by several thousand times—enabling the execution of complex CNNs on a single chip of small form factor. A silicon implementation in a 12-nm technology shows that the technique achieves a total memory-footprint reduction of up to 374× compared to the best previously published technique at a negligible area overhead.
Lennart Bamberg, Arash Pourtaherian, Luc Waeijen, Anupam Chahar, Orlando Moreira
IEEE Trans. Parallel Distributed Syst.2
2022 ARTS: An adaptive regularization training schedule for activation sparsity exploration
abstract
Brain-inspired event-based processors have attracted considerable attention for edge deployment because of their ability to efficiently process Convolutional Neural Networks (CNNs) by exploiting sparsity. On such processors, one critical feature is that the speed and energy consumption of CNN inference are approximately proportional to the number of non-zero values in the activation maps. Thus, to achieve top performance, an efficient training algorithm is required to largely suppress the activations in CNNs. We propose a novel training method, called Adaptive-Regularization Training Schedule (ARTS), which dramatically decreases the non-zero activations in a model by adaptively altering the regularization coefficient through training. We evaluate our method across an extensive range of computer vision applications, including image classification, object recognition, depth estimation, and semantic segmentation. The results show that our technique can achieve 1.41 × to 6.00 × more activation suppression on top of ReLU activation across various networks and applications, and outperforms the state-of-the-art methods in terms of training time, activation suppression gains, and accuracy. A case study for a commercially-available event-based processor, Neuronflow, shows that the activation suppression achieved by ARTS effectively reduces CNN inference latency by up to 8.4 × and energy consumption by up to 14.1 ×.
Zeqi Zhu, Arash Pourtaherian, Luc Waeijen, Lennart Bamberg, Egor Bondarev, Orlando Moreira
DSD2
2021 How to exploit sparsity in RNNs on event-driven architectures
abstract
Event-driven architectures have been shown to provide low-power, low-latency artificial neural network (ANN) inference. This is especially beneficial on Edge devices, particularly when combined with sparse execution. Recurrent neural networks (RNNs) are ANNs that emulate memory. Their recurrent connection enables the reuse of previous output for the generation of new output. However, when trying to use RNNs in a sparse context on event-driven architectures, novel challenges in synchronization and the usage of sparse data are encountered. In this work, these challenges are systematically analyzed, and mechanisms to overcome them are proposed. Experimental results of a monocular depth estimation use case on the NeuronFlow architecture show that sparsity in RNNs can be exploited effectively on event-driven architectures.
Jarno Brils, Luc Waeijen, Arash Pourtaherian
SCOPES3
2021 Mask-MCNet: Tooth instance segmentation in 3D point clouds of intra-oral scans
abstract
Computational dentistry uses computerized methods and mathematical models for dental image analysis. One of the fundamental problems in computational dentistry is accurate tooth instance segmentation in high-resolution mesh data of intra-oral scans (IOS). This paper presents a new computational model based on deep neural networks, called Mask-MCNet, for end-to-end learning of tooth instance segmentation in 3D point cloud data of IOS. The proposed Mask-MCNet localizes each tooth instance by predicting its 3D bounding box and simultaneously segments the points that belong to each individual tooth instance. The proposed model processes the input raw 3D point cloud in its original spatial resolution without employing a voxelization or down-sampling technique. Such a characteristic preserves the finely detailed context in data like fine curvatures in the border between adjacent teeth and leads to a highly accurate segmentation as required for clinical practice (e.g. orthodontic planning). The experiments show that the Mask-MCNet outperforms state-of-the-art models by achieving 98% Intersection over Union (IoU) score on tooth instance segmentation which is very close to human expert performance.
Farhad G. Zanjani, Arash Pourtaherian, Svitlana Zinger, David Anssari Moin, Frank Claessen, Teo Cherici, Sarah Parinussa, Peter H. N. de With
Neurocomputing2
2021 Infant Facial Expression Analysis: Towards a Real-Time Video Monitoring System Using R-CNN and HMM
abstract
The manual monitoring of young infants suffering from diseases like reflux is significant, since infants can hardly articulate their feelings. In this work, we propose a video-based infant monitoring system for the analysis of infant expressions and states, approaching real-time performance. The expressions of interest consist of discomfort, unhappy, joy and neutral, whereas states include sleep, pacifier and open mouth. Benefiting from the expression analysis, the discomfort moments can also be used and correlated with a symptom-related disease, such as a reflux measurement for the diagnosis of gastroesophageal reflux. The system consists of three components: infant expressions and states detection, object tracking and detection compensation. The proposed system is based on combining expression detection using Fast R-CNN with a compensated detection using analyzing information from the previous frame and utilizing a Hidden Markov Model. The experimental results show a mean average precision of 81.9% and 84.8% for 4 infant expressions and 3 states evaluated with both clinical and daily datasets. Meanwhile, the average precision for discomfort detection achieves up to 90%.
Cheng Li 0042, Arash Pourtaherian, Lonneke van Onzenoort, Walther E. Tjon a Ten, Peter H. N. de With
IEEE J. Biomed. Health Informatics2
2019 Mask-MCNet: Instance Segmentation in 3D Point Cloud of Intra-oral Scans
Farhad G. Zanjani, David Anssari Moin, Frank Claessen, Teo Cherici, Sarah Parinussa, Arash Pourtaherian, Svitlana Zinger, Peter H. N. de With
MICCAI (5)6
2019 Video-based discomfort detection for infants
Yue Sun 0001, Caifeng Shan, Tao Tan 0002, Xi Long 0001, Arash Pourtaherian, Svitlana Zinger, Peter H. N. de With
Mach. Vis. Appl.5
2018 Enhanced face alignment using an unsupervised roll estimation initialization
abstract
We propose a novel and efficient initialization method for generalized facial landmark localization with an unsupervised roll-angle estimation based on B-spline models. We first show that the roll angle is crucial for an accurate landmark localization. Therefore, we develop an unsupervised roll-angle estimation by adopting a joint 1st -order B-spline model, which is robust to intensity variations and generic for application to various face detectors. The method consists of three steps. First, the scaled-normalized Laplacian of Gaussian operator is applied to a bounding box generated by a face detector for extracting facial feature segments. Second, a joint 1 st -order B-spline model is fitted to the extracted facial feature segments, using an iterative optimization method. Finally, the roll angle is estimated through the aligned segments. We evaluate four state-of-the-art landmark localization schemes with the proposed roll-angle estimation initialization in the benchmark dataset. The proposed method boosts the performance of landmark localization in general, especially for cases with large head pose. Moreover, the proposed unsupervised roll-angle estimation method outperforms the standard supervised methods, such as random forest and support vector regression by 41.6% and 47.2%, respectively.
Cheng Li 0042, Arash Pourtaherian, Walther E. Tjon a Ten, Peter H. N. de With
ICMV2
2017 Improving Needle Detection in 3D Ultrasound Using Orthogonal-Plane Convolutional Networks
Arash Pourtaherian, Farhad G. Zanjani, Svitlana Zinger, Nenad Mihajlovic, Gary C. Ng, Hendrikus H. M. Korsten, Peter H. N. de With
MICCAI (2)1
2017 Medical Instrument Detection in 3-Dimensional Ultrasound Data Volumes
abstract
Ultrasound-guided medical interventions are broadly applied in diagnostics and therapy, e.g., regional anesthesia or ablation. A guided intervention using 2-D ultrasound is challenging due to the poor instrument visibility, limited field of view, and the multi-fold coordination of the medical instrument and ultrasound plane. Recent 3-D ultrasound transducers can improve the quality of the image-guided intervention if an automated detection of the needle is used. In this paper, we present a novel method for detecting medical instruments in 3-D ultrasound data that is solely based on image processing techniques and validated on various ex vivo and in vivo data sets. In the proposed procedure, the physician is placing the 3-D transducer at the desired position, and the image processing will automatically detect the best instrument view, so that the physician can entirely focus on the intervention. Our method is based on the classification of instrument voxels using volumetric structure directions and robust approximation of the primary tool axis. A novel normalization method is proposed for the shape and intensity consistency of instruments to improve the detection. Moreover, a novel 3-D Gabor wavelet transformation is introduced and optimally designed for revealing the instrument voxels in the volume, while remaining generic to several medical instruments and transducer types. Experiments on diverse data sets, including in vivo data from patients, show that for a given transducer and an instrument type, high detection accuracies are achieved with position errors smaller than the instrument diameter in the 0.5-1.5-mm range on average.
Arash Pourtaherian, Harm J. Scholten, Lieneke Kusters, Svitlana Zinger, Nenad Mihajlovic, Alexander F. Kolen, Fei Zuo, Gary C. Ng, Hendrikus H. M. Korsten, Peter H. N. de With
IEEE Trans. Medical Imaging1
2015 Multi-resolution Gabor wavelet feature extraction for needle detection in 3D ultrasound
abstract
Ultrasound imaging is employed for needle guidance in various minimally invasive procedures such as biopsy guidance, regional anesthesia and brachytherapy. Unfortunately, a needle guidance using 2D ultrasound is very challenging, due to a poor needle visibility and a limited field of view. Nowadays, 3D ultrasound systems are available and more widely used. Consequently, with an appropriate 3D image-based needle detection technique, needle guidance and interventions may significantly be improved and simplified. In this paper, we present a multi-resolution Gabor transformation for an automated and reliable extraction of the needle-like structures in a 3D ultrasound volume. We study and identify the best combination of the Gabor wavelet frequencies. High precision in detecting the needle voxels leads to a robust and accurate localization of the needle for the intervention support. Evaluation in several ex-vivo cases shows that the multi-resolution analysis significantly improves the precision of the needle voxel detection from 0.23 to 0.32 at a high recall rate of 0.75 (gain 40%), where a better robustness and confidence were confirmed in the practical experiments.
Arash Pourtaherian, Svitlana Zinger, Nenad Mihajlovic, Peter H. N. de With, Gary C. Ng, Hendrikus H. M. Korsten
ICMV1
2014 Gabor-based needle detection and tracking in three-dimensional ultrasound data volumes
abstract
During needle interventions for e.g. regional anaesthesia or biopsy, it is very important to visualize the needle and its tip with respect to important structures in the body. In this work, we propose a novel image-based needle detection technique in a 3D ultrasound volume dataset, which can improve the intervention. We present a novel application of the 3D Gabor transformation, which exploits needle-like structures with appropriate designs. Furthermore, we introduce a needle tracking algorithm based on Gradient Descent and show that it limits the computational complexity and detection error. Finally, we visualize the needle on 2D cross-sections of the volume in order to be presented to the physician. Evaluation of our system in challenging cases shows a high detection score (up to 100% but needs larger sets) and accurate visualization.
Arash Pourtaherian, Svitlana Zinger, Peter H. N. de With, Hendrikus H. M. Korsten, Nenad Mihajlovic
ICIP1
2013 TROD: Tracking with occlusion handling and drift correction
abstract
We present a tracking framework in which we learn a HOG-based object detector in the first video frame and use this detector to localize the object in subsequent frames. We contribute and improve the tracking on the three following points. First, an occlusion-handling algorithm exploits discriminative information from the detector by dividing the object bounding box into patches and comparing each patch to the object model. Second, a drift-correction technique uses descriptive information of the object by calculating the similarity between the object in the previous frame and its shifted versions in the current frame. Third, a stochastic learning algorithm updates the object detector using single object and single background samples for selected frames only. Experiments with benchmark sequences show that the proposed tracker outperforms state-of-the-art methods on several sequences and has the smallest average location error.
Arash Pourtaherian, Rob G. J. Wijnhoven, Peter H. N. de With
ICIP1