Hamed Tabkhi

dblp:47/3258 · DBLP profile ↗
← Back
40ranked-venue papers
9as first author
16since 2021 · last 2026
0000-0001-5420-1121ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 26 · 9 first-author · 5 since 2021Computer networks · 5 · 2 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 From Offline to Periodic Adaptation for Pose-Based Shoplifting Detection in Real-World Retail Security
abstract
Shoplifting is a growing operational and economic challenge for retailers, with incidents rising and losses increasing despite extensive video surveillance. Continuous human monitoring is infeasible, motivating automated, privacy-preserving, and resource-aware detection solutions. In this paper, we cast shoplifting detection as a pose-based, unsupervised video anomaly detection problem and introduce a periodic adaptation framework designed for on-site Internet of Things (IoT) deployment. Our approach enables edge devices in smart retail environments to adapt from streaming, unlabeled data, supporting scalable and low-latency anomaly detection across distributed camera networks. To support reproducibility, we introduceRetailS1, a new large-scale real-world shoplifting dataset collected from a retail store under multi-day, multi-camera conditions, capturing unbiased shoplifting behavior in realistic IoT settings. For deployable operation, thresholds are selected using both F1 andHPRSscores, the harmonic mean of precision, recall, and specificity, during data filtering and training. In periodic adaptation experiments, our framework consistently outperformed offline baselines on AUC-ROC and AUC-PR in 91.6% of evaluations, with each training update completing in under 30 minutes on edge-grade hardware, demonstrating the feasibility and reliability of our solution for IoT-enabled smart retail deployment.
Shanle Yao, Narges Rashvand, Armin Danesh Pazho, Hamed Tabkhi
IEEE Internet Things J.4
2026 Toward Adaptive Human-Centric Video Anomaly Detection: A Comprehensive Framework and a New Benchmark
abstract
Human-centric Video Anomaly Detection (VAD) aims to identify human behaviors that deviate from normal. At its core, human-centric VAD faces substantial challenges, such as the complexity of diverse human behaviors, the rarity of anomalies, and ethical constraints. These challenges limit access to high-quality datasets and highlight the need for a dataset and framework supporting continual learning. Moving towards adaptive human-centric VAD, we introduce the HuVAD (Human-centric privacy-enhanced Video Anomaly Detection) dataset and a novel Unsupervised Continual Anomaly Learning (UCAL) framework. UCAL enables incremental learning, allowing models to adapt over time, bridging traditional training and real-world deployment. HuVAD prioritizes privacy by providing deidentified annotations and includes seven indoor/outdoor scenes, offering over 5× more pose-annotated frames than previous datasets. Our standard and continual benchmarks, utilize a comprehensive set of metrics, demonstrating that UCAL-enhanced models achieve superior performance in 83.33% of cases, setting a new state-of-the-art (SOTA). The dataset can be accessed at https://github.com/TeCSAR-UNCC/HuVAD.
Armin Danesh Pazho, Shanle Yao, Ghazal Alinezhad Noghre, Babak Rahimi Ardabili, Vinit Katariya, Hamed Tabkhi
IEEE Trans. Circuits Syst. Video Technol.6
2025 Adversarially-Refined VQ-GAN with Dense Motion Tokenization for Spatio-Temporal Heatmaps
abstract
Continuous human motion understanding remains a core challenge in computer vision due to its high dimensionality and inherent redundancy. Efficient compression and representation are crucial for analyzing complex motion dynamics. In this work, we introduce an adversarially-refined VQ-GAN framework with dense motion tokenization for compressing spatio-temporal heatmaps while preserving the fine-grained traces of human motion. Our approach combines dense motion tokenization with adversarial refinement, which eliminates reconstruction artifacts like motion smearing and temporal misalignment observed in non-adversarial baselines. Our experiments on the CMU Panoptic dataset [7] provide conclusive evidence of our method’s superiority, outperforming the dVAE baseline by 9.31% SSIM and reducing temporal instability by 37.1%. Furthermore, our dense tokenization strategy enables a novel analysis of motion complexity, revealing that 2D motion can be optimally represented with a compact 128-token vocabulary, while 3D motion’s complexity demands a much larger 1024-token codebook for faithful reconstruction. These results establish practical deployment feasibility across diverse motion analysis applications. The code base for this work is available at https://github.com/TeCSAR-UNCC/Pose-Quantization.
Gabriel Maldonado, Narges Rashvand, Armin Danesh Pazho, Ghazal Alinezhad Noghre, Vinit Katariya, Hamed Tabkhi
ICMLA6
2024 Real-Time Bus Arrival Prediction: A Deep Learning Approach for Enhanced Urban Mobility
Narges Rashvand, Sanaz Sadat Hosseini, Mona Azarbayjani, Hamed Tabkhi
ICORES4
2024 Expanding hardware accelerator system design space exploration with gem5-SALAMv2
Zephaniah Spencer, Samuel Rogers, Joshua Slycord, Hamed Tabkhi
J. Syst. Archit.4
2024 A POV-Based Highway Vehicle Trajectory Dataset and Prediction Architecture
abstract
Vehicle Trajectory datasets that provide multiple point-of-views (POVs) can be valuable for various traffic safety and management applications. Despite the abundance of trajectory datasets, few offer a comprehensive and diverse range of driving scenes, capturing multiple viewpoints of various highway layouts, merging lanes, and configurations. This limits their ability to capture the nuanced interactions between drivers, vehicles, and the roadway infrastructure. We introduce the Carolinas Highway Dataset (CHD) (CHD available at:https://github.com/TeCSAR-UNCC/Carolinas_Dataset), a vehicle trajectory, detection, and tracking dataset. CHD is a collection of 1.6 million frames captured in highway-based videos from eye-level and high-angle POVs at eight locations across Carolinas with 338,000 vehicle trajectories. The locations, timing of recordings, and camera angles were carefully selected to capture various road geometries, traffic patterns, lighting conditions, and driving behaviors. We also present PishguVe (PishguVe code available at:https://github.com/TeCSAR-UNCC/PishguVe), a novel vehicle trajectory prediction architecture that uses attention-based graph isomorphism and convolutional neural networks. The results demonstrate that PishguVe outperforms existing algorithms with better ADE and FDE in eye-level, and high-angle POV trajectory datasets. Compared to best-performing models on CHD, PishguVe achieves lower ADE and FDE on eye-level data by 14.58% and 27.38%, respectively, and improves ADE and FDE on high-angle data by 8.3% and 6.9%, respectively.
Vinit Katariya, Ghazal Alinezhad Noghre, Armin Danesh Pazho, Hamed Tabkhi
IEEE Trans. Intell. Transp. Syst.4
2024 A Survey of Graph-Based Deep Learning for Anomaly Detection in Distributed Systems
abstract
Anomaly detection is a crucial task in complex distributed systems. A thorough understanding of the requirements and challenges of anomaly detection is pivotal to the security of such systems, especially for real-world deployment. While there are many works and application domains that deal with this problem, few have attempted to provide an in-depth look at such systems. In this survey, we explore the potentials of graph-based algorithms to identify anomalies in distributed systems. These systems can be heterogeneous or homogeneous, which can result in distinct requirements. One of our objectives is to provide an in-depth look at graph-based approaches to conceptually analyze their capability to handle real-world challenges such as heterogeneity and dynamic structure. This study gives an overview of the State-of-the-Art (SotA) research articles in the field and compare and contrast their characteristics. To facilitate a more comprehensive understanding, we present three systems with varying abstractions as use cases. We examine the specific challenges involved in anomaly detection within such systems. Subsequently, we elucidate the efficacy of graphs in such systems and explicate their advantages. We then delve into the SotA methods and highlight their strength and weaknesses, pointing out the areas for possible improvements and future works.
Armin Danesh Pazho, Ghazal Alinezhad Noghre, Arnab A. Purkayastha, Jagannadh Vempati, Martin Otto 0002, Hamed Tabkhi
IEEE Trans. Knowl. Data Eng.6
2023 Real-World Community-in-the-Loop Smart Video Surveillance System
abstract
In recent years, smart video surveillance (SVS) systems have become essential in maintaining public safety and security, particularly in smart city environments. We propose an SVS system that uses advanced technologies such as artificial intelligence and computer vision to ensure the timely detection of anomalous behaviors and suspicious objects. The system’s performance is demonstrated through a smartphone application and real-world scenario videos, highlighting its effectiveness in enhancing citizen security with low latency. This paper represents a demonstration of such a system for implementing community-in-the-loop smart video surveillance systems and emphasizes their practicality in improving public safety in various settings. The study adds to the growing research on deploying smart video surveillance systems and underscores the importance of engaging local communities in these projects.
Shanle Yao, Babak Rahimi Ardabili, Armin Danesh Pazho, Ghazal Alinezhad Noghre, Christopher Neff, Hamed Tabkhi
SMARTCOMP6
2023 Ancilia: Scalable Intelligent Video Surveillance for the Artificial Intelligence of Things
abstract
With the advancement of vision-based artificial intelligence, the proliferation of the Internet of Things connected cameras, and the increasing societal need for rapid and equitable security, the demand for accurate real-time intelligent surveillance has never been higher. This article presents Ancilia, an end-to-end scalable, intelligent video surveillance system for the Artificial Intelligence of Things. Ancilia brings state-of-the-art artificial intelligence to real-world surveillance applications while respecting ethical concerns and performing high-level cognitive tasks in real time. Ancilia aims to revolutionize the surveillance landscape, to bring more effective, intelligent, and equitable security to the field, resulting in safer and more secure communities without requiring people to compromise their right to privacy.
Armin Danesh Pazho, Christopher Neff, Ghazal Alinezhad Noghre, Babak Rahimi Ardabili, Shanle Yao, Mohammadreza Baharani, Hamed Tabkhi
IEEE Internet Things J.7
2022 ATCN: Resource-efficient Processing of Time Series on Edge
abstract
This article presents a scalable deep learning model called Agile Temporal Convolutional Network (ATCN) for highly accurate fast classification and time series prediction in resource-constrained embedded systems. ATCN is a family of compact networks with formalized hyperparameters that enable application-specific adjustments to be made to the model architecture. It is primarily designed for embedded edge devices with very limited performance and memory, such as wearable biomedical devices and real-time reliability monitoring systems. ATCN makes fundamental improvements over the mainstream temporal convolutional neural networks, including residual connections to increase the network depth and accuracy and the incorporation of separable depth-wise convolution to reduce the computational complexity of the model. As part of the present work, two ATCN families, namely T0 and T1, are also presented and evaluated on different ranges of embedded processors: Cortex-M7 and Cortex-A57 processors. An evaluation of the ATCN models against the best-in-class InceptionTime and MiniRocket shows that ATCN almost maintains accuracy while improving the execution time on a broad range of embedded and cyber-physical applications with demand for real-time processing on the embedded edge. At the same time, in contrast to existing solutions, ATCN is the first time series classifier based on deep learning that can be run bare-metal on embedded microcontrollers (Cortex-M7) with limited computational performance and memory capacity while delivering state-of-the-art accuracy.
Mohammadreza Baharani, Hamed Tabkhi
ACM Trans. Embed. Comput. Syst.2
2022 DeepTrack: Lightweight Deep Learning for Vehicle Trajectory Prediction in Highways
abstract
Vehicle trajectory prediction is essential for enabling safety-critical intelligent transportation systems (ITS) applications used in management and operations. While there have been some promising advances in the field, there is a need for modern deep learning algorithms that allow real-time trajectory prediction on embedded IoT devices. This article presents DeepTrack, a novel deep learning algorithm customized for real-time vehicle trajectory prediction and monitoring applications in arterial management, freeway management, traffic incident management, and work zone management for high-speed incoming traffic. In contrast to previous methods, the vehicle dynamics are encoded using Temporal Convolutional Networks (TCNs) to provide more robust time prediction with less computation. DeepTrack also uses depthwise convolution, which reduces the complexity of models compared to existing approaches in terms of model size and operations. Overall, our experimental results demonstrate that DeepTrack achieves comparable accuracy to state-of-the-art trajectory prediction models but with smaller model sizes and lower computational complexity, making it more suitable for real-world deployment.
Vinit Katariya, Mohammadreza Baharani, Nichole Morris, Omidreza Shoghli, Hamed Tabkhi
IEEE Trans. Intell. Transp. Syst.5
2021 CARPe Posterum: A Convolutional Approach for Real-Time Pedestrian Path Prediction
abstract
Pedestrian path prediction is an essential topic in computer vision and video understanding. Having insight into the movement of pedestrians is crucial for ensuring safe operation in a variety of applications including autonomous vehicles, social robots, and environmental monitoring. Current works in this area utilize complex generative or recurrent methods to capture many possible futures. However, despite the inherent real-time nature of predicting future paths, little work has been done to explore accurate and computationally efficient approaches for this task. To this end, we propose a convolutional approach for real-time pedestrian path prediction, CARPe. It utilizes a variation of Graph Isomorphism Networks in combination with an agile convolutional neural network design to form a fast and accurate path prediction approach. Notable results in both inference speed and prediction accuracy are achieved, improving FPS considerably in comparison to current state-of-the-art methods while delivering competitive accuracy on well-known path prediction datasets.
Matías Mendieta, Hamed Tabkhi
AAAI2
2021 MG-DmDSE: Multi-Granularity Domain Design Space Exploration Considering Function Similarity
abstract
Heterogeneous accelerator-rich (ACC-rich) platforms combining general-purpose cores and specialized HW accelerators (ACCs) promise high-performance and low-power streaming application deployments in a variety of domains such as video analytics and software-defined radio. In order to benefit a domain of applications, a domain platform exploration tool must take advantage of structural and functional similarities across applications by allocating a common set of ACCs. A previous approach [1] proposed a GenetIc Domain Exploration tool (GIDE) that applied a restrictive binding algorithm that mapped applications functions to monolithic accelerators. This approach suffered from lower average application throughput across and reduced platform generality. This paper introduces a Multi-Granularity based Domain Design Space Exploration tool (MG-DmDSE) to improve both average application throughput as well as platform generality. The key contributions of MG-DmDSE are: (1) Applying a multi-granular decomposition of coarse grain application functions into more granular compute kernels. (2) Examining compute similarity between functions in order to produce more generic functions. (3) Configuring monolithic ACCs by selectively bypassing compute elements within them during DSE to expose more functionality. To assess MG-DmDSE, both GIDE and MG-DmDSE were applied to applications in the OpenVX library. MG-DmDSE achieves an average 2.84x greater application throughput compared to GIDE. Additionally, 87.5% of applications benefited from running on the platform produced by MG-DmDSE vs 50% from GIDE, which indicated increase platform generality.
Jinghan Zhang 0001, Aly Sultan, Hamed Tabkhi, Gunar Schirner
DATE3
2021 DeepDive: An Integrative Algorithm/Architecture Co-Design for Deep Separable Convolutional Neural Networks
abstract
Deep Separable Convolutional Neural Network (DSCNN) has become the emerging paradigm by offering modular networks with structural sparsity to achieve higher accuracy with relatively lower operations and parameters. However, there is a lack of customized architectures that can provide flexible solutions that fit the sparsity of the DSCNNs. This paper introduces DeepDive, a fully-functional vertical co-design framework, for power-efficient implementation of DSCNNs on edge FPGAs. DeepDive's architecture supports crucial heterogeneous Compute Units (CUs) to fully support DSCNNs with various convolutional operators interconnected with structural sparsity. It offers FPGA-aware training and online quantization combined with modular synthesizable C++ CUs, customized for DSCNNs. The execution results on Xilinx's ZCU102 FPGA board demonstrate 47.4 and 233.3 FPS/Watt for MobileNet-V2 and a compact version of EfficientNet, respectively, as two state-of-the-art depthwise separable CNNs. These comparisons showcase how DeepDive improves FPS/Watt by 2.2× and 1.51× over Jetson Nano high and low power modes, respectively. It also enhances FPS/Watt by about 2.27× and 37.25× over two other FPGA implementations.
Mohammadreza Baharani, Ushma Sunil, Kaustubh Manohar, Steven Furgurson, Hamed Tabkhi
ACM Great Lakes Symposium on VLSI5
2021 Real-World Graph Convolution Networks (RW-GCNs) for Action Recognition in Smart Video Surveillance
Justin Sanchez, Christopher Neff, Hamed Tabkhi
SEC3
2021 Toward AI-enabled augmented reality to enhance the safety of highway work zones: Feasibility, requirements, and challenges
Sepehr Sabeti, Omidreza Shoghli, Mohammadreza Baharani, Hamed Tabkhi
Adv. Eng. Informatics4
2020 gem5-SALAM: A System Architecture for LLVM-based Accelerator Modeling
abstract
With the prevalence of hardware accelerators as an integral part of the modern systems on chip (SoCs), the ability to quickly and accurately model accelerators within the system it operates is critical. This paper presents gem5-SALAM as a novel system architecture for LLVM-based modeling and simulation of custom hardware accelerators integrated into the gem5 framework. gem5-SALAM overcomes the inherent limitations of state-of-the-art trace-based pre-register-transfer level (RTL) simulators by offering a truly "execute-in-execute" LLVM-based model. It enables scalable modeling of multiple dynamically interacting accelerators with full-system simulation support. To create sustainable long-term expansion compatible with the gem5 system framework, gem5-SALAM offers a general-purpose and modular communication interface and memory hierarchy integrated into the gem5 ecosystem which streamlines designing and modeling accelerators for new and emerging applications. Validation on the MachSuite [17] benchmarks present a timing estimation error of less than 1% against Vivado High-Level Synthesis (HLS) tool. Results also show less than a 4% area and power estimation error against Synopsys Design Compiler. Additionally, system validation against implementations on a Ultrascale+ ZCU102 shows an average end-to-end timing error of less than 2%. Lastly, this paper presents the capabilities of gem5-SALAM in cycle-level profiling and full system design space exploration of accelerator-rich systems.
Samuel Rogers, Joshua Slycord, Mohammadreza Baharani, Hamed Tabkhi
MICRO4
2020 REVAMP2T: Real-Time Edge Video Analytics for Multicamera Privacy-Aware Pedestrian Tracking
abstract
This article presents real-time edge video analytics for multicamera privacy-aware pedestrian tracking (REVAMP2T), as an integrated end-to-end Internet of Things (IoT) system for privacy built-in decentralized situational awareness. REVAMP2T presents novel algorithmic and system constructs to push deep learning and video analytics next to IoT devices (i.e., video cameras). On the algorithm side, REVAMP2T proposes a unified integrated computer vision pipeline for detection, reidentification, and tracking across multiple cameras without the need for storing the streaming data. At the same time, it avoids facial recognition and tracks and reidentifies the pedestrians based on their key features at runtime. On the IoT system side, REVAMP2T provides an infrastructure to maximize the hardware utilization on the edge, orchestrates global communications, and provides system-wide reidentification, without the use of personally identifiable information, for a distributed IoT network. For the results and evaluation, this article also proposes a new metric, accuracy-efficiency (Æ), for holistic evaluation of IoT systems for real-time video analytics based on accuracy, performance, and power efficiency. REVAMP2T outperforms the current state of the art by as much as 13-fold Æ improvement.
Christopher Neff, Matías Mendieta, Shrey Mohan, Mohammadreza Baharani, Samuel Rogers, Hamed Tabkhi
IEEE Internet Things J.6
2020 AWARE-CNN: Automated Workflow for Application-Aware Real-Time Edge Acceleration of CNNs
abstract
This article presents the application-aware real-time edge acceleration of CNNs (AWARE-CNNs) accelerators, which is a novel architecture design methodology for real-time execution of deep learning algorithms on IoT devices. AWARE leverages the reconfigurability of field-programmable gate arrays (FPGAs) to create application-specific architectures customized to match the inherent dataflow of targeted deep neural networks and user-specified real-time requirements. The customized datapath is combined with a customized memory path to guarantee deterministic latency-aware execution over streaming data. For results and evaluation, we have developed a Chisel-based implementation of AWARE-CNN with a full integrative framework for application-specific architecture generation and synthesis (AWARE-CNN architecture compiler). Our results demonstrate the ability to execute Tiny DarkNet and shallow MobileNet inference at 120 frames/s (FPS) and 75 FPS, using only 2.8 and 3.4 W, respectively, on a Xilinx XCZU9EG FPGA. In addition, AWARE-CNN framework's flexibility with respect to the targeted convolutional neural networks and user constraints is validated by targeting additional design points for AlexNet (as a baseline network) and Tiny YOLOv2.
Justin Sanchez, Adarsh Sawant, Christopher Neff, Hamed Tabkhi
IEEE Internet Things J.4
2020 Allocating One Common ACC-Rich Platform for Many Streaming Applications
abstract
Many demanding streaming applications share functional and structural similarities with other apps in their respective domain, e.g., video analytics, software-defined radio, and radar. This opens the opportunity for specialization (e.g., heterogeneous computing) to achieve the needed efficiency and/or performance. However, current design space exploration (DSE) focuses on an individual application in isolation (e.g., one particular vision flow), but not a set of similar applications. Hence, optimizations that occur due to considering multiple applications simultaneously are missed. New DSE methodologies and tools are needed with a broader scope of application sets instead of individual applications. This article introduces a novel domain-specific DSE (DS-DSE) approach focusing on streaming applications. Key contributions are: 1) a formalized method to extract the functional and structural similarities of domain applications; 2) a rapid platform performance estimation and comparison at two abstraction levels: domain score (DS) and analytic performance estimation (APE) model; 3) two novel algorithms, dynamic score selection (DSS), and GenetIc domain exploration (GIDE), for hardware/software partitioning of a domain-specific platform to maximize the throughput across domain applications (under certain constraints); and 4) a methodology to evaluate a platform's benefit for a set of applications. We demonstrate DSS's and GIDE's benefits using OpenVX applications and synthetic domains. The DSS and GIDE generated domain-specific platforms improve performance over application-specific platforms by 58% and 75% for OpenVX, as well as by 23% and 48% for synthetic applications. GIDE's platforms reach 99.8% (OpenVX) and 97.6% (synthetic) throughput of the domain optimal platform obtained through exhaustive search.
Jinghan Zhang 0001, Hamed Tabkhi, Gunar Schirner
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2019 Mitigating Application Diversity for Allocating a Unified ACC-Rich Platform
abstract
Heterogeneous accelerator-rich (ACC-rich) platforms combining general-purpose cores and specialized HW accelerators (ACCs) promise high-performance and low-power streaming application (app) deployments, e.g. for video analytics, software-defined radio, and radar. In order to recover NRE, a unified platform for a set of applications (apps) is desirable. When apps have functional and structural similarities, they can benefit from common ACCs. Identifying the most beneficial set of common ACCs is challenging. However, current allocation strategies mostly focus on one app in isolation. Automatically allocating a unified platform requires simultaneously considering many apps, an efficient design space traversal and a fair evaluation across diverse apps. This paper introduces a Unified ACC-rich Platform Allocation (UPA) methodology for sets of data flow apps. Key contributions are: (1) a genetic algorithm (GA) guided by a fair and efficient evaluation to allocate one unified platform for many apps, (2) defining relative efficiency for fair comparison across diverse apps, and (3) defining metrics to quantify many app platform efficiency. This paper demonstrates UPA's benefits using OpenVX apps. A 12-ACCs-UPA improves average efficiency 4.59x over app-dedicated platforms. The UPA platform enables more apps (55% of OpenVX apps) to be efficiently deployed (≥ 60% of optimal app-dedicated platform). The benefits increase even further with increasing ACC budget.
Jinghan Zhang 0001, Hamed Tabkhi, Gunar Schirner
ICCD2
2019 Real-Time Deep Learning at the Edge for Scalable Reliability Modeling of Si-MOSFET Power Electronics Converters
abstract
With the significant growth of advanced high-frequency power converters, online monitoring and active reliability assessment of power electronic devices are extremely crucial. This paper presents a transformative approach, named deep learning reliability awareness of converters at the edge (Deep RACE), for real-time reliability modeling and prediction of high-frequency MOSFET power electronic converters. Deep RACE offers a holistic solution which comprises algorithm advances, and full system integration (from the cloud down to the edge node) to create a near real-time reliability awareness. On the algorithm side, this paper proposes a deep learning algorithmic solution based on stacked long short-term memory for collective reliability training and inference across collective MOSFET converters based on device resistance changes. Deep RACE also proposes an integrative edge-to-cloud solution to offer a scalable decentralized devices-specific reliability monitoring, awareness, and modeling. The MOSFET convertors are Internet-of-Things (IoT) devices which have been empowered with edge real-time deep learning processing capabilities. The proposed Deep RACE solution has been prototyped and implemented through learning from MOSFET data set provided by NASA. Our experimental results show an average miss prediction of 8.9% over five different devices which is a much higher accuracy compared to well-known classical approaches (Kalman filter and particle filter). Deep RACE only requires 26-ms processing time and 1.87-W computing power on edge IoT device.
Mohammadreza Baharani, Mehrdad Biglarbegian, Babak Parkhideh, Hamed Tabkhi
IEEE Internet Things J.4
2019 Alleviating Scalability Limitation of Accelerator-Based Platforms
abstract
Accelerator-based chip multiprocessors (ACMPs), which combine application-specific HW accelerators (ACCs) with host processor core(s), are promising architectures for high-performance and power-efficient computing. However, ACMPs with many ACCs have scalability limitations. The ACCs' performance benefits can be overshadowed by bottlenecks on shared resources of processor core(s), communication fabric/DMA, and on-chip memory. Primarily, this is rooted in the ACCs' data access and the orchestration dependency. Due to very loosely defined ACC communication semantics, and relying on general architectures, the resources bottlenecks hamper performance. This paper explores and alleviates the scalability limitations of ACMPs. To this end, this paper first proposes ACMPerf, an analytical model to capture the impact of the resources bottlenecks on the achievable ACCs' benefits. Then, this paper identifies and formalizes ACC communication semantics which paves the path toward a more scalable integration of ACCs. The semantics describe four primary aspects: 1) data access; 2) data granularity; 3) data marshalling; and 4) synchronization. Finally, this paper proposes a novel architecture of transparent self-synchronizing accelerators (TSS). TSS efficiently realizes our identified communication semantics of direct ACC-to-ACC connections often occurring in streaming applications. TSS delivers more of the ACCs' benefits than conventional ACMP architectures. Given the same set of ACCs, TSS has up to 130× higher throughput and 78× lower energy consumption, mainly due to reducing the load on shared architectural resources by 78.3×.
Nasibeh Teimouri, Hamed Tabkhi, Gunar Schirner
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2018 Locality aware memory assignment and tiling
abstract
With the trend toward specialization, an efficient memory-path design is vital to capitalize customization in data-path. A monolithic memory hierarchy is often highly inefficient for irregular applications, traditionally targeted for CPUs. New approaches and tools are required to offer application-specific memory customization combining the benefits of cache and scratchpad memory simultaneously.
Samuel Rogers, Hamed Tabkhi
DAC2
2018 DS-DSE: Domain-specific design space exploration for streaming applications
abstract
Domain-specific computing is promising for high-performance low-power execution of applications with similar functionality. In particular, streaming applications with significant functional and structural similarities can tremendously benefit. However, current Design Space Exploration (DSE) focuses on individual applications in isolation. Hence, much of the domain optimization opportunities are missed. DSE methodologies need to broaden the scope from individual applications in isolation to optimizing across applications within a domain. This paper introduces a novel Domain-Specific DSE (DS-DSE) approach for domain-specific computing with a focus on streaming applications. Key contributions are: (1) a formalized method to extract the functional and structural similarities of domain applications, (2) a novel algorithm for hardware/software partitioning of a domain-specific platform to maximize the throughput across domain applications (under certain constraints) and (3) a methodology to evaluate a domain platform. This paper demonstrates the benefits using 4 domains: OpenVX (vision processing), and 3 synthetic domains (with greater complexity). Our experiments demonstrate a performance improvement (average throughput) of 36.8% for OpenVX and 46.2% for synthetic domains of the DS-DSE generated platform compared to an application-specific platform.
Jinghan Zhang 0001, Hamed Tabkhi, Gunar Schirner
DATE2
2016 Improving scalability of CMPs with dense ACCs coverage
Nasibeh Teimouri, Hamed Tabkhi, Gunar Schirner
DATE2
2016 Guiding Power/Quality Exploration for Communication-Intense Stream Processing
abstract
In this paper, we explore the power/quality trade-off for streaming applications with a shift from the computation to the communication aspects of the design. The paper proposes a systematic exploration methodology to formulate and traverse power/quality trade-off for the class of adaptive streaming applications. The formalization enables to procedurally transition from a set of design requirements to architecture goals. The architecture goals can then be realized through design choices yielding system designs that meet the initial requirements. The reported results are based on an actual implementation of Mixture of Gaussian (MoG) background subtraction on Xilinx Zynq platform.
Hamed Tabkhi, Majid Sabbagh, Gunar Schirner
ACM Great Lakes Symposium on VLSI1
2016 Hardware thread reordering to boost OpenCL throughput on FPGAs
abstract
Availability of OpenCL for FPGAs has raised new questions about the efficiency of massive thread-level parallelism on FPGAs. The general trend is toward creating deep pipelining and in-order execution of many OpenCL threads across a shared data-path. While this can be a very effective approach for regular kernels, its efficiency significantly diminishes for irregular kernels with runtime-dependent control flow. We need to look for new approaches to improve execution efficiency of FPGAs when targeting irregular OpenCL kernels. This paper proposes a novel solution, called Hardware Thread Reordering (HTR), to boost the throughput of the FPGAs when executing irregular kernels possessing non-deterministic runtime control flow. The key insight of HRT is out-of-order OpenCL thread execution over a shared data-path to achieve significantly higher throughput. The thread reordering is performed at a basic-block level granularity. The synthesized basic-blocks are extended with independent pipeline control signals and context registers to bypass the live values of reordered threads. We demonstrate the efficiency of our proposed solution on three parallel irregular kernels. For the experiments, we utilize the LegUp tool to compare the baseline (in-order) data-path with HTR-enhanced data-path. Our RTL simulation results demonstrate that HTR-enhanced data-path achieves up to 11× increase in kernels throughput at a very low overhead (less than 2× increase in FPGA resources).
Amir Momeni, Hamed Tabkhi, Gunar Schirner, David R. Kaeli
ICCD2
2015 An efficient architecture solution for low-power real-time background subtraction
abstract
Embedded vision is a rapidly growing market with a host of challenging algorithms. Among vision algorithms, Mixture of Gaussian (MoG) background subtraction is a frequently used kernel involving massive computation and communication. Tremendous challenges need to be reolved to provide MoG's high computation and communication demands with minimal power consumption allowing its embedded deployment. This paper proposes a customized architecture for power-efficient realization of MoG background subtraction operating at Full-HD resolution. Our design process benefits from system-level design principles. An SLDL-captured specification (result of high-level explorations) serves as a specification for architecture realization and hand-crafted RTL design. To optimize the architecture, this paper employs a set of optimization techniques including parallelism extraction, algorithm tuning, operation width sizing and deep pipelining. The final MoG implementation consists of 77 pipeline stages operating at 148.5 MHz implemented on a Zynq-7000 SoC. Furthermore, our background subtraction solution is flexible allowing end users to adjust algorithm parameters according to scene complexity. Our results demonstrate a very high efficiency for both indoor and outdoor scenes with 145 mW on-chip power consumption and more than 600× speedup over software execution on ARM Cortex A9 core.
Hamed Tabkhi, Majid Sabbagh, Gunar Schirner
ASAP1
2015 Revisiting accelerator-rich CMPs: challenges and solutions
abstract
Heterogeneous Chip Multiprocessors (CMP)s, which combine processor cores with specialized HW accelerators, are one main approach to high-performance low-power computing. While it is promising for few accelerators, the scalability is a major challenge with increasing number of accelerators. Resources including memory, communication fabric and processor turn into bottlenecks and result in accelerator under-utilization and cripple the performance.
Nasibeh Teimouri, Hamed Tabkhi, Gunar Schirner
DAC2
2015 Bridging Architecture and Programming for Throughput-Oriented Vision Processing (Abstract Only)
abstract
With the expansion of OpenCL support across many heterogeneous devices (including FPGAs, GPUs and CPUs), the programmability of these systems has been significantly increased. At the same time, new questions arise about which device should be targeted for each OpenCL software kernel. Once we select a device, then we are left to customize the application, selecting the right granularity of parallelism and frequency of host-to-device communication. In this paper, we study the impact of source-level decisions on the overall execution time when developing OpenCL program across different heterogeneous devices. We focus on two mainstream architecture classes (GPUs and FPGAs), and consider throughput-oriented advanced vision processing. To guide this exploration, we propose a new vertical classification for selecting the grain of parallelism for advanced vision processing applications. To carry out this study we have selected the Mean-shift object tracking algorithm as a representative candidate of advanced vision algorithms. Overall, our evaluation demonstrates that fine-grained parallelism can greatly benefit FPGA execution (up to a 4X speed-up), while a combination of coarse-grained and fine-grained parallelism achieves the best performance on a GPU (up to a 6X speed-up). Also, there can be a large benefit if we can execute both the parallel and serial parts of the program on a FPGA (up to a 21X speed-up).
Amir Momeni, Hamed Tabkhi, Gunar Schirner, David R. Kaeli
FPGA2
2015 A Joint SW/HW Approach for Reducing Register File Vulnerability
abstract
The Register File (RF) is a particularly vulnerable component within processor core and at the same time a hotspot with high power density. To reduce RF vulnerability, conventional HW-only approaches such as Error Correction Codes (ECCs) or modular redundancies are not suitable due to their significant power overhead. Conversely, SW-only approaches either have limited improvement on RF reliability or require considerable performance overhead. As a result, new approaches are needed that reduce RF vulnerability with minimal power and performance overhead. This article introduces Application-guided Reliability-enhanced Register file Architecture (ARRA), a novel approach to reduce RF vulnerability of embedded processors. Taking advantage of uneven register utilization, ARRA mirrors, guided by a SW instrumentation, frequently used active registers into passive registers. ARRA is particularly suitable for control applications, as they have a high reliability demand with fairly low (uneven) RF utilization. ARRA is a cross-layer joint HW/SW approach based on an ARRA-extended RF microarchitecture, an ISA extension, as well as static binary analysis and instrumentation. We evaluate ARRA benefits using an ARRA-enhanced Blackfin processor executing a set of DSPBench and MiBench benchmarks. We quantify the benefits using RF Vulnerability Factor (RFVF) and Mean Work To Failure (MWTF). ARRA significantly reduces RFVF from 35% to 6.9% in cost of 0.5% performance lost for control applications. With ARRA’s register mirroring, it can also correct Multiple Bit Upsets (MBUs) errors, achieving an 8x increase in MWTF. Compared to a partially ECC-protected RF approach, ARRA demonstrates higher efficiency by achieving comparable vulnerability reduction at much lower power consumption.
Hamed Tabkhi, Gunar Schirner
ACM Trans. Archit. Code Optim.1
2014 Function-Level Processor (FLP): Raising efficiency by operating at function granularity for market-oriented MPSoC
abstract
The exponential growth in computation demand drives chip vendors to heterogeneous architectures combining Instruction-Level Processors (ILPs) and custom HW Accelerators (HWACCs) in an attempt to provide the needed processing capabilities while meeting power/energy requirements. ILPs, on one hand, are highly flexible, but power inefficient. Custom HWACCs, on the other hand, are inflexible (focusing on dedicated kernels), but highly power efficient. Since, designing HWACCs for every application is cost prohibitive, large portions of applications still run inefficiently on ILPs. New processing architectures are needed that combine the power efficiency of HWACCs while still retaining sufficient flexibility to realize applications across targeted market segments. This paper introduces Function-Level Processors (FLPs) to fill the gap between ILPs and dedicated HWACCs. FLPs are comprised of configurable Function Blocks (FBs) implementing selected functions which are then interconnected via programmable point-to-point connections constructing an extensible/configurable macro data-path. An FLP raises programming abstraction to a Function-Set Architecture (FSA) controlling FBs allocation, configuration and scheduling. We demonstrate FLP benefits with an industry example of the Pipeline-Vision Processor (PVP). We highlight the gained flexibility by mapping 10 embedded vision applications entirely to the FLP-PVP offering up to 22.4 GOPs/s with average power of 120 mW. The results also demonstrate that our FLP-PVP solution consumes 14×-18× less power than an ILP and 5x less power than a hybrid ILP+HWACCs solution.
Hamed Tabkhi, Robert Bushey, Gunar Schirner
ASAP1
2014 A Power-Efficient FPGA-Based Mixture-of-Gaussian (MoG) Background Subtraction for Full-HD Resolution
abstract
This short paper briefly describes an FPGA-based realization of MoG background subtraction operating at fullHD frame resolution. Our HW hand-crafted MoG consists of 77 pipeline stages operating at 148.5 MHz implemented on a Zynq-7000 SoC. The results very high efficiency with a power consumption of less than 500 mW which is 600X more efficient than an embedded software solution.
Hamed Tabkhi, Majid Sabbagh, Gunar Schirner
FCCM1
2014 A GPU-Based Algorithm-Specific Optimization for High-Performance Background Subtraction
abstract
Background subtraction is an essential first stage in many vision applications differentiating foreground pixels from the background scene, with Mixture of Gaussians (MoG) being a widely used implementation choice. MoG's high computation demand renders a real-time single threaded realization infeasible. With it's pixel level parallelism, deploying MoG on top of parallel architectures such as a Graphics Processing Unit (GPU) is promising. However, MoG poses many challenges having a significant control flow (potentially reducing GPU efficiency) as well as a significant memory bandwidth demand. In this paper, we propose a GPU implementation of Mixture of Gaussians (MoG) that surpasses real-time processing for full HD (1080p 60 Hz). This paper describes step-wise optimizations starting from general GPU optimizations (such as memory coalescing, computation & communication overlapping), via algorithm-specific optimizations including control flow reduction and register usage optimization, to windowed optimization utilizing shared memory. For each optimization, this paper evaluates the performance potential and identifies architectural bottlenecks. Our CUDA-based implementation improves performance over sequential implementation by 57×, 97× and 101× through general, algorithm-specific, and windowed optimizations respectively, without impact to the output quality.
Chulian Zhang, Hamed Tabkhi, Gunar Schirner
ICPP2
2014 Application-Guided Power Gating Reducing Register File Static Power
abstract
Power and energy efficiency are on the top priority list in embedded computing. Embedded processors taped out in deep submicron technology have a high contribution of static power to overall power consumption. At the same time, current embedded processors often include a large register file (RF) to increase performance. However, a larger RF aggravates the static power issues associated with technology shrinking. Therefore, approaches to improve static power consumption of large RFs are in high demand. In this paper, we introduce an application-guided function-level register file power-gating (AFReP) approach to efficiently manage and reduce the RF's static power consumption. The AFReP is an interplay of automatic binary analysis and instrumentation at function-level granularity supported by instruction-set architecture and microarchitecture extensions. The AFReP enables runtime power-gating of registers during unutilized periods, whereas applications can fully benefit from a large RF during utilized periods. To demonstrate the AFReP's potential for reducing static power consumption, we have enhanced a Blackfin processor with the AFReP technology. Using the AFReP, the RF static power is reduced on average by 64% and 39% for control and DSP applications, respectively. At the same time, the AFReP only induces a very minimal overhead of 0.4% and 0.6%.
Hamed Tabkhi, Gunar Schirner
IEEE Trans. Very Large Scale Integr. Syst.1
2012 Application-specific power-efficient approach for reducing register file vulnerability
abstract
This paper introduces a power efficient approach for improving reliability of heterogeneous register files in embedded processors. The approach is based on the fact that control applications have high demands in reliability, while many special-purpose register are unused in a considerable portion of execution. The paper proposes a static application binary analysis which is applied at function-level granularity and offers a systematic way to manage the RF's protection by mirroring the content of used registers into unused ones. The simulation results on an enhanced Blackfin processor demonstrate that Register File Vulnerability Factor (RFVF) is reduced from 35% to 6.9% in cost of 1% performance lost on average for control applications from Mibench suite.
Hamed Tabkhi, Gunar Schirner
DATE1
2012 AFReP: Application-guided Function-level Registerfile power-gating for embedded processors
abstract
With shrinking CMOS feature size, static power is growing significantly and power density has emerged as an increasing concern. At the same time, one trend of embedded processors is toward larger Register Files (RFs) which further increases static power dissipation and aggravating the issue. This paper introduces an Application-guided Function-level Register file Power-gating (AFReP) that reduces static power of RFs in embedded processors. Our AFReP approach is based on a automatic analysis of register lifetime in the application binary, followed an automatic binary instrumentation for runtime RF power-gating. The instrumented code executes on a processor with ISA and micro-architecture extension for power-gating control over individual registers. Our application binary analysis/instrumentation operates at function-level granularity, automatically gating the registers that do not contribute to program outcome. Our experimental results using an AFReP-enhanced Blackfin processor demonstrate average RF static power reduction by 60% and 52% for control and DSP applications from Mibench and DSPstone suites, respectively. The added instructions for run-time power-gating increase execution time by only 1% on average.
Hamed Tabkhi, Gunar Schirner
ICCAD1
2012 ARRA: Application-guided reliability-enhanced registerfile architecture for embedded processors
Hamed Tabkhi, Gunar Schirner
VLSI-SoC1
2010 An Efficient Method to Reliable Data Transmission in Network-on-Chips
abstract
Data transmission in Network-on-Chips (NoCs) is a serious problem due to cross talk faults happening in adjacent communication links. This paper proposes an efficient flow-control method to enhance the reliability of packet transmission in Network-on-Chips. The method investigates the opposite direction transitions appearing between flits of a packet to reorder the flits in the packet. Flits are reordered in a fixed-size window to reduce: 1) the probability of cross talk occurrence, and 2) the total power consumed for packet delivery. The proposed flow-control method is evaluated by a VHDL-based simulator under different window sizes and various channel widths. Simulation results enable NoC designers to make a trade-off between window size, reliability and power consumption of packet delivery. This method is also compared with other cross talk tolerant methods in terms of reliability and power consumption. Comparison results confirm that the method is a cost efficient solution to overcome the cross talk problem.
Ahmad Patooghy, Hamed Tabkhi, Seyed Ghassem Miremadi
DSD2