EDBT 2026 Demo / reviewers in the wild / expert
Saeid Gorgin 0001
dblp:92/547 · also Saeed Gorgin 0001
· DBLP profile ↗
39ranked-venue papers
10as first author
21since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 18 · 6 first-author · 12 since 2021Theory of computation · 9 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 7 · 2 since 2021Computer networks · 3 · 2 since 2021Security and privacy · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TRM: An Efficient Hypervisor-Based Framework For Malware Analysis and Memory ReconstructionabstractModern rootkits leverage kernel privileges to hide from analysis tools, while obfuscation techniques render them resistant to static analysis. Reverse engineering malware requires observing memory usage and reconstructing data structures. Existing tools rely on instrumentation or emulation, which introduce high overhead, leave detectable artifacts, and cannot reliably analyze kernel-level malware. We present The Reversing Machine (TRM), a hypervisor-based framework for high-performance introspection of evasive malware. It is the first system to support selective memory tracing and structure reconstruction in the hypervisor. TRM repurposes hardware virtualization to efficiently detect user-kernel mode transitions and obtain memory traces independently of a potentially compromised guest kernel. TRM introduces new insights into leveraging hardware virtualization for runtime memory reconstruction and analysis of data structures while remaining invisible to malware. We demonstrate automatic reconstruction of function signatures and data structures, reduce system call latency overhead from 142% to 57% compared to prior work, and accelerate manual reverse engineering by 43% on average, even for complex kernel objects. TRM shows that hypervisor-level memory tracing makes data structure reconstruction practical in hostile environments, bridging the gap between prior feasibility studies and real-world malware analysis. Mohammad Sina Karvandi, Soroush Meghdadi Zanjani, Sima Arasteh, Saleh Khalaj Monfared, Mohammad K. Fallah, Saeid Gorgin 0001, Jeong-A Lee, Asia Slowinska, Erik van der Kouwe |
AsiaCCS | 6 |
| 2026 | RowArmor: Efficient and Comprehensive Protection Against DRAM Disturbance Attacks
Minbok Wi, Yoonyul Yoo, Yoojin Kim, Jumin Kim, Yesin Ryu, Saeid Gorgin 0001, Jung Ho Ahn, Jungrae Kim |
ASPLOS (2) | 7 |
| 2026 | Cerberus: Cross-Layer ECC Co-Design for Robust and Efficient Memory Protection
Junhwan Kim, Yesin Ryu, Saeid Gorgin 0001, Jungrae Kim |
ISCA | 4 |
| 2026 | Efficient Modular Addition for FPGA-Based Cryptographic OperationsabstractModular adders are essential components in finite field arithmetic, serving as key components in public-key cryptographic algorithms like Elliptic Curve Cryptography (ECC) and Post-Quantum Cryptography (PQC). Naïve implementation of modular adders, due to two cascaded adders with large operand bit widths struggle to meet high-frequency requirements. On the other hand, parallel implementations, while faster, demand excessive resources and power, making them impractical for many applications. This paper introduces a novel modular addition algorithm leveraging a novel operand representation based on the two-valued digit encoding (Twit). In this approach, each operand is represented as an n-bit unsigned number augmented by a Twit value {0,±δ}. The algorithm efficiently computes modular addition by speculating and dynamically adjusting the twit value in the result, achieving both computational and resource efficiency. The proposed design has been implemented on a Xilinx 7-series FPGA, demonstrating superior performance in achieving high operating frequencies (i.e., 8% to 36% depending on operand bit widths) while significantly reducing resource utilization (i.e., >36%). In addition to extensive analytical and synthesis-based evaluations, we further demonstrate the benefits of the proposed adder within application-level cryptographic datapaths (ECC and PQC). Saeid Gorgin 0001, Amirhossein Sadr, Dara Rahmati, Ali Jahanian 0001, Jungrae Kim |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2026 | Efficient neural network acceleration using redundant residue number systems
Soudabeh Mousavi, Dara Rahmati, Amirhossein Sadr, Saeid Gorgin 0001, Jungrae Kim |
J. Supercomput. | 4 |
| 2026 | FPGA-accelerated real-time DCGANs via Xilinx DPUs and Vitis AI
Amirhossein Sadr, Aida Pakniyat, Dara Rahmati, Saeid Gorgin 0001 |
J. Supercomput. | 4 |
| 2025 | A Generic Modulo-(2n±δ) Addition Algorithm via Two-Valued Digit EncodingabstractModular adders are essential arithmetic components in Residue Number System (RNS)-based applications, including digital signal processing, cryptography, and machine learning. These applications consistently push the boundaries of dynamic range (DR) and operating frequency, making the design of efficient generic modular adders a critical and evolving challenge. This paper presents a novel algorithm for modulo$-(2^{n}\pm \delta)$addition, where$\delta$is an integer within the range$0\leq\delta\leq 2^{n-1}-1$. The proposed approach leverages a two-valued digit (twit) for encoding the value of$\pm\delta$and uses a faithful representation of operands. In this representation, each operand is encoded as an n-bit unsigned number augmented by a twit value$\{0,\pm\delta\}$. The algorithm efficiently performs modular addition by speculating and adjusting the twit value in the addition result. When the result exceeds the modulus, it subtracts$\mathrm{z}^{n}\pm\delta$by ignoring the carry-out and adjusting the speculated twit value. This adjustment is achieved through an XOR operation between the carry-out and the speculated twit value, simplifying the modular reduction process. The proposed design has been synthesized for practical n$(4\leq n\leq 16$using a FreePDK 45 nm process. The results demonstrate superior performance across key metrics such as delay, area, and power consumption compared to previous designs, highlighting the efficacy and scalability of the approach. Saeid Gorgin 0001, Amirhossein Sadr, Dara Rahmati, Jungrae Kim |
ARITH | 1 |
| 2025 | PoP-ECC: Robust and Flexible Error Correction against Multi-Bit Upsets in DNN AcceleratorsabstractDeep Neural Networks (DNNs) in safety-critical systems require high reliability. Many systems deploy Error Correction Codes (ECCs) to protect DNNs from memory errors. However, continuous process scaling increases memory errors in severity and frequency, necessitating strong protection against Multi-Bit Upsets (MBUs). This paper proposes Parities of Parities ECC (PoP-ECC), a novel two-tier memory protection scheme designed to provide robust, efficient, and flexible protection against MBUs. PoP-ECC generates Virtual Parities (VPs), which are used to compute secondlevel parities called Parities of Parities (PPs). This two-level ECC structure allows for dynamic error correction tailored to varying error patterns, ensuring system reliability with minimal memory overhead. Our evaluation demonstrates that PoP-ECC can tolerate significantly higher MBU ratios compared to state-of-the-art solutions, with negligible delay, area, and power overhead. Taewon Park, Saeid Gorgin 0001, Dongwhee Kim, Michael B. Sullivan 0001, Jungrae Kim |
DAC | 2 |
| 2025 | Scaling Out Chip Interconnect Networks with Implicit Sequence NumbersabstractAs AI models outpace the capabilities of single processors, interconnects across chips have become a critical enabler for scalable computing. These processors exchange massive amounts of data at cache-line granularity, prompting the adoption of new interconnect protocols like CXL, NVLink, and UALink, designed for high bandwidth and small payloads. However, the increasing transfer rates of these protocols heighten susceptibility to errors. While mechanisms like Cyclic Redundancy Check (CRC) and Forward Error Correction (FEC) are standard for reliable data transmission, scaling chip interconnects to multi-node configurations introduces new challenges, particularly in managing silently dropped flits in switching devices. Giyong Jung, Saeid Gorgin 0001, John Kim 0001, Jungrae Kim |
SC | 2 |
| 2025 | Poster: Integration of Wearable and Affective Computing via Abstraction and Decision Fusion ArchitectureabstractThis paper introduces an efficient emotion detection method to integrate wearable and affective computing paradigms. Our research contributes to advancing emotion detection technologies, offering potential applications in diverse domains such as healthcare, human-computer interaction, and personalized computing experiences. Our approach addresses the increasing need for real-time emotion recognition while minimizing computational demands. By leveraging low-computation techniques, we propose a novel framework that achieves high accuracy in emotion detection. Besides, advanced data abstraction methods are developed to reduce data workload keeping detection performance. Experimental results demonstrate a notable accuracy rate of $89 .77$%, affirming the efficacy of our proposed method. Mohammadreza Najafi, Mohammad K. Fallah, Saeid Gorgin 0001, Ghassem Jaberipur, Jeong-A Lee |
WoWMoM | 3 |
| 2025 | Efficient hardware accelerators for k-nearest neighbors classification using most significant digit first arithmetic
Saeid Gorgin 0001, Malik Zohaib Nisar, Jeong-A Lee |
J. Supercomput. | 1 |
| 2024 | An ultra-low-computation model for understanding sign languages
Mohammad K. Fallah, Mohammadreza Najafi, Saeid Gorgin 0001, Jeong-A Lee |
Expert Syst. Appl. | 3 |
| 2023 | Modulo-(2q - 3) Multiplication with Fully Modular Partial Product Generation and ReductionabstractGiven the residue number systems that contain moduli of the form 2q± 1 and 2q± 3, it is desirable to employ delay-balanced adders and multipliers, in order to synchronize the operation of parallel residue channels. The required modulo-(2q± 3) adders, with compatible speed with modulo- (2q± 1) adders, already exist with parallel prefix architectures. However, the previously reported modulo-(2q± 3) multipliers, in one way or another, produce the non-modular products of the residues at the outset and work towards yielding the final modular product. This seems to be the main source of incompatible performance with the existing modulo- (2q± 1) fully modular multipliers. Therefore, as the first endeavor, we were motivated to design and implement efficient modulo-(2q− 3) multipliers with fully modular partial product generation and reduction that are more compatible with their modulo- (2q− 1) counterparts. However, unlike the case of modulo 2q− 1, it turns out that the straightforward modulo-(2q− 3) partial product reduction (e.g., via Wallace-tree reduction with greedy use of full adders and half adders) falls into an infinite loop of reduction stages. Therefore, we undertake a modified reduction algorithm that requires at most two reduction levels more than that of the modulo- (2q− 1) case to converge. To ensure the correct operation of the algorithm and ease the design process, an in-house software program produces the exact composition of reduction cells in each level of partial product reduction. Analytical and synthesis-based evaluations of the proposed design, and the previous ones, exhibit better figures of merit, as regards the delay (≥ 24%), area-delay (≥ 6%) and energy (≥ 10%) measures. Ghassem Jaberipur, Saeid Gorgin 0001, Navid Ahamadian, Jeong-A Lee |
ARITH | 2 |
| 2023 | A new energy-efficient and temperature-aware routing protocol based on fuzzy logic for multi-WBANs
Danial Javaheri, Pooia Lalbakhsh, Saeid Gorgin 0001, Jeong-A Lee, Mohammad Masdari |
Ad Hoc Networks | 3 |
| 2023 | Fuzzy logic-based DDoS attacks and network traffic anomaly detection methods: Classification, overview, and future perspectives
Danial Javaheri, Saeid Gorgin 0001, Jeong-A Lee, Mohammad Masdari |
Inf. Sci. | 2 |
| 2023 | A Generalized Residue Number System Design Approach for Ultralow-Power Arithmetic Circuits Based on Deterministic Bit-StreamsabstractThe peak power consumption has become an important concern in the hardware design process of some of today’s applications, such as energy harvesting (EH) and bio-implantable (BI) electronic devices. The limited peak harvested power in EH devices and heating concerns in BI devices are the main reasons for power control’s importance in these devices. This article proposes a generalized design approach for ultralow-power arithmetic circuits. The proposed circuits are based on residue number system (RNS) combined with deterministic bit-streams. The resulting circuits can be used in systems with a restricted power budget. We suggest several approaches to design generic hardware-efficient adders, multipliers, multiply-accumulate (MAC) unit, forward converters (FCs), and reverse converters (RCs). Using the proposed approach, designing these components for any moduli of the RNS can be performed through simple bit-width adjustments in the circuits. The synthesis results show that the proposed adder achieves, on average, 69% and 2% lower area compared to the bit-serial and a state-of-the-art RNS adder, respectively. Furthermore, the proposed multiplier outperforms the bit-serial, interleaved, and a state-of-the-art design for multiplying RNS numbers by, on average, 57%, 60%, and 77% in terms of power consumption, respectively. The efficiency of our approach is shown via two essential applications, digital signal processing, and machine learning. We implement an FFT engine using the proposed method. Compared to prior RNS implementations, our design achieves 47% lower power consumption. We also implement a CNN accelerator’s processing element (PE) with the proposed computation elements. Our design provides considerable speedup and lower power consumption compared to a state-of-the-art ultralower-power design. Kamyar Givaki, Ahmad Khonsari, MohammadHosein Gholamrezaei, Saeid Gorgin 0001, M. Hassan Najafi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | HyperDbg: Reinventing Hardware-Assisted DebuggingabstractSoftware analysis, debugging, and reverse engineering have a crucial impact in today's software industry. Efficient and stealthy debuggers are especially relevant for malware analysis. However, existing debugging platforms fail to address a transparent, effective, and high-performance low-level debugger due to their detectable fingerprints, complexity, and implementation restrictions. Mohammad Sina Karvandi, MohammadHosein Gholamrezaei, Saleh Khalaj Monfared, Soroush Meghdadi Zanjani, Behrooz Abbassi, Reza Mortazavi, Saeid Gorgin 0001, Dara Rahmati, Michael Schwarz 0001 |
CCS | 8 |
| 2022 | An Efficient FPGA Implementation of k-Nearest Neighbors via Online Arithmeticabstractk-NN, as one of the well-employed classification algorithms, severely suffers from a computationally intensive nature. This paper exploits the parallelism and digit level pipelining opportunities via FPGA devices and Online arithmetic to offer an efficient k-NN FPGA implementation. All the required operations for computing distances and sorting are applied to serially coming data. Moreover, we dynamically terminate the unnecessary computations once they are detected. To the best of our knowledge, the proposed k-NN implementation is the first one that used FPGA and Online arithmetic effectively. It provides up to 34% speedup compared to the best state-of-the-art design. Saeid Gorgin 0001, MohammadHosein Gholamrezaei, Danial Javaheri, Jeong-A Lee |
FCCM | 1 |
| 2022 | An Energy-Efficient K-means Clustering FPGA Accelerator via Most-Significant Digit First ArithmeticabstractK-means clustering is the most well-known unsupervised learning method that partitions the input dataset into$K$clusters based on the similarity between the data samples. In this paper, to achieve an energy-efficient implementation without sacrificing performance, we take advantage of massive parallelism and digit-level pipelining via FPGA and the most-significant digit first arithmetic. Having the result of the most-significant digits in advance provides the possibility of early termination for unnecessary computations and fetching just the required most-significant part of data points from memory. This early termination technique significantly increases performance and decreases energy consumption. Our experimental results from various datasets and comparisons with the state-of-the-art FPGA accelerators indicate that our proposed design has effectively reduced energy consumption without any performance loss. Saeid Gorgin 0001, MohammadHosein Gholamrezaei, Danial Javaheri, Jeong-A Lee |
FPT | 1 |
| 2022 | Hardware Efficient FIR Filter Architectures Using Accurate Unary Stochastic ComputingabstractFinite Impulse Response (FIR) filters are commonly used due to lower sensitivity to noise than their recursive counterparts. Computations of FIR filters require numerous multiply-and-accumulate (MAC) operations. Therefore, hard-ware implementation of high-order adaptive FIR filters results in a considerable area and power consumption. This paper proposes a hardware-efficient FIR engine based on the integration of deterministic approaches to Stochastic Computing (SC) with Residue Number Systems (RNS). The design inherits the intrinsic simplicity and low hardware requirements of SC circuits. As a contribution of our work, in contrast to other SC-based methods that impose errors on computations, our proposed method offers exact results like the binary implementations of FIR filters. Furthermore, the design decreases the required clock cycles, which can be translated to higher throughput in comparison with its SC predecessor (for example, 4× for 8-bit computations) at the cost of acceptable hardware overhead. Kamyar Givaki, Ahmad Khonsari, M. Hossein Gholamrezaei, Dara Rahmati, Saeid Gorgin 0001 |
ICCD | 5 |
| 2021 | A TSX-Based KASLR Break: Bypassing UMIP and Descriptor-Table Exiting
Mohammad Sina Karvandi, Saleh Khalaj Monfared, Sina Kiarostami, Dara Rahmati, Saeid Gorgin 0001 |
CRiSIS | 5 |
| 2020 | On the Resilience of Deep Learning for Reduced-voltage FPGAsabstractDeep Neural Networks (DNNs) are inherently computation-intensive and also power-hungry. Hardware accelerators such as Field Programmable Gate Arrays (FPGAs) are a promising solution that can satisfy these requirements for both embedded and High-Performance Computing (HPC) systems. In FPGAs, as well as CPUs and GPUs, aggressive voltage scaling below the nominal level is an effective technique for power dissipation minimization. Unfortunately, bit-flip faults start to appear as the voltage is scaled down closer to the transistor threshold due to timing issues, thus creating a resilience issue.This paper experimentally evaluates the resilience of the training phase of DNNs in the presence of voltage underscaling related faults of FPGAs, especially in on-chip memories. Toward this goal, we have experimentally evaluated the resilience of LeNet-5 and also a specially designed network for CIFAR-10 dataset with different activation functions of Rectified Linear Unit (Relu) and Hyperbolic Tangent (Tanh). We have found that modern FPGAs are robust enough in extremely low-voltage levels and that low-voltage related faults can be automatically masked within the training iterations, so there is no need for costly software-or hardware-oriented fault mitigation techniques like ECC. Approximately 10% more training iterations are needed to fill the gap in the accuracy. This observation is the result of the relatively low rate of undervolting faults, i.e., <0.1%, measured on real FPGA fabrics. We have also increased the fault rate significantly for the LeNet-5 network by randomly generated fault injection campaigns and observed that the training accuracy starts to degrade. When the fault rate increases, the network with Tanh activation function outperforms the one with Relu in terms of accuracy, e.g., when the fault rate is 30% the accuracy difference is 4.92%. Kamyar Givaki, Behzad Salami 0001, Reza Hojabr, S. M. Reza Tayaranian, Ahmad Khonsari, Dara Rahmati, Saeid Gorgin 0001, Adrián Cristal, Osman S. Unsal |
PDP | 7 |
| 2020 | A fuzzy irregular cellular automata-based method for the vertex colouring problemabstractVertex colouring is among the most important problems in graph theory which has been widely applied across different real-world problems. In vertex colouring problem (VCP), the goal is to assign a distinct colour to each vertex of the graph in such a way that no two adjacent vertices have the same colour. This paper presents a fuzzy irregular cellular automaton (FICA) for finding a near-optimal solution for the VCP. FICA is an extension fuzzy cellular automaton (FCA) in which the cells of the automaton can be arranged in an irregular structure. The aim of the proposed method is to reap the benefits of both FCA and irregular cellular automata while minimising their drawbacks. To evaluate the proposed method, various computer simulations have been conducted on a variety of graphs. The results suggest that the proposed method is able to achieve better results in terms of the minimum number of required colours and the execution time of the algorithm, compared to other peer algorithms. Mostafa Kashani, Saeid Gorgin 0001, Seyed Vahab Shojaedini |
Connect. Sci. | 2 |
| 2019 | Using Residue Number Systems to Accelerate Deterministic Bit-stream MultiplicationabstractInaccuracy of computations is an important challenge with Stochastic Computing (SC). Deterministic approaches are proposed to produce completely accurate results with SC circuits. Current deterministic methods need a large number of clock cycles to produce exact result. This directly translates to a very high energy consumption. We propose a method based on the Residue Number Systems (RNS) to mitigate the high processing time of the deterministic methods. Compared to the state-of-the-art deterministic methods of SC, our approach delivers 760x and 170x improvement in terms of processing time and energy consumption. Kamyar Givaki, Reza Hojabr, M. Hassan Najafi, Ahmad Khonsari, M. Hossein Gholamrezayi, Saeid Gorgin 0001, Dara Rahmati |
ASAP | 6 |
| 2019 | Multi-Agent non-Overlapping Pathfinding with Monte-Carlo Tree SearchabstractIn this work, we propose a novel implementation of Monte-Carlo Tree Search (MCTS) algorithm to solve a multiagent pathfinding (MAPF) problem. We employ an optimization of MCTS with low time-complexity and acceptable reliability to approach the MAPF problems with no time constraint. To examine the efficiency and performance of the proposed approach, the NumberLink problem as a MAPF is investigated. We show that the addressed problem could be characterized as multi-agent pathfinding problem with no overlapping paths for the agents. Furthermore, we define this problem to be a simplified and special case of Multi-commodity flow problem (MCFP). Our MCTS solution utilizes a modified search-tree structure to efficiently solve the problem based on a 2-dimensional search space which performs in quadratic time complexity (O(m4) where input size is m2) and linear memory complexity (O(m2)). To evaluate our algorithm, we investigate the efficiency of the proposed solution for the well-known Flow Free puzzle. Our implementation solves a large 40 × 40 Numberlink puzzle in 21 minutes. To the best of our knowledge, there is no other efficient solution for this puzzle where the size of the problem is considerably large. Sina Kiarostami, Mohammad Reza Daneshvaramoli, Saleh Khalaj Monfared, Dara Rahmati, Saeid Gorgin 0001 |
CoG | 5 |
| 2019 | Fast AES Implementation: A High-Throughput Bitsliced ApproachabstractIn this work, a high-throughput bitsliced AES implementation is proposed, which builds upon a new data representation scheme that exploits the parallelization capability of modern multi/many-core platforms. This representation scheme is employed as a building block to redesign all of the AES stages to tailor them for multi/many-core AES implementation. With the proposed bitsliced approach, each parallelization unit processes an unprecedented number of thirty-two 128-bit input data. Hence, a high order of prallelization is achieved by the proposed implementation technique. Based on the characteristics of this new implementation model, the ShiftRows stage can be implicitly handled through input rearrangement and is simplified to the point where its computing process can be neglected. In this implementation, costly Byte-wise operations are performed through register shift and swapping. In addition, the need for look-up table based I/O operations, which are used by the Substitute Bytes stage is eliminated through using S-box logic circuit. The S-box logic circuit is optimized to simultaneously process 32 chunks of 128-bit input data. We develop high-throughput CTR and ECB AES encryption/decryption on 6 CUDA-enabled GPUs, which achieve 1.47 and 1.38 Tbps of encryption throughput on Tesla V100 GPU, respectively. Omid Hajihassani, Saleh Khalaj Monfared, Seyed Hossein Khasteh, Saeid Gorgin 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2019 | Accuracy and availability modeling of social networks for Internet of Things event detection applications
Meghdad Aynehband, Mehdi Hosseinzadeh 0001, Houman Zarrabi, Saeid Gorgin 0001 |
Wirel. Networks | 4 |
| 2018 | A high-performance and energy-efficient exhaustive key search approach via GPU on DES-like cryptosystems
Armin Ahmadzadeh, Omid Hajihassani, Saeid Gorgin 0001 |
J. Supercomput. | 3 |
| 2017 | Sign-Magnitude Encoding for Efficient VLSI Realization of Decimal MultiplicationabstractDecimal X × Y multiplication is a complex operation, where intermediate partial products (IPPs) are commonly selected from a set of precomputed radix-10 X multiples. Some works require only [0, 5] × X via recoding digits of Y to one-hot representation of signed digits in [-5,5]. This reduces the selection logic at the cost of one extra IPP. Two's complement signed-digit (TCSD) encoding is often used to represent IPPs, where dynamic negation (via one xor per bit of X multiples) is required for the recoded digits of Y in [-5, -1]. In this paper, despite generation of 17 IPPs, for 16-digit operands, we manage to start the partial product reduction (PPR) with 16 IPPs that enhance the VLSI regularity. Moreover, we save 75% of negating xors via representing precomputed multiples by sign-magnitude signed-digit (SMSD) encoding. For the first-level PPR, we devise an efficient adder, with two SMSD input numbers, whose sum is represented with TCSD encoding. Thereafter, multilevel TCSD 2:1 reduction leads to two TCSD accumulated partial products, which collectively undergo a special early initiated conversion scheme to get at the final binary-coded decimal product. As such, a VLSI implementation of 16 × 16-digit parallel decimal multiplier is synthesized, where evaluations show some performance improvement over previous relevant designs. Saeid Gorgin 0001, Ghassem Jaberipur |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2015 | Efficient continuous skyline computation on multi-core processors based on Manhattan distanceabstractThe continuous Skyline query has recently become the subject of the several researches due to its wide spectrum of applications such as multi-criteria decision making, graph analysis network, wireless sensor network and data exploration. In these applications, the datasets are huge and have various dimensions. Moreover, they constantly change as time passes. Therefore, this query is considered as a computation intensive operation that finding the result in a reasonable time is a challenge. In this paper, we present an efficient parallel continuous Skyline approach. In our suggested method, the dataset points are sorted and pruned based on Manhattan distance. Moreover, we use several optimization methods to optimize memory usage in comparison with naïve implementation. In addition, besides the applied conventional parallelization methods, we partition the time steps based on the number of available cores. The experimental results for a dataset that contains 800k points with 7 dimensions show considerable speedup. Ehsan Montahaie, Milad Ghafouri, Saied Rahmani, Hanie Ghasemi, Farzad Sharif Bakhtiar, Rashid Zamanshoar, Kianoush Jafari, Mohsen Gavahi, Reza Mirzaei, Armin Ahmadzadeh, Saeid Gorgin 0001 |
MEMOCODE | 11 |
| 2015 | Comment on "High Speed Parallel Decimal Multiplication With Redundant Internal Encodings"abstractHanpropose a new method for parallel decimal multiplication with redundant partial products. They compare the performance of their multiplier with some previous relevant works, based on analytical and synthesis results. We have noted that the claimed critical delay path in (IEEE Trans. Computers, vol. 62, no. 5, pp. 956–968, May 2013) is faster than the actual critical delay path. Therefore, comparison results seem to be deceptive. For example, our accurate analytical evaluation devaluated the claimed speed advantage over the multiplier of (Microelectronics J., vol. 40, no. 10, pp. 1471–1481, Oct. 2009). Furthermore, we synthesized both multipliers, to show synthesis results confirm those of analytical evaluation. Saeid Gorgin 0001, Ghassem Jaberipur |
IEEE Trans. Computers | 1 |
| 2014 | Cost-efficient implementation of k-NN algorithm on multi-core processorsabstractk-nearest neighbor's algorithm plays a significant role in the processing time of many applications in a variety of fields such as pattern recognition, data mining and machine learning. In this paper, we present an accurate parallel method for implementing k-NN algorithm in multi-core platforms. Based on the problem definition we used Mahalanobis distance and developed mathematic techniques and deployed best programming experiences to accelerate contest reference implementation. Our method makes exhaustive use of CPU and minimizes memory access. This method is the winner of cost-adjust-performance of MEMOCODE contest design 2014 and is 616× faster than the reference implementation of the contest. Armin Ahmadzadeh, Reza Mirzaei, Hatef Madani, Mohammad Shobeiri, Mahsa Sadeghi, Mohsen Gavahi, Kianoush Jafari, Mohsen Mahmoudi Aznaveh, Saeid Gorgin 0001 |
MEMOCODE | 9 |
| 2014 | A fast emulator for ARM-based embedded systemsabstractThis paper presents a high-performance implementation for an Intel 8080 emulator on a Raspberry Pi device. The problem was defined as a software contest in MEMOCODE 2014 and this implementation took the second place in this contest. We deployed several optimization techniques and employed best programming practices to increase the performance of the naïve reference implementation. Improving data structure usage and modifying function calls are the techniques that resulted in higher performance of this implementation. Our implementation has about 2.5 times speedup over the reference code of the contest. Nariman Eskandari, Hatef Madani, Armin Ahmadzadeh, Mohsen Mahmoudi Aznaveh, Saeid Gorgin 0001 |
MEMOCODE | 5 |
| 2013 | Fast and adaptive BP-based multi-core implementation for stereo matching
Armin Ahmadzadeh, Hatef Madani, Kianoush Jafari, Farzad Salimi Jazi, Shervin Daneshpajouh, Saeid Gorgin 0001 |
MEMOCODE | 6 |
| 2011 | A Family of High Radix Signed Digit AddersabstractSigned digit (SD) number systems allow for high performance carry-free adders. Maximally redundant SD (MRSD) alternatives provide maximal encoding efficiency among Radix-2hSD number systems, whereby value of h tunes the area-time trade-off. Straightforward implementation of the conventional carry-free addition algorithm requires three O(log h) addition-like operations in sequence. However, there are several MRSD implementations with only one such operation. Some of them are delay optimized, but suffer from extensive hardware redundancy, while some other equally fast adders show less power/area consumption. A careful study of the latter cases hints on variety of improvement options, based on which and a new transfer computation technique, we develop a family of faster MRSD adders that consume less power/area than all the previous relevant works. They also fit efficiently within the redundant digit floating point addition scheme. However, similar to their relevant ancestor designs, suffer from an inherent property of MRSD adders, i.e., difficulty of handling hidden leading zero-digits. To remedy this problem, we use less redundant SD representations, where our transfer extraction method applies efficiently and leads to far less complex leading zero-digit detection. All the presented designs are supported by exhaustive correctness tests and performance evaluation via 0.13 micrometer CMOS technology synthesis. Saeid Gorgin 0001, Ghassem Jaberipur |
IEEE Symposium on Computer Arithmetic | 1 |
| 2011 | GPU-based NoC simulatorabstractIn this paper, we present a design and implementation of a NoC simulator, which was the subject of the MEMOCODE 2011 hardware/software co-design competition. Our design is based on the GPU platform using CUDA. For this purpose, we used an NVIDIA GeForce GTX 480 and could achieve a factor of 18 faster over the reference code-platform of the contest. Mahdy Zolghadr, Koosha Mirhosseini, Saeid Gorgin 0001, Abbas Nayebi |
MEMOCODE | 3 |
| 2010 | Redundant-Digit Floating-Point Addition Scheme Based on a Stored Rounding ValueabstractDue to the widespread use and inherent complexity of floating-point addition, much effort has been devoted to its speedup via algorithmic and circuit techniques. We propose a new redundant-digit representation for floating-point numbers that leads to computation speedup in two ways: (1) Reducing the per-operation latency when multiple floating-point additions are performed before result conversion to nonredundant format and (2) Removing the addition associated with rounding. While the first of these advantages is offered by other redundant representations, the second one is unique to our approach, which replaces the power- and area-intensive rounding addition by low-latency insertion of a rounding two-valued digit, or twit, in a position normally assigned to a redundant twit within the redundant-digit format. Instead of conventional sign-magnitude representation, we use a sign-embedded encoding that leads to lower hardware redundancy, and thus, reduced power dissipation. While our intermediate redundant representations remain incompatible with the IEEE 754-2008 standard, many application-specific systems, such as those in DSP and graphics domains, can benefit from our designs. Description of our radix-16 redundant representation and its addition algorithm is followed by the architecture of a floating-point adder based on this representation. Detailed circuit designs are provided for many of the adder's critical subfunctions. Simulation and synthesis based on a 0.13 ¿m CMOS standard process show a latency reduction of 15 percent or better, and both area and power savings of around 58 percent, compared with the best designs reported in the literature. Ghassem Jaberipur, Behrooz Parhami, Saeid Gorgin 0001 |
IEEE Trans. Computers | 3 |
| 2009 | Fully Redundant Decimal ArithmeticabstractHardware implementation of all the basic radix-10 arithmetic operations is evolving as a new trend in the design and implementation of general purpose digital processors. Redundant representation of partial products and remainders is common in the multiplication and division hardware algorithms, respectively. Carry-free implementation of the more frequent add/subtract operations, with the byproduct of enhancing the speed of multiplication and division, is possible with redundant number representation. However, conversion of redundant results to conventional representations entails slow carry propagation that can be avoided if the results are kept in redundant format for later use as operands of other arithmetic operations. Given that redundant decimal representations, contrary to redundant binary, do not necessarily require extra storage, we are motivated to develop a framework for fully redundant decimal arithmetic, where all operands and results belong to the same redundant decimal number system and can be stored and later used as operands of further decimal operations. In this paper, we present a new faster decimal signed digit add/sub unit and show how it can be efficiently used in the design of decimal multipliers and dividers, where all operands and results are represented with the same redundant digit set [-7, 7]. Saeid Gorgin 0001, Ghassem Jaberipur |
IEEE Symposium on Computer Arithmetic | 1 |
| 2007 | Reversible Barrel ShiftersabstractData shifting is required in many key computer operations from address decoding to computer arithmetic. Full barrel shifters are often on the critical path, which has led most research to be directed toward speed optimizations. With the advent of quantum computer and reversible logic, design and implementation of all devices in this logic has received more attention. This paper proposes a reversible implementation of a barrel shifter, and also evaluation of its quantum cost is presented. Saeid Gorgin 0001, Amir Kaivani |
AICCSA | 1 |