Bobin Deng

dblp:58/10076 · DBLP profile ↗
← Back
23ranked-venue papers
3as first author
17since 2021 · last 2026
0000-0001-8361-9025ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 3 first-author · 3 since 2021Computer networks · 7 · 7 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Federated Spiking Neural Networks With Top-κ Vector-Wise Trimming for Byzantine-Robust and Communication-Efficient Edge Intelligence
Manh V. Nguyen, Liang Zhao 0024, Bobin Deng, Jian Zhang 0028, Shaoen Wu
IEEE Internet Things J.3
2025 A Sparsity Predicting Approach for Large Language Models via Activation Pattern Clustering
Nobel Dhar, Bobin Deng, Md. Romyull Islam, Xinyue Zhang 0001, Kazi Fahim Ahmad Nasif, Kun Suo
Euro-Par (1)2
2025 Automatic Molecular Dynamics Simulation with Binding Affinity Prediction Using Deep Learning
abstract
In this study, we developed an automatic pipeline for molecular dynamics (MD) simulations and a deep learning model for predicting protein-protein binding affinities. The framework will output common post-simulation analyses, including total energy, Root Mean Square Deviation (RMSD), Root Mean Square Fluctuation (RMSF), hydrogen bonds, and salt bridges. It will minimize manual interaction. This tool also integrates a 3D Convolutional Neural Network (CNN) that processes structural features from simulation trajectories to predict binding affinity (pKd and$\Delta G$). We evaluate several model versions with different feature map sizes to balance between predictive accuracy wit hardware requirements. A case study on the human urate transporter GLUT9 (PDB: 8Y65) demonstrates the pipeline's ability to provide rapid, initial insights into protein stability and dynamics. This automated approach serves as a powerful complement to traditional analysis tools. It allows researchers to quickly screen simulation results and identify systems for more detailed investigation.
Lingtao Chen, Kazi Fahim Ahmad Nasif, Shuteng Niu, Bobin Deng, Chloe Yixin Xie
ICTAI5
2025 Multimodal Deep Learning for Alzheimer's Disease Classification: A Practical Fusion Framework for Clinical Deployment
abstract
Deploying multimodal deep learning for Alzheimer's disease (AD) classification in real-world clinical practice remains challenging due to class imbalance, limited data, and computational constraints. We present a robust and efficient CNN-Transformer hybrid framework that integrates cross-modal attention, self-supervised contrastive pretraining, and progressive imbalance mitigation, tailored for clinical neuroimaging. Using the ADNI dataset (744 subjects; CN: 48%, MCI: 34%, AD: 17%), our approach employs adaptive class weighting, label smoothing, and targeted augmentation to address severe imbalance, while an optimized preprocessing pipeline reduces training time by over 97 %. The model achieves 84.3 % accuracy, 88.2 % ROC-AUC, and 85.0 % macro F1, outperforming realistic baselines such as InterFusion$(82.7 \%)$, with sub-2-hour training and 0.05 s inference per case on standard hardware. Extensive ablation confirms the necessity of each component for robust generalization. Limitations include reliance on paired MRI-PET data and single-cohort validation. This study demonstrates that thoughtfully adapted, resource-conscious AI can bridge the gap between research and clinical deployment in neuroimaging.
Zakaria O. Elghazzali, Chen Zhao 0022, Lingtao Chen, Bobin Deng, Kazi Fahim Ahmad Nasif, Chloe Yixin Xie
ICTAI5
2025 AccelMD: A Self-Adaptive AI-Enabled Framework for Accelerating Molecular Dynamics Simulations
abstract
Molecular Dynamics (MD) simulation is a fundamental exploring approach for numerous scientific fields, including but not limited to drug discovery, biology, material science, and chemistry. Unfortunately, large-scale MD simulation generally requires long-period processing, even in resource-intensive supercomputers. Parallel computing is the primary methodology to accelerate MD simulation. However, parallel computing has a low theoretical improvable ceiling for MD simulation because the atom statuses of one timestep depend on its previous timestep. To unlock the full potential of MD simulation in drug discovery and biology, we propose an AccelMD framework for reducing processing time and computing costs. AccelMD is an AI-enabled framework that accelerates MD simulation while maintaining high predictive accuracy. In order to extend AccelMD's applicability and flexibility, we also developed an extra preprocessing stage, which allows the framework to adapt to various protein inputs automatically. Compared to conventional MD, the AccelMD framework achieves 47.59 X speedup on average (max: 63.78 X, min: 30.61 X). Regarding the prediction accuracy of protein structures, the average Mean Absolute Error (MAE) and TMScore values of AccelMD are 0.0058 and 0.99980, respectively. Our empirical experiments exihibit the robustness of AccelMD framework for the study of complex, long-timescale molecular interactions.
Kazi Fahim Ahmad Nasif, Bobin Deng, Lingtao Chen, Chloe Yixin Xie, Shaolei Teng, Liang Zhao 0024, Syed Md Shamsul Alam, Nobel Dhar, Kun Suo, Dan Chia-Tien Lo
ICTAI2
2025 Assessing and Visualizing Completeness, Co-Coverage, and Scalability in Multivariate Time-Series Data
abstract
Assessing data quality in multivariate time-series datasets is crucial for reliable analysis, particularly when dealing with missing values, inconsistent feature availability, and massive records in large-scale edge computing and IoT clusters. Existing methods often fall short of capturing intricate patterns of missingness and co-coverage, restricting the capacity to make well-informed decisions regarding the usability of the data. In order to systematically extract reliable data segments, this paper presents a comprehensive framework that combines a heuristic model with temporal coverage, period-specific missingness, and co-coverage metrics. By integrating these metrics with visualizations such as temporal coverage heatmaps and parallel coordinates plots, the framework reveals complex patterns of missingness while supporting human involvement in validating data subsets. Our approach effectively balances automation with expert judgment, enhancing the interpretability of data quality assessments. The findings show that the proposed methods satisfy the design specifications for revealing patterns, quantifying missingness impact, measuring feature availability, guiding feature selection, and facilitating scalable, multi-scale data summarization. The framework offers a solid way to improve the quality of data in multivariate time-series analysis, opening the door to more precise and trustworthy insights for assessing data gathered from edge computing infrastructures and large-scale, heterogeneous IoT deployments, where data consistency and completeness are frequently very variable.
Long Vu, Madeline Frank, Honghui Xu 0001, Sisi Chen, Tu N. Nguyen 0001, Selena He, Bobin Deng, Kun Suo
IPCCC7
2025 Characterizing and Understanding Energy Footprint and Efficiency of Small Language Model on Edges
abstract
Cloud-based large language models (LLMs) and their variants have significantly influenced real-world applications. Deploying smaller models (i.e., small language models (SLMs)) on edge devices offers additional advantages, such as reduced latency and independence from network connectivity. However, edge devices’ limited computing resources and constrained energy budgets challenge efficient deployment. This study evaluates the power efficiency of five representative SLMs — Llama 3.2, Phi-3 Mini, TinyLlama, and Gemma 2 on Raspberry Pi 5, Jetson Nano, and Jetson Orin Nano (CPU and GPU configurations). Results show that Jetson Orin Nano with GPU acceleration achieves the highest energy-to-performance ratio, significantly outperforming CPU-based setups. Llama 3.2 provides the best balance of accuracy and power efficiency, while TinyLlama is well-suited for low-power environments at the cost of reduced accuracy. In contrast, Phi-3 Mini consumes the most energy despite its high accuracy. In addition, GPU acceleration, memory bandwidth, and model architecture are key in optimizing inference energy efficiency. Our empirical analysis offers practical insights for AI, smart systems, and mobile ad-hoc platforms to leverage tradeoffs from accuracy, inference latency, and power efficiency in energy-constrained environments.
Md. Romyull Islam, Bobin Deng, Nobel Dhar, Tu N. Nguyen 0001, Selena He, Yong Shi 0002, Kun Suo
MASS2
2025 FL-SNNs: Benchmarking the Byzantine-Robustness of Uniquely-Shaped Surrogate Gradients
abstract
The rise of Edge AI necessitates energy-efficient models like Spiking Neural Networks (SNNs), often trained using Federated Learning (FL) to preserve data privacy. However, FL is vulnerable to Byzantine attacks, where malicious clients disrupt training. While SNNs offer potential energy benefits due to their event-driven nature, their unique training mechanisms, particularly the use of surrogate gradients to handle non-differentiable spike events, raise questions about their inherent robustness in adversarial FL settings. We evaluate the robustness of SNNs employing 5 surrogate gradients (distinct by function shape) against 7 diverse Byzantine attacks and assess recovery potential using 5 robust aggregation rules (AGRs). Our extensive experiments (1032 runs) reveal that SNNs are not universally more robust than ANNs; they show resilience to certain structured attacks (e.g., MinMax) but vulnerability to others (e.g., Label Flip). We find a moderate positive correlation between surrogate gradient choice and recovery effectiveness using AGRs, with Triangle and Rectangle surrogates often enabling better recovery, though this advantage is context-dependent. Our results underscore that robust AGRs (like DnC and RFA) are essential for mitigating attacks in SNN-based FL, regardless of the surrogate gradient used. We conclude that achieving reliable SNN deployment in adversarial FL requires a holistic, context-aware approach, carefully considering the interplay between network type, surrogate gradient, threat model, and defense mechanisms. Our code is open-sourced for reproducibility1.
Manh V. Nguyen, Liang Zhao 0024, Bobin Deng, Shaoen Wu
MASS3
2025 Sparsified Federated Learning With Spiking Neural Networks: Resistance Against Byzantine Attacks While Lowering Communication Traffics
abstract
Spiking Neural Networks (SNN), which offer exceptional energy efficiency for inference, and Federated Learning (FL), which offers privacy-preserving training, is a rising area of interest that highly beneficial towards Internet of Things (IoT) devices. Despite this, research that tackles Byzantine attacks and bandwidth limitation in FL-SNN, both poses significant threats on model convergence and training times, still remains largely unexplored. In this paper, we first systematically evaluate the robustness of ANN and SNN in the FL context under four model-poisoning Byzantine attacks. We find that FL-SNN demonstrate better reliability than FL-ANN against most Byzantine attacks except MinMax. We then propose the$Top-\kappa$sparsification approach for better robustness in FL-SNN and to reduce communication overhead. Using this simple compression method, we observe ~40% accuracy enhancement in FL-SNN training under the lethal MinMax attack, leading to FL-SNN being more robust than FL-ANN in all four model-poisoning Byzantine attacks. This study highlights the dual benefits of FL-SNN with$Top-\kappa$sparsification in significantly reduce energy consumption and provide better robustness to Byzantine attacks for edge AI applications.
Manh V. Nguyen, Liang Zhao 0024, Bobin Deng, Shaoen Wu
VTC2025-Spring3
2024 Practical Considerations of Fully Homomorphic Encryption in Privacy-Preserving Machine Learning
abstract
Machine learning has been successfully applied to big data analytics across various disciplines. However, as data is collected from diverse sectors, much of it is private and confidential. At the same time, one of the major challenges in machine learning is the slow training speed of large models, which often requires high-performance servers or cloud services. To protect data privacy while still allowing model training on such servers, privacy-preserving machine learning using Fully Homomorphic Encryption (FHE) has gained significant attention. However, its widespread adoption is hindered by performance degradation. This paper presents our experiments on training models over encrypted data using FHE. The results show that while FHE ensures privacy, it can significantly degrade performance, requiring complex tuning to optimize.
Dan Chia-Tien Lo, Yong Shi 0002, Hossain Shahriar, Bobin Deng, Xinyue Zhang 0001, Mei-Lan Chen
IEEE Big Data4
2024 Predicting Protein-Protein Binding Affinity with Deep Learning: A Comparative Analysis of CNN and Transformer Models
abstract
Binding affinity (BA) prediction is important for drug discovery and protein engineering. It seeks to understand the interaction strength between proteins and their ligands (or proteins). This information assists in the design of proteins with enhanced or novel functions, as well as understanding the molecular mechanisms of drug action. This paper presents the development and comparative analysis of two deep learning models, a convolutional neural network (CNN) and a transformer model. Many variants of models in this research were developed using TensorFlow. One model that utilizes ProteinBERT was developed using PyTorch. The CNN model captures local sequence features effectively, while the Transformer model leverages self-attention mechanisms to learn long-range dependencies within the sequences. Protein sequences are the inputs for the models. The sequences are processed using various encoders, like One-hot encoding, Sequence-Statistics-Content, and Position Specific Scoring Matrix. The predicted outputs are Gibbs free energy changes, a key indicator of binding affinity. From this study, both the CNN and transformer models can achieve the same level of accuracy under different conditions. For the CNN model, it can handle full data without sacrificing performance, but it takes much more time to preprocess the features from the protein sequences. The transformer model can achieve the same level of accuracy as the CNN model with no big predictive errors for each protein, but it requires the model to run on less data, which removes some rarely long protein sequences. This study emphasizes the potential of advanced deep learning architectures to enhance the predictive strengths of binding affinity models.
Lingtao Chen, Kazi Fahim Ahmad Nasif, Bobin Deng, Shuteng Niu, Chloe Yixin Xie
ICTAI3
2024 Activation Sparsity Opportunities for Compressing General Large Language Models
abstract
Deploying local AI models, such as Large Language Models (LLMs), to edge devices can substantially enhance devices’ independent capabilities, alleviate the server’s burden, and lower the response time. Owing to these tremendous potentials, many big tech companies have been actively promoting edge LLM evolution and released several lightweight Small Language Models (SLMs) to bridge this gap. However, SLMs currently only work well on limited real-world applications. We still have huge motivations to deploy more powerful (larger-scale) AI models on edge devices and enhance their smartness level. Unlike the conventional approaches for AI model compression, we investigate from activation sparsity. The activation sparsity method is orthogonal and combinable with existing techniques to maximize compression rate while maintaining great accuracy. According to statistics of open-source LLMs, their Feed-Forward Network (FFN) components typically comprise a large proportion of parameters (around $\tfrac{2}{3}$). This internal feature ensures that our FFN optimizations would have a better chance of achieving effective compression. Moreover, our findings are beneficial to general LLMs and are not restricted to ReLU-based models.This work systematically investigates the tradeoff between enforcing activation sparsity and perplexity (accuracy) on state-of-the-art LLMs. Our empirical analysis demonstrates that we can obtain around 50% of main memory and computing reductions for critical FFN components with negligible accuracy degradation. This extra 50% sparsity does not naturally exist in the current LLMs, which require tuning LLMs’ activation outputs by injecting zero-enforcing thresholds. To obtain the benefits of activation sparsity, we provide a guideline for the system architect for LLM prediction and prefetching. Moreover, we further verified the predictability of activation patterns in recent LLMs. The success prediction allows the system to prefetch the necessary weights while omitting the inactive ones and their successors (compress models from the memory’s perspective), therefore lowering cache/memory pollution and reducing LLM execution time on resource-constraint edge devices.
Nobel Dhar, Bobin Deng, Md. Romyull Islam, Kazi Fahim Ahmad Nasif, Liang Zhao 0024, Kun Suo
IPCCC2
2024 Characterizing and Understanding the Performance of Small Language Models on Edge Devices
abstract
In recent years, significant advancements in computing power, data richness, algorithmic development, and the growing demand for applications have catalyzed the rapid emergence and proliferation of large language models (LLMs) across various scenarios. Concurrently, factors such as computing resource limitations, cost considerations, real-time application requirements, task-specific customization, and privacy concerns have also driven the development and deployment of small language models (SLMs). Unlike extensively researched and widely deployed LLMs in the cloud, the performance of SLM workloads and their resource impact on edge environments remain poorly understood. More detailed studies will have to be carried out to understand the advantages, constraints, performances, and resource consumption in different settings of the edge.This paper addresses this gap by comprehensively analyzing representative SLMs on edge platforms. Initially, we provide a summary of contemporary edge hardware and popular SLMs. Subsequently, we quantitatively evaluate several widely used SLMs, including TinyLlama, Phi-3, Llama-3, etc., on popular edge platforms such as Raspberry Pi, Nvidia Jetson Orin, and Mac mini. Our findings reveal that the interaction between different hardware and SLMs can significantly impact edge AI workloads while introducing non-negligible overhead. Our experiments demonstrate that variations in performance and resource usage might constrain the workload capabilities of specific models and their feasibility on edge platforms. Therefore, users must judiciously match appropriate hardware and models based on the requirements and characteristics of the edge environment to avoid performance bottlenecks and optimize the utility of edge computing capabilities.
Md. Romyull Islam, Nobel Dhar, Bobin Deng, Tu N. Nguyen 0001, Selena He, Kun Suo
IPCCC3
2024 The Robustness of Spiking Neural Networks in Communication and its Application towards Network Efficiency in Federated Learning
abstract
Spiking Neural Networks (SNNs) have recently gained significant interest in on-chip learning in embedded devices and emerged as an energy-efficient alternative to conventional Artificial Neural Networks (ANNs). However, to extend SNNs to a Federated Learning (FL) setting involving collaborative model training, the communication between the local devices and the remote server remains the bottleneck, which is often restricted and costly. In this paper, we first explore the inherent robustness of SNNs under noisy communication in FL. Building upon this foundation, we propose a novel Federated Learning with Top-κ Sparsification (FLTS) algorithm to reduce the bandwidth usage for FL training. We discover that the proposed scheme with SNNs allows more bandwidth savings compared to ANNs without impacting the model’s accuracy. Additionally, the number of parameters to be communicated can be reduced to as low as 6% of the size of the original model. We further improve the communication efficiency by enabling dynamic parameter compression during model training. Extensive experiment results demonstrate that our proposed algorithms significantly outperform the baselines in terms of communication cost and model accuracy and are promising for practical network-efficient FL with SNNs.
Manh V. Nguyen, Liang Zhao 0024, Bobin Deng, William Severa, Honghui Xu 0001, Shaoen Wu
IPCCC3
2024 Fixed-point Encoding and Architecture Exploration for Residue Number Systems
abstract
Residue Number Systems (RNS) demonstrate the fascinating potential to serve integer addition/ multiplication-intensive applications. The complexity of Artificial Intelligence (AI) models has grown enormously in recent years. From a computer system’s perspective, ensuring the training of these large-scale AI models within an adequate time and energy consumption has become a big concern. Matrix multiplication is a dominant subroutine in many prevailing AI models, with an addition/multiplication-intensive attribute. However, the data type of matrix multiplication within machine learning training typically requires real numbers, which indicates that RNS benefits for integer applications cannot be directly gained by AI training. The state-of-the-art RNS real-number encodings, including floating-point and fixed-point, have defects and can be further enhanced. To transform default RNS benefits to the efficiency of large-scale AI training, we propose a low-cost and high-accuracy RNS fixed-point representation: Single RNS Logical Partition (S-RNS-Logic-P) representation with Scaling-down Postprocessing Multiplication (SD-Post-Mul) . Moreover, we extend the implementation details of the other two RNS fixed-point methods: Double RNS Concatenation and S-RNS-Logic-P representation with Scaling-down Preprocessing Multiplication . We also design the architectures of these three fixed-point multipliers. In empirical experiments, our S-RNS-Logic-P representation with SD-Post-Mul method achieves less latency and energy overhead while maintaining good accuracy. Furthermore, this method can easily extend to the Redundant Residue Number System to raise the efficiency of error-tolerant domains, such as improving the error correction efficiency of quantum computing.
Bobin Deng, Bhargava Nadendla, Kun Suo, Chloe Yixin Xie, Dan Chia-Tien Lo
ACM Trans. Archit. Code Optim.1
2023 Deep Machine Learning on Segmenting and Classifying Crop Images Taken by Unmanned Aerial Vehicle
abstract
In the realm of precision agriculture, a crucial element involves the precise quantification or estimation of seedlings, fruits, and other agricultural produce on expansive multi-acre farms at various stages of cultivation. With the advent of unmanned aerial vehicles (UAVs), capturing images of watermelon fields has become a straightforward task. These images can be subsequently processed, segmented, and categorized to determine the total count of watermelons. Currently, conventional methods are employed to address this challenge, but they have their limitations. The field has benefited from the evolution of machine learning, which has the potential to streamline the process. Nevertheless, the training phase is intricate, and achieving a valuable model can be demanding. This research delves into an examination and presentation of the existing pre-trained models for image processing in this context.
Dan Chia-Tien Lo, Bobin Deng, Yong Shi 0002
IEEE Big Data2
2022 Scalable Energy-Efficient Microarchitectures With Computational Error Tolerance Via Redundant Residue Number Systems
abstract
Due to high leakage current and threshold voltage, Dennard scaling has reached its limit on conventional semiconductor technology. Energy reduction at the transistor level by simply lowering supply voltage has proven to be infeasible for these devices (e.g., MOSFETs). Some recently proposed millivolt switch techniques aim to mitigate these issues, by maintaining a high on/off ratio of drain currents with a much lower supply voltage. However,$V_{dd}$reduction is constrained by high intermittent error probabilities in millivolt switches. Energy-efficient microarchitectures that are computationally error-tolerant are therefore urgently needed. This article systematically leverages the error correction and checkpointing properties of Redundant Residue Number Systems (RRNS) by varying the number of non-redundant ($n$) and redundant ($r$) residues. The state-of-the-art of RRNS microarchitecture is confined to a fixed configuration point within such a($n$n,$r$r)-RRNSdesign plane, as it supports single error correction alone. Being able to efficiently handle resilience in this($n$n,$r$r)-RRNSplane significantly improves reliability, allowing further${V_{dd}}$reduction to save energy. To this end, first, we propose a scalable RRNS microarchitecture that simultaneously supports both, error-correction, as well as checkpointing with restart capabilities upon detecting uncorrectable errors. Second, we design a novel RRNS-based adaptive checkpointing&restart mechanisms that automatically guarantees reliability while minimizing the energy-delay product (EDP). To the best of our knowledge, these are the first set of checkpointing mechanisms targeting the RRNS infrastructure. Moreover, these mechanisms optimize the usage efficiency of memory capacity. Third, we systematically explore the RRNS design space to find the best ($n$,$r$) configuration point. For similar reliability when compared to a conventional binary core without computationally error-tolerant (runs at high$V_{dd}$), the proposed RRNS scalable microarchitecture reduces EDP by 53 percent on average for memory-intensive workloads and by 67 percent on average for non-memory-intensive workloads.
Bobin Deng, Sriseshan Srikanth, Anirudh Jain, Thomas M. Conte, Erik DeBenedictis, Jeanine E. Cook
IEEE Trans. Computers1
2018 Memory System Design for Ultra Low Power, Computationally Error Resilient Processor Microarchitectures
abstract
Dennard scaling ended a decade ago. Energy reduction by lowering supply voltage has been limited because of guard bands and a subthreshold slope of over 60mV/decade in MOSFETs. On the other hand, newly-proposed logic devices maintain a high on/off ratio for drain currents even at significantly lower operating voltages. However, such ultra low power technology would eventually suffer from intermittent errors in logic as a result of operating close to the thermal noise floor. Computational error correction mitigates this issue by efficiently correcting stochastic bit errors that may occur in computational logic operating at low signal energies, thereby allowing for energy reduction by lowering supply voltage to tens of millivolts. Cores based on a Redundant Residual Number System (RRNS), which represents a number using a tuple of smaller numbers, are a promising candidate for implementing energyefficient computational error correction. However, prior RRNS core microarchitectures abstract away the memory hierarchy and do not consider the power-performance impact of RNS-based memory addressing. When compared with a non-error-correcting core addressing memory in binary, naive RNS-based memory addressing schemes cause a slowdown of over 3x/2x for inorder/out-of-order cores respectively. In this paper, we analyze RNS-based memory access pattern behavior and provide solutions in the form of novel schemes and the resulting design space exploration, thereby, extending and enabling a tangible, ultra low power RRNS based architecture.
Sriseshan Srikanth, Paul G. Rabbat, Eric R. Hein, Bobin Deng, Thomas M. Conte, Erik DeBenedictis, Jeanine E. Cook, Michael P. Frank
HPCA4
2018 Extending Moore's Law via Computationally Error-Tolerant Computing
abstract
Dennard scaling has ended. Lowering the voltage supply ( V dd ) to sub-volt levels causes intermittent losses in signal integrity, rendering further scaling (down) no longer acceptable as a means to lower the power required by a processor core. However, it is possible to correct the occasional errors caused due to lower V dd in an efficient manner and effectively lower power. By deploying the right amount and kind of redundancy, we can strike a balance between overhead incurred in achieving reliability and energy savings realized by permitting lower V dd . One promising approach is the Redundant Residue Number System (RRNS) representation. Unlike other error correcting codes, RRNS has the important property of being closed under addition, subtraction and multiplication, thus enabling computational error correction at a fraction of an overhead compared to conventional approaches. We use the RRNS scheme to design a Computationally-Redundant, Energy-Efficient core, including the microarchitecture, Instruction Set Architecture (ISA) and RRNS centered algorithms. From the simulation results, this RRNS system can reduce the energy-delay-product by about 3× for multiplication intensive workloads and by about 2× in general, when compared to a non-error-correcting binary core.
Bobin Deng, Sriseshan Srikanth, Eric R. Hein, Thomas M. Conte, Erik DeBenedictis, Jeanine E. Cook, Michael P. Frank
ACM Trans. Archit. Code Optim.1
2014 Efficient execution of speculative threads and transactions with hardware transactional memory
Gongming Li, Hong An, Qi Li 0034, Bobin Deng, Wenbo Dai
Future Gener. Comput. Syst.4
2012 SeTM: Efficient Execution of Speculative Threads with Hardware Transactional Memory
abstract
Thread-Level Speculation (TLS) was researched to automatically parallelize portions of serial programs for execution, and transactional memory (TM) was studied as a promising alternative of lock for parallel programming due to its simplicity. Both TLS and TM require similar underlying support. In the paper, we present SeTM (Sequential Transactional Memory), a hardware enhanced TM system which supports TLS at minor extra cost. Signature is an effective way to buffer speculative states in TM and TLS. But it cripples TM and TLS performance due to its false-positive in terms of conflict detection, especially for conflict-intensive TLS. SeTM adopts R/W bits and signature concurrently to ameliorate this bad influence. Additionally, SeTM introduces fast rollback mechanism, which provides fast abort recovery for eager log-based HTM and TLS. The most important contribution of SeTM is conflict-tolerant mechanism, which tolerates some ambiguous data conflicts in TLS. Six representative benchmarks have been adopted to evaluate our model. Our experimental results show that our scheme improves the execution performance of most tested codes at a modest hardware cost. For a set of important scientific loops, we report the highest speedup of 6.5 with 15 cores. Besides, experimental results also show good scalability of SeTM system.
Gongming Li, Hong An, Qi Li 0034, Bobin Deng, Wenbo Dai
ICPADS4
2012 Distributed replay protocol for distributed uniprocessors
abstract
Data speculation technique has been heavily exploited in various scenarios of architecture design. It bridges the time or space gap between data producer and data consumer, which gives opportunities to processors to gain significant speedups. However, large instruction windows, deep pipeline and increasing latency of on-chip communication make data misspeculation very expensive in modern processors.
Mengjie Mao, Hong An, Bobin Deng, Xuechao Wei, Wenting Han
ICS3
2011 A Priority-Aware NoC to Reduce Squashes in Thread Level Speculation for Chip Multiprocessors
abstract
Thread Level Speculation (TLS) is a technique aims at boosting the performance of sequential programs running on Chip Multiprocessors (CMPs) by automatically parallelizing them. It exempts programmers from the heavy task of parallel programming. But its performance may suffer from frequent squashing caused by inter-thread data dependency violation. In this paper, we propose a Network-on-Chip (NoC) in CMP that employs a priority-aware packet arbitration policy. Packet scheduling guided by such policy reduces the occurrence of TLS squashes. Simulation results with 5 applications show that our policy reduces squashes by 22% in best case and 15% on average. Moreover, our priority aware approach could be generalized to similar scenarios in which different threads running on CMP manifest different priorities.
Wenbo Dai, Hong An, Qi Li 0034, Gongming Li, Bobin Deng, Shilei Wu
ISPA5