EDBT 2026 Demo / reviewers in the wild / expert
Bita Darvish Rouhani
dblp:157/9224 · also Bita Darvish Rohani, Bita Rouhani
· DBLP profile ↗
25ranked-venue papers
14as first author
3since 2021 · last 2025
0000-0002-8412-4320ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 18 · 13 first-author · 1 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Security and privacy · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
10 papers |
Hardware accelerators and domain-specific architectures · 63% Reconfigurable computing and FPGAs · 18% Parallel and multicore computing · 11% | |
| Network and information security
5 papers |
Security and privacy of machine learning · 55% Cryptographic protocols and secure computation · 18% Hardware security and side channels · 18% | |
| Artificial intelligence
7 papers |
Efficient and distributed learning · 64% Language models and text generation · 22% Deep learning architectures and training · 8% | |
| Databases, data mining, and information retrieval
1 paper |
Machine learning and data management · 100% |
Topics — the 28 heaviest of 33, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures
machine learning accelerator |
2.2 | 6 | 2023 | With Shared Microexponents, A Little Shifting Goes a Long Way · ISCA 2023 Pushing the Limits of Narrow Precision Inferencing at Cloud Scale with Microsoft Floating Point · NeurIPS 2020 Deepsecure: scalable provably-secure deep learning · DAC 2018 |
Machine learning › Efficient and distributed learning
model compression |
1.3 | 2 | 2025 | ResMoE: Space-efficient Compression of Mixture of Experts LLMs via Residual Restoration · KDD (1) 2025 Pushing the Limits of Narrow Precision Inferencing at Cloud Scale with Microsoft Floating Point · NeurIPS 2020 |
Natural language and speech › Language models and text generation
large language model inference |
0.9 | 1 | 2025 | ResMoE: Space-efficient Compression of Mixture of Experts LLMs via Residual Restoration · KDD (1) 2025 |
Machine learning › Efficient and distributed learning › model compression › parameter compression
mixture-of-experts compression |
0.9 | 1 | 2025 | ResMoE: Space-efficient Compression of Mixture of Experts LLMs via Residual Restoration · KDD (1) 2025 |
Reconfigurable computing and FPGAs
FPGA accelerator |
0.7 | 2 | 2018 | CausaLearn: Automated Framework for Scalable Streaming-based Causal Bayesian Learning using FPGAs · FPGA 2018 MAXelerator: FPGA accelerator for privacy preserving multiply-accumulate (MAC) on cloud servers · DAC 2018 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
low-precision arithmetic |
0.7 | 1 | 2023 | With Shared Microexponents, A Little Shifting Goes a Long Way · ISCA 2023 |
Security and privacy of machine learning
adversarial defense |
0.5 | 1 | 2021 | CuRTAIL: ChaRacterizing and Thwarting AdversarIal Deep Learning · IEEE Trans. Dependable Secur. Comput. 2021 |
Security and privacy of machine learning › adversarial defense
adversarial example detection |
0.5 | 1 | 2021 | CuRTAIL: ChaRacterizing and Thwarting AdversarIal Deep Learning · IEEE Trans. Dependable Secur. Comput. 2021 |
Security and privacy of machine learning
adversarial robustness |
0.5 | 1 | 2021 | CuRTAIL: ChaRacterizing and Thwarting AdversarIal Deep Learning · IEEE Trans. Dependable Secur. Comput. 2021 |
Cryptographic protocols and secure computation
garbled circuits |
0.4 | 2 | 2018 | Deepsecure: scalable provably-secure deep learning · DAC 2018 MAXelerator: FPGA accelerator for privacy preserving multiply-accumulate (MAC) on cloud servers · DAC 2018 |
Security and privacy of machine learning › model intellectual property protection › model watermarking
deep neural network watermarking |
0.4 | 1 | 2019 | DeepSigns: An End-to-End Watermarking Framework for Ownership Protection of Deep Neural Networks · ASPLOS 2019 |
Hardware security and side channels
intellectual property protection |
0.4 | 1 | 2019 | DeepAttest: an end-to-end attestation framework for deep neural networks · ISCA 2019 |
Security and privacy of machine learning › model intellectual property protection
model watermarking |
0.4 | 1 | 2019 | DeepSigns: An End-to-End Watermarking Framework for Ownership Protection of Deep Neural Networks · ASPLOS 2019 |
Hardware security and side channels
trusted execution environments |
0.4 | 1 | 2019 | DeepAttest: an end-to-end attestation framework for deep neural networks · ISCA 2019 |
Digital forensics and information hiding
watermarking |
0.4 | 1 | 2019 | DeepSigns: An End-to-End Watermarking Framework for Ownership Protection of Deep Neural Networks · ASPLOS 2019 |
Cryptographic protocols and secure computation
secure multiparty computation |
0.3 | 1 | 2018 | Deepsecure: scalable provably-secure deep learning · DAC 2018 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › trustworthy machine learning accelerator
privacy-preserving machine learning accelerator |
0.3 | 1 | 2018 | MAXelerator: FPGA accelerator for privacy preserving multiply-accumulate (MAC) on cloud servers · DAC 2018 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN training |
0.3 | 1 | 2017 | Deep3: Leveraging Three Levels of Parallelism for Efficient Deep Learning · DAC 2017 |
Parallel and multicore computing › parallelization strategies
model and data parallelism |
0.3 | 1 | 2017 | Deep3: Leveraging Three Levels of Parallelism for Efficient Deep Learning · DAC 2017 |
Parallel and multicore computing
parallel programming models and runtimes |
0.3 | 1 | 2017 | Deep3: Leveraging Three Levels of Parallelism for Efficient Deep Learning · DAC 2017 |
Machine learning and data management
scalable machine learning |
0.2 | 1 | 2015 | Flexible Transformations For Learning Big Data · SIGMETRICS 2015 |
Machine learning and data management
sparse representation |
0.2 | 1 | 2015 | Flexible Transformations For Learning Big Data · SIGMETRICS 2015 |
Distributed systems
distributed data processing |
0.2 | 1 | 2015 | Flexible Transformations For Learning Big Data · SIGMETRICS 2015 |
Hardware reliability and fault tolerance › redundancy
modular redundancy |
0.1 | 1 | 2021 | CuRTAIL: ChaRacterizing and Thwarting AdversarIal Deep Learning · IEEE Trans. Dependable Secur. Comput. 2021 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference |
0.1 | 1 | 2018 | CausaLearn: Automated Framework for Scalable Streaming-based Causal Bayesian Learning using FPGAs · FPGA 2018 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
markov chain monte carlo |
0.1 | 1 | 2018 | CausaLearn: Automated Framework for Scalable Streaming-based Causal Bayesian Learning using FPGAs · FPGA 2018 |
Security and privacy of machine learning
privacy-preserving inference |
0.1 | 1 | 2018 | Deepsecure: scalable provably-secure deep learning · DAC 2018 |
Machine learning › Efficient and distributed learning
distributed training |
0.1 | 1 | 2017 | Deep3: Leveraging Three Levels of Parallelism for Efficient Deep Learning · DAC 2017 |
Methods — techniques the papers use, named apart from their topics
quantization · 2.2yao's garbled circuit · 1.3block floating point · 1.3unsupervised learning · 1.0modular redundancy training · 1.0wasserstein barycenter · 0.9residual approximation · 0.9microsoft floating point · 0.9probability density function estimation · 0.8fingerprint embedding · 0.8activation map analysis · 0.8preprocessing optimization · 0.7logic synthesis · 0.7automated design space customization · 0.7hamiltonian markov chain monte carlo · 0.3HW/SW co-design · 0.3sparse representation · 0.2domain-specific transformation · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ResMoE: Space-efficient Compression of Mixture of Experts LLMs via Residual RestorationabstractMixture-of-Experts (MoE) Transformer, the backbone architecture of multiple phenomenal language models, leverages sparsity by activating only a fraction of model parameters for each input token. The sparse structure, while allowing constant time costs, results in space inefficiency: we still need to load all the model parameters during inference. We introduce ResMoE, an innovative MoE approximation framework that utilizes Wasserstein barycenter to extract a common expert (barycenter expert) and approximate the residuals between this barycenter expert and the original ones. ResMoE enhances the space efficiency for inference of large-scale MoE Transformers in a one-shot and data-agnostic manner without retraining while maintaining minimal accuracy loss, thereby paving the way for broader accessibility to large language models. We demonstrate the effectiveness of ResMoE through extensive experiments on Switch Transformer, Mixtral, and DeepSeekMoE models. The results show that ResMoE can reduce the number of parameters in an expert by up to 75% while maintaining comparable performance. The code is available at https://github.com/iDEA-iSAIL-Lab-UIUC/ResMoE, and the supplementary appendix is available at https://famous-blue-raincoat.github.io/mengtingai/files/ResMoE_Appendix.pdf. Mengting Ai, Tianxin Wei, Yifan Chen 0004, Zhichen Zeng 0001, Ritchie Zhao, Girish Varatkar, Bita Darvish Rouhani, Xianfeng Tang, Hanghang Tong, Jingrui He |
KDD (1) | 7 |
| 2023 | With Shared Microexponents, A Little Shifting Goes a Long WayabstractThis paper introduces Block Data Representations (BDR), a framework for exploring and evaluating a wide spectrum of narrow-precision formats for deep learning. It enables comparison of popular quantization standards, and through BDR, new formats based on shared microexponents (MX) are identified, which outperform other state-of-the-art quantization approaches, including narrow-precision floating-point and block floating-point. MX utilizes multiple levels of quantization scaling with ultra-fine scaling factors based on shared microexponents in the hardware. The effectiveness of MX is demonstrated on real-world models including large-scale generative pretraining and inferencing, and production-scale recommendation systems. Bita Darvish Rouhani, Ritchie Zhao, Venmugil Elango, Rasoul Shafipour, Mathew Hall, Maral Mesmakhosroshahi, Ankit More, Levi Melnick, Maximilian Golub, Girish Varatkar, Lai Shao, Gaurav Kolhe, Dimitry Melts, Jasmine Klar, Renee L'Heureux, Matt Perry, Doug Burger, Eric S. Chung, Zhaoxia Deng, Sam Naghshineh, Jongsoo Park, Maxim Naumov |
ISCA | 1 |
| 2021 | CuRTAIL: ChaRacterizing and Thwarting AdversarIal Deep LearningabstractRecent advances in adversarial Deep Learning (DL) have opened up a new and largely unexplored surface for malicious attacks jeopardizing the integrity of autonomous DL systems. This article introduces CuRTAIL, a novel end-to-end computing framework to characterize and thwart potential adversarial attacks and significantly improve the reliability (safety) of a victim DL model. We formalize the goal of preventing adversarial attacks as an optimization problem to minimize the rarely observed regions in the latent feature space spanned by a DL network. To solve the aforementioned minimization problem, a set of complementary but disjoint modular redundancies are trained to validate the legitimacy of the input samples. The proposed countermeasure is unsupervised, meaning that no adversarial sample is leveraged to train modular redundancies. This, in turn, ensures the effectiveness of the defense in the face of generic attacks. We evaluate the robustness of our proposed methodology against the state-of-the-art adaptive attacks in a white-box setting considering that the adversary knows everything about the victim model and its defenders. Extensive evaluations for analyzing MNIST, CIFAR10, and ImageNet data corroborate the effectiveness of CuRTAIL framework against adversarial samples. The computations in each modular redundancy can be performed independently of the other redundancy modules. As such, CuRTAIL detection algorithm can be completely parallelized among multiple hardware settings to achieve maximum throughput. We further provide an open-source Application Programming Interface (API) to facilitate the adoption of the proposed framework for various applications. Mojan Javaheripi, Mohammad Samragh Razlighi, Bita Darvish Rouhani, Tara Javidi, Farinaz Koushanfar |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2020 | SpecMark: A Spectral Watermarking Framework for IP Protection of Speech Recognition Systems
Huili Chen, Bita Darvish Rouhani, Farinaz Koushanfar |
INTERSPEECH | 2 |
| 2020 | Pushing the Limits of Narrow Precision Inferencing at Cloud Scale with Microsoft Floating PointabstractIn this paper, we explore the limits of Microsoft Floating Point (MSFP), a new class of datatypes developed for production cloud-scale inferencing on custom hardware. Through the co-evolution of hardware design and algorithms, MSFP achieves accuracy comparable to or better than industry standards Bfloat16 and INT8 at 3x and 4x lower cost, respectively. MSFP incurs negligible impact to accuracy (<1%), requires no changes to the model topology, and is integrated with a mature cloud production pipeline. MSFP supports various classes of deep learning models including CNNs, RNNs, and Transformers without modification. Finally, we characterize the accuracy and implementation of MSFP and demonstrate its efficacy on a number of production scenarios, including models that power major online scenarios such as web search, question-answering, and image classification. Bita Darvish Rouhani, Daniel Lo, Ritchie Zhao, Jeremy Fowers, Kalin Ovtcharov, Anna Vinogradsky, Sarah Massengill, Lita Yang, Ray Bittner, Alessandro Forin, Haishan Zhu, Taesik Na, Prerak Patel, Shuai Che, Lok Chand Koppaka, Subhojit Som, Kaustav Das, Saurabh Tiwary, Steven K. Reinhardt, Sitaram Lanka, Eric S. Chung, Doug Burger |
NeurIPS | 1 |
| 2019 | DeepSigns: An End-to-End Watermarking Framework for Ownership Protection of Deep Neural NetworksabstractDeep Learning (DL) models have created a paradigm shift in our ability to comprehend raw data in various important fields, ranging from intelligence warfare and healthcare to autonomous transportation and automated manufacturing. A practical concern, in the rush to adopt DL models as a service, is protecting the models against Intellectual Property (IP) infringement. DL models are commonly built by allocating substantial computational resources that process vast amounts of proprietary training data. The resulting models are therefore considered to be an IP of the model builder and need to be protected to preserve the owner's competitive advantage. We propose DeepSigns, the first end-to-end IP protection framework that enables developers to systematically insert digital watermarks in the target DL model before distributing the model. DeepSigns is encapsulated as a high-level wrapper that can be leveraged within common deep learning frameworks including TensorFlow and PyTorch. The libraries in DeepSigns work by dynamically learning the Probability Density Function (pdf) of activation maps obtained in different layers of a DL model. DeepSigns uses the low probabilistic regions within the model to gradually embed the owner's signature (watermark) during DL training while minimally affecting the overall accuracy and training overhead. DeepSigns can demonstrably withstand various removal and transformation attacks, including model pruning, model fine-tuning, and watermark overwriting. We evaluate DeepSigns performance on a wide variety of DL architectures including wide residual convolution neural networks, multi-layer perceptrons, and long short-term memory models. Our extensive evaluations corroborate DeepSigns' effectiveness and applicability. We further provide a highly-optimized accompanying API to facilitate training watermarked neural networks with a training overhead as low as 2.2%. Bita Darvish Rouhani, Huili Chen, Farinaz Koushanfar |
ASPLOS | 1 |
| 2019 | SemiHD: Semi-Supervised Learning Using Hyperdimensional ComputingabstractIn the Internet of Things (IoT), the large volume of data generated by sensors poses significant computational challenges in resource-constrained environments. Most existing machine learning algorithms are unable to train a proper model using a significantly small amount of labeled data available in practice. In this paper, we propose SemiHD, a novel semi-supervised algorithm based on brain-inspired HyperDimensional (HD) computing. SemiHD performs the cognitive task by emulating neuron's activity in high-dimensional space. SemiHD maps data points into high-dimensional space and trains a model based on the available labeled data. To improve the quality of the model, SemiHD iteratively expands the training data by labeling data points which can be classified by the current model with high confidence. We also proposed a framework which enables users to trade accuracy for efficiency and select the desired reliability of the model in detecting out of scope data. We have evaluated SemiHD's accuracy and efficiency on a wide range of classification applications and two types of embedded devices: Raspberry Pi 3 and Kintex-7 FPGA. Our evaluation shows that SemiHD can improve the classification accuracy of supervised HD by 10.2% on average (up to 27.3%). In addition, we observe that SemiHD FPGA implementation achieves 7.11× faster and 12.6× energy efficiency as compared to the CPU implementation. Mohsen Imani, Samuel Bosch, Mojan Javaheripi, Bita Darvish Rouhani, Farinaz Koushanfar, Tajana Rosing |
ICCAD | 4 |
| 2019 | DeepAttest: an end-to-end attestation framework for deep neural networksabstractEmerging hardware architectures for Deep Neural Networks (DNNs) are being commercialized and considered as the hardware-level Intellectual Property (IP) of the device providers. However, these intelligent devices might be abused and such vulnerability has not been identified. The unregulated usage of intelligent platforms and the lack of hardware-bounded IP protection impair the commercial advantage of the device provider and prohibit reliable technology transfer. Our goal is to design a systematic methodology that provides hardware-level IP protection and usage control for DNN applications on various platforms. To address the IP concern, we present DeepAttest, the first on-device DNN attestation method that certifies the legitimacy of the DNN program mapped to the device. DeepAttest works by designing a device-specific fingerprint which is encoded in the weights of the DNN deployed on the target platform. The embedded fingerprint (FP) is later extracted with the support of the Trusted Execution Environment (TEE). The existence of the pre-defined FP is used as the attestation criterion to determine whether the queried DNN is authenticated. Our attestation framework ensures that only authorized DNN programs yield the matching FP and are allowed for inference on the target device. DeepAttest provisions the device provider with a practical solution to limit the application usage of her manufactured hardware and prevents unauthorized or tampered DNNs from execution. Huili Chen, Cheng Fu 0002, Bita Darvish Rouhani, Jishen Zhao, Farinaz Koushanfar |
ISCA | 3 |
| 2019 | DeepMarks: A Secure Fingerprinting Framework for Digital Rights Management of Deep Learning ModelsabstractDeep Neural Networks (DNNs) are revolutionizing various critical fields by providing an unprecedented leap in terms of accuracy and functionality. Due to the costly training procedure, high-performance DNNs are typically considered as the Intellectual Property (IP) of the model builder and need to be protected. While DNNs are increasingly commercialized, the pre-trained models might be illegally copied or redistributed after they are delivered to malicious users. In this paper, we introduce DeepMarks, the first end-to-end collusion-secure fingerprinting framework that enables the owner to retrieve model authorship information and identification of unique users in the context of deep learning (DL). DeepMarks consists of two main modules: (i) Designing unique fingerprints using anti-collusion codebooks for individual users; and (ii) Encoding each constructed fingerprint (FP) in the probability density function (pdf) of the weights by incorporating an FP-specific regularization loss during DNN re-training. We investigate the performance of DeepMarks on various datasets and DNN architectures. Experimental results show that the embedded FP preserves the accuracy of the host DNN and is robust against different model modifications that might be conducted by the malicious user. Furthermore, our framework is scalable and yields perfect detection rates and no false alarms when identifying the participants of FP collusion attacks under theoretical guarantee. The runtime overhead of retrieving the embedded FP from the marked DNN can be as low as 0.056%. Huili Chen, Bita Darvish Rouhani, Cheng Fu 0002, Jishen Zhao, Farinaz Koushanfar |
ICMR | 2 |
| 2018 | MAXelerator: FPGA accelerator for privacy preserving multiply-accumulate (MAC) on cloud serversabstractThis paper presents MAXelerator, the first hardware accelerator for privacy-preserving machine learning (ML) on cloud servers. Cloud-based ML is being increasingly employed in various data sensitive scenarios. While it enhances both efficiency and quality of the service, it also raises concern about privacy of the users' data. We create a practical privacy-preserving solution for matrix-based ML on cloud servers. We show that for the majority of the ML applications, the privacy-sensitive computation boils down to either matrix multiplication, which is a repetition of Multiply-Accumulate (MAC) or the MAC itself. We design an FPGA architecture for privacy-preserving MAC to accelerate the ML computation based on the well known Secure Function Evaluation protocol named Yao's Garbled Circuit. MAXelerator demonstrates up to 57× improvement in throughput per core compared to the fastest existing GC framework. We corroborate the effectiveness of the accelerator with real-world case studies in privacy-sensitive scenarios. Siam U. Hussain, Bita Darvish Rouhani, Mohammad Ghasemzadeh 0002, Farinaz Koushanfar |
DAC | 2 |
| 2018 | Deepsecure: scalable provably-secure deep learningabstractThis paper presents DeepSecure, the an scalable and provably secure Deep Learning (DL) framework that is built upon automated design, efficient logic synthesis, and optimization methodologies. DeepSecure targets scenarios in which neither of the involved parties including the cloud servers that hold the DL model parameters or the delegating clients who own the data is willing to reveal their information. Our framework is the first to empower accurate and scalable DL analysis of data generated by distributed clients without sacrificing the security to maintain efficiency. The secure DL computation in DeepSecure is performed using Yao's Garbled Circuit (GC) protocol. We devise GC-optimized realization of various components used in DL. Our optimized implementation achieves up to 58-fold higher throughput per sample compared with the best prior solution. In addition to the optimized GC realization, we introduce a set of novel low-overhead pre-processing techniques which further reduce the GC overall runtime in the context of DL. Our extensive evaluations demonstrate up to two orders-of-magnitude additional runtime improvement achieved as a result of our pre-processing methodology. Bita Darvish Rouhani, M. Sadegh Riazi, Farinaz Koushanfar |
DAC | 1 |
| 2018 | CausaLearn: Automated Framework for Scalable Streaming-based Causal Bayesian Learning using FPGAsabstractThis paper proposes CausaLearn, the first automated framework that enables real-time and scalable approximation of Probability Density Function (PDF) in the context of causal Bayesian graphical models. CausaLearn targets complex streaming scenarios in which the input data evolves over time and independence cannot be assumed between data samples (e.g., continuous time-varying data analysis). Our framework is devised using a HW/SW co-design approach. We provide the first implementation of Hamiltonian Markov Chain Monte Carlo on FPGA that can efficiently sample from the steady state probability distribution at scales while considering the correlation between the observed data. CausaLearn is customizable to the limits of the underlying resource provisioning in order to maximize the effective system throughput. It uses physical profiling to abstract high-level hardware characteristics. These characteristics are integrated into our automated customization unit in order to tile, schedule, and batch the PDF approximation workload corresponding to the pertinent platform resources and constraints. We benchmark the design performance for analyzing various massive time-series data on three FPGA platforms with different computational budgets. Our extensive evaluations demonstrate up to two orders-of-magnitude runtime and energy improvements compared to the best-known prior solution. We provide an accompanying API that can be leveraged by data scientists and practitioners to automate and abstract hardware design optimization. Bita Darvish Rouhani, Mohammad Ghasemzadeh 0002, Farinaz Koushanfar |
FPGA | 1 |
| 2018 | Assured deep learning: practical defense against adversarial attacksabstractDeep Learning (DL) models have been shown to be vulnerable to adversarial attacks. In light of the adversarial attacks, it is critical to reliably quantify the confidence of the prediction in a neural network to enable safe adoption of DL models in autonomous sensitive tasks (e.g., unmanned vehicles and drones). This article discusses recent research advances for unsupervised model assurance against the strongest adversarial attacks known to date and quantitatively compare their performance. Given the widespread usage of DL models, it is imperative to provide model assurance by carefully looking into the feature maps automatically learned within D1 models instead of looking back with regret when deep learning systems are compromised by adversaries. Bita Darvish Rouhani, Mohammad Samragh Razlighi, Mojan Javaheripi, Tara Javidi, Farinaz Koushanfar |
ICCAD | 1 |
| 2018 | DeepFense: online accelerated defense against adversarial deep learningabstractRecent advances in adversarial Deep Learning (DL) have opened up a largely unexplored surface for malicious attacks jeopardizing the integrity of autonomous DL systems. With the wide-spread usage of DL in critical and time-sensitive applications, including unmanned vehicles, drones, and video surveillance systems, online detection of malicious inputs is of utmost importance. We propose DeepFense, the first end-to-end automated framework that simultaneously enables efficient and safe execution of DL models. DeepFense formalizes the goal of thwarting adversarial attacks as an optimization problem that minimizes the rarely observed regions in the latent feature space spanned by a DL network. To solve the aforementioned minimization problem, a set of complementary but disjoint modular redundancies are trained to validate the legitimacy of the input samples in parallel with the victim DL model. DeepFense leverages hardware/software/algorithm co-design and customized acceleration to achieve just-in-time performance in resource-constrained settings. The proposed countermeasure is unsupervised, meaning that no adversarial sample is leveraged to train modular redundancies. We further provide an accompanying API to reduce the non-recurring engineering cost and ensure automated adaptation to various platforms. Extensive evaluations on FPGAs and GPUs demonstrate up to two orders of magnitude performance improvement while enabling online adversarial sample detection. Bita Darvish Rouhani, Mohammad Samragh Razlighi, Mojan Javaheripi, Tara Javidi, Farinaz Koushanfar |
ICCAD | 1 |
| 2018 | ReDCrypt: Real-Time Privacy-Preserving Deep Learning Inference in Clouds Using FPGAsabstractArtificial Intelligence (AI) is increasingly incorporated into the cloud business in order to improve the functionality (e.g., accuracy) of the service. The adoption of AI as a cloud service raises serious privacy concerns in applications where the risk of data leakage is not acceptable. Examples of such applications include scenarios where clients hold potentially sensitive private information such as medical records, financial data, and/or location. This article proposes ReDCrypt, the first reconfigurable hardware-accelerated framework that empowers privacy-preserving inference of deep learning models in cloud servers. ReDCrypt is well-suited for streaming (a.k.a., real-time AI) settings where clients need to dynamically analyze their data as it is collected over time without having to queue the samples to meet a certain batch size. Unlike prior work, ReDCrypt neither requires to change how AI models are trained nor relies on two non-colluding servers to perform. The privacy-preserving computation in ReDCrypt is executed using Yao’s Garbled Circuit (GC) protocol. We break down the deep learning inference task into two phases: (i) privacy-insensitive (local) computation, and (ii) privacy-sensitive (interactive) computation. We devise a high-throughput and power-efficient implementation of GC protocol on FPGA for the privacy-sensitive phase. ReDCrypt’s accompanying API provides support for seamless integration of ReDCrypt into any deep learning framework. Proof-of-concept evaluations for different DL applications demonstrate up to 57-fold higher throughput per core compared to the best prior solution with no drop in the accuracy. Bita Darvish Rouhani, Siam U. Hussain, Kristin E. Lauter, Farinaz Koushanfar |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2017 | Deep3: Leveraging Three Levels of Parallelism for Efficient Deep LearningabstractThis paper proposes Deep3 an automated platform-aware Deep Learning (DL) framework that brings orders of magnitude performance improvement to DL training and execution. Deep3 is the first to simultaneously leverage three levels of parallelism for performing DL: data, network, and hardware. It uses platform profiling to abstract physical characterizations of the target platform. The core of Deep3 is a new extensible methodology that enables incorporation of platform characteristics into the higher-level data and neural network transformation. We provide accompanying libraries to ensure automated customization and adaptation to different datasets and platforms. Proof-of-concept evaluations demonstrate 10-100 fold physical performance improvement compared to the state-of-the-art DL frameworks, e.g., TensorFlow. Bita Darvish Rouhani, Azalia Mirhoseini, Farinaz Koushanfar |
DAC | 1 |
| 2017 | TinyDL: Just-in-time deep learning solution for constrained embedded systemsabstractThis work proposes TinyDL, an automated end-to-end framework that aims to integrate the state-of-the-art Deep Learning (DL) models into embedded systems. TinyDL enables efficient training and execution of DL models as data is collected over time while adhering to the underlying physical resources and constraints. The constraints can be characterized in terms of memory bandwidth, energy resources (e.g., battery life), and/or real-time requirement. TinyDL takes advantage of platform profiling to abstract the physical characteristics of the target embedded device. We introduce a platform-aware signal transformation methodology to enable DL training and execution within the confine of the available resources. Our approach balances the trade-off between data movements and computations to improve the performance of costly iterative DL training/execution. Proof-of-concept implementation on NVIDIA Jetson TK1 embedded platform demonstrates up to two orders of magnitude energy improvement over the previous DL solutions, none of which had been amenable to constrained devices. Bita Darvish Rouhani, Azalia Mirhoseini, Farinaz Koushanfar |
ISCAS | 1 |
| 2017 | RISE: An Automated Framework for Real-Time Intelligent Video Surveillance on FPGAabstractThis paper proposes RISE, an automated Reconfigurable framework for real-time background subtraction applied to Intelligent video SurveillancE. RISE is devised with a new streaming-based methodology that adaptively learns/updates a corresponding dictionary matrix from background pixels as new video frames are captured over time. This dictionary is used to highlight the foreground information in each video frame. A key characteristic of RISE is that it adaptively adjusts its dictionary for diverse lighting conditions and varying camera distances by continuously updating the corresponding dictionary. We evaluate RISE on natural-scene vehicle images of different backgrounds and ambient illuminations. To facilitate automation, we provide an accompanying API that can be used to deploy RISE on FPGA-based system-on-chip platforms. We prototype RISE for end-to-end deployment of three widely-adopted image processing tasks used in intelligent transportation systems: License Plate Recognition (LPR), image denoising/reconstruction, and principal component analysis. Our evaluations demonstrate up to 87-fold higher throughput per energy unit compared to the prior-art software solution executed on ARM Cortex-A15 embedded platform. Bita Darvish Rouhani, Azalia Mirhoseini, Farinaz Koushanfar |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2016 | Perform-ML: performance optimized machine learning by platform and content aware customizationabstractWe propose Perform-ML, the first Machine Learning (ML) framework for analysis of massive and dense data which customizes the algorithm to the underlying platform for the purpose of achieving optimized resource efficiency. Perform-ML creates a performance model quantifying the computational cost of iterative analysis algorithms on a pertinent platform in terms of FLOPs, communication, and memory, which characterize runtime, storage, and energy. The core of Perform-ML is a novel parametric data projection algorithm, called Elastic Dictionary (ExD), that enables versatile and sparse representations of the data which can help in minimizing performance cost. We show that Perform-ML can achieve the optimal performance objective, according to our cost model, by platform aware tuning of the ExD parameters. An accompanying API ensures automated applicability of Perform-ML to various algorithms, datasets, and platforms. Proof-of-concept evaluations of massive and dense data on different platforms demonstrate more than an order of magnitude improvements in performance compared to the state of the art, within guaranteed user-defined error bounds. Azalia Mirhoseini, Bita Darvish Rouhani, Ebrahim M. Songhori, Farinaz Koushanfar |
DAC | 2 |
| 2016 | DeLight: Adding Energy Dimension To Deep Neural NetworksabstractPhysical viability, in particular energy efficiency, is a key challenge in realizing the true potential of Deep Neural Networks (DNNs). In this paper, we aim to incorporate the energy dimension as a design parameter in the higher-level hierarchy of DNN training and execution to optimize for the energy resources and constraints. We use energy characterization to bound the network size in accordance to the pertinent physical resources. An automated customization methodology is proposed to adaptively conform the DNN configurations to the underlying hardware characteristics while minimally affecting the inference accuracy. The key to our approach is a new context and resource aware projection of data to a lower-dimensional embedding by which learning the correlation between data samples requires significantly smaller number of neurons. We leverage the performance gain achieved as a result of the data projection to enable the training of different DNN architectures which can be aggregated together to further boost the inference accuracy. Accompanying APIs are provided to facilitate rapid prototyping of an arbitrary DNN application customized to the underlying platform. Proof-of-concept evaluations for deployment of different visual, audio, and smart-sensing benchmarks demonstrate up to 100-fold energy improvement compared to the prior-art DL solutions. Bita Darvish Rouhani, Azalia Mirhoseini, Farinaz Koushanfar |
ISLPED | 1 |
| 2016 | Automated Real-Time Analysis of Streaming Big and Dense Data on Reconfigurable PlatformsabstractWe propose SSketch, a novel automated framework for efficient analysis of dynamic big data with dense (non-sparse) correlation matrices on reconfigurable platforms. SSketch targets streaming applications where each data sample can be processed only once and storage is severely limited. Our framework adaptively learns from the stream of input data and updates a corresponding ensemble of lower-dimensional data structures, a.k.a., a sketch matrix . A new sketching methodology is introduced that tailors the problem of transforming the big data with dense correlations to an ensemble of lower-dimensional subspaces such that it is suitable for hardware-based acceleration performed by reconfigurable hardware. The new method is scalable, while it significantly reduces costly memory interactions and enhances matrix computation performance by leveraging coarse-grained parallelism existing in the dataset. SSketch provides an automated optimization methodology for creating the most accurate data sketch for a given set of user-defined constraints, including runtime and power as well as platform constraints such as memory. To facilitate automation, SSketch takes advantage of a Hardware/Software (HW/SW) co-design approach: It provides an Application Programming Interface that can be customized for rapid prototyping of an arbitrary matrix-based data analysis algorithm. Proof-of-concept evaluations on a variety of visual datasets with more than 11 million non-zeros demonstrate up to a 200-fold speedup on our hardware-accelerated realization of SSketch compared to a software-based deployment on a general-purpose processor. Bita Darvish Rouhani, Azalia Mirhoseini, Ebrahim M. Songhori, Farinaz Koushanfar |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2015 | SSketch: An Automated Framework for Streaming Sketch-Based Analysis of Big Data on FPGAabstractThis paper proposes SSketch, a novel automated computing framework for FPGA-based online analysis of big data with dense (non-sparse) correlation matrices. SSketch targets streaming applications where each data sample can be processed only once and storage is severely limited. The stream of input data is used by SSketch for adaptive learning and updating a corresponding ensemble of lower dimensional data structures, a.k.a., A sketch matrix. A new sketching methodology is introduced that tailors the problem of transforming the big data with dense correlations to an ensemble of lower dimensional subspaces such that it is suitable for hardware-based acceleration performed by reconfigurable hardware. The new method is scalable, while it significantly reduces costly memory interactions and enhances matrix computation performance by leveraging coarse-grained parallelism existing in the dataset. To facilitate automation, SSketch takes advantage of a HW/SW co-design approach: It provides an Application Programming Interface (API) that can be customized for rapid prototyping of an arbitrary matrix-based data analysis algorithm. Proof-of-concept evaluations on a variety of visual datasets with more than 11 million non-zeros demonstrates up to 200 folds speedup on our hardware-accelerated realization of SSketch compared to a software-based deployment on a general purpose processor. Bita Darvish Rouhani, Ebrahim M. Songhori, Azalia Mirhoseini, Farinaz Koushanfar |
FCCM | 1 |
| 2015 | Flexible Transformations For Learning Big DataabstractThis paper proposes a domain-specific solution for iterative learning of big and dense (non-sparse) datasets. A large host of learning algorithms, including linear and regularized regression techniques, rely on iterative updates on the data connectivity matrix in order to converge to a solution. The performance of such algorithms often severely degrade when it comes to large and dense data. Massive dense datasets not only induce obligatory large number of arithmetics, but they also incur unwanted message passing cost across the processing nodes. Our key observation is that despite the seemingly dense structures, in many applications, data can be transformed into a new space where sparse structures become revealed. We propose a scalable data transformation scheme that enables creating versatile sparse representations of the data. The transformation can be tuned to benefit the underlying platform's cost and constraints. Our evaluations demonstrate significant improvement in energy usage, runtime, and mem Azalia Mirhoseini, Ebrahim M. Songhori, Bita Darvish Rouhani, Farinaz Koushanfar |
SIGMETRICS | 3 |
| 2015 | Agent-Oriented Based Enterprise Architecture Implementation Methodology
Babak Darvish Rouhani, Mohd Naz'ri Mahrin, Fatemeh Nikpay, Pourya Nikfard, Bita Darvish Rouhani |
WorldCIST (1) | 5 |
| 2014 | Current Issues on Enterprise Architecture Implementation Methodology
Babak Darvish Rouhani, Mohd Naz'ri Mahrin, Fatemeh Nikpay, Bita Darvish Rouhani |
WorldCIST (2) | 4 |