EDBT 2026 Demo / reviewers in the wild / expert
Maede Hemmat
dblp:191/6807 · also Maedeh Hemmat
· DBLP profile ↗
9ranked-venue papers
9as first author
3since 2021 · last 2022
0000-0002-0085-4589ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 9 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | $\text{Edge}^{n}$ AI: Distributed Inference with Local Edge Devices and Minimal LatencyabstractWe propose$\text{Edge}^{n}$AI, a framework to decompose a complex deep neural networks (DNN) over$n$available local edge devices with minimal communication overhead and overall latency. Our framework creates small DNNs (SNNs) from an original DNN by partitioning its classes across the edge devices, while taking into account their available resources. Class-aware pruning is applied to aggressively reduce the size of the SNN on each edge device. The SNNs perform inference in parallel, and are configured to generate a ‘Don't Know’ response when an unassigned class is identified. Our experiments show up to 17X inference speedup compared to a recent work, on devices of at most 150 MB memory when distributing a variant of VGG-16 over 20 parallel edge devices. Maede Hemmat, Azadeh Davoodi, Yu Hen Hu |
ASP-DAC | 1 |
| 2022 | CAP'NN: A Class-aware Framework for Personalized Neural Network InferenceabstractWe propose a framework for Class-aware Personalized Neural Network Inference (CAP’NN), which prunes an already-trained neural network model based on the preferences of individual users. Specifically, by adapting to the subset of output classes that each user is expected to encounter, CAP’NN is able to prune not only ineffectual neurons but also miseffectual neurons that confuse classification, without the need to retrain the network. CAP’NN also exploits the similarities among pruning requests from different users to minimize the timing overheads of pruning the network. To achieve this, we propose a clustering algorithm that groups similar classes in the network based on the firing rates of neurons for each class and then implement a lightweight cache architecture to store and reuse information from previously pruned networks. In our experiments with VGG-16, AlexNet, and ResNet-152 networks, CAP’NN achieves, on average, up to 47% model size reduction while actually improving the top-1(5) classification accuracy by up to 3.9%(3.4%) when the user only encounters a subset of the trained classes in these networks. Maede Hemmat, Joshua San Miguel, Azadeh Davoodi |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2021 | AirNN: A Featherweight Framework for Dynamic Input-Dependent Approximation of CNNsabstractIn this work, we propose AirNN, a novel framework which enables dynamic approximation of an already-trained convolutional neural network (CNN) in hardware during inference. AirNN enables input-dependent approximation of the CNN to achieve energy saving without much degradation in its classification accuracy at runtime. For each input, AirNN uses only a fraction of the CNN’s weights based on that input (with the rest remaining 0) to conduct the inference. Consequently, energy saving is possible due to fewer number of fetches from off-chip memory as well as fewer multiplications for majority of the inputs. To achieve per-input approximation, we propose a clustering algorithm that groups similar weights in the CNN based on their importance, and design an iterative framework that decides dynamically how many clusters of weights should be fetched from off-chip memory for each individual input. We also propose new hardware structures to implement our framework on top of a recently proposed FPGA-based CNN accelerator. In our experiments with popular CNNs, we, on average, show 49% energy saving with less than 3% degradation in classification accuracy due to doing inference with only a fraction of the weights for the majority of the inputs. We also propose a greedy interleaving scheme, implemented in hardware, in order to improve the performance of the iterative procedure and compensate for its latency overhead. Maede Hemmat, Joshua San Miguel, Azadeh Davoodi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2020 | CRANIA: Unlocking Data and Value Reuse in Iterative Neural Network ArchitecturesabstractA common inefficiency in traditional Convolutional Neural Network (CNN) architectures is that they do not adapt to variations in inputs. Not all inputs require the same amount of computation to be correctly classified, and not all of the weights in the network contribute equally to generate the output. Recent work introduces the concept of iterative inference, enabling per-input approximation. Such an iterative CNN architecture clusters weights based on their importance and saves significant power by incrementally fetching weights from off-chip memory until the classification result is accurate enough. Unfortunately, this comes at a cost of increased execution time since some inputs need to go through multiple rounds of inference, negating the savings in energy. We propose Cache Reuse Approximation for Neural Iterative Architectures (CRANIA) to overcome this inefficiency. We recognize that the re-execution and clustering built into these iterative CNN architectures unlock significant temporal data reuse and spatial value reuse, respectively. CRANIA introduces a lightweight cache+compression architecture customized to the iterative clustering algorithm, enabling up to 9 × energy savings and speeding up inference by 5.8 × with only 0.3% area overhead. Maede Hemmat, Tejas Shah, Joshua San Miguel |
ASP-DAC | 1 |
| 2020 | CAP'NN: Class-Aware Personalized Neural Network InferenceabstractWe propose CAP’NN, a framework for Class-Aware Personalized Neural Network Inference. CAP’NN prunes an already-trained neural network model based on the preferences of individual users. Specifically, by adapting to the subset of output classes that each user is expected to encounter, CAP’NN is able to prune not only ineffectual neurons but also miseffectual neurons that confuse classification, without the need to retrain the network. CAP’NN achieves up to 50% model size reduction while actually improving the top-l(5) classification accuracy by up to 2.3%(3.2%) when the user only encounters a subset of VGG-16 classes. Maede Hemmat, Joshua San Miguel, Azadeh Davoodi |
DAC | 1 |
| 2019 | Power-efficient ReRAM-aware CNN model generation
Maede Hemmat, Azadeh Davoodi |
Integr. | 1 |
| 2018 | Power-Efficient ReRAM-Aware CNN Model GenerationabstractThis is the first work to propose generation of the network model of a convolutional neural network (CNN) specifically for power-efficient implementation using the ReRAM technology. State-of-the-art in this area is based on implementation of an already fixed CNN model. It uses parallel crossbar structures to achieve a desired precision quantization of the edge weights using the limited precision provided by individual ReRAM devices. In contrast, in this work we keep the ReRAM crossbar structure in mind during model generation, and target altering a base CNN model such that the resulting implementation will be more power efficient. This is by means of eliminating the parallel crossbars as much as possible, thus reducing the number of required analog-to-digital and digital-to-analog converters which are dominant sources of power consumption. We propose four architectural techniques which are applicable to convolutional and fully-connected layers in a CNN and alter the network model to work with lower precision of the weights with negligible loss in accuracy. In our experiments, we achieve at least 50% power savings with almost no loss in accuracy for popular CNNs compared to ReRAM implementation of the base model. Maede Hemmat, Azadeh Davoodi |
ICCD | 1 |
| 2017 | Hybrid TFET-MOSFET circuit: A solution to design soft-error resilient ultra-low power digital circuit
Maede Hemmat, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
Integr. | 1 |
| 2016 | Hybrid TFET-MOSFET circuits: An approach to design reliable ultra-low power circuits in the presence of process variationabstractIn this work, to increase the timing yield of Tunnel Field Effect Transistor (TFET) circuits in the presence of the process variation, we propose to use MOSFET-based gates instead of some TFET-based gates in the TFET circuits. This hybridization approach originates from the fact that TFETs are more sensitive to process variation, when compared to conventional MOSFETs. First, we investigate the impact of process variations on Homojunction InAs TFETs by extracting the distributions of electrical parameters such as threshold voltage. Then, a hybrid TFET-MOSFET circuit design approach for increasing the reliability of the TFET circuits is introduced. The power consumptions of hybrid circuits are considerably smaller than the corresponding ones realized using CMOS circuits. In the proposed hybrid approach, the circuit is basically implemented in TFET to reduce the power and energy consumption while the gates whose their variations may lead to the timing violation, are implemented using MOSFET-based gates. The decision on replacing the TFET-based gates by their corresponding MOSFET-based gates during the hybrid design is made through a heuristic algorithm. The proposed algorithm considers the sensitivity of each TFET-based gate to the process variation. To assess the efficacy of the proposed approach, the proposed algorithm is applied to some circuits of the ISCAS'85 and ISCAS'89 benchmark packages. The results show that the reliability of the TFET-MOSFET-based circuits are up to 74% larger than that of the pure TFET-based circuits. Furthermore, the energy and leakage power consumptions of the proposed hybrid circuits are up to 56% and 80%, respectively, smaller than those of the pure MOSFET-based design. Maede Hemmat, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
VLSI-SoC | 1 |