VLDB 2026 Research / reviewers in the wild / expert
Rickard Ewetz
dblp:127/9041
· DBLP profile ↗
82ranked-venue papers
14as first author
54since 2021 · last 2026
0000-0002-4183-6926ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 62 · 14 first-author · 34 since 2021Artificial intelligence and machine learning · 19 · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TRACK: Robust Path-based In-Memory Computing for Efficient Execution of Boolean LogicabstractIn-memory computing (IMC) with non-volatile memory offers a promising pathway to overcome the von Neumann bottleneck [14]. Among digital IMC paradigms, path-based computing has emerged as an efficient approach for executing Boolean logic. However, existing synthesis methodologies rely on digital abstractions that ignore non-ideal analog effects, which can compromise functional correctness. This paper introduces TRACK, a robust framework for path-based IMC that preserves functional correctness under non-ideal analog effects. First, TRACK provides a high-fidelity simulator for analog verification, which leverages a worst-case approximation technique to enable scalable analysis without requiring exhaustive simulations. Second, TRACK introduces new hardware architectures and mapping directives, which ensures that the digital abstraction holds for large-scale applications. We evaluate TRACK on nine Revlib circuits, eight EPFL control benchmarks, and eight ISCAS85 designs. Our fast simulation technique accelerates analog verification by over 10,000 × for large Boolean functions, making the analysis of complex applications practical. Compared to existing IMC baselines, TRACK maintains functional correctness while delivering up to 78% energy savings and 28% lower latency. Venkata Nithin Kamineni, Jinam Modasiya, Nathaniel Cady, Rickard Ewetz |
ACM Great Lakes Symposium on VLSI | 4 |
| 2025 | Explaining ViTs Using Information FlowabstractComputer vision models can be explained by attributing the output decision to the input pixels. While effective methods for explaining convolutional neural networks have been proposed, these methods often produce low-quality attributions when applied to vision transformers (ViTs). State-of-the-art methods for explaining ViTs capture the flow of patch information using transition matrices. However, we observe that transition matrices alone are not sufficiently expressive to accurately explain ViT models. In this paper, we define a theoretical approach to creating explanations for ViTs called InFlow. The framework models the patch-to-patch information flow using a combination of transition matrices and patch embeddings. Moreover, we define an algebra for updating the transition matrices of series connected components, diverging paths, and converging paths in the ViT model. This algebra allows the InFlow framework to produce high quality attributions which explain ViT decision making. In experimental evaluation on ImageNet, with three models, InFlow outperforms six ViT attribution methods in the standard insertion, deletion, SIC and AIC metrics by up to 18%. Qualitative results demonstrate InFlow produces more relevant and sharper explanations. Code is publicly available at \url{https://github.com/chasewalker26/InFlow-ViT-Explanation.} Chase Walker, Md Rubel Ahmed, Sumit Kumar Jha 0001, Rickard Ewetz |
AISTATS | 4 |
| 2025 | Metric-Driven Attributions for Vision TransformersabstractAttribution algorithms explain computer vision models by attributing the model response to pixels within the input. Existing attribution methods generate explanations by combining transformations of internal model representations such as class activation maps, gradients, attention, or relevance scores. The effectiveness of an attribution map is measured using attribution quality metrics. This leads us to pose the following question: if attribution methods are assessed using attribution quality metrics, why are the metrics not used to generate the attributions? In response to this question, we propose a Metric-Driven Attribution for explaining Vision Transformers (ViT) called MDA. Guided by attribution quality metrics, the method creates attribution maps by performing patch order and patch magnitude optimization across all patch tokens. The first step orders the patches in terms of importance and the second step assigns the magnitude to each patch while preserving the patch order. Moreover, MDA can provide a smooth trade-off between sparse and dense attributions by modifying the optimization objective. Experimental evaluation demonstrates the proposed MDA method outperforms $7$ existing ViT attribution methods by an average of $12\%$ across $12$ attribution metrics on the ImageNet dataset for the ViT-base $16 \times 16$, ViT-tiny $16 \times 16$, and ViT-base $32 \times 32$ models. Code is publicly available at https://github.com/chasewalker26/MDA-Metric-Driven-Attributions-for-ViT. Chase Walker, Sumit Kumar Jha 0001, Rickard Ewetz |
ICLR | 3 |
| 2025 | Grammar-Forced Translation of Natural Language to Temporal Logic using LLMsabstractTranslating natural language (NL) into a formal language such as temporal logic (TL) is integral for human communication with robots and autonomous systems. State-of-the-art approaches decompose the task into a grounding of atomic propositions (APs) phase and a translation phase. However, existing methods struggle with accurate grounding, the existence of co-references, and learning from limited data. In this paper, we propose a framework for NL to TL translation called Grammar Forced Translation (GraFT). The framework is based on the observation that previous work solves both the grounding and translation steps by letting a language model iteratively predict tokens from its full vocabulary. In contrast, GraFT reduces the complexity of both tasks by restricting the set of valid output tokens from the full vocabulary to only a handful in each step. The solution space reduction is obtained by exploiting the unique properties of each problem. We also provide a theoretical justification for why the solution space reduction leads to more efficient learning. We evaluate the effectiveness of GraFT using the CW, GLTL, and Navi benchmarks. Compared with state-of-the-art translation approaches, it can be observed that GraFT improves the end-to-end translation accuracy by 5.49% and out-of-domain translation accuracy by 14.06% on average. William English 0001, Dominic Simon, Sumit Kumar Jha 0001, Rickard Ewetz |
ICML | 4 |
| 2025 | Street2Air: A Framework for Synthesizing Aerial Vehicle Views from Ground ImagesabstractAnnotated aerial view images are often missing from fine-grained vehicle type classification datasets. This lack of data limits both the accuracy and robustness of models when applied to top-down views, which are essential for applications such as autonomous drones and aerial surveillance. Models trained only on street-level images often fail to generalize to aerial perspectives, requiring more time and multiple observations to recognize vehicles accurately. In contrast, models trained with both street-level and aerial views can perform more reliably and with faster inference in drone-based systems. However, collecting real aerial data at scale can be costly and logistically challenging. In this paper, we propose AVA (Automated Aerial View Augmentation), a framework for aerial data augmentation via 3D asset generation and contextual scene synthesis. Since standalone 3D vehicle models from 2D images are not directly usable for detection, we embed them in realistic backgrounds to enable learning of both object features and scene context. AVA first constructs 3D vehicle models from street-view images. To ensure data quality, we introduce a realism checker that discards incomplete or distorted assets. We then apply geometric transformations to generate aerial 2D views. The 2D views pass through a text-to-video generator that adds background context, mimicking typical drone imagery. We evaluate our data augmentation approach by fine-tuning several object detection backbones. Notably, the pretrained YOLOv11 model, when fine-tuned with AVA augmented data, achieves a significant [email protected] improvement from 0.06 to 0.51 in classifying previously unseen vehicles from aerial perspectives. Md Rubel Ahmed, Fazle Rahat, M. Shifat Hossain, Sumit Kumar Jha 0001, Rickard Ewetz |
ICMLA | 5 |
| 2025 | Multitask Contrastive Learning using Task-Wise Training and Partitioned Embedding SpaceabstractMany real-world computer vision tasks require learning to associate multiple properties of different modalities with the same image. Multi-task learning enables a single model to learn these properties simultaneously by leveraging shared knowledge across related tasks to enhance generalization with single-modal data. On the other hand, contrastive learning effectively captures robust multi-modal features by aligning similar representations and distinguishing dissimilar ones. However, state-of-the-art methods struggle with combining these two learning approaches due to the difficulty in optimizing both shared and task-specific objectives. In this paper, we introduce a Multi-Task Contrastive Learning (MTCL) framework that partitions the embedding space to support both classification and regression tasks within a multi-task paradigm. By batching samples with tasks and structuring the embedding space to accommodate diverse task-specific requirements, our method retains the advantages of contrastive learning while addressing the unique challenges of multi-task learning. We evaluate our approach on three benchmark multi-task datasets—Zappos50K, CUB200, and MEDIC. We also introduce a multi-task Vehicles dataset that includes orientation. On the benchmark datasets, our model shows 24.5%, 17.2%, and 30.0% increase in overall classification accuracy compared to the SOTA methods. M. Shifat Hossain, Sumit Kumar Jha 0001, Hao Zheng 0005, Rickard Ewetz |
ICMLA | 4 |
| 2025 | Attr-RAG: Attribution-Guided Retrieval-Augmented Generation for Scientific Experiment DesignabstractEvidence-based science depends on the iterative integration of experimentation, a process traditionally driven by slow and error-prone human effort. This has inspired the vision of an automated "robot scientist" capable of conducting end-to-end experimentation. While Large Language Models (LLMs) can generate procedural instructions, they often struggle to accurately describe scientific experiments due to the limited availability of high-quality, domain-specific examples in their training data. Retrieval-Augmented Generation (RAG) helps bridge this gap by allowing LLMs to access up-to-date external information. However, despite being effective for short questions, RAG struggles with long-form scientific experimental queries due to information loss from chunk fragmentation and retrieval of irrelevant information. In this paper, we propose Attr-RAG, an attribution-guided RAG framework to remove irrelevant or misleading context and retaining only complete, relevant information. Unlike traditional RAG methods that rely solely on vector similarity, Attr-RAG introduces a refinement stage using occlusion-based attribution to identify which retrieved chunks truly influence the LLM’s response. This attribution-guided filtering ensures that only contextually coherent chunks are used for accurate and grounded final answer generation. Attr-RAG demonstrated superior performance in 9 out of 10 chemistry lab experiment tasks of the ChemEx dataset and outperformed baselines across most quantitative evaluation metrics. In qualitative evaluations conducted by state-of-the-art LLM judges (GPT-4o, Gemini 2.5, and Grok 3), the top mean scores of 27.8, 27.1, and 22.9, respectively, were achieved across six key evaluation criteria. Fazle Rahat, M. Shifat Hossain, Arvind Ramanathan, Sumit Kumar Jha 0001, Hao Zheng 0005, Rickard Ewetz |
ICMLA | 6 |
| 2025 | Knowledge Editing for Multi-Hop Question Answering Using Semantic AnalysisabstractLarge Language Models (LLMs) require lightweight avenues of updating stored information that has fallen out of date. Knowledge Editing (KE) approaches have been successful in updating model knowledge for simple factual queries but struggle with handling tasks that require compositional reasoning such as multi-hop question answering (MQA). We observe that existing knowledge editors leverage decompositional techniques that result in illogical reasoning processes. In this paper, we propose a knowledge editor for MQA based on semantic analysis called CHECK. Our framework is based on insights from an analogy between compilers and reasoning using LLMs. Similar to how source code is first compiled before being executed, we propose to semantically analyze reasoning chains before executing the chains to answer questions. Reasoning chains with semantic errors are revised to ensure consistency through logic optimization and re-prompting the LLM model at a higher temperature. We evaluate the effectiveness of CHECK against five state-of-the-art frameworks on four datasets and achieve an average 22.8% improved MQA accuracy. Dominic Simon, Rickard Ewetz |
IJCAI | 2 |
| 2025 | Detecting and Removing Adversarial Patches using Frequency SignaturesabstractComputer vision systems deployed in safety-critical applications have proven to be susceptible to adversarial patches. The patches can cause catastrophic outcomes within autonomous driving scenarios. Existing defense techniques learn discriminative patch features or trigger patterns, which leave the defenses vulnerable to unseen patch attacks. In this paper, we propose Corner Cutter, a defense against adversarial patches that is robust to unseen patches and adaptive attacks. The framework is based on the insight that the construction process of adversarial patches leaves an attack signature in the frequency domain. The signature can be detected in different adversarial patches, including the LaVAN patch, the adversarial patch, the naturalistic patch, and a projected gradient descent-based patch. The framework neutralizes identified patches by isolating the high frequency signals and removing the corresponding pixels in the image domain. Corner Cutter is able to achieve an 11% increase in adversarial accuracy for the image classification task and an 8% increase in mean average precision on the Naturalistic patch over other defenses. The evaluations also demonstrate that the framework is robust to unseen patches and adaptive attacks. Dominic Simon, Chase Walker, Sumit Kumar Jha 0001, Rickard Ewetz |
IJCNN | 4 |
| 2025 | GAMMA: Gated Multi-hop Message Passing for Homophily-Agnostic Node Representation in GNNsabstractThe success of Graph Neural Networks (GNNs) leverages the homophily principle, where connected nodes share similar features and labels. However, this assumption breaks down in heterophilic graphs, where same-class nodes are often distributed across distant neighborhoods rather than immediate connections. Recent attempts expand the receptive field through multi-hop aggregation schemes that explicitly preserve intermediate representations from each hop distance. While effective at capturing heterophilic patterns, these methods require separate weight matrices per hop and feature concatenation, causing parameters to scale linearly with hop count. This leads to high computational complexity and GPU memory consumption. We propose Gated Multi-hop Message Passing (GAMMA), where nodes assess how relevant the aggregated information is from their k-hop neighbors. This assessment occurs through multiple refinement steps where the node compares each hop's embedding with its current representation, allowing it to focus on the most informative hops. During the forward pass, GAMMA finds the optimal mix of multi-hop information local to each node using a single feature vector without needing separate representations for each hop, thereby maintaining dimensionality comparable to single hop GNNs. In addition, we propose a weight sharing scheme that leverages a unified transformation for aggregated features from multiple hops so the global heterophilic patterns specific to each hop are learned during training. As such, GAMMA captures both global (per-hop) and local (per-node) heterophily patterns without high computation and memory overhead. Experiments show GAMMA matches or exceeds state-of-the-art heterophilic GNN accuracy, achieving up to $\approx20\times$ faster inference. Our code is publicly available at \url{https://github.com/amir-ghz/GAMMA}. Amir Ghazizadeh Ahsaei, Rickard Ewetz, Hao Zheng 0005 |
NeurIPS | 2 |
| 2025 | Data Augmentation for Image Classification Using Generative AIabstractScaling laws dictate that the performance of AI models is proportional to the amount of available data. Data augmentation is a promising solution to expanding the dataset size. Traditional approaches focused on augmentation using rotation, translation, and resizing. Recent approaches use generative AI models to improve dataset diversity. However, the generative methods struggle with issues such as subject corruption and the introduction of irrelevant artifacts. In this paper, we propose the Automated Generative Data Augmentation (AGA). The framework combines the utility of large language models (LLMs), diffusion models, and segmentation models to augment data. AGA preserves foreground authenticity while ensuring background diversity. Specific contributions include: i) segment and superclass based object extraction, ii) prompt diversity with combinatorial complexity using prompt decomposition, and iii) affine subject manipulation. We evaluate AGA against state-of-the-art (SOTA) techniques on three representative datasets, ImageNet, CUB and iWildCam. The experimental evaluation demonstrates an accuracy improvement of 15.6% and 23.5% for in and out-of-distribution data compared to baseline models respectively. There is also 64.3% improvement in SIC score compared to the baselines. Fazle Rahat, M. Shifat Hossain, Md Rubel Ahmed, Sumit Kumar Jha 0001, Rickard Ewetz |
WACV | 5 |
| 2025 | PATCHOUT: Adversarial Patch Detection and Localization using Semantic ConsistencyabstractAbstract Computer vision systems are actively deployed in safety-critical applications such as autonomous vehicles. Real-world adversarial patches are capable of compromising the artificial intelligence (AI) systems with catastrophic outcomes. Existing defenses against patch attacks are based on identifying neurons, features, or gradients of high intensity. However, these defenses are vulnerable to weaker attacks that have less obvious attack signatures. In this paper, we propose the PATCHOUT framework that detects and locates adversarial patches using semantic consistency. Within patch detection, the key insight is that the top class predictions for an entity are semantically consistent for benign images, whereas they are inconsistent for attacked images. Within patch localization, it is observed that patches are semantically consistent with a coarse grained segmentation of the image. This allows the PATCHOUT framework to detect and remove adversarial patches using a class consistency checker as well as image segmentation, attribution analysis, and image restoration techniques. The experimental evaluation demonstrates that PATCHOUT can detect a broad range of adversarial patches with over 90% accuracy. The framework achieves 20% higher accuracy than other defenses. The framework is also evaluated against unseen attacks and adaptive attacks, reducing the success rate of adaptive attacks from 56% to 24%. Dominic Simon, Sumit Kumar Jha 0001, Rickard Ewetz |
Neural Process. Lett. | 3 |
| 2025 | LOGIC: Logic Synthesis for Digital In-Memory ComputingabstractIn-memory processing offers a promising solution for enhancing the performance of data-intensive applications. While analog in-memory computing demonstrates remarkable efficiency, its limited precision is suitable only for approximate computing tasks. In contrast, digital in-memory computing delivers the deterministic precision necessary to accelerate high-assurance applications. Current digital in-memory computing methods typically involve manually breaking down arithmetic operations into in-memory compute kernels. In contrast, traditional digital circuits are synthesized through intricate and automated design workflows. In this article, we introduce a logic synthesis framework called LOGIC, which facilitates the translation of high-level applications into digital in-memory compute kernels that can be executed using non-volatile memory. We propose techniques for decomposing element-wise arithmetic operations into in-memory kernels while minimizing the number of in-memory operations. Additionally, we optimize the sequence of in-memory operations to reduce non-volatile memory utilization. To address the NP-hard execution sequencing optimization problem, we have developed two look-ahead algorithms that offer practical solutions. Additionally, we leverage data layout reorganization to efficiently accelerate applications that heavily rely on sparse matrix-vector multiplication operations. Our experimental evaluations demonstrate that our proposed synthesis approach improves the area and latency of fixed-point multiplication by 84% and 20% compared to the state-of-the-art, respectively. Moreover, when applied to scientific computing applications sourced from the SuiteSparse Matrix Collection, our design achieves remarkable improvements in area, latency, and energy efficiency by factors of 4.8×, 2.6×, and 11×, respectively. Muhammad Rashedul Haq Rashed, Sven Thijssen, Sumit Kumar Jha 0001, Rickard Ewetz |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2024 | Integrated Decision Gradients: Compute Your Attributions Where the Model Makes Its DecisionabstractAttribution algorithms are frequently employed to explain the decisions of neural network models. Integrated Gradients (IG) is an influential attribution method due to its strong axiomatic foundation. The algorithm is based on integrating the gradients along a path from a reference image to the input image. Unfortunately, it can be observed that gradients computed from regions where the output logit changes minimally along the path provide poor explanations for the model decision, which is called the saturation effect problem. In this paper, we propose an attribution algorithm called integrated decision gradients (IDG). The algorithm focuses on integrating gradients from the region of the path where the model makes its decision, i.e., the portion of the path where the output logit rapidly transitions from zero to its final value. This is practically realized by scaling each gradient by the derivative of the output logit with respect to the path. The algorithm thereby provides a principled solution to the saturation problem. Additionally, we minimize the errors within the Riemann sum approximation of the path integral by utilizing non-uniform subdivisions determined by adaptive sampling. In the evaluation on ImageNet, it is demonstrated that IDG outperforms IG, Left-IG, Guided IG, and adversarial gradient integration both qualitatively and quantitatively using standard insertion and deletion metrics across three common models. Chase Walker, Sumit Kumar Jha 0001, Kenny Chen, Rickard Ewetz |
AAAI | 4 |
| 2024 | Towards Area-Efficient Path-Based In-Memory Computing using Graph IsomorphismsabstractIn-memory computing has attracted significant attention due to its potential to alleviate the issues caused by the von Neumann bottleneck. Path-based computing is a recently proposed in-memory computing paradigm for evaluating Boolean functions using nanoscale crossbars. Unlike state-of-the-art paradigms that use expensive WRITE operations to execute functions, path-based computing only relies on READ operations, which translates into benefits of low power consumption and low computational delay. Unfortunately, path-based computing comes with the penalty of substantial area overhead. In this paper, we introduce the ISO framework, a hardware-software solution for minimizing the area overhead of path-based computing systems. The framework is based on mapping computation to in-memory kernels using an intermediate k-LUT representation. The k-LUTs facilitate reusing hardware resources that realize the same computational structures. The reuse is performed by detecting identical subfunctions using isomorphic graphs. We also present a set of program instruction and scheduling algorithms to facilitate the hardware reuse. We have evaluated our proposed ISO framework on the 10 ISCAS85 benchmarks. Our experimental evaluation indicates that our proposed architecture improves energy consumption, latency, and area by $1.30\times, 76.59\times$, and $2.79\times$ on the average compared with previous state-of-the-art methods for path-based computing. Sven Thijssen, Muhammad Rashedul Haq Rashed, Hao Zheng 0005, Sumit Kumar Jha 0001, Rickard Ewetz |
ASPDAC | 5 |
| 2024 | READ-based In-Memory Computing using Sentential Decision DiagramsabstractProcessing-in-memory (PIM) has the potential to unleash unprecedented computing capabilities. While most in-memory computing paradigms rely on repeatedly programming the non-volatile memory devices, recent computing paradigms are capable of evaluating Boolean functions by simply observing the flow of electrical currents within a crossbar of non-volatile memory. Synthesizing Boolean functions into such crossbar designs is a fundamental problem for next-generation in-memory computing systems. The selection of the data structure used to guide the synthesis process has a first-order impact on the overall system performance. State-of-the-art in-memory computing paradigms leverage representations such as majority inverter graphs (MIGs), and binary decision diagrams (BDDs). In this paper, we propose the Cascading Crossbar Synthesis using SDDs (C2S2) framework for automatically synthesizing Boolean logic into crossbar designs. The cornerstone of the C2S2framework is a newly invented data structure called sentential decision diagrams (SDDs). It has been proved that SDDs are more succinct than binary decision diagrams (BDDs). To minimize expensive data transfer on the system bus, C2S2maps computation to multiple crossbars that are connected together in series. The C2S2framework is evaluated using 13 benchmark circuits. Compared with state-of-the-art paradigms such as CONTRA, FLOW, and PATH, C2S2improves energy-efficiency by $6.8 \times$ while maintaining similar latency. Sven Thijssen, Muhammad Rashedul Haq Rashed, Sumit Kumar Jha 0001, Rickard Ewetz |
ASPDAC | 4 |
| 2024 | On the Design of Novel Attention Mechanism for Enhanced Efficiency of TransformersabstractWe present a new xor-based attention function for efficient hardware implementation of transformers. While the standard attention mechanism relies on matrix multiplication between the key and the transpose of the query, we propose replacing the computation of this attention function with bitwise xor operations. We mathematically analyze the information-theoretic properties of the standard multiplication-based attention, demonstrating that it preserves input entropy, and then computationally show that the xor-based attention approximately preserves the entropy of its input despite small variations in correlations between the inputs. Across various admittedly simple tasks, including arithmetic, sorting, and text generation, we show comparable performance to baseline methods using scaled GPT models. The xor-based computation of the attention function shows substantial improvement in power consumption, latency, and circuit area compared to the corresponding multiplication-based attention function. This hardware efficiency makes xor-based attention more compelling for the deployment of transformers under tight resource constraints, opening new application domains in sustainable energy-efficient computing. Additional optimizations to the xor-based attention function can further improve efficiency of transformers. Sumit Kumar Jha 0001, Susmit Jha, Rickard Ewetz, Alvaro Velasquez |
DAC | 3 |
| 2024 | Execution Sequence Optimization for Processing In-Memory using Parallel Data PreparationabstractProcessing in-memory (PIM) promises to unleash unprecedented computing capabilities for high-data-rate applications. Computation using PIM is performed by breaking down computationally expensive operations into in-memory kernels that can be efficiently executed using non-volatile memory. Logic styles such as MAGIC require that each output memory cell is prepared for evaluation before executing the functional logic operation. State-of-the-art synthesis algorithms perform the preparation immediately after memory cells have expired. Unfortunately, this results in that columns of cells are prepared greedily, instead of leveraging efficient parallel data preparation instructions. In this paper, we propose the PREP framework that maximizes the opportunities for parallel column preparation using execution sequence optimization. The key idea of the framework is to postpone data preparation instructions until there are no available prepared cells. Next, the accumulated memory cells are prepared in parallel to release the memory for functional evaluations. The framework is capable of exploring a frontier of area-performance solutions. The PREP framework is evaluated using 15 benchmarks from the SuiteSparse library. Compared with state-of-the-art synthesis tools, energy consumption and latency are respectively reduced by 27% and 25% with no additional cost in crossbar memory. Muhammad Rashedul Haq Rashed, Sven Thijssen, Dominic Simon, Sumit Kumar Jha 0001, Rickard Ewetz |
DAC | 5 |
| 2024 | Synthesis of Compact Flow-based Computing Circuits from Boolean ExpressionsabstractProcessing in-memory has the potential to accelerate high-data-rate applications beyond the limits of modern hardware. Flow-based computing is a computing paradigm for executing Boolean logic within nanoscale memory arrays by leveraging the natural flow of electric current. Previous approaches of mapping Boolean logic onto flow-based computing circuits have been constrained by their reliance on binary decision diagrams (BDDs), which translates into high area overhead. In this paper, we introduce a novel framework called FACTOR for mapping logic functions into dense flow-based computing circuits. The proposed methodology introduces Boolean connectivity graphs (BCGs) as a more versatile representation, capable of producing smaller crossbar circuits. The framework constructs concise BCGs using factorization and expression trees. Next, the BCGs are modified to be amenable for mapping to crossbar hardware. We also propose a time multiplexing strategy for sharing hardware between different Boolean functions. Compared with the state-of-the-art approach, the experimental evaluation using 14 circuits demonstrates that FACTOR reduces area, speed, and energy with 80%, 2%, and 12%, respectively, compared with the state-of-the-art synthesis method for flow-based computing. Sven Thijssen, Muhammad Rashedul Haq Rashed, Sumit Kumar Jha 0001, Rickard Ewetz |
DAC | 4 |
| 2024 | Equivalence Checking for Flow-Based Computing using Iterative SAT SolvingabstractProcessing in-memory is projected to shatter the von Neumann bottleneck and enable acceleration of data-intensive applications. Flow-based computing is an efficient in-memory computing paradigm for accelerating the execution of Boolean logic. While recent synthesis algorithms can map complex functions into flow-based computing circuits, the functional correctness cannot be verified using state-of-the-art equivalence checking techniques. The challenge is that non-volatile memory devices are intrinsically bi-directional, which introduces cycles in the computational graph. These cycles break traditional equivalence checking methods that are based on SAT formulations. In this paper, we propose a framework for equivalence checking of flow-based computing circuits that is called FlowSAT. The framework captures each circuit using an undirected computational graph. The key idea of FlowSAT is to introduce helper variables, in the form of arrows, that dynamically convert the undirected graph into a directed graph. This facilitates equivalence checking to be performed using traditional SAT formulations. However, it is prohibitively expensive to ban all possible cycles using arrow variables. Therefore, we propose to eliminate cycles by iteratively adding constraints to the SAT formulation. Our experimental evaluation demonstrates that FlowSAT is up to an order of magnitude faster than state-of-the-art methods. The framework is capable of verifying all 20/20 benchmark circuits, while the previous state-of-the-art technique is only capable of verifying 12/20 circuits within a time limit of one hour. Sven Thijssen, Muhammad Rashedul Haq Rashed, Md Rubel Ahmed, Suraj Singireddy, Sumit Kumar Jha 0001, Rickard Ewetz |
ICCAD | 6 |
| 2024 | NSP: A Neuro-Symbolic Natural Language Navigational PlannerabstractPath planners that can interpret free-form natural language instructions hold promise to automate a wide range of robotics applications. These planners simplify user interactions and enable intuitive control over complex semi-autonomous systems. While existing symbolic approaches offer guarantees on the correctness and efficiency, they struggle to parse free-form natural language inputs. Conversely, neural approaches based on pre-trained Large Language Models (LLMs) can manage natural language inputs but lack performance guaran-tees. In this paper, we propose a neuro-symbolic framework for path planning from natural language inputs called NSP. The framework leverages the neural reasoning abilities of LLMs to i) craft symbolic representations of the environment and ii) a symbolic path planning algorithm. Next, a solution to the path planning problem is obtained by executing the algorithm on the environment representation. The framework uses a feedback loop from the symbolic execution environment to the neural generation process to self-correct syntax errors and satisfy execution time constraints. We evaluate our neuro-symbolic approach using a benchmark suite with 1500 path-planning problems. The experimental evaluation shows that our neuro-symbolic approach produces 90.1% valid paths that are on average 19-77% shorter than state-of-the-art neural approaches. William English 0001, Dominic Simon, Sumit Kumar Jha 0001, Rickard Ewetz |
ICMLA | 4 |
| 2024 | Out-of-Distribution Detection for Contrastive Models Using Angular Distance MeasuresabstractVision-language models have demonstrated extraordinary zero-shot image classification capabilities. Out-of-distribution (OOD) detection is the problem of determining if a model is operating within its knowledge limits. While distance-based detection algorithms have emerged as a promising approach to OOD detection, we observe that there is a disparity between the distance measures used for OOD detection and model training. Recent studies have attempted to mitigate this shortcoming by modifying the contrastive learning process, which is highly undesirable for foundation models. In this paper, we propose an Angular distance-based out-of-distribution detection method for Contrastive models (AEC), an OOD detection framework for foundational contrastive models based on an angular distance measure. The angular distance-based score is compliant with the standard training process and circumvents the need to modify the training process of the model. We also formulate a distance transformation and solve an optimization problem to determine an OOD score threshold value for in-distribution and out-of-distribution data. The experimental evaluation demonstrates that AEC outperforms state-of-the-art OOD detection models in terms of AUROC, FPR@95TPR, accuracy, and correct ID metrics. We obtained an overall AUROC and FPR@95TPR of 70.42 and 83.98 from the proposed algorithm, which is significantly better compared to the SOTA OOD detection algorithms. M. Shifat Hossain, Sumit Kumar Jha 0001, Chase Walker, Rickard Ewetz |
ICMLA | 4 |
| 2024 | CLE: Context-Aware Local Explanations for High Dimensional Tabular DataabstractExplainable artificial intelligence (XAI) seeks to enhance the transparency, interpretability, and trustworthiness of AI models. One solution strategy for explaining complex AI models for high-dimensional tabular data is to approximate them locally using surrogate models. Surrogate models such as linear regression and decision trees are inherently interpretable and can be used as an explanation. However, it is challenging for linear regression and decision trees to provide meaningful explanations for data points far from the decision boundary. In this paper, we propose a framework that provides Context-aware Local Explanations for high-dimensional tabular data called CLE. We observe that the quality of explanations from different local models varies depending on the data point. The CLE framework uses the context around a data point to select the type of symbolic explanation. Moreover, we propose to utilize feature attributions to explain data points that are far from the decision boundary. The proposed method is evaluated using high-dimensional tabular datasets from the domains of power systems, breast cancer detection, heart disease detection, and website phishing detection. The experimental results show that CLE can provide meaningful local explanations for data points far from the decision boundary. The framework explains data points using three different types of local models and demonstrates a smooth trade-off between explanation accuracy and interpretability. It can be observed that a relatively simple decision tree can explain a data point with 92.31 % accuracy. Fazle Rahat, M. Shifat Hossain, Md Rubel Ahmed, Rickard Ewetz |
ICMLA | 4 |
| 2024 | Attribution Quality Metrics with Magnitude Alignment
Chase Walker, Dominic Simon, Kenny Chen, Rickard Ewetz |
IJCAI | 4 |
| 2024 | PATH: Evaluation of Boolean Logic Using Path-Based In-Memory Computing SystemsabstractIn-memory computing using non-volatile memory is a promising pathway to accelerate data-intensive applications. While substantial research efforts have been dedicated to executing Boolean logic using digital in-memory computing, the limitation of state-of-the-art paradigms is that they heavily rely on repeatedly switching the state of the non-volatile resistive devices using expensive WRITE operations. In this paper, we propose a new in-memory computing paradigm called path-based computing for evaluating Boolean logic. Computation within the paradigm is performed using a one-time expensive compilation phase and a fast and efficient evaluation phase. The key property of the paradigm is that the execution phase only involves cheap READ operations. First, we define an analogy between binary decision diagrams (BDDs) and one-transistor one-memristor (1T1M) crossbars that allows Boolean functions to be mapped into crossbar designs. When such crossbar design becomes too large to be physically realizable, we propose to synthesize the Boolean function into a path-based computing system. A path-based computing system consists of a topology of staircase structures. A staircase structure is a cascade of hardwired crossbars, which minimizes inter-crossbar communication. We evaluate the proposed paradigm using ten circuits from the Revlib benchmark suite, eight control circuits of the EPFL benchmark suite, and eight ISCAS85 benchmarks. Compared with state-of-the-art digital in-memory computing paradigms, path-based computing improves energy and latency with 1006× and 10× on average, respectively. Sven Thijssen, Muhammad Rashedul Haq Rashed, Sumit Kumar Jha 0001, Rickard Ewetz |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2023 | Discovering the in-Memory Kernels of 3D Dot-Product EnginesabstractThe capability of resistive random access memory (ReRAM) to implement multiply-and-accumulate operations promises unprecedented efficiency in the design of scientific computing applications. While the use of two-dimensional (2D) ReRAM crossbar has been well investigated in the last few years, the design of in-memory dot-product engines using three-dimensional (3D) ReRAM crossbars remains a topic of active investigations. In this paper, we holistically explore how to leverage 3D ReRAM crossbars with several (2 to 7) stacked crossbar layers. In contrast, previous studies have focused on 3D ReRAM with at most 2 stacked crossbar layers. We first discover the in-memory compute kernels that can be realized using 3D ReRAM with multiple stacked crossbar layers. We discover that matrices with different sparsity patterns can be realized by appropriately assigning the inputs and outputs to the perpendicular metal wires within the 3D stack. We present a design automation tool to map sparse matrices within scientific computing applications to the discovered 3D kernels. The proposed framework is evaluated using 20 applications from the SuitSparse Matrix Collection. Compared with 2D crossbars, the proposed approach using 3D crossbars improves area, energy, and latency with 2.02X, 2.37X, 2.45X, respectively. Muhammad Rashedul Haq Rashed, Sumit Kumar Jha 0001, Rickard Ewetz |
ASP-DAC | 3 |
| 2023 | FLOW-3D: Flow-Based Computing on 3D Nanoscale Crossbars with Minimal SemiperimeterabstractThe emergence of data-intensive applications has spurred the interest for in-memory computing using nanoscale crossbars. Flow-based in-memory computing is a promising approach for evaluating Boolean logic using the natural flow of electrical currents. While automated synthesis approaches have been developed for 2D crossbars, 3D crossbars have advantageous properties in terms of density, area, and performance. In this paper, we propose the first framework for performing flow-based computing using 3D crossbars. The framework, FLOW-3D, automatically synthesizes a Boolean function into a crossbar design. FLOW-3D is based on an analogy between BDDs and crossbars, resulting in the synthesis of 3D crossbar designs with minimal semiperimeter. A BDD with n nodes is mapped to a 3D crossbar with (n + k) metal wires. The k extra metal wires are needed to handle hardware-imposed constraints. Compared with the state-of-the-art synthesis tool for 2D crossbars, FLOW-3D improves semiperimeter, area, energy consumption, and latency up to 61%, 84%, 37%, and 41% on 15 Revlib benchmarks. Sven Thijssen, Sumit Kumar Jha 0001, Rickard Ewetz |
ASP-DAC | 3 |
| 2023 | UpTime: Towards Flow-based In-Memory Computing with High Fault-ToleranceabstractProcessing in-memory promises to accelerate data-intensive applications by breaking von-Neumann based design principles. Flow-based computing is an in-memory computing paradigm that has shown immense potential for executing Boolean logic. Unfortunately, the immature fabrication processes for nanoscale memristor crossbars still struggle with yield challenges and run-time defects, which may render the computing system non-functional. Even worse, no previous studies have investigated the fault-tolerance of flow-based computing systems, which could potentially limit the capabilities of the entire paradigm. In this paper, we propose the UpTime framework to provide guaranties on the functional correctness and to maximize the lifetime of flow-based computing systems. The framework utilizes data layout organization to mitigate errors from faults with known type and location. To handle defects occurring at run-time, we propose the use of an error detection signal that can be evaluated with low overhead. The experimental evaluation demonstrates that the UpTime framework is capable of guaranteeing functional correctness for an average of 15.24 years. The up-time to down-time ratio is 99.9992%. Compared with utilizing the state-of-the-art write-verify scheme, the proposed error signal reduces power consumption by 25% and increases throughput by 6%, respectively. Sven Thijssen, Muhammad Rashedul Haq Rashed, Sumit Kumar Jha 0001, Rickard Ewetz |
DAC | 4 |
| 2023 | Automated Synthesis for In-Memory ComputingabstractProcessing in-memory has the potential to break von-Neumann based design principles and unleash exascale computing capabilities. A rudimentary problem for in-memory paradigms is to decompose mathematical operations into in-memory compute kernels. In this paper, we propose the AUTO framework that automatically maps arithmetic operations into in-memory compute kernels that can be executed using non-volatile memory. The AUTO framework is based on defining semantically complete custom adders optimized for in-memory computing. Using a library of such adders and a projection of the partial product space, we discover decomposition that enable fixed-point multiplication to be executed with fewer steps. The framework also directly applies the technique to dot-product operations to further improve performance. Compared with state-of-the-art, the experimental results demonstrate that AUTO can perform fixed-point multiplication and dot-product operations with 16% and 19% fewer steps, respectively. For a library of scientific computing applications, this translates into energy and latency improvements of 15 % and 17 %, respectively. Muhammad Rashedul Haq Rashed, Sven Thijssen, Sumit Kumar Jha 0001, Rickard Ewetz |
ICCAD | 4 |
| 2023 | Path-Based Processing using In-Memory Systolic Arrays for Accelerating Data-Intensive ApplicationsabstractThe next wave of scientific discovery is predicated on unleashing beyond-exascale simulation capabilities using in-memory computing. Path-based computing is a promising in-memory logic style for accelerating Boolean logic with deterministic precision. However, existing studies on path-based computing are limited to executing small combinational circuits. In this paper, we propose a framework called PSYS to accelerate data-intensive scientific computing applications using path-based in-memory systolic arrays. The approach leverages path-based computing for multiplying known constants with an unknown operand, which substantially reduces the computational complexity compared with general purpose multiplication of two unknown operands. The systolic arrays minimize data movement by storing the matrix elements using non-volatile memory and performing processing in-place. The framework decomposes unstructured computations to the systolic arrays while considering the non-regular computational patterns of the applications. Our experimental evaluations employ applications from the domains of engineering, physics, and mathematics. The experimental results demonstrate that compared with the state-of-the-art, the PSYS framework improves energy and latency by a factor of 101x and 23x, respectively. Muhammad Rashedul Haq Rashed, Sven Thijssen, Sumit Kumar Jha 0001, Hao Zheng 0005, Rickard Ewetz |
ICCAD | 5 |
| 2023 | Verification of Flow-Based Computing Systems Using Bounded Model CheckingabstractFlow-based computing is a digital in-memory computing paradigm with tremendous potential. Its favorable characteristics, such as high robustness, low energy consumption and small computational delay make it a strong contender for integration into future computing systems. While most studies on emerging computing paradigms are focused on synthesis, it is crucial to develop methods to verify the functional correctness of the resulting designs. Flow-based computing is based on an undirected computational graph, which prevents equivalence checking to be performed by solving SAT formulations. In this paper, we propose a framework called XSAT for equivalence checking of crossbar designs for flow-based computing. The XSAT framework draws on bounded model checking (BMC) to convert the undirected computational graph into a directed acyclic computational graph (DAG). The conversion allows traditional SAT-based equivalence checking techniques to be used at the expense of increasing the size of the problem. We further introduce a divide-and-conquer technique to accelerate the verification process. The technique divides the main problem into many subproblems of smaller size, which can be executed in parallel using multiple cores or nodes. From the experimental evaluation, it can be observed that the XSAT framework can solve all nineteen MCNC benchmarks whereas previous SOTA techniques can only solve eleven out of the nineteen benchmarks within one hour, i.e., with speed-ups of one to two orders of magnitude. Moreover, the divide-and-conquer technique results in speed-ups of up to 93× on large benchmark circuits. Sven Thijssen, Suraj Singireddy, Muhammad Rashedul Haq Rashed, Sumit Kumar Jha 0001, Rickard Ewetz |
ICCAD | 5 |
| 2023 | Input-Aware Flow-Based In-Memory ComputingabstractIn-memory computing using nanoscale crossbar arrays is a promising solution strategy to overcome the limitations of the von Neumann architecture. Flow-based computing is an emerging in-memory computing paradigm for evaluating Boolean logic using the natural flow of electrical currents. Previous studies on flow-based computing have focused on synthesizing crossbar designs with small dimensions to improve various performance metrics. In this paper, we observe that the latency and energy of evaluating a Boolean input vector is dependent on the state of the crossbar design (or the previous input vector). To take advantage of this observation, we propose the REORDER framework that reorders the sequence of input vectors to improve performance. The reordering reduces the overall number of WRITE operations to the non-volatile memory devices, which has a first-order impact on the overall performance of flow-based computing systems. The optimal input sequence can be obtained by formulating and solving a traveling salesman problem (TSP). The REORDER framework leverages a heuristic solution to balance pre-processing overhead with reduction in device switching. We evaluate the REORDER framework on image processing applications that allow input vector reordering. Compared with a naïve input sequence, the framework improves time and energy efficiency by 78% and 69% respectively for image filtering and by 94% and 72% respectively for feature extraction. Suraj Singireddy, Muhammad Rashedul Haq Rashed, Sven Thijssen, Rickard Ewetz, Sumit Kumar Jha 0001 |
ICCD | 4 |
| 2023 | STREAM: Toward READ-Based In-Memory Computing for Streaming-Based Processing for Data-Intensive ApplicationsabstractWith the rise of data-intensive applications, traditional computing paradigms have hit the memory-wall. In-memory computing using emerging nonvolatile memory (NVM) technology is a promising solution strategy to overcome the limitations of the von-Neumann architecture. In-memory computing using NVM devices has been explored in both analog and digital domains. Analog in-memory computing can perform matrix–vector multiplication (MVM) in an extremely energy-efficient manner. However, analog in-memory computing is prone to errors and resulting precision is therefore low. On the contrary, digital in-memory computing is a viable option for accelerating scientific computations that require deterministic precision. In recent years, several digital in-memory computing styles have been proposed. Unfortunately, state-of-the-art digital in-memory computing styles rely on repeated WRITE operations which involve switching of NVM devices. WRITE operations in NVM cells are expensive in terms of energy, latency, and device endurance. In this article, we propose a READ-based in-memory computing framework called STREAM. The framework performs streaming-based data processing for data-intensive applications. The STREAM framework consists of a synthesis tool that decomposes an arbitrary Boolean function into in-memory compute kernels. Two synthesis approaches are proposed to generate READ-based in-memory compute kernels using data structures from logic synthesis. A hardware/software co-design technique is developed to minimize the intercrossbar data communication. The STREAM framework is evaluated using circuits from the ISCAS85 benchmark suite, and Suite-Sparse applications to scientific computing. Compared with state-of-the-art in-memory computing framework, the proposed framework improves latency and energy performance with up to$200 \times $and$20\times $, respectively. Muhammad Rashedul Haq Rashed, Sven Thijssen, Sumit Kumar Jha 0001, Fan Yao 0001, Rickard Ewetz |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2022 | Shaping Noise for Robust Attributions in Neural Stochastic Differential EquationsabstractNeural SDEs with Brownian motion as noise lead to smoother attributions than traditional ResNets. Various attribution methods such as saliency maps, integrated gradients, DeepSHAP and DeepLIFT have been shown to be more robust for neural SDEs than for ResNets using the recently proposed sensitivity metric. In this paper, we show that neural SDEs with adaptive attribution-driven noise lead to even more robust attributions and smaller sensitivity metrics than traditional neural SDEs with Brownian motion as noise. In particular, attribution-driven shaping of noise leads to 6.7%, 6.9% and 19.4% smaller sensitivity metric for integrated gradients computed on three discrete approximations of neural SDEs with standard Brownian motion noise: stochastic ResNet-50, WideResNet-101 and ResNeXt-101 models respectively. The neural SDE model with adaptive attribution-driven noise leads to 25.7% and 4.8% improvement in the SIC metric over traditional ResNets and Neural SDEs with Brownian motion as noise. To the best of our knowledge, we are the first to propose the use of attributions for shaping the noise injected in neural SDEs, and demonstrate that this process leads to more robust attributions than traditional neural SDEs with standard Brownian motion as noise. Sumit Kumar Jha 0001, Rickard Ewetz, Alvaro Velasquez, Arvind Ramanathan, Susmit Jha |
AAAI | 2 |
| 2022 | STREAM: Towards READ-based In-Memory Computing for Streaming based Data ProcessingabstractProcessing in-memory breaks von-Neumann based design principles to accelerate data-intensive applications. While analog in-memory computing is extremely energy-efficient, the low precision narrows the spectrum of viable applications. In contrast, digital in-memory computing has deterministic precision and can therefore be used to accelerate a broad range of high assurance applications. Unfortunately, the state-of-the-art digital in-memory computing paradigms rely on repeatedly switching the non-volatile memory devices using expensive WRITE operations. In this paper, we propose a framework called STREAM that performs READ-based in-memory computing for streaming-based data processing. The framework consists of a synthesis tool that decomposes high-level programs into in-memory compute kernels that are executed using non-volatile memory. The paper presents hardware/software co-design techniques to minimize the data movement between different nanoscale crossbars within the platform. The framework is evaluated using circuits from ISCAS85 benchmark suite and Suite-Sparse applications to scientific computing. Compared with WRITE-based in-memory computing, the READ-based in-memory computing improves latency and power consumption up to 139X and 14X, respectively. Muhammad Rashedul Haq Rashed, Sven Thijssen, Sumit Kumar Jha 0001, Fan Yao 0001, Rickard Ewetz |
ASP-DAC | 5 |
| 2022 | Towards resilient analog in-memory deep learning via data layout re-organizationabstractProcessing in-memory paves the way for neural network inference engines. An arising challenge is to develop the software/hardware interface to automatically compile deep learning models onto in-memory computing platforms. In this paper, we observe that the data layout organization of a deep neural network (DNN) model directly impacts the model's classification accuracy. This stems from that the resistive parasitics within a crossbar introduces a dependency between the matrix data and the precision of the analog computation. To minimize the impact of the parasitics, we first perform a case study to understand the underlying matrix properties that result in computation with low and high precision, respectively. Next, we propose the XORG framework that performs data layout organization for DNNs deployed on in-memory computing platforms. The data layout organization improves precision by optimizing the weight matrix to crossbar assignments at compile time. The experimental results show that the XORG framework improves precision with up to 3.2X and 31% on the average. When accelerating DNNs using XORG, the write bit-accuracy requirements are relaxed with 1-bit and the robustness to random telegraph noise (RTN) is improved. Muhammad Rashedul Haq Rashed, Amro Awad, Sumit Kumar Jha 0001, Rickard Ewetz |
DAC | 4 |
| 2022 | PATH: evaluation of boolean logic using path-based in-memory computingabstractProcessing in-memory breaks von Neumann-based constructs to accelerate data-intensive applications. Noteworthy efforts have been devoted to executing Boolean logic using digital in-memory computing. The limitation of state-of-the-art paradigms is that they heavily rely on repeatedly switching the state of the non-volatile resistive devices using expensive WRITE operations. In this paper, we propose a new in-memory computing paradigm called path-based computing for evaluating Boolean logic. Computation within the paradigm is performed using a one-time expensive compile phase and a fast and efficient evaluation phase. The key property of the paradigm is that the execution phase only involves cheap READ operations. Moreover, a synthesis tool called PATH is proposed to automatically map computation to a single crossbar design. The PATH tool also supports the synthesis of path-based computing systems where the total number of crossbars and the number of inter-crossbar connections are minimized. We evaluate the proposed paradigm using 10 circuits from the RevLib benchmark suite. Compared with state-of-the-art digital in-memory computing paradigms, path-based computing improves energy and latency up to 4.7X and 8.5X, respectively. Sven Thijssen, Sumit Kumar Jha 0001, Rickard Ewetz |
DAC | 3 |
| 2022 | Hybrid Digital-Digital In-Memory ComputingabstractIn-memory computing (IMC) using emerging non-volatile memory promises exascale computing capabilities for a number of data-intensive workloads. The state-of-the-art solution to accelerating high assurance applications is based on digital in-memory computing. Digital in-memory computing can be WRITE-based or READ-based, i.e., logic is evaluated while switching or without switching the state of the non-volatile resistive devices. All prominent studies for accelerating matrix-vector multiplication (MVM) based applications utilize a single digital logic style. However, we observe that WRITE-based and READ-based digital in-memory computing are advantageous for dense and sparse matrices, respectively. In this paper, we propose a new computing paradigm called hybrid digital-digital in-memory computing paradigm. The paper also introduces automated synthesis tool for mapping computation to a hybrid architecture. The key idea is to first decompose the matrix into dense and sparse blocks. Next, bit-slicing is used to further decompose the dense blocks into sparse and dense parts. The dense (sparse) blocks are mapped to WRITE-based (READ-based) digital in-memory accelerators. The proposed paradigm is evaluated using 12 applications from various domains. Compared with WRITE-based IMC, the hybrid digital-digital paradigm improves energy and speed with 13X and 20X at the expense of increasing the area with 151X. Compared with READ-based IMC, the hybrid paradigms improves energy, speed, and area with 264X, 198X, and 2996X, respectively. Muhammad Rashedul Haq Rashed, Sumit Kumar Jha 0001, Fan Yao 0001, Rickard Ewetz |
DATE | 4 |
| 2022 | Logic Synthesis for Digital In-Memory ComputingabstractProcessing in-memory is a promising solution strategy for accelerating data-intensive applications. While analog in-memory computing is extremely efficient, the limited precision is only acceptable for approximate computing applications. Digital in-memory computing provides the deterministic precision required to accelerate high assurance applications. State-of-the-art digital in-memory computing schemes rely on manually decomposing arithmetic operations into in-memory compute kernels. In contrast, traditional digital circuits are synthesized using complex and automated design flows. In this paper, we propose a logic synthesis framework called LOGIC for mapping high-level applications into digital in-memory compute kernels that can be executed using non-volatile memory. We first propose techniques to decompose element-wise arithmetic operations into in-memory kernels while minimizing the number of in-memory operations. Next, the sequence of the in-memory operation is optimized to minimize non-volatile memory utilization. Lastly, data layout re-organization is used to efficiently accelerate applications dominated by sparse matrix-vector multiplication operations. The experimental evaluations show that the proposed synthesis approach improves the area and latency of fixed-point multiplication by 77% and 20% over the state-of-the-art, respectively. On scientific computing applications from Suite Sparse Matrix Collection, the proposed design improves the area, latency and, energy by 3.6X, 2.6X, and 8.3X, respectively. Muhammad Rashedul Haq Rashed, Sumit Kumar Jha 0001, Rickard Ewetz |
ICCAD | 3 |
| 2022 | Equivalence Checking for Flow-Based ComputingabstractThe rapid growth of data-intensive applications has spurred the interest for novel in-memory computing paradigms. With the recent innovations within flow-based computing, complex circuit specification can automatically be compiled into crossbar designs. This has raised the important question of verifying the functional correctness of the synthesized crossbars. Unfortunately, the traditional equivalence checking techniques based on SAT formulations cannot directly be applied to flow-based computing. This explains why the existing techniques are rather naive and have exponential runtime complexity. In this paper, we present a framework called CHECK that casts the equivalence checking problem into a problem of detecting simple paths in an undirected graph, which enables verification to be performed using efficient graph algorithms. Moreover, the scaleability of the equivalence checking is further improved by dynamically shrinking the size of the graph using logic rules. The experimental results demonstrate the proposed graph-based approach is one to two orders of magnitude faster than brute-force enumeration. This translates into that CHECK is capable of verifying 25 designs from the RevLib suite. In contrast, naive enumeration is only capable of verifying 18 out of the 25 designs. Sven Thijssen, Sumit Kumar Jha 0001, Rickard Ewetz |
ICCD | 3 |
| 2022 | Classifying the Ideological Orientation of User-Submitted Texts in Social MediaabstractWith the long-term goal of understanding how language is used and evolves within online communities, this work explores the application of natural language processing techniques to classify text articles according to their ideological orientation (i.e., conservative or liberal). We first collect a balanced corpus of text articles posted to the online communities r/Liberal and r/Conservative from the social media website Reddit. Using the corpus, we develop and apply three classifiers. The baseline classifier is a Bayes model that accounts for each text article’s web domain, as such, classification is independent of content. Next, we develop a support vector machine (SVM) model with term frequency-inverse document frequency (TF-IDF) features; this approach highlight differences in language using a count-based feature-space to differentiate text articles. Last, we evaluate the context-based transformer (RoBERTa) model and discuss its under-performance relative to the baseline and SVM models. Kamalakkannan Ravi, Adan Ernesto Vela, Rickard Ewetz |
ICMLA | 3 |
| 2022 | ExplainIt!: A Tool for Computing Robust Attributions of DNNsabstractResponsible integration of deep neural networks into the design of trustworthy systems requires the ability to explain decisions made by these models. Explainability and transparency are critical for system analysis, certification, and human-machine teaming. We have recently demonstrated that neural stochastic differential equations (SDEs) present an explanation-friendly DNN architecture. In this paper, we present ExplainIt, an online tool for explaining AI decisions that uses neural SDEs to create visually sharper and more robust attributions than traditional residual neural networks. Our tool shows that the injection of noise in every layer of a residual network often leads to less noisy and less fragile integrated gradient attributions. The discrete neural stochastic differential equation model is trained on the ImageNet data set with a million images, and the demonstration produces robust attributions on images in the ImageNet validation library and on a variety of images in the wild. Our online tool is hosted publicly for educational purposes. Sumit Kumar Jha 0001, Alvaro Velasquez, Rickard Ewetz, Laura L. Pullum, Susmit Jha |
IJCAI | 3 |
| 2022 | COMPACT: Flow-Based Computing on Nanoscale Crossbars With Minimal Semiperimeter and Maximum DimensionabstractIn-memory computing is a promising solution strategy for data-intensive applications to circumvent the von Neumann bottleneck. Flow-based computing is the concept of performing in-memory computing using sneak paths in nanoscale crossbar arrays. The limitation of the previous work is that the resulting crossbar representations have large size. In this article, we present a framework called COMPACT for mapping Boolean functions to crossbar representations with a minimal semiperimeter (the number of wordlines plus bitlines) and/or maximum dimension (the maximum of the wordlines or bitlines). The COMPACT framework is based on an analogy between binary decision diagrams (BDDs) and nanoscale memristor crossbar arrays. More specifically, nodes and edges in a BDD correspond to wordlines/bitlines and memristors in a crossbar array, respectively. The relation enables a Boolean function represented by a BDD with$n$nodes and an odd cycle transversal of size$k$to be mapped to a crossbar with a semiperimeter of$n+k$. The$k$extra wordlines/bitlines are introduced due to crossbar connection constraints, i.e., wordlines (bitlines) cannot directly be connected to wordlines (bitlines). Moreover, there exists a tradeoff between the semiperimeter and maximum dimension. Consequently, COMPACT can sometimes reduce the maximum dimension by slightly increasing the length of the semiperimeter. We also extend COMPACT to handle multioutput functions using shared BDD (SBDDs) and alignment constraints on the inputs and outputs. Compared with the state-of-the-art mapping technique, the semiperimeter and maximum dimension are reduced by 55% and 85%, respectively. The area, power consumption, and computation delay are reduced by 89%, 19%, 56%, respectively. Sven Thijssen, Sumit Kumar Jha 0001, Rickard Ewetz |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | XMAP: Programming Memristor Crossbars for Analog Matrix-Vector Multiplication: Toward High Precision Using Representable MatricesabstractLinear transformations are the dominating computation within many important applications. The natural multiply-and-accumulate feature of memristor crossbar arrays promise unprecedented processing capabilities to resistive dot-product engines (DPEs), which can accelerate approximate matrix–vector multiplication (MVM). Unfortunately, the precision of the analog computation may be degraded by parasitics, nonlinear device characteristics, and variations. In this article, we propose a framework, called XMAP, for mapping an arbitrary matrix into appropriate memristor conductance values (or state variables for nonlinear devices). The specified conductance values are next programmed to the memristor hardware using accurate closed-loop tuning. XMAP is based on formulating the mapping problem as a mathematical optimization problem, which can be elegantly minimized using the concept of representable matrices, i.e., the matrices that can be represented on a crossbar. Compared to the state-of-the-art conversion algorithm, the computational accuracy is improved with up to$3.29 \times $at the expense of overhead in runtime. The precision improvements translate into noteworthy application-level benefits within signal compression and neural network inference. Necati Uysal, Baogang Zhang, Sumit Kumar Jha 0001, Rickard Ewetz |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | Synthesis of Clock Networks with a Mode-Reconfigurable TopologyabstractModern digital circuits are often required to operate in multiple modes to cater to variable frequency and power requirements. Consequently, the clock networks for such circuits must be synthesized, meeting different timing constraints in different operational modes. The overall power consumption and robustness to variations of a clock network are determined by the topology. However, state-of-the-art clock networks use the same topology in every mode, despite that timing constraints in low- and high-performance modes can be very different. In this article, we propose a clock network with a mode-reconfigurable topology (MRT) for circuits with positive-edge-triggered sequential elements. In high-performance modes, the MRT structure is reconfigured into a near-tree to provide the required robustness to variations. In low-performance modes, the MRT structure is reconfigured into a tree to save power. Non-tree (or near-tree) structures provide robustness to variations by appropriately constructing multiple alternative paths from the clock source to the clock sinks, which neutralizes the negative impact of variations. In MRT structures, OR-gates are used to join multiple alternative paths into a single path. Hence, the MRT structures consume no short-circuit power because there is only one gate driving each net. Moreover, it is straightforward to reconfigure an MRT structure into a tree topology using a single clock gate. In high-performance modes, the experimental results demonstrate that MRT structures have \( 25\% \) lower power consumption than state-of-the-art near-tree structures. In low-performance modes, the power consumption of the MRT structure is similar to the power consumption of a clock tree. Necati Uysal, Rickard Ewetz |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2021 | Towards Resilient Deployment of In-Memory Neural Networks with High ThroughputabstractResistive computing systems (RCSs) promise exascale computing capabilities to inference engines for deep learning. However, the classification accuracy of the accelerated neural networks may be degraded by defects. While hardware-aware training schemes can restore the accuracy of convolutional neural networks (CNNs) with low throughput, the schemes are rendered futile when weights are replicated to improve throughput. On the other hand, we discover that weight replication provides new opportunities for data layout organization. In this paper, we propose a framework for resilient deployment of high throughput CNNs to RCSs. The framework is based on integrating a data layout organization step and a distribution guided training step into the flow for mapping CNNs to RCSs. The data layout organization step involves modifying the weight matrix to crossbar assignments using channel, pixel, and hybrid channel-pixel data layout transformations. The distribution guided training is focused on training CNNs with weights that are amenable for data layout organization. The experimental results demonstrate that the proposed techniques expand the average solution space for data layout organization with 1.4× 1014X. This translates into that CNNs with high throughput can be deployed onto RCS with up to 10% defects and still attain high classification accuracy. Baogang Zhang, Rickard Ewetz |
DAC | 2 |
| 2021 | COMPACT: Flow-Based Computing on Nanoscale Crossbars with Minimal SemiperimeterabstractIn-memory computing is a promising solution strategy for data-intensive applications to circumvent the von Neumann bottleneck. Flow-based computing is the concept of performing in-memory computing using sneak paths in nanoscale crossbar arrays. The limitation of previous work is that the resulting crossbar representations have large dimensions. In this paper, we present a framework called COMPACT for mapping Boolean functions to crossbar representations with minimal semiperimeter (the number of wordlines plus bitlines). The COMPACT framework is based on an analogy between binary decision diagrams (BDDs) and nanoscale memristor crossbar arrays. More specifically, nodes and edges in a BDD correspond to wordlines/bitlines and memristors in a crossbar array, respectively. The relation enables a function represented by a BDD with$n$nodes and an odd cycle transversal of size$k$to be mapped to a crossbar with a semiperimeter of n+k. The$k$extra wordlines/bitlines are introduced due to crossbar connection constraints, i.e. wordlines (bitlines) cannot directly be connected to wordlines (bitlines). For multi-input multi-output functions, COMPACT can also be applied to shared binary decision diagrams (SBDDs), which further reduces the size of the crossbar representations. Compared with the state-of-the-art mapping technique, the semiperimeter is reduced from 2.13n to 1.09n on the average, which translates into crossbar representations with 78% smaller area. The power consumption and the computation delay are on the average reduced by 7% and 52%, respectively. Sven Thijssen, Sumit Kumar Jha 0001, Rickard Ewetz |
DATE | 3 |
| 2021 | Accelerating AI Applications using Analog In-Memory Computing: Challenges and OpportunitiesabstractLinear transformations are the dominating computation within many artificial intelligence (AI) applications. The natural multiply and accumulate feature of resistive crossbar arrays promise unprecedented processing capabilities to resistive dot-product engines (DPEs), which can accelerate approximate matrix-vector multiplication using analog in-memory computing. Unfortunately, the functional correctness of the accelerated AI applications may be compromised by various sources of errors. In this paper, we will outline the most pressing robustness challenges, the limitations of state-of-the-art solutions, and future opportunities for research. Shravya Channamadhavuni, Sven Thijssen, Sumit Kumar Jha 0001, Rickard Ewetz |
ACM Great Lakes Symposium on VLSI | 4 |
| 2021 | Hybrid Analog-Digital In-Memory ComputingabstractToday's high performance computing (HPC) systems are limited by the expensive data movement between processing and memory units. An emerging solution strategy is to perform in-memory computing (IMC) using non-volatile memory. However, state-of-the-art in-memory computing paradigms fail to simultaneously deliver high precision and high energy-efficiency. Analog in-memory computing is extremely energy-efficient but inherently vulnerable to errors. In contrast, digital in-memory computing based on Boolean logic is robust to errors but less energy-efficient. In this paper, we propose a new paradigm called hybrid analog-digital in-memory computing. The paper also proposes the associated in-memory computing platform and design automation tool chain needed to perform computation using the paradigm. The paradigm is capable of performing matrix-vector multiplication with both high energy-efficiency and precision. The key idea of the paradigm is to first decompose the most significant bits (MSBs) of the desired computation into Boolean functions and the least significant bits (LSBs) into matrix-vector multiplication operations. Next, the operations are mapped to digital and analog in-memory computing hardware, respectively. The proposed paradigm is evaluated using applications from the domains of structural engineering, mathematics, and statistics. Compared with analog in-memory computing, the proposed paradigm is capable of meeting the constraints on the computational accuracy. Compared with digital in-memory computing, systems, power, speed, and area are respectively improved with 2.44X, 2.45X and 2.32X. Muhammad Rashedul Haq Rashed, Sumit Kumar Jha 0001, Rickard Ewetz |
ICCAD | 3 |
| 2021 | An OCV-Aware Clock Tree Synthesis MethodologyabstractClosing timing after clock tree synthesis (CTS) is very challenging in the presence of on-chip variations (OCVs). State-of-the-art design flows first synthesize an initial clock tree that contains timing violations introduced by OCVs. Next, aggressive clock tree optimization (CTO) is applied to eliminate the timing violations. Unfortunately, it may be impossible to eliminate all violations given the structure of the initial clock tree. In this paper, we propose an OCV-aware clock tree synthesis methodology that aims to rethink how to account for OCVs. The key idea is to predict the impact of OCVs early in the synthesis process, which allows the variations to be compensated for using non-uniform safety margins. This results in a synthesis flow that is almost correct-by-design. In contrast, state-of-the-art design flows often have an unpredictable success rate because the OCVs are considered too late in the synthesis process. Concretely, this is achieved by top-down constructing a virtual clock tree that is refined bottom-up into a real clock tree implementation. To balance the quality of results (QoR) and runtime, multiple top-level tree topologies are enumerated and pruned in the synthesis process. Compared with the CTO based approach, the experimental results demonstrate that the proposed methodology reduces the total negative slack (TNS) and worst negative slack (WNS) with 90% and 75%, respectively. Necati Uysal, Rickard Ewetz |
ICCAD | 2 |
| 2021 | On Smoother Attributions using Neural Stochastic Differential EquationsabstractSeveral methods have recently been developed for computing attributions of a neural network's prediction over the input features. However, these existing approaches for computing attributions are noisy and not robust to small perturbations of the input. This paper uses the recently identified connection between dynamical systems and residual neural networks to show that the attributions computed over neural stochastic differential equations (SDEs) are less noisy, visually sharper, and quantitatively more robust. Using dynamical systems theory, we theoretically analyze the robustness of these attributions. We also experimentally demonstrate the efficacy of our approach in providing smoother, visually sharper and quantitatively robust attributions by computing attributions for ImageNet images using ResNet-50, WideResNet-101 models and ResNeXt-101 models. Sumit Kumar Jha 0001, Rickard Ewetz, Alvaro Velasquez, Susmit Jha |
IJCAI | 2 |
| 2021 | Automated Synthesis of Quantum Circuits Using Symbolic Abstractions and Decision ProceduresabstractQuantum algorithms are notoriously hard to design and require significant human ingenuity and insight. We present a new methodology called Quantum Automated Synthesizer (QUASH) that can automatically synthesize quantum circuits using decision procedures that perform symbolic reasoning for combinatorial search. Our automated synthesis approach constructs finite symbolic abstract models of the quantum gates automatically and discovers a quantum circuit as a composition of quantum gates using these symbolic models. Our key insight is that most current quantum algorithms work on a finite number of classical inputs, and hence, their correctness proof relies only on reasoning about a finite set of quantum states that can be represented using finite symbolic systems. We demonstrate the potential of our approach by automatically synthesizing four quantum circuits and re-discovering the Bernstein-Vazirani quantum algorithm using state-of-the-art decision procedures. Our synthesis approach only requires distinguishing between a finite set of symbolic quantum states; for example, the synthesis of the Bernstein-Vazirani quantum algorithm only requires reasoning about the following qubit states: |0, |1, -i|0, i|1, |+, |-, e1/2|1i, eiπ/4|1 and a remaining symbolic state representing all other possible quantum states. Our approach leverages decision procedures and theorem provers to assist in the discovery of new quantum algorithms and is a step towards the automation of quantum algorithm design. Alvaro Velasquez, Sumit Kumar Jha 0001, Rickard Ewetz, Susmit Jha |
ISCAS | 3 |
| 2021 | LADDER: Architecting Content and Location-aware Writes for Crossbar Resistive MemoriesabstractResistive memories (ReRAM) organized in the form of crossbars are promising for main memory integration. While offering high cell density, crossbar-based ReRAMs suffer from variable write latency requirement for RESET operations due to the varying impact of IR drop, which jointly depends on the data pattern of the crossbar and the location of target cells being RESET. The exacerbated worst-case RESET latencies can significantly limit system performance. Md Hafizul Islam Chowdhuryy, Muhammad Rashedul Haq Rashed, Amro Awad, Rickard Ewetz, Fan Yao 0001 |
MICRO | 4 |
| 2021 | Computational Restructuring: Rethinking Image Compression Using Resistive Crossbar ArraysabstractImage compression is performed on billions of edge devices deployed in the Internet of Things (IoT). The bottleneck of the compression is the 2-D discrete cosine transform (2D DCT), which involves performing two matrix-matrix multiplications in series. Earlier studies have explored directly mapping the 2D DCT computation to emerging resistive crossbar arrays (RCAs), which promise to perform matrix-vector multiplication (MVM) with extremely small energy-delay product. The main drawback is that the series computation is inherently vulnerable to errors. In this article, we propose to fundamentally rethink how to perform image compression using RCAs. The key idea is to restructure the computation to natively match the properties of the underlying resistive hardware. This allows three of the main design steps within image compression (2D DCT, quantization, and zig-zag reordering) to be integrated into a single analog MVM operation. The integration is facilitated by the development of a 2D DCT reconstruction technique, a frequency spectrum optimization technique, and a quantization optimization technique. The techniques improve the robustness to errors, eliminates the storage of intermediate data, enables processing of small image blocks, facilitates the utilization of large-scale RCAs, and reduces the requirements on the expensive domain interfaces. Compared with the previous work, the experimental results demonstrate significant improvements in image quality while reducing power and latency with up to 62% and 21%, respectively. Baogang Zhang, Necati Uysal, Rickard Ewetz |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2020 | Representable Matrices: Enabling High Accuracy Analog Computation for Inference of DNNs using MemristorsabstractAnalog computing based on memristor technology is a promising solution to accelerating the inference phase of deep neural networks (DNNs). A fundamental problem is to map an arbitrary matrix to a memristor crossbar array (MCA) while maximizing the resulting computational accuracy. The state-of-the-art mapping technique is based on a heuristic that only guarantees to produce the correct output for two input vectors. In this paper, a technique that aims to produce the correct output for every input vector is proposed, which involves specifying the memristor conductance values and a scaling factor realized by the peripheral circuitry. The key insight of the paper is that the conductance matrix realized by an MCA is only required to be proportional to the target matrix. The selection of the scaling factor between the two regulates the utilization of the programmable memristor conductance range and the representability of the target matrix. Consequently, the scaling factor is set to balance precision and value range errors. Moreover, a technique of converting conductance values into state variables and vice versa is proposed to handle memristors with non-ideal device characteristics. Compared with the state-of-the-art technique, the proposed mapping results in 4X-9X smaller errors. The improvements translate into that the classification accuracy of a seven-layer convolutional neural network (CNN) on CIFAR-10 is improved from 20.5% to 71.8%. Baogang Zhang, Necati Uysal, Deliang Fan, Rickard Ewetz |
ASP-DAC | 4 |
| 2020 | Computational Restructuring: Rethinking Image Processing using Memristor Crossbar ArraysabstractImage processing is a core operation performed on billions of sensor-devices in the Internet of Things (IoT). Emerging memristor crossbar arrays (MCAs) promise to perform matrix-vector multiplication (MVM) with extremely small energy-delay product, which is the dominating computation within the two-dimensional Discrete Cosine Transform (2D DCT). Earlier studies have directly mapped the digital implementation to MCA based hardware. The drawback is that the series computation is vulnerable to errors. Moreover, the implementation requires the use of large image block sizes, which is known to degrade the image quality. In this paper, we propose to restructure the 2D DCT into an equivalent single linear transformation (or MVM operation). The reconstruction eliminates the series computation and reduces the processed block sizes from N×N to √N×√N Consequently, both the robustness to errors and the image quality is improved. Moreover, the latency, power, and area is reduced with 2X while eliminating the storage of intermediate data, and the power and area can be further reduced with up to 62% and 74% using frequency spectrum optimization. Baogang Zhang, Necati Uysal, Rickard Ewetz |
DATE | 3 |
| 2020 | Redundant Neurons and Shared Redundant Synapses for Robust Memristor-based DNNs with Reduced OverheadabstractThe dominating computational workload in the inference phase of deep neural networks (DNNs) is matrix-vector multiplication. An arising solution to accelerate the inference phase is to perform analog matrix-vector multiplication using memristor crossbar arrays (MCAs). A key challenge is that stuck-at-fault defects may degrade the classification accuracy of the memristor-based DNNs. A common technique to reduce the negative impact of stuck-at-faults is to utilize redundant synapses, i.e, each row in a weight matrix is realized using two (or r) parallel rows in an MCA. In this paper, we propose to handle stuck-at-faults by inserting redundant neurons and by sharing redundant synapses. The first technique is based on inserting redundant neurons to surgically repair neurons connected to rows and columns in the MCAs with many stuck-at-faults. The second technique is focused on sharing redundant synapses between different neurons to reduce the hardware overhead, which generalizes (1:r) synapse redundancy in previous studies to (q:r) synapse redundancy. The experimental results demonstrate new trade-offs between robustness and hardware overhead without requiring the neural networks to be retrained. Compared with state-of-the-art, the power and area overhead for a neural network can be reduced with up to 16% and 25%, respectively. Baogang Zhang, Necati Uysal, Deliang Fan, Rickard Ewetz |
ACM Great Lakes Symposium on VLSI | 4 |
| 2020 | DP-MAP: Towards Resistive Dot-Product Engines with Improved PrecisionabstractThe natural multiply and accumulate feature of memristor crossbar arrays promises unprecedented processing capabilities to resistive dot-product engines (DPEs), which can accelerate approximate matrix-vector multiplication. To overcome the challenges of low-precision devices and voltage drop over non-zero array parasitics, each matrix element can be represented using two memristors. In this paper, we propose differential pair map (DP-MAP) - the first matrix to memristor conductance mapping algorithm specifically designed for crossbars with a differential pair configuration. In contrast, previous works consider the differential pair configuration as an afterthought, which limits the achievable precision. The specified conductance values are next programmed to the memristor hardware using accurate closed-loop tuning. Analog computation with high precision is attained by judiciously selecting the conductance range and avoiding to explicitly decompose each matrix into a positive and negative component. Short run-time is achieved using a hierarchical optimization algorithm and two speed-up techniques. Compared with earlier studies, the computational accuracy is improved with 3.36X. This translates into signal and image compression with 61% and 94% higher quality, respectively. The simulation time of complex physical systems modeled using partial differential equations (PDEs) is reduced with 5.87X. Necati Uysal, Baogang Zhang, Sumit Kumar Jha 0001, Rickard Ewetz |
ICCAD | 4 |
| 2020 | Synthesis of Clock Networks with a Mode Reconfigurable Topology and No Short Circuit CurrentabstractCircuits deployed in the Internet of Things operate in low and high performance modes to cater to variable frequency and power requirements. Consequently, the clock networks for such circuits must be synthesized meeting drastically different timing constraints under variations in the different modes. The overall power consumption and robustness to variations of a clock network is determined by the topology. However, state-of-the-art clock networks use the same topology in every mode, despite that the timing constraints in the low and high performance modes are very different. In this paper, we propose a clock network with a mode reconfigurable topology (MRT) for circuits with positive-edge triggered sequential elements. In high performance modes, the required robustness to variations is provided by reconfiguring the MRT structure into a near-tree. In low performance modes, the MRT structure is reconfigured into a tree to save power. Non-tree (or near-tree) structures provide robustness to variations by appropriately constructing multiple alternative paths from the clock source to the clock sinks, which neutralizes the negative impact of variations. In MRT structures, OR-gates are used to join multiple alternative paths into a single path. Consequently, the MRT structures consume no short circuit power because there is only one gate driving each net. Moreover, it is straightforward to reconfigure MRT structures into a tree by gating the clock signal in part of the structure. Compared with state-of-the-art near-tree structures, MRT structures have 8% lower power consumption and similar robustness to variations in high performance modes. In low performance modes, the power consumption is 16% smaller when reconfiguration is used. Necati Uysal, Juan Ariel Cabrera, Rickard Ewetz |
ISPD | 3 |
| 2020 | Handling Stuck-at-Fault Defects Using Matrix Transformation for Robust Inference of DNNsabstractMatrix-vector multiplication is the dominating computational workload in the inference phase of deep neural networks (DNNs). Memristor crossbar arrays (MCAs) can efficiently perform matrix-vector multiplication in the analog domain. A key challenge is that memristor devices may suffer stuck-at-fault defects, which can severely degrade the classification accuracy. Earlier studies have shown that the accuracy loss can be recovered by utilizing additional hardware or hardware aware training. In this article, we propose a framework that handles stuck-at-faults using matrix transformations, which is called the MT framework. The framework is based on introducing a cost metric that captures the negative impact of the stuck-atfault defects. Next, the cost metric is minimized by applying matrix transformations T. A transformation T changes a weight matrix W into a new weight matrix W̃ = T(W). In particular, a row flipping transformation, a permutation transformation, and a value range transformation are proposed. The row flipping transformation results in that stuck-off (stuck-on) faults are translated into stuck-on (stuck-off) faults. The permutation transformation maps small (large) weights to memristors stuck-off (stuck-on). The value range transformation is based on reducing the magnitude of the smallest and largest elements in the weight matrices, which results in that the stuck-at-faults introduce smaller errors. The experimental results demonstrate that the MT framework is capable of recovering 99% of the accuracy loss on both the MNIST and CIFAR-10 datasets without utilizing hardware aware training. The accuracy improvements come at the expense of an 8.19× and 9.23× overhead in power and area, respectively. Nevertheless, the overhead can be reduced with up to 50% by leveraging hardware aware training. Baogang Zhang, Necati Uysal, Deliang Fan, Rickard Ewetz |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2019 | Latency constraint guided buffer sizing and layer assignment for clock trees with useful skewabstractClosing timing using clock tree optimization (CTO) is a tremendously challenging problem that may require designer intervention. CTO is performed by specifying and realizing delay adjustments in an initially constructed clock tree. Delay adjustments are typically realized by inserting delay buffers or detour wires. In this paper, we propose a latency constraint guided buffer sizing and layer assignment framework for clock trees with useful skew, called the (BLU) framework. The BLU framework realizes delay adjustments during CTO by performing buffer sizing and layer assignment. Given an initial clock tree, the BLU framework first predicts the final timing quality and specifies a set of delay adjustments, which are translated into latency constraints. Next, buffer sizing and layer assignment is performed with respect to the latency constraints using an extension of van Ginneken's algorithm. Moreover, the framework includes a feature of reducing the power consumption by relaxing the latency constraints and a method of improving the timing performance by tightening the latency constraints. The experimental results demonstrate that the proposed framework is capable of reducing the capacitive cost with 13% on the average. The total negative slack (TNS) and worst negative slack (WNS) are reduced with up to 58% and 20%, respectively. Necati Uysal, Wen-Hao Liu 0001, Rickard Ewetz |
ASP-DAC | 3 |
| 2019 | Handling stuck-at-faults in memristor crossbar arrays using matrix transformationsabstractMatrix-vector multiplication is the dominating computational workload in the inference phase of neural networks. Memristor crossbar arrays (MCAs) can inherently execute matrix-vector multiplication with low latency and small power consumption. A key challenge is that the classification accuracy may be severely degraded by stuck-at-fault defects. Earlier studies have shown that the accuracy loss can be recovered by retraining each neural network or by utilizing additional hardware. In this paper, we propose to handle stuck-at-faults using matrix transformations. A transformation T changes a weight matrix W into a weight matrix, @ = T(W), which is more robust to stuck-at-faults. In particular, we propose a row flipping transformation, a permutation transformation, and a value range transformation. The row flipping transformation results in that stuck-off (stuck-on) faults are translated into stuck-on (stuck-off) faults. The permutation transformation maps small (large) weights to memristors stuck-off (stuck-on). The value range transformation is based on reducing the magnitude of the smallest and largest elements in the matrix, which results in that each stuck-at-fault introduces an error of smaller magnitude. The experimental results demonstrate that the proposed framework is capable of recovering 99% of the accuracy loss introduced by stuck-at-faults without requiring the neural network to be retrained. Baogang Zhang, Necati Uysal, Deliang Fan, Rickard Ewetz |
ASP-DAC | 4 |
| 2019 | Noise Injection Adaption: End-to-End ReRAM Crossbar Non-ideal Effect Adaption for Neural Network MappingabstractIn this work, we investigate various non-ideal effects (Stuck-At-Fault (SAF), IR-drop, thermal noise, shot noise, and random telegraph noise)of ReRAM crossbar when employing it as a dot-product engine for deep neural network (DNN) acceleration. In order to examine the impacts of those non-ideal effects, we first develop a comprehensive framework called PytorX based on main-stream DNN pytorch framework. PytorX could perform end-to-end training, mapping, and evaluation for crossbar-based neural network accelerator, considering all above discussed non-ideal effects of ReRAM crossbar together. Experiments based on PytorX show that directly mapping the trained large scale DNN into crossbar without considering these non-ideal effects could lead to a complete system malfunction (i.e., equal to random guess) when the neural network goes deeper and wider. In particular, to address SAF side effects, we propose a digital SAF error correction algorithm to compensate for crossbar output errors, which only needs one-time profiling to achieve almost no system accuracy degradation. Then, to overcome IR drop effects, we propose a Noise Injection Adaption (NIA) methodology by incorporating statistics of current shift caused by IR drop in each crossbar as stochastic noise to DNN training algorithm, which could efficiently regularize DNN model to make it intrinsically adaptive to non-ideal ReRAM crossbar. It is a one-time training method without the request of retraining for every specific crossbar. Optimizing system operating frequency could easily take care of rest non-ideal effects. Various experiments on different DNNs using image recognition application are conducted to show the efficacy of our proposed methodology. Zhezhi He, Jie Lin 0004, Rickard Ewetz, Jiann-Shiun Yuan, Deliang Fan |
DAC | 3 |
| 2019 | STAT: Mean and Variance Characterization for Robust Inference of DNNs on Memristor-based PlatformsabstractAn emerging solution to accelerate the inference phase of deep neural networks (DNNs) is to utilize memristor crossbar arrays (MCAs) to perform highly efficient matrix-vector multiplication in the analog domain. An adverse challenge is that memristor devices may suffer stuck-at-fault defects, which may compromise the classification accuracy. Stuck-at-fault defects have previously been handled by neuron permutation or by retraining neural networks. In this paper, we propose the STAT framework that utilizes statistics to guide optimization techniques that provide robustness to stuck-at-fault defects. In particular, bias weights are modified to minimize the input error to each neuron with respect to an input vector. The input vector is selected to be equal to the mean from a statistical characterization. Variance statistics are used to define a weight significance metric, which is used to prioritize assigning weights connected to neurons with large (small) variance to non-defective (defective) memristors using neuron permutation, as errors introduced by neurons with small variance can be eliminated by modifying the bias weights. The experimental results demonstrate that the STAT framework improves the normalized classification accuracy from 62.1% to 96.1% without any hardware overhead. Baogang Zhang, Necati Uysal, Rickard Ewetz |
ACM Great Lakes Symposium on VLSI | 3 |
| 2019 | Scalable Construction of Clock Trees With Useful Skew and High Timing QualityabstractClock trees can be constructed based on static arrival time constraints or dynamic implied skew constraints. Dynamic implied skew constraints allow the full timing margins to be utilized. However, the dynamic skew constraints require a high run-time complexity to be evaluated. In contrast, static arrival time constraints are more restrictive but can be evaluated in constant time. Consequently, there is a tradeoff between timing margin utilization and run-time. In this paper, a scalable clock tree synthesis (CTS) framework is proposed for the construction of low-cost useful skew trees (USTs) with high timing quality. The scalability is based on combining the use of arrival time constraints with virtual minimum and maximum delay offsets, which facilitates that a pair of smaller subtrees can be joined into a larger subtree in constant time. The ability to quickly join subtrees is leveraged to perform a high degree of solution space exploration, which translates into the construction of USTs with low-cost. In particular, clock trees with various routing tree topologies, buffer tree topologies, buffer sizes, and stem wire lengths are explored. Moreover, the arrival time constraints are specified with the objective of being the least restrictive to reduce cost. Furthermore, the constraints are respecified throughout the tree construction process using a slack graph (SG) to expose additional timing margins. The high timing quality is obtained by seamlessly integrating arbitrary timing models using the SG. Finally, the proposed CTS framework is integrated with a clock tree optimization framework to demonstrate that the constructed USTs are capable of meeting timing constraints under the influence of on-chip variations. Rickard Ewetz, Cheng-Kok Koh |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2018 | Clustering of flip-flops for useful-skew clock tree synthesisabstractThe clock network of a circuit is a main contributor to the power consumption of any ASIC design. A key technique that is used to reduce power consumption is to cluster flipflops or latches into groups and to place each group of flipflops close together to reduce the clock wire length. In this paper, we introduce a clock tree synthesis methodology that incorporates clustering with a previously published useful-skew clock tree synthesis technique to minimize the clock wire length. The clustering process is guided by bounded arrival time constraints, which enable its efficiency. Experimental results show that the proposed methodology reduces up to 34% of the total power consumption while meeting all timing constraints. Chuan Yean Tan, Rickard Ewetz, Cheng-Kok Koh |
ASP-DAC | 2 |
| 2018 | OCV guided clock tree topology reconstructionabstractThe timing performance of clock trees in scaled technology nodes may be severely degraded by on-chip variations (OCV). Clock tree optimization (CTO) is employed to eliminate timing violations by specifying a set of non-negative delay adjustments using a linear programming (LP) formulation. Next, the delay adjustments are realized in the clock tree by inserting delay buffers and detour wires. The drawback is that given the topology of the initial clock tree, it may be impossible to remove all timing violations. In this paper, a framework that performs OCV guided clock tree topology reconstruction is proposed. The framework reconstructs the topology of a clock tree while improving the lower bounds on the worst negative slack (WNS) and the total negative slack (TNS). Next, traditional CTO is employed to reduce WNS and TNS to the improved lower bounds. The reconstruction of the clock tree topology is guided by a predicted leaf buffer slack graph (pLB-SG). The leaf buffers that must be placed closer in the tree topology are identified by detecting cycles (or strongly connected components) in the pLB-SG. The experimental results demonstrate that the proposed framework can on the average reduce WNS and TNS with 84% and 80%, respectively. Necati Uysal, Rickard Ewetz |
ASP-DAC | 2 |
| 2018 | Software and Hardware Techniques for Reducing the Impact of Quantization Errors in Memristor Crossbar ArraysabstractMatrix-vector multiplication is the dominating computational workload in the evaluation of neural networks. It has recently been demonstrated that memristor crossbar arrays (MCAs) can perform matrix-vector multiplication with small power consumption and low latency. However, the computational accuracy may be degraded by quantization errors. The quantization errors of mapping a matrix to an MCA are proportional to the the number of distinguishable states of each memristor and the difference between the largest and smallest element in the matrix. In this paper, we propose a framework for mapping an arbitrary matrix A to a grid of MCAs (or a single MCA) while minimizing the negative impact of quantization errors. The framework is guided by a total quantization error bound (TQEB) metric, which is an upper bound on the total quantization errors (TQE). Using the proposed TQEB metric, three techniques of reducing TQE are proposed. The first method is based on scaling and shifting the rows in A with different factors to improve the memristor conductance band utilization. The second technique is based on representing a single column in a matrix A using multiple columns in an MCA, to reduce the magnitude of the smallest and largest elements in A. The third technique is based on permuting the order of the columns in A when the matrix A is required to be mapped to a grid of MCAs. The quantization errors are reduced by assigning matrix values of similar magnitude to the same MCAs in the grid. The experimental results demonstrate that the proposed metric and techniques are capable of greatly reducing the negative impact of quantization errors. Baogang Zhang, Rickard Ewetz |
ICCD | 2 |
| 2017 | Delay-driven layer assignment for advanced technology nodesabstractThis paper addresses a delay-driven layer assignment problem with consideration of via delay and coupling effect in the global routing stage. A negotiation-based framework is proposed to balance delay, congestion, and via count. Coupling capacitance is considered using a probabilistic look-up table. Finally, the proposed algorithm uses both parallel wires and wide wires to reduce wire delay. The effectiveness of our layer assignment algorithm is supported by extensive experimental results. Szu-Yuan Han, Wen-Hao Liu 0001, Rickard Ewetz, Cheng-Kok Koh, Kai-Yuan Chao, Ting-Chi Wang |
ASP-DAC | 3 |
| 2017 | A Clock Tree Optimization Framework with Predictable Timing QualityabstractEliminating timing violations using clock tree optimization (CTO) persist to be a tedious problem in ultra scaled technologies. State-of-the-art CTO techniques are based on predicting the final timing quality by specifying a set of delay adjustments in the form of delay adjustment points (DAPs). Next, the DAPs are realized to eliminate the timing violations. Unfortunately, it is difficult to realize delay adjustments of exact magnitudes. In this paper, the correlation between the predicted and achieved timing quality is improved by specifying delay adjustments in the form of delay adjustment ranges (DARs). The DARs are formed such that the predicted timing quality is achieved if each delay adjustment is realized within the respective DAR. The framework first predicts the final timing quality. Next, the DARs are specified and optimized while treating the predicted timing quality as a constraint. The optimization is a trade-off between the ease of delay adjustment realization, the total amount of delay adjustment, and the number of delay adjustments. Moreover, the framework accounts for delay adjustment induced on-chip variations (OCV) and transition time constraints. On a set of synthesized circuits, it is demonstrated that the framework improves the correlation between the predicted and achieved timing quality compared with in earlier studies. The average total negative slack (TNS) and worst negative slack (WNS) are improved by 91% and 85%, respectively. Rickard Ewetz |
DAC | 1 |
| 2017 | Clock Tree Construction based on Arrival Time ConstraintsabstractThere are striking differences between constructing clock trees based on dynamic implied skew constraints and based on static arrival time constraints. Dynamic implied skew constraints allow the full timing margins to be utilized, but the constraints are required to be updated (with high time complexity). In contrast, static arrival time constraints are decoupled and are not required to be updated. Therefore, the constraints can be obtained in constant time, which facilitates the exploration of various tree topologies. On the other hand, arrival time constraints do not allow the full timing margins to be utilized. Consequently, there is a trade-off between topology exploration and timing margin utilization. In this paper, the advantages of static arrival time constraints are leveraged to construct clock trees with useful skew while exploring various tree topologies. Moreover, the constraints are specified and respecified throughout the synthesis process reduce the cost of the constructed clock trees. It is experimentally demonstrated that the proposed approach results in clock trees with 16% lower average capacitive cost compared with clock trees constructed based on dynamic implied skew constraints. Rickard Ewetz, Cheng-Kok Koh |
ISPD | 1 |
| 2017 | Fast clock scheduling and an application to clock tree synthesis
Rickard Ewetz, Cheng-Kok Koh |
Integr. | 1 |
| 2016 | MCMM clock tree optimization based on slack redistribution using a reduced slack graphabstractModern clock networks are required to operate in multiple corners and in multiple modes (MCMM). An initially constructed clock tree may contain different timing violations in different mode and corner combinations. Clock tree optimization (CTO) is employed to remove these timing violations. We propose a CTO framework based on slack redistribution using a reduced slack graph. The main idea is to reduce the MCMM problem to an equivalent single-corner single-mode (SCSM) problem using delay adjustment linearization. Using the equivalent SCSM problem, a linear program is solved to determine a set of delay adjustments to remove the timing violations. Next, the delay adjustments are realized using feasible delay adjustment ranges. The experimental results show that the proposed framework obtains average reductions of 84% and 83% in the total negative slack and the worst negative slack, respectively, at the expense of a 4% capacitive overhead. Rickard Ewetz, Cheng-Kok Koh |
ASP-DAC | 1 |
| 2016 | Construction of Latency-Bounded Clock TreesabstractClock trees must be constructed to function even under the influence of on-chip variations (OCV). Bounding the latency of a clock tree, i.e., the maximum delay from the tree root to any sequential element, is important because the latency correlates with the maximum magnitude of the skews caused by OCV. In this paper, a latency constraint graph (LCG) that captures the latencies of a set of subtrees and the skew constraints between the subtrees is introduced. The minimum latency of a clock tree that can be constructed from the corresponding subtrees is equal to the (negative of the) length of a shortest path in the LCG, which can be computed in $O(VE)$. Based on the LCG, we propose a framework that consists of a latency-aware clock tree synthesis (CTS) phase and a clock tree optimization (CTO) phase to construct latency-bounded clock trees. When applied to a set of synthesized circuits, the framework is capable of constructing latency-bounded clock trees that have higher yield compared to clock trees constructed in previous studies. Rickard Ewetz, Chuan Yean Tan, Cheng-Kok Koh |
ISPD | 1 |
| 2016 | Construction of Reconfigurable Clock Trees for MCMM Designs Using Mode Separation and Scenario CompressionabstractThe clock networks of many modern circuits have to operate in multiple corners and multiple modes (MCMM). We propose to construct mode-reconfigurable clock trees (MRCTs) based on mode separation and scenario compression. The technique of scenario compression is proposed to consider the timing constraints in multiple scenarios at the same time, compressing the MCMM problem into an equivalent single-corner multiple-mode (SCMM), or single-corner single-mode (SCSM) problem. The compression is performed by combining the skew constraints of the different scenarios in skew constraint graphs based on delay linearization and dominating skew constraints. An MRCT consists of several clock trees and mode separation involves, depending on the active mode, selecting one of the clock trees to deliver the clock signal. To limit the overhead, the bottom part (closer to the clock sinks) of all the different clock trees are shared and only the top part (closer to the clock source) of the clock network is mode reconfigurable. The reconfiguration is realized using OR-gates and a one-input-multiple-output demultiplexer. The experimental results show that for a set of synthesized MCMM circuits, with 715 to 13, 216 sequential elements, the proposed approach can achieve high yield. Rickard Ewetz, Cheng-Kok Koh |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2015 | Fast clock skew scheduling based on sparse-graph algorithmsabstractIncorporating timing constraints explicitly imposed by the data and control paths during clock network synthesis can enhance the robustness of the synthesized clock networks. With these constraints, a clock scheduler can be used to guide the synthesis of a clock network by specifying a set of feasible arrival times at the respective sequential elements. Clock scheduling can be either static or dynamic. In static clock scheduling, a clock schedule is first specified; next, a clock network is constructed realizing the prescribed schedule. Clock trees constructed using this approach may consume significant routing resources. In dynamic clock scheduling, the clock tree and clock schedule are both simultaneously constructed and determined, respectively. In earlier studies, the scalability of dynamic clock scheduling, which is essentially a shortest path problem, has been limited. The bottleneck is in finding the shortest paths between different vertices in an incrementally changing weighted graph. In this work, we present two clock schedulers that address the scalability issues by exploiting the sparsity of this weighted graph. Experimental results show that the proposed clock schedulers are one to two orders of magnitude faster compared to a published scheduler in an earlier work. The proposed clock schedulers are scalable, and are tested on a synthesized circuit with 348 710 cells, 57 491 sequential elements, and 496 727 explicit timing constraints. Rickard Ewetz, Shankarshana Janarthanan, Cheng-Kok Koh |
ASP-DAC | 1 |
| 2015 | Construction of reconfigurable clock trees for MCMM designsabstractThe clock networks of modern circuits must be able to operate in multiple corners and multiple modes (MCMM). Earlier studies on clock network synthesis for MCMM designs focus on the legalization of an initial clock network that has timing violations in different corners or modes. We propose a mode reconfigurable clock tree (MRCT) that is based on a correct-by-construction approach. An MRCT consists of multiple clock trees. Depending on the active mode, the MRCT is reconfigured such that one of the clock trees is activated to deliver the clock signal. To limit the overhead, the bottom part of the network (closer to the clock sinks) is shared among all of the clock trees, and only the top part of the network (closer to the clock source) is mode reconfigurable. The reconfiguration is realized using or-gates and a single one-input-multiple-output demultiplexer. The MRCT is constructed in a bottom-up fashion by iteratively merging subtrees to form larger subtrees. When two subtrees cannot be merged because of mode-incompatible constraints, an or-gate is inserted to separate the incompatible modes. Corner-incompatible constraints are resolved by reducing safety margins of appropriate skew constraints. The experimental results show that for a set of synthesized MCMM circuits with 715 to 13; 216 sequential elements, the proposed approach can achieve high yield. Rickard Ewetz, Shankarshana Janarthanan, Cheng-Kok Koh |
DAC | 1 |
| 2015 | A Useful Skew Tree Framework for Inserting Large Safety MarginsabstractThe construction of clock trees for modern designs is challenging because the clock trees need to be constructed with adequate safety margins such that the skew constraints are satisfied even under variations. The amount of safety margin required in a skew constraint is dependent on the distance of the corresponding sequential elements in the tree topology. In certain cases, the amount of safety margin that can be inserted may be limited. Consequently, the corresponding sequential elements should be placed close in the topology, i.e., the point of divergence to these elements is low in the clock tree, in order to reduce the influence of variations. By using safety margins and lowering the point of divergence, we present a framework for the construction of useful skew trees with large safety margins inserted in the skew constraints. The framework, called UST-LSM, first identifies tight skew constraints by the detection of negative cycles in a weighted skew constraint graph. Next, the corresponding sequential elements of these skew constraints are clustered early in tree topology. Compared to earlier studies, we can allow larger safety margins in skew constraints spanning between sequential elements within a subtree. This translates into an improvement of yield from 46.8% to 98.8% on a synthesized benchmark with 7,674 sequential elements and 63,440 skew constraints. Rickard Ewetz, Cheng-Kok Koh |
ISPD | 1 |
| 2015 | Cost-Effective Robustness in Clock Networks Using Near-Tree StructuresabstractClock trees are commonly used to deliver clock signals to sequential elements in circuits. However, by construction, tree structures are inherently prone to failure caused by variations. The robustness of a clock tree can be improved by inserting redundancy in the form of cross links or multilevel fusion trees. Such near-tree structures can provide robustness at low cost. In this paper, we establish that the locations of the inserted redundancy are crucial in providing cost-effective robustness. We present two methods to systematically insert redundancy. The redundancy is realized by either inserting cross links or performing local merges. Moreover, we present a vertex reduction method that reduces the amount of redundancy that needs to be inserted in our near-tree structures. Empirical results show that our structures are more robust to variations and have lower power consumption compared to the state-of-the-art clock networks. Furthermore, our near-tree structures provide smooth trade-offs between cost and robustness, reducing clock skews by 11%-39% at an expense of 3%-68% higher power consumption. Rickard Ewetz, Cheng-Kok Koh |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2014 | A study on the use of parallel wiring techniques for sub-20nm designsabstractWire sizing can be used to reduce the delays of critical nets. However, because of the forbidden pitch issue in sub-20nm designs, wide wires may no longer be an attractive solution because of the restrictive wire spacing requirement from advanced lithography. In this work, we investigate the suitability of the parallel wiring technique, in which multiple parallel wires are used to route the same net, as an alternative to routing a net using a single wide wire. In particular, we study the trade offs between parasitics, timing, power, and routing resources. Our study reveals that wire sizing using both parallel wires and wide wires can be advantageous. Moreover, if high layout densities are required, parallel wiring can be a viable approach in solving timing problems for sub-20nm designs. Rickard Ewetz, Wen-Hao Liu 0001, Kai-Yuan Chao, Ting-Chi Wang, Cheng-Kok Koh |
ACM Great Lakes Symposium on VLSI | 1 |
| 2014 | A TSV-cross-link-based approach to 3D-clock network synthesis for improved robustnessabstractTo obtain high yield for 3D ICs, random open defects, process variations, and thermal induced stress are key issues that must be addressed when synthesizing 3D clock networks. Current research on 3D clock synthesis often focuses on the construction and optimization of a 3D clock tree topology. Moreover, extra circuitry has been proposed to enable pre-bond testing and substitution of through silicon vias (TSVs) with random open defects. However, tree structures inherently have limited robustness to variations and may suffer failures arising from defects and/or process variations. To counter such problems, we propose to use TSVs to add redundancy in a 3D clock network. The proposed 3D network would have a complete 2D clock network on each die, facilitating pre-bond testing. Also, cross links would be inserted within each die using wires and across dies using TSVs to improve timing robustness within each die and across dies, respectively. Moreover, clock buffers are placed outside of zones that have high TSV-induced stress that could influence carrier mobility. Experimental results show that the proposed 3D clock networks have no failures due to random open defects, and on the average have 53% lower skew compared to 3D tree structures. Rickard Ewetz, Anirudh Udupa, Ganesh Subbarayan, Cheng-Kok Koh |
ACM Great Lakes Symposium on VLSI | 1 |
| 2013 | Local merges for effective redundancy in clock networksabstractProcess and environmental variations affect the reliability of clock networks. By synthesizing non-tree structures, the robustness of clock networks can be improved at the expense of higher capacitance. A cheap way of converting a tree structure to a non-tree structure is to insert cross links. Unfortunately, the robustness seems to improve only when the links are sufficiently short. Other non-tree structures such as meshes and multilevel fusion trees improve the robustness more effectively, but with much higher cost. In this work, we develop a new non-tree topology by merging a sub-clock tree with all other sub-clock trees that contain sequential elements that require strict synchronization. Results show that when compared with the state-of-the-art solutions, clock networks constructed with the proposed structure have similar capacitance but notable improved robustness. moreover, the clock networks can satisfy tight skew constraints even when simulated under a more stringent variations model, with 22% lower capacitance when compared to solutions in earlier studies. Rickard Ewetz, Cheng-Kok Koh |
ISPD | 1 |