VLDB 2026 Research / reviewers in the wild / expert
Chaitali Chakrabarti
dblp:45/2824
· DBLP profile ↗
165ranked-venue papers
12as first author
34since 2021 · last 2026
0000-0002-9859-7778ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 109 · 6 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 45 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 13 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 1 since 2021Computer networks · 6 · 3 since 2021Software engineering, systems software and programming languages · 5 · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Chiplet-NAS: Chiplet-aware Neural Architecture Search for Efficient AI Inference on 2.5D IntegrationabstractThe co-design of neural network architectures and their target chiplet-based hardware systems presents a significant challenge due to the vast and combinatorial design space. Identifying solutions that are Pareto-optimal across competing objectives of task accuracy, system latency, and power consumption requires solutions beyond manual design and brute-force methods. This paper proposes a closed-loop chiplet-aware neural architecture search (Chiplet-NAS) framework to automate the exploration and discover hardware-optimized models for efficient AI inference on 2.5 D chiplet-based systems. The framework integrates a Tree-structured Parzen Estimator (TPE) for sampleefficient search with CLAIRE, a chiplet-based library and fast performance benchmarking tool, to provide direct hardware feedback on latency and energy consumption, along with accuracy optimization. The framework is evaluated by co-designing ResNet-based model architectures with chiplet based hardware systems. Compared to a baseline NAS that optimizes only on the task accuracy, our Chiplet-NAS achieves significant power and performance benefits at the iso-accuracy. Pragnya Sudershan Nalla, Nikhil K. Cherukuri, Sachin S. Sapatnekar, Chaitali Chakrabarti, Yu Cao 0001, Jeff Zhang 0001 |
ASP-DAC | 6 |
| 2026 | A portable framework with generalized runtime features for task graph execution and concurrent multi-application deployment on heterogeneous systems
Serhan Gener, Md Sahil Hassan, Hasan Umut Suluhan, Liangliang Chang, Chaitali Chakrabarti, Tsung-Wei Huang, Ümit Y. Ogras, Ali Akoglu |
Future Gener. Comput. Syst. | 5 |
| 2026 | CADAS: Communication-Aware Dynamic Scheduler on CGRAs for Large-Volume and Real-Time ProcessingabstractModern data-intensive applications demand accelerators that can adapt to dynamic and high-throughput workloads. Coarse-Grained Reconfigurable Arrays (CGRAs) have emerged as promising candidates for such workloads due to their spatial architecture and run-time reconfigurability. However, ad-hoc hardware configurations and traditional static compilation techniques struggle to cope with the run-time irregularity and control-flow dynamism. This article first presents a systematic design space exploration (DSE) to identify the optimized hardware configurations tailored to application-specific constraints, such as area budget, throughput requirement, and throughput efficiency. Then, it proposes a communication-aware dynamic scheduling approach built on a hardware/software co-design that combines preloading and scoreboard mechanisms to minimize reconfiguration overhead while maximizing interconnect bandwidth utilization. Evaluated on the optimized configurations and the respective spectrum sensing benchmarks, the proposed scheduling method achieves up to 1.6× performance improvement over a baseline and 1.3× over an adapted state-of-the-art (SOTA) dynamic scheduling strategy. Hasan Umut Suluhan, Chaitali Chakrabarti, Ali Akoglu, Ümit Y. Ogras |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2025 | Invited: EDA for Heterogeneous IntegrationabstractThe advent of heterogeneous integration (HI) places new demands on EDA tooling. Building large systems requires (1) methods for chiplet disaggregation that map the system to smaller chiplets, working in conjunction with system-technology co-optimization to determine the right design decisions that optimize computation and communication, together with the choice of substrate and chiplet technologies; (2) multiphysics and multiscale analyses that incorporate thermomechanical aspects into performance analysis, ranging from fast machine-learningdriven analyses in early stages to signoff-quality multiphysics-based analysis; (3) physical design techniques for placing and routing chiplets and embedded active/passive elements on and within the substrate, including the design of thermal and power delivery solutions; and (4) underlying infrastructure required to facilitate HI-based design, including the design and characterization of chiplet libraries and the establishment of data formats and standards. This paper overviews these issues and lays out a set of EDA needs for HI designs. Emad Haque, Pragnya Sudershan Nalla, Chetal Choppali Sudarshan, Divya Yogi, Chaitali Chakrabarti, Vidya A. Chhabria, Ramesh Harjani, Jeff Zhang 0001, Sachin S. Sapatnekar |
DAC | 6 |
| 2025 | CLAIRE: Composable Chiplet Libraries for AI InferenceabstractArtificial intelligence has made a significant impact on fields like computer vision, Natural Language Processing (NLP), healthcare, and robotics. However, recent AI models, such as GPT-4 and LLaMAv3, demand significant number of computational resources, pushing monolithic chips to their technological and practical limits. 2.5D chiplet-based heterogeneous architectures have been proposed to address these technological and practical limits. While chiplet optimization for models like Convolutional Neural Networks (CNNs) is well-established, scaling this approach to accommodate diverse AI inference models with different computing primitives, data volumes, and different chiplet sizes is very challenging. A set of hardened IPs and chiplet libraries optimized for a broad range of AI applications is proposed in this work. We derive the set of chiplet configurations that are composable, scalable and reusable by employing an analytical framework trained on a diverse set of AI algorithms. Testing these set of library synthesized configurations on a different set of algorithms, we achieve a$1.99\times-3.99\times$improvement in non-recurring engineering (NRE) chiplet design costs, with minimal performance overhead compared to custom chiplet-based ASIC designs. Similar to soft IPs for SoC development, the library of chiplets improves flexibility, reusability, and efficiency for AI hardware designs. Pragnya Sudershan Nalla, Emad Haque, Yaotian Liu, Sachin S. Sapatnekar, Jeff Zhang 0001, Chaitali Chakrabarti, Yu Cao 0001 |
DATE | 6 |
| 2025 | ML4SODA: A Decision Tree Guided Design Space Exploration for Fast and High Quality MLIR-based HLS
Darshith Manjunath, Nicolas Bohm Agostini, Antonino Tumeo, Jeff Zhang 0001, Chaitali Chakrabarti |
ACM Great Lakes Symposium on VLSI | 5 |
| 2025 | K-PACT: Kernel Planning for Adaptive Context Switching - A Framework for Clustering, Placement, and Prefetching in Spectrum SensingabstractEfficient wideband spectrum sensing requires rapid evaluation and re-evaluation of signal presence and type across multiple subchannels. These tasks involve multiple hypothesis testing, where each hypothesis is implemented as decision tree workflow with compute-intensive kernels, including FFT, matrix operations, and signal-specific analyses. Given the dynamic nature of the spectrum environment, the ability to quickly switch between hypotheses is essential for maintaining low-latency, high-throughput operation. This work assumes a coarse-grained reconfigurable architecture consisting of an array of processing elements (PEs), each equipped with a local instruction memory (IMEM) capable of storing and executing kernels used in spectrum sensing applications. We propose a planner tool that efficiently maps hypothesis workflows onto this architecture to enable fast runtime context switching with minimal overhead. The planner performs two key tasks: clustering temporally non-overlapping kernels to share IMEM resources within a PE sub-array, and placing these clusters onto hardware to ensure efficient scheduling and data movement. By preloading kernels that are not simultaneously active into the same IMEM, our tool enables low-latency reconfiguration without runtime conflicts. It models the planning process as a multi-objective optimization, balancing trade-offs among context switch overhead, scheduling latency, and dataflow efficiency. We evaluate the proposed tool in simulated spectrum sensing scenario with 48 concurrent subchannels. Results show that our approach reduces off-chip binary fetches by 207.81×, lowers average switching time by 98.24×, and improves per-subband execution time by 132.92× over baseline without preloading. These improvements demonstrate that intelligent planning is critical for adapting to fast-changing spectrum environments in next-generation radio frequency systems. Hasan Umut Suluhan, Serhan Gener, Chaitali Chakrabarti, Ümit Y. Ogras, Ali Akoglu |
ICCAD | 4 |
| 2025 | Mitigating Overfitting During Speech Foundation Model Fine-tuning: Applications to Dysarthric Speech Detection
Yan Xiong 0002, Visar Berisha, Julie M. Liss, Chaitali Chakrabarti |
INTERSPEECH | 4 |
| 2025 | TEE-SFL: Time and Energy-efficient solution for addressing communication heterogeneity in Split Federated Learning SchemesabstractSplit Federated Learning (SFL) is a privacy-preserving distributed machine learning framework suitable for resource-constrained clients. It splits the neural network model between client and server such that most of the computations are offloaded to the server, without sharing private client raw data to server. One drawback of the system is the large communication cost caused by frequent activations/gradients exchanges between client and server. Previous SFL works have focused on reducing the communication cost of homogeneous clients without considering communication heterogeneity due to variations in channel conditions. In this paper, we focus on a more realistic scenario, where clients have to communicate under different channel conditions and also handle non-iid data. We address the heterogeneity in the channel conditions of different clients by applying different degrees of compression on the data to be communicated and also adding an auxiliary network at the client-end to facilitate local training during times when the clients cannot communicate to the server. The experimental results on VGG11 and ResNet18 show that our method significantly reduces the training time and energy consumption with very little accuracy drop, compared with other competing methods. Deliang Fan, Chaitali Chakrabarti |
ISLPED | 5 |
| 2025 | Phantom: Privacy-Preserving Deep Neural Network Model Obfuscation in Heterogeneous TEE and GPU System
Juyang Bai, Md Hafizul Islam Chowdhuryy, Fan Yao 0001, Chaitali Chakrabarti, Deliang Fan |
USENIX Security Symposium | 5 |
| 2025 | HeteroSFL: Split Federated Learning With Heterogeneous Clients and Non-IID DataabstractSplit Federated Learning (SFL) is an emerging privacy-preserving decentralized learning scheme which splits a machine learning model between client and server such that most of the computations are offloaded to the server. While SFL has low computation cost on the client side, it has high communication cost. Existing SFL schemes focus on reducing the communication cost for homogeneous clients. However, a more realistic scenario is when clients are heterogeneous and process data with different distributions. In this paper, we focus on client-level heterogeneity caused by different communication data rates. We propose HeteroSFL, the first SFL framework with heterogeneous clients that handles non-IID data with label distribution skew across groups of clients. HeteroSFL compresses data with different compression factors in low-end and highend groups using narrow and wide bottleneck layers (BL), respectively. It provides a mechanism to address the challenge of aggregating different-sized BL models and utilizes bi-directional knowledge sharing (BDKS) to address the overfitting issue caused by different label distributions across high-and low-end groups of clients. Our experimental results show that HeteroSFL achieves significant training time reduction with minimum accuracy loss compared to competing methods. Specifically, it can reduce the training time of SFL by 16× to 256× with 1.24% to 5.59% accuracy loss for VGG11 on CIFAR10 for non-IID data. Xing Chen 0009, Deliang Fan, Chaitali Chakrabarti |
IEEE Internet Things J. | 4 |
| 2025 | HISIM: Analytical Performance Modeling and Design Space Exploration of 2.5D/3D Integration for AI ComputingabstractMonolithic designs face significant fabrication cost and data movement challenges, especially when executing complex and diverse AI models. Advanced 2.5D/3D packaging promises high bandwidth and connection density to overcome these challenges, yet it also introduces new electro-thermal constraints. This article develops a suite of analytical performance models to enable efficient benchmarking of a 2.5D/3D heterogeneous system for energy-efficient AI computing. These models encompass various performance metrics related to computing units, network-on-chip (NoC), and network-on-package (NoP). The results are summarized into a new tool, HISIM, which is$10^{4} \times $–$10^{6} \times $faster than state-of-the-art AI benchmark tools. Furthermore, HISIM integrates rapid thermal simulation for the 2.5D/3D system, helping shed light on both the potential and limitations of 2.5D/3D heterogeneous integration (HI) on representative AI algorithms. The code of HISIM is available athttps://github.com/mec-UMN/HISIM. Zhenyu Wang 0016, Pragnya Sudershan Nalla, Jingbo Sun 0003, A. Alper Goksoy, Sumit K. Mandal, Jae-sun Seo, Vidya A. Chhabria, Jeff Zhang 0001, Chaitali Chakrabarti, Ümit Y. Ogras, Yu Cao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 9 |
| 2025 | Coarse-Grained Task Parallelization by Dynamic Profiling for Heterogeneous SoC-Based Embedded SystemabstractIn this study, we introduce a methodology for automatically transforming user applications written in C/C++ to a parallel representation consisting of coarse-grained tasks based on dynamic profiling. Such a parallel representation is suitable for mapping applications onto heterogeneous SoCs. We present our approach for instrumenting the user application binary during the compilation process with parallel primitives that enable the runtime system to schedule and execute independent computation-intensive coarse-grained tasks concurrently. We use the proposed compilation and code transformation methodology to retarget each application for execution on a heterogeneous SoC composed of processor cores and accelerators. We demonstrate the capabilities of our integrated compile time and runtime flow through task-level parallelization and functionally correct execution of real-world applications in the communication systems and radar processing domains. We demonstrate the functionality of our integrated system by executing six distinct applications with different degrees of parallelism on four different platforms: an eight-core general-purpose processor, a heterogeneous SoC simulator, and two heterogeneous SoCs utilizing the Xilinx Zynq UltraScale+ FPGA and the Nvidia Jetson AGX board. Our integrated approach offers a path forward for application developers to take full advantage of the target SoC without requiring users to become hardware or parallel programming experts. Liangliang Chang, Serhan Gener, Joshua Mack, Hasan Umut Suluhan, Ali Akoglu, Chaitali Chakrabarti |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2025 | RIMMS: Runtime Integrated Memory Management System for Heterogeneous ComputingabstractEfficient memory management in heterogeneous systems is increasingly challenging due to diverse compute architectures (e.g., CPU, GPU, and FPGA) and dynamic task mappings not known at compile time. Existing approaches often require programmers to manage data placement and transfers explicitly, or assume static mappings that limit portability and scalability. This article introduces RIMMS (Runtime Integrated Memory Management System), a lightweight, runtime-managed, hardware-agnostic memory abstraction layer that decouples application development from low-level memory operations. RIMMS transparently tracks data locations, manages consistency, and supports efficient memory allocation across heterogeneous compute elements without requiring platform-specific tuning or code modifications. We integrate RIMMS into a baseline runtime and evaluate with complete radar signal processing applications across CPU+GPU and CPU+FPGA platforms. RIMMS delivers up to 2.43× speedup on GPU-based and 1.82× on FPGA-based systems over the baseline. Compared to IRIS, a recent heterogeneous runtime system, RIMMS achieves up to 3.08X speedup and matches the performance of native CUDA implementations while significantly reducing programming complexity. Despite operating at a higher abstraction level, RIMMS incurs only 1–2 cycles of overhead per memory management call, making it a low-cost solution. These results demonstrate RIMMS’s ability to deliver high performance and enhanced programmer productivity in dynamic, real-world heterogeneous environments. Serhan Gener, Aditya Ukarande, Shilpa Mysore Srinivasa Murthy, Md Sahil Hassan, Joshua Mack, Chaitali Chakrabarti, Ümit Y. Ogras, Ali Akoglu |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2024 | EMGAN: Early-Mix-GAN on Extracting Server-Side Model in Split Federated LearningabstractSplit Federated Learning (SFL) is an emerging edge-friendly version of Federated Learning (FL), where clients process a small portion of the entire model. While SFL was considered to be resistant to Model Extraction Attack (MEA) by design, a recent work shows it is not necessarily the case. In general, gradient-based MEAs are not effective on a target model that is changing, as is the case in training-from-scratch applications. In this work, we propose a strong MEA during the SFL training phase. The proposed Early-Mix-GAN (EMGAN) attack effectively exploits gradient queries regardless of data assumptions. EMGAN adopts three key components to address the problem of inconsistent gradients. Specifically, it employs (i) Early-learner approach for better adaptability, (ii) Multi-GAN approach to introduce randomness in generator training to mitigate mode collapse, and (iii) ProperMix to effectively augment the limited amount of synthetic data for a better approximation of the target domain data distribution. EMGAN achieves excellent results in extracting server-side models. With only 50 training samples, EMGAN successfully extracts a 5-layer server-side model of VGG-11 on CIFAR-10, with 7% less accuracy than the target model. With zero training data, the extracted model achieves 81.3% accuracy, which is significantly better than the 45.5% accuracy of the model extracted by the SoTA method. The code is available at "https://github.com/zlijingtao/SFL-MEA". Xing Chen 0009, Li Yang 0009, Adnan Siraj Rakin, Deliang Fan, Chaitali Chakrabarti |
AAAI | 6 |
| 2024 | Exploiting 2.5D/3D Heterogeneous Integration for AI ComputingabstractThe evolution of AI algorithms has not only revolutionized many application domains, but also posed tremendous challenges on the hardware platform. Advanced packaging technology today, such as 2.5D and 3D interconnection, provides a promising solution to meet the ever-increasing demands of bandwidth, data movement, and system scale in AI computing. This work presents HISIM, a modeling and benchmarking tool for chiplet-based heterogeneous integration. HISIM emphasizes the hierarchical interconnection that connects various chiplets through network-on-package. It further integrates technology roadmap, power/latency prediction, and thermal analysis together to support electro-thermal co-design. Leveraging HISIM with in-memory computing chiplets, we explore the advantages and limitations of 2.5D and 3D heterogenous integration on representative AI algorithms, such as DNNs, transformers, and graph neural networks. Zhenyu Wang 0016, Jingbo Sun 0003, A. Alper Goksoy, Sumit K. Mandal, Yaotian Liu, Jae-sun Seo, Chaitali Chakrabarti, Ümit Y. Ogras, Vidya A. Chhabria, Jeff Zhang 0001, Yu Cao 0001 |
ASPDAC | 7 |
| 2024 | Improving Speech-Based Dysarthria Detection using Multi-task Learning with Gradient Projection
Yan Xiong 0002, Visar Berisha, Julie M. Liss, Chaitali Chakrabarti |
INTERSPEECH | 4 |
| 2024 | PED: Probabilistic Energy-efficient Deadline-aware scheduler for heterogeneous SoCs
Xing Chen 0009, Anish Krishnakumar, Ümit Y. Ogras, Chaitali Chakrabarti |
J. Syst. Archit. | 4 |
| 2024 | Cyclebite: Extracting Task Graphs From Unstructured Compute-ProgramsabstractExtracting portable performance in an application requires structuring that program into a data-flow graph of coarse-grained tasks (CGTs). Structuring applications that interconnect multiple external libraries and custom code (i.e., “Code From The Wild” (CFTW)) is challenging. When experts manually restructure a program, they trivialize the extraction of structure; however, this expertise is not broadly available. Automatic structuring approaches focus on the intersection of hot code and static loops, ignoring the data dependencies between tasks and significantly reducing the scope of analyzeable programs. This work addresses the problem of extracting the data-flow graph of CGTs from CFTW. To that end, we present Cyclebite. Our approach extracts CGTs from unstructured compute-programs by detecting CGT candidates in the simplified Markov Control Graph (MCG), and localizing CGTs in an epoch profile. Additionally, the epoch profile extracts the data dependence between CGTs required to build the data-flow graph of CGTs. Cyclebite demonstrates a robust selectivity for critical CGTs relative to the state-of-the-art (SoA), leading to a potential speedup of 12x on average and thread-scaling of 24x on average compared to modern compiler optimizers. We validate the results of Cyclebite and compare them to two SoA techniques using an input corpus of 25 open-source C/C++ libraries with 2,019 unique execution profiles. Benjamin R. Willis, Aviral Shrivastava, Joshua Mack, Shail Dave, Chaitali Chakrabarti, John S. Brunhaver |
IEEE Trans. Computers | 5 |
| 2023 | MocoSFL: enabling cross-client collaborative self-supervised learning
Lingjuan Lyu, Daisuke Iso, Chaitali Chakrabarti, Michael Spranger |
ICLR | 4 |
| 2023 | Aligning Speech Enhancement for Improving Downstream Classification Performance
Yan Xiong 0002, Visar Berisha, Chaitali Chakrabarti |
INTERSPEECH | 3 |
| 2022 | ResSFL: A Resistance Transfer Framework for Defending Model Inversion Attack in Split Federated LearningabstractThis work aims to tackle Model Inversion (MI) attack on Split Federated Learning (SFL). SFL is a recent distributed training scheme where multiple clients send intermediate activations (i. e., feature map), instead of raw data, to a central server. While such a scheme helps reduce the computational load at the client end, it opens itself to reconstruction of raw data from intermediate activation by the server. Existing works on protecting SFL only consider inference and do not handle attacks during training. So we propose ResSFL, a Split Federated Learning Framework that is designed to be MI-resistant during training. It is based on deriving a resistant feature extractor via attacker-aware training, and using this extractor to initialize the client-side model prior to standard SFL training. Such a method helps in reducing the computational complexity due to use of strong inversion model in client-side adversarial training as well as vulnerability of attacks launched in early training epochs. On CIFAR-100 dataset, our proposed framework successfully mitigates MI attack on a VGG-11 model with a high reconstruction Mean-Square-Error of 0.050 compared to 0.005 obtained by the baseline system. The frame-work achieves 67.5% accuracy (only 1 % accuracy drop) with very low computation overhead. Code is released at: https://github.com/zlijingtao/ResSFL. Adnan Siraj Rakin, Xing Chen 0009, Zhezhi He, Deliang Fan, Chaitali Chakrabarti |
CVPR | 6 |
| 2022 | Deep Learning for Moving Blockage Prediction using Real mmWave MeasurementsabstractMillimeter wave (mmWave) communication is a key component of 5G systems and beyond. Such systems provide high bandwidth and high data rate but are sensitive to blockages. A sudden blockage in the line of sight (LOS) link leads to abrupt disconnection. Thus addressing blockage problems is essential for enhancing the reliability and latency of mmWave communication networks. In this paper, we propose a novel solution that relies only on in-band mmWave wireless measurements to proactively predict future dynamic line-of-sight (LOS) link blockages. The proposed solution utilizes deep neural networks and special patterns of received signal power, which we call pre-blockage wireless signatures, to infer future blockages. Specifically, the machine learning models attempt to predict: (i) Whether a blockage will occur in the next few seconds? (ii) At what time instance will this blockage occur? To evaluate our proposed approach, we build a mmWave communication setup with moving blockage in an indoor scenario and collect received power sequences. Simulation results on a real dataset show that blockage occurrence can be predicted with more than 85% accuracy, and the exact time instance of blockage occurrence can be obtained with less than 2 time instances (1.66s) error for prediction interval of 10 time instances (8.8s). This demonstrates the potential of the proposed solution for dynamic blockage prediction and proactive hand-off. Shunyao Wu, Muhammad Alrabeiah, Andrew Hredzak, Chaitali Chakrabarti, Ahmed Alkhateeb |
ICC | 4 |
| 2022 | Big-Little Chiplets for In-Memory Acceleration of DNNs: A Scalable Heterogeneous ArchitectureabstractMonolithic in-memory computing (IMC) architectures face significant yield and fabrication cost challenges as the complexity of DNNs increases. Chiplet-based IMCs that integrate multiple dies with advanced 2.5D/3D packaging offers a low-cost and scalable solution. They enable heterogeneous architectures where the chiplets and their associated interconnection can be tailored to the non-uniform algorithmic structures to maximize IMC utilization and reduce energy consumption. This paper proposes a heterogeneous IMC architecture with big-little chiplets and a hybrid network-on-package (NoP) to optimize the utilization, interconnect bandwidth, and energy efficiency. For a given DNN, we develop a custom methodology to map the model onto the big-little architecture such that the early layers in the DNN are mapped to the little chiplets with higher NoP bandwidth and the subsequent layers are mapped to the big chiplets with lower NoP bandwidth. Furthermore, we achieve a scalable solution by incorporating a DRAM into each chiplet to support a wide range of DNNs beyond the area limit. Compared to a homogeneous chiplet-based IMC architecture, the proposed big-little architecture achieves up to 329× improvement in the energy-delay-area product (EDAP) and up to 2× higher IMC utilization. Experimental evaluation of the proposed big-little chiplet-based RRAM IMC architecture for ResNet-50 on ImageNet shows 259×, 139×, and 48× improvement in energy-efficiency at lower area compared to Nvidia V100 GPU, Nvidia T4 GPU, and SIMBA architecture, respectively. A. Alper Goksoy, Sumit K. Mandal, Zhenyu Wang 0016, Chaitali Chakrabarti, Jae-sun Seo, Ümit Y. Ogras, Yu Cao 0001 |
ICCAD | 5 |
| 2022 | Enabling Software-Defined RF Convergence with a Novel Coarse-Scale Heterogeneous ProcessorabstractRF system development is traditionally constrained by a restrictive trade-off between power efficiency and programmatic flexibility. We outline a path towards achieving both, thereby enabling a range of new system concepts that better utilize limited resources. As an example, for many future applications, we consider RF convergence – reusing the same spectrum and waveforms to achieve multiple distributed system functions and goals, simultaneously. To enable this next step in processing, we develop a novel framework that includes both software and the system-on-chip (SoC) design. Daniel W. Bliss, Tutu Ajayi, Ali Akoglu, Ilkin Aliyev, Toygun Basaklar, Leul Belayneh, David T. Blaauw, John S. Brunhaver, Chaitali Chakrabarti, Liangliang Chang, Kuan-Yu Chen 0001, Ming-Hung Chen, Xing Chen 0004, Alex R. Chiriyath, Alhad Daftardar, Ronald G. Dreslinski, Arindam Dutta, Allen-Jasmin Farcas, Yukang Fu, A. Alper Goksoy, Xin He 0011, Md Sahil Hassan, Andrew Herschfelt, Jacob Holtom, Hun-Seok Kim, Anish Krishnakumar, Owen Ma, Joshua Mack, Saurav Mallik, Sumit K. Mandal, Radu Marculescu, Brittany M. McCall, Trevor N. Mudge, Ümit Y. Ogras, Vishrut Pandey, Saquib Ahmad Siddiqui, Yu-Hsiu Sun, Adarsh A. Venkataramani, Xiangdong Wei, Benjamin R. Willis, Hanguang Yu, Yufan Yue |
ISCAS | 9 |
| 2022 | Improving Energy Efficiency of Convolutional Neural Networks on Multi-core Architectures through Run-time ReconfigurationabstractConvolutional neural networks (CNNs) are built with convolution layers which account for most of their computation time. The differences in the convolution kernel types (2D, point-wise, depth-wise), and input sizes lead to significant differences in their computation and memory demands. In this work, we exploit run-time reconfiguration to adapt to the differences in the characteristics of different convolution kernels on a low-power reconfigurable architecture, Transmuter. The architecture consists of light-weight cores interconnected by caches and crossbars that support run-time reconfiguration between different cache modes - shared or private, different dataflow modes - systolic or parallel, and different computation mapping schemes. To achieve run-time reconfiguration, we propose a decision-tree-based engine that selects the optimal Transmuter configuration at a low cost. The proposed method is evaluated on commonly-used CNN models such as ResNetl8, VGGII, AlexNet and MobileNetV3. Simulation results show that run-time reconfiguration helps improve the energy efficiency of Transmuter in the range of 3.1$\times-13.7\times$ across all networks. Yan Xiong 0002, David T. Blaauw, Hun-Seok Kim, Trevor N. Mudge, Ronald G. Dreslinski, Chaitali Chakrabarti |
ISCAS | 7 |
| 2022 | LiDAR-Aided Mobile Blockage Prediction in Real-World Millimeter Wave SystemsabstractLine-of-sight link blockages represent a key challenge for the reliability and latency of millimeter wave (mmWave) and terahertz (THz) communication networks. This paper proposes to leverage LiDAR sensory data to provide awareness about the communication environment and proactively predict dynamic link blockages before they happen. This allows the network to make proactive decisions for hand-off/beam switching which enhances its reliability and latency. We formulate the LiDAR-aided blockage prediction problem and present the first real-world demonstration for LiDAR-aided blockage prediction in mmWave systems. In particular, we construct a large-scale real-world dataset, based on the DeepSense 6G structure, that comprises co-existing LiDAR and mmWave communication measurements in outdoor vehicular scenarios. Then, we develop an efficient LiDAR data denoising (static cluster removal) algorithm and a machine learning model that proactively predicts dynamic link blockages. Based on the real-world dataset, our LiDAR-aided approach is shown to achieve 95% accuracy in predicting blockages happening within 100ms and more than 80% prediction accuracy for blockages happening within one second. If used for proactive hand-off, the proposed solutions can potentially provide an order of magnitude saving in the network latency, which highlights a promising direction for addressing the blockage challenges in mmWave/sub-THz networks. Shunyao Wu, Chaitali Chakrabarti, Ahmed Alkhateeb |
WCNC | 2 |
| 2022 | Impact of On-chip Interconnect on In-memory Acceleration of Deep Neural NetworksabstractWith the widespread use of Deep Neural Networks (DNNs), machine learning algorithms have evolved in two diverse directions—one with ever-increasing connection density for better accuracy and the other with more compact sizing for energy efficiency. The increase in connection density increases on-chip data movement, which makes efficient on-chip communication a critical function of the DNN accelerator. The contribution of this work is threefold. First, we illustrate that the point-to-point (P2P)-based interconnect is incapable of handling a high volume of on-chip data movement for DNNs. Second, we evaluate P2P and network-on-chip (NoC) interconnect (with a regular topology such as a mesh) for SRAM- and ReRAM-based in-memory computing (IMC) architectures for a range of DNNs. This analysis shows the necessity for the optimal interconnect choice for an IMC DNN accelerator. Finally, we perform an experimental evaluation for different DNNs to empirically obtain the performance of the IMC architecture with both NoC-tree and NoC-mesh. We conclude that, at the tile level, NoC-tree is appropriate for compact DNNs employed at the edge, and NoC-mesh is necessary to accelerate DNNs with high connection density. Furthermore, we propose a technique to determine the optimal choice of interconnect for any given DNN. In this technique, we use analytical models of NoC to evaluate end-to-end communication latency of any given DNN. We demonstrate that the interconnect optimization in the IMC architecture results in up to 6 × improvement in energy-delay-area product for VGG-19 inference compared to the state-of-the-art ReRAM-based IMC architectures. Sumit K. Mandal, Chaitali Chakrabarti, Jae-sun Seo, Ümit Y. Ogras, Yu Cao 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2022 | T-BFA: Targeted Bit-Flip Adversarial Weight AttackabstractTraditional Deep Neural Network (DNN) security is mostly related to the well-known adversarial input example attack.Recently, another dimension of adversarial attack, namely, attack on DNN weight parameters, has been shown to be very powerful. Asa representative one, the Bit-Flip based adversarial weight Attack (BFA) injects an extremely small amount of faults into weight parameters to hijack the executing DNN function. Prior works of BFA focus on un-targeted attacks that can hack all inputs into a random output class by flipping a very small number of weight bits stored in computer memory. This paper proposes the first work oftargetedBFA based (T-BFA) adversarial weight attack on DNNs, which can intentionally mislead selected inputs to a target output class. The objective is achieved by identifying the weight bits that are highly associated with classification of a targeted output through a class-dependent weight bit searching algorithm. Our proposed T-BFA performance is successfully demonstrated on multiple DNN architectures for image classification tasks. For example, by merely flipping 27 out of 88 million weight bits of ResNet-18, our T-BFA can misclassify all the images from Hen class into Goose class (i.e., 100% attack success rate) in ImageNet dataset, while maintaining 59.35% validation accuracy. Adnan Siraj Rakin, Zhezhi He, Fan Yao 0001, Chaitali Chakrabarti, Deliang Fan |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Probabilistic Risk-Aware Scheduling with Deadline Constraint for Heterogeneous SoCsabstractHardware Trojans can compromise System-on-Chip (SoC) performance. Protection schemes implemented to combat these threats cannot guarantee 100% detection rate and may also introduce performance overhead. This paper defines the risk of running a job on an SoC as a function of the misdetection rate of the hardware Trojan detection methods implemented on the cores in the SoC. Given the user-defined deadlines of each job, our goal is to minimize the job-level risk as well as the deadline violation rate for both static and dynamic scheduling scenarios. We assume that there is no relationship between the execution time and risk of a task executed on a core. Our risk-aware scheduling algorithm first calculates the probability of possible task allocations and then uses it to derive the task-level deadlines. Each task is then allocated to the core with minimum risk that satisfies the task-level deadline. In addition, in dynamic scheduling, where multiple jobs are injected randomly, we propose to explicitly operate with a reduced virtual deadline to avoid possible future deadline violations. Simulations on randomly generated graphs show that our static scheduler has no deadline violations and achieves 5.1%–17.2% lower job-level risk than the popular Earliest Time First (ETF) algorithm when the deadline constraint is 1.2×–3.0× the makespan of ETF. In the dynamic case, the proposed algorithm achieves a violation rate comparable to that of Earliest Deadline First (EDF) , an algorithm optimized for dynamic scenarios. Even when the injection rate is high, it outperforms EDF with 8.4%–10% lower risk when the deadline is 1.5×–3.0× the makespan of ETF. Xing Chen 0004, Ümit Y. Ogras, Chaitali Chakrabarti |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2021 | CoSPARSE: A Software and Hardware Reconfigurable SpMV Framework for Graph AnalyticsabstractSparse matrix-vector multiplication (SpMV) is a critical building block for iterative graph analytics algorithms. Typically, such algorithms have a varying active vertex set across iterations. This variability has been used to improve performance by either dynamically switching algorithms between iterations (software) or designing custom accelerators (hardware) for graph analytics algorithms. In this work, we propose a novel framework, CoSPARSE, that employs hardware and software reconfiguration as a synergistic solution to accelerate SpMV-based graph analytics algorithms. Building on previously proposed general-purpose reconfigurable hardware, we implement CoSPARSE as a software layer, abstracting the hardware as a specialized SpMV accelerator. CoSPARSE dynamically selects software and hardware configurations for each iteration and achieves a maximum speedup of 2.0 × compared to the naïve implementation with no reconfiguration. Across a suite of graph algorithms, CoSPARSE outperforms a state-of-the-art shared memory framework, Ligra, on a Xeon CPU with up to 3.51 × better performance and 877 × better energy efficiency. Siying Feng, Jiawen Sun, Subhankar Pal, Xin He 0011, Kuba Kaszyk, Dong-Hyeon Park, John Magnus Morton, Trevor N. Mudge, Murray Cole, Michael F. P. O'Boyle, Chaitali Chakrabarti, Ronald G. Dreslinski |
DAC | 11 |
| 2021 | RADAR: Run-time Adversarial Weight Attack Detection and Accuracy RecoveryabstractAdversarial attacks on Neural Network weights, such as the progressive bit-flip attack (PBFA), can cause a catastrophic degradation in accuracy by flipping a very small number of bits. Furthermore, PBFA can be conducted at run time on the weights stored in DRAM main memory. In this work, we propose RADAR, a Run-time adversarial weight Attack Detection and Accuracy Recovery scheme to protect DNN weights against PBFA. We organize weights that are interspersed in a layer into groups and employ a checksum-based algorithm on weights to derive a 2-bit signature for each group. At run time, the 2-bit signature is computed and compared with the securely stored golden signature to detect the bit-flip attacks in a group. After successful detection, we zero out all the weights in a group to mitigate the accuracy drop caused by malicious bit-flips. The proposed scheme is embedded in the inference computation stage. For the ResNet-18 ImageNet model, our method can detect 9.6 bit-flips out of 10 on average. For this model, the proposed accuracy recovery scheme can restore the accuracy from below 1% caused by 10 bit flips to above 69%. The proposed method has extremely low time and storage overhead. System-level simulation on gem5 shows that RADAR only adds < 1% to the inference time, making this scheme highly suitable for run-time attack detection and mitigation. Adnan Siraj Rakin, Zhezhi He, Deliang Fan, Chaitali Chakrabarti |
DATE | 5 |
| 2021 | SIAM: Chiplet-based Scalable In-Memory Acceleration with Mesh for Deep Neural NetworksabstractIn-memory computing (IMC) on a monolithic chip for deep learning faces dramatic challenges on area, yield, and on-chip interconnection cost due to the ever-increasing model sizes. 2.5D integration or chiplet-based architectures interconnect multiple small chips (i.e., chiplets) to form a large computing system, presenting a feasible solution beyond a monolithic IMC architecture to accelerate large deep learning models. This paper presents a new benchmarking simulator, SIAM, to evaluate the performance of chiplet-based IMC architectures and explore the potential of such a paradigm shift in IMC architecture design. SIAM integrates device, circuit, architecture, network-on-chip (NoC), network-on-package (NoP), and DRAM access models to realize an end-to-end system. SIAM is scalable in its support of a wide range of deep neural networks (DNNs), customizable to various network structures and configurations, and capable of efficient design space exploration. We demonstrate the flexibility, scalability, and simulation speed of SIAM by benchmarking different state-of-the-art DNNs with CIFAR-10, CIFAR-100, and ImageNet datasets. We further calibrate the simulation results with a published silicon result, SIMBA. The chiplet-based IMC architecture obtained through SIAM shows 130 and 72 improvement in energy-efficiency for ResNet-50 on the ImageNet dataset compared to Nvidia V100 and T4 GPUs. Sumit K. Mandal, Manvitha Pannala, Chaitali Chakrabarti, Jae-sun Seo, Ümit Y. Ogras, Yu Cao 0001 |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2021 | Front-End Architecture Design for Low-Complexity 3-D Ultrasound Imaging Based on Synthetic Aperture Sequential BeamformingabstractThe 3-D ultrasound imaging provides distinct advantages over its 2-D counterpart leading to a more accurate analysis of tumors and cysts. However, the front end of a 3-D system must receive and process data at prodigious rates, making it impractical for power-constrained portable systems. Synthetic aperture sequential beamforming (SASB) is an ultrasound beamforming technique that splits the computation into two stages, such that the computation in Stage 1 can be completed in the power-constrained front end while the remaining computation can be done elsewhere. In this article, we present several algorithmic and architectural techniques to enable efficient computation of Stage 1 processing without compromising imaging quality. Specifically, we present algorithmic techniques that reduce the computational complexity in Stage 1 by 17× through a systematic reduction in the number of apodization coefficients. We propose a 3-D die stacked architecture where the signals received by 961 active transducers are digitized, routed by a network-onchip, and processed in parallel. This architecture does not require the explicit storage of incoming data samples. We synthesize the architecture using TSMC 28-nm technology node. The front-end power consumption is around 1.5 W, making it suitable for portable applications. Jian Zhou 0012, Sumit K. Mandal, Brendan L. West, Siyuan Wei, Ümit Y. Ogras, Oliver Kripfgans, J. Brian Fowlkes, Thomas F. Wenisch, Chaitali Chakrabarti |
IEEE Trans. Very Large Scale Integr. Syst. | 9 |
| 2020 | Transmuter: Bridging the Efficiency Gap using Memory and Dataflow ReconfigurationabstractWith the end of Dennard scaling and Moore's law, it is becoming increasingly difficult to build hardware for emerging applications that meet power and performance targets, while remaining flexible and programmable for end users. This is particularly true for domains that have frequently changing algorithms and applications involving mixed sparse/dense data structures, such as those in machine learning and graph analytics. To overcome this, we present a flexible accelerator called Transmuter, in a novel effort to bridge the gap between General-Purpose Processors (GPPs) and Application-Specific Integrated Circuits (ASICs). Transmuter adapts to changing kernel characteristics, such as data reuse and control divergence, through the ability to reconfigure the on-chip memory type, resource sharing and dataflow at run-time within a short latency. This is facilitated by a fabric of light-weight cores connected to a network of reconfigurable caches and crossbars. Transmuter addresses a rapidly growing set of algorithms exhibiting dynamic data movement patterns, irregularity, and sparsity, while delivering GPU-like efficiencies for traditional dense applications. Finally, in order to support programmability and ease-of-adoption, we prototype a software stack composed of low-level runtime routines, and a high-level language library called TransPy, that cater to expert programmers and end-users, respectively. Subhankar Pal, Siying Feng, Dong-Hyeon Park, Aporva Amarnath, Chi-Sheng Yang, Xin He 0011, Jonathan Beaumont, Kyle May, Yan Xiong 0002, Kuba Kaszyk, John Magnus Morton, Jiawen Sun, Michael F. P. O'Boyle, Murray Cole, Chaitali Chakrabarti, David T. Blaauw, Hun-Seok Kim, Trevor N. Mudge, Ronald G. Dreslinski |
PACT | 16 |
| 2020 | Defending and Harnessing the Bit-Flip Based Adversarial Weight AttackabstractRecently, a new paradigm of the adversarial attack on the quantized neural network weights has attracted great attention, namely, the Bit-Flip based adversarial weight attack, aka. Bit-Flip Attack (BFA). BFA has shown extraordinary attacking ability, where the adversary can malfunction a quantized Deep Neural Network (DNN) as a random guess, through malicious bit-flips on a small set of vulnerable weight bits (e.g., 13 out of 93 millions bits of 8-bit quantized ResNet-18). However, there are no effective defensive methods to enhance the fault-tolerance capability of DNN against such BFA. In this work, we conduct comprehensive investigations on BFA and propose to leverage binarization-aware training and its relaxation - piece-wise clustering as simple and effective countermeasures to BFA. The experiments show that, for BFA to achieve the identical prediction accuracy degradation (e.g., below 11% on CIFAR-10), it requires 19.3× and 480.1× more effective malicious bit-flips on ResNet-20 and VGG-11 respectively, compared to defend-free counterparts. Zhezhi He, Adnan Siraj Rakin, Chaitali Chakrabarti, Deliang Fan |
CVPR | 4 |
| 2020 | Defending Bit-Flip Attack through DNN Weight ReconstructionabstractRecent studies show that adversarial attacks on neural network weights, aka, Bit-Flip Attack (BFA), can degrade Deep Neural Network’s (DNN) prediction accuracy severely. In this work, we propose a novel weight reconstruction method as a countermeasure to such BFAs. Specifically, during inference, the weights are reconstructed such that the weight perturbation due to BFA is minimized or diffused to the neighboring weights. We have successfully demonstrated that our method can significantly improve the DNN robustness against random and gradient-based BFA variants. Even under the most aggressive attacks (i.e., greedy progressive bit search), our method maintains a test accuracy of 60% on ImageNet after 5 iterations while the baseline accuracy drops to below 1%. Adnan Siraj Rakin, Yan Xiong 0002, Liangliang Chang, Zhezhi He, Deliang Fan, Chaitali Chakrabarti |
DAC | 7 |
| 2020 | Accelerating Linear Algebra Kernels on a Massively Parallel Reconfigurable ArchitectureabstractMuch of the recent work on domain-specific architectures has focused on bridging the gap between performance/efficiency and programmability. We consider one such example architecture, Transformer, consisting of light-weight cores interconnected by caches and crossbars that supports run-time reconfiguration between shared and private cache mode operations. We present customized implementation of a select set of linear algebra kernels, namely, triangular matrix solver, LU decomposition, QR decomposition and matrix in-version, on Transformer. The performance of the kernel algorithms is evaluated with respect to execution time and energy efficiency. Our study shows that each kernel achieves high performance for a certain cache mode and that this cache mode can change when the matrix size changes, making a case for run-time reconfiguration. A. Soorishetty, Jian Zhou 0012, Subhankar Pal, David T. Blaauw, Trevor N. Mudge, Ronald G. Dreslinski, Chaitali Chakrabarti |
ICASSP | 8 |
| 2020 | Compressing LSTM Networks with Hierarchical Coarse-Grain Sparsity
Deepak Kadetotad, Jian Meng, Visar Berisha, Chaitali Chakrabarti, Jae-sun Seo |
INTERSPEECH | 4 |
| 2020 | Accelerating Deep Neural Network Computation on a Low Power Reconfigurable ArchitectureabstractRecent work on neural network architectures has focused on bridging the gap between performance/efficiency and programmability. We consider implementations of three popular neural networks, ResNet, AlexNet and ASGD weight-dropped Recurrent Neural Network (AWD RNN) on a low power programmable architecture, Transformer. The architecture consists of light-weight cores interconnected by caches and crossbars that support run-time reconfiguration between shared and private cache mode operations. We present efficient implementations of key neural network kernels and evaluate the performance of each kernel when operating in different cache modes. The best-performing cache modes are then used in the implementation of the end-to-end network. Simulation results show superior performance with ResNet, AlexNet and AWD RNN achieving 188.19 GOPS/W, 150.53 GOPS/W and 120.68 GOPS/W, respectively, in the 14 nm technology node. Yan Xiong 0002, Jian Zhou 0012, Subhankar Pal, David T. Blaauw, Hun-Seok Kim, Trevor N. Mudge, Ronald G. Dreslinski, Chaitali Chakrabarti |
ISCAS | 8 |
| 2020 | Tetris: Using Software/Hardware Co-Design to Enable Handheld, Physics-Limited 3D Plane-Wave Ultrasound ImagingabstractHigh volume acquisition rates are imperative for certain medical ultrasound imaging applications, such as 3D elastography and 3D vector flow imaging. As ultrasound imaging transitions from 2D to 3D, the massive data bandwidth and billions of trigonometric operations required to reconstruct each volume leaves conventional computer architectures falling short. Despite recent algorithmic improvements, high-volume-rate ultrasound imaging remains computationally infeasible on known platforms. In this article, we expand our previous work on Tetris, a novel hardware accelerator for separable ultrasound beamforming that enables volume acquisition rates up to the physics limits of acoustic propagation delay. Through algorithmic and hardware optimizations, we enable an image reconstruction system design outclassing previously proposed accelerators in performance while lowering hardware complexity, storage, and power requirements. Tetris operates in a streaming fashion-without requiring on-chip storage of the entire receive signal-reconstructing volumes in real-time. For a representative imaging task, our proposed system generates physics-limited 13,000 volumes per second in a 2 watt power budget. The Tetris beamformer has an unprecedented power efficiency of 2.03 tera-beamforming operations per watt-an increase in efficiency of nearly 3× compared to the prior work. Brendan L. West, Jian Zhou 0012, Ronald G. Dreslinski, Oliver Kripfgans, J. Brian Fowlkes, Chaitali Chakrabarti, Thomas F. Wenisch |
IEEE Trans. Computers | 6 |
| 2019 | Tetris: A Streaming Accelerator for Physics-Limited 3D Plane-Wave Ultrasound ImagingabstractHigh volume acquisition rates are imperative for medical ultrasound imaging applications, such as 3D elastography and 3D vector flow imaging. Unfortunately, despite recent algorithmic improvements, high-volume-rate imaging remains computationally infeasible on known platforms. Brendan L. West, Jian Zhou 0012, Ronald G. Dreslinski, J. Brian Fowlkes, Oliver Kripfgans, Chaitali Chakrabarti, Thomas F. Wenisch |
DAC | 6 |
| 2019 | Joint Optimization of Quantization and Structured Sparsity for Compressed Deep Neural NetworksabstractThe usage of Deep Neural Networks (DNN) on resource-constrained edge devices has been limited due to their high computation and large memory requirement. In this work, we propose an algorithm to compress DNNs by jointly optimizing structured sparsity and quantization constraints in a single DNN training framework. The proposed algorithm has been extensively validated on high/low capacity DNNs and wide/deep sparse DNNs. Further, we perform Pareto-optimal analysis to extract optimal DNN models from a large set of trained DNN models. The optimal structurally-compressed DNN model achieves ~50X weight memory reduction without test accuracy degradation, compared to floating-point uncompressed DNN. Gaurav Srivastava 0001, Deepak Kadetotad, Shihui Yin, Visar Berisha, Chaitali Chakrabarti, Jae-sun Seo |
ICASSP | 5 |
| 2019 | Residual + Capsule Networks (ResCap) for Simultaneous Single-Channel Overlapped Keyword Recognition
Yan Xiong 0002, Visar Berisha, Chaitali Chakrabarti |
INTERSPEECH | 3 |
| 2019 | Configurable-ECC: Architecting a Flexible ECC Scheme to Support Different Sized Accesses in High Bandwidth Memory SystemsabstractDesigning error correction code (ECC) to guarantee strong reliability for high bandwidth memory (HBM) is imperative in high performance computers, especially for systems equipped with graphics processing units (GPUs). The design of ECC is challenging because future GPUs are expected to implement a memory subsystem supporting fine and coarse-grained data accesses to match the difference in the spatial locality of GPGPU applications. Current ECC designs, however, are developed for a fixed data fetch granularity. To have a more flexible design, we propose a novel memory protection scheme, called Config(urable)-ECC, which provides strong reliability for both fine and coarse-grained data accesses. Config-ECC consists of two tiers of ECC protection. The tier-1 code is a strong product code that can correct errors due to small granularity faults and detect errors caused by large granularity faults. The tier-2 code is an XOR-based code that is employed to correct errors incurred by large granularity faults. Config-ECC provides stronger reliability and/or lower energy consumption compared to state-of-the-art fixed 32B and 64B ECC schemes. It reduces the HBM energy by 17-21 percent while reducing the failure in time (FIT) rate by 20 times compared to a state-of-the-art fixed 64B ECC scheme with an insignificant 1.2 percent performance overhead. Hsing Min Chen, Shin-Ying Lee, Trevor N. Mudge, Carole-Jean Wu, Chaitali Chakrabarti |
IEEE Trans. Computers | 5 |
| 2019 | Low Complexity, Hardware-Efficient Neighbor-Guided SGM Optical Flow for Low-Power Mobile Vision ApplicationsabstractAccurate, low-latency, and energy-efficient optical flow estimation is a fundamental kernel function to enable several real-time vision applications on mobile platforms. This paper presents neighbor-guided semi-global matching (NG-fSGM), a new low-complexity optical flow algorithm tailored for low-power mobile applications. NG-fSGM obtains high accuracy optical flow by aggregating local matching costs over a semiglobal region, successfully resolving local ambiguity in textureless and occluded regions. The proposed NG-fSGM aggressively prunes the search space based on neighboring pixels' information to significantly lower the algorithm complexity from the original fSGM. As a result, NG-fSGM achieves 17.9× reduction in the number of computations and 8.37× reduction in memory space compared to the original fSGM without compromising its algorithm accuracy. A multicore architecture for NG-fSGM is implemented in hardware to quantify algorithm complexity and power consumption. The proposed architecture realizes NG-fSGM with overlapping blocks processed in parallel to enhance throughput and to lower power consumption. The eightcore architecture achieves 20 M pixel/s (66 frames/s for VGA) throughput with 9.6 mm2area at 679.2-mW power consumption in 28-nm node. Ziyun Li 0001, Jiang Xiang, Luyao Gong, David T. Blaauw, Chaitali Chakrabarti, Hun-Seok Kim |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2018 | OuterSPACE: An Outer Product Based Sparse Matrix Multiplication AcceleratorabstractSparse matrices are widely used in graph and data analytics, machine learning, engineering and scientific applications. This paper describes and analyzes OuterSPACE, an accelerator targeted at applications that involve large sparse matrices. OuterSPACE is a highly-scalable, energy-efficient, reconfigurable design, consisting of massively parallel Single Program, Multiple Data (SPMD)-style processing units, distributed memories, high-speed crossbars and High Bandwidth Memory (HBM). We identify redundant memory accesses to non-zeros as a key bottleneck in traditional sparse matrix-matrix multiplication algorithms. To ameliorate this, we implement an outer product based matrix multiplication technique that eliminates redundant accesses by decoupling multiplication from accumulation. We demonstrate that traditional architectures, due to limitations in their memory hierarchies and ability to harness parallelism in the algorithm, are unable to take advantage of this reduction without incurring significant overheads. OuterSPACE is designed to specifically overcome these challenges. We simulate the key components of our architecture using gem5 on a diverse set of matrices from the University of Florida's SuiteSparse collection and the Stanford Network Analysis Project and show a mean speedup of 7.9× over Intel Math Kernel Library on a Xeon CPU, 13.0× against cuSPARSE and 14.0× against CUSP when run on an NVIDIA K40 GPU, while achieving an average throughput of 2.9 GFLOPS within a 24 W power budget in an area of 87 mm2. Subhankar Pal, Jonathan Beaumont, Dong-Hyeon Park, Aporva Amarnath, Siying Feng, Chaitali Chakrabarti, Hun-Seok Kim, David T. Blaauw, Trevor N. Mudge, Ronald G. Dreslinski |
HPCA | 6 |
| 2018 | Design and Analysis of Energy-Efficient and Reliable 3-D ReRAM Cross-Point Array System
Manqing Mao, Shimeng Yu, Chaitali Chakrabarti |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2017 | Low-power neuromorphic speech recognition engine with coarse-grain sparsityabstractIn recent years, we have seen a surge of interest in neuromorphic computing and its hardware design for cognitive applications. In this work, we present new neuromorphic architecture, circuit, and device co-designs that enable spike-based classification for speech recognition task. The proposed neuromorphic speech recognition engine supports a sparsely connected deep spiking network with coarse granularity, leading to large memory reduction with minimal index information. Simulation results show that the proposed deep spiking neural network accelerator achieves phoneme error rate (PER) of 20.5% for TIMIT database, and consume 2.57mW in 40nm CMOS for real-time performance. To alleviate the memory bottleneck, the usage of non-volatile memory is also evaluated and discussed. Shunti Yin, Deepak Kadetotad, Bonan Yan, Chang Song 0001, Yiran Chen 0001, Chaitali Chakrabarti, Jae-sun Seo |
ASP-DAC | 6 |
| 2017 | A Multilayer Approach to Designing Energy-Efficient and Reliable ReRAM Cross-Point Array SystemabstractIn this paper, we study the 1-selector1-resistor (1S1R) cross-point resistive random access memory (ReRAM) array because of its high density, fast access time, and ultralow stand-by power. Specifically, we focus on an access scheme where a data line is parallelly accessed from multiple subarrays with multibits accessed per subarray. A direct implementation of such a scheme has high energy efficiency but lower reliability compared with a single bit per subarray baseline scheme. So this paper proposes a low cost multilayer approach to improve energy-efficiency of multibits per access scheme without compromising reliability. At the cell level, we show how proper choices of bit-line and source-line voltage and SET recovery help reduce error rate by ten times. At the system level, we propose a new rotated multiarray access scheme where the average error rate of every accessed data line is one order of magnitude lower than the worst case, making it possible to achieve block failure rate of 10-10 with a simple Bose, Chaudhuri, and Hocquenghem t = 4 code. We show that for a 1 GB 1S1R ReRAM, the proposed approach can reduce energy by 41% with 2% extra area while maintaining latency and reliability compared with the baseline system. Manqing Mao, Pai-Yu Chen, Shimeng Yu, Chaitali Chakrabarti |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2016 | Efficient memory compression in deep neural networks using coarse-grain sparsification for speech applicationsabstractRecent breakthroughs in deep neural networks have led to the proliferation of its use in image and speech applications. Conventional deep neural networks (DNNs) are fully-connected multi-layer networks with hundreds or thousands of neurons in each layer. Such a network requires a very large weight memory to store the connectivity between neurons. In this paper, we propose a hardware-centric methodology to design low power neural networks with significantly smaller memory footprint and computation resource requirements. We achieve this by judiciously dropping connections in large blocks of weights. The corresponding technique, termed coarse-grain sparsification (CGS), introduces hardware-aware sparsity during the DNN training, which leads to efficient weight memory compression and significant computation reduction during classification without losing accuracy. We apply the proposed approach to DNN design for keyword detection and speech recognition. When the two DNNs are trained with 75% of the weights dropped and classified with 5–6 bit weight precision, the weight memory requirement is reduced by 95% compared to their fully-connected counterparts with double precision, while maintaining similar performance in keyword detection accuracy, word error rate, and sentence error rate. To validate this technique in real hardware, a time-multiplexed architecture using a shared multiply and accumulate (MAC) engine was implemented in 65nm and 40nm low power (LP) CMOS. In 40nm at 0.6 V, the keyword detection network consumes 36µW and the speech recognition network consumes 552µW, making this technique highly suitable for mobile and wearable devices. Deepak Kadetotad, Sairam Arunachalam, Chaitali Chakrabarti, Jae-sun Seo |
ICCAD | 3 |
| 2016 | Low complexity optical flow using neighbor-guided semi-global matchingabstractThis paper presents Neighbor-Guided SemiGlobal Matching (NG-fSGM), a new method for optical flow. It is based on SGM, a popular dynamic programming algorithm for stereo vision, where the disparity of each pixel is calculated by aggregating local matching costs over the entire image to resolve local ambiguity in texture-less and occluded regions. Unlike conventional SGM, NG-fSGM operates on a subset of the search space that has been aggressively pruned based on neighboring pixels' information. Our proposed method achieves a fast approximation of SGM with significantly simpler cost aggregation and flow computation. Compared to a prior SGM extension for optical flow, the proposed NG-fSGM provides about 9x reduction in the number of computations and 5x reduction in the memory requirement with only 0.17% accuracy degradation when evaluated with Middlebury benchmark test cases. Jiang Xiang, Ziyun Li 0001, David T. Blaauw, Hun-Seok Kim, Chaitali Chakrabarti |
ICIP | 5 |
| 2016 | Design of a reliable RRAM-based PUF for compact hardware security primitivesabstractPhysical Unclonable Functions (PUF) have to be highly reliable especially when it is being used along with cryptographic hash modules for key generation. To achieve ultrahigh reliability, the conventional approach employs error correction codes (ECC) based on helper data input. Such an approach not only increases the hardware overhead of the PUF but also reduces the entropy of the system, resulting in both hardware and software security issues. In this paper we design a compact and highly reliable PUF architecture based on resistive random access memory (RRAM). We propose a new design where the sum of the read-out currents of multiple RRAM cells is used for generating one response bit. This method statistically minimizes any early-lifetime failure due to RRAM retention degradation at high temperature or under voltage stress. We employ a device model that is calibrated with IMEC HfOx RRAM experimental data and show that with 8 cells per bit, we can ensure99.9999% reliability) for a lifetime >10 years at 125°C. We embed the RRAM PUF into SHA-256 and show that the hardware overhead of the proposed RRAM PUF based architecture is significantly lower than one that uses a traditional RRAM PUF with ECC. Ayush Shrivastava, Pai-Yu Chen, Yu Cao 0001, Shimeng Yu, Chaitali Chakrabarti |
ISCAS | 5 |
| 2016 | RATT-ECC: Rate Adaptive Two-Tiered Error Correction Codes for Reliable 3D Die-Stacked MemoryabstractThis article proposes a rate-adaptive, two-tiered error-correction scheme (RATT-ECC) that provides strong reliability (10 10 x reduction in raw FIT rate) for an HBM-like 3D DRAM system. The tier-1 code is a strong symbol-based code that can correct errors due to small granularity faults and detect errors caused by large granularity faults; the tier-2 code is an XOR-based code that corrects errors detected by the tier-1 code. The rate-adaptive feature of RATT-ECC enables permanent bank failures to be handled through sparing. It can also be used to significantly reduce the refresh power consumption without decreasing reliability and timing performance. Hsing Min Chen, Carole-Jean Wu, Trevor N. Mudge, Chaitali Chakrabarti |
ACM Trans. Archit. Code Optim. | 4 |
| 2016 | Using Low Cost Erasure and Error Correction Schemes to Improve Reliability of Commodity DRAM SystemsabstractMost server-grade systems provide Chipkill-Correct error protection at the expense of power and performance. In this paper we present a low overhead solution to improving the reliability of commodity DRAM systems with no change in the existing memory architecture. Specifically, we propose five erasure and error correction (E-ECC) schemes that provide at least Chipkill-Correct protection for x4 (Schemes 1, 2 and 3), x8 (Scheme 4) and x16 (Scheme 5) DRAM systems. All schemes have superior error correction performance due to the use of strong symbol-based codes. Synthesis results in 28 nm node show that the decoding latency of these codes is negligible compared to the DRAM access latency. In addition, we make use of erasure codes to extend the lifetime of the DRAM systems. Specifically, once a chip is marked faulty due to persistent errors, all E-ECC schemes correct erasures due to that faulty chip and also correct an additional random error in a second chip. Evaluation with SPEC2006 workloads show that compared to x4 Chipkill-Correct schemes, Scheme 5 has the highest IPC improvement (mean of 7 percent) and Scheme 4 has the largest power reduction (mean of 18 percent) and the largest increase in energy efficiency (mean of 25 percent). Hsing Min Chen, Supreet Jeloka, Akhil Arunkumar, David T. Blaauw, Carole-Jean Wu, Trevor N. Mudge, Chaitali Chakrabarti |
IEEE Trans. Computers | 7 |
| 2015 | Optimizing latency, energy, and reliability of 1T1R ReRAM through appropriate voltage settingsabstractResistive RAM (ReRAM) has fast access time, ultra-low stand-by power and high reliability, making it a viable memory technology to replace DRAM for main memory. The 1-transistor-1-resistor (1T1R) ReRAM array has density comparable to that of a DRAM array and the advantages of lower programming energy and higher reliability compared to the ultrahigh density ReRAM cross-point array. In this paper, we show how circuit operation parameters, such as the pulse amplitude and pulse widths of word-line (WL) voltage, bit-line (BL) voltage, and source-line (SL) voltage can be used to lower latency, lower power and improve reliability. SPICE simulation results demonstrate that appropriate choice of voltage settings can be used to reduce the write latency of the 1T1R cell by 29.4% and reduce write energy by 46.7% over the DRAM cell. Next, we show how the endurance of ReRAM cell can be improved by increasing the ratio between OFF and ON resistances and reducing SL voltage. We find that of these, reducing the SL voltage results in significant improvement in endurance with smaller energy overhead. Next, we evaluate the system-level performance of a 1GB ReRAM and DRAM memory system using CACTI and GEM5. Simulation results using SPEC CPU INT 2006 and DaCapo-9.12 benchmarks show that the ReRAM based main memory can improve IPC by 4.2% and energy by up to 77.8% compared to a DRAM system. Manqing Mao, Yu Cao 0001, Shimeng Yu, Chaitali Chakrabarti |
ICCD | 4 |
| 2014 | A hybrid approach to offloading mobile image classificationabstractCurrent mobile devices are unable to execute complex vision applications in a timely and power efficient manner without offloading some of the computation. This paper examines the tradeoffs that arise from executing some of the workload onboard and some remotely. Feature extraction and matching play an essential role in image classification and have the potential to be executed locally. Along with advances in mobile hardware, understanding the computation requirements of these applications is essential to realize their full potential in mobile environments. We analyze the ability of a mobile platform to execute feature extraction and matching, and prediction workloads under various scenarios. The best configuration for optimal runtime (11% faster) executes feature extraction with a GPU onboard and offloads the rest of the pipeline. Alternatively, compressing and sending the image over the network achieves lowest data transferred (2.5× better) and lowest energy usage (3.7× better) than the next best option. Johann Hauswald, Thomas Manville, Ronald G. Dreslinski, Chaitali Chakrabarti, Trevor N. Mudge |
ICASSP | 5 |
| 2014 | A multi-modal approach to emotion recognition using undirected topic modelsabstractA multi-modal framework for emotion recognition using bag-of-words features and undirected, replicated softmax topic models is proposed here. Topic models ignore the temporal information between features, allowing them to capture the complex structure without a brute-force collection of statistics. Experiments are performed over face, speech and language features extracted from the USC IEMOCAP database. Performance on facial features yields an unweighted average recall of 60.71%, a relative improvement of 8.89% over state-of-the-art approaches. A comparable performance is achieved when considering only speech (57.39%) or a fusion of speech and face information (66.05%). Individually, each source is shown to be strong at recognizing either sadness (speech) or happiness (face) or neutral (language) emotions, while, a multi-modal fusion retains these properties and improves the accuracy to 68.92%. Implementation time for each source and their combination is provided. Results show that a turn of 1 second duration can be classified in approximately 666.65ms, thus making this method highly amenable for real-time implementation. Mohit Shah, Chaitali Chakrabarti, Andreas Spanias |
ISCAS | 2 |
| 2014 | Image processing using approximate datapath unitsabstractIn this paper we present approximate adders and multipliers to reduce the datapath complexity of image processing systems with only a small degradation in PSNR performance. We build upon the approximate circuits proposed in [8] and [9]. We show that selective application of accurate and approximate adders can significantly improve the accuracy of a 2D DCT system. For instance, our implementation of 2D DCT has comparable PSNR performance compared to [8] with 34-50% reduction in area. We also propose an approximate multiplier where the partial products have varying degrees of approximation. Such a multiplier helps improve the accuracy of the system as demonstrated through FFT and Gaussian filter case studies. Madhu Vasudevan, Chaitali Chakrabarti |
ISCAS | 2 |
| 2014 | A Distributed Canny Edge Detector: Algorithm and FPGA ImplementationabstractThe Canny edge detector is one of the most widely used edge detection algorithms due to its superior performance. Unfortunately, not only is it computationally more intensive as compared with other edge detection algorithms, but it also has a higher latency because it is based on frame-level statistics. In this paper, we propose a mechanism to implement the Canny algorithm at the block level without any loss in edge detection performance compared with the original frame-level Canny algorithm. Directly applying the original Canny algorithm at the block-level leads to excessive edges in smooth regions and to loss of significant edges in high-detailed regions since the original Canny computes the high and low thresholds based on the frame-level statistics. To solve this problem, we present a distributed Canny edge detection algorithm that adaptively computes the edge detection thresholds based on the block type and the local distribution of the gradients in the image block. In addition, the new algorithm uses a nonuniform gradient magnitude histogram to compute block-based hysteresis thresholds. The resulting block-based algorithm has a significantly reduced latency and can be easily integrated with other block-based image codecs. It is capable of supporting fast edge detection of images and videos with high resolutions, including full-HD since the latency is now a function of the block size instead of the frame size. In addition, quantitative conformance evaluations and subjective tests show that the edge detection performance of the proposed algorithm is better than the original frame-based algorithm, especially when noise is present in the images. Finally, this algorithm is implemented using a 32 computing engine architecture and is synthesized on the Xilinx Virtex-5 FPGA. The synthesized architecture takes only 0.721 ms (including the SRAM READ/WRITE time and the computation time) to detect edges of 512 × 512 images in the USC SIPI database when clocked at 100 MHz and is faster than existing FPGA and GPU implementations. Qian Xu 0009, Srenivas Varadarajan, Chaitali Chakrabarti, Lina J. Karam |
IEEE Trans. Image Process. | 3 |
| 2013 | Sonic Millip3De: A massively parallel 3D-stacked accelerator for 3D ultrasoundabstractThree-dimensional (3D) ultrasound is becoming common for non-invasive medical imaging because of its high accuracy, safety, and ease of use. Unlike other modalities, ultrasound transducers require little power, which makes hand-held imaging platforms possible, and several low-resolution 2D devices are commercially available today. However, the extreme computational requirements (and associated power requirements) of 3D ultrasound image formation has, to date, precluded hand-held 3D capable devices. We describe the Sonic Millip3De, a new system architecture and accelerator for 3D ultrasound beamformation-the most computationally intensive aspect of image formation. Our three-layer die-stacked design features a custom beamsum accelerator that employs massive data parallelism and a streaming transform-select-reduce pipeline architecture enabled by our new iterative beamsum delay calculation algorithm. Based on RTL-level design and floorplanning for an industrial 45nm process, we show Sonic Millip3De can enable 3D ultrasound with a fully sampled 128×96 transducer array within a 16W full-system power budget (400× less than a conventional DSP solution) and will meet a 5W safe power target by the 11nm node. Richard Sampson, Ming Yang 0004, Siyuan Wei, Chaitali Chakrabarti, Thomas F. Wenisch |
HPCA | 4 |
| 2013 | A speech emotion recognition framework based on latent Dirichlet allocation: Algorithm and FPGA implementationabstractIn this paper, we present a speech-based emotion recognition framework based on a latent Dirichlet allocation model. This method assumes that incoming speech frames are conditionally independent and exchangeable. While this leads to a loss of temporal structure, it is able to capture significant statistical information between frames. In contrast, a hidden Markov model-based approach captures the temporal structure in speech. Using the German emotional speech database EMO-DB for evaluation, we achieve an average classification accuracy of 80.7% compared to 73% for hidden Markov models. This improvement is achieved at the cost of a slight increase in computational complexity. We map the proposed algorithm onto an FPGA platform and show that emotions in a speech utterance of duration 1.5s can be identified in 1.8ms, while utilizing 70% of the resources. This further demonstrates the suitability of our approach for real-time applications on hand-held devices. Mohit Shah, Lifeng Miao, Chaitali Chakrabarti, Andreas Spanias |
ICASSP | 3 |
| 2013 | Data storage time sensitive ECC schemes for MLC NAND Flash memoriesabstractErrors in MLC NAND Flash can be classified into retention errors and program interference (PI) errors. While retention errors are dominant when the data storage time is greater than 1 day, PI errors are dominant for short data storage times. Furthermore these two types of errors have different probabilities of 0->1 or 1->0 bit flips. We utilize the characteristics of the two types of errors in the development of ECC schemes for applications that have different storage times. In both cases, we first apply Gray coding and 2-bit interleaving. The corresponding most significant bit (MSB) and least significant bit (LSB) sub-page has only one type of dominating error (0->1 or 1->0). Next we form a product code using linear block code along rows and even parity check along columns to detect all the possible error locations. We develop an algorithm to choose errors among the possible error locations based on the dominant error type. Performance simulation and hardware implementation results show that the proposed solutions have the same performance as BCH codes with larger error correction capability but with significantly lower hardware overhead. For instance, for a 2KB MLC Flash used in long storage time applications, the proposed ECC scheme has 50% lower energy and 60% lower decoding latency compared to the BCH scheme. Chengen Yang, Deepak Muckatira, Chaitali Chakrabarti |
ICASSP | 4 |
| 2013 | Parallelization techniques for implementing trellis algorithms on graphics processorsabstractIn this paper, we study different schemes to parallelize trellis algorithms for efficient implementation on a GPU. We consider parallelization schemes at the packet-level, subblock-level and trellis-level to increase the number of threads in a GPU implementation. At the trellis-level, we consider state-level, forward-backward traversal and branch-metric parallelism. To evaluate the performance of the different schemes, an LTE uplink Turbo decoder is implemented on an NVIDIA GTX470 GPU. Tradeoffs between throughput, latency and bit error rate are presented. Our most balanced configuration is simultaneously processing multiple subblocks in a packet in conjunction with recovery schemes and trellis-level parallelism, which can achieve a throughput of 19.65 Mbps with a latency of 0.56 ms at bit error rate of 10-5for 1.3 dB channel SNR. We also show how different combinations of parallelization schemes can be used to satisfy systems with widely varying requirements of throughput, latency and bit error rate. Yen-Po Chen, Ronald G. Dreslinski, Chaitali Chakrabarti, Achilleas Anastasopoulos, Scott A. Mahlke, Trevor N. Mudge |
ISCAS | 4 |
| 2013 | Exploring DRAM organizations for energy-efficient and resilient exascale memoriesabstractThe power target for exascale supercomputing is 20MW, with about 30% budgeted for the memory subsystem. Commodity DRAMs will not satisfy this requirement. Additionally, the large number of memory chips (>10M) required will result in crippling failure rates. Although specialized DRAM memories have been reorganized to reduce power through 3D-stacking or row buffer resizing, their implications on fault tolerance have not been considered. We show that addressing reliability and energy is a co-optimization problem involving tradeoffs between error correction cost, access energy and refresh power---reducing the physical page size to decrease access energy increases the energy/area overhead of error resilience. Additionally, power can be reduced by optimizing bitline lengths. The proposed 3D-stacked memory uses a page size of 4kb and consumes 5.1pJ/bit based on simulations with NEK5000 benchmarks. Scaling to 100PB, the memory consumes 4.7MW at 100PB/s which, while well within the total power budget (20MW), is also error-resilient. Bharan Giridhar, Michael Cieslak, Deepankar Duggal, Ronald G. Dreslinski, Hsing Min Chen, Robert Patti, Betina Hold, Chaitali Chakrabarti, Trevor N. Mudge, David T. Blaauw |
SC | 8 |
| 2013 | Energy and Quality-Aware Multimedia Signal ProcessingabstractThis paper presents techniques to reduce energy with minimal degradation in system performance for multimedia signal processing algorithms. It first provides a survey of energy-saving techniques such as those based on voltage scaling, reducing number of computations and reducing dynamic range. While these techniques reduce energy, they also introduce errors that affect the performance quality. To compensate for these errors, techniques that exploit algorithm characteristics are presented. Next, several hybrid energy-saving techniques that further reduce the energy consumption with low performance degradation are presented. For instance, a combination of voltage scaling and dynamic range reduction is shown to achieve 85% energy saving in a low pass FIR filter for a fairly low noise level. A combination of computation reduction and dynamic reduction for Discrete Cosine Transform shows, on average, 33% to 46% reduction in energy consumption while incurring 0.5 dB to 1.5 dB loss in PSNR. Both of these techniques have very little overhead and achieve significant energy reduction with little quality degradation. Yunus Emre, Chaitali Chakrabarti |
IEEE Trans. Multim. | 2 |
| 2013 | Techniques for Compensating Memory Errors in JPEG2000abstractThis paper presents novel techniques to mitigate the effects of SRAM memory failures caused by low voltage operation in JPEG2000 implementations. We investigate error control coding schemes, specifically single error correction double error detection code based schemes, and propose an unequal error protection scheme tailored for JPEG2000 that reduces memory overhead with minimal effect in performance. Furthermore, we propose algorithm-specific techniques that exploit the characteristics of the discrete wavelet transform coefficients to identify and remove SRAM errors. These techniques do not require any additional memory, have low circuit overhead, and more importantly, reduce the memory power consumption significantly with only a small reduction in image quality. Yunus Emre, Chaitali Chakrabarti |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2012 | Accelerating neuromorphic vision algorithms for recognitionabstractVideo analytics introduce new levels of intelligence to automated scene understanding. Neuromorphic algorithms, such as HMAX, are proposed as robust and accurate algorithms that mimic the processing in the visual cortex of the brain. HMAX, for instance, is a versatile algorithm that can be repurposed to target several visual recognition applications. This paper presents the design and evaluation of hardware accelerators for extracting visual features for universal recognition. The recognition applications include object recognition, face identification, facial expression recognition, and action recognition. These accelerators were validated on a multi-FPGA platform and significant performance enhancement and power efficiencies were demonstrated when compared to CMP and GPU platforms. Results demonstrate as much as 7.6X speedup and 12.8X more power-efficient performance when compared to those platforms. Ahmed Al-Maashri, Michael DeBole, Matthew Cotter, Nandhini Chandramoorthy, Yang Xiao 0002, Narayanan Vijaykrishnan, Chaitali Chakrabarti |
DAC | 7 |
| 2012 | Process variation in near-threshold wide SIMD architecturesabstractNear-threshold operation has emerged as a competitive approach for energy-efficient architecture design. In particular, a combination of near-threshold circuit techniques and parallel SIMD computations achieves excellent energy efficiency for easy-to-parallelize applications. However, near-threshold operations suffer from delay variations due to increased process variability. This is exacerbated in wide SIMD architectures where the number of critical paths are multiplied by the SIMD width. This paper provides a systematic in-depth study of delay variations in near-threshold operations and shows that simple techniques such as structural duplication and supply voltage/frequency margining are sufficient to mitigate the timing variation problems in wide SIMD architectures at the cost of marginal area and power overhead. Sangwon Seo, Ronald G. Dreslinski, Mark Woh, Yongjun Park 0001, Chaitali Chakrabarti, Scott A. Mahlke, David T. Blaauw, Trevor N. Mudge |
DAC | 5 |
| 2012 | Neural activity tracking using spatial compressive particle filteringabstractWe investigate and demonstrate the sparsity of electroencephalography (EEG) signals in the spatial domain by incorporating grid spacing in the area of the head enclosing the brain volume. We exploit this spatial sparsity and propose a new approach for tracking neural activity that is based on compressive particle filtering. Our approach results in reducing the number of EEG channels required to be stored and processed for neural tracking using particle filtering. Simulations using both synthetic and real EEG signals illustrate that the proposed algorithm has tracking performance comparable to existing methods while using only a reduced set of EEG channels. Lifeng Miao, Jun Jason Zhang, Antonia Papandreou-Suppappola, Chaitali Chakrabarti |
ICASSP | 4 |
| 2012 | Hierarchical modeling of Phase Change memory for reliable designabstractAs CMOS based memory devices near their end, memory technologies, such as Phase Change Random Access Memory (PRAM), have emerged as viable alternatives. This work develops a hierarchical modeling framework that connects the unique device physics of PRAM with its circuit and state transition properties. Such an approach enables design exploration at various levels in order to optimize the performance and yield. By providing a complete set of compact models, it supports SPICE simulation of PRAM in the presence of process variations and temporal degradation. Furthermore, this work proposes a new metric, State Transition Curve (STC) that supports the assessment of other performance metrics (e.g., power, speed, yield, etc.), helping gain valuable insights on PRAM reliability. Ketul Sutaria, Chengen Yang, Chaitali Chakrabarti, Yu Cao 0001 |
ICCD | 4 |
| 2012 | Design of orthogonal coded excitation for synthetic aperture imaging in ultrasound systemsabstractThis paper presents a new digital front-end architecture for synthetic aperture ultrasound (SAU) imaging using orthogonal chirps and orthogonal Golay codes. Compared to existing systems that perform decoding before beamforming, the proposed systems have comparable performance and significantly lower computation and space complexity. Unfortunately the proposed systems suffer loss in performance in the presence of body motion. To address this problem, we propose a simple motion compensation scheme that improves both the SNR and the RSLL performance. A comparison of the complexity of both the systems shows that while Golay code-based system has lower computation complexity than chirp-based system, if motion compensation is included, then the complexity of the two systems are comparable. Ming Yang 0004, Chaitali Chakrabarti |
ISCAS | 2 |
| 2012 | Transpose-free SAR imaging on FPGA platformabstractRange-Doppler Algorithm (RDA) and Chirp Scaling Algorithm (CSA) are two widely used Synthetic Aperture Radar (SAR) imaging schemes. Both require multiple transpose operations which increase the total processing time significantly. In this paper, we propose transpose-free flow for both RDA and CSA. This is achieved by modifying the existing flows in order to utilize the access patterns favored by the external memory. As a result, the peak performance of the memory is sustained and the processing time shortened. The proposed Field Programmable Gate Array (FPGA)-based implementation outperforms the existing SAR accelerators; it computes RDA and CSA on data size of 4, 096 × 4, 096 in 323ms and 162ms, respectively. Chi-Li Yu, Chaitali Chakrabarti |
ISCAS | 2 |
| 2012 | Product Code Schemes for Error Correction in MLC NAND Flash MemoriesabstractError control coding (ECC) is essential for correcting soft errors in Flash memories. In this paper we propose use of product code based schemes to support higher error correction capability. Specifically, we propose product codes which use Reed-Solomon (RS) codes along rows and Hamming codes along columns and have reduced hardware overhead. Simulation results show that product codes can achieve better performance compared to both Bose-Chaudhuri-Hocquenghem codes and plain RS codes with less area and low latency. We also propose a flexible product code based ECC scheme that migrates to a stronger ECC scheme when the numbers of errors due to increased program/erase cycles increases. While these schemes have slightly larger latency and require additional parity bit storage, they provide an easy mechanism to increase the lifetime of the Flash memory devices. Chengen Yang, Yunus Emre, Chaitali Chakrabarti |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2011 | Low energy motion estimation via selective aproximationsabstractThis paper presents a novel sum of absolute difference (SAD) scheme that significantly reduces the energy consumption of the motion estimation kernel in video coders. The proposed scheme exploits the facts that most of the absolute difference (AD) calculations result in small values, and most of the large AD values do not contribute to the SAD values of the blocks that are selected. Thus the large AD values can be approximated, resulting in lower critical path delay and lower energy consumption of the SAD unit. In addition, the proposed scheme truncates one lower order bit and variants of this scheme implement sub-sampling to further reduce the energy consumption. Performance of the proposed technique is evaluated using the H.264 video encoding framework. Simulation results show 37.5% reduction in energy consumption at nominal voltage and 68% energy reduction for iso-throughput compared to the conventional implementation with only 0.06% drop in PSNR and 1.8% increase in compressed data rate. With an additional ½ sub-sampling, the proposed scheme achieves 90% reduction in energy consumption with 0.6% reduction in PSNR and 6.5% increase in compressed data rate. Yunus Emre, Chaitali Chakrabarti |
ASAP | 2 |
| 2011 | An algorithm-architecture co-design framework for gridding reconstruction using FPGAsabstractGridding is a method of interpolating irregularly sampled data on to a uniform grid and is a critical image reconstruction step in several applications which operate on non-Cartesian sampled data. In this paper, we present an algorithm architecture co-design framework for accelerating gridding using FPGAs. We present a parameterized hardware library for accelerating gridding to support both arbitrary and regular trajectories. We further describe our kernel automation framework which supports several kernel functions through look-up-table (LUT) based Taylor polynomial evaluation. This framework is integrated using an in-house multi-FPGA development platform which provides hardware infrastructure for integrating custom accelerators. Design-space exploration is enabled by an automation flow which allows system generation from an algorithm specification. We further provide several case studies by realizing systems for nonuniform fast Fourier transform (NuFFT) with different parameter sets and porting them on to the BEE3 platform. Results show speedups of more than 16X and 2X over existing CPU and FPGA implementations respectively, and up to 5.5 times higher performance-per-watt over a comparable GPU implementation. Srinidhi Kestur, Kevin M. Irick, Ahmed Al-Maashri, Narayanan Vijaykrishnan, Chaitali Chakrabarti |
DAC | 6 |
| 2011 | Pipeline strategy for improving optimal energy efficiency in ultra-low voltage designabstractThis paper investigates pipelining methodologies for the ultra low voltage regime. Based on an analytical model and simulations, we propose a pipelining technique that provides higher energy efficiency and performance than conventional approaches to ultra low voltage design. Two-phase latch based design and sequential circuit optimizations are also proposed to further improve energy efficiency and performance. Silicon results demonstrate a 16b multiplier using the approaches in 65nm CMOS improve energy efficiency by 30% and performance by 60%. Mingoo Seok, Dongsuk Jeon, Chaitali Chakrabarti, David T. Blaauw, Dennis Sylvester |
DAC | 3 |
| 2011 | Data-path and memory error compensation technique for low power JPEG implementationabstractThis paper presents a novel technique to mitigate effects of data-path and memory errors in JPEG implementations. These errors are mainly caused by voltage scaling and process variation in scaled technologies. We characterize the data-path and memory errors and derive a probability distribution of the total number of errors. We propose an algorithm-specific technique that corrects most errors after quantization in the JPEG encoder by exploiting the characteristics of the quantized coefficients. The technique achieves high performance with small circuit overhead. Simulation results show that the proposed technique has a PSNR performance degradation of around 1.5 dB compared to the error-free case, and 4 dB improvement compared to the no correction case at compression rate of 0.75 bpp when BER = 10-4. Yunus Emre, Chaitali Chakrabarti |
ICASSP | 2 |
| 2011 | Energy-optimized high performance FFT processorabstractThis paper proposes an ultra low energy FFT processor suitable for sensor applications. The processor is based on R4MDC but achieves full utilization of computational elements. It has two parallel datapaths that increase throughput by a factor of 2 and also enable high memory utilization. The proposed design is implemented in 65nm CMOS technology and post-layout simulation including parasitic capacitances shows it achieves 9.25× higher energy efficiency than state-of-the-art FFT processors and high throughput relative to past subthreshold circuit implementations. Dongsuk Jeon, Mingoo Seok, Chaitali Chakrabarti, David T. Blaauw, Dennis Sylvester |
ICASSP | 3 |
| 2011 | A framework for accelerating neuromorphic-vision algorithms on FPGAsabstractImplementations of neuromorphic algorithms are traditionally implemented on platforms which consume significant power, falling short of their biologically underpinnings. Recent improvements in FPGA technology have led to FPGAs becoming a platform in which these rapidly evolving algorithms can be implemented. Unfortunately, implementing designs on FPGAs still prove challenging for nonexperts, limiting their use in the neuroscience domain. In this paper, a FPGA framework is presented which enables neuroscientists to compose multi-FPGA systems for a cortical object classification model. This is demonstrated by mapping this algorithm onto two distinct platforms providing speedups of up to ~28X over a reference CPU implementation. Michael DeBole, Ahmed Al-Maashri, Matthew Cotter, Chi-Li Yu, Chaitali Chakrabarti, Narayanan Vijaykrishnan |
ICCAD | 5 |
| 2010 | A special-purpose compiler for look-up table and code generation for function evaluationabstractElementary functions are extensively used in computer graphics, signal and image processing, and communication systems. This paper presents a special-purpose compiler that automatically generates customized look-up tables and implementations for elementary functions under user given constraints. The generated implementations include a C/C++ code that can be used directly by applications running on multicores, as well as a MATLAB-like code that can be translated directly to a hardware module on FPGA platforms. The experimental results show that our solutions for function evaluation bring significant performance improvements to applications on multicores as well as significant resource savings to designs on FPGAs. Lanping Deng, Praveen Yedlapalli, Sai Prashanth Muralidhara, Hui Zhao 0013, Mahmut T. Kandemir, Chaitali Chakrabarti, Nikos Pitsianis, Xiaobai Sun |
DATE | 7 |
| 2010 | Energy-aware adaptive OFDM systemsabstractAn effective way of reducing the energy consumption without affecting the quality of service is by adapting to the channel conditions. In this paper we describe one such scheme for an OFDM system using space division multiplexing technology. We consider tuning parameters such as modulation level and number of antennas that are considered by WiMax and LTE, as well as number of sub-carriers, peak to average power ratio, interpolation, pilot length and cyclic prefix. We describe a two-phase procedure for determining the parameter settings for minimizing energy consumption given the transmission rate, error performance and channel conditions for frequency selective fading channels. Our results show that we can reduce energy by an additional 5%-30% compared to systems that can only adapt modulation order and number of antennas. Yunus Emre, Chaitali Chakrabarti |
ICASSP | 2 |
| 2010 | A distributed psycho-visually motivated Canny edge detectorabstractThis paper proposes a distributed Canny edge detection algorithm which can be mapped onto multi-core architectures for high throughput applications. In contrast to the conventional Canny edge detection algorithm which makes use of the global image gradient histogram to determine the threshold for edge detection, the proposed algorithm adaptively computes the edge detection threshold based on the local distribution of the gradients in the considered image block. The efficacy of the distributed Canny in detecting psycho-visually important edges is validated using a visual sharpness metric. The proposed distributed Canny edge detection algorithm has the capacity to scale up the throughput adaptively, based on the number of computing engines. The algorithm achieves about 72 times speed up for a 16-core architecture, without any change in performance. Furthermore, the internal memory requirements are significantly reduced especially for smaller block sizes. For instance, if a 512×512 image is processed in 64×64 blocks using the proposed scheme, the memory is reduced by a factor of 70 as compared to the original Canny edge detector. Srenivas Varadarajan, Chaitali Chakrabarti, Lina J. Karam, Judit Martinez Bauza |
ICASSP | 2 |
| 2010 | Bandwidth-intensive FPGA architecture for multi-dimensional DFTabstractMulti-dimensional (MD) Discrete Fourier Transform (DFT) is a key kernel algorithm in many signal processing algorithms, including radar data processing and medical imaging. Although there are many efficient software solutions, they are not suitable for applications that require fast response time. In this paper we focus on FPGA-based implementation of MDDFT. The proposed architecture is based on a decomposition algorithm that takes into account FPGA resources and the characteristics of off-chip memory access, namely, the burst access pattern of the Synchronous Dynamic RAM (SDRAM). The architecture can support 2D, 3D, and even higher dimensional DFT with high performance. It has been implemented on a Xilinx Virtex-5 FPGA platform and its performance for 2D and 3D DFT measured and analyzed. Chi-Li Yu, Chaitali Chakrabarti, Narayanan Vijaykrishnan |
ICASSP | 2 |
| 2010 | Diet SODA: a power-efficient processor for digital camerasabstractPower has become the most critical design constraint for embedded handheld devices. This paper proposes a power-efficient SIMD architecture, referred to as Diet SODA, for DSP applications. The key design idea is to apply near-threshold operation on a single instruction and multiple data (SIMD) architecture to significantly lower the power consumption. The major features of Diet SODA are very wide SIMD width, scatter/gather data prefetcher, and dual mode operation. A case study was performed on digital still camera (DSC) applications; the results show that Diet SODA achieves ∼130x better performance and ∼340x better energy efficiency than a DSP solution. Sangwon Seo, Ronald G. Dreslinski, Mark Woh, Chaitali Chakrabarti, Scott A. Mahlke, Trevor N. Mudge |
ISLPED | 4 |
| 2010 | A Low-Power DSP for Wireless CommunicationsabstractThis paper proposes a low-power high-throughput digital signal processor (DSP) for baseband processing in wireless terminals. It builds on our earlier architecture-Signal processing On Demand Architecture (SODA)-which is a four-processor, 32-lane SIMD machine that was optimized for WCDMA 2 Mbps and IEEE 802.11a. SODA has several shortcomings including large register file power, wasted cycles for data alignment, etc., and cannot satisfy the higher throughput and lower power requirements of emerging standards. We propose SODA-II, which addresses these problems by deploying the following schemes: operation chaining, pipelined execution of SIMD units, staggered memory access, and multicycling of computation units. Operation chaining involves chaining the primitive instructions, thereby eliminating unnecessary register file accesses and saving power. Pipelined execution of the vector instructions through the SIMD units improves the system throughput. Staggered execution of computation units helps simplify the data alignment networks. It is implemented in conjunction with multicycling so that the computation units are busy most of the time. The proposed architecture is evaluated with an in-house architecture emulator which uses component-level area and power models built with Synopsys and Artisan tools. Our results show that for WCDMA 2 Mbps, the proposed architecture uses two processors and consumes only 120 mW while SODA uses four processors and consumes 210 mW when implemented in 0.13-μm technology and clocked at 300 MHz. Hyunseok Lee, Chaitali Chakrabarti, Trevor N. Mudge |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2009 | Energy-aware error control coding for Flash memoriesabstractThe use of Flash memories in portable embedded systems is ever increasing. This is because of the multi-level storage capability that makes them excellent candidates for high density memory devices. However, cost of writing or programming Flash memories is an order of magnitude higher than traditional memories. In this paper, we design an algorithm to reduce both average write energy and latency in Flash memories. We achieve this by reducing the number of expensive '01' and '10' bit-patterns during error control coding. We show that the algorithm does not change the error correction capability and moreover improves endurance. Simulations results on representative bit-stream traces show that the use of the proposed algorithm saves, on average, 33% of write energy and 31% of latency of Intel MLC NOR Flash memory, and improves the endurance by 24%. Veera Papirla, Chaitali Chakrabarti |
DAC | 2 |
| 2009 | An H.264/SVC memory architecture supporting spatial and course-grained quality scalabilitiesabstractThe standardized scalable video coding (SVC) extension of H.264/AVC achieves significant improvements in coding efficiency relative to the scalable profiles of prior video coding standards, but its computational complexity and memory access requirements make the design of a low power hardware architecture a challenging task. This paper presents an SVC decoder architecture supporting spatial and coarse-grained quality scalability. The architecture optimizes the size of the on-chip memory and reduces the power-consuming and time-intensive external memory accesses. Niranjan D. Narvekar, Bharatan Konnanath, Shalin M. Mehta, Santosh Chintalapati, Ismail AlKamal, Chaitali Chakrabarti, Lina J. Karam |
ICIP | 6 |
| 2009 | AnySP: anytime anywhere anyway signal processingabstractIn the past decade, the proliferation of mobile devices has increased at a spectacular rate. There are now more than 3.3 billion active cell phones in the world-a device that we now all depend on in our daily lives. The current generation of devices employs a combination of general-purpose processors, digital signal processors, and hardwired accelerators to provide giga-operations-per-second performance on milliWatt power budgets. Such heterogeneous organizations are inefficient to build and maintain, as well as waste silicon area and power. Looking forward to the next generation of mobile computing, computation requirements will increase by one to three orders of magnitude due to higher data rates, increased complexity algorithms, and greater computation diversity but the power requirements will be just as stringent. Scaling of existing approaches will not suffice instead the inherent computational efficiency, programmability, and adaptability of the hardware must change. To overcome these challenges, this paper proposes an example architecture, referred to as AnySP, for the next generation mobile signal processing. AnySP uses a co-design approach where the next generation wireless signal processing and high-definition video algorithms are analyzed to create a domain specific programmable architecture. At the heart of AnySP is a configurable single-instruction multiple-data datapath that is capable of processing wide vectors or multiple narrow vectors simultaneously. In addition, deeper computation subgraphs can be pipelined across the single-instruction multiple-data lanes. These three operating modes provide high throughput across varying application types. Results show that AnySP is capable of sustaining 4G wireless processing and high-definition video throughput rates, and will approach the 1000 Mops/mW efficiency barrier when scaled to 45nm. Mark Woh, Sangwon Seo, Scott A. Mahlke, Trevor N. Mudge, Chaitali Chakrabarti, Krisztián Flautner |
ISCA | 5 |
| 2009 | Low power robust signal processingabstractVoltage scaling has proven to be very effective in reducing the power consumption of digital systems. However, voltage overscaling, ie., reducing the voltage below the critical voltage, introduces errors which have to be compensated by additional computations. In this paper, we propose the use of radix-2 redundant binary arithmetic (RBR) as an alternative for designing low power robust systems. We show that for large data widths, such systems have superior energy-delay product (EDP) and error performance compared to 2's complement based systems. However, for smaller data widths, the 2's complement system has better EDP performance, and in such cases, we propose a low complexity prediction technique to compensate for voltage overscaled errors. We evaluate the performance of the RBR system and the 2's complement system for two image processing kernels, namely, the Gaussian filter and DCT/IDCT. We show that the RBR system is the low power design choice for the Gaussian filter with 16 bits precision while the 2's complement system with most significant bit prediction is the low power design choice for the IDCT kernel. Veera Papirla, Aarul Jain, Chaitali Chakrabarti |
ISLPED | 3 |
| 2009 | An Automated Framework for Accelerating Numerical Algorithms on Reconfigurable Platforms Using Algorithmic/Architectural OptimizationabstractThis paper describes TANOR, an automated framework for designing hardware accelerators for numerical computation on reconfigurable platforms. Applications utilizing numerical algorithms on large-size data sets require high-throughput computation platforms. The focus is on N-body interaction problems which have a wide range of applications spanning from astrophysics to molecular dynamics. The TANOR design flow starts with a MATLAB description of a particular interaction function, its parameters, and certain architectural constraints specified through a graphical user interface. Subsequently, TANOR automatically generates a configuration bitstream for a target FPGA along with associated drivers and control software necessary to direct the application from a host PC. Architectural exploration is facilitated through support for fully custom fixed-point and floating-point representations in addition to standard number representations such as single-precision floating point. Moreover, TANOR enables joint exploration of algorithmic and architectural variations in realizing efficient hardware accelerators. TANOR's capabilities have been demonstrated for three different N-body interaction applications: the calculation of gravitational potential in astrophysics, the diffusion or convolution with Gaussian kernel common in image processing applications, and the force calculation with vector-valued kernel function in molecular dynamics simulation. Experimental results show that TANOR-generated hardware accelerators achieve lower resource utilization without compromising numerical accuracy, in comparison to other existing custom accelerators. Jungsub Kim, Lanping Deng, Prasanth Mangalagiri, Kevin M. Irick, Kanwaldeep Sobti, Mahmut T. Kandemir, Narayanan Vijaykrishnan, Chaitali Chakrabarti, Nikos Pitsianis, Xiaobai Sun |
IEEE Trans. Computers | 8 |
| 2009 | Design Methodology for Low Power and Parametric Robustness Through Output-Quality Modulation: Application to Color-Interpolation FilteringabstractPower dissipation and robustness to process variation have conflicting design requirements. Scaling of voltage is associated with larger variations, while Vdd upscaling or transistor up-sizing for parametric-delay variation tolerance can be detrimental for power dissipation. However, for a class of signal-processing systems, effective tradeoff can be achieved between Vdd scaling, variation tolerance, and ldquooutput quality.rdquo In this paper, we develop a novel low-power variation-tolerant algorithm/architecture for color interpolation that allows a graceful degradation in the peak-signal-to-noise ratio (PSNR) under aggressive voltage scaling as well as extreme process variations. This feature is achieved by exploiting the fact that all computations used in interpolating the pixel values do not equally contribute to PSNR improvement. In the presence of Vdd scaling and process variations, the architecture ensures that only the ldquoless important computationsrdquo are affected by delay failures. We also propose a different sliding-window size than the conventional one to improve interpolation performance by a factor of two with negligible overhead. Simulation results show that, even at a scaled voltage of 77% of nominal value, our design provides reasonable image PSNR with 40% power savings. Nilanjan Banerjee, Georgios Karakonstantis, Jung Hwan Choi, Chaitali Chakrabarti, Kaushik Roy 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2009 | Maximizing the Lifetime of Embedded Systems Powered by Fuel Cell-Battery HybridsabstractFuel cell (FC) is a viable alternative power source for portable applications; it has higher energy density than traditional Li-ion battery and thus can achieve longer lifetime for the same weight or volume. However, because of its limited power density, it can hardly track fast fluctuations in the load current of digital systems. A hybrid power source, which consists of a FC and a Li-ion battery, has the advantages of long lifetime and good load following capabilities. In this paper, we consider the problem of extending the lifetime of a fuel-cell-based hybrid source that is used to provide power to an embedded system which supports dynamic voltage scaling (DVS). We propose an energy-based optimization framework that considers the characteristics of both the energy consumer (the embedded system) and the energy provider (the hybrid power source). We use this framework to develop algorithms that determine the output power level of the FC and the scaling factor of the DVS processor during task scheduling. Simulations on task traces based on a real-application (Path Finder) and a randomized version demonstrate significant superiority of our algorithms with respect to a conventional DVS algorithm which only considers energy minimization of the embedded system. Jianli Zhuo, Chaitali Chakrabarti, Kyungsoo Lee, Naehyuck Chang, Sarma B. K. Vrudhula |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2008 | Accurate models for estimating area and power of FPGA implementationsabstractThis paper presents accurate area and power estimation models for implementations using FPGAs from the Xilinx Virtex-2Pro family. These models are designed to facilitate efficient design space exploration in an automated algorithm-architecture codesign framework. Detailed models for accurately estimating the number of slices, block RAMs and 18times18-bit multipliers for fixed point and floating-point IP cores have been developed. These models are also utilized to develop accurate power models that consider the effect of logic power, signal power, clock power and I/O power. In all cases, the model coefficients have been derived by using curve fitting or regression analysis. The modeling error for the IP cores is very small (average 0.95%). The error for fairly large examples such as floating point implementation of 8-point FFTs is also quite small; it is 1.87% for estimation of number of slices and 3.48% for estimation of power consumption. Lanping Deng, Kanwaldeep Sobti, Chaitali Chakrabarti |
ICASSP | 3 |
| 2008 | Extending the lifetime of media recorders constrained by battery and flash memory sizeabstractThe lifetime of a stand-alone media recorder is a function of both the battery size and flash memory size. In this paper, we present a power management framework for media recorders that significantly enhances their lifetime while minimizing the flash memory usage and maintaining the same level of recording quality. This is achieved by implementing a mixture of encoding algorithms of different complexities that generate data with different compression ratios, and in turn balancing the energy consumption and the flash memory usage. Younghyun Kim 0001, Youngjin Cho, Naehyuck Chang, Chaitali Chakrabarti, Nam Ik Cho |
ISLPED | 4 |
| 2008 | From SODA to scotch: The evolution of a wireless baseband processorabstractWith the multitude of existing and upcoming wireless standards, it is becoming increasingly difficult for hardware-only baseband processing solutions to adapt to the rapidly changing wireless communication landscape. Software defined radio (SDR) promises to deliver a cost effective and flexible solution by implementing a wide variety of wireless protocols in software. In previous work, a fully programmable multicore architecture, SODA, was proposed that was able to meet the real-time requirements of 3G wireless protocols. SODA consists of one ARM control processor and four wide single instruction multiple data (SIMD) processing elements. Each processing element consists of a scalar and a wide 512-bit 32-lane SIMD datapath. A commercial prototype based on the SODA architecture, Ardbeg (named after a brand of Scotch whisky), has been developed. In this paper, we present the architectural evolution of going from a research design to a commercial prototype, including the goals, tradeoffs, and final design choices. Ardbegpsilas redesign process can be grouped into the following three major areas: optimizing the wide SIMD datapath, providing long instruction word (LIW) support for SIMD operations, and adding application-specific hardware accelerators. Because SODA was originally designed with 180 nm technology, the wide SIMD datapath is re-optimized in Ardbeg for 90 nm technology. This includes re-evaluating the most efficient SIMD width, designing a wider SIMD shuffle network, and implementing faster SIMD arithmetic units. Ardbeg also provides modest LIW support by allowing two SIMD operations to issue in the same cycle. This LIW execution supports SDR algorithmspsila most common parallel SIMD execution patterns with minimal hardware overhead. A viable commercial SDR solution must be competitive with existing ASIC solutions. Therefore, algorithm-specific hardware is added for performance bottleneck algorithms while still maintaining enough flexibility to support multiple wireless protocols. The combination of these architectural improvements allows Ardbeg to achieve 1.5-7x speedup over SODA across multiple wireless algorithms while consuming less power. Mark Woh, Yuan Lin 0002, Sangwon Seo, Scott A. Mahlke, Trevor N. Mudge, Chaitali Chakrabarti, Richard Bruce, Danny Kershaw, Alastair Reid 0001, Mladen Wilder, Krisztián Flautner |
MICRO | 6 |
| 2008 | Energy-efficient dynamic task scheduling algorithms for DVS systemsabstractDynamic voltage scaling (DVS) is a well-known low-power design technique that reduces the processor energy by slowing down the DVS processor and stretching the task execution time. However, in a DVS system consisting of a DVS processor and multiple devices, slowing down the processor increases the device energy consumption and thereby the system-level energy consumption. In this paper, we first use system-level energy consideration to derive the “optimal ” scaling factor by which a task should be scaled if there are no deadline constraints. Next, we develop dynamic task-scheduling algorithms that make use of dynamic processor utilization and optimal scaling factor to determine the speed setting of a task. We present algorithm duEDF , which reduces the CPU energy consumption and algorithm duSYS and its reduced preemption version, duSYS_PC , which reduce the system-level energy. Experimental results on the video-phone task set show that when the CPU power is dominant, algorithm duEDF results in up to 45% energy savings compared to the non-DVS case. When the CPU power and device power are comparable, algorithms duSYS and duSYS_PC achieve up to 25% energy saving compared to CPU energy-efficient algorithm duEDF , and up to 12% energy saving over the non-DVS scheduling algorithm. However, if the device power is large compared to the CPU power, then we show that a DVS scheme does not result in lowest energy. Finally, a comparison of the performance of algorithms duSYS and duSYS_PC show that preemption control has minimal effect on system-level energy reduction. Jianli Zhuo, Chaitali Chakrabarti |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2008 | A fuel-cell-battery hybrid for portable embedded systemsabstractThis article presents our work on the development of a fuel cell (FC) and battery hybrid (FC-Bh) system for use in portable microelectronic systems. We describe the design and control of the hybrid system, as well as a dynamic power management (DPM)-based energy management policy that extends its operational lifetime. The FC is of the proton exchange membrane (PEM) type, operates at room temperature, and has an energy density which is 4--6 times that of a Li-ion battery. The FC cannot respond to sudden changes in the load, and so a system powered solely by the FC is not economical. An FC-Bh power source, on the other hand, can provide the high energy density of the FC and the high power density of a battery. In this work we first describe the prototype FC-Bh system that we have built. Such a prototype helps to characterize the performance of a hybrid power source, and also helps explore new energy management strategies for embedded systems powered by hybrid sources. Next we describe a Matlab/Simulink-based FC-Bh system simulator which serves as an alternate experimental platform and that enables quick evaluation of system-level control policies. Finally, we present an optimization framework that explicitly considers the characteristics of the FC-Bh system and is aimed at minimizing the fuel consumption. This optimization framework is applied on top of a prediction-based DPM policy and is used to derive a new fuel-efficient DPM scheme. The proposed scheme demonstrates up to 32% system lifetime extension compared to a competing scheme when run on a real trace-based MPEG encoding example. Kyungsoo Lee, Naehyuck Chang, Jianli Zhuo, Chaitali Chakrabarti, Sudheendra Kadri, Sarma B. K. Vrudhula |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2007 | Dynamic Power Management with Hybrid Power SourcesabstractDPM (Dynamic Power Management) is an effective technique for reducing the energy consumption of embedded systems that is based on migrating to a low power state when possible. While conventional DPM minimizes the energy consumption of the embedded system, it does not utilize the properties of the power source. Alternative power sources such as fuel cells (PCs) have substantially different power and efficiency characteristics that have to be taken into account while developing policies that maximize their operational lifetime. In this paper, we present a new DPM policy for embedded systems powered by FC based hybrid source. We develop an optimization framework that explicitly considers the FC system efficiency and is aimed at minimizing the fuel consumption. Next we apply this optimization framework on top of a prediction based DPM policy to develop a new fuel-efficient DPM scheme. The proposed algorithm was applied to a real trace based MPEG encoding example and demonstrated up to 32% more system lifetime extension compared to a competing scheme. Jianli Zhuo, Chaitali Chakrabarti, Kyungsoo Lee, Naehyuck Chang |
DAC | 2 |
| 2007 | TANOR: A Tool for Accelerating N-Body Simulations on Reconfigurable PlatformsabstractAlgorithm-architecture co-exploration is hindered by the lack of efficient tools. As a consequence, designers are currently able to explore only a limited set of points in the whole design space. Therefore, a tool that can allow fast exploration of algorithmic and architectural tradeoffs in an automated manner is highly desired. In this paper, we describe TANOR an automated tool targeted for designing hardware accelerators for the class of N-body interaction problems. The design flow, starting from a high level (MATLAB) description, configures the entire system automatically. We describe the design of TANOR and demonstrate the effectiveness and adaptability of our tool using three different target applications, namely, the gravitational kernel used in astrophysics, the gaussian kernel common in image processing applications, and a force calculation kernel applied in molecular dynamics. Our results demonstrate that TANOR generates hardware accelerator that are competitive with existing custom accelerator. Jungsub Kim, Prasanth Mangalagiri, Kevin M. Irick, Mahmut T. Kandemir, Narayanan Vijaykrishnan, Kanwaldeep Sobti, Lanping Deng, Chaitali Chakrabarti, Nikos Pitsianis, Xiaobai Sun |
FPL | 8 |
| 2007 | Memory Efficient LDPC Code Design for High Throughput Software Defined Radio (SDR) systemsabstractLow-density parity-check (LDPC) codes have been adopted in the physical layer protocol of many communication systems because of their superior performance. A direct implementation of the LDPC decoder on an existing platform, such as a software defined radio (SDR), is likely to be inefficient. Our approach is to design the LDPC code in a way that takes into account the constraints imposed by the existing architecture, without compromising the communication performance. In this paper, a procedure for architecture-aware LDPC code design which minimize the number of global memory accesses in a memory constrained system is derived. The procedure is built on top of existing super-code based LDPC code design. The proposed code construction procedure also results in reduction in the number of iterations and thereby increases the throughput significantly. Yuming Zhu, Chaitali Chakrabarti |
ICASSP (2) | 2 |
| 2007 | Design methodology to trade off power, output quality and error resiliency: application to color interpolation filteringabstractPower dissipation and tolerance to process variations pose conflicting design requirements. Scaling of voltage is associated with larger variations, while Vdd upscaling or transistor up-sizing for process tolerance can be detrimental for power dissipation. However, for certain signal processing systems such as those used in color image processing, we noted that effective trade-offs can be achieved between Vdd scaling, process tolerance and “output quality”. In this paper we demonstrate how these tradeoffs can be effectively utilized in the development of novel low-power variation tolerant architectures for color interpolation. The proposed architecture supports a graceful degradation in the PSNR (Peak Signal to Noise Ratio) under aggressive voltage scaling as well as extreme process variations in sub-70nm technologies. This is achieved by exploiting the fact that some computations are more important and contribute more to the PSNR improvement compared to the others. The computations are mapped to the hardware in such a way that only the less important computations are affected by Vdd-scaling and process variations. Simulation results show that even at a scaled voltage of 60% of nominal Vdd value, our design provides reasonable image PSNR with 69% power savings Georgios Karakonstantis, Nilanjan Banerjee, Kaushik Roy 0001, Chaitali Chakrabarti |
ICCAD | 4 |
| 2007 | Throughput of multi-core processors under thermal constraintsabstractWe analyze the effect of thermal constraints on the performance and power of multi-core processors. We propose system-level power and thermal models, and derive expressions for (a) the maximum number of cores that can be activated, with and without throttling, (b) the speedup (multi-core over single core), and the total power consumption, both as functions of the number of active cores. These expressions involve parameters like power per core, thermal resistance of hottest die block and package, and leakage dependence on temperature. We also computed the above metrics (a) and (b) numerically by solving the detailed Hotspot circuit of an multicore processor driven by a block-level exponential temperaturedependent leakage model. When compared to these numerical results, we found that the above expressions for (a) were at most 8% underpredicted, while those for (b) were accurately predicted. The proposed analytical approach is the first of its kind to relate metrics of interest in multi-core processors to high-level design parameters. Compared to numerical approaches, it provides much faster computation time, and valuable insight for processor designers. Ravishankar Rao, Sarma B. K. Vrudhula, Chaitali Chakrabarti |
ISLPED | 3 |
| 2007 | Energy management of DVS-DPM enabled embedded systems powered by fuel cell-battery hybrid sourceabstractDynamic voltage scaling (DVS) and dynamic power management (DPM) are the two main techniques for reducing the energy consumption of embedded systems. The effectiveness of both DVS and DPMneeds to be considered in the development of an energy management policy for a system that consists of both DVS-enabled and DPM-enabled components. The characteristics of the power source also have to be explicitly taken into account. In this paper, we propose a policy to maximize the operational lifetime of a DVS-DPM enabled embedded system powered by a fuel cell-battery (FC-B) hybrid source. We show that the lifetime of the system is determined by the fuel consumption of the fuel cell (FC), and that the fuel consumption can be minimized by a combination of a load energy minimization policy and an optimal fuel flow control policy. The proposed method, when applied to a randomized task trace, demonstrated superior performance compared to competing policies based on DVS and/or DPM. Jianli Zhuo, Chaitali Chakrabarti, Naehyuck Chang |
ISLPED | 2 |
| 2007 | A System Level Energy Model and Energy-Quality Evaluation for Integrated Transceiver Front-EndsabstractAs CMOS technology scales down, digital supply voltage and digital power consumption goes down. However, the supply voltage and power consumption of the RF front-end and analog sections do not scale in a similar fashion. In fact, in many state-of-the-art communication transceivers, RF and analog sections can consume more energy compared to the digital part. In this paper, first, a system level energy model for all the components in the RF and analog front-end is presented. Next, the RF and analog front-end energy consumption and communication quality of three representative systems are analyzed: a single user point-to-point wireless data communication system, a multi-user code division multiple access (CDMA)-based system and a receive-only video distribution system. For the single user system, the effect of occupied signal bandwidth, peak-to-average ratio (PAR), symbol rate, constellation size, and pulse-shaping filter roll-off factor is analyzed; for the CDMA-based multi-user system, the effect of the number of users in the cell and multiple access interference (MAI) along with the PAR and filter roll-off factor is studied; for the receive-only system, the effect of 1/f noise for direct-conversion receiver and the effect of IF frequency for low-IF architecture on the RF front-end power consumption is analyzed. For a given communication quality specification, it is shown that the energy consumption of a wireless communication front-end can be scaled down by adjusting parameters such as the pulse shaping filter roll-off factor, constellation size, symbol rate, number of users in the cell, and signal center frequency Bertan Bakkaloglu, Chaitali Chakrabarti |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2007 | Automatic antenna-tuning unit for software-defined and cognitive radioabstractAbstract This paper discusses the implementation of an automatic antenna tuning unit system (ATU) for software‐defined and cognitive radio. The ATU simplifies radio frequency (RF) front‐end design for multi‐band, multi‐mode radios by allowing an electrically small reconfigurable antenna to become a frequency‐agile selective component (essentially a tunable filter). In implementing the ATU, impedance synthesizers (tunable matching networks) using RF MEMS switches as the control elements are used to match a more or less arbitrary load to a convenient impedance value. To generate feedback data that can be used to optimize the impedance synthesizer, the incident and reflected powers at the input to the impedance synthesizer are sampled using a power sensor block comprising directional coupler, logarithmic RF power detectors, analog‐to‐digital converters (ADCS), and return loss computation algorithms running on a field programmable gate array (FPGA). From a practical point of view, in order to be compatible with commercial wireless handset devices, we propose to design and ultimately implement a fully integrated ATU system. In this paper, a hard‐ ware implementation of the ATU prototype has been demonstrated to verify narrowband automatic tuning ability of the ATU under constantly changing environment conditions, and we present simulated results for a logarithmic power detector designed with 55 dB dynamic range over the 800 MHz–2 GHz frequency band using 0.25 µ CMOS technology. Copyright © 2007 John Wiley & Sons, Ltd. Sung-Hoon Oh, James T. Aberle, Bertan Bakkaloglu, Chaitali Chakrabarti |
Wirel. Commun. Mob. Comput. | 5 |
| 2006 | High-level power management of embedded systems with application-specific energy cost functionsabstractMost existing dynamic voltage scaling (DVS) schemes for multiple tasks assume an energy cost function (energy consumption versus execution time) that is independent of the task characteristics. In practice the actual energy cost functions vary significantly from task to task. Different tasks running on the same hardware platform can exhibit different memory and peripheral access patterns, cache miss rates, etc. These effects results in a distinct energy cost function for each task.We present a new formulation and solution to the problem of minimizing the total (dynamic and static) system energy while executing a set of tasks under DVS. First, we demonstrate and quantify the dependence of the energy cost function on task characteristics by direct measurements on a real hardware platform (the TI OMAP processor) using real application programs. Next, we present simple analytical solutions to the problem of determining energy-optimal voltage scale factors for each task, while allowing each task to be preempted and to have its own energy cost function. Based on these solutions, we present simple and efficient algorithms for implementing DVS with multiple tasks. We consider two cases: (1) all tasks have a single deadline, and (2) each task has its own deadline. Experiments on a real hardware platform using real applications demonstrate a 10% additional saving in total system energy compared to previous leakage-aware DVS schemes. Youngjin Cho, Naehyuck Chang, Chaitali Chakrabarti, Sarma B. K. Vrudhula |
DAC | 3 |
| 2006 | Extending the lifetime of fuel cell based hybrid systemsabstractFuel cells are clean power sources that have much higher energy densities and lifetimes compared to batteries. However, fuel cells have limited load following capabilities and cannot be efficiently utilized if used in isolation. In this work, we consider a hybrid system where a fuel cell based hybrid power source is used to provide power to a DVFS processor. The hybrid power source consists of a room temperature fuel cell operating as the primary power source and a Li-ion battery (that has good load following capability) operating as the secondary source. Our goal is to develop polices to extend the lifetime of the fuel cell based hybrid system. First, we develop a charge based optimization framework which minimizes the charge loss of the hybrid system (and not the energy consumption of the DVFS processor). Next, we propose a new algorithm to minimize the charge loss by judiciously scaling the load current. We compare the performance of this algorithm with one that has been optimized for energy, and demonstrate its superiority. Finally, we evaluate the performance of the hybrid system under different system configurations and show how to determine the best combination of fuel cell size and battery capacity for a given embedded application. Jianli Zhuo, Chaitali Chakrabarti, Naehyuck Chang, Sarma B. K. Vrudhula |
DAC | 2 |
| 2006 | Aggregated Circulant Matrix Based LDPC CodesabstractThis paper presents a variation of circulant matrix based LDPC codes which allows more than one circulant identity matrix in a submatrix of the parity check matrix. The aggregated LDPC supports higher decoding throughput with small increase in datapath complexity. The construction algorithm, bit error rate (BER) performance, information update rule and the architecture for high decoding throughput are also presented. Yuming Zhu, Chaitali Chakrabarti |
ICASSP (3) | 2 |
| 2006 | SODA: A Low-power Architecture For Software RadioabstractThe physical layer of most wireless protocols is traditionally implemented in custom hardware to satisfy the heavy computational requirements while keeping power consumption to a minimum. These implementations are time consuming to design and difficult to verify. A programmable hardware platform capable of supporting software implementations of the physical layer, or software defined radio, has a number of advantages. These include support for multiple protocols, faster time-to-market, higher chip volumes, and support for late implementation changes. The challenge is to achieve this without sacrificing power. In this paper, we present a design study for a fully programmable architecture, SODA, that supports software defined radio a high-end signal processing application. Our design achieves high performance, energy efficiency, and programmability through a combination of features that include single-instruction multiple- data (SIMD) parallelism, and hardware optimized for 16bit computations. The basic processing element is an asymmetric processor consisting of a scalar and SIMD pipeline, and a set of distributed scratchpad memories that are fully managed in software. Results show that a four processor design is capable of meeting the throughput requirements of theW-CDMA and 802.11a protocols, while operating within the strict power constraints of a mobile terminal. Yuan Lin 0002, Hyunseok Lee, Mark Woh, Yoav Harel, Scott A. Mahlke, Trevor N. Mudge, Chaitali Chakrabarti, Krisztián Flautner |
ISCA | 7 |
| 2006 | Reducing idle mode power in software defined radio terminalsabstractIn this paper, we propose a processor which is optimized for idle mode operation of a software defined radio (SDR) terminal. Since a SDR terminal spends most of its time in the idle mode, reducing the power consumption in this mode directly translates to longer terminal standby time. Workload analysis of idle mode operations of contemporary standards showed that these are dominated by FIR filtering, which can be easily parallelized. This analysis was used in the design of the idle mode processor. The key architectural components are an SIMD unit for the parallel computations that dominate the workload, a conventional scalar unit for the sequential computations, and a control unit which supports efficient data memory access and loop control. The idle mode processor was modeled with Verilog and synthesized using standard cells in 0.13 micron technology. It consumes about 9mW at 1.08V. Hyunseok Lee, Trevor N. Mudge, Chaitali Chakrabarti |
ISLPED | 3 |
| 2006 | An optimal analytical solution for processor speed control with thermal constraintsabstractAs semiconductor manufacturing technology scales to smaller device sizes, the power consumption of clocked digital ICs begins to increase. Dynamic voltage and frequency scaling (DVFS) is a well-known technique for conserving energy. Recently, it has also been used to control the CPU temperature as part of Dynamic Thermal Management (DTM) techniques. Most works in these areas assume that the optimum speed profile (for either minimizing energy or maximizing performance) is a constant profile. However, in the presence of thermal constraints, we show that the optimal profile is in general, a time-varying function. We formulate the problem of maximizing the average throughput of a processor over a given time period, subject to thermal and speed constraints, as a problem in the calculus of variations. The variational approach provides a powerful framework for precisely specifying and solving the speed control problem, and allows us to obtain an exact analytical solution. The solution methodology is very general, and works for any convex power model, and simple lumped RC thermal models. The resulting speed profiles were found to consist of up to three segments, of which one of them is a decreasing function of time, and the others are constant. We analyze the effect of different parameters like the initial temperature, thermal capacitance and the maximum rated speed on the nature and the cost of the optimum solution. We also propose a two-speed solution that approximates the optimal speed curve. This solution was found to achieve a performance close to that of the optimum, and is also easier to implement in real processors. Ravishankar Rao, Sarma B. K. Vrudhula, Chaitali Chakrabarti, Naehyuck Chang |
ISLPED | 3 |
| 2006 | Maximizing the lifetime of embedded systems powered by fuel cell-battery hybridsabstractFuel cells are a viable alternative power source for portable applications. They have higher energy density than traditional Li-ion batteries and can achieve longer lifetime for the same weight or volume. However, because of their limited power density, they can not track fluctuations in the load current fast. A hybrid power source, that consists of a fuel cell and a Li-ion battery, has the advantages of long lifetime and good load following capabilities. In this work, we consider the problem of extending the lifetime of a fuel-cell based hybrid source that is used to provide power to a DVFS processor. We propose a new algorithm that is built on top of an energy based optimization framework. The algorithm simultaneously adjusts the fuel flow rate (at the producer end), and judiciously scales the load current (at the consumer end) to minimize the energy loss of the hybrid system. Simulations on randomly generated task sets demonstrate the superiority of this algorithm with respect to an algorithm that does not allow adjustment of the fuel flow rate. Jianli Zhuo, Chaitali Chakrabarti, Naehyuck Chang, Sarma B. K. Vrudhula |
ISLPED | 2 |
| 2006 | A coprocessor architecture for fast protein structure prediction
Rahim Khoja, Mehul Marolia, Tinku Acharya, Chaitali Chakrabarti |
Pattern Recognit. | 4 |
| 2006 | Study of energy and performance of space-time decoding systems in concatenation with turbo decodingabstractRecent studies have shown that using space-time code is an effective approach to increase the data rate over wireless channels. Space-time turbo (ST-Turbo) codes formed by concatenating space-time codes with turbo codes, take advantage of both the high diversity order of space-time systems and the randomness of the turbo codes. In this paper, we compare two ST-Turbo codes, i.e., simple space-time turbo codes (SiSTT) and turbo trellis-coded modulation space-time block codes (TTCM-STBCs), and their approximate versions with respect to performance and energy consumption for both general-purpose processor and synthesized implementations. The approximations are aimed at reducing the computational complexity and include reduction in the number of paths, number of iterations, and datapath computations. Analysis of the simulation results show that SiSTT-based versions should be used for higher SNR applications where low energy consumption is the primary design objective, and TTCM-STBC-based versions should be used where performance is the primary design objective. Finally, four ST-Turbo algorithms (i.e., baseline SiSTT and its energy-efficient approximate version and the baseline TTCM-STBC and its energy-efficient approximate version) have been synthesized in 0.18-/spl mu/m CMOS technology and the implementations compared with respect to area, power, and latency. Yuming Zhu, Chaitali Chakrabarti |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2005 | An efficient dynamic task scheduling algorithm for battery powered DVS systemsabstractBattery lifetime enhancement is a critical design parameter for mobile computing devices. Maximizing battery life-time is a particularly difficult problem due to the non-linearity of the battery behavior and its dependence on the characteristics of the discharge profile. In this paper we address the problem of dynamic task scheduling with voltage scaling in a battery-powered DVS system. The objective is to maximize the battery performance measured in terms of charge consumption during execution of the tasks. We present a new battery-aware dynamic task scheduling algorithm, darEDF, based on an efficient slack utilization scheme that employs dynamic speed setting of tasks in run queue. We compare darEDF with three state of the art energy-efficient algorithms, lpfpsEDF, lppsEDF, lpSEH, with respect to battery performance and energy consumption. We show that darEDF has better performance than lpSEH (which has close to optimal energy value), and has lower run-time complexity. Jianli Zhuo, Chaitali Chakrabarti |
ASP-DAC | 2 |
| 2005 | System-level energy-efficient dynamic task schedulingabstractDynamic voltage scaling (DVS) is a well-known low power design technique that reduces the processor energy by slowing down the DVS processor and stretching the task execution time. But in a DVS system consisting of a DVS processor and multiple devices, slowing down the processor increases the device energy consumption and thereby the system-level energy consumption. In this paper, we present dynamic task scheduling algorithms for periodic tasks that minimize the system-level energy (CPU energy + device standby energy). The algorithms use a combination of (i) optimal speed setting, which is the speed that minimizes the system energy for a specific task, and (ii) limited preemption which reduces the numbers of possible preemptions. For the case when the CPU power and device power are comparable, these algorithms achieve up to 43 % energy savings compared to [1], but only up to 12 % over the non-DVS scheduling. If the device power is large compared to the CPU power, we show that DVS should not be employed. Jianli Zhuo, Chaitali Chakrabarti |
DAC | 2 |
| 2005 | Static task-scheduling algorithms for battery-powered DVS systemsabstractBattery lifetime enhancement is a critical design parameter for mobile computing devices. Maximizing the battery lifetime is a particularly difficult problem due to the nonlinearity of the battery behavior and its dependence on the characteristics of the discharge profile. In this paper, we address the problem of task scheduling with voltage scaling in a battery-powered single and multiprocessor system such that the residual charge or the battery voltage (the parameters for evaluating battery performance) is maximized. We propose an efficient heuristic algorithm using a charge-based cost function derived from the analytical battery model. Our algorithm first creates a task sequence that ensures battery survival, and then distributes the available delay slack so that the cost function is maximized. The effectiveness of the algorithm has been verified using DUALFOIL, a low-level Li-ion battery simulator. The algorithm has been validated on synthetic examples created from applications running on Compaq's handheld computing research platform, ITSY Princey Chowdhury, Chaitali Chakrabarti |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2005 | Memory sub-banking scheme for high throughput MAP-based SISO decodersabstractThe sliding window (SW) approach has been proposed as an effective means of reducing the memory requirements as well as the decoding latency of the maximum a posteriori (MAP) based soft-input soft-output (SISO) decoder in a Turbo decoder. In this paper, we present sub-banked memory implementations (both single port and dual port) of the SW SISO decoder that achieves high throughput, low decoding latency, and reduced memory energy consumption. Our contributions include derivation of the optimal memory sub-banked structure for different SW configurations, study of the relationship between memory size and energy consumption for different SW configurations and study of the effect of number of sub-banks on the throughput/decoding latency for a given SW configuration. Mayank Tiwari 0002, Yuming Zhu, Chaitali Chakrabarti |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2004 | Memory sub-banking scheme for high throughput turbo decoderabstractTurbo codes have revolutionized the world of coding theory with their superior performance. However, the implementation of these codes is both computationally and memory-intensive. Recently, the sliding window (SW) approach has been proposed as an effective means of reducing the decoding delay as well as the memory requirements of turbo implementations. In this paper, we present a sub-banked implementation of the SW-based approach that achieves high throughput, low decoding latency and reduced memory energy consumption. Our contributions include derivation of the optimal memory sub-banked structure for different SW configurations, study of the relationship between memory size, energy consumption and decoding latency for different SW configurations and study of the effect of number of sub-banks on the throughput and decoding latency of a given SW configuration. The theoretical study has been validated by SimpleScalar for a rate 1/3 MAP decoder. Mayank Tiwari 0002, Yuming Zhu, Chaitali Chakrabarti |
ICASSP (5) | 3 |
| 2004 | Design and implementation of low-energy turbo decodersabstractTurbo codes have been chosen in the third generation cellular standard for high-throughput data communication. These codes achieve remarkably low bit error rates at the expense of high-computational complexity. Thus for hand held communication devices, designing energy efficient Turbo decoders is of great importance. In this paper, we present a suite of MAP-based Turbo decoding algorithms with energy-quality tradeoffs for additive white Gaussian noise (AWGN) and fading channels. We derive these algorithms by applying approximation techniques such as pruning the trellis, reducing the number of states, scaling the extrinsic information, applying sliding window, and early termination on the MAP-based algorithm. We show that a combination of these techniques can result in energy savings of 53.2%(50.0%) on a general purpose processor and energy savings of 80.66%(80.81%) on a hardware implementation for AWGN (fading) channels if a drop of 0.35 dB in SNR can be tolerated, at a bit error rate (BER) of 10/sup -5/. We also propose an adaptive Turbo decoding technique that is suitable for low power operation in noisy environments. J. Kaza, Chaitali Chakrabarti |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2003 | A high-performance JPEG2000 architectureabstractJPEG2000 is an upcoming compression standard for still images that has a feature set well tuned for diverse data dissemination. These features are possible due to adaptation of the discrete wavelet transform, intra-subband bit-plane coding, and binary arithmetic coding in the standard. We propose a system-level architecture capable of encoding and decoding the JPEG2000 core algorithm that has been defined in Part I of the standard. The key components include dedicated architectures for wavelet, bit plane, and arithmetic coders and memory interfacing between the coders. The system architecture has been implemented in VHDL and its performance evaluated for a set of images. The estimated area of the architecture, in 0.18-/spl mu/ technology, is 3-mm square and the estimated frequency of operation is 200 MHz. Kishore Andra, Chaitali Chakrabarti, Tinku Acharya |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2003 | Variable voltage task scheduling algorithms for minimizing energy/powerabstractIn this paper, we propose variable voltage task scheduling algorithms that minimize energy or minimize peak power for the case when the task arrival times, deadline times, execution times, periods, and switching activities are given. We consider aperiodic (earliest due date, earliest deadline first), as well as periodic (rate monotonic, earliest deadline first) scheduling algorithms. We use the Lagrange multiplier method to theoretically determine the relation between the task voltages such that the energy or peak power is minimum, and then develop an iterative algorithm that satisfies the relation. The asymptotic complexity of the existing scheduling algorithms change very mildly with the application of the proposed algorithms. We show experimentally (random experiments as well as real-life cases), that the voltage assignment obtained by the proposed low-complexity algorithm is very close to that of the optimal energy (0.1% error) and optimal peak power (1% error) assignment. Ali Manzak, Chaitali Chakrabarti |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2002 | Battery-conscious task sequencing for portable devices including voltage/clock scalingabstractOperation of battery-powered portable systems can no longer be sustained once a battery becomes discharged. Maximization of the battery lifetime is a difficult task due to nonlinearity of battery behavior that depends on the characteristics of the system load profile. We address the problem of task sequencing without and with voltage/clock scaling that shapes the profile so that the battery lifetime is maximized. We developed an accurate analytical battery model and validated it with measurements taken on a real lithium-ion battery used in a pocket computer. We use the model as a basis for a unique battery conscious cost function and utilize its properties to develop several novel algorithms, including insertion of recovery periods and voltage/clock scaling for delay slack distribution. Daler N. Rakhmatov, Sarma B. K. Vrudhula, Chaitali Chakrabarti |
DAC | 3 |
| 2002 | Energy-efficient turbo decoderabstractTurbo codes have been recently adopted in the next generation of wideband CDMA standards. These codes achieve superior performance at the expense of high computational complexity. This makes their low energy implementation a very important yet challenging problem. In this paper we study the effect of different approximation techniques such as pruning the trellis, reducing the number of states, sliding window, early termination on the Bit Error Rate (BER) and energy consumption for a Turbo decoder implemented on a general purpose processor. We show that a combination of these techniques can result in 66.5% energy reduction for a log-MAP based Turbo decoder if a loss of 1.4 dB in SNR can be tolerated at BER = 10−5. Jagadeesh Kaza, Chaitali Chakrabarti |
ICASSP | 2 |
| 2002 | Low-power approach for decoding convolutional codes with adaptive viterbi algorithm approximationsabstractSignificant power reduction can be achieved by exploiting real-time variation in system characteristics while decoding convolutional codes.The approach proposed herein adaptively approximates Viterbi decoding by varying truncation length and pruning threshold of the T-algorithm while employing trace-back memory management. Adaptation is performed according to variations in signal-to-noise ratio, code rate, and maximum acceptable bit error rate.Potential energy reduction of 70 to 97.5% compared to Viterbi decoding is demonstrated.Superiority of adaptive T-algorithm decoding compared to fixed T-algorithm decoding is studied.General conclusions about when applications can particularly benefit from this approach are given. Russell E. Henning, Chaitali Chakrabarti |
ISLPED | 2 |
| 2002 | A low power scheduling scheme with resources operating at multiple voltagesabstractThis paper presents resource and latency constrained scheduling algorithms to minimize power/energy consumption when the resources operate at multiple voltages (5 V, 3.3 V, 2.4 V, and 1.5 V). The proposed algorithms are based on efficient distribution of slack among the nodes in the data-flow graph. The distribution procedure tries to implement the minimum energy relation derived using the Lagrange multiplier method in an iterative fashion. Two algorithms are proposed, 1) a low complexity O(n/sup 2/) algorithm and 2) a high complexity O(n/sup 2/ log(L)) algorithm, where n is the number of nodes and L is the latency. Experiments with some HLS benchmark examples show that the proposed algorithms achieve significant power/energy reduction. For instance, when the latency constraint is 1.5 times the critical path delay, the average reduction is 39%. Ali Manzak, Chaitali Chakrabarti |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2001 | Address Code Generation for Digital Signal ProcessorsabstractA wM %&Qp x q _ LU, % 8 71S v7 8AOy z){+ |}d ~ ~d~ v7 / %) 8 # 9 0! " A8;:L b @ 8:9 4 Sathishkumar Udayanarayanan, Chaitali Chakrabarti |
DAC | 2 |
| 2001 | Efficient implementation of a set of lifting based wavelet filtersabstractLifting based wavelet transform implementation not only helps in reducing the number of computations but also achieves lossy to lossless performance with finite precision. We first do a precision analysis for the set of seven filters proposed by the JPEG2000 verification model. We determine the precision required to implement the filters using fixed point 2's complement arithmetic for lossless as well as lossy coding. Next we propose a unified architecture for implementing this set of filters for both the forward and the inverse transform. Kishore Andra, Chaitali Chakrabarti, Tinku Acharya |
ICASSP | 2 |
| 2001 | An approach for enabling DCT/IDCT energy reduction scalability in MPEG-2 video codecsabstractIt would be desirable, in terms of energy conservation, to use a low complexity approximate algorithm to do all DCT and IDCT computation in an MPEG-2 video codec. However, there is a significant quality penalty associated with this approach that may not always be acceptable. A practical algorithmic method is studied for achieving scalable energy reduction during DCT and IDCT computation in MPEG-2 video codecs at the expense of reasonable amounts of quality. For example, by applying exact and approximate DCT/IDCT algorithms appropriately, the energy consumption of DCT and IDCT execution in two video codecs communicating with one another can be reduced by 8% for quality reduction of 0.4 dB average PSNR, 14% for 0.8 dB; reduction, or 22% for 1.4 dB reduction. Russell E. Henning, Chaitali Chakrabarti |
ICASSP | 2 |
| 2001 | Voltage Scaling for Energy Minimization with QoS ConstraintsabstractWe propose variable voltage scheduling algorithms that minimize energy while satisfying the quality of service (QoS) requirements. We consider the case when multiple applications are running on a single processor equipped with a limited sized buffer and each application has a different computational load and timing constraint. We use the Lagrange multiplier method to theoretically determine the relation between the application voltages such that the energy is minimum, and then develop iterative algorithms to satisfy the relation. The iterative algorithms find the minimum energy solution with polynomial time complexity for both the off-line case and the online case. We show the effect of buffer size and application deadline times on the ability of the system to reduce energy. Furthermore, we consider the effect of discharge current on battery life and show that the voltage assignment for maximum battery capacity is very similar to the voltage assignment for maximum energy. Ali Manzak, Chaitali Chakrabarti |
ICCD | 2 |
| 2001 | Variable voltage task scheduling algorithms for minimizing energyabstractIn this paper we propose variable voltage task scheduling algorithms (periodic as well as aperiodic) that minimize energy. We rst apply the existing task scheduling algorithms to obtain a feasible schedule and then distribute the available slack using an iterative algorithm that satis es the theoretically obtained relation for minimum energy. Weshow experimentally that the voltage assignment obtained by our algorithm is very close (0.1% error) to that of the optimal assignment. Ali Manzak, Chaitali Chakrabarti |
ISLPED | 2 |
| 2001 | Data memory design and exploration for low-power embedded systemsabstractIn embedded system design, the designer has to choose an on-chip memory configuration that is suitable for a specific application. To aid in this design choice, we present a memory exploration procedure based on three performance metrics, namely, cache size, the memory access time and the energy consumption. We show the importance of including energy in the performance metrics, since an increase in the cache size and line size reduces the memory access time but does not necessarily reduce the energy consumption. The memory exploration procedures enable us to find the cache configuration (cache size, line size) that satisfies the area and time constraints while minimizing the energy consumption, and the cache configuration that satisfies the area and energy constraints while minimizing the memory access time. The exploration procedures for cache configuration is very efficient since it considers only a selected set of candidate points. Finally, we validate our exploration procedures by running simulation experiments on MediaBench applications. Wen-Tsong Shiue, Sathishkumar Udayanarayanan, Chaitali Chakrabarti |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2000 | Variable voltage task scheduling for minimizing energy or minimizing powerabstractWe propose task scheduling algorithms that minimize energy or minimize power for the case when the tasks have different arrival times, deadline times, execution times and switching activities. We theoretically determine the relation between the operating voltages for the minimum energy (power) assignment and develop a polynomial time scheduling algorithm that uses this relation. We show experimentally that the voltage assignment obtained by our algorithm is very close to that of the optimal assignment. Ali Manzak, Chaitali Chakrabarti |
ICASSP | 2 |
| 2000 | A Multi-Bit Binary Arithmetic Coding TechniqueabstractWe propose a new methodology for binary arithmetic coding which reduces the number of arithmetic operations significantly at the expense of a mild reduction in compression ratio. We achieve this by (i) considering a two-symbol nonoverlapping window and not coding the second symbol if both of them are most probable symbols and (ii) moving the majority of computations to the least probable symbol path. As a result, we reduce the additions/subtractions required by 60-70%, with a loss of compression ratio of about 1-3% compared to the Q-coder. This reduction in computational complexity makes the proposed technique particularly suitable for low-power VLSI implementation. We have described the proposed algorithm and analyzed the results. We have also described a VLSI architecture capable of carrying out the algorithm. Kishore Andra, Tinku Acharya, Chaitali Chakrabarti |
ICIP | 3 |
| 2000 | A programmable processor for cryptographyabstractCryptography has numerous applications in today's world, the most prevalent one being transferring messages safely over the network. Cryptographic algorithms are either implemented in software on a general-purpose processor or in hardware on an application-specific processor. While the software implementations tend to be time consuming, the hardware implementations are too specific and cannot even support small modifications. In this paper, a programmable architecture that can handle a large number of algorithms including DES, RSA, Blowfish, SAFER, et cetera has been developed. The architecture consists of addition, subtraction, modular multiplication, exponentiation and XOR units and thus can support a majority of the cryptographic algorithms. A high data rate is achieved by applying loop unrolling to the Montgomery algorithm that is used for modular multiplication and exponentiation. The differences in the number of bits, key length, and sequence of operations is handled by the microprogrammed control unit. A VHDL model has been developed and synthesized using AutoLogic II from Mentor Graphics. The results show a frequency of operation of 77 Megahertz and an area of 23,000 "Optimization COST" units. Sukumar S. Raghuram, Chaitali Chakrabarti |
ISCAS | 2 |
| 2000 | ILP-based scheme for low power scheduling and resource bindingabstractIn this paper, we present an ILP based scheme for high-level synthesis for low power applications. Specifically, we present (i) an ILP-based model for latency constrained scheduling that minimizes the number of resources, the peak power consumption and peak area, and (ii) a LP-based model for resource binding that minimizes the amount of switching at the input of the functional units. The ILP based scheduler is very flexible since it allows the relative importance of the three objectives (number of resources, peak power, peak area) to be determined by user-defined weighting factors. The LP-based method for resource binding consists of creating a multistage graph with m stages (corresponding to m cycles in the schedule) and n nodes per stage (corresponding to n functional units of the same type) and finding n disjoint paths such that the total cost (corresponding to the switching activity) of these paths is minimum. Wen-Tsong Shiue, Chaitali Chakrabarti |
ISCAS | 2 |
| 2000 | Energy-efficient code generation for DSP56000 family (poster session)abstractThis paper presents a procedure to generate energy-efficient code for the Motorola DSP56K processor based on increasing the packing efficiency and minimizing the number of address instructions. The key features are a novel scheduling algorithm that reduces the dependencies between instructions, a register allocation algorithm that spills variables based on their packability, and an address code generation algorithm that minimizes the number of additional instructions. The size of the code generated by this procedure is on the average 45% (25%) smaller than that generated by Motorola's g56K (SPAM). Sathishkumar Udayanarayanan, Chaitali Chakrabarti |
ISLPED | 2 |
| 2000 | VLSI architectures for weighted order statistic (WOS) filters
Chaitali Chakrabarti, Lori E. Lucke |
Signal Process. | 1 |
| 1999 | Memory Exploration for Low Power, Embedded SystemsabstractArticle Memory exploration for low power, embedded systems Share on Authors: Wen-Tsong Shiue Arizona State University, Department of Electrical Engineering, Tempe, AZ Arizona State University, Department of Electrical Engineering, Tempe, AZView Profile , Chaitali Chakrabarti Arizona State University, Department of Electrical Engineering, Tempe, AZ Arizona State University, Department of Electrical Engineering, Tempe, AZView Profile Authors Info & Claims DAC '99: Proceedings of the 36th annual ACM/IEEE Design Automation ConferenceJune 1999 Pages 140–145https://doi.org/10.1145/309847.309902Online:01 June 1999Publication History 136citation707DownloadsMetricsTotal Citations136Total Downloads707Last 12 Months5Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Wen-Tsong Shiue, Chaitali Chakrabarti |
DAC | 2 |
| 1999 | Activity models for use in low power, high-level synthesisabstractCharacteristics of the data being processed can be used to reduce the power consumption in the data path of a VLSI circuit by exploiting their relationship with transition activity during high-level synthesis. Important relationships between fixed-point, two's complement data characteristics and 0/spl rarr/1 transition activity in static CMOS circuits are presented in this paper. Models for computing transition activity in terms of a new set of transition parameters are developed. Propagation of data characteristics through multiplication and addition functional units is discussed. The use of the relationships and models to analyze and significantly reduce 0/spl rarr/1 transition activity with little computational effort is illustrated with examples. Russell E. Henning, Chaitali Chakrabarti |
ICASSP | 2 |
| 1999 | Efficient realizations of encoders and decoders based on the 2-D discrete wavelet transformabstractIn this paper, we present architectures and scheduling algorithms for encoders and decoders that are based on the two-dimensional discrete wavelet transform. We consider the design of encoders and decoders individually, as well as in an integrated encoder-decoder system. We propose architectures ranging from a single-instruction multiple-data processor arrays to folded architectures that are suitable for single-chip implementations. The scheduling algorithms for the folded architectures range from those that try to minimize the latency to those that try to minimize the storage and keep the data flow regular. We include a comparison of the performance of these algorithms to aid the designer in choosing one that is best suited for a specific application. Chaitali Chakrabarti, Clint Mumford |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 1997 | High-Level Design Synthesis of a Low Power, VLIW Processor for the IS-54 VSELP Speech EncoderabstractGeneral purpose DSPs typically used to implement speech coders in digital cellular phones do not allow enough exploitation of the speech coding algorithm itself for power reduction. In this paper, high-level design synthesis of a low power, VLIW (very long instruction word) processor dedicated to implementing the IS-54 VSELP speech encoding algorithm is presented. Significant power reduction is achieved through algorithm dependent techniques, including application specific hardware design, supply voltage reduction through highly parallel execution, and exploitation of data correlation inherent to the algorithm. Preliminary estimates indicate that the design could result in a 5.35 mm/sup 2/ processor that executes in real-time with an average power dissipation of about 28 mW. Russell E. Henning, Chaitali Chakrabarti |
ICCD | 2 |
| 1996 | Efficient realizations of analysis and synthesis filters based on the 2-D discrete wavelet transformabstractThis paper presents folded architectures and scheduling algorithms for computing the 2-D DWT for analysis and synthesis filters. The folded architectures consist of two parallel computation units (one for computations along the rows and the other for computations along the columns) and two storage units to store the intermediate outputs that are generated by the two units. The scheduling algorithms range from those that try to minimise the latency to those that try to minimise the control unit complexity and keep the data flow regular. A comparison of the scheduling algorithms has been included to aid the designer in choosing an algorithm that is best suited for a particular application. Chaitali Chakrabarti, Clint Mumford |
ICASSP | 1 |
| 1996 | Motion estimation of two-dimensional objects based on the straight line hough transform: A new approach
Hsiang-Ling Li, Chaitali Chakrabarti |
Pattern Recognit. | 2 |
| 1996 | A new architecture for the Viterbi decoder for code rate k/nabstractA novel VLSI architecture is proposed for implementing a long constraint length Viterbi decoder (VD) for code rate k/n. This architecture is based on the encoding structure where k input bits are shifted into k shift registers in each cycle. The architecture is designed in a hierarchical manner by breaking the system into several levels and designing each level independently. The tasks in the design of each level range from determining the number of computation units, and the interconnection between the units, to the allocation and scheduling of operations. Additional design issues such as in-place storage of accumulated path metrics and trace back implementation of the survivor memory have also been addressed. The resulting architecture is regular, has a foldable global topology and is very flexible. It also achieves a better than linear trade-off between hardware complexity and computation time. Hsiang-Ling Li, Chaitali Chakrabarti |
IEEE Trans. Commun. | 2 |
| 1995 | A survey of architectures for the discrete and continuous wavelet transformsabstractWavelet transforms have proven to be useful tools for several applications, including signal analysis, signal coding, and image compression. This paper surveys the VLSI architectures that have been proposed for computing the discrete and continuous wavelet transforms for 1-D and 2-D signals. The proposed architectures range from SIMD arrays to folded architectures such as systolic arrays and parallel filters. The SIMD arrays have a size that is proportional to that of the data sequence and are optimal with respect to time. The folded architectures, on the other hand, support single chip implementations and are optimal with respect to both area and time under the word-serial model. Chaitali Chakrabarti, Mohan Vishwanath, Robert Michael Owens |
ICASSP | 1 |
| 1995 | A new Viterbi decoder design for code rate k/nabstractA novel VLSI architecture is proposed for implementing a long constraint length Viterbi decoder (VD) for code rate k/n. This architecture is based on the encoding structure where k input bits are shifted into k shift registers in each cycle. The architecture is designed in a hierarchical manner by breaking the system into several levels and designing each level independently. At each level, the number of computation units, the interconnection between the units as well as allocation and scheduling issues have been determined. In-place storage of accumulated path metrics and trace back implementation of the survivor memory have also been addressed. The resulting architecture is regular, flexible and achieves and better than linear tradeoff between hardware complexity and computation time. Hsiang-Ling Li, Chaitali Chakrabarti |
ICASSP | 2 |
| 1995 | Low power data format converter design using semi-static register allocationabstractIn many applications, such as digital signal processing, data format converters are used to reformat the data transferred between processing modules. In VLSI implementations, these converters consume a large portion of the available resources. Various methods have been proposed to synthesize data format converter architectures while optimizing the number of registers used to store the data. In this paper, we present a new register allocation scheme which not only minimizes the number of resistors, but also minimizes the power consumption in the data format converter. Low power data format converters are synthesized by minimizing the transitions and interconnections between the registers used to store the data. We present both a heuristic and an integer linear programming formulation to solve the allocation problem. Our method shows significant improvement over previous techniques. Kala Srivatsan, Chaitali Chakrabarti, Lori E. Lucke |
ICCD | 2 |
| 1995 | A New Viterbi Decoder Design for Code Rate K/NabstractA novel VLSI architecture is proposed for implementing a long constraint length Viterbi Decoder (VD) for code rate k/n. This architecture is based on the encoding structure where k input bits are shifted into k shift registers in each cycle. The architecture is designed in a hierarchical manner by breaking the system into several levels and designing each level independently. At each level, the number of computation units, the interconnection between the units as well as allocation and scheduling issues have been determined. In-place storage of accumulated path metrics and trace back implementation of the survivor memory have also been addressed. The resulting architecture is regular, flexible and achieves a better than linear tradeoff between hardware complexity and computation time. Hsiang-Ling Li, Chaitali Chakrabarti |
ISCAS | 2 |
| 1995 | A New Architecture for the Viterbi Decoder for Code Rate k/n1abstractA novel VLSI architecture is proposed for implementing a long constraint length Viterbi decoder (VD) for code rate k/n. This architecture is based on the encoding structure where k input bits are shifted into k shift registers in each cycle. The architecture is designed in a hierarchical manner by breaking the system into several levels and designing each level independently. The tasks in the design of each level range from determining the number of computation units, and the interconnection between the units, to the allocation and scheduling of operations. Additional design issues such as in-place storage of accumulated path metrics and trace back implementation of the survivor memory have also been addressed. The resulting architecture is regular, has a foldable global topology and is very flexible. It also achieves a better than linear trade-off between hardware complexity and computation time. Hsiang-Ling Li, Chaitali Chakrabarti |
IEEE Trans. Commun. | 2 |
| 1995 | Architectures for hierarchical and other block matching algorithmsabstractHierarchical block matching is an efficient motion estimation technique which provides an adaptation of the block size and the search area to the properties of the image. In this paper, we propose two novel special-purpose architectures to implement hierarchical block matching for real-time applications. The first architecture is memory-efficient, but requires a large external memory bandwidth and a large number of processors. The second architecture requires significantly fewer processors, but additional on-chip memory. We describe in details the processor architecture, the memory organization and the scheduling for both these architectures. We also show how the second architecture can be modified to handle full-search and 3-step hierarchical search block matching algorithms, with significant reduction in the hardware complexity as compared to existing architectures. Gagan Gupta 0002, Chaitali Chakrabarti |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1995 | A digit-serial architecture for gray-scale morphological filteringabstractWe present a digit-serial architecture for gray-scale morphological operations that operate on radix-2 redundant numbers. We present new implementations for a redundant number adder and a maximum unit that are used in the morphological dilation unit. These new designs have areas comparable to 2's complement implementations but have significantly smaller latencies. Lori E. Lucke, Chaitali Chakrabarti |
IEEE Trans. Image Process. | 2 |
| 1994 | A VLSI architecture for real-time hierarchical encoding/decoding of video using the wavelet transformabstractNovel online algorithms and architectures for hierarchical coding using the wavelet transform are presented. These algorithms/architectures compute the decomposition-reconstruction cycle of the wavelet transform with minimum latency and buffering for any given blocking factor. A hierarchical scheme which enables cheap video conferencing and/or multicast over a heterogeneous network is also presented. This architecture supports single chip implementations of the encoder, the decoder, and the transcoder for some choices of the wavelet filter and vector quantization schemes.> Mohan Vishwanath, Chaitali Chakrabarti |
ICASSP (2) | 2 |
| 1994 | Efficient Architectures for Hidden Surface RemovalabstractWe present several new efficient architectures to solve the hidden surface problem in the feature domain. All the architectures operate on segments (instead of pixels) and create a list of visible segments for each scan line. We present two new semi-systolic architectures consisting of an array of M processors, where M is the maximum number of overlapping segments. Both architectures require presorting of the segment endpoints and have a latency of O(N), where N is the number of input segments. We present two new sorting network architectures which do not require any presorting of the endpoints. These architectures consist of O(log N) stages of segment merge units and are based on odd-even merge sort and running merge sort.> Chaitali Chakrabarti, Lori E. Lucke |
ICIP (1) | 1 |
| 1994 | Novel Sorting Netowrk-Based Architectures for Rank Order FiltersabstractThis paper presents two novel sorting network-based architectures for computing high sample rate non recursive rank order filters. The proposed architectures consist of significantly fewer comparators than existing architectures that are based on bubble-sort and Batcher's odd-even merge sort. The reduction in the number of comparators is obtained by sorting the columns of the window only once, and by merging the sorted columns in a way such that the set of candidate elements for the output is very small. The number of comparators per output are reduced even further by processing a block of outputs at a time.> Chaitali Chakrabarti, Li-Yu Wang |
ISCAS | 1 |
| 1994 | VLSI Architectures for Hierarchical Block MatchingabstractHierarchical block matching is an efficient motion estimation technique which provides an adaptation of the block size and the search area, to the properties of the image. In this work, we propose two novel special-purpose architectures for implementing hierarchical block matching. The first architecture is memory-efficient, but requires a large external memory bandwidth and a large number of processors. The second architecture requires significantly fewer processors, but additional on-chip memory. We describe the processor architecture, the memory organization and the scheduling details for both the architectures.> Gagan Gupta 0002, Chaitali Chakrabarti |
ISCAS | 2 |
| 1994 | High Sample Rate Architectures for Block Adaptive FiltersabstractIn this paper we propose a variety of architectures for implementing block adaptive filters in the time-domain. These filters are based on a block implementation of the least mean squares (BLMS) algorithm. First, we present an architecture which directly maps the BLMS algorithm into an array of processors. Next, we describe an architecture where the weight vector is updated without explicitly computing the filter error. Third, we describe an architecture which exploits the redundant computations of overlapping windows. All the architectures have a significantly smaller sample period compared to frequency domain implementations. Moreover, the sample periods can be reduced even further by applying relaxed look-ahead techniques.> Srikanth Karkada, Chaitali Chakrabarti, Andreas Spanias |
ISCAS | 2 |
| 1994 | A Digit-Serial Architecture for Gray-Scale Morphological FilteringabstractWe present a digit-serial architecture for gray-scale morphological operations which operates on radix-2 redundant numbers. We present new implementations of a redundant number adder and maximum unit used in the morphological dilation unit. These new designs have areas comparable to 2's complement implementations, but have significantly smaller latencies.> Lori E. Lucke, Chaitali Chakrabarti |
ISCAS | 2 |
| 1994 | Novel sorting network-based architectures for rank order filtersabstractThis paper presents two novel sorting network-based architectures for computing high sample rate nonrecursive rank order filters. The proposed architectures consist of significantly fewer comparators than existing sorting network-based architectures that are based on bubble-sort and Batcher's odd-even merge sort. The reduction in the number of comparators is obtained by sorting the columns of the window only once, and by merging the sorted columns in a way such that the number of candidate elements for the output is very small. The number of comparators per output is reduced even further by processing a block of outputs at a time. Block processing procedures that exploit the computational overlap between consecutive windows are developed for both the proposed networks.> Chaitali Chakrabarti, Li-Yu Wang |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 1993 | VLSI architectures for recursive median filters
Chaitali Chakrabarti |
ICASSP (1) | 1 |
| 1993 | Efficient stack filter implementations of rank order filters
Chaitali Chakrabarti |
ISCAS | 1 |
| 1991 | VLSI Architectures for Multidimensional TransformsabstractThe authors propose a family of VLSI architectures with area-time tradeoffs for computing (N*N* . . . *N) d-dimensional linear separable transforms. For fixed-precision arithmetic with b bits, the architectures have an area A=O(N/sup d+2a/) and computation time T=O(dN/sup d/2-a/b), and achieve the AT/sup 2/ bound of AT/sup 2/=O(n/sup 2/b/sup 2/) for constant d, where n=N/sup d/ and O> Chaitali Chakrabarti, Joseph F. JáJá |
IEEE Trans. Computers | 1 |
| 1990 | A parallel algorithm for template matching on an SIMD mesh connected computerabstractAn efficient parallel algorithm to compute template matching of an N$0N input image with an M*M template on a single-instruction multiple-data (SIMD) mesh-connected computer with P processors is proposed. The input image is mapped into the processor array such that each processor stores N/sup 2//P data in the cyclic mode. The template values are circulated among the processors instead of being broadcast or stored in the processor memory. There is no movement of the intermediate results. The computation and the communication time complexity of the algorithm is O(M/sup 2/N/sup 2//P) for all P in the range M/sup 2/> Chaitali Chakrabarti, Joseph F. JáJá |
ICPR (2) | 1 |
| 1990 | Systolic Architectures for the Computation of the Discrete Hartley and the Discrete Cosine Transforms Based on Prime Factor DecompositionabstractTwo-dimensional systolic array implementations for computing the discrete Hartley transform (DHT) and the discrete cosine transform (DCT) when the transform size N is decomposable into mutually prime factors are proposed. The existing two-dimensional formulations for DHT and DCT are modified, and the corresponding algorithms are mapped into two-dimensional systolic arrays. The resulting architecture is fully pipelined with no control units. The hardware design is based on bit serial left to right MSB (most significant bit) to LSB (least significant bit) binary arithmetic.> Chaitali Chakrabarti, Joseph F. JáJá |
IEEE Trans. Computers | 1 |