EDBT 2026 Demo / reviewers in the wild / expert
Sujit Dey
dblp:d/SujitDey
· DBLP profile ↗
174ranked-venue papers
19as first author
9since 2021 · last 2026
0000-0001-9671-3950ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 119 · 17 first-author · 4 since 2021Computer networks · 30Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 4Artificial intelligence and machine learning · 2Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient SoC Power Estimation With Machine LearningabstractWe propose machine learning (ML) power, the first ML-based framework that: 1) accelerates end-to-end RTL power estimation by addressing both its key bottlenecks—RTL simulation and power-model evaluation—and 2) extends ML-based power models to system-on-chips (SoCs) with configurable intellectual property (IP) blocks. ML-Power builds models that predict power versus time traces of each SoC block using a very small subset of internal signals called power proxies. The framework is composed of three components: ML-based power proxy activity estimation (PACE), exploiting spatio-temporal correlations for power proxy selection and ML-based power model evaluation (SCOPE), and representative configuration selection using active learning (RECAL). PACE trains a sequence-to-sequence (seq2seq) ML model to translate a transaction-level execution trace into a cycle-level trace of the power proxy signals, allowing much faster simulation models to be used in place of RTL simulation. SCOPE selects power proxies and trains ML-based power models for each block within the SoC. RECAL enables ML-Power to handle SoCs with configurable intellectual property (IP) blocks by using active learning to select a small subset of representative configurations, which are then used to train a unified power model that generalizes across the entire design space. We evaluate ML-Power on an ARM SSE-300 SoC and two RISC-V-based SoCs. Compared to prior state-of-the-art ML-based power estimation frameworks, SCOPE trains power models$\sim 65\times $faster, picks 40% fewer proxies, and achieves 2% lower estimation error. For the two RISC-V SoCs with configurable IP blocks, ML-Power achieves less than 10% error in per-cycle power across the entire design space using only 10–14 training configurations. ML-Power also achieves$\sim 300\times $inference-time speedup over a commercial RTL power estimation tool with less than 7% error in per-cycle power estimates. Sujay Pandit, Sujit Dey, Anand Raghunathan |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2025 | ML-Power: Machine Learning based Power Estimation for SoCsabstractAccurate power estimation early in the design cycle is crucial for the design of power-efficient System-on-Chips (SoCs). Power estimation has been researched at various levels of abstraction, with a well-known trade-off between efficiency and accuracy. Recent work has shown great promise for machine learning (ML) techniques to advance the state-of-the-art in power estimation. We propose ML-Power, the first ML-based framework that can address both key bottlenecks involved in power estimation, viz. simulation and power model evaluation. ML-Power builds and uses models that estimate power in each SoC block using a very small subset of internal signals, known as power proxies. ML-Power consists of two key components, PACE and SCOPE. PACE trains a sequence-to-sequence ML model to translate a transaction-level execution trace into a cycle-level trace of the power proxy signals, allowing much faster simulation models to be used in place of RTL simulation. SCOPE selects power proxies and trains ML-based power models for blocks (IPs) within the SoC. SCOPE improves upon prior work in ML-based power modeling by exploiting spatio-temporal correlations among SoC blocks to minimize the number of power proxies while also improving estimation accuracy. We evaluate ML-Power for an ARM SSE-300 SoC and two RISC-V based SoCs. Compared to prior state-of-the-art ML-based power estimation frameworks, SCOPE trains power models ∼65× faster, picks 40% fewer proxies, and achieves 2% lower estimation error. ML-Power also achieves ∼300× speedup over a commercial RTL power estimation tool with less than 7% error in per-cycle power estimates. Sujay Pandit, Sujit Dey, Anand Raghunathan |
ISLPED | 2 |
| 2024 | SLEXNet: Adaptive Inference Using Slimmable Early Exit Neural NetworksabstractDeep learning is a proven method in many applications. However, it requires high computation resources and usually has a constant architecture. Mobile systems are good candidates to benefit from deep learning applications since they are closely integrated in people’s life. However, mobile systems experience varying conditions for the same reason. Constant deep learning architectures against varying resources cannot satisfy the requirements of the applications, so dynamic deep learning architectures are needed. In this work, we propose SLEXNet, a slimmable early exit neural network architecture. SLEXNet combines dynamic depth and width architectures to adapt to varying time and power conditions. Moreover, we propose a runtime scheduling algorithm that can estimate inference time and power consumption of SLEXNet variations on runtime. We train SLEXNet on real aerial drone images and implement the runtime on NVIDIA Jetson Orin. We show that our approach achieves significantly better responses to time and power requirements in varying conditions than baseline dynamic depth and width techniques in a wide range of experiments. Basar Kütükçü, Sabur Baidya, Sujit Dey |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2023 | Classification of Patient Recovery From COVID-19 Symptoms Using Consumer Wearables and Machine LearningabstractCurrent remote monitoring of COVID-19 patients relies on manual symptom reporting, which is highly dependent on patient compliance. In this research, we present a machine learning (ML)-based remote monitoring method to estimate patient recovery from COVID-19 symptoms using automatically collected wearable device data, instead of relying on manually collected symptom data. We deploy our remote monitoring system, namely eCOVID, in two COVID-19 telemedicine clinics. Our system utilizes a Garmin wearable and symptom tracker mobile app for data collection. The data consists of vitals, lifestyle, and symptom information which is fused into an online report for clinicians to review. Symptom data collected via our mobile app is used to label the recovery status of each patient daily. We propose a ML-based binary patient recovery classifier which uses wearable data to estimate whether a patient has recovered from COVID-19 symptoms. We evaluate our method using leave-one-subject-out (LOSO) cross-validation, and find that Random Forest (RF) is the top performing model. Our method achieves an F1-score of 0.88 when applying our RF-based model personalization technique using weighted bootstrap aggregation. Our results demonstrate that ML-assisted remote monitoring using automatically collected wearable data can supplement or be used in place of manual daily symptom tracking which relies on patient compliance. Jared Leitner, Alexander Behnke, Po-Han Chiang, Michele Ritter, Marlene Millen, Sujit Dey |
IEEE J. Biomed. Health Informatics | 6 |
| 2022 | Adaptive C-V2X Sidelink Communications for Vehicular Applications Beyond Safety MessagesabstractThe current Cellular Vehicle-to-Everything (C-V2X) Sidelink communication protocol provides a low latency interface for sharing short safety messages among Road-Side Units and vehicles. However, while its packets are broadcasted in the channel, the throughput is vulnerable to channel conditions and cannot meet the needs of the emerging connected and autonomous vehicles applications (e.g. vehicular fusion tasks), which require multimodal and multi-source sensor data sharing. In this work, we establish a C-V2X testbed on the campus of the University of California, San Diego, to study the feasibility of using C-V2X Sidelink communications for transmitting sensor data in real-time. We implement an end-to-end RGB sensor data (i.e. camera image frames) transmission mechanism on the C-V2X Sidelink testbed and explore the corresponding Quality of Service (QoS) characteristics under two configurable link-level parameters, the Modulation and Coding Scheme (MCS) and packet size. We then propose a cross-layer predictive and adaptive framework which adjusts, in real-time, the MCS and packet size settings based on side-channel information to optimize the QoS for image frame transmission. The real-world trace-driven emulation shows that the proposed policy improves the average frame goodput performance by 28% compared to fixed configuration policies that are used in current Sidelink communications. Yu-Jen Ku, Bryse Flowers, Samuel Thornton, Sabur Baidya, Sujit Dey |
VTC Spring | 5 |
| 2022 | Contention Grading and Adaptive Model Selection for Machine Vision in Embedded SystemsabstractReal-time machine vision applications running on resource-constrained embedded systems face challenges for maintaining performance. An especially challenging scenario arises when multiple applications execute at the same time, creating contention for the computational resources of the system. This contention results in increase in inference delay of the machine vision applications, which can be unacceptable for time-critical tasks. To address this challenge, we propose an adaptive model selection framework that mitigates the impact of system contention and prevents unexpected increases in inference delay by trading off the application accuracy minimally. The framework has two parts, which are performed pre-deployment and at runtime. The pre-deployment part profiles the system for contention in a black-box manner and produces a model set that is specifically optimized for the contention levels observed in the system. The runtime part predicts the inference delays of each model considering the system contention and selects the best model according to the predictions for each frame. Compared to a fixed individual model with similar accuracy, our framework improves the performance by significantly reducing the inference delay violations against a specified threshold. We implement our framework on the Nvidia Jetson TX2 platform and show that our approach achieves greater than 20% reductions in delay violations over the individual baseline models. Basar Kütükçü, Sabur Baidya, Anand Raghunathan, Sujit Dey |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2022 | Personalized Blood Pressure Estimation Using Photoplethysmography: A Transfer Learning ApproachabstractIn this paper, we present a personalized deep learning approach to estimate blood pressure (BP) using the photoplethysmogram (PPG) signal. We propose a hybrid neural network architecture consisting of convolutional, recurrent, and fully connected layers that operates directly on the raw PPG time series and provides BP estimation every 5 seconds. To address the problem of limited personal PPG and BP data for individuals, we propose a transfer learning technique that personalizes specific layers of a network pre-trained with abundant data from other patients. We use the MIMIC III database which contains PPG and continuous BP data measured invasively via an arterial catheter to develop and analyze our approach. Our transfer learning technique, namely BP-CRNN-Transfer, achieves a mean absolute error (MAE) of 3.52 and 2.20 mmHg for SBP and DBP estimation, respectively, outperforming existing methods. Our approach satisfies both the BHS and AAMI blood pressure measurement standards for SBP and DBP. Moreover, our results demonstrate that as little as 50 data samples per person are required to train accurate personalized models. We carry out Bland-Altman and correlation analysis to compare our method to the invasive arterial catheter, which is the gold-standard BP measurement method. Jared Leitner, Po-Han Chiang, Sujit Dey |
IEEE J. Biomed. Health Informatics | 3 |
| 2021 | Low Overhead Codebook Design for mmWave Roadside Units Placed at Smart IntersectionsabstractIn order to meet the high data rate requirements of emerging roadway use cases, mmWave vehicular communications will be needed. This work studies the ability of vehicles to communicate with a Roadside Unit (RSU) placed at an intersection. Practical mmWave radios utilize a codebook, a discrete set of analog beams, that is periodically searched during runtime to find the optimal beam to use for each receiver. This search creates overhead as the wireless channel is not used for communication while this beam search is happening. This work focuses on reducing the overhead of beam training by optimizing the site-specific codebook design of a RSU. Owing to the sparsity of the mmWave channel and the user distribution for vehicles, it is found that 85% of beams can be removed from the codebook with zero-impact. By carefully selecting the usage of wide beams the codebook size can be further reduced to just 64 beams while still providing omni-directional coverage for an intersection. Other research thrusts have focused on attempting to augment or remove beam training entirely; however, this necessitates a change to the PHY layer. Codebook optimization achieves approximately 80% of the communications performance that would be achieved if beam training overhead could be completely removed while only requiring a radio configuration update. Thus, this work finds that today’s commercial mmWave radios are sufficient for deployments in RSUs. To validate the proposed codebook optimization algorithm, a detailed mmWave ray tracing framework that encompasses 3D environmental information and material properties of reflectors is developed. Bryse Flowers, Xinyu Zhang 0003, Sujit Dey |
PIMRC | 3 |
| 2021 | Predictive Adaptive Streaming to Enable Mobile 360-Degree and VR ExperiencesabstractAs 360-degree videos and virtual reality (VR) applications become popular for consumer and enterprise use cases, the desire to enable truly mobile experiences also increases. Delivering 360-degree videos and cloud/edge-based VR applications require ultra-high bandwidth and ultra-low latency[1], challenging to achieve with mobile networks. A common approach to reduce bandwidth is streaming only the field of view (FOV). However, extracting and transmitting the FOV in response to user head motion can add high latency, adversely affecting user experience. In this paper, we propose a predictive adaptive streaming approach, where the predicted view with high predictive probability is adaptively encoded in relatively high quality according to bandwidth conditions and transmitted in advance, leading to a simultaneous reduction in bandwidth and latency. The predictive adaptive streaming method is based on a deep-learning-based viewpoint prediction model we develop, which uses past head motions to predict where a user will be looking in the 360-degree view. Using a very large dataset consisting of head motion traces from over 36,000 viewers for nineteen 360-degree/VR videos, we validate the ability of our predictive adaptive streaming method to offer high-quality view while simultaneously significantly reducing bandwidth. Xueshi Hou, Sujit Dey, Jianzhong Zhang 0002, Madhukar Budagavi |
IEEE Trans. Multim. | 2 |
| 2020 | Vehicular and Edge Computing for Emerging Connected and Autonomous Vehicle ApplicationsabstractEmerging connected and autonomous vehicles involve complex applications requiring not only optimal computing resource allocations but also efficient computing architectures. In this paper, we unfold the critical performance metrics required for emerging vehicular computing applications and show with preliminary experimental results, how optimal choices can be made to satisfy the static and dynamic computing requirements in terms of the performance metrics. We also discuss the feasibility of edge computing architectures for vehicular computing and show tradeoffs for different offloading strategies. The paper shows directions for light weight, high performance and low power computing paradigms, architectures and design-space exploration tools to satisfy evolving applications and requirements for connected and autonomous vehicles. Sabur Baidya, Yu-Jen Ku, Hengyu Zhao, Jishen Zhao, Sujit Dey |
DAC | 5 |
| 2020 | X-Array: approximating omnidirectional millimeter-wave coverage using an array of phased arraysabstractMillimeter-wave (mmWave) networks are conventionally considered to bear a fundamental coverage limitation, due to the directional beams and limited field-of-view (FoV) of the phased array antennas. In this paper, we explore an array of phased arrays (APA) architecture, which aggregates co-located phased arrays with complementary FoVs to approximate WiFi-like omni-directional coverage. We found that straightforwardly activating all the arrays may even hamper network performance. To fully exploit the APA's potential, we propose X-Array, which jointly selects the arrays and beams, and applies a dynamic co-phasing mechanism to ensure different arrays' signals enhance each other. X-Array also incorporates a link recovery mechanism to identify alternative arrays/beams that can efficiently recover the link from outage. We have implemented X-Array on a commodity 802.11ad APA radio. Our experiments demonstrate that X-Array can approach omni-directional coverage and maintain high performance in spite of link dynamics. Jingqi Huang, Xinyu Zhang 0003, Hyoil Kim, Sujit Dey |
MobiCom | 5 |
| 2019 | Head and Body Motion Prediction to Enable Mobile VR Experiences with Low LatencyabstractAs virtual reality (VR) applications become popular, the desire to enable high-quality, lightweight and mobile VR leads to various edge/cloud-based techniques. This paper introduces a predictive pre-rendering approach to address the ultra-low latency challenge in edge/cloud-based six Degrees of Freedom (6DoF) VR. Compared to 360-degree videos and 3DoF (head motion only) VR, 6DoF VR supports both head and body motions, thus not only viewing direction, but also viewing position changes. In our approach, the predictive view is rendered in advance based on the predicted viewing direction and position, leading to a reduction in latency. The key to achieving this efficient predictive pre-rendering approach is to predict the head and body motion accurately using past head and body motion traces. We develop a deep learning-based model and validate its ability using a dataset of over 840,000 samples for head and body motion. Xueshi Hou, Jianzhong Zhang 0002, Madhukar Budagavi, Sujit Dey |
GLOBECOM | 4 |
| 2019 | Personalized Blood Pressure Estimation using Photoplethysmography and Wavelet DecompositionabstractBlood pressure (BP) is the most important indicator of cardiovascular diseases. Traditional cuff-based methods for measuring BP require manual intervention and time. These methods may lead to inaccurate measurements and are not practical for continuous BP monitoring, which is crucial for detecting abnormal BP fluctuations. In this study, we propose a personalized machine learning model to estimate BP using his/her previous BPs and the photoplethysmogram (PPG) signal, the simplest and most popular tool for non-invasive diagnosis. To best utilize the information contained in the PPG signal, we propose to apply wavelet decomposition to extract features from the PPG signal. The arterial blood pressure (ABP) time series is processed with an exponentially weighted moving average (EWMA) and a peak detection technique to derive the SBP, DBP, and their corresponding trends. Finally, a random forest model is used to construct a predictive model based on these features. The MIMIC dataset is used for analysis and comparison with other BP estimation methods. Our experimental results demonstrate the proposed approach has smaller estimation error than existing methods, with mean average errors (MAE) for SBP and DBP equal to 3.43 and 1.73, respectively. Jared Leitner, Po-Han Chiang, Sujit Dey |
HealthCom | 3 |
| 2019 | Sustainable Vehicular Edge Computing Using Local and Solar-Powered Roadside Unit ResourcesabstractIn this paper, we explore a sustainable solution to growing vehicular computing and communication needs by utilizing edge computing and communication resources of a network of Solar- powered Roadside Units (SRSUs), along with any vehicular computing resources available. The solution ensures no additional grid energy expended while minimizing any QoS loss for the vehicle users (VUs). An SRSU consists of a small cell base station (SBS) and a road-edge computing (REC) node, is powered by a low-cost solar system. VUs can offload their vehicular application tasks to SRSUs to receive high throughput and low latency services. However, the limited capacity of solar energy, REC computing, and bandwidth resources may cause service disruption and affect the Quality of Service (QoS) that VUs receive. To minimize such QoS loss, we formulate a dynamic offloading QoS loss minimization problem, where the different subtasks of a VU application are optimally executed either locally using the VU computing resource or remotely using the REC resource, considering the energy, computing and bandwidth constraints of the SRSU network. We then propose to solve it by a heuristic algorithm which jointly makes the optimal user association, subtask offloading, and SRSU resource allocation decisions. To evaluate the proposed algorithm, we build a simulation framework consisting of a dense SRSU network using real-world solar energy generation and urban vehicular traffic data. The simulation results show that our proposed approach can significantly reduce QoS loss compared to other best effort strategies. Yu-Jen Ku, Sujit Dey |
VTC Fall | 2 |
| 2019 | User performance evaluation and real-time guidance in cloud-based physical therapy monitoring and guidance system
Wenchuan Wei, Yao Lu 0006, Eric Rhoden, Sujit Dey |
Multim. Tools Appl. | 4 |
| 2018 | Personalized Effect of Health Behavior on Blood Pressure: Machine Learning Based Prediction and RecommendationabstractBlood pressure (BP) is one of the most important indicator of human health. In this paper, we investigate the relationship between BP and health behavior (e.g. sleep and exercise). Using the data collected from off-the-shelf wearable devices and wireless home BP monitors, we propose a data driven personalized model to predict daily BP level and provide actionable insight into health behavior and daily BP. In the proposed machine learning model using Random Forest (RF), trend and periodicity features of BP time-series are extracted to improve prediction. To further enhance the performance of the prediction model, we propose RF with Feature Selection (RFFS), which performs RF-based feature selection to filter out unnecessary features. Our experimental results demonstrate that the proposed approach is robust to different individuals and has smaller prediction error than existing methods. We also validate the effectiveness of personalized recommendation of health behavior generated by RFFS model. Po-Han Chiang, Sujit Dey |
HealthCom | 2 |
| 2018 | Quality of Service Optimization for Vehicular Edge Computing with Solar-Powered Road Side UnitsabstractThis paper shows the viability of Solar-powered Road Side Units (SRSU), consisting of small cell base stations and Mobile Edge Computing (MEC) servers, and powered solely by solar panels with battery, to provide connected vehicles with a low- latency, easy-to-deploy and energy-efficient communication and edge computing infrastructure. However, SRSU may entail a high risk of power deficiency, leading to severe Quality of Service (QoS) loss due to spatial and temporal fluctuation of solar power generation. Meanwhile, the data traffic demand also varies with space and time. The mismatch between solar power generation and SRSU power consumption makes optimal use of solar power challenging. In this paper, we model the above problem with three sub-problems, the SRSU power consumption minimization problem, the temporal energy balancing problem and spatial energy balancing problem. Three algorithms are proposed to solve the above sub-problems, and they together provide a complete joint battery charging and user association control algorithm to minimize the QoS loss under delay constraint of the computing tasks. Results with a simulated urban environment using actual solar irradiance and vehicular traffic data demonstrates that the proposed solution reduces the QoS loss significantly compared to greedy approaches. Yu-Jen Ku, Po-Han Chiang, Sujit Dey |
ICCCN | 3 |
| 2018 | Novel Hybrid-Cast Approach to Reduce Bandwidth and Latency for Cloud-Based Virtual SpaceabstractIn this article, we explore the possibility of enabling cloud-based virtual space applications for better computational scalability and easy access from any end device, including future lightweight wireless head-mounted displays. In particular, we investigate virtual space applications such as virtual classroom and virtual gallery, in which the scenes and activities are rendered in the cloud, with multiple views captured and streamed to each end device. A key challenge is the high bandwidth requirement to stream all the user views, leading to high operational cost and potential large delay in a bandwidth-restricted wireless network. We propose a novel hybrid-cast approach to save bandwidth in a multi-user streaming scenario. We identify and broadcast the common pixels shared by multiple users, while unicasting the residual pixels for each user. We formulate the problem of minimizing the total bitrate needed to transmit the user views using hybrid-casting and describe our approach. A common view extraction approach and a smart grouping algorithm are proposed and developed to achieve our hybrid-cast approach. Simulation results show that the hybrid-cast approach can significantly reduce total bitrate by up to 55% and avoid congestion-related latency, compared to traditional cloud-based approach of transmitting all the views as individual unicast streams, hence addressing the bandwidth challenges of the cloud, with additional benefits in cost and delay. Xueshi Hou, Yao Lu 0006, Sujit Dey |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2017 | Wireless VR/AR with Edge/Cloud ComputingabstractTriggered by several head-mounted display (HMD) devices that have come to the market recently, such as Oculus Rift, HTC Vive, and Samsung Gear VR, significant interest has developed in virtual reality (VR) systems, experiences and applications. However, the current HMD devices are still very heavy and large, negatively affecting user experience. Moreover, current VR approaches perform rendering locally either on a mobile device tethered to an HMD, or on a computer/console tethered to the HMD. In this paper, we discuss how to enable a truly portable and mobile VR experience, with light weight VR glasses wirelessly connecting with edge/cloud computing devices that perform the rendering remotely. We investigate the challenges associated with enabling the new wireless VR approach with edge/cloud computing with different application scenarios that we implement. Specifically, we analyze the challenging bitrate and latency requirements to enable wireless VR, and investigate several possible solutions. Xueshi Hou, Yao Lu 0006, Sujit Dey |
ICCCN | 3 |
| 2017 | Asymmetric and selective object rendering for optimized Cloud Mobile 3D Display Gaming user experience
Yao Lu 0006, Yao Liu 0008, Sujit Dey |
Multim. Tools Appl. | 3 |
| 2016 | A Novel Hyper-Cast Approach to Enable Cloud-Based Virtual Classroom ApplicationsabstractIn this paper, we explore the possibility of enabling cloud-based virtual classroom applications providing the advantages of computational scalability and access from any end device. In particular, we investigate a virtual classroom application in which the classroom including teacher, students and activities are rendered on the cloud, with each student view captured and streamed to students' end devices. By identifying that many student views may share common pixels, we design a novel hyper-cast system so that common pixels can be transmitted by a single broadcast stream while the residual pixels for individual student view can be transmitted by unicast in cellular networks. The hyper-cast approach can significantly reduce total bitrate needed, and therefore decrease cloud cost and cellular bandwidth. Simulation results show that the proposed hyper-cast technique can save up to 44.6% bitrate compared to traditional cloud-based approach. Xueshi Hou, Yao Lu 0006, Sujit Dey |
ISM | 3 |
| 2016 | Enhancing Mobile Video Capacity and Quality Using Rate Adaptation, RAN Caching and ProcessingabstractAdaptive Bit Rate (ABR) streaming has become a popular video delivery technique, credited with improving Quality of Experience (QoE) of videos delivered on wireless networks. Recent independent research reveals video caching in the Radio Access Network (RAN) holds promise for increasing the network capacity and improving video QoE. In this paper, we investigate opportunities and challenges of combining the advantages of ABR and RAN caching to increase the video capacity and QoE of the wireless networks. While each ABR video is divided into multiple chunks that can be requested at different bit rates, a cache hit requires the presence of a specific chunk at a desired bit rate, making ABR-aware RAN caching challenging. To address this without having to cache all bit rate versions of a video, we propose adding limited processing capacity to each RAN cache. This enables transrating a higher rate version that may be available in the cache, to satisfy a request for a lower rate version, and joint caching and processing policies that leverage the backhaul, caching, and processing resources most effectively, thereby maximizing video capacity of the network. We also propose a novel rate adaptation algorithm that uses video characteristics to simultaneously change the video encoding and transmission rate. The results of extensive statistical simulations demonstrate the effectiveness of our approaches in achieving significant capacity gain over ABR or RAN caching alone, as well as other ways of enabling ABR-aware RAN caching, while improving video QoE. Hasti A. Pedersen, Sujit Dey |
IEEE/ACM Trans. Netw. | 2 |
| 2016 | Emulation-Based Analysis of System-on-Chip Performance Under VariationsabstractThe scaling of integrated circuits into the nanometer regime has led to variations emerging as a primary design concern. Most efforts in the area of variation-tolerant design have focused on the physical, circuit, and logic levels of abstraction. However, inevitable increases in the magnitude of variations with scaling have elevated them to a design concern that must be addressed starting at the system level. We address the problem of analyzing the performance of system-onchip (SoC) architectures in the presence of variations. A modern SoC is a complex ensemble of components that are organized into multiple voltage and frequency domains or islands. The impact of variations on the clock frequencies of individual SoC components may be analyzed using existing tools, such as circuit-level statistical timing analysis. However, the key challenge that needs to be addressed is how to translate these component-level clock frequency distributions into a system-level performance distribution. This task is particularly complex and challenging due to the interdependences between components' execution, indirect effects of shared resources, and interactions between multiple system-level execution paths. We argue that an accurate variation-aware performance analysis requires Monte Carlo-based repeated system execution. We describe a framework variability emulation for SoC performance analysis (VESPA)-that leverages emulation to significantly speed up the performance analysis without sacrificing the generality and accuracy achieved by Monte Carlo-based simulation. We further improve the efficiency of VESPA by utilizing correlated sampling to reduce the number of samples needed for Monte Carlo simulations. We demonstrate the utility of VESPA by applying it to design variation-tolerant architectures for three example SoCs. Our experiments show the performance improvements of ~180× compared with the state-of-the-art hardware-software cosimulation tools and also underscore the potential of VESPA to enable variation-aware design and exploration at the system level. Vivek Joy Kozhikkottu, Rangharajan Venkatesan, Anand Raghunathan, Sujit Dey |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2015 | Motion data alignment for real-time guidance in avatar based physical therapy training systemabstractIn this paper, we propose an Avatar based Virtual Reality user training system that efficiently trains users in performing a variety of activities using a pre-recorded avatar. To evaluate and monitor the user's adherence to the avatar's instructions, the system compares the user's motion data against the avatar's motion data, with the latter established as the ground truth dataset. Unfortunately, human reaction delay may cause the motion sequences between the user and the avatar to be misaligned. Consequently, to enable accurate comparison, we analyze four signal processing time delay estimation methods—an existing method and three proposed methods—to align the motion sequences between the user and pre-recorded avatar, allowing the correct frames to be compared. Our experiments demonstrate that the proposed methods perform better data alignment than the existing method and the fourth method, which employs a novel spatial-temporal segmentation algorithm, has the highest potential to be the optimal delay estimation approach. Further, to provide real-time guidance to the user, we determine a unique tolerance threshold for each activity such that a user accuracy value below the threshold value prompts real-time guidance to correct the user and an accuracy value above the threshold is tolerated. We perform an experiment with the assistance of a physical trainer and use the experimental data to design a histogram-based method using Bayesian decision theory to determine the threshold values. Dennis Shen, Yao Lu 0006, Sujit Dey |
HealthCom | 3 |
| 2015 | Power-efficient base station operation through user QoS-aware adaptive RF chain switching techniqueabstractIn this paper, we develop a user Quality of Service (QoS) aware adaptive Radio Frequency (RF) chain switching technique that dynamically adapts the number of active RF chains to minimize the total power consumption of the base station (BS). Specifically, we first formulate an optimization problem to minimize the total power consumption under the throughput and Block Error Rate (BLER) requirements of the users of the BS. Due to the prohibitive complexity of exhaustive algorithm achieving an optimal solution, we propose a practically implementable heuristic algorithm, which jointly optimizes the number of active RF chains, frequency and time resources. Simulation results demonstrate that the proposed algorithm can achieve significant savings in total power consumption of the BSs compared to the conventional static All-On scheme wherein the activity of RF chains is not adapted. Ranjini Guruprasad, Kyuho Son, Sujit Dey |
ICC | 3 |
| 2015 | Optimizing Cloud Mobile 3D Display Gaming user experience by asymmetric object of interest renderingabstractThe growing popularity of auto-stereoscopic 3D displays for mobile devices, together with ubiquitous wireless networks, have fueled an increasing user expectation for rich 3D mobile multimedia experiences, including 3D display gaming. However, rendering 3D games on mobile devices requires high computational power and battery and thus may restrict users from enjoying true 3D experience for a long time. In this paper, we explore the possibility of developing Cloud Mobile 3D Display Gaming, where the 3D video rendering and encoding are performed on cloud servers, with the resulting 3D video streamed wirelessly to mobile devices with 3D displays. However, with the significantly higher bitrate requirement for 3D video, ensuring user experience may be a challenge considering the bandwidth constraints and fluctuations of mobile networks. In this paper, we propose a novel asymmetric Object of Interest (OOI) rendering approach, which adapts the rendering richness of different objects according to their importance in order to reduce the video encoding bitrate needed while maintaining a satisfactory video quality, thereby making it easier to transmit the 3D game video over wireless network. Specifically, we first develop a model to quantitatively measure the user experience by different OOI rendering settings. We also develop a model to relate the bitrate of the resulting video with the changes of different OOI Rendering settings. We further propose an optimization algorithm which uses the above two models to automatically decide the optimal OOI rendering settings for left view and right view to ensure the best user experience given the network conditions. Experiments conducted using real 4G-LTE network profiles on commercial cloud service demonstrate the improvement in user experience when the proposed optimization algorithm is applied. Yao Lu 0006, Yao Liu 0008, Sujit Dey |
ICC | 3 |
| 2015 | A Joint Asymmetric Graphics Rendering and Video Encoding Approach for Optimizing Cloud Mobile 3D Display Gaming User ExperienceabstractWith the development and deployment of ubiquitous wireless network together with the growing popularity of mobile auto-stereoscopic 3D displays, more and more applications have been developed to enable rich 3D mobile multimedia experiences, including 3D display gaming. Simultaneously, with the emergence of cloud computing, more mobile applications are being developed to take advantage of the elastic cloud resources. In this paper, we explore the possibility of developing Cloud Mobile 3D Display Gaming, where the 3D video rendering and encoding are performed on cloud servers, with the resulting 3D video streamed to mobile devices with 3D displays through wireless network. However, with the significantly higher bitrate requirement for 3D videos, ensuring user experience may be a challenge considering the bandwidth constraints of mobile networks. In order to address this challenge, different techniques have been proposed including asymmetric graphics rendering and asymmetric video encoding. In this paper, for the first time, we propose a joint asymmetric graphics rendering and video encoding approach, where both the encoding quality and rendering richness of left view and right view are asymmetric, to enhance the user experience of the cloud mobile 3D display gaming system. Specifically, we first conduct extensive user studies to develop a user experience model that takes into account both video encoding impairment and graphics rendering impairment. We also develop a model to relate the bitrate of the resulting video with the video encoding settings and graphics rendering settings. Finally we propose an optimization algorithm that can automatically choose the video encoding settings and graphics rendering settings for left view and right view to ensure the best user experience given the network conditions. Experiments conducted using real 4G-LTE network profiles on commercial cloud service demonstrate the improvement in user experience when the proposed optimization algorithm is applied. Yao Liu 0008, Sujit Dey |
ISM | 3 |
| 2015 | Renewable energy-aware video download in cellular networksabstractThe use of renewable energy (RE) sources is a promising solution to reduce grid power consumption and carbon dioxide emissions (CO2e) of cellular networks. However, the benefit of utilizing RE is limited by its highly intermittent and unreliable nature leading to mismatch between generation and base station (BS) needs, resulting in low savings in grid power. To address the above challenges, we propose a novel base station time resource allocation technique using data storage at user equipments (UEs) to optimally utilize renewable energy to reduce grid power consumption. The proposed approach transforms the surplus RE (in excess of the BS power requirements) to excess data delivered to the users and stored in UE data storages, to draw from in deficit periods (when RE generation is lesser than BS requirements) to reduce grid power. Though the proposed approach is applicable to any application that utilizes UE data buffer, we formulate the problem for mobile video and propose a algorithm for RE aware BS resource allocation during mobile video download, as the latter will dominate wireless traffic and hence BS power consumption. Our experimental results using sample solar and BS utilization traces demonstrate the ability of the proposed approach to reduce grid power consumption by increasing solar power utilization while satisfying user QoS requirements. Po-Han Chiang, Ranjini Guruprasad, Sujit Dey |
PIMRC | 3 |
| 2015 | Construction and evaluation of ontological tag trees
Chetan Kumar Verma, Vijay Mahadevan, Nikhil Rasiwasia, Gaurav Aggarwal, Alejandro Jaimes, Sujit Dey |
Expert Syst. Appl. | 7 |
| 2015 | Enhancing Video Encoding for Cloud Gaming Using Rendering InformationabstractCloud gaming allows games to be rendered on the cloud server and allows the rendered videos to be encoded and streamed in real time to the player's devices. Compared with other video streaming applications, cloud gaming offers a unique opportunity to enhance the video encoding process by exploiting rendering information. In this paper, we propose two techniques to improve cloud gaming video encoding, aiming at enhancing the perceived video quality and reducing the computational complexity, respectively. First, we develop a rendering-based prioritized encoding technique to improve the perceived game video quality according to network bandwidth constraints. We first propose a technique to generate a macroblock (MB)-level saliency map for every game video frame using rendering information. Furthermore, based on such a saliency map, a prioritized rate allocation scheme is proposed to dynamically adjust the value of quantization parameter of each MB. The experimental results indicate that the perceptual quality can be greatly improved using the proposed technique. We also develop a rendering-based encoding acceleration technique that utilizes rendering information to reduce the computational complexity of video encoding. This technique mainly consists of two parts. First, we propose a method to directly calculate the motion vectors (MVs) without employing the compute intensive motion search procedure. Second, based on the computed MVs, we propose a fast mode selection algorithm to reduce the number of candidate modes of each MB. The experimental results show that the proposed technique can achieve more than 42% saving in encoding time with very limited degradation in video quality. Yao Liu 0008, Sujit Dey, Yao Lu 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2015 | Battery Aware Video Delivery Techniques Using Rate Adaptation and Base Station ReconfigurationabstractWith mobile video increasingly becoming an important driver of mobile device usage, the battery consumption of mobile devices will be dominated by video delivery and playback . In this paper, we develop battery efficient video download techniques that vary video download rate dynamically, including stopping video download at times, depending on mobile device buffer levels and the channel conditions experienced, to maximize battery life while ensuring no degradation in user experience. The proposed dynamic download rate adaptation techniques enable the base station to adapt the MIMO transceiver configurations to reduce battery load required by MIMO components on the mobile device. In order to further enhance battery life, we propose to utilize video bit rate adaptation, in addition to download rate adaptation and MIMO reconfiguration. The proposed battery aware bit rate adaptation techniques take into account the mobile device battery and buffer levels, and network load and channel conditions experienced, to maximize battery lifetime (hence video viewing time) while ensuring desired level of video experience (measured in terms of video quality and stalls experienced). We propose a new metric termed “video experience longevity (VEL)” which quantifies the performance of the proposed bit rate adaptation techniques in terms of video viewing time and video experience. Extensive experiments conducted under variable channel conditions and network load demonstrate that the proposed battery aware video delivery techniques can significantly outperform other video delivery techniques in terms of battery lifetime and VEL metric (for bit rate adaptation techniques) while ensuring desired level of video experience. Ranjini Guruprasad, Sujit Dey |
IEEE Trans. Multim. | 2 |
| 2015 | An Application Adaptation Approach to Mitigate the Impact of Dynamic Thermal Management on Video EncodingabstractDue to limitations of cooling methods such as using fan and heat sink, dynamic thermal management (DTM) is being widely adopted to manage the temperature of computing systems. However, application of DTM can reduce the system performance and thereby affect the quality of real-time applications. Real-time video encoding, which has high computational need and hard deadlines, is a commonly used application that can be severely affected by the usage of DTM. We study the effect of DTM on a widely used H.264 video encoder and formulate a multidimensional optimization problem to maximize video quality and minimize bit rate while ensuring that the video encoder can run in real time in spite of DTM effects. We model the effects of adapting encoding parameters on video quality, bit rate, and encoder speed. We propose a dynamic application adaptation method to efficiently solve the optimization problem by optimally adapting the encoding parameters in response to DTM effects. In addition, we show that the proposed dynamic application adaptation method would reduce the need for cooling methods such as forced convection cooling. We implement the proposed approach on an Intel® Core™ 2 Duo platform where dynamic voltage and frequency scaling (DVFS) is used for DTM. Our measurements with several videos reveal that when DTM is applied, the video quality is affected significantly. However, using the proposed adaptation algorithm, the encoder can run in real time, and the quality loss is minimized with only a marginal increase in the bit rate. Ali Mirtar, Sujit Dey, Anand Raghunathan |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2015 | Joint Work and Voltage/Frequency Scaling for Quality-Optimized Dynamic Thermal ManagementabstractDynamic thermal management (DTM) is commonly used to ensure reliable and safe operation in modern computing systems. DTM techniques are based on slowing down or shutting down parts of a system; hence, they effectively reduce system performance and thereby adversely impact applications. In this paper, we focus on real-time applications in which degradation in performance translates to a loss in application quality, and address the problem of quality-optimized DTM, wherein the objective of DTM is to satisfy specified temperature constraints while optimizing application quality metrics. We first introduce a new DTM method called dynamic work scaling (DWS), which is based on modulating an application's computational requirements. Next, we observe that application quality and platform temperature are effectively determined by two key parameters, viz., the application's computational requirement and the platform's computing capacity, and formulate the relationship between them. Finally, we propose a quality-optimized DTM based on joint dynamic work and voltage/frequency scaling (DWVFS). We have implemented the proposed DTM technique and evaluated it for two applications: 1) H.264 video encoding and 2) turbo decoding. Our results demonstrate that DWVFS can provide superior results in terms of application quality compared with both DVFS and DWS-based DTM at identical temperature constraints. Ali Mirtar, Sujit Dey, Anand Raghunathan |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2014 | Variation Aware Cache Partitioning for Multithreaded ProgramsabstractMultithreaded programs are commonly written and optimized for homogeneous multi-core processors assuming equal performance from all the cores. This assumption greatly simplifies the partitioning and balancing of an application's workload across threads; however, it no longer holds when the frequencies of the cores differ due to within-die variations, leading to a degradation in performance. We observe that, in addition to the frequency of the core that it executes on, the performance of a thread is also dependent on the share of shared system resources, such as last-level cache, that it receives. We propose variation-aware cache partitioning as an approach to redress the variation-induced imbalance in the execution times of threads, thereby improving the performance of multi-threaded programs. We discuss the challenges involved in realizing our proposal, including synchronization (e.g., barriers) across threads, which results in faster threads being limited by slower threads, the complex and non-linear relationship between a thread's performance and the cache capacity allocated to it, and the fact that different program phases, can respond quite differently to varying cache capacity. We propose a runtime scheme to perform spatio-temporal cache partitioning while considering both chip characteristics (frequency variations) and program characteristics. We evaluate the proposed technique by applying it to an ensemble of variation-impacted multi-cores executing multi-threaded programs from the PARSEC and SPEC-OMP suites, and demonstrate that it results in an average performance improvement of 15% by mitigating the impact of frequency variations. Vivek Joy Kozhikkottu, Abhisek Pan, Vijay S. Pai, Sujit Dey, Anand Raghunathan |
DAC | 4 |
| 2014 | Variation tolerant design of a vector processor for recognition, mining and synthesisabstractVariations have emerged as one of the most significant challenges facing the design of integrated circuits in nanoscale technologies. As a consequence, variation tolerant design has become essential at all levels of design abstraction. Vivek Joy Kozhikkottu, Swagath Venkataramani, Sujit Dey, Anand Raghunathan |
ISLPED | 3 |
| 2014 | Mobile device video caching to improve video qoe and cellular network capacityabstractAs the video resolution, storage, and rendering capabilities of mobile devices improve, consumers tend to watch more videos on their devices. To address this increasing demand, there is a need to improve the capacity of cellular networks while also improving video Quality of Experience (QoE). Ensuring video QoE via wireless links remains a challenge, not only due to the backhaul demand, but also due to the varying wireless channel conditions caused by fading, multiuser interference, and peak traffic loads. Here, we introduce a reactive Mobile Device Caching (rMDC) framework to improve video capacity (number of concurrent video viewing sessions) and QoE in the presence of the challenges explained above. Using rMDC, a mobile device caches videos reactively as it requests them, evicts the videos least likely to be requested according to its neighbors aggregate User Preference Profile (UPP), and sharing video contents using D2D communication. We use a discrete event statistical simulation framework to study the performance of rMDC. Our simulation results demonstrate the effectiveness of rMDC, along with UPP-based caching, in achieving higher capacity and better QoE compared to no mobile device caching. Hasti A. Pedersen, Sujit Dey |
MSWiM | 2 |
| 2014 | Video-Aware Scheduling and Caching in the Radio Access NetworkabstractIn this paper, we introduce distributed caching of videos at the base stations of the Radio Access Network (RAN) to significantly improve the video capacity and user experience of mobile networks. To ensure effectiveness of the massively distributed but relatively small-sized RAN caches, unlike Internet content delivery networks (CDNs) that can store millions of videos in a relatively few large-sized caches, we propose RAN-aware reactive and proactive caching policies that utilize User Preference Profiles (UPPs) of active users in a cell. Furthermore, we propose video-aware backhaul and wireless channel scheduling techniques that, in conjunction with edge caching, ensure maximizing the number of concurrent video sessions that can be supported by the end-to-end network while satisfying their initial delay requirements and minimize stalling. To evaluate our proposed techniques, we developed a statistical simulation framework using MATLAB and performed extensive simulations under various cache sizes, video popularity and UPP distributions, user dynamics, and wireless channel conditions. Our simulation results show that RAN caches using UPP-based caching policies, together with video-aware backhaul scheduling, can improve capacity by 300% compared to having no RAN caches, and by more than 50% compared to RAN caches using conventional caching policies. The results also demonstrate that using UPP-based RAN caches can significantly improve the probability that video requests experience low initial delays. In networks where the wireless channel bandwidth may be constrained, application of our video-aware wireless channel scheduler results in significantly (up to 250%) higher video capacity with very low stalling probability. Hasti Ahlehagh, Sujit Dey |
IEEE/ACM Trans. Netw. | 2 |
| 2013 | Adaptive Bit Rate capable video caching and schedulingabstractAdaptive Bit Rate Streaming (ABR) has become a popular video delivery technique, credited to improving the quality of delivered video on wireless networks. At the same time, recent research has shown video caching in the Radio Access Network (RAN) can be a promising way to increase capacity of the network while reducing video latency and improving video quality of experience. In this work, we investigate the opportunities and challenges of combining the advantages of ABR streaming and RAN caching to maximize the video capacity of wireless networks. Since with ABR, each video is divided into multiple segments, chunks, and each chunk can be requested at different bit rates, the caching requirements are very different and challenging; a cache hit will require not only the presence of a specific video chunk, but also the availability of the desired bit rate. One way to solve this problem is to cache all variants of a video, but this approach may significantly increase storage and backhaul requirements or reduce the number of unique videos that can be cached. In this paper, we introduce a framework consisting of rate adaptation algorithm, video caching and processing within the wireless cloud, with the aim to improve video capacity of the wireless network and satisfy or exceed QoE of each video request. To achieve this goal, we propose a new ABR algorithm, along with an ABR aware Least Recently Used (LRU) caching policy with Processing (ABRLRU-P) to support caching of video chunks with different bit rates. Using our MATLAB statistical simulation framework, we demonstrate a capacity improvement of up to 83% when ABR is used with RAN caching and processing using the ABR-LRU-P policy compared with using no RAN caching and video processing. Further, using ABR along with ABRLRU-P caching policy can improve the capacity by 68% compared with using ABR and a straightforward static LRU caching policy which fetches all video bit rate versions of a video chunk upon a cache miss. Hasti Ahlehagh, Sujit Dey |
WCNC | 2 |
| 2013 | Rate adaptation and base station reconfiguration for battery efficient video downloadabstractWith increasing usage of mobile devices to watch mobile video, the battery consumption of mobile devices will be dominated by video delivery and playback. In this paper, we develop battery efficient video delivery techniques that vary video transmission rate dynamically depending on battery and buffer levels of the mobile device, the channel conditions experienced, and video quality requirements while maximizing battery life and ensuring user experience. The proposed dynamic video rate adaptation techniques enable the base station to adapt the Multi Input Multi Output (MIMO) transceiver configurations to reduce battery current required by MIMO components on the mobile device. The proposed techniques also make the base station stop video transmission opportunistically, thereby eliminating the battery load imposed by the mobile MIMO components. Experiments conducted under various network conditions for various video download profiles show that more than 50% improvement in battery lifetime is possible in comparison to conventional video download techniques not aware of battery. Ranjini Guruprasad, Sujit Dey |
WCNC | 2 |
| 2013 | QoS-aware dynamic cell reconfiguration for energy conservation in cellular networksabstractGiven the significant energy consumption in operating base stations (BSs), improving their energy efficiency is an important problem in cellular networks. To this end, this paper proposes a novel framework, called DCR (dynamic cell reconfiguration) that dynamically adjust the set of active BSs and user association according to user traffic demand for energy conservation. In order to overcome prohibitive computational complexity in finding an optimal solution, we take an approach to design simple yet effective algorithms. We demonstrate that the proposed framework is not only computationally efficient but also can achieve the performance close to the optimum solution from an exhaustive search. Through simulations based on a real dataset of BS topology and utilization, we show that DCR can yield about a 30-40% reduction compared to the conventional static scheme where all BSs are always turned on. Kyuho Son, Santosh V. Nagaraj, Mahasweta Sarkar, Sujit Dey |
WCNC | 4 |
| 2013 | Fully Automated Learning for Application-Specific Web Video ClassificationabstractPersonalization applications such as content recommendations, product recommendations and advertisements, and social network related recommendations, can be quite beneficial for both service providers and users. Such applications need to understand user preferences in order to provide customized services. As user engagement with web videos has grown significantly, understanding user preferences based on videos viewed looks promising. The above requires ability to classify web videos into a set of categories appropriate for the personalization application. However, such categories may be substantially different from common categories like Sports, Music, Comedy, etc. used by video sharing websites, leading to lack of labeled training videos for such categories. In this paper, we study the feasibility and effectiveness of a fully automated framework to obtain training videos to enable classification of web videos to any arbitrary set of categories, as desired by the personalization application. We investigate the desired properties in training data that can lead to high performance of the trained classification models. We then develop an approach to identify and score keywords based on their suitability to retrieve training videos, with the desired properties, for the specified set of categories. Experimental results on several sets of categories demonstrate the ability of the proposed approach to obtain effective training data, and hence achieve high video classification performance. Chetan Kumar Verma, Sujit Dey |
Web Intelligence | 2 |
| 2013 | Adaptive Mobile Cloud Computing to Enable Rich Mobile Multimedia ApplicationsabstractWith worldwide shipments of smartphones (487.7 million) exceeding PCs (414.6 million including tablets) in 2011, and in the US alone, more users predicted to access the Internet from mobile devices than from PCs by 2015, clearly there is a desire to be able to use mobile devices and networks like we use PCs and wireline networks today. However, in spite of advances in the capabilities of mobile devices, a gap will continue to exist, and may even widen, with the requirements of rich multimedia applications. Mobile cloud computing can help bridge this gap, providing mobile applications the capabilities of cloud servers and storage together with the benefits of mobile devices and mobile connectivity, possibly enabling a new generation of truly ubiquitous multimedia applications on mobile devices: Cloud Mobile Media (CMM) applications. Shaoxuan Wang, Sujit Dey |
IEEE Trans. Multim. | 2 |
| 2012 | Recovery-based design for variation-tolerant SoCsabstractParameter variations have emerged as a significant threat to continued CMOS scaling in the nanometer regime. Due to increasing performance penalties associated with worst-case design, recovery based design has emerged as a promising approach for dealing with the impact of variations. Previous work has applied recovery based design at the circuit and micro-architecture levels of abstraction. In this work, we address the problem of designing variation-tolerant SoCs using the recovery based design paradigm. We demonstrate that a monolithic implementation of recovery based design fails to scale for large SoCs. We propose the concept of recovery islands, wherein each island consists of one or more SoC components that can recover independent of the rest of the SoC, and demonstrate how our proposal can be easily realized via minor changes to a traditional SoC design flow. We study the tradeoffs involved in applying recovery based design at the system level. We demonstrate that it is critical to account for (i) the inherent diversity of the error-voltage profiles among various components in an SoC, and (ii) the impact of error recovery in a component on overall system performance. We then propose a systematic recovery-based SoC design methodology that partitions a given SoC into recovery islands and also computes the optimal operating points for each island, taking into account the various system level trade-offs involved. We evaluate our framework on three different SoC designs, an 802.11b MAC processor, an MPEG encoder and a Wireless Video Capture system and demonstrate an average of 32% energy savings over conventional designs. Vivek Joy Kozhikkottu, Sujit Dey, Anand Raghunathan |
DAC | 2 |
| 2012 | Hierarchical video caching in wireless cloud: Approaches and algorithmsabstractWe had introduced video caching techniques in the Radio Access Network (RAN) in [1] as a way to reduce the need to bring requested videos from Internet CDNs, thus reducing overall backhaul traffic, improving video quality of experience and increasing network capacity to support more simultaneous video requests. In this paper, we investigate supplementing the resulting wireless cloud with a hierarchical caching scheme, where the gateways in the Core Network (CN) also have video caches. The hierarchical caching approach further improves network capacity by enabling multiple cell sites to share caches at higher levels of the hierarchy, thereby improving overall cache hit ratio, without increasing the total cache size used. In addition, we exploit hierarchical caching to better accommodate mobility, so that when a user with an active video session moves from one cell to a neighboring cell, it is likely that the video currently being downloaded is already in a cache within the RAN or CN network associated with the new cell. To achieve the goal of improving capacity and supporting mobility, we extend our User Preference Profile (UPP) based caching policies [1] to accommodate the hierarchical caching structure introduced in this paper. For all the videos that miss the cache in any layer of hierarchy, we propose a scheduling approach to allocate RAN and CN backhaul resources judiciously so as to maximize the capacity of the wireless network. We extend our discrete event statistical simulation framework developed in [1] to study the performance of the proposed hierarchical caching approach. Our simulation results show that using hierarchical caching can enhance cache hit ratio by 24% and network capacity by up to 45% compared to caching only in the RAN. Significant capacity gains are also observed when additionally considering user mobility. Hasti Ahlehagh, Sujit Dey |
ICC | 2 |
| 2012 | Wireless network aware cloud scheduler for scalable cloud mobile gamingabstractCloud Mobile Gaming (CMG) [1][10][11] has been proposed as an approach to enable rich Internet games on mobile devices, where the rendering of the games is performed on cloud servers, as opposed to on mobile devices. Though promising, the CMG approach may require significant cloud computing resources for the concurrent gaming sessions, and even more critically, significant bandwidth for delivering the rendered videos back to mobile devices, leading to high cloud costs, and questions regarding resource constrained wireless networks. This paper addresses the problem of making the CMG approach scalable and economically feasible by proposing a novel Wireless Cloud Scheduler (WCS), which can increase the number of simultaneous CMG sessions that can be supported while ensuring Mobile Gaming User Experience (MGUE)[1] with the available wireless network resources, while minimizing the cloud service cost incurred by the CMG provider. Unlike conventional network schedulers, WCS considers simultaneously the constraints of the wireless networks that may be available to each CMG user, including cellular and WiFi, as well as the cost of available cloud resources, while scheduling the most optimal wireless link and cloud server for each CMG session. To further enhance the performance of WCS, we also propose a joint scheduling-adaptation algorithm, that can systematically leverage adaptation techniques introduced in [10][11] to adapt the communication needs of in-service users if the available wireless network bandwidth is not sufficient for a new CMG user. Our simulation results demonstrate that the use of WCS, and the joint scheduling-adaptation algorithm, can significantly improve the performance of the CMG approach, increasing the number of simultaneous CMG sessions that can be supported, while maximizing aggregate MGUE and minimizing the average cloud service cost. Shaoxuan Wang, Yao Liu 0008, Sujit Dey |
ICC | 3 |
| 2012 | Video caching in Radio Access Network: Impact on delay and capacityabstractIn this paper, we introduce distributed caching of videos at the base-stations of the Radio Access Network (RAN) as a way to reduce the need to bring requested videos from Internet CDNs, thereby reducing backhaul transmission, improving video quality of experience - delay and video stalling - and increasing overall network capacity to support more number of simultaneous video requests. Unlike Internet CDNs that can store millions of videos in a relatively few large sized caches, our proposed caching architecture consists of a very large number of micro-caches, with each base-station micro-cache being able to store only 1000s of videos, and hence may not be able to have high cache hit ratio. To address this challenge, we propose two new caching policies based on the User Preference Profile (UPP) of users in a cell: R-UPP (Reactive UPP) and P-UPP (Proactive UPP). Further, for videos that result in cache misses and need to be fetched from Internet CDNs, we develop a video scheduling approach that allocates the RAN backhaul resources to the video requests so as to reduce video latency and increase network capacity. We develop a discrete event statistical simulation framework using MATLAB to study the performance of RAN caching. Our simulation results show that RAN micro-caches with the proposed UPP-based caching policies, together with the proposed scheduling approach, can improve the probability of video requests that can meet initial delay requirements by almost 60%, and the number of concurrent video requests that can be served by up to 100%. The results also show that UPP based policies can enhance network capacity by up to 30% compared to conventional caching policies. Hasti Ahlehagh, Sujit Dey |
WCNC | 2 |
| 2012 | Variation-Aware Voltage Level SelectionabstractIn this paper, we address the problem of variation-aware selection of voltage levels for system-on-chips (SoCs) that are organized into multiple frequency and voltage domains. Conventionally, the voltage levels for each domain, as well as the mapping between frequencies and voltages, are determined without considering variations. In the presence of variations, these choices are often suboptimal since the frequency versus voltage characteristics vary from one SoC instance to another and across different voltage domains within an instance. We present a two-pronged approach to address this problem. First, we propose breaking the conventional fixed coupling between voltage levels and frequencies and demonstrate that performing this association based on the characteristics of individual chip instances can lead to significant improvements in power and performance. Second, we show that voltage levels that are computed while accounting for variations can lead to further improvements. We present a methodology to determine a set of discrete voltage levels in a variation-aware manner by generating and quantizing the ideal voltage distribution for a given SoC. Our experiments on an 802.11 MAC processor SoC indicate that the proposed techniques lead to significant improvements in power and performance characteristics in the presence of variations. We obtained an improvement of up to 68% in parametric yield (number of chips meeting power and performance targets) compared to conventional voltage scaling. Saumya Chandra, Anand Raghunathan, Sujit Dey |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2011 | VESPA: Variability emulation for System-on-Chip performance analysisabstractWe address the problem of analyzing the performance of System-on-chip (SoC) architectures in the presence of variations. Existing techniques such as gate-level statistical timing analysis compute the distributions of clock frequencies of SoC components. However, we demonstrate that translating component-level characteristics into a system-level performance distribution is a complex and challenging problem due to the inter-dependencies between components' execution, indirect effects of shared resources, and interactions between multiple system-level “execution paths”. We argue that accurate variation-aware system-level performance analysis requires repeated system execution, which is prohibitively slow when based on simulation. Emulation is a widely-used approach to drastically speedup system-level simulation, but it has not been hitherto applied to variation analysis. We describe a framework - Variability Emulation for SoC Performance Analysis (VESPA) - that adapts and applies emulation to the problem of variation aware SoC performance analysis. The proposed framework consists of three phases: component variability characterization, variation-aware emulation setup, and Monte-carlo driven emulation. We demonstrate the utility of the proposed framework by applying it to design variation-aware architectures for two example SoCs - an 802.11 MAC processor and an MPEG encoder. Our results suggest that variability emulation has great potential to enable variation-aware design and exploration at the system level. Vivek Joy Kozhikkottu, Rangharajan Venkatesan, Anand Raghunathan, Sujit Dey |
DATE | 4 |
| 2010 | Rendering Adaptation to Address Communication and Computation Constraints in Cloud Mobile GamingabstractA new Cloud Mobile Gaming (CMG) approach, where the responsibility of executing the gaming engines, including the most compute intensive tasks of graphic rendering, is put on cloud servers instead of the mobile devices, has the potential for enabling mobile users to play the same rich Internet games available to PC users. However, the mobile gaming user experience may be limited by the communication constraint imposed by available wireless networks and computation constraint imposed by the cost and availability of cloud servers. In this paper, we propose a rendering adaptation technique which can adapt the game rendering parameters to satisfy CMG communication and computation constraints, such that the overall mobile gaming user experience is maximized. Experiments conducted on a commercial UMTS network demonstrate that the proposed rendering adaptation techniques can make the CMG approach feasible: ensuring protection against wireless network conditions, and ensuring server computation scalability, thereby ensuring acceptable mobile gaming user experience. Shaoxuan Wang, Sujit Dey |
GLOBECOM | 2 |
| 2010 | Addressing Response Time and Video Quality in Remote Server Based Internet Mobile GamingabstractA new remote server based gaming approach, where the responsibility of executing the gaming engines is put on remote servers instead of the mobile devices, has the potential for enabling mobile users to play the same rich Internet games available to PC users. However, the mobile gaming user experience may be limited by risks of unacceptably high response time, and low gaming video quality, as gaming control commands, and the resulting gaming video, have to travel through wireless networks characterized by high bandwidth fluctuations, latency and packet loss. In this paper, we propose a set of application layer optimization techniques to ensure acceptable gaming response time and video quality in the remote server based approach. The techniques include downlink gaming video rate adaptation, uplink delay optimization, and client play-out delay adaptation. Experiments conducted on a commercial HSDPA network demonstrate that the proposed optimization techniques can make the remote server based approach feasible, allowing rich Internet games to be played on mobile devices, while ensuring acceptable user experience. Shaoxuan Wang, Sujit Dey |
WCNC | 2 |
| 2010 | Variation-Aware System-Level Power AnalysisabstractThe operational characteristics of integrated circuits in nanoscale semiconductor technology are expected to be increasingly affected by variations in the manufacturing process and the operating environment. In this paper, we address the problem of incorporating the effects of variations into system-level power analysis tools. We consider both manufacturing-induced (die-to-die and within-die) variations in device characteristics, and operation-induced dynamic variations in on-chip temperature. To motivate our work, we first analyze the impact of variations on the power consumption of an example System-on-Chip (SoC). We show how simple extensions of current approaches to system-level power estimation (based on spreadsheets or system-level simulation) are not well-suited to performing variation-aware power-estimation. We propose a system-level power estimation methodology that accurately and efficiently analyzes the impact of variations on SoC power consumption. The proposed methodology combines fast trace analysis, power-state based leakage modeling, efficient thermal analysis, and Monte Carlo sampling to generate SoC power distributions, and power variability traces over time. The key benefit of the methodology is that it captures critical inter-dependencies between component workload profiles, leakage power, and variations in temperature and device parameters, while avoiding time-consuming iterative simulations. Our implementation of the proposed methodology within an in-house system-level power estimation framework indicates speedups of up to 4 orders of magnitude with negligible loss in accuracy as compared to Monte Carlo techniques. We also illustrate the application of our analysis framework can be used to explore a new class of “variation-aware” system-level power management techniques. Saumya Chandra, Kanishka Lahiri, Anand Raghunathan, Sujit Dey |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2009 | Modeling and Characterizing User Experience in a Cloud Server Based Mobile Gaming ApproachabstractWith the evolution of mobile devices and networks, and the growing trend of mobile Internet access, rich, multi-player gaming using mobile devices, similar to PC-based Internet games, has tremendous potential and interest. However, the current client-server architecture for PC-based Internet games, where most of the storage and computational burden of the game lies with the client device, does not work with mobile devices, constraining mobile gaming to either downloadable, single player games, or very light non-interactive versions of the rich multi-player Internet games. In this paper, we study a cloud server based approach, termed Cloud Mobile Gaming (CMG), where the burden of executing the gaming engines is put on cloud servers, and the mobile devices just communicate the users' gaming commands to the servers. We analyze the factors affecting the quality of user experience using the CMG approach, including the game genres, video encoding factors, and the conditions of the wireless network. Based on the above analysis, we develop a model for Mobile Gaming User Experience (MGUE), and develop a prototype for real-time measurement of MGUE that can be used in real networks. We validate MGUE model using controlled subjective testing, and then use it to characterize user experience achievable using the CMG approach in wireless networks. Shaoxuan Wang, Sujit Dey |
GLOBECOM | 2 |
| 2009 | Model-Based Techniques for Data Reliability in Wireless Sensor NetworksabstractWireless sensor networks are a fast-growing class of systems. They offer many new design challenges, due to stringent requirements like tight energy budgets, low-cost components, limited processing resources, and small footprint devices. Such strict design goals call for technologies like nanometer-scale semiconductor design and low-power wireless communication to be used. But using them would also make the sensor data more vulnerable to errors, within both the sensor nodes' hardware and the wireless communication links. Assuring the reliability of the data is going to be one of the major design challenges of future sensor networks. Traditional methods for reliability cannot always be used, because they introduce overheads at different levels, from hardware complexity to amount of data transmitted. This paper presents a new method that makes use of the properties of sensor data to enable reliable data collection. The approach consists of creating predictive models based on the temporal correlation in the data and using them for real-time error correction. This method handles multiple sources of errors together without imposing additional complexity or resource overhead at the sensor nodes. We demonstrate the ability to correct transient errors arising in sensor node hardware and wireless communication channels through simulation results on real sensor data. Shoubhik Mukhopadhyay, Curt Schurgers, Debashis Panigrahi, Sujit Dey |
IEEE Trans. Mob. Comput. | 4 |
| 2009 | Variation-Tolerant Dynamic Power Management at the System-LevelabstractThe power characteristics of system-on-chips (SoCs) in nanoscale technologies are significantly impacted by manufacturing process variations, making it important to consider these effects during system-level power analysis and optimization. In this paper, we identify and address the problem of designing effective power management schemes in the presence of such variations. In particular, we demonstrate that conventional power management schemes, which are designed without considering the impact of variations, can result in substantial power wastage. We therefore propose two approaches to variation-aware power management, namely, design-specific and chip-specific approaches. In each of these approaches, the goal is to consider the impact of variations while deriving power management policy parameters, in order to optimize metrics that are relevant under variations. We motivate and introduce these metrics, and present both exact and heuristic approaches to optimize them. The methods are designed and implemented in the context of two power management frameworks, namely an ideal oracle-based framework and a timeout-based framework. We experimentally evaluate the proposed ideas using an ARM946 processor core model. For the oracle-based framework, variation-aware power management can result in improvements of upto 59% for$\mu +\sigma$, and upto 55% for 95th percentile of the energy distribution, over conventional power management schemes that do not consider variations. For the timeout-based framework, we obtain reductions of upto 43% in$\mu+\sigma$and upto 55% in the 99th percentile of the energy distribution. Saumya Chandra, Kanishka Lahiri, Anand Raghunathan, Sujit Dey |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2008 | Dynamically Configurable Bus Topologies for High-Performance On-Chip CommunicationabstractThe on-chip communication architecture is a primary determinant of overall performance in complex system-on-chip (SoC) designs. Since the communication requirements of SoC components can vary significantly over time, communication architectures that dynamically detect and adapt to such variations can substantially improve system performance. In this paper, we propose Flexbus, a new architecture that can efficiently adapt thelogicalconnectivityof the communication architecture and the components connected to it. Flexbus achieves this by dynamically controlling both the communication architecture topology, as well as the mapping of SoC components to the communication architecture. This is achieved through newdynamicbridgeby-pass, andcomponentremappingtechniques. In this paper, we introduce these techniques, describe how they can be realized within modern on-chip buses, and discuss policies for run-time reconfiguration of Flexbus-based architectures.The techniques underlying Flexbus are general, and are applicable to a variety of bus standards. We have implemented Flexbus as an extension of the popular AMBA AHB bus, and have evaluated it using a commercial design flow. We report on experiments conducted to analyze its area, timing, and performance under a wide variety of system-level traffic profiles. We have applied Flexbus to two example SoC designs: 1) an IEEE 802.11 MAC processor and 2) an UMTS turbo decoder. Our results show that Flexbus provides gains of up to 34.55 % in application data-rates over conventional architectures, with negligible area overhead and a 3.2% penalty in clock period. Krishna Sekar, Kanishka Lahiri, Anand Raghunathan, Sujit Dey |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2008 | Intelligent Robustness Insertion for Optimal Transient Error Tolerance Improvement in VLSI CircuitsabstractDue to aggressive technology scaling, VLSI circuits are becoming increasingly susceptible to radiation-induced single-event-upsets (SEUs). Redundancy insertion has been adopted to provide the circuit with additional transient error resiliency. However, its applicability and efficiency are limited by the tight design constraints and budgets. In this paper, we present an intelligent ldquoconstraint-aware robustness insertionrdquo methodology. By selectively protecting sequential elements in static CMOS digital circuits, it is able to maximally improve the SEU tolerance while keeping the incurred design overhead within acceptable range. Our technique consists of three major components. The first one is a configurable hardening sequential cell design that serves as the basic building block of the framework; the second one is a robustness calibration technique that evaluates the relative error tolerance of all sequential elements and provides guidelines to the redundancy insertion; the third one is an optimization algorithm that searches for the optimal protection scheme under given design constraints and budgets. Simulation results show that the intelligent robustness insertion reduced the error rate by 46% with zero timing penalty and 10% area increase. Furthermore, by exploring the tradeoffs between reliability and design overhead, we also demonstrate the proposed technique can help achieve high reliability improvement while keeping the design overhead within acceptable range. Sujit Dey |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2007 | System-on-Chip Power Management Considering Leakage Power VariationsabstractThe power characteristics of System-on-chips (SoCs) in nanoscale technologies are significantly impacted by process variations, making it important to consider these effects during system-level power analysis and optimization. In this paper, we identify and address the problem of designing effective power management schemes in the presence of such variations. In particular, we demonstrate that conventional power management schemes, which are designed without considering the impact of variations, can result in substantial power wastage. We therefore propose two approaches to variation-aware power management, namely, design-specific and chip-specific approaches. In each of these approaches, the goal is to consider the impact of variations while deriving the values of parameters that govern popular power management policies. The policy parameters are derived so as to optimize metrics that are relevant under variations. We motivate and introduce these metrics, and use a combination of analytical and empirical approaches to optimize them. We experimentally evaluate the proposed ideas in the context of an ARM processor core, and demonstrate that variation-aware power management can result in improvements of upto 59% in μ + σ of the energy distribution, and upto 55% for the 95th percentile of the energy distribution, with respect to conventional power management schemes that do not consider variations. Saumya Chandra, Kanishka Lahiri, Anand Raghunathan, Sujit Dey |
DAC | 4 |
| 2007 | Joint Computation and Communication Scheduling to Enable Rich Mobile ApplicationsabstractIncreasing interest in sensor networking and ubiquitous computing has created a trend towards embedding more and more intelligence into our surroundings. This enables thin wireless clients to support powerful applications by using processing resources that are available around them. One of the major challenges is how to schedule both the embedded computing and wireless communication resources to support a broad set of client nodes. This scheduling has to take into account the constraints on the wireless capacity, as well as the resource limitations on the processing elements. We have developed a set of fast algorithms to perform this scheduling, based on an LP-approachand a greedy solution. These algorithms are able to perform the scheduling with a performance close to the optimal exhaustive solution, but with an execution time that is reduced by three or more orders of magnitude (from hours to seconds). Shoubhik Mukhopadhyay, Curt Schurgers, Sujit Dey |
GLOBECOM | 3 |
| 2007 | Modeling soft error effects considering process variationsabstractThis paper addresses the aggregated effects of two types of variations that contribute to the reliability degradation. The first one is the increasing level of process variation; the second one is one particular type of environmental variation - the radiation-induced soft error. Their simultaneous presence can cause large negative performance impact. We present a statistical approach to model the generation and propagation of a transient soft error inside combinational circuits considering the existence of inter-die channel length variation in CMOS digital circuits. Experiment results have demonstrated that channel length variation can significantly aggravate the soft error effect, which can be accurately evaluated using the proposed methodology. Sujit Dey |
ICCD | 2 |
| 2007 | A Device and Network-Aware Scaling Framework for Efficient Delivery of Scalable Video over Wireless NetworksabstractIn recent years, the number of devices that can support multimedia-content over wireless networks has increased significantly. With devices ranging from cellular phones to TV's, there is a challenge faced by content service providers to support multimedia delivery for a diverse set of devices. Scalable video coding (SVC) may provide a solution by enabling devices with techniques to extract video streams scaled to match their capabilities. However, SVC can be significantly bandwidth inefficient. By sending the highest possible quality stream over the network, irrespective of the capabilities of the end-devices and the network, SVC can limit the number of devices being served and reduce the network's average quality satisfaction. In this paper, we present a device and network- aware scaling (DeNAS) framework to address these problems. Based on device capability information, network capacity, and the desired service objective, DeNAS finds the suitable scaling level for each video stream prior to its delivery over the wireless network. We demonstrate the effectiveness of DeNAS through simulation and show that DeNAS is able to improve quality satisfaction, increase bandwidth efficiency, and satisfy a greater number of clients simultaneously. Naomi Ramos, Sujit Dey |
PIMRC | 2 |
| 2007 | Evaluating Transient Error Effects in Digital Nanometer CircuitsabstractRadiation-induced transient errors have become a great threat to the reliability of nanometer circuits. The need for cost-effective robust circuit design mandates the development of efficient reliability metrics. We present a novel ldquonoise impact analysisrdquo methodology to evaluate the transient error effects in static CMOS digital circuits. With both the circuit, and the transient noise abstracted in the format of matrices, the circuit-noise interaction is modeled by a series of matrix transformations. During the transformation, factors that potentially affect the propagation & capture of transient errors are modeled as matrix operations. Finally, a ldquonoise capture ratiordquo is computed as the probability of a sequential element capturing transient noise inside the combinational logics, It is used as a measure of the transient noise effects in the circuit. Comparison with SPICE simulation demonstrates that our technique can accurately, yet quickly estimate the probability of transient errors causing observable error effects. The proposed methodology will greatly facilitate the economic design of robust nanometer circuits. Xiaoliang Bai, Sujit Dey |
IEEE Trans. Reliab. | 3 |
| 2007 | Dynamic adaptation policies to improve quality of service of real-time multimedia applications in IEEE 802.11e WLAN Networks
Naomi Ramos, Debashis Panigrahi, Sujit Dey |
Wirel. Networks | 3 |
| 2006 | Integrated data relocation and bus reconfiguration for adaptive system-on-chip platformsabstractDynamic variations in application functionality and performance requirements can lead to the imposition of widely disparate requirements on system-on-chip (SoC) platform hardware over time. This has led to interest in the design and use of adaptive SoC platforms that are capable of providing high performance in the face of such variations. Recent advances in circuits and architectures are enabling platforms that contain various mechanisms for runtime adaptation. However, the problem of exploiting such configurability in a coordinated manner at the system level remains a challenging task. In this work, we focus on two configurable subsystems of SoC platforms that play a crucial role in determining overall system performance, namely, the on-chip communication architecture, and the on-chip memory architecture. Using detailed case studies, we demonstrate the limitations of designs in which the architectural configuration of a bus-based communication architecture and the placement of data in memory are statically optimized, and those in which each is customized separately, without considering their interdependence. We propose an integrated methodology for dynamically relocating on-chip data and reconfiguring the communication architecture, and discuss the necessary hardware support. Experiments conducted on an SoC platform that integrates decoders for the UMTS (3G) and IEEE 802.11a (wireless LAN) standards demonstrate that the proposed integrated adaptation technique helps boost the maximum achievable performance by up to 32% over the best statically optimized design Krishna Sekar, Kanishka Lahiri, Anand Raghunathan, Sujit Dey |
DATE | 4 |
| 2006 | Considering process variations during system-level power analysisabstractProcess variations will increasingly impact the operational characteristics of integrated circuits in nanoscale semiconductor technologies. Researchers have proposed various design techniques to address process variations at the mask, circuit, and logic levels. However, as the magnitude of process variations increases, their effects will need to be addressed earlier in the design cycle.In this paper, we propose techniques for accurately and efficiently incorporating the effects of process variations into system-level power estimation tools. To motivate our work, we first study the impact of process variations on the power consumption of an example System-on-Chip (SoC). We consider simple extensions of current approaches to system-level power estimation (spreadsheet-based and simulation-based power estimation), and demonstrate their limitations in performing variation-aware power estimation. We propose a system-level power estimation methodology that can accurately and efficiently analyze the impact of process variations on SoC power. The proposed methodology combines efficient trace-based analysis, power-state based leakage modeling, and Monte Carlo sampling. The key benefit of the proposed methodology is that it captures the necessary inter-dependencies while avoiding iterative system-level simulation. Our implementation of the proposed techniques within an in-house system-level power estimation framework indicates 2-5 orders of magnitude efficiency gains, with negligible loss in accuracy, compared to direct Monte Carlo techniques that require iterative system simulation. Saumya Chandra, Kanishka Lahiri, Anand Raghunathan, Sujit Dey |
ISLPED | 4 |
| 2006 | Evaluating and Improving Transient Error Tolerance of CMOS Digital VLSI CircuitsabstractVLSI circuits are becoming increasingly susceptible to radiation-induced "single event upset (SEU)". This paper focuses on one type of SEU caused by particle strikes inside the combinational logics, called "single event transient (SET)". We study various factors affecting SET effects in CMOS digital circuits and present a static method of analyzing the circuit's SET tolerance. We also propose a heuristic cell resizing process to effectively improve the circuit SET tolerance with limited design overhead. Experimental results have shown that our analysis can accurately evaluate SET effects and the cell resizing process is able to significantly reduce the probability of SETs becoming stable errors with no timing cost and negligible area penalty Sujit Dey |
ITC | 2 |
| 2005 | FLEXBUS: a high-performance system-on-chip communication architecture with a dynamically configurable topologyabstractIn this paper, we describe FLEXBUS, a flexible, high-performance onchip communication architecture featuring a dynamically configurable topology. FLEXBUS is designed to detect run-time variations in communication traffic characteristics, and efficiently adapt the topology of the communication architecture, both at the system-level, through dynamic bridge by-pass, as well as at the component-level, using component re-mapping. We describe the FLEXBUS architecture in detail and present techniques for its run-time configuration based on the characteristics of the on-chip communication traffic. The techniques underlying FLEXBUS can be used in the context of a variety of on-chip communication architectures. In particular, we demonstrate its application to AMBA AHB, a popular commercial on-chip bus. Detailed experiments conducted on the FLEXBUS architecture using a commercial design flow, and its application to an IEEE 802.11 MAC processor design, demonstrate that it can provide significant performance gains as compared to conventional architectures (up to 31.5% in our experiments), with negligible hardware overhead. Krishna Sekar, Kanishka Lahiri, Anand Raghunathan, Sujit Dey |
DAC | 4 |
| 2005 | Constraint-aware robustness insertion for optimal noise-tolerance enhancement in VLSI circuitsabstractReliability of nanometer circuits is becoming a major concern in today's VLSI chip design due to interferences from multiple noise sources as well as radiation-induced soft errors. Traditional noise analysis/avoidance and manufacturing testing are no longer sufficient to handle the dynamic interactions between various noise sources and unpredictable operational variations. Therefore, "robustness insertion" has been adopted as the supplementary approach to ensure high circuit reliability through on-line protections. However, the related design overhead is not always acceptable, especially for cost/timing-sensitive designs. In this paper, we present a novel "constraint-aware robustness insertion" methodology protect the sequential elements in digital circuits against various noise effects. Based on a configurable hardening sequential cell design and an efficient sequential cell robustness estimation technique, an optimization algorithm is developed to search for the optimal protection scheme under given timing and area constraints. Experiment results demonstrate that the proposed methodology is able to achieve a high degree of noise-tolerance while keeping the protection cost within limit. Sujit Dey |
DAC | 3 |
| 2005 | A static noise impact analysis methodology for evaluating transient error effects in digital VLSI circuitsabstractSingle-event-upset (SEU) has become a great threat to the reliability of nanometer circuits. The need for cost-effective robust circuit design mandates the development of efficient reliability analysis. In this paper, a static "noise impact analysis" methodology is developed to estimate the circuit vulnerability. First, both the circuit elements and the transient noise are abstracted in the format of matrices. Then the circuit-noise interaction is modeled by a series of matrix transformations, which jointly considers three masking effects that can potentially prevent transient noise from causing observable errors. Finally, the error-resiliency of the sequential elements is considered in determining the impact of transient noise on the circuit. Experiment results demonstrate that our technique can accurately yet quickly estimate the circuit failure rate by comparing with HSPICE simulation. The proposed methodology has greatly facilitate the economic design of robust nanometer circuit. Xiaoliang Bai, Sujit Dey |
ITC | 3 |
| 2005 | Dynamic end-to-end image adaptation for guaranteed quality of service in wireless image data servicesabstractThe paper investigates total end-to-end service variability issues in providing dynamic image adaptation services to wireless end users. While most existing image adaptation techniques focus on improving wireless transmission latency by adapting content-rich images to the time varying bandwidth availability of wireless networks, we have found dynamic service load variations at the server also significantly affect end-to-end service latency. We introduce a dynamic end-to-end image adaptation technique that optimizes the runtime image adaptation process so as to provide guaranteed total service latency (including both server processing latency and wireless transmission latency) with a maximized image quality. Through judicious tuning of application-layer image compression parameters, the proposed technique dynamically optimizes both the image quality and each component of service latency required for wireless image data services simultaneously depending on service variability issues existing at both server-side and network-side. Experimental results under synthetic service workloads demonstrate that our proposed dynamic end-to-end image adaptation can provide guaranteed service latency with a minimal impact on the image quality loss. Dong-Gi Lee, Sujit Dey |
WCNC | 2 |
| 2004 | A scalable soft spot analysis methodology for compound noise effects in nano-meter circuitsabstractCircuits using nano-meter technologies are becoming increasingly vulnerable to signal interference from multiple noise sources as well as radiation-induced soft errors. One way to ensure reliable functioning of chips is to be able to analyze and identify the spots in the circuit which are susceptible to such effects (called “soft spots ” in this paper), and to make sure such soft spots are “hardened ” so as to resist multiple noise effects and soft errors. In this paper, we present a scalable soft spot analysis methodology to study the vulnerability of digital ICs exposed to nano-meter noise and transient soft errors. First, we define “softness ” as an important characteristic to gauge system vulnerability. Then several key factors affecting softness are examined. Finally an efficient Automatic Soft Spot Analyzer (ASSA) is developed to obtain the softness distribution which reflects the unbalanced noise-tolerant capability of different regions in a design. The proposed methodology provides guidelines to reduction of severe nano-meter noise effects caused by aggressive design in the premanufacturing phase, and guidelines to selective insertion of online protection schemes to achieve higher robustness. The quality of the proposed soft-spot analysis technique is validated by HSPICE simulation, and its scalability is demonstrated on a commercial embedded processor. Xiaoliang Bai, Sujit Dey |
DAC | 3 |
| 2004 | VSHAPER: an efficient method of serving video streams shaped for diverse wireless communication conditionsabstractThe transmission of video over wireless channels is a challenging problem, in part due to the rapid and wide variations in wireless channel bandwidth and error conditions. In this paper, we present a run-time method for assembling a video stream, VSHAPER, that shapes the video stream depending on the current wireless channel conditions. Previous methods for shaping video streams have either concentrated on bandwidth adaptation only, or have required the video stream to be compressed at run-time. VSHAPER assembles the output video stream from several pre-compressed videos, requiring minimal computational overhead at the video server. A modified rate-distortion optimization method for deciding how to shape the video stream is also used to minimize computational overhead. The output of VSHAPER is a video stream that can be read by any standard video decoder. Significant gains in the video quality of received video during varying wireless conditions are observed. In addition, a video server's capacity (number of video streams it can simultaneously serve) is increased significantly. Clark N. Taylor, Sujit Dey |
GLOBECOM | 2 |
| 2004 | Run-time allocation of buffer resources for maximizing video clip quality in a wireless last-hop systemabstractThe transmission of digital video over wireless channels faces many problems common Io both wired and wireless networks such as bandwidth and buffer restrictions and low energy constraints, together with problems unique to wireless technologies such as extreme and rapidly varying channel noise. In this paper, we focus on enabling (he wireless communication of video clips despite the noisy and dynamic nature of wireless channels. Many efforts have been made in the past to increase the error resiliency of compressed video, but these approaches have not considered the effects of buffering on the received video quality. Traditional buffering (where a "sliding window" of the next few frames is kept in the buffer at all times), while useful for overcoming latency jitter, cannot be used Io guarantee against poor video quality and stalls due to wireless channel fluctuations. To help overcome the weaknesses of traditional buffering, we introduced out-of-order buffering (ORBit), a method for selecting important frames from throughout a video clip and buffering them before playback begins. ORBit can lead to significant gains in video quality despite variations in the wireless channel during the communication of a video clip, but requires the judicious selection of ORBit frames, with their corresponding quantization and channel coding levels. To enable run-time selection of ORBit frames, we introduce a fast video distortion estimation technique. Using our estimation technique, we present an ORBiI decision algorithm that selects ORBit frames depending on the wireless technology and end-user appliance constraints. Using the ORBit has to wait for a few seconds before the video display begins. During decision algorithm, significant gains of up to 9 dB in video quality are obtained. Clark N. Taylor, Sujit Dey |
ICC | 2 |
| 2004 | Dynamic image adaptation technique and architecture to enhance server performance in wireless image servicesabstractImage-based data services, like wireless photos with camera phones and multi-media message services, are gaining in popularity and wireless data market share. Dynamic image adaptation techniques have been shown to be effective to provide required tradeoffs between image quality, network bandwidth availability, and wireless transmission latency, enabling low-cost, best-effort wireless image data transmission (D.G. Lee et al., 2002; D.G. Lee et al., 2003). This paper investigates the effect of image adaptation techniques on image servers, including server latency and server capacity, and end-to-end service latency using a proposed performance evaluation and exploration framework. We first motivate the need for dynamic adaptation techniques by demonstrating the greatly reduced end-to-end service latency using the proposed dynamic adaptation technique over the existing methods. We next investigate the impact of image adaptation techniques under overloaded server conditions. While the proposed dynamic adaptation technique can provide server performance improvement (8/spl times/ lower maximum server latency and 6/spl times/ server capacity increase), even more significant server performance increases can be achieved through the use of a novel configurable hardware/software (HW/SW) architecture for the image adaptation technique. The proposed HW/SW architecture is capable of achieving significantly higher server performance (server capacity increase of 18/spl times/ and server latency decrease of 34/spl times/), while accommodating several different types of image adaptation techniques. Experimental results show that customized image-based data services can be enabled at significantly reduced server costs, and thereby reducing overall wireless service cost. Dong-Gi Lee, Sujit Dey |
PIMRC | 2 |
| 2004 | CHASER: content and channel aware object scheduling and error control for wireless Web access in 3G networksabstractA major bottleneck in satisfying the increasing demand for wireless multimedia access is the dynamic error conditions requiring a very high error protection overhead in terms of energy consumption and access delay. In this paper, we explore ways to use application-level data properties to reduce the error protection cost. We observe that the objects constituting a wireless Web-based access differ in terms of the inherent resiliency to residual error, and the object's importance to the overall content quality. Additionally, since the cost of achieving a target residual error level depends on the experienced channel conditions, the access cost can be affected by changing the transmission schedule of objects according to the channel conditions. Based on the above observations, in this paper, we present a framework that exploits the application properties, namely content importance and error resiliency, to reduce the cost of error protection by scheduling objects for transmission and adapting the lower-layer error control mechanisms based on the data properties as well as wireless channel conditions. The approach is evaluated under diverse channel conditions using a MATLAB-based simulation environment. Our results show that the proposed approach leads to significant savings in access cost (average 37.4%) with minimal degradation in application quality. Debashis Panigrahi, Sujit Dey |
PIMRC | 2 |
| 2004 | Model based error correction for wireless sensor networksabstractOne of the main challenges in wireless sensor networks is to provide low-cost, low-energy reliable data collection. Reliability against transient errors in sensor data can be provided using the model-based error correction described in (S. Mukhopadhyay et al., Mar. 2004), in which temporal correlation in the data is used to correct errors without any overheads at the sensor nodes. In the above work it is assumed that a perfect model of the data is available. However, as variations in the physical process are context-dependent and time-varying in a real sensor network, it is infeasible to have an accurate model of the data properties a priori, thus leading to reduced correction efficiency. In this paper, we address this issue by presenting a scalable methodology for improving the accuracy of data modeling through on-line estimation and model updates. Additionally, we propose enhancements to the data correction algorithm to incorporate robustness against dynamic model changes and potential modeling errors. We evaluate our system through simulations using real sensor data collected from different sources. Experimental results demonstrate that the proposed enhancements lead to an improvement of up to a factor of 10 over the earlier approach. Shoubhik Mukhopadhyay, Debashis Panigrahi, Sujit Dey |
SECON | 3 |
| 2004 | Data aware, low cost error correction for wireless sensor networksabstractOne or the main challenges in adoption and deployment of wireless networked sensing applications is ensuring reliable sensor data collection and aggregation, while satisfying the low-cost, low-energy operating constraints of such applications. A wireless sensor network is inherently vulnerable to different sources of unreliability resulting in transient failures. Existing reliability techniques that address transient failures in circuits and communication channels incur prohibitively high energy, bandwidth and cost overheads in the sensor nodes. In this paper we investigate application-level error correction techniques for sensor networks that exploit the properties of sensor data to eliminate any overhead on the sensor nodes, at the expense of nominal buffer requirements at the data aggregator nodes, which are much less cost/energy constrained. Our approach involves use of spatio-temporal correlations in sensor data, the goals of the application, and its vulnerability to various errors. We present our error-correction algorithm and evaluate it through simulations using real and synthetic sensor data. Experimental results validate the feasibility of our approach to provide high degree of reliability in sensor data aggregation, without imposing overheads on sensor nodes. Shoubhik Mukhopadhyay, Debashis Panigrahi, Sujit Dey |
WCNC | 3 |
| 2004 | Interconnect coupling-aware driver modeling in static noise analysis for nanometer circuitsabstractWith geometries shrinking in nanometer technologies, crosstalk noise becomes a critical issue. Modern designs like system-on-chips have millions of noise-prone nodes, mandating fast yet accurate crosstalk noise analysis techniques. Using linear circuit model, static noise analysis can efficiently estimate crosstalk noise. Traditionally in static noise analysis, drivers' holding resistances are precharacterized without considering the potential impact of crosstalk noise. However, crosstalk induced voltage fluctuation strongly affects the behavior of nonlinear drivers. When facing different coupling interconnects and hence crosstalk noise, a driver's holding resistance can change dramatically. In nanometer circuits, this substantial variation of nonlinear drivers cannot be totally ignored. To achieve high-quality in noise estimation yet maintain the efficiency of linear circuit model, we propose a novel interconnect coupling-aware driver modeling method. Based on layout-extracted interconnect parameters and precharacterized driver models, an effective holding resistance is calculated to capture the impact of the nonlinear driver. Multiple aggressors with synchronous and asynchronous switching activities are also considered. The proposed method is simple, efficient, and enables on-the-fly calculation of the effective holding resistance. Experiments show that with negligible computation overhead, the coupling-aware driver modeling methodology can significantly improve the quality of static noise analysis. Xiaoliang Bai, Rajit Chandra, Sujit Dey, P. V. Srinivas |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2004 | High-level crosstalk defect Simulation methodology for system-on-chip interconnectsabstractFor system-on-chips (SoC) using nanometer technologies, buses and long interconnects are susceptible to crosstalk defects that may lead to functional and timing failures. Testing for crosstalk defects is becoming important to ensure error-free operation of an SoC. To efficiently evaluate crosstalk-defect coverage of existing tests and facilitate the development of new crosstalk test methodologies, effective crosstalk-defect coverage-analysis techniques are needed. In this paper, we present an efficient high-level crosstalk-defect simulation methodology for interconnects dominated by capacitive coupling effects. A novel coupling defect-simulation model was developed and implemented in hardware description languages. The high-level crosstalk-defect simulation methodology was examined by SPICE simulations. Experimental results show the crosstalk defect simulation methodology efficiently provides high-fidelity defect-coverage results. The proposed methodology enables fast exploration and evaluation of different tests, leading to high-quality, low-cost manufacturing tests for crosstalk-induced ac failures. Xiaoliang Bai, Sujit Dey |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2004 | Efficient power profiling for battery-driven embedded system designabstractThe ability to efficiently and accurately estimate battery life under different design choices at the system level is an important aid in designing battery-efficient systems. Recently developed battery models help by estimating battery life under given profiles of the battery discharge current over time. However, existing techniques for energy (or average power) estimation do not provide sufficient information (such as time profiles of system power consumption) to drive battery-life estimation. Techniques that are capable of generating such profiles often lack the efficiency required to support exploration at the system level. In this paper, we describe techniques for efficient generation of system-level power profiles, for use in a battery-life estimation framework. Our power profiling technique allows a designer to experiment with: 1) the mapping of system tasks to a set of architectural components and 2) the mapping of system communications to a specified communication architecture, and efficiently generate system power profiles for each alternative. The resulting profiles can then be analyzed using existing battery models to estimate battery lifetime and capacity. Extensive experiments conducted on an IEEE 802.11 MAC processor design demonstrate that our power profiler offers orders of magnitude improvement in runtimes over state-of-the-art cosimulation-based power estimation techniques, while suffering minimal loss of accuracy (average profiling error was 3.8%). Kanishka Lahiri, Anand Raghunathan, Sujit Dey |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2004 | Design space exploration for optimizing on-chip communication architecturesabstractRapid growth in the complexity of system-on-chips is being accompanied by increasing volume and diversity of on-chip communication traffic, which in turn, is driving the development of advanced system-level communication architectures. While these architectures have the potential to improve system performance, they pose significant new challenges to the system designer, owing to the complex design space defined by the availability of numerous network topologies, communication protocols, and mapping alternatives for system communications. In this paper, we address the problem of mapping a system's communication requirements to a given communication architecture template. We illustrate the nature of the communication architecture design space, and describe an exploration methodology that uses efficient algorithms to help automate the process of mapping the system communications to the selected template. In addition, we demonstrate the importance of simultaneously optimizing the on-chip communication protocols in order to maximize system performance. Experiments conducted on example systems, including a cell forwarding unit of an ATM switch, indicate that the proposed techniques aid in automatically constructing communication architectures that have high performance. For the systems we considered, the solutions generated using our methodology had 53% superior performance (on average), over those based on conventional architectures and mapping approaches. The algorithms used in the proposed methodology are computationally efficient, and scale well with increasing communication architecture complexity. Kanishka Lahiri, Anand Raghunathan, Sujit Dey |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2004 | Design of high-performance system-on-chips using communication architecture tunersabstractIn this paper, we present a methodology for the design of high-performance system-on-chip communication architectures. The approach is based on the addition of a layer of circuitry called the communication architecture tuner (CAT) layer around an existing communication architecture topology. The added layer provides a system with the capability of adapting to runtime variability in the communication needs of its constituent components. For example, more critical data may be handled differently, leading to lower communication latencies. The CAT associated with each component monitors its internal state, analyzes the communication transactions it generates, and "predicts" the relative importance of the transactions in terms of their impact on system-level performance metrics. It then configures the protocol parameters of the underlying communication architecture (e.g., priorities, burst modes, etc.) to best suit the system's changing communication needs. We illustrate the issues and tradeoffs involved in the design of CAT-based communication architectures, and present algorithms that automate the key steps. Experiments with example systems indicate that performance metrics (e.g., number of missed deadlines, average processing time) for systems with CAT-based communication architectures are significantly (sometimes over an order of magnitude) better than those with conventional communication architectures. Kanishka Lahiri, Anand Raghunathan, Ganesh Lakshminarayana, Sujit Dey |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2004 | Common-case computation: a high-level energy and performance optimization techniqueabstractThis paper proposes a novel circuit design methodology, called common-case computation (CCC)-based design, and new design automation algorithms for optimizing energy consumption and performance. The proposed techniques are applicable in conjunction with any high-level design methodology, where a structural register-transfer level (RTL) description and its corresponding scheduled behavioral (cycle-accurate functional) description are available. It is a well-known fact that in behavioral descriptions of hardware circuits (and also in software programs), a small set of computations often account for most of the computational complexity. However, in the hardware implementations (structural RTL or lower level), the common cases and the remaining computations are typically treated alike. This paper shows that identifying and exploiting common cases during the design process can lead to implementations that are much more efficient in terms of energy consumption and performance. Ganesh Lakshminarayana, Anand Raghunathan, Kamal S. Khouri, Niraj K. Jha, Sujit Dey |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2004 | Resource budgeting for Multiprocess High-level synthesisabstractThis paper presents a new high-level synthesis methodology to generate optimized register-transfer level (RTL) implementations for multiprocess behavioral descriptions. The concurrent communicating processes specification paradigm is widely used in digital circuit and system design, and is employed in all popular hardware description languages. It has been shown that interprocess communication and synchronization can result in complex timing interdependencies, which significantly affect the performance of a multiprocess system. In this paper, we demonstrate that state-of-the-art high-level synthesis tools can generate significantly suboptimal implementations for behaviors that contain concurrent communicating processes. We present an analysis of how interprocess communication impacts high-level synthesis steps, and describe a new methodology to adapt existing high-level synthesis tools to optimize multiprocess descriptions. Our methodology is based on executing multiprocess performance analysis and process-by-process scheduling in an iterative manner. We present algorithms for key steps in the proposed methodology. We have performed extensive experiments in the context of a commercial high-level design flow to evaluate the proposed techniques. The results clearly demonstrate the utility of our techniques in synthesizing implementations with superior area, performance, and energy consumption. For example, up to 40% performance improvement (average of 35.6%) was achieved with little or no area overhead (average of 4.8%). In effect, the proposed techniques lead to a shift of the entire area-delay tradeoff curve for a design, to include superior designs that were hitherto infeasible. Our techniques also simultaneously result in up to 50% (average of 33.5%) improvement in energy and up to 69% (average of 58.3%) improvement in the energy-delay product. Anand Raghunathan, Niraj K. Jha, Sujit Dey |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2004 | Optimizing designs using the addition of deflection operationsabstractThis paper introduces hot potato behavioral synthesis transformation techniques. These techniques add deflection operations in the behavioral description of a computation in such a way that the requirements for two important components of the final implementation cost, the number of registers and the number of interconnects, are significantly reduced. Moreover, we demonstrate how hot potato techniques can be effectively used during behavioral synthesis to minimize the partial scan overhead to make the synthesized design testable. Jennifer Wong-Ma, Miodrag Potkonjak, Sujit Dey |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2004 | Double sampling data checking technique: an online testing solution for multisource noise-induced errors on on-chip interconnects and busesabstractWith processors and system-on-chips using nano-meter technologies, several design and test efforts have been recently developed to eliminate and test for many emerging DSM noise effects. In this paper, we show the emergence of multisource noise effects, where multiple DSM noise sources combine to produce functional and timing errors even when each separate noise source itself does not. We show the dynamic nature of multisource noise, and the need for online testing to detect such noise errors. We propose an online approach based on low-cost double-sampling data checking circuit to test for such noise effects in on-chip buses. Based on the proposed circuit, an effective and efficient testing methodology has been developed to facilitate online testing for generic on-chip buses. The applicability of this methodology is demonstrated through embedding the online detection circuit in a bus design. The validated design shows the effectiveness of the proposed testing methodology for multisource noise-induced errors in global interconnects and buses. Sujit Dey |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2003 | A scalable software-based self-test methodology for programmable processorsabstractSoftware-based self-test (SBST) is an emerging approach to address the challenges of high-quality, at-speed test for complex programmable processors and systems-on chips (SoCs) that contain them. While early work on SBST has proposed several promising ideas, many challenges remain in applying SBST to realistic embedded processors. We propose a systematic scalable methodology for SBST that automates several key steps. The proposed methodology consists of (i) identifying test program templates that are well suited for test delivery to each module within the processor, (ii) extracting input/output mapping functions that capture the controllability/observability constraints imposed by a test program template for a specific module-under-test, (iii) generating module-level tests by representing the input/output mapping functions as virtual constraint circuits, and (iv) automatic synthesis of a software self-test program from the module-level tests. We propose novel RTL simulation-based techniques for template ranking and selection, and techniques based on the theory of statistical regression for extraction of input/output mapping functions. An important advantage of the proposed techniques is their scalability, which is necessitated by the significant and growing complexity of embedded processors.To demonstrate the utility of the proposed methodology, we have applied it to a commercial state-of-the-art embedded processor (Xtensa™ from Tensilica Inc.). We believe this is the first practical demonstration of software-based self-test on a processor of such complexity. Experimental results demonstrate that software self-test programs generated using the proposed methodology are able to detect most (95.2%) of the functionally testable faults, and achieve significant simultaneous improvements in fault coverage and test length compared with conventional functional test. Srivaths Ravi 0001, Anand Raghunathan, Sujit Dey |
DAC | 4 |
| 2003 | Dynamic Platform Management for Configurable Platform-Based System-on-Chips
Krishna Sekar, Kanishka Lahiri, Sujit Dey |
ICCAD | 3 |
| 2003 | Separate Dual-Transistor Registers - A Circuit Solution for On-line Testing of Transient Error in UDSM-ICabstractThis paper addresses the soft-error problem in UDSM circuits by presenting on-line fault-tolerant circuit design techniques. In our scheme, separate dual transistor (SDT) structure is introduced into the register design as a key component to increase the input-signal stability as well as the robustness of the circuit against the effects of ionizing particles. Our work not only demonstrates the feasibility of its physical implementation, but also shows the cost effectiveness. To compare with other fault-tolerant techniques, ISCAS89 circuits have been synthesized with the SDT standard cells to investigate its cost/timing overheads. Our benchmark comparison reveals its better applicability over two representative techniques (TMR and ECC) for the logic circuits in digital systems. Sujit Dey |
IOLTS | 2 |
| 2003 | HyAC: A Hybrid Structural SAT Based ATPG for CrosstalkabstractAs technology evolves into the deep sub-micron era, signal integrity problems are growing into a major challenge. An important source of signal integrity problems is the crosstalk noise generated by coupling capacitances between wires. Test vectors that activate and propagate crosstalk noise effects are becoming an essential part of design verification and manufacturing test. However, deriving such vectors is a complex task. In this paper, we propose HyAC, a fast yet accurate hybrid ATPG method targeting multiple-aggressor induced crosstalk errors. Given a victim and a set of aggressors, the proposed ATPG method searches for test vectors to activate and propagate a crosstalk error for the victim. Due to logic constraints, it may not be possible to trigger all aggressors simultaneously. Therefore, firstly we use an implication graph (IG) that consists of logic variables and structural information to check for logic conflicts. If the current set of aggressors is not feasible, our algorithm automatically searches for the next-best subset of aggressors (resulting in the largest noise). After a set of feasible aggressors is identified, we use a modified PODEM [21] algorithm to search for test vectors. This hybrid structural SAT-based ATPG method inherits advantages from both Boolean– satisfiabilitity based methods and structural-based methods to achieve flexibility and efficiency. We demonstrate the accuracy, high quality, and run time efficiency of HyAC through experiments conducted on several benchmark circuits as well as a circuit from a commercial processor. 1. Xiaoliang Bai, Sujit Dey, Angela Krstic |
ITC | 2 |
| 2003 | LI-BIST: A Low-Cost Self-Test Scheme for SoC Logic Cores and Interconnects
Krishna Sekar, Sujit Dey |
J. Electron. Test. | 2 |
| 2003 | Fault-coverage analysis techniques of crosstalk in chip interconnectsabstractThis paper addresses the problem of evaluating the effectiveness of test sets to detect crosstalk defects in system-level interconnects and buses of deep submicron (DSM) chips. The fast and accurate estimation technique will enable: 1) evaluation of different existing tests, like functional, scan, logic built-in self-test (BIST), and delay tests, for effective testing of crosstalk defects in core-to-core interconnects and 2) development of crosstalk tests if the existing tests are not sufficient, thereby minimizing the cost of interconnect testing. Based on a covering relationship we distinguish between transition tests in detecting crosstalk defects and develop an abstract crosstalk fault model for chip interconnects. With this fault model and the covering relationship, we develop a fast and efficient method to estimate the fault coverage of any general test set. We also develop a simulation-based technique to calculate the probability of occurrence of the defects corresponding to each fault, which enables the fault-coverage analysis technique to produce accurate estimates of the actual crosstalk defect coverage of a given test set. The crosstalk test and fault properties, as well as the accuracy of the proposed crosstalk coverage analysis techniques, have been validated through extensive simulation experiments. The experiments also demonstrate that the proposed crosstalk techniques are orders of magnitude faster than the alternative method of SPICE-level simulation. Finally, we demonstrate the practical applicability of the proposed fault-coverage analysis technique by using it to evaluate the crosstalk fault coverage of logic BIST tests for the system-level interconnects and buses in a digital signal processor core. Sujit Dey |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2003 | High-level macro-modeling and estimation techniques for switching activity and power consumptionabstractWe present efficient techniques for estimating switching activity and power consumption at the register-transfer level (RTL), using a combination of macro-modeling for datapath blocks, and control logic analysis techniques based on partial delay information. Previous work on estimating switching activity and power at the RTL has ignored the presence of glitches at various datapath and control signals. We demonstrate that glitches can form a significant component of the switching activity at signals in typical RTL circuits. In particular, for control-flow intensive designs, we show that the controller substantially affects the activity and power consumption in the datapath due to the presence of glitches at control signals. Since the final implementation of the controller is not available during high-level design iterations, we develop techniques that estimate glitching activity at control signals using control expressions and partial delay information. For datapath blocks that operate on word-level data, we construct piecewise linear models that capture the variation of output glitching activity and power consumption with various word-level parameters like mean, standard deviation, spatial and temporal correlations, and glitching activity at the block's inputs. For RTL blocks that operate on bit vectors that need not have an associated word-level value, we present accurate bit-level modeling techniques for glitching activity as well as power consumption. This allows us to perform accurate power estimation for control-flow intensive circuits, where most of the power consumed is dissipated in non-arithmetic components like multiplexers, registers, vector logic operators, etc. Experimental results on several RTL designs demonstrate the accuracy of the proposed estimation techniques. Our RTL power estimator produced estimates that were within 7% of those produced by an in-house power analysis tool on the final gate-level implementation, while being over 50/spl times/ faster than its gate-level counterpart. Anand Raghunathan, Sujit Dey, Niraj K. Jha |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2002 | Software-based diagnosis for processorsabstractSoftware-based self-test (SBST) is emerging as a promising technology for enabling at-speed test of high-speed microprocessors using low-cost testers. We explore the fault diagnosis capability of SBST, in which functional information can be used to guide and facilitate the generation of diagnostic tests. By using a large number of carefully constructed diagnostic test programs, the fault universe can be divided into fine-grained partitions, each corresponding to a unique pass/fail pattern. We evaluate the quality of diagnosis by constructing diagnostic-tree-based fault dictionaries. We demonstrate the feasibility of the proposed method by applying it to a processor example. Experimental results show its potential as an effective method for diagnosing larger processors. Sujit Dey |
DAC | 2 |
| 2002 | Embedded software-based self-testing for SoC designabstractAt-speed testing of high-speed circuits is becoming increasingly difficult with external testers due to the growing gap between design and tester performance, growing cost of high-performance testers and increasing yield loss caused by inherent tester inaccuracy. Therefore, empowering the chip to test itself seems like a natural solution. Hardware-based self-testing techniques have limitations due to performance and area overhead and problems caused by the application of non-functional patterns.Embedded software-based self-testing has recently become focus of intense research. In this methodology, the programmable cores are used for on-chip test generation, measurement, response analysis and even diagnosis. After the programmable core on a System-on-Chip (SoC) has been self-tested, it can be reused for testing on-chip buses, interfaces and other non-programmable cores. The advantages of this methodology include at-speed testing, low design-for-testability overhead and application of functional patterns in the functional environment. In this paper, we give a survey and outline the roadmap and challenges of this emerging embedded software-based self-testing paradigm. Angela Krstic, Wei-Cheng Lai, Kwang-Ting Cheng, Sujit Dey |
DAC | 5 |
| 2002 | Communication architecture based power management for battery efficient system designabstractCommunication-based power management (CBPM) is a new battery-driven system-level power management methodology in which the system-level communication architecture regulates the execution of various system components, with the aim of improving battery efficiency, and hence, battery life. Unlike conventional power management policies (that attempt to efficiently shut down idle components), CBPM may delay the execution of selected system components even when they are active, in order to adapt the system-level current discharge profile to suit the battery's characteristics.In this paper, we present a methodology for the design of CBPM based systems, which consists of system-level performance and power profiling, battery discharge analysis, instrumentation of system components to facilitate CBPM, definition of CBPM policies, and generation of the CBPM-based system architecture. We present extensive evaluations of CBPM, and demonstrate its application to the design of an IEEE 801.11 Wireless LAN MAC processor system. Our results indicate that CBPM based systems are significantly more battery efficient than those based on conventional power management techniques. Further, we demonstrate that CBPM enables design-time as well as run-time tradeoffs between system performance and battery life. Kanishka Lahiri, Sujit Dey, Anand Raghunathan |
DAC | 2 |
| 2002 | Battery-efficient architecture for an 802.11 MAC processorabstractRapid growth in the complexity of wireless devices, communication protocols, and applications, combined with slow improvements in battery technologies, have created a "battery gap" that is only projected to increase with advances in wireless communication technologies and applications. Conventional approaches to bridging this gap exploit low-power network protocols and handset architectures. However, it is now well known that minimizing the total energy or average power drawn from a battery does not necessarily lead to maximizing battery life, calling for new battery-driven approaches to protocol and hardware design. We present a battery-efficient architecture for an 802.11 MAC processor, which incorporates a new battery-driven approach to power management. The MAC processor employs a novel on-chip bus architecture that is capable of regulating the profile of the current drawn by the system, enabling battery discharge at high efficiencies. The proposed battery friendly MAC processor architecture enables significant increases in battery capacity and lifetime, while minimizing performance impacts. Further, the developed architecture provides mechanisms that allow for trade-offs between battery life and performance, and can be configured to adapt the power management techniques based on the network traffic characteristics. Kanishka Lahiri, Anand Raghunathan, Sujit Dey |
ICC | 3 |
| 2002 | Adaptive and energy efficient wavelet image compression for mobile multimedia data servicesabstractTo enable wireless Internet and other data services using mobile appliances, there is a critical need to support content-rich cellular data communication, including voice, text, image and video. However, mobile communication of multimedia content has several bottlenecks, including limited bandwidth of cellular networks, channel noise, and battery constraints of the appliances. We address the energy and bandwidth bottlenecks of image data communication. We present an energy efficient, adaptive data codec for still images that can significantly minimize the energy required for wireless image communication, while meeting bandwidth constraints of the wireless network, the image quality, and latency constraints of the wireless service. Based on wavelet image compression, we propose an energy efficient wavelet image transform algorithm (EEWITA) for lossy compression of still images, enabling significant reductions in computation as well as communication energy needed, with minimal degradation an image quality. Additionally, we identify, wavelet image compression parameters that can be used to effect trade-offs between the energy savings, quality of the image, and required communication bandwidth. We also present a dynamic configuration methodology that selects the optimal set of parameters to minimize energy under network, service, and appliance constraints. We demonstrate the significant energy and air time (service cost) savings possible by using the proposed energy efficient, adaptive image codec under different cellular access technologies. Dong-Gi Lee, Sujit Dey |
ICC | 2 |
| 2002 | On-Line Testing of Multi-Source Noise-Induced Errors on the Interconnects and Buses of System-on-ChipsabstractWith processors and system-on-chips using nano-meter technologies, several design and test efforts have been recently developed to eliminate and test for many emerging DSM (deep sub-micron) noise effects. In this paper, we show the emergence of multi-source noise effects, where multiple DSM noise sources combine to produce functional and timing errors even when each separate noise source itself does not. We show the dynamic nature of multi-source noise, and the need for on-line testing to detect such noise errors. We propose a double-sampling data checking based low-cost on-line error detection circuit to test for such noise effects in on-chip buses. Based on the proposed circuit, an effective and efficient testing methodology has been developed to facilitate online testing for generic on-chip buses. The applicability of this methodology is demonstrated through embedding the on-line detection circuit in a bus design. The validated design shows the effectiveness of the proposed testing methodology for multi-source noise-induced errors in global interconnects and buses. Sujit Dey |
ITC | 3 |
| 2002 | Validation and Test of Network Processors and ASICsabstractRapid growth in the internet and enterprise network traffic, need by both subscribers and operators to provide for differentiated services, and the constant evolution of wireline and wireless networking functions, algorithms and standards, is leading to a high demand for ultra high speed, yet programmable, network processors and ASICs. This session will introduce the design of network processors, and highlight important validation and test challenges associated with network processors and ASICs. The first speaker in this session will review the architecture of a typical OC-192 network processor – a highly complex system-on-chip, consisting of multiple processors, specialized hardware blocks, multiple high-speed embedded memory units, and a highly concurrent and complex on-chip interconnect structure. The talk will analyze the validation and manufacturing test problems of hardware-software GHz network processor chips, which have to use aggressive architecture designs and nano-meter technologies to obtain the required multi-GHz speed, while relying on the presence of multiple processors to provide the necessary flexibility. The second talk will address testing of embedded GHz network processors. In particular, it will describe the challenges faced in testing a specific high-speed low-power network processor chip, which includes two custom-design 64-bit MIPS processors, 0.5MB of L2 cache, and multiple high speed communication blocks, such as gigabit-ethernet, hypertransport, and double-data-rate memory controllers. Some test solutions will be presented, and major test challenges that need to be addressed will be listed. The third talk in this session will address the problem of delay testing in communications chips. In communication devices such as network switches and routers, there usually are multiple unrelated clocks controlling various interfaces and logic. These multiple clock domains presents unique challenges in production testing since it becomes very difficult to predict when the expected output will show up at the pin boundary. The talk will show why delay defects need to be addressed, and how to deal with cycle uncertainty problem during production testing. It will present a novel compaction technique used to deal with the test data volume explosion resulting from including of delay defect screening tests. Proceedings of the 20 th IEEE VLSI Test Symposium (VTS02) 1093-0167/02 $17.00 © 2002 IEEE C.-H. Chia, Sujit Dey, Faraydon Karim, Haluk Konuk, Keesup Kim |
VTS | 2 |
| 2002 | LI-BIST: A Low-Cost Self-Test Scheme for SoC Logic Cores and InterconnectsabstractFor deep sub-micron system-on-chips (SoC), interconnects are critical determinants of performance, reliability and power Buses and long interconnects being susceptible to crosstalk noise, may lead to functional and timing failures. Existing at-speed interconnect crosstalk test methods are based on either (i) inserting dedicated interconnect selftest structures (leading to significant area overhead), or (ii) using existing logic BIST structures (e.g., LFSRs), which often result in poor defect coverage. Additionally, it has been shown that the power consumed during testing can potentially become a significant concerti. In this paper we present Logic-Interconnect BIST (LI-BIST), a comprehensive self-test solution for both the logic of the cores and the SoC interconnects. LI-BIST reuses existing LFSR structures but generates high-quality tests for interconnect crosstalk defects, while minimizing area overhead and interconnect power consumption. On applying LI-BIST to a DSP chip, we achieved crosstalk defect coverage of 99.7% for the interconnects and single stuck-at-fault coverage of 91.36% for the logic cores, while incurring an area overhead of only 4% over conventional logic BIST. Krishna Sekar, Sujit Dey |
VTS | 2 |
| 2002 | Testing for Interconnect Crosstalk Defects Using On-Chip Embedded Processor Cores
Xiaoliang Bai, Sujit Dey |
J. Electron. Test. | 3 |
| 2002 | Cosimulation-based power estimation for system-on-chip designabstractWe present efficient power estimation techniques for hardware-software (HW-SW) system-on-chip (SoC) designs. Our techniques are based on concurrent and synchronized execution of multiple power estimators that analyze different parts of the SoC (we refer to this as coestimation), driven by a system-level simulation master. We motivate the need for power coestimation, and demonstrate that performing independent power estimation for the various system components can lead to significant errors in the power estimates, especially for control-intensive and reactive-embedded systems. We observe that the computation time for performing power coestimation is dominated by: i) the requirement to analyze/simulate some parts of the system at lower levels of abstraction in order to obtain accurate estimates of timing and switching activity information and ii) the need to communicate between and synchronize the various simulators. Thus, a naive implementation of power coestimation may be too inefficient to be used in an iterative design exploration framework. To address this issue, we present several acceleration (speed-up) techniques for power coestimation. The acceleration techniques are energy caching, software power macro-modeling, and statistical sampling. Our speed-up techniques reduce the workload of the power estimators for the individual SoC components, as well as their communication/synchronization overhead. Experimental results indicate that the use of the proposed acceleration techniques results in significant (8/spl times/ to 87/spl times/) speed-ups in SOC power estimation time, with minimal impact on accuracy. We also show the utility of our coestimation tool to explore system-level power tradeoffs for a TCP/IP check-sum engine subsystem. Marcello Lajolo, Anand Raghunathan, Sujit Dey, Luciano Lavagno |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2001 | Testing for Interconnect Crosstalk Defects Using On-Chip Embedded Processor CoresabstractCrosstalk effects degrade the integrity of signals traveling on long interconnects and must be addressed during manufacturing testing. External testing for crosstalk is expensive due to the need for high-speed testers. Built-in self-test, while eliminating the need for a high-speed tester, may lead to excessive test overhead as well as overly aggressive testing. To address this problem, we propose a new software-based self-test methodology for system-on-chips (SoC) based on embedded processors. It enables an on-chip embedded processor core to test for crosstalk in system-level interconnects by executing a self-test program in the normal operational mode of the SoC. We have demonstrated the feasibility of this method by applying it to test the interconnects of a processor-memory system. The defect coverage was evaluated using a system-level crosstalk defect simulation method. Xiaoliang Bai, Sujit Dey |
DAC | 3 |
| 2001 | On-Chip Communication Architecture for OC-768 Network ProcessorsabstractThe need for network processors capable of forwarding IP packets at OC-192 and higher data rates has been well established. At the same time, there is a growing need for complex tasks, like packet classification and differentiated services, to be performed by network processors. At OC-768 data rate, a network processor has 9 nanoseconds to process a minimum-size IP packet. Such ultra high-speed processing, involving complex memory-intensive tasks, can only be achieved by multi-CPU distributed memory systems, using very high performance on-chip communication architectures. In this paper, we propose a novel communication network architecture for 8-CPU distributed-memory systems that has the potential to deliver the throughput required in next generation routers. We then show that our communication architecture can easily scale to accommodate much greater number of network nodes. Our network architecture yields higher performance than the traditional bus and crossbar yet has low implementation cost. It is quite flexible and can be implemented in either packet or circuit switched mode. We will compare and contrast our proposed architecture with busses and crossbars using metrics such as throughput and physical layout cost. Faraydon Karim, Sujit Dey, Ramesh R. Rao |
DAC | 3 |
| 2001 | Modeling and Minimization of Interconnect Energy Dissipation in Nanometer TechnologiesabstractAs the technology sizes of semiconductor devices continue to decrease, the effect of nanometer technologies on interconnects, such as crosstalk glitches and timing variations, become more significant. In this paper, we study the effect of nanometer technologies on energy dissipation in interconnects. We propose a new power estimation technique which considers DSM effects, resulting in significantly more accurate energy dissipation estimates than transition-count based methods for on-chip interconnects. We also introduce an energy minimization technique which attempts to minimize large voltage swings across the cross-coupling capacitances between interconnects. Even though the number of transitions may increase, our method yields a decrease in power consumption of up to 50%. Clark N. Taylor, Sujit Dey |
DAC | 2 |
| 2001 | Adaptive image compression for wireless multimedia communicationabstractTo enable ubiquitous wireless multimedia communication, the bottlenecks to communicating multimedia data over wireless channels must be addressed. Two significant bottlenecks which need to be overcome are the bandwidth and energy consumption requirements for mobile multimedia communication. In this paper, we address the bandwidth and energy dissipation bottlenecks by adapting the image compression parameters to current communication conditions and constraints. We focus on the JPEG image compression algorithm, and present the results of varying some image compression parameters on energy dissipation, bandwidth required, and quality of image received. We present a methodology for selecting the JPEG image compression parameters in order to minimize energy consumption while meeting latency, bandwidth, and quality of image constraints. Clark N. Taylor, Sujit Dey |
ICC | 2 |
| 2001 | High-level Crosstalk Defect Simulation for System-on-Chip InterconnectsabstractFor system-on-chip (SoC) devices using deep submicron (DSM) technologies, the interconnects are becoming critical determinants for performance and reliability. Buses and long interconnects are susceptible to crosstalk defects and may lead to functional and timing failure. Hence, testing for crosstalk errors on interconnects and buses in a SoC has become critical. To facilitate development of new crosstalk test methodologies and to efficiently evaluate crosstalk defect coverage for existing tests there is a need for efficient crosstalk defect coverage analysis techniques. In this paper, we present an efficient high-level crosstalk defect simulation methodology. By using a novel high-level DSM error model for the interconnects, together with HDL models for the cores, our methodology enables fast crosstalk defect simulation to be conducted at high level. We validate the high-level interconnect DSM error model by comparing its outputs with HSPICE simulation results. The fast and accurate high-level crosstalk defect simulation methodology will enable evaluation and exploration of new crosstalk test techniques, as well as existing tests, leading to the development of a low-cost crosstalk test. Xiaoliang Bai, Sujit Dey |
VTS | 2 |
| 2001 | Software-based self-testing methodology for processor coresabstractAt-speed testing of gigahertz processors using external testers may not be technically and economically feasible. Hence, there is an emerging need for low-cost high-quality self-test methodologies that can be used by processors to test themselves at-speed. Currently, built-in self-test (BIST) is the primary self-test methodology available. While memory BIST is commonly used for testing embedded memory cores, complex logic designs such as microprocessors are rarely tested with logic BIST. In this paper, we first analyze the issues associated with current hardware-based logic-BIST methodologies by applying a commercial logic-BIST tool to two processor cores. We then propose a new software-based self-testing methodology for processors, which uses a software tester embedded in the processor memory as a vehicle for applying structural tests. The software tester consists of programs for test generation and test application. Prior to the test, structural tests are prepared for processor components in the form of self-test signatures. During the process of self-test, the test generation program expands the self-test signatures into test sets and the test application program applies the tests to the components under test at the speed of the processor. Application of the novel software-based self-test method demonstrates its significant cost/fault coverage benefits and its ability to apply at-speed test while alleviating the need for high-speed testers. Sujit Dey |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2001 | System-level performance analysis for designing on-chipcommunication architecturesabstractThis paper presents a novel system-level performance analysis technique to support the design of custom communication architectures for system-on-chip integrated circuits. Our technique fills a gap in existing techniques for system-level performance analysis, which are either too slow to use in an iterative communication architecture design framework (e.g., simulation of the complete system) or are not accurate enough to drive the design of the communication architecture (e.g., techniques that perform a "static" analysis of the system performance). Our technique is based on a hybrid trace-based performance-analysis methodology in which an initial cosimulation of the system is performed with the communication described in an abstract manner (e.g., as events or abstract data transfers). An abstract set of traces are extracted from the initial cosimulation containing necessary and sufficient information about the computations and communications of the system components. The system designer then specifies a communication architecture by: 1) selecting a topology consisting of dedicated as well as shared communication channels (shared buses) interconnected by bridges; 2) mapping the abstract communications to paths in the communication architecture; and 3) customizing the protocol used for each channel. The traces extracted in the initial step are represented as a communication analysis graph (CAG) and an analysis of the CAG provides an estimate of the system performance as well as various statistics about the components and their communication. Experimental results indicate that our performance-analysis technique achieves accuracy comparable to complete system simulation (an average error of 1.88%) while being over two orders of magnitude faster. Kanishka Lahiri, Anand Raghunathan, Sujit Dey |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2000 | Self-test methodology for at-speed test of crosstalk in chip interconnectsabstractThe effect of crosstalk errors is most significant in high-performance circuits, mandating at-speed testing for crosstalk defects. This paper describes a self-test methodology that we have developed to enable on-chip at-speed testing of crosstalk defects in System-on-Chip interconnects. The self-test methodology is based on the Maximal Aggressor Fault Model [13], that enables testing of the interconnect with a linear number of test patterns. To enable self-testing of the interconnects, we have designed efficient on-chip test generators and error detectors to be embedded in necessary cores; while the test generators generate test vectors for crosstalk faults, the error detectors analyze the transmission of the test sequences received from the interconnects, and detect any transmission errors. We have also designed test controllers to initiate and manage test transactions by activating the appropriate test generators and error detectors, and having error diagnosis capability. We have developed, simulated, and synthesized parameterized HDL models of the self-test structures. We have applied the self-test methodology to test crosstalk defects in the buses of a DSP chip. Using a new high-level crosstalk simulation technique, we have validated the self-test methodology, including the self-test structures inserted in the DSP chip. Xiaoliang Bai, Sujit Dey, Janusz Rajski |
DAC | 2 |
| 2000 | Embedded hardware and software self-testing methodologies for processor coresabstractAt-speed testing of GHz processors using external testers may not be technically and economically feasible. Hence, there is an emerging need for low-cost, high-quality self-test methodologies, which can be used by processors to test themselves at-speed. Currently, Built-In Self-Test (BIST) is the primary self-test methodology available and is widely used for testing embedded memory cores. In this paper, we report our experiences in applying a commercial BIST methodology to two processor cores and analyze the problems associated with the current hardware-based BIST methodologies. We propose a new software-based self-testing methodology for processors, which uses a software tester embedded in the processor memory as a vehicle for applying structural tests. The software tester consists of programs for test generation and test application. Prior to the test, structural tests are prepared for processor components in the form of self-test signatures. During the process of self-test, the test generation program expands the self-test signatures into test sets, and the test application program applies the tests to the components-under-test at the speed of the processor. Application of the novel software-based self-test method demonstrates its significant cost/fault coverage benefits and its ability to apply at-speed test while alleviating the need for high-speed testers. Sujit Dey, Pablo Sanchez, Krishna Sekar |
DAC | 2 |
| 2000 | Test challenges for deep sub-micron technologiesabstractThe use of deep submicron process technologies presents several new challenges in the area of manufacturing test. While a significant body of work has been devoted to identifying and investigating design challenges in nanometer technologies, the impact on test strategies and methodologies is still not well understood. This paper highlights the challenges to current test methodologies arising from technology driven trends, and will present an overview of emerging techniques that address deep submicron test challenges. Kwang-Ting Cheng, Sujit Dey, Mike Rodgers, Kaushik Roy 0001 |
DAC | 2 |
| 2000 | Communication architecture tuners: a methodology for the design of high-performance communication architectures for systems-on-chipsabstractIn this chapter, we present a general methodology for the design of custom system-on-chip communication architectures. Our technique is based on the addition of a layer of circuitry, called the Communication Architecture Tuner (CAT), around any existing communication architecture topology. The added layer enhances the ability of the system to adapt to changing communication needs of its constituent components. For example, more critical data may be handled differently, leading to lower communication latencies. The CAT monitors the internal state and communication transactions of each component, and “predicts” the relative importance of each communication transaction in terms of its potential impact on different system-level performance metrics. It then configures the protocol parameters of the underlying communication architecture (e.g., priorities, DMA modes,etc.) to best suit the system's changing communication needs. Kanishka Lahiri, Anand Raghunathan, Ganesh Lakshminarayana, Sujit Dey |
DAC | 4 |
| 2000 | Efficient Power Co-Estimation Techniques for System-on-Chip DesignabstractWe present efficient power estimation techniques for HW/SW System-On-Chip (SOC) designs. Our techniques are based on concurrent and synchronized execution of multiple power estimators that analyze different parts of the SOC (we refer to this as co-estimation), driven by a system-level simulation master. We motivate the need for power co-estimation, and demonstrate that performing independent power estimation for the various system components can lead to significant errors in the power estimates, especially for control-intensive and reactive embedded systems. We observe that the computation time for performing power co-estimation is dominated by: (i) the requirement to analyze/simulate some parts of the system at lower levels of abstraction in order to obtain accurate estimates of timing and switching activity information and (ii) the need to communicate between and synchronize the various simulators. Thus, a naive implementation of power co-estimation may be too inefficient to be used in an iterative design exploration framework. To address this issue, we present several acceleration (speedup) techniques for power co-estimation. The acceleration techniques are energy caching, software power macromodeling, and statistical sampling. Our speedup techniques reduce the workload of the power estimators for the individual SOC components, as well as their communication/synchronization overhead. Experimental results indicate that the use of the proposed acceleration techniques results in significant (8/spl times/ to 87/spl times/) speedups in SOC power estimation time, with minimal impact on accuracy. We also show the utility of our co-estimation tool to explore system-level power tradeoffs for a TCP/IP network interface card sub-system and an automotive controller. Marcello Lajolo, Anand Raghunathan, Sujit Dey, Luciano Lavagno |
DATE | 3 |
| 2000 | Efficient Exploration of the SoC Communication Architecture Design SpaceabstractIn this paper, we present a methodology and efficient algorithms for the design of high-performance system-on-chip communication architectures. Our methodology automatically and optimally maps the various communications between system components onto a target communication architecture template that can consist of an arbitrary interconnection of shared or dedicated channels. In addition, our techniques simultaneously configure the communication protocols of each channel in the architecture in order to optimize system performance. We motivate the need for systematic exploration of the communication architecture design space, and highlight the issues involved through illustrative examples. We present a methodology and algorithms that address these issues, including the size and complexity of the design space. We present experimental results on example systems, including a cell forwarding unit of an ATM switch, that demonstrate the benefits of using the proposed techniques. Experimental results indicate that our techniques are successful in achieving significant improvements in system performance over conventional communication architectures (observed speedups over typical architectures such as single shared buses averaged 53%). Moreover, we demonstrate that our design space exploration methodology and optimization algorithms are efficient (low CPU times), underlining their usefulness as part of any system design flow. Kanishka Lahiri, Anand Raghunathan, Sujit Dey |
ICCAD | 3 |
| 2000 | Test of Future System-on-ChipsabstractSpurred by technology leading to the availability of millions of gates per chip, system-level integration is evolving as a new paradigm, allowing entire systems to be built on a single chip. Being able to rapidly develop, manufacture, test, debug and verify complex SOCs is crucial for the continued success of the electronics industry. This growth is expected to continue full force at least for the next decade, while making possible the production of multimillion transistor chips. However, to make its production practical and cost effective the industry road maps identify a number of major hurdles to be overcome. The key hurdle is related to test and diagnosis. This embedded tutorial analyzes these hurdles, relates them to the advancements in semiconductor technology and presents potential solutions to address them. These solutions are meant to ensure that test and diagnosis contribute to the overall growth of the SOC industry and do not slow it down. This embedded tutorial in addition presents the state of the art in system-level integration and addresses the strategies and current industrial practices in the test of system-on-chip. It discusses the requirements for test reuse in hierarchical design, such as embedded test strategies for individual cores, test access mechanisms, optimizing test resource partitioning, and embedded test management and integration at the System-on-Chip level. Processor cores being one of the most common cores embedded in a SOC, issues related to self-testing embedded processor cores are addressed. Future research challenges and opportunities are discussed in enabling testing of future SOCs which use deep submicron technologies. Yervant Zorian, Sujit Dey, Mike Rodgers |
ICCAD | 2 |
| 2000 | Analysis of interconnect crosstalk defect coverage of test setsabstractThis paper addresses the problem of evaluating the effectiveness of test sets to detect crosstalk defects in interconnects of deep sub-micron circuits. The fast and accurate estimation technique will enable: (a) evaluation of different existing tests, like functional, scan, logic BIST, and delay tests, for effective testing of crosstalk defects in interconnects, and (b) development of crosstalk tests if the existing tests are not sufficient, thereby minimizing the cost of interconnect testing. Based on a covering relationship we establish between transition tests in detecting crosstalk defects, we develop an abstract crosstalk fault model for circuit interconnects. Based on this fault model, and the covering relationship, we develop a fast and efficient method to estimate the fault coverage of any general test set. We also develop a simulation-based technique to calculate the probability of occurrence of the defects corresponding to each fault, which enables the fault coverage analysis technique to produce accurate estimates of the actual crosstalk defect coverage of a given test set. The crosstalk test and fault properties, as well as the accuracy of the proposed crosstalk coverage analysis techniques, have been validated through extensive simulation experiments. The experiments also demonstrate that the proposed crosstalk techniques are orders of magnitude faster than the alternative method of SPICE-level simulation. Finally, we demonstrate the practical applicability of the proposed fault coverage analysis technique by using it to evaluate the crosstalk fault coverage of logic BIST tests for the buses in a DSP core. Sujit Dey |
ITC | 2 |
| 2000 | DEFUSE: A Deterministic Functional Self-Test Methodology for ProcessorsabstractAt-speed testing is becoming increasingly difficult with external testers as the speed of microprocessors approaches the GHz range. One solution to this problem is built-in self-test. However, due to their reliance on random patterns, current logic BIST techniques are not able to deal with large designs without adding high test overhead. In this paper, we propose a functional self-test technique that is deterministic in nature. By targeting the structural test need of manageable components with the aid of processor functionality, this technique has the fault coverage advantage of deterministic structural testing and the at-speed advantage of functional testing. Most importantly, by relieving testers from test application, it enables at-speed testing of GHz processors with low speed testers. We have demonstrated our methodology on a simple accumulator-based microprocessor. The results show that with the proposed technique, we are able to apply high-quality at-speed tests with no test overhead. Sujit Dey |
VTS | 2 |
| 2000 | A fast and low-cost testing technique for core-based system-chipsabstractThis paper proposes a new methodology for testing a core-based system chip, targeting the simultaneous reduction of test area overhead and test application time. At the core level, testability and transparency can be achieved by the core provider by reusing existing logic inside the core, providing different versions of the core having different area overheads and transparency latencies. The technique analyzes the topology of the system-chip to select the core versions that best meet the user's desired test area overhead and test application time objectives. Application of the method to example system-chips demonstrates the ability to design highly testable system-chips with minimized test area overhead, minimized test application time, or a desired tradeoff between the two. Significant reduction in area overhead and test application time compared to existing system chip testing techniques is also demonstrated. Indradeep Ghosh, Sujit Dey, Niraj K. Jha |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1999 | Common-Case Computation: A High-Level Technique for Power and Performance OptimizationabstractThis paper presents a design methodology, called common-case computation (CCC), and new design automation algorithms for optimizing power consumption or performance. The proposed techniques are applicable in conjunction with any high-level design methodology where a structural register-transfer level (RTL) description and its corresponding scheduled behavioral (cycle-accurate functional RTL) description are available. It is a well-known fact that in behavioral descriptions of hardware (also in software), a small set of computations (CCCs) often accounts for most of the computational complexity. However, in hardware implementations (structural RTL or lower level), CCCs and the remaining computations a typically treated alike. This paper shows that identifying and exploiting CCCs during the design process can lead to implementations that are much more efficient in terms of power consumption or performance. We propose a CCC-based high-level design methodology with the following steps: extraction of common-case behaviors and execution conditions from the scheduled description, simplification of the common-case behaviors in a stand-alone manner, synthesis of common-case detection and execution circuits from the common-case behaviors, and composing the original design with the common-case circuits, resulting in a CCC-optimized design. We demonstrate that CCC-optimized designs reduce power consumption by up to 91.5%, or improve performance by up to 76.6% compared to designs derived without special regard for CCCs. Ganesh Lakshminarayana, Anand Raghunathan, Kamal S. Khouri, Niraj K. Jha, Sujit Dey |
DAC | 5 |
| 1999 | Fault modeling and simulation for crosstalk in system-on-chip interconnectsabstractSystem-on-chips (SOCs) using ultra deep sub-micron (DSM) technologies and GHz clock frequencies have been predicted by the 1997 SIA Road Map. Recent studies, as well as experiments reported in this paper, show significant crosstalk effects in long on-chip interconnects of GHz DSM chips. Recognizing the importance of high-speed, reliable interconnects in GHz SOCs, we address in this paper the problem of testing for glitch and delay errors caused by crosstalk in buses and interconnects between components of a SOC. Since it is not possible to explicitly test for all the possible process variations and defects that can lead to crosstalk errors in SOC interconnects, we present an abstract model, Maximum Aggressor (MA) fault model, and its test requirements. The attractiveness of the model is that it can abstract crosstalk defects in interconnects with a linear number of faults, while the corresponding MA tests provide complete coverage for all level defects related to cross-coupling capacitance the interconnects. A SPICE-level fault simulation methodology is presented which allows simulation of a small subset of the potentially exponential number of defects. The simulation methodology also enables validation of the proposed fault model and the resulting test set. Michael Cuviello, Sujit Dey, Xiaoliang Bai |
ICCAD | 2 |
| 1999 | Fast performance analysis of bus-based system-on-chip communication architecturesabstractThis paper addresses the problem of efficient and accurate performance analysis to drive the exploration and design of bus-based system-on-chip (SOC) communication architectures. Our technique fills a gap in existing techniques for system-level performance analysis, which are either too slow to use in an iterative communication architecture design framework (e.g., simulation of the complete system), or are not accurate enough to drive the design of the communication architecture (e.g., techniques that perform a static analysis of the system performance). The proposed system-level performance analysis technique consists of: initial co-simulation performed after HW/SW partitioning and mapping, with the communication between components modeled in an abstract manner (e.g., as events or data transfers); extraction of abstracted symbolic traces, represented as a bus and synchronization event (BSE) graph, that captures the activity of the various system components and their communication over time; and manipulation of the BSE graph using the bus parameters, to derive the behavior of the system accounting for effects of the bus architecture. We present experimental results on several example systems, including a TCP/IP network interface card sub-system. The results indicate that our performance estimation technique is over two orders of magnitude faster than performing a complete system simulation, while being very accurate (within 2.2% of performance estimates derived from accurate HW/SW co-simulation). Kanishka Lahiri, Anand Raghunathan, Sujit Dey |
ICCAD | 3 |
| 1999 | Resynthesis and retiming for optimum partial scanabstractAn effective partial scan approach selects flip-flops (FPs) in the minimum feedback vertex set (MFVS) of the FF dependency graph, so that all loops, except self-loops, are broken. However, the MFVS of the circuit (the minimum number of gates whose removal makes the circuit acyclic) is a lower bound and in many cases, significantly smaller than the MFVS of the FF dependency graph. Since only FFs can be considered for scan, this paper investigates the possibility of repositioning FFs so that, in the modified circuit, every circuit MFVS gate drives at least one FF that can be scanned. We show that resynthesis and retiming can always transform any circuit into an equivalent circuit whose FF dependency graph MFVS is equal to the MFVS of the original circuit. Therefore, the MVFS of a circuit is a tight lower bound on the number of scan FFs needed. We first identify the necessary and sufficient conditions under which legal retiming can produce the desired FF repositioning. We show that circuits that do not satisfy these conditions can always be suitably modified using two new resynthesis transformations. The modified circuit can always be retimed to achieve the desired FF repositioning. Experimental results for several large sequential benchmarks show that the number of scan FFs required for the resynthesized and retimed circuit is significantly smaller than that required for the original circuit. Srimat T. Chakradhar, Sujit Dey |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1999 | Controller-based power management for control-flow intensive designsabstractThis paper presents a low-overhead controller-based power management technique that redesigns control logic to reconfigure the existing data path components under idle conditions so as to minimize unnecessary activity. Controller-based power management exploits the fact that though the control signals in a register-transfer level implementation are fully specified, they can be respecified under certain states/conditions when the data path components that they control need not be active. We demonstrate that controller-based power management is often better-suited to control-flow intensive designs than comparable conventional power management techniques such as operand isolation. We present an algorithm to perform power management through controller redesign that consists of constructing an activity graph for each data path component, identifying conditions under which the component need not be active, and relabeling the activity graph resulting in redesign of the corresponding control expressions. We provide a comprehensive analysis of the potential side effects of controller-based power management on circuit delay, glitching activity at control and data path signals, and formation of false combinational cycles. Our algorithm avoids the above negative effects of controller-based power management to maximize power savings and minimize overheads. We present experimental results which demonstrate that (1) controller-based power management results in large power savings at minimal overheads for control-flow intensive designs, which pose several challenges to conventional power management techniques and (2) it is important to consider the various potential negative effects while performing controller-based power management in order to obtain maximal power savings. Sujit Dey, Anand Raghunathan, Niraj K. Jha, Kazutoshi Wakabayashi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1999 | A low overhead design for testability and test generation technique for core-based systems-on-a-chipabstractIn a fundamental paradigm shift in system design, entire systems are being built on a single chip, using multiple embedded cores. Though the newest system design methodology has several advantages in terms of time-to-market and system cost, testing such core-based systems is difficult, mainly due to the problem of justifying test sequences at the inputs of a core embedded deep in the circuit and propagating test responses from the core outputs. In this paper, we first present a design for testability technique for testing such core-based systems. In this scheme, untestable cores are first made testable using hierarchical testability analysis techniques. If necessary, additional testability hardware is added to the cores to make them transparent so that they can propagate test data without information loss. This testability and transparency technique is currently applicable to cores of the following types: application-specific integrated circuits, application-specific programmable processors, and application-specific instruction processors. Other core types can be made testable and transparent using traditional techniques. The testable and transparent cores can then he integrated together with some system-level testability hardware to ensure justification of precomputed test sequences of each core from system primary inputs to the core inputs and propagation of test responses from core outputs to system primary outputs. Justification and propagation of test sequences are done at the system level by extending and suitably modifying the symbolic hierarchical testability analysis method that has been successfully applied to register-transfer level circuits. Since the testability analysis method is symbolic, the system test generation method is independent of the bit-width of the cores. The system-level test set is obtained as a byproduct of the testability analysis and insertion method without further search. The test methodology was applied to six example systems. Besides the proposed test method, the two methods that are currently used in the industry were also evaluated: (1) FScan-BScan, where each core is full-scanned, and system test is performed using boundary scan and (2) FScan-TBus, where each core is full-scanned, and system test is performed using a test bus. The experiments show that the proposed scheme has significantly lower area overhead, delay overhead, and test application time compared to FScan-BScan and FScan-TBus, without any compromise in the system fault coverage. Indradeep Ghosh, Niraj K. Jha, Sujit Dey |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 1999 | Register transfer level power optimization with emphasis on glitch analysis and reductionabstractWe present design-for-low-power techniques for register-transfer level (RTL) controller/data path circuits. We analyze the generation and propagation of glitches in both the control and data path parts of the circuit. In data-flow intensive designs, glitching power is primarily due to the chaining of arithmetic functional units. In control-flow intensive designs, on the other hand, multiplexer networks and registers dominate the total circuit power consumption, and the control logic can generate a significant amount of glitches at its outputs, which in turn propagate through the data path to account for a large portion of the glitching power in the entire circuit. Our analysis also highlights the relationship between the propagation of glitches from control signals and the bit-level correlation between data signals. Based on the analysis, we develop techniques that attempt to reduce glitching power consumption by minimizing propagation of glitches in the RTL circuit. Our techniques include restructuring multiplexer networks (to enhance data correlations and eliminate glitchy control signals), clocking control signals, and inserting selective rising/falling delays, in order to kill the propagation of glitches from control as well as data signals. In addition, we present a procedure to automatically perform the well-known power-reduction technique of clock gating through an efficient structural analysis of the RTL circuit, while avoiding the introduction of glitches on the clock signals. Application of the proposed power optimization techniques to several RTL circuits shows significant power savings, with negligible area and delay overheads. Anand Raghunathan, Sujit Dey, Niraj K. Jha |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1999 | Power management in high-level synthesisabstractIn this paper, we present a power-management methodology targeted toward high-level synthesis of data-dominated behavioral descriptions. It is founded on the observation that variable assignment can significantly affect power-management opportunities in the synthesized architecture, i.e., variable assignment determines whether or not spurious operations get executed by functional units in the architecture. We introduce perfectly power managed architectures, whose functional units do not execute any spurious operations. We present a variable assignment technique which, when used in high-level synthesis, produces architectures which are perfectly power-managed. Unlike many previously proposed power-management techniques, our method does not add latches or any other circuitry in front of functional units or registers and is, therefore, free of the attendant performance penalty. Experimental results indicate savings of up to 52.5% (average 23.0%) in power consumption over already power-optimized architectures. The area overheads due to our technique are also low and averaged 2.5% for our examples. Ganesh Lakshminarayana, Anand Raghunathan, Niraj K. Jha, Sujit Dey |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 1998 | Considering Testability during High-level Design (Embedded Tutorial)abstractConsidering testability during the early stages of the design flow can have several benefits, including significantly improved fault coverage, reduced test hardware overheads, and reduced design iteration times. This paper presents an overview of high level design methodologies that consider testability during the early (behavior and architecture) stages of the design flow, and their testability benefits. The topics reviewed include behavioral and RTL test synthesis approaches that generate easily testable implementations targeting ATPG (full and partial scan) and BIST methodologies, and techniques to use high-level information for ATPG. Sujit Dey, Anand Raghunathan, Rabindra K. Roy |
ASP-DAC | 1 |
| 1998 | A Fast and Low Cost Testing Technique for Core-Based System-on-ChipabstractThis paper proposes a new methodology for testing a core-based system-on-chip (SOC), targeting the simultaneous reduction of test area overhead and test application time. Testing of embedded cores is achieved using the transparency properties of surrounding cores. At the core level, testability and transparency can be achieved by reusing existing logic inside the core, and providing different versions of the core having different area overheads and transparency latencies. At the chip level, the technique analyzes the topology of the SOC to select the core versions that best meet the user's desired test area overhead and test application time objectives. Application of the method to example SOCs demonstrates the ability to design highly testable SOCs with minimized test area overhead, minimized test application time, or a desired trade-off between the two. Significant reduction in area overhead and test application time compared to an existing SOC testing technique is also demonstrated. Indradeep Ghosh, Sujit Dey, Niraj K. Jha |
DAC | 2 |
| 1998 | High-level design validation and testabstractNo abstract available. Sujit Dey, Jacob A. Abraham, Yervant Zorian |
ICCAD | 1 |
| 1998 | Transforming control-flow intensive designs to facilitate power managementabstractWe present techniques to transform scheduled descriptions of control-flow intensive designs to facilitate power management. We investigate the factors that inhibit the application of power management in synthesized RTL implementations. Based on these insights, we present transformation techniques based on the concepts of variable protection, variable re-naming and re-assignment, and limited controller state memory insertion that result in inherently powermanaged architectures. Our transformation techniques can be easily used in conjunction with any existing resource sharing algorithm or in the framework of existing high-level synthesis tools. Experimental results on control-flow intensive designs indicated reductions of upto 76.6% (35.6% on average) in power consumption at area overheads not exceeding 10.1% (1.1% on average) over already poweroptimized designs. Ganesh Lakshminarayana, Anand Raghunathan, Niraj K. Jha, Sujit Dey |
ICCAD | 4 |
| 1998 | Testing embedded-core based system chipsabstractAdvances in semiconductor process and design technology enable the design of complex system chips. Traditional IC design in which every circuit is designed from scratch and reuse is limited to standard-cell libraries, is more and more replaced by a design style based on embedding large reusable modules, the so-called cores. This core-based design poses a series of new challenges, especially in the domains of manufacturing test and design validation and debug. This paper provides an overview of current industrial practices as well as academic research in these areas. We also discuss industry-wide efforts by VSIA and IEEE P1500 and describe the challenges for future research. Yervant Zorian, Erik Jan Marinissen, Sujit Dey |
ITC | 3 |
| 1998 | Design for Testability Techniques at the Behavioral and Register-Transfer Levels
Sujit Dey, Anand Raghunathan, Kenneth D. Wagner |
J. Electron. Test. | 1 |
| 1998 | Controller Resynthesis for Testability Enhancement of RTL Controller/Data Path Circuits
Srivaths Ravi 0001, Indradeep Ghosh, Rabindra K. Roy, Sujit Dey |
J. Electron. Test. | 4 |
| 1998 | A controller redesign technique to enhance testability of controller-data path circuitsabstractWe study the effect of the controller on the testability of sequential circuits composed of controllers and data paths. We show that even when all the loops of the circuit have been broken by using scan flip-flops (FF's) and the control and data path parts are individually 100% testable, the composite circuit may not be easily testable by gate-level sequential automatic test pattern generation (ATPG). Analysis shows that a primary problem in test pattern generation of combined controller-data path circuits is the correlation of control signals due to implications imposed by the controller specification. A design-for-testability (DFT) technique is developed to redesign the controller such that the implications which may produce conflicts during test pattern generation are eliminated. The DFT technique involves adding extra control vectors to the controller. Experimental results show the ability of the controller DFT technique to produce highly testable controller-data path circuits, with nominal hardware overhead. Sujit Dey, Vijay Gangaram, Miodrag Potkonjak |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1998 | Effects of resource sharing on circuit delay: an assignment algorithm for clock period optimizationabstractThis paper analyzes the effect of resource sharing and assignment on the clock period of the synthesized circuit. The assignment phase assigns or binds operations of the scheduled behavioral description to a set of allocated resources. We focus on control-flow intensive descriptions, characterized by the presence of mutually exclusive paths due to the presence of nested conditional branches and loops. We show that clustering multiple operations in the same state of the schedule, possibly leading to chaining of functional units (FUs) in the RTL circuit, is an effective way to minimize the total number of clock cycles, and hence total execution time. We present an assignment algorithm that is particularly effective for such design styles by minimizing data chaining and hence the clock period of the circuit, thereby leading to further reduction in total execution time. Existing resource sharing and assignment approaches for reducing the clock period of the resulting circuit either increase the resource allocation or use faster modules, both leading to leading to larger area requirements. In this paper we show that even when the type of available resource units and the number of resource units of each type is fixed, different assignments may lead to circuits with significant differences in clock period. We provide a comprehensive analysis of how resource sharing and assignment introduces long paths in the circuit. Based on the analysis, we develop an assignment algorithm that uses a high-level delay estimator to asign operations to a fixed set of available resources so as to minimize the clock period of the resultant circuit, with no or minimal effect on the area of the circuit. Experimental results on several conditional-intensive designs demonstrate the effectiveness of the assignment algorithm. Subhrajit Bhattacharya, Sujit Dey, Franc Brglez |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 1997 | Power Management Techniques for Control-Flow Intensive DesignsabstractThis paper presents a low-overhead controller-based powermanagement technique that re-specifies control signals to reconfigureexisting multiplexer networks and functional units to minimizeunnecessary activity. We demonstrate that conventional powermanagement techniques may often not be suited to control-flowintensive designs, and provide a comprehensive analysis of thepotential negative effects of power management on circuit delay,glitching activity at control and data path signals, and formationof false combinational cycles. We present techniques to performpower management through controller re-specification while avoidingthe above negative effects, and demonstrate the efficiency ofthe techniques through experiments. Anand Raghunathan, Sujit Dey, Niraj K. Jha, Kazutoshi Wakabayashi |
DAC | 2 |
| 1997 | Performance analysis of a system of communicating processesabstractEfficient exploration of the system design space necessitates fast and accurate performance estimation as opposed to the computationally prohibitive alternative of exhaustive simulation. The paper addresses the issue of worst case performance analysis of a system described as a set of concurrent communicating processes. We show that the synchronization overhead associated with inter process communication can contribute significantly to the overall system performance. Application of existing performance analysis techniques, which target single process descriptions, lead to inaccurate performance estimates as the synchronization overhead is not accounted for. We present PERC, a fast and accurate worst case performance analysis technique which analyzes inter process communication, and accounts for synchronization overhead while computing the worst case performance estimate of a given system implementation. Application of PERC to example systems described as multiple communicating processes shows the ability of the proposed method to accurately estimate the worst case performance of the system implementation. Sujit Dey, Surendra Bommu |
ICCAD | 1 |
| 1997 | H-SCAN+: A Practical Low-Overhead RTL Design-for-Testability Technique for Industrial DesignsabstractH-SCAN (1996) was presented as a low overhead design-for-testability strategy which is applicable to RT-level controller-data path circuits. However, from the view-point of practical use, there is a possibility that the area overhead of H-SCAN is larger than that of full-scan. Moreover, H-SCAN is unable to handle many features present in actual designs. In this paper, we propose a modified H-SCAN scheme, called "H-SCAN+", as an improved solution for actual designs. H-SCAN+ consists of several enhancements, including techniques to minimize scan design area overhead, handling of features present in actual designs, and techniques to significantly minimize the running time. We provide comprehensive results of applying H-SCAN+ to several actual RT-level designs. Toshiharu Asaka, Masaaki Yoshida, Subhrajit Bhattacharya, Sujit Dey |
ITC | 4 |
| 1997 | A Low-Overhead Design for Testability and Test Generation Technique for Core-Based SystemsabstractIn a fundamental paradigm shift in system design, entire systems are being built on a single chip, using multiple embedded cores. Though the newest system design methodology has several advantages in terms of time-to-market and system cost, testing such core-based systems is difficult due to the problem of justifying test sequences at the inputs of a core embedded deep in the system, and propagating test responses from the core outputs. In this paper, we present a design for testability and symbolic test generation technique for testing such core-based systems on a chip. The proposed method consists of two parts: (i) core-level DFT to make each core testable and transparent, the latter needed to propagate test data through the cores, and (ii) system-level DFT and test generation to ensure the justification and propagation of the precomputed test sequences and test responses of the core. Since the hierarchical testability analysis technique used to tackle the above problem is symbolic, the system test generation method is independent of the bit-width of the cores. The system-level test set is obtained as a by-product of the testability analysis and insertion method without further search. Besides the proposed test method, the two methods that are currently used in the industry were also evaluated on two example systems: (i) FScan-BScan, where each core is full-scanned, and system test is performed using boundary scan, and (ii) FScan-TBus, where each core is full-scanned, and system test is performed using a test bus. The experiments show that the proposed scheme has significantly lower area overhead, delay overhead, and test application time compared to FScan-BScan and FScan-TBus, without any compromise in the system fault coverage. Indradeep Ghosh, Niraj K. Jha, Sujit Dey |
ITC | 3 |
| 1997 | Nonscan design-for-testability techniques using RT-level design informationabstractThis paper presents nonscan design-for-testability (DFT) techniques applicable to register-transfer (RT)-level data path circuits. Knowledge of high-level design information, in the form of the RT-level structure, as well as the functions of the RT-level components is utilized to develop effective nonscan DFT techniques. Instead of conventional techniques of selecting flip-flops (FF's) to make systems controllable/observable, execution units (EXU's) are selected using the EXU S-graph introduced in this paper. Controllability/observability points can be implemented using register files and constants. We introduce the notion of k-level controllable and observable loops and demonstrate that it suffices to make all the loops k-level controllable/observable, k>0, to achieve very high test efficiency. The new testability measure eliminates the need by traditional DFT techniques to make all loops directly (zero-level) controllable/observable, reducing significantly the hardware overhead required and making the nonscan DFT approach feasible and effective. We discuss ways of avoiding the formation of reconvergent regions while adding test points to make loops k-level controllable/observable. We introduce dual points, which utilize the different controllability/observability levels of loops, to make one loop controllable while making another loop observable. We present efficient algorithms to add the minimal hardware possible to make all loops in the data path k-level controllable/observable, without the use of scan FF's. The nonscan DFT techniques were applied to several data path circuits. The experimental results demonstrate the effectiveness of the k-level testability measure, and the use of distributed and dual points, to generate easily testable data paths with reduced hardware overhead. The hardware overhead and the test application time required for the nonscan designs are significantly lower than for the partial scan designs. Most significantly, the experimental results demonstrate the ability of the RT-level DFT techniques to produce nonscan testable data paths, which can be tested at-speed. Sujit Dey, Miodrag Potkonjak |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1996 | Glitch Analysis and Reduction in Register Transfer LevelabstractWe present design-for-low-power techniques based on glitch reduction for register-transfer level circuits.We analyze the generation and propagation of glitches in both the control and data path parts of the circuit.Based on the analysis, we develop techniques that attempt to reduce glitching power consumption by minimizing generation and propagation of glitches in the RTL circuit.Our techniques include restructuring multiplexer networks (to enhance data correlations, eliminate glitchy control signals, and reduce glitches on data signals), clocking control signals, and inserting selective rising/falling delays.Our techniques are suited to control-flow intensive designs, where glitches generated at control signals have a significant impact on the circuit's power consumption, and multiplexers and registers often account for a major portion of the total power.Application of the proposed techniques to several examples shows significant power savings, with negligible area and delay overheads. Anand Raghunathan, Sujit Dey, Niraj K. Jha |
DAC | 2 |
| 1996 | High-Level Synthesis for Testability: A Survey and PerspectiveabstractWe review behavioral and RTL test synthesis and synthesis for testability approaches that generate easily testable implementations. We also include an overview of high-level synthesis techniques to assist high-level ATPG. Kenneth D. Wagner, Sujit Dey |
DAC | 2 |
| 1996 | Register-transfer level estimation techniques for switching activity and power consumptionabstractWe present techniques for estimating switching activity and power consumption in register-transfer level (RTL) circuits. Previous work on this topic has ignored the presence of glitching activity at various data path and control signals, which can lead to significant underestimation of switching activity. For data path blocks that operate on word-level data, we construct piecewise linear models that capture the variation of output glitching activity and power consumption with various word-level parameters like mean, standard deviation, spatial and temporal correlations, and glitching activity at the block's inputs. For RTL blocks that operate on data that need not have an associated word-level value, we present accurate bit-level modeling techniques for glitching activity as well as power consumption. This allows us to perform accurate power estimation for control-flow intensive circuits, where most of the power consumed is dissipated in non-arithmetic components like multiplexers, registers, vector logic operators, etc. Since the final implementation of the controller is not available during high-level design iterations, we develop techniques that estimate glitching activity at control signals using control expressions and partial delay information. Experiments on example RTL designs resulted in power estimates that were within 7% of those produced by an inhouse power analysis tool on the final gate-level implementation. Anand Raghunathan, Sujit Dey, Niraj K. Jha |
ICCAD | 2 |
| 1996 | Controller re-specification to minimize switching activity in controller/data path circuitsabstractThis paper proposes a controller-based technique for minimizing switching activity in controller/data path circuits. Though the control signals in a register transfer level (RTL) implementation are fully specified, they can be respecified under certain states/conditions when the data path components that they control need not be active. Unlike techniques that insert extra circuitry like transparent latches, controller re-specification is a low-overhead technique that merely reconfigures existing multiplexer networks and functional units to minimize activity in the data path. Hence, it is well suited to control-flow intensive designs, where power consumption in multiplexer networks forms a major component of the total power consumption. Our controller re-specification algorithm consists of constructing an activity graph for each data path component, identifying conditions under which the component need not be active, and re-labeling the activity graph resulting in re-specification of the corresponding control expressions. Application of the proposed technique to several RTL circuits demonstrated the ability to reduce the total (controller+data path) power consumption by up to 51.8% compared to the initial area-optimized implementations, with nominal area and delay overheads. Anand Raghunathan, Sujit Dey, Niraj K. Jha, Kazutoshi Wakabayashi |
ISLPED | 2 |
| 1996 | H-SCAN: A high level alternative to full-scan testing with reduced area and test application overheadsabstractThis paper presents H-SCAN, a practical testing methodology that can be easily applied to a high-level design specification. H-SCAN allows the use of combinational test patterns without the high area and test application time overheads associated with full-scan testing. Connectivities between registers existing in an RT-level design are exploited to reduce the area overhead associated with implementing a scan scheme. Test application time is significantly reduced by using the parallelism inherent in the design, and eliminating the pin constraint of parallel scan schemes by analyzing the test responses on-chip using existing comparators. The proposed method also includes generating appropriate sequential test vectors from combinational test vectors generated by a combinational ATPG program. Application of H-SCAN to RT-level designs and fault simulation using the test patterns generated by H-SCAN shows fault coverage comparable to full-scan testing, with significant reduction in test area overhead and test application time when compared to a traditional gate-level full-scan implementation. Subhrajit Bhattacharya, Sujit Dey |
VTS | 2 |
| 1996 | Fast true delay estimation during high level synthesisabstractThis paper addresses the problem of true delay estimation during high level design. The true delay is the delay of the longest sensitizable path in the resulting circuit, as opposed to the topological delay which is the delay of the longest path in the circuit. The existing delay estimation techniques either estimate the topological delay, which may be pessimistic if the longest path is unsensitizable or false, or estimate the true delay using gate-level timing analysis which may be prohibitively expensive. Resource sharing in high level synthesis can create false paths in the circuit implementation. Hence, determining the clock period using topological delay can be unduly conservative, resulting in excessive hardware to meet tight timing specifications. In this paper, we introduce an efficient technique to compute an estimate of the true delay. The proposed technique relies on partitioning the paths in the circuit and topological delay computation, and not on path sensitization. The paths in the implementation are partitioned into two sets given the high level information on scheduling and resource sharing: the complete determining path set (CDP/sub R/) and the nondetermining path set (NDP/sub R/). We prove that the delay of the longest path in CDP/sub R/ is lower bounded by the true delay and upper bounded by the topological delay of the circuit. Consequently, an estimate of the true delay of the resulting circuit can be computed by measuring the topological delay of the longest path in CDP/sub R/. We have developed a Functional delay ESTimation tool (FEST). Experimental results on a set of benchmarks reveal the following: approximately 50% of all paths are in NDP/sub R/ and can be ignored for true delay estimation, and the true delay estimates are on the average 15% less than the topological delay. The high level true delay estimates are accurate, as verified by comparing with the true delays obtained by gate-level timing analysis on actual implementations. Furthermore, results reveal that high level true delay estimation can be done very fast, even when gate-level true delay estimation becomes infeasible. Subhrajit Bhattacharya, Sujit Dey, Franc Brglez |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1995 | Synthesis-for-testability using transformationsabstractWe address the problem of transforming a behavioral specification so that synthesis of a testable implementation from the new specification requires significantly less area and partial scan cost than synthesis from the original specification. A two-stage objective function, that estimates the area and testability of the final implementation, and also captures enabling effects of the transformations, is developed. Optimization is done using a new randomized branch and bound steepest descent algorithm. Application of the transformation algorithm on several examples demonstrates significant simultaneous improvement in both area and testability of the final implementations. Miodrag Potkonjak, Sujit Dey, Rabindra K. Roy |
ASP-DAC | 2 |
| 1995 | A controller-based design-for-testability technique for controller-data path circuitsabstractThis paper investigates the effect of the controller on the testability of sequential circuits composed of controllers and data paths. It is shown that even when both the controller and the data path parts are individually 100% testable, the composite circuit may not be easily testable by gate-level sequential ATPG. Analysis shows that a primary problem in test pattern generation of combined controller-data path circuits is the correlation of control signals due to implications imposed by the controller specification. A design-for-testability technique is developed to re-design the controller such that the implications which may produce conflicts during test pattern generation are eliminated. The DFT technique involves adding extra control vectors to the controller. Experimental results show the ability of the controller DFT technique to produce highly testable controller-data path circuits, with nominal hardware overhead. Sujit Dey, Vijay Gangaram, Miodrag Potkonjak |
ICCAD | 1 |
| 1995 | Design-for-debugging of application specific designsabstractWe address the problem of considering debugging requirements during high level synthesis by providing low-cost hardware support and scheduling and assignment methods for ensuring controllability and observability of the user specified variables. Two key conceptually new design ideas that enable efficient debugging are developed: pipelining of debugging variables for improving their scheduling and assignment freedom and use of I/O buffers for improving resource utilization of I/O pins. The provably optimal bounds for the maximum cardinality of the set of controllable and observable variables for a given design specification are derived. A polynomial time complexity synthesis algorithm for achieving the bounds is developed. The minimization of hardware overhead gives rise to a combinatorial optimization problem which is solved using a non-greedy heuristic algorithm. The effectiveness of the proposed Design-for-Debugging approach is demonstrated on several examples. Miodrag Potkonjak, Sujit Dey, Kazutoshi Wakabayashi |
ICCAD | 2 |
| 1995 | Design of testable sequential circuits by repositioning flip-flops
Sujit Dey, Srimat T. Chakradhar |
J. Electron. Test. | 1 |
| 1995 | Exploiting multicycle false paths in the performance optimization of sequential logic circuitsabstractThis paper addresses the performance optimization problem for sequential logic circuits. It is shown how the notion of false paths, traditionally defined for combinational logic circuits, can be extended to the sequential context by considering the operation of the circuit over multiple clock-cycles. These multicycle false paths can be removed from the circuit using techniques similar to those proposed for combinational logic circuits. This observation offers new techniques to improve the performance of sequential logic circuits. An implementation of an algorithm that uses these ideas shows significant performance improvement on some typical benchmark circuits at a modest area overhead.> Pranav Ashar, Sujit Dey, Sharad Malik |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1995 | Considering testability at behavioral level: use of transformations for partial scan cost minimization under timing and area constraintsabstractWe address the problem of transforming a behavioral specification so that synthesis of a testable implementation from the new specification requires significantly less area and partial scan cost than synthesis from the original specification. The proposed approach has three components: a library of relevant transformation mechanisms, an objective function, and an optimization algorithm. The most effective transformations for testability optimization are identified by analyzing the fundamental relationship between transformational mechanisms and topological and functional properties of the computations that affect testability. A dynamic, two-stage objective function that estimates the area and testability of the final implementation, and also captures enabling and disabling effects of the transformations, is developed. Optimization is done using a new randomized branch and bound steepest descent algorithm. Application of the transformation algorithm on several benchmark examples demonstrates significant simultaneous improvement in both area and testability of the final implementations.> Miodrag Potkonjak, Sujit Dey, Rabindra K. Roy |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1995 | Behavioral synthesis of area-efficient testable designs using interaction between hardware sharing and partial scanabstractWe introduce BETS, a behavioral test synthesis system, for the synthesis of high-throughput, area-efficient testable designs. While hardware sharing is a powerful technique to achieve area efficiency, it may adversely affect the testability of the synthesized design by introducing new loops. Besides CDFG loops, hardware sharing introduces three other types of loops: assignment loops, sequential false loops, and register files cliques. We provide a comprehensive analysis and a formal grammar characterization of the sources of loops in the data path during behavioral synthesis. Partial scan is a cost-effective technique for sequential circuit testing. Hardware sharing of scan registers can be used to minimize the number of scan registers required to synthesize data paths with minimal number of loops. The scan registers can be shared amongst several variables of the CDFG, to break not only the loops in the CDFG, but also the very loops introduced in the data path by hardware sharing. A new random walk based algorithm is proposed to break all CDFG loops using a minimal number of scan registers. The subsequent scheduling and assignment phase avoids formation of loops in the data path by reusing the scan registers, while ensuring high resource utilization. The experimental results demonstrate the effectiveness of the new technique to synthesize easily testable data paths, with nominal hardware overhead, while maintaining the performance of the designs. The partial scan overhead incurred by the technique is significantly less than that of a gate-level partial scan approach.> Miodrag Potkonjak, Sujit Dey, Rabindra K. Roy |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1994 | Clock Period Optimization During Resource Sharing and AssignmentabstractAbstract- This paper analyzes the effect of resource sharing and assignment on the clock period of the synthesized circuit. We focus on behavioral specifications with mutually exclusive paths, due to the presence of nested conditional branches and loops. It is shown that even when the set of available resources is fixed, different assignments may lead to circuits with significant differences in clock period. We provide a comprehensive analysis of how resource sharing and assignment introduces long paths in the circuit. Based on the analysis, we develop an assignment algorithm which uses a high-level delay estimator to assign operations to a fixed set of available resources so as to minimize the clock period of the resultant circuit. Experimental results on several conditionalintensive designs demonstrate the effectiveness of the assignment algorithm. I. Subhrajit Bhattacharya, Sujit Dey, Franc Brglez |
DAC | 2 |
| 1994 | Performance Analysis and Optimization of Schedules for Conditional and Loop-Intensive SpecificationsabstractAbstract- This paper presents a new method,based on Markov chain analysis, to evaluate the performance of schedules of behavioral specifications. The proposed performance measure is the expected number of clock cycles required by the schedule for a complete execution of the behavioral specification for any distribution of inputs. The measure considers both the repetition of operations (due to loops) and their conditional execution (due to conditional branches). We propose an efficient technique to calculate the metric. We introduce a loop-directed scheduling algorithm (LDS). The algorithm produces schedules such that the expected number of clock cycles, required by the schedule for a complete execution of the behavioral specification, is minimized. Experimental results on several conditional and loop-intensive specifications demonstrate the relevance and effectiveness of both the performance measure and the scheduling algorithm. I. Subhrajit Bhattacharya, Sujit Dey, Franc Brglez |
DAC | 2 |
| 1994 | Resynthesis and Retiming for Optimum Partial ScanabstractArticle Free Access Share on Resynthesis and retiming for optimum partial scan Authors: Srimat T. Chakradhar C&C Research Laboratories, NEC, 4 Independence Way, Princeton, NJ C&C Research Laboratories, NEC, 4 Independence Way, Princeton, NJView Profile , Sujit Dey C&C Research Laboratories, NEC, 4 Independence Way, Princeton, NJ C&C Research Laboratories, NEC, 4 Independence Way, Princeton, NJView Profile Authors Info & Claims DAC '94: Proceedings of the 31st annual Design Automation ConferenceJune 1994 Pages 87–93https://doi.org/10.1145/196244.196288Published:06 June 1994Publication History 36citation145DownloadsMetricsTotal Citations36Total Downloads145Last 12 Months6Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Srimat T. Chakradhar, Sujit Dey |
DAC | 2 |
| 1994 | Optimizing Resource Utilization and Testability Using Hot Potato TechniquesabstractThis paper introduces hot potato high level synthesis transformation techniques. These techniques add deflection operations in a computation in such a way that a specific goal is optimized. We demonstrate how the requirements for two important components of the final implementation cost, registers and interconnects, are significantly reduced using new technique. It is also demonstrated how hot potato techniques can be effectively used during high level synthesis to minimize the partial scan overhead to make the synthesized design testable. Miodrag Potkonjak, Sujit Dey |
DAC | 2 |
| 1994 | Behavioral synthesis of low-cost partial scan designs for DSP applicationsabstractPartial scan is a popular design for testability technique for cost-effective sequential automatic test pattern generation (ATPG). An efficient partial scan approach selects flip-flops (FFs) in the minimum feedback vertex set (MFVS) of the FF dependency graph, so that loops are broken. Through an analysis of the sources of loops in the data path, this paper proposes a new high-level synthesis methodology to synthesize DSP designs which have low-cardinality MFVS, thereby reducing the cost of partial scan significantly. A test efficiency of 100% could be achieved for all designs synthesized by the proposed approach, requiring a significantly less number of FFs to be scanned compared to the original implementations.> Sujit Dey, Miodrag Potkonjak, Rabindra K. Roy |
ICASSP (2) | 1 |
| 1994 | Provably correct high-level timing analysis without path sensitization
Subhrajit Bhattacharya, Sujit Dey, Franc Brglez |
ICCAD | 2 |
| 1994 | Non-scan design-for-testability of RT-level data paths
Sujit Dey, Miodrag Potkonjak |
ICCAD | 1 |
| 1994 | Transforming Behavioral Specifications to Facilitate Synthesis of Testable DesignsabstractRecently, several high level synthesis approaches have been proposed to synthesize testable data paths from behavioral specifications. This paper introduces a novel technique to transform behavioral specifications, such that an existing behavioral test synthesis system can generate area-efficient, testable designs with significantly lower partial scan overhead. Experimental results demonstrate the significant savings in partial scan overhead when the transformation is applied before using the behavioral test synthesis system to synthesize 100% test-efficient designs. Sujit Dey, Miodrag Potkonjak |
ITC | 1 |
| 1994 | Retiming sequential circuits to enhance testabilityabstractThis paper presents a technique to enhance the testability of sequential circuits by repositioning registers. A novel retiming for testability technique is proposed that reduces cycle lengths in the dependency graph, converts sequential redundancies into combinational redundancies, and yields retimed circuits that usually require fewer scan registers to break all cycles (except self-loops) as compared to the original circuit. The retiming technique is based on a new minimum cost flow formulation that simultaneously considers the interactions among all strongly connected components (SCCs) of the circuit to minimize the number of registers in the SCCs. Experimental results on several large sequential circuits demonstrate the effectiveness of the proposed retiming for testability technique.> Sujit Dey, Srimat T. Chakradhar |
VTS | 1 |
| 1994 | Synthesizing designs with low-cardinality minimum feedback vertex set for partial scan applicationabstractAn efficient partial scan approach for cost-effective sequential ATPG is to select flip-flops (FFs) in the minimum feedback vertex set (MFVS) of the FF dependency graph, so that loops are broken. Through a comprehensive analysis of the sources of loops in the data path, this paper proposes a new high-level synthesis methodology to synthesize data paths which have low-cardinality MFVS, thereby reducing the cost of partial scan significantly. A test efficiency of 100% could be achieved for all designs synthesized by the proposed approach, requiring a significantly less number of FFs to be scanned compared to the original implementations.> Sujit Dey, Miodrag Potkonjak, Rabindra K. Roy |
VTS | 1 |
| 1993 | Sequential Circuit Delay optimization Using Global Path DelaysabstractABSTRACT: We propose a novel sequential delay op-timization technique based on network flow methods that simultaneously exploits delays on all paths in the circuit. We view the sequential circuit as an intercon-nection of path segments with pre-specified delays. Path segments are bounded by flip-flops, primary inputs or primary outputs. Recognizing that a delay optimizer can satisfy certain delay constraints more easily than others, we first propose a measure of difficulty for the delay optimizer. Our measure is based on explicit path delays to be satisfied by the delay optimizer. Also, our measure induces a partial order on the set of possible delay constraints. We then compute a set of delay con-straints that is optimal with respect to our measure. The delay constraint set is optimal in the sense that it is the easiest constraint that can be specified to the delay optimizer. We formulate the delay constraint cal-culation problem as a minimum cost network flow prob-lem. If the delay optimizer satisfies the optimal delay constraint set, then the resynthesized circuit may have several pat hs exceeding the desired clock period. How-ever, we show that the resynthesized circuit can always be retimed to achieve the desired clock period. Exper-imental results on MCNC synthesis benchmarks show that our method improves the performance of circuits beyond what is achievable using optimal retiming and conventional combinational logic synthesis. 1. Srimat T. Chakradhar, Sujit Dey, Miodrag Potkonjak, Steven G. Rothweiler |
DAC | 2 |
| 1993 | Critical Path Minimization Using Retiming and Algebraic Speed-UpabstractThe power of retiming is often limited by the underlying topology of a computational structure.We combine the power of retiming with a complete set of algebraic transformations in an iterative improvement framework, where retiming and algebraic speed-up algorithms are successively applied, so that the latter enables the former.The key part of the approach is a new algebraic speed-up algorithm being used for the first time in high-level synthesis for transformations of algebraic expressions so that an arbitrary set of input arrival times and output required times are satisfied.Since the new method moves delays forward only and retiming is done locally and very infrequently, it also always calculates the new initial state efficiently.The proposed approach has yielded results better or equal to the best previously published on all benchmark examples and on several novel real-life examples. Zia Iqbal, Miodrag Potkonjak, Sujit Dey, Alice C. Parker |
DAC | 3 |
| 1993 | Exploiting hardware sharing in high-level synthesis for partial scan optimizationabstractA new approach to high level synthesis, which simultaneously addresses testability and resource utilization, is presented. We explore the relationship between hardware sharing, loops in the synthesized data-path, and partial scan overhead. Since loops make a circuit hard to test, a comprehensive analysis of the sources of loops in the data path, created during high level synthesis, is provided. The paper introduces the problem of breaking CDFG loops with a minimal number of scan registers. Subsequent scheduling and assignment avoid formation of loops in the data path by sharing the scan registers, while ensuring high resource utilization. Experimental results demonstrate the effectiveness of the technique to synthesize easily testable data paths, with significantly less partial scan cost than a gate-level partial scan approach. Sujit Dey, Miodrag Potkonjak, Rabindra K. Roy |
ICCAD | 1 |
| 1993 | High Performance Embedded System Optimization Using Algebraic and Generalized Retiming TechniquesabstractRetiming, algebraic and redundancy manipulation transformations are widely used in both the high level synthesis and the compilers fields. We present a new approach on how these powerful transformations can be applied to improve the performance of embedded systems, by optimizing their latency and throughput. A simple modification is sufficient to adapt both the Leiserson-Saxe retiming algorithm and the recently introduced ERB algorithm for the new task. We introduce a new negative retiming technique and the algorithm which coordinates this technique with both algebraic and redundancy manipulation techniques for latency optimization. The effectiveness of all discussed techniques is demonstrated on a set of "real-life" examples. Latency and throughput are improved by factors of 7.06 and 2.83 respectively, often with minimal or no additional hardware overhead.> Miodrag Potkonjak, Sujit Dey, Zia Iqbal, Alice C. Parker |
ICCD | 2 |
| 1993 | Transformations and resynthesis for testability of RT-level control-data path specificationsabstractThis paper introduces a technique to transform a given register-transfer level (RT-level) design, consisting of control logic and data path, into a functionally equivalent, minimized design which is 100% testable under full-scan at the gate level. The proposed RT-level optimization technique uses the RT-level structure and exploits the interaction between the control and the data path. Our approach maintains the RT-level design hierarchy while performing RT-level transformations of initially specified data path, followed by resynthesis of control using don't cares extracted from the data path. Experiments with several RTL benchmarks demonstrate the effectiveness of the technique in generating fully testable designs. In addition, comparison with logic-level techniques show the advantages of the proposed technique as an optimizing tool to produce circuits with reduced area and delay.> Subhrajit Bhattacharya, Franc Brglez, Sujit Dey |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 1992 | Exploiting multi-cycle false paths in the performance optimization of sequential circuitsabstractIt is shown how the notion of false paths, traditionally defined for combinational logic circuits, can be extended to the sequential context by considering the operation of the circuit over multiple clock-cycles. Multicycle false paths can be removed from the circuit using techniques similar to those proposed for combinational logic circuits. This observation offers techniques to improve the performance of sequential logic circuits. A preliminary implementation of an algorithm that uses these ideas shows significant performance improvement on some typical benchmark circuits at a very modest area overhead.> Pranav Ashar, Sujit Dey, Sharad Malik |
ICCAD | 2 |
| 1992 | Performance optimization of sequential circuits by eliminating retiming bottlenecksabstractA method to improve the effectiveness of retiming by transforming the sequential circuit is proposed. Bottlenecks which prevent retiming to achieve a desired clock period are identified. Conditions to eliminate the retiming bottlenecks are derived. These conditions are satisfied by a process of identifying subcircuits and satisfying a set of timing constraints on the subcircuits. The transformed circuit, which satisfies the timing constraints, can be retimed to achieve the desired clock period. If the original circuit has its initial state specified, the method always generates the final circuit with an equivalent initial state. Experimental results on a variety of sequential benchmark circuits demonstrate significant performance improvement.> Sujit Dey, Miodrag Potkonjak, Steven G. Rothweiler |
ICCAD | 1 |
| 1991 | Partitioning Sequential Circuits for Logic OptimizationabstractThe concepts of corolla partitioning based on an analysis of signal reconvergence to cyclic sequential circuits are extended. The sequential circuit is partitioned into corollas that will contain latches but can be peripherally retimed and resynthesized using combinational techniques. Cycles are broken in the circuit by ensuring that the partitions that are formed are acyclic. Application of the proposed partitioning, retiming and resynthesis approach to a set of large sequential benchmarks has shown considerable gains after resynthesis.> Sujit Dey, Franc Brglez, Gershon Kedem |
ICCD | 1 |
| 1990 | Corolla Based Circuit Partitioning and ResynthesisabstractThis paper introduces a circuit partitioning method based on analysis of reconvergent fanout. We consider a DAG model for a circuit. We define a corolla as a set of overlapping reconvergent fanout regions. We partition the DAG into a set of non-overlapping corollas and use the corollas to resynthesize the circuit. We show that resynthesis of large benchmark circuits consistently reduces transistor pairs and layout area while improving delay and testability. Sujit Dey, Franc Brglez, Gershon Kedem |
DAC | 1 |
| 1990 | A New Parallel Sorting Algorithm and its Efficient VLSI ImplementationabstractIn this paper we develop a new parallel algorithm for sorting which has a time complexity of O(log n) and requires n2/log n processors. The algorithm can be readily mapped on an SIMD mesh connected array of processors which has all the features of efficient VLSI implementation. The corresponding hardware algorithm maintains the O(Log n) execution time and has a low O(n) interprocessor communication time. Sujit Dey, Pradip K. Srimani |
Comput. J. | 1 |