Vivek Chaturvedi

dblp:72/8378 · DBLP profile ↗
← Back
25ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0003-1358-0107ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 19 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DAPO: Design Structure-Aware Pass Ordering for HLS via Contrastive and Reinforcement Learning
abstract
High-Level Synthesis (HLS) tools are widely adopted in FPGA-based domain-specific accelerator design. However, existing tools rely on fixed optimization strategies inherited from software compilations, limiting their effectiveness. Tailoring optimization strategies to specific designs requires deep semantic understanding, accurate hardware metric estimation, and advanced search algorithms - capabilities that current approaches lack.We propose DAPO, a design structure-aware pass ordering framework that extracts program semantics from control and data flow graphs, employs contrastive learning to generate rich embeddings, and leverages an analytical model for accurate hardware metric estimation. These components jointly guide a reinforcement learning agent to discover design-specific optimization strategies. Evaluations on standard HLS benchmarks demonstrate that our end-to-end flow delivers 1.67× speedup on pragma-free designs and a 2.36× speedup on designs with pragmas over Vitis HLS with comparable resource usage.
Jinming Ge, Linfeng Du, Likith Anaparty, Shangkun Li, Tingyuan Liang, Afzal Ahmad, Vivek Chaturvedi, Sharad Sinha, Zhiyao Xie, Jiang Xu 0001, Wei Zhang 0012
DATE7
2026 ViT-EBoT: Vision Transformer for Encrypted Botnet Detection in Resource-Constrained Edge Devices
Akshara Ravi, Vivek Chaturvedi, Muhammad Shafique 0001
IEEE Internet Things J.2
2025 ADVeRL-ELF: ADVersarial ELF Malware Generation using Reinforcement Learning
abstract
Deep learning models are now pervasive in the malware detection domain owing to their high accuracy and performance efficiency. However, it is critical to analyze the robustness of these models by introducing adversarial attacks that can expose their vulnerabilities. Nevertheless, adversarial malware generation problem for Linux has not been well-investigated. In this work, we propose a novel reinforcement learning framework, ADVeRL-ELF to generate adversarial ELF malware by adding semantic NOPs within the executable region. Experimental results show that ADVeRL-ELF achieved an attack success rate of 59.5%. These adversarial malware can be leveraged to harden the Linux based malware detection systems.
Akshara Ravi, Vivek Chaturvedi, Muhammad Shafique 0001
DAC2
2025 Hy-Deft: A Hybrid Defense Technique for Vision Transformers against Adversarial Attacks in Medical Imaging
abstract
Vision Transformers (ViTs) have emerged as an essential tool in medical imaging applications, offering superior performance compared to Convolutional Neural Networks (CNNs). However, by altering input labels with subtle perturbations, adversarial attacks can seriously compromise the accuracy of ViTs. Adversarial training is a popular countermeasure for such attacks, but results in extended training time and reduction in the accuracy of clean Images. In this work, we propose a hybrid defense technique namely Hy-Deft that improves the resilience of ViT models against adversarial attacks by combining sophisticated feature extraction with adversarial training. For feature extraction, we use techniques such as Histogram of Oriented Gradients (HOG), Local Binary Patterns (LBP), and ScaleInvariant Feature Transform (SIFT), which are subsequently processed through a feature projection layer. The adversarial training part is implemented using a CNN and the features are extracted before the classification phase. Our approach integrates feature vectors derived from raw images and those extracted from the CNN to the ViT for comprehensive implementation. Extensive experiments are performed using three datasets from MedMNIST - namely, BreastMNIST, BloodMNIST, and OrganCMNIST against prominent adversarial attack algorithms, including FGSM, PGD, BIM, and C&W. The proposed method reduces training overhead as adversarial training is performed in a separate lightweight CNN and the clean accuracy is increased through the incorporation of robust feature extraction techniques. Experiential results show that the clean image accuracy is increased by a maximum of 4% and the accuracy after applying FGSM, PGD, BIM, and C&W attacks by 79%, 77%, 72%, and 12%, respectively.
Neha A S, Vivek Chaturvedi, Muhammad Shafique 0001
IJCNN2
2025 CapsBeam: Accelerating Capsule Network-Based Beamformer for Ultrasound Nonsteered Plane-Wave Imaging on Field-Programmable Gate Array
abstract
In recent years, there has been a growing trend in accelerating computationally complex nonreal-time beamforming algorithms in ultrasound imaging using deep learning models. However, due to the large size and complexity, these state-of-the-art deep learning techniques pose significant challenges when deploying on resource-constrained edge devices. In this work, we propose a novel capsule network-based beamformer called CapsBeam, designed to operate on raw radio frequency data and provide an envelope of beamformed data through nonsteered plane-wave insonification. In experiments on in vivo data, CapsBeam reduced artifacts compared to the standard Delay-and-Sum (DAS) beamforming. For in vitro data, CapsBeam demonstrated a 32.31% increase in contrast, along with gains of 16.54% and 6.7% in axial and lateral resolution compared to the DAS. Similarly, in silico data showed a 26% enhancement in contrast, along with improvements of 13.6% and 21.5% in axial and lateral resolution, respectively, compared to the DAS. To reduce the parameter redundancy and enhance the computational efficiency, we pruned the model using our multilayer look-ahead kernel pruning (LAKP-ML) methodology, achieving a compression ratio of 85% without affecting the image quality. Additionally, the hardware complexity of the proposed model is reduced by applying quantization, simplification of nonlinear operations, and parallelizing operations. Finally, we proposed a specialized accelerator architecture for the pruned and optimized CapsBeam model, implemented on a Xilinx ZU7EV FPGA. The proposed accelerator achieved a throughput of 30 GOPS for the convolution operation and 17.4 GOPS for the dynamic routing operation.
T. P. Abdul Rahoof, Vivek Chaturvedi, Mahesh Raveendranatha Panicker, Muhammad Shafique 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2024 Tiny- VBF: Resource-Efficient Vision Transformer based Lightweight Beamformer for Ultrasound Single-Angle Plane Wave Imaging
abstract
Accelerating compute intensive non-real-time beam-forming algorithms in ultrasound imaging using deep learning architectures has been gaining momentum in the recent past. Nonetheless, the complexity of the state-of-the-art deep learning techniques poses challenges for deployment on resource-constrained edge devices. In this work, we propose a novel vision transformer based tiny beamformer (Tiny-VBF), which works on the raw radio-frequency channel data acquired through single-angle plane wave insonification. The output of our Tiny-VBF provides fast envelope detection requiring very low frame rate, i.e. 0.34 GOPs/Frame for a frame size of 368 × 128 in comparison to the state-of-the-art deep learning models. It also exhibited an 8% increase in contrast and gains of 5% and 33% in axial and lateral resolution respectively when compared to Tiny-CNN on in-vitro dataset. Additionally, our model showed a 4.2% increase in contrast and gains of 4% and 20% in axial and lateral resolution respectively when compared against conventional Delay-and-Sum (DAS) beamformer. We further propose an accelerator architecture and implement our Tiny-VBF model on a Zynq UltraScale+ MPSoC ZCU104 FPGA using a hybrid quantization scheme with 50% less resource consumption compared to the floating-point implementation, while preserving the image quality.
T. P. Abdul Rahoof, Vivek Chaturvedi, Mahesh Raveendranatha Panicker, Muhammad Shafique 0001
DATE2
2024 S-E Pipeline: A Vision Transformer (ViT) based Resilient Classification Pipeline for Medical Imaging Against Adversarial Attacks
abstract
Vision Transformer (ViT) is becoming widely popular in automating accurate disease diagnosis in medical imaging owing to its robust self-attention mechanism. However, ViTs remain vulnerable to adversarial attacks that may thwart the diagnosis process by leading it to intentional misclassification of critical disease. In this paper, we propose a novel image classification pipeline, namely, S-E Pipeline, that performs multiple pre-processing steps that allow ViT to be trained on critical features so as to reduce the impact of input perturbations by adversaries. Our method uses a combination of segmentation and image enhancement techniques such as Contrast Limited Adaptive Histogram Equalization (CLAHE), Unsharp Masking (UM), and High-Frequency Emphasis filtering (HFE) as preprocessing steps to identify critical features that remain intact even after adversarial perturbations. The experimental study demonstrates that our novel pipeline helps in reducing the effect of adversarial attacks by 72.22% for the ViT-b32 model and 86.58% for the ViT-l32 model. Furthermore, we have shown an end-to-end deployment of our proposed method on the NVIDIA Jetson Orin Nano board to demonstrate its practical use case in modern hand-held devices that are usually resource-constrained.
Neha A S, Vivek Chaturvedi, Muhammad Shafique 0001
IJCNN2
2023 FastCaps: A Design Methodology for Accelerating Capsule Network on Field Programmable Gate Arrays
abstract
Capsule Network (CapsNet) has shown significant improvement in understanding the variation in images along with better generalization ability compared to traditional Convolutional Neural Network (CNN). CapsNet preserves spatial relationship among extracted features and apply dynamic routing to efficiently learn the internal connections between capsules. However, due to the capsule structure and the complexity of the routing mechanism, it is non-trivial to accelerate CapsNet performance in its original form on Field Programmable Gate Array (FPGA). Most of the existing works on CapsNet have achieved limited acceleration as they implement only the dynamic routing algorithm on FPGA, while considering all the processing steps synergistically is important for real-world applications of Capsule Networks. Towards this, we propose a novel two-step approach that deploys a full-fledged CapsNet on FPGA. First, we prune the network using a novel Look-Ahead Kernel Pruning (LAKP) methodology that uses the sum of look-ahead scores of the model parameters. Next, we simplify the nonlinear operations, reorder loops, and parallelize operations of the routing algorithm to reduce CapsNet hardware complexity. To the best of our knowledge, this is the first work accelerating a full-fledged CapsNet on FPGA. Experimental results on the MNIST and F-MNIST datasets (typical in Capsule Network community) show that the proposed LAKP approach achieves an effective compression rate of 99.26% and 98.84%, and achieves a throughput of 82 FPS and 48 FPS on Xilinx PYNQ-Z1 FPGA, respectively. Furthermore, reducing the hardware complexity of the routing algorithm increases the throughput to 1351 FPS and 934 FPS respectively. As corroborated by our results, this work enables highly performance-efficient deployment of CapsNets on low-cost FPGA that are popular in modern edge devices.
T. P. Abdul Rahoof, Vivek Chaturvedi, Muhammad Shafique 0001
IJCNN2
2023 ViT4Mal: Lightweight Vision Transformer for Malware Detection on Edge Devices
abstract
There has been a tremendous growth of edge devices connected to the network in recent years. Although these devices make our life simpler and smarter, they need to perform computations under severe resource and energy constraints, while being vulnerable to malware attacks. Once compromised, these devices are further exploited as attack vectors targeting critical infrastructure. Most existing malware detection solutions are resource and compute-intensive and hence perform poorly in protecting edge devices. In this paper, we propose a novel approach ViT4Mal that utilizes a lightweight vision transformer (ViT) for malware detection on an edge device. ViT4Mal first converts executable byte-code into images to learn malware features and later uses a customized lightweight ViT to detect malware with high accuracy. We have performed extensive experiments to compare our model with state-of-the-art CNNs in the malware detection domain. Experimental results corroborate that ViTs don’t demand deeper networks to achieve comparable accuracy of around 97% corresponding to heavily structured CNN models. We have also performed hardware deployment of our proposed lightweight ViT4Mal model on the Xilinx PYNQ Z1 FPGA board by applying specialized hardware optimizations such as quantization, loop pipelining, and array partitioning. ViT4Mal achieved an accuracy of ~94% and a 41x speedup compared to the original ViT model.
Akshara Ravi, Vivek Chaturvedi, Muhammad Shafique 0001
ACM Trans. Embed. Comput. Syst.2
2021 Longevity Framework: Leveraging Online Integrated Aging-Aware Hierarchical Mapping and VF-Selection for Lifetime Reliability Optimization in Manycore Processors
abstract
Rapid device aging in the nano era threatens system lifetime reliability, posing a major intrinsic threat to system functionality. Traditional techniques to overcome the aging-induced device slowdown, such as guardbanding are static and incur performance, power, and area penalties. In a manycore processor, the system-level design abstraction offers dynamic opportunities through the control of task-to-core mappings and per-core operation frequency towards more balanced core aging profile across the chip, optimizing the system lifetime reliability while meeting the application performance requirements. This article presents Longevity Framework (LF) that leverages online integrated aging-aware hierarchical mapping and voltage frequency (VF)-selection for lifetime reliability optimization in manycore processors. The mapping exploration is hierarchical to achieve scalability. The VF-selection builds on the trade-offs involved between power, performance, and aging as the VF is scaled while leveraging the per-core DVFS capabilities. The methodology takes the chip-wide process variation into account. Extensive experimentation, comparing the proposed approach with two state-of-the-art methods, for 64-core and 256-core systems running applications from PARSEC and SPLASH-2 benchmark suites, show an improvement of up to 3.2 years in the system lifetime reliability and 4× improvement in the average core health.
Vijeta Rathore, Vivek Chaturvedi, Amit Kumar Singh 0002, Thambipillai Srikanthan, Muhammad Shafique 0001
IEEE Trans. Computers2
2020 Thermal Aware Lifetime Reliability Optimization for Automotive Distributed Computing Applications
abstract
As the automotive industry is moving towards electric and self-driving vehicles, how to ensure the high degree of reliability for electronic control systems (ECS) has emerged as a serious concern. Temperature plays a vital role in the reliability of ECS because vehicles are subjected to high chip temperature due to harsh operating conditions and high integrated circuit (IC) on-chip power density. In this paper, we study the problem on how to optimize the lifetime reliability and guarantee the chip's peak temperature of ECS by judiciously allocating the application to ECS. We first propose a simple mathematical programming based thermal aware approach, assuming temperature can reach a stable status immediately. We then present a genetic algorithm approach based on effective and computationally efficient methods for peak temperature identification and system-wide lifetime reliability calculation, by taking advantage of the periodicity of vehicle applications. Our experimental results, based on both synthetic test cases and practical benchmarks demonstrate the significant improvement in lifetime reliability and CPU time for automotive ECS achieved by our proposed algorithms compared to the state-of-the-art results.
Ajinkya S. Bankar, Shi Sha, Vivek Chaturvedi, Gang Quan
ICCD3
2020 Energy Minimization for Multicore Platforms Through DVFS and VR Phase Scaling With Comprehensive Convex Model
abstract
Energy management is a critical challenge in multicore processors due to continuous technology scaling. Previous methods have mostly focused on the energy minimization of the processor cores. However, energy overhead of the off-chip voltage regulator (VR) has recently shown to be a nontrivial part of the total energy consumption and has been previously overlooked. In this paper, we propose an overall energy optimization method for the system that minimizes both per-core energy consumption and VR energy consumption using dynamic voltage frequency scaling and VR phase scaling by solving a comprehensive convex model. In order to improve the accuracy of the task latency model, a new task model considering both computation and memory access of the task is also developed. Furthermore, for better scalability and lower online overhead, we decompose our proposed convex method into two stages: 1) an offline stage and 2) an online stage. During the offline stage, we explore the convex model by assuming different numbers of active phases of the VR, various workload pressures and workload characteristics to collect the optimal frequency assignments under different scenarios. During the online stage, the specific frequency assignment for cores and optimal active phase number of the VR are selected and applied based on the actual workload pressure and its characteristics running on the cores. Experiments on real benchmarks show that when compared with the state-of-the-art approaches, which are oblivious to VR overheads and exploit slack time to achieve energy minimization, our method can achieve a significant energy saving of up to 22.4% with negligible online overhead.
Zuomin Zhu, Wei Zhang 0012, Vivek Chaturvedi, Amit Kumar Singh 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2019 LifeGuard: A Reinforcement Learning-Based Task Mapping Strategy for Performance-Centric Aging Management
abstract
Device scaling to subdeca nanometer has pushed device aging as a primary design concern. In manycore systems, inevitable process variation further adds to delay degradation and, coupled with the scalability issues in manycores, makes aging management, while meeting performance demands, a complex problem. LifeGuard is a performance-centric reinforcement learning-based task mapping strategy that leverages the different impact of applications on aging for improving system health. Experimental results, comparing LifeGuard with two state-of-the-art aging optimizing techniques, on a 256-core system, showed that LifeGuard led to improved health for, respectively, 57% and 74% of the cores, and also an enhanced aggregate core frequency.
Vijeta Rathore, Vivek Chaturvedi, Amit Kumar Singh 0002, Thambipillai Srikanthan, Muhammad Shafique 0001
DAC2
2019 Towards Scalable Lifetime Reliability Management for Dark Silicon Manycore Systems
abstract
Aggressive technology scaling enabled very high integration density. Unfortunately, it also led to issues such as process variation, increased power density and consequently rising chip temperature resulting in accelerated device aging and poor lifetime reliability of different components in a manycore system. Moreover, thermal and power limitations let only a fraction of the chip function at full speed; the rest is the dark silicon. Most of the lifetime reliability enhancement solutions for the multi-/manycore systems in the literature are heuristic-based, while some use standard compute-intensive methods to solve the optimization problem making them not scale well with the manycore size. The heuristic-based solutions are formulated to search through the design space of a fine granularity making it huge, limiting their scalability. Also, these approaches do not account for the impact of different applications' execution behavior on the aging of the underlying cores, and their performance requirement distribution across the cores to their advantage. In this paper, we present our resource management strategies towards building scalable lifetime reliability enhancement solutions for dark silicon manycore systems. The first technique, Hierarchical Mapping approach (HiMap), maps a periodic workload employing a block-based hierarchical method that leverages dark cores for thermal mitigation. The second approach, LifeGuard, uses reinforcement learning to learn the applications' aging behavior, and is aware of the performance requirement pattern onto the core frequencies. It maps randomly arriving requests and is scalable to the number of applications and the size of a manycore.
Vijeta Rathore, Vivek Chaturvedi, Amit Kumar Singh 0002, Thambipillai Srikanthan, Muhammad Shafique 0001
IOLTS2
2018 HiMap: A hierarchical mapping approach for enhancing lifetime reliability of dark silicon manycore systems
abstract
Technology scaling into the nano-scale CMOS regime has resulted in increased leakage and roadblock on voltage scaling, which has led to several issues like high power density and elevated on-chip temperature. This consequently aggravates device aging, compromising lifetime reliability of the manycore systems. This paper proposes HiMap, a dynamic hierarchical mapping approach to maximize lifetime reliability of manycore systems while satisfying performance, power, and thermal constraints. HiMap is process variation- and aging-aware. It comprises of two levels: (1) it identifies a region of cores suitable for mapping, and (2) it maps threads in the region and intersperses dark cores for thermal mitigation while considering the current health of the cores. Both the levels strive to reduce aging variance across the chip. We evaluated HiMap for 64-core and 256-core systems. Results demonstrate an improved system lifetime reliability by up to 2 years at the end of 3.25 years of use, as compared to the state-of-the-art.
Vijeta Rathore, Vivek Chaturvedi, Amit Kumar Singh 0002, Thambipillai Srikanthan, R. Rohith, Siew-Kei Lam, Muhammad Shafique 0001
DATE2
2017 Two-stage thermal-aware scheduling of task graphs on 3D multi-cores exploiting application and architecture characteristics
abstract
In this paper, we propose a two-stage thermal-aware task scheduling policy which exploits the application and system architecture characteristics to decouple the mapping of task-graphs for the performance and peak temperature optimization into two stages. At the first stage, the algorithm collects the best mapping of task-graphs exploiting the application and architecture characteristics to minimize the makespan of the task-graphs. At the second stage, a light-weight online algorithm comprised of efficient thermal rank and combined power models is performed to map the task nodes to the real cores for temperature minimization while maintaining the best possible performance achieved in the first stage. Compared to the previous approaches which perform the performance and temperature optimization together, our method can reduce the online mapping algorithm complexity and improve its efficiency. Experiments on real benchmarks show that an average of 6.3°C peak temperature reduction and 6.8% performance improvement can be achieved compared to other existing methods.
Zuomin Zhu, Vivek Chaturvedi, Amit Kumar Singh 0002, Wei Zhang 0012, Yingnan Cui
ASP-DAC2
2016 Performance Constraint-Aware Task Mapping to Optimize Lifetime Reliability of Manycore Systems
abstract
Negative bias temperature instability (NBTI) has emerged as a critical challenge to lifetime reliability of computing systems. Traditionally, temperature-aware methodologies are used to mitigate the impact of NBTI on aging and degradation of computing systems. However, in the presence of process variation, which is the norm in manycore processors, temperature-aware techniques are inefficient in improving lifetime reliability and can result in poor performance. In this paper, we propose a novel performance constraint-aware task mapping technique to improve lifetime reliability by mitigating NBTI considering on-chip process variation. Our approach consists of two phases, namely design-time and run-time. During design time, we generate Pareto-optimal mappings. Following which, our run-time technique judiciously intervenes to perform workload migration to save the weakest processing core. We compare our approach with performance-greedy and thermal-aware task mapping techniques. The experiment results demonstrate that our approach outperforms other two techniques and improves lifetime reliability of a manycore system as much as 54% without violating the throughput constraint.
Vijeta Rathore, Vivek Chaturvedi, Thambipillai Srikanthan
ACM Great Lakes Symposium on VLSI2
2016 Cost-efficient Acceleration of Hardware Trojan Detection Through Fan-Out Cone Analysis and Weighted Random Pattern Technique
abstract
Fabless semiconductor industry and government agencies have raised serious concerns about tampering with inserting hardware Trojans (HTs) in an integrated circuit supply chain in recent years. In this paper, a low hardware overhead acceleration method of the detection of HTs based on the insertion of 2-to-1 MUXs as test points is proposed. In the proposed method, the fact that one logical gate has a significant impact on the transition probability of the logical gates in its logical fan-out cone is utilized to optimize the number of the inserted MUXs. The nets which have smaller transition probability than the user-specified threshold and minimal logical depth from the primary inputs are selected as the candidate nets. As for each candidate net, only its input net with smallest signal probability is required to be inserted the MUXs-based test points. The procedure repeats until the minimal transition probability of the entire circuit is not smaller than the threshold value. In order to further optimize the number of required insertions and reduce the overhead, the weighted random pattern technique is also applied. Experiment results on ISCAS'89 benchmark circuits show that our proposed method can achieve remarkable improvement of transition probability with on average 9.50% power, 2.37% delay, and 10.26% area penalty.
Wei Zhang 0012, Thambipillai Srikanthan, Jason Teo Kian Jin, Vivek Chaturvedi, Tao Luo 0014
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2016 Decentralized Thermal-Aware Task Scheduling for Large-Scale Many-Core Systems
abstract
Technology scaling has enabled fast increase in the number of cores integrated in many-core systems. However, feature size shrinking also makes large-scale many-core systems vulnerable to thermal failures. Thermal-aware task scheduling is an efficient technique to reduce the run-time temperatures of many-core processors. Most existing thermal-aware task scheduling algorithms leverage centralized scheduling schemes to gather the overall information and generate the task schedule at a center scheduler. Although that scheme can achieve the optimal temperature reduction, however, it faces severe computation bottleneck and communication congestion when the many-core processors evolve to large-scale with hundreds or thousands of cores. In this paper, we propose a decentralized thermal-aware scheduling algorithm to address this problem in large-scale systems. Experiment results on various benchmarks show that our decentralized algorithm achieves significant improvement on scalability (up to 84.3% reduction in monitoring traffic) and similar benefits on temperature reduction (by 5%) when compared with the state-of-the-art thermal-aware scheduling algorithm.
Yingnan Cui, Wei Zhang 0012, Vivek Chaturvedi, Bingsheng He
IEEE Trans. Very Large Scale Integr. Syst.3
2014 Thermal-aware task scheduling for peak temperature minimization under periodic constraint for 3D-MPSoCs
abstract
3D-MPSoC offer great performance and scalability benefits. However, due to strong vertical thermal correlation and increased power density, thermal challenges in 3D-MPSoC are critical. In this paper, we propose a novel thermal aware task scheduling technique that combine intelligent task mapping with DVFS to minimize the peak temperature of the system. Particularly, our approach leverages on the fundamental thermal characteristics of 3D architecture when mapping tasks to processing cores and employing DVFS at design time followed by a simple thermal optimization step at run time. Our experiments validate the efficiency of our approach in peak temperature minimization up to 14°C compared to other existing methods.
Vivek Chaturvedi, Amit Kumar Singh 0002, Wei Zhang 0012, Thambipillai Srikanthan
RSP1
2014 Throughput maximization for periodic real-time systems under the maximal temperature constraint
abstract
In this article, we study the problem of how to maximize the throughput of a periodic real-time system under a given peak temperature constraint. We assume that different tasks in our system may have different power and thermal characteristics. Two scheduling approaches are presented. The first is built upon processors that can be in either active or sleep mode. By judiciously selecting tasks with different thermal characteristics as well as alternating the processor's active / sleep mode, the sleep period required to cool down the processor is kept at a minimum level, and, as the result, the throughput is maximized. We further extend this approach for processors with dynamic voltage/frequency scaling (DVFS) capability. Our experiments on a large number of synthetic test cases as well as real benchmark programs show that the proposed methods not only consistently outperform the existing approaches in terms of throughput maximization, but also significantly improve the feasibility of tasks when a more stringent temperature constraint is imposed.
Vivek Chaturvedi, Gang Quan, Jeffrey Fan, Meikang Qiu
ACM Trans. Embed. Comput. Syst.2
2013 An analytical solution for multi-core energy calculation with consideration of leakage and temperature dependency
abstract
Energy minimization is a critical issue and challenge when considering the cyclic dependency of leakage power and temperature as IC technology reaches deep sub-micron level. In this paper, we present an analytical method to calculate the energy consumption efficiently and effectively for a given voltage schedule on a multi-core platform, with the leakage/temperature dependency taken into consideration. Our experiments show that the proposed method can achieve a speedup of 15 times compared with the numerical method, with a relative error of no more than 1.5%.
Ming Fan 0001, Vivek Chaturvedi, Shi Sha, Gang Quan
ISLPED2
2012 On the fundamentals of leakage aware real-time DVS scheduling for peak temperature minimization
Vivek Chaturvedi, Shangping Ren, Gang Quan
J. Syst. Archit.1
2011 Leakage conscious DVS scheduling for peak temperature minimization
abstract
In this paper, we incorporate the dependencies among the leakage, the temperature and the supply voltage into the theoretical analysis and explore the fundamental characteristics on how to employ dynamic voltage scaling (DVS) to reduce the peak operating temperature. We find that, for a specific interval, a real-time schedule using the lowest constant speed is not necessarily the optimal choice any more in minimizing the peak temperature. We identify the scenarios when a schedule using two different speeds can outperform the one using the constant speed. In addition, we find that the constant speed schedule is still the optimal one to minimize the peak temperature at the temperature stable status when scheduling a periodic task set. We formulate our conclusions into several theorems with formal proofs.
Vivek Chaturvedi, Gang Quan
ASP-DAC1
2010 Feasibility Analysis for Temperature-Constraint Hard Real-Time Periodic Tasks
abstract
While the dynamic thermal management problem is closely related to the dynamic power management problem, it has its own distinct features. In this paper, we study the feasibility checking problem for real-time periodic task sets under the peak temperature constraint. We show that the traditional scheduling approach, i.e. to repeat the schedule that is feasible through the range of one hyperperiod, does not apply any more. We then present new necessary and sufficient conditions to check the feasibility of real-time schedules. We further incorporate the close relationship of leakage, temperature, and supply voltage into our feasibility analysis, and develop more elaborated feasibility conditions. Our experiments, based on technical parameters derived from a processor using the 65 nm IC technology, demonstrate the effectiveness of our feasibility conditions and, at the same time, highlight the fact that a power/thermal-aware computing technique becomes ineffective at the submicron scale if the inter dependency of leakage, temperature, and supply voltage is not properly addressed.
Gang Quan, Vivek Chaturvedi
IEEE Trans. Ind. Informatics2