VLDB 2026 Research / reviewers in the wild / expert
Radu Marculescu
dblp:88/3494
· DBLP profile ↗
182ranked-venue papers
21as first author
38since 2021 · last 2026
0000-0003-1826-7646ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 129 · 20 first-author · 7 since 2021Artificial intelligence and machine learning · 21 · 1 first-author · 17 since 2021Software engineering, systems software and programming languages · 21 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 19 · 14 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 9 · 1 first-author · 5 since 2021Computer networks · 7 · 1 since 2021Human-computer interaction and ubiquitous computing · 6 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NeuS-QA: Grounding Long-Form Video Understanding in Temporal Logic and Neuro-Symbolic ReasoningabstractWhile vision-language models (VLMs) excel at tasks involving single images or short videos, they still struggle with Long Video Question Answering (LVQA) due to its demand for complex multi-step temporal reasoning. Vanilla approaches, which simply sample frames uniformly and feed them to a VLM along with the question, incur significant token overhead. This forces aggressive downsampling of long videos, causing models to miss fine-grained visual structure, subtle event transitions, and key temporal cues. Recent works attempt to overcome these limitations through heuristic approaches; however, they lack explicit mechanisms for encoding temporal relationships and fail to provide any formal guarantees that the sampled context actually encodes the compositional or causal logic required by the question. To address these foundational gaps, we introduce NeuS-QA, a training-free, plug-and-play neuro-symbolic pipeline for LVQA. NeuS-QA first translates a natural language question into a logic specification that models the temporal relationship between frame-level events. Next, we construct a video automaton to model the video's frame-by-frame event progression, and finally employ model checking to compare the automaton against the specification to identify all video segments that satisfy the question's logical requirements. Only these logic-verified segments are submitted to the VLM, thus improving interpretability, reducing hallucinations, and enabling compositional reasoning without modifying or fine-tuning the model. Experiments on the LongVideoBench and CinePile benchmarks show that NeuS-QA significantly improves performance by over 10%, particularly on questions involving event ordering, causality, and multi-step reasoning. Sahil Shah, S. P. Sharan, Harsh Goel, Minkyu Choi 0001, Mustafa Munir, Manvik Pasula, Radu Marculescu, Sandeep Chinchali |
AAAI | 7 |
| 2026 | AdaptViG: Adaptive Vision GNN with Exponential Decay GatingabstractVision Graph Neural Networks (ViGs) offer a new direction for advancements in vision architectures. While powerful, ViGs often face substantial computational challenges stemming from their graph construction phase, which can hinder their efficiency. To address this issue we propose AdaptViG, an efficient and powerful hybrid Vision GNN that introduces a novel graph construction mechanism called Adaptive Graph Convolution. This mechanism builds upon a highly efficient static axial scaffold and a dynamic, content-aware gating strategy called Exponential Decay Gating. This gating mechanism selectively weighs long-range connections based on feature similarity. Furthermore, AdaptViG employs a hybrid strategy, utilizing our efficient gating mechanism in the early stages and a full Global Attention block in the final stage for maximum feature aggregation. Our method achieves a new state-of-the-art trade-off between accuracy and efficiency among Vision GNNs. For instance, our AdaptViG-M achieves 82.6% top-1 accuracy, outperforming ViG-B by 0.3% while using 80% fewer parameters and 84% fewer GMACs. On downstream tasks, AdaptViG-M obtains 45.8 mIoU, 44.8 APbox, and 41.1 APmask, surpassing the much larger EfficientFormer-L7 by 0.7 mIoU, 2.2 APbox, and 2.1 APmask, respectively, with 78% fewer parameters. Code: https://github.com/mmunir127/AdaptViG. Mustafa Munir, Md Mostafijur Rahman, Radu Marculescu |
WACV | 3 |
| 2026 | SmoothDiffusion-VE: Real-time Generative Video Editing Using Adaptive Feature CacheabstractVideo editing with diffusion models presents significant challenges, especially under real-time constraints. Current methods either enhance temporal consistency at the cost of slow processing or rely on frame-by-frame editing, leading to flickering and temporal artifacts. To address both challenges, we propose SmoothDiffusion-VE, a streaming-based editing approach that improves temporal consistency and processing speed through our proposed Adaptive Feature Cache (AFC) and motion-guided attention. The AFC dynamically adjusts the caching behavior based on perceptual similarity (LPIPS) between frames, i.e., shifting to a mini-cache mode for similar frames to reduce computational load. Conversely, significant frame changes trigger deeper caching to maintain robust temporal coherence. Our motion-guided attention selectively focuses on dynamic regions using optical flow, reducing unnecessary computations in static areas and accelerating processing. SmoothDiffusion-VE can run 28 FPS on one RTX 4090 GPU, achieving a 1564× speedup over Plug-and-Play Diffusion (PNP) and a 1916× speedup over Diffusion Motion Transfer (DMT), delivering a powerful solution for fast and consistent video editing. Mustafa Munir, Sophia Zalewski, Shiqiu Liu, David Tarjan, Sushmitha Belede, Anjul Patney, Radu Marculescu |
WACV | 7 |
| 2025 | Skip2-LoRA: A Lightweight On-device DNN Fine-tuning Method for Low-cost Edge DevicesabstractThis paper proposes Skip2-LoRA as a lightweight fine-tuning method for deep neural networks to address the gap between pre-trained and deployed models. In our approach, trainable LoRA (low-rank adaptation) adapters are inserted between the last layer and every other layer to enhance the network expressive power while keeping the backward computation cost low. This architecture is well-suited to cache intermediate computation results of the forward pass and then can skip the forward computation of seen samples as training epochs progress. We implemented the combination of the proposed architecture and cache, denoted as Skip2-LoRA, and tested it on a $15 single board computer. Our results show that Skip2-LoRA reduces the fine-tuning time by 90.0% on average compared to the counterpart that has the same number of trainable parameters while preserving the accuracy, while taking only a few seconds on the microcontroller board. Hiroki Matsutani, Masaaki Kondo, Kazuki Sunaga, Radu Marculescu |
ASP-DAC | 4 |
| 2025 | Batteryless Gesture Recognition via Learned SamplingabstractBatteryless wearables offer the potential for continuous and maintenance-free operation through ambient energy harvesting. However, the constrained and uncertain nature of harvested energy makes it impossible to continuously sense, process, and transmit data. In this batteryless setting, we no longer have access to complete and continuous data streams. Thus, we must adapt machine learning models to this new paradigm where only a limited subset of the data can be sampled and processed. In this work, we consider solar-powered gesture recognition as a target application where sensing and processing are constrained to a strict energy budget. We propose a learning-based approach where a sampling policy and gesture classifier are jointly trained via a shared representation. This enables the system to learn which sensor samples are informative while simultaneously optimizing the classifier for sparse inputs. By actively deciding when to sample, we improve gesture recognition accuracy by up to 10% compared to fixed-rate subsampling across a wide range of energy budgets. This closes the gap to a standard battery-powered approach by up to 40%. Geffen Cooper, Radu Marculescu |
BSN | 2 |
| 2025 | EffiDec3D: An Optimized Decoder for High-Performance and Efficient 3D Medical Image SegmentationabstractRecent 3D deep networks such as SwinUNETR, Swin-Unetrv2, and 3D UX-Net have shown promising performance by leveraging self-attention and large-kernel convolutions to capture the volumetric context. However, their substantial computational requirements limit their use in real-time and resource-constrained environments. The high #FLOPs and #Params in these networks stem largely from complex decoder designs with high-resolution layers and excessive channel counts. In this paper, we propose EffiDec3D, an optimized 3D decoder that employs a channel reduction strategy across all decoder stages, which sets the number of channels to the minimum needed for accurate feature representation. Additionally, EffiDec3D removes the high-resolution layers when their contribution to segmentation quality is minimal. Our optimized EffiDec3D decoder achieves a 96.4% reduction in #Params and a 93.0% reduction in #FLOPs compared to the decoder of original 3D UX-Net. Similarly, for SwinUNETR and SwinUNETRv2 (which share an identical decoder), we observe reductions of 94.9% in #Params and 86.2% in #FLOPs. Our extensive experiments on 12 different medical imaging tasks confirm that EffiDec3D not only significantly reduces the computational demands, but also maintains a performance level comparable to original models, thus establishing a new standard for efficient 3D medical image segmentation. Our implementation is available at https://github.com/SLDGroup/EffiDec3D. Md Mostafijur Rahman, Radu Marculescu |
CVPR | 2 |
| 2025 | News Source Credibility Assessment: A Reddit Case StudyabstractWe present a transformer-based model for credibility assessment, CREDiBERT (CREDibility assessment using Bi-directional Encoder Representations from Transformers), fine-tuned for Reddit submissions focusing on political discourse. We adopt a semi-supervised training approach for CREDiBERT, leveraging the community structure of Reddit. By encoding submission content using CREDiBERT and integrating it with a classification neural network, we improve the credibility assessment for Reddit submission by 3% in F1 score compared to existing methods. Additionally, we introduce a new version of the post-to-post network in Reddit that efficiently encodes user interactions to enhance the credibility assessment task by 8% in the F1 score. We demonstrate CREDiBERT's applicability by evaluating the susceptibility of Reddit communities to different topics and assessing the credibility score of unseen sources. Arash Amini, Yigit E. Bayiz, Ashwin Ram 0003, Radu Marculescu, Ufuk Topcu |
ICWSM | 4 |
| 2025 | Susceptibility of Communities Against Low-Credibility Content in Social News WebsitesabstractSocial news websites, such as Reddit, have evolved into prominent platforms for sharing and discussing news. A key issue on social news websites is the formation of low-credibility communities, which often lead to the spread of highly biased or uncredible news. We develop a method to identify communities prone to uncredible or highly biased news within a social news website. We employ a user embedding pipeline that detects user communities based on their stances toward posts and news sources. We then project each community onto a credibility-bias space and analyze the distributional characteristics of each projected community to identify those that have a high risk of adopting beliefs with low credibility or high bias. This approach also enables the prediction of individual users' susceptibility to low-credibility content based on their community affiliation. Our results show that latent space clusters effectively indicate the credibility and bias levels of their users, with significant variance observed across clusters---a 34% difference in the users' susceptibility to low-credibility content and a 8.3% difference in the users' susceptibility to high political bias. Yigit E. Bayiz, Arash Amini, Radu Marculescu, Ufuk Topcu |
ICWSM | 3 |
| 2025 | EfficientMedNeXt: Multi-receptive Dilated Convolutions for Medical Image Segmentation
Md Mostafijur Rahman, Mustafa Munir, Radu Marculescu |
MICCAI (4) | 3 |
| 2025 | LoMix: Learnable Weighted Multi-Scale Logits Mixing for Medical Image SegmentationabstractU-shaped networks output logits at multiple spatial scales, each capturing a different blend of coarse context and fine detail. Yet, training still treats these logits in isolation—either supervising only the final, highest-resolution logits or applying deep supervision with identical loss weights at every scale—without exploring mixed-scale combinations. Consequently, the decoder output misses the complementary cues that arise only when coarse and fine predictions are fused. To address this issue, we introduce LoMix (Logits Mixing), a Neural Architecture
Search (NAS)-inspired, differentiable plug-and-play module that generates new mixed-scale outputs and learns how exactly each of them should guide the training process. More precisely, LoMix mixes the multi-scale decoder logits with four lightweight fusion operators: addition, multiplication, concatenation, and attention-based weighted fusion, yielding a rich set of synthetic “mutant” maps. Every original or mutant map is given a softplus loss weight that is co-optimized with network parameters, mimicking a one-step architecture search that automatically discovers the most useful scales, mixtures, and operators. Plugging LoMix into recent U-shaped architectures (i.e., PVT-V2-B2 backbone with EMCAD decoder) on Synapse 8-organ dataset improves DICE by +4.2% over single-output supervision, +2.2% over deep supervision, and +1.5% over equally weighted additive fusion, all with zero inference overhead. When training data are scarce (e.g., one or two labeled scans, 5% of the trainset), the advantage grows to +9.23%, underscoring LoMix’s data efficiency. Across four benchmarks and diverse U-shaped networks, LoMiX improves DICE by up to +13.5% over single-output supervision, confirming that learnable weighted mixed-scale fusion generalizes broadly while remaining data efficient, fully interpretable, and overhead-free at inference. Our implementation is available at https://github.com/SLDGroup/LoMix. Md Mostafijur Rahman, Radu Marculescu |
NeurIPS | 2 |
| 2025 | Mitigating Intra- and Inter-modal Forgetting in Continual Learning of Unified Multimodal ModelsabstractUnified Multimodal Generative Models (UMGMs) unify visual understanding and image generation within a single autoregressive framework. However, their ability to continually learn new tasks is severely hindered by catastrophic forgetting, both within a modality (intra-modal) and across modalities (inter-modal). While intra-modal forgetting has been studied in prior continual learning (CL) work, inter-modal forgetting remains largely unexplored. In this paper, we identify and empirically validate this phenomenon in UMGMs and provide a theoretical explanation rooted in gradient conflict between modalities. To address both intra- and inter-modal forgetting, we propose Modality-Decoupled Experts (MoDE), a lightweight and scalable architecture that isolates modality-specific updates to mitigate the gradient conflict and leverages knowledge distillation to prevent catastrophic forgetting and preserve pre-trained capabilities. Unlike previous CL methods that remain modality-coupled and suffer from modality gradient conflict, MoDE explicitly decouples modalities to prevent interference. Experiments across diverse benchmarks demonstrate that MoDE significantly mitigates both inter- and intra-modal forgetting, outperforming prior CL baselines in unified multimodal generation settings. Xiwen Wei, Mustafa Munir, Radu Marculescu |
NeurIPS | 3 |
| 2025 | Demo Abstract: Lightweight Training and Inference for Self-Supervised Depth Estimation on Edge DevicesabstractDeploying monocular depth estimation on resource-constrained edge devices is a significant challenge, particularly when attempting to perform both training and inference concurrently. Current lightweight, self-supervised approaches typically rely on complex frameworks that are hard to implement and deploy in real-world settings. To address this gap, we introduce the first framework for Lightweight Training and Inference (LITI) that combines ready-to-deploy models with streamlined code and fully functional, parallel training and inference pipelines. Our experiments show various models being deployed for inference, training, or both inference and training, leveraging inputs from a real-time RGB camera sensor. Thus, our framework enables training and inference on resource-constrained edge devices for complex applications such as depth estimation. Our framework is available at: https://github.com/SLDGroup/LITI. Allen-Jasmin Farcas, Radu Marculescu |
SenSys | 2 |
| 2025 | Ada-VE: Training-Free Consistent Video Editing Using Adaptive Motion PriorabstractVideo-to- video synthesis poses significant challenges in maintaining character consistency, smooth temporal tran-sitions, and preserving visual quality during fast motion. While recent fully cross-frame self-attention mechanisms have improved character consistency across multiple frames, they come with high computational costs and often include re-dundant operations, especially for videos with higher frame rates. To address these inefficiencies, we propose an adaptive motion-guided cross-frame attention mechanism that selectively reduces redundant computations. This enables a greater number of cross-frame attentions over more frames within the same computational budget, thereby enhancing both video quality and temporal coherence. Our method leverages optical flow to focus on moving regions while sparsely attending to stationary areas, allowing for the joint editing of more frames without increasing computational demands. Traditional frame interpolation techniques struggle with motion blur and flickering in intermediate frames, which compromises visual fidelity. To mitigate this, we intro-duce KV-caching for jointly edited frames, reusing keys and values across intermediate frames to preserve visual quality and maintain temporal consistency throughout the video. With our adaptive cross-frame self-attention approach, we achieve a threefold increase in the number of keyframes processed compared to existing methods, all within the same computational budget as fully cross-frame attention base-lines. This results in significant improvements in prediction accuracy and temporal consistency, outperforming state-of-the-art approaches. Code is made publicly available at https://github.com/tanvir-utexaslAdaVEltree/main. Tanvir Mahmud, Mustafa Munir, Radu Marculescu, Diana Marculescu |
WACV | 3 |
| 2025 | RapidNet: Multi-Level Dilated Convolution Based Mobile BackboneabstractVision transformers (ViTs) have dominated computer vision in recent years. However, ViTs are computationally expensive and not well suited for mobile devices; this led to the prevalence of convolutional neural network (CNN) and ViT-based hybrid models for mobile vision applications. Recently, Vision GNN (ViG) and CNN hybrid models have also been proposed for mobile vision tasks. However, all of these methods remain slower compared to pure CNN-based models. In this work, we propose Multi-Level Dilated Convolutions to devise a purely CNN-based mobile backbone. Using Multi-Level Dilated Convolutions allows for a larger theoretical receptive field than standard convolutions. Different levels of dilation also allow for interactions between the short-range and long-range features in an image. Experiments show that our proposed model outperforms state-of-the-art (SOTA) mobile CNN, ViT, ViG, and hybrid architectures in terms of accuracy and/or speed on image classification, object detection, instance segmentation, and semantic segmentation. Our fastest model, RapidNet-Ti, achieves 76.3% top-1 accuracy on ImageNet-1K with 0.9 ms inference latency on an iPhone 13 mini NPU, which is faster and more accurate than MobileNetV2×1.4 (74.7% top-1 with 1.0 ms latency). Our work shows that pure CNN architectures can beat SOTA hybrid and ViT models in terms of accuracy and speed when designed properly11Code: https://github.com/mmunir127/RapidNet-Official: Mustafa Munir, Md Mostafijur Rahman, Radu Marculescu |
WACV | 3 |
| 2025 | Online-LoRA: Task-Free Online Continual Learning via Low Rank AdaptationabstractCatastrophic forgetting is a significant challenge in online continual learning (OCL), especially for nonstationary data streams that do not have well-defined task boundaries. This challenge is exacerbated by the memory constraints and privacy concerns inherent in rehearsal buffers. To tackle catastrophic forgetting, in this paper, we introduce Online-LoRA, a novel framework for task-free OCL. Online-LoRA allows to finetune pre-trained Vision Transformer (ViT) models in real-time to address the limitations of rehearsal buffers and leverage pre-trained models' performance benefits. As the main contribution, our approach features a novel online weight regularization strategy to identify and consolidate important model parameters. Moreover, Online-LoRA leverages the training dynamics of loss values to enable the automatic recognition of the data distribution shifts. Extensive experiments across many task-free OCL scenarios and benchmark datasets (including CIFAR-100, ImageNet-R, ImageNet-S, CUB-200 and CORe50) demonstrate that Online-LoRA can be robustly adapted to various ViT architectures, while achieving better performance compared to SOTA methods11Code: https://github.com/Christina200/online-LoRA-official.git. Xiwen Wei, Guihong Li, Radu Marculescu |
WACV | 3 |
| 2024 | Packet Pruning: Finding Better Energy Spending Policies for Batteryless Human Activity RecognitionabstractBatteryless sensing devices which rely on energy harvesting can enable more sustainable and long-lasting Internet of Things (IoT) based wearables. While it has become feasible to implement energy-harvesting based wearables for digital health applications, it remains challenging to integrate such devices and the data they collect into machine learning pipelines for tasks such as human activity recognition (HAR). A key obstacle is uncertainty in the data acquisition process. Given the discontinuous and uncertain availability of harvested energy, when should a sensor spend energy to sample and transmit data packets for processing? A common approach is to spend energy opportunistically by sending packets whenever sufficient energy is available. However, when considering a specific task, namely HAR with kinetic energy harvesting based sensors, this approach unfairly prioritizes data from activities where more energy can be harvested (e.g., running). In this work, we improve the opportunistic energy spending policy by pruning redundant packets to reallocate energy towards activities where less energy is harvested. Our approach results in an increase in the F1-score of ‘lower energy’ activities while having a minimal impact on the F1-score of ‘higher energy’ activities.11Code: https://github.com/SLDGroup/PacketPruning Geffen Cooper, Radu Marculescu |
BSN | 2 |
| 2024 | A Tiny Supervised ODL Core with Auto Data Pruning for Human Activity RecognitionabstractIn this paper, we introduce a low-cost and lowpower tiny supervised on-device learning (ODL) core that can address the distributional shift of input data for human activity recognition. Although ODL for resource-limited edge devices has been studied recently, how exactly to provide the training labels to these devices at runtime remains an open-issue. To address this problem, we propose to combine an automatic data pruning with supervised ODL to reduce the number queries needed to acquire predicted labels from a nearby teacher device and thus save power consumption during model retraining. The data pruning threshold is automatically tuned, eliminating a manual threshold tuning. As a tinyML solution at a few mW for the human activity recognition, we design a supervised ODL core that supports our automatic data pruning using a 45nm CMOS process technology. We show that the required memory size for the core is smaller than the same-shaped multilayer perceptron (MLP) and the power consumption is only 3.39m W. Experiments using a human activity recognition dataset show that the proposed automatic data pruning reduces the communication volume by 55.7% and power consumption accordingly with only 0.9 % accuracy loss. Hiroki Matsutani, Radu Marculescu |
BSN | 2 |
| 2024 | GreedyViG: Dynamic Axial Graph Construction for Efficient Vision GNNsabstractVision graph neural networks (ViG) offer a new avenue for exploration in computer vision. A major bottleneck in ViGs is the inefficient k-nearest neighbor (KNN) operation used for graph construction. To solve this issue, we propose a new method for designing ViGs, Dynamic Axial Graph Construction (DAGC), which is more efficient than KNN as it limits the number of considered graph connections made within an image. Additionally, we propose a novel CNN-GNN architecture, GreedyViG, which uses DAGC. Extensive experiments show that GreedyViG beats existing ViG, CNN, and ViT architectures in terms of accuracy, GMACs, and parameters on image classification, object detection, instance segmentation, and semantic segmentation tasks. Our smallest model, GreedyViG-S, achieves 81.1% top-1 accuracy on ImageNet-1K, 2.9% higher than Vision GNN and 2.2% higher than Vision HyperGraph Neural Network (ViHGNN), with less GMACs and a similar number of parameters. Our largest model, GreedyViG-B obtains 83.9% top-1 accuracy, 0.2% higher than Vision GNN, with a 66.6% decrease in parameters and a 69% decrease in GMACs. GreedyViG-B also obtains the same accuracy as ViHGNN with a 67.3% decrease in parameters and a 71.3% decrease in GMACs. Our work shows that hybrid CNN-GNN architectures not only provide a new avenue for designing efficient models, but that they can also exceed the performance of current state-of-the-art models11Code: https://github.com/SLDGroup/GreedyViG.. Mustafa Munir, William Avery, Md Mostafijur Rahman, Radu Marculescu |
CVPR | 4 |
| 2024 | EMCAD: Efficient Multi-Scale Convolutional Attention Decoding for Medical Image SegmentationabstractAn efficient and effective decoding mechanism is crucial in medical image segmentation, especially in scenarios with limited computational resources. However, these decoding mechanisms usually come with high computational costs. To address this concern, we introduce EMCAD, a new efficient multi-scale convolutional attention decoder, designed to optimize both performance and computational efficiency. EMCAD leverages a unique multi-scale depth-wise convolution block, significantly enhancing feature maps through multi-scale convolutions. EMCAD also employs channel, spatial, and grouped (large-kernel) gated attention mechanisms, which are highly effective at capturing intricate spatial relationships while focusing on salient regions. By employing group and depth-wise convolution, EMCAD is very efficient and scales well (e.g., only 1.91M parameters and 0.381G FLOPs are needed when using a standard encoder). Our rigorous evaluations across 12 datasets that belong to six medical image segmentation tasks reveal that EMCAD achieves state-of-the-art (SOTA) performance with 79.4% and 80.3% reduction in #Params and #FLOPs, respectively. Moreover, EMCAD's adaptability to different encoders and versatility across segmentation tasks further establish EMCAD as a promising tool, advancing the field towards more efficient and accurate medical image analysis. Our implementation is available at https://github.com/SLDGroupIEMCAD. Md Mostafijur Rahman, Mustafa Munir, Radu Marculescu |
CVPR | 3 |
| 2024 | Machine Unlearning for Image-to-Image Generative ModelsabstractMachine unlearning has emerged as a new paradigm to deliberately forget data samples from a given model in order to adhere to stringent regulations.
However, existing machine unlearning methods have been primarily focused on classification models, leaving the landscape of unlearning for generative models relatively unexplored.
This paper serves as a bridge, addressing the gap by providing a unifying framework of machine unlearning for image-to-image generative models.
Within this framework, we propose a computationally-efficient algorithm, underpinned by rigorous theoretical analysis, that demonstrates negligible performance degradation on the retain samples, while effectively removing the information from the forget samples.
Empirical studies on two large-scale datasets, ImageNet-1K and Places-365, further show that our algorithm does not rely on the availability of the retain samples, which further complies with data retention policy.
To our best knowledge, this work is the first that represents systemic, theoretical, empirical explorations of machine unlearning specifically tailored for image-to-image generative models. Guihong Li, Hsiang Hsu, Chun-Fu Chen 0001, Radu Marculescu |
ICLR | 4 |
| 2024 | G-CASCADE: Efficient Cascaded Graph Convolutional Decoding for 2D Medical Image SegmentationabstractIn this paper, we are the first to propose a new graph convolution-based decoder namely, Cascaded Graph Convolutional Attention Decoder (G-CASCADE), for 2D medical image segmentation. G-CASCADE progressively refines multi-stage feature maps generated by hierarchical transformer encoders with an efficient graph convolution block. The encoder utilizes the self-attention mechanism to capture long-range dependencies, while the decoder refines the feature maps preserving long-range information due to the global receptive fields of the graph convolution block. Rigorous evaluations of our decoder with multiple transformer encoders on five medical image segmentation tasks (i.e., Abdomen organs, Cardiac organs, Polyp lesions, Skin lesions, and Retinal vessels) show that our model outperforms other state-of-the-art (SOTA) methods. We also demonstrate that our decoder achieves better DICE scores than the SOTA CASCADE decoder with 80.8% fewer parameters and 82.3% fewer FLOPs. Our decoder can easily be used with other hierarchical encoders for general-purpose semantic and medical image segmentation tasks. The implementation can be found at: https://github.com/SLDGroup/G-CASCADE. Md Mostafijur Rahman, Radu Marculescu |
WACV | 2 |
| 2024 | Zero-Shot Neural Architecture Search: Challenges, Solutions, and OpportunitiesabstractRecently, zero-shot (or training-free) Neural Architecture Search (NAS) approaches have been proposed to liberate NAS from the expensive training process. The key idea behind zero-shot NAS approaches is to design proxies that can predict the accuracy of some given networks without training the network parameters. The proxies proposed so far are usually inspired by recent progress in theoretical understanding of deep learning and have shown great potential on several datasets and NAS benchmarks. This paper aims to comprehensively review and compare the state-of-the-art (SOTA) zero-shot NAS approaches, with an emphasis on their hardware awareness. To this end, we first review the mainstream zero-shot proxies and discuss their theoretical underpinnings. We then compare these zero-shot proxies through large-scale experiments and demonstrate their effectiveness in both hardware-aware and hardware-oblivious NAS scenarios. Finally, we point out several promising ideas to design better proxies. Guihong Li, Duc Hoang, Kartikeya Bhardwaj, Ming Lin 0002, Zhangyang Wang, Radu Marculescu |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Quarantine in Motion: A Graph Learning and Multi-Agent Reinforcement Learning Framework to Reduce Disease Transmission Without LockdownabstractExposure notification applications are designed to help trace disease spreading by alerting exposed individuals to get tested. However, false alarms can cause users to become hesitant to respond, making the applications ineffective. To address the shortcomings of slow manual contact tracing, costly lockdowns, and unreliable exposure notification applications, better disease mitigation strategies are needed. In this paper, we propose a new disease mitigation paradigm where people can reduce infection spreading while maintaining some mobility (i.e., Quarantine in Motion). Our approach utilizes Graph Neural Networks (GNNs) to predict disease hotspots such as restaurants, shops and parks, and Multi-Agent Reinforcement Learning (MARL) to collaboratively manage human mobility to reduce disease transmission. As proof of concept, we simulate an infection using real-world mobility data from New York City (over 200,000 devices) and Austin (over 36,000 devices) and train 10,000 agents from each city to manage disease dynamics. Through simulation, we show that a trained population suppresses their reproduction rate below 1, thereby mitigating the outbreak. Sofia Hurtado, Radu Marculescu |
ASONAM | 2 |
| 2023 | Efficient On-Device Training via Gradient FilteringabstractDespite its importance for federated learning, continuous learning and many other applications, on-device training remains an open problem for EdgeAI. The problem stems from the large number of operations (e.g., floating point multiplications and additions) and memory consumption required during training by the back-propagation algorithm. Consequently, in this paper, we propose a new gradient filtering approach which enables on-device CNN model training. More precisely, our approach creates a special structure with fewer unique elements in the gradient map, thus significantly reducing the computational complexity and memory consumption of back propagation during training. Extensive experiments on image classification and semantic segmentation with multiple CNN models (e.g., MobileNet, DeepLabV3, UPerNet) and devices (e.g., Raspberry Pi and Jetson Nano) demonstrate the effectiveness and wide applicability of our approach. For example, compared to SOTA, we achieve up to 19× speedup and 77.1% memory savings on ImageNet classification with only 0.1% accuracy loss. Finally, our method is easy to implement and deploy; over 20× speedup and 90% energy savings have been observed compared to highly optimized baselines in MKLDNN and CUDNN on NVIDIA Jetson Nano. Consequently, our approach opens up a new direction of research with a huge potential for on-device training.11Code: https://github.com/SLDGroup/GradientFilter-CVPR23 Yuedong Yang, Guihong Li, Radu Marculescu |
CVPR | 3 |
| 2023 | Revisiting Pruning at Initialization Through the Lens of Ramanujan Graph
Duc N. M. Hoang, Shiwei Liu 0003, Radu Marculescu, Zhangyang Wang |
ICLR | 3 |
| 2023 | ZiCo: Zero-shot NAS via inverse Coefficient of Variation on Gradients
Guihong Li, Yuedong Yang, Kartikeya Bhardwaj, Radu Marculescu |
ICLR | 4 |
| 2023 | TIPS: Topologically Important Path Sampling for Anytime Neural NetworksabstractAnytime neural networks (AnytimeNNs) are a promising solution to adaptively adjust the model complexity at runtime under various hardware resource constraints. However, the manually-designed AnytimeNNs are biased by designers’ prior experience and thus provide sub-optimal solutions. To address the limitations of existing hand-crafted approaches, we first model the training process of AnytimeNNs as a discrete-time Markov chain (DTMC) and use it to identify the paths that contribute the most to the training of AnytimeNNs. Based on this new DTMC-based analysis, we further propose TIPS, a framework to automatically design AnytimeNNs under various hardware constraints. Our experimental results show that TIPS can improve the convergence rate and test accuracy of AnytimeNNs. Compared to the existing AnytimeNNs approaches, TIPS improves the accuracy by 2%-6.6% on multiple datasets and achieves SOTA accuracy-FLOPs tradeoffs. Guihong Li, Kartikeya Bhardwaj, Yuedong Yang, Radu Marculescu |
ICML | 4 |
| 2023 | Efficient Low-rank Backpropagation for Vision Transformer AdaptationabstractThe increasing scale of vision transformers (ViT) has made the efficient fine-tuning of these large models for specific needs a significant challenge in various applications. This issue originates from the computationally demanding matrix multiplications required during the backpropagation process through linear layers in ViT.
In this paper, we tackle this problem by proposing a new Low-rank BackPropagation via Walsh-Hadamard Transformation (LBP-WHT) method. Intuitively, LBP-WHT projects the gradient into a low-rank space and carries out backpropagation. This approach substantially reduces the computation needed for adapting ViT, as matrix multiplication in the low-rank space is far less resource-intensive. We conduct extensive experiments with different models (ViT, hybrid convolution-ViT model) on multiple datasets to demonstrate the effectiveness of our method. For instance, when adapting an EfficientFormer-L1 model on CIFAR100, our LBP-WHT achieves 10.4\% higher accuracy than the state-of-the-art baseline, while requiring 9 MFLOPs less computation.
As the first work to accelerate ViT adaptation with low-rank backpropagation, our LBP-WHT method is complementary to many prior efforts and can be combined with them for better performance. Yuedong Yang, Hung-Yueh Chiang, Guihong Li, Diana Marculescu, Radu Marculescu |
NeurIPS | 5 |
| 2023 | Medical Image Segmentation via Cascaded Attention DecodingabstractTransformers have shown great promise in medical image segmentation due to their ability to capture long-range dependencies through self-attention. However, they lack the ability to learn the local (contextual) relations among pixels. Previous works try to overcome this problem by embedding convolutional layers either in the encoder or decoder modules of transformers thus ending up sometimes with inconsistent features. To address this issue, we propose a novel attention-based decoder, namely CASCaded Attention DEcoder (CASCADE), which leverages the multi-scale features of hierarchical vision transformers. CASCADE consists of i) an attention gate which fuses features with skip connections and ii) a convolutional attention module that enhances the long-range and local context by suppressing background information. We use a multi-stage feature and loss aggregation framework due to their faster convergence and better performance. Our experiments demonstrate that transformers with CASCADE significantly outperform state-of-the-art CNN- and transformer-based approaches, obtaining up to 5.07% and 6.16% improvements in DICE and mIoU scores, respectively. CASCADE opens new ways of designing better attention-based decoders. Md Mostafijur Rahman, Radu Marculescu |
WACV | 2 |
| 2023 | SUGAR: Efficient Subgraph-Level Training via Resource-Aware Graph PartitioningabstractGraph Neural Networks (GNNs) have demonstrated a great potential in a variety of graph-based applications, such as recommender systems, drug discovery, and object recognition. Nevertheless, resource-efficient GNN learning is a rarely explored topic despite its many benefits for edge computing and Internet of Things (IoT) applications. To improve this state of affairs, this work proposes efficientsubgraph-level training viaresource-aware graph partitioning (SUGAR). SUGAR first partitions the initial graph into a set of disjoint subgraphs and then performs local training at the subgraph-level We provide a theoretical analysis and conduct extensive experiments on five graph benchmarks to verify its efficacy in practice. Our results across five different hardware platforms demonstrate great runtime speedup and memory reduction of SUGAR on large-scale graphs. We believe SUGAR opens a new research direction towards developing GNN methods that are resource-efficient, hence suitable for IoT deployment. Zihui Xue, Yuedong Yang, Radu Marculescu |
IEEE Trans. Computers | 3 |
| 2023 | Domain-Specific Architectures: Research Problems and Promising ApproachesabstractProcess technology-driven performance and energy efficiency improvements have slowed down as we approach physical design limits. General-purpose manycore architectures attempt to circumvent this challenge, but they have a significant performance and energy-efficient gap compared to special-purpose solutions. Domain-specific architectures (DSAs), an instance of heterogeneous architectures, efficiently combine general-purpose cores and specialized hardware accelerators to boost energy efficiency and provide programming flexibility. Indeed, the hardware, software, and systems aspects in DSAs are highly tailored to maximize the energy efficiency of applications in a target domain. As DSAs and their conceptualization advance rapidly, there is a strong need to understand the research problems that need immediate attention. This article discusses the primary research directions in the design and runtime management of DSAs. Then, it surveys some promising approaches and highlights the outstanding research needs. Anish Krishnakumar, Ümit Y. Ogras, Radu Marculescu, Michael Kishinevsky, Trevor N. Mudge |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2023 | Introduction to the Special Issue on Domain-Specific System-on-Chip Architectures and Run-Time Management TechniquesabstractDomain-specific systems-on-chip (DSSoCs), a class of heterogeneous many-core systems, are recognized as a promising approach to narrowing down the performance and energy-efficiency gap between custom hardware accelerators and programmable processors.However, fulfilling this promise depends on successfully addressing a number of fundamental research questions.For instance, given a target domain, a designer must develop a suitable architecture and determine the set of appropriate hardware accelerators.While integrating too many accelerators would increase the design cost, missing critical accelerators can undermine the system performance and energy efficiency.Typically, a rich set of accelerators can dramatically lower the processing times.Hence, the rest of the system components, such as the on-chip communication, must also match the high performance requirements and enable nanosecond-level latencies between the IP blocks and accelerators.DSSoCs must also provide software tools, application programming interfaces (APIs), and accelerator interfaces such that application developers can utilize them efficiently.Finally, a range of runtime management methodologies and algorithms are required to make the best use of the DSSoC resources and power budgets.This Special Issue presents eleven research papers and a survey targeting these topics selected from over 40 submissions.It represents a remarkable collective effort involving both the academic and industrial research communities.The articles in this issue present novel and impactful solutions to important research problems, including novel device technologies, hardware accelerators, high-level synthesis techniques, design space exploration, scheduling, virtualization, compiler, and test techniques for domain-specific designs.The survey article titled "Domain-Specific Architectures (DSAs): Research Problems and Promising Approaches" provides a comprehensive overview of various research directions, outstanding challenges, and promising approaches in DSSoC system design.Starting from the lowest level of abstraction, the article "Experimental Demonstration of STT-MRAM-based Nonvolatile Instantly On/Off System: Case Studies" presents a solution for IoT applications using nonvolatile STT-MRAMs.Experimental results show 15.1% lower power consumption with two orders of magnitude faster data restore time.At the hardware design level, the paper titled "SHARP: An Adaptable, Energy-Efficient Accelerator for Recurrent Neural Network" identifies adaptiveness as a key feature missing from existing RNN accelerators.It proposes an intelligent tiled-based dispatching mechanism to efficiently handle the data dependencies.The Ümit Y. Ogras, Radu Marculescu, Trevor N. Mudge, Michael Kishinevsky |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2022 | Quarantine in Motion: A Graph Learning Framework to Reduce Disease Transmission Without LockdownabstractExposure notification applications are developed to increase the scale and speed of disease contact tracing. Indeed, by taking advantage of Bluetooth technology, they track the infected population's mobility and then inform close contacts to get tested. In this paper, we ask whether these applications can extend from reactive to preemptive risk management tools? To this end, we propose a new framework that utilizes graph neural networks (GNN) and real-world Foursquare mobility data to predict high risk locations on an hourly basis. As a proof of concept, we then simulate a risk-informed Foursquare population of over 36,000 people in Austin TX after the peak of an outbreak. We find that even after 50% of the population has been infected with COVID-19, they can still maintain their mobility, while reducing the new infections by 13%. Consequently, these results are a first step towards achieving what we call Quarantine in Motion. Sofia Hurtado, Radu Marculescu, Justin A. Drake |
ASONAM | 2 |
| 2022 | INDENT: Incremental Online Decision Tree Training for Domain-Specific Systems-on-ChipabstractThe performance and energy efficiency potential of heterogeneous architectures has fueled domain-specific systems-on-chip (DSSoCs) that integrate general-purpose and domain-specialized hardware accelerators. Decision trees (DTs) perform high-quality, low-latency task scheduling to utilize the massive parallelism and heterogeneity in DSSoCs effectively. However, offline trained DT scheduling policies can quickly become ineffective when applications or hardware configurations change. There is a critical need for runtime techniques to train DTs incrementally without sacrificing accuracy since current training approaches have large memory and computational power requirements. To address this need, we propose INDENT, an incremental online DT framework to update the scheduling policy and adapt it to unseen scenarios. INDENT updates DT schedulers at runtime using only 1--8% of the original training data embedded during training. Thorough evaluations with hardware platforms and DSSoC simulators demonstrate that INDENT performs within 5% of a DT trained from scratch using the entire dataset and outperforms current state-of-the-art approaches. Anish Krishnakumar, Radu Marculescu, Ümit Y. Ogras |
ICCAD | 2 |
| 2022 | Enabling Software-Defined RF Convergence with a Novel Coarse-Scale Heterogeneous ProcessorabstractRF system development is traditionally constrained by a restrictive trade-off between power efficiency and programmatic flexibility. We outline a path towards achieving both, thereby enabling a range of new system concepts that better utilize limited resources. As an example, for many future applications, we consider RF convergence – reusing the same spectrum and waveforms to achieve multiple distributed system functions and goals, simultaneously. To enable this next step in processing, we develop a novel framework that includes both software and the system-on-chip (SoC) design. Daniel W. Bliss, Tutu Ajayi, Ali Akoglu, Ilkin Aliyev, Toygun Basaklar, Leul Belayneh, David T. Blaauw, John S. Brunhaver, Chaitali Chakrabarti, Liangliang Chang, Kuan-Yu Chen 0001, Ming-Hung Chen, Xing Chen 0004, Alex R. Chiriyath, Alhad Daftardar, Ronald G. Dreslinski, Arindam Dutta, Allen-Jasmin Farcas, Yukang Fu, A. Alper Goksoy, Xin He 0011, Md Sahil Hassan, Andrew Herschfelt, Jacob Holtom, Hun-Seok Kim, Anish Krishnakumar, Owen Ma, Joshua Mack, Saurav Mallik, Sumit K. Mandal, Radu Marculescu, Brittany M. McCall, Trevor N. Mudge, Ümit Y. Ogras, Vishrut Pandey, Saquib Ahmad Siddiqui, Yu-Hsiu Sun, Adarsh A. Venkataramani, Xiangdong Wei, Benjamin R. Willis, Hanguang Yu, Yufan Yue |
ISCAS | 32 |
| 2021 | Pruning digital contact networks for meso-scale epidemic surveillance using foursquare dataabstractWith the recent advances in human sensing, the push to integrate human mobility tracking with epidemic modeling highlights the lack of groundwork at the mesoscale (e.g., city-level) for both contact tracing and transmission dynamics. Although GPS data has been used to study city-level outbreaks in the past, existing approaches fail to capture the path of infection at the individual level. Consequently, in this paper, we extend epidemics prediction from estimating the size of an outbreak at the population level to estimating the individuals who may likely get infected within a finite period of time. To this end, we propose a network science based method to first build and then prune the dynamic contact networks for recurring interactions; these networks can serve as the backbone topology for mechanistic epidemics modeling. We test our method using Foursquare's Points of Interest (POI) smart phone geolocation data from over 1.3 million devices to better approximate the COVID-19 infection curves for two major (yet very different) US cities, (i.e., Austin and New York City), while maintaining the granularity of individual transmissions and reducing model uncertainty. Our method provides a foundation for building a disease prediction framework at the mesoscale that can help both policy makers and individuals better understand their estimated state of health and help the pandemic mitigation efforts. Sofia Hurtado, Radu Marculescu, Justin A. Drake, Ravi Srinivasan |
ASONAM | 2 |
| 2021 | How Does Topology Influence Gradient Propagation and Model Performance of Deep Networks With DenseNet-Type Skip Connections?abstractDenseNets introduce concatenation-type skip connections that achieve state-of-the-art accuracy in several computer vision tasks. In this paper, we reveal that the topology of the concatenation-type skip connections is closely related to the gradient propagation which, in turn, enables a predictable behavior of DNNs’ test performance. To this end, we introduce a new metric called NN-Mass to quantify how effectively information flows through DNNs. Moreover, we empirically show that NN-Mass also works for other types of skip connections, e.g., for ResNets, Wide-ResNets (WRNs), and MobileNets, which contain addition-type skip connections (i.e., residuals or inverted residuals). As such, for both DenseNet-like CNNs and ResNets/WRNs/MobileNets, our theoretically grounded NN-Mass can identify models with similar accuracy, despite having significantly different size/compute requirements. Detailed experiments on both synthetic and real datasets (e.g., MNIST, CIFAR-10, CIFAR-100, ImageNet) provide extensive evidence for our insights. Finally, the closed-form equation of our NN-Mass enables us to design significantly compressed DenseNets (for CIFAR-10) and MobileNets (for ImageNet) directly at initialization without time-consuming training and/or searching.1 Kartikeya Bhardwaj, Guihong Li, Radu Marculescu |
CVPR | 3 |
| 2021 | FLASH: Fast Neural Architecture Search with Hardware OptimizationabstractNeural architecture search (NAS) is a promising technique to design efficient and high-performance deep neural networks (DNNs). As the performance requirements of ML applications grow continuously, the hardware accelerators start playing a central role in DNN design. This trend makes NAS even more complicated and time-consuming for most real applications. This paper proposes FLASH, a very fast NAS methodology that co-optimizes the DNN accuracy and performance on a real hardware platform. As the main theoretical contribution, we first propose the NN-Degree, an analytical metric to quantify the topological characteristics of DNNs with skip connections (e.g., DenseNets, ResNets, Wide-ResNets, and MobileNets). The newly proposed NN-Degree allows us to do training-free NAS within one second and build an accuracy predictor by training as few as 25 samples out of a vast search space with more than 63 billion configurations. Second, by performing inference on the target hardware, we fine-tune and validate our analytical models to estimate the latency, area, and energy consumption of various DNN architectures while executing standard ML datasets. Third, we construct a hierarchical algorithm based on simplicial homology global optimization (SHGO) to optimize the model-architecture co-design process, while considering the area, latency, and energy consumption of the target hardware. We demonstrate that, compared to the state-of-the-art NAS approaches, our proposed hierarchical SHGO-based algorithm enables more than four orders of magnitude speedup (specifically, the execution time of the proposed algorithm is about 0.1 seconds). Finally, our experimental evaluations show that FLASH is easily transferable to different hardware architectures, thus enabling us to do NAS on a Raspberry Pi-3B processor in less than 3 seconds. Guihong Li, Sumit K. Mandal, Ümit Y. Ogras, Radu Marculescu |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2020 | INVITED: New Directions in Distributed Deep Learning: Bringing the Network at Forefront of IoT DesignabstractIn this paper, we first highlight three major challenges to large-scale adoption of deep learning at the edge: (i) Hardware-constrained IoT devices, (ii) Data security and privacy in the IoT era, and (iii) Lack of network-aware deep learning algorithms for distributed inference across multiple IoT devices. We then provide a unified view targeting three research directions that naturally emerge from the above challenges: (1) Federated learning for training deep networks, (2) Data-independent deployment of learning algorithms, and (3) Communication-aware distributed inference. We believe that the above research directions need a network-centric approach to enable the edge intelligence and, therefore, fully exploit the true potential of IoT. Kartikeya Bhardwaj, Wei Chen 0124, Radu Marculescu |
DAC | 3 |
| 2020 | On Network Science and Mutual Information for Explaining Deep Neural NetworksabstractIn this paper, we present a new approach to interpret deep learning models. By coupling mutual information with network science, we explore how information flows through feedforward networks. We show that efficiently approximating mutual information allows us to create an information measure that quantifies how much information flows between any two neurons of a deep learning model. To that end, we propose NIF, Neural Information Flow, a technique for codifying information flow that exposes deep learning model internals and provides feature attributions. Umang Bhatt, Kartikeya Bhardwaj, Radu Marculescu, José M. F. Moura |
ICASSP | 4 |
| 2020 | Edge AI: Systems Design and ML for IoT Data AnalyticsabstractWith the explosion in Big Data, it is often forgotten that much of the data nowadays is generated at the edge. Specifically, a major source of data is users' endpoint devices like phones, smart watches, etc., that are connected to the internet, also known as the Internet-of-Things (IoT). This "edge of data" faces several new challenges related to hardware-constraints, privacy-aware learning, and distributed learning (both training as well as inference). So what systems and machine learning algorithms can we use to generate or exploit data at the edge? Can network science help us solve machine learning (ML) problems? Can IoT-devices help people who live with some form of disability and many others benefit from health monitoring? Radu Marculescu, Diana Marculescu, Ümit Y. Ogras |
KDD | 1 |
| 2020 | FedMAX: Mitigating Activation Divergence for Accurate and Communication-Efficient Federated Learning
Wei Chen 0124, Kartikeya Bhardwaj, Radu Marculescu |
ECML/PKDD (2) | 3 |
| 2020 | DS3: A System-Level Domain-Specific System-on-Chip Simulation FrameworkabstractHeterogeneous systems-on-chip (SoCs) are highly favorable computing platforms due to their superior performance and energy efficiency potential compared to homogeneous architectures. They can be further tailored to a specific domain of applications by incorporating processing elements (PEs) that accelerate frequently used kernels in these applications. However, this potential is contingent upon optimizing the SoC for the target domain and utilizing its resources effectively at runtime. To this end, system-level design - including scheduling, power-thermal management algorithms and design space exploration studies - plays a crucial role. This article presents a system-level domain-specific SoC simulation (DS3) framework to address this need. DS3 enables both design space exploration and dynamic resource management for power-performance optimization of domain applications. We showcase DS3 using six real-world applications from wireless communications and radar processing domain. DS3, as well as the reference applications, is shared as open-source software to stimulate research in this area. Samet E. Arda, Anish Krishnakumar, A. Alper Goksoy, Nirmal Kumbhare, Joshua Mack, Anderson Luiz Sartor, Ali Akoglu, Radu Marculescu, Ümit Y. Ogras |
IEEE Trans. Computers | 8 |
| 2020 | Runtime Task Scheduling Using Imitation Learning for Heterogeneous Many-Core SystemsabstractDomain-specific systems-on-chip, a class of heterogeneous many-core systems, is recognized as a key approach to narrow down the performance and energy-efficiency gap between custom hardware accelerators and programmable processors. Reaching the full potential of these architectures depends critically on optimally scheduling the applications to available resources at runtime. Existing optimization-based techniques cannot achieve this objective at runtime due to the combinatorial nature of the task scheduling problem. As the main theoretical contribution, this article poses scheduling as a classification problem and proposes a hierarchical imitation learning (IL)-based scheduler that learns from an Oracle to maximize the performance of multiple domain-specific applications. Extensive evaluations with six streaming applications from wireless communications and radar domains show that the proposed IL-based scheduler approximates an offline Oracle policy with more than 99% accuracy for performance- and energy-based optimization objectives. Furthermore, it achieves almost identical performance to the Oracle with a low runtime overhead and successfully adapts to new applications, many-core system configurations, and runtime variations in application characteristics. Anish Krishnakumar, Samet E. Arda, A. Alper Goksoy, Sumit K. Mandal, Ümit Y. Ogras, Anderson Luiz Sartor, Radu Marculescu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2019 | MetaNN: accurate classification of host phenotypes from metagenomic data using neural networksabstractBACKGROUND: Microbiome profiles in the human body and environment niches have become publicly available due to recent advances in high-throughput sequencing technologies. Indeed, recent studies have already identified different microbiome profiles in healthy and sick individuals for a variety of diseases; this suggests that the microbiome profile can be used as a diagnostic tool in identifying the disease states of an individual. However, the high-dimensional nature of metagenomic data poses a significant challenge to existing machine learning models. Consequently, to enable personalized treatments, an efficient framework that can accurately and robustly differentiate between healthy and sick microbiome profiles is needed. RESULTS: In this paper, we propose MetaNN (i.e., classification of host phenotypes from Metagenomic data using Neural Networks), a neural network framework which utilizes a new data augmentation technique to mitigate the effects of data over-fitting. CONCLUSIONS: We show that MetaNN outperforms existing state-of-the-art models in terms of classification accuracy for both synthetic and real metagenomic data. These results pave the way towards developing personalized treatments for microbiome related diseases. Chieh Lo, Radu Marculescu |
BMC Bioinform. | 2 |
| 2019 | Learning-Based Application-Agnostic 3D NoC Design for Heterogeneous Manycore SystemsabstractThe rising use of deep learning and other big-data algorithms has led to an increasing demand for hardware platforms that are computationally powerful, yet energy-efficient. Due to the amount of data parallelism in these algorithms, high-performance three-dimensional (3D) manycore platforms that incorporate both CPUs and GPUs present a promising direction. However, as systems use heterogeneity (e.g., a combination of CPUs, GPUs, and accelerators) to improve performance and efficiency, it becomes more pertinent to address the distinct and likely conflicting communication requirements (e.g., CPU memory access latency or GPU network throughput) that arise from such heterogeneity. Unfortunately, it is difficult to quickly explore the hardware design space and choose appropriate tradeoffs between these heterogeneous requirements. To address these challenges, we propose the design of a 3D Network-on-Chip (NoC) for heterogeneous manycore platforms that considers the appropriate design objectives for a 3D heterogeneous system and explores various tradeoffs using an efficient machine learning (ML)-based multi-objective optimization (MOO) technique. The proposed design space exploration considers the various requirements of its heterogeneous components and generates a set of 3D NoC architectures that efficiently trades off these design objectives. Our findings show that by jointly considering these requirements (latency, throughput, temperature, and energy), we can achieve 9.6 percent better Energy-Delay Product on average at nearly iso-temperature conditions when compared to a thermally-optimized design for 3D heterogeneous NoCs. More importantly, our results suggest that our 3D NoCs optimized for a few applications can be generalized for unknown applications as well. Our results show that these generalized 3D NoCs only incur a 1.8 percent (36-tile system) and 1.1 percent (64-tile system) average performance loss compared to application-specific NoCs. Biresh Kumar Joardar, Ryan Gary Kim, Janardhan Rao Doppa, Partha Pratim Pande, Diana Marculescu, Radu Marculescu |
IEEE Trans. Computers | 6 |
| 2019 | Memory- and Communication-Aware Model Compression for Distributed Deep Learning Inference on IoTabstractModel compression has emerged as an important area of research for deploying deep learning models on Internet-of-Things (IoT). However, for extremely memory-constrained scenarios, even the compressed models cannot fit within the memory of a single device and, as a result, must be distributed across multiple devices. This leads to a distributed inference paradigm in which memory and communication costs represent a major bottleneck. Yet, existing model compression techniques are not communication-aware. Therefore, we propose Network of Neural Networks (NoNN), a new distributed IoT learning paradigm that compresses a large pretrained ‘teacher’ deep network into several disjoint and highly-compressed ‘student’ modules, without loss of accuracy. Moreover, we propose a network science-based knowledge partitioning algorithm for the teacher model, and then train individual students on the resulting disjoint partitions. Extensive experimentation on five image classification datasets, for user-defined memory/performance budgets, show that NoNN achieves higher accuracy than several baselines and similar accuracy as the teacher model, while using minimal communication among students. Finally, as a case study, we deploy the proposed model for CIFAR-10 dataset on edge devices and demonstrate significant improvements in memory footprint (up to 24×), performance (up to 12×), and energy per node (up to 14×) compared to the large teacher model. We further show that for distributed inference on multiple edge devices, our proposed NoNN model results in up to 33× reduction in total latency w.r.t. a state-of-the-art model compression baseline. Kartikeya Bhardwaj, Chingyi Lin, Anderson Luiz Sartor, Radu Marculescu |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2018 | Hybrid on-chip communication architectures for heterogeneous manycore systemsabstractThe widespread adoption of big data has led to the search for highperformance and low-power computational platforms. Emerging heterogeneous manycore processing platforms consisting of CPU and GPU cores along with various types of accelerators offer power and area-efficient trade-offs for running these applications. However, heterogeneous manycore architectures need to satisfy the communication and memory requirements of the diverse computing elements that conventional Network-on-Chip (NoC) architectures are unable to handle effectively. Further, with increasing system sizes and level of heterogeneity, it becomes difficult to quickly explore the large design space and establish the appropriate design trade-offs. To address these challenges, machine learning-inspired heterogeneous manycore system design is a promising research direction to pursue. In this paper, we highlight various salient features of heterogeneous manycore architectures enabled by emerging interconnect technologies and machine learning techniques. Biresh Kumar Joardar, Janardhan Rao Doppa, Partha Pratim Pande, Diana Marculescu, Radu Marculescu |
ICCAD | 5 |
| 2018 | Dimensionality Reduction via Community Detection in Small Sample Datasets
Kartikeya Bhardwaj, Radu Marculescu |
PAKDD (3) | 2 |
| 2018 | On-Chip Communication Network for Efficient Training of Deep Convolutional Networks on Heterogeneous Manycore SystemsabstractConvolutional Neural Networks (CNNs) have shown a great deal of success in diverse application domains including computer vision, speech recognition, and natural language processing. However, as the size of datasets and the depth of neural network architectures continue to grow, it is imperative to design high-performance and energy-efficient computing hardware for training CNNs. In this paper, we consider the problem of designing specialized CPU-GPU based heterogeneous manycore systems for energy-efficient training of CNNs. It has already been shown that the typical on-chip communication infrastructures employed in conventional CPU-GPU based heterogeneous manycore platforms are unable to handle both CPU and GPU communication requirements efficiently. To address this issue, we first analyze the on-chip traffic patterns that arise from the computational processes associated with training two deep CNN architectures, namely, LeNet and CDBNet, to perform image classification. By leveraging this knowledge, we design a hybrid Network-on-Chip (NoC) architecture, which consists of both wireline and wireless links, to improve the performance of CPU-GPU based heterogeneous manycore platforms running the above-mentioned CNN training workloads. The proposed NoC achieves 1.8× reduction in network latency and improves the network throughput by a factor of 2.2 for training CNNs, when compared to a highly-optimized wireline mesh NoC. For the considered CNN workloads, these network-level improvements translate into 25 percent savings in full-system energy-delay-product (EDP). This demonstrates that the proposed hybrid NoC for heterogeneous manycore architectures is capable of significantly accelerating training of CNNs while remaining energy-efficient. Wonje Choi 0001, Karthi Duraisamy, Ryan Gary Kim, Janardhan Rao Doppa, Partha Pratim Pande, Diana Marculescu, Radu Marculescu |
IEEE Trans. Computers | 7 |
| 2017 | Discovering Hidden Knowledge in Carbon Emissions Data: A Multilayer Network Approach
Kartikeya Bhardwaj, HingOn Miu, Radu Marculescu |
DS | 3 |
| 2017 | 3D NoC-Enabled Heterogeneous Manycore Architectures for Accelerating CNN Training: Performance and Thermal Trade-offsabstractAs deep learning technology is increasingly employed in diverse applications domains, the demand for computational power to enable these algorithms also increases. In this respect, high-performance three-dimensional (3D) heterogeneous manycore systems present a promising direction. However, deep learning on these systems pose several design challenges. First, the network-on-chip (NoC) must handle the traffic requirements of both CPU and GPU communications. Second, 3D system designs must address thermal issues resulting from high-power density. In this work, we propose a design methodology for a heterogeneous 3D NoC architecture that not only satisfies the traffic requirements of both CPUs and GPUs, but also reduces thermal hotspots. To this end, we target the training of two widely employed convolutional neural networks (CNN), namely, LeNet and CIFAR. By using our joint performance-thermal optimization methodology to create a 3D NoC for training CNNs, we reduce the maximum temperature by 22% while incurring only 5% full-system energy-delay-product degradation over a solely performance optimized 3D NoC. This demonstrates that, our design methodology achieves considerable temperature reduction with negligible loss in performance. Biresh Kumar Joardar, Wonje Choi 0001, Ryan Gary Kim, Janardhan Rao Doppa, Partha Pratim Pande, Diana Marculescu, Radu Marculescu |
NOCS | 7 |
| 2017 | MPLasso: Inferring microbial association networks using prior microbial knowledgeabstractDue to the recent advances in high-throughput sequencing technologies, it becomes possible to directly analyze microbial communities in human body and environment. To understand how microbial communities adapt, develop, and interact with the human body and the surrounding environment, one of the fundamental challenges is to infer the interactions among different microbes. However, due to the compositional and high-dimensional nature of microbial data, statistical inference cannot offer reliable results. Consequently, new approaches that can accurately and robustly estimate the associations (putative interactions) among microbes are needed to analyze such compositional and high-dimensional data. We propose a novel framework called Microbial Prior Lasso (MPLasso) which integrates graph learning algorithm with microbial co-occurrences and associations obtained from scientific literature by using automated text mining. We show that MPLasso outperforms existing models in terms of accuracy, microbial network recovery rate, and reproducibility. Furthermore, the association networks we obtain from the Human Microbiome Project datasets show credible results when compared against laboratory data. Chieh Lo, Radu Marculescu |
PLoS Comput. Biol. | 2 |
| 2017 | Non-Stationary Bayesian Learning for Global SustainabilityabstractAn increasingly warming planet calls for widespread use of sustainable energy sources like solar energy. To meet the rising energy demand, the focus of state-of-the-art solar energy models on local predictions is no longer sufficient as it only leads to local optimization of solar resources. Hence, a new class of models is needed that can provide a global response towards sustainability. In this paper, therefore, we propose a new approach that models cloud movement as a multilayer network and then performs parameter learning on it to generate short-term predictions of cloud fraction/solar irradiance simultaneously at a large number of locations. These learned parameters capture the spatio-temporal interdependencies of solar energy which can allow power-grid operators and policy-makers at different locations to know who impacts the solar energy of whom. Our results indicate a Root Mean Square Error (RMSE) of 8-18% in one-hour cloud fraction prediction. Finally, using our network approach, we show that the cloud movement likely follows a power law distribution, an important domain knowledge discovery that may be useful for future models. A major consequence of our approach is that it can enable power-grid operators/policy-makers to see beyond the local boundaries of their respective geographical locations. Kartikeya Bhardwaj, Radu Marculescu |
IEEE Trans. Sustain. Comput. | 2 |
| 2017 | Imitation Learning for Dynamic VFI Control in Large-Scale Manycore SystemsabstractManycore chips are widely employed in high-performance computing and large-scale data analysis. However, the design of high-performance manycore chips is dominated by power and thermal constraints. In this respect, voltage-frequency island (VFI) is a promising design paradigm to create scalable energy-efficient platforms. By dynamically tailoring the voltage and frequency of each island, we can further improve the energy savings within given performance constraints. Inspired by the recent success of imitation learning (IL) in many application domains and its significant advantages over reinforcement learning (RL), we propose the first architecture-independent IL-based methodology for dynamic VFI (DVFI) control in manycore systems. Due to its popularity in the EDA community, we consider an RL-based DVFI control methodology as a strong baseline. Our experimental results demonstrate that IL is able to obtain higher quality policies than RL (on average, 5% less energy with the same level of performance) with significantly less computation time and hardware area overheads (3.1X and 8.8X, respectively). Ryan Gary Kim, Wonje Choi 0001, Janardhan Rao Doppa, Partha Pratim Pande, Diana Marculescu, Radu Marculescu |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2016 | Hybrid network-on-chip architectures for accelerating deep learning kernels on heterogeneous manycore platformsabstractIn recent years, designing specialized manycore heterogeneous architectures for deep learning kernels has become an area of great interest. However, the typical on-chip communication infrastructures employed on conventional manycore platforms are unable to handle both CPU and GPU communication requirements efficiently. Hence, in this paper, our aim is to enhance the performance of heterogeneous manycore architectures through the design of a hybrid NoC consisting of both wireline and wireless links. To this end, we specifically target the resource-intensive backpropagation algorithm commonly used as the training method in deep learning. For backpropagation, the proposed hybrid NoC achieves 1.9X reduction in network latency and improves the network throughput by a factor of 2 with respect to a highly optimized mesh NoC. These network level improvements translate into 25% savings in full system energy-delay-product (EDP). This demonstrates the capability of the proposed hybrid and heterogeneous manycore architecture in accelerating deep learning kernels in an energy-efficient manner. Wonje Choi 0001, Karthi Duraisamy, Ryan Gary Kim, Janardhan Rao Doppa, Partha Pratim Pande, Radu Marculescu, Diana Marculescu |
CASES | 6 |
| 2016 | nOS: A nano-sized distributed operating system for many-core embedded systemsabstractWe introduce nOS, a “nano-sized” fully distributed operating system aimed at large-scale, many-core embedded systems. nOS enables dynamic runtime optimisation of energy and execution time through lightweight and scalable distributed protocols. nOS implements new dynamic resource optimisation algorithms, and provides an intuitive and easy-to-use programmer API that supports runtime task energy optimisation through dynamic frequency scaling, transparent task communication tracking, and automatic task mapping. Critically, nOS has a completely distributed implementation, providing excellent scalability. Contrary to other approaches, the dynamic runtime optimisations require no a priori knowledge of workload or communication patterns. By generating runtime measurements of thread performance, core load, and process communication, we show that nOS can deliver improvements that would not be possible with only static analysis. Using a many-core system called Swallow, we show a <;3kB fullstack implementation of nOS together with application, OS and hardware. Using two applications with different communication patterns, we illustrate the power and flexibility of our approach, as well as various tradeoffs in energy and performance from making better mapping choices than would be available offline. Simon J. Hollis, Edward Ma, Radu Marculescu |
ICCD | 3 |
| 2016 | Wireless NoC for VFI-Enabled Multicore Chip Design: Performance Evaluation and Design Trade-OffsabstractMultiple Voltage Frequency Island (VFI)-based designs can reduce the energy dissipation in multicore chips. Indeed, by tailoring the voltages and frequencies of each VFI domain, we can achieve significant energy savings subject to specific performance constraints. The achievable performance of VFI-based multicore platforms depends on the overall communication backbone, which relies predominantly on Networks-on-Chip (NoCs). Traditionally mesh-based NoCs have been used in VFI-based systems. However, the mesh-based NoCs have large latency and energy overheads due to their inherently long multihop paths. Emerging paradigms such as the millimeter (mm)-wave small-world wireless Networks-on-Chip (mSWNoCs) have lately been observed to help reduce the impact of the communication backbone on the performance of the multicore chips. In this work, we demonstrate that not only do mSWNoC-enabled VFI designs mitigate some of the full-system performance degradation inherent in VFI-partitioned multicore designs, but they also help in eliminating it entirely for certain applications. We also demonstrate that the VFI-partitioned designs used in conjunction with a novel NoC architecture like mSWNoC can achieve significant energy savings while minimizing the impact on the performance for each application under consideration. Ryan Gary Kim, Wonje Choi 0001, Guangshuo Liu, Ehsan Mohandesi, Partha Pratim Pande, Diana Marculescu, Radu Marculescu |
IEEE Trans. Computers | 7 |
| 2016 | A Support Vector Regression (SVR)-Based Latency Model for Network-on-Chip (NoC) ArchitecturesabstractIn this paper, we propose SVR-NoC, a network-on-chip (NoC) latency model using support vector regression (SVR). More specifically, based on the application communication information and the NoC routing algorithm, the channel and source queue waiting times are first estimated using an analytical queuing model with two equivalent queues. To improve the prediction accuracy, the queuing theory-based delay estimations are included as features in the learning process. We then propose a learning framework that relies on SVR to collect training data and predict the traffic flow latency. The proposed learning methods can be used to analyze various traffic scenarios for the target NoC platform. Experimental results on both synthetic and real-application traffic demonstrate on average less than 12% prediction error in network saturation load, as well as more than 100× speedup compared to cycle-accurate simulations can be achieved. Zhiliang Qian, Da-Cheng Juan, Paul Bogdan, Chi-Ying Tsui, Diana Marculescu, Radu Marculescu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2016 | Performance Evaluation of NoC-Based Multicore Systems: From Traffic Analysis to NoC Latency ModelingabstractIn this survey, we review several approaches for predicting performance of Network-on-Chip (NoC)-based multicore systems, starting from the traffic models to the complex NoC models for latency evaluation. We first review typical traffic models to represent the application workloads in NoC. Specifically, we review Markovian and non-Markovian (e.g., self-similar or long-range memory-dependent) traffic models and discuss their applications on multicore platform design. Then, we review the analytical techniques to predict NoC performance under given input traffic. We investigate analytical models for average as well as maximum delay evaluation. We also review the developments and design challenges of NoC simulators. One interesting research direction in NoC performance evaluation consists of combining simulation and analytical models in order to exploit their advantages together. Toward this end, we discuss several newly proposed approaches that use hardware-based or learning-based techniques. Finally, we summarize several open problems and our perspective to address these challenges. Zhiliang Qian, Paul Bogdan, Chi-Ying Tsui, Radu Marculescu |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2016 | Wireless NoC and Dynamic VFI Codesign: Energy Efficiency Without Performance PenaltyabstractMultiple voltage frequency island (VFI)-based designs can reduce the energy dissipation in multicore platforms by taking advantage of the varying nature of the application workloads. Indeed, the voltage/frequency (V/F) levels of the VFIs can be dynamically tailored by considering the workload-driven variations in the application. Traditionally, mesh-based networks-on-chip (NoCs) have been used in VFI-based systems; however, they have large latency and energy overheads due to the inherently long multihop paths. Consequently, in this paper, we explore the emerging paradigm of wireless NoC (WiNoC) and demonstrate that by incorporating WiNoC, VFI, and dynamic V/F tuning in a synergistic manner, we can design energy-efficient multicore platforms without introducing noticeable performance penalty. Our experimental results show that for the benchmarks considered, the proposed approach can achieve between 5.7% and 46.6% energy-delay product (EDP) savings over the state-of-the-art system and 26.8% and 60.5% EDP savings over a standard baseline non-VFI mesh-based system. This opens up a new of class of codesign approaches that can make WiNoCs the communication technology of choice for future multicore platforms. Ryan Gary Kim, Wonje Choi 0001, Partha Pratim Pande, Diana Marculescu, Radu Marculescu |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2015 | Energy efficient MapReduce with VFI-enabled multicore platformsabstractIn an era when power constraints and data movement are proving to be significant barriers for high-end computing, multicore architectures offer a low-power and highly scalable platform suitable for both data- and compute-intensive applications. MapReduce is a popular framework to facilitate the management and development of big-data workloads. In this work, we demonstrate that by using a wireless NoC-enabled Voltage Frequency Island (VFI)-based multicore platform it is possible to enhance the energy efficiency of MapReduce implementations without paying significant execution time penalties. Our experimental results show that for the benchmarks considered, the designed VFI system can achieve an average of 33.7% energy-delay product (EDP) savings over the standard baseline non-VFI mesh-based system while paying a maximum of 3.22% execution time penalty. Karthi Duraisamy, Ryan Gary Kim, Wonje Choi 0001, Guangshuo Liu, Partha Pratim Pande, Radu Marculescu, Diana Marculescu |
DAC | 6 |
| 2015 | Statistical Learning in Chip (SLIC)abstractDespite best efforts, integrated systems are “born” (manufactured) with a unique `personality' that stems from our inability to precisely fabricate their underlying circuits, and create software a priori for controlling the resulting uncertainty. It is possible to use sophisticated test methods to identify the best-performing systems but this would result in unacceptable yields and correspondingly high costs. The system personality is further shaped by its environment (e.g., temperature, noise and supply voltage) and usage (i.e., the frequency and type of applications executed), and since both can fluctuate over time, so can the system's personality. Systems also “grow old” and degrade due to various wear-out mechanisms (e.g., negative-bias temperature instability), and unexpectedly due to various early-life failure sources. These “nature and nurture” influences make it extremely difficult to design a system that will operate optimally for all possible personalities. To address this challenge, we propose to develop statistical learning in-chip (SLIC). SLIC is a holistic approach to integrated system design based on continuously learning key personality traits on-line, for self-evolving a system to a state that optimizes performance hierarchically across the circuit, platform, and application levels. SLIC will not only optimize integrated-system performance but also reduce costs through yield enhancement since systems that would have before been deemed to have weak personalities (unreliable, faulty, etc.) can now be recovered through the use of SLIC. R. D. (Shawn) Blanton, Xin Li 0001, Ken Mai, Diana Marculescu, Radu Marculescu, Jeyanandh Paramesh, Jeff G. Schneider, Donald E. Thomas |
ICCAD | 5 |
| 2015 | The (Low) Power of Less Wiring: Enabling Energy Efficiency in Many-Core Platforms Through Wireless NoCabstractDuring the last decade, we have witnessed a major transition from computation- to communication-centric design of integrated circuits and systems. In particular, the network-on-chip (NoC) approach has emerged as the major design paradigm for multicore systems-on-chip (SoC). The major challenges in traditional wire-based NoCs are the high latency and power consumption of the multi-hop links. By inserting single-hop long-range wireless links in place of multi-hop wired links, the overall system performance can be significantly improved. We should adopt novel architectures inspired by the on-chip wireless links to design high-performance multi-core chips. In this regard, the small-world network-inspired wireless NoC (WiNoC) has emerged as an enabling interconnection infrastructure to design high-bandwidth and energy-efficient multicore chips. In this paper we present the various challenges and possible solutions for designing energy-efficient massive multicore chips enabled by the WiNoC paradigm. Partha Pratim Pande, Ryan Gary Kim, Wonje Choi 0001, Diana Marculescu, Radu Marculescu |
ICCAD | 6 |
| 2014 | A comprehensive and accurate latency model for Network-on-Chip performance analysisabstractIn this work, we propose a new, accurate, and comprehensive analytical model for Network-on-Chip (NoC) performance analysis. Given the application communication graph, the NoC architecture, and the routing algorithm, the proposed framework analyzes the links dependency and then determines the ordering of queuing analysis for performance modeling. The channel waiting times in the links are estimated using a generalized G/G/1/K queuing model, which can tackle bursty traffic and dependent arrival times with general service time distributions. The proposed model is general and can be used to analyze various traffic scenarios for NoC platforms with arbitrary buffer and packet lengths. Experimental results on both synthetic and real applications demonstrate the accuracy and scalability of the newly proposed model. Zhiliang Qian, Da-Cheng Juan, Paul Bogdan, Chi-Ying Tsui, Diana Marculescu, Radu Marculescu |
ASP-DAC | 6 |
| 2014 | Energy-efficient VFI-partitioned multicore design using wireless NoC architecturesabstractIn recent years, multiple Voltage Frequency Island (VFI)-based designs have increasingly made their way into both commercial and research multicore platforms. On the other hand, the wireless Network-on-Chip (WiNoC) architecture has emerged as an energy-efficient and high bandwidth communication backbone for massively integrated multicore platforms. It becomes therefore possible to exploit the small-world effects induced by the wireless links of a WiNoC to achieve efficient inter-VFI data exchanges. In this work, we demonstrate that WiNoCs can provide better latency and energy profiles compared to traditional mesh-like architecture for VFI-partitioned multicore designs. The performance gains and energy efficiency are achieved due to the low-power wireless shortcuts in conjunction with the small-world architecture. Indeed, our experimental results show energy improvements as large as 40% for multithreaded application benchmarks. Ryan Gary Kim, Guangshuo Liu, Paul Wettin, Radu Marculescu, Diana Marculescu, Partha Pratim Pande |
CASES | 4 |
| 2014 | Low-latency wireless 3D NoCs via randomized shortcut chipsabstractIn this paper, we demonstrate that we can reduce the communication latency significantly by inserting a fraction of randomness into a wireless 3D NoC (where CMOS wireless links are used for vertical inter-chip communication) when considering the physical constraints of the 3D design space. Towards this end, we consider two cases, namely 1) replacing existing horizontal 2D links in a wireless 3D NoC with randomized shortcut NoC links and 2) enabling full connectivity by adding a randomized NoC layer to a wireless 3D platform with partial or no horizontal connectivity. Consequently, the packet routing is optimized by exploiting both the existing and the newly added random NoC. At the same time, by adding randomly wired shortcut NoCs to a wireless 3D platform, a good balance can be established between the modularity of the design and the minimum randomness needed to achieve low latency, and experimental results show that by adding a random NoC chip to wireless 3D CMPs without built-in horizontal connectivity, the communication latency can be reduced by as much as 26.2% when compared to adding a 2D mesh NoC. Also, the application execution time and average flit transfer energy can be improved accordingly. Hiroki Matsutani, Michihiro Koibuchi, Ikki Fujiwara, Takahiro Kagami, Yasuhiro Take, Tadahiro Kuroda, Paul Bogdan, Radu Marculescu, Hideharu Amano |
DATE | 8 |
| 2014 | Introduction to the special session on "Interconnect enhances architecture: Evolution of wireless NoC from planar to 3D"abstractContinuing progress and unprecedented integration levels in current silicon technologies make possible complete end-user systems consisting of an extremely high number of cores integrated on a single chip for embedded or high-performance computing. However, without developing new paradigms for energy- and thermally-efficient design, meeting the computing, storage, and communication demands of the emerging applications is highly unlikely. Moreover, in order to sustain the predicted growth of number of embedded cores on a single die, it is extremely important to have a scalable, low power, and high bandwidth on-chip communication infrastructure. Towards this end, wireless Network-on-Chip (WiNoC) represents an emerging paradigm to design a low power yet high bandwidth interconnect infrastructure for multicore chips. Radu Marculescu, Partha Pratim Pande, Deuk Hyoun Heo, Hiroki Matsutani |
NOCS | 1 |
| 2014 | An efficient Network-on-Chip (NoC) based multicore platform for hierarchical parallel genetic algorithmsabstractIn this work, we propose a new Network-on-Chip (NoC) architecture for implementing the hierarchical parallel genetic algorithm (HPGA) on a multi-core System-on-Chip (SoC) platform. We first derive the speedup metric of an NoC architecture which directly maps the HPGA onto NoC in order to identify the main sources of performance bottlenecks. Specifically, it is observed that the speedup is mostly affected by the fixed bandwidth that a master processor can use and the low utilization of slave processor cores. Motivated by the theoretical analysis, we propose a new architecture with two multiplexing schemes, namely dynamic injection bandwidth multiplexing (DIBM) and time-division based island multiplexing (TDIM), to improve the speedup and reduce the hardware requirements. Moreover, a task-aware adaptive routing algorithm is designed for the proposed architecture, which can take advantage of the proposed multiplexing schemes to further reduce the hardware overhead. We demonstrate the benefits of our approach using the problem of protein folding prediction, which is a process of importance in biology. Our experimental results show that the proposed NoC architecture achieves up to 240X speedup compared to a single island design. The hardware cost is also reduced by 50% compared to a direct NoC-based HPGA implementation. Yuankun Xue, Zhiliang Qian, Guopeng Wei, Paul Bogdan, Chi-Ying Tsui, Radu Marculescu |
NOCS | 6 |
| 2014 | Miniature Devices in the Wild: Modeling Molecular Communication in Complex Extracellular SpacesabstractMiniature devices voyaging inside the human body for diagnostic and drug delivery purposes is no longer a wild dream. At the very heart of such an endeavor lies the capability of miniature devices like synthetic cells and microrobots to achieve complex tasks collectively by exchanging information molecules. Towards this end, we model the spatiotemporal dynamics of the molecular transport process in complex extracellular spaces (ECSs) such that the signaling delay can be accurately predicted. More precisely, we use parameters like ECS volume fraction, tortuosity, and cross-section area of diffusion paths to capture the physicochemical features of the ECS. Based on these parameters, we propose a new algorithm to calculate the directional diffusion coefficient, which is then used in an effective diffusion equation to describe the molecular transport process across the region of interest. Our modeling results show good agreement with detailed 3D simulations in complex ECSs, while the classical diffusion and previous approaches fail to capture the heterogeneity and directionality of the transport process. Consequently, the proposed approach represents a major step towards characterizing the interaction of cooperative miniature devices that can achieve complex tasks via diffusion-based molecular communication. Guopeng Wei, Radu Marculescu |
IEEE J. Sel. Areas Commun. | 2 |
| 2014 | Exploiting Emergence in On-Chip InterconnectsabstractTo solve the grand challenges in contemporary chip design, such as process-to-core mapping, energy reduction, and maintenance of programmer/hardware abstraction, we advocate for self-optimizing (emergent) networks-on-chip (NoC). In these networks, topology and information flow adapt dynamically to maximize the network throughput or minimize the network latency via distributed application of microrules. In this paper, we introduce the concept of emergent small-world NoCs and discuss novel design decisions, e.g., Skip-links, that improve performance and reduce energy consumption of multicore systems. More precisely, we demonstrate that our proposed solution is able to adapt to a wide range of traffic patterns and provide reductions in data hop count of up to 20 percent while maintaining energy and area costs. We show how emergent networks can be useful for on-chip processor-to-processor communications, and also demonstrate how SoC and off-chip I/O traffic may be optimized for latency and critical load. Simon J. Hollis, Chris Jackson, Paul Bogdan, Radu Marculescu |
IEEE Trans. Computers | 4 |
| 2013 | A case for wireless 3D NoCs for CMPsabstractInductive-coupling is yet another 3D integration technique that can be used to stack more than three known-good-dies in a SiP without wire connections. We present a topology-agnostic 3D CMP architecture using inductive-coupling that offers great flexibility in customizing the number of processor chips, SRAM chips, and DRAM chips in a SiP after chips have been fabricated. In this paper, first, we propose a routing protocol that exchanges the network information between all chips in a given SiP to establish efficient deadlock-free routing paths. Second, we propose its optimization technique that analyzes the application traffic patterns and selects different spanning tree roots so as to minimize the average hop counts and improve the application performance. Hiroki Matsutani, Paul Bogdan, Radu Marculescu, Yasuhiro Take, Daisuke Sasaki, Hao Zhang 0020, Michihiro Koibuchi, Tadahiro Kuroda, Hideharu Amano |
ASP-DAC | 3 |
| 2013 | Identifying dynamics and collective behaviors in microblogging tracesabstractMicroblogging disseminates realtime information through dynamic user interactions. While it is intuitive that such interactions may generate patterns, it is difficult to identify and characterize them in satisfactory detail. In this paper, we propose using a combination of dynamic graphs and time-series to study the dynamics and collective behaviors in microblogging. To enable automatic pattern identification, a distance metric is developed to incorporate the heterogeneous aspects of the dynamical interactions. We demonstrate the effectiveness of the proposed approach using a month long Twitter dataset and show that the new representation and distance metric are both essential for discovering the patterns of collective microblogging, such as propagation of breaking news, advertisement, social movement, and interest group formation. Huan-Kai Peng, Radu Marculescu |
ASONAM | 2 |
| 2013 | Closed-loop control for power and thermal management in multi-core processors: formal methods and industrial practiceabstractThe need to use feedback to come up with context-dependent and workload-aware strategies for runtime power and thermal management (PTM) in high-end and mobile processors has been advocated since the early 2000. Two seminal papers that appeared in 2002 [1], [2] defined a framework for the use of feedback mechanisms for power and temperature control. In [1], the focus was on power management with the goal being to extend battery life on the AMD Mobile Athlon. This was one of the earliest papers to use DVFS settings as actuators to guarantee a given energy level in the battery at the end of a given time interval. The controller was implemented using a combination of OS files and Linux kernel modules. Almost simultaneously, [2] posed the dynamic thermal management task as a formal control-theoretic problem requiring the thermal modeling of the processor and the use of the established control structures of classical feedback theory. Some of the defining features of [2] include the development of layout-based thermal RC models for the processor; the use of an architecturally-driven control mechanism, namely, the instruction fetching rate; and the use of the SPEC2000 benchmarks to illustrate temperature control action under various workloads. The controller used in [2] is a Proportional-Integral-Differential (PID) structure whose input is the deviation of the sensed temperature from the target temperature and whose output is the toggle rate of the instruction fetching mechanism. Ibrahim M. Elfadel, Radu Marculescu, David Atienza 0001 |
DATE | 2 |
| 2013 | SVR-NoC: a performance analysis tool for network-on-chips using learning-based support vector regression modelabstractIn this work, we propose SVR-NoC, a learning-based support vector regression (SVR) model for evaluating Network-on-Chip (NoC) latency performance. Different from the state-of-the-art NoC analytical model, which uses classical queuing theory to directly compute the average channel waiting time, the proposed SVR-NoC model performs NoC latency analysis based on learning the typical training data. More specifically, we develop a systematic machine-learning framework that uses the kernel-based support vector regression method to predict the channel average waiting time and the traffic flow latency. Experimental results show that SVR-NoC can predict the average packet latency accurately while achieving about 120X speed-up over simulation-based evaluation methods. Zhiliang Qian, Da-Cheng Juan, Paul Bogdan, Chi-Ying Tsui, Diana Marculescu, Radu Marculescu |
DATE | 6 |
| 2013 | Performance evaluation of multicore systems: from traffic analysis to latency predictions (embedded tutorial)abstractAs technology scaling down allows multiple processing components to be integrated on a single chip, the modern computing systems led to the advent of Multiprocessor System-on-Chip (MPSoC) and Chip Multiprocessor (CMP) design. Network-on-Chips (NoCs) have been proposed as a promising solution to tackle the complex on-chip communication problems on these multicore platforms. In order to optimize the NoC-based multicore system design, it is essential to evaluate the NoC performance with respect to numerous configurations in a large design space. Taking the traffic characteristics into account and using an appropriate latency model become crucially important to provide an accurate and fast evaluation. In this tutorial, we survey the current progresses in these aspects. We first review the NoC workload modeling and traffic analysis techniques. Then, we discuss the mathematical formalisms of evaluating the performance under a given traffic model, for both the average and worst-case latency predictions. Finally, the advantages of combining the analytical and simulation-based techniques are discussed and new attempts for bridging these two approaches are reviewed. Zhiliang Qian, Paul Bogdan, Chi-Ying Tsui, Radu Marculescu |
ICCAD | 4 |
| 2013 | Efficient Modeling and Simulation of Bacteria-Based Nanonetworks with BNSimabstractBacteria-based networks are formed using native or engineered bacteria that communicate at nano-scale. This definition includes the micro-scale molecular transportation system which uses chemotactic bacteria for targeted cargo delivery, as well as genetic circuits for intercellular interactions like quorum sensing or light communication. To characterize the dynamics of bacterial networks accurately, we introduce BNSim, an open-source, parallel, stochastic, and multiscale modeling platform which integrates various simulation algorithms, together with genetic circuits and chemotactic pathway models in a complex 3D environment. Moreover, we show how this platform can be used to model synthetic bacterial consortia which implement a XOR function and aggregate nearby bacteria using light communication. Consequently, the results demonstrate how BNSim can predict various properties of realistic bacterial networks and provide guidance for their actual wet-lab implementations. Guopeng Wei, Paul Bogdan, Radu Marculescu |
IEEE J. Sel. Areas Commun. | 3 |
| 2013 | Bumpy Rides: Modeling the Dynamics of Chemotactic Interacting BacteriaabstractRecent advances in synthetic biology have brought the fantasy of having synthetic multicellular systems working at nanoscale closer to reality. Indeed, such systems consisting of networked biological nanomachines have a great potential for many novel applications like environmental monitoring and healthcare. In this paper, we consider dense networks of interacting bacteria capable of monitoring and treating diseases that affect microscopic regions in the human body. Towards this end, we propose a modeling approach for capturing the dynamics of such dense populations of interacting bacteria and estimate their performance for diagnostic and drug delivery purposes. Consequently, our approach can be used to identify various design trade-offs for dense networks of bacteria and design predictable and reliable synthetic multicellular systems. Guopeng Wei, Paul Bogdan, Radu Marculescu |
IEEE J. Sel. Areas Commun. | 3 |
| 2013 | Pacemaker control of heart rate variability: A cyber physical system perspectiveabstractCardiac diseases, like those related to abnormal heart rate activity, have an enormous economic and psychological impact worldwide. The approaches used to control the behavior of modern pacemakers ignore the fractal nature of heart rate activity. The purpose of this article is to present a Cyber Physical System approach to pacemaker design that exploits precisely the fractal properties of heart rate activity in order to design the pacemaker controller. Towards this end, we solve a finite horizon optimal control problem based on the heartbeat time series and show that this control problem can be converted into a system of linear equations. We also compare and contrast the performance of the fractal optimal control problem under six different cost functions. Finally, to get an idea of hardware complexity, we implement the fractal optimal controller on a Virtex4 FPGA and report some preliminary results in terms of area overhead. Paul Bogdan, Radu Marculescu |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2013 | Dynamic power management for multidomain system-on-chip platforms: An optimal control approachabstractReducing energy consumption in multiprocessor systems-on-chip (MPSoCs) where communication happens via the network-on-chip (NoC) approach calls for multiple voltage/frequency island (VFI)-based designs. In turn, such multi-VFI architectures need efficient, robust, and accurate runtime control mechanisms that can exploit the workload characteristics in order to save power. Despite being tractable, the linear control models for power management cannot capture some important workload characteristics (e.g., fractality, nonstationarity) observed in heterogeneous NoCs; if ignored, such characteristics lead to inefficient communication and resources allocation, as well as high power dissipation in MPSoCs. To mitigate such limitations, we propose a new paradigm shift from power optimization based on linear models to control approaches based on fractal-state equations. As such, our approach is the first to propose a controller for fractal workloads with precise constraints on state and control variables and specific time bounds. Our results show that significant power savings can be achieved at runtime while running a variety of benchmark applications. Paul Bogdan, Radu Marculescu |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2012 | Modeling populations of micro-robots for biological applicationsabstractIn order to detect and cure diseases affecting microscopic regions of the human body, swarms of micro-robots can harness the motility of bacteria and swim towards inaccessible regions of the body to detect abnormal behavior and/or deliver various drugs. In this paper, we propose a modeling approach for the dynamics of dense networks of swarms and estimate their performance via stochastic metrics. Paul Bogdan, Guopeng Wei, Radu Marculescu |
ICC | 3 |
| 2012 | An Optimal Control Approach to Power Management for Multi-Voltage and Frequency Islands Multiprocessor Platforms under Highly Variable WorkloadsabstractReducing energy consumption in multi-processor systems-on-chip (MPSoCs) where communication happens via the network-on-chip (NoC) approach calls for multiple voltage/frequency island (VFI)-based designs. In turn, such multi-VFI architectures need efficient, robust, and accurate run-time control mechanisms that can exploit the workload characteristics in order to save power. Despite being tractable, the linear control models for power management cannot capture some important workload characteristics (e.g., fractality, non-stationarity) observed in heterogeneous NoCs, if ignored, such characteristics lead to inefficient communication and resources allocation, as well as high power dissipation in MPSoCs. To mitigate such limitations, we propose a new paradigm shift from power optimization based on linear models to control approaches based on fractal-state equations. As such, our approach is the first to propose a controller for fractal workloads with precise constraints on state and control variables and specific time bounds. Our results show that significant power savings (about 70%) can be achieved at run-time while running a variety of benchmark applications. Paul Bogdan, Radu Marculescu, Rafael Tornero |
NOCS | 2 |
| 2012 | Dynamic power management for multicores: Case study using the intel SCCabstractIn this paper, a dynamic voltage and frequency scaling (DVFS) based power management approach is implemented using the Intel SCC platform. To validate our framework, we show various power and performance trade-offs using the SCC platform, while running a computationally intensive parallel image-processing application under different power management conditions. David Radu, Paul Bogdan, Radu Marculescu |
VLSI-SoC | 3 |
| 2012 | Technology-driven limits on runtime power management algorithms for multiprocessor systems-on-chipabstractRuntime power management is a critical technique for reducing the energy footprint of digital electronic devices and enabling sustainable computing, since it allows electronic devices to dynamically adapt their power and energy consumption to meet performance requirements. In this article, we consider the case of MultiProcessor Systems-on-Chip (MPSoC) implemented using multiple Voltage and Frequency Islands (VFIs) relying on fine-grained Dynamic Voltage and Frequency Scaling (DVFS) to reduce the system power dissipation. In particular, we present a framework to theoretically analyze the impact of three important technology-driven constraints; (i) reliability-driven upper limits on the maximum supply voltage; (ii) inductive noise-driven constraints on the maximum rate of change of voltage/frequency; and (iii) the impact of manufacturing process variations on the performance of DVFS control for multiple VFI MPSoCs. The proposed analysis is general, in the sense that it is not bound to a specific DVFS control algorithm, but instead focuses on theoretically bounding the performance that any DVFS controller can possibly achieve. Our experimental results on real and synthetic benchmarks show that in the presence of reliability- and temperature-driven constraints on the maximum frequency and maximum frequency increment, any DVFS control algorithm will lose up to 87% performance in terms of the number of steps required to reach a reference steady state. In addition, increasing process variations can lead to up to 60% of fabricated chips being unable to meet the specified DVFS control specifications, irrespective of the DVFS algorithm used. Nonetheless, we note that although conventional DVFS might become less effective with technology scaling, it will continue to play an important role in the context of emerging power management techniques, for example, for massively parallel multiprocessor systems where only a subset of cores can be turned on at any given point of time due to total power constraints. Siddharth Garg, Diana Marculescu, Radu Marculescu |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2011 | FARM: Fault-aware resource management in NoC-based multiprocessor platformsabstractIn this paper, we address the problem of run-time resource management in non-ideal multiprocessor platforms where communication happens via the Network-on-chip (NoCs) approach. More precisely, we propose a system-level fault-tolerant technique for application mapping which aims at optimizing the entire system performance and communication energy consumption, while considering the occurrence of permanent, transient, and intermittent faults in the system. As the main theoretical contribution, we address the problem of spare core placement and its impact on system fault-tolerance (FT) properties. Then, we investigate several metrics and provide insight into the fault-aware resource management process for such non-ideal multiprocessor platforms. Experimental results show that our proposed resource management technique is efficient and highly scalable and significant throughput improvements can be achieved compared to the existing solutions that do not consider failures in the system. Chen-Ling Chou, Radu Marculescu |
DATE | 2 |
| 2011 | Sustainability through massively integrated computing: Are we ready to break the energy efficiency wall for single-chip platforms?abstractWhile traditional cluster computers are more constrained by power and cooling costs for solving extreme-scale (or exascale) problems, the continuing progress and integration levels in silicon technologies make possible complete end-user systems on a single chip. This massive level of integration makes modern multicore chips all pervasive in domains ranging from climate forecasting and astronomical data analysis, to consumer electronics, smart phones, and biological applications. Consequently, designing multicore chips for exascale computing while using the embedded systems design principles looks like a promising alternative to traditional cluster-based solutions. This paper aims to present an overview of new, far-reaching design methodologies and run-time optimization techniques that can help breaking the energy efficiency wall in massively integrated single-chip computing platforms. Partha Pratim Pande, Fabien Clermidy, Diego Puschini, Imen Mansouri, Paul Bogdan, Radu Marculescu, Amlan Ganguly |
DATE | 6 |
| 2011 | Dynamic power management of voltage-frequency island partitioned Networks-on-Chip using Intel's Single-chip Cloud ComputerabstractContinuous technology scaling has enabled the integration of multiple cores on the same chip. To overcome the disadvantages of buses, the Network-on-Chip (NoC) architecture has been proposed as a new communication paradigm. To further mitigate the tradeoff between performance and power consumption, dynamic voltage and frequency scaling (DVFS) became the de facto approach in multi-core design. DVFS-based NoC communication was implemented in Intel's most recent Singlechip Cloud Computer (SCC). Using the SCC we demonstrate a power management algorithm that runs in real time and dynamically adjusts the performance of the islands to reduce power consumption while maintaining the same level of performance. David Radu, Paul Bogdan, Radu Marculescu, Ümit Y. Ogras |
NOCS | 3 |
| 2011 | A software framework for trace analysis targeting multicore platforms designabstractThis demonstration presents a complete software framework for dynamically mapping multi-threaded applications on a cycle accurate Network-on-Chip (NoC) architecture, analyzing the statistics of network workloads and drawing general guidelines regarding the design and optimization of NoCs. Guopeng Wei, Paul Bogdan, Radu Marculescu |
NOCS | 3 |
| 2011 | Non-Stationary Traffic Analysis and Its Implications on Multicore Platform DesignabstractNetworks-on-chip (NoCs) have been proposed as a viable solution to solving the communication problem in multicore systems. In this new setup, mapping multiple applications on available computational resources leads to interaction and contention at various network resources. Consequently, taking into account the traffic characteristics becomes of crucial importance for performance analysis and optimization of the communication infrastructure, as well as proper resource management. Although queuing-based approaches have been traditionally used for performance analysis purposes, they cannot properly account for many of the traffic characteristics (e.g., non-stationarity, self-similarity) that are crucial for multicore platform design. To overcome these limitations, we propose a statistical physics inspired approach to capture the traffic dynamics in multicore systems. As shown later in this paper, this is of fundamental significance for re-thinking the very basis of multicore systems design; it also opens up new research directions into NoC optimization which require accurate models of time-dependent and space-dependent traffic behavior. Paul Bogdan, Radu Marculescu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2011 | Hitting Time Analysis for Fault-Tolerant Communication at Nanoscale in Future Multiprocessor PlatformsabstractThis paper investigates the on-chip stochastic communication and proposes an analytical model for computing its mean hitting time. Toward this end, we model the on-chip stochastic communication of any source-destination pair as a branching and annihilating random walk taking place on a finite mesh. The evolution of this branching process is studied via a master equation which helps us estimate the mean number of communication rounds needed to reach a destination node from a particular source node. Besides the probabilistic performance analysis, we also present experimental results for two concrete platforms and assess the potential of stochastic communication for future nanotechnologies. Paul Bogdan, Radu Marculescu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2010 | Find your flow: increasing flow experience by designing "human" embedded systemsabstractIn this paper, we argue that future systems need to be designed using a flexible user-centric design methodology geared primarily toward maximizing the user satisfaction (i.e., flow experience) rather than seeking mainly optimization of performance and power consumption. Compared to the traditional design, we aim at re-focusing the current design paradigm by placing the user behavior at the center of the design process and by using psychological variables such as user ability and motivation as the main drivers of this process. This allows systems to become more capable of promptly adapting to different users' needs and of enhancing short- and long-term user satisfaction. Preliminary results show the potential of this methodology for maximizing the users' positive experience when running embedded applications on multiprocessor platforms. Chen-Ling Chou, Anca M. Miron, Radu Marculescu |
DAC | 3 |
| 2010 | Custom feedback control: enabling truly scalable on-chip power management for MPSoCsabstractIn this paper, we propose Custom Feedback Control, a new dynamic voltage and frequency control architecture for MP-SoC designs that bridges the gap between the two extreme points on the performance versus implementation cost trade-off curve, i.e., fully-centralized and full-decentralized control architectures. We outline a methodology to efficiently explore the vast design space of Custom Feedback control architectures, enabling designers to synthesize controllers that meet both the performance and implementation cost criteria. Our experimental results on an MPSoC platform running a video-encoding application demonstrate that, for the same energy dissipation, Custom Feedback control can achieve within 5% of the performance of a fully-centralized controller with only 17% of the implementation cost. In contrast, the performance of a fully-decentralized controller can be up to 2.5X worse than that of the fully-centralized controller. Siddharth Garg, Diana Marculescu, Radu Marculescu |
ISLPED | 3 |
| 2010 | QuaLe: A Quantum-Leap Inspired Model for Non-stationary Analysis of NoC Traffic in Chip Multi-processorsabstractThis paper identifies non-stationary effects in grid like Network-on-Chip (NoC) traffic and proposes QuaLe, a novel statistical physics-inspired model, that can account for non-stationarity observed in packet arrival processes. Using a wide set of real application traces, we demonstrate the need for a multi-fractal approach and analyze various packet arrival properties accordingly. As a case study, we show the benefits of our multifractal approach in estimating the probability of missing deadlines in packet scheduling for chip multiprocessors (CMPs). Paul Bogdan, Miray Kas, Radu Marculescu, Onur Mutlu |
NOCS | 3 |
| 2010 | Run-Time Task Allocation Considering User Behavior in Embedded Multiprocessor Networks-on-ChipabstractIn this paper, we propose a run-time strategy for allocating application tasks to embedded multiprocessor systems-on-chip platforms where communication happens via the network-on-chip approach. As a novel contribution, we incorporate the user behavior information in the resource allocation process; this allows the system to better respond to real-time changes and to adapt dynamically to different user needs. Several algorithms are proposed for solving the task allocation problem while minimizing the communication energy consumption and network contention. When the user behavior is taken into consideration, we observe more than 70% communication energy savings (with negligible energy and run-time overhead) compared to an arbitrary contiguous task allocation strategy. Chen-Ling Chou, Radu Marculescu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2010 | Designing Heterogeneous Embedded Network-on-Chip Platforms With Users in MindabstractIn this paper, we propose a user-centric design methodology targeting heterogeneous embedded systems-on-chip where communication happens via the network-on-chip approach. More precisely, in this new design methodology, we consider explicitly the information about the user experience and apply machine learning techniques to develop a design flow which aims at minimizing the workload variance; this allows the system to better adapt to different types of user needs and workload variations. Our experimental results show that by considering the user experience into the design space exploration step, the system platforms generated by our approach achieve more than 30% energy savings, on average, compared to the single platform derived from the traditional design flow; this implies that each system configuration we generate is highly suitable for the targeted class of user and workload behaviors. Chen-Ling Chou, Radu Marculescu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2010 | Guest Editorial: Special Section on the ACM/IEEE Symposium on Networks-on-Chip 2009abstractThe four papers in this special section are extended versions of papers presented at the 3rd ACM/IEEE Symposium on Networks-on-Chip (NOCS) in San Diego, CA, in 2009. Radu Marculescu, Axel Jantsch |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2010 | An Analytical Approach for Network-on-Chip Performance AnalysisabstractNetworks-on-chip (NoCs) have recently emerged as a scalable alternative to classical bus and point-to-point architectures. To date, performance evaluation of NoC designs is largely based on simulation which, besides being extremely slow, provides little insight on how different design parameters affect the actual network performance. Therefore, it is practically impossible to use simulation for optimization purposes. In this paper, we present a mathematical model for on-chip routers and utilize this new model for NoC performance analysis. The proposed model can be used not only to obtain fast and accurate performance estimates, but also to guide the NoC design process within an optimization loop. The accuracy of our approach and its practical use is illustrated through extensive simulation results. Ümit Y. Ogras, Paul Bogdan, Radu Marculescu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2009 | Technology-driven limits on DVFS controllability of multiple voltage-frequency island designs: a system-level perspectiveabstractIn this paper, we consider the case of network-on-chip (NoC) based multiple-processor systems-on-chip (MPSoCs) implemented using multiple voltage and frequency islands (VFIs) that rely on fine-grained dynamic voltage and frequency scaling (DVFS) for run-time control of the system power dissipation. Specifically, we present a framework to compute theoretical bounds on the performance of DVFS controllers for such systems under the impact of three important technology driven constraints: (i) reliability and temperature driven upper limits on the maximum supply voltage; (ii) inductive noise driven constraints on the maximum rate of change of voltage/frequency; and (iii) increasing manufacturing process variations. Our experimental results show that, for the benchmarks considered, any DVFS control algorithm will lose up to 87% performance, measured in terms of the number of steps required to reach a reference steady state, in the presence of maximum frequency and maximum frequency increment constraints. In addition, increasing process variations can lead to up to 60% of fabricated chips being unable to meet the specified DVFS control specifications, irrespective of the DVFS algorithm used. Siddharth Garg, Diana Marculescu, Radu Marculescu, Ümit Y. Ogras |
DAC | 3 |
| 2009 | User-centric design space exploration for heterogeneous Network-on-Chip platformsabstractIn this paper, we present a design methodology for automatic platform generation of future heterogeneous systems where communication happens via the network-on-chip (NoC) approach. As a novel contribution, we consider explicitly the information about the user experience into a design flow which aims at minimizing the workload variance; this allows the system to better adapt to different types of user needs and workload variations. More specifically, we first collect various user traces from various applications and generate specific clusters using machine learning techniques. For each cluster of such user traces, depending on the architectural parameters extracted from high-level specifications, we propose an optimization method to generate the NoC system architecture. Finally, we validate the user-centric design space exploration using realistic traces and compare it to the traditional NoC design methodology. Chen-Ling Chou, Radu Marculescu |
DATE | 2 |
| 2009 | Outstanding Research Problems in NoC Design: System, Microarchitecture, and Circuit PerspectivesabstractTo alleviate the complex communication problems that arise as the number of on-chip components increases, network-on-chip (NoC) architectures have been recently proposed to replace global interconnects. In this paper, we first provide a general description of NoC architectures and applications. Then, we enumerate several related research problems organized under five main categories: Application characterization, communication paradigm, communication infrastructure, analysis, and solution evaluation. Motivation, problem description, proposed approaches, and open issues are discussed for each problem from system, microarchitecture, and circuit perspectives. Finally, we address the interactions among these research problems and put the NoC design process into perspective. Radu Marculescu, Ümit Y. Ogras, Li-Shiuan Peh, Natalie D. Enright Jerger, Yatin Vasant Hoskote |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2009 | Design and Management of Voltage-Frequency Island Partitioned Networks-on-ChipabstractThe design of many core systems-on-chip (SoCs) has become increasingly challenging due to high levels of integration, excessive energy consumption and clock distribution problems. To deal with these issues, we consider network-on-chip (NoC) architectures partitioned into several voltage-frequency islands (VFIs) and propose a design methodology for runtime energy management. The proposed approach minimizes the energy consumption subject to performance constraints. Then, we present efficient techniques for on-the-fly workload monitoring and management to ensure that the system can cope with variability in the workload and various technology-related parameters. Simulation results demonstrate the effectiveness of our approach in reducing the overall system energy consumption for a real video application. Finally, the results and functional correctness are validated using an field-programmable gate-array (FPGA) prototype for an NoC with multiple VFIs. Ümit Y. Ogras, Radu Marculescu, Diana Marculescu, Eun-Gu Jung |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2008 | Predictive Energy-Efficient Multicast for Large-Scale Mobile Ad Hoc NetworksabstractEnergy-efficient multicast routing is of primary concern for mobile ad hoc networks (MANET). However, none of existing energy-efficient multicast algorithms is applicable to large-scale MANETs, either due to their complexity (which is either NP-hard or polynomial with respect to the network size), or due to the huge overhead caused by frequent exchanges of location information. To solve the scalability and overhead issues, we propose the Predictive .Energy-efficient Multicast Algorithm (PEMA) which exploits statistical properties of the network, as opposed to relying on route details or network topology. The running time of PEMA depends on the multicast group size, not network size; this makes PEMA fast enough even for MANETs consisting of 1000 or more nodes. Simulation results show that PEMA not only results in significant energy savings compared to other existing algorithms, but also attains good packet delivery ratio in mobile environments. Jung-Chun Kao, Radu Marculescu |
CCNC | 2 |
| 2008 | Variation-adaptive feedback control for networks-on-chip with multiple clock domainsabstractThis paper discusses the use of networks-on-chip (NoCs) consisting of multiple voltage-frequency islands to cope with power consumption, clock distribution and parameter variation problems in future multiprocessor systems-on-chip (MPSoCs). In this architecture, communication within each island is synchronous, while communication across different islands is achieved via mixed-clock, mixed-voltage queues. In order to dynamically control the speed of each domain in the presence of parameter and workload variations, we propose a robust feedback control methodology. Towards this end, we first develop a state-space model based on the utilization of the inter-domain queues. Then, we identify the theoretical conditions under which the network is controllable. Finally, we synthesize state feedback controllers to cope with workload variations and minimize power consumption. Experimental results demonstrate robustness to parameter variations and more than 40 % energy savings by exploiting workload variations through dynamic voltagefrequency scaling (DVFS) for a hardware MPEG-2 encoder design. Ümit Y. Ogras, Radu Marculescu, Diana Marculescu |
DAC | 2 |
| 2008 | User-Aware Dynamic Task Allocation in Networks-on-ChipabstractIn this paper, we propose a run-time strategy for allocating the application tasks to platform resources in homogeneous networks-on-chip (NoCs). As novel contribution, we incorporate the user behavior information in the resource allocation process; this allows system to better respond to real-time changes and adapt dynamically to user needs. Several algorithms are then proposed for solving the task allocation problem, while minimizing the communication energy consumption and network contention. If user behavior is taken into consideration, we observe about 60% communication energy savings (with negligible and energy runtime overhead) compared to an arbitrary task allocation strategy. Chen-Ling Chou, Radu Marculescu |
DATE | 2 |
| 2008 | Contention-aware application mapping for Network-on-Chip communication architecturesabstractIn this paper, we analyze the impact of network contention on the application mapping for tile-based network-on-chip (NoC) architectures. Our main theoretical contribution consists of an integer linear programming (ILP) formulation of the contention-aware application mapping problem which aims at minimizing the inter-tile network contention. To solve the scalability problem caused by ILP formulation, we propose a linear programming (LP) approach followed by an mapping heuristic. Taken together, they provide near-optimal solutions while reducing the runtime significantly. Experimental results show that, compared to other existing mapping approaches based on communication energy minimization, our contention-aware mapping technique achieves a significant decrease in packet latency (and implicitly, a throughput increase) with a negligible communication energy overhead. Chen-Ling Chou, Radu Marculescu |
ICCD | 2 |
| 2008 | Communication-Aware Face Detection Using Noc Architecture
Hung-Chih Lai, Radu Marculescu, Marios Savvides, Tsuhan Chen |
ICVS | 2 |
| 2008 | Energy- and Performance-Aware Incremental Mapping for Networks on Chip With Multiple Voltage LevelsabstractAchieving effective run-time mapping on multiprocessor systems-on-chip (MPSoCs) is a challenging task, particularly since the arrival order of the target applications is not known a priori. This paper targets real-time applications which are dynamically mapped onto embedded MPSoCs, where communication happens via the Network-on-Chip (NoC) approach, and resources connected to the NoC have multiple voltage levels. We address precisely the energy- and performance-aware incremental mapping problem for NoCs with multiple voltage levels and propose an efficient technique (consisting of region selection and node allocation) to solve it. Moreover, the proposed technique allows for new applications to be added to the system with minimal in- terprocessor communication overhead. Experimental results show that the proposed technique is very fast, and as much as 50% communication energy savings can be achieved compared to using an arbitrary allocation scheme. Chen-Ling Chou, Ümit Y. Ogras, Radu Marculescu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2008 | Analysis and optimization of prediction-based flow control in networks-on-chipabstractNetworks-on-Chip (NoC) communication architectures have emerged recently as a scalable solution to on-chip communication problems. While the NoC architectures may offer higher bandwidth compared to traditional bus-based communication, their performance can degrade significantly in the absence of effective flow control algorithms. Unfortunately, flow control algorithms developed for macronetworks, either rely on local information, or suffer from large communication overhead and unpredictable delays. Hence, using them in the NoC context is problematic at best. For this reason, we propose a predictive closed-loop flow control mechanism and make the following contributions: First, we develop traffic source and router models specifically targeted to NoCs. Then, we utilize these models to predict the possible congestion in the network. Based on this information, the proposed scheme controls the packet injection rate at traffic sources in order to regulate the total number of packets in the network. We also illustrate the proposed traffic source model and the applicability of the proposed flow controller to actual designs using real NoC implementations. Finally, simulations and experimental study using our FPGA prototype show that the proposed controller delivers a better performance compared to the traditional switch-to-switch flow control algorithms under various real and synthetic traffic patterns. Ümit Y. Ogras, Radu Marculescu |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2008 | Enabling multimedia using resource-constrained video processing techniques: A node-centric perspectiveabstractSuccessful proliferation of multimedia-enabled devices and advances in very large-scale integration (VLSI) technology has spawned new research efforts in migrating video processing applications onto ever smaller and more inexpensive devices. This article focuses on the technical challenges associated with that migration. Due to limitations in size, battery lifetime, and, ultimately, cost, mapping complex video applications onto resource-constrained systems is a very challenging proposition. To this end, we first consider a technique, region-of-interest (ROI) processing, of defining a window within a video frame and only operating on the data inside that window, ignoring the rest of the frame. By using this lossy technique, the processing requirements can be reduced by roughly 80% while the error introduced in the quality of the results is roughly 10%. The other technique is adaptive data partitioning (ADP) combined with a content-based power management algorithm. By distributing video processing among multiple processors and shutting them down when they are not needed, the energy consumed per processor can be reduced by 60% without sacrificing the performance of the underlying video-based application. Taken together, these novel techniques enable ambient multimedia systems and maintain the needed overall efficiency in video processing. Nicholas H. Zamora, Xiaoping Hu 0003, Ümit Y. Ogras, Radu Marculescu |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2007 | Quantum-Like Effects in Network-on-Chip Buffers BehaviorabstractThis paper proposes a paradigm shift in modeling and optimization of NoCs by identifying quantum-like effects in buffers behavior. The key idea of our proposed approach involves identifying a virtual random growing network (VRGN) which describes the NoC buffers evolution as a function of the packet injection rate. Paul Bogdan, Radu Marculescu |
DAC | 2 |
| 2007 | Voltage-Frequency Island Partitioning for GALS-based Networks-on-ChipabstractDue to high levels of integration and complexity, the design of multi-core SoCs has become increasingly challenging. In particular, energy consumption and distributing a single global clock signal throughout a chip have become major design bottlenecks. To deal with these issues, a globally asynchronous, locally synchronous (GALS) design is considered for achieving low power consumption and modular design. Such a design style fits nicely with the concept of voltage-frequency islands (VFIs) which has been recently introduced for achieving fine-grain system-level power management. This paper proposes a design methodology for partitioning an NoC architecture into multiple VFIs and assigning supply and threshold voltage levels to each VFI. Simulation results show about 40% savings for a real video application and demonstrate the effectiveness of our approach in reducing the overall system energy consumption. The results and functional correctness are validated using an FPGA prototype for an NoC with multiple VFIs. Ümit Y. Ogras, Radu Marculescu, Puru Choudhary, Diana Marculescu |
DAC | 2 |
| 2007 | Analytical router modeling for networks-on-chip performance analysis
Ümit Y. Ogras, Radu Marculescu |
DATE | 2 |
| 2007 | Distributed power-management techniques for wireless network video systemsabstractWireless sensor networks operating on limited energy resources need to be power efficient to extend the system lifetime. This is especially challenging for video sensor networks due to the large volumes of data they need to process in short periods of time. Towards this end, this paper proposes two coordinated power management policies for video sensor networks. These policies are scalable as the system grows and flexible to video parameters and network characteristics. In addition to simulation results, the prototype demonstrates the feasibility of implementing these policies. Finally, the analytical framework we provide gives an upper bound for the achievable sleep fraction and insight into how adjusting select parameters will affect the performance of the power management policies Nicholas H. Zamora, Jung-Chun Kao, Radu Marculescu |
DATE | 3 |
| 2007 | Energy-efficient anonymous multicast in mobile ad-hoc networksabstractProtecting personal privacy and energy efficiency are two primary concerns for mobile ad hoc networks. However, no energy-efficient multicast algorithm designed for preserving anonymity has been proposed to date. At the same time, existing approaches cannot be applied to anonymous routing due to their incapability of preserving anonymity. To solve this critical issue, we propose an energy-efficient anonymous multicast algorithm (EEAMA), which relies only on the statistical properties of the wireless network. This not only makes EEAMA suitable to preserving anonymity, but also reduces its execution time significantly. The complexity of EEAMA increases polynomially with the size of the multicast group, as opposed to the size of the network which determines the complexity of all approaches in the literature. Extensive simulation results show that compared to anonymous unicast, EEAMA offers both better performance (in terms of packet delivery ratio, end-to-end delay and network throughput) and significant energy savings. Jung-Chun Kao, Radu Marculescu |
ICPADS | 2 |
| 2007 | Towards Open Network-on-Chip BenchmarksabstractMeasuring and comparing performance, cost, and other features of advanced communication architectures for complex multi core/multiprocessor systems on chip is a significant challenge which has hardly been addressed so far. This document outlines the top-level view on a system of benchmarks for networks on chip (NoC), which intends to cover a wide spectrum of NoC design aspects, from application modeling to performance evaluation and post-manufacturing test and reliability. For performance benchmarking, requirements and features are described for application programs, synthetic micro-benchmarks, and abstract benchmark applications. Then, it proposes ways to measure and benchmark reliability, fault tolerance and testability of the on-chip communication fabric. This paper introduces the main concepts and ideas for benchmarking NoCs in a systematic and comparable way. It will be followed up by a report that will define a benchmark framework and the syntax of interfaces for benchmark programs that will allow the community to build-up a benchmark suite Cristian Grecu, André Ivanov, Partha Pratim Pande, Axel Jantsch, Erno Salminen, Ümit Y. Ogras, Radu Marculescu |
NOCS | 7 |
| 2007 | Real-Time Anonymous Routing for Mobile Ad Hoc NetworksabstractWe propose the anonymous symmetrically cryptographic (ASC) routing protocol which is entirely based on a symmetric cryptosystem. The ASC protocol preserves the identity privacy, location privacy and route anonymity with only a negligible overhead in terms of processing requirements and packet size. Furthermore, the ASC protocol does not rely on any trusted agent or centralized mechanism (both of which are impractical in hostile environments). Compared with other anonymous routing protocols, the ASC protocol reduces the end-to-end delay by orders of magnitude, while performing comparably well in terms of packet delivery ratio. These features enable anonymity for applications with strict QoS requirements (e.g. video and audio streaming) over mobile ad hoc networks. Jung-Chun Kao, Radu Marculescu |
WCNC | 2 |
| 2007 | Minimizing Eavesdropping Risk by Transmission Power Control in Multihop Wireless NetworksabstractTo defend against reconnaissance activity in ad hoc wireless networks, we propose transmission power control as an effective mechanism for minimizing the eavesdropping risk. Our main contributions are given as follows. First, we cast the wth-order eavesdropping risk as the maximum probability of packets being eavesdropped when there are w adversarial nodes in the network. Second, we derive the closed-form solution of the first-order eavesdropping risk as a polynomial function of the normalized transmission radius. This derivation assumes a uniform distribution of user nodes. Then, we generalize the model to allow arbitrary user nodes distribution and prove that the uniform user distribution minimizes the first-order eavesdropping risk. This result plays an essential role in deriving analytical bounds for the eavesdropping risk given arbitrary user distributions. Our simulation results show that, for a wide range of nonuniform traffic patterns, the difference in their eavesdropping risk values from the corresponding lower bounds is 3 dB or less. Jung-Chun Kao, Radu Marculescu |
IEEE Trans. Computers | 2 |
| 2007 | On-chip communication architecture exploration: A quantitative evaluation of point-to-point, bus, and network-on-chip approachesabstractTraditionally, design-space exploration for systems-on-chip (SoCs) has focused on the computational aspects of the problem at hand. However, as the number of components on a single chip and their performance continue to increase, a shift from computation-based to communication-based design becomes mandatory. As a result, the communication architecture plays a major role in the area, performance, and energy consumption of the overall system. This article presents a comprehensive evaluation of three on-chip communication architectures targeting multimedia applications. Specifically, we compare and contrast the network-on-chip (NoC) with point-to-point (P2P) and bus-based communication architectures in terms of area, performance, and energy consumption. As the main contribution, we present complete P2P, bus-, and NoC-based implementations of a real multimedia application (i. e. the MPEG-2 encoder), and provide direct measurements using an FPGA prototype and actual video clips, rather than simulation and synthetic workloads. We also support the experimental findings through a theoretical analysis. Both experimental and analysis results show that the NoC architecture scales very well in terms of area, performance, energy, and design effort, while the P2P and bus-based architectures scale poorly on all accounts except for performance and area, respectively. Hyung Gyu Lee, Naehyuck Chang, Ümit Y. Ogras, Radu Marculescu |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2007 | System-level performance/power analysis for platform-based design of multimedia applicationsabstractThe objective of this article is to introduce the use of Stochastic Automata Networks (SANs) as an effective formalism for application-architecture modeling in system-level average-case analysis for platform-based design. By platform, we mean a family of heterogeneous architectures that satisfy a set of architectural constraints imposed to allow re-use of hardware and software components. More precisely, we show how SANs can be used early in the design cycle to identify the best performance/power trade-offs among several application-architecture combinations. Having this information available not only helps avoid lengthy simulations for predicting power and performance figures, but also enables efficient mapping of different applications onto a chosen platform. We illustrate the benefits of our methodology by using the “Picture-in-Picture” video decoder as a driver application. Nicholas H. Zamora, Xiaoping Hu 0003, Radu Marculescu |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2006 | Design space exploration and prototyping for on-chip multimedia applicationsabstractTraditionally, design space exploration for Systems-on-Chip (SoCs) has focused on the computational aspects of the problem at hand. However, as the number of components on a single chip and their performance continue to increase, a shift from computation-bound to communication-bound design becomes mandatory. Towards this end, this paper presents a comprehensive evaluation of two communication architectures targeting multimedia applications. Specifically, we compare and contrast the Network-on-Chip (NoC) and Point-to-Point (P2P) communication architectures in terms of power, performance, and area. As the main contribution, we present complete P2P and NoC-based implementations of a real multimedia application (MPEG-2 encoder), and provide direct measurements using a FPGA prototype and actual video clips, rather than simulation and synthetic workload. From an experi-mental standpoint, we show that the NoC architecture scales very well in terms of area, performance, power and design effort, while the P2P architecture scales poorly on all accounts except performance. Hyung Gyu Lee, Ümit Y. Ogras, Radu Marculescu, Naehyuck Chang |
DAC | 3 |
| 2006 | Prediction-based flow control for network-on-chip trafficabstractNetworks-on-Chip (NoC) architectures provide a scalable solution to on-chip communication problem but the bandwidth offered by NoCs can be utilized efficiently only in presence of effective flow control algorithms. Unfortunately, the flow control algorithms pub-lished to date for macronetworks, either rely on local information, or suffer from large communication overhead and unpredictable delays. Hence, using them in the NoC context is problematic at best. For this reason, we propose a predictive closed-loop flow con-trol mechanism and make the following contributions: First, we develop traffic source and router models specifically targeted to NoCs. Then, we utilize these models to predict the cases of possible congestion in the network. Based on this information, the proposed scheme controls the packet injection rate at traffic sources in order to regulate the total number of packets in the network. Evaluations involving real and synthetic traffic patterns show that the proposed controller delivers a superior performance compared to the traditional switch-to-switch flow control algorithms. Ümit Y. Ogras, Radu Marculescu |
DAC | 2 |
| 2006 | Is "Network" the next "Big Idea" in design?abstractAs the complexity of nowadays systems continues to grow, we are moving away from creating individual components from scratch, toward methodologies that emphasize composition of re-usable components via the network paradigm. Complex component interactions can create a range of amazing behaviors, some useful, some unwanted, some even dangerous. To manage them, a "science" for network design is evolving, applicable in some surprising areas. In this paper, we consider a few application domains and discus the design challenges involved from a methodology standpoint. From large-scale hardware/software systems, to dynamically adaptive sensor networks, and network-on-chip architectures, these ideas find wide application Radu Marculescu, Jan M. Rabaey, Alberto L. Sangiovanni-Vincentelli |
DATE | 1 |
| 2006 | Communication architecture optimization: making the shortest path shorter in regular networks-on-chipabstractNetwork-on-chip (NoC)-based communication represents a promising solution to complex on-chip communication problems. Due to their regular structure, mesh-like NoC architectures have become very popular recently. However, they have poor topological properties such as long inter-node distances. In this paper, we address this very issue and explore the potential of partial NoC customization to improve both static and dynamic properties of the network significantly, while minimally affecting its regularity. Precise energy measurements on an FPGA prototype show that the improvement in network properties is achieved without a significant penalty in area and communication energy consumption. Ümit Y. Ogras, Radu Marculescu, Hyung Gyu Lee, Naehyuck Chang |
DATE | 2 |
| 2006 | Generalized Rate Analysis for Media-Processing PlatformsabstractIn this paper we address the "rate analysis" problem for media-processing pla$brnzs consisting oJ'rnultiple processor cores connected zn a pipelined fashion. More precisely, we aim at determining tight bounds on the rates at which multimedia streams can be fed into such urchitectures. These bounds depend on urchitectuml constt-uints (e.g. the available on-chip memory, bus urhitrution policies, etc.), as well as the upplicution churacteristics (e.g. application partitioning und mapping, workloud rutes generaled by different tasks, etc.). The proposed frurnework for rate analysis can be used for fast design space exploration to determine how these bounds change with different architrctural parameters, mapping of the application, or chunging the QoS requirements associated with the input strrarns. Samarjit Chakraborty, Radu Marculescu |
RTCSA | 3 |
| 2006 | Eavesdropping Minimization via Transmission Power Control in Ad-Hoc Wireless NetworksabstractReconnaissance activity is the most frequent incident on computer networks since 2002. In fact, most attacks (including DoS attacks) are usually preceded by reconnaissance activity. In order to defend against reconnaissance activity in ad-hoc wireless networks, we propose to use transmission power control as an effective mean to minimize the eavesdropping risk. Our main contributions are as follows: first, we cast the w-th order eavesdropping risk as the maximum probability of packets being eavesdropped when there are w adversarial nodes in the network. Second, we derive the closed-form solution of the 1st order eavesdropping risk as a 3rd-order polynomial function of normalized transmission radius. This derivation is based on the recently proposed model by El Gamal which assumes a uniform distribution of user nodes. Then we generalize the model to allow arbitrary user nodes distribution and prove that the uniform user distribution actually minimizes the 1st order eavesdropping risk. This result plays an essential role in deriving the first analytical bounds for the eavesdropping risk given arbitrary user distribution. Our simulation results show that for a wide range of non-uniform traffic patterns, the eavesdropping risk has the same order of magnitude as the corresponding uniform traffic cases Jung-Chun Kao, Radu Marculescu |
SECON | 2 |
| 2006 | On Optimization of E-Textile Systems Using Redundancy and Energy-Aware RoutingabstractRecent advances in the electronic device manufacturing technology have opened many research opportunities in pervasive computing. Among the emerging design platforms, "electronic textiles" (or e-textiles) make possible a wide variety of novel applications, ranging from consumer electronics to aerospace devices. Due to the harsh environment of e-textile components and battery size limitations, low-power and redundancy techniques are critical for obtaining successful e-textile applications. In this paper, we consider a platform which consists of dedicated components for e-textiles, including computational modules, dedicated transmission lines, and thin-film batteries on fiber substrates. As a theoretical contribution, we address the issue of the energy-aware routing for e-textile platforms and propose an efficient algorithm to solve it. Furthermore, we derive an analytical upper bound for determining the maximum number of achievable jobs over all possible e-textile routing frameworks. From a practical standpoint, for the Advanced Encryption Standard (AES) cipher, the routing technique we propose achieves close to or more than 75 percent of this theoretical upper bound. Moreover, compared to the non-energy-aware counterpart, the new routing technique increases the number of encryption jobs by one order of magnitude. Jung-Chun Kao, Radu Marculescu |
IEEE Trans. Computers | 2 |
| 2006 | System-Level Buffer Allocation for Application-Specific Networks-on-Chip Router DesignabstractIn this paper, a novel system-level buffer planning algorithm that can be used to customize the router design in networks-on-chip (NoCs) is presented. More precisely, given the traffic characteristics of the target application and the total budget of the available buffering space, the proposed algorithm automatically assigns the buffer depth for each input channel, in different routers across the chip, such that the overall performance is maximized. This is in deep contrast with the uniform assignment of buffering resources (currently used in NoC design), which can significantly degrade the overall system performance. Indeed, the experimental results show that while the proposed algorithm is very fast, significant performance improvements can be achieved compared to the uniform buffer allocation. For instance, for a complex audio/video application, about 80% savings in buffering resources, can be achieved by smart buffer allocation using the proposed algorithm Jingcao Hu, Ümit Y. Ogras, Radu Marculescu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2006 | Computation and communication refinement for multiprocessor SoC design: A system-level perspectiveabstractContinuous advancements in semiconductor technology enable the design of complex systems-on-chips (SoCs) composed of tens or hundreds of IP cores. At the same time, the applications that need to run on such platforms have become increasingly complex and have tight power and performance requirements. Achieving a satisfactory design quality under these circumstances is only possible when both computation and communication refinement are performed efficiently, in an automated and synergistic manner. Consequently, formal and disciplined system-level design methodologies are in great demand for future multiprocessor design. This article provides a broad overview of some fundamental research issues and state-of-the-art solutions concerning both computation and communication aspects of system-level design. The methodology we advocate consists of developing abstract application and platform models, followed by application mapping onto the target platform, and then optimizing the overall system via performance analysis. In addition, a communication refinement step is critical for optimizing the communication infrastructure in this multiprocessor setup. Finally, simulation and prototyping can be used for accurate performance evaluation purposes. Radu Marculescu, Ümit Y. Ogras, Nicholas H. Zamora |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2006 | "It's a small world after all": NoC performance optimization via long-range link insertionabstractNetworks-on-chip (NoCs) represent a promising solution to complex on-chip communication problems. The NoC communication architectures considered so far are based on either completely regular or fully customized topologies. In this paper, we present a methodology to automatically synthesize an architecture which is neither regular nor fully customized. Instead, the communication architecture we propose is a superposition of a few long-range links and a standard mesh network. The few application-specific long-range links we insert significantly increase the critical traffic workload at which the network transitions from a free to a congested state. This way, we can exploit the benefits offered by both complete regularity and partial topology customization. Indeed, our experimental results demonstrate a significant reduction in the average packet latency and a major improvement in the achievable network through with minimal impact on network topology Ümit Y. Ogras, Radu Marculescu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2005 | Communication-Centric SoC Design for Nanoscale DomainabstractIn the realm of 35nm technology, it becomes possible to have thousands of IP blocks that need to communicate efficiently. Large-scale integration of these blocks onto a single chip makes the use of truly scalable networks-on-chips (NoC) communication architectures inevitable. This paper provides an overview of the outstanding research issues involved in designing application-specific NoC architectures by considering explicitly the level of customization envisioned in the communication architecture. For each category of approaches, we discuss the significance of the problem, provide a problem statement and survey the relevant solutions to date. Ümit Y. Ogras, Jingcao Hu, Radu Marculescu |
ASAP | 3 |
| 2005 | Energy-Aware Routing for E-Textile ApplicationsabstractAs the scale of electronic devices shrinks, "electronic textiles" (e-textiles) will make possible a wide variety of novel applications which are currently infeasible. Due to the wearability concerns, low-power techniques are critical for e-textile applications. In this paper, we address the issue of energy-aware routing for e-textile platforms and propose an efficient algorithm to solve it. The platform we consider consists of dedicated components for e-textiles, including computational modules, dedicated transmission lines and thin-film batteries on fiber substrates. Furthermore, we derive an analytical upper bound for the achievable number of jobs completed over all possible routing strategies. From a practical standpoint, for the advanced encryption standard (AES) cipher, the routing technique we propose achieves about fifty percent of this analytical upper bound. Moreover, compared to the non-energy-aware counterpart, our routing technique increases the number of encryption jobs completed by one order of magnitude. Jung-Chun Kao, Radu Marculescu |
DATE | 2 |
| 2005 | Energy- and Performance-Driven NoC Communication Architecture Synthesis Using a Decomposition ApproachabstractIn this paper, we present a methodology for customized communication architecture synthesis that matches the communication requirements of the target application. This is an important problem, particularly for network-based implementations of complex applications. Our approach is based on using frequently encountered generic communication primitives as an alphabet capable of characterizing any given communication pattern. The proposed algorithm searches through the entire design space for a solution that minimizes the system total energy consumption, while satisfying the other design constraints. Compared to the standard mesh architecture, the customized architecture generated by the newly proposed approach shows about 36% throughput increase and 51% reduction in the energy required to encrypt 128 bits of data with a standard encryption algorithm. Ümit Y. Ogras, Radu Marculescu |
DATE | 2 |
| 2005 | Application-specific network-on-chip architecture customization via long-range link insertionabstractNetworks-on-chip (NoCs) represent a promising solution to complex on-chip communication problems. The NoC communication architectures considered so far are based on either completely regular or fully customized topologies. In this paper, we present a methodology to automatically synthesize an architecture where a few application-specific long-range links are inserted on top of a regular mesh network. This way, we can better exploit the benefits of both complete regularity and partial customization. Indeed, our experimental results show that inserting application-specific long-range links significantly increases the critical traffic workload at which the network state transits from a free to a congested regime. This, in turn, results in a significant reduction in the average packet latency and a major improvement in the network achievable throughput. Ümit Y. Ogras, Radu Marculescu |
ICCAD | 2 |
| 2005 | Hierarchical Adaptive Dynamic Power ManagementabstractDynamic power management aims at extending battery life by switching devices to lower-power modes when there is a reduced demand for service. Static power management strategies can lead to poor performance or unnecessary power consumption when there are wide variations in the rate of requests for service. This paper presents a hierarchical scheme for adaptive dynamic power management (DPM) under nonstationary service requests. As the main theoretical contribution, we model the nonstationary request process as a Markov-modulated process with a collection of modes, each corresponding to a particular stationary request process. Optimal DPM policies are precalculated offline for selected modes using standard algorithms available for stationary Markov decision processes (MDPs). The power manager then switches online among these policies to accommodate the stochastic mode-switching request dynamics using an adaptive algorithm to determine the optimal switching rule based on the observed sample path. As a target application, we present simulations of hierarchical DPM for hard disk drives where the read/write request arrivals are modeled as a Markov-modulated Poisson process. Simulation results show that the power consumption of our approach under highly nonstationary request arrivals is less than that of a previously proposed heuristic approach and is even comparable to that of the optimal policy under stationary Poisson request process with the same arrival rate as the average arrival rate of the nonstationary request process. Bruce H. Krogh, Radu Marculescu |
IEEE Trans. Computers | 3 |
| 2005 | Energy- and performance-aware mapping for regular NoC architecturesabstractIn this paper, we present an algorithm which automatically maps a given set of intellectual property onto a generic regular network-on-chip (NoC) architecture and constructs a deadlock-free deterministic routing function such that the total communication energy is minimized. At the same time, the performance of the resulting communication system is guaranteed to satisfy the specified design constraints through bandwidth reservation. As the main theoretical contribution, we first formulate the problem of energy- and performance-aware mapping in a topological sense, and show how the routing flexibility can be exploited to expand the solution space and improve the solution quality. An efficient branch-and-bound algorithm is then proposed to solve this problem. Experimental results show that the proposed algorithm is very fast, and significant communication energy savings can be achieved. For instance, for a complex video/audio application, 51.7% communication energy savings have been observed, on average, compared to an ad hoc implementation. Jingcao Hu, Radu Marculescu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2004 | Enabling on-chip diversity through architectural communication design
Tudor Dumitras, Sam Kerner, Radu Marculescu |
ASP-DAC | 3 |
| 2004 | DyAD: smart routing for networks-on-chipabstractIn this paper, we present and evaluate a novel routing scheme called DyAD which combines the advantages of both deterministic and adaptive routing schemes. More precisely, we envision a new routing technique which judiciously switches between deterministic and adaptive routing based on the network's congestion conditions. The simulation results show the effectiveness of DyAD by comparing it with purely deterministic and adaptive routing schemes under different traffic patterns. Moreover, a prototype router based on the DyAD idea has been designed and evaluated. Compared to purely adaptive routers, the overhead of implementing DyAD is negligible (less than 7%), while the performance is consistently better. Jingcao Hu, Radu Marculescu |
DAC | 2 |
| 2004 | Adaptive data partitioning for ambient multimediaabstractIn the near future, Ambient Intelligence (AmI) will become part of everyday life. Combining feature-rich multimedia with AmI (dubbed Ambient Multimedia for short) has the potential of changing the way we perceive and interact with our environment. One major difficulty, however, in designing Ambient Multimedia Systems (AMS) comes from the strong constraints imposed on system resources by the AmI application requirements. In this paper, we propose a method for mapping multimedia applications on systems with very limited resources (i.e. memory, computing capability and battery lifetime) by combining adaptive data partitioning with dynamic power management. The potential of the approach is illustrated through a case study of an object tracking application running on a resource-constrained platform. Xiaoping Hu 0003, Radu Marculescu |
DAC | 2 |
| 2004 | Energy-Aware Communication and Task Scheduling for Network-on-Chip Architectures under Real-Time ConstraintsabstractIn this paper, we present a novel energy-aware scheduling (EAS) algorithm which statically schedules both communication transactions and computation tasks onto heterogeneous network-on-chip (NoC) architectures under real-time constraints. Our algorithm automatically assigns tasks onto different processing elements and then schedules their execution. At the same time, the algorithm also takes into consideration the exact communication delay by scheduling communication transactions in parallel. As the main contribution, we first formulate the problem of concurrent communication and task scheduling for heterogeneous NoC architectures and then propose an efficient heuristic to solve it. Experimental results show that significant energy savings can be achieved by using our energy-aware scheduler while meeting the specified performance constraints. For instance, for a complex multimedia application, 44% energy savings have been observed, on average, compared to the schedules generated by a standard earliest-deadline-first scheduler. Jingcao Hu, Radu Marculescu |
DATE | 2 |
| 2004 | Distributed Multimedia System Design: A Holistic PerspectiveabstractMultimedia systems play a central part in many human activities. Due to the significant advances in the VLSI technology, there is an increasing demand for portable multimedia appliances capable of handling advanced algorithms required in all forms of communication. Over the years, we have witnessed a steady move from standalone (or desktop) multimedia to deeply distributed multimedia systems. Whereas desktop-based systems are mainly optimized based on the performance constraints, power consumption is the key design constraint for multimedia devices that draw their energy from batteries. The overall goal of successful design is then to find the best mapping of the target multimedia application onto the architectural resources, while satisfying an imposed set of design constraints (e.g. minimum power dissipation, maximum performance) and specified QoS metrics (e.g. end-to-end latency, jitter, loss rate) which directly impact the media quality. This paper addresses a few fundamental issues that make the design process particularly challenging and offers a holistic perspective towards a coherent design methodology. Radu Marculescu, Massoud Pedram, Jörg Henkel |
DATE | 1 |
| 2004 | Hierarchical Adaptive Dynamic Power ManagementabstractThe main contribution of this paper is a novel hierarchical scheme for adaptive dynamic power management (DPM) under nonstationary service requests. We model the nonstationary arrival process of service requests as a Markov-modulated stochastic process in which the stochastic process for each modulation state models a particular stationary mode of the arrival process. The bottom layer of our hierarchical architecture is a set of stationary optimal DPM policies, pre-calculated off-line for selected modes from policy optimization in Markov decision processes. The supervisory power manager at the top layer adaptively and optimally switches among these stationary policies on-line to accommodate the actual mode-switching arrival dynamics. Simulation results show that our approach, under highly nonstationary requests, can lead to significant power savings compared to previously proposed heuristic approaches. Bruce H. Krogh, Radu Marculescu |
DATE | 3 |
| 2004 | Application-specific buffer space allocation for networks-on-chip router designabstractWe present a system-level buffer planning algorithm that can be used to customize the router design in networks-on-chip (NoCs). More precisely, given the traffic characteristics of the target application and the buffering space budget, our algorithm automatically assigns the buffer depth for each input channel, in different routers across the chip, to match the communication pattern, such that the overall performance is maximized. This is in deep contrast with the uniform assignment of buffering resources (currently used in NoC design) which can significantly degrade the overall system performance. For instance, for a complex audio/video application, about 85% savings in buffering resources can be achieved by smart buffer allocation using our algorithm without any reduction in performance. Jingcao Hu, Radu Marculescu |
ICCAD | 2 |
| 2004 | Toward an Integrated Design Methodology for Fault-Tolerant, Multiple Clock/Voltage Integrated SystemsabstractThis paper describes a communication-centric design methodology that addresses the fundamental challenges induced by the emergence of truly heterogeneous systems-on-chip (SoCs). For such systems, the globally asynchronous design paradigm seems to be the most promising (if not the only) solution for providing an underlying substrate for cost-effective and power efficient on-chip communication among diverse, mixed technology IPs. Additional challenges are related to reliability and error resilience of on-chip communication architectures. The proposed on-chip communication methodology targets all levels of abstraction, from circuit, to microarchitecture and system-level by seamlessly integrating solutions for robust and efficient globally asynchronous communication among diverse IPs. Radu Marculescu, Diana Marculescu, Lawrence T. Pileggi |
ICCD | 1 |
| 2004 | Data partitioning techniques for pervasive multimedia platformsabstractIn this paper, we propose a method for mapping multimedia applications on systems with very limited resources (i.e. memory, computing capability and battery lifetime) by combining adaptive data partitioning with content-based dynamic power management. The potential of the approach is illustrated through a case study of an object tracking application running on a resource constrained system which can be embedded in the environment (e.g. offices, home or conference rooms) to offer significantly more opportunities for ubiquitous information, seamless communication, enhanced security, etc. compared to today's portable or stationary devices. Besides power and performance trade-offs, we also explore the scaling effects on data partitioning and provide insights for possible optimization when designing such systems Xiaoping Hu 0003, Ümit Y. Ogras, Nicholas H. Zamora, Radu Marculescu |
ICME | 4 |
| 2004 | Resource-aware video processing techniques for ambient multimedia systemsabstractAmbient intelligence (AmI) is the inconspicuous presence of computing into every facet of our lives. AmI systems of the future will contain devices with highly limited resources in terms of processing power, memory, and battery lifetime. Contrary to this are the memory, power, and cycle-hungry video processing applications which are required to provide the level of utility demanded by users. The paper introduces the idea of processing a portion of a video frame with the intent of achieving high levels of video processing performance while reducing considerably the hardware requirements for that processing. The techniques we propose operate only on a portion of the video frame and optionally adjust that portion's size dynamically to match the video content reasonably well. Our results show that this technique can save roughly 75% in both memory and processing cycle requirements and up to 87% in energy consumption, while inducing less than 10% error in XY processing of the video frame Nicholas H. Zamora, Xiaoping Hu 0003, Ümit Y. Ogras, Radu Marculescu |
ICME | 4 |
| 2004 | Architecting voltage islands in core-based system-on-a-chip designsabstractVoltage islands enable core-level power optimization for System-on-Chip (SoC) designs by utilizing a unique supply voltage for each core. Architecting voltage islands involves island partition creation, voltage level assignment and floorplanning. The task of island partition creation and level assignment have to be done simultaneously in a floorplanning context due to the physical constraints involved in the design process. This leads to a floorplanning problem formulation that is very different from the traditional floorplanning for ASIC-style design.In this paper, we define the problem of architecting voltage islands in core-based designs and present a new algorithm for simultaneous voltage island partitioning, voltage level assignment and physical-level floorplanning. Application of the proposed algorithm to a few benchmark and industrial examples is demonstrated using a prototype tool. Results show power savings of 14%--28%, depending on the constraints imposed on the number of voltage islands and other physical-level parameters. Jingcao Hu, Youngsoo Shin, Nagu R. Dhanwada, Radu Marculescu |
ISLPED | 4 |
| 2004 | On-chip traffic modeling and synthesis for MPEG-2 video applicationsabstractThe objective of this paper is to introduce self-similarity as a fundamental property exhibited by the bursty traffic between on-chip modules in typical MPEG-2 video applications. Statistical tests performed on relevant traces extracted from common video clips establish unequivocally the existence of self-similarity in video traffic. Using a generic tile-based communication architecture, we discuss the implications of our findings on on-chip buffer space allocation and present quantitative evaluations for typical video streams. We also describe a technique for synthetically generating traces having statistical properties similar to those obtained from real video clips. Our proposed technique speeds up buffer simulations, allows media system designers to explore architectures rapidly and use large media data benchmarks more efficiently. We believe that our findings open new directions of research with deep implications on some fundamental issues in on-chip networks design for multimedia applications. Girish Varatkar, Radu Marculescu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2003 | Towards on-chip fault-tolerant communicationabstractAs CMOS technology scales down into the deep-submicron (DSM) domain, devices and interconnects are subject to new types of malfunctions and failures that are harder to predict and avoid with the current system-on-chip (SoC) design methodologies. Relaxing the requirement of 100% correctness in operation drastically reduces the costs of design but, at the same time, requires SoCs be designed with some degree of system-level fault-tolerance. In this paper, we introduce a high-level model of DSM failure patterns and propose a new communication paradigm for SoCs, namely stochastic communication. Specifically, for a generic tile-based architecture, we propose a randomized algorithm which not only separates computation from communication, but also provides the required fault-tolerance to on-chip failures. This new technique is easy and cheap to implement in SoCs that integrate a large number of communicating IP cores. Tudor Dumitras, Sam Kerner, Radu Marculescu |
ASP-DAC | 3 |
| 2003 | Energy-aware mapping for tile-based NoC architectures under performance constraintsabstractIn this paper, we present an algorithm which automatically maps the IPs/cores onto a generic regular Network on Chip (NoC) architecture such that the total communication energy is minimized. At the same time, the performance of the mapped system is guaranteed to satisfy the specified constraints through bandwidth reservation. As the main contribution, we first formulate the problem of energy-aware mapping, in a topological sense, and then propose an efficient branch-and-bound algorithm to solve it. Experimental results show that the proposed algorithm is very fast and robust, and significant energy savings can be achieved. For instance, for a complex video/audio SoC design, on average, 60.4% energy savings have been observed compared to an ad-hoc implementation. Jingcao Hu, Radu Marculescu |
ASP-DAC | 2 |
| 2003 | On-Chip Stochastic Communication
Tudor Dumitras, Radu Marculescu |
DATE | 2 |
| 2003 | Exploiting the Routing Flexibility for Energy/Performance Aware Mapping of Regular NoC Architectures
Jingcao Hu, Radu Marculescu |
DATE | 2 |
| 2003 | Ambient Intelligence Visions and Achievements: Linking Abstract Ideas to Real-World Concepts
Menno Lindwer, Diana Marculescu, Twan Basten, Rainer Zimmermann, Radu Marculescu, Stefan Jung, Eugenio Cantatore |
DATE | 5 |
| 2003 | Fault-Tolerant Techniques for Ambient Intelligent Distributed Systems
Diana Marculescu, Nicholas H. Zamora, Phillip Stanley-Marbell, Radu Marculescu |
ICCAD | 4 |
| 2003 | Communication-Aware Task Scheduling and Voltage Selection for Total Systems Energy Minimization
Girish Varatkar, Radu Marculescu |
ICCAD | 2 |
| 2003 | Designing Application Specific Networks-On-Chip: Five easy pieces
Radu Marculescu |
VLSI-SOC | 1 |
| 2003 | Electronic textiles: a platform for pervasive computing
Diana Marculescu, Radu Marculescu, Nicholas H. Zamora, Phillip Stanley-Marbell, Pradeep K. Khosla, Sungmee Park, Sundaresan Jayaraman, Stefan Jung, L. Weber, K. Cottet, Janus Grzyb, Gerhard Tröster, Mark T. Jones, Thomas Martin 0001, Zahi Nakad |
Proc. IEEE | 2 |
| 2003 | Electronic textiles: A platform for pervasive computingabstractThe invention of the Jacquard weaving machine led to the concept of a stored "program" and "mechanized" binary information processing. This development served as the inspiration for C. Babbage's analytical engine-the precursor to the modern-day computer. Today, more than 200 years later, the link between textiles and computing is more realistic than ever. In this paper, we look at the synergistic relationship between textiles and computing and identify the need for their "integration" using tools provided by an emerging new field of research that combines the strengths and capabilities of electronics and textiles into one: electronic textiles, or e-textiles. E-textiles, also called smart fabrics, have not only "wearable" capabilities like any other garment, but also have local monitoring and computation, as well as wireless communication capabilities. Sensors and simple computational elements are embedded in e-textiles, as well as built into yarns, with the goal of gathering sensitive information, monitoring vital statistics, and sending them remotely (possibly over a wireless channel) for further processing. The paper provides an overview of existing efforts and associated challenges in this area, while describing possible venues and opportunities for future research. Diana Marculescu, Radu Marculescu, Nicholas H. Zamora, Phillip Stanley-Marbell, Pradeep K. Khosla, Sungmee Park, Sundaresan Jayaraman, Stefan Jung, Christel Lauterbach, Werner Weber, Tünde Kirstein, Didier Cottet, Janus Grzyb, Gerhard Tröster, Mark T. Jones, Thomas Martin 0001, Zahi Nakad |
Proc. IEEE | 2 |
| 2003 | Modeling, Analysis, and Self-Management of Electronic TextilesabstractScaling in CMOS device technology has made it possible to cheaply embed intelligence in a myriad of devices. In particular, it has become feasible to fabricate flexible materials (e.g., woven fabrics) with large numbers of computing and communication elements embedded into them. Such computational fabrics, electronic textiles, or e-textiles have applications ranging from smart materials for aerospace applications to wearable computing. This paper addresses the modeling of computation, communication and failure in e-textiles and investigates the performance of two techniques, code migration and remote execution, for adapting applications executing over the hardware substrate, to failures in both devices and interconnection links. The investigation is carried out using a cycle-accurate simulation environment developed to model computation, power consumption, and node/link failures for large numbers of computing elements in configurable network topologies. A detailed analysis of the two techniques for adapting applications to the error prone substrate is presented, as well as a study of the effects of parameters, such as failure rates, communication speeds, and topologies, on the efficacy of the techniques and the performance of the system as a whole. It is shown that code migration and remote execution provide feasible methods for adapting applications to take advantage of redundancy in the presence of failures and involve trade offs in communication versus memory requirements in processing elements. Phillip Stanley-Marbell, Diana Marculescu, Radu Marculescu, Pradeep K. Khosla |
IEEE Trans. Computers | 3 |
| 2002 | Challenges and opportunities in electronic textiles modeling and optimizationabstractThis paper addresses an emerging new field of research that combines the strengths and capabilities of electronics and textiles in one: electronic textiles, or e-textiles. E-textiles, also called Smart Fabrics, have not only "wearable" capabilities like any other garment, but also local monitoring and computation, as well as wireless communication capabilities. Sensors and simple computational elements are embedded in e-textiles, as well as built into yarns, with the goal of gathering sensitive information, monitoring vital statistics and sending them remotely (possibly over a wireless channel) for further processing. Possible applications include medical (infant or patient) monitoring, personal information processing systems, or remote monitoring of deployed personnel in military or space applications. We illustrate the challenges imposed by the dual textile/electronics technology on their modeling and optimization methodology. Diana Marculescu, Radu Marculescu, Pradeep K. Khosla |
DAC | 2 |
| 2002 | Traffic analysis for on-chip networks design of multimedia applicationsabstractThe objective of this paper is to introduce self-similarity as a fundamental property exhibited by the bursty traffic between on-chip modules in typical MPEG-2 video applications. Statistical tests performed on relevant traces extracted from common video clips establish unequivocally the existence of self-similarity in video traffic. Using a generic communication architecture, we also discuss the implications of our findings on on-chip buffer space allocation and present quantitative evaluations for typical video streams. We believe that our findings open up new directions of research with deep implications on some fundamental issues in on-chip network design for multimedia applications. Girish Varatkar, Radu Marculescu |
DAC | 2 |
| 2002 | On-chip communication analysis for multimedia applicationsabstractThe objective of this paper is to introduce self-similarity as a fundamental property exhibited by the bursty traffic behavior between different on-chip modules in typical MPEG-2 video applications. Statistical tests performed on relevant traces extracted from common video clips establish unequivocally the existence of self-similarity in on-chip video traffic. Using a generic on-chip communication architecture, we discuss the implications of our findings on on-chip buffer space allocation. We also describe a synthetic trace generation procedure for speeding up the buffer simulation process. Girish Varatkar, Radu Marculescu |
ICME (2) | 2 |
| 2001 | System-Level Power/Performance Analysis for Embedded Systems DesignabstractThis paper presents a formal technique for system-level power/performance analysis that can help the designer to select the right platform starting from a set of target applications. By platform we mean a family of heterogeneous architectures that satisfy a set of architectural constraints imposed to allow re-use of hardware and software components. More precisely, we introduce the Stochastic Automata Networks (SANs) as an effective formalism for average-case analysis that can be used early in the design cycle to identify the best power/performance figure among several appli-cation-architecture combinations. This information not only helps avoid lengthy profiling simulations, but also enables efficient map-pings of the applications onto the chosen platform. We illustrate the features of our technique through the design of an MPEG-2 video decoder application. Amit Nandi, Radu Marculescu |
DAC | 2 |
| 2001 | Probabilistic application modeling for system-level perfromance analysisabstractThe objective of this paper is to introduce the Stochastic Automata Networks (SANs) as an effective formalism for application modeling in system-level analysis. More precisely we present a methodology for application modeling for system-level power/performance analysis that can help the designer to select the right platform and implement a set of target multimedia applications. We also show that, under various input traces, the steady-state behavior of the application itself is characterized by very different 'clusterings' of the probability distributions. Having this information available, not only helps to avoid lengthy profiling simulations for predicting power and performance figures, but also enables efficient mappings of the applications onto a chosen platform. We illustrate the benefits of our methodology using the MPEG-2 video decoder as the driver application. Radu Marculescu, Amit Nandi |
DATE | 1 |
| 2001 | System-Level Power/Performance Analysis of Portable Multimedia Systems Communicating over Wireless ChannelsabstractThis paper presents a new methodology for system-level power and performance analysis of wireless multimedia systems. More precisely, we introduce an analytical approach based on concurrent processes modeled as Stochastic Automata Networks (SANs) that can be effectively used to integrate power and performance metrics in system-level design. We show that 1) under various input traces and wireless channel conditions, the average-case behavior of a multimedia system consisting of a video encoder/decoder pair is characterized by very different probability distributions and power consumption values and 2) in order to identify the best trade-off between power and performance figures, one must take into consideration the entire environment (i.e., encoder, decoder and channel) for which the system is being designed. Compared to using simulation, our analytical technique reduces the time needed to find the steady-state behavior by orders of magnitude, with some limited loss in accuracy compared to the exact solution. We illustrate the potential of our methodology using the MPEG-2 video as the driver application. Radu Marculescu, Amit Nandi, Luciano Lavagno, Alberto L. Sangiovanni-Vincentelli |
ICCAD | 1 |
| 2000 | Improving simulation efficiency for circuit-level power estimation [CMOS]abstractIn this paper we present an effective technique for compacting a large sequence of input vectors into a much shorter one so as to reduce the circuit-level simulation time by orders of magnitude and maintain the accuracy of the power estimates. In particular, we model the effects of complex spatiotemporal correlations and rise/fall time slopes on total power dissipation. As the results demonstrate, large compaction ratios of orders of magnitude can be obtained without significant loss (about 5%, on average) in the accuracy of power estimates. Radu Marculescu, Cristinel Ababei |
ISCAS | 1 |
| 2000 | Stochastic sequential machine synthesis with application to constrained sequence generationabstractIn power estimation, one is faced with two problems: (1) generating input vector sequences that satisfy a given statistical behavior (in terms of signal probabilities and correlations among bits); (2) making these sequences as short as possible so as to improve the efficiency of power simulators. Stochastic sequential machines (SSMs) can be used to solve both problems. In particular, this paper presents a general procedure for SSM synthesis and describes a new framework for sequence characterization to match designers' needs for sequence generation or compaction. Experimental results demonstrate that compaction ratios of 1–3 orders of magnitude can be obtained without much loss in accuracy of total power estimates. Diana Marculescu, Radu Marculescu, Massoud Pedram |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2000 | Theoretical bounds for switching activity analysis in finite-state machinesabstractThe objective of this paper is to provide lower and upper bounds for the switching activity on the state lines in finite state machines (FSMs). Using a Markov chain model for the behavior of the FSM states, we derive theoretical bounds for the average Hamming distance on the state lines which are valid irrespective of the state encoding used in the final implementation. Such lower and upper bounds, in addition to providing a target for any state assignment algorithm, can also be used as parameters in a high-level power model and thus provide an early indication about the performance limits of the target FSM. Diana Marculescu, Radu Marculescu, Massoud Pedram |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1999 | Non-stationary effects in trace-driven power analysisabstractThe objective of this paper is to present an analytic technique for power analysis under non-stationary conditions. We use the transitive closure calculation to identify the transient component in the behavior of the target machine and then, based on the fundamental matrix and a symbolic approach (or support from simulation), we find the actual power distribution that corresponds to the transient regime. The present technique complements the current techniques (either for average or peak power estimation) to handle the case when transient effects exist and cannot be ignored. 1.1 Keywords power consumption, transient regime, Markov chains Radu Marculescu, Diana Marculescu, Massoud Pedram |
ISLPED | 1 |
| 1999 | Sequence compaction for power estimation: theory and practiceabstractPower estimation has become a critical step in the design of today's integrated circuits (ICs). Power dissipation is strongly input pattern dependent and, hence, to obtain accurate power values one has to simulate the circuit with a large number of vectors that typify the application data. The goal of this paper is to present an effective and robust technique for compacting large sequences of input vectors into much smaller ones such that the power estimates are as accurate as possible and the simulation time is reduced by orders of magnitude. Specifically, this paper introduces the hierarchical modeling of Markov chains as a flexible framework for capturing not only complex spatiotemporal correlations, but also dynamic changes in the sequence characteristics. In addition to this, we introduce and characterize a family of variable-order dynamic Markov models which provide an effective way for accurate modeling of external input sequences that affect the behavior of finite state machines. The new framework is very effective and has a high degree of adaptability. As the experimental results show, large compaction ratios of orders of magnitude can be obtained without significant loss in accuracy (less than 5% on average) for power estimates. Radu Marculescu, Diana Marculescu, Massoud Pedram |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1998 | Trace-Driven Steady-State Probability Estimation in FSMs with Application to Power EstimationabstractThis paper illustrates, analytically and quantitatively, the effect of high-order temporal correlations on steady-state and transition probabilities in finite state machines (FSMs). As the main theoretical contribution, we extend the previous work done on steady-state probability calculation in FSMs to account for complex spatiotemporal correlations which are present at the primary inputs when the target machine models real hardware and receives data from real applications. More precisely: (1) using the concept of constrained reachability analysis, the correct set of Chapman-Kolmogorov equations is constructed; and (2) based on stochastic complementation and iterative aggregation/disaggregation techniques, exact and approximate methods for finding the state occupancy probabilities in the target machine are presented. From a practical point of view, we show that assuming temporal independence or even using first-order temporal models is not sufficient due to the inaccuracies induced in steady-state and transition probability calculations. Experimental results show that, if the order of the source is underestimated, not only the set of reachable sets is incorrectly determined, but also the steady-state probability values can be more than 100% off from the correct ones. This strongly impacts the accuracy of the total power estimates that can be obtained via probabilistic approaches. Diana Marculescu, Radu Marculescu, Massoud Pedram |
DATE | 2 |
| 1998 | Theoretical bounds for switching activity analysis in finite-state machinesabstractThe objective of this paper is to provide lower and upper bounds for the switching activity on the state lines in Finite State Machines (FSMs). Using a Markov chain model for the behavior of the states of the FSM, we derive theoretical bounds for the average Hamming distance on the state lines which are valid irrespective of the state encoding used in the final implementation. Such lower and upper bounds, in addition to providing a target for any state assignment algorithm, can also be used as parameters in a high-level model of power, and thus provide an early indication about the performance limits of the target FSM. Experimental results obtained for the mcnc'91 benchmark suite show that our bounds are tighter than the bounds reported previously by other researchers and can be effectively used in a high-level power estimation framework. Diana Marculescu, Radu Marculescu, Massoud Pedram |
ISLPED | 2 |
| 1998 | Probabilistic modeling of dependencies during switching activity analysisabstractThis paper addresses, from a probabilistic point of view, the issue of switching activity estimation in combinational circuits under the zero-delay model. As the main theoretical contribution, we extend the previous work done on switching activity estimation to explicitly account for complex spatiotemporal correlations which occur at the primary inputs when the target circuit receives data from real applications. More precisely, using lag-one Markov chains, two new concepts-conditional independence and signal isotropy-are brought into attention and based on them, sufficient conditions for exact analysis of complex dependencies are given. From a practical point of view, it is shown that the relative error in calculating the switching activity of a logic gate using only pairwise probabilities can be upper-bounded. It is proved that the conditional independence problem is NP-complete and thus, relying on the concept of signal isotropy, approximate techniques with bounded error are proposed for estimating the switching activity. Evaluations of the model and a comparative analysis on benchmark circuits show that node-by-node switching activities are strongly pattern dependent and therefore, accounting for spatiotemporal dependencies is mandatory if accuracy is a major concern. Radu Marculescu, Diana Marculescu, Massoud Pedram |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1997 | Adaptive models for input data compaction for power simulatorsabstractPresents an effective and robust technique for compacting a large sequence of input vectors into a much smaller input sequence so as to reduce the circuit/gate-level simulation time by orders of magnitude and maintain the accuracy of the power estimates. In particular, this paper introduces and characterizes a family of dynamic Markov trees that can model complex the spatiotemporal correlations which occur during power estimation in both combinational and sequential circuits. As the results demonstrate, large compaction ratios of 1-2 orders of magnitude can be obtained without a significant loss (less than 5% on average) in the accuracy of the power estimates. Radu Marculescu, Diana Marculescu, Massoud Pedram |
ASP-DAC | 1 |
| 1997 | Sequence Compaction for Probabilistic Analysis of Finite-State MachinesabstractThe objective of this paper is to provide aneffective technique for accurate modeling of the externalinput sequences that affect the behavior of Finite StateMachines (FSMs). The proposed approach relies on adaptivemodeling of binary input streams as Markov sources of fixed-order.The input model itself is derived through a one-passtraversal of the input sequence and can be used to generatean equivalent sequence, much shorter in length compared tothe original sequence. The compacted sequence can besubsequently used with any available simulator to derive thesteady-state and transition probabilities, and the total powerconsumption in the target circuit. As the results demonstrate,large compaction ratios of orders of magnitude can beobtained without a significant loss (less than 3% on average)in the accuracy of estimated values. Diana Marculescu, Radu Marculescu, Massoud Pedram |
DAC | 2 |
| 1997 | Hierarchical Sequence Compaction for Power EstimationabstractAbstract- This paper presents an effective technique for compacting a large sequence of input vectors into a much smaller one such that when the two sequences are applied to any circuit, the resulting power dissipations are nearly the same. Specifically, this paper introduces the hierarchical modeling of Markov chains as a flexible framework for capturing not only complex spatiotemporal correlations, but also dynamic changes in the sequence characteristics. The new framework has a high degree of adaptability, i.e. the hierarchical model is dynamically grown according to the sequence behavior. Experimental results demonstrate that large compaction ratios can be obtained without significant loss in accuracy (less than 5 % on average) for power estimates. I. Radu Marculescu, Diana Marculescu, Massoud Pedram |
DAC | 1 |
| 1997 | Composite sequence compaction for finite-state machines using block entropy and high-order Markov modelsabstractThe objective of this paper is to provide an eflective technique for accurate modeling of the external input sequences that affect the behavior of Finite State Machines (FSMs). Based on the block entropy concept, we present a technique for identifying the order of variableorderMarkov sources of information.Furthermore, using dynamic Markov modeling, we propose an eflective approach to compact an initial sequence into a much shortel; equivalent one.The compacted sequence, can be subsequently used with any available simulator to derive the steady-state and transition probabilities, and the total power consumption in the target circuit.As the results demonstrate, large compaction ratios of orders of magnitude can be obtained without significant loss (less than 5% on average) in the accuracy of estimated values. Radu Marculescu, Diana Marculescu, Massoud Pedram |
ISLPED | 1 |
| 1996 | Stochastic Sequential Machine Synthesis Targeting Constrained Sequence GenerationabstractAbstract- The problem of stochastic sequential machines (SSM) synthesis is addressed and its relationship with the constrained sequence generation problem which arises during power estimation is discussed. In power estimation, one has to generate input vector sequences that satisfy a given statistical behavior (in terms of transition probabilities and correlations among bits) and/or to make these sequences as short as possible so as to improve the efficiency of power simulators. SSMs can be used to solve both problems. Based on Moore-type machines, a general procedure for SSM synthesis is revealed and a new framework for sequence characterization is built to match designer’s needs for sequence generation or compaction. As results demonstrate, compaction ratios of 1-2 orders of magnitude can be obtained without much loss in accuracy of total power estimates. I. Diana Marculescu, Radu Marculescu, Massoud Pedram |
DAC | 2 |
| 1996 | Improving the Efficiency of Power Simulators by Input Vector CompactionabstractAccurate power estimation is essential for low power digital CMOS circuit design.Power dissipation is input pattern dependent.To obtain an accurate power estimate, a large input vector set must be used which leads to very long simulation time.One solution is to generate a compact vector set that is representative of the original input vector set and can be simulated in a reasonable time.In this paper, we propose an input vector compaction technique that preserves the statistical properties of the original sequence.Experimental results show that a compaction ratio of 100X is achieved with less than 2% average error in the power estimates. Chi-Ying Tsui, Radu Marculescu, Diana Marculescu, Massoud Pedram |
DAC | 2 |
| 1996 | Information theoretic measures for power analysis [logic design]abstractThis paper considers the problem of estimating the power consumption at logic and register-transfer levels of design from an information theoretical point of view. In particular, it is demonstrated that the average switching activity in the circuit can be calculated using either entropy or informational energy averages. For control circuits and random logic, the output entropy (informational energy) per bit is calculated as a function of the input entropy (informational energy) per bit and an implementation dependent information scaling factor. For data-path circuits, the output entropy (informational energy) is calculated from the input entropy (informational energy) using a compositional technique which has linear complexity in terms of the circuit size. Finally, from these input and output values, the entropy (informational energy) per circuit line is calculated and used as an estimate for the average switching activity. The proposed switching activity estimation technique does not require simulation and is thus extremely fast, yet produces sufficiently accurate estimates. Diana Marculescu, Radu Marculescu, Massoud Pedram |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1995 | Efficient Power Estimation for Highly Correlated Input StreamsabstractAbstract- Power estimation in combinational modules is addressed from a probabilistic point of view. The zero-delay hypothesis is considered and under highly correlated input streams, the activities at the primary outputs and all internal nodes are estimated. For the first time, the relationship between logic and probabilistic domains is investigated and two new concepts- conditional independence and isotropy of signals- are brought into attention. Based on them, a sufficient condition for analyzing complex dependencies is given. In the most general case, the conditional independence problem has been shown to be NP-complete and thus appropriate heuristics are presented to estimate switching activity. Detailed experiments demonstrate the accuracy and efficiency of the method. The results reported here are useful in low power design. I. Radu Marculescu, Diana Marculescu, Massoud Pedram |
DAC | 1 |
| 1994 | Switching activity analysis considering spatiotemporal correlations
Radu Marculescu, Diana Marculescu, Massoud Pedram |
ICCAD | 1 |
| 1993 | Worst-case analysis for pseudorandom testingabstractThe testing problem of combinational circuits with pseudorandom patterns is investigated. Work by previous workers for single faults is extended to multiple faults situations; in addition, masking effects between disjoint/conjoint faults are considered. An analytical model based on Markov chains with any number of states is proposed for random/pseudorandom testing and relationships between test length and test confidence are developed. Evaluations of the model for double and triple faults are presented using well-known examples. The results presented in this paper are useful for BIST systems that use random/pseudorandom input patterns.> Radu Marculescu |
VTS | 1 |