EDBT 2026 Demo / reviewers in the wild / expert
Paul Bogdan
dblp:05/5539
· DBLP profile ↗
84ranked-venue papers
16as first author
27since 2021 · last 2026
0000-0003-2118-0816ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 52 · 13 first-author · 2 since 2021Artificial intelligence and machine learning · 18 · 17 since 2021Software engineering, systems software and programming languages · 8 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Computer networks · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Security and privacy · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Uncertainty Quantification in LLM Agents: Foundations, Emerging Challenges, and OpportunitiesabstractChangdae Oh, Seongheon Park, To Eun Kim, Jiatong Li, Wendi Li, Samuel Yeh, Sean Du, Hamed Hassani, Paul Bogdan, Dawn Song, Sharon Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Changdae Oh, Seongheon Park, To Eun Kim, Wendi Li, Samuel Yeh 0001, Sean Du, Seyed Hamed Hassani, Paul Bogdan, Dawn Song, Yixuan Li 0001 |
ACL (1) | 9 |
| 2025 | Controllable Generative Model for Brain EvolutionabstractToday’s generative models can synthesize magnetic resonance images (MRIs) of the brain at specific ages. However, such models can neither map the aging process longitudinally within subjects, nor accommodate its variability across subjects. Such approaches also cannot predict anatomic features of aging in ways that can be validated retrospectively or trusted prospectively. We introduce a three-dimensional hybrid ControlNet + diffusion model that uses the baseline T1-weighted MRIs of healthy adults to predict individual neuroanatomic aging trajectories, as reflected by follow-up MRIs. The approach captures individual anatomical changes with an average predicted voxelwise intensity error of 15% and structural similarity index of 93%. Unlike methods relying on qualitative validation, our approach quantifies the fidelity of prospective MRI synthesis using FreeSurfer volumetrics. Because brain atrophy reflects risk for Alzheimer’s disease (AD), our model’s ability to generate individual-specific prospective MRIs suggests its clinical potential to assist AD risk estimation. Gengshuo Liu, Nikhil N. Chaudhari, Nikos Kanakaris, Chenzhong Yin, Paul Bogdan, Andrei Irimia |
ICASSP | 5 |
| 2025 | Exploiting Application-to-Architecture Dependencies for Designing Scalable OSabstractWith the advent of hundreds of cores on a chip to accelerate applications, the operating system (OS) needs to exploit the existing parallelism provided by the underlying hardware resources to determine the right amount of processes to be mapped on the multi-core systems. However, the existing OS is not scalable and is oblivious to applications. We address these issues by adopting a multi-layer network representation of the dynamic application-to-OS-to-architecture dependencies, namely the NetworkedOS. We adopt a compile-time analysis and construct a network representing the dependencies between dynamic instructions translated from the applications and the kernel and services. We propose an overlapping partitioning scheme to detect the clusters or processes that can potentially run in parallel to be mapped onto cores while reducing the number of messages transferred. At run time, processes are mapped onto the multi-core systems, taking into consideration the process affinity. Our experimental results indicate that NetworkedOS achieves performance improvement as high as 7.11x compared to Linux running on a 128-core system and 2.01x to Barrelfish running on a 64-core system. Nikos Kanakaris, Anzhe Cheng, Chenzhong Yin, Nesreen K. Ahmed, Shahin Nazarian, Andrei Irimia, Paul Bogdan |
ICASSP | 8 |
| 2025 | Edge Probability Graph Models Beyond Edge Independency: Concepts, Analyses, and AlgorithmsabstractDesirable random graph models (RGMs) should (i) reproduce common patterns in real-world graphs (e.g., powerlaw degrees, small diameters, and high clustering), (ii) generate variable (i.e., not overly similar) graphs, and (iii) remain tractable to compute and control graph statistics. A common class of RGMs (e.g., Erdős-Rényi and stochastic Kronecker) outputs edge probabilities, so we need to realize (i.e., sample from) the output edge probabilities to generate graphs. Typically, the existence of each edge is assumed to be determined independently, for simplicity and tractability. However, with edge independency, RGMs provably cannot produce high subgraph densities and high output variability simultaneously. In this work, we explore RGMs beyond edge independence that can better reproduce common patterns while maintaining high tractability and variability. Theoretically, we propose an edge-dependent realization (i.e., sampling) framework called binding that provably preserves output variability, and derive closed-form tractability results on subgraph (e.g., triangle) densities. Practically, we propose algorithms for graph generation with binding and parameter fitting of binding. Our empirical results demonstrate that RGMs with binding exhibit high tractability and well reproduce common patterns, significantly improving upon edge-independent RGMs. Fanchen Bu, Ruochen Yang, Paul Bogdan, Kijung Shin |
ICDM | 3 |
| 2025 | Neuron-based Multifractal Analysis of Neuron Interaction Dynamics in Large ModelsabstractIn recent years, there has been increasing attention on the capabilities of large-scale models, particularly in handling complex tasks that small-scale models are unable to perform. Notably, large language models (LLMs) have demonstrated ``intelligent'' abilities such as complex reasoning and abstract language comprehension, reflecting cognitive-like behaviors. However, current research on emergent abilities in large models predominantly focuses on the relationship between model performance and size, leaving a significant gap in the systematic quantitative analysis of the internal structures and mechanisms driving these emergent abilities. Drawing inspiration from neuroscience research on brain network structure and self-organization, we propose (i) a general network representation of large models, (ii) a new analytical framework — *Neuron-based Multifractal Analysis (NeuroMFA)* - for structural analysis, and (iii) a novel structure-based metric as a proxy for emergent abilities of large models. By linking structural features to the capabilities of large models, *NeuroMFA* provides a quantitative framework for analyzing emergent phenomena in large models. Our experiments show that the proposed method yields a comprehensive measure of the network's evolving heterogeneity and organization, offering theoretical foundations and a new perspective for investigating emergence in large models. Xiongye Xiao, Heng Ping, Defu Cao, Yaxing Li, Yizhuo Zhou, Nikos Kanakaris, Paul Bogdan |
ICLR | 9 |
| 2025 | End-to-End Learning Framework for Solving Non-Markovian Optimal ControlabstractInteger-order calculus fails to capture the long-range dependence (LRD) and memory effects found in many complex systems. Fractional calculus addresses these gaps through fractional-order integrals and derivatives, but fractional-order dynamical systems pose substantial challenges in system identification and optimal control tasks. In this paper, we theoretically derive the optimal control via linear quadratic regulator (LQR) for fractional-order linear time-invariant (FOLTI) systems and develop an end-to-end deep learning framework based on this theoretical foundation. Our approach establishes a rigorous mathematical model, derives analytical solutions, and incorporates deep learning to achieve data-driven optimal control of FOLTI systems. Our key contributions include: (i) proposing a novel method for system identification and optimal control strategy in FOLTI systems, (ii) developing the first end-to-end data-driven learning framework, Fractional-Order Learning for Optimal Control (FOLOC), that learns control policies from observed trajectories, and (iii) deriving theoretical bounds on the sample complexity for learning accurate control policies under fractional-order dynamics. Experimental results indicate that our method accurately approximates fractional-order system behaviors without relying on Gaussian noise assumptions, pointing to promising avenues for advanced optimal control. Xiaole Zhang, Peiyu Zhang 0002, Xiongye Xiao, Vasileios Tzoumas, Vijay Gupta 0001, Paul Bogdan |
ICML | 7 |
| 2025 | MIHC: Multi-View Interpretable Hypergraph Neural Networks with Information Bottleneck for Chip Congestion PredictionabstractWith AI advancement and increasing circuit complexity, efficient chip design through Electronic Design Automation (EDA) is critical. Fast and accurate congestion prediction in chip layout and routing can significantly enhance automated design performance. Existing congestion modeling methods are limited by **(i)** ineffective processing and fusion of multi-view circuit data information, and **(ii)** insufficient reliability and interpretability in the prediction process. To address these challenges, We propose **M**ulti-view **I**nterpretable **H**ypergraph for **C**hip (**MIHC**), a trustworthy 'multi-view hypergraph neural network'-based framework that **(i)** processes both graph and image information in unified hypergraph representations, capturing topological and geometric circuit data, and **(ii)** implements a novel subgraph Information Bottleneck mechanism identifying critical congestion-correlated regions to guide predictions. This represents the first attempt to incorporate such interpretability into congestion prediction through informative graph reasoning. Experiments show our model reduces NMAE by 16.67% and 8.57% in cell-based and grid-based predictions on ISPD2015, and 5.26% and 2.44% on CircuitNet-N28, respectively, compared to state-of-the-art methods. Rigorous cross-design generalization experiments further validate our method’s capability to handle entirely unseen circuit designs. Zeyue Zhang, Heng Ping, Peiyu Zhang 0002, Nikos Kanakaris, Xiaoling Lu, Paul Bogdan, Xiongye Xiao |
NeurIPS | 6 |
| 2024 | Unlocking Deep Learning: A BP-Free Approach for Parallel Block-Wise Training of Neural NetworksabstractBackpropagation (BP) has been a successful optimization technique for deep learning models. However, its limitations, such as backward- and update-locking, and its biological implausibility, hinder the concurrent updating of layers and do not mimic the local learning processes observed in the human brain. To address these issues, recent research has suggested using local error signals to asynchronously train network blocks. However, this approach often involves extensive trial-and-error iterations to determine the best configuration for local training. This includes decisions on how to decouple network blocks and which auxiliary networks to use for each block. In our work, we introduce a novel BP-free approach: a block-wise BP-free (BWBPF) neural network that leverages local error signals to optimize distinct sub-neural networks separately, where the global loss is only responsible for updating the output layer. The local error signals used in the BP-free model can be computed in parallel, enabling a potential speed-up in the weight update process through parallel implementation. Our experimental results consistently show that this approach can identify transferable decoupled architectures for VGG and ResNet variations, outperforming models trained with end-to-end backpropagation and other state-of-the-art block-wise learning techniques on datasets such as CIFAR-10 and Tiny-ImageNet. The code is released at https://github.com/Belis0811/BWBPF. Anzhe Cheng, Heng Ping, Zhenkun Wang 0008, Xiongye Xiao, Chenzhong Yin, Shahin Nazarian, Mingxi Cheng, Paul Bogdan |
ICASSP | 8 |
| 2024 | Discovering Malicious Signatures in Software from Structural InteractionsabstractMalware represents a significant security concern in today’s digital landscape, as it can destroy or disable operating systems, steal sensitive user information, and occupy valuable disk space. However, current malware detection methods, such as static-based and dynamic-based approaches, struggle to identify newly developed ("zero-day") malware and are limited by customized virtual machine (VM) environments. To overcome these limitations, we propose a novel malware detection approach that leverages deep learning, mathematical techniques, and network science. Our approach focuses on static and dynamic analysis and utilizes the Low-Level Virtual Machine (LLVM) to profile applications within a complex network. The generated network topologies are input into the GraphSAGE architecture to efficiently distinguish between benign and malicious software applications, with the operation names denoted as node features. Importantly, the GraphSAGE models analyze the network’s topological geometry to make predictions, enabling them to detect state-of-the-art malware and prevent potential damage during execution in a VM. To evaluate our approach, we conduct a study on a dataset comprising source code from 24,376 applications, specifically written in C/C++, sourced directly from widely-recognized malware and various types of benign software. The results show a high detection performance with an Area Under the Receiver Operating Characteristic Curve (AUROC) of 99.85%. Our approach marks a substantial improvement in malware detection, providing a notably more accurate and efficient solution when compared to current state-of-the-art malware detection methods. The code is released at https://github.com/HantangZhang/MGN. Chenzhong Yin, Hantang Zhang, Mingxi Cheng, Xiongye Xiao, Xinghe Chen, Paul Bogdan |
ICASSP | 7 |
| 2024 | Neuro-Inspired Information-Theoretic Hierarchical Perception for Multimodal LearningabstractIntegrating and processing information from various sources or modalities are critical for obtaining a comprehensive and accurate perception of the real world in autonomous systems and cyber-physical systems. Drawing inspiration from neuroscience, we develop the Information-Theoretic Hierarchical Perception (ITHP) model, which utilizes the concept of information bottleneck. Different from most traditional fusion models that incorporate all modalities identically in neural networks, our model designates a prime modality and regards the remaining modalities as detectors in the information pathway, serving to distill the flow of information. Our proposed perception model focuses on constructing an effective and compact information flow by achieving a balance between the minimization of mutual information between the latent state and the input modal state, and the maximization of mutual information between the latent states and the remaining modal states. This approach leads to compact latent state representations that retain relevant information while minimizing redundancy, thereby substantially enhancing the performance of multimodal representation learning. Experimental evaluations on the MUStARD, CMU-MOSI, and CMU-MOSEI datasets demonstrate that our model consistently distills crucial information in multimodal learning scenarios, outperforming state-of-the-art benchmarks. Remarkably, on the CMU-MOSI dataset, ITHP surpasses human-level performance in the multimodal sentiment binary classification task across all evaluation metrics (i.e., Binary Accuracy, F1 Score, Mean Absolute Error, and Pearson Correlation). Xiongye Xiao, Gengshuo Liu, Defu Cao, Yaxing Li, Tianqing Fang, Mingxi Cheng, Paul Bogdan |
ICLR | 9 |
| 2024 | A Structure-Aware Framework for Learning Device Placements on Computation GraphsabstractComputation graphs are Directed Acyclic Graphs (DAGs) where the nodes correspond to mathematical operations and are used widely as abstractions in optimizations of neural networks. The device placement problem aims to identify optimal allocations of those nodes to a set of (potentially heterogeneous) devices. Existing approaches rely on two types of architectures known as grouper-placer and encoder-placer, respectively. In this work, we bridge the gap between encoder-placer and grouper-placer techniques and propose a novel framework for the task of device placement, relying on smaller computation graphs extracted from the OpenVINO toolkit. The framework consists of five steps, including graph coarsening, node representation learning and policy optimization. It facilitates end-to-end training and takes into account the DAG nature of the computation graphs. We also propose a model variant, inspired by graph parsing networks and complex network analysis, enabling graph representation learning and jointed, personalized graph partitioning, using an unspecified number of groups. To train the entire framework, we use reinforcement learning using the execution time of the placement as a reward. We demonstrate the flexibility and effectiveness of our approach through multiple experiments with three benchmark models, namely Inception-V3, ResNet, and BERT. The robustness of the proposed framework is also highlighted through an ablation study. The suggested placements improve the inference speed for the benchmark models by up to $58.2\%$ over CPU execution and by up to $60.24\%$ compared to other commonly used baselines. Shukai Duan 0002, Heng Ping, Nikos Kanakaris, Xiongye Xiao, Panagiotis Kyriakis, Nesreen K. Ahmed, Peiyu Zhang 0002, Guixiang Ma, Mihai Capota, Shahin Nazarian, Theodore L. Willke, Paul Bogdan |
NeurIPS | 12 |
| 2024 | Phasor Measurement Unit Change-Point Detection of Frequency Hurst Exponent Anomaly With Time-to-EventabstractThe objective of this article is real-time detection of a change-point in the baseline distribution of the frequency signal generated by Phasor Measurement Units (PMUs) that could indicate potential for voltage collapse, false data injection, or other security threats. The alarm flags an anomaly event when the cumulative sum (CUSUM) of the log-likelihood departures of the Hurst exponent of the frequency from its baseline statistic exceeds a threshold set up as a compromise between the conflicting objectives of minimum detection delay and acceptable False Alarm Rate. As main theoretical contribution, an extra protection layer is developed that provides an estimate of the time to the alarm event, giving the Transmission System Operator (TSO) or an autonomous agent more time to deploy proactive measures to avoid large catastrophic system states. The proposed method is illustrated by a retrospective analysis of the 2012 Indian blackout that reveals that 10-12 minutes before the voltage collapse, a significant increase in the Hurst exponent of the frequency could have been detected. Conceptually, this shows that PMUs provide significant value towards ensuring autonomous cognition in the power grid by enabling the abnormal events to be fault-detected in an unsupervised fashion. Jayson Sia, Edmond A. Jonckheere, Laith Shalalfeh, Paul Bogdan |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2023 | Raising The Limit of Image Rescaling Using Auxiliary EncodingabstractNormalizing flow models using invertible neural networks (INN) have been widely investigated for successful generative image super-resolution (SR) by learning the transformation between the normal distribution of latent variable z and the conditional distribution of high-resolution (HR) images gave a low-resolution (LR) input. Recently, image rescaling models like IRN utilize the bidirectional nature of INN to push the performance limit of image upscaling by optimizing the downscaling and upscaling steps jointly. While the random sampling of latent variable z is useful in generating diverse photo-realistic images, it is not desirable for image rescaling when accurate restoration of the HR image is more important. Hence, in places of random sampling of z, we propose auxiliary encoding modules to further push the limit of image rescaling performance. Two options to store the encoded latent variables in downscaled LR images, both readily supported in existing image file format, are proposed. One is saved as the alpha-channel, the other is saved as meta-data in the image header, and the corresponding modules are denoted as suffixes -A and -M respectively. Optimal network architectural changes are investigated for both options to demonstrate their effectiveness in raising the rescaling performance limit on different baseline models including IRN and DLV-IRN. Chenzhong Yin, Zhihong Pan 0001, Xin Zhou 0017, Paul Bogdan |
ICASSP | 5 |
| 2023 | Coupled Multiwavelet Operator Learning for Coupled Differential Equations
Xiongye Xiao, Defu Cao, Ruochen Yang, Gengshuo Liu, Chenzhong Yin, Radu Balan, Paul Bogdan |
ICLR | 8 |
| 2023 | Generative Decoding of Visual StimuliabstractReconstructing natural images from fMRI recordings is a challenging task of great importance in neuroscience. The current architectures are bottlenecked because they fail to effectively capture the hierarchical processing of visual stimuli that takes place in the human brain. Motivated by that fact, we introduce a novel neural network architecture for the problem of neural decoding. Our architecture uses Hierarchical Variational Autoencoders (HVAEs) to learn meaningful representations of natural images and leverages their latent space hierarchy to learn voxel-to-image mappings. By mapping the early stages of the visual pathway to the first set of latent variables and the higher visual cortex areas to the deeper layers in the latent hierarchy, we are able to construct a latent variable neural decoding model that replicates the hierarchical visual information processing. Our model achieves better reconstructions compared to the state of the art and our ablation study indicates that the hierarchical structure of the latent space is responsible for that performance. Eleni Miliotou, Panagiotis Kyriakis, Jason D. Hinman, Andrei Irimia, Paul Bogdan |
ICML | 5 |
| 2023 | Robust Learning Under Label Noise by Optimizing the Tails of the Loss DistributionabstractWe introduce a novel optimization strategy that combines empirical risk minimization (ERM) with a sample weighting scheme to decrease the bias induced by training with noisy labels. We accomplish this by penalizing the tails of the loss distribution, as samples associated with the right tail (high loss values) are more likely to be mislabeled. We implement the proposed sample weighting scheme by minimizing a discriminatory risk that downweighs mislabeled samples. Moreover, we show that the proposed method enables us to optimize the location and shape of the loss distribution simultaneously. In addition, we couple our method with a distributionally robust optimization stage. The goal of this coupling is to increase the weight of correctly labeled but hard-to-Iearn samples that exhibit a high loss value. This coupling also improves model accuracy on underrepresented minority subsets when training with imbalanced classes. Experimental results show that for 60% noise contamination, our method, when compared to the ERM, improves the final accuracy on both Fashion-MNIST and CIFAR-10 by 18.4% and 8.5%, respectively. Valeriu Balaban, Jayson Sia, Paul Bogdan |
ICMLA | 3 |
| 2022 | Non-Linear Operator Approximations for Initial Value Problems
Xiongye Xiao, Radu Balan, Paul Bogdan |
ICLR | 4 |
| 2022 | Pareto Policy Adaptation
Panagiotis Kyriakis, Jyotirmoy V. Deshmukh, Paul Bogdan |
ICLR | 3 |
| 2022 | Improving Robustness: When and How to Minimize or Maximize the Loss VarianceabstractWe introduce distributional variance penalization, a strategy for learning with limited and/or mislabeled data. While minimizing the loss function currently stands as the training objective for many machine learning applications, it suffers from poor robustness. In this paper, we show that we can improve upon robustness issues by minimizing the average loss along with penalizing the variance. In particular, we expand on past studies of directly penalizing the variance which adjusts the weights of individual samples, resulting in improved robustness. However, the weights can take negative values and lead to unstable behavior. We introduce distributional variance penalization, which solves the issue of negative weights. Distributional variance penalization minimizes the expectation with respect to a distinct distribution that achieves a similar weighting scheme as direct variance penalization. We study the impact of both positive and negative variance penalization in the context of classification, and show that the generalization and the robustness against mislabeled data can be improved for a broad class of loss functions. Experimental results show that test accuracy improves by up to 20% compared to ERM when training with limited data or mislabeled data. Valeriu Balaban, Hoda Bidkhori, Paul Bogdan |
ICMLA | 3 |
| 2022 | Secure Distributed/Federated Learning: Prediction-Privacy Trade-Off for Multi-Agent SystemabstractDecentralized learning is an efficient emerging paradigm for boosting the computing capability of multiple bounded computing agents. In the big data era, performing inference within the distributed and federated learning (DL and FL) frameworks, the central server needs to process a large amount of data while relying on various agents to perform multiple distributed training tasks. Considering the decentralized computing topology, privacy has become a first-class concern. Moreover, assuming limited information processing capability for the agents calls for a sophisticated privacy-preserving decentralization that ensures efficient computation. Towards this end, we study the privacy-aware server to multi-agent assignment problem subject to information processing constraints associated with each agent, while maintaining the privacy and assuring learning informative messages received by agents about a global terminal through the distributed private federated learning (DPFL) approach. To find a decentralized scheme for a two-agent system, we formulate an optimization problem that balances privacy and accuracy, taking into account the quality of compression constraints associated with each agent. We propose an iterative converging algorithm by alternating over self-consistent equations. We also numerically evaluate the proposed solution to show the privacy-prediction trade-off and demonstrate the efficacy of the novel approach in ensuring privacy in DL and FL. Mohamed Ridha Znaidi, Paul Bogdan |
ISIT | 3 |
| 2022 | Introduction to the Special Issue on Internet-of-Medical-ThingsabstractNo abstract available. Paul Bogdan, Radu Grosu, Insup Lee 0001 |
ACM Trans. Comput. Heal. | 1 |
| 2022 | A Design Methodology for Energy-Aware Processing in Unmanned Aerial VehiclesabstractUnmanned Aerial Vehicles (UAVs) have rapidly become popular for monitoring, delivery, and actuation in many application domains such as environmental management, disaster mitigation, homeland security, energy, transportation, and manufacturing. However, the UAV perception and navigation intelligence (PNI) designs are still in their infancy and demand fundamental performance and energy optimizations to be eligible for mass adoption. In this article, we present a generalizable three-stage optimization framework for PNI systems that (i) abstracts the high-level programs representing the perception, mining, processing, and decision making of UAVs into complex weighted networks tracking the interdependencies between universal low-level intermediate representations; (ii) exploits a differential geometry approach to schedule and map the discovered PNI tasks onto an underlying manycore architecture. To mine the complexity of optimal parallelization of perception and decision modules in UAVs, this proposed design methodology relies on an Ollivier-Ricci curvature-based load-balancing strategy that detects the parallel communities of the PNI applications for maximum parallel execution, while minimizing the inter-core communication; and (iii) relies on an energy-aware mapping scheme to minimize the energy dissipation when assigning the communities onto tile-based networks-on-chip. We validate this approach based on various drone PNI designs including flight controller, path planning, and visual navigation. The experimental results confirm that the proposed framework achieves 23% flight time reduction and up to 34% energy savings for the flight controller application. In addition, the optimization on a 16-core platform improves the on-time visit rate of the path planning algorithm by 14% while reducing 81% of run time for ConvNet visual navigation. Jingyu He, Ioana Corina Bogdan, Shahin Nazarian, Paul Bogdan |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2021 | Learning Code Representations Using Multifractal-based Graph NetworksabstractLearning representations of software codes is a critical problem for a wide range of system applications, e.g., compiler optimization, software classification, malicious software detection, and performance optimization. Recently, learning graph-based representations of software programs has been used to model the inherent structural dependencies in programming languages (e.g., C++, Python). In this paper, we propose a novel graph neural network framework that utilizes multifractal analysis for LLVM intermediate representations (IR). We then show empirically that the proposed framework is capable of capturing long-range structural dependencies that appear in software codes. We conduct experiments and comparisons on two downstream system applications: (1) predicting heterogeneous compute device mappings (graph classification), and (2) compiler reachability analysis (node classification). We observe that introducing a structural inductive bias through multifractal topological features enables GNNs to capture long-range dependencies among nodes, thus, it improves the accuracy of GNN models for applications that require learning code representations. Guixiang Ma, Mihai Capota, Theodore L. Willke, Shahin Nazarian, Paul Bogdan, Nesreen K. Ahmed |
IEEE BigData | 6 |
| 2021 | Learning Hyperbolic Representations of Topological Features
Panagiotis Kyriakis, Iordanis Fostiropoulos, Paul Bogdan |
ICLR | 3 |
| 2021 | Trust-aware Control for Intelligent Transportation SystemsabstractMany intelligent transportation systems are multiagent systems, i.e., both the traffic participants and the subsystems within the transportation infrastructure can be modeled as interacting agents. The use of AI-based methods to achieve coordination among the different agents systems can provide greater safety over transportation systems containing only human-operated vehicles, and also improve the system efficiency in terms of traffic throughput, sensing range, and enabling collaborative tasks. However, increased autonomy makes the transportation infrastructure vulnerable to compromised vehicular agents or infrastructure. This paper proposes a new framework by embedding the trust authority into transportation infrastructure to systematically quantify the trustworthiness of agents using an epistemic logic known as subjective logic. In this paper, we make the following novel contributions: (i) We propose a framework for using the quantified trustworthiness of agents to enable trust-aware coordination and control. (ii) We demonstrate how to synthesize trust-aware controllers using an approach based on reinforcement learning. (iii) We comprehensively analyze an autonomous intersection management (AIM) case study and develop a trust-aware version called AIM-Trust that leads to lower accident rates in scenarios consisting of a mixture of trusted and untrusted agents. Mingxi Cheng, Junyao Zhang 0003, Shahin Nazarian, Jyotirmoy V. Deshmukh, Paul Bogdan |
IV | 5 |
| 2021 | Multiwavelet-based Operator Learning for Differential EquationsabstractThe solution of a partial differential equation can be obtained by computing the inverse operator map between the input and the solution space. Towards this end, we introduce a $\textit{multiwavelet-based neural operator learning scheme}$ that compresses the associated operator's kernel using fine-grained wavelets. By explicitly embedding the inverse multiwavelet filters, we learn the projection of the kernel onto fixed multiwavelet polynomial bases. The projected kernel is trained at multiple scales derived from using repeated computation of multiwavelet transform. This allows learning the complex dependencies at various scales and results in a resolution-independent scheme. Compare to the prior works, we exploit the fundamental properties of the operator's kernel which enable numerically efficient representation. We perform experiments on the Korteweg-de Vries (KdV) equation, Burgers' equation, Darcy Flow, and Navier-Stokes equation. Compared with the existing neural operator approaches, our model shows significantly higher accuracy and achieves state-of-the-art in a range of datasets. For the time-varying equations, the proposed method exhibits a ($2X-10X$) improvement ($0.0018$ ($0.0033$) relative $L2$ error for Burgers' (KdV) equation). By learning the mappings between function spaces, the proposed method has the ability to find the solution of a high-resolution input after learning from lower-resolution data. Xiongye Xiao, Paul Bogdan |
NeurIPS | 3 |
| 2021 | Plasticity-on-Chip Design: Exploiting Self-Similarity for Data CommunicationsabstractWith the increasing demand for distributed big data analytics and data-intensive programs which contribute to large volumes of packets among processing elements (PEs) and memory banks, we witness a pressing need for new mathematical models and algorithms that can engineer a brain-inspired plasticity into the computing platforms by mining the topological complexity of high-level programs (HLPs) and exploiting their self-similar and fractal characteristics for designing reconfigurable domain-specific computing architectures. In this article, we present Plasticity-on-Chip (PoC) by engineering plasticity into ”artificial brains” to mine and exploit the self-similarity of HLPs. First, we present a communication modeling of HLPs (e.g., C/C++ implementations of various applications) that relies on static and dynamic compiler analysis of programs with varying input seeds, performing comprehensive program analysis of all traces, and representing the HLPs as weighted directed acyclic graphs while capturing the intrinsic timing constraints and data/control flow requirements. Second, we propose a rigorous mathematical framework for determining the optimal parallel degree of executing a set of interacting HLPs (by partitioning them into clusters of densely interconnected supernodes - tasks) which helps us decide the number of available heterogeneous PEs, the amount of required memory and the structure of the synthesized deadlock-free irregular NoC topology that offers an efficient communication medium. These clusters serve as abstract models of computation for the synthesized PEs within the parallel execution model. Finally, exploiting the fractal and complex networks concepts, we extract in-depth features from graphs that serve as inputs for distributed reinforcement learning. Our experimental results on synthesized PEs and NoCs show performance improvements as high as 7.61x when compared to the traditional NoC and 2.6x compared to gem5-Aladdin. Shahin Nazarian, Paul Bogdan |
IEEE Trans. Computers | 3 |
| 2020 | S4oC: A Self-Optimizing, Self-Adapting Secure System-on-Chip Design Framework to Tackle Unknown Threats - A Network Theoretic, Learning ApproachabstractWe propose a framework for the design and optimization of a secure self-optimizing, self-adapting system-on-chip (S4oC) architecture. The goal is to minimize the impact of attacks such as hardware Trojan and side-channel, by making real-time adjustments. S4oC learns to reconfigure itself, subject to various security measures and attacks, some of which possibly unknown at design time. Furthermore, the data types and patterns of the target applications, environmental conditions, and sources of variations are incorporated. S4oC is a manycore system, modeled as a four-layer graph, representing the model of computation (MoCp), model of connection (MoCn), model of memory (MoM) and model of storage (MoS), with a large number of elements including heterogeneous reconfigurable processing elements in MoCp, and memory elements in the MoM layer. Security driven community detection, and neural networks are utilized for application task clustering, and distributed reinforcement learning (RL) for task mapping. Shahin Nazarian, Paul Bogdan |
ISCAS | 2 |
| 2020 | An Efficient Task Mapping for Manycore SystemsabstractSystem-on-chip (SoC) has migrated from single core to manycore architectures to cope with the increasing complexity of real-life applications. Application task mapping has a significant impact on the efficiency of manycore system (MCS) computation and communication. We present WAANSO, a scalable framework that incorporates a Wavelet Clustering based approach to cluster application tasks. We also introduce Ant Swarm Optimization (ASO) based on iterative execution of Ant Colony Optimization (ACO) and Particle Swarm Optimization (PSO) for task clustering and mapping to the MCS processing elements. We have shown that WAANSO can significantly increase the MCS energy and performance efficiencies. Based on our experiments on a 64-core system, WAANSO improves energy efficiency by 19%, compared to baseline approaches, namely DPSO, ACO and branch and bound (B&B). Additionally, the performance improves by 65.86% compared to Density-Based Spatial Clustering of Applications with Noise (DBSCAN) baseline. Xiqian Wang, Jiajin Xi, Yinghao Wang, Paul Bogdan, Shahin Nazarian |
ISCAS | 4 |
| 2020 | VRoC: Variational Autoencoder-aided Multi-task Rumor Classifier Based on TextabstractSocial media became popular and percolated almost all aspects of our daily lives. While online posting proves very convenient for individual users, it also fosters fast-spreading of various rumors. The rapid and wide percolation of rumors can cause persistent adverse or detrimental impacts. Therefore, researchers invest great efforts on reducing the negative impacts of rumors. Towards this end, the rumor classification system aims to to detect, track, and verify rumors in social media. Such systems typically include four components: (i) a rumor detector, (ii) a rumor tracker, (iii) a stance classifier, and (iv) a veracity classifier. In order to improve the state-of-the-art in rumor detection, tracking, and verification, we propose VRoC, a tweet-level variational autoencoder-based rumor classification system. VRoC consists of a co-train engine that trains variational autoencoders (VAEs) and rumor classification components. The co-train engine helps the VAEs to tune their latent representations to be classifier-friendly. We also show that VRoC is able to classify unseen rumors with high levels of accuracy. For the PHEME dataset, VRoC consistently outperforms several state-of-the-art techniques, on both observed and unobserved rumors, by up to 26.9%, in terms of macro-F1 scores. Mingxi Cheng, Shahin Nazarian, Paul Bogdan |
WWW | 3 |
| 2020 | H₂O-Cloud: A Resource and Quality of Service-Aware Task Scheduling Framework for Warehouse-Scale Data CentersabstractCloud computing has attracted both end-users and cloud service providers (CSPs) in recent years. Improving resource utilization rate (RUtR), such as CPU and memory usages on servers, while maintaining quality of service (QoS) is one key challenge faced by CSPs with warehouse-scale datacenters. Prior works proposed various algorithms to reduce energy cost or to improve RUtR, which either lack the fine-grained task scheduling capabilities, or fail to take a comprehensive system model into consideration. This article presents H2O-Cloud, a Hierarchical and Hybrid Online task scheduling framework for warehouse-scale Cloud service providers, to improve resource usage effectiveness while maintaining QoS. H2O-Cloud is highly scalable and considers comprehensive information, such as various workload scenarios, cloud platform configurations, user request information, and dynamic pricing model. The hierarchy and hybridity of the framework, combined with its deep reinforcement learning (DRL) engines, enable H2O-Cloud to efficiently start on-the-go scheduling and learning in an unpredictable environment without pretraining. Our experiments confirm the high efficiency of the proposed H2O-Cloud when compared to baseline approaches, in terms of energy and cost while maintaining QoS. Compared with a state-of-the-art DRL-based algorithm, H2O-Cloud achieves up to 201.17% energy cost efficiency improvement, 47.88% energy efficiency improvement, and 551.76% reward rate improvement. Mingxi Cheng, Ji Li 0006, Paul Bogdan, Shahin Nazarian |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2019 | Specification Mining and Robust Design under Uncertainty: A Stochastic Temporal Logic ApproachabstractIn this paper, we propose Stochastic Temporal Logic (StTL) as a formalism for expressing probabilistic specifications on time-varying behaviors of controlled stochastic dynamical systems. To make StTL a more effective specification formalism, we introduce the quantitative semantics for StTL to reason about the robust satisfaction of an StTL specification by a given system. Additionally, we propose using the robustness value as the objective function to be maximized by a stochastic optimization algorithm for the purpose of controller design. Finally, we formulate an algorithm for parameter inference for Parameteric-StTL specifications, which allows specifications to be mined from output traces of the underlying system. We demonstrate and validate our framework on two case studies inspired by the automotive domain. Panagiotis Kyriakis, Jyotirmoy V. Deshmukh, Paul Bogdan |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2019 | Self-Optimizing and Self-Programming Computing Systems: A Combined Compiler, Complex Networks, and Machine Learning ApproachabstractThere exists an urgent need for determining the right amount and type of specialization while making a heterogeneous system as programmable and flexible as possible. Therefore, in this paper, we pioneer a self-optimizing and selfprogramming computing system (SOSPCS) design framework that achieves both programmability and flexibility and exploits computing heterogeneity [e.g., CPUs, GPUs, and hardware accelerators (HWAs)]. First, at compile time, we form a task pool consisting of hybrid tasks with different processing element (PE) affinities according to target applications. Tasks preferred to be executed on GPUs or accelerators are detected from target applications by neural networks. Tasks suitable to run on CPUs are formed by community detection to minimize data movement overhead. Next, a distributed reinforcement learning-based approach is used at runtime to allow agents to map the tasks onto the network-on-chip-based heterogeneous PEs by learning an optimal policy based on Q values in the environment. We have conducted experiments on a heterogeneous platform consisting of CPUs, GPUs, and HWAs with deep learning algorithms such as matrix multiplication, ReLU, and sigmoid functions. We concluded that SOSPCS provides performance improvement up to 4.12× and energy reduction up to 3.24× compared to the state-of-the-art approaches. Shahin Nazarian, Paul Bogdan |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2018 | Prometheus: Processing-in-memory heterogeneous architecture design from a multi-layer network theoretic strategyabstractWith increasing demand for distributed intelligent physical systems performing big data analytics on the field and in real-time, processing-in-memory (PIM) architectures integrating 3D-stacked memory and logic layers could provide higher performance and energy efficiency. Towards this end, the PIM design requires principled and rigorous optimization strategies to identify interactions and manage data movement across different vaults. In this paper, we introduce Prometheus, a novel PIM-based framework that constructs a comprehensive model of computation and communication (MoCC) based on a static and dynamic compilation of an application. Firstly, by adopting a low level virtual machine (LLVM) intermediate representation (IR), an input application is modeled as a two-layered graph consisting of (i) a computation layer in which the nodes denote computation IR instructions and edges denote data dependencies among instructions, and (ii) a communication layer in which the nodes denote memory operations (e.g., load/store) and edges represent memory dependencies detected by alias analysis. Secondly, we develop an optimization framework that partitions the multi-layer network into processing communities within which the computational workload is maximized while balancing the load among computational clusters. Thirdly, we propose a community-to-vault mapping algorithm for designing a scalable hybrid memory cube (HMC)-based system where vaults are interconnected through a network-on-chip (NoC) approach rather than a crossbar architecture. This ensures scalability to hundreds of vaults in each cube. Experimental results demonstrate that Prometheus consisting of 64 HMC-based vaults improves system performance by 9.8x and achieves 2.3x energy reduction, compared to conventional systems. Shahin Nazarian, Paul Bogdan |
DATE | 3 |
| 2018 | Prediction-based fast thermoelectric generator reconfiguration for energy harvesting from vehicle radiatorsabstractThermoelectric generation (TEG) has increasingly drawn attention for being environmentally friendly. A few researches have focused on improving TEG efficiency at system level on vehicle radiators. The most recent reconfiguration algorithm shows improvement on performance but suffers from major drawback on computational time and energy overhead, and non-scalability in terms of array size and processing frequency. In this paper, we propose a novel TEG array reconfiguration algorithm that determines near-optimal configuration with an acceptable computational time. More precisely, with O(N) time complexity, our prediction-based fast TEG reconfiguration algorithm enables all modules to work at or near their maximum power points (MPP). Additionally, we incorporate prediction methods to further reduce the runtime and switching overhead during the reconfiguration process. Experimental results present 30% performance improvement, almost 100 χ reduction on switching overhead and 13 χ enhancement on computational speed compared to the baseline and prior work. The scalability of our algorithm makes it applicable to larger scale systems such as industrial boilers and heat exchangers. Feiyang Kang, Caiwen Ding, Ji Li 0006, Donkyu Baek, Shahin Nazarian, Xue Lin 0001, Paul Bogdan, Naehyuck Chang |
DATE | 9 |
| 2018 | Accelerating Coverage Directed Test Generation for Functional Verification: A Neural Network-based FrameworkabstractWith increasing design complexity, the correlation between test transactions and functional properties becomes non-intuitive, hence impacting the reliability of test generation. This paper presents a modified coverage directed test generation based on an Artificial Neural Network (ANN). The ANN extracts features of test transactions and only those which are learned to be critical, will be sent to the design under verification. Furthermore, the priority of coverage groups is dynamically learned based on the previous test iterations. With ANN-based screening, low-coverage or redundant assertions will be filtered out, which helps accelerate the verification process. This allows our framework to learn from the results of the previous vectors and use that knowledge to select the following test vectors. Our experimental results confirm that our learning-based framework can improve the speed of existing function verification techniques by 24.5x and also also deliver assertion coverage improvement, ranging from 4.3x to 28.9x, compared to traditional coverage directed test generation, implemented in UVM. Fanchao Wang, Hanbin Zhu, Pranjay Popli, Paul Bogdan, Shahin Nazarian |
ACM Great Lakes Symposium on VLSI | 5 |
| 2018 | Toward Enabling Automated Cognition and Decision-Making in Complex Cyber-Physical SystemsabstractThis article presents a framework for empowering automated cognition and decision making in complex cyber-physical systems (CPS). The key idea is for each cyber-physical component in the system to be able to construct from observations and actions a reduced model of its behavior and respond to external stimuli in the form of a (timed) probabilistic automaton such as a semi-Markov decision process. The reduced order behavioral models of various cyber-physical components can then be integrated into an optimization and decision-making module to determine the best actions for the components under different operating scenarios and cost/payoff functions by solving a bounded-rationality global decision-making problem. Paul Bogdan, Massoud Pedram |
ISCAS | 1 |
| 2018 | Scalable Network-on-Chip Architectures for Brain-Machine Interface Applications
Karthi Duraisamy, Paul Bogdan, Janardhan Rao Doppa, Partha Pratim Pande |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2017 | A load balancing inspired optimization framework for exascale multicore systems: A complex networks approachabstractMany-core multi-threaded performance is plagued by on-chip communication nonidealities, limited memory bandwidth, and critical sections. Inspired by complex network theory of social communities, we propose a novel methodology to model the dynamic execution of an application and partition the application into an optimal number of clusters for parallel execution. We first adopt an LLVM IR compiler analysis of a specific application and construct a dynamic application dependency graph encoding its computational and memory operations. Next, based on this graph, we propose an optimization model to find the optimal clusters such that (1) the intra-cluster edges are maximized, (2) the execution times of the clusters are nearly equalized, for load balancing, and (3) the cluster size does not exceed the core count. Our novel approach confines data movement to be mainly inside a cluster for power reduction and congestion prevention. Finally, we propose an algorithm to sort the graph of connected clusters topologically and map the clusters onto NoC. Experimental results on a 32-core NoC demonstrate a maximum speedup of 131.82% when compared to thread-based execution. Furthermore, the scalability of our framework makes it a promising software design automation platform. Yuankun Xue, Shahin Nazarian, Paul Bogdan |
ICCAD | 4 |
| 2017 | A Reconfigurable Wireless NoC for Large Scale Microbiome Community AnalysisabstractUnderstanding the role of competition and cooperation among multiple interacting species of microorganisms that constitute the microbiome and decipher how they enforce homeostasis or trigger diseases requires the development of multi-scale computational models capable of capturing both intra-cell processing (i.e., gene-to-protein interactions) and inter-cell interactions. The multi-scale interdependency that governs the interactions from genes to proteins within a cell and from molecular messengers to cells to microbial communities within the environment raises numerous computation and communication challenges. Internal cell processing cannot be simulated without knowledge of the surroundings. Similarly, cell-cell communication cannot be fully abstracted without stated of internal processing and diffusion effects of molecular messengers. To address the computeand communication-intensive nature of modeling microbial communities, in this paper, we propose a novel reconfigurable NoC-based manycore architecture capable of simulating a large scale microbial community. The reconfiguration of the NoC topology is achieved through the fractal analysis of NoC traffic and use of the on-chip wireless interfaces. More precisely, we analyze the computational and communication workloads and exploit the observed fractal characteristics for proposing a mathematical strategy for NoC reconfiguration. Experimental results demonstrate that the proposed NoC architecture achieves 56.6 and 62.8 percent improvement in energy delay product over the conventional wireline mesh and flatten butterfly-based high radix NoC architectures, respectively. Karthi Duraisamy, Joe Baylon, Turbo Majumder, Guopeng Wei, Paul Bogdan, Deuk Hyoun Heo, Partha Pratim Pande |
IEEE Trans. Computers | 6 |
| 2017 | Fundamental Challenges Toward Making the IoT a Reachable Reality: A Model-Centric InvestigationabstractConstantly advancing integration capability is paving the way for the construction of the extremely large scale continuum of the Internet where entities or things from vastly varied domains are uniquely addressable and interacting seamlessly to form a giant networked system of systems known as the Internet-of-Things (IoT). In contrast to this visionary networked system paradigm, prior research efforts on the IoT are still very fragmented and confined to disjoint explorations of different applications, architecture, security, services, protocol, and economical domains, thus preventing design exploration and optimization from a unified and global perspective. In this context, this survey article first proposes a mathematical modeling framework that is rich in expressivity to capture IoT characteristics from a global perspective. It also sets forward a set of fundamental challenges in sensing, decentralized computation, robustness, energy efficiency, and hardware security based on the proposed modeling framework. Possible solutions are discussed to shed light on future development of the IoT system paradigm. Yuankun Xue, Ji Li 0006, Shahin Nazarian, Paul Bogdan |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2017 | Multicast-Aware High-Performance Wireless Network-on-Chip ArchitecturesabstractToday's multiprocessor platforms employ the network-on-chip (NoC) architecture as the preferable communication backbone. Conventional NoCs are designed predominantly for unicast data exchanges. In such NoCs, the multicast traffic is generally handled by converting each multicast message to multiple unicast transmissions. Hence, applications dominated by multicast traffic experience high queuing latencies and significant performance penalties when running on systems designed with unicast-based NoC architectures. Various multicast mechanisms such as XY-tree multicast and path multicast have already been proposed to enhance the performance of the traditional wireline mesh NoC incorporating multicast traffic. However, even with such added features, the multihop nature of the wireline mesh NoC leads to high network latencies and thus limits the achievable system performance. In this paper, to sustain the high-bandwidth and high-throughput requirements of emerging applications, we propose the design of a wireless NoC (WiNoC) architecture incorporating necessary multicast support. By integrating congestion-aware multicast routing with network coding, the WiNoC is able to efficiently handle heavy multicast injections. For applications running with a broadcast-heavy Hammer cache coherence protocol, the proposed multicast-aware WiNoC achieves an average of 47% reduction in message latency compared with the XY-tree-based multicast-aware mesh NoC. This network level improvement translates into a 26% saving in full-system energy delay product. Karthi Duraisamy, Yuankun Xue, Paul Bogdan, Partha Pratim Pande |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2016 | Power and thermal management in massive multicore chips: theoretical foundation meets architectural innovation and resource allocationabstractContinuing progress and integration levels in silicon technologies make possible complete end-user systems consisting of extremely high number of cores on a single chip targeting either embedded or high-performance computing. However, without new paradigms of energy- and thermally-efficient designs, producing information and communication systems capable of meeting the computing, storage and communication demands of the emerging applications will be unlikely. The broad topic of power and thermal management of massive multicore chips is actively being pursued by a number of researchers worldwide, from a variety of different perspectives, ranging from workload modeling to efficient on-chip network infrastructure design to resource allocation. Successful solutions will likely adopt and encompass elements from all or at least several levels of abstraction. Starting from these ideas, we consider a holistic approach in establishing the Power-Thermal-Performance (PTP) trade-offs of massive multicore processors by considering three inter-related but varying angles, viz., on-chip traffic modeling, novel Networks-on-Chip (NoC) architecture and resource allocation/mapping Paul Bogdan, Partha Pratim Pande, Hussam Amrouch, Muhammad Shafique 0001, Jörg Henkel |
CASES | 1 |
| 2016 | A spatio-temporal fractal model for a CPS approach to brain-machine-body interfaces
Yuankun Xue, Paul Bogdan |
DATE | 3 |
| 2016 | Power-aware virtual machine mapping in the data-center-on-a-chip paradigmabstractIt is projected that hundreds of cores can be integrated into a chip at the sub-20nm technology nodes. However, some challenges exist in the many-core architecture such as maintaining memory coherence, underutilized parallelism, and increased inter-core communication delay. This work proposes the data-center-on-a-chip (DCoC) paradigm employing virtualization technologies commonly used in today's data centers to reduce the overhead of maintaining memory coherence and inter-core communication and improve parallelism. In the DCoC paradigm, user applications with specific resource requirements need to be mapped onto different chips of a data center and different cores of a chip in the form of virtual machines (VMs). By a judicious VM mapping method, the data center performance can be maximized while satisfying the power budget and power density constraints of the chips and the resource requirements of VMs. To tackle the NP-hardness of the VM mapping problem, we propose a two-tier algorithm, which effectively solves the mapping problem with polynomial time complexity. Xue Lin 0001, Yuankun Xue, Paul Bogdan, Yanzhi Wang 0001, Siddharth Garg, Massoud Pedram |
ICCD | 3 |
| 2016 | Improving NoC performance under spatio-temporal variability by runtime reconfiguration: a general mathematical frameworkabstractReal-world applications exhibit time varying traffic volumes and patterns (structures). However, regular or application specific networks-on-chip (NoC) optimized at design time are predominantly developed either under the assumption of time independent application structures and dynamics, or without considering them at all. The limited adaptability of the monolithic NoCs to spatio-temporal application variability results in poor performance and power inefficient designs. Most of prior reconfiguration approaches are evaluated and validated through a subset of experimental instances out of the entire problem space, resulting in a constrained applicability to general cases. Towards this end, we propose a general mathematical framework for reconfigurable NoC design and runtime optimization under spatio-temporal varying workloads. Our proposed mathematical modeling framework can be applied to NoCs of arbitrary topologies, sizes and routing protocols. More precisely, we first propose the general mathematical modeling for reconfigurable NoC systems. Then, as a case study, we formulate the runtime NoC reconfiguration as an optimization problem assuming applications modeled via time dependent graphical models. We show this problem is NP-hard and prove its submodularity. This submodularity property provides guarantees that our proposed greedy algorithms attaints optimality region. To further justify the mathematical framework, we investigate a reconfigurable NoC architecture and evaluate it under real-world applications. Experimental results demonstrate a 52.3% latency improvement and a 30.2% energy reduction compared to baseline design. Yuankun Xue, Paul Bogdan |
NOCS | 2 |
| 2016 | A Support Vector Regression (SVR)-Based Latency Model for Network-on-Chip (NoC) ArchitecturesabstractIn this paper, we propose SVR-NoC, a network-on-chip (NoC) latency model using support vector regression (SVR). More specifically, based on the application communication information and the NoC routing algorithm, the channel and source queue waiting times are first estimated using an analytical queuing model with two equivalent queues. To improve the prediction accuracy, the queuing theory-based delay estimations are included as features in the learning process. We then propose a learning framework that relies on SVR to collect training data and predict the traffic flow latency. The proposed learning methods can be used to analyze various traffic scenarios for the target NoC platform. Experimental results on both synthetic and real-application traffic demonstrate on average less than 12% prediction error in network saturation load, as well as more than 100× speedup compared to cycle-accurate simulations can be achieved. Zhiliang Qian, Da-Cheng Juan, Paul Bogdan, Chi-Ying Tsui, Diana Marculescu, Radu Marculescu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2016 | Performance Evaluation of NoC-Based Multicore Systems: From Traffic Analysis to NoC Latency ModelingabstractIn this survey, we review several approaches for predicting performance of Network-on-Chip (NoC)-based multicore systems, starting from the traffic models to the complex NoC models for latency evaluation. We first review typical traffic models to represent the application workloads in NoC. Specifically, we review Markovian and non-Markovian (e.g., self-similar or long-range memory-dependent) traffic models and discuss their applications on multicore platform design. Then, we review the analytical techniques to predict NoC performance under given input traffic. We investigate analytical models for average as well as maximum delay evaluation. We also review the developments and design challenges of NoC simulators. One interesting research direction in NoC performance evaluation consists of combining simulation and analytical models in order to exploit their advantages together. Toward this end, we discuss several newly proposed approaches that use hardware-based or learning-based techniques. Finally, we summarize several open problems and our perspective to address these challenges. Zhiliang Qian, Paul Bogdan, Chi-Ying Tsui, Radu Marculescu |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2016 | Network-on-Chip-Enabled Multicore Platforms for Parallel Model Predictive ControlabstractInternet-of-Things architecture aims to provide smart connectivity not only with existing computers, but also with new context-aware computing resources, extending soon beyond von Neumann devices for the purpose of mining, prediction, and control of cyber and physical components. These cyber-physical systems (CPSs) not only lead to the accumulation of large amounts of data that can be used to build comprehensive mathematical models, but also raise the quest for real-time analysis and control in diverse application domains, such as environment, healthcare, avionics, smart interconnected automobiles, and smart buildings. Endowing the CPS with a higher degree of distributed smartness and cognition (adaptation) to process massive amounts of data requires efficient control modules. In addition, the prohibitive nature of power consumption, data movement, and memory bandwidth issues calls for a shift of processing the decision-making strategies from within large supercomputing centers closer to the actual sensing site via many distributed networks-on-chip (NoCs)-based multicore platforms. Toward this end, in this paper, we propose an efficient NoC-based multicore architecture capable of solving large-scale nonlinear model predictive control (NMPC) problems. By carefully analyzing the spatiotemporal workload characteristics of the NMPC problems, we propose the design of an efficient NoC architecture. Our proposed NoC architecture achieves up to 29% improvement in latency and 28% improvement in energy dissipation over the conventional mesh NoC-based counterpart. Karthi Duraisamy, Paul Bogdan, Turbo Majumder, Partha Pratim Pande |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2015 | A cyber-physical systems approach to personalized medicine: challenges and opportunities for noc-based multicore platforms
Paul Bogdan |
DATE | 1 |
| 2015 | NoC-enabled multicore architectures for stochastic analysis of biomolecular reactions
Turbo Majumder, Paul Bogdan, Partha Pratim Pande |
DATE | 3 |
| 2015 | Analyzing the Dark Silicon Phenomenon in a Many-Core Chip Multi-Processor under Deeply-Scaled Process TechnologiesabstractThe impact of dark silicon phenomenon on multicore processors under deeply-scaled FinFET technologies is investigated in this paper. To do this accurately, a cross-layer framework, spanning device, circuit, and architecture levels is initially introduced. Using this framework, leakage and dynamic power consumptions as well as frequency levels of in-order and out-of-order (OoO) processor cores, and on-chip cache memories and routers in a network-on-chip-based chip multiprocessor system synthesized in 7nm FinFET technology and operating in both super- and near-threshold voltage regimes are presented. Subsequently, total power consumptions of multicore chips manufactured with (i) OoO and (ii) in-order processor cores are reported and compared. According to our results, for a 64-core chip and 15W thermal design power budget, 64% and 39% dark silicon are observed in OoO and in-order multicores, respectively, under super-threshold regime. These percentages drop to 19% and 0% for OoO and in-order multicores operating in the near-threshold regime, respectively. Furthermore, the highest energy efficiencies are achieved by operating in the near-threshold regime, which points to the effectiveness of near-threshold computing in mitigating the effect of dark silicon phenomenon under deeply-scaled technologies. Alireza Shafaei, Yanzhi Wang 0001, Srikanth Ramadurgam, Yuankun Xue, Paul Bogdan, Massoud Pedram |
ACM Great Lakes Symposium on VLSI | 5 |
| 2015 | Multiscale modeling of biological communicationabstractTo capture the nonlinear and stochastic features of biological communication, we propose a non-equilibrium statistical physics inspired model describing the information processing within a cell, communication processes among cells and behavioral population response in an interkingdom biological environment. We propose new metrics to quantify information processing and transmission at various biological scales. We highlight the need for taking the biological nonlinear dynamics into account and evaluate the newly proposed information theoretic metrics for a Vibrio fischeri population capable of producing bioluminiscence. Navid Azizan, Paul Bogdan |
ICC | 2 |
| 2015 | Mathematical Models and Control Algorithms for Dynamic Optimization of Multicore Platforms: A Complex Dynamics ApproachabstractThe continuous increase in integration densities contributed to a shift from Dennard's scaling to a parallelization era of multi-/many-core chips. However, for multicores to rapidly percolate the application domain from consumer multimedia to high-end functionality (e.g., security, healthcare, big data), power/energy and thermal efficiency challenges must be addressed. Increased power densities can raise on-chip temperatures, which in turn decrease chip reliability and performance, and increase cooling costs. For a dependable multicore system, dynamic optimization (power / thermal management) has to rely on accurate yet low complexity workload models. Towards this end, we present a class of mathematical models that generalize prior approaches and capture their time dependence and long-range memory with minimum complexity. This modeling framework serves as the basis for defining new efficient control and prediction algorithms for hierarchical dynamic power management of future data-centers-on-a-chip. Paul Bogdan, Yuankun Xue |
ICCAD | 1 |
| 2015 | Machine Learning-Based Energy Management in a Hybrid Electric Vehicle to Minimize Total Operating CostabstractThis paper investigates the energy management problem in hybrid electric vehicles (HEVs) focusing on the minimization of the operating cost of an HEV, including both fuel and battery replacement cost. More precisely, the paper presents a nested learning framework in which both the optimal actions (which include the gear ratio selection and the use of internal combustion engine versus the electric motor to drive the vehicle) and limits on the range of the state-of-charge of the battery are learned on the fly. The inner-loop learning process is the key to minimization of the fuel usage whereas the outer-loop learning process is critical to minimization of the amortized battery replacement cost. Experimental results demonstrate a maximum of 48% operating cost reduction by the proposed HEV energy management policy. Xue Lin 0001, Paul Bogdan, Naehyuck Chang, Massoud Pedram |
ICCAD | 2 |
| 2015 | Optimizing fuel economy of hybrid electric vehicles using a Markov decision process modelabstractIn contrast to conventional internal combustion engine (ICE) propelled vehicles, hybrid electric vehicles (HEVs) can achieve both higher fuel economy and lower pollutant emissions. The HEV features a hybrid propulsion system consisting of one ICE and one or more electric motors (EMs). The use of both ICE and EM increases the complexity of HEV power management, and so advanced power management policy is required for achieving higher performance and lower fuel consumption. This work aims at minimizing the HEV fuel consumption over any driving cycles, about which no complete information is available to the HEV controller in advance. Therefore, this work proposes to model the HEV power management problem as a Markov decision process (MDP) and derives the optimal power management policy using the policy iteration technique. Simulation results over real-world and testing driving cycles demonstrate that the proposed optimal power management policy improves HEV fuel economy by 23.9% on average compared to the rule-based policy. Xue Lin 0001, Yanzhi Wang 0001, Paul Bogdan, Naehyuck Chang, Massoud Pedram |
Intelligent Vehicles Symposium | 3 |
| 2015 | Mathematical Modeling and Control of Multifractal Workloads for Data-Center-on-a-Chip OptimizationabstractBuilding autonomous data-centers-on-chip (DCoC) for exascale computing requires mathematical frameworks that account and exploit the non-stationary and multi-fractal characteristics of computation and communication workloads. Towards this end, relying on DCoC (Intel's SCC) measurements, we propose a complex dynamical modeling approach that captures the observed multi-fractal characteristics of inter-event times between successive workload changes and the magnitude of the increments in DCoC workloads. Our novel mathematical framework allows for the analysis of higher order moments and enables the formulation of more accurate model predictive control strategies for multi-fractal dynamics. We investigate the impact of the multi-fractal spectrum richness on the performance of the control algorithm. Our mathematical formalism can further be used to model, analyze and solve DCoC design problems (e.g., topology reconfiguration, buffer sizing, mapping, scheduling, resource management, congestion control). Paul Bogdan |
NOCS | 1 |
| 2015 | NoC Architectures as Enablers of Biological Discovery for Personalized and Precision MedicineabstractThis paper overviews the main computational issues in personalized and precision medicine (PPM), and present a cogent case for network-on-chip (NoC)-based multicore platforms as enablers in the process. We identify a series of challenges for the design and optimization of NoC-based solutions for PPM. To capture the characteristics of the cyber-physical sensing and processing, we propose a new computational model built on a dynamical heterogeneous hyper-graph description of application-to-architecture interactions. Starting from these premises, we summarize a few implications on NoC design methodologies, present some NoC-based solutions that deal with some of the challenges, and outline a few open problems. Paul Bogdan, Turbo Majumder, Arvind Ramanathan, Yuankun Xue |
NOCS | 1 |
| 2015 | User Cooperation Network Coding Approach for NoC Performance ImprovementabstractThe astonishing rate of sensing modalities and data generation poses a tremendous impact on computing platforms for providing real-time mining and prediction capabilities. We are capable of monitoring thousands of genes and their interactions, but we lack efficient computing platforms for large-scale (exa-scale) data processing. Towards this end, we propose a novel hierarchical Network-on-Chip (NoC) architecture that exploits user-cooperated network coding (NC) concepts for improving system throughput. Our proposed architecture relies on a light-weighted subnet of cooperation unit routers (CUR) for multicast traffic. Coding network interface (CNI) performs encoding/decoding of NC symbols and shares the data flows among cooperation units(CUs). We endow our proposed NC-based NoC architecture with: (i) a corridor routing algorithm (CRA) for maximizing network throughput and (ii) an adaptive flit dropping (AFD) scheme to mitigate congestion, branch-blocking and deadlock at run-time. The experimental results demonstrate that our proposed platform offers up to 127X multicast throughput improvement over multiple-unicast and XY tree-based multicast under synthetic collective traffic scenario. We have evaluated the proposed platform with different realworld benchmarks under network sizes of 4x4 to 32x32. Simulation results show 21%--91% latency improvement and up to 25X runtime reduction over conventional mesh NoC performing genetic-algorithm based protein folding analysis. FPGA implementation results show minimal overhead. Yuankun Xue, Paul Bogdan |
NOCS | 2 |
| 2014 | A comprehensive and accurate latency model for Network-on-Chip performance analysisabstractIn this work, we propose a new, accurate, and comprehensive analytical model for Network-on-Chip (NoC) performance analysis. Given the application communication graph, the NoC architecture, and the routing algorithm, the proposed framework analyzes the links dependency and then determines the ordering of queuing analysis for performance modeling. The channel waiting times in the links are estimated using a generalized G/G/1/K queuing model, which can tackle bursty traffic and dependent arrival times with general service time distributions. The proposed model is general and can be used to analyze various traffic scenarios for NoC platforms with arbitrary buffer and packet lengths. Experimental results on both synthetic and real applications demonstrate the accuracy and scalability of the newly proposed model. Zhiliang Qian, Da-Cheng Juan, Paul Bogdan, Chi-Ying Tsui, Diana Marculescu, Radu Marculescu |
ASP-DAC | 3 |
| 2014 | Disease Diagnosis-on-a-Chip: Large Scale Networks-on-Chip based Multicore Platform for Protein Folding AnalysisabstractProtein folding is critical for many biological processes. In this work, we propose an NoC-based multi-core platform for protein folding computation. We first identify the speedup bottleneck for applying conventional genetic algorithm on a mesh-based multi-core platform. Then, we address this computation- and communication- intensive problem while taking into account both hardware and software aspects. Specifically, we group the processing cores into islands and propose an NoC-based multicore architecture for intra- and inter-island communication. The high scalability of the proposed platform allows us to integrate from 100 to 1200 cores for the folding computation. We then propose a genetic migration algorithm to take advantage of the massive parallel platform. Our simulation results show that the proposed platform offers near-linear speedup as the number of cores increases. We also report the hardware cost in area and power based on a 100-core FPGA prototype. Yuankun Xue, Zhiliang Qian, Paul Bogdan, Fan Ye 0001, Chi-Ying Tsui |
DAC | 3 |
| 2014 | Low-latency wireless 3D NoCs via randomized shortcut chipsabstractIn this paper, we demonstrate that we can reduce the communication latency significantly by inserting a fraction of randomness into a wireless 3D NoC (where CMOS wireless links are used for vertical inter-chip communication) when considering the physical constraints of the 3D design space. Towards this end, we consider two cases, namely 1) replacing existing horizontal 2D links in a wireless 3D NoC with randomized shortcut NoC links and 2) enabling full connectivity by adding a randomized NoC layer to a wireless 3D platform with partial or no horizontal connectivity. Consequently, the packet routing is optimized by exploiting both the existing and the newly added random NoC. At the same time, by adding randomly wired shortcut NoCs to a wireless 3D platform, a good balance can be established between the modularity of the design and the minimum randomness needed to achieve low latency, and experimental results show that by adding a random NoC chip to wireless 3D CMPs without built-in horizontal connectivity, the communication latency can be reduced by as much as 26.2% when compared to adding a 2D mesh NoC. Also, the application execution time and average flit transfer energy can be improved accordingly. Hiroki Matsutani, Michihiro Koibuchi, Ikki Fujiwara, Takahiro Kagami, Yasuhiro Take, Tadahiro Kuroda, Paul Bogdan, Radu Marculescu, Hideharu Amano |
DATE | 7 |
| 2014 | Reinforcement learning based power management for hybrid electric vehiclesabstractCompared to conventional internal combustion engine (ICE) propelled vehicles, hybrid electric vehicles (HEVs) can achieve both higher fuel economy and lower pollution emissions. The HEV consists of a hybrid propulsion system containing one ICE and one or more electric motors (EMs). The use of both ICE and EM increases the complexity of HEV power management, and therefore requires advanced power management policies to achieve higher performance and lower fuel consumption. Towards this end, our work aims at minimizing the HEV fuel consumption over any driving cycle (without prior knowledge of the cycle) by using a reinforcement learning technique. This is in clear contrast to prior work, which requires deterministic or stochastic knowledge of the driving cycles. In addition, the proposed reinforcement learning technique enables us to (partially) avoid reliance on complex HEV modeling while coping with driver specific behaviors. To our knowledge, this is the first work that applies the reinforcement learning technique to the HEV power management problem. Simulation results over real-world and testing driving cycles demonstrate the proposed HEV power management policy can improve fuel economy by 42%. Xue Lin 0001, Yanzhi Wang 0001, Paul Bogdan, Naehyuck Chang, Massoud Pedram |
ICCAD | 3 |
| 2014 | An efficient Network-on-Chip (NoC) based multicore platform for hierarchical parallel genetic algorithmsabstractIn this work, we propose a new Network-on-Chip (NoC) architecture for implementing the hierarchical parallel genetic algorithm (HPGA) on a multi-core System-on-Chip (SoC) platform. We first derive the speedup metric of an NoC architecture which directly maps the HPGA onto NoC in order to identify the main sources of performance bottlenecks. Specifically, it is observed that the speedup is mostly affected by the fixed bandwidth that a master processor can use and the low utilization of slave processor cores. Motivated by the theoretical analysis, we propose a new architecture with two multiplexing schemes, namely dynamic injection bandwidth multiplexing (DIBM) and time-division based island multiplexing (TDIM), to improve the speedup and reduce the hardware requirements. Moreover, a task-aware adaptive routing algorithm is designed for the proposed architecture, which can take advantage of the proposed multiplexing schemes to further reduce the hardware overhead. We demonstrate the benefits of our approach using the problem of protein folding prediction, which is a process of importance in biology. Our experimental results show that the proposed NoC architecture achieves up to 240X speedup compared to a single island design. The hardware cost is also reduced by 50% compared to a direct NoC-based HPGA implementation. Yuankun Xue, Zhiliang Qian, Guopeng Wei, Paul Bogdan, Chi-Ying Tsui, Radu Marculescu |
NOCS | 4 |
| 2014 | Exploiting Emergence in On-Chip InterconnectsabstractTo solve the grand challenges in contemporary chip design, such as process-to-core mapping, energy reduction, and maintenance of programmer/hardware abstraction, we advocate for self-optimizing (emergent) networks-on-chip (NoC). In these networks, topology and information flow adapt dynamically to maximize the network throughput or minimize the network latency via distributed application of microrules. In this paper, we introduce the concept of emergent small-world NoCs and discuss novel design decisions, e.g., Skip-links, that improve performance and reduce energy consumption of multicore systems. More precisely, we demonstrate that our proposed solution is able to adapt to a wide range of traffic patterns and provide reductions in data hop count of up to 20 percent while maintaining energy and area costs. We show how emergent networks can be useful for on-chip processor-to-processor communications, and also demonstrate how SoC and off-chip I/O traffic may be optimized for latency and critical load. Simon J. Hollis, Chris Jackson, Paul Bogdan, Radu Marculescu |
IEEE Trans. Computers | 3 |
| 2013 | A case for wireless 3D NoCs for CMPsabstractInductive-coupling is yet another 3D integration technique that can be used to stack more than three known-good-dies in a SiP without wire connections. We present a topology-agnostic 3D CMP architecture using inductive-coupling that offers great flexibility in customizing the number of processor chips, SRAM chips, and DRAM chips in a SiP after chips have been fabricated. In this paper, first, we propose a routing protocol that exchanges the network information between all chips in a given SiP to establish efficient deadlock-free routing paths. Second, we propose its optimization technique that analyzes the application traffic patterns and selects different spanning tree roots so as to minimize the average hop counts and improve the application performance. Hiroki Matsutani, Paul Bogdan, Radu Marculescu, Yasuhiro Take, Daisuke Sasaki, Hao Zhang 0020, Michihiro Koibuchi, Tadahiro Kuroda, Hideharu Amano |
ASP-DAC | 2 |
| 2013 | SVR-NoC: a performance analysis tool for network-on-chips using learning-based support vector regression modelabstractIn this work, we propose SVR-NoC, a learning-based support vector regression (SVR) model for evaluating Network-on-Chip (NoC) latency performance. Different from the state-of-the-art NoC analytical model, which uses classical queuing theory to directly compute the average channel waiting time, the proposed SVR-NoC model performs NoC latency analysis based on learning the typical training data. More specifically, we develop a systematic machine-learning framework that uses the kernel-based support vector regression method to predict the channel average waiting time and the traffic flow latency. Experimental results show that SVR-NoC can predict the average packet latency accurately while achieving about 120X speed-up over simulation-based evaluation methods. Zhiliang Qian, Da-Cheng Juan, Paul Bogdan, Chi-Ying Tsui, Diana Marculescu, Radu Marculescu |
DATE | 3 |
| 2013 | Performance evaluation of multicore systems: from traffic analysis to latency predictions (embedded tutorial)abstractAs technology scaling down allows multiple processing components to be integrated on a single chip, the modern computing systems led to the advent of Multiprocessor System-on-Chip (MPSoC) and Chip Multiprocessor (CMP) design. Network-on-Chips (NoCs) have been proposed as a promising solution to tackle the complex on-chip communication problems on these multicore platforms. In order to optimize the NoC-based multicore system design, it is essential to evaluate the NoC performance with respect to numerous configurations in a large design space. Taking the traffic characteristics into account and using an appropriate latency model become crucially important to provide an accurate and fast evaluation. In this tutorial, we survey the current progresses in these aspects. We first review the NoC workload modeling and traffic analysis techniques. Then, we discuss the mathematical formalisms of evaluating the performance under a given traffic model, for both the average and worst-case latency predictions. Finally, the advantages of combining the analytical and simulation-based techniques are discussed and new attempts for bridging these two approaches are reviewed. Zhiliang Qian, Paul Bogdan, Chi-Ying Tsui, Radu Marculescu |
ICCAD | 2 |
| 2013 | Balls into Bins via Local SearchabstractWe propose a natural process for allocating n balls into n bins that are organized as the vertices of an undirected graph G.Each ball first chooses a vertex u in G uniformly at random.Then the ball performs a local search in G starting from u until it reaches a vertex with local minimum load, where the ball is finally placed on.In our main result, we prove that this process yields a maximum load of only Θ(log log n) on expander graphs.In addition, we show that for d-dimensional grids the maximum load is Θ log n log log n 1 d+1 .Finally, for almost regular graphs with minimum degree Ω(log n), we prove that the maximum load is constant and also reveal a fundamental difference between random and arbitrary tie-breaking rules. Paul Bogdan, Thomas Sauerwald, Alexandre Stauffer, He Sun 0001 |
SODA | 1 |
| 2013 | Efficient Modeling and Simulation of Bacteria-Based Nanonetworks with BNSimabstractBacteria-based networks are formed using native or engineered bacteria that communicate at nano-scale. This definition includes the micro-scale molecular transportation system which uses chemotactic bacteria for targeted cargo delivery, as well as genetic circuits for intercellular interactions like quorum sensing or light communication. To characterize the dynamics of bacterial networks accurately, we introduce BNSim, an open-source, parallel, stochastic, and multiscale modeling platform which integrates various simulation algorithms, together with genetic circuits and chemotactic pathway models in a complex 3D environment. Moreover, we show how this platform can be used to model synthetic bacterial consortia which implement a XOR function and aggregate nearby bacteria using light communication. Consequently, the results demonstrate how BNSim can predict various properties of realistic bacterial networks and provide guidance for their actual wet-lab implementations. Guopeng Wei, Paul Bogdan, Radu Marculescu |
IEEE J. Sel. Areas Commun. | 2 |
| 2013 | Bumpy Rides: Modeling the Dynamics of Chemotactic Interacting BacteriaabstractRecent advances in synthetic biology have brought the fantasy of having synthetic multicellular systems working at nanoscale closer to reality. Indeed, such systems consisting of networked biological nanomachines have a great potential for many novel applications like environmental monitoring and healthcare. In this paper, we consider dense networks of interacting bacteria capable of monitoring and treating diseases that affect microscopic regions in the human body. Towards this end, we propose a modeling approach for capturing the dynamics of such dense populations of interacting bacteria and estimate their performance for diagnostic and drug delivery purposes. Consequently, our approach can be used to identify various design trade-offs for dense networks of bacteria and design predictable and reliable synthetic multicellular systems. Guopeng Wei, Paul Bogdan, Radu Marculescu |
IEEE J. Sel. Areas Commun. | 2 |
| 2013 | Pacemaker control of heart rate variability: A cyber physical system perspectiveabstractCardiac diseases, like those related to abnormal heart rate activity, have an enormous economic and psychological impact worldwide. The approaches used to control the behavior of modern pacemakers ignore the fractal nature of heart rate activity. The purpose of this article is to present a Cyber Physical System approach to pacemaker design that exploits precisely the fractal properties of heart rate activity in order to design the pacemaker controller. Towards this end, we solve a finite horizon optimal control problem based on the heartbeat time series and show that this control problem can be converted into a system of linear equations. We also compare and contrast the performance of the fractal optimal control problem under six different cost functions. Finally, to get an idea of hardware complexity, we implement the fractal optimal controller on a Virtex4 FPGA and report some preliminary results in terms of area overhead. Paul Bogdan, Radu Marculescu |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2013 | Dynamic power management for multidomain system-on-chip platforms: An optimal control approachabstractReducing energy consumption in multiprocessor systems-on-chip (MPSoCs) where communication happens via the network-on-chip (NoC) approach calls for multiple voltage/frequency island (VFI)-based designs. In turn, such multi-VFI architectures need efficient, robust, and accurate runtime control mechanisms that can exploit the workload characteristics in order to save power. Despite being tractable, the linear control models for power management cannot capture some important workload characteristics (e.g., fractality, nonstationarity) observed in heterogeneous NoCs; if ignored, such characteristics lead to inefficient communication and resources allocation, as well as high power dissipation in MPSoCs. To mitigate such limitations, we propose a new paradigm shift from power optimization based on linear models to control approaches based on fractal-state equations. As such, our approach is the first to propose a controller for fractal workloads with precise constraints on state and control variables and specific time bounds. Our results show that significant power savings can be achieved at runtime while running a variety of benchmark applications. Paul Bogdan, Radu Marculescu |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2012 | Modeling populations of micro-robots for biological applicationsabstractIn order to detect and cure diseases affecting microscopic regions of the human body, swarms of micro-robots can harness the motility of bacteria and swim towards inaccessible regions of the body to detect abnormal behavior and/or deliver various drugs. In this paper, we propose a modeling approach for the dynamics of dense networks of swarms and estimate their performance via stochastic metrics. Paul Bogdan, Guopeng Wei, Radu Marculescu |
ICC | 1 |
| 2012 | An Optimal Control Approach to Power Management for Multi-Voltage and Frequency Islands Multiprocessor Platforms under Highly Variable WorkloadsabstractReducing energy consumption in multi-processor systems-on-chip (MPSoCs) where communication happens via the network-on-chip (NoC) approach calls for multiple voltage/frequency island (VFI)-based designs. In turn, such multi-VFI architectures need efficient, robust, and accurate run-time control mechanisms that can exploit the workload characteristics in order to save power. Despite being tractable, the linear control models for power management cannot capture some important workload characteristics (e.g., fractality, non-stationarity) observed in heterogeneous NoCs, if ignored, such characteristics lead to inefficient communication and resources allocation, as well as high power dissipation in MPSoCs. To mitigate such limitations, we propose a new paradigm shift from power optimization based on linear models to control approaches based on fractal-state equations. As such, our approach is the first to propose a controller for fractal workloads with precise constraints on state and control variables and specific time bounds. Our results show that significant power savings (about 70%) can be achieved at run-time while running a variety of benchmark applications. Paul Bogdan, Radu Marculescu, Rafael Tornero |
NOCS | 1 |
| 2012 | Dynamic power management for multicores: Case study using the intel SCCabstractIn this paper, a dynamic voltage and frequency scaling (DVFS) based power management approach is implemented using the Intel SCC platform. To validate our framework, we show various power and performance trade-offs using the SCC platform, while running a computationally intensive parallel image-processing application under different power management conditions. David Radu, Paul Bogdan, Radu Marculescu |
VLSI-SoC | 2 |
| 2011 | Sustainability through massively integrated computing: Are we ready to break the energy efficiency wall for single-chip platforms?abstractWhile traditional cluster computers are more constrained by power and cooling costs for solving extreme-scale (or exascale) problems, the continuing progress and integration levels in silicon technologies make possible complete end-user systems on a single chip. This massive level of integration makes modern multicore chips all pervasive in domains ranging from climate forecasting and astronomical data analysis, to consumer electronics, smart phones, and biological applications. Consequently, designing multicore chips for exascale computing while using the embedded systems design principles looks like a promising alternative to traditional cluster-based solutions. This paper aims to present an overview of new, far-reaching design methodologies and run-time optimization techniques that can help breaking the energy efficiency wall in massively integrated single-chip computing platforms. Partha Pratim Pande, Fabien Clermidy, Diego Puschini, Imen Mansouri, Paul Bogdan, Radu Marculescu, Amlan Ganguly |
DATE | 5 |
| 2011 | Dynamic power management of voltage-frequency island partitioned Networks-on-Chip using Intel's Single-chip Cloud ComputerabstractContinuous technology scaling has enabled the integration of multiple cores on the same chip. To overcome the disadvantages of buses, the Network-on-Chip (NoC) architecture has been proposed as a new communication paradigm. To further mitigate the tradeoff between performance and power consumption, dynamic voltage and frequency scaling (DVFS) became the de facto approach in multi-core design. DVFS-based NoC communication was implemented in Intel's most recent Singlechip Cloud Computer (SCC). Using the SCC we demonstrate a power management algorithm that runs in real time and dynamically adjusts the performance of the islands to reduce power consumption while maintaining the same level of performance. David Radu, Paul Bogdan, Radu Marculescu, Ümit Y. Ogras |
NOCS | 2 |
| 2011 | A software framework for trace analysis targeting multicore platforms designabstractThis demonstration presents a complete software framework for dynamically mapping multi-threaded applications on a cycle accurate Network-on-Chip (NoC) architecture, analyzing the statistics of network workloads and drawing general guidelines regarding the design and optimization of NoCs. Guopeng Wei, Paul Bogdan, Radu Marculescu |
NOCS | 2 |
| 2011 | Non-Stationary Traffic Analysis and Its Implications on Multicore Platform DesignabstractNetworks-on-chip (NoCs) have been proposed as a viable solution to solving the communication problem in multicore systems. In this new setup, mapping multiple applications on available computational resources leads to interaction and contention at various network resources. Consequently, taking into account the traffic characteristics becomes of crucial importance for performance analysis and optimization of the communication infrastructure, as well as proper resource management. Although queuing-based approaches have been traditionally used for performance analysis purposes, they cannot properly account for many of the traffic characteristics (e.g., non-stationarity, self-similarity) that are crucial for multicore platform design. To overcome these limitations, we propose a statistical physics inspired approach to capture the traffic dynamics in multicore systems. As shown later in this paper, this is of fundamental significance for re-thinking the very basis of multicore systems design; it also opens up new research directions into NoC optimization which require accurate models of time-dependent and space-dependent traffic behavior. Paul Bogdan, Radu Marculescu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2011 | Hitting Time Analysis for Fault-Tolerant Communication at Nanoscale in Future Multiprocessor PlatformsabstractThis paper investigates the on-chip stochastic communication and proposes an analytical model for computing its mean hitting time. Toward this end, we model the on-chip stochastic communication of any source-destination pair as a branching and annihilating random walk taking place on a finite mesh. The evolution of this branching process is studied via a master equation which helps us estimate the mean number of communication rounds needed to reach a destination node from a particular source node. Besides the probabilistic performance analysis, we also present experimental results for two concrete platforms and assess the potential of stochastic communication for future nanotechnologies. Paul Bogdan, Radu Marculescu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2010 | QuaLe: A Quantum-Leap Inspired Model for Non-stationary Analysis of NoC Traffic in Chip Multi-processorsabstractThis paper identifies non-stationary effects in grid like Network-on-Chip (NoC) traffic and proposes QuaLe, a novel statistical physics-inspired model, that can account for non-stationarity observed in packet arrival processes. Using a wide set of real application traces, we demonstrate the need for a multi-fractal approach and analyze various packet arrival properties accordingly. As a case study, we show the benefits of our multifractal approach in estimating the probability of missing deadlines in packet scheduling for chip multiprocessors (CMPs). Paul Bogdan, Miray Kas, Radu Marculescu, Onur Mutlu |
NOCS | 1 |
| 2010 | An Analytical Approach for Network-on-Chip Performance AnalysisabstractNetworks-on-chip (NoCs) have recently emerged as a scalable alternative to classical bus and point-to-point architectures. To date, performance evaluation of NoC designs is largely based on simulation which, besides being extremely slow, provides little insight on how different design parameters affect the actual network performance. Therefore, it is practically impossible to use simulation for optimization purposes. In this paper, we present a mathematical model for on-chip routers and utilize this new model for NoC performance analysis. The proposed model can be used not only to obtain fast and accurate performance estimates, but also to guide the NoC design process within an optimization loop. The accuracy of our approach and its practical use is illustrated through extensive simulation results. Ümit Y. Ogras, Paul Bogdan, Radu Marculescu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2007 | Quantum-Like Effects in Network-on-Chip Buffers BehaviorabstractThis paper proposes a paradigm shift in modeling and optimization of NoCs by identifying quantum-like effects in buffers behavior. The key idea of our proposed approach involves identifying a virtual random growing network (VRGN) which describes the NoC buffers evolution as a function of the packet injection rate. Paul Bogdan, Radu Marculescu |
DAC | 1 |