Vladimir Stojanovic

dblp:17/6937 · also Vladimir Marko Stojanovic · DBLP profile ↗
← Back
61ranked-venue papers
7as first author
25since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 26 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 14 · 13 since 2021Computer networks · 10 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 3Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 A UCIe Optical I/O Retimer Chiplet for AI Scale-up
Vladimir Stojanovic
HCS1
2025 Pseudo-label guided dual classifier domain adversarial network for unsupervised cross-domain fault diagnosis with small samples
Yawei Sun, Hongfeng Tao, Vladimir Stojanovic
Adv. Eng. Informatics3
2025 Open-set classification method via latent representation prompt and time-frequency fusion toward unknown fault recognition
Yawei Sun, Hongfeng Tao, Vladimir Stojanovic
Adv. Eng. Informatics3
2025 A Generic Single-Source Domain Generalization Framework for Fault Diagnosis via Wavelet Packet Augmentation and Pseudo-Domain Generation
abstract
During real-time production in industrial Internet of Things systems, equipment changes its operating speed due to changing operating conditions. And dynamic speed changes of rotating machinery under fluctuating workloads often lead to domain changes of vibration signals, which will directly lead to degradation of fault diagnostic model performance. Furthermore, the acquisition of data from multiple domains in real industrial scenarios is challenging due to the expense of collecting data from all possible working conditions. Consequently, applying diagnostic models trained using a single-source domain directly to an unknown target domain is a very challenging single domain generalization problem. Therefore, a generic single-source domain generalization framework via wavelet packet augmentation and pseudo-domain generation for fault diagnosis under unknown operating conditions is proposed in this paper. Pseudo-domain generation involves augmenting single-source domain by integrating data generetion model, thereby enhancing prediction accuracy. Furthermore, a wavelet packet augmentation method is proposed. Initially, the original signal is decomposed to obtain high and low frequency information. Subsequently, the high and low frequency information within the batch are linearly interpolated, respectively. Consequently, the interpolated high and low frequency information is then reconstructed to yield enhanced samples. The experimental results on four datasets show that the proposed framework can effectively improve the robustness of the generalization ability of fault diagnosis under unknown operating environments.
Yawei Sun, Hongfeng Tao, Yuanzhi Ni, Vladimir Stojanovic
IEEE Internet Things J.4
2025 Multi-domain weakly decoupled domain generalization network for fault diagnosis under unknown operating conditions
Yawei Sun, Hongfeng Tao, Vladimir Stojanovic
Knowl. Based Syst.3
2025 Composite neural learning-based adaptive actuator failure compensation control for full-state constrained autonomous surface vehicle
Shuai Song, Xiaona Song, Vladimir Stojanovic
Neural Comput. Appl.4
2025 PID-fuzzy switching-based strategy to heading control for remote operated vehicle
Baolong Xie, Shuping He, Honghai Wang, Vladimir Stojanovic, Kaibo Shi
Neural Comput. Appl.5
2025 Enhanced Control of Nonlinear Cyber-Physical Systems: A Higher-Order Relaxation Matrix-Based Convex Optimization Algorithm
abstract
This paper investigates the static output feedback (SOF) control problem of T-S fuzzy systems based on higher-order Lyapunov functions (HO-LFs). Firstly, to alleviate the conservatism of controller design, an advanced high-order time-varying relaxation matrix (HO-TVRM) approach is proposed. The HO-TVRM approach can completely adapt to the homogeneous polynomial analysis framework and effectively utilize the potential characteristics of fuzzy control to broaden the feasible domain of the reference problem. Secondly, to deal with the non-convex constraints caused by SOF control, a more flexible switching sequence convex optimization (SSCO) algorithm is developed, which adopts a switching-type optimization variable structure and combines variable-weight-based switching mechanism to alternately linearize non-convex terms during the iteration process. As a result, a convex optimization problem that approximates the original problem is obtained. Finally, simulations are presented to testify the effectiveness of the relaxed SOF control strategy.
Xiangpeng Xie 0001, Jianwei Xia, Vladimir Stojanovic
IEEE Trans Autom. Sci. Eng.4
2025 ADP-Based Prescribed-Time Control for Nonlinear Time-Varying Delay Systems With Uncertain Parameters
abstract
In this paper, we investigate the problem of prescribed-time optimal control using reinforcement learning technology. Unlike finite/fixed-time control methods that only achieve stability within specified time bounds, we propose a prescribed-time adaptive dynamic programming (ADP) control approach that ensures both optimality and prescribed-time stability. To address the challenge of solving the nonlinear Hamilton-Jacobi-Bellman (HJB) equation in finite horizon, we construct an actor-critic neural network (NN) with a time-varying activation function. The novel weight update laws are derived from the system’s terminal error and the approximate error of the HJB equation. This derivation eliminates the need for knowledge of dynamic conditions while ensuring compliance with terminal constraints. Based on the proposed prescribed time stability criterion, the control scheme is proven to satisfy prescribed time stability while also ensuring optimal system performance index and bounded weights. We apply the designed control scheme in a time-varying delay system and simulation examples validate the efficacy of the strategy.Note to Practitioners—Many industrial processes are nonlinear time-varying delay systems, which brings great challenges to solving the HJB equation of the prescribed-time control problem. Therefore, it is of great value to solve the prescribed-time control problem by using the advantages of finite time ADP in solving the time-varying HJB equation. The ADP-based prescribed-time optimal control (ADPPTC) scheme designed in this paper aims to optimize the steady-state performance of nonlinear time-varying delay systems, and because the user can prescribe the stability time, the system can accurately complete a task within a specific time range. At the same time, the prior demand for the system state can be eliminated by the proposed actor-critic neural network.
Kun Zhang 0005, Xiangpeng Xie 0001, Vladimir Stojanovic
IEEE Trans Autom. Sci. Eng.4
2024 Fuzzy wavelet neural adaptive finite-time self-triggered fault-tolerant control for a quadrotor unmanned aerial vehicle with scheduled performance
Xiaona Song, Chenglin Wu 0002, Shuai Song, Vladimir Stojanovic, Inés Tejado
Eng. Appl. Artif. Intell.4
2024 Autoregressive data generation method based on wavelet packet transform and cascaded stochastic quantization for bearing fault diagnosis under unbalanced samples
Yawei Sun, Hongfeng Tao, Vladimir Stojanovic
Eng. Appl. Artif. Intell.3
2024 Saturated-threshold event-triggered adaptive global prescribed performance control for nonlinear Markov jumping systems and application to a chemical reactor model
Xiaona Song, Shuai Song, Vladimir Stojanovic
Expert Syst. Appl.4
2024 Dynamic Event-Triggered Consensus Control for Interval Type-2 Fuzzy Multi-Agent Systems
abstract
This study investigates a fuzzy consensus problem for nonlinear multi-agent systems (MASs) with parameter uncertainties via dynamic event-triggered mechanism. The interval type-2 (IT2) fuzzy systems are used to model MASs, and a dynamic event-triggered fuzzy controller is presented. A Lyapunov-Krasovskii functional is constructed in which the sampling property is included, based on which the sufficient condition is determined for dynamic event-triggered fuzzy controller, while all agents are guaranteed to be consensus exponentially. Finally, the MASs of tunnel diode circuit and one-link manipulator are used to verify the efficiency of the presented control method, respectively.
Zhenbin Du, Xiangpeng Xie 0001, Zifang Qu, Yangyang Hu, Vladimir Stojanovic
IEEE Trans. Circuits Syst. I Regul. Pap.5
2024 Relaxed Model Predictive Control of T-S Fuzzy Systems via a New Switching-Type Homogeneous Polynomial Technique
abstract
This paper proposes a novel approach to model predictive control (MPC) for Takagi-Sugeno (T-S) fuzzy systems by combining the homogeneous polynomial technique and a switching mechanism. Aiming at improving the overall system performance by dynamically switching between different controllers according to system dynamics changing, the proposed switching MPC (SMPC) scheme formulates a group of optimization problems to determine the optimal switching strategy and corresponding control actions. To be specific, the optimization problem is solved at each moment to generate a control sequence that optimizes the gain matrices to achieve the minimized performance index. Building upon the traditional MPC structure, the proposed SMPC method forms a new framework by reducing the conservatism of controller design in the solving process based on linear matrix inequality (LMI) conditions via introducing the homogeneous polynomial method. As a result, the proposed SMPC method enhances the feasibility of controller solving, improves control performance and expands the domain of attraction by utilizing the switching mechanism to handle the system dynamics more freely. Finally, the effectiveness and superiority of the proposed method are validated through different examples.
Wenwen You, Xiangpeng Xie 0001, Hui Wang 0109, Jianwei Xia, Vladimir Stojanovic
IEEE Trans. Fuzzy Syst.5
2023 Driving Compute Scale-out Performance with Optical I/O Chiplets in Advanced System-in-Package Platforms
Mark Wade, Chen Sun 0003, Matthew N. Sysak, Vladimir Stojanovic, Pooya Tadayon, Ravi Mahajan, Babak Sabi
HCS4
2023 Bipartite synchronization for cooperative-competitive neural networks with reaction-diffusion terms via dual event-triggered mechanism
Xiaona Song, Nana Wu, Shuai Song, Yijun Zhang 0001, Vladimir Stojanovic
Neurocomputing5
2023 Self-triggered finite-time control for discrete-time Markov jump systems
Haiying Wan, Xiaoli Luan, Vladimir Stojanovic, Fei Liu 0001
Inf. Sci.3
2023 Quantized neural adaptive finite-time preassigned performance control for interconnected nonlinear systems
Xiaona Song, Shuai Song, Vladimir Stojanovic
Neural Comput. Appl.4
2023 Switching-Like Event-Triggered State Estimation for Reaction-Diffusion Neural Networks Against DoS Attacks
Xiaona Song, Nana Wu, Shuai Song, Vladimir Stojanovic
Neural Process. Lett.4
2023 Pretraining Graph Neural Networks for Few-Shot Analog Circuit Modeling and Design
abstract
Being able to predict the performance of circuits without running expensive simulations is a desired capability that can catalyze automated design. In this article, we present a supervised pretraining approach to learn circuit representations that can be adapted to new circuit topologies or unseen prediction tasks. We hypothesize that if we train a neural network (NN) that can predict the output direct current (dc) voltages of a wide range of circuit instances it will be forced to learn generalizable knowledge about the role of each circuit element and how they interact with each other. The dataset for this supervised learning objective can be easily collected at scale since the required dc simulation to get ground truth labels is relatively cheap. This representation would then be helpful for few-shot generalization to unseen circuit metrics that require more time-consuming simulations for obtaining the ground-truth labels. To cope with the variable topological structure of different circuits we describe each circuit as a graph and use graph NNs (GNNs) to learn node embeddings. We show that pretraining GNNs on prediction of output node voltages can encourage learning representations that can be adapted to new unseen topologies or prediction of new circuit-level properties with up to 10x more sample efficiency compared to a randomly initialized model. We further show that we can improve the sample efficiency of prior SoTA model-based optimization methods by$2\times $(almost as good as using an oracle model) via fintuning pretrained GNNs as the feature extractor of the learned models.
Kourosh Hakhamaneshi, Marcel Nassar, Mariano Phielipp, Pieter Abbeel, Vladimir Stojanovic
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2023 An Optimal Iterative Learning Control Approach for Linear Systems With Nonuniform Trial Lengths Under Input Constraints
abstract
In practical applications of iterative learning control (ILC), the repetitive process may end up early by accident during the performance improvement along the trial axis, which yields the nonuniform trial length problem. For such practical systems, input signals are usually constrained because of some certain physical limitations. This article proposes an optimal ILC algorithm for linear time-invariant multiple-input–multiple-output (MIMO) systems with nonuniform trial lengths under input constraints. The optimal ILC framework is specifically modified for the nonuniform trial length problem, where the primal–dual interior point method is introduced to deal with the input constraints. Hence, the constraint handling capability are improved compared with the conventional counterparts for nonuniform trial lengths. Also, the monotonic convergence property of the proposed optimal ILC algorithm is obtained in the sense of mathematical expectation. Finally, the effectiveness of the proposed algorithm is verified on the numerical simulation of a mobile robot.
Zhihe Zhuang, Hongfeng Tao, Yiyang Chen 0001, Vladimir Stojanovic, Wojciech Paszke
IEEE Trans. Syst. Man Cybern. Syst.4
2022 Fuzzy Fault Detection for Markov Jump Systems With Partly Accessible Hidden Information: An Event-Triggered Approach
abstract
This article addresses the design issue of fuzzy asynchronous fault detection filter (FAFDF) for a class of nonlinear Markov jump systems by an event-triggered (ET) scheme. The ET scheme can be applied to cut down the transmission times from the system to FAFDF. It is assumed that the system modes cannot be obtained synchronously by the filter, and instead, there is a detector that can measure the estimated modes of the system. The asynchronous phenomenon between the system and the filter is characterized via a hidden Markov model with partly accessible mode detection probabilities. Applying the Lyapunov function methods, sufficient conditions for the presence of FAFDF are obtained. Finally, an application of a wheeled mobile manipulator with hybrid joints is employed to clarify that the devised FAFDF can detect the faults without any incorrect alarm.
Peng Cheng 0010, Shuping He, Vladimir Stojanovic, Xiaoli Luan, Fei Liu 0001
IEEE Trans. Cybern.3
2022 Asynchronous Fault Detection Observer for 2-D Markov Jump Systems
abstract
In this article, the problem of the asynchronous fault detection (FD) observer design is discussed for 2-D Markov jump systems (MJSs) expressed by a Roesser model. In general, the FD observer cannot work synchronously with the system, that is, the mode of the observer varies with the mode of the system in line with some conditional transitional probabilities. For dealing with this difficult point, a hidden Markov model (HMM) is employed. Then, combining the$H_{\infty }$attenuation index and$H_{\_{}}$increscent index, a multiobjective solution to the FD problem is formed. In terms of linear matrix inequality technology, sufficient conditions are gained to guarantee the existence of the asynchronous FD. Simultaneously, an asynchronous FD algorithm is generated to acquire the optimal performance indices. Finally, a numerical example concerned with the Darboux equation is demonstrated to exhibit the soundness of the developed approach.
Peng Cheng 0010, Hai Wang 0004, Vladimir Stojanovic, Shuping He, Kaibo Shi, Xiaoli Luan, Fei Liu 0001, Changyin Sun 0001
IEEE Trans. Cybern.3
2022 Asynchronous Fault Detection for Interval Type-2 Fuzzy Nonhomogeneous Higher Level Markov Jump Systems With Uncertain Transition Probabilities
abstract
Based on the interval type-2 fuzzy (IT2F) approach, this article investigates the fault detection filter design problem for a class of nonhomogeneous higher level Markov jump systems with uncertain transition probabilities. Considering that the mode information of the system cannot be obtained synchronously by the filter, the hidden Markov model can be seen as a detector to handle this asynchronous problem, and the parameter uncertainty can be processed by the IT2F approach with the lower and upper membership functions. Then, the asynchronous IT2F filter is designed to deal with the fault detection problem. Furthermore, the Gaussian transition probability density function is introduced to describe the uncertainty transition probabilities of the system and the filter. Based on the Lyapunov theory, the existence of the designed asynchronous IT2F filter and the dissipativity of the filter error system can be well ensured. In this article, the simulation study on a quarter-car suspension system verifies that the designed asynchronous IT2F filter can detect faults without error alarms.
Hai Wang 0004, Vladimir Stojanovic, Peng Cheng 0010, Shuping He, Xiaoli Luan, Fei Liu 0001
IEEE Trans. Fuzzy Syst.3
2021 Finite-time asynchronous dissipative filtering of conic-type nonlinear Markov jump systems
Shuping He, Vladimir Stojanovic, Xiaoli Luan, Fei Liu 0001
Sci. China Inf. Sci.3
2020 Event-based fuzzy control for T-S fuzzy networked systems with various data missing
Ziran Chen, Baoyong Zhang, Vladimir Stojanovic, Yijun Zhang 0001, Zhengqiang Zhang
Neurocomputing3
2019 Analog Circuit Generator based on Deep Neural Network enhanced Combinatorial Optimization
abstract
A deep neural network (DNN) based stochastic combinatorial optimization framework is presented that can find the optimal sizing of circuits in a sample-efficient manner. This sample efficiency allows us to unify this framework with generator-based tools like Berkeley Analog Generator (BAG) [1] to directly optimize layout, given the high level circuit specifications. We use this tool to design an optical link receiver layout, satisfying high-level design specifications, using post-layout simulations of only 348 design instances. Compared to an evolutionary algorithm without our DNN-based discriminator, our framework improves the sample efficiency and run time by more than 200x.
Kourosh Hakhamaneshi, Nick Werblun, Pieter Abbeel, Vladimir Stojanovic
DAC4
2019 BagNet: Berkeley Analog Generator with Layout Optimizer Boosted with Deep Neural Networks
abstract
The discrepancy between post-layout and schematic simulation results continues to widen in analog design due in part to the domination of layout parasitics. This paradigm shift is forcing designers to adopt design methodologies that seamlessly integrate layout effects into the standard design flow. Hence, any simulation-based optimization framework should take into account time-consuming post-layout simulation results. This work presents a learning framework that learns to reduce the number of simulations of evolutionary-based combinatorial optimizers, using a DNN that discriminates against generated samples, before running simulations. Using this approach, the discriminator achieves at least two orders of magnitude improvement on sample efficiency for several large circuit examples including an optical link receiver layout.
Kourosh Hakhamaneshi, Nick Werblun, Pieter Abbeel, Vladimir Stojanovic
ICCAD4
2018 Optimal Spectral Estimation and System Trade-Off in Long-Distance Frequency-Modulated Continuous-Wave Lidar
abstract
Frequency-modulated continuous-wave (FMCW) LIDAR is a promising technology for next-generation integrated 3D imaging systems. However, it has been considered difficult to apply FMCW LIDAR for long-distance (> 100m) targets, such as those in automotive and airborne applications. Maintaining coherence between the reflected beam from the target and locally forwarded beam becomes a significant challenge for tunable laser design. This paper demonstrates the possibility of extending the detection range of FMCW LIDAR beyond the coherence range of its laser by improving the spectral estimation algorithm. By exploiting the Lorentzian prior of the received signal in the spectral domain, > 10x improvement in ranging accuracy is achieved compared to traditional algorithms that do not consider phase noise in the signal model. In light of this finding, the end-to-end modeling framework is presented to examine true system-level trade-offs of FMCW LIDAR and the feasibility of long -distance measurement.
Taehwan Kim 0010, Pavan Bhargava, Vladimir Stojanovic
ICASSP3
2015 DELPHI: a framework for RTL-based architecture design evaluation using DSENT models
abstract
Computer architects are increasingly interested in evaluating their ideas at the register-transfer level (RTL) to gain more precise insights on the key characteristics (frequency, area, power) of a micro/architectural design proposal. However, the RTL synthesis process is notoriously tedious, slow, and errorprone and is often outside the area of expertise of a typical computer architect, as it requires familiarity with complex CAD flows, hard-to-get tools and standard cell libraries. The effort is further multiplied when targeting multiple technology nodes and standard cell variants to study technology dependence. This paper presents DELPHI, a flexible, open framework that leverages the DSENT modeling engine for faster, easier, and more efficient characterization of RTL hardware designs. DELPHI first synthesizes a Verilog or VHDL RTL design (either using the industry-standard Synopsys Design Compiler tool or a combination of open-source tools) to an intermediate structural netlist. It then processes the resulting synthesized netlist to generate a technology-independent DSENT design model. This model can then be used within a modified version of the DSENT flow to perform very fast—one to two orders of magnitude faster than full RTL synthesis—estimation of hardware performance characteristics, such as frequency, area, and power across a variety of DSENT technology models (e.g., 65nm Bulk, 32nm SOI, 11nm Tri-Gate, etc.). In our evaluation using 26 RTL design examples, DELPHI and DSENT were consistently able to closely track and capture design trends of conventional RTL synthesis results without the associated delay and complexity. We are releasing the full DELPHI framework (including a fully open-source flow) at http://www.ece.cmu.edu/CALCM/delphi/.
Michael Papamichael, Cagla Cakir, Chen Sun 0003, Chia-Hsin Owen Chen, James C. Hoe, Ken Mai, Li-Shiuan Peh, Vladimir Stojanovic
ISPASS8
2014 Serbia Forum - Digital Cultural Heritage Portal
Aleksandar Mihajlovic, Vladisav Jelisavcic, Bojan Marinkovic, Milan Todorovic, Zoran Ognjanovic, Sinisa Tomovic, Vladimir Stojanovic, Veljko M. Milutinovic
ICISP7
2013 Relays do not leak: CMOS does
abstract
This paper describes the micro-architectural and circuit design techniques for building complex VLSI circuits with micro-electromechanical (MEM) relays and presents experimental results to demonstrate the viability of this technology. By tailoring the circuits and micro-architecture to the relay device characteristics, the performance of the relay-based multiplier is improved by an order of magnitude over any known static CMOS style implementation, and by ~4x over CMOS pass-gate equivalent implementations. A 16-bit relay multiplier is shown to offer ~10x lower energy per operation at sub-10 MOPS throughputs when compared to an optimized CMOS multiplier at an equivalent 90 nm technology node. The functionality of the primary multiplier building block, a full (7:3) compressor built with 46 scaled MEM-relays, which is the largest working MEM-relay circuit reported to date, is also demonstrated.
Hossein Fariborzi, Fred Chen, Rhesa Nathanael, I-Ru Chen, Louis Hutin, Rinus Lee, Tsu-Jae King Liu, Vladimir Stojanovic
DAC8
2013 Breaking the energy barrier in fault-tolerant caches for multicore systems
abstract
Balancing cache energy efficiency and reliability is a major challenge for future multicore system design. Supply voltage reduction is an effective tool to minimize cache energy consumption, usually at the expense of increased number of errors. To achieve substantial energy reduction without degrading reliability, we propose an adaptive fault-tolerant cache architecture, which provides appropriate error control for each cache line based on the number of faulty cells detected at reduced supply voltages. Our experiments show that the proposed approach can improve energy efficiency by more than 25% and energy-execution time product by over 10%, while improving reliability up to 4X using Mean-Error-To-Failure (METF) metric, compared to the next-best solution at the cost of 0.08% storage overhead.
Paul Ampadu, Meilin Zhang, Vladimir Stojanovic
DATE3
2013 Hardware-Software Codesign for Embedded Numerical Acceleration
abstract
In this work we aim to strike a balance between performance, power consumption and design effort for complex digital signal processing within the power and size constraints of embedded systems. Looking across the design stack, from algorithm formulation down to accelerator microarchitecture, we show that a high degree of flexibility and design reuse can be achieved without much performance sacrifice. The foundation of our design is a numerical accelerator template. Extensively parameterized, it allows us to develop the design while postponing microarchitectural decisions until program is known. Statically scheduling compiler provides a link between the algorithm and template instantiation parameters. Results show that the derived design can significantly outperform embedded processors for similar power cost and also approach the high-performance processor performance for a fraction of the power cost.
Ranko Sredojevic, Andrew Wright, Vladimir Stojanovic
FCCM3
2012 Performance trade-offs and design limitations of analog-to-information converter front-ends
abstract
This paper evaluates the impact of circuit impairments on the energy cost and performance limitations of analog-to-information converters (AIC). In applications where signal frequencies are high, but information bandwidths are low, AICs have been proposed as a potential solution to overcome the resolution and performance limitations of sampling jitter in high-speed analog-to-digital converters (ADC). Although the AIC architecture facilitates slower ADCs, the signal encoding, typically realized with a mixer-like circuit, still occurs at the Nyquist frequency of the input to avoid aliasing. We show that the jitter of this mixing stage limits the achievable AIC resolution. In this work, the end-to-end system evaluation framework is designed to analyze these limitations as well as the relative energy-efficiency of AICs versus ADCs across the resolution, receiver gain and signal sparsity. The evaluation shows that AICs improve the resolution by 1 bit when the signal of interest is very sparse, and enable 2× in energy savings when no pre-amplification is required.
Omid Abari, Fred Chen, Fabian Lim, Vladimir Stojanovic
ICASSP4
2012 Non-asymptotic analysis of compressed sensing random matrices: An U-statistics approach
abstract
We apply Heoffding's U-statistics to obtain non-asymptotic analysis for compressed sensing (CS) random matrices. These powerful (U-statistics) tools appear to apply naturally to CS theory, in particular here we focus on one particular large deviation result. We chose two applications to outline how U-statistics may apply to various CS recovery guarantees. Pros, cons, and further directions of the approach, are discussed. Restricted isometries of random matricies have well-regarded importance in CS. They guarantee i) uniqueness of sparse solutions, and ii) robust recovery. The fraction of size-k submatrices (out of all (n/k) of them), that satisfy CS-type restricted isometeries, is an U-statistic. Concentration of U-statistics predict the “average-case” behavior of such isometries. U-statistics related to Fuchs' conditions for ℓ1-minimization support recovery, are derived. This leads to bounds on the fraction of recoverable k-supports. Empirically, we observe significant improvement over a recent large deviation (non-asymptotic) bound by Donoho & Tanner, for some practical system sizes with large undersampling. The results apply regardless of column distribution, e.g. Gaussian, Bernoulli, etc. Similar concentration behavior has been empirically observed, when the sampling matrix is constructed using pseudorandom sequences (important in practice).
Fabian Lim, Vladimir Stojanovic
ICC2
2012 Cross-layer Energy and Performance Evaluation of a Nanophotonic Manycore Processor System Using Real Application Workloads
abstract
Recent advances in nanophotonic device research have led to a proliferation of proposals for new architectures that employ optics for on-chip communication. However, since standard simulation tools have not yet caught up with these advances, the quality and thoroughness of the evaluations of these architectures have varied widely. This paper provides the first complete end-to-end analysis of an architecture using on-chip optical interconnect. This analysis incorporates realistic performance and energy models for both electrical and optical devices and circuits into a full-fledged functional simulator, thus enabling detailed analyses when running actual applications. Since on-chip optics is not yet mature and unlikely to see widespread use for several more years, we perform our analysis on a future 1000-core processor implemented in an 11nm technology node. We find that the proposed optical interconnect can provide between 1.8x and 4.8x better energy-delay product than conventional electrical-only interconnects. In addition, based on a detailed energy breakdown of all processor components, we conclude that a thermal ring resonators and on-chip lasers that allow rapid power gating are key areas worthy of additional nanophotonic research. This will help guide future optical device research to the areas likely to provide the best payoff.
George Kurian, Chen Sun 0003, Chia-Hsin Owen Chen, Jason E. Miller, Jürgen Michel, Dimitri A. Antoniadis, Li-Shiuan Peh, Lionel C. Kimerling, Vladimir Stojanovic, Anant Agarwal
IPDPS10
2012 DSENT - A Tool Connecting Emerging Photonics with Electronics for Opto-Electronic Networks-on-Chip Modeling
abstract
With the rise of many-core chips that require substantial bandwidth from the network on chip (NoC), integrated photonic links have been investigated as a promising alternative to traditional electrical interconnects. While numerous opto-electronic NoCs have been proposed, evaluations of photonic architectures have thus-far had to use a number of simplifications, reflecting the need for a modeling tool that accurately captures the tradeoffs for the emerging technology and its impacts on the overall network. In this paper, we present DSENT, a NoC modeling tool for rapid design space exploration of electrical and opto-electrical networks. We explain our modeling framework and perform an energy-driven case study, focusing on electrical technology scaling, photonic parameters, and thermal tuning. Our results show the implications of different technology scenarios and, in particular, the need to reduce laser and thermal tuning power in a photonic network due to their non-data-dependent nature.
Chen Sun 0003, Chia-Hsin Owen Chen, George Kurian, Jason E. Miller, Anant Agarwal, Li-Shiuan Peh, Vladimir Stojanovic
NOCS8
2010 Silicon photonics and memories
Vladimir Stojanovic
Hot Chips Symposium1
2010 Re-architecting DRAM memory systems with monolithically integrated silicon photonics
abstract
The performance of future manycore processors will only scale with the number of integrated cores if there is a corresponding increase in memory bandwidth. Projected scaling of electrical DRAM architectures appears unlikely to suffice, being constrained by processor and DRAM pin-bandwidth density and by total DRAM chip power, including off-chip signaling, cross-chip interconnect, and bank access energy. In this work, we redesign the DRAM main memory system using a proposed monolithically integrated silicon photonics technology and show that our photonically interconnected DRAM (PIDRAM) provides a promising solution to all of these issues. Photonics can provide high aggregate pin-bandwidth density through dense wavelength-division multiplexing. Photonic signaling provides energy-efficient communication, which we exploit to not only reduce chip-to-chip interconnect power but to also reduce cross-chip interconnect power by extending the photonic links deep into the actual PIDRAM chips. To complement these large improvements in interconnect bandwidth and power, we decrease the number of bits activated per bank to improve the energy efficiency of the PIDRAM banks themselves. Our most promising design point yields approximately a 10x power reduction for a single-chip PIDRAM channel with similar throughput and area as a projected future electrical-only DRAM. Finally, we propose optical power guiding as a new technique that allows a single PIDRAM chip design to be used efficiently in several multi-chip configurations that provide either increased aggregate capacity or bandwidth.
Scott Beamer, Chen Sun 0003, Yong-Jin Kwon, Ajay Joshi, Christopher Batten, Vladimir Stojanovic, Krste Asanovic
ISCA6
2010 Compact Modeling of Nonlinear Analog Circuits Using System Identification via Semidefinite Programming and Incremental Stability Certification
abstract
This paper presents a system identification technique for generating stable compact models of typical analog circuit blocks in radio frequency systems. The identification procedure is based on minimizing the model error over a given training data set subject to an incremental stability constraint, which is formulated as a semidefinite optimization problem. Numerical results are presented for several analog circuits, including a distributed power amplifier, as well as a MEM device. It is also shown that our dynamical models can accurately predict important circuit performance metrics, and may thus, be useful for design optimization of analog systems.
Bradley N. Bond, Zohaib Mahmood, Yan Li 0029, Ranko Sredojevic, Alexandre Megretski, Vladimir Stojanovic, Yehuda Avniel, Luca Daniel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2009 Yield-driven iterative robust circuit optimization algorithm
abstract
This paper proposes an equation-based multi-scenario iterative robust optimization methodology for analog/mixed-signal circuits. We show that due to local circuit performance monotonicity in random variations constraint maximization can be used to efficiently find critical constraints and worst-case scenarios of random process variations and populate them into a multi-scenario optimization. This algorithm scales gracefully with circuit size and is tested on both two-stage and fully differential folded-cascode operational amplifiers with a 90 nm predictive model. The improving yield-trends are confirmed across process and random variations with Hspice Monte-Carlo simulations.
Yan Li 0029, Vladimir Stojanovic
DAC2
2009 Designing multi-socket systems using silicon photonics
abstract
Future single-board multi-socket systems may be unable to deliver the needed memory bandwidth electrically due to power limitations, which will hurt their ability to drive performance improvements. Energy efficient off-chip silicon photonics could be used to deliver the needed bandwidth, and it could be extended on-chip to create a relatively flat network topology. That flat network may make it possible to implement the same number of cores with a greater number of smaller dies for a cost advantage with negligible performance degradation.
Scott Beamer, Krste Asanovic, Christopher Batten, Ajay Joshi, Vladimir Stojanovic
ICS5
2009 Silicon-photonic clos networks for global on-chip communication
abstract
Future manycore processors will require energy-efficient, high-throughput on-chip networks. Silicon-photonics is a promising new interconnect technology which offers lower power, higher bandwidth density, and shorter latencies than electrical interconnects. In this paper we explore using photonics to implement low-diameter non-blocking crossbar and Clos networks. We use analytical modeling to show that a 64-tile photonic Clos network consumes significantly less optical power, thermal tuning power, and area compared to global photonic crossbars over a range of photonic device parameters. Compared to various electrical on-chip networks, our simulation results indicate that a photonic Clos network can provide more uniform latency and throughput across a range of traffic patterns while consuming less power. These properties will help simplify parallel programming by allowing the programmer to ignore network topology during optimization.
Ajay Joshi, Christopher Batten, Yong-Jin Kwon, Scott Beamer, Imran Shamim, Krste Asanovic, Vladimir Stojanovic
NOCS7
2009 A Modeling and exploration framework for interconnect network design in the nanometer era
abstract
As we approach serious scaling roadblocks in the next few process nodes, it is imperative to identify new emerging technologies that can complement or supplant CMOS in the future. We present an integrated cyclic approach to explore new interconnect technologies in the nanometer era for manycore systems, where on-chip interconnects are jointly optimized at all the levels in the design hierarchy to develop a complete interconnect solution - from interconnect technology to network topology.
Ajay Joshi, Fred Chen, Vladimir Stojanovic
NOCS3
2008 Low-Complexity Pattern-Eliminating Codes for ISI-Limited Channels
abstract
This paper introduces low-complexity block codes, termed pattern-eliminating codes (PEC), which achieve a potentially large performance improvement over channels with residual inter-symbol interference (ISI). The codes are systematic, require no decoding and allow for simple encoding. They operate by prohibiting the occurrence of harmful symbol patterns. On some discrete-time communication channels, the (n, n - 1) PEC can prohibit all occurrences of symbol patterns causing worst- case ISI. The effectiveness of a PEC is shown to be uniquely determined by the sign-signature of the channel response, and a simple criterion is given for identifying channels for which the (n, n - 1) code is effective. It is also shown that for most channel signatures, the (n, n - 1) PEC can be augmented by a (O, n - 1) runlength-limiting (RLL) code at no additional coding overhead. This paper also explores properties of the (n, n-b) PEC for b > 1. The simulation results show that the (n, n - 1) PEC can provide error-rate reductions of several orders of magnitude, even with rate penalty taken into account. It is also shown that channel conditioning, such as equalization, can have a large effect on the code performance and potentially large gains can be derived from optimizing the equalizer jointly with a pattern-eliminating code.
Natasha Blitvic, Lizhong Zheng, Vladimir Stojanovic
ICC3
2008 Integrated circuit design with NEM relays
abstract
To overcome the energy-efficiency limitations imposed by finite sub-threshold slope in CMOS transistors, this paper explores the design of integrated circuits based on nano-electro-mechanical (NEM) relays. A dynamical Verilog-A model of the NEM relay is described and correlated to device measurements. Using this model we explore NEM relay design strategies for digital logic and I/O that can significantly improve the energy efficiency of the whole VLSI system. By exploiting the low effective threshold voltage and zero leakage achievable with these relays, we show that NEM relay-based adders can achieve an order of magnitude or more improvement in energy efficiency over CMOS adders with ns-range delays and with no area penalty. By applying parallelism, this improvement in energy-efficiency can be achieved at higher throughputs as well, at the cost of increased area. Similar improvements in high-speed I/O energy are also predicted by making use of the relays to implement highly energy-efficient digital-to-analog and analog-to-digital converters.
Fred Chen, Hei Kam, Dejan Markovic, Tsu-Jae King Liu, Vladimir Stojanovic, Elad Alon
ICCAD5
2008 Optimization-based framework for simultaneous circuit-and-system design-space exploration: a high-speed link example
abstract
Connecting system-level performance models with circuit information has been a long-standing problem in analog/mixed-signal front-ends, like radios and high-speed links. High-speed links are particularly hard to analyze because of the complex interplay of device/circuit parasitics and channel filtering operation. In this paper we introduce optimization-based framework for link design-space exploration, connecting the link transmission quality and top-level filter settings with circuit power, sizing and biasing. We derive a special analytical discrete time representation that avoids the size explosion of the symbolic problem description improving the parsing and solver time by orders of magnitude and making this joint optimization possible in real-time. This robust and accurate problem formulation is derived in signomial form and is compatible with existing optimization approaches to circuit sizing. We demonstrate this optimization framework on a link design-space exploration example, investigating trade-offs between the transmit pre-emphasis and linear receiver equalizer and their impact on overall link power vs. data rate.
Ranko Sredojevic, Vladimir Stojanovic
ICCAD2
2007 Practical Limits of Multi-Tone Signaling Over High-Speed Backplane Electrical Links
abstract
Application of Discrete Multi-tone (DMT) signaling to high-speed backplane interconnects requires major modifications to the well-known analysis methods applied to wireline communication systems. Tight power budgets in backplane links impose severe constraints on DMT block size and use of channel shortening filters in the system. Consequently, maximum throughput is achieved in a DMT system that is (residual) interference limited and water-filling is not applicable in its original form. In this paper, the DMT system is cast as a Second Order Conic (SOQ problem with peak transmit power as the constraint and optimum integer bit-loading and power allocation are achieved through a novel incremental integer bit-loading algorithm. The convex framework is subsequently used to find "practical" upper bounds on the performance of multi-tone signaling over highspeed links. The results indicate that DMT has the potential to achieve 15-23 Gbps over typical backplane channels and system requirements in terms of FFT block size and prefix length are obtained.
Amir Amirkhany, Aliazam Abbasfar, Vladimir Stojanovic, Mark Horowitz
ICC3
2007 Statistical Simulator for Block Coded Channels with Long Residual Interference
abstract
In this paper, we present a simple statistical simulation technique for channels with long memory operating under coded data. The proposed technique employs a divide-and-conquer approach, both at the level of the channel response and the individual codeword. We introduce an efficient algorithm for parity tracking during the generation and combining of the voltage distributions computed in individual sub-problems. The resulting computational complexity increases linearly both with the channel and the codeword length, keeping the number of parity bits constant. The complexity increases exponentially with the number of parity bits in a codeword. Thus, the technique is of most use for high-rate codes. Finally, the technique is applied to examine the effect of coding on high-speed links which operate at BERs smaller than 1e-15 and where residual interference typically spans several hundreds of bits. This simulator enables systematic evaluation of power-performance trade-offs in introducing error correction/detection codes into high-speed link systems.
Natasha Blitvic, Vladimir Stojanovic
ICC2
2007 Equalized interconnects for on-chip networks: modeling and optimization framework
abstract
This paper presents a modeling framework for fast design space exploration and optimization of equalized on-chip interconnects. The exploration is enabled by cross-layer modeling that connects the transistor and wire parameters to link performance, equalization coefficients, and architecture- friendly metrics (delay, energy-per-bit, and throughput density). Appropriate models are derived to speed-up the search by more than two orders of magnitude and make a million point design space searchable in less than two hours on a standard machine. With this approach we are able to find the best link design for target throughput, power and area constraints, thus enabling the architectural optimization of energy-efficient on-chip networks. For the same latency and throughput density, equalized interconnects optimized using the new methodology have up to 10x better energy-efficiency than optimized repeater interconnects.
Byungsub Kim, Vladimir Stojanovic
ICCAD2
2006 Power-centric design of high-speed I/Os
abstract
With increasing aggregate off-chip bandwidths exceeding terabits/second (Tb/s), the power dissipation is a serious design consideration. Additionally, design of I/O links is constrained by a complex set of specifications such as voltage levels, voltage noise, signal deterministic jitter, random jitter, slew rate, BER etc. These specifications lead to complex tradeoffs for both circuits and circuit architecture in order to minimize power. This paper presents a design framework that enables the analysis of tradeoffs in the design of an I/O transmitter. The design framework includes BER analysis with a channel model coupled with logic sizing optimization that is constrained by the desired signaling specification.
Hamid Hatamkhani, Frank Lambrecht, Vladimir Stojanovic, Chih-Kong Ken Yang
DAC3
2006 A cost-effective implementation of an ECC-protected instruction queue for out-of-order microprocessors
abstract
Major sources of transient errors in microprocessors today include noise and single event upsets. As feature sizes and voltages are reduced to create faster, more efficient, and computationally more powerful processors, these errors will increase significantly. We show that (contrary to conventional wisdom) error correction codes (ECC) can be efficiently utilized to handle these errors as instructions are being processed through the microprocessor pipeline. We will analyze some of the tradeoffs involved in a hardware implementation of ECC for the instruction queue with respect to performance, power, area, and reliability. Specifically, for an environment with high error rates, we show that we can correct all single bit errors with a negligible drop in performance. Our approach can be generalized to other data structures within the microprocessor, including the register file and reorder buffer.
Vladimir Stojanovic, R. Iris Bahar, Jennifer Dworak, Richard Weiss 0001
DAC1
2006 Analog Multi-Tone Signaling for High-Speed Backplane Electrical Links
abstract
Implementing a multi-tone (MT) architecture for high-speed backplane electrical links is difficult given the tight power and complexity constraints in this application. This paper proposes an approach that incorporates a baseband (BB) channel and a few passband (PB) channels. In this MT system inter- channel interference (ICI) and inter-symbol interference (ISI) are eliminated through fractionally spaced equalization at the transmitter and feedback equalization at the receiver. The design is modeled as a MIMO system, and optimal equalizer coefficients to minimize the transmit peak voltage are found by casting the optimization as a Second Order Conic (SOC) problem. In addition, for systems that need adaptation, we show how equalizer and power allocation coefficients can be obtained (sub optimally) using Zero Forcing (ZF) optimization. The effect of transmitter and receiver clock jitter are modeled in a way that can be included in both SOC and ZF optimizations, and the performance of this system is compared to more conventional baseband examples. It is shown that this AMT system can be built with complexity/power similar to a comparable performance baseband system, but has the ability to scale to higher bit rates.
Amir Amirkhany, Aliazam Abbasfar, Vladimir Stojanovic, Mark Horowitz
GLOBECOM3
2004 Equalization of modal dispersion in multimode fiber using spatial light modulators
abstract
Intersymbol interference (ISI) due to modal dispersion is the dominant limitation to the bit rate-distance product in multimode fiber-optic communication systems. If the light launched into the fiber excites only the desired principal modes, modal dispersion can be eliminated. We can achieve this by using spatial light modulators (SLMs) to perform adaptive spatial filtering on the electric fields of the light. In this paper, we develop an optimization framework for setting the SLMs to obtain an upper bound on the achievable performance and develop heuristics that nearly reach this upper bound Using this framework, we show that both a sophisticated semidefinite programming-based algorithm and a simple adaptive algorithm achieve performance close to the upper bound. Performance and system complexity tradeoff curves are constructed, showing that a 20/spl times/20 array of SLM pixels with binary phase control performs within 15% of more complex implementations. Finally, we extend the framework and present preliminary results showing the promise of further increases in the capabilities of multimode fiber by using the fiber as a multiple-input multiple-output (MIMO) transmission medium.
Elad Alon, Vladimir Stojanovic, Joseph M. Kahn, Stephen P. Boyd, Mark Horowitz
GLOBECOM2
2004 Multi-tone signaling for high-speed backplane electrical links
abstract
A multi-tone architecture is proposed for high-speed backplane serial links. To limit complexity, the links use analog multi-tone rather than the more modern DMT. The tradeoffs involved in the design of such a system are examined and the performance of a serial link based on this approach is compared to a baseband architecture in terms of data rate and complexity using a convex optimization framework. Slightly less than 2/spl times/ improvement in data rate at reasonable complexity is shown to be achievable with the proposed architecture.
Amir Amirkhany, Vladimir Stojanovic, Mark Horowitz
GLOBECOM2
2004 Optimal linear precoding with theoretical and practical data rates in high-speed serial-link backplane communication
abstract
Multi-Gb/s high-speed links face significant challenges in keeping up with the increase in desired data rates. In the evaluation of achievable data rates, it is necessary to include both link-specific noise sources and implementation driven constraints. We construct models of these noise sources and constraints in order to estimate the theoretical limits of typical high-speed link channels. In order to estimate the data rates of practical baseband architectures, we solve the power constrained optimal linear precoding problem and formulate a bit-error rate (BER) driven optimization, including all link-specific noise sources. The problem is shown to be quasiconcave, hence, a globally optimal solution is guaranteed. Using this optimization framework, we show that practical data rates are mainly limited by inter-symbol interference (ISI) due to complexity constraints on the number of precoder and equalizer taps. After these constraints are removed, we further show that slicer resolution and sampling jitter are limiting the higher bandwidth utilization provided by multi-level modulations. Better circuits are needed to improve the bandwidth utilization to more than 2bits/dimension in baseband.
Vladimir Stojanovic, Amir Amirkhany, Mark Horowitz
ICC1
2002 Transmit pre-emphasis for high-speed time-division-multiplexed serial-link transceiver
abstract
Time division multiplexing (TDM) must be employed in multi-Gb/s transceivers in order to overcome onchip clock frequency limitations. This paper describes a transmit pre-emphasis filter for a multi-level transceiver making use of TDM. The possible applications of such a transceiver include serial links and chip-to-chip communication. The requirement of very low probability of error in the absence of coding, and the need for an adaptive solution impose a peak transmit power constraint. The TDM system is mapped to a multiple-input-multiple-output (MIMO) system, and the noise sources are analyzed. The design of the pre-emphasis filter is shown to be a non-convex optimization problem, whose optimal solution is very difficult to obtain. Still, sub-optimal solutions are derived in closed form and adaptive implementations are described. Simulation results using parameters obtained from an experimental testbed indicate that these sub-optimal solutions actually achieve very good performance.
Vladimir Stojanovic, George Ginis, Mark Horowitz
ICC1
2002 Methods for true power minimization
abstract
This paper presents methods for efficient power minimization at circuit and micro-architectural levels. The potential energy savings are strongly related to the energy profile of a circuit. These savings are obtained by using gate sizing, supply voltage, and threshold voltage optimization, to minimize energy consumption subject to a delay constraint. The true power minimization is achieved when the energy reduction potentials of all tuning variables are balanced. We derive the sensitivity of energy to delay for each of the tuning variables connecting its energy saving potential to the physical properties of the circuit. This helps to develop understanding of optimization performance and identify the most efficient techniques for energy reduction. The optimizations are applied to some examples that span typical circuit topologies including inverter chains, SRAM decoders, and adders. At a delay of 20% larger than the minimum, energy savings of 40% to 70% are possible, indicating that achieving peak performance is expensive in terms of energy. Energy savings of about 50% can be achieved without delay penalty with the balancing of sizes, supplies, and thresholds.
Robert W. Brodersen, Mark Horowitz, Dejan Markovic, Borivoje Nikolic, Vladimir Stojanovic
ICCAD5
1998 Comparative analysis of latches and flip-flops for high-performance systems
abstract
In this paper we propose a set of rules for consistent estimation of the real performance and power features of the latch and flip-flop structures. A new simulation and optimization approach is presented, targeting both high-performance and power budget issues. The analysis approach reveals the sources of performance and power consumption bottlenecks in different design styles. Certain misleading parameters have been properly modified and weighted to reflect the real properties of the compared structures. Furthermore, the results of the comparison of representative latches and flip-flops illustrate the advantages of our approach and the suitability of different design styles for high-performance applications.
Vladimir Stojanovic, Vojin G. Oklobdzija, Raminder Singh Bajwa
ICCD1
1998 A unified approach in the analysis of latches and flip-flops for low-power systems
abstract
In this paper we propose a set of rules for consistent estimation of the real performance and power features of the latch and flip-flop structures. A new simulation and optimization approach is presented, targeting both high-performance and power budget issues. The analysis approach reveals the sources of performance and power consumption bottlenecks in different design styles. Certain misleading parameters have been properly modified and weighted to reflect the real properties of the compared structures. Furthermore, the results of the comparison of representative latches and flip-flops illustrate the advantages of our approach and the suitability of different design styles for low-power and high-performance applications.
Vladimir Stojanovic, Vojin G. Oklobdzija, Raminder Singh Bajwa
ISLPED1