EDBT 2026 Demo / reviewers in the wild / expert
Srinivasan Murali
dblp:43/6289
· DBLP profile ↗
40ranked-venue papers
16as first author
5since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 34 · 14 first-author · 1 since 2021Software engineering, systems software and programming languages · 8 · 4 first-authorSecurity and privacy · 4 · 2 first-author · 3 since 2021Computer networks · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OptiVibe: Keystroke Inference Attacks Through a New Optical-Vibration Side Channel
YoungTak Cho, Sanket Suresh Badgujar, Srinivasan Murali, Xuhao Xie, Ming Li 0006 |
ICDCS | 3 |
| 2025 | SnoopDog: Detecting USB Bus Sniffers Using Responsive EMRabstractThe lack of encryption and authentication mechanisms in USB standards renders USB traffic susceptible to sniffing attacks. This paper presents an initial effort to detect USB bus sniffing through the development of a detection system, SnoopDog. It does not require hardware redesign of USB devices or modifications to the kernel or USB protocol stack. The system utilizes a probe-and-detect strategy. The host PC generates bait traffic with a dummy endpoint address. While benign devices discard this traffic due to the address mismatch, a sniffer captures the data, consequently emitting responsive electromagnetic radiation (EMR). To determine whether a USB device is a sniffer, SnoopDog calculates the correlation between the bait traffic and the responsive EMR signals captured near the target device. A high correlation indicates the presence of a sniffer. Recognizing that sniffer's EMR signals can be weak, we introduce a novel temporal folding scheme to improve the signal-to-noise ratio (SNR). To evaluate the performance, we build a prototype of SnoopDog and conduct comprehensive evaluations under a variety of settings, where SnoopDog delivers a promising detection accuracy with minor system overhead. Srinivasan Murali, YoungTak Cho, Huadi Zhu, Pan Li 0001, Ming Li 0006 |
ACSAC | 1 |
| 2025 | Continuous User Authentication for Extended Reality Using Pupil Reflexive Mechanisms as a BiometricabstractWith the rapid adoption of extended reality (XR) technologies in both consumer and enterprise domains, continuous and unobtrusive user authentication has become increasingly important. Existing authentication methods are often intrusive, static, or insufficiently secure for immersive environments. In this work, we propose a novel passive authentication framework that leverages users' real-time pupil light reflex (PLR) in response to visual stimuli rendered in XR. By treating screen brightness as a natural, time-varying challenge and modeling the user's pupil response as the biometric signal, our system learns to extract identity-specific features that are invariant to environmental content. We implement our prototype on two commercial XR headsets and evaluate it through a user study involving eight participants across diverse XR applications. Our system achieves an equal error rate (EER) of 0.093 with a 2-minute prediction window. These results demonstrate the feasibility of pupillary dynamics as a behavioral biometric for secure, continuous authentication in immersive environments. This study lays the foundation for future work on scalable, multimodal, and adaptive biometric authentication in XR. Shuaikang Hou, Muyao Tang, Srinivasan Murali, Huadi Zhu |
MobiHoc | 3 |
| 2023 | Continuous Authentication Using Human-Induced Electric PotentialabstractMost terminal devices authenticate users only once at the time of initial login, leaving the terminal unprotected during an active session when the original user leaves it unattended. To address this issue, continuous authentication has been proposed by automatically locking the terminal after a period of inactivity. However, it does not fully eliminate the risk of unauthorized access before the session expires. Recent research has also investigated the feasibility of using physiological and behavioral patterns as biometrics. This study presents a novel two-factor continuous authentication that explores a new form of signal called human-induced electric potential captured by wearables in contact with the user’s body. By analyzing this signal, we can determine the time of user-terminal interactions and compare it with information recorded by the terminal’s OS. If the original user remains on the same terminal, the two-source readings would match. Additionally, the proposed scheme includes an extra layer of protection by extracting terminal’s physical fingerprints from the human-induced electric potential to defend against advanced mimicry attacks. To test the effectiveness of our design, a low-cost wearable prototype is developed. Through extensive experiments, it is found that the proposed scheme has a low error rate of 2.3%, with minimal computational and energy requirements. Srinivasan Murali, Wenqiang Jin, Vighnesh Sivaraman, Huadi Zhu, Tianxi Ji, Pan Li 0001, Ming Li 0006 |
ACSAC | 1 |
| 2021 | Periscope: A Keystroke Inference Attack Using Human Coupled Electromagnetic EmanationsabstractThis study presents Periscope, a novel side-channel attack that exploits human-coupled electromagnetic (EM) emanations from touchscreens to infer sensitive inputs on a mobile device. Periscope is motivated by the observation that finger movement over the touchscreen leads to time-varying coupling between these two. Consequently, it impacts the screen's EM emanations that can be picked up by a remote sensory device. We intend to map between EM measurements and finger movements to recover the inputs. As the significant technical contribution of this work, we build an analytic model that outputs finger movement trajectories based on given EM readings. Our approach does not need a large amount of labeled dataset for offline model training, but instead a couple of samples to parameterize the user-specific analytic model. We implement Periscope with simple electronic components and conduct a suite of experiments to validate this attack's impact. Experimental results show that Periscope achieves a recovery rate over 6-digit PINs of 56.2% from a distance of 90 cm. Periscope is robust against environment dynamics and can well adapt to different device models and setting contexts. Wenqiang Jin, Srinivasan Murali, Huadi Zhu, Ming Li 0006 |
CCS | 2 |
| 2020 | Harnessing the Ambient Radio Frequency Noise for Wearable Device PairingabstractWearable devices that capture user's rich information regarding their health conditions and daily activities have unmet pairing needs. Today's solutions, which primarily rely on human involvement, are cumbersome, error-prone, and do not scale well. Despite some prior efforts trying to fill this gap, they either rely on some sophisticated sensors, such as electromyogram (EMG) or electrocardiogram (ECG) pads that may not universally exist, or non-trivial design of communication transceivers that cannot be found easily on current commercial devices. Therefore, a pairing scheme for wearable devices that is secure, practical, and convenient is in dire need. In this paper, we propose a novel approach that leverages ambient radio frequency (RF) noise. Our design is based on a key observation that received RF noise power measured in the logarithmic scale at different parts of a human body surface experience the same variation trend, whereas those from different human bodies or off the body are distinct. Wearables make use of the observed noise as the entropy source for the proposed pairing protocol. Extensive experiments show that our scheme has an equal error rate (EER) as low as 1.4% for pairing. Its key generation rate reaches 138 bits/sec, which beats so-far existing pairing schemes. Besides, our scheme can be efficiently executed within 0.97 s. Its incurred energy consumption is as low as 0.27 J for the entire pairing procedure. Wenqiang Jin, Ming Li 0006, Srinivasan Murali, Linke Guo |
CCS | 3 |
| 2020 | Towards 3D human pose construction using wifiabstractThis paper presents WiPose, the first 3D human pose construction framework using commercial WiFi devices. From the pervasive WiFi signals, WiPose can reconstruct 3D skeletons composed of the joints on both limbs and torso of the human body. By overcoming the technical challenges faced by traditional camera-based human perception solutions, such as lighting and occlusion, the proposed WiFi human sensing technique demonstrates the potential to enable a new generation of applications such as health care, assisted living, gaming, and virtual reality. WiPose is based on a novel deep learning model that addresses a series of technical challenges. First, WiPose can encode the prior knowledge of human skeleton into the posture construction process to ensure the estimated joints satisfy the skeletal structure of the human body. Second, to achieve cross environment generalization, WiPose takes as input a 3D velocity profile which can capture the movements of the whole 3D space, and thus separate posture-specific features from the static objects in the ambient environment. Finally, WiPose employs a recurrent neural network (RNN) and a smooth loss to enforce smooth movements of the generated skeletons. Our evaluation results on a real-world WiFi sensing testbed with distributed antennas show that WiPose can localize each joint on the human skeleton with an average error of 2.83cm, achieving a 35% improvement in accuracy over the state-of-the-art posture construction model designed for dedicated radar sensors. Hongfei Xue, Chenglin Miao, Sen Lin 0009, Chong Tian, Srinivasan Murali, Haochen Hu, Lu Su 0001 |
MobiCom | 7 |
| 2016 | Touch-based system for beat-to-beat impedance cardiogram acquisition and hemodynamic parameters estimation
Dionisije Sopic, Srinivasan Murali, Francisco J. Rincón, David Atienza 0001 |
DATE | 2 |
| 2016 | Ultra-Low Power Estimation of Heart Rate Under Physical Activity Using a Wearable Photoplethysmographic SystemabstractIn the last years, the need for enhancing health and preventing problems with remote monitoring is increasing. A non-invasive low-cost technique for processing bio-signals and monitoring vital parameters, at rest and during physical activity, is the use of wearable PhotoPlethysmoGraphic (PPG) systems. However, in order to detect a relevant vital parameter, such as the heart rate during demanding exercises, motion artifacts must be removed from the signals retrieved. In this paper, we present a fast and easy to implement algorithm to estimate the heart rate value which does not need to reconstruct the noise-free signal nor does it apply adaptive filtering as existing algorithms, thus gaining computational time and stored memory space. The method consists of applying the Fast Fourier Transform on short windows of data and removing motion artifacts relying on single-sided amplitude spectrum analysis of PPG and 3-axis accelerometer signals. The results show that our algorithm manages to remove a wide range of motion artifacts achieving an average absolute error of only 1.27 BPM between the heart rate estimated by the algorithm every second and the ground-truth value. The method was successfully implemented on a wearable PPG device achieving an execution time of 226 ms per second, hence obtaining a battery lifetime of 9.37 days. Elisabetta De Giovanni, Srinivasan Murali, Francisco J. Rincón, David Atienza 0001 |
DSD | 2 |
| 2015 | Estimation of Blood Pressure and Pulse Transit Time Using Your SmartphoneabstractIt is widely recognized today that there is an alarming rise of lifestyle-induced chronic diseases (e.g., type II diabetes) in our society. Therefore, a strong need exists for cost-effective and non-invasive devices that can measure blood pressure (BP) to monitor, diagnose and follow-up patients at risk, but also healthy population in general. One promising method for arterial BP estimation is to measure a surrogate marker of it, such as, Pulse Transit Time (PTT) and derive pressure values from it. However, current methods for measuring PTT require complex sensing and analysis circuitry and the related medical devices are expensive and inconvenient for the user to wear. In this paper, we present a new smartphone-based method to estimate PTT reliably and subsequently BP from the baseline sensors on smartphones. This new approach involves determining PTT by simultaneously measuring the time the blood leaves the heart, by recording the heart sound using the standard microphone of the phone and the time it reaches the finger, by measuring the pulse wave using the phone's camera. Moreover, we also describe algorithms that can be executed directly on current smartphones to obtain clean and robust heart sound signals and to extract the pulse wave characteristics using smartphones. We also present methods to ensure a synchronous capture of the waveforms, which is essential to obtain reliable PTT values with inexpensive sensors. Our experiments show that the computational overhead of the proposed two-phase processing method is minimum, with the ability to reliably measure the PTT values in a fully accurate (beat-to-beat) fashion using directly state-of-the-art smartphones as medical devices. Alair Dias Junior, Srinivasan Murali, Francisco J. Rincón, David Atienza 0001 |
DSD | 2 |
| 2014 | Ultra-Low Power Design of Wearable Cardiac Monitoring SystemsabstractThis paper presents the system-level architecture of novel ultra-low power wireless body sensor nodes (WBSNs) for real-time cardiac monitoring and analysis, and discusses the main design challenges of this new generation of medical devices. In particular, it highlights first the unsustainable energy cost incurred by the straightforward wireless streaming of raw data to external analysis servers. Then, it introduces the need for new cross-layered design methods (beyond hardware and software boundaries) to enhance the autonomy of WBSNs for ambulatory monitoring. In fact, by embedding more onboard intelligence and exploiting electrocardiogram (ECG) specific knowledge, it is possible to perform real-time compressive sensing, filtering, delineation and classification of heartbeats, while dramatically extending the battery lifetime of cardiac monitoring systems. The paper concludes by showing the results of this new approach to design ultra-low power wearable WBSNs in a real-life platform commercialized by SmartCardia. This wearable system allows a wide range of applications, including multi-lead ECG arrhythmia detection and autonomous sleep monitoring for critical scenarios, such as monitoring of the sleep state of airline pilots. Rubén Braojos, Hossein Mamaghanian, Alair Dias Junior, Giovanni Ansaloni, David Atienza 0001, Francisco J. Rincón, Srinivasan Murali |
DAC | 7 |
| 2013 | Computing Accurate Performance Bounds for Best Effort Networks-on-ChipabstractReal-time (RT) communication support is a critical requirement for many complex embedded applications which are currently targeted to Network-on-chip (NoC) platforms. In this paper, we present novel methods to efficiently calculate worst case bandwidth and latency bounds for RT traffic streams on wormhole-switched NoCs with arbitrary topology. The proposed methods apply to best-effort NoC architectures, with no extra hardware dedicated to RT traffic support. By applying our methods to several realistic NoC designs, we show substantial improvements (more than 30 percent in bandwidth and 50 percent in latency, on average) in bound tightness with respect to existing approaches. Dara Rahmati, Srinivasan Murali, Luca Benini, Federico Angiolini, Giovanni De Micheli, Hamid Sarbazi-Azad |
IEEE Trans. Computers | 2 |
| 2013 | Designing best effort networks-on-chip to meet hard latency constraintsabstractMany classes of applications require Quality of Service (QoS) guarantees from the system interconnect. In Networks-on-Chip (NoC) QoS guarantees usually translate into bandwidth and latency constraints for the traffic flows and require hardware support in the NoC fabric and its interfaces. In this article we present a novel NoC synthesis framework to automatically build networks that meet hard latency constraints of end-to-end traffic streams without requiring specialized hardware for the network components. The hard latency constraints are met by carefully designing the NoC topology and selecting the appropriate routes for flow using lean best-effort network components. We perform experiments on several System on Chip (SoC) benchmarks. We compared against a topology synthesis method with no support for real-time constraints and we show that the proposed method can produce topologies that can meet significantly tighter worst case latency constraints (on average 44%). We also show that the tightest worst case latency can be provided with little overhead on power consumption (on average 8.5%). Ciprian Seiculescu, Dara Rahmati, Srinivasan Murali, Hamid Sarbazi-Azad, Luca Benini, Giovanni De Micheli |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2010 | Design of networks on chips for 3D ICsabstractThree-dimensional integrated circuits, where multiple silicon layers are stacked vertically have emerged recently. The 3DICs have smaller form factor, shorter and efficient use of wires and allow integration of diverse technologies in the same device. The use of Networks on Chips (NoCs) to connect components in a 3D chip is a necessity. In this short paper, we present an outline on designing application-specific NoCs for 3D ICs. Srinivasan Murali, Luca Benini, Giovanni De Micheli |
ASP-DAC | 1 |
| 2010 | Networks on Chips: from research to productsabstractResearch on Networks on Chips (NoCs) has spanned over a decade and its results are now visible in some products. Thus the seminal idea of using networking technology to address the chip-level interconnect problem has been shown to be correct. Moreover, as technology scales down in geometry and chips scale up in complexity, NoCs become the essential element to achieve the desired levels of performance and quality of service while curbing power consumption levels. Design and timing closure can only be achieved by a sophisticated set of tools that address NoC synthesis, optimization and validation. Giovanni De Micheli, Ciprian Seiculescu, Srinivasan Murali, Luca Benini, Federico Angiolini, Antonio Pullini |
DAC | 3 |
| 2010 | A method to remove deadlocks in Networks-on-Chips with Wormhole flow controlabstractNetworks-on-Chip (NoCs) are a promising interconnect paradigm to address the communication bottleneck of Systems-on-Chip (SoCs). Wormhole flow control is widely used as the transmission protocol in NoCs, as it offers high throughput and low latency. To match the application characteristics, customized irregular topologies and routing functions are used. With wormhole flow control and custom irregular NoC topologies, deadlocks can occur during system operation. Ensuring a deadlock free operation of custom NoCs is a major challenge. In this paper, we address this important issue and present a method to remove deadlocks in application-specific NoCs. Our method can be applied to any NoC topology and routing function, and the potential deadlocks are removed by adding minimal number of virtual or physical channels. Experiments on a variety of realistic benchmarks show that our method results in a large reduction in the number of resources needed (88% on average) and NoC power consumption, area reduction (66% area savings on average) when compared to the state-of-the-art deadlock removal methods. Ciprian Seiculescu, Srinivasan Murali, Luca Benini, Giovanni De Micheli |
DATE | 2 |
| 2010 | SunFloor 3D: A Tool for Networks on Chip Topology Synthesis for 3-D Systems on ChipsabstractThree-dimensional integrated circuits (3D-ICs) are a promising approach to address the integration challenges faced by current systems on chips (SoCs). Designing an efficient network on chip (NoC) interconnect for a 3-D SoC that meets not only the application performance constraints but also the constraints imposed by the 3-D technology is a significant challenge. In this paper, we present a design tool, SunFloor 3D, to synthesize application-specific 3-D NoCs. The proposed tool determines the best NoC topology for the application, finds paths for the communication flows, assigns the network components to the 3-D layers, and places them in each layer. We perform experiments on several SoC benchmarks and present a comparative study between 3-D and 2-D NoC designs. Our studies show large improvements in interconnect power consumption (average of 38%) and delay (average of 13%) for the 3-D NoC when compared to the corresponding 2-D implementation. Our studies also show that the synthesized topologies result in large power (average of 54%) and delay savings (average of 21%) when compared to standard topologies. Ciprian Seiculescu, Srinivasan Murali, Luca Benini, Giovanni De Micheli |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2009 | Synthesis of networks on chips for 3D systems on chipsabstractThree-dimensional stacking of silicon layers is emerging as a promising solution to handle the design complexity and heterogeneity of Systems on Chips (SoCs). Networks on Chips (NoCs) are necessary to efficiently handle the 3D interconnect complexity. Designing power efficient NoCs for 3D SoCs that satisfy the application performance requirements, while satisfying the 3D technology constraints is a big challenge. In this work, we address this problem and present a synthesis approach for designing power-performance efficient 3D NoCs. We present methods to determine the best topology, compute paths and perform placement of the NoC components in each 3D layer. We perform experiments on varied, realistic SoC benchmarks to validate the methods and also perform a comparative study of the resulting 3D NoC designs with 3D optimized mesh topologies. The NoCs designed by our synthesis method results in large interconnect power reduction (average of 38%) and latency reduction (average of 25%) when compared to traditional NoC designs. Srinivasan Murali, Ciprian Seiculescu, Luca Benini, Giovanni De Micheli |
ASP-DAC | 1 |
| 2009 | NoC topology synthesis for supporting shutdown of voltage islands in SoCsabstractIn many Systems on Chips (SoCs), the cores are clustered in to voltage islands. When cores in an island are unused, the entire island can be shutdown to reduce the leakage power consumption. However, today, the interconnect architecture is a bottleneck in allowing the shutdown of the islands. In this paper, we present a synthesis approach to obtain customized application-specific Networks on Chips (NoCs) that can support the shutdown of voltage islands. Our results on realistic SoC benchmarks show that the resulting NoC designs only have a negligible overhead in SoC active power consumption (average of 3%) and area (average of 0.5%) to support the shutdown of islands. The shutdown support provided can lead to a significant leakage and hence total power savings. Ciprian Seiculescu, Srinivasan Murali, Luca Benini, Giovanni De Micheli |
DAC | 2 |
| 2009 | SunFloor 3D: A tool for Networks On Chip topology synthesis for 3D systems on chipsabstractThree-dimensional integrated circuits are a promising approach to address the integration challenges faced by current Systems on Chips (SoCs). Designing an efficient Network on Chip (NoC) interconnect for a 3D SoC that not only meets the application performance constraints, but also the constraints imposed by the 3D technology, is a significant challenge. In this work we present a design tool, SunFloor 3D, to synthesize application-specific 3D NoCs. The proposed tool determines the best NoC topology for the application, finds paths for the communication flows, assigns the network components on to the 3D layers and performs a placement of them in each layer. We perform experiments on several SoC benchmarks and present a comparative study between 3D and 2D NoC designs. Our studies show large improvements in interconnect power consumption (average of 38%) and delay (average of 13%) for the 3D NoC when compared to the corresponding 2D implementation. Our studies also show that the synthesized topologies result in large power (average of 54%) and delay savings (average of 21%) when compared to standard topologies. Ciprian Seiculescu, Srinivasan Murali, Luca Benini, Giovanni De Micheli |
DATE | 2 |
| 2009 | A method for calculating hard QoS guarantees for Networks-on-ChipabstractMany Networks-on-Chip (NoC) applications exhibit one or \nmore critical traffic flows that require hard Quality of Service \n(QoS). Guaranteeing bandwidth and latency for such real time \nflows is crucial. In this paper, we present novel methods to \nefficiently calculate worst-case bandwidth and latency bounds \nand thereby provide hard QoS guarantees. Importantly, the \nproposed methods apply even to best-effort NoC architectures, \nwith no extra hardware dedicated to QoS support. By applying \nour methods to several realistic NoC designs, we show \nsubstantial improvements (on average, more than 30% in \nbandwidth and 50% in latency) in bound tightness with respect \nto existing approaches.1 Dara Rahmati, Srinivasan Murali, Luca Benini, Federico Angiolini, Giovanni De Micheli, Hamid Sarbazi-Azad |
ICCAD | 2 |
| 2008 | Temperature Control of High-Performance Multi-core Platforms Using Convex OptimizationabstractWith technology advances, the number of cores integrated on a chip and their speed of operation is increasing. This, in turn is leading to a significant increase in chip temperature. Temperature gradients and hot-spots not only affect the performance of the system, but also lead to unreliable circuit operation and affect the life-time of the chip. Meeting the temperature constraints and reducing the hot-spots are critical for achieving reliable and efficient operation of complex multi-core systems. In this work, we present Pro-Temp, a convex optimization based method that pro-actively controls the temperature of the cores, while minimizing the power consumption and satisfying application performance constraints. The method guarantees that the temperature of the cores are below a user- defined threshold at all instances of operation, while also reducing the hot-spots. We perform experiments on several realistic multi-core benchmarks, which show that the proposed method guarantees that the cores never exceed the maximum temperature limit, while matching the application performance requirements. We compare this to traditional methods, where we find several temperature violations during the operation of the system. Srinivasan Murali, Almir Mutapcic, David Atienza 0001, Rajesh K. Gupta 0001, Stephen P. Boyd, Luca Benini, Giovanni De Micheli |
DATE | 1 |
| 2008 | Network-on-Chip design and synthesis outlook
David Atienza 0001, Federico Angiolini, Srinivasan Murali, Antonio Pullini, Luca Benini, Giovanni De Micheli |
Integr. | 3 |
| 2007 | NoC Design and Implementation in 65nm TechnologyabstractAs embedded computing evolves towards ever more powerful architectures, the challenge of properly interconnecting large numbers of on-chip computation blocks is becoming prominent. Networks-on-chip (NoCs) have been proposed as a scalable solution to both physical design issues and increasing bandwidth demands. However, this claim has not been fully validated yet, since the design properties and tradeoffs of NoCs have not been studied in detail below the 100 nm threshold. This work is aimed at shedding light on the opportunities and challenges, both expected and unexpected, of NoC design in nanometer CMOS. We present fully working 65 nm NoC designs, a complete NoC synthesis flow and detailed scalability analysis Antonio Pullini, Federico Angiolini, Paolo Meloni, David Atienza 0001, Srinivasan Murali, Luigi Raffo, Giovanni De Micheli, Luca Benini |
NOCS | 5 |
| 2007 | An Application-Specific Design Methodology for On-Chip Crossbar GenerationabstractDesigning a power-efficient interconnection architecture for multiprocessor systems-on-chips (MPSoCs) satisfying the application performance constraints is a nontrivial task. In order to meet the tight time-to-market constraints and to effectively handle the design complexity, it is essential to provide a computer-aided design tool support for automating this task. In this paper, we address the issue of ldquoapplication-specific design of optimal crossbar architecturerdquo satisfying the performance requirements of the application and optimal binding of the cores onto the crossbar resources. We present a simulation-based design approach that is based on the analysis of the actual traffic trace of the application, considering local variations in traffic rates, temporal overlap among traffic streams, and criticality of traffic streams. Our approach is physical design aware, where the wiring complexity of the crossbar architecture is also considered during the design process. This leads to detecting timing violations on the wires early in the design cycle and to having accurate estimates of the power consumption on the wires. We apply our methodology onto several MPSoC designs, and the synthesized crossbar platforms are validated for performance by cycle-accurate SystemC simulation of the designs. The crossbar matrix power consumption values are based on the synthesis of the register transfer level models of the designs, obtained using industry standard tools. The experimental case studies show large reduction in communication architecture power consumption (45.3% on average) and total wirelength (38% on average) for the MPSoC designs when compared with traditional design approaches. The synthesized crossbar designs also lead to large reduction in transaction latencies (up to 7 ) when compared with the existing design approaches. Srinivasan Murali, Luca Benini, Giovanni De Micheli |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2007 | Timing-Error-Tolerant Network-on-Chip Design MethodologyabstractWith technology scaling, the wire delay as a fraction of the total delay is increasing, and the communication architecture is becoming a major bottleneck for system performance in systems on chip (SoCs). A communication-centric design paradigm, networks on chip (NoCs), has been proposed recently to address the communication issues of SoCs. As the geometries of devices approach the physical limits of operation, NoCs will be susceptible to various noise sources such as crosstalk, coupling noise, process variations, etc. Designing systems under such uncertain conditions become a challenge, as it is harder to predict the timing behavior of the system. The use of conservative design methodologies that consider all possible delay variations due to the noise sources, targeting safe system operation under all conditions will result in poor system performance. An aggressive design approach that provides resilience against such timing errors is required for maximizing system performance. In this paper, we present T-error, which is a timing-error-tolerant aggressive design method to design the individual components of the NoC (such as switches, links, and network interfaces), so that the communication subsystem can be clocked at a much higher frequency than a traditional conservative design (up to 1.5x increase in frequency). The NoC is designed to tolerate timing errors that arise from overclocking without substantially affecting the latency for communication. We also present a way to dynamically configure the NoC between the overclocked mode and the normal mode, where the frequency of operation is lower than or equal to the traditional design's frequency, so that the error recovery penalty is completely hidden under normal operation. Experiments on several benchmark applications show large performance improvement (up to 33% reduction in average packet latency) for the proposed system when compared to traditional systems. Rutuparna Tamhankar, Srinivasan Murali, Stergios Stergiou, Antonio Pullini, Federico Angiolini, Luca Benini, Giovanni De Micheli |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2007 | Synthesis of Predictable Networks-on-Chip-Based Interconnect Architectures for Chip MultiprocessorsabstractToday, chip multiprocessors (CMPs) that accommodate multiple processor cores on the same chip have become a reality. As the communication complexity of such multicore systems is rapidly increasing, designing an interconnect architecture with predictable behavior is essential for proper system operation. In CMPs, general-purpose processor cores are used to run software tasks of different applications and the communication between the cores cannot be precharacterized. Designing an efficient network-on-chip (NoC)-based interconnect with predictable performance is thus a challenging task. In this paper, we address the important design issue of synthesizing the most power efficient NoC interconnect for CMPs, providing guaranteed optimum throughput and predictable performance for any application to be executed on the CMP. In our synthesis approach, we use accurate delay and power models for the network components (switches and links) that are obtained from layouts of the components using industry standard tools. The synthesis approach utilizes the floorplan knowledge of the NoC to detect timing violations on the NoC links early in the design cycle. This leads to a faster design cycle and quicker design convergence across the high-level synthesis approach and the physical implementation of the design. We validate the design flow predictability of our proposed approach by performing a layout of the NoC synthesized for a 25-core CMP. Our approach maintains the regular and predictable structure of the NoC and is applicable in practice to existing NoC architectures. Srinivasan Murali, David Atienza 0001, Paolo Meloni, Salvatore Carta, Luca Benini, Giovanni De Micheli, Luigi Raffo |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2006 | Mapping and configuration methods for multi-use-case networks on chipsabstractTo provide a scalable communication infrastructure for systems on chips (SoCs), networks on chips (NoCs), a communication centric design paradigm is needed. To be cost effective, SoCs are often programmable and integrate several different applications or use-cases on to the same chip. For the SoC platform to support the different use-cases, the NoC architecture should satisfy the performance constraints of each individual use-case. In this work we motivate the need to consider multiple use-cases during the NoC design process. We present a method to efficiently map the applications on to the NoC architecture, satisfying the design constraints of each individual use-case. We also present novel ways to dynamically reconfigure the network across the different use-cases and explore the possibility of integrating dynamic voltage and frequency scaling (DVS/DFS) techniques with the use-case centric NoC design methodology. We validate the performance of the design methodology on several SoC applications. The dynamic reconfiguration of the NoC integrated with DVS/DFS schemes results in large power savings for the resulting NoC systems Srinivasan Murali, Martijn Coenen, Andrei Radulescu, Kees Goossens, Giovanni De Micheli |
ASP-DAC | 1 |
| 2006 | A multi-path routing strategy with guaranteed in-order packet delivery and fault-tolerance for networks on chipabstractIn this work we present a multi-path routing strategy that guaran-tees in-order packet delivery for Networks on Chips (NoCs). We present a design methodology that uses the routing strategy to opti-mally spread the traffic in the NoC to minimize the network band-width needs and power consumption. We also integrate support for tolerance against transient and permanent failures in the NoC links in the methodology by utilizing spatial and temporal redundancy for transporting packets. Our experimental studies show large re-duction in network bandwidth requirements (36.86% on average) and power consumption (30.51% on average) compared to single-path systems. The area overhead of the proposed scheme is small (a modest 5% increase in network area). Hence, it is practical to be used in the on-chip domain. Srinivasan Murali, David Atienza 0001, Luca Benini, Giovanni De Micheli |
DAC | 1 |
| 2006 | A methodology for mapping multiple use-cases onto networks on chipsabstractA communication-centric design approach, networks on chips (NoCs), has emerged as the design paradigm for designing a scalable communication infrastructure for future systems on chips (SoCs). As technology advances, the number of applications or use-cases integrated on a single chip increases rapidly. The different use-cases of the SoC have different communication requirements (such as different bandwidth, latency constraints) and traffic patterns. The underlying NoC architecture has to satisfy the constraints of all the use-cases. In this work, we present a methodology to map multiple use-cases onto the NoC architecture, satisfying the constraints of each use-case. We present dynamic re-configuration mechanisms that match the NoC configuration to the communication characteristics of each use-case, also accounting for use-cases that can run in parallel. The methodology is applied to several real and synthetic SoC benchmarks, which result in a large reduction in NoC area (an average of 80%) and power consumption (an average of 54%) compared to traditional design approaches Srinivasan Murali, Martijn Coenen, Andrei Radulescu, Kees Goossens, Giovanni De Micheli |
DATE | 1 |
| 2006 | Designing application-specific networks on chips with floorplan informationabstractWith increasing communication demands of processor and memory cores in Systems on Chips (SoCs), scalable Networks on Chips (NoCs) are needed to interconnect the cores. For the use of NoCs to be feasible in today's industrial designs, a custom-tailored, application-specific NoC that satisfies the design objectives and constraints of the targeted application domain is required. In this work, we present a design methodology that automates the synthesis of such application-specific NoC architectures. We present a floorplan aware design method that considers the wiring complexity of the NoC during the topology synthesis process. This leads to detecting timing violations on the NoC links early in the design cycle and to have accurate power estimations of the interconnect. We incorporate mechanisms to prevent deadlocks during routing, which is critical for proper operation of NoCs. We integrate the NoC synthesis method with an existing design flow, automating NoC synthesis, generation, simulation and physical design processes. We also present ways to ensure design convergence across the levels. Experiments on several SoC benchmarks are presented, which show that the synthesized topologies provide a large reduction in network power consumption (2.78x on average) and improvement in performance (1.59x on average) over the best mesh and mesh-based custom topologies. An actual layout of a multimedia SoC with the NoC designed using our methodology is presented, which shows that the designed NoC supports the required frequency of operation (close to 900 MHz) without any timing violations. We could design the NoC from input specifications to layout in 4 hours, a process that usually takes several weeks. Srinivasan Murali, Paolo Meloni, Federico Angiolini, David Atienza 0001, Salvatore Carta, Luca Benini, Giovanni De Micheli, Luigi Raffo |
ICCAD | 1 |
| 2006 | Reliability Support for On-Chip Memories Using Networks-on-ChipabstractAs the geometries of the transistors reach the physical limits of operation, one of the main design challenges of systems-on-chips (SoCs) will be to provide dynamic (run-time) support against permanent and intermittent faults that can occur in the system. One of the most critical elements that affect the correct behavior of the system is the unreliable operation of on-chip memories. In this paper we present a novel solution to enable fault tolerant on-chip memory design at the system level for multimedia applications, based on the network-on-chip (NoC) interconnection paradigm. We transparently keep backup copies of critical data on a reliable memory; upon a fault event, data is fetched from the backup copy in hardware, without any software intervention. The use of a NoC backbone enables an efficient design which is modular, scalable and efficient. We proceed to demonstrating its effectiveness with two real-life application case studies, and explore the performance under varying architectural configurations. The overhead to support the proposed approach is very small compared to non-fault tolerant systems, i.e. no negative performance impact and an area increase dominated by that of just the backup storage itself. Federico Angiolini, David Atienza 0001, Srinivasan Murali, Luca Benini, Giovanni De Micheli |
ICCD | 3 |
| 2006 | Designing Message-Dependent Deadlock Free Networks on Chips for Application-Specific Systems on ChipsabstractNetworks on chip (NoC) has emerged as the paradigm for designing scalable communication architecture for systems on chips (SoCs). Avoiding the conditions that can lead to deadlocks in the network is critical for using NoCs in real designs. Methods that can lead to deadlock-free operation with minimum power and area overhead are important for designing application-specific NoCs. A major class of deadlocks that occur in NoCs are due to the dependencies among the resources shared by different message types. In this work, we consider the problem of avoiding message-dependent deadlocks during the NoC topology synthesis phase. We show that by considering this issue during topology synthesis, we can obtain a significantly better NoC design than traditional methods, where the deadlock avoidance issue is dealt with separately. Our experiments on several SoC benchmarks show that our proposed scheme provides large reduction in NoC power consumption (an average of 38.5%) and NoC area (an average of 30.7%) when compared to traditional approaches Srinivasan Murali, Paolo Meloni, Federico Angiolini, David Atienza 0001, Salvatore Carta, Luca Benini, Giovanni De Micheli, Luigi Raffo |
VLSI-SoC | 1 |
| 2005 | Mapping and physical planning of networks-on-chip architectures with quality-of-service guaranteesabstractNetworks on Chips (NoCs) have evolved as the communication design paradigm of future Systems on Chips (SoCs). In this work we target the NoC design of complex SoCs with heterogeneous processor/memory cores, providing Quality-of-Service (QoS) for the application. We present an integrated approach to mapping of cores onto NoC topologies and physical planning of NoCs, where the position and size of the cores and network components are computed. Our design methodology automates NoC mapping, physical planning, topology selection, topology optimization and instantiation, bridging an important design gap in building application specific NoCs. We also present a methodology to guarantee QoS for the application during the mapping-physical planning process by satisfying the delay/jitter constraints and real-time constraints of the traffic streams. Experimental studies show large area savings (up to 2x), bandwidth savings (up to 5x) and network component savings (up to 2.2x in buffer count, 3.8x in number of wires, 1.6x in switch ports) compared to traditional design approaches. Srinivasan Murali, Luca Benini, Giovanni De Micheli |
ASP-DAC | 1 |
| 2005 | Performance driven reliable link design for networks on chipsabstractWith decreasing feature size of transistors, the interconnect wire delay is becoming a major bottleneck in current Systems on Chips (SoCs). Another effect of shrinking feature size is that the wires are becoming unreliable as they are increasingly susceptible to various noise sources such as cross-talk, coupling noise, soft errors etc. Increasing importance of wire delay and reliability has lead to a communication centric design approach, Networks on Chip (NoC), for building complex SoCs. Current NoC communication design methodologies are based on conservative design approaches and consider worst case operating conditions for link design, resulting in large latency penalty for data transmission. In order to sub-stantially decrease the link delay and thereby increase system performance an aggressive design approach is needed. In this work we present Terror, timing error tolerant communication system, for aggressively designing the links of NoCs. In our methodology, instead of avoiding timing errors by a worst-case design, we do aggressive design by tolerating timing errors. Simulation results show large latency savings (up to 35%) for the Terror based system compared to traditional design methodology. Rutuparna Tamhankar, Srinivasan Murali, Giovanni De Micheli |
ASP-DAC | 2 |
| 2005 | An Application-Specific Design Methodology for STbus Crossbar GenerationabstractAs the communication requirements of current and future Multiprocessor Systems on Chips (MPSoCs) continue to increase, scalable communication architectures are needed to support the heavy communication demands of the system. This is reflected in the recent trend that many of the standard bus products such as STbus, have now introduced the capability of designing a crossbar with multiple buses operating in parallel. The crossbar configuration should be designed to closely match the application traffic characteristics and performance requirements. In this work we address this issue of application-specific design of optimal crossbar (using STbus crossbar architecture), satisfying the performance requirements of the application and optimal binding of cores onto the crossbar resources. We present a simulation based design approach that is based on analysis of actual traffic trace of the application, considering local variations in traffic rates, temporal overlap among traffic streams and criticality of traffic streams. Our methodology is applied to several MPSoC designs and the resulting crossbar platforms are validated for performance by cycle-accurate SystemC simulation of the designs. The experimental case studies show large reduction in packet latencies (up to 7×) and large crossbar component savings (up to 3.5×) compared to traditional design approaches. Srinivasan Murali, Giovanni De Micheli |
DATE | 1 |
| 2005 | NoC Synthesis Flow for Customized Domain Specific Multiprocessor Systems-on-ChipabstractThe growing complexity of customizable single-chip multiprocessors is requiring communication resources that can only be provided by a highly-scalable communication infrastructure. This trend is exemplified by the growing number of network-on-chip (NoC) architectures that have been proposed recently for system-on-chip (SoC) integration. Developing NoC-based systems tailored to a particular application domain is crucial for achieving high-performance, energy-efficient customized solutions. The effectiveness of this approach largely depends on the availability of an ad hoc design methodology that, starting from a high-level application specification, derives an optimized NoC configuration with respect to different design objectives and instantiates the selected application specific on-chip micronetwork. Automatic execution of these design steps is highly desirable to increase SoC design productivity. This work illustrates a complete synthesis flow, called Netchip, for customized NoC architectures, that partitions the development work into major steps (topology mapping, selection, and generation) and provides proper tools for their automatic execution (SUNMAP, xpipescompiler). The entire flow leverages the flexibility of a fully reusable and scalable network components library called xpipes, consisting of highly-parameterizable network building blocks (network interface, switches, switch-to-switch links) that are design-time tunable and composable to achieve arbitrary topologies and customized domain-specific NoC architectures. Several experimental case studies are presented In the work, showing the powerful design space exploration capabilities of the proposed methodology and tools. Davide Bertozzi, Antoine Jalabert, Srinivasan Murali, Rutuparna Tamhankar, Stergios Stergiou, Luca Benini, Giovanni De Micheli |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2004 | SUNMAP: a tool for automatic topology selection and generation for NoCsabstractIncreasing communication demands of processor and memory cores in Systems on Chips (SoCs) necessitate the use of Networks on Chip (NoC) to interconnect the cores. An important phase in the design of NoCs is he mapping of cores onto the most suitable opology for a given application. In this paper, we present SUNMAP a tool for automatically selecting he best topology for a given application and producing a mapping of cores onto that topology. SUNMAP explores various design objectives such as minimizing average communication delay, area, power dissipation subject to bandwidth and area constraints. The tool supports different routing functions (dimension ordered, minimum-path, traffic splitting) and uses floorplanning information early in the topology selection process to provide feasible mappings. The network components of the chosen NoC are automatically generated using cycle-accurate SystemC soft macros from X-pipes architecture. SUNMAP automates NoC selection and generation, bridging an important design gap in building NoCs. Several experimental case studies are presented in the paper, which show the rich design space exploration capabilities of SUNMAP. Srinivasan Murali, Giovanni De Micheli |
DAC | 1 |
| 2004 | ×pipesCompiler: A Tool for Instantiating Application Specific Networks on ChipabstractFuture systems on chips (SoCs) will integrate a large number of processor and storage cores onto a single chip and require networks on chip (NoC) to support the heavy communication demands of the system. The individual components of the SoCs will be heterogeneous in nature with widely varying functionality and communication requirements. The communication infrastructure should optimally match communication patterns among these components accounting for the individual component needs. In this paper we present /spl times/pipesCompiler, a tool for automatically instantiating an application-specific NoC for heterogeneous multi-processor SoCs. The /spl times/pipesCompiler instantiates a network of building blocks from a library of composable soft macros (switches, network interfaces and links) described in SystemC at the cycle-accurate level. The network components are optimized for that particular network and support reliable, latency-insensitive operation. Example systems with application-specific NoCs built using the /spl times/pipesCompiler show large savings in area (factor of 6.5), power (factor of 2.4) and latency (factor of 1.42) when compared to a general-purpose mesh-based NoC architecture. Antoine Jalabert, Srinivasan Murali, Luca Benini, Giovanni De Micheli |
DATE | 2 |
| 2004 | Bandwidth-Constrained Mapping of Cores onto NoC ArchitecturesabstractWe address the design of complex monolithic systems, where processing cores generate and consume a varying and large amount of data, thus bringing the communication links to the edge of congestion. Typical applications are in the area of multi-media processing. We consider a mesh-based networks on chip (NoC) architecture, and we explore the assignment of cores to mesh cross-points so that the traffic on links satisfies bandwidth constraints. A single-path deterministic routing between the cores places high bandwidth demands on the links. The bandwidth requirements can be significantly reduced by splitting the traffic between the cores across multiple paths. In this paper, we present NMAP, a fast algorithm that maps the cores onto a mesh NoC architecture under bandwidth constraints, minimizing the average communication delay. The NMAP algorithm is presented for both single minimum-path routing and split-traffic routing. The algorithm is applied to a benchmark DSP design and the resulting NoC is built and simulated at cycle accurate level in SystemC using macros from the /spl times/pipes library. Also, experiments with six video processing applications show significant savings in bandwidth and communication cost for NMAP algorithm when compared to existing algorithms. Srinivasan Murali, Giovanni De Micheli |
DATE | 1 |