VLDB 2026 Research / reviewers in the wild / expert
Murugan Sankaradass
dblp:83/1079 · also Murugan Sankaradas
· DBLP profile ↗
20ranked-venue papers
2as first author
11since 2021 · last 2025
0000-0002-4608-1630ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 7 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Software engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Robot navigation and mapping · 42% Video understanding and tracking · 28% Reinforcement learning · 28% | |
| Computer networks
3 papers |
Edge and fog computing · 67% Internet of things and sensor networks · 17% Wireless networking · 17% | |
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Hardware accelerators and domain-specific architectures · 34% Processor architecture and microarchitecture · 23% Parallel and multicore computing · 17% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Smart cities and intelligent transportation · 100% |
Topics — the 19 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Edge and fog computing › video analytics
video analytics pipeline |
1.0 | 2 | 2025 | CamTuner: Adaptive Video Analytics Pipelines via Real-Time Automated Camera Parameter Tuning · IEEE Trans. Mob. Comput. 2025 Enhancing Video Analytics Accuracy via Real-time Automated Camera Parameter Tuning · SenSys 2022 |
Robotics › Robot navigation and mapping
sensor fusion |
0.9 | 1 | 2025 | Roadside Multi-LiDAR Data Fusion for Enhanced Traffic Safety · KDD (1) 2025 |
Smart cities and intelligent transportation
traffic management |
0.9 | 1 | 2025 | Roadside Multi-LiDAR Data Fusion for Enhanced Traffic Safety · KDD (1) 2025 |
Computer vision › Video understanding and tracking
video analytics |
0.6 | 1 | 2022 | Enhancing Video Analytics Accuracy via Real-time Automated Camera Parameter Tuning · SenSys 2022 |
Internet of things and sensor networks › wireless sensor network › sensor network management
sensor calibration |
0.3 | 1 | 2025 | Roadside Multi-LiDAR Data Fusion for Enhanced Traffic Safety · KDD (1) 2025 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.2 | 2 | 2010 | A dynamically configurable coprocessor for convolutional neural networks · ISCA 2010 A Massively Parallel Digital Learning Processor · NIPS 2008 |
Processor architecture and microarchitecture › many-core architecture
intel xeon phi |
0.2 | 1 | 2013 | COSMIC: middleware for high performance and reliable multiprocessing on xeon phi coprocessors · HPDC 2013 |
Parallel and multicore computing
parallel programming runtimes |
0.2 | 1 | 2013 | COSMIC: middleware for high performance and reliable multiprocessing on xeon phi coprocessors · HPDC 2013 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
CNN accelerator |
0.1 | 1 | 2010 | A dynamically configurable coprocessor for convolutional neural networks · ISCA 2010 |
Reconfigurable computing and FPGAs › reconfigurable computing
reconfigurable accelerator |
0.1 | 1 | 2010 | A dynamically configurable coprocessor for convolutional neural networks · ISCA 2010 |
Processor architecture and microarchitecture
chip multiprocessor |
0.1 | 1 | 2006 | Software architecture exploration for high-performance security processing on a multiprocessor mobile SoC · DAC 2006 |
Embedded and real-time systems › embedded processor
code compression |
0.0 | 1 | 2003 | CoCo: a hardware/software platform for rapid prototyping of code compression technologies · DAC 2003 |
Cryptographic primitives and cryptanalysis
cryptographic implementation |
0.0 | 1 | 2002 | System design methodologies for a wireless security processing platform · DAC 2002 |
Hardware accelerators and domain-specific architectures › domain-specific accelerator
application-specific processor |
0.0 | 1 | 2002 | System design methodologies for a wireless security processing platform · DAC 2002 |
Electronic design automation
hardware/software co-design |
0.0 | 1 | 2002 | System design methodologies for a wireless security processing platform · DAC 2002 |
Electronic design automation
system-level design |
0.0 | 1 | 2002 | System design methodologies for a wireless security processing platform · DAC 2002 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.0 | 1 | 2010 | A dynamically configurable coprocessor for convolutional neural networks · ISCA 2010 |
Reconfigurable computing and FPGAs › reconfigurable computing
FPGA-based machine learning accelerators |
0.0 | 1 | 2008 | A Massively Parallel Digital Learning Processor · NIPS 2008 |
Reconfigurable computing and FPGAs
FPGA prototyping |
0.0 | 1 | 2003 | CoCo: a hardware/software platform for rapid prototyping of code compression technologies · DAC 2003 |
Methods — techniques the papers use, named apart from their topics
spatiotemporal calibration · 2.6LiDAR data fusion · 2.6reinforcement learning · 1.1SARSA · 1.1deep learning · 0.9SARSA reinforcement learning · 0.9VLIW compilation · 0.2middleware · 0.2data structure access profiling · 0.1branch-and-bound · 0.1multiply-accumulate · 0.1SIMD · 0.1performance macro-modeling · 0.1architecture refinement · 0.1algorithmic exploration · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Roadside Multi-LiDAR Data Fusion for Enhanced Traffic SafetyabstractRoadside LiDAR (Light Detection and Ranging) sensors promise safer and faster traffic management and vehicular operations. However, occlusion and small view angles are significant challenges to widespread use of roadside LiDARs. We consider fusing data from multiple LiDARs at a traffic intersection to better estimate traffic parameters than one can estimate from a single LiDAR. The key challenge is to calibrate multiple LiDARs both in time and space. The problem is more complex when heterogeneous sensors differ in resolution and are positioned arbitrarily on a traffic intersection. Md. Parvez Mollah, Biplob Debnath, Murugan Sankaradass, Srimat T. Chakradhar, Abdullah Mueen |
KDD (1) | 3 |
| 2025 | Real-Time Network-Aware Roadside LiDAR Data Compression
Md. Parvez Mollah, Murugan Sankaradass, Ravi K. Rajendran, Srimat T. Chakradhar |
VEHITS | 2 |
| 2025 | CamTuner: Adaptive Video Analytics Pipelines via Real-Time Automated Camera Parameter TuningabstractIn Video Analytics Pipelines (VAP), Analytics Units (AUs) such as object detection and face recognition operating on remote servers rely heavily on surveillance cameras to capture high-quality video streams to achieve high accuracy. Modern network cameras offer an array of parameters that directly influence video quality. While a few of such parameters, e.g., exposure, focus and white balance, are automatically adjusted by the camera internally, the others are not. We denote such camera parameters as non-automated (NAUTO) parameters. In this work, we first show that in a typical surveillance camera deployment, environmental condition changes can have significant adverse effect on the accuracy of insights from the AUs, but such adverse impact can potentially be mitigated by dynamically adjusting NAUTO camera parameters in response to changes in environmental conditions. Second, since most end-users lack the skill or understanding to appropriately configure these parameters and typically use a fixed parameter setting, we presentCamTuner, to our knowledge, the first framework that dynamically adapts NAUTO camera parameters to optimize the accuracy of AUs in a VAP in response to adverse changes in environmental conditions.CamTuneris based on SARSA reinforcement learning and it incorporates two novel components: a light-weight analytics quality estimator and a virtual camera that drastically speed up offline RL training. Our controlled experiments and real-world VAP deployment show that compared to a VAP using the default camera setting,CamTunerenhances VAP accuracy by detecting 15.9% additional persons and 2.6% –4.2% additional cars (without any false positives) in a large enterprise parking lot.CamTuneropens up new avenues for elevating video analytics accuracy, transcending mere incremental enhancements achieved through refining deep-learning models. Sibendu Paul, Kunal Rao, Giuseppe Coviello, Murugan Sankaradass, Y. Charlie Hu, Srimat T. Chakradhar |
IEEE Trans. Mob. Comput. | 4 |
| 2023 | Elixir: A System to Enhance Data Quality for Multiple Analytics on a Video StreamabstractIoT sensors, especially video cameras, are ubiquitously deployed around the world to perform a variety of computer vision tasks in several verticals including retail, health-care, safety and security, transportation, manufacturing, etc. To amortize their high deployment effort and cost, it is desirable to perform multiple video analytics tasks, which we refer to as Analytical Units (AUs), off the video feed coming out of every camera. As AUs typically use deep-learning based AI/ML models, their performances depend on the quality of the input video. The most recent work has shown that dynamically adjusting the camera setting exposed by popular network cameras can help improve the quality of the video feed and hence the AU accuracy, in a single AU setting. In this paper, we first show that in a multi-AU setting, changing the camera setting has disproportionate impact on different AUs performance. In particular, the optimal setting for one AU may severely degrade the performance for another AU, and further, the impact on different AUs varies as the environmental condition changes. We then present Elixir, a system to enhance the video stream quality for multiple analytics on a video stream. Elixir leverages Multi-Objective Reinforcement Learning (MORL), where the RL agent caters to the objectives from different AUs and adjusts the camera setting to simultaneously enhance the performance of all AUs. To define the multiple objectives in MORL, we develop new AU-specific quality estimator values for each individual AU. We evaluate Elixir through real-world experiments on a testbed with three cameras deployed next to each other (overlooking a large enterprise parking lot) running Elixir and two baseline approaches, respectively. Elixir correctly detects 7.1% (22,068) and 5.0% (15,731) more cars, 94% (551) and 72% (478) more faces, and 670.4% (4975) and 158.6% (3507) more persons than the default-setting and time-sharing approaches, respectively. It also detects 115 license plates, far more than the time-sharing approach (7) and the default setting (0). Sibendu Paul, Kunal Rao, Giuseppe Coviello, Murugan Sankaradass, Y. Charlie Hu, Srimat T. Chakradhar |
SMARTCOMP | 4 |
| 2023 | AnB: Application-in-a-Box to Rapidly Deploy and Self-optimize 5G AppsabstractWe present "Application in a Box" (AnB) product concept aimed at simplifying the deployment and operation of remote 5G applications. AnB comes pre-configured with all necessary hardware and software components, including sensors like cameras, hardware and software components for a local 5G wireless network, and 5G-ready apps. Enterprises can easily download additional apps from an App Store. Setting up a 5G infrastructure and running applications on it is a significant challenge, but AnB is designed to make it fast, convenient, and easy, even for those without extensive knowledge of software, computers, wireless networks, or AI-based analytics. With AnB, customers only need to open the box, set up the sensors, turn on the 5G networking and edge computing devices, and start running their applications. Our system software automatically deploys and optimizes the pipeline of microservices in the application on a tiered computing infrastructure that includes device, edge, and cloud computing. Application scalability, dynamic resource management, placement of critical tasks for low-latency response, and dynamic network bandwidth allocation for efficient 5G network usage are all automatically orchestrated.AnB offers cost savings, simplified setup and management, and increased reliability and security. We’ve implemented several real-world applications, such as collision prediction at busy traffic light intersections and remote construction site monitoring using video analytics. With AnB, deployment and optimization effort can be reduced from several months to just a few minutes. This is the first-of-its-kind approach to easing deployment effort and automating self-optimization of the application during system operation. Kunal Rao, Murugan Sankaradass, Giuseppe Coviello, Ciro Giuseppe De Vita, Gennaro Mellone, Wang-Pin Hsiung, Srimat T. Chakradhar |
SMARTCOMP | 2 |
| 2022 | Efficient Compression Method for Roadside LiDAR DataabstractRoadside LiDAR (Light Detection and Ranging) sensors are recently being explored for intelligent transportation systems aiming at safer and faster traffic management and vehicular operations. A key challenge in such systems is to efficiently transfer massive point-cloud data from the roadside LiDAR devices to the edge connected through a 5G network for real-time processing. In this paper, we consider the problem of compressing roadside (i.e. static) LiDAR data in real-time that provides a unique condition unexplored by current methods. Existing point-cloud compression methods assume moving LiDARs (that are mounted on vehicles) and do not exploit spatial consistency across frames over time. Md. Parvez Mollah, Biplob Debnath, Murugan Sankaradass, Srimat T. Chakradhar, Abdullah Mueen |
CIKM | 3 |
| 2022 | ROMA: Resource Orchestration for Microservices-based 5G ApplicationsabstractWith the growth of 5G, Internet of Things (IoT), edge computing and cloud computing technologies, the infrastructure (compute and network) available to emerging applications (AR/VR, autonomous driving, industry 4.0, etc.) has become quite complex. There are multiple tiers of computing (IoT devices, near edge, far edge, cloud, etc.) that are connected with different types of networking technologies (LAN, LTE, 5G, MAN, WAN, etc.). Deployment and management of applications in such an environment is quite challenging. In this paper, we propose ROMA, which performs resource orchestration for microservices-based 5G applications in a dynamic, heterogeneous, multi-tiered compute and network fabric. We assume that only application-level requirements are known, and the detailed requirements of the individual microservices in the application are not specified. As part of our solution, ROMA identifies and leverages the coupling relationship between compute and network usage for various microservices and solves an optimization problem in order to appropriately identify how each microservice should be deployed in the complex, multi-tiered compute and network fabric, so that the end-to-end application requirements are optimally met. We implemented two real-world 5G applications in video surveillance and intelligent transportation system (ITS) domains. Through extensive experiments, we show that ROMA is able to save up to 90%, 55% and 44% compute and up to 80%, 95% and 75% network bandwidth for the surveillance (watchlist) and transportation application (person and car detection), respectively. This improvement is achieved while honoring the application performance requirements, and it is over an alternative scheme that employs a static and overprovisioned resource allocation strategy by ignoring the resource coupling relationships. Anousheh Gholami, Kunal Rao, Wang-Pin Hsiung, Oliver Po, Murugan Sankaradass, Srimat T. Chakradhar |
NOMS | 5 |
| 2022 | Enhancing Video Analytics Accuracy via Real-time Automated Camera Parameter TuningabstractIn Video Analytics Pipelines (VAP), Analytics Units (AUs) such as object detection and face recognition running on remote servers critically rely on surveillance cameras to capture high-quality video streams in order to achieve high accuracy. Modern IP cameras come with a large number of camera parameters that directly affect the quality of the video stream capture. While a few of such parameters, e.g., exposure, focus, white balance are automatically adjusted by the camera internally, the remaining ones are not. We denote such camera parameters as non-automated (NAUTO) parameters. In this paper, we first show that environmental condition changes can have significant adverse effect on the accuracy of insights from the AUs, but such adverse impact can potentially be mitigated by dynamically adjusting NAUTO camera parameters in response to changes in environmental conditions. We then present CamTuner, to our knowledge, the first framework that dynamically adapts NAUTO camera parameters to optimize the accuracy of AUs in a VAP in response to adverse changes in environmental conditions. CamTuner is based on SARSA reinforcement learning and it incorporates two novel components: a light-weight analytics quality estimator and a virtual camera that drastically speed up offline RL training. Our controlled experiments and real-world VAP deployment show that compared to a VAP using the default camera setting, CamTuner enhances VAP accuracy by detecting 15.9% additional persons and 2.6%--4.2% additional cars (without any false positives) in a large enterprise parking lot and 9.7% additional cars in a 5G smart traffic intersection scenario, which enables a new usecase of accurate and reliable automatic vehicle collision prediction (AVCP). CamTuner opens doors for new ways to significantly enhance video analytics accuracy beyond incremental improvements from refining deep-learning models. Sibendu Paul, Kunal Rao, Giuseppe Coviello, Murugan Sankaradass, Oliver Po, Y. Charlie Hu, Srimat T. Chakradhar |
SenSys | 4 |
| 2022 | DyCo: Dynamic, Contextualized AI ModelsabstractDevices with limited computing resources use smaller AI models to achieve low-latency inferencing. However, model accuracy is typically much lower than the accuracy of a bigger model that is trained and deployed in places where the computing resources are relatively abundant. We describe DyCo, a novel system that ensures privacy of stream data and dynamically improves the accuracy of small models used in devices. Unlike knowledge distillation or federated learning, DyCo treats AI models as black boxes. DyCo uses a semi-supervised approach to leverage existing training frameworks and network model architectures to periodically train contextualized, smaller models for resource-constrained devices. DyCo uses a bigger, highly accurate model in the edge-cloud to auto-label data received from each sensor stream. Training in the edge-cloud (as opposed to the public cloud) ensures data privacy, and bespoke models for thousands of live data streams can be designed in parallel by using multiple edge-clouds. DyCo uses the auto-labeled data to periodically re-train, stream-specific, bespoke small models. To reduce the periodic training costs, DyCo uses different policies that are based on stride, accuracy, and confidence information. We evaluate our system, and the contextualized models, by using two object detection models for vehicles and people, and two datasets (a public benchmark and another real-world proprietary dataset). Our results show that DyCo increases the mAP accuracy measure of small models by an average of 16.3% (and up to 20%) for the public benchmark and an average of 19.0% (and up to 64.9%) for the real-world dataset. DyCo also decreases the training costs for contextualized models by more than an order of magnitude. Yi Yang 0018, Murugan Sankaradass, Srimat T. Chakradhar |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2021 | Edge-based fever screening system over private 5G
Murugan Sankaradass, Kunal Rao, Ravi K. Rajendran, Amit Redkar, Srimat T. Chakradhar |
SEC | 1 |
| 2021 | F3S: Free Flow Fever ScreeningabstractIdentification of people with elevated body temperature can reduce or dramatically slow down the spread of infectious diseases like COVID-19. We present a novel fever-screening system, F3S, that uses edge machine learning techniques to accurately measure core body temperatures of multiple individuals in a free-flow setting. F3S performs real-time sensor fusion of visual camera with thermal camera data streams to detect elevated body temperature, and it has several unique features: (a) visual and thermal streams represent very different modalities, and we dynamically associate semantically-equivalent regions across visual and thermal frames by using a new, dynamic alignment technique that analyzes content and context in real-time, (b) we track people through occlusions, identify the eye (inner canthus), forehead, face and head regions where possible, and provide an accurate temperature reading by using a prioritized refinement algorithm, and (c) we robustly detect elevated body temperature even in the presence of personal protective equipment like masks, or sunglasses or hats, all of which can be affected by hot weather and lead to spurious temperature readings. F3S has been deployed at over a dozen large commercial establishments, providing contact-less, free-flow, real-time fever screening for thousands of employees and customers in indoors and outdoor settings. Kunal Rao, Giuseppe Coviello, Min Feng 0001, Biplob Debnath, Wang-Pin Hsiung, Murugan Sankaradass, Yi Yang 0018, Oliver Po, Utsav Drolia, Srimat T. Chakradhar |
SMARTCOMP | 6 |
| 2013 | COSMIC: middleware for high performance and reliable multiprocessing on xeon phi coprocessors
Srihari Cadambi, Giuseppe Coviello, Cheng-Hong Li, Rajat Phull, Kunal Rao, Murugan Sankaradass, Srimat T. Chakradhar |
HPDC | 6 |
| 2010 | A dynamically configurable coprocessor for convolutional neural networksabstractConvolutional neural networks (CNN) applications range from recognition and reasoning (such as handwriting recognition, facial expression recognition and video surveillance) to intelligent text applications such as semantic text analysis and natural language processing applications. Two key observations drive the design of a new architecture for CNN. First, CNN workloads exhibit a widely varying mix of three types of parallelism: parallelism within a convolution operation, intra-output parallelism where multiple input sources (features) are combined to create a single output, and inter-output parallelism where multiple, independent outputs (features) are computed simultaneously. Workloads differ significantly across different CNN applications, and across different layers of a CNN. Second, the number of processing elements in an architecture continues to scale (as per Moore's law) much faster than off-chip memory bandwidth (or pin-count) of chips. Based on these two observations, we show that for a given number of processing elements and off-chip memory bandwidth, a new CNN hardware architecture that dynamically configures the hardware on-the-fly to match the specific mix of parallelism in a given workload gives the best throughput performance. Our CNN compiler automatically translates high abstraction network specification into a parallel microprogram (a sequence of low-level VLIW instructions) that is mapped, scheduled and executed by the coprocessor. Compared to a 2.3 GHz quad-core, dual socket Intel Xeon, 1.35 GHz C870 GPU, and a 200 MHz FPGA implementation, our 120 MHz dynamically configurable architecture is 4x to 8x faster. This is the first CNN architecture to achieve real-time video stream processing (25 to 30 frames per second) on a wide range of object detection and recognition tasks. Srimat T. Chakradhar, Murugan Sankaradass, Venkata Jakkula, Srihari Cadambi |
ISCA | 2 |
| 2009 | A Massively Parallel Coprocessor for Convolutional Neural NetworksabstractWe present a massively parallel coprocessor for accelerating Convolutional Neural Networks (CNNs), a class of important machine learning algorithms. The coprocessor functional units, consisting of parallel 2D convolution primitives and programmable units performing sub-sampling and non-linear functions specific to CNNs, implement a ldquometa-operatorrdquo to which a CNN may be compiled to. The coprocessor is serviced by distributed off-chip memory banks with large data bandwidth. As a key feature, we use low precision data and further increase the effective memory bandwidth by packing multiple words in every memory operation, and leverage the algorithmpsilas simple data access patterns to use off-chip memory as a scratchpad for intermediate data, critical for CNNs. A CNN is mapped to the coprocessor hardware primitives with instructions to transfer data between the memory and coprocessor. We have implemented a prototype of the CNN coprocessor on an off-the-shelf PCI FPGA card with a single Xilinx Virtex5 LX330T FPGA and 4 DDR2 memory banks totaling 1 GB. The coprocessor prototype can process at the rate of 3.4 billion multiply accumulates per second (GMACs) for CNN forward propagation, a speed that is 31x faster than a software implementation on a 2.2 GHz AMD Opteron processor. For a complete face recognition application with the CNN on the coprocessor and the rest of the image processing tasks on the host, the prototype is 6-10times faster, depending on the host-coprocessor bandwidth. Murugan Sankaradass, Venkata Jakkula, Srihari Cadambi, Srimat T. Chakradhar, Igor Durdanovic, Eric Cosatto, Hans Peter Graf |
ASAP | 1 |
| 2009 | A Massively Parallel FPGA-Based Coprocessor for Support Vector MachinesabstractWe present a massively parallel FPGA-based coprocessor for Support Vector Machines (SVMs), a machine learning algorithm whose applications include recognition tasks such as learning scenes, situations and concepts, and reasoning tasks such as analyzing the recognized scenes and semantics. The coprocessor architecture, targeted at both SVM training and classification, is based on clusters of vector processing elements (VPEs) operating in single-instruction multiple data (SIMD) mode to take advantage of large amounts of data parallelism in the application. We use the FPGA's DSP elements as parallel multiply-accumulators (MACs), a core computation in SVMs. A key feature of the architecture is that it is customized to low precision arithmetic which permits one DSP unit to perform two or more MACs in parallel. Low precision also reduces the required number of parallel off-chip memory accesses by packing multiple data words on the FPGA-memory bus. We have built a prototype using an off-the-shelf PCI-based FPGA card with a Xilinx Virtex 5 FPGA and 1 GB DDR2 memory. For SVM training, we observe application-level end-to-end computation speeds of over 9 billion multiply-accumulates per second (GMACs). For SVM classification, using data packing, the application speed increases to 14 GMACs. The FPGA-based system is about 20times faster than a dual Opteron 2.2 GHz processor CPU, and dissipates around 10 W of power. Srihari Cadambi, Igor Durdanovic, Venkata Jakkula, Murugan Sankaradass, Eric Cosatto, Srimat T. Chakradhar, Hans Peter Graf |
FCCM | 4 |
| 2008 | A Massively Parallel Digital Learning ProcessorabstractWe present a new, massively parallel architecture for accelerating machine learning algorithms, based on arrays of variable-resolution arithmetic vector processing elements (VPE). Groups of VPEs operate in SIMD (single instruction multiple data) mode, and each group is connected to an independent memory bank. In this way memory bandwidth scales with the number of VPE, and the main data flows are local, keeping power dissipation low. With 256 VPEs, implemented on two FPGA (field programmable gate array) chips, we obtain a sustained speed of 19 GMACS (billion multiply-accumulate per sec.) for SVM training, and 86 GMACS for SVM classification. This performance is more than an order of magnitude higher than that of any FPGA implementation reported so far. The speed on one FPGA is similar to the fastest speeds published on a Graphics Processor for the MNIST problem, despite a clock rate of the FPGA that is six times lower. High performance at low clock rates makes this massively parallel architecture particularly attractive for embedded applications, where low power dissipation is critical. Tests with Convolutional Neural Networks and other learning algorithms are under way now. Hans Peter Graf, Srihari Cadambi, Igor Durdanovic, Venkata Jakkula, Murugan Sankaradass, Eric Cosatto, Srimat T. Chakradhar |
NIPS | 5 |
| 2007 | Exploring Software Partitions for Fast Security Processing on a Multiprocessor Mobile SoCabstractThe functionality of mobile devices, such as cell phones and personal digital assistants (PDAs), has evolved to include various applications where security is a critical concern (secure web transactions, mobile commerce, download and playback of protected audio/video content, connection to corporate private networks, etc.). Security mechanisms (e.g., secure communication protocols) involve cryptographic algorithms, and are often quite computationally intensive, challenging the constrained processing and battery resources of mobile devices. Extensive design effort and aggressive hardware and software optimizations are required to address this challenge. Previous work has addressed the design of hardware architectures (custom accelerators, domain-specific processors, etc.) to accelerate security processing, and many emerging systems-on-chip (SoCs) feature some form of hardware support for security. In this paper, we address the complementary problem of mapping a complex security software library to an SoC platform with security hardware enhancements. We present a systematic methodology for exploring the software architecture for security processing for a commercial heterogeneous multiprocessor SoC for mobile devices. The SoC contains multiple host processors executing applications and a dedicated programmable security processing engine. We developed an exploration methodology to map the code and data of security software libraries onto the platform, with the objective of maximizing the overall application-visible performance. The salient features of the methodology include: 1) the use of real performance measurements from a prototyping board, which contains the target platform, to drive the exploration; 2) a new data structure access profiling framework that allows us to accurately model the communication overheads involved in off loading a given set of functions to the security processor; and 3) an exact branch-and-bound-based design space exploration algorithm that determines the best mapping of security library functions and data structures to the host and security processors. We used the proposed framework to map a commercial security library to the target mobile application SoC. The resulting optimized software architecture outperformed several manually designed software architectures, resulting in up to 12.5 times speed-up for individual cryptographic operations (encryption, hashing) and 2.2-6.2 times speed-up for applications such as a digital rights management (DRM) agent and secure sockets layer (SSL) client. We also demonstrate the applicability of our framework to software architecture exploration in other multiprocessor scenarios. Divya Arora 0001, Anand Raghunathan, Srivaths Ravi 0001, Murugan Sankaradass, Niraj K. Jha, Srimat T. Chakradhar |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2006 | Software architecture exploration for high-performance security processing on a multiprocessor mobile SoCabstractWe present a systematic methodology for exploring the security processing software architecture for a commercial heterogeneous multiprocessor system-on-chip (SoC) for mobile devices. The SoC contains multiple host processors executing applications and a dedicated programmable security processing engine. We developed an exploration methodology to map the code and data of security software libraries onto the platform, with the objective of maximizing the overall application-visible performance. The salient features of the methodology include (i) the use of real performance measurements from a prototyping board that contains the target platform to drive the exploration, (ii) a new data structure access profiling framework that allows us to accurately model the communication overheads involved in offloading a given set of functions to the security processor, and (iii) an exact branch-and-bound based design space exploration algorithm that determines the best mapping of security library functions and data structures to the host and security processors.We used the proposed framework to map a commercial security library to the target mobile application SoC. The resulting optimized software architecture outperformed several manually-designed software architectures, resulting in upto 12.5X speedup for individual cryptographic operations (encryption, hashing) and 2.2X-6.2X speedup for applications such as a Digital Rights Management (DRM) agent and Secure Sockets Layer (SSL) client. We also demonstrate the applicability of our framework to software architecture exploration in other multiprocessor scenarios. Divya Arora 0001, Anand Raghunathan, Srivaths Ravi 0001, Murugan Sankaradass, Niraj K. Jha, Srimat T. Chakradhar |
DAC | 4 |
| 2003 | CoCo: a hardware/software platform for rapid prototyping of code compression technologiesabstractIn recent years instruction code compression/decompression technologies have emerged as an efficient way to a) reduce the memory usage of an embedded system, b) to improve performance through effectively higher bandwidths and/or to c) reduce the overall power consumption of a system processing compressed code. We have presented efficient code compression/decompression techniques and architectures in the past. For the commercialization phase, we designed a novel hardware/software code compression/decompression platform (CoCo). It consists of a software platform that prepares, optimizes, compresses and compiles instruction code and a generic, parameterizable FPGA-based hardware architecture in form of a hardware platform that allows to rapidly evaluate prototypes of diverse compression/decompression technologies. We show the flexibility of CoCo, its ability to achieve code compression ratios (parameterizable) of up to 50% with a slight system performance gain and its ability to apply compression on real-world compiled code without any limitations where others have made implicit software-restrictive assumptions. Haris Lekatsas, Jörg Henkel, Srimat T. Chakradhar, Venkata Jakkula, Murugan Sankaradass |
DAC | 5 |
| 2002 | System design methodologies for a wireless security processing platformabstractSecurity protocols are critical to enabling the growth of a wide range of wireless data services and applications. However, they impose a high computational burden that is mismatched with the modest processing capabilities and battery resources available on wireless clients. Bridging the security processing gap, while retaining sufficient programmability in order to support a wide range of current and future security protocol standards, requires the use of novel system architectures and design methodologies.We present the system-level design methodology used to design a programmable security processor platform for next-generation wireless handsets. The platform architecture is based on (i) a configurable and extensible processor that is customized for efficient domain-specific processing, and (ii) layered software libraries implementing cryptographic algorithms that are optimized to the hardware platform. Our system-level design methodology enables the efficient co design of optimal cryptographic algorithms and an optimized system architecture. It includes novel techniques for algorithmic exploration and tuning, performance characterization and macro-modeling of software libraries, and architecture refinement based on selection of instruction extensions to accelerate performance-critical, computation-intensive operations. We have designed a programmable security processor platform to support both public-key and private key operations using the proposed methodology, and have evaluated its performance through extensive system simulations as well as hardware prototyping. Our experiments demonstrate large performance improvements (e.g., 31.0X for DES, 33.9X for 3DES, 17.4X for AES, and upto 66.4X for RSA) compared to well-optimized software implementations on a state-of-the-art embedded processor. Srivaths Ravi 0001, Anand Raghunathan, Nachiketh R. Potlapally, Murugan Sankaradass |
DAC | 4 |