Vanchinathan Venkataramani

dblp:130/1363 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
1since 2021 · last 2022
0000-0002-0259-6456ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 4 first-author · 1 since 2021Software engineering, systems software and programming languages · 3 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Parallel and multicore computing · 65% Embedded and real-time systems · 13% Processor architecture and microarchitecture · 10%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing › parallel scheduling
communication scheduling
0.612022
ASCENT: Communication Scheduling for SDF on Bufferless Software-Defined NoC · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2022
Embedded and real-time systems
real-time scheduling
0.612022
ASCENT: Communication Scheduling for SDF on Bufferless Software-Defined NoC · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2022
Parallel and multicore computing › dataflow computing › dataflow scheduling
synchronous dataflow scheduling
0.612022
ASCENT: Communication Scheduling for SDF on Bufferless Software-Defined NoC · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2022
Parallel and multicore computing
parallel scheduling
0.522017
Optimal Greedy Algorithm for Many-Core Scheduling · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017
Distributed scheduling for many-cores using cooperative game theory · DAC 2016
Processor architecture and microarchitecture
many-core architecture
0.422017
Defragmentation of Tasks in Many-Core Architecture · ACM Trans. Archit. Code Optim. 2017
Distributed scheduling for many-cores using cooperative game theory · DAC 2016
Parallel and multicore computing › task scheduling
many-core scheduling
0.312017
Optimal Greedy Algorithm for Many-Core Scheduling · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017
Parallel and multicore computing › parallel scheduling
runtime scheduling
0.312017
Optimal Greedy Algorithm for Many-Core Scheduling · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2017
Parallel and multicore computing
task allocation
0.312017
Defragmentation of Tasks in Many-Core Architecture · ACM Trans. Archit. Code Optim. 2017
Parallel and multicore computing
task scheduling
0.312017
Defragmentation of Tasks in Many-Core Architecture · ACM Trans. Archit. Code Optim. 2017
Energy-efficient computing › power management › system-level power management
hierarchical power management
0.212013
Hierarchical power management for asymmetric multi-core in dark silicon era · DAC 2013
Energy-efficient computing
power management
0.212013
Hierarchical power management for asymmetric multi-core in dark silicon era · DAC 2013
Parallel and multicore computing › task allocation
task-to-core mapping
0.112016
Distributed scheduling for many-cores using cooperative game theory · DAC 2016
Processor architecture and microarchitecture › multicore design › heterogeneous multicore
asymmetric multicore
0.012013
Hierarchical power management for asymmetric multi-core in dark silicon era · DAC 2013
Processor architecture and microarchitecture
multicore design
0.012013
Hierarchical power management for asymmetric multi-core in dark silicon era · DAC 2013
Energy-efficient computing
thermal design power
0.012013
Hierarchical power management for asymmetric multi-core in dark silicon era · DAC 2013
Energy-efficient computing
thermal management
0.012013
Hierarchical power management for asymmetric multi-core in dark silicon era · DAC 2013

Methods — techniques the papers use, named apart from their topics

task-to-core mapping · 0.6synchronous dataflow model · 0.6offline scheduling · 0.6polynomial-time algorithm · 0.3greedy algorithm · 0.3dynamic programming comparison · 0.3concavity analysis · 0.3NP-hardness analysis · 0.3cooperative game theory · 0.2control theory · 0.2
YearPublicationVenuePosition
2022 ASCENT: Communication Scheduling for SDF on Bufferless Software-Defined NoC
abstract
Bufferless software-defined network-on-chip (NoC) is a promising alternative to conventional dynamic routing as it offers predictable data movement with real-time guarantees. Existing time-division multiplexing (TDM)-based mechanisms for predictability assume the worst-case communication pattern (e.g., all-to-all) and compute a fixed schedule wherein the cores can only communicate during the allocated time slots. These approaches lead to low application throughput as they cannot adapt to application characteristics. In this article, we present an application specific, non-TDM-based communication scheduling mechanism for bufferless software-defined NoCs. We choose the synchronous dataflow (SDF) model of computation to represent the input streaming applications. We propose ASCENT, a novel offline approach that takes the SDF-specified streaming application and the NoC architecture as input, exploits the task interactions and the timing information in the SDF, and generates the task-to-core mapping and communication schedule that is represented compactly in hardware. ASCENT achieves$5.8\times $better performance on average than existing TDM-based NoCs and manages to achieve the performance of an ideal dynamically routed NoC, yet ensuring predictability.
Vanchinathan Venkataramani, Bruno Bodin, Aditi Kulkarni Mohite, Tulika Mitra, Li-Shiuan Peh
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2020 Unified Thread- and Data-Mapping for Multi-Threaded Multi-Phase Applications on SPM Many-Cores
abstract
Scratchpad Memories (SPMs) are more scalable than caches as they offer better performance with lower power and area overheads. This scalability advocates their suitability as on-chip memory in many-cores. However, SPM many-cores delegate the responsibility of thread- and data-mapping to the software. The mapping is especially challenging in the case of multi-threaded multi-phase applications. Threads from these applications exhibit both inter- and intra-phase data-sharing patterns. These patterns intricately intertwine thread- and data- mapping across phases. The accompanying qualitative mapping is the key to extract application performance on SPM many-cores.State-of-the-art framework for SPM many-cores performs thread- and data-mapping independently. Furthermore, it can only operate with single-phase multi-threaded applications. We are the first to propose in this work, a unified thread- and data-mapping framework for NoC-based SPM many-cores when executing multi-threaded multi-phase applications. Experimental evaluations show, on average, 1.36x performance improvement compared to the state-of-the-art framework for multi-threaded multi-phase applications.
Vanchinathan Venkataramani, Anuj Pathania, Tulika Mitra
DATE1
2020 Time-Predictable Software-Defined Architecture with Sdf-Based Compiler Flow for 5g Baseband Processing
abstract
The advent of 5G networks motivates the need for high-performance, low-power, time-predictable hardware that can handle the aggressive real-time latency and throughput requirements of baseband processing. With newer generations like 5G, programmable hardware that can adapt readily to network specification updates becomes a critical requirement. We introduce a software-defined array-based many-core architecture, called SPECTRUM, that couples lightweight predictable hardware components with a compiler flow that orchestrates the on-chip hardware resources. This design, by construction, provides timing guarantees with a programmable architecture. Our architecture and compiler flow are designed to support basestation baseband processing computation represented using deterministic Synchronous Data Flow (SDF) model of computation. SDF is commonly used to represent signal processing applications and fits well with real-time systems requirements. We demonstrate substantial power savings with SPECTRUM compared to existing DSPs while meeting the performance requirements.
Vanchinathan Venkataramani, Bruno Bodin, Aditi Kulkarni Mohite, Tulika Mitra, Li-Shiuan Peh
ICASSP1
2020 Simultaneous Progressing Switching Protocols for Timing Predictable Real-Time Network-on-Chips
abstract
Inter-core communication is a central challenge in many-core systems for which Network-on-chips (NoCs) have been demonstrated to scale well and to provide good overall performance. However, not only the distributed structure but also the link switching of NoCs have imposed a great challenge in the design and analysis for real-time systems where timing verification is mandatory. NoC protocols like worm-hole switching are designed with scalability and flexibility in mind, thus the existing link switching protocols usually consider each single link to be scheduled independently. The flexibility of such link-based arbitrations allows each packet to be distributed over multiple switches but also increases the number of possible link states (the number of flits in a buffer) that have to be considered in the worst-case timing analysis for real-time systems. To achieve timing predictability by design, we propose a family of less flexible switching protocols, called Simultaneous Progressing Switching Protocols (SP2), in which the links used by a flow either all simultaneously transmit one flit (if it exists) of this flow or none of them transmits any flit of this flow. Based on the all-or-nothing property of Sp2, we reduce the schedulability of the NoC to the uniprocessor self-suspension scheduling problem. Moreover, the proposed approach is not limited to any specific underlying routing protocols, which are usually constructed for deadlock avoidance instead of timing predictability.
Niklas Ueter, Jian-Jia Chen, Georg von der Brüggen, Vanchinathan Venkataramani, Tulika Mitra
RTCSA4
2020 SPECTRUM: A Software-defined Predictable Many-core Architecture for LTE/5G Baseband Processing
abstract
Wireless communication standards such as Long-term Evolution (LTE) are rapidly changing to support the high data-rate of wireless devices. The physical layer baseband processing has strict real-time deadlines, especially in the next-generation applications enabled by the 5G standard. Existing basestation transceivers utilize customized DSP cores or fixed-function hardware accelerators for physical layer baseband processing. However, these approaches incur significant non-recurring engineering costs and are inflexible to newer standards or updates. Software-programmable processors offer more adaptability. However, it is challenging to sustain guaranteed worst-case latency and throughput at reasonably low-power on shared-memory many-core architectures featuring inherently unpredictable design choices, such as caches and Network-on-chip (NoC). We propose SPECTRUM , a predictable, software-defined many-core architecture that exploits the massive parallelism of the LTE/5G baseband processing workload. The focus is on designing scalable lightweight hardware that can be programmed and defined by sophisticated software mechanisms. SPECTRUM employs hundreds of lightweight in-order cores augmented with custom instructions that provide predictable timing, a purely software-scheduled NoC that orchestrates the communication to avoid any contention, and per-core software-controlled scratchpad memory with deterministic access latency. Compared to many-core architecture like Skylake-SP (average power 215 W) that drops 14% packets at high-traffic load, 256-core SPECTRUM by definition has zero packet drop rate at significantly lower average power of 24 W. SPECTRUM consumes 2.11× lower power than C66x DSP cores+accelerator platform in baseband processing. We also enable SPECTRUM to handle dynamic workloads with multiple service categories present in 5G mobile network (Enhanced Mobile Broadband (eMBB), Ultra-reliable and Low-latency Communications (URLLC), and Massive Machine Type Communications (mMTC)), using a run-time scheduling and mapping algorithm. Experimental evaluations show that our algorithm performs task/NoC mapping at run-time on fewer cores compared to the static mapping (that reserves cores exclusively for each service category) while still meeting the differentiated latency and reliability requirements.
Vanchinathan Venkataramani, Aditi Kulkarni Mohite, Tulika Mitra, Li-Shiuan Peh
ACM Trans. Embed. Comput. Syst.1
2019 SPECTRUM: a software defined predictable many-core architecture for LTE baseband processing
abstract
Wireless communication standards such as Long Term Evolution (LTE) are rapidly changing to support the high data rate of wireless devices. The physical layer baseband processing has strict real-time deadlines, especially in the next-generation applications enabled by the 5G standard. Existing base station transceivers utilize customized Digital Signal Processing (DSP) cores or fixed-function hardware accelerators for physical layer baseband processing. However, these approaches incur significant non-recurring engineering costs and are inflexible to newer standards or updates. Software programmable processors offer more adaptability. However, it is challenging to sustain guaranteed worst-case latency and throughput at reasonably low-power on shared-memory many-core architectures featuring inherently unpredictable design choices, such as caches and network-on chip. We propose SPECTRUM, a predictable software defined many-core architecture that exploits the massive parallelism of the LTE baseband processing. The focus is on designing a scalable lightweight hardware that can be programmed and defined by sophisticated software mechanisms. SPECTRUM employs hundreds of lightweight in-order cores augmented with custom instructions that provide predictable timing, a purely software-scheduled on-chip network that orchestrates the communication to avoid any contention and per-core software controlled scratchpad memory with deterministic access latency. Compared to a many-core architecture like Skylake-SP (average power 215W) that drops 14% packets at high traffic load, 256-core SPECTRUM by definition has zero packet drop rate at significantly lower average power of 24W. SPECTRUM consumes 2.11x lower power than C66x DSP cores+accelerator platform in baseband processing. SPECTRUM is also well-positioned to support future 5G workloads.
Vanchinathan Venkataramani, Aditi Kulkarni Mohite, Tulika Mitra, Li-Shiuan Peh
LCTES1
2019 Scratchpad-Memory Management for Multi-Threaded Applications on Many-Core Architectures
abstract
Contemporary many-core architectures, such as Adapteva Epiphany and Sunway TaihuLight, employ per-core software-controlled Scratchpad Memory (SPM) rather than caches for better performance-per-watt and predictability. In these architectures, a core is allowed to access its own SPM as well as remote SPMs through the Network-On-Chip (NoC). However, the compiler/programmer is required to explicitly manage the movement of data between SPMs and off-chip memory. Utilizing SPMs for multi-threaded applications is even more challenging, as the shared variables across the threads need to be placed appropriately. Accessing variables from remote SPMs with higher access latency further complicates this problem as certain links in the NoC may be heavily contended by multiple threads. Therefore, certain variables may need to be replicated in multiple SPMs to reduce the contention delay and/or the overall access time. We present Coordinated Data Management (CDM), a compile-time framework that automatically identifies shared/private variables and places them with replication (if necessary) to suitable on-chip or off-chip memory, taking NoC contention into consideration. We develop both an exact Integer Linear Programming (ILP) formulation as well as an iterative, scalable algorithm for placing the data variables in multi-threaded applications on many-core SPMs. Experimental evaluation on the Parallella hardware platform confirms that our allocation strategy reduces the overall execution time and energy consumption by 1.84× and 1.83× , respectively, when compared to the existing approaches.
Vanchinathan Venkataramani, Mun Choon Chan, Tulika Mitra
ACM Trans. Embed. Comput. Syst.1
2018 LOCUS: Low-Power Customizable Many-Core Architecture for Wearables
abstract
Application requirements, such as real-time response, are pushing wearable devices to leverage more powerful processors inside the SoC (system on chip). However, existing wearable devices are not well suited for such challenging applications due to poor performance, and the conventional powerful many-core architectures are not appropriate either due to the stringent power budget in this domain. We propose LOCUS—a low-power, customizable, many-core processor for next-generation wearable devices. LOCUS combines customizable processor cores with a customizable network on a message-passing architecture to deliver very competitive performance/watt—an average 3.1× compared to quad-core ARM processors used in state-of-the-art wearable devices. A combination of full system simulation with representative applications from the wearable domain and RTL synthesis of the architecture show that 16-core LOCUS achieves an average 1.52× performance/watt improvement over a conventional 16-core shared memory many-core architecture. A dynamic power management mechanism is proposed to further decrease the power consumption in both computation and communication, which improves the performance/watt of LOCUS by 1.17×.
Cheng Tan 0002, Aditi Kulkarni Mohite, Vanchinathan Venkataramani, Manupa Karunaratne, Tulika Mitra, Li-Shiuan Peh
ACM Trans. Embed. Comput. Syst.3
2017 Defragmentation of Tasks in Many-Core Architecture
abstract
Many-cores can execute multiple multithreaded tasks in parallel. A task performs most efficiently when it is executed over a spatially connected and compact subset of cores so that performance loss due to communication overhead imposed by the task’s threads spread across the allocated cores is minimal. Over a span of time, unallocated cores can get scattered all over the many-core, creating fragments in the task mapping. These fragments can prevent efficient contiguous mapping of incoming new tasks leading to loss of performance. This problem can be alleviated by using a task defragmenter, which consolidates smaller fragments into larger fragments wherein the incoming tasks can be efficiently executed. Optimal defragmentation of a many-core is an NP-hard problem in the general case. Therefore, we simplify the original problem to a problem that can be solved optimally in polynomial time. In this work, we introduce a concept of exponentially separable mapping (ESM), which defines a set of task mapping constraints on a many-core. We prove that an ESM enforcing many-core can be defragmented optimally in polynomial time.
Anuj Pathania, Vanchinathan Venkataramani, Muhammad Shafique 0001, Tulika Mitra, Jörg Henkel
ACM Trans. Archit. Code Optim.2
2017 Optimal Greedy Algorithm for Many-Core Scheduling
abstract
In this paper, we propose an optimal greedy algorithm for the problem of run-time many-core scheduling. The previously best known centralized optimal algorithm proposed for the problem is based on dynamic programming. A dynamic programming-based scheduler has high overheads which grow fast with increase in both the number of cores in the many-cores as well as number of tasks independently executing on them. We show in this paper that the inherent concavity of extractable instructions per cycle in tasks with increase in number of allocated cores allows for an alternative greedy algorithm. The proposed algorithm significantly reduces the run-time scheduling overheads, while maintaining theoretical optimality. In practice, it reduces the problem solving time 10 000x to provide near-optimal solutions.
Anuj Pathania, Vanchinathan Venkataramani, Muhammad Shafique 0001, Tulika Mitra, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2016 LOCUS: low-power customizable many-core architecture for wearables
abstract
The requirements' demands of applications, such as real-time response, are pushing the wearable devices to leverage more power-efficient processors inside the SoC (System-on-chip). However, existing wearable devices are not well suited for such challenging applications due to poor performance, while the conventional powerful many-core architectures are not appropriate either due to the stringent power budget in this domain. We propose LOCUS - a low-power, customizable, many-core processor for next-generation wearable devices. LOCUS combines customizable processor cores with a customizable network on a message-passing architecture to deliver very competitive performance/watt - an average 3.1x compared to quad-core ARM processors used in the state-of-the-art wearable devices. A combination of full-system simulation with representative applications from wearable domain and RTL synthesis of the architecture show that 16-core LOCUS achieves an average 1.52x performance/watt improvement over a conventional 16-core shared-memory many-core architecture.
Cheng Tan 0002, Aditi Kulkarni Mohite, Vanchinathan Venkataramani, Manupa Karunaratne, Tulika Mitra, Li-Shiuan Peh
CASES3
2016 Distributed scheduling for many-cores using cooperative game theory
abstract
Many-cores are envisaged to include hundreds of processing cores etched on to a single die and will execute tens of multi-threaded tasks in parallel to exploit their massive parallel processing potential. A task can be sped up by assigning it to more than one core. Moreover, processing requirements of tasks are in a constant state of flux and some of the cores assigned to a task entering a low processing requirement phase can be transferred to a task entering high requirement phase, maximizing overall performance of the system.
Anuj Pathania, Vanchinathan Venkataramani, Muhammad Shafique 0001, Tulika Mitra, Jörg Henkel
DAC2
2016 Distributed fair scheduling for many-cores
Anuj Pathania, Vanchinathan Venkataramani, Muhammad Shafique 0001, Tulika Mitra, Jörg Henkel
DATE2
2014 Design space exploration of multiple loops on FPGAs using high level synthesis
abstract
Real-world applications such as image processing, signal processing, and others often contain a sequence of computation intensive kernels, each represented in the form of a nested loop. High-level synthesis (HLS) enables efficient hardware implementation of these loops using high-level programming languages. HLS tools also allow the designers to evaluate design choices with different trade-offs through pragmas/directives. Prior design space exploration techniques for HLS primarily focus on either single nested loop or multiple loops without consideration to the data dependencies among them. In this paper, we propose efficient design space exploration techniques for applications that consist of multiple nested loops with or without data dependencies. In particular, we develop an algorithm to derive the Pareto-optimal curve (performance versus area) of the application when mapped onto FPGAs using HLS. Our algorithm is efficient as it effectively prunes the dominated points in the design space. We also develop accurate performance and area models to assist the design space exploration process. Experiments on various scientific kernels and real-world applications demonstrate that our design space exploration technique is accurate and efficient.
Guanwen Zhong, Vanchinathan Venkataramani, Yun Liang 0001, Tulika Mitra, Smaïl Niar
ICCD2
2013 Power-performance modeling on asymmetric multi-cores
abstract
Asymmetric multi-core architectures have recently emerged as a promising alternative in a power and thermal constrained environment. They typically integrate cores with different power and performance characteristics, which makes mapping of workloads to appropriate cores a challenging task. Limited number of performance counters and heterogeneous memory hierarchy increase the difficulty in predicting the performance and power consumption across cores in commercial asymmetric multi-core architectures. In this work, we propose a software-based modeling technique that can estimate performance and power consumption of workloads for different core types. We evaluate the accuracy of our technique on ARM big. LITTLE asymmetric multi-core platform.
Mihai Pricopi, Thannirmalai Somu Muthukaruppan, Vanchinathan Venkataramani, Tulika Mitra, Sanjay Vishin
CASES3
2013 Hierarchical power management for asymmetric multi-core in dark silicon era
abstract
Asymmetric multi-core architectures integrating cores with diverse power-performance characteristics is emerging as a promising alternative in the dark silicon era where only a fraction of the cores on chip can be powered on due to thermal limits. We introduce a hierarchical power management framework for asymmetric multi-cores that builds on control theory and coordinates multiple controllers in a synergistic manner to achieve optimal power-performance efficiency while respecting the thermal design power budget. We integrate our framework within Linux and implement/evaluate it on real ARM big.LITTLE asymmetric multi-core platform.
Thannirmalai Somu Muthukaruppan, Mihai Pricopi, Vanchinathan Venkataramani, Tulika Mitra, Sanjay Vishin
DAC3