Deepak Gangadharan

dblp:98/7452 · DBLP profile ↗
← Back
29ranked-venue papers
8as first author
15since 2021 · last 2025
0000-0001-6630-0012ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 4 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 RL-ALMS: Reinforcement Learning-based Adaptive and Lightweight Model Selection Framework for Energy-Efficient Lane Detection in Electric Vehicles
abstract
Lane detection is critical for Advanced Driver Assistance Systems (ADAS) in autonomous vehicles. Deep learning models offer high accuracy but are computationally demanding for resource-constrained edge platforms, while lightweight image processing algorithms are energy-efficient but often falter in complex scenarios. This paper presents RL-ALMS (Reinforcement Learning-based Adaptive and Lightweight Model Selection), a framework that dynamically switches between deep learning and image processing lane detection models in electric vehicles (EVs). RL-ALMS employs a contextual multi-armed bandit with Thompson sampling to intelligently select the optimal model based on road conditions, battery state, and accuracy needs. The novelty lies in its system-level integration of a battery-aware reward function and a sustainability-driven selection mechanism, specifically designed to balance real-time detection accuracy with long-term EV energy constraints. Experimental results on an edge computing platform demonstrate RL-ALMS achieves 90.3% average accuracy over a 50 km journey while maintaining $\mathbf{8 1 . 2 \%}$ battery reserve, significantly outperforming fixed highaccuracy (battery depletion) and energy-efficient ($64.5 \%$ accuracy) strategies. RL-ALMS maintains consistent performance across diverse conditions, representing a practical advancement for energy-constrained, real-time lane detection in EVs.
Rhuthik P, Deepak Gangadharan
DSD2
2025 Network Traversal Time (NTT) Analysis of ST Flows with Non-zero Arrival Jitter in TSN Networks
Pavan Kumar Kondooru, Deepak Gangadharan
VECoS2
2025 HAPTA: Adaptive Learning Based Horizontal Computation Offloading over RSUs in VEC
abstract
The rapid evolution of vehicular edge computing (VEC) has made task offloading to roadside units (RSUs) essential for addressing the computational demands of modern vehicular applications. However, efficiently partitioning and scheduling tasks across RSUs in dynamic environments remains challenging. This paper introduces a novel algorithm, HAPTA (Horizontal, Adaptive Learning-based, Partitioned, Timeslot-based Algorithm), a time-slot based task partitioning and offloading approach using the Upper Confidence Bound (UCB) algorithm, to adaptively allocate tasks to RSUs. We consider task offloading horizontally across RSU nodes. Unlike static or heuristic-based scheduling methods, the proposed HAPTA framework dynamically selects RSUs with optimal resource availability while considering vehicular mobility and task deadlines. The framework supports splitting tasks into subtasks, enabling efficient resource utilization and meeting stringent latency constraints. Experimental evaluations on real-world vehicular datasets demonstrate that the HAPTA approach outperforms previously introduced heuristic-based methods, achieving higher task completion rates and lower latency, particularly in high-density scenarios.
Tanniru Abhinav Siddharth, Deepak Gangadharan
VTC2025-Spring2
2025 Reinforcement learning for traffic signal control: advancing efficiency through hybrid exploration strategies
Saidulu Thadikamalla, Piyush Joshi, Deepak Gangadharan
J. Supercomput.3
2024 Checkpointing-Aware End-to-End Data Age Analysis of Task Chains under Transient Faults
abstract
Safety-critical real-time systems are conventionally designed to meet specific hard real-time deadline requirements. Many applications in such systems consist of task chains, which exhibit complex temporal dependencies based on the activation rates of the tasks. As a result, there has been an increasing interest in analyzing the age of data (i.e., the time elapsed since the data was first updated) that propagates through the task chains as an important system performance parameter. Transient faults in safety-critical systems can lead to undesirable consequences. One of the most common methods used for fault tolerance in such systems is checkpointing. However, the interference due to checkpointing can significantly impact the data age, reducing the predictability of the system’s performance.This paper proposes an analytical framework for calculating the data age in real-time task chains with checkpointing for transient faults in both single-core and multi-core platforms. The proposed analysis has been validated by comparing it with extensive simulations of the execution of real-time task chain-based systems under faults. It shows comparable data age values with low time complexity.
Sridhar Mallareddy, Pavan Kumar Kondooru, Deepak Gangadharan
ISORC3
2024 Preliminary Modeling of Energy-Aware Integrated Allocations in Robotic Mobile Fulfillment Systems
abstract
The Robotic Mobile Fulfillment System (RMFS) is an automated technology to fulfill various types of orders (e.g., vehicle assembly) in modern warehouses or factories. In particular, battery-equipped mobile robots move racks that contain various components (Stock Keeping Unit) by visiting replenishment or picking stations to complete orders. We formulate the preliminary system model to optimize the order throughput and energy consumption of mobile robots, which consists of (i) rack allocation, (ii) task allocation, and (iii) route routing.
Kyujin Kyung, Deepak Gangadharan, BaekGyu Kim
RTCSA2
2024 Grid LSTM based Attention Modelling for Traffic Flow Prediction
abstract
Traffic flow prediction is an important task that can directly impact the control of traffic flow positively and improve the overall traffic throughput. Although a large number of studies have been performed to improve traffic flow prediction, there are very few works on purely temporal prediction models, which is important for execution on an edge device that does not have access to spatial flow information. In order to explore the temporal prediction models further, we propose an innovative hybrid long short-term memory (LSTM) model, which we call Grid LSTM based Attention Modelling for Traffic Flow Prediction (GLSTM-A), that helps to encode temporal information better at various levels/scenarios. The proposed architecture incorporates a Grid LSTM to capture historical dependencies and a simple LSTM layer dedicated to the short-term analysis of recent data. Moreover, an innovative attention mechanism is designed to focus on the importance of data features automatically for further enhancing the model's predictive capabilities. Our proposed GLSTM-A outperforms other popular temporal prediction models such as temporal convolutional network (TCN), Bi-LSTM and LSTM, in terms of prediction accuracy and memory efficiency as mentioned in the experimental results. Experimental results and ablation studies on benchmark datasets demonstrate the superior performance of the proposed model over existing state-of-the-art models in various time series prediction tasks.
Rahul Biju, Sai Usha Nagasri Goparaju, Deepak Gangadharan, Bappaditya Mandal
VTC Spring3
2024 Online Partitioned Scheduling over RSU for Computation Offloading in Vehicular Edge Computing
abstract
With advances in vehicular communications technology such as Vehicle-to-Infrastructure (V2I) and Vehicle-to-Vehicle (V2V), the computation of vehicular tasks has become distributed facilitated by computation offloading from the vehicular platform to road side units (RSUs) or the cloud infrastructure. In this work, due to communication overheads, we do not consider the cloud and only consider offloading horizontally across RSU nodes. However, the advantage of offloading depends upon how the tasks are scheduled across the RSUs. Unlike existing horizontal offloading works, this work explores the benefits of online task partitioned scheduling for computation offloading from multiple vehicles to RSUs or edge nodes. Specifically, we propose a new efficient timeslot-based online hybrid partitioned scheduling algorithm, which splits some tasks into subtasks and schedules them across RSU nodes while considering vehicle flow constraints. We compared and evaluated the effectiveness of our proposed hybrid partitioned scheduling algorithm with the fully partitioned algorithm. We also compared the performance of the aforementioned algorithms with an optimal scheduling algorithm utilizing several experiments conducted on a real-world vehicular dataset.
Tanniru Abhinav Siddharth, Kethu Sesha Sarath Reddy, Joseph John Cherukara, Deepak Gangadharan
VTC Fall4
2023 Time Series-based Driving Event Recognition for Two Wheelers
abstract
Classification of a motorcycle's driving events can provide deep insights to detect issues related to driver safety. In order to perform the above, we developed a hardware system with 3-D accelerometer/gyroscope sensors that can be deployed on a motorcycle. The data obtained from these sensors is used to identify various driving events. We firstly investigated several machine learning (ML) models to classify driving events. However, in this process, we identified that though the overall accuracy of these traditional ML models is decent enough, the class-wise accuracy of these models is poor. Hence, we have developed time-series-based classification algorithms using LSTM and Bi-LSTM to classify various driving events. The experiments conducted have demonstrated that the proposed models have surpassed the state-of-the-art models in the context of driving event recognition with better class-wise accuracies. We have also deployed these models on an edge device (Raspberry Pi) with similar prediction accuracies. The experiments demonstrated that the proposed Bi-LSTM model showed a minimum of 86% accuracy in the case of a Left Turn (LT) event and a maximum of 99% accuracy for the event Stop (ST) in class-wise prediction when implemented on Raspberry Pi for a two wheeler driving dataset.
Sai Usha Nagasri Goparaju, L. Lakshmanan, Abhinav Navnit, Rahul Biju, Lovish B, Deepak Gangadharan, Aftab M. Hussain
DATE6
2023 Dynamic Data Delivery Framework for Connected Vehicles via Edge Nodes with Variable Routes
abstract
With increasing connectivity and sophisticated software, modern vehicles are able to leverage different kinds of services provided by the environment. One such service recommended by the Automotive Edge Computing Consortium (AECC) is downloading high-definition map data by vehicles. This high volume of data can be provided to the vehicles when moving by pre-allocating resources on edge server nodes or roadside units if the routes are known apriori. However, this is not a realistic assumption to make in general. Therefore, in this work, we propose a two-stage optimization framework for efficient data delivery to connected vehicles via edge nodes while considering dynamic route changes. We have evaluated the efficiency of this proposed approach (considering a real-world dataset) with respect to (a) offline optimization strategies considering fixed routes and (b) a greedy approach considering route changes. Our proposed approach works considerably better than the existing approaches in the context of dynamic route changes.
Joseph John Cherukara, SVSLN Surya Suhas Vaddhiparthy, Deepak Gangadharan, BaekGyu Kim
VTC Fall3
2023 Optimization and Performance Evaluation of Hybrid Deep Learning Models for Traffic Flow Prediction
abstract
Traffic flow prediction has been regarded as a critical problem in intelligent transportation systems. An accurate prediction can help mitigate congestion and other societal problems while facilitating safer, cost and time-efficient travel. However, this requires the prediction algorithm to consider several complex characteristics of traffic flow data. These complex characteristics are an amalgamation of the spatial, temporal and periodic features exhibited by traffic flow data. To extract and leverage these features for traffic flow prediction, several hybrid deep learning models have been developed recently; however, there are still some challenges to determine the optimal architecture considering both spatial and temporal features. In this work, we perform an extensive comparison of hybrid deep learning models with and without periodicity to understand the prediction accuracy of these popular models. We propose an optimization framework that unifies genetic algorithm (GA) embedded optimization with prediction models in order to derive optimized deep learning architectures for hybrid traffic flow prediction exploring 2D spatial and temporal information. The framework enables the improvement of prediction performance and eliminates the hand-tuning process. An improved temporal convolutional network (TCN) architecture is derived using the GA driven optimization, which achieves superior traffic flow prediction accuracy compared to all other existing hybrid deep learning models on the freeway and urban traffic data from the PeMS traffic data set. We also evaluate the performance of the derived hybrid deep learning algorithms on the Raspberry PI embedded platform.
Sai Usha Nagasri Goparaju, Rahul Biju, Pravalika M, Bhavana MC, Deepak Gangadharan, Bappaditya Mandal, Pradeep C
VTC2023-Spring5
2023 Time-Series based Fall Detection in Two-Wheelers
abstract
Driving event recognition plays a crucial role in understanding and enhancing road safety. This research focuses on developing efficient time-series based models for Fall detection in two-wheelers. Traditional machine learning models proved inadequate in accurately classifying Fall scenarios due to their inability to capture temporal transitions in kinematic states. To address this limitation, time-series based Deep Learning (DL) models are proposed, utilizing Long Short-Term Memory (LSTM) networks. These networks enable direct learning from raw time series data, eliminating the need for manual feature engineering. Additionally, Bi-LSTMs were employed to capture contextual information from both past and future timesteps, further improving the model’s understanding of driving events. The architecture was enhanced with an attention mechanism to boost accuracy. Experimental results showcased that the proposed Bi-LSTM model achieved an overall accuracy of 97%, with a specific accuracy of approximately 92% in detecting Fall scenarios. This research contributes to the development of an accurate Time-series based system for Fall detection, facilitating improved road safety in the context of two-wheelers.
Sai Usha Nagasri Goparaju, Keerthi Pothalaraju, Shriya Dullur, Arihant Jain, Deepak Gangadharan
VTC Fall5
2023 Collision-Aware Data Delivery Framework for Connected Vehicles via Edges
abstract
With the rapid advancements in communication technologies, the paradigm of connected vehicles is drastically transforming the automotive industry, enabling efficient data and service delivery to vehicles via edges. Various works have considered data delivery without a Medium Access Layer (MAC), which can result in multiple data frame collisions in the network. The time-slot-based MAC layer strategy uses slot assignment to ensure collision-free data delivery for multiple vehicles across various transmission channels at each edge. However, the increasing requests from various vehicular nodes can increase network congestion, thus servicing fewer vehicles. In the current work, we propose an optimization framework for collision-aware data delivery considering two state-of-art MAC protocols, HCCA and VeMAC. The proposed framework minimizes the global slot utilization cost for edge-to-vehicle data delivery, considering vehicle flow, edge resources, and vehicle overlaps while avoiding possible data transmission collisions. Further, we demonstrate the practicality of the framework in terms of the number of vehicles served, global slot utilization cost, and bandwidth cost. We further analyze the framework with differences in vehicle densities for various problem sizes using a real-world traffic scenario.
SVSLN Surya Suhas Vaddhiparthy, Joseph John Cherukara, Deepak Gangadharan, BaekGyu Kim
VTC Fall3
2022 Global Edge Bandwidth Cost Gradient-based Heuristic for Fast Data Delivery to Connected Vehicles under Vehicle Overlaps
abstract
The emergence of vehicle connectivity technologies and associated applications have paved the way for increased consumer interest in connected vehicles. These modern day vehicles are now capable of sending/receiving vast amounts of data and offloading computation (which is one possible service) to servers thereby improving safety, comfort, driving experience, etc. In the early stages of connectivity, all the data communication and computation offloading happened between the cloud server and the vehicles. However, this is not feasible in scenarios having strict timing requirements and bandwidth cost constraints. Vehicular Edge Computing (VEC) demonstrated an efficient way to tackle the above problem. In order to optimally utilize the resources of the edge servers for data delivery, an efficient edge resource allocation framework needs to be developed. In a recent work, data/service delivery to connected vehicles assumed a worst-case scenario that all vehicles with routes passing through an edge appear in the edge coverage region simultaneously. However, this worst-case scenario is very pessimistic, which results in overestimation of edge resources. We address this by precisely computing the set of vehicles which simultaneously appear in the coverage region of an edge (which we call vehicle overlaps). In this work, we first propose an optimization framework for edge resource allocation that minimizes the bandwidth cost of data delivery to connected vehicles while considering the traffic flow and vehicle overlaps. Then, we propose an efficient heuristic to deliver data based on minimizing global edge bandwidth cost gradient under vehicle overlaps. We demonstrate the improvement in resource allocation considering vehicle overlaps. Using real world traffic data, we also demonstrate reduction in data delivery times using the proposed heuristic.
Akshaj Gupta, Joseph John Cherukara, Deepak Gangadharan, BaekGyu Kim, Oleg Sokolsky, Insup Lee 0001
VTC Spring3
2021 E-PODS: A Fast Heuristic for Data/Service Delivery in Vehicular Edge Computing
abstract
With the rise in state-of-the-art communication modes for vehicles such as vehicle to vehicle (V2V), vehicle to infrastructure (V2I) and vehicle to cloud (V2C), modern vehicles are increasingly being connected to cloud and fog/edge nodes. These vehicle connectivity modes have enabled the realization of Vehicular Edge Computing (VEC) paradigm, whereby vehicles can leverage fog/edge node resources for storage/computation. In a VEC system, vehicles receive very important and large quantity of data from edge nodes, which is termed as data delivery. In addition, edge nodes can execute some services and send the results back to the vehicle, which is called service delivery. Fast and efficient edge resource allocation for data/service delivery is important in order to serve as many vehicles as possible in the VEC system. However, edge resource allocation is complex with large number of edges and vehicles, while also considering vehicle flow parameters. In this work, we propose Edge-Pairwise Optimal Data/Service Delivery (E-PODS), which is a fast and efficient heuristic for data/service delivery. Through experiments with synthetic and real vehicular traces, we demonstrate that E-PODS is considerably faster than the optimal approach, while making resource allocations that are close to optimal in terms of total edge bandwidth cost and number of serviced vehicles.
Akshaj Gupta, Joseph John Cherukara, Deepak Gangadharan, BaekGyu Kim, Oleg Sokolsky, Insup Lee 0001
VTC Spring3
2018 Bandwidth Optimal Data/Service Delivery for Connected Vehicles via Edges
abstract
The paradigm of connected vehicles is fast gaining lot of attraction in the automotive industry. Recently, a lot of technological innovation has been pushed through to realize this paradigm using vehicle to cloud (V2C), infrastructure (V2I) and vehicle (V2V) communications. This has also opened the doors for efficient delivery of data/service to the vehicles via edge devices that are closer to the vehicles. In this work, we propose an optimization framework that can be used to deliver data/service to the connected vehicles such that a bandwidth cost objective is optimized. For the first time, we also integrate a vehicle flow model in the optimization framework to model the traffic flow in the coverage area of the edges. Using the optimization framework, we study the variation of the optimal bandwidth cost for varying problem sizes and vehicle flow model parameter values for both data and service delivery.
Deepak Gangadharan, Oleg Sokolsky, Insup Lee 0001, BaekGyu Kim, Chung-Wei Lin, Shinichi Shiraishi
IEEE CLOUD1
2018 Data Freshness Over-Engineering: Formulation and Results
abstract
In many application scenarios, data consumed by real-time tasks are required to meet a maximum age, or freshness, guarantee. In this paper, we consider the end-to-end freshness constraint of data that is passed along a chain of tasks in a uniprocessor setting. We do so with few assumptions regarding the scheduling algorithm used. We present a method for selecting the periods of tasks in chains of length two and three such that the end-to-end freshness requirement is satisfied, and then extend our method to arbitrary chains. We perform evaluations of both methods using parameters from an embedded benchmark suite (E3S) and several schedulers to support our result.
Dagaen Golomb, Deepak Gangadharan, Sanjian Chen, Oleg Sokolsky, Insup Lee 0001
ISORC2
2018 A Design-Time/Run-Time Application Mapping Methodology for Predictable Execution Time in MPSoCs
abstract
Executing multiple applications on a single MPSoC brings the major challenge of satisfying multiple quality requirements regarding real-time, energy, and so on. Hybrid application mapping denotes the combination of design-time analysis with run-time application mapping. In this article, we present such a methodology, which comprises a design space exploration coupled with a formal performance analysis. This results in several resource reservation configurations, optimized for multiple objectives, with verified real-time guarantees for each individual application. The Pareto-optimal configurations are handed over to run-time management, which searches for a suitable mapping according to this information. To provide any real-time guarantees, the performance analysis needs to be composable and the influence of the applications on each other has to be bounded. We achieve this either by spatial or a novel temporal isolation for tasks and by exploiting composable networks-on-chip (NoCs). With the proposed temporal isolation, tasks of different applications can be mapped to the same resource, while, with spatial isolation, one computing resource can be exclusively used by only one application. The experiments reveal that the success rate in finding feasible application mappings can be increased by the proposed temporal isolation by up to 30% and energy consumption can be reduced compared to spatial isolation.
Andreas Weichslgartner, Stefan Wildermann, Deepak Gangadharan, Michael Glaß, Jürgen Teich
ACM Trans. Embed. Comput. Syst.3
2017 Extensible Energy Planning Framework for Preemptive Tasks
abstract
Cyber-physical systems (CSPs) are demanding energy-efficient design not only of hardware (HW), but also of software (SW). Dynamic Voltage and Frequency Scaling (DVFS) and Dynamic Power Manage (DPM) are most popular techniques to improve the energy efficiency. However, contemporary complicated HW and SW designs requires more elaborate and sophisticated energy management and efficiency evaluation techniques. This paper is concerned about energy supply planning for real-time scheduling systems (units) of which tasks need to meet deadlines. This paper presents a model-based compositional energy planning technique that computes a minimal ratio of processor frequency that preserves schedulability of independent and preemptive tasks. The minimal ratio of processor frequency can be used to plan the energy supply of real-time components. Our model-based technique is extensible by refining our model with additional features so that energy management techniques and their energy efficiency can be evaluated by model checking techniques. We exploit the compositional framework for hierarchical scheduling systems and provide a new resource model for the frequency computation. As results, our use-case for avionics software components shows that our new method outperforms the classical real-time calculus (RTC) method, requiring 36.21% less frequency ratio on average for scheduling units under RM than the RTC method.
Jin Hyun Kim, Deepak Gangadharan, Oleg Sokolsky, Axel Legay, Insup Lee 0001
ISORC2
2016 Platform-Based Plug and Play of Automotive Safety Features: Challenges and Directions (Invited Paper)
abstract
Optional software-based features are increasingly becoming an important cost driver in automotive systems. These include features pertaining to active safety, infotainment, etc. Currently, these optional features are integrated into the vehicles at the factory during assembly. This severely restricts the flexibility of the customer to select and use features on-demand and therefore, the customer will either have to be satisfied with an available set of feature options or pre-order a car with the required features from the manufacturer resulting in considerable delay. In order to increase flexibility and reduce the delay, it is necessary to provide the option to configure the vehicle on-demand at the dealership or remotely. In this paper, we present our vision and challenges involved in developing a platform infrastructure that allows on-demand deployment of automotive safety features and ensures their correct execution.
Deepak Gangadharan, Jin Hyun Kim, Oleg Sokolsky, BaekGyu Kim, Chung-Wei Lin, Shinichi Shiraishi, Insup Lee 0001
RTCSA1
2014 Quality-aware video decoding on thermally-constrained MPSoC platforms
abstract
Current mobile devices extensively run video players that are power hungry. Further, higher power densities as a result of technology scaling results in higher on-chip temperatures. Unlike general purpose computer systems, mobile devices that run on batteries cannot afford to have expensive cooling mechanisms. Therefore, in order to satisfy thermal constraints while running power hungry applications, dynamic thermal management (DTM) techniques have been employed. For multimedia applications, the techniques primarily relied on dynamic voltage and frequency scaling (DVFS) and dynamic power management (DPM) while taking care that maximum video quality is achieved. However, no prior work has exploited frame drops to lower the inserted idle times under predetermined quality constraints. In this work, we propose a DPM framework that utilizes frame drops to dynamically insert low idle times in order to satisfy a peak temperature constraint under a given quality constraint. This also reduces the end-to-end latency. The latencies are further reduced by maintaining lightweight workload histories. For the videos used in our experiments, it was observed that a small reduction in quality of 2 dB (reduction from 32 dB to 30 dB) due to frame drops in motion videos results in a maximum latency reduction of 1.7 sec.
Deepak Gangadharan, Jürgen Teich, Samarjit Chakraborty
ASAP1
2014 Runtime Reconfigurable Bus Arbitration for Concurrent Applications on Heterogeneous MPSoC Architectures
abstract
This paper describes a runtime reconfigurable bus arbitration technique for concurrent applications on heterogeneous MPSoC architectures. Here, a hardware/software approach is introduced as part of a runtime framework that enables selecting and adapting different policies (i. e., fixed-priority, TDMA, and Round-Robin) such that the performance goals of concurrent applications can be satisfied. To evaluate the hardware cost, we compare our proposed solution with respect to a well-known SPARC V8 architecture supporting fixed-priority arbitration. Notably, even providing the flexibility for selecting up to three different policies, our reconfigurable arbiter needs only 25% and 7% more LUTs and slices registers, respectively. The reconfiguration overhead for changing between different policies is 56 cycles and for programming new time slots, only 28 cycles are necessary. For demonstrating the benefits of this reconfiguration framework, we setup a mixed hard/soft real-time scenario by considering four applications with different timeliness requirements. The experimental results show that by reconfiguring the arbiter, less processing elements can be used for achieving a specific target frame rate. Moreover, adjusting the time slots for TDMA, we can speedup a soft real-time algorithm while still satisfying the deadline for hard real-time applications.
Éricles Sousa, Deepak Gangadharan, Frank Hannig, Jürgen Teich
DSD2
2013 Quality-aware media scheduling on MPSoC platforms
abstract
Applications that stream multiple video/audio or video+audio clips are being implemented in embedded devices. A Picture-in-Picture (PiP) application is one such application scenario, where two videos are played simultaneously. Although the PiP application is very efficiently handled in televisions and personal computers by providing maximum quality of service to the multiple streams, it is a difficult task in devices with resource constraints. In order to efficiently utilize the resources, it is essential to derive the necessary processor cycles for multiple video streams such that they are displayed with some prespecified quality constraint. Therefore, we propose a network calculus based formal framework to help schedule multiple media streams in the presence of buffer contraints. Further, our framework also presents a schedulability analysis condition to check if the multimedia streams can be scheduled such that a prespecified quality constraint is satisfied with the available service. We present this framework in the context of a PiP application, but it is applicable in general for multiple media streams. The results obtained using the formal framework were further verified using experiments involving system simulation.
Deepak Gangadharan, Samarjit Chakraborty, Roger Zimmermann
DATE1
2013 Multi-ASIP platform synthesis for Event-Triggered applications with cost/performance trade-offs
abstract
In this paper, we propose a technique to synthesize a cost-efficient distributed platform consisting of multiple Application Specific Instruction Set Processors (multi-ASIPs) running applications with strict timing constraints. Multi-ASIP platform synthesis is a non-trivial task for two reasons. Firstly, we need to know the WCET of tasks in target applications to derive platforms (including synthesized ASIPs) in which the tasks are schedulable. However, the WCET of tasks can be known only after the ASIPs are synthesized. We break this circular dependency by using a probability distribution of the WCET of a task (further referred to as the WCET uncertainty model), which takes into account the underlying microarchitectural configurations for the ASIP implementation. Secondly, the datapath area of the multi-ASIPs synthesized is an important design factor that contributes significantly towards the overall cost of the platform. We propose an area estimation model and a WCET uncertainty model that consider the effect of task datapath similarity. Based on these two models, we support the designer in exploring cost/performance trade-offs during the platform synthesis. We propose an Evolutionary Algorithm-based approach to solve this multiobjective optimization problem. The proposed approach has been evaluated using several benchmarks and it provides a number of multi-ASIP platform solutions exploring the trade-offs in the cost/performance design space.
Deepak Gangadharan, Laura Micconi, Paul Pop, Jan Madsen
RTCSA1
2012 ASAM: Automatic Architecture Synthesis and Application Mapping
abstract
This paper focuses on mastering the automatic architecture synthesis and application mapping for heterogeneous massively-parallel MPSoCs based on customizable application-specific instruction-set processors (ASIPs). It presents an over-view of the research being currently performed in the scope of the European project ASAM of the ARTEMIS program. The paper briefly presents the results of our analysis of the main problems to be solved and challenges to be faced in the design of such heterogeneous MPSoCs. It explains which system, design, and electronic design automation (EDA) concepts seem to be adequate to resolve the problems and address the challenges. Finally, it introduces and briefly discusses the ASAM design-flow and its main stages.
Lech Józwiak, Menno Lindwer, Rosilde Corvino, Paolo Meloni, Laura Micconi, Jan Madsen, Erkan Diken, Deepak Gangadharan, Roel Jordans, Sebastiano Pomata, Paul Pop, Giuseppe Tuveri, Luigi Raffo
DSD8
2011 Fast hybrid simulation for accurate decoded video quality assessment on MPSoC platforms with resource constraints
abstract
Multimedia decoders mapped onto MPSoC platforms exhibit degraded video quality when the critical system resources such as buffer and processor frequency are constrained. Hence, it is essential for system designers to find the appropriate mix of resources, living within the constraints, for a desired output video quality. A naive approach to do this would be to run expensive system simulations of the decoder tasks mapped onto a model of the underlying MPSoC architecture. This turns out to be inefficient when the input video library set has a large number of video clips. We propose a fast hybrid simulation framework to quantitatively estimate decoded video quality in the context of an MPEG-2 decoder. Here, the workload of simulation heavy tasks are estimated using accurate analytical models. The workload of other light (but difficult to analytically model) tasks are obtained from system simulations. This framework enables the system designer to perform a fast trade-off analysis of the system resources in order to choose the optimal combination of resources for the desired video quality. When compared to a naive system simulation approach, the hybrid simulation-based framework shows speed-up factors of about 5× for motion and 8× for still videos. The results obtained using this framework highlight important trade-offs such as the decoded video quality (measured in terms of the peak signal to noise ratio (PSNR)) vs buffer size and PSNR vs processor frequency.
Deepak Gangadharan, Samarjit Chakraborty, Roger Zimmermann
ASP-DAC1
2011 Video quality-driven buffer dimensioning in MPSoC platforms via prioritized frame drops
abstract
We study the impact of a novel prioritized frame dropping scheme in buffer-constrained multiprocessor system-on-chip (MPSoC) platforms. Accurate buffer dimensioning has attracted lot of research interest as large on-chip buffers result in increased silicon area and higher costs. Multimedia applications present the flexibility of trading off quality for buffer space without any noticeable deterioration in video quality. The frame dropping scheme is crucial here to drop frames appropriately such that the required buffer size is reduced and target quality requirement is satisfied. Towards this, we propose a simple prioritized frame dropping mechanism which reduces the required buffer space more than existing frame dropping policies. We also provide a fast iterative procedure to find the minimum buffer size for a video clip with O(log(Ndrop)) number of iterations, where Ndropis the maximum number of frames that can be dropped for a video clip so that a prespecified quality in terms of peak signal to noise ratio (PSNR) value is satisfied.
Deepak Gangadharan, Haiyang Ma, Samarjit Chakraborty, Roger Zimmermann
ICCD1
2011 Energy-aware complexity adaptation for mobile video calls
abstract
High energy consumption has become a challenge for multimedia applications on mobile platforms. We propose a cross layer framework that integrates complexity adaptation and energy conservation for mobile video calls. First we select the most utility-aware encoding and decoding parameters for videos of different motion levels through extensive offline profiling and analysis. Then we design a feedback algorithm to adaptively apply different coding parameters while monitoring the system performance online during a video call. To minimize energy consumption, we utilize Dynamic Voltage and Frequency Scaling (DVFS) for the CPU and buffered transmissions for the network. Our experimental results show an effective saving on energy consumption, with on average 52% savings on the CPU and 30% on the wireless network, while still maintaining high quality service.
Haiyang Ma, Deepak Gangadharan, Nalini Venkatasubramanian, Roger Zimmermann
ACM Multimedia2
2011 Video Quality Driven Buffer Sizing via Frame Drops
abstract
We study the impact of video frame drops in buffer constrained multiprocessor system-on-chip (MPSoC) platforms. Since on-chip buffer memory occupies a significant amount of silicon area, accurate buffer sizing has attracted a lot of research interest lately. However, all previous work studied this problem with the underlying assumption that no video frame drops can be tolerated. In reality, multimedia applications can often tolerate some frame drops without significantly deteriorating their output quality. Although system simulations can be used to perform video quality driven buffer sizing, they are time consuming. In this paper, we first demonstrate a dual-buffer management scheme to drop only the less significant frames. Based on this scheme, we then propose a formal framework to evaluate the buffer size vs. video quality trade-offs, which in turn will help a system designer to perform quality driven buffer sizing. In particular, we mathematically characterize the maximum numbers of frame drops for various buffer sizes and evaluate how they affect the worst-case PSNR value of the decoded video. We evaluate our proposed framework with anMPEG-2 decoder and compare the obtained results with that of a cycle-accurate simulator. Our evaluations show that for an acceptable quality of 30 dB, it is possible to reduce the buffer size by up to 28.6% which amounts to 25.88 megabits.
Deepak Gangadharan, Linh T. X. Phan, Samarjit Chakraborty, Roger Zimmermann, Insup Lee 0001
RTCSA (1)1