Sobhan Niknam

dblp:167/5771 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
3since 2021 · last 2025
0000-0001-6146-363XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Platform Performance Suite (PPS): A Framework for Performance Analysis & Diagnosis of Complex Cyber-physical Systems
abstract
The performance of cyber-physical systems (CPS) is a determining factor for their success and often needs to be guaranteed. When performance issues occur, their analysis and the identification of the root-cause should be fast. Typically, the analysis requires the association of system observations (i.e., tracing) to design and implementation artifacts. For this association, multidisciplinary domain knowledge is required, which typically resigns in the workforce but is unmanageable due to the immense and continuously increasing system complexity.This paper proposes a model-based method, the Platform Performance Suite (PPS), supported by corresponding tooling. PPS enables automated analysis for the identification of the root-cause when performance issues, like system throughput loss, occur. The cornerstone of the approach is the Lifecycle Software Architecture model which captures the domain knowledge in models by modeling all the relationships among the system artifacts from specification up to the runtime phase. The performance analysis at runtime phase is carried out on TMSC models guaranteeing the generic applicability of the method, while the connection to the other lifecycle phases enables the automated identification of the root-cause by pinpointing specific artifacts. The evaluation of the method has been carried out on a world-leading complex and high-performing CPS, the ASML TWINSCAN system, where a throughput loss needs to be addressed promptly. PPS has shown remarkable speed-up in the identification of the root-cause compared to the state-of-practice method.
Konstantinos Triantafyllidis, Yuri Blankenstein, Sobhan Niknam, Jos Hegge
ICPE3
2023 Thermal Management for S-NUCA Many-Cores via Synchronous Thread Rotations
abstract
On-chip thermal management is quintessential to a thermally safe operation of a many-core processor. The presence of a physically distributed logically shared Last-Level Cache (LLC) significantly reduces the performance penalty of migrating threads within the cores of an S-NUCA many-core. This cost reduction allows novel thermal management of these many-cores via synchronous thread migration. Synchronous thread migration provides a viable alternative to Dynamic Voltage and Frequency Scaling (DVFS) and asynchronous thread migration used traditionally to manage thermals of S-NUCA many-cores. We present a theoretical method to compute the peak tem-perature in many-cores with synchronous thread migrations. We use the method to create a thermal management heuristic called HotPotato that maximizes the performance of S-NUCA many-cores under a peak temperature constraint. We implement HotPotato within the state-of-the-art HotSniper simulator. Detailed interval thermal simulations with HotSniper show an average 10.72% improvement in response time of S-NUCA many-cores when scheduling with HotPotato compared to a state-of-the-art thermal-aware S-NUCA scheduler.
Yixian Shen, Sobhan Niknam, Anuj Pathania, Andy D. Pimentel
DATE2
2021 T-TSP: Transient-Temperature Based Safe Power Budgeting in Multi-/Many-Core Processors
abstract
Power budgeting techniques allow thermally safe operation in multi-/many-core processors while still allowing for efficient exploitation of available thermal headroom. Core-level power budgeting techniques like Thermal Safe Power (TSP) have allowed for more efficient operations than chip-level power budgeting techniques like Thermal Design Power (TDP) since the liner granularity permits operations closer to the threshold temperature without thermal violations.State-of-the-art TSP bases its power budgeting calculations on the long-term steady-state temperature of cores while ignoring trends in their short-term transient temperature. In this paper, we propose a new power budgeting technique called T-TSP (Transient-Temperature-based Safe Power) that bases its calculation on the current temperature of the core, a detail ignored by TSP. T-TSP provides a dynamic power budget to a core, which inversely correlates with the core’s thermal headroom. Dynamic power budgeting with T-TSP allows cores to reach the threshold temperature faster than TSP and operate safely close to it in perpetuity. Therefore, it provides the same thermal guarantees as TSP but enables even more efficient exploitation of thermal headroom.We integrate T-TSP with a state-of-the-art thermal interval simulation toolchain. Our detailed evaluations show that benchmarks execute faster by up to 17.94% and 8.37% on average when we do power budgeting with T-TSP instead of the state-of- the-art TSP. Finally, we make T-TSP publicly available in both its integrated and stand-alone forms.
Sobhan Niknam, Anuj Pathania, Andy D. Pimentel
ICCD1
2020 On the implementation and execution of adaptive streaming applications modeled as MADF
abstract
It has been shown that the mode-aware dataflow (MADF) is an advantageous analysis model for adaptive streaming applications. However, no attention has been paid on how to implement and execute an application, modeled and analyzed with the MADF model, on a Multi-Processor System-on-Chip, such that the properties of the analysis model are preserved. Therefore, in this paper, we consider this matter and propose a generic parallel implementation and execution approach for adaptive streaming applications modeled with MADF. Our approach can be easily realized on top of existing operating systems while supporting the utilization of a wider range of schedules. In particular, we demonstrate our approach on LITMUSRT as one of the existing real-time extensions of the Linux kernel. Finally, to show the practical applicability of our approach and its conformity to the analysis model, we present a case study using a real-life adaptive streaming application.
Sobhan Niknam, Peng Wang 0036, Todor P. Stefanov
SCOPES1
2019 Surf-Bless: A Confined-interference Routing for Energy-Efficient Communication in NoCs
abstract
In this paper, we address the problem of how to achieve energy-efficient confined-interference communication on a bufferless NoC taking advantage of the low power consumption of such NoC. We propose a novel routing approach called Surfing on a Bufferless NoC (Surf-Bless) where packets are assigned to domains and Surf-Bless guarantees that interference between packets is confined within a domain, i.e., there is no interference between packets assigned to different domains. By experiments, we show that our Surf-Bless routing approach is effective in supporting confined-interference communication and consumes much less energy than the related approaches.
Peng Wang 0036, Sobhan Niknam, Sheng Ma, Zhiying Wang 0003, Todor P. Stefanov
DAC2
2019 Hard Real-Time Scheduling of Streaming Applications Modeled as Cyclic CSDF Graphs
abstract
Recently, it has been shown that the classical hard real-time scheduling theory can be applied to streaming applications modeled as acyclic Cyclo-Static Dataflow (CSDF) graphs. However, many streaming applications are modeled as cyclic CSDF graphs, thus they are not supported by such scheduling theory. Therefore, in this paper, we propose an approach which enables to apply the classical hard real-time scheduling theory on streaming applications modeled as cyclic CSDF graphs. The proposed approach converts each task in a cyclic CSDF graph to a constrained-deadline periodic task. This conversion enables the utilization of many hard real-time scheduling algorithms which offer properties such as temporal isolation and fast calculation of the required number of processors for scheduling the tasks. We evaluate the performance of our approach in comparison to existing scheduling approaches. The evaluation, on a set of real-life benchmarks, demonstrates that our approach can schedule the tasks in an application, modeled as a cyclic CSDF graph, with guaranteed throughput equal or comparable to the throughput obtained by existing scheduling approaches while providing hard real-time guarantees for every task in the application thereby enabling temporal isolation among concurrently running tasks/applications on a multi-processor platform.
Sobhan Niknam, Peng Wang 0036, Todor P. Stefanov
DATE1
2019 EVC-Based Power Gating Approach to Achieve Low-Power and High Performance NoC
abstract
High power consumption becomes the major bottleneck that prevents applying Network-on-Chips (NoCs) on future many-core systems. Power gating is an effective way to reduce the power consumption of a NoC. However, conventional power gating approaches cause significant packet latency increase as well as additional power consumption overhead due to the power gating mechanism. One comprehensive way to reduce these negative impacts is to bypass powered-off routers in a NoC when transferring packets. Therefore, in this paper, we propose an express virtual channel based (EVC-based) power gating approach. In our approach, packets can take pre-defined virtual bypass paths to bypass intermediate routers that can be powered-on or powered-off. Furthermore, based on our extended router structure, a certain transmission ability of the powered-off routers is kept to transfer packets going through the normal paths. Thus, even though some packets do not take a virtual bypass path, they still have less probability to be blocked by the powered-off routers. Compared with a conventional NoC without power gating, our EVC-based power gating approach causes only 2.67% performance penalty, which is less than 28.67%, 7.24%, and 5.69% penalties in related approaches. With small hardware overhead, our approach reduces on average 68.29% of the total power consumption in a NoC, which is comparable with the 72.94%, 73.56%, and 75.3% reduction of the total power consumption in related approaches.
Peng Wang 0036, Sobhan Niknam, Sheng Ma, Zhiying Wang 0003, Todor P. Stefanov
DSD2
2019 Enabling Cognitive Autonomy on Small Drones by Efficient On-Board Embedded Computing: An ORB-SLAM2 Case Study
abstract
In this paper, we present a case study which investigates whether/how Simultaneous Localization and Mapping (SLAM), e.g., the ORB-SLAM2 application, can be executed on a small, energy-efficient, multi-processor embedded platform with an ARM big.LITTLE architecture, e.g., the ODROID-XU4 platform, mounted on a small drone with a limited energy budget while meeting real-time performance requirements. More specifically, we model and implement ORB-SLAM2 as a Kahn Process Network (KPN) which exploits pipeline parallelism and enables efficient mapping and execution of ORB-SLAM2 onto ODROID-XU4. Moreover, our KPN model enables the application of generic model transformations to exploit data-level parallelism as well. Then, we propose and implement, on top of the Linux operating system, an environment for efficient execution of applications modeled as KPNs. Finally, we perform a simple design space exploration (DSE) to investigate the trade-off between system performance and power consumption when alternative ORB-SLAM2 KPNs are executed on different configurations of the ODROID-XU4 platform. The obtained results of this DSE clearly show the feasibility of running ORB-SLAM2 on ODROID-XU4 in real time with a limited power budget for a given range of flying time, thereby enabling cognitive autonomy on small drones.
Erqian Tang, Sobhan Niknam, Todor P. Stefanov
DSD2
2018 Resource Optimization for Real-Time Streaming Applications Using Task Replication
abstract
In this paper, we study the problem of exploiting parallelism in a hard real-time streaming application modeled as an acyclic synchronous data flow (SDF) graph and scheduled on a heterogeneous multiprocessor system-on-chip platform to alleviate the capacity fragmentation due to partitioned scheduling algorithms and reduce the number of required processors when a throughput requirement is satisfied. As the main contribution in this paper, we propose a method to determine a replication factor for each task in an acyclic SDF graph such that by distributing the workloads among more parallel tasks with lower utilization in the obtained transformed graph, the left capacity on the processors can be efficiently exploited, hence reducing the number of required processors. The experimental results, on a set of real-life streaming applications, demonstrate that our approach can reduce the minimum number of processors required to schedule an application and considerably improve the memory requirements and application latency compared to related approaches while meeting the same throughput constraint.
Sobhan Niknam, Peng Wang 0036, Todor P. Stefanov
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2018 Modeling, Analysis, and Hard Real-Time Scheduling of Adaptive Streaming Applications
abstract
In real-time systems, the application's behavior has to be predictable at compile-time to guarantee timing constraints. However, modern streaming applications which exhibit adaptive behavior due to mode switching at run-time, may degrade system predictability due to unknown behavior of the application during mode transitions. Therefore, proper temporal analysis during mode transitions is imperative to preserve system predictability. To this end, in this paper, we initially introduce mode-aware data flow (MADF) which is our new predictable model of computation to efficiently capture the behavior of adaptive streaming applications. Then, as an important part of the operational semantics of MADF, we propose the maximum-overlap offset which is our novel protocol for mode transitions. The main advantage of this transition protocol is that, in contrast to self-timed transition protocols, it avoids timing interference between modes upon mode transitions. As a result, any mode transition can be analyzed independently from the mode transitions that occurred in the past. Based on this transition protocol, we propose a hard real-time analysis as well to guarantee timing constraints by avoiding processor overloading during mode transitions. Therefore, using this protocol, we can derive a lower bound and an upper bound on the earliest starting time of the tasks in the new mode during mode transitions in such a way that hard real-time constraints are respected.
Jiali Teddy Zhai, Sobhan Niknam, Todor P. Stefanov
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2017 A Novel Approach to Reduce Packet Latency Increase Caused by Power Gating in Network-on-Chip
abstract
The power gating technique is an effective way to reduce the high static power consumption in a Network-on-Chip (NoC). However, with notable wakeup delay, the power gating technique incurs significant packet latency increase. In this paper, we propose a novel Duty Buffer (DB) structure and an efficient DB-based power gating scheme to overcome this drawback. By keeping minimal number of DB active to replace any sleeping virtual channel in a router, our approach can efficiently reduce the packet latency increase along the whole routing path. Compared with a conventional five-stage pipeline router without power gating, our approach, with only one flit depth of the DB, increases the average packet latency by only 9.67%, which is much less than 57% and 21.75% latency increase in related approaches. With small hardware overhead, our approach can save on average 52.19% of the total power consumption in a NoC, which is comparable with 59.39% and 57.05% power savings in related approaches.
Peng Wang 0036, Sobhan Niknam, Zhiying Wang 0003, Todor P. Stefanov
NOCS2