VLDB 2026 Research / reviewers in the wild / expert
John Jose
dblp:123/7042
· DBLP profile ↗
40ranked-venue papers
4as first author
23since 2021 · last 2026
0000-0002-0314-8778ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 36 · 4 first-author · 20 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Topology and Reliability Aware Qubit Mapping for Quantum Core SystemsabstractScaling quantum processors motivates multicore (modular) quantum architectures, where inter-core teleportation and state transfer are substantially slower and less reliable than intra-core operations. A key compilation/runtime challenge in such systems is inter-core qubit mapping: assigning logical qubits to cores across circuit timeslices to minimize non-local transfers under core-capacity and network constraints. Existing approaches such as FGP-rOEE, QUBO-based mapping, and Hungarian Qubit Assignment (HQA) reduce inter-core interactions, but are often evaluated under simplified interconnect assumptions and do not explicitly optimize for topology and reliability-dependent transfer costs. In this work, we present TR-HQA, a topology and reliability aware extension of HQA that minimizes expected inter-core transfer cost over a weighted interconnect graph, where edge weights capture hop distance and link success probability. We also propose GTR-QA, a greedy approximation that reduces placement time while maintaining competitive transfer cost. Using a multicore mapping simulator built on a pytket front-end, we evaluate TR-HQA and GTR-QA across circuit families (structured benchmarks and random circuits) and multiple interconnect topologies. Our results show that TR-HQA consistently reduces expected inter-core communication cost compared to topology-agnostic baselines, while GTR-QA provides a favorable quality–runtime trade-off. Across standard quantum benchmarks (QFT, Cuccaro Adder, and Draper Adder), TR-HQA reduces inter-core communication cost by 2.15 × on average (1.70 × –2.67 ×) and improves runtime by 11.8 × on average (1.02 × –30 ×). Rajeswari Suance P. S, Akansh Khandelwal, Surankan De, Satyajit Das, John Jose, Maurizio Palesi |
CF | 5 |
| 2026 | AxLEA: Approximate ARX-based Lightweight Encryption Algorithm for Resource Constrained DevicesabstractOur daily lives are increasingly dependent on small, resource-limited devices that often handle sensitive information, emphasizing the need for robust security measures. However, many lightweight cryptographic algorithms compromise speed and efficiency to conserve resources. To overcome this challenge, we introduce the AxLEA cryptosystem, the first to apply approximation techniques to a reversible cryptographic system. AxLEA improves computational efficiency by mitigating carry propagation overhead through approximation. We design an invertible 32-bit approximate adder and integrate it into the round function and round key generation process. This integration reduces delay by 80% and decreases area and power consumption by 9% in rolled implementations, while unrolled implementations achieve overall 50% performance improvement. Also, the proposed design achieves a 5 \(\times\) reduction in energy consumption in both rolled and unrolled implementations, demonstrating improved efficiency over LEA across different architectural configurations. AxLEA meets critical security requirements, satisfying the avalanche effect and randomness tests specified by the NIST and ENT test suites. Additionally, it demonstrates strong resistance to linear and differential cryptanalysis. These results make AxLEA an efficient and secure solution for highly resource-constrained devices. Vivekananda G, Thejaswini P, Sukumar Nandi, John Jose |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2025 | CosMoS: Architectural Support for Cost-Effective Data Movement in a Disaggregated Memory SystemsabstractMemory disaggregation has emerged as a strong alternative to traditional server systems for improved memory utilization and scalability. The compute nodes with a small local memory are connected to disaggregated memory pools through memory-semantic interconnects, such as CXL, that support high-bandwidth and low-latency memory access. A primary concern with these systems is high remote memory latency due to the presence of a network interconnect between CPU and memory. A hot page migration is generally deployed in a hybrid memory system to improve the locality of memory access. In this article, we identify the major challenges in the present OS-based or hardware-based techniques for page migration in disaggregated memory. To this end, we propose CosMoS , an architectural solution for Cos t-effective data Mo vement in a S calable disaggregated memory system that can predict, schedule, and optimize the movement of hot pages between local and remote memory. Our mechanism also eliminates long delays on the critical demand memory accesses to other remote pages that are obstructed during hot page movement. We evaluate CosMoS over many data-centric workloads. Our results show a 20% performance improvement with CosMoS compared to the state-of-the-art mechanism and an 86% improvement compared to the baseline for a large-scale disaggregated memory system. Amit Puri, John Jose, Venkatesh Tamarapalli |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2025 | Securing Network-on-Chips against Trojan-Induced Packet Duplication AttacksabstractThe third-party Intellectual Property (IP) supply chain exposes System-on-Chip designs to malicious implants like Hardware Trojans (HTs). With extremely rare trigger conditions, some HTs can evade conventional and even machine learning-based validation methods. Current detection and mitigation approaches fall short, especially against HTs capable of creating detrimental effects on cache and Network-on-Chip (NoC) performance. In this article, we present a novel intermittent and robust HT called LOKI, which primarily operates within the Network Adapter (NA) but can simultaneously impact the performance of the NoC, the shared cache, and the cores. LOKI is implanted in a malicious IP’s NA and triggers packet duplication attacks, leading to increased latency and performance degradation across the system. This action has cascading effects on the system, including increased latency in the NoC, adversely impacting communication efficiency. The duplication also leads to increased cache misses and longer miss penalties, further degrading cache performance. Additionally, LOKI affects the Instruction-Per-Cycle (IPC) of the cores, thus influencing overall processing performance. Our evaluation demonstrates that LOKI causes a 3.53× increase in packet latency, a 15% increase in miss penalty, and a 10% decrease in overall IPC. To neutralize the effects of HT-induced packet duplication, we propose a ubiquitous mitigation framework called HULK. HULK is installed in the NA and monitors all messages going in and out of the NoC, allowing it to address anomalies occurring in the NA, routers, and links. Experimental evaluation shows that HULK can effectively mitigate LOKI’s impact, achieving baseline system-like performance with negligible hardware overhead. Unlike existing HT-specific mitigation proposals, HULK serves as a generic solution to neutralize all types of packet duplication attacks. To promote reproducibility and community adoption, we have open sourced the implementation at https://github.com/itsmanju/hulk . Manju Rajan, Abhijit Das 0002, John Jose |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2024 | DRackSim: Simulating CXL-enabled Large-Scale Disaggregated Memory SystemsabstractMemory disaggregation has emerged as an alternative to traditional server architecture in data centers to target better memory utilization and higher scalability. It involves multiple independent compute nodes and remote memory pools that get hardware support through high-speed cache-coherent interconnects such as CXL. This paper introduces DRackSim, a simulation infrastructure for scalable disaggregated memory systems. DRackSim primarily models multiple compute nodes, memory pools, local/global memory managers, and a network interconnect for coherent memory access. An application-level simulation approach simulates an out-of-order x86 multi-core processor and a multi-level cache hierarchy at compute nodes. The network interface is simulated through a queue-based approach to handle remote memory access at multiple granularity. It also models a global memory manager for remote address space at the memory pools. Finally, we integrate a modified DRAMSim2 to perform local/remote memory simulation by declaring multiple instances of DRAMSim2. We rigorously validate DRackSim subsystems against Gem5 and a hardware prototype. Finally, we explore the design space by modeling various use-case scenarios for disaggregated memory systems and evaluate their performance over various HPC and cloud benchmarks. Amit Puri, Kartheek Bellamkonda, Kailash Narreddy, John Jose, Venkatesh Tamarapalli, Narayanan Vijaykrishnan |
SIGSIM-PADS | 4 |
| 2024 | TROP: TRust-aware OPportunistic Routing in NoC with Hardware TrojansabstractMultiple software and hardware intellectual property (IP) components are combined on a single chip to form Multi-Processor Systems-on-Chips (MPSoCs). Due to the rigid time-to-market constraints, some of the IPs are from outsourced third parties. Due to the supply-chain management of IP blocks being handled by unreliable third-party vendors, security has grown as a crucial design concern in the MPSoC. These IPs may get exposed to certain unwanted practises like the insertion of malicious circuits called Hardware Trojan (HT) leading to security threats and attacks, including sensitive data leakage or integrity violations. A Network-on-Chip (NoC) connects various units of an MPSoC. Since it serves as the interface between various units in an MPSoC, it has complete access to all the data flowing through the system. This makes NoC security a paramount design issue. Our research focuses on a threat model where the NoC is infiltrated by multiple HTs that can corrupt packets. Data integrity verified at the destination’s network interface (NI) triggers re-transmissions of packets if the verification results in an error. In this article, we propose an opportunistic trust-aware routing strategy that efficiently avoids HT while ensuring that the packets arrive at their destination unaltered. Experimental results demonstrate the successful movement of packets through opportunistically selected neighbours along a trust-aware path free from the HT effect. We also observe a significant reduction in the rate of packet re-transmissions and latency at the expense of incurring minimum area and power overhead. Syam Sankar, Ruchika Gupta, John Jose, Sukumar Nandi |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2024 | Enhancing Lifetime and Performance of MLC NVM Caches Using Embedded Trace BuffersabstractLarge volumes of on-chip and off-chip memory are required by contemporary applications. Emerging non-volatile memory technologies including STT-RAM, PCM, and ReRAM are becoming popular for on-chip and off-chip memories as a result of their desirable properties. Compared to traditional memory technologies such as SRAM and DRAM, they have minimal leakage current and high packing density. Non Volatile Memories (NVM), however, have a low write endurance, a high write latency, and high write energy. Non-volatile Single Level Cell (SLC) memories can store a single bit of data in each memory cell, whereas Multi Level Cells (MLC) can store two or more bits in each memory cell. Although MLC NVMs have substantially higher packing density than SLCs, their lifetime and access speed are key concerns. For a given cache size, MLC caches consume 1.84× less space and 2.62× less leakage power than SLC caches. We propose Trace buffer Assisted Non-volatile Memory Cache (TANC), an approach that increases the lifespan and performance of MLC-based last-level caches using the underutilized Embedded Trace Buffers (ETB). TANC improves the lifetime of MLC LLCs up to 4.36× and decreases average memory access time by 4% compared to SLC NVM LLCs and by 6.41× and 11%, respectively, compared to baseline MLC LLCs. John Jose, Narayanan Vijaykrishnan |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2023 | A Practical Approach For Workload-Aware Data Movement in Disaggregated Memory SystemsabstractMemory disaggregation is a solid alternative to traditional server systems that can overcome memory scalability issues in next-generation HPC data centers. In a rack-level disaggregated system, multiple compute nodes with small local memory rely on remote memory pools (memory nodes) to fulfill their memory demands. An in-network memory manager manages remote memory address space and allocates it to compute nodes which can access the memory at cache-line granularity using coherent interconnects such as CXL (or GenZ). However, the memory access cost is significantly increased due to the presence of the network. Even though a page migration system can exploit the locality of memory accesses, accessing a remote page starves the block-level requests. Further, page migrations introduce additional overheads which combined with starvation may even degrade the performance. All these issues require systematic evaluation of disaggregated memory systems to achieve improved designs. This paper presents a hardware mechanism for workload-aware data movement between compute and memory pools that significantly reduces the memory access cost. Firstly, our design enables centralized hot-page migration in a multi-tiered disaggregated memory that is aware of access patterns for individual compute nodes. Secondly, we analyze the complexities of accessing a remote memory page and propose a novel solution to eliminate starvation by serving all the remote memory requests at cache block granularity and by sharing bandwidth between page and block memory requests. Lastly, we add extra hardware support to get rid of additional overheads in a page migration system. We evaluate our designs over a variety of multi-threaded benchmarks using a cycle-level simulator which is specially designed to simulate a disaggregated memory system. Our design performs 10% to 100% better than traditional RDMA-based disaggregated systems that access remote memory at page granularity and 5% to 35% better than baseline disaggregated systems that use coherent interconnects for block-level access. Amit Puri, Kartheek Bellamkonda, Kailash Narreddy, John Jose, Venkatesh Tamarapalli |
SBAC-PAD | 4 |
| 2023 | Strengthening NoC Security: Leveraging Hybrid Encryption for Data Packet ProtectionabstractThe increasing complexity and scale of modern computing systems have led to the emergence of System-on-Chip (SoC) architectures, with Network-on-Chip (NoC) architectures serving as a communication infrastructure within these systems. Due to rigid time-to-market constraints, recent SoC designs involve using third-party IPs. This outsourcing can lead to security vulnerabilities in SoC, especially in N oC packet transmission. Considering security attacks, this paper addresses the need for cost-effective secure packet transmission in N oC. We propose a hybrid encryption framework by discussing the challenges associated with symmetric and asymmetric encryption techniques to achieve a balance between security and efficient data transmission in NoC. By combining the strengths of both encryption methods, our approach aims to provide enhanced protection against security threats while minimizing performance overhead. Thejaswini P, Sahana A. R., Shankar Singh C., John Jose |
TENCON | 4 |
| 2023 | Secure Routing Framework for Mitigating Time-Delay Trojan Attack in System-on-ChipabstractIn order to meet the complex requirements of the semiconductor market, the packet-based on-chip interconnect IP; Network-on-Chip (NoC) used in System-on-Chips (SoCs), provides provisions for in-house designers to customise the NoC design. This opens a backdoor for the adversary to insert malicious circuits to deploy attacks like exposing cryptography keys , resource depletion attacks, etc. A malicious implant, such as a Hardware Trojan (HT) on NoC that initiates a delay-of-service attack, can tamper with the system and the application performance. In this work, we model an HT that mounts a time-delay attack in a NoC by violating the path selection strategy used by the route compute unit of the adaptive NoC router. Our experimental analysis shows that the proposed HT increases the packet latency by 14.5% and degrades the system performance (IPC) by 15% over the Baseline. For HT detection, we propose a framework that uses packet traffic analysis and path monitoring to localise the HT. We also propose a security wrapper module for the route compute unit that suppresses the effect of HT with an average IPC reduction of only 1.8% over the Baseline. Manju Rajan, Mayank Choksey, John Jose |
J. Syst. Archit. | 3 |
| 2023 | ZPP: A Dynamic Technique to Eliminate Cache Pollution in NoC based MPSoCsabstractData prefetching efficiently reduces the memory access latency in NUCA architectures as the Last Level Cache (LLC) is shared and distributed across multiple cores. But cache pollution generated by prefetcher reduces its efficiency by causing contention for shared resources such as LLC and the underlying network. The paper proposes Zero Pollution Prefetcher (ZPP) that eliminates cache pollution for NUCA architecture. For this purpose, ZPP uses L1 prefetcher and places the prefetched blocks in the data locations of LLC where modified blocks are stored. Since modified blocks in LLC are stale and request for such blocks are served from the exclusively owned private cache, their space unnecessary consumes power to maintain such stale data in the cache. The benefits of ZPP are (a) Eliminates cache pollution in L1 and LLC by storing prefetched blocks in LLC locations where stale blocks are stored. (b) Insufficient cache space is solved by placing prefetched blocks in LLC as LLCs are larger in size than L1 cache. This helps in prefetching more cache blocks, thereby increasing prefetch aggressiveness. (c) Increasing prefetch aggressiveness increases its coverage. (d) It also maintains an equivalent lookup latency to L1 cache for prefetched blocks. Experimentally it has been found that ZPP increases weighted speedup by 2.19x as compared to a system with no prefetching while prefetch coverage and prefetch accuracy increases by 50%, and 12%, respectively compared to the baseline. 1 Dipika Deb, John Jose |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2023 | Self Adaptive Logical Split Cache Techniques for Delayed Aging of NVM LLCabstractDue to the technological advancements in the last few decades, several applications have emerged that demand more computing power and on-chip and off-chip memories. However, the scaling of memory technologies is not at par with computing throughput of modern day multi-core processors. Conventional memory technologies such as SRAM and DRAM have technological limitations to meet large on-chip memory requirements owing to their low packaging density and high leakage power. In order to meet the ever-increasing demand for memory, researchers came up with alternative solutions, such as emerging non-volatile memory technologies such as STT-RAM, PCM, and ReRAM. However, these memory technologies have limited write endurance and high write energy. This emphasizes the need for a policy that will reduce the writes or distribute the writes uniformly across the memory thereby enhancing its lifetime by delaying the early wear out of memory cells due to frequent writes. We propose two techniques, Enhanced-Virtually Split Cache (E-ViSC) and Protean-Virtually Split Cache (P-ViSC), which dynamically adjust the cache configuration to distribute the writes uniformly across the memory to enhance the lifetime. Experimental studies show that E-ViSC and P-ViSC improve lifetime of NVM L2 caches by upto 2.31× and 1.97× respectively. John Jose |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2022 | Securing On-chip Interconnect against Delay Trojan using Dynamic Adaptive CagingabstractWith the progressive innovation of VLSI technology, Tiled Chip Multicore Processors (TCMP) have surfaced up as the backbone of the modern data intensive parallel multi-core systems. Network-on-Chip (NoC) is considered as the most preferred choice for on-chip communication. Manufacturers have begun to investigate the prospects of using third-party IP in sophisticated TCMP designs due to strict time-to-market limitations. The inflated reliance over third party IPs induced security vulnerabilities in inter-tile communication. In this paper, we implement a novel Hardware Trojan (HT) called as Delay Trojan (DT) placed in an NoC router. Proposed DT adds random delay to flits going through it, while other NoC routers merely experience regular congestion, making DT detection difficult. As a result, packets of latency-critical applications stalls impacting system performance and throughput. Further, we propose a dynamic adaptive learning framework embedded in NoC routers that detects DT with reasonable accuracy and alerts neighboring routers. We also propose a caging technique to re-route packets. Our experimental study evaluates the impact of DT and the effectiveness of the proposed solution. Ruchika Gupta, Vedika J. Kulkarni, John Jose, Sukumar Nandi |
ACM Great Lakes Symposium on VLSI | 3 |
| 2022 | DAReS: Deflection Aware Rerouting between Subnetworks in Bufferless On-Chip NetworksabstractNetwork on Chip (NoC) is an effective intercommunication structure used in the design of efficient Tiled Chip Multi Processor (TCMP) systems as they improve system performance manifold. Bufferless NoC has emerged as a popular design choice to address area and energy concerns associated with buffered NoC systems. For low to medium injection rate applications, both bufferless and buffered routers show similar network performance. As the network load rises, network performance of bufferless router based designs deteriorate due to increased deflections. This paper proposes a subnetwork based bufferless design, DAReS, to minimize deflections by redirecting contending flit in one subnetwork to unoccupied productive ports of other subnetwork without incurring any extra cycle delay. From evaluations, we observe that our proposed design approach improves network performance by minimizing deflection rate, power dissipation and shows better throughput in comparison to state-of-the-art bufferless router. Rose George Kunthara, Rekha K. James, Simi Zerine Sleeba, John Jose |
ACM Great Lakes Symposium on VLSI | 4 |
| 2022 | Hardware Trojan Mitigation for Securing On-chip Networks from Dead Flit AttacksabstractWith the advancements in VLSI technology, Tiled Chip Multicore Processors (TCMP) with packet switched Network-on-Chip (NoC) have emerged as the backbone of the modern data intensive parallel multi-core systems. Tight time-to-market and cost constraints have forced chip manufacturers to use third-party IPs in sophisticated TCMP designs. This dependence over third party IPs has instigated security vulnerabilities in inter-tile communication that cannot be detected at manufacturing and testing phases. This includes possibility of having malicious circuits like Hardware Trojans (HT). NoC is the likely target of HT insertion due to its significance and positional advantage from system and communication standpoints. Recent research shows that HTs can manipulate control fields of NoC packets and leads to dead flit attacks that has the potential to disrupt the on-chip communication resulting in application level stalling. In this paper, we propose run time detection of such dead flit attacks by analyzing packet movement behaviours. We also propose a cost effective mitigation mechanism by re-routing the packets around the HT infected router. Our experimental study with real benchmarks on 8x8 mesh TCMP evaluates the effectiveness of the proposed solution. Mohammad Humam Khan, Ruchika Gupta, Vedika J. Kulkarni, John Jose, Sukumar Nandi |
VLSI-SoC | 4 |
| 2022 | ENDURA : Enhancing Durability of Multi Level Cell STT-RAM based Non Volatile Memory Last Level CachesabstractWith high packing density and low leakage power, Spin Transfer Torque Random Access Memories (STT-RAM) are a promising alternative to replace traditional memory technologies such as SRAM and DRAM. Applications will continue to demand more memory for processing in the coming decades. To achieve higher cell density, Multi-Level Cell STT-RAM (MLC STT-RAM) that can store two or more bits in a single memory cell is preferred over Single Level Cell STT-RAM (SLC STT-RAM). But their multistep read and write operations lead to significant read and write latency. The multistep write operations are also affecting the durability of MLC STT-RAM. Specialised wear levelling techniques are not available for MLC STT-RAM. We propose ENDURA, a technique that could improve the lifetime and latency of MLC STT-RAMs. ENDURA extends the lifetime of 2MB and 4MB MLC STT-RAM L2 caches by 2.05x and 2.59x on a single-core system and reduces write latency by 9.6% and 8.06% respectively with minimal overhead. Yogesh Kumar 0002, John Jose |
VLSI-SoC | 3 |
| 2022 | RIBiT: Reduced Intra-flit Bit Transitions for Bufferless NoCabstractIn modern Tiled Chip Multicore Processor (TCMP) systems, Network on Chip (NoC) is the preferred interconnect solution to overcome scalability and performance bottleneck issues that conventional bus-based architectures face. For low to medium NoC traffic, the energy and area efficient bufferless router is a better design choice compared to buffered structures. Dynamic power contributes to the majority of total power dissipation during data transmission whereas only a fraction of it is due to leakage power. Self-switching and cross-coupling activities across NoC links are responsible for total dynamic power, of which latter is the prime contributor. In any NoC system, data encoding techniques are generally employed at Network Interface (NI) level to minimize power dissipation across NoC links. We propose a data encoding mechanism for bufferless NoCs to minimize bit transitions within the flit which will result in reduced dynamic link power. Our suggested approach leverages a modified version of Delta encoding technique where the flit is encoded into data differences by a configurable module placed inside NI of each core. No additional control lines and hence no changes to the network are required for our proposed encoding scheme. Experimental analysis done using Xilinx Vivado shows that our proposed design approach has significant reduction in intra-flit bit transitions in comparison to the baseline designs. Akshay Sarman, Alwin Shaju, Rose George Kunthara, K. Neethu, Rekha K. James, John Jose |
VLSI-SoC | 6 |
| 2022 | FlitZip: Effective Packet Compression for NoC in MultiProcessor System-on-ChipabstractApplications running on Network on Chip (NoC) based multicore systems demand increased on-chip network bandwidth that can cater to the need for intensive communication among the cores and caches. Due to strict area and power budget, the bandwidth offered by NoC is very limited. Data-intensive and communication-centric applications on encountering a cache miss lead to a considerable burden on the underlying network for transferring blocks from multiple cache hierarchies to the requesting core as packets. This increases the packet transmission latency, thereby slowing down the system performance. Also, NoC being the highest component of power consumption after the cores, an increase in packets increases the dynamic power consumption of NoC. The article proposes FlitZip that addresses the problem by reducing on-chip traffic through compressing network packets. Hence, the compressed packet requires less bandwidth during its transfer, reducing the network's power consumption. Experimental analysis shows that FlitZip achieves a better compression ratio of 52 percent, reduces packet latency and bandwidth utilization by 19.28 and 27 percent, respectively. It also reduces the area and power consumption of the de/compression units by 53.33 and 62.3 percent, respectively, compared to the state-of-the-art packet compression technique, NoΔ. Dipika Deb, Rohith M. K., John Jose |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2021 | Improving Lifetime of Non-Volatile Memory Caches by Logical PartitioningabstractWe are in an era of highly data-intensive applications, and the existing memory technologies are inadequate to meet their challenges. Non-Volatile Memories (NVMs) have emerged as a cost-effective alternative to the conventional SRAM based Last Level Caches (LLC) and DRAM-based main memories; however, they suffer from limited write endurance. Applications having non-uniform writes will cause heavily written blocks to fail faster than lightly written blocks, thereby reducing the lifetime of NVMs. Most of the modern processors use split organization in the first level cache and unified organization in the subsequent cache levels. Our proposed approach, ViSC (Virtually Split Cache) explores the write variation across the data and instruction blocks by virtually splitting unified LLC for wear-leveling. The logical mapping of LLC ways into instruction and data is interchanged periodically to distribute the writes uniformly. Our experimental results show that ViSC reduces the write variations significantly and improves the lifetime of NVMs by 1.94, 2.06, and 1.72 times for unicore, dual-core, and quad-core, respectively, by incurring negligible power and area overheads. Abdul Khader Thalakkattu Moosa, John Jose |
ACM Great Lakes Symposium on VLSI | 3 |
| 2021 | Packet header attack by hardware trojan in NoC based TCMP and its impact analysisabstractWith the advancement of VLSI technology, Tiled Chip Multicore Processors (TCMP) with packet switched Network-on-Chip (NoC) have been emerged as the backbone of the modern data intensive parallel systems. Due to tight time-to-market constraints, manufacturers are exploring the possibility of integrating several third-party Intellectual Property (IP) cores in their TCMP designs. Presence of malicious Hardware Trojan (HT) in the NoC routers can adversely affect communication between tiles leading to degradation of overall system performance. In this paper, we model an HT mounted on the input buffers of NoC routers that can alter the destination address field of selected NoC packets. We study the impact of such HTs and analyse its first and second order impacts at the core level, cache level, and NoC level both quantitatively and qualitatively. Our experimental study shows that the proposed HT can bring application to a complete halt by stalling instruction issue and can significantly impact the miss penalty of L1 caches. The impact of re-transmission techniques in the context of HT impacted packets getting discarded is also studied. We also expose the unrealistic assumptions and unacceptable latency overheads of existing mitigation techniques for packet header attacks and emphasise the need for alternative cost effective HT management techniques for the same. Vedika J. Kulkarni, Manju Rajan, Ruchika Gupta, John Jose, Sukumar Nandi |
NOCS | 4 |
| 2021 | Opportunistic Caching in NoC: Exploring Ways to Reduce Miss PenaltyabstractDue to limited on-chip caching, data-driven applications with large memory footprint encounter frequent cache misses. Such applications suffer from recurring miss penalty when they re-reference recently evicted cache blocks. To meet the worst-case performance requirements, Network-on-Chip (NoC) routers are provisioned with input port buffers. However, recent studies reveal that these buffers remain underutilised except during network congestion. Trace buffers are Design-for-Debug (DfD) hardware employed in NoC routers for post-silicon debug and validation. Nevertheless, they become non-functional once a design goes into production and remain in the routers left unused. In this article, we exploit the underutilised NoC router buffers and the unused trace buffers to store recently evicted cache blocks. While these blocks are stored in the buffers, future re-reference to these blocks can be replied from the NoC router. Such an opportunistic caching of evicted blocks in NoC routers significantly reduce the miss penalty. Experimental analysis shows that the proposed architectures can achieve up to 21 percent (16 percent on average) reduction in miss penalty and 19 percent (14 percent on average) improvement in overall system performance. While we have a negligible area and leakage power overhead of 2.58 and 3.94 percent, respectively, dynamic power reduces by 6.12 percent due to the improvement in performance. Abhijit Das 0002, John Jose, Maurizio Palesi |
IEEE Trans. Computers | 3 |
| 2021 | COPE: Reducing Cache Pollution and Network Contention by Inter-tile Coordinated Prefetching in NoC-based MPSoCsabstractPrefetching helps in reducing the memory access latency in multi-banked NUCA architecture, where the Last Level Cache (LLC) is shared. In such systems, an application running on core generates significant traffic on the shared resources, the underlying network and LLC. While prefetching helps to increase application performance, but an inaccurate prefetcher can cause harm by generating unwanted traffic that additionally increases network and LLC contention. Increased network contention results in untimely prefetching of cache blocks, thereby reducing the effectiveness of a prefetcher. Prefetch accuracy is extensively used to reduce unwanted prefetches that can mitigate the prefetcher caused contention. However, the conventional prefetch accuracy parameter has major limitations in NUCA architectures. The article exposes that prefetch accuracy can create two major false-positive cases of prefetching, Under-estimation and Over-estimation problems, and false feedback loop that can mislead a prefetcher in generating more unwanted traffic. We propose a novel technique, Coordinated Prefetching for Efficient (COPE), which addresses these issues by redefining prefetch accuracy for such architectures and identifies additional parameters that can avoid generating unwanted prefetch requests. Experiment conducted using PARSEC benchmark on a 64-core system shows that COPE achieve 3% reduction in L1 cache miss rate, 12.64% improvement in IPC, 23.2% reduction in average packet latency and 18.56% reduction in dynamic power consumption of the underlying network. Dipika Deb, John Jose, Maurizio Palesi |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2021 | Data Criticality in Multithreaded Applications: An Insight for Many-Core SystemsabstractMultithreaded applications are capable of exploiting the full potential of many-core systems. However, network-on-chip (NoC)-based intercore communication in many-core systems is responsible for 60%-75% of the miss latency experienced by multithreaded applications. Delay in the arrival of critical data at the requesting core severely hampers performance. This brief presents some interesting insights about how critical data are requested from the memory by multithreaded applications. Then it investigates the cause of delay in NoC and how it affects the performance. Finally, this brief shows how NoC-aware memory access optimizations can significantly improve performance. Our experimental evaluation considers Early Restart memory access optimization and demonstrates that by exploiting available NoC resources, critical data can be prioritized to reduce miss penalty by 11% and improve overall system performance by 9%. Abhijit Das 0002, John Jose, Prabhat Mishra 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2020 | Router Buffer Caching for Managing Shared Cache Blocks in Tiled Multi-Core ProcessorsabstractMultiple cores in a tiled multi-core processor are connected using a network-on-chip mechanism. All these cores share the last-level cache (LLC). For large-sized LLCs, generally, non-uniform cache architecture design is considered, where the LLC is split into multiple slices. Accessing highly shared cache blocks from an LLC slice by several cores simultaneously results in congestion at the LLC, which in turn increases the access latency. To deal with this issue, we propose a congestion management technique in the LLC that equips the NoC router with small storage to keep a copy of heavily shared cache blocks. To identify highly shared cache blocks, we also propose a prediction classifier in the LLC controller. We implement our technique in Sniper, an architectural simulator for multi-core systems, and evaluate its effectiveness by running a set of parallel benchmarks. Our experimental results show that the proposed technique is effective in reducing the LLC access time. Joe Augustine, Kanakagiri Raghavendra, John Jose, Madhu Mutyam |
ICCD | 3 |
| 2020 | Reducing Off-Chip Miss Penalty by Exploiting Underutilised On-Chip Router BuffersabstractThe era of data driven applications expose limited on-chip caching in modern Tiled Chip Multi-Processors (TCMPs). Some applications suffer from frequent last level cache (LLC) miss and travel off-chip to fetch data and instructions. Off-chip miss penalty is very expensive as it severely hampers application execution time. Modern Network-on-Chip (NoC) based TCMPs employ input buffered routers for scalable communication bandwidth. In this work, we exploit underutilised buffers of NoC routers to store recently evicted LLC blocks. While these blocks are locally stored, future data requests for such blocks are directly replied from the NoC router. Local reply from routers avoid off-chip travel and significantly reduces LLC miss penalty. To make sure that such storage of evicted LLC blocks does not create NoC congestion, we incorporate block forwarding and dropping using dynamic router buffer contention updates. We experimentally validate that our proposed optimisations significantly reduces LLC miss penalty and improves overall system performance. We achieve a maximum system speedup of up to 13% and an average system speedup of 7%. Abhijit Das 0002, John Jose |
ICCD | 3 |
| 2020 | SECTAR: Secure NoC using Trojan Aware RoutingabstractSystem-on-Chips (SoCs) are designed using different Intellectual Property (IP) blocks from multiple third-party vendors to reduce design cost while meeting aggressive time-to-market constraints. Designing trustworthy SoCs need to address the increasing concerns related to supply-chain security vulnerabilities. Malicious implants on IPs, such as Hardware Trojans (HTs) are one of the significant security threats in designing trustworthy SoCs. It is a major challenge to detect Trojans in complex multi-processor SoCs using conventional pre- and post-silicon validation methodologies. Packet-based Network-on-Chip (NoC) is a widely used solution for on-chip communication between IPs in complex SoCs. The focus of this paper is to enable trusted NoC communication in the presence of potentially untrusted IPs. This paper makes three key contributions. (1) We model an HT in NoC router that activates misrouting of the packets to initiate a denial of service, delay of service, and injection suppression. (2) We propose a dynamic shielding technique that isolates the identified HT infected IP. (3) We present a secure routing algorithm to bypass the HT infected NoC router. Experimental results on HT infected NoC demonstrate that the proposed method reduces effective average packet latency by 38% in real benchmarks and 48% in synthetic traffic patterns. Our method also increases throughput and reduces effective average deflected packet latency by 62% in real benchmarks and 97% in synthetic traffic patterns. Manju Rajan, Abhijit Das 0002, John Jose, Prabhat Mishra 0001 |
NOCS | 3 |
| 2020 | Exploiting Data Resilience in Wireless Network-on-chip ArchitecturesabstractThe emerging wireless Network-on-Chip (WiNoC) architectures are a viable solution for addressing the scalability limitations of manycore architectures in which multi-hop long-range communications strongly impact both the performance and energy figures of the system. The energy consumption of wired links as well as that of radio communications account for a relevant fraction of the overall energy budget. In this article, we extend the approximate computing paradigm to the case of the on-chip communication system in manycore architectures. We present techniques, circuitries, and programming interfaces aimed at reducing the energy consumption of a WiNoC by exploiting the trade-off energy saving vs. application output degradation. The proposed platform—namely, xWiNoC—uses variable voltage swing links and tunable transmitting power wireless interfaces along with a programming interface that allows the programmer to specify those data structures that are error-resilient. Thus, communications induced by the access to such error-resilient data structures are carried out by using links and radio channels that are configured to work in a low energy mode, albeit by exposing a higher bit error rate. xWiNoC is assessed on a set of applications belonging to different domains in which the trade-off energy vs. performance vs. application result quality is discussed. We found that up to 50% of communication energy saving can be obtained with a negligible impact on the application output quality and 3% in application performance degradation. Giuseppe Ascia, Vincenzo Catania, Salvatore Monteleone, Maurizio Palesi, Davide Patti, John Jose, Valerio Mario Salerno |
ACM J. Emerg. Technol. Comput. Syst. | 6 |
| 2019 | Analyzing networks-on-chip based deep neural networksabstractOne of the most promising architectures for performing deep neural network inferences on resource-constrained embedded devices is based on massive parallel and specialized cores interconnected by means of a Network-on-Chip (NoC). In this paper, we extensively evaluate NoC-based deep neural network accelerators by exploring the design space spanned by several architectural parameters. We show how latency is mainly dominated by the on-chip communication whereas energy consumption is mainly accounted by memory (both on-chip and off-chip). Giuseppe Ascia, Vincenzo Catania, Salvatore Monteleone, Maurizio Palesi, Davide Patti, John Jose |
NOCS | 6 |
| 2019 | DoLaR: Double Layer Routing for Bufferless Mesh Network-on-ChipabstractNetwork on Chip (NoC) is embraced as an interconnect solution for the design of large tiled chip multiprocessors (TCMP). Bufferless NoC router is a promising approach due to its simple router design, energy and hardware efficiency. NoC, which rely on underlying network architecture, is characterized by performance measures like latency, deflection rate, throughput and power. In this paper, we come up with DoLaR architecture to raise performance of standard bufferless 2D mesh NoC by stacking two similar layers of 8×8 meshes one above the other. DoLaR employs standard 5-port bufferless router architecture and the unused ports of edge routers are utilized to make vertical interconnections between the layers. Simulation results show that our proposed design surpasses existing state-of-the-art 5-port 2D mesh and torus bufferless router designs in terms of better network saturation point and minimized deflection rate, average flit latency and power consumption. Rose George Kunthara, K. Neethu, Rekha K. James, Simi Zerine Sleeba, John Jose |
TENCON | 5 |
| 2019 | Cost effective routing techniques in 2D mesh NoC using on-chip transmission lines
Dipika Deb, John Jose, Shirshendu Das, Hemangee K. Kapoor |
J. Parallel Distributed Comput. | 2 |
| 2018 | Critical Packet Prioritisation by Slack-Aware Re-Routing in On-Chip NetworksabstractPacket based Network-on-Chip (NoC) connect tens to hundreds of components in a multi-core system. The routing and arbitration policies employed in traditional NoCs treat all application packets equally. However, some packets are critical as they stall application execution whereas others are not. We differentiate packets based on a metric called slack that captures a packet's criticality. We observe that majority of NoC packets generated by standard application based benchmarks do not have slack and hence are critical. Prioritising these critical packets during routing and arbitration will reduce application stall and improve performance. We study the diversity and interference of packets to propose a policy that prioritises critical packets in NoC. This paper presents a slack-aware re-routing (SAR) technique that prioritises lower slack packets over higher slack packets and explores alternate minimal path when two no-slack packets compete for same output port. Experimental evaluation on a 64-core Tiled Chip Multi-Processor (TCMP) with 8×8 2D mesh NoC using both multiprogrammed and multithreaded workloads show that our proposed policy reduces application stall time by upto 22% over traditional round-robin policy and 18% over state-of-the-art slack-aware policy. Abhijit Das 0002, Sarath Babu 0003, John Jose, Sangeetha Jose, Maurizio Palesi |
NOCS | 3 |
| 2018 | Traffic Aware Deflection Rerouting Mechanism for Mesh Network on ChipabstractIn two dimensional mesh Network on Chips (NoC), efficient routing algorithms route majority of the flits through the central routers of the network, whereas routers at the edges and corners experience relatively lesser flit flow. This in turn leads to higher traffic towards central routers than to edge and corner routers. Such uneven traffic distribution causes thermal hot-spots at the center of the chip where the load is high, and reduces the average life-time of the chip. In existing buffer-less deflection routing techniques, load balanced traffic distribution is not considered as a factor during assignment of links to mis-routed flits. Devising deflection routing techniques with greater load balancing capability is a major challenge for efficient thermal management of the chip. This paper proposes an adaptive routing mechanism that can provide a more balanced traffic profile in a deflection router based mesh NoC. Significant number of deflected flits are rerouted towards the edges/corners of the mesh, thereby reducing the load on the central routers. From evaluations, it is seen that the proposed technique reduces traffic variance compared to NoCs using baseline deflection routers. Transient temperature variation studies using Hotspot tool substantiate our findings. Simi Zerine Sleeba, John Jose, Maurizio Palesi, Rekha K. James, Maniyelil Govindankutty Mini |
VLSI-SoC | 2 |
| 2017 | Implementation and analysis of hotspot mitigation in mesh NoCs by cost-effective deflection routing techniqueabstractNetwork-on-Chip (NoC) serves as an efficient communication framework among the components of Chip MultiProcessors (CMPs). With increasing number of computation intensive applications communication between cores also increases, which creates high congestion resulting in network performance degradation. Handling congestion is a key network management issue in NoC. Hotspots are non-uniform traffic formation near cores where some cores need to handle a relatively higher traffic compared to others. Prolonged presence of these hotspots increases the communication latency of packets flowing through them. This work proposes a novel approach to identify destination hotspots and upon identification, packets are de-routed away from these hotspot cores using a cost-effective deflection routing technique. Experimental results show that in highly congested networks, our approach detects destination hotspots with great accuracy. De-routing of packets away from hotspots help them achieve congestion relief thereby decreasing the average latency of packets flowing in the network. R. S. Reshma Raj, Abhijit Das 0002, John Jose |
VLSI-SoC | 3 |
| 2015 | Dynamic migratory selection strategy for adaptive routing in mesh NoCsabstractNoC architectures are the most commonly used communication framework for multicore processors. Few factors that affect the performance of an on-chip network include the efficiency of the routing algorithm and the effectiveness of the output selection strategy used. All popular selection strategies use a static technique that behaves uniformly across various traffic patterns. This paper proposes a cost effective adaptive model of Regional Congestion Awareness (RCA) selection strategy. The traffic analyser incorporated in the proposed model learns the flit-flow pattern at each router and makes RCA behaves like the local best selection strategy under local traffic. Under non-local traffic, the normal RCA selection strategy works as it is. This switching (migration) between two selection strategies is done by proper controlling of aggregation and propagation mechanisms of RCA. As only in-router information is used for this switching, the design has no additional communication overhead. This dynamic switching decreases average packet latency and effectively optimises the network resources depending on traffic pattern, thereby, reducing power consumption. Our experiments on 8°8 mesh NoC with various synthetic and real traffic patterns show promising improvements compared to the existing baseline adaptive selection strategies. John Jose, Joe Augustine, Sijin Sebastian |
VLSI-SoC | 1 |
| 2014 | Minimally buffered single-cycle deflection routerabstractWith the drift from computation centric designs to communication centric designs in the Chip Multi Processor (CMP) era, the interconnect fabric is gaining more importance. An efficient NoC in terms of power, area and average flit latency has a huge impact on the overall performance of a CMP. In the current work, we propose MinBSD - a minimally buffered, single cycle, deflection router. It incorporates different operations (Injection, Ejection, Preemption, Re-injection) in a single module to handle the traffic effectively and ensures smooth flow of flits through router pipeline. It performs overlapped execution of independent operations. These factors not only make MinBSD to operate in a single cycle but also to reduce the critical path latency resulting in a faster interconnect network. Experimental results show that MinBSD reduces the average flit latency on real work loads, reduces die area and power consumption when compared to the existing state-of-the-art minimally buffered deflection routers. Gnaneswara Rao Jonna, John Jose, Rachana Radhakrishnan, Madhu Mutyam |
DATE | 2 |
| 2014 | WeDBless: weighted deflection bufferless router for mesh NoCsabstractBufferless NoC routers employing deflection routing are gaining popularity due to their power and area efficiency. We propose WeDBless, a bufferless deflection router that reduces deflection rate of flits by employing port allocation based on weighted deflection of flits. The proposed method directs the frequently misrouted flits towards their destination by increasing their probability of getting a productive output port. Our evaluations on synthetic traffic patterns show that WeDBless achieves significant reduction in deflection rate, average flit latency and improvement in network saturation point compared to the state-of-the-art bufferless router and reduced complexity in route computing logic. Simi Zerine Sleeba, John Jose, Maniyelil Govindankutty Mini |
ACM Great Lakes Symposium on VLSI | 2 |
| 2014 | Implementation and Analysis of History-Based Output Channel Selection Strategies for Adaptive Routers in Mesh NoCsabstractThe efficiency and effectiveness of an adaptive router in an NoC-based multicore system is evaluated by the performance it achieves under varying inter-core communication traffic. A well-designed selection strategy plays an important role in an adaptive router to act upon dynamic traffic variations. The effectiveness of a selection strategy depends on what metric is used to represent congestion, how precisely this metric captures the actual congestion, and how much cost is involved in capturing the congestion on a real-time scale. Congestion is formed over a period of time due to cumulative and chain reaction effects. We propose novel history-based selection strategies that could be used with any adaptive, deadlock-free, minimal routing in mesh NoCs. Buffer occupancy time and rate of flit flow across reachable ports of neighboring routers in the recent past are captured, propagated, and maintained in a cost-effective way to compute the selection metric. Experimental results on real and synthetic workloads show that our proposed selection strategies significantly outperform state-of-the-art techniques. John Jose, Madhu Mutyam |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2013 | DeBAR: deflection based adaptive router with minimal bufferingabstractEnergy efficiency of the underlying communication framework plays a major role in the performance of multicore systems. NoCs with buffer-less routing are gaining popularity due to simplicity in the router design, low power consumption, and load balancing capacity. With minimal number of buffers, deflection routers evenly distribute the traffic across links. In this paper, we propose an adaptive deflection router, DeBAR, that uses a minimal set of central buffers to accommodate a fraction of mis-routed flits. DeBAR incorporates a hybrid flit ejection mechanism that gives the effect of dual ejection with a single ejection port, an innovative adaptive routing algorithm, and a selective flit buffering based on flit marking. Our proposed router design reduces the average flit latency and the deflection rate, and improves the throughput with respect to the existing minimally buffered deflection routers without any change in the critical path. John Jose, Bhawna Nayak, Damarla Kranthi Kumar, Madhu Mutyam |
DATE | 1 |
| 2013 | SLIDER: Smart Late Injection DEflection Router for mesh NoCsabstractNetwork-on-Chip (NoC) provides a scalable communication interface for processing cores in large multicore systems. An efficient NoC router should not only minimize the average packet latency of the network but also have minimum pipeline latency, area, and power. Area and power overheads are affecting the scalability and popularity of traditional input buffered routers. In this context minimally buffered deflection routers are emerging as a cost effective alternative. We propose SLIDER, Smart Late Injection DEflection Router, that uses side buffers for accommodating a fraction of deflected flits. The main contributions of this work are smart late injection and selective flit preemption. In SLIDER the injection stage is kept at the end of the router pipeline. This reduces the contention in the arbitration stage, eliminates unwanted intra-router movement of flits and effectively utilizes the idle output channels. We parallelize independent operations in the router pipeline and reduce the pipeline latency by 25%. Experimental results on synthetic and real workloads show that SLIDER reduces average flit latency, channel wastage, and deflection rate, and increases throughput in the network when compared to the state-of-the-art minimally buffered deflection routers. Bhawna Nayak, John Jose, Madhu Mutyam |
ICCD | 2 |
| 2012 | TRACKER: A low overhead adaptive NoC router with load balancing selection strategyabstractThe effectiveness of an adaptive router in a Network on Chip (NoC) is evaluated by the selection metric it uses and its impact on overall performance. In this paper, we propose a flit flow history based load balancing selection strategy that can be used in any adaptive routers for output port selection. Using this selection strategy, we propose an adaptive router TRACKER, that keeps track of flow of flits through all its ports and updates this tracked information to its neighbors in a cost effective manner. Routers make use of these flit flow estimates to compute a novel selection metric for output port selection of incoming flits. TRACKER outperforms the baseline adaptive router architectures using odd-even routing model with conventional selection metrics like count of free virtual channels, count of fluid buffers and buffer occupancy time at reachable downstream neighbors. John Jose, K. V. Mahathi, J. Shiva Shankar, Madhu Mutyam |
ICCAD | 1 |