Debendra Das Sharma

dblp:79/2400 · DBLP profile ↗
← Back
14ranked-venue papers
14as first author
7since 2021 · last 2026
0000-0003-1530-7788ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 14 first-author · 7 since 2021
YearPublicationVenuePosition
2026 Guest Editors' Introduction: Special Issue on CXL(®)
Debendra Das Sharma, Gustavo Alonso, Guangyu Sun 0003
IEEE Trans. Computers1
2025 On-Package Memory with Universal Chiplet Interconnect Express (UCIe): A Low Power, High Bandwidth, Low Latency and Low Cost Approach
abstract
Emerging computing applications such as Artificial Intelligence (AI) are facing a memory wall with existing on-package memory solutions that are unable to meet the power-efficient bandwidth demands. We propose to enhance UCIe with memory semantics to deliver power-efficient bandwidth and cost-effective on-package memory solutions applicable across the entire computing continuum. We propose approaches by reusing existing LPDDR6 and HBM memory through a logic die that connects to the SoC using UCIe. We also propose an approach where the DRAM die natively supports UCIe instead of the LPDDR6 bus interface. Our approaches result in significantly higher bandwidth density (up to 10x), lower latency (up to 3x), lower power (up to 3x), and lower cost compared to existing HBM4 and LPDDR on-package memory solutions.
Debendra Das Sharma, Swadesh Choudhary, Peter Z. Onufryk, Rob Pelt
HOTI1
2023 Invited: Compute Express Link™ (CXL™): An Open Interconnect for Cloud Infrastructure
abstract
Compute Express Link is an open industry standard interconnect offering caching and memory semantics on top of PCI-Express (PCIe)®, with resource pooling and fabric capabilities. This paper delves into the role of CXL to solve some of the challenges in the cloud infrastructure.
Debendra Das Sharma
DAC1
2023 Pipelined and Partitionable Forward Error Correction and Cyclic Redundancy Check Circuitry Implementation for PCI Express® 6.0
abstract
PCI Express®(PCIe®) specification has been doubling the data rate every generation in a backward compatible manner every three years. PCIe 6.0 specification will adopt PAM-4 signaling at 64.0 GT/s for maintaining the same channel reach, cost, and power profile as prior generations. A Forward Error Correction (FEC) mechanism will offset the high BER of PAM-4. A strong Cyclic Redundancy Check (CRC) and a link level replay mechanism will deliver a low-latency, high bandwidth efficiency, and highly reliable solution expected of a Load-Store interconnect. We propose a non-pipelined implementation of the FEC and CRC that is part of the PCIe 6.0 base specification. We also propose a partitionable and pipelined implementation for FEC and CRC for lowering gate count and latency. We have tested the correctness of our register transfer logic (RTL) implementation in a field programmable gate array (FPGA) implementation in addition to simulation. Synthesis results from the Synopsys DC compiler demonstrates that for a x16 PCIe Link partitionable to up to x4s, with 4 independent controllers using independently partitionable logic, we achieve a gate count of about 100,000 for the transmit and receive side with a FEC + CRC delay of less than 1 nano second (nsec) in each direction.
Debendra Das Sharma, Swadesh Choudhary
HOTI1
2022 Compute Express Link®: An open industry-standard interconnect enabling heterogeneous data-centric computing
abstract
Compute Express Link is an open industry standard interconnect offering caching and memory semantics on top of PCI Express®. In addition to providing high-bandwidth and low-latency connectivity between host processor and accelerators, smart network interface card, and memory expansion devices, it also enables resource pooling across multiple systems for scalable, power-efficient and cost-effective computing. This paper delves into the micro-architectural design to deliver power-efficient performance based on our experience designing a Xeon CPU and FPGA with this technology with demonstrated silicon interoperability.
Debendra Das Sharma
HOTI1
2021 Keynote 1: Compute Express Link (CXL) changing the game for Cloud Computing
abstract
Summary form only given, as follows. The complete presentation was not made available for publication as part of the conference proceedings. High-performance workloads demand heterogeneous processing, tiered memory architecture, persistent memory support, infrastructure accelerators such as smart NICs, and infrastructure processing units to meet the demands of the emerging compute landscape. Applications such as Artificial Intelligence, Machine Learning, Analytics, 5G, automotive, and high-performance computing are driving significant change in infrastructure such as Cloud, Edge, and client computing. Interconnect is a key pillar in this evolving computational landscape. The recent advent of Compute Express Link (CXL), a new open standard for cache-coherent interconnect, with its memory and coherency semantics has made it possible to pool computational and memory resources at the rack level using low-latency, higher-throughput, and memorycoherent access mechanisms. CXL is adopting networking features such as multi-host connectivity, pooled memory, persistence flows, and fabric manager while keeping its low-latency load-store semantics intact. The load-store I/O interconnects such as PCI Express (PCIe) and CXL are evolving to provide efficient access mechanisms across multiple nodes with advanced atomics, acceleration, smart NICs, persistent memory support etc. In this talk we will explore how synergistic evolution across load-store interconnects and fabrics can benefit the compute infrastructure of the future.
Debendra Das Sharma
HOTI1
2021 A low latency approach to delivering alternate protocols with coherency and memory semantics using PCI Express® 6.0 PHY at 64.0 GT/s
abstract
PCI Express®(PCIe®) specification has been doubling the data rate every generation in a backward compatible manner every two to three years, while maintaining the same latency, power, channel reach, cost, and reliability profile. It also provides architected support for alternate protocols such as those with coherency or memory semantics to run on PCIe PHY, in addition to the PCIe protocol. As a result, protocols such as Compute Express Link® (CXL®) as well as several CPU-CPU proprietary cache coherency links such as Ultra-Path Interconnect (UPI) run on PCIe PHY due to its low-latency and power-efficiency characteristics. PCIe 6.0 specification is adopting PAM-4 signaling to deliver 64.0 GT/s per Lane. A Flit-based approach with a light weight Forward Error Correction, strong CRC, and a low-latency link retry mechanism is used to meet the stringent low-latency, high bandwidth efficiency, scalability, and high reliability goals of PCIe applications. In this paper, we propose enhancements to the above mechanisms to have 0 latency impact while maintaining the bandwidth efficiency and high reliability for alternate protocols involving coherency and memory semantics where any latency adder may have an adverse impact on performance, while using the same PCIe 6.0 PHY. Our results demonstrate that these goals are achievable with PCIe 6.0 PHY while meeting the latency, bandwidth, reliability goals with multiple protocol support.
Debendra Das Sharma
HOTI1
2009 Intel® 5520 chipset: An I / O hub chipset for server, workstation, and high end desktop
Debendra Das Sharma
Hot Chips Symposium1
1998 Job Scheduling in Mesh Multicomputers
abstract
A new approach for dynamic job scheduling in mesh-connected multiprocessor systems, which supports a multiuser environment, is proposed in this paper. Our approach combines a submesh reservation policy with a priority-based scheduling policy to obtain high performance in terms of high throughput, high utilization, and low turn-around times for jobs. This high performance is achieved at the expense of scheduling jobs in a strictly fair, FCFS fashion; in fact, the algorithm is parameterized to allow trade-offs between performance and (short-term) POPS fairness. The proposed scheduler can be used with any submesh allocation policy. A fast and efficient implementation of the proposed scheduler has also been presented. The performance of the proposed scheme has been compared with the FCFS policy, the only existing scheduling strategy for meshes, to demonstrate the effectiveness of the proposed approach. Simulation results indicate that our scheduling strategy outperforms the FCFS policy significantly. Specifically, our strategy significantly reduces the average waiting delay of jobs over the FCFS policy. The fast implementation of the proposed scheduler results in low allocation and deallocation time overhead, as well as low space overhead.
Debendra Das Sharma, Dhiraj K. Pradhan
IEEE Trans. Parallel Distributed Syst.1
1996 Submesh Allocation in Mesh Multicomputers Using Busy-List: A BestFit Approach with Complete Recognition Capability
Debendra Das Sharma, Dhiraj K. Pradhan
J. Parallel Distributed Comput.1
1995 Processor Allocation in Hypercube Multicomputers: Fast and Efficient Strategies for Cubic and Noncubic Allocation
abstract
A new approach for dynamic processor allocation in hypercube multicomputers which supports a multi-user environment is proposed. A dynamic binary tree is used for processor allocation along with an array of free lists. Two algorithms are proposed based on this approach, capable of efficiently handling cubic as well as noncubic allocation. Time complexities for both allocation and deallocation are shown to be polynomial, a significant improvement over the existing exponential and even super-exponential algorithms. Unlike existing schemes, the proposed strategies are best-fit strategies within their search space. Simulation results indicate that the proposed strategies outperform the existing ones in terms of parameters such as average delay in honoring a request, average allocation time, average deallocation time, and memory overhead.>
Debendra Das Sharma, Dhiraj K. Pradhan
IEEE Trans. Parallel Distributed Syst.1
1994 Subcube Level Time-Sharing in Hypercube Multicomputers
abstract
A novel approach for subcube level time-sharing in hypercube multicomputers is proposed. Using this approach, multiple tasks may execute on the same processors. The tasks may be completely or partially overlapped with synchronous or asynchronous context switching. A dynamic binary tree is used for subcube allocation and deallocation. An incoming task is allocated to a subcube within which it will encounter minimum interference from other tasks. The allocation and deallocation time complexities are shown to be 0(n2). The proposed strategy has been implemented on an nCUBE 2. Measurement results indicate that the proposed strategy outperforms the FCFS-based batch-scheduling policy and an existing time-sharing policy by significantly reducing the average turn-around times.
Debendra Das Sharma, G. D. Holland, Dhiraj K. Pradhan
ICPP (2)1
1994 Job Scheduling in Mesh Multicomputers
abstract
A new approach for dynamic job scheduling in mesh-connected multiprocessor system, which supports a multi-user environment, is proposed. The proposed job scheduler combines a priority-based scheduling policy with a submesh reservation policy to obtain high performance in terms of high throughput, high utilization and low turn-around times for jobs. The proposed scheduling strategy offers the flexibility of achieving high performance at the expense of short-term 'fairness' towards certain jobs. A fast and efficient implementation of the proposed scheduler has also been presented. Simulation results indicate that our scheduling strategy outperforms the FCFS policy significantly by reducing the average waiting delay significantly.
Debendra Das Sharma, Dhiraj K. Pradhan
ICPP (2)1
1993 Fast and Efficient Strategies for Cubic and Non-Cubic Allocation in Hypercube Multiprocessors
abstract
A new approach for dynamic processor allo cation in hypercube multiprocessors which sup ports a multi-user environment is proposed. A dynamic binary tree is used for processor allo cation along with an array of free lists. Two al gorithms are proposed based on this approach that are capable of handling cubic as well as non-cubic allocation efficiently. The time com plexities for both allocation and deallocation are shown to be polynomial; orders of mag nitude improvement over the existing expo nential and even super-exponential algorithms. Unlike the existing strategies, the proposed strategies are best-fit strategies and do not ex cessively fragment the hypercube. Simulation results indicate that the proposed strategies outperform the existing ones in terms of pa rameters such as average delay in honoring a request, average allocation time and average deallocation time.
Debendra Das Sharma, Dhiraj K. Pradhan
ICPP (1)1