Keren Bergman

dblp:34/3257 · DBLP profile ↗
← Back
45ranked-venue papers
2as first author
10since 2021 · last 2025
0000-0001-8580-1728ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 33 · 2 first-author · 7 since 2021Computer networks · 10 · 3 since 2021Software engineering, systems software and programming languages · 3Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Silicon Photonic Accelerated Memory Pooling for Efficient Compute Resource Allocation
abstract
Training and inference of large-scale machine learning models place diverse demands on compute systems, with increasingly higher requirements for processing power, memory capacity, and memory bandwidth. Current compute systems mainly comprise CUs tightly coupled to a fixed amount of high-bandwidth memory (HBM), which needs to be placed in close proximity to the compute die due to constraints related to signal integrity, energy consumption and limited die periphery. This tight coupling limits the system's ability to flexibly adapt to varying compute and memory demands. In this work, we start by exploring the design space of co-packaged photonic I/Os and identifying compute die's shoreline width as a critical architectural constraint. We then present SiPAM, a silicon photonic accelerated memory pooling architecture that integrates dense wavelength division multiplexed (DWDM) photonic I/Os directly along the compute die's shoreline. By replacing local HBMs and conventional electrical networking interfaces with photonic I/Os, we build an unified high-bandwidth photonic communication domain that enables reconfigurable access to both memory and network resources. This architectural design decouples memory and network scaling from compute packaging constraints and allows the system to flexibly adapt to workload demands. We then develop a quantitative optimization methodology to efficiently allocate compute and memory resources based on workloads and parallelism strategies. Our evaluations show that SiPAM scales effectively with model size and I/O bandwidth, improves compute utilization across GPU and memory generations, and achieves up to 3.5×speedup in end-to-end iteration time and up to 5.46× improvement in system efficiency compared to state-of-the-art Nvidia B100 GPU systems.
Zhenguo Wu, Keren Bergman
HOTI2
2025 Scaling Co-Packaged Optical Interconnects Using Hybrid 2.5D/3D Integration
abstract
Tightly integrated optical interconnects can provide high-bandwidth, energy-efficient inter-node communication. We describe a novel system which uses hybrid 2.5D/3D integration to compose a state-of-the-art FPGA compute chiplet, three electrical interface chiplets, and three photonic interface chiplets. We use register-transfer-, gate-, transistor-, and device-level simulations to demonstrate the potential for this system to achieve 96Tb/s of bi-directional bandwidth, and we experimentally demonstrate key components including a complete opto-electrical channel. Our results provide a strong case for hybrid 2.5D/3D integration as the key enabler for scaling co-packaged optical interconnects.
Austin Rovinski, Yanghui Ou, Christine Ou, Devesh Khilwani, Yuyang Wang 0003, Songli Wang, Sunwoo Lee 0002, Keren Bergman, Alyosha C. Molnar, Christopher Batten
ISCAS8
2025 ACTINA: Adapting Circuit-Switching Techniques for AI Networking Architectures
abstract
While traditional datacenters rely on static, electrically switched fabrics, Optical Circuit Switch (OCS)-enabled reconfigurable networks offer dynamic bandwidth allocation and lower power consumption. This work introduces a quantitative framework for evaluating reconfigurable networks in large-scale AI systems, guiding the adoption of various OCS and link technologies by analyzing trade-offs in reconfiguration latency, link bandwidth provisioning, and OCS placement. Using this framework, we develop two in-workload reconfiguration strategies and propose an OCS-enabled, multi-dimensional all-to-all topology that supports hybrid parallelism with improved energy efficiency. Our evaluation demonstrates that with state-of-the-art per-GPU bandwidth, the optimal in-workload strategy achieves up to 2.3 × improvement over the commonly used one-shot approach when reconfiguration latency is low (<100 μ s). However, with sufficiently high bandwidth, one-shot reconfiguration can achieve comparable performance without requiring in-workload reconfiguration. Additionally, our proposed topology improves performance–power efficiency, achieving up to 1.75 × better trade-offs than Fat-Tree and 3D-Torus–based OCS network architectures.
Zhenguo Wu, Benjamin Klenk, Larry Dennison, Keren Bergman
SC4
2024 3D-Integrated, Low Power, High Bandwidth Density Opto-Electronic Transceiver
abstract
We demonstrate a dense, highly parallel, and scalable multi-channel transceiver array for photonic chip-to-chip links. A CMOS electronic chip is flip-chipped onto a silicon photonic chip with 25 μm-pitch bumps to enable 0.8 Tbps data transmission through a single fiber utilizing a comb laser at a record bandwidth density of 5.3 Tbps/mm2. The unit transmitter and receiver cells inside the 80 channel-array incorporate high speed circuitry to drive, read from, and tune the photonic elements, while only occupying 25 μm × 75 μm each.
Devesh Khilwani, Sunwoo Lee 0002, Christine Ou, Stuart Daudlin, Anthony Rizzo, Songli Wang, Michael Cullen, Keren Bergman, Alyosha C. Molnar
ISCAS8
2023 Efficient Intra-Rack Resource Disaggregation for HPC Using Co-Packaged DWDM Photonics
abstract
The diversity of workload requirements and increasing hardware heterogeneity in emerging high performance computing (HPC) systems motivate resource disaggregation. Resource disaggregation allows compute and memory resources to be allocated individually as required to each workload. However, it is unclear how to efficiently realize this capability and cost-effectively meet the stringent bandwidth and latency requirements of HPC applications. To that end, we describe how modern photonics can be co-designed with modern HPC racks to implement flexible intra-rack resource disaggregation and fully meet the bit error rate (BER) and high escape bandwidth of all chip types in modern HPC racks. Our photonic-based disaggregated rack provides an average application speedup of 11% (46% maximum) for 25 CPU and 61% for 24 GPU benchmarks compared to a similar system that instead uses modern electronic switches for disaggregation. Using observed resource usage from a production system, we estimate that an iso-performance intra-rack disaggregated HPC system using photonics would require 4× fewer memory modules and 2× fewer NICs than a non-disaggregated baseline.
George Michelogiannakis, Yehia Arafa, Brandon Cook 0001, Liang Yuan Dai, Abdel-Hameed A. Badawy, Madeleine Glick, Yuyang Wang 0003, Keren Bergman, John Shalf
CLUSTER8
2023 Enabling Quasi-Static Reconfigurable Networks With Robust Topology Engineering
abstract
Many optical circuit switched data center networks (DCN) have been proposed in the last decade to attain higher capacity and topology reconfigurability, though commercial adoption of these architectures have been minimal. One major challenge these architectures face is the difficulty of handling uncertain traffic demands using commercial optical circuit switches (OCS) with high switching latency. Prior works have generally focused on developing fast-switching OCS prototypes to quickly react to traffic variations through frequent reconfigurations. This approach, however, adds tremendous complexity overhead to the control plane, and raises the barrier for commercial adoption of optical circuit switched data center networks. We propose, a robust topology and routing optimization framework for reconfigurable optical circuit switched data centers. co-optimizes topology and routing based on a convex set of traffic matrices, and offers strict throughput guarantees for any future traffic matrices bounded by the convex set. For the bursty traffic demands that are unbounded by the convex set, we employ a desensitization technique to reduce performance hit. This enables to generate topology and routing solutions capable of handling unexpected traffic changes without relying on frequent topology reconfigurations. Our extensive evaluations based on Facebook’s production DCN traces show that, even with daily reconfigurations which could be realized by current commercial MEMS-based OCSs from Calient Technologies, achieves about 20% lower max link utilization, and about 32% lower average hop count compared to cost-equivalent static topologies. Our work shows that adoption of reconfigurable topologies in commercial DCNs is feasible even without fast OCSs.
Min Yee Teh, Shizhen Zhao, Peirui Cao, Keren Bergman
IEEE/ACM Trans. Netw.4
2022 Optically connected memory for disaggregated data centers
Mauricio G. Palma, Maarten Hattink, Ruth Rubio-Noriega, Lois Orosa 0001, Onur Mutlu, Keren Bergman, Rodolfo Azevedo
J. Parallel Distributed Comput.7
2022 A Case For Intra-rack Resource Disaggregation in HPC
abstract
The expected halt of traditional technology scaling is motivating increased heterogeneity in high-performance computing (HPC) systems with the emergence of numerous specialized accelerators. As heterogeneity increases, so does the risk of underutilizing expensive hardware resources if we preserve today’s rigid node configuration and reservation strategies. This has sparked interest in resource disaggregation to enable finer-grain allocation of hardware resources to applications. However, there is currently no data-driven study of what range of disaggregation is appropriate in HPC. To that end, we perform a detailed analysis of key metrics sampled in NERSC’s Cori, a production HPC system that executes a diverse open-science HPC workload. In addition, we profile a variety of deep-learning applications to represent an emerging workload. We show that for a rack (cabinet) configuration and applications similar to Cori, a central processing unit with intra-rack disaggregation has a 99.5% probability to find all resources it requires inside its rack. In addition, ideal intra-rack resource disaggregation in Cori could reduce memory and NIC resources by 5.36% to 69.01% and still satisfy the worst-case average rack utilization.
George Michelogiannakis, Benjamin Klenk, Brandon Cook 0001, Min Yee Teh, Madeleine Glick, Larry Dennison, Keren Bergman, John Shalf
ACM Trans. Archit. Code Optim.7
2021 Designing data center networks using bottleneck structures
abstract
This paper provides a mathematical model of data center performance based on the recently introduced Quantitative Theory of Bottleneck Structures (QTBS). Using the model, we prove that if the traffic pattern is \textit{interference-free}, there exists a unique optimal design that both minimizes maximum flow completion time and yields maximal system-wide throughput. We show that interference-free patterns correspond to the important set of patterns that display data locality properties and use these theoretical insights to study three widely used interconnects---fat-trees, folded-Clos and dragonfly topologies. We derive equations that describe the optimal design for each interconnect as a function of the traffic pattern. Our model predicts, for example, that a 3-level folded-Clos interconnect with radix 24 that routes 10\% of the traffic through the spine links can reduce the number of switches and cabling at the core layer by 25\% without any performance penalty. We present experiments using production TCP/IP code to empirically validate the results and provide tables for network designers to identify optimal designs as a function of the size of the interconnect and traffic pattern.
Jordi Ros-Giralt, Noah Amsel, Sruthi Yellamraju, James R. Ezick, Richard A. Lethin, Yuang Jiang, Aosong Feng, Leandros Tassiulas, Zhenguo Wu, Min Yee Teh, Keren Bergman
SIGCOMM11
2021 SiP-ML: high-bandwidth optical network interconnects for machine learning training
abstract
This paper proposes optical network interconnects as a key enabler for building high-bandwidth ML training clusters with strong scaling properties. Our design, called SiP-ML, accelerates the training time of popular DNN models using silicon photonics links capable of providing multiple terabits-per-second of bandwidth per GPU. SiP-ML partitions the training job across GPUs with hybrid data and model parallelism while ensuring the communication pattern can be supported efficiently on the network interconnect. We develop task partitioning and device placement methods that take the degree and reconfiguration latency of optical interconnects into account. Simulations using real DNN models show that, compared to the state-of-the-art electrical networks, our approach improves training time by 1.3--9.1x.
Mehrdad Khani Shirkoohi, Manya Ghobadi, Mohammad Alizadeh, Madeleine Glick, Keren Bergman, Amin Vahdat, Benjamin Klenk, Eiman Ebrahimi
SIGCOMM6
2020 Optically Connected Memory for Disaggregated Data Centers
abstract
Recent advances in integrated photonics enable the implementation of reconfigurable, high-bandwidth, and low energy-per-bit interconnects in next-generation data centers. We propose and evaluate an Optically Connected Memory (OCM) architecture that disaggregates the main memory from the computation nodes in data centers. OCM is based on micro-ring resonators (MRRs), and it does not require any modification to the DRAM memory modules. We calculate energy consumption from real photonic devices and integrate them into a system simulator to evaluate performance. Our results show that (1) OCM is capable of interconnecting four DDR4 memory channels to a computing node using two fibers with 1.07 pJ energy-per-bit consumption and (2) OCM performs up to 5.5x faster than a disaggregated memory with 40G PCIe NIC connectors to computing nodes.
Alexander Gazman, Maarten Hattink, Mauricio G. Palma, Meisam Bahadori, Ruth Rubio-Noriega, Lois Orosa 0001, Madeleine Glick, Onur Mutlu, Keren Bergman, Rodolfo Azevedo
SBAC-PAD10
2020 TAGO: rethinking routing design in high performance reconfigurable networks
abstract
Many reconfigurable network topologies have been proposed in the past. However, efficient routing on top of these flexible interconnects still presents a challenge. In this work, we reevaluate key principles that have guided the designs of many routing protocols on static networks, and see how well those principles apply on reconfigurable network topologies. Based on a theoretical analysis of key properties that routing in a reconfigurable network should satisfy to maximize performance, we propose a topology-aware, globally-direct oblivious (TAGO) routing protocol for reconfigurable topologies. Our proposed routing protocol is simple in design and yet, when deployed in conjunction with a reconfigurable network topology, improves throughput by up to 2.2× compared to established routing protocols and even comes within 10% of the throughput of impractical adaptive routing that has instant global congestion information.
Min Yee Teh, Yu-Han Hung, George Michelogiannakis, Shijia Yan, Madeleine Glick, John Shalf, Keren Bergman
SC7
2020 Silicon Photonics Codesign for Deep Learning
abstract
Deep learning is revolutionizing many aspects of our society, addressing a wide variety of decision-making tasks, from image classification to autonomous vehicle control. Matrix multiplication is an essential and computationally intensive step of deep-learning calculations. The computational complexity of deep neural networks requires dedicated hardware accelerators for additional processing throughput and improved energy efficiency in order to enable scaling to larger networks in the upcoming applications. Silicon photonics is a promising platform for hardware acceleration due to recent advances in CMOS-compatible manufacturing capabilities, which enable efficient exploitation of the inherent parallelism of optics. This article provides a detailed description of recent implementations in the relatively new and promising platform of silicon photonics for deep learning. Opportunities for multiwavelength microring silicon photonic architectures codesigned with field-programmable gate array (FPGA) for pre- and postprocessing are presented. The detailed analysis of a silicon photonic integrated circuit shows that a codesigned implementation based on the decomposition of large matrix-vector multiplication into smaller instances and the use of nonnegative weights could significantly simplify the photonic implementation of the matrix multiplier and allow increased scalability. We conclude this article by presenting an overview and a detailed analysis of design parameters. Insights for ways forward are explored.
Qixiang Cheng, Jihye Kwon, Madeleine Glick, Meisam Bahadori, Luca P. Carloni, Keren Bergman
Proc. IEEE6
2019 Bandwidth steering in HPC using silicon nanophotonics
abstract
As bytes-per-FLOP ratios continue to decline, communication is becoming a bottleneck for performance scaling. This paper describes bandwidth steering in HPC using emerging reconfigurable silicon photonic switches. We demonstrate that placing photonics in the lower layers of a hierarchical topology efficiently changes the connectivity and consequently allows operators to recover from system fragmentation that is otherwise hard to mitigate using common task placement strategies. Bandwidth steering enables efficient utilization of the higher layers of the topology and reduces cost with no performance penalties. In our simulations with a few thousand network endpoints, bandwidth steering reduces static power consumption per unit throughput by 36% and dynamic power consumption by 14% compared to a reference fat tree topology. Such improvements magnify as we taper the bandwidth of the upper network layer. In our hardware testbed, bandwidth steering improves total application execution time by 69%, unaffected by bandwidth tapering.
George Michelogiannakis, Yiwen Shen 0002, Min Yee Teh, Xiang Meng 0003, Benjamin Aivazi, Taylor L. Groves, John Shalf, Madeleine Glick, Manya Ghobadi, Larry Dennison, Keren Bergman
SC11
2018 Low-Power Optical Interconnects based on Resonant Silicon Photonic Devices: Recent Advances and Challenges
abstract
The progressive blooming of silicon photonics technology (SiP) over the last decade has indicated that optical interconnects may substitute the electrical wires for data movement over short distances in the future. A key enabler is the resonant structures that can participate in both modulation and demultiplexing of a high throughput wavelength division multiplexed (WDM) photonic link. The optical and electro-optical properties of such devices are subject to various design considerations, operation conditions, and optimization procedures. We present recent technological advances in photonic links based on resonant structures and highlight the key challenges that must be overcome at a large scale. Furthermore, we discuss how the design space of these resonant devices, down to the geometrical parameters and fabrication errors, can affect the performance and reliability of a photonic link.
Meisam Bahadori, Keren Bergman
ACM Great Lakes Symposium on VLSI2
2018 Empowering Flexible and Scalable High Performance Architectures with Embedded Photonics
abstract
The recent explosive growth in data analytics applications that rely on machine and deep learning techniques are seismically changing the landscape of high performance architectures. These techniques rely on graphics processing units (GPU) and manycore (CPU) technologies whose need for intense performance is pushing current interconnect networks to their limits. Driven by these applications, the execution performance along with the energy consumption of massive parallel systems is increasingly determined by how data is moved among the numerous compute and memory resources. Embedded photonic interconnect technologies can address critical data-movement challenges by delivering higher communication bandwidth densities at significantly improved energy efficiencies. New disaggregated architectures enabled by embedded photonics and optical bandwidth steering can reduce the system-wide energy consumption.
Keren Bergman
IPDPS1
2017 Energy-performance optimized design of silicon photonic interconnection networks for high-performance computing
abstract
We present detailed electrical and optical models of the elements that comprise a WDM silicon photonic link. The electronics is assumed to be based on 65 nm CMOS node and the optical modulators and demultiplexers are based on microring resonators. The goal of this study is to analyze the energy consumption and scalability of the link by finding the right combination of (number of channels × data rate per channel) that fully covers the available optical power budget. Based on the set of empirical and analytical models presented in this work, a maximum capacity of 0.75 Tbps can be envisioned for a point-to-point link with an energy consumption of 1.9 pJ/bit. Sub-pJ/bit energy consumption is also predicted for aggregated bitrates up to 0.35 Tbps.
Meisam Bahadori, Sébastien Rumley, Robert P. Polster, Alexander Gazman, Matt Traverso, Mark Webster, Kaushik Patel, Keren Bergman
DATE8
2017 Optical interconnects for extreme scale computing systems
Sébastien Rumley, Meisam Bahadori, Robert P. Polster, Simon D. Hammond, David M. Calhoun, Arun Rodrigues, Keren Bergman
Parallel Comput.8
2016 Flexfly: enabling a reconfigurable dragonfly through silicon photonics
abstract
The Dragonfly topology provides low-diameter connectivity for high-performance computing with all-to-all global links at the inter-group level. Our traffic matrix characterization of various scientific applications shows consistent mismatch between the imbalanced group-to-group traffic and the uniform global bandwidth allocation of Dragonfly. Though adaptive routing has been proposed to utilize bandwidth of non-minimal paths, increased hops and cross-group interference lower efficiency. This work presents a photonic architecture, Flexfly, which “trades” global links among groups using low-radix Silicon photonic switches. With transparent optical switching, Flexfly reconfigures the inter-group topology based on traffic pattern, stealing additional direct bandwidth for communication-intensive group pairs. Simulations with applications such as GTC, Nekbone and LULESH show up to 1.8x speedup over Dragonfly paired with UGAL routing, along with halved hop count and latency for cross-group messages. We built a 32-node Flexfly prototype using a Silicon photonic switch connecting four groups and demonstrated 820 ns interconnect reconfiguration time.
Payman Samadi, Sébastien Rumley, Christine P. Chen, Yiwen Shen 0002, Meisam Bahadori, Keren Bergman, Jeremiah J. Wilke
SC7
2014 Accelerating incast and multicast traffic delivery for data-intensive applications using physical layer optics
abstract
We present a control plane architecture to accelerate multicast and incast traffic delivery for data-intensive applications in cluster-computing interconnection networks. The architecture is experimentally examined by enabling physical layer optical multicasting on-demand for the application layer to achieve non-blocking performance.
Payman Samadi, Varun Gupta 0002, Berk Birand, Howard Wang, Gil Zussman, Keren Bergman
SIGCOMM6
2014 Real-Time Power Control for Dynamic Optical Networks - Algorithms and Experimentation
abstract
Core and aggregation optical networks are remarkably static, despite the emerging dynamic capabilities of the individual optical devices. This stems from the inability to address optical impairments in real-time. As a result, tasks such as adding and removing wavelengths take a substantial amount of time, and therefore, optical networks are over-provisioned and inefficient in terms of capacity and energy. Optical Performance Monitors (OPMs) that assess the Quality of Transmission (QoT) in real-time can be used to overcome these inefficiencies. However, prior work mostly focused on the single link level. In this paper, we present a network-wide optimization algorithm that leverages OPM measurements to dynamically control the wavelengths' power levels. Hence, it allows adding and dropping wavelengths quickly while mitigating the impacts of impairments caused by these actions, thereby facilitating efficient operation of higher layer protocols. We evaluate the algorithm's performance using a network-scale optical simulator under real-world scenarios and show that the ability to add and drop wavelengths dynamically can lead to significant power savings. Moreover, we experimentally evaluate the algorithm in an optical testbed and discuss the practical implementation issues. To the best of our knowledge, this paper is the first attempt at providing a global power control algorithm that uses live OPM measurements to enable dynamic optical networking.
Berk Birand, Howard Wang, Keren Bergman, Daniel C. Kilper, Thyaga Nandagopal, Gil Zussman
IEEE J. Sel. Areas Commun.3
2013 Real-time power control for dynamic optical networks - Algorithms and experimentation
abstract
Core and aggregation optical networks are remarkably static, despite the emerging dynamic capabilities of the individual optical devices. This stems from the inability to address optical impairments in real-time. As a result, tasks such as adding and removing wavelengths take a substantial amount of time, and therefore, optical networks are over-provisioned and inefficient in terms of capacity and energy. Optical Performance Monitors (OPMs) that assess the Quality of Transmission (QoT) in real-time can be used to overcome these inefficiencies. However, prior work mostly focused on the single link level. In this paper, we present a network-wide optimization algorithm that leverages OPM measurements to dynamically control the wavelengths' power levels. Hence, it allows adding and dropping wavelengths quickly while mitigating the impacts of impairments caused by these actions, thereby facilitating efficient operation of higher layer protocols. We evaluate the algorithm's performance using a network-scale optical simulator under real-world scenarios and show that the ability to add and drop wavelengths dynamically can lead to significant power savings. Moreover, we experimentally evaluate the algorithm in an optical testbed and discuss the practical implementation issues. To the best of our knowledge, this paper is the first attempt at providing a global power control algorithm that uses live OPM measurements to enable dynamic optical networking.
Berk Birand, Howard Wang, Keren Bergman, Daniel C. Kilper, Thyaga Nandagopal, Gil Zussman
ICNP3
2013 P-sync: A Photonically Enabled Architecture for Efficient Non-local Data Access
abstract
Communication in multi- and many-core processors has long been a bottleneck to performance due to the high cost of long-distance electrical transmission. This difficulty has been partially remedied by architectural constructs such as caches and novel interconnect topologies, albeit at a steep cost in terms of complexity. Unfortunately, even these measures are rendered ineffective by certain kinds of communication, most notably scatter and gather operations that exhibit highly nonlocal data access patterns. Much work has gone into examining how the increased bandwidth density afforded by chip-scale silicon photonic interconnect technologies affects computing, but photonics have additional properties that can be leveraged to greatly accelerate performance and energy efficiency under such difficult loads. This paper describes a novel synchronized global photonic bus and system architecture called P-sync that uses photonics' distance independence to greatly improve performance on many important applications previously limited by electronic interconnect. The architecture is evaluated in the context of a non-local yet common application: the distributed Fast Fourier Transform. We show that it is possible to achieve high efficiency by tightly balancing computation and communication latency in P-sync and achieve upwards of a 6× performance increase on gather patterns, even when bandwidth is equalized.
David Whelihan, Jeffrey J. Hughes, Scott M. Sawyer, Eric Robinson, Michael M. Wolf, Sanjeev Mohindra, Julie Mullen, Anna Klein, Michelle S. Beard, Nadya Bliss, Johnnie Chan, Robert Hendry, Keren Bergman, Luca P. Carloni
IPDPS13
2012 Cross-layer enabled translucent optical network with real-time impairment awareness
abstract
The existing dimensioning strategy for translucent, sub-wavelength switching architectures relies on over-provisioning, and consequently, overuse of costly, power-consuming optical-electrical-optical (O/E/O) regenerators. In addition, due to a variety of external phenomena, many physical layer impairments are time-varying, and hence, can strongly degrade network performance. In this work, we introduce a Cross-Layer Optical Network Element (CLONE) used for the dynamic management of physical layer impairments in the network. We investigate the impact of real-time impairment aware routing in a CLONE-enabled optical network with sub-wavelength switching. Simulation results show that the CLONE-enabled network architecture provides improvements in: (1) energy efficiency by optimizing the usage of regenerators, and (2) network performance in terms of the packet loss probability.
Oscar Pedrola, Balagangadhar G. Bathula, Michael S. Wang, Atiyah Ahsan, Davide Careglio, Keren Bergman
GLOBECOM6
2011 Burst-Mode Transmission and Data Recovery for Multi-GHz Optical Packet Switching Network Testing
abstract
This paper describes the challenges associated with recovery of multi-GHz short burst communications in multi-channel systems, especially DWDM optical packet switching networks. Burst-mode data transmission complicates the standard methods typically used to recover embedded clock information and recover the serial data. Previously implemented solutions to this problem either constrain the test capability or are not extensible to higher signaling rates. A new approach, capable of locking in a single cycle at rates up to 10 Gbps is described in this paper. This solution applies to both the testing strategies and a receiver circuit used for an end-application optical switching network interface.
Carl Edward Gray, David C. Keezer, Howard Wang, Keren Bergman
Asian Test Symposium4
2011 VANDAL: A tool for the design specification of nanophotonic networks
abstract
Continuing to scale CMP performance at reasonable power budgets has forced chip designers to consider emerging silicon-photonic technologies as the primary means of on- and off-chip communication. Different designs for chip-scale photonic interconnects have been proposed, and system-level simulations have shown them to be far superior to purely electronic network solutions. However, specifying the exact geometries for all the photonic devices used in these networks is currently a time-consuming and difficult manual process. We present VANDAL, a layout tool which provides a user with semi-automatic assistance for placing silicon photonic devices, modifying their geometries, and routing waveguides for hierarchically building photonic networks. VANDAL also includes SCILL, a scripting language that can be used to automate photonic device place and route for repeatability, automation, verification, and scaling. We demonstrate some of the features and flexibility of the CAD environment with a case study, designing modulator and detector banks for integrated photonic links.
Gilbert Hendry, Johnnie Chan, Luca P. Carloni, Keren Bergman
DATE4
2011 Load-Aware Anycast Routing in IP-over-WDM Networks
abstract
In this work we propose anycast routing methods to improve the performance of reconfigurable WDM networks under the variations in the IP traffic. We first investigate anycast communication via impairment-aware anycast routing (IAAR); our simulation results show significant improvement in the blocking probability. We also investigate the proposed load-aware anycast routing (LAAR) for the varying traffic model. From the results we observe that LAAR minimizes the lightpath request loss, by dynamically choosing the anycast configuration based on the network load.
Balagangadhar G. Bathula, Vinod Vokkarane, Caroline P. Lai, Keren Bergman
ICC4
2011 Photonic network-on-chip architectures using multilayer deposited silicon materials for high-performance chip multiprocessors
abstract
Integrated photonics has been slated as a revolutionary technology with the potential to mitigate the many challenges associated with on- and off-chip electrical interconnection networks. To date, all proposed chip-scale photonic interconnects have been based on the crystalline silicon platform for CMOS-compatible fabrication. However, maintaining CMOS compatibility does not preclude the use of other CMOS-compatible silicon materials such as silicon nitride and polycrystalline silicon. In this work, we investigate utilizing devices based on these deposited materials to design photonic networks with multiple layers of photonic devices. We apply rigorous device optimization and insertion loss analysis on various network architectures, demonstrating that multilayer photonic networks can exhibit dramatically lower total insertion loss, enabling unprecedented bandwidth scalability. We show that significant improvements in waveguide propagation and waveguide crossing insertion losses resulting from using these materials enables the realization of topologies that were previously not feasible using only the single-layer crystalline silicon approaches.
Aleksandr Biberman, Kyle Preston, Gilbert Hendry, Nicolás Sherwood-Droz, Johnnie Chan, Jacob S. Levy, Michal Lipson, Keren Bergman
ACM J. Emerg. Technol. Comput. Syst.8
2011 Time-division-multiplexed arbitration in silicon nanophotonic networks-on-chip for high-performance chip multiprocessors
Gilbert Hendry, Eric Robinson, Vitaliy Gleyzer, Johnnie Chan, Luca P. Carloni, Nadya Bliss, Keren Bergman
J. Parallel Distributed Comput.7
2011 Physical-Layer Modeling and System-Level Design of Chip-Scale Photonic Interconnection Networks
abstract
Photonic technology is becoming an increasingly attractive solution to the problems facing today's electronic chip-scale interconnection networks. Recent progress in silicon photonics research has enabled the demonstration of all the necessary optical building blocks for creating extremely high-bandwidth density and energy-efficient links for on-chip and off-chip communications. From the feasibility and architecture perspective however, photonics represents a dramatic paradigm shift from traditional electronic network designs due to fundamental differences in how electronics and photonics function and behave. As a result of these differences, new modeling and analysis methods must be employed in order to properly realize a functional photonic chip-scale interconnect design. In this paper, we present a methodology for characterizing and modeling fundamental photonic building blocks which can subsequently be combined to form full photonic network architectures. We also describe a set of tools which can be utilized to assess the physical-layer and system-level performance properties of a photonic network. The models and tools are integrated in a novel open-source design and simulation environment. We present a case study of two different photonic networks-on-chip to demonstrate how our improved understanding and modeling of the physical-layer details of photonic communications can be used to better understand the system-level performance impact.
Johnnie Chan, Gilbert Hendry, Keren Bergman, Luca P. Carloni
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2010 PhoenixSim: A simulator for physical-layer analysis of chip-scale photonic interconnection networks
abstract
Recent developments have shown the possibility of leveraging silicon nanophotonic technologies for chip-scale interconnection fabrics that deliver high bandwidth and power efficient communications both on- and off-chip. Since optical devices are fundamentally different from conventional electronic interconnect technologies, new design methodologies and tools are required to exploit the potential performance benefits in a manner that accurately incorporates the physically different behavior of photonics. We introduce PhoenixSim, a simulation environment for modeling computer systems that incorporates silicon nanophotonic devices as interconnection building blocks. PhoenixSim has been developed as a cross-discipline platform for studying photonic interconnects at both the physical-layer level and at the architectural and system levels. The broad scope at which modeled systems can be analyzed with PhoenixSim provides users with detailed information into the physical feasibility of the implementation, as well as the network and system performance. Here, we describe details about the implementation and methodology of the simulator, and present two case studies of silicon nanophotonic-based networks-on-chip.
Johnnie Chan, Gilbert Hendry, Aleksandr Biberman, Keren Bergman, Luca P. Carloni
DATE4
2010 Hybrid on-chip data networks
Gilbert Hendry, Keren Bergman
Hot Chips Symposium2
2010 Tools and methodologies for designing energy-efficient photonic networks-on-chip for highperformance chip multiprocessors
abstract
Photonic interconnection networks have recently been proposed as a replacement to conventional electronic network-on-chip solutions in delivering the ever increasing communication requirements of future chip multiprocessors. While photonics offers superior bandwidth density, lower latencies, and improvements in energy efficiency over electronics, the photonic network designs that can leverage these benefits cannot be easily derived by simply mimicking electronic layouts. In fact, proper implementations of photonic interconnects will require the careful consideration of a variety new physical-layer metrics and design requirements that did not exist with electronics. Here, we review some of the currently proposed designs for chip-scale photonic interconnection networks, the design methodologies required to produce viable network topologies, and a simulation environment, called PhoenixSim, that we have developed to accurately model and study those metrics and designs.
Johnnie Chan, Gilbert Hendry, Aleksandr Biberman, Keren Bergman
ISCAS4
2010 Photonic Chip-Scale Interconnection Networks for Performance-Energy Optimized Computing
abstract
As chip multiprocessors (CMPs) scale to increasing numbers of cores and greater on-chip computational power, the gap between the available off-chip bandwidth and that which is required to appropriately feed the processors continues to widen under current memory access architectures. For many high-performance computing applications, the bandwidth available for both on- and off-chip communications can play a vital role in efficient execution due to the use of data-parallel or data-centric algorithms. Electronic interconnected systems are increasingly bound by their communications infrastructure and the associated power dissipation of high-bandwidth data movement. Recent advances in chip-scale silicon photonic technologies have created the potential for developing optical interconnection networks that can offer highly energy efficient communications and significantly improve computing performance-per-Watt. This talk will examine the design and performance of photonic networks-on-chip architectures that support both on-chip communication and off-chip memory access in an energy efficient manner.
Keren Bergman
NOCS1
2010 Circuit-Switched Memory Access in Photonic Interconnection Networks for High-Performance Embedded Computing
abstract
As advancements in CMOS technology trend toward ever increasing core counts in chip multiprocessors for high-performance embedded computing, the discrepancy between on- and off-chip communication bandwidth continues to widen due to the power and spatial constraints of electronic off-chip signaling. Silicon photonics-based communication offers many advantages over electronics for network-on-chip design, namely power consumption that is effectively agnostic to distance traveled at the chip- and board-scale, even across chip boundaries. In this work we develop a design for a photonic network-on-chip with integrated DRAM I/O interfaces and compare its performance to similar electronic solutions using a detailed network-on-chip simulation. When used in a circuit-switched network, silicon nanophotonic switches offer higher bandwidth density and low power transmission, adding up to over 10x better performance and 3-5x lower power over the baseline for projective transform, matrix multiply, and Fast Fourier Transform (FFT), all key algorithms in embedded real-time signal and image processing.
Gilbert Hendry, Eric Robinson, Vitaliy Gleyzer, Johnnie Chan, Luca P. Carloni, Nadya Bliss, Keren Bergman
SC7
2009 Networking hardware: what drives innovation?
Jack Brassil, Jonathan M. Smith, Flavio Bonomi, Keren Bergman, Paul Congdon, Ivan Seskar, Steve Muir
ANCS4
2009 Analysis of photonic networks for a chip multiprocessor using scientific applications
abstract
As multiprocessors scale to unprecedented numbers of cores in order to sustain performance growth, it is vital that these gains are not nullified by high energy consumption from inter-core communication. With recent advances in 3D Integration CMOS technology, the possibility for realizing hybrid photonic-electronic networks-on-chip warrants investigating real application traces on functionally comparable photonic and electronic network designs. We present a comparative analysis using both synthetic benchmarks as well as real applications, run through detailed cycle accurate models implemented under the OMNeT++ discrete event simulation environment. Results show that when utilizing standard process-to-processor mapping methods, this hybrid network can achieve 75times improvement in energy efficiency for synthetic benchmarks and up to 37times improvement for real scientific applications, defined as network performance per energy spent, over an electronic mesh for large messages across a variety of communication patterns.
Gilbert Hendry, Shoaib Kamil 0001, Aleksandr Biberman, Johnnie Chan, Benjamin G. Lee, Marghoob Mohiyuddin, Keren Bergman, Luca P. Carloni, John Kubiatowicz, Leonid Oliker, John Shalf
NOCS8
2008 Photonic networks-on-chip: Opportunities and challenges
abstract
As the number of processing cores that are integrated into a chip multiprocessors (CMP) continues to grow, the network-on-chip paradigm has emerged as a promising solution to address the problem of providing a robust interconnect network among them. In future high-performance CMPs, however, the high bandwidth requirements for both intra-chip and off-chip communication are severely challenging the electronic communications infrastructure to meet these demands without consuming a large fraction of the overall on-chip power-dissipation budget. The introduction of photonic technology for on-chip communication holds the promise of delivering scalable bandwidth-per-watt performance that cannot be achieved using only electronic communication. After reviewing the key recent technologies advances that are making possible the integration of photonic devices with CMOS processes, we describe a hybrid micro-architecture for NoCs that combines a broadband photonic circuit-switched network with an electronic packet-switched control network and we discuss the pros and cons of using two different network topologies to implement it.
Michele Petracca, Keren Bergman, Luca P. Carloni
ISCAS2
2008 Photonic Networks-on-Chip for Future Generations of Chip Multiprocessors
abstract
The design and performance of next-generation chip multiprocessors (CMPs) will be bound by the limited amount of power that can be dissipated on a single die. We present photonic networks-on-chip (NoC) as a solution to reduce the impact of intra-chip and off-chip communication on the overall power budget. A photonic interconnection network can deliver higher bandwidth and lower latencies with significantly lower power dissipation. We explain why on-chip photonic communication has recently become a feasible opportunity and explore the challenges that need to be addressed to realize its implementation. We introduce a novel hybrid micro-architecture for NoCs combining a broadband photonic circuit-switched network with an electronic overlay packet-switched control network. We address the critical design issues including: topology, routing algorithms, deadlock avoidance, and path-setup/tear-down procedures. We present experimental results obtained with POINTS, an event-driven simulator specifically developed to analyze the proposed idea, as well as a comparative power analysis of a photonic versus an electronic NoC. Overall, these results confirm the unique benefits for future generations of CMPs that can be achieved by bringing optics into the chip in the form of photonic NoCs.
Assaf Shacham, Keren Bergman, Luca P. Carloni
IEEE Trans. Computers2
2007 The Case for Low-Power Photonic Networks on Chip
abstract
Packet-switched networks on chip (NoC) have been advocated as a natural communication mechanism among the processing cores in future chip multiprocessors (CMP). However, electronic NoCs do not directly address the power budget problem that limits the design of high-performance chips in nanometer technologies. We make the case for a hybrid approach to NoC design that combines a photonic transmission layer with an electronic control layer. A comparative power analysis with a fully-electronic NoC shows that large bandwidths can be exchanged at dramatically lower power consumption.
Assaf Shacham, Keren Bergman, Luca P. Carloni
DAC2
2007 Co-development of test electronics and PCI Express interface for a multi-Gbps optical switching network
abstract
This paper presents the design and performance characteristics of a system designed to interface between a PCI Express port and an optical packet switched network as well as provide inline test capability for the whole system. A single lane of PCI Express traffic is inverse multiplexed across eight parallel channels and retransmitted in a burst packet at aggregate data rates of 20 to 36 Gbps. Testing options include loopback self-test, data synthesis and substitution in-line with or in place of system data, and variable channel-to-channel skew. The I/O interfaces also support a range of variable analog parameters such as peak output amplitude, output amplitude swing, and common mode ranges for the input and output to evaluate the performance of or adapt to changes in the opto-electronic components. This design flexibility also allows for use of the system in more conventional electronic applications with little or no required modifications.
Carl Edward Gray, Odile Liboiron-Ladouceur, David C. Keezer, Keren Bergman
ITC4
2007 On the Design of a Photonic Network-on-Chip
abstract
Recent remarkable advances in nanoscale silicon-photonic integrated circuitry specifically compatible with CMOS fabrication have generated new opportunities for leveraging the unique capabilities of optical technologies in the on-chip communications infrastructure. Based on these nano-photonic building blocks, we consider a photonic network-on-chip architecture designed to exploit the enormous transmission bandwidths, low latencies, and low power dissipation enabled by data exchange in the optical domain. The novel architectural approach employs a broadband photonic circuit-switched network driven in a distributed fashion by an electronic overlay control network which is also used for independent exchange of short messages. We address the critical network design issues for insertion in chip multiprocessors (CMP) applications, including topology, routing algorithms, path-setup and tear-down procedures, and deadlock avoidance. Simulations show that this class of photonic networks-on-chip offers a significant leap in the performance for CMP intrachip communication systems delivering low-latencies and ultra-high throughputs per core while consuming minimal power
Assaf Shacham, Keren Bergman, Luca P. Carloni
NOCS2
2007 The Data Vortex, an All Optical Path Multicomputer Interconnection Network
abstract
All optical path interconnection networks employing dense wavelength division multiplexing can provide vast improvements in supercomputer performance. However, the lack of efficient optical buffering requires investigation of new topologies and routing techniques. This paper introduces and evaluates the data vortex optical switching architecture which uses cylindrical routing paths as a packet buffering alternative. In addition, the impact of the number of angles on the overall network performance is studied through simulation. Using optimal topology configurations, the data vortex is compared to two existing switching architectures-butterfly and omega networks. The three networks are compared in terms of throughput, accepted traffic ratio, and average packet latency. The data vortex is shown to exhibit comparable latency and a higher acceptance rate (2times at 50 percent load) than the butterfly and omega topologies
Cory Hawkins, Benjamin A. Small, D. Scott Wills, Keren Bergman
IEEE Trans. Parallel Distributed Syst.4
2003 Application and Demonstration of a Digital Test Core: Optoelectronic Test Bed and Wafer-level Prober
abstract
Abstract A multi-purpose digital test core utilizing programmable logic has been introduced [1,2] to implement many of the functions of traditional automated test equipment (ATE). While previous papers have described the theory, this paper quantifies the results and presents additional applications with improved methods operating up to 4.4Gpbs. The digital test core provides a substantial number of programmable I/O for testing circuits and systems. It may be used either to enhance the capabilities of ATE or to provide autonomous testing within large systems or arrays of components. This technique has been expanded upon to produce greater functionality at higher frequencies. Based upon limitations of current ATE and BIST, the need for the digital test core is described. The test core concept is reviewed within an opto-electronic pattern generator and sampler with an eventual goal of terabit-per-second aggregate data rate. The performance of the device is discussed, and a second application of the digital test core is introduced as a nano-scale wafer-level embedded tester.
John S. Davis, David C. Keezer, Odile Liboiron-Ladouceur, Keren Bergman
ITC4
1996 Transparent Optical Networks with Time-Division Multiplexing (Invited Paper)
abstract
Most research efforts to date on optical networks have concentrated on wavelength-division multiplexing (WDM) techniques where the information from different channels is routed via separate optical wavelengths. The data corresponding to a particular channel is selected at the destination node by a frequency filter. Optical time-division multiplexing (OTDM) has been considered as an alternative to WDM for future networks operating in excess of 10 Gb/s. Systems based on TDM techniques rely upon a synchronized clock frequency and timing to separate the multiplexed channels. Advances in device technologies have opened new opportunities for implementing OTDM in very high-speed long-haul transmission as well as networking. The multiterahertz bandwidth made available with the advent of optical fibers has spurred investigation and development of transparent all-optical networks that may overcome the bandwidth bottlenecks caused by electro-optic conversion. This paper presents an overview of current OTDM networks and their supporting technologies. A novel network architecture is introduced, aimed at offering both ultra-high speed (up to 100 Gb/s) and maximum parallelism for future terabit data communications. Our network architecture is based on several key state-of-the-art optical technologies that we have demonstrated.
Seung-Woo Seo, Keren Bergman, Paul R. Prucnal
IEEE J. Sel. Areas Commun.2