Ray T. Chen

dblp:53/4318 · DBLP profile ↗
← Back
26ranked-venue papers
2as first author
15since 2021 · last 2026
0000-0002-9181-4266ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 20 · 1 first-author · 10 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Software engineering, systems software and programming languages · 4 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 ENLighten: Lighten the Transformer, Enable Efficient Optical Acceleration
abstract
Photonic computing has emerged as a promising substrate for accelerating the dense linear-algebra operations at the heart of AI, but its adoption for large Transformer models remains in its infancy. In supporting these massive models, we identify two key bottlenecks: (1) costly electro-optic conversions and data-movement overheads that erode energy efficiency as model sizes scale; (2) a mismatch between limited on-chip photonic resources and the scale of Transformer workloads, which forces frequent reuse of photonic tensor cores and dilutes throughput gains. To address these challenges, we introduce a hardware–software co-design framework. First, we propose Lighten, a PTC-aware compression flow that post-hoc decomposes each Transformer weight matrix into a low-rank component plus a structured sparse component aligned to photonic tensor-core granularity, all without lengthy retraining. Second, we present ENLighten, a reconfigurable photonic accelerator architecture featuring dynamically adaptive tensor cores, driven by a broadband light redistribution, for fine-grained sparsity support and full power gating of inactive parts. On ImageNet, Lighten prunes Base-scale vision transformer by 50% with a ~1% accuracy drop after only 3 epochs of fine-tuning within 1 hour, and when deployed on ENLighten it achieves a $2.5 \times$ improvement in energy-delay product over the state-of-the-art photonic Transformer accelerator.
Hanqing Zhu, Zhican Zhou, Shupeng Ning, Xuhao Wu, Ray T. Chen, Yating Wan, David Z. Pan
ASP-DAC5
2024 Lightening-Transformer: A Dynamically-Operated Optically-Interconnected Photonic Transformer Accelerator
abstract
The wide adoption and significant computing resource cost of attention-based transformers, e.g., Vision Transformers and large language models, have driven the demand for efficient hardware accelerators. While electronic accelerators have been commonly used, there is a growing interest in exploring photonics as an alternative technology due to its high energy efficiency and ultra-fast processing speed. Photonic accelerators have demonstrated promising results for convolutional neural networks (CNNs) workloads, which predominantly rely on weight-static linear operations. However, they encounter challenges when it comes to efficiently supporting attention-based Transformer architectures, raising questions about the applicability of photonics to advanced machine-learning tasks. The primary hurdle lies in their inefficiency in handling the unique workloads inherent to Transformers, i.e., dynamic and full-range tensor multiplication. In this work, we propose Lightening-Transformer, the first light-empowered, high-performance, and energy-efficient photonic Transformer accelerator. To overcome the fundamental limitation of existing photonic tensor core designs, we introduce a novel dynamically-operated photonic tensor core, DPTC, consisting of a crossbar array of interference-based optical vector dot-product engines, supporting highly parallel, dynamic, and full-range matrix multiplication. Furthermore, we design a dedicated accelerator that integrates our novel photonic computing cores with photonic interconnects for inter-core data broadcast, fully unleashing the power of optics. The comprehensive evaluation demonstrates that Lightening-Transformer achieves >2.6x energy and > 12 x latency reductions compared to prior photonic accelerators and delivers the lowest energy cost and 2 to 3 orders of magnitude lower energy-delay product compared to the electronic Transformer accelerator, all while maintaining digital-comparable accuracy. Our work highlights the immense potential of photonics for efficient hardware accelerators, particularly for advanced machine-learning workloads, such as Transformer-backboned large language models (LLM). Our implementation is available at https://github.com/zhuhanqing/Lightening-Transformer.
Hanqing Zhu, Jiaqi Gu 0002, Hanrui Wang 0002, Zixuan Jiang, Zhekai Zhang, Rongxing Tang, Chenghao Feng, Song Han 0003, Ray T. Chen, David Z. Pan
HPCA9
2024 PACE: Pacing Operator Learning to Accurate Optical Field Simulation for Complicated Photonic Devices
abstract
Electromagnetic field simulation is central to designing, optimizing, and validating photonic devices and circuits. However, costly computation associated with numerical simulation poses a significant bottleneck, hindering scalability and turnaround time in the photonic circuit design process. Neural operators offer a promising alternative, but existing SOTA approaches, Neurolight, struggle with predicting high-fidelity fields for real-world complicated photonic devices, with the best reported 0.38 normalized mean absolute error in Neurolight. The interplays of highly complex light-matter interaction, e.g., scattering and resonance, sensitivity to local structure details, non-uniform learning complexity for full-domain simulation, and rich frequency information, contribute to the failure of existing neural PDE solvers. In this work, we boost the prediction fidelity to an unprecedented level for simulating complex photonic devices with a novel operator design driven by the above challenges. We propose a novel cross-axis factorized PACE operator with a strong long-distance modeling capacity to connect the full-domain complex field pattern with local device structures. Inspired by human learning, we further divide and conquer the simulation task for extremely hard cases into two progressively easy tasks, with a first-stage model learning an initial solution refined by a second model. On various complicated photonic device benchmarks, we demonstrate one sole PACE model is capable of achieving 73% lower error with 50% fewer parameters compared with various recent ML for PDE solvers. The two-stage setup further advances high-fidelity simulation for even more intricate cases. In terms of runtime, PACE demonstrates 154-577x and 11.8-12x simulation speedup over numerical solver using scipy or highly-optimized pardiso solver, respectively. We open-sourced the code and *complicated* optical device dataset at [PACE-Light](https://github.com/zhuhanqing/PACE-Light).
Hanqing Zhu, Wenyan Cong, Guojin Chen, Shupeng Ning, Ray T. Chen, Jiaqi Gu 0002, David Z. Pan
NeurIPS5
2023 SqueezeLight: A Multi-Operand Ring-Based Optical Neural Network With Cross-Layer Scalability
abstract
Optical neural networks (ONNs) are promising hardware platforms for next-generation artificial intelligence acceleration with ultrafast speed and low-energy consumption. However, previous ONN designs are bounded by one multiply–accumulate operation per device, showing unsatisfying scalability. In this work, we propose a scalable ONN architecture, dubbedSqueezeLight. We propose a nonlinear optical neuron based on multioperand ring resonators (MORRs) to squeeze vector dot-product into a single device with low wavelength usage and built-in nonlinearity. A block-level squeezing technique with structured sparsity is exploited to support higher scalability. We adopt a robustness-aware training algorithm to guarantee variation tolerance. To enable a truly scalable ONN architecture, we extendSqueezeLightto a separable optical CNN architecture that further squeezes in the layer level. Two orthogonal convolutional layers are mapped to one MORR array, leading to order-of-magnitude higher software training scalability. We further explore augmented representability forSqueezeLightby introducing parametric MORR neurons with trainable nonlinearity, together with a nonlinearity-aware initialization method to stabilize convergence. Experimental results show thatSqueezeLightachieves one-order-of-magnitude better compactness and efficiency than previous designs with high fidelity, trainability, and robustness. Our open-source codes are available athttps://github.com/JeremieMelo/SqueezeLight.
Jiaqi Gu 0002, Chenghao Feng, Hanqing Zhu, Zheng Zhao 0003, Zhoufeng Ying, Ray T. Chen, David Z. Pan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2023 ELight: Toward Efficient and Aging-Resilient Photonic In-Memory Neurocomputing
abstract
Optical phase change material (PCM) has emerged promising to enable photonic in-memory neurocomputing in optical neural network (ONN) designs. However, massive photonic tensor core (PTC) reuse is required to implement large matrix multiplication due to the limited single-core scale. The resultant large number of PCM writes during inference incurs serious dynamic energy costs and overwhelms the fragile PCM with limited write endurance, causing the severe aging issue. Moreover, the aged PCM would distort the stored value and significantly degrade the reliability of PTC. In this work, we propose a holistic solution,ELight, to tackle both the aging issue and the post-aging reliability issue, where a proactive aging-aware optimization framework minimizes the overall PCM write cost and a post-aging tolerance scheme overcomes the effect of aged PCM. Specifically, in the aging-aware optimization part, we propose write-aware training to encourage the similarity among weight blocks and combine it with a post-training optimization technique to reduce programming efforts by eliminating redundant writes. Next, an efficient groupwise row-based weight-PTC remapping scheme is introduced to tolerate the reprogrammability degradation due to the aged PCM. Experiments show thatELightcan achieve over$20 \times $reductions in the total number of write operations and dynamic energy cost with comparable accuracy. Moreover,ELightcan guarantee significant accuracy recovery under the aged PCM within photonic memories. With ourELight, photonic in-memory neurocomputing will step forward toward practical applications in machine learning with order-of-magnitude longer lifetime, lower programming energy cost, and significant resilience against PCM aging effects.
Hanqing Zhu, Jiaqi Gu 0002, Chenghao Feng, Zixuan Jiang, Ray T. Chen, David Z. Pan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2022 ELight: Enabling Efficient Photonic In-Memory Neurocomputing with Life Enhancement
abstract
With the recent advances in optical phase change material (PCM), photonic in-memory neurocomputing has demonstrated its superiority in optical neural network (ONN) designs with near-zero static power consumption, time-of-light latency, and compact footprint. However, photonic tensor cores require massive hardware reuse to implement large matrix multiplication due to the limited single-core scale. The resultant large number of PCM writes leads to serious dynamic power and overwhelms the fragile PCM with limited write endurance. In this work, we propose a synergistic optimization framework, ELight, to minimize the overall write efforts for efficient and reliable optical in-memory neurocomputing. We first propose write-aware training to encourage the similarity among weight blocks, and combine it with a post-training optimization method to reduce programming efforts by eliminating redundant writes. Experiments show that ELight can achieve over$20\times$reduction in the total number of writes and dynamic power with comparable accuracy. With our ELight, photonic in-memory neurocomputing will step forward towards viable applications in machine learning with preserved accuracy, order-of-magnitude longer lifetime, and lower programming energy.
Hanqing Zhu, Jiaqi Gu 0002, Chenghao Feng, Zixuan Jiang, Ray T. Chen, David Z. Pan
ASP-DAC6
2022 ADEPT: automatic differentiable DEsign of photonic tensor cores
abstract
Photonic tensor cores (PTCs) are essential building blocks for optical artificial intelligence (AI) accelerators based on programmable photonic integrated circuits. PTCs can achieve ultra-fast and efficient tensor operations for neural network (NN) acceleration. Current PTC designs are either manually constructed or based on matrix decomposition theory, which lacks the adaptability to meet various hardware constraints and device specifications. To our best knowledge, automatic PTC design methodology is still unexplored. It will be promising to move beyond the manual design paradigm and "nurture" photonic neurocomputing with AI and design automation. Therefore, in this work, for the first time, we propose a fully differentiable framework, dubbed ADEPT, that can efficiently search PTC designs adaptive to various circuit footprint constraints and foundry PDKs. Extensive experiments show superior flexibility and effectiveness of the proposed ADEPT framework to explore a large PTC design space. On various NN models and benchmarks, our searched PTC topology outperforms prior manually-designed structures with competitive matrix representability, 2×-30× higher footprint compactness, and better noise robustness, demonstrating a new paradigm in photonic neural chip design. The code of ADEPT is available at link using the TorchONN library.
Jiaqi Gu 0002, Hanqing Zhu, Chenghao Feng, Zixuan Jiang, Ray T. Chen, David Z. Pan
DAC7
2022 Fuse and Mix: MACAM-Enabled Analog Activation for Energy-Efficient Neural Acceleration
abstract
Analog computing has been recognized as a promising low-power alternative to digital counterparts for neural network acceleration. However, conventional analog computing is mainly in a mixed-signal manner. Tedious analog/digital (A/D) conversion cost significantly limits the overall system's energy efficiency. In this work, we devise an efficient analog activation unit with magnetic tunnel junction (MTJ)-based analog content-addressable memory (MACAM), simultaneously realizing nonlinear activation and A/D conversion in a fused fashion. To compensate for the nascent and therefore currently limited representation capability of MACAM, we propose to mix our analog activation unit with digital activation dataflow. A fully differential framework, SuperMixer, is developed to search for an optimized activation workload assignment, adaptive to various activation energy constraints. The effectiveness of our proposed methods is evaluated on a silicon photonic accelerator. Compared to standard activation implementation, our mixed activation system with the searched assignment can achieve competitive accuracy with >60% energy saving on A/D conversion and activation.
Hanqing Zhu, Keren Zhu 0001, Jiaqi Gu 0002, Harrison Jin, Ray T. Chen, Jean Anne C. Incorvia, David Z. Pan
ICCAD5
2022 NeurOLight: A Physics-Agnostic Neural Operator Enabling Parametric Photonic Device Simulation
abstract
Optical computing has become emerging technology in next-generation efficient artificial intelligence (AI) due to its ultra-high speed and efficiency. Electromagnetic field simulation is critical to the design, optimization, and validation of photonic devices and circuits.However, costly numerical simulation significantly hinders the scalability and turn-around time in the photonic circuit design loop. Recently, physics-informed neural networks were proposed to predict the optical field solution of a single instance of a partial differential equation (PDE) with predefined parameters. Their complicated PDE formulation and lack of efficient parametrization mechanism limit their flexibility and generalization in practical simulation scenarios. In this work, for the first time, a physics-agnostic neural operator-based framework, dubbed NeurOLight, is proposed to learn a family of frequency-domain Maxwell PDEs for ultra-fast parametric photonic device simulation. Specifically, we discretize different devices into a unified domain, represent parametric PDEs with a compact wave prior, and encode the incident light via masked source modeling. We design our model to have parameter-efficient cross-shaped NeurOLight blocks and adopt superposition-based augmentation for data-efficient learning. With those synergistic approaches, NeurOLight demonstrates 2-orders-of-magnitude faster simulation speed than numerical solvers and outperforms prior NN-based models by ~54% lower prediction error using ~44% fewer parameters.
Jiaqi Gu 0002, Zhengqi Gao, Chenghao Feng, Hanqing Zhu, Ray T. Chen, Duane S. Boning, David Z. Pan
NeurIPS5
2021 Efficient On-Chip Learning for Optical Neural Networks Through Power-Aware Sparse Zeroth-Order Optimization
abstract
Optical neural networks (ONNs) have demonstrated record-breaking potential in high-performance neuromorphic computing due to their ultra-high execution speed and low energy consumption. However, current learning protocols fail to provide scalable and efficient solutions to photonic circuit optimization in practical applications. In this work, we propose a novel on-chip learning framework to release the full potential of ONNs for power-efficient in situ training. Instead of deploying implementation-costly back-propagation, we directly optimize the device configurations with computation budgets and power constraints. We are the first to model the ONN on-chip learning as a resource-constrained stochastic noisy zeroth-order optimization problem, and propose a novel mixed-training strategy with two-level sparsity and power-aware dynamic pruning to offer a scalable on-chip training solution in practical ONN deployment. Compared with previous methods, we are the first to optimize over 2,500 optical components on chip. We can achieve much better optimization stability, 3.7x-7.6x higher efficiency, and save >90% power under practical device variations and thermal crosstalk.
Jiaqi Gu 0002, Chenghao Feng, Zheng Zhao 0003, Zhoufeng Ying, Ray T. Chen, David Z. Pan
AAAI5
2021 SqueezeLight: Towards Scalable Optical Neural Networks with Multi-Operand Ring Resonators
abstract
Optical neural networks (ONNs) have demonstrated promising potentials for next-generation artificial intelligence acceleration with ultra-low latency, high bandwidth, and low energy consumption. However, due to high area cost and lack of efficient sparsity exploitation, previous ONN designs fail to provide scalable and efficient neuromorphic computing, which hinders the practical implementation of photonic neural accelerators. In this work, we propose a novel design methodology to enable a more scalable ONN architecture. We propose a nonlinear optical neuron based on multi-operand ring resonators to achieve neuromorphic computing with a compact footprint, low wavelength usage, learnable neuron balancing, and built-in nonlinearity. The structured sparsity is exploited to support more efficient ONN engines via a fine-grained structured pruning technique. A robustness-aware learning method is adopted to guarantee the variation-tolerance of our ONN. Simulation and experimental results show that the proposed ONN achieves one-order-of-magnitude improvement in compactness and efficiency over previous designs with high fidelity and robustness.
Jiaqi Gu 0002, Chenghao Feng, Zheng Zhao 0003, Zhoufeng Ying, Ray T. Chen, David Z. Pan
DATE6
2021 O2NN: Optical Neural Networks with Differential Detection-Enabled Optical Operands
abstract
Optical neuromorphic computing has demonstrated promising performance with ultra-high computation speed, high bandwidth, and low energy consumption. The traditional optical neural network (ONN) architectures realize neuromorphic computing via electrical weight encoding. However, previous ONN design methodologies can only handle static linear projection with stationary synaptic weights, thus fail to support efficient and flexible computing when both operands are dynamically-encoded light signals. In this work, we propose a novel ONN engine O2NN based on wavelength-division multiplexing and differential detection to enable high-performance, robust, and versatile photonic neural computing with both light operands. Balanced optical weights and augmented quantization are introduced to enhance the representability and efficiency of our architecture. Static and dynamic variations are discussed in detail with a knowledge-distillation-based solution given for robustness improvement. Discussions on hardware cost and efficiency are provided for a comprehensive comparison with prior work. Simulation and experimental results show that the proposed ONN architecture provides flexible, efficient, and robust support for high-performance photonic neural computing with fully-optical operands under low-bit quantization and practical variations.
Jiaqi Gu 0002, Zheng Zhao 0003, Chenghao Feng, Zhoufeng Ying, Ray T. Chen, David Z. Pan
DATE5
2021 Towards Memory-Efficient Neural Networks via Multi-Level in situ Generation
abstract
Deep neural networks (DNN) have shown superior performance in a variety of tasks. As they rapidly evolve, their escalating computation and memory demands make it challenging to deploy them on resource-constrained edge devices. Though extensive efficient accelerator designs, from traditional electronics to emerging photonics, have been successfully demonstrated, they are still bottlenecked by expensive memory accesses due to tremendous gaps between the bandwidth/power/latency of electrical memory and computing cores. Previous solutions fail to fully-leverage the ultra-fast computational speed of emerging DNN accelerators to break through the critical memory bound. In this work, we propose a general and unified framework to trade expensive memory transactions with ultra-fast on-chip computations, directly translating to performance improvement. We are the first to jointly explore the intrinsic correlations and bit-level redundancy within DNN kernels and propose a multi-level in situ generation mechanism with mixed-precision bases to achieve on-the-fly recovery of high-resolution parameters with minimum hardware overhead. Extensive experiments demonstrate that our proposed joint method can boost the memory efficiency by 10-20× with comparable accuracy over four state-of-the-art designs, when benchmarked on ResNet-18/DenseNet-121/MobileNetV2/V3 with various tasks.
Jiaqi Gu 0002, Hanqing Zhu, Chenghao Feng, Zixuan Jiang, Ray T. Chen, David Z. Pan
ICCV6
2021 L2ight: Enabling On-Chip Learning for Optical Neural Networks via Efficient in-situ Subspace Optimization
abstract
Silicon-photonics-based optical neural network (ONN) is a promising hardware platform that could represent a paradigm shift in efficient AI with its CMOS-compatibility, flexibility, ultra-low execution latency, and high energy efficiency. In-situ training on the online programmable photonic chips is appealing but still encounters challenging issues in on-chip implementability, scalability, and efficiency. In this work, we propose a closed-loop ONN on-chip learning framework L2ight to enable scalable ONN mapping and efficient in-situ learning. L2ight adopts a three-stage learning flow that first calibrates the complicated photonic circuit states under challenging physical constraints, then performs photonic core mapping via combined analytical solving and zeroth-order optimization. A subspace learning procedure with multi-level sparsity is integrated into L2ight to enable in-situ gradient evaluation and fast adaptation, unleashing the power of optics for real on-chip intelligence. Extensive experiments demonstrate our proposed L2ight outperforms prior ONN training protocols with 3-order-of-magnitude higher scalability and over 30x better efficiency, when benchmarked on various models and learning tasks. This synergistic framework is the first scalable on-chip learning solution that pushes this emerging field from intractable to scalable and further to efficient for next-generation self-learnable photonic neural chips. From a co-design perspective, L2ight also provides essential insights for hardware-restricted unitary subspace optimization and efficient sparse training. We open-source our framework at the link.
Jiaqi Gu 0002, Hanqing Zhu, Chenghao Feng, Zixuan Jiang, Ray T. Chen, David Z. Pan
NeurIPS5
2021 Toward Hardware-Efficient Optical Neural Networks: Beyond FFT Architecture via Joint Learnability
abstract
As a promising neuromorphic framework, the optical neural network (ONN) demonstrates ultrahigh inference speed with low energy consumption. However, the previous ONN architectures have high area overhead which limits their practicality. In this article, we propose an area-efficient ONN architecture based on structured neural networks, leveraging optical fast Fourier transform for efficient computation. A two-phase software training flow with structured pruning is proposed to further reduce the optical component utilization. Experimental results demonstrate that the proposed architecture can achieve 2.2- 3.7× area cost improvement compared with the previous singular value decomposition-based architecture with comparable inference accuracy. A novel optical microdisk-based convolutional neural network architecture with joint learnability is proposed as an extension to move beyond Fourier transform and multilayer perception, enabling hardware-aware ONN design space exploration with lower area cost, higher power efficiency, and better noise-robustness.
Jiaqi Gu 0002, Zheng Zhao 0003, Chenghao Feng, Zhoufeng Ying, Ray T. Chen, David Z. Pan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2020 Towards Area-Efficient Optical Neural Networks: An FFT-based Architecture
abstract
As a promising neuromorphic framework, the optical neural network (ONN) demonstrates ultra-high inference speed with low energy consumption. However, the previous ONN architectures have high area overhead which limits their practicality. In this paper, we propose an area-efficient ONN architecture based on structured neural networks, leveraging optical fast Fourier transform for efficient computation. A two-phase software training flow with structured pruning is proposed to further reduce the optical component utilization. Experimental results demonstrate that the proposed architecture can achieve 2.2~3.7× area cost improvement compared with the previous singular value decomposition-based architecture with comparable inference accuracy.
Jiaqi Gu 0002, Zheng Zhao 0003, Chenghao Feng, Ray T. Chen, David Z. Pan
ASP-DAC5
2020 FLOPS: EFficient On-Chip Learning for OPtical Neural Networks Through Stochastic Zeroth-Order Optimization
abstract
Optical neural networks (ONNs) have attracted extensive attention due to its ultra-high execution speed and low energy consumption. The traditional software-based ONN training, however, suffers the problems of expensive hardware mapping and inaccurate variation modeling while the current on-chip training methods fail to leverage the self-learning capability of ONNs due to algorithmic inefficiency and poor variation- robustness. In this work, we propose an on-chip learning method to resolve the aforementioned problems that impede ONNs' full potential for ultra-fast forward acceleration. We directly optimize optical components using stochastic zeroth-order optimization on-chip, avoiding the traditional high-overhead back-propagation, matrix decomposition, or in situ devicelevel intensity measurements. Experimental results demonstrate that the proposed on-chip learning framework provides an efficient solution to train integrated ONNs with 3~4× fewer ONN forward, higher inference accuracy, and better variation-robustness than previous works.
Jiaqi Gu 0002, Zheng Zhao 0003, Chenghao Feng, Wuxi Li, Ray T. Chen, David Z. Pan
DAC5
2020 ROQ: A Noise-Aware Quantization Scheme Towards Robust Optical Neural Networks with Low-bit Controls
abstract
Optical neural networks (ONNs) demonstrate orders-of-magnitude higher speed in deep learning acceleration than their electronic counterparts. However, limited control precision and device variations induce accuracy degradation in practical ONN implementations. To tackle this issue, we propose a quantization scheme that adapts a full-precision ONN to low-resolution voltage controls. Moreover, we propose a protective regularization technique that dynamically penalizes quantized weights based on their estimated noise-robustness, leading to an improvement in noise robustness. Experimental results show that the proposed scheme effectively adapts ONNs to limited-precision controls and device variations. The resultant four-layer ONN demonstrates higher inference accuracy with lower variances than baseline methods under various control precisions and device noises.
Jiaqi Gu 0002, Zheng Zhao 0003, Chenghao Feng, Hanqing Zhu, Ray T. Chen, David Z. Pan
DATE5
2019 Hardware-software co-design of slimmed optical neural networks
abstract
Optical neural network (ONN) is a neuromorphic computing hardware based on optical components. Since its first on-chip experimental demonstration, it has attracted more and more research interests due to the advantages of ultra-high speed inference with low power consumption. In this work, we design a novel slimmed architecture for realizing optical neural network considering both its software and hardware implementations. Different from the originally proposed ONN architecture based on singular value decomposition which results in two implementation-expensive unitary matrices, we show a more area-efficient architecture which uses a sparse tree network block, a single unitary block and a diagonal block for each neural network layer. In the experiments, we demonstrate that by leveraging the training engine, we are able to find a comparable accuracy to that of the previous architecture, which brings about the flexibility of using the slimmed implementation. The area cost in terms of the Mach-Zehnder interferometers, the core optical components of ONN, is 15%-38% less for various sizes of optical neural networks.
Zheng Zhao 0003, Derong Liu 0002, Meng Li 0004, Zhoufeng Ying, Biying Xu, Bei Yu 0001, Ray T. Chen, David Z. Pan
ASP-DAC8
2019 Exploiting Wavelength Division Multiplexing for Optical Logic Synthesis
abstract
Photonic integrated circuit (PIC), as a promising alternative to traditional CMOS circuit, has demonstrated the potential to accomplish on-chip optical interconnects and computations in ultra-high speed and/or low power consumption. Wavelength division multiplexing (WDM) is widely used in optical communication for enabling multiple signals being processed and transferred independently. In this work, we apply WDM to optical logic PIC synthesis to reduce the PIC area.
Zheng Zhao 0003, Derong Liu 0002, Zhoufeng Ying, Biying Xu, Chenghao Feng, Ray T. Chen, David Z. Pan
DATE6
2019 Design Technology for Scalable and Robust Photonic Integrated Circuits: Invited Paper
abstract
Photonic integrated circuit (PIC), as a promising alternative to traditional CMOS circuit, has demonstrated the potential to accomplish on-chip optical signal transmission and computations in ultra-high speed and/or low power consumption. One of the critical challenges of PIC, however, is that its scalability and robustness are limited by cascaded optical power loss and noise error. In this paper, we analyze the scalability and noise robustness challenges facing photonic integrated circuits, for two representative PIC applications: logic computing and neural networks. Automated design algorithms and learning methodologies are proposed to resolve these issues.
Zheng Zhao 0003, Jiaqi Gu 0002, Zhoufeng Ying, Chenghao Feng, Ray T. Chen, David Z. Pan
ICCAD5
2018 Logic synthesis for energy-efficient photonic integrated circuits
abstract
The development of photonic integrated circuits (PICs) has made it possible to accomplish on-chip optical interconnects and computations. As a promising alternative to traditional CMOS circuits, optics has demonstrated the ability to realize ultra-high speed and low-power information processing and communications. In this work, we propose a logic synthesis methodology for PICs. For the first time, practical issues including the insertion losses from optical combiners and switches are considered. Two optimization techniques based on binary decision diagram, combiner elimination and coupler assignment, are proposed to improve the power efficiency for PICs. Experimental results of MCNC and IWLS combinational benchmarks showed our method could efficiently generate quality PICs with a 27.02X better optical power efficiency on average, and greatly reduce the optical power depletion and facilitate large-scale on-chip optical computation.
Zheng Zhao 0003, Zheng Wang 0036, Zhoufeng Ying, Shounak Dhar, Ray T. Chen, David Z. Pan
ASP-DAC5
2018 OPERON: optical-electrical power-efficient route synthesis for on-chip signals
abstract
As VLSI technology scales to deep sub-micron, optical interconnect becomes an attractive alternative for on-chip communication. The traditional optical routing works mainly optimize the path loss, and few works explicitly exploit the optical-electrical co-design of on-chip interconnects. To overcome these limitations, we present an efficient framework that directs the hybrid optical and electrical routes with a global view of power optimization. In this framework, on-chip signal bits are processed as hyper nets; the combination of optical and electrical routes are designed for hyper nets; then a formulation is given to find the appropriate solution of each hyper net and follows a speed-up algorithm; a min-cost max-flow network is utilized to reduce the consumed optical waveguides. Experimental results demonstrate the effectiveness of the proposed framework.
Derong Liu 0002, Zheng Zhao 0003, Zheng Wang 0036, Zhoufeng Ying, Ray T. Chen, David Z. Pan
DAC5
2009 O-Router: an optical routing framework for low power on-chip silicon nano-photonic integration
abstract
In this work, we present a new optical routing framework, O-Router for future low-power on-chip optical interconnect integration utilizing silicon compatible nano-photonic devices. We formulate the optical layer routing problem as the minimization of total on-chip optical modulator cost (laser power consumption) with Integer Linear Programming technique under various detection constraints. Key techniques for variable number reduction and routing speedup are also explored and utilized. O-Router is tested on optical netlist benchmarks modified from top global nets of ISPD98/08 routing benchmarks. O-Router experimental results are compared with conventional minimum spanning tree algorithm, demonstrating an average of over 50% improvement in terms of total on-chip optical layer power reduction.
Duo Ding, Haiyu Huang 0001, Ray T. Chen, David Z. Pan
DAC4
2007 Optical interconnects: a viable solution for interconnection beyond 10 gbit/sec
abstract
The speed and complexity of integrated circuits are increasing rapidly as integrated circuit technology advances from very-large-scale integrated (VLSI) circuits to ultra-large-scale integrated (ULSI) circuits. As the number of devices per chip, the number of chips per board, the modulation speed, and the degree of integration continue to increase, electrical interconnects are facing their fundamental bottlenecks, such as speed, packaging, fan-out, and power dissipation. In the quest for high-density packaging of electronic circuits, the construction of multichip modules (MCM), which decrease the surface area by removing package walls between chips, improved signal integrity by shortening interconnection distances and removing impedance problems and capacitances. The employment of copper and materials with lower dielectric constant materials can release the bottleneck in a chip level for the next several years. The International Technology Roadmap for Semiconductors (ITRS) expects that on-chip local clock speed will constantly increase to 10 GHz by the year 2011. Electrical interconnects operating at a high-frequency region have many problems to be solved, such as crosstalk, impedance matching, power dissipation, skew, and packing density. Optical interconnection has several advantages, such as immunity to the electromagnetic interference, independency to impedance mismatch, less power consumption, and high-speed operation. Although the optical interconnects have great advantages compared with the copper/low K interconnection, they still have some difficulties regarding packaging, multilayer technology, signal tapping, and reworkability. In this presentation, the progress of optical interconnect for intra and inter-board levels will be presented with both passive and active components suitable for system integration including thin film planar waveguides, vertical cavity surface emitting lasers (VCSELs), PIN photodiode array and silicon nano-photonic crystal waveguide modulators.
Ray T. Chen
ISPD1
2000 Fully embedded board-level guided-wave optoelectronic interconnects
abstract
A fully embedded board-level guided-wave optical interconnection is presented to solve the packaging compatibility problem. All elements involved in providing high-speed optical communications within one board are demonstrated. Experimental results on a 12-channel linear array of thin-film polyimide waveguides, vertical-cavity surface-emitting lasers (VCSELs) (42 /spl mu/m), and silicon MSM photodetectors (10 /spl mu/m) suitable for a fully embedded implementation are provided. Two types of waveguide couplers, titled gratings and 45/spl deg/ total internal reflection mirrors, are fabricated within the polyimide waveguides. Thirty-five to near 100% coupling efficiencies are experimentally confirmed. By doing so, all the real estate of the PC board surface are occupied by electronics, and therefore one only observes the performance enhancement due to the employment of optical interconnection but does not worry about the interface problem between electronic and optoelectronic components unlike conventional approaches. A high speed 1-48 optical clock signal distribution network for Cray T-90 super computer is demonstrated. A waveguide propagation loss of 0.21 dB/cm at 850 nm was experimentally confirmed for the 1-48 clock signal distribution and for point-to-point interconnects. The feasibility of using polyimide as the interlayer dielectric material to form hybrid three-dimensional interconnects is also demonstrated. Finally, a waveguide bus architecture is presented, which provides a realistic bidirectional broadcasting transmission of optical signals. Such a structure is equivalent to such IEEE standard bus protocols as VME bus and FutureBus.
Ray T. Chen, Chulchae Choi, Yujie J. Liu, Bipin Bihari, Suning Tang, R. Wickman, B. Picor, Mary K. Hibbs-Brenner, J. Bristow, Y. S. Liu
Proc. IEEE1