EDBT 2026 Demo / reviewers in the wild / expert
Hanqing Zhu
dblp:164/8690
· DBLP profile ↗
22ranked-venue papers
9as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VideoLifter: Lifting Videos to 3D with Fast and Efficient Hierarchical Stereo AlignmentabstractEfficiently reconstructing 3D scenes from monocular video remains a core challenge in computer vision, vital for applications in virtual reality, robotics, and scene understanding. Recently, frame-by-frame progressive reconstruction without camera poses is commonly adopted, incurring high computational overhead and compounding errors when scaling to longer videos. To overcome these issues, we introduce VideoLifter, a novel video-to-3D pipeline that leverages a local-to-global strategy on a fragment basis, achieving both extreme efficiency and SOTA quality. Locally, VideoLifter leverages learnable 3D priors to register fragments, extracting essential information for subsequent 3D Gaussian initialization with enforced inter-fragment consistency and optimized efficiency. Globally, it employs a tree-based hierarchical merging method with key frame guidance for inter-fragment alignment, pairwise merging with Gaussian point pruning, and subsequent joint optimization to ensure global consistency while efficiently mitigating cumulative errors. This approach significantly accelerates the reconstruction process, reducing training time by over 82 % while achieving better visual quality than SOTA methods. Wenyan Cong, Hanqing Zhu, Jiahui Lei, Colton Stearns, Yuanhao Cai, Dilin Wang, Matt Feiszli, Leonidas J. Guibas, Zhangyang Wang, Weiyao Wang 0001, Zhiwen Fan |
3DV | 2 |
| 2026 | ENLighten: Lighten the Transformer, Enable Efficient Optical AccelerationabstractPhotonic computing has emerged as a promising substrate for accelerating the dense linear-algebra operations at the heart of AI, but its adoption for large Transformer models remains in its infancy. In supporting these massive models, we identify two key bottlenecks: (1) costly electro-optic conversions and data-movement overheads that erode energy efficiency as model sizes scale; (2) a mismatch between limited on-chip photonic resources and the scale of Transformer workloads, which forces frequent reuse of photonic tensor cores and dilutes throughput gains. To address these challenges, we introduce a hardware–software co-design framework. First, we propose Lighten, a PTC-aware compression flow that post-hoc decomposes each Transformer weight matrix into a low-rank component plus a structured sparse component aligned to photonic tensor-core granularity, all without lengthy retraining. Second, we present ENLighten, a reconfigurable photonic accelerator architecture featuring dynamically adaptive tensor cores, driven by a broadband light redistribution, for fine-grained sparsity support and full power gating of inactive parts. On ImageNet, Lighten prunes Base-scale vision transformer by 50% with a ~1% accuracy drop after only 3 epochs of fine-tuning within 1 hour, and when deployed on ENLighten it achieves a $2.5 \times$ improvement in energy-delay product over the state-of-the-art photonic Transformer accelerator. Hanqing Zhu, Zhican Zhou, Shupeng Ning, Xuhao Wu, Ray T. Chen, Yating Wan, David Z. Pan |
ASP-DAC | 1 |
| 2025 | INSIGHT: A Universal Neural Simulator Framework for Analog Circuits with Autoregressive TransformersabstractThe compute-intensive nature of SPICE simulations hinders effective analog design automation. This paper introduces INSIGHT, a data-efficient, adaptive, high-fidelity, technologyagnostic universal neural simulator framework that formulates analog performance prediction as an autoregressive sequence generation task to accurately predict performance across diverse circuits. INSIGHT achieves test $\mathbf{R}^{\mathbf{2}}$ scores $\geq \mathbf{0. 9 5}$, outperforming existing neural surrogates. Cross-technology transfer learning experiments show that INSIGHT can preserve model performance with $\sim \mathbf{6 0 \%}$ less training data. Low-Rank Adaptation (LoRA) integration further reduces memory footprint by $\sim 42 \%$ and training time by $\sim 25 \%$, maintaining high performance. Our experiments show that INSIGHT-based RL sizing framework achieves $100-1000 \times$ lower simulation costs over existing sizing methods for identical benchmarks and target specifications. Souradip Poddar, Yao Lai, Hanqing Zhu, Bosun Hwang, David Z. Pan |
DAC | 4 |
| 2025 | PPAAS: PVT and Pareto Aware Analog Sizing via Goal-conditioned Reinforcement LearningabstractDevice sizing is a critical yet challenging step in analog and mixed-signal circuit design, requiring careful optimization to meet diverse performance specifications. This challenge is further amplified under process, voltage, and temperature (PVT) variations, which cause circuit behavior to shift across different corners. While reinforcement learning (RL) has shown promise in automating sizing for fixed targets, training a generalized policy that can adapt to a wide range of design specifications under PVT variations requires much more training samples and resources. To address these challenges, we propose a Goal-conditioned RL framework that enables efficient policy training for analog device sizing across PVT corners, with strong generalization capability. To improve sample efficiency, we introduce Pareto-front Dominance Goal Sampling, which constructs an automatic curriculum by sampling goals from the Pareto frontier of previously achieved goals. This strategy is further enhanced by integrating Conservative Hindsight Experience Replay to stabilize training and accelerate convergence. To reduce simulation overhead, our framework incorporates a Skip-on-Fail simulation strategy. Experiments on benchmark circuits demonstrate ∼1.6× improvement in sample efficiency and ∼4.1× improvement in simulation efficiency compared to existing sizing methods. Code and benchmarks are publicly available HERE. Seunggeun Kim, Ziyi Wang 0010, Sungyoung Lee 0004, Hanqing Zhu, Doyun Kim, David Z. Pan |
ICCAD | 5 |
| 2024 | Lightening-Transformer: A Dynamically-Operated Optically-Interconnected Photonic Transformer AcceleratorabstractThe wide adoption and significant computing resource cost of attention-based transformers, e.g., Vision Transformers and large language models, have driven the demand for efficient hardware accelerators. While electronic accelerators have been commonly used, there is a growing interest in exploring photonics as an alternative technology due to its high energy efficiency and ultra-fast processing speed. Photonic accelerators have demonstrated promising results for convolutional neural networks (CNNs) workloads, which predominantly rely on weight-static linear operations. However, they encounter challenges when it comes to efficiently supporting attention-based Transformer architectures, raising questions about the applicability of photonics to advanced machine-learning tasks. The primary hurdle lies in their inefficiency in handling the unique workloads inherent to Transformers, i.e., dynamic and full-range tensor multiplication. In this work, we propose Lightening-Transformer, the first light-empowered, high-performance, and energy-efficient photonic Transformer accelerator. To overcome the fundamental limitation of existing photonic tensor core designs, we introduce a novel dynamically-operated photonic tensor core, DPTC, consisting of a crossbar array of interference-based optical vector dot-product engines, supporting highly parallel, dynamic, and full-range matrix multiplication. Furthermore, we design a dedicated accelerator that integrates our novel photonic computing cores with photonic interconnects for inter-core data broadcast, fully unleashing the power of optics. The comprehensive evaluation demonstrates that Lightening-Transformer achieves >2.6x energy and > 12 x latency reductions compared to prior photonic accelerators and delivers the lowest energy cost and 2 to 3 orders of magnitude lower energy-delay product compared to the electronic Transformer accelerator, all while maintaining digital-comparable accuracy. Our work highlights the immense potential of photonics for efficient hardware accelerators, particularly for advanced machine-learning workloads, such as Transformer-backboned large language models (LLM). Our implementation is available at https://github.com/zhuhanqing/Lightening-Transformer. Hanqing Zhu, Jiaqi Gu 0002, Hanrui Wang 0002, Zixuan Jiang, Zhekai Zhang, Rongxing Tang, Chenghao Feng, Song Han 0003, Ray T. Chen, David Z. Pan |
HPCA | 1 |
| 2024 | Inductive Modeling for Realtime Cold Start RecommendationsabstractIn recommendation systems, the timely delivery of new content to their relevant audiences is critical for generating a growing and high quality collection of content for all users. The nature of this problem requires retrieval models to be able to make inferences in real time and with high relevance. There are two specific challenges for cold start contents. First, the information loss problem in a standard Two Tower model, due to the limited feature interactions between the user and item towers, is exacerbated for cold start items due to training data sparsity. Second, the huge volume of user-generated content in industry applications today poses a big bottleneck in the end-to-end latency of recommending new content. To overcome the two challenges, we propose a novel architecture, the Item History Model (IHM). IHM directly injects user-interaction information into the item tower to overcome information loss. In addition, IHM incorporates an inductive structure using attention-based pooling to eliminate the need for recurring training, a key bottleneck for the real-timeness. On both public and industry datasets, we demonstrate that IHM can not only outperform baselines in recommending cold start contents, but also achieves SoTA real-timeness in industry applications. Chandler Zuo, Jonathan Castaldo, Hanqing Zhu, Ji Liu 0002, Yangpeng Ou |
KDD | 3 |
| 2024 | PACE: Pacing Operator Learning to Accurate Optical Field Simulation for Complicated Photonic DevicesabstractElectromagnetic field simulation is central to designing, optimizing, and validating photonic devices and circuits.
However, costly computation associated with numerical simulation poses a significant bottleneck, hindering scalability and turnaround time in the photonic circuit design process.
Neural operators offer a promising alternative, but existing SOTA approaches, Neurolight, struggle with predicting high-fidelity fields for real-world complicated photonic devices, with the best reported 0.38 normalized mean absolute error in Neurolight.
The interplays of highly complex light-matter interaction, e.g., scattering and resonance, sensitivity to local structure details, non-uniform learning complexity for full-domain simulation, and rich frequency information, contribute to the failure of existing neural PDE solvers.
In this work, we boost the prediction fidelity to an unprecedented level for simulating complex photonic devices with a novel operator design driven by the above challenges.
We propose a novel cross-axis factorized PACE operator with a strong long-distance modeling capacity to connect the full-domain complex field pattern with local device structures.
Inspired by human learning, we further divide and conquer the simulation task for extremely hard cases into two progressively easy tasks, with a first-stage model learning an initial solution refined by a second model.
On various complicated photonic device benchmarks, we demonstrate one sole PACE model is capable of achieving 73% lower error with 50% fewer parameters compared with various recent ML for PDE solvers.
The two-stage setup further advances high-fidelity simulation for even more intricate cases.
In terms of runtime,
PACE demonstrates 154-577x and 11.8-12x simulation speedup over numerical solver using scipy or highly-optimized pardiso solver, respectively.
We open-sourced the code and *complicated* optical device dataset at [PACE-Light](https://github.com/zhuhanqing/PACE-Light). Hanqing Zhu, Wenyan Cong, Guojin Chen, Shupeng Ning, Ray T. Chen, Jiaqi Gu 0002, David Z. Pan |
NeurIPS | 1 |
| 2023 | A High Resolution SAR Imaging Method for Moving Target Based on Range Doppler and Particle Swarm Optimization AlgorithmabstractSynthetic aperture radar (SAR) imaging for moving target can obtain complete situational awareness information of the detection area, and can realize the monitoring and control for moving target in the region of interest, which has important military and civilian dual-use value. However, due to the complex motion of target, the processing results of the existing SAR imaging methods severly defocused. In this paper, a high resolution SAR imaging method for moving target is proposed. First, we eliminate the coupling induced by linear range cell migration (RCM) by keystone transform. Then, the particle swarm optimization algorithm (PSO) is utilized to estimate the Doppler frequency rate, which can solve the problem of Doppler frequency rate mismatching when azimuth compression. Simulation results verifies the effectiveness of the proposed method. Dajiang Zhou, Hanqing Zhu, Yulin Huang 0001, Yongchao Zhang 0001, Jianyu Yang 0001, Qingying Yi |
IGARSS | 2 |
| 2023 | Moving Target Detection Method for Passive Radar Using LEO Communication Satellite ConstellationabstractIn recent years, many countries are actively deploying Low-Earth-Orbit (LEO) communication satellite constellations, which have the advantages of both high power flux density (PFD) on the surface of the earth and large signal bandwidth. From the perspective of radar application, these new LEO constellations are very suitable as opportunity of illuminator for target detection in passive radar systems. In this paper, the echo signal using LEO communication satellite is analyzed, and a moving target detection method is proposed. Hanqing Zhu, Dajiang Zhou, Zhongyu Li 0001, Hongyang An, Jianyu Yang 0001 |
IGARSS | 1 |
| 2023 | Pre-RMSNorm and Pre-CRMSNorm Transformers: Equivalent and Efficient Pre-LN TransformersabstractTransformers have achieved great success in machine learning applications.
Normalization techniques, such as Layer Normalization (LayerNorm, LN) and Root Mean Square Normalization (RMSNorm), play a critical role in accelerating and stabilizing the training of Transformers.
While LayerNorm recenters and rescales input vectors, RMSNorm only rescales the vectors by their RMS value.
Despite being more computationally efficient, RMSNorm may compromise the representation ability of Transformers.
There is currently no consensus regarding the preferred normalization technique, as some models employ LayerNorm while others utilize RMSNorm, especially in recent large language models.
It is challenging to convert Transformers with one normalization to the other type.
While there is an ongoing disagreement between the two normalization types,
we propose a solution to unify two mainstream Transformer architectures, Pre-LN and Pre-RMSNorm Transformers.
By removing the inherent redundant mean information in the main branch of Pre-LN Transformers, we can reduce LayerNorm to RMSNorm, achieving higher efficiency.
We further propose the Compressed RMSNorm (CRMSNorm) and Pre-CRMSNorm Transformer based on a lossless compression of the zero-mean vectors.
We formally establish the equivalence of Pre-LN, Pre-RMSNorm, and Pre-CRMSNorm Transformer variants in both training and inference.
It implies that Pre-LN Transformers can be substituted with Pre-(C)RMSNorm counterparts at almost no cost, offering the same arithmetic functionality along with free efficiency improvement.
Experiments demonstrate that we can reduce the training and inference time of Pre-LN Transformers by 1% - 10%. Zixuan Jiang, Jiaqi Gu 0002, Hanqing Zhu, David Z. Pan |
NeurIPS | 3 |
| 2023 | SqueezeLight: A Multi-Operand Ring-Based Optical Neural Network With Cross-Layer ScalabilityabstractOptical neural networks (ONNs) are promising hardware platforms for next-generation artificial intelligence acceleration with ultrafast speed and low-energy consumption. However, previous ONN designs are bounded by one multiply–accumulate operation per device, showing unsatisfying scalability. In this work, we propose a scalable ONN architecture, dubbedSqueezeLight. We propose a nonlinear optical neuron based on multioperand ring resonators (MORRs) to squeeze vector dot-product into a single device with low wavelength usage and built-in nonlinearity. A block-level squeezing technique with structured sparsity is exploited to support higher scalability. We adopt a robustness-aware training algorithm to guarantee variation tolerance. To enable a truly scalable ONN architecture, we extendSqueezeLightto a separable optical CNN architecture that further squeezes in the layer level. Two orthogonal convolutional layers are mapped to one MORR array, leading to order-of-magnitude higher software training scalability. We further explore augmented representability forSqueezeLightby introducing parametric MORR neurons with trainable nonlinearity, together with a nonlinearity-aware initialization method to stabilize convergence. Experimental results show thatSqueezeLightachieves one-order-of-magnitude better compactness and efficiency than previous designs with high fidelity, trainability, and robustness. Our open-source codes are available athttps://github.com/JeremieMelo/SqueezeLight. Jiaqi Gu 0002, Chenghao Feng, Hanqing Zhu, Zheng Zhao 0003, Zhoufeng Ying, Ray T. Chen, David Z. Pan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | ELight: Toward Efficient and Aging-Resilient Photonic In-Memory NeurocomputingabstractOptical phase change material (PCM) has emerged promising to enable photonic in-memory neurocomputing in optical neural network (ONN) designs. However, massive photonic tensor core (PTC) reuse is required to implement large matrix multiplication due to the limited single-core scale. The resultant large number of PCM writes during inference incurs serious dynamic energy costs and overwhelms the fragile PCM with limited write endurance, causing the severe aging issue. Moreover, the aged PCM would distort the stored value and significantly degrade the reliability of PTC. In this work, we propose a holistic solution,ELight, to tackle both the aging issue and the post-aging reliability issue, where a proactive aging-aware optimization framework minimizes the overall PCM write cost and a post-aging tolerance scheme overcomes the effect of aged PCM. Specifically, in the aging-aware optimization part, we propose write-aware training to encourage the similarity among weight blocks and combine it with a post-training optimization technique to reduce programming efforts by eliminating redundant writes. Next, an efficient groupwise row-based weight-PTC remapping scheme is introduced to tolerate the reprogrammability degradation due to the aged PCM. Experiments show thatELightcan achieve over$20 \times $reductions in the total number of write operations and dynamic energy cost with comparable accuracy. Moreover,ELightcan guarantee significant accuracy recovery under the aged PCM within photonic memories. With ourELight, photonic in-memory neurocomputing will step forward toward practical applications in machine learning with order-of-magnitude longer lifetime, lower programming energy cost, and significant resilience against PCM aging effects. Hanqing Zhu, Jiaqi Gu 0002, Chenghao Feng, Zixuan Jiang, Ray T. Chen, David Z. Pan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | The Design, Education and Evolution of a Robotic BabyabstractInspired by Alan Turing's idea of a child machine, in this article, we introduce the formal definition of a robotic baby, an integrated system with minimal world knowledge at birth, capable of learning incrementally and interactively, and adapting to the world. Within the definition, fundamental capabilities and system characteristics of the robotic baby are identified and presented as the system-level requirements. As a minimal viable prototype, theBabyarchitecture is proposed with a systems engineering design approach to satisfy the system-level requirements, which has been verified and validated with simulations and experiments on a robotic system. We demonstrate the capabilities of the robotic baby in natural language acquisition and semantic parsing in English and Chinese, as well as in natural language grounding, natural language reinforcement learning, natural language programming, and system introspection for explainability. The education and evolution of the robotic baby are illustrated with real-world robotic demonstrations. Inspired by the genetic inheritance in human beings, knowledge inheritance in robotic babies and its benefits regarding evolution are discussed. Hanqing Zhu, Sean Wilson, Eric Feron |
IEEE Trans. Robotics | 1 |
| 2022 | ELight: Enabling Efficient Photonic In-Memory Neurocomputing with Life EnhancementabstractWith the recent advances in optical phase change material (PCM), photonic in-memory neurocomputing has demonstrated its superiority in optical neural network (ONN) designs with near-zero static power consumption, time-of-light latency, and compact footprint. However, photonic tensor cores require massive hardware reuse to implement large matrix multiplication due to the limited single-core scale. The resultant large number of PCM writes leads to serious dynamic power and overwhelms the fragile PCM with limited write endurance. In this work, we propose a synergistic optimization framework, ELight, to minimize the overall write efforts for efficient and reliable optical in-memory neurocomputing. We first propose write-aware training to encourage the similarity among weight blocks, and combine it with a post-training optimization method to reduce programming efforts by eliminating redundant writes. Experiments show that ELight can achieve over$20\times$reduction in the total number of writes and dynamic power with comparable accuracy. With our ELight, photonic in-memory neurocomputing will step forward towards viable applications in machine learning with preserved accuracy, order-of-magnitude longer lifetime, and lower programming energy. Hanqing Zhu, Jiaqi Gu 0002, Chenghao Feng, Zixuan Jiang, Ray T. Chen, David Z. Pan |
ASP-DAC | 1 |
| 2022 | ADEPT: automatic differentiable DEsign of photonic tensor coresabstractPhotonic tensor cores (PTCs) are essential building blocks for optical artificial intelligence (AI) accelerators based on programmable photonic integrated circuits. PTCs can achieve ultra-fast and efficient tensor operations for neural network (NN) acceleration. Current PTC designs are either manually constructed or based on matrix decomposition theory, which lacks the adaptability to meet various hardware constraints and device specifications. To our best knowledge, automatic PTC design methodology is still unexplored. It will be promising to move beyond the manual design paradigm and "nurture" photonic neurocomputing with AI and design automation. Therefore, in this work, for the first time, we propose a fully differentiable framework, dubbed ADEPT, that can efficiently search PTC designs adaptive to various circuit footprint constraints and foundry PDKs. Extensive experiments show superior flexibility and effectiveness of the proposed ADEPT framework to explore a large PTC design space. On various NN models and benchmarks, our searched PTC topology outperforms prior manually-designed structures with competitive matrix representability, 2×-30× higher footprint compactness, and better noise robustness, demonstrating a new paradigm in photonic neural chip design. The code of ADEPT is available at link using the TorchONN library. Jiaqi Gu 0002, Hanqing Zhu, Chenghao Feng, Zixuan Jiang, Ray T. Chen, David Z. Pan |
DAC | 2 |
| 2022 | Fuse and Mix: MACAM-Enabled Analog Activation for Energy-Efficient Neural AccelerationabstractAnalog computing has been recognized as a promising low-power alternative to digital counterparts for neural network acceleration. However, conventional analog computing is mainly in a mixed-signal manner. Tedious analog/digital (A/D) conversion cost significantly limits the overall system's energy efficiency. In this work, we devise an efficient analog activation unit with magnetic tunnel junction (MTJ)-based analog content-addressable memory (MACAM), simultaneously realizing nonlinear activation and A/D conversion in a fused fashion. To compensate for the nascent and therefore currently limited representation capability of MACAM, we propose to mix our analog activation unit with digital activation dataflow. A fully differential framework, SuperMixer, is developed to search for an optimized activation workload assignment, adaptive to various activation energy constraints. The effectiveness of our proposed methods is evaluated on a silicon photonic accelerator. Compared to standard activation implementation, our mixed activation system with the searched assignment can achieve competitive accuracy with >60% energy saving on A/D conversion and activation. Hanqing Zhu, Keren Zhu 0001, Jiaqi Gu 0002, Harrison Jin, Ray T. Chen, Jean Anne C. Incorvia, David Z. Pan |
ICCAD | 1 |
| 2022 | Uni-Retriever: Towards Learning the Unified Embedding Based Retriever in Bing Sponsored SearchabstractEmbedding based retrieval (EBR) is a fundamental building block in many web applications. However, EBR in sponsored search is distinguished from other generic scenarios and technically challenging due to the need of serving multiple retrieval purposes: firstly, it has to retrieve high-relevance ads, which may exactly serve user's search intent; secondly, it needs to retrieve high-CTR ads so as to maximize the overall user clicks. In this paper, we present a novel representation learning framework Uni-Retriever developed for Bing Search, which unifies two different training modes knowledge distillation and contrastive learning to realize both required objectives. On one hand, the capability of making high-relevance retrieval is established by distilling knowledge from the "relevance teacher model''. On the other hand, the capability of making high-CTR retrieval is optimized by learning to discriminate user's clicked ads from the entire corpus. The two training modes are jointly performed as a multi-objective learning process, such that the ads of high relevance and CTR can be favored by the generated embeddings. Besides the learning strategy, we also elaborate our solution for EBR serving pipeline built upon the substantially optimized DiskANN, where massive-scale EBR can be performed with competitive time and memory efficiency, and accomplished in high-quality. We make comprehensive offline and online experiments to evaluate the proposed techniques, whose findings may provide useful insights for the future development of EBR systems. Uni-Retriever has been mainstreamed as the major retrieval path in Bing's production thanks to the notable improvements on the representation and EBR serving quality. Jianjin Zhang, Zheng Liu 0011, Weihao Han, Shitao Xiao, Ruicheng Zheng, Yingxia Shao, Hao Sun 0015, Hanqing Zhu, Premkumar Srinivasan, Qi Zhang 0066, Xing Xie 0001 |
KDD | 8 |
| 2022 | NeurOLight: A Physics-Agnostic Neural Operator Enabling Parametric Photonic Device SimulationabstractOptical computing has become emerging technology in next-generation efficient artificial intelligence (AI) due to its ultra-high speed and efficiency. Electromagnetic field simulation is critical to the design, optimization, and validation of photonic devices and circuits.However, costly numerical simulation significantly hinders the scalability and turn-around time in the photonic circuit design loop. Recently, physics-informed neural networks were proposed to predict the optical field solution of a single instance of a partial differential equation (PDE) with predefined parameters. Their complicated PDE formulation and lack of efficient parametrization mechanism limit their flexibility and generalization in practical simulation scenarios. In this work, for the first time, a physics-agnostic neural operator-based framework, dubbed NeurOLight, is proposed to learn a family of frequency-domain Maxwell PDEs for ultra-fast parametric photonic device simulation. Specifically, we discretize different devices into a unified domain, represent parametric PDEs with a compact wave prior, and encode the incident light via masked source modeling. We design our model to have parameter-efficient cross-shaped NeurOLight blocks and adopt superposition-based augmentation for data-efficient learning. With those synergistic approaches, NeurOLight demonstrates 2-orders-of-magnitude faster simulation speed than numerical solvers and outperforms prior NN-based models by ~54% lower prediction error using ~44% fewer parameters. Jiaqi Gu 0002, Zhengqi Gao, Chenghao Feng, Hanqing Zhu, Ray T. Chen, Duane S. Boning, David Z. Pan |
NeurIPS | 4 |
| 2021 | Towards Memory-Efficient Neural Networks via Multi-Level in situ GenerationabstractDeep neural networks (DNN) have shown superior performance in a variety of tasks. As they rapidly evolve, their escalating computation and memory demands make it challenging to deploy them on resource-constrained edge devices. Though extensive efficient accelerator designs, from traditional electronics to emerging photonics, have been successfully demonstrated, they are still bottlenecked by expensive memory accesses due to tremendous gaps between the bandwidth/power/latency of electrical memory and computing cores. Previous solutions fail to fully-leverage the ultra-fast computational speed of emerging DNN accelerators to break through the critical memory bound. In this work, we propose a general and unified framework to trade expensive memory transactions with ultra-fast on-chip computations, directly translating to performance improvement. We are the first to jointly explore the intrinsic correlations and bit-level redundancy within DNN kernels and propose a multi-level in situ generation mechanism with mixed-precision bases to achieve on-the-fly recovery of high-resolution parameters with minimum hardware overhead. Extensive experiments demonstrate that our proposed joint method can boost the memory efficiency by 10-20× with comparable accuracy over four state-of-the-art designs, when benchmarked on ResNet-18/DenseNet-121/MobileNetV2/V3 with various tasks. Jiaqi Gu 0002, Hanqing Zhu, Chenghao Feng, Zixuan Jiang, Ray T. Chen, David Z. Pan |
ICCV | 2 |
| 2021 | L2ight: Enabling On-Chip Learning for Optical Neural Networks via Efficient in-situ Subspace OptimizationabstractSilicon-photonics-based optical neural network (ONN) is a promising hardware platform that could represent a paradigm shift in efficient AI with its CMOS-compatibility, flexibility, ultra-low execution latency, and high energy efficiency. In-situ training on the online programmable photonic chips is appealing but still encounters challenging issues in on-chip implementability, scalability, and efficiency. In this work, we propose a closed-loop ONN on-chip learning framework L2ight to enable scalable ONN mapping and efficient in-situ learning. L2ight adopts a three-stage learning flow that first calibrates the complicated photonic circuit states under challenging physical constraints, then performs photonic core mapping via combined analytical solving and zeroth-order optimization. A subspace learning procedure with multi-level sparsity is integrated into L2ight to enable in-situ gradient evaluation and fast adaptation, unleashing the power of optics for real on-chip intelligence. Extensive experiments demonstrate our proposed L2ight outperforms prior ONN training protocols with 3-order-of-magnitude higher scalability and over 30x better efficiency, when benchmarked on various models and learning tasks. This synergistic framework is the first scalable on-chip learning solution that pushes this emerging field from intractable to scalable and further to efficient for next-generation self-learnable photonic neural chips. From a co-design perspective, L2ight also provides essential insights for hardware-restricted unitary subspace optimization and efficient sparse training. We open-source our framework at the link. Jiaqi Gu 0002, Hanqing Zhu, Chenghao Feng, Zixuan Jiang, Ray T. Chen, David Z. Pan |
NeurIPS | 2 |
| 2020 | ROQ: A Noise-Aware Quantization Scheme Towards Robust Optical Neural Networks with Low-bit ControlsabstractOptical neural networks (ONNs) demonstrate orders-of-magnitude higher speed in deep learning acceleration than their electronic counterparts. However, limited control precision and device variations induce accuracy degradation in practical ONN implementations. To tackle this issue, we propose a quantization scheme that adapts a full-precision ONN to low-resolution voltage controls. Moreover, we propose a protective regularization technique that dynamically penalizes quantized weights based on their estimated noise-robustness, leading to an improvement in noise robustness. Experimental results show that the proposed scheme effectively adapts ONNs to limited-precision controls and device variations. The resultant four-layer ONN demonstrates higher inference accuracy with lower variances than baseline methods under various control precisions and device noises. Jiaqi Gu 0002, Zheng Zhao 0003, Chenghao Feng, Hanqing Zhu, Ray T. Chen, David Z. Pan |
DATE | 4 |
| 2015 | MDTC: An efficient approach to TCAM-based multidimensional table compressionabstractTernary Content Addressable Memory(TCAM)-based multidimensional tables are widely used to implement Access Control Lists (ACLs) for Internet packet classification and filtering, and have also become attractive for constructing the forwarding tables of Internet routers and the flow tables of Openflow switches, where multiple fields are generally used to match incoming packets. However, as such tables can grow quickly as the Internet develops fast, and sometimes even expand in size because of TCAMs limitation in storing rules with range fields, it becomes imperative to compress these tables. In this paper, we propose a fast and efficient approach to multidimensional table compression. We divide the multidimensional space iteratively to obtain a series of cells, and then combine those cells that are associated with the same action. Our approach applies to tables of any dimension, addresses the range expansion problem, and provides efficient compression for TCAM-based tables. The experiments show that our approach has low computing cost in time, which is significant for the online update of tables. On average, it reduces 23.0% entries of the real-life ACLs, 25.8% to 55.1% of the generated two-dimension tables, 55.1% of the generated ACLs, and 28.4% of the generated Openflow flow tables. Hanqing Zhu, Mingwei Xu 0001, Qing Li 0006, Jun Li 0001, Yuan Yang 0001, Suogang Li |
Networking | 1 |