VLDB 2026 Research / reviewers in the wild / expert
Jun Feng 0008
dblp:00/4883-8
· DBLP profile ↗
15ranked-venue papers
0as first author
7since 2021 · last 2022
0000-0002-2313-7144ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 7 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | A Reliability Concern on Photonic Neural NetworksabstractEmerging integrated photonic neural networks have experimentally proved to achieve an ultra-high speedup of deep neural network training and inference in the optical domain. However, photonic devices suffer from the inherent crosstalk noise and loss, inevitably leading to reliability concerns. This paper systematically analyzes the impacts of crosstalk and loss on photonic computing systems. We propose a crosstalk-aware model for reliability estimation and find out the worst-case bounds as we increase the footprints and scales of the photonic chips. Our evaluations show that −30dB crosstalk noise can cause maximal photonic chip integration to a sharp drop by 109x. To facilitate very-large-scale photonic integration for future computing, we further propose multiple heterogeneous bijou photonic-cores to address the crosstalk-aware reliability concern. Yinyi Liu, Jun Feng 0008, Shixi Chen, Jiang Xu 0001 |
DATE | 3 |
| 2022 | Improving the thermal reliability of photonic chiplets on multicore processors
Xuanqi Chen, Jun Feng 0008, Shixi Chen, Jiang Xu 0001 |
Integr. | 3 |
| 2022 | HERO: Pbit High-Radix Optical Switch Based on Integrated Silicon Photonics for Data CenterabstractTo establish flatten networks and accomplish rapid and efficient communications in the future hyper-scale data centers, HERO, a high-radix optical switch based on integrated silicon photonics, is proposed in this work. The architecture of HERO, including the switch fabric, switch interface, and switch controller, is described in detail. Two new switch control approaches: 1) split-transaction predictive control and 2) wavelength-group switching, are developed. The efficient control together with the optimized high-radix integrated optical switch fabrics help HERO achieve over 1 Pbps switching capacity. Even for small packets, such as 64–256 B Ethernet packets, the maximal utilization of the switch can reach up to 83%, and the throughput can approximate 1 Pbps. Further design explorations on the packet length and some key configurations, including the number of wavelengths and wavelength groups, are also conducted in this work, paving the way to the design of high-performance flatten data-center networks in the future. Jun Feng 0008, Jiang Xu 0001, Xuanqi Chen, Shixi Chen, Yinyi Liu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | Fast and Accurate Statistical Simulation of Shared-Memory Applications on Multicore SystemsabstractDetailed cycle-accurate simulation of multicore systems is naturally slow. Statistical simulation is one alternative that permits trading off simulation speed for accuracy. However, there is a lack of effective memory locality models for multicore applications. Hence, existing statistical simulators neglect data-sharing between threads. Additionally, the standard method to speed up statistical simulations is to blindly reduce the trace length to be synthesized. While this gives good control over the speedup, it leaves the simulation error unbounded. In this work, we introduce a novel statistical simulation methodology for exploration of shared-memory multicore systems. It includes a newsharing-localitymodel (Shalom) that captures and reproduces data-sharing in multithread applications. Furthermore, we propose a method to bound the simulation error for a particular metric while maximizing speedup. The technique works by monitoring the convergence of the statistical synthesis. It is referred to asconvergence-deterministicsimulation (Condens). The combination ofShalomandCondensis around 130x faster than cycle-accurate simulations with reasonable accuracy loss. Our approach is also 5x faster than state-of-the-art sampling simulation under the same accuracy level. Compared to previous statistical simulators ignoring sharing, our approach is 2x more accurate for performance metrics and 8x more accurate for cache miss estimations. Fan Jiang 0015, Rafael Kioji Vivas Maeda, Jun Feng 0008, Shixi Chen, Lin Chen 0029, Xiao Li 0038, Jiang Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2021 | Multi-Core Power Management through Deep Reinforcement LearningabstractAchieving high energy efficiency is a primary design objective for multi-core systems. Dynamic voltage and frequency scaling (DVFS) is one of the most widely-adopted low-power techniques. In this paper, we present a reinforcement learning-based DVFS control approach to reduce energy consumption under user-specified performance requirements. The learning agent periodically selects the voltage and frequency level for all cores based on observations of their computation intensiveness, memory behaviors as well as synchronization among cores. Experimental results on multiple real applications show that the proposed method can achieve significant energy reduction. Zhongyuan Tian, Lin Chen 0029, Xiao Li 0038, Jun Feng 0008, Jiang Xu 0001 |
ISCAS | 4 |
| 2021 | Simultaneously Tolerate Thermal and Process Variations Through Indirect Feedback Tuning for Silicon Photonic NetworksabstractSilicon photonics is the leading candidate technology for high-speed and low-energy-consumption networks. Thermal and process variations are the two main challenges of achieving high-reliability photonic networks. Thermal variation is due to the heat issues created by application, floorplan, and environment, while process variation is caused by fabrication variability in the deposition, masking, exposition, etching, and doping. Tuning techniques are then required to overcome the impact of the variations and efficiently stabilize the performance of silicon photonic networks. We extend our previous optical switch integration model, BOSIM, to support the variation and thermal analyses. Based on device properties, we propose indirect feedback tuning (IFT) to simultaneously alleviate thermal and process variations. IFT can improve the BER of silicon photonic networks to 10-9under different variation situations. Compared to state-of-the-art techniques, IFT can achieve an up to 1.52 ×108times bit-error-rate improvement and 4.11X better heater energy efficiency. Indirect feedback does not require high-speed optical signal detection, and thus, the circuit design of IFT saves up to 61.4% of the power and 51.2% of the area compared to state-of-the-art designs. Xuanqi Chen, Jun Feng 0008, Jiang Xu 0001, Shixi Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2021 | Reduce Loss and Crosstalk in Integrated Silicon-Photonic Multistage Switching Fabrics Through Multichip PartitionabstractWith the increasing popularity of data-intensive applications in data centers, the switching fabric in the internode network becomes significant. Silicon-photonic switching fabrics have a bright future in data centers, which offer high bandwidth, high energy efficiency, and low latency. However, integrating a high radix multistage switching fabric in a single chip faces challenges. A large number of waveguide crossings on the silicon photonic die causes massive power loss and introduces a tremendous amount of crosstalk noise. In this article, we propose a chip partition optimization platform (POP), which can decrease the number of waveguide crossings and shorten the on-chip traversal distance of optical signals. Our algorithms can effectively reduce the power loss and crosstalk noise in silicon-photonic multistage switching fabrics, and help to improve the signal integrity. For example, compared with the common design, POP can achieve 33-dB improvement on average power loss, 42-dB improvement on the worst-case power loss, and 39-dB improvement on the worst-case signal to noise ratio, in a$1024\times1024$butterfly based silicon-photonic switching fabric. Zhehui Wang, Jiang Xu 0001, Jun Feng 0008, Shixi Chen, Xuanqi Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2020 | Efficient Optical Power Delivery System for Hybrid Electronic-Photonic Manycore ProcessorsabstractA lot of efforts have been devoted to optically enabled high-performance communication infrastructures for future manycore processors. Silicon photonic network promises high bandwidth, high energy efficiency and low latency. However, the ever-increasing network complexity results in high optical power demands, which stress the optical power delivery and affect delivery efficiency. Facing these challenges, we propose Ring-based Optical Active Delivery (ROAD) system, to effectively manage and efficiently deliver high optical power throughout photonic-electronic hybrid systems. Experimental results demonstrate up to 5.49X energy efficiency improvement compared to traditional design without affecting processor performance. Shixi Chen, Jiang Xu 0001, Xuanqi Chen, Jun Feng 0008, Zhongyuan Tian, Xiao Li 0038 |
DATE | 5 |
| 2020 | Modeling and Analysis of Optical Modulators Based on Free-Carrier Plasma Dispersion EffectabstractSilicon photonic networks are revolutionizing computing systems by improving the energy efficiency, bandwidth, and latency of data movements. Optical modulators, such as microresonators (MRs) and Mach–Zehnder interferometers (MZIs), are the basic building blocks of silicon photonic networks. This paper proposes a SPICE-compatible electro-optical co-simulation model, basic optical switch integration model (BOSIM), to systematically study optical modulators using PN, PIN, and metal–insulator–silicon (MIS) capacitor device technologies. BOSIM holistically models both transient and steady state properties, such as switching speed, power, transmission spectrum, area, and carrier distribution. BOSIM is validated by the measured data from eight research groups and companies. Compared to MRs, BOSIM shows MZIs are fast, with a high extinction ratio and large bandwidth but in the sacrifice of loss, energy, and area. Using a PIN diode over a PN diode can save area, but retain the loss and energy, while an MIS capacitor has the shock response of carrier distribution in a narrow range and is marginalized gradually. For instance, an MZI can achieve a$2.5 {\times }$bit rate,$6.06{\times }$extinction ratio,$71.04 {\times }\,\,3$-dB bandwidth, but costs at least$1.93 {\times }$passing loss,$1.46 {\times }$energy consumption, and$16.67 {\times }$area, compared with MR. Xuanqi Chen, Yi-Shing Chang, Jiang Xu 0001, Jun Feng 0008, Peng Yang 0003, Zhehui Wang, Luan H. K. Duong |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2020 | CAMON: Low-Cost Silicon Photonic Chiplet for Manycore ProcessorsabstractWhile many new applications prefer manycore processor with a large number of cores, the exploding communications among multiple cores, caches, and off-chip memories is posing a fundamental challenge on manycore designs. Silicon photonics-based interconnection network promises high bandwidth, low latency, and high energy efficiency, and can potentially meet the communication requirements of manycore processors. In this paper, we propose CAMON, a small low-cost silicon photonic chiplet integrated into the manycore processor package. CAMON chiplet can effectively alleviate the communication bottlenecks of manycore processors and improve the energy efficiency of data movement, especially for large-scale systems. We develop a distributed arbitration system, a low-power low-latency optical interface, and an off-chip laser preactivation mechanism for CAMON. The experimental results show that compared with the electrical network, CAMON can improve the full-system performance per energy by 4.6×, speedup the manycore processor by 2.7×, and save the area of the processor die by 3%, in a 512-core system. Zhehui Wang, Jiang Xu 0001, Yi-Shing Chang, Jun Feng 0008, Xuanqi Chen, Shixi Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2020 | A Cross-Layer Optimization Framework for Integrated Optical Switches in Data CentersabstractThe advancement of silicon photonics promises integrated optical switches to provide high-bandwidth, low-latency, and low-power communications in data centers. An optical switch’s loss limits its scale and affects the energy efficiency of the switch system. In this paper, we present cross-layer optical switch optimization (CLOSO), a cross-layer optimization (CLO) framework, based on not only photonic device models at the physical layer but also optical switch models at the fabric layer. With the proposed framework, optimal losses of optical switches can be evaluated efficiently, and the corresponding losses and design parameters of photonic devices can be obtained. Using CLOSO, we optimize four categories of integrated optical switches, Crossbar, PILOSS, DRAGON, and FODON, and compare them regarding their optimal worst-case loss with variation of the switch scale and data rate of signals. Furthermore, system-level evaluations of the optimized optical switches are performed, demonstrating a significant improvement of energy efficiency from the CLO. For instance, CLOSO helps to reduce the energy consumption of a 64-port DRAGON and FODON to as low as 6 pJ/bit and that of a 128-port DRAGON and FODON to as low as 10 pJ/bit. The investigation of 128-port switches also shows the necessity of adaptive power control on lasers for high-radix integrated optical switches. Through quantitative analyses and comparisons, CLOSO shows the capability of facilitating initial design exploration of optical switches and paves the way to fair evaluations and comparisons of switch systems in data centers. Peng Yang 0003, Yi-Shing Chang, Jiang Xu 0001, Xuanqi Chen, Zhehui Wang, Jun Feng 0008 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2020 | Multidomain Inter/Intrachip Silicon Photonic Networks for Energy-Efficient Rack-Scale Computing SystemsabstractRack-scale computing systems are promising to undertake the emerging large-scale applications by distributing massive tasks to processing cores. The communication and coordination efficiency of these tasks and resources directly affect the system performance and energy consumption. Silicon photonic interconnects are expected to address the communication and system power consumption challenges imposed on rack-scale systems. However, the control for optical interconnects can cause server performance degradation if not properly designed, especially for the complicated and time-consuming multidomain networks. In this paper, we study the optical interconnects for rack-scale computing systems and propose a new communication flow and control scheme for the efficient coordination of distributed resources. Particularly, we first propose a forward propagation strategy that parallels the path reservation process with the distributed tasks connection setup. Second, we develop a pre-emptive chain feedback (PCF) scheme to optimize multidomain path reservation. The PCF scheme pre-emptively allocates network resources with the help of multicell reservation window and quickly releases resources with a feedback mechanism. This solution increases the network resources utilization and task coordination efficiency while minimizing path reservation overheads. Comparing to the baseline InfiniBand network fabric and handshake scheme, PCF can improve network throughput greatly under uniform and hotspot traffic patterns. Realistic benchmark results show that the PCF scheme on average reduces 52% and 60% energy consumption per unit system performance than InfiniBand and the handshake scheme for a 256-node rack system. Peng Yang 0003, Zhehui Wang, Jiang Xu 0001, Yi-Shing Chang, Xuanqi Chen, Rafael Kioji Vivas Maeda, Jun Feng 0008 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2019 | Systematic Exploration of High-Radix Integrated Silicon Photonic Switches for DatacentersabstractHigh-radix integrated silicon photonic switches promise ultrahigh bandwidth communications required by next generation data centers. To holistically explore the characteristics of high-radix integrated optical switches, this work systematically studies the latency, throughput and energy consumption, with detailed models and various system configurations. Three categories of space switches, blocking, rearrangeable non-blocking and strictly non-blocking switches, are investigated, together with one of the widely used wavelength switches, arrayed waveguide grating router (AWGR). The work paves the ways to automatically optimize high-radix integrated silicon photonic switches. Jun Feng 0008, Xuanqi Chen, Zhehui Wang, Shixi Chen, Jiang Xu 0001 |
ICCAD | 2 |
| 2018 | Co-manage power delivery and consumption for manycore systems using reinforcement learningabstractMaintaining high energy efficiency has become a critical design issue for high-performance systems. Many power management techniques have been proposed for the processor cores such as dynamic voltage and frequency scaling (DVFS). However, very few solutions consider the power losses suffered on the power delivery system (PDS), despite the fact that they have a significant impact on the system overall energy efficiency. With the explosive growth of system complexity and highly dynamic workloads variations, it is challenging to find the optimal power management policies which can effectively match the power delivery with the power consumption. To tackle the above problems, we propose a reinforcement learning-based power management scheme for manycore systems to jointly monitor and adjust both the PDS and the processor cores aiming to improve system overall energy efficiency. The learning agents distributed across power domains not only manage the power states of processor cores but also control the on/off states of on-chip VRs to proactively adapt to the workload variations. Experimental results with realistic applications show that when the proposed approach is applied to a large-scale system with a hybrid PDS, it lowers the system overall energy-delay-product (EDP) by 41% than a traditional monolithic DVFS approach with a bulky off-chip VR. Haoran Li 0002, Zhongyuan Tian, Rafael Kioji Vivas Maeda, Xuanqi Chen, Jun Feng 0008, Jiang Xu 0001 |
ICCAD | 5 |
| 2018 | Decentralized Collaborative Power Management through Multi-Device Knowledge SharingabstractBattery-powered mobile devices have limited energy capacity, urging the development of efficient power management approaches. Reinforcement learning (RL) algorithms are adaptive to the changing environment and have been widely used for runtime power management. Recently, collaborative RL-based approaches have been explored to accelerate the learning process. Nonetheless, existing works usually require a cloud service provider to achieve centralized multi-device knowledge sharing, which suffers from the single point of failure and does not guarantee users' privacy. To address this issue, we propose a decentralized multi-device collaborative power management approach in this work, where devices directly share their knowledge with their trusted neighbors. Experimental results show that the proposed method can achieve an up to 21% energy reduction with a 4× speedup over the individual learning-based approach. Zhongyuan Tian, Haoran Li 0002, Rafael Kioji Vivas Maeda, Jun Feng 0008, Jiang Xu 0001 |
ICCD | 4 |