VLDB 2026 Research / reviewers in the wild / expert
Shixi Chen
dblp:47/1018
· DBLP profile ↗
22ranked-venue papers
5as first author
14since 2021 · last 2024
0000-0001-7282-0495ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 1 first-author · 13 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-author · 1 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Towards Scalable GPU System with Silicon Photonic ChipletabstractGPU-based computing has emerged as a predominant solution for high-performance computing and machine learning applications. The continuously escalating computing demand fore-sees a requirement for larger-scale GPU systems in the future. However, this expansion is constrained by the finite number of transistors per die. Although chip let technology shows potential for building large-scale systems, current chiplet interconnection technologies suffer from limitations in both bandwidth and en-ergy efficiency. In contrast, optical interconnect has ultra-high bandwidth and energy efficiency, and thereby is promising for constructing chiplet-based GPU systems. Yet, previously proposed optical networks lack scalability and cannot be directly applied to existing chiplet-based GPU systems. In this work, we address the challenges of designing large-scale G PU systems with silicon photonic chiplets. We propose GROOT, a group-based optical network that divides the entire system into groups and facilitates resource sharing among the chiplets within each group. Additionally, we design dedicated channel mapping and allocation policies tailored for the request network and the reply network, respectively. Experimental results show that GROOT achieves 48% improvement on performance and 24.5% reduction on system energy consumption over the baseline. Chengeng Li, Shixi Chen |
DATE | 3 |
| 2024 | PhotonNTT: Energy-Efficient Parallel Photonic Number Theoretic Transform AcceleratorabstractFully homomorphic encryption (FHE) presents a promising opportunity to remove privacy barriers in various scenarios including cloud computing and secure database search, by enabling computation on encrypted data. However, integrating FHE with real-world applications remains challenging due to its significant computational overhead. In the FHE scheme, Number Theoretic Transform (NTT) consumes the primary computing resources and has great potential for acceleration. For the first time, we present a photonic NTT accelerator, PhotonNTT, with high energy efficiency and parallelism to address the above challenge. Our approach involves formulating the NTT into matrix-vector multiplication (MVM) operations and mapping the data flow into parallel photonic MVM units. A dedicated data mapping scheme is proposed to introduce free spectral range (FSR) and distributed RAM design into the system, which enables a high bit-wise parallelism level. The system's reliability is validated through the Monte-Carlo BER analysis. The experimen-tal evaluation shows that the proposed architecture outperforms SOTA CiM-based NTT accelerators with an improvement of 50x in throughput and 63x improvement in energy efficiency. Yinyi Liu, Chengeng Li, Shixi Chen, Fengshi Tian, Wei Zhang 0012, Jiang Xu 0001 |
DATE | 7 |
| 2024 | PC-oriented Prediction-based Runtime Power Management for GPGPU using Knowledge TransferabstractAs Moore's law slows down, computing systems must prioritize higher energy efficiency to sustain performance scaling. GPUs have emerged as the primary workhorses of computing resources, making the achievement of high energy efficiency in GPUs a critical concern. However, implementing runtime power management on GPUs poses significant challenges due to the high variations and complexities arising from workloads and hardware configurations, which render offline optimization and reactive-based methods less effective. In this paper, we present a program counter (PC)-oriented prediction-based power management approach for GPGPUs. Our approach leverages the benefits of prediction to address online variations while enhancing prediction capability through knowledge transfer across different levels of resources. Experiments conducted on realistic applications demonstrate that our proposed method achieves the maximum energy savings under a user-defined performance constraint compared to state-of-the-art designs. Lin Chen 0029, Xiao Li 0038, Shixi Chen, Fan Jiang 0015, Chengeng Li, Wei Zhang 0012, Jiang Xu 0001 |
SPAA | 3 |
| 2024 | Deep Reinforcement Learning-Based Power Management for Chiplet-Based Multicore SystemsabstractChiplet technology has emerged as a promising solution to address the increasing demand for high-performance computing in light of the slowdown of Moore’s law. While chiplet-based multicore systems offer higher performance through heterogeneous integration, they also pose challenges for power delivery system (PDS) design. The integration of additional vertical and inter-chiplet connections, along with higher power density, impose stringent requirements on power delivery. Moreover, PDS efficiency is affected by workload variations at runtime, necessitating the need to design and manage PDSs and processors as a whole to improve system energy efficiency while balancing performance. In this article, we propose an offline-online co-design optimization methodology that combines offline PDS design optimization with online power management. To address the power consumption and delivery mismatch, we introduce a centralized deep Q-network (DQN)-based online control scheme for power co-management in chiplet-based multicore systems. By carefully designing the state space and reward functions, our approach achieves workload-aware adaptive control to reduce the energy-delay-product (EDP) while maintaining PDS efficiency under a given performance target (PT). We conduct evaluations on realistic applications to validate the effectiveness of our approach. For 64-core systems, our method achieves an average EDP reduction of 67% while meeting a 90% PT, surpassing state-of-the-art modular Q-learning (MQL)-based and heuristic-based approaches by up to 4% and 16%, respectively. Additionally, our approach demonstrates wiser action selection policies, higher control stability, and lower implementation overhead compared to the MQL-based approach. Xiao Li 0038, Lin Chen 0029, Shixi Chen, Fan Jiang 0015, Chengeng Li, Wei Zhang 0012, Jiang Xu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2023 | RONet: Scaling GPU System with Silicon Photonic ChipletabstractModern GPU systems integrate hundreds of SMs on a single die, and future scaling envisions even more SMs being incorporated. However, the limited number of transistors per die constrains this growth. While current chiplet technology shows promise, its performance is limited by the bandwidth and energy efficiency of existing chiplet interconnect technologies. In contrast, optical interconnects offer ultra-high bandwidth and energy efficiency, making them ideal for high-performance chiplet-based GPUs. This work proposes a novel region-based optical network, called RONet, that divides a chiplet-based GPU system with a 2D Mesh layout into multiple row and column regions, where each region is connected by a separate optical link. Additionally, RONet employs a tuning-free transmission mechanism to further enhance inter-chiplet bandwidth. Experimental results show that RONet achieves 43% improvement on performance and 25.4% reduction on system energy consumption over the baseline. Chengeng Li, Fan Jiang 0015, Shixi Chen, Yinyi Liu, Lin Chen 0029, Xiao Li 0038, Jiang Xu 0001 |
ICCAD | 3 |
| 2022 | PHANES: ReRAM-based photonic accelerator for deep neural networksabstractResistive random access memory (ReRAM) has demonstrated great promises of in-situ matrix-vector multiplications to accelerate deep neural networks. However, subject to the intrinsic properties of analog processing, most of the proposed ReRAM-based accelerators require excessive costly ADC/DAC to avoid distortion of electronic analog signals during inter-tile transmission. Moreover, due to bit-shifting before addition, prior works require longer cycles to serially calculate partial sum compared to multiplications, which dramatically restricts the throughput and is more likely to stall the pipeline between layers of deep neural networks. Yinyi Liu, Shixi Chen, Jiang Xu 0001 |
DAC | 4 |
| 2022 | A Reliability Concern on Photonic Neural NetworksabstractEmerging integrated photonic neural networks have experimentally proved to achieve an ultra-high speedup of deep neural network training and inference in the optical domain. However, photonic devices suffer from the inherent crosstalk noise and loss, inevitably leading to reliability concerns. This paper systematically analyzes the impacts of crosstalk and loss on photonic computing systems. We propose a crosstalk-aware model for reliability estimation and find out the worst-case bounds as we increase the footprints and scales of the photonic chips. Our evaluations show that −30dB crosstalk noise can cause maximal photonic chip integration to a sharp drop by 109x. To facilitate very-large-scale photonic integration for future computing, we further propose multiple heterogeneous bijou photonic-cores to address the crosstalk-aware reliability concern. Yinyi Liu, Jun Feng 0008, Shixi Chen, Jiang Xu 0001 |
DATE | 4 |
| 2022 | Accelerating Cache Coherence in Manycore Processor through Silicon Photonic ChipletabstractCache coherence overhead in manycore systems is becoming prominent with the increase of system scale. However, traditional electrical networks restrict the efficiency of cache coherence transactions in the system due to the limited bandwidth and long latency. Optical network promises high bandwidth and low latency, and supports both efficient unicast and multicast transmission, which can potentially accelerate cache coherence in manycore systems. This work proposes a novel photonic cache coherence network with a physically centralized logically distributed directory called PCCN for chiplet-based manycore systems. PCCN adopts a channel sharing method with a contention solving mechanism for efficient long-distance coherence-related packet transmission. Experiment results show that compared to state-of-the-art proposals, PCCN can speed up application execution time by 1.32x, reduce memory access latency by 26%, and improve energy efficiency by 1.26x, on average, in a 128-core system. Chengeng Li, Fan Jiang 0015, Shixi Chen, Yinyi Liu, Jiang Xu 0001 |
ICCAD | 3 |
| 2022 | Improving the thermal reliability of photonic chiplets on multicore processors
Xuanqi Chen, Jun Feng 0008, Shixi Chen, Jiang Xu 0001 |
Integr. | 5 |
| 2022 | HERO: Pbit High-Radix Optical Switch Based on Integrated Silicon Photonics for Data CenterabstractTo establish flatten networks and accomplish rapid and efficient communications in the future hyper-scale data centers, HERO, a high-radix optical switch based on integrated silicon photonics, is proposed in this work. The architecture of HERO, including the switch fabric, switch interface, and switch controller, is described in detail. Two new switch control approaches: 1) split-transaction predictive control and 2) wavelength-group switching, are developed. The efficient control together with the optimized high-radix integrated optical switch fabrics help HERO achieve over 1 Pbps switching capacity. Even for small packets, such as 64–256 B Ethernet packets, the maximal utilization of the switch can reach up to 83%, and the throughput can approximate 1 Pbps. Further design explorations on the packet length and some key configurations, including the number of wavelengths and wavelength groups, are also conducted in this work, paving the way to the design of high-performance flatten data-center networks in the future. Jun Feng 0008, Jiang Xu 0001, Xuanqi Chen, Shixi Chen, Yinyi Liu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2022 | Fast and Accurate Statistical Simulation of Shared-Memory Applications on Multicore SystemsabstractDetailed cycle-accurate simulation of multicore systems is naturally slow. Statistical simulation is one alternative that permits trading off simulation speed for accuracy. However, there is a lack of effective memory locality models for multicore applications. Hence, existing statistical simulators neglect data-sharing between threads. Additionally, the standard method to speed up statistical simulations is to blindly reduce the trace length to be synthesized. While this gives good control over the speedup, it leaves the simulation error unbounded. In this work, we introduce a novel statistical simulation methodology for exploration of shared-memory multicore systems. It includes a newsharing-localitymodel (Shalom) that captures and reproduces data-sharing in multithread applications. Furthermore, we propose a method to bound the simulation error for a particular metric while maximizing speedup. The technique works by monitoring the convergence of the statistical synthesis. It is referred to asconvergence-deterministicsimulation (Condens). The combination ofShalomandCondensis around 130x faster than cycle-accurate simulations with reasonable accuracy loss. Our approach is also 5x faster than state-of-the-art sampling simulation under the same accuracy level. Compared to previous statistical simulators ignoring sharing, our approach is 2x more accurate for performance metrics and 8x more accurate for cache miss estimations. Fan Jiang 0015, Rafael Kioji Vivas Maeda, Jun Feng 0008, Shixi Chen, Lin Chen 0029, Xiao Li 0038, Jiang Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2021 | A general framework for privacy-preserving of data publication based on randomized response techniquesabstractPrivacy preserving is a paramount concern in publishing datasets that contain sensitive information. Preventing privacy disclosure and providing useful information to legitimate users for data analyzing/mining are conflicting goals. Randomized response is a class of techniques that perturbs each sensitive value in a certain way, so that personal privacy is protected while the large-trend of the entire dataset is still recoverable. However, existing randomized response techniques do not allow to flexibly configure the level of privacy protection, support only a few types of aggregate queries, and cannot achieve the best answer accuracy from perturbed data. These drawbacks impair the effectiveness of those techniques. This paper proposes a general framework based on randomized response techniques, which has good flexibility and extensibility, and can improve the effectiveness of randomized response methods. Our approach is validated by extensive experiments and comparison with existing randomized response and generalization methods. Chaobin Liu, Shixi Chen, Shuigeng Zhou, Jihong Guan |
Inf. Syst. | 2 |
| 2021 | Simultaneously Tolerate Thermal and Process Variations Through Indirect Feedback Tuning for Silicon Photonic NetworksabstractSilicon photonics is the leading candidate technology for high-speed and low-energy-consumption networks. Thermal and process variations are the two main challenges of achieving high-reliability photonic networks. Thermal variation is due to the heat issues created by application, floorplan, and environment, while process variation is caused by fabrication variability in the deposition, masking, exposition, etching, and doping. Tuning techniques are then required to overcome the impact of the variations and efficiently stabilize the performance of silicon photonic networks. We extend our previous optical switch integration model, BOSIM, to support the variation and thermal analyses. Based on device properties, we propose indirect feedback tuning (IFT) to simultaneously alleviate thermal and process variations. IFT can improve the BER of silicon photonic networks to 10-9under different variation situations. Compared to state-of-the-art techniques, IFT can achieve an up to 1.52 ×108times bit-error-rate improvement and 4.11X better heater energy efficiency. Indirect feedback does not require high-speed optical signal detection, and thus, the circuit design of IFT saves up to 61.4% of the power and 51.2% of the area compared to state-of-the-art designs. Xuanqi Chen, Jun Feng 0008, Jiang Xu 0001, Shixi Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2021 | Reduce Loss and Crosstalk in Integrated Silicon-Photonic Multistage Switching Fabrics Through Multichip PartitionabstractWith the increasing popularity of data-intensive applications in data centers, the switching fabric in the internode network becomes significant. Silicon-photonic switching fabrics have a bright future in data centers, which offer high bandwidth, high energy efficiency, and low latency. However, integrating a high radix multistage switching fabric in a single chip faces challenges. A large number of waveguide crossings on the silicon photonic die causes massive power loss and introduces a tremendous amount of crosstalk noise. In this article, we propose a chip partition optimization platform (POP), which can decrease the number of waveguide crossings and shorten the on-chip traversal distance of optical signals. Our algorithms can effectively reduce the power loss and crosstalk noise in silicon-photonic multistage switching fabrics, and help to improve the signal integrity. For example, compared with the common design, POP can achieve 33-dB improvement on average power loss, 42-dB improvement on the worst-case power loss, and 39-dB improvement on the worst-case signal to noise ratio, in a$1024\times1024$butterfly based silicon-photonic switching fabric. Zhehui Wang, Jiang Xu 0001, Jun Feng 0008, Shixi Chen, Xuanqi Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2020 | Efficient Optical Power Delivery System for Hybrid Electronic-Photonic Manycore ProcessorsabstractA lot of efforts have been devoted to optically enabled high-performance communication infrastructures for future manycore processors. Silicon photonic network promises high bandwidth, high energy efficiency and low latency. However, the ever-increasing network complexity results in high optical power demands, which stress the optical power delivery and affect delivery efficiency. Facing these challenges, we propose Ring-based Optical Active Delivery (ROAD) system, to effectively manage and efficiently deliver high optical power throughout photonic-electronic hybrid systems. Experimental results demonstrate up to 5.49X energy efficiency improvement compared to traditional design without affecting processor performance. Shixi Chen, Jiang Xu 0001, Xuanqi Chen, Jun Feng 0008, Zhongyuan Tian, Xiao Li 0038 |
DATE | 1 |
| 2020 | CAMON: Low-Cost Silicon Photonic Chiplet for Manycore ProcessorsabstractWhile many new applications prefer manycore processor with a large number of cores, the exploding communications among multiple cores, caches, and off-chip memories is posing a fundamental challenge on manycore designs. Silicon photonics-based interconnection network promises high bandwidth, low latency, and high energy efficiency, and can potentially meet the communication requirements of manycore processors. In this paper, we propose CAMON, a small low-cost silicon photonic chiplet integrated into the manycore processor package. CAMON chiplet can effectively alleviate the communication bottlenecks of manycore processors and improve the energy efficiency of data movement, especially for large-scale systems. We develop a distributed arbitration system, a low-power low-latency optical interface, and an off-chip laser preactivation mechanism for CAMON. The experimental results show that compared with the electrical network, CAMON can improve the full-system performance per energy by 4.6×, speedup the manycore processor by 2.7×, and save the area of the processor die by 3%, in a 512-core system. Zhehui Wang, Jiang Xu 0001, Yi-Shing Chang, Jun Feng 0008, Xuanqi Chen, Shixi Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2019 | Systematic Exploration of High-Radix Integrated Silicon Photonic Switches for DatacentersabstractHigh-radix integrated silicon photonic switches promise ultrahigh bandwidth communications required by next generation data centers. To holistically explore the characteristics of high-radix integrated optical switches, this work systematically studies the latency, throughput and energy consumption, with detailed models and various system configurations. Three categories of space switches, blocking, rearrangeable non-blocking and strictly non-blocking switches, are investigated, together with one of the widely used wavelength switches, arrayed waveguide grating router (AWGR). The work paves the ways to automatically optimize high-radix integrated silicon photonic switches. Jun Feng 0008, Xuanqi Chen, Zhehui Wang, Shixi Chen, Jiang Xu 0001 |
ICCAD | 6 |
| 2019 | A novel privacy preserving method for data publicationabstractPrivacy has received increasing concerns in publication of datasets that contain sensitive information. Preventing privacy disclosure and providing useful information to legitimate users for data mining are conflicting goals. Generalization and randomized response methods were proposed in database community to tackle this problem. However, both of them have postulated the same prior belief for all transactions, which might be wrong modeling and lead to privacy breach. Besides, generalization and randomized response methods usually require a privacy controlling parameter to control the tradeoff between privacy and data quality, which may put the data publishers in a dilemma. In this paper, a novel privacy preserving method for data publication is proposed based on conditional probability distribution and machine learning techniques, which can achieve different prior beliefs for different transactions. A basic cross sampling algorithm and a complete cross sampling algorithm are designed respectively for the settings of single sensitive attribute and multiple sensitive attributes, and an improved complete algorithm is developed by using Gibbs sampling, in order to enhance data utility when data are not sufficient. Our method can offer stronger privacy guarantee, while, as shown in the extensive experiments, retaining better data utility. Chaobin Liu, Shixi Chen, Shuigeng Zhou, Jihong Guan |
Inf. Sci. | 2 |
| 2013 | Recursive mechanism: towards node differential privacy and unrestricted joinsabstractExisting differential privacy (DP) studies mainly consider aggregation on data sets where each entry corresponds to a particular participant to be protected. In many situations, a user may pose a relational algebra query on a database with sensitive data, and desire differentially private aggregation on the result of the query. However, no existing work is able to release such aggregation when the query contains unrestricted join operations. This severely limits the applications of existing DP techniques because many data analysis tasks require unrestricted joins. One example is subgraph counting on a graph. Furthermore, existing methods for differentially private subgraph counting support only edge DP and are subject to very simple subgraphs. Until recent, whether any nontrivial graph statistics can be released with reasonable accuracy for arbitrary kind of input graphs under node DP was still an open problem. Shixi Chen, Shuigeng Zhou |
SIGMOD Conference | 1 |
| 2012 | Integrating historical noisy answers for improving data utility under differential privacyabstractDifferential privacy is a robust principle for privacy preserving data analysis tasks, and has been successfully applied to a variety of applications. However, the number of queries that can be answered is limited for preventing privacy disclosure. Once the privacy budget is exhausted, all succeeding queries must be rejected. Therefore, each of the historical query answers is valuable and it is important to exploit them together to learn more about the data. We propose to integrate all available linear query answers into a consistent form that embodies our knowledge learned from the noisy answers, obtaining more accurate answers to past queries and even new queries, improving the data utility. Two distinct approaches are developed for this purpose, one via principle component analysis, and another via maximum entropy method. The second approach also generates a synthetic database, which is useful for differentially private data publishing. One important goal of our work is to ensure that the running time of our approaches does not grow with the cardinality of the universe of a data tuple, so that high-dimensional data with very large domain can still be tackled efficiently. Shixi Chen, Shuigeng Zhou, Sourav S. Bhowmick |
EDBT | 1 |
| 2009 | Concept Clustering of Evolving DataabstractMuch work has focused on mining evolving data, and most approaches learn the latest model from the latest data. The problem with these approaches is that the learned model is always of low quality. In this paper, we propose a clustering approach to find hidden concepts that control data generation. Unlike traditional clustering methods that are based on data similarity (measured by Euclidean distance, e.g.), we devise a new similarity metric for concept similarity. We propose a two step algorithm, which uses dynamic programming and hierarchical clustering to find concepts in the data. Shixi Chen, Haixun Wang, Shuigeng Zhou |
ICDE | 1 |
| 2008 | Stop Chasing Trends: Discovering High Order Models in Evolving DataabstractMany applications are driven by evolving data - patterns in Web traffic, program execution traces, network event logs, etc., are often non-stationary. Building prediction models for evolving data becomes an important and challenging task. Currently, most approaches work by "chasing trends", that is, they keep learning or updating models from the evolving data, and use these impromptu models for online prediction. In many cases, this proves to be both costly and ineffective - much time is wasted on re-learning recurring concepts, yet the classifier may remain one step behind the current trend all the time. In this paper, we propose to mine high-order models in evolving data. More often than not, there are a limited number of concepts, or stable distributions, in the data stream, and concepts switch between each other constantly. We mine all such concepts offline from a historical stream, and build high quality models for each of them. At run time, combining historical concept change patterns and cues provided by an online training stream, we find the most likely current concept and use its corresponding models to classify data in an unlabeled stream. The primary advantage of the high-order model approach is its high accuracy. Experiments show that in benchmark datasets, classification error of the high-order model is only a small fraction of that of the current best approaches. Another important benefit is that, unlike state-of-the-art approaches, our approach does not require users to tune any parameters to achieve a satisfying result on streams of different characteristics. Shixi Chen, Haixun Wang, Shuigeng Zhou, Philip S. Yu |
ICDE | 1 |