EDBT 2026 Demo / reviewers in the wild / expert
Yuechen Chen
dblp:166/1870
· DBLP profile ↗
10ranked-venue papers
8as first author
3since 2021 · last 2024
0000-0001-6671-8443ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 7 first-author · 3 since 2021Security and privacy · 3 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A Balanced Sparse Matrix Convolution Accelerator for Efficient CNN TrainingabstractSparse Convolutional Neural Network (CNN) training is well known to be time-consuming due to significant off-chip memory traffic. To effectively deploy sparse training, existing accelerators store matrices in a compressed format to eliminate memory accesses for zeros; hence, accelerators are designed to process compressed matrices to avoid zero computations. We have observed that the compression rate is greatly affected by the sparsity in the matrices with different formats. Given the varying levels of sparsity in activations, weights, errors, and gradients matrices throughout the sparse training process, it becomes impractical to achieve consistently high compression rates using a singular compression method for the entire duration of the training. Moreover, random zeros in the matrices result in irregular computation patterns, further increasing execution time. To address these issues, we propose a balanced sparse matrix convolution accelerator design for efficient CNN training. Specifically, a dual matrix compression technique is developed that seamlessly combines two widely used sparse matrix compression formats with a control algorithm for lower memory traffic during training. Based on this compression technique, a two-level workload balancing technique is then designed to further reduce the execution time and energy consumption. Finally, an accelerator is implemented to support the proposed techniques. The cycle-accurate simulation results show that the proposed accelerator reduces the execution time by 34% and the energy consumption by 24% on average compared to existing sparse training accelerators. Yuechen Chen, Ahmed Louri, Shanshan Liu 0001, Fabrizio Lombardi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2023 | Slack-Aware Packet Approximation for Energy-Efficient Network-on-ChipsabstractNetwork-on-Chips (NoCs) are the standard on-chip communication fabrics for connecting cores, caches, and memory controllers in multi/many-core systems. With the increase in communication load introduced by emerging parallel computing applications, on-chip communication is becoming more costly than computation in terms of energy consumption. This paper contributes to existing research on approximate communication by proposing a slack-aware packet approximation technique to reduce the energy consumed by NoCs for sustainable parallel computation. The proposed approximation technique lowers both the execution time and NoC power consumption by reducing the packet size based on slack. The slack is the number of cycles by which a packet can be delayed in the network with no effect on execution time. Thus, low-slack packets are considered critical to system performance, and prioritizing these packets during the transmission will significantly reduce execution time. The proposed technique includes a slack-aware control policy to identify low-slack packets and accelerates these packets using two packet approximation mechanisms, namely, an in-network approximation (INAP) and a network interface approximation (NIAP). INAP mechanism prioritizes low-slack packets during the arbitration phase of the router by approximating packets with high-slack. NIAP mechanism reduces the latency of the network links and switch traversals by truncating data for the low-slack packets. An approximate network interface and router are implemented to support the proposed technique with lightweight packet approximation hardware for lower power consumption and execution time. Cycle-accurate simulations using the AxBench and PARSEC benchmark suites show that the proposed approximate communication technique achieves reductions of up to 24% in execution time and 38% in energy consumption with 1.1% less accuracy loss on average compared to existing approximate communication techniques. Yuechen Chen, Ahmed Louri, Shanshan Liu 0001, Fabrizio Lombardi |
IEEE Trans. Sustain. Comput. | 1 |
| 2022 | Approximate Network-on-Chips with Application to Image ClassificationabstractApproximation is an emerging design methodology for reducing power consumption and latency of on-chip communication in many computing applications. However, existing approximation techniques either achieve modest improvements in these metrics or require retraining after approximation. Since classifying many images introduces intensive on-chip communication, reductions in both network latency and power consumption are highly desired. In this paper, we propose an approximate communication technique (ACT) to improve the efficiency of on-chip communications for image classification applications. The proposed technique exploits the error-tolerance of the image classification process to reduce power consumption and latency of on-chip communications, resulting in better overall performance for image classification. This is achieved by incorporating novel quality control and data approximation mechanisms that reduce the packet size. In particular, the proposed quality control mechanisms identify the error-resilient variables and automatically adjust the error thresholds of the variables based on the image classification accuracy. The proposed data approximation mechanisms significantly reduce packet size when the variables are transmitted. The proposed technique reduces the number of flits in each data packet as well as the on-chip communication while maintaining an excellent image classification accuracy. Cycle-accurate simulation results show that ACT achieves 27% in network latency reduction and 28% in dynamic power reduction as compared to existing approximate communication techniques with less than 0.85% classification accuracy loss. Yuechen Chen, Ahmed Louri, Shanshan Liu 0001, Fabrizio Lombardi |
NAS | 1 |
| 2020 | Leakage-Resilient Inner-Product Functional Encryption in the Bounded-Retrieval Model
Linru Zhang, Xiangning Wang, Yuechen Chen, Siu-Ming Yiu |
ICICS | 3 |
| 2020 | Learning-Based Quality Management for Approximate Communication in Network-on-ChipsabstractCurrent multi/many-core systems spend large amounts of time and power transmitting data across on-chip interconnects. This problem is aggravated when data-intensive applications, such as machine learning and pattern recognition, are executed in these systems. Recent studies show that some data-intensive applications can tolerate modest errors, thus opening a new design dimension, namely, trading result quality for better system performance. In this article, we explore application error tolerance and propose an approximate communication framework to reduce the power consumption and latency of network-on-chips (NoCs). The proposed framework incorporates a quality control method and a data approximation mechanism to reduce the packet size to decrease network power consumption and latency. The quality control method automatically identifies the error-resilient variables that can be approximated during transmission and calculates their error thresholds based on the quality requirements of the application by analyzing the source code. The data approximation method includes a lightweight lossy compression scheme, which significantly reduces packet size when the error-resilient variables are transmitted. This framework results in fewer flits in each data packet and reduces traffic in NoCs while guaranteeing the quality requirements of applications. Our cycle-accurate simulation using the AxBench benchmark suite shows that the proposed approximate communication framework achieves 62% latency reduction and 43% dynamic power reduction compared to previous approximate communication techniques while ensuring 95% result quality. Yuechen Chen, Ahmed Louri |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2020 | An Approximate Communication Framework for Network-on-ChipsabstractCurrent multi-/many-core systems spend large amounts of time and power transmitting data across on-chip interconnects. This problem is aggravated when data-intensive applications, such as machine learning and pattern recognition, are executed in these systems. Recent studies show that some data-intensive applications can tolerate modest errors, thus opening a new design dimension, namely, trading result quality for better system performance. In this article, we explore application error tolerance and propose an approximate communication framework to reduce the power consumption and latency of network-on-chips (NoCs). The proposed framework incorporates a quality control method and a data approximation mechanism to reduce the packet size to decrease network power consumption and latency. The quality control method automatically identifies the error-resilient variables that can be approximated during transmission and calculates their error thresholds based on the quality requirements of the application by analyzing the source code. The data approximation method includes a lightweight lossy compression scheme, which significantly reduces packet size when the error-resilient variables are transmitted. This framework results in fewer flits in each data packet and reduces traffic in NoCs while guaranteeing the quality requirements of applications. Our cycle-accurate simulation using the AxBench benchmark suite shows that the proposed approximate communication framework achieves 62 percent latency reduction and 43 percent dynamic power reduction compared to previous approximate communication techniques while ensuring 95 percent result quality. Yuechen Chen, Ahmed Louri |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2019 | An online quality management framework for approximate communication in network-on-chipsabstractApproximate communication is being seriously considered as an effective technique for reducing power consumption and improving the communication efficiency of network-on-chips (NoCs). A major problem faced by these techniques is quality control: how do we ensure that the network will transmit data with sufficient accuracy for applications to produce acceptable results? Previous methods that addressed this issue require each application to calculate the approximation level for every piece of approximable data, which takes hundreds of cycles. So the approximation information is often not available when a request packet is transmitted. Therefore, the reply packet with the approximable data is transmitted with unnecessarily absolute accuracy, reducing the effectiveness of approximate communication. Yuechen Chen, Ahmed Louri |
ICS | 1 |
| 2018 | DEC-NoC: An Approximate Framework Based on Dynamic Error Control with Applications to Energy-Efficient NoCsabstractNetwork-on-Chips (NoCs) have emerged as the standard on-chip communication fabrics for multi/many core systems and system on chips. However, as the number of cores on chip increases, so does power consumption. Recent studies have shown that NoC power consumption can reach up to 40% of the overall chip power [1]-[3]. Considerable research efforts have been deployed to significantly reduce NoC power consumption. In this paper, we build on approximate computing techniques and propose an approximate communication methodology called DEC-NoC for reducing NoC power consumption. The proposed DEC-NoC leverages applications' error tolerance and dynamically reduces the amount of error checking and correction in packet transmission, which results in a significant reduction in the number of retransmitted packets. The reduction in packet retransmission results in reduced power consumption. Our cycle accurate simulation using PARSEC benchmark suites shows that DEC-NoC achieves up to 56% latency reduction and up to 58% dynamic power reduction compared to NoC architectures with conventional error control techniques. Yuechen Chen, Md Farhadur Reza, Ahmed Louri |
ICCD | 1 |
| 2018 | Secure Compression and Pattern Matching Based on Burrows-Wheeler TransformabstractSearchable compressed data structures (e.g.Burrows-Wheeler Transform) enable one to create a memory-efficient index for large datasets such as human genomes. On the other hand, storing such an index in a third-party server, e.g., cloud, may have the privacy and confidentiality issues. An open problem in the community is to construct a secure variant of such a data structure. This problem is challenging as most of the existing works were shown to be insecure and none of them is able to perform pattern matching. In this paper, we provide the first solution based on Burrows-Wheeler Transform (BWT) to solve this problem (our scheme can do both compression and pattern matching). A new security definition, called isomophism-restricted IND-CPA security, is proposed. We show that our scheme is secure under this definition and our scheme is practical by experiments. Gongxian Zeng, Meiqi He, Linru Zhang, Jun Zhang 0049, Yuechen Chen, Siu-Ming Yiu |
PST | 5 |
| 2014 | Fully Secure Ciphertext-Policy Attribute Based Encryption with Security Mediator
Yuechen Chen, Zoe Lin Jiang, Siu-Ming Yiu, Joseph K. Liu, Man Ho Au, Xuan Wang 0002 |
ICICS | 1 |