Shengyu Duan

dblp:184/5391 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
7since 2021 · last 2026
0009-0000-9321-8380ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 4 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AURORA - AUtomated 8T SRAM Wired-OR Logic Array for Boolean-Based Machine Learning
Komal Krishnamurthy, Marcos L. L. Sartori, Shengyu Duan, Alexandre Yakovlev, Rishad A. Shafik
DATE3
2025 FPGA-Based Processing-In-Memory with Optimized Multi-BRAM Reduction for Improved Latency-Cost Trade-Off
Shengyu Duan, Chen Zhan, Xiaoli Zhi
ICIC (10)2
2023 FHC-DQP: Federated Hierarchical Clustering for Distributed QoS Prediction
abstract
With the overwhelming explosion of Web services, how to effectively predict unknown QoS has become a key issue of differentiating large-scale similar or functionally equivalent Web services. However, current state-of-the-art QoS prediction approaches based on deep learning still suffer from two deficiencies. First, they mainly focus on predicting vacant QoS in a centralized manner and scarcely take into account distributed QoS prediction, which makes difficult to protect the privacy information of users invoking Web services. Second, they have ignored the hierarchical collaborative relationship to better extract latent features of users and services, reducing the accuracy of QoS prediction. To address these two issues, we propose a novel framework calledFederatedHierarchicalClustering forDistributedQoSPrediction(FHC-DQP). It collaboratively performs distributed federated training on independent users’ QoS invocations, and then the extracted federated users’ private features are fed to clustering algorithm for partitioning them into a set of clusters. By iteratively federated hierarchical clustering, users are fine-grained partitioned together and those users within the same cluster have stronger collaborative relevance for more effectively learning the latent features of users and services leading to the performance improvement of distributed QoS prediction, where contextual-aware deep neural network is designed for personalized QoS prediction. Extensive experiments are conducted based on a public real-world benchmarking dataset called WS-DREAM with almost 2,000,000 user-service historical QoS invocations. Compared with both centralized and federated competing baselines, the results demonstrate FHC-DQP receives superior performance for distributed QoS prediction, when it provides privacy-preserving of users’ QoS invocations.
Guobing Zou, Shengxiang Hu 0002, Shengyu Duan, Yanglan Gan, Bofeng Zhang, Yixin Chen 0001
IEEE Trans. Serv. Comput.4
2022 A secure authentication scheme based on differential public PUF
abstract
Physical Unclonable Functions (PUFs) have emerged as a promising primitive to provide hardware security services like authentication for integrated circuit applications. Public PUFs (PPUFs) address the crucial PUF vulnerability of requiring databases to store the secrete reference information, by using publicly accessible models. PPUFs apply highly complicated structures to cause noticeable time differences between actual execution and simulation, preventing authentication fraud. This paper investigates one of the PPUF families, differential PPUF (dPPUF), and the resistance of dPPUF against prediction attacks. A prediction attack is presented by exploiting the relationship between response, challenge and the initial states of arbiter circuits in dPPUFs. We show the problem is caused by the circuit function of dPPUF, and is independent of the structure of dPPUF. A new authentication scheme is thereby proposed, which challenges a dPPUF subsequently with multiple correlated challenges. We show the above mentioned security flaw is eliminated by our proposed scheme. This scheme further improves uniformity and uniqueness of dPPUFs by 5.81% and 1.78%, respectively, and leads to negligible reductions for CRP space, compared with a conventional authentication process.
Shengyu Duan, Gaole Sai
CF1
2022 Work-in-Progress: Accelerated Matrix Factorization by Approximate Computing for Recommendation System
abstract
Matrix factorization (MF) is widely used in collaborative filtering-based recommendation systems, but the computational complexity greatly increases for larger scaled recommendation systems. We propose to accelerate MF by performing approximate matrix multiplications, considering the joint sparsity of the decomposed matrices. We show our method realizes a more than 1.1 speedup with a minimal error, and the speedup can be higher for the recommendation systems with larger scales.
Yining Wu, Gaole Sai, Shengyu Duan
EMSOFT3
2022 Hardware Acceleration for 1D-CNN Based Real-Time Edge Computing
Gaole Sai, Shengyu Duan
NPC3
2022 DeepLTSC: Long-Tail Service Classification via Integrating Category Attentive Deep Neural Network and Feature Augmentation
abstract
With the explosive growth in the number and diversity of Web services, correlative research has been investigated on Web service classification, as it fundamentally promotes advanced service-oriented applications, such as service discovery, selection, composition and recommendation. However, conventional approaches are restricted to indiscriminatingly classify Web services, which can trigger many challenges. First, they have not made full advantage of the implicit relationships among multi-dimensional information of Web services, such as the increasing number of service categories. Thus, it leads to low effectiveness of learning and representing service features, failing to ensure the overall accuracy of service classification. Second, the imbalance of service distributions has been ignored, while it is observed that service categories reveal distinct long-tail characteristics. That results in low accuracy on service classification for those categories that contain fewer Web services. To handle the challenges of more effectively learning implicit service features across the service repository, and with a particular concentration on those tail categories that contain fewer Web services, we propose a novel framework called DeepLTSC to more accurately perform the task of Web service classification under long-tail distributions. In DeepLTSC, we first present an improved label attentive convolutional deep neural network (LACNN) with service categories, which can generate deep service features to improve the overall classification performance. Then, a proposed service feature augmentation model (SFA) together with focal loss function is integrated into DeepLTSC to further optimize service features, aiming to boost the classification accuracy on tail service categories. Extensive experiments are conducted on three large-scale real-world services datasets with different long-tail distributions. The results demonstrate that DeepLTSC significantly outperforms state-of-the-art approaches for Web service classification on both overall and tail categories.
Guobing Zou, Song Yang 0003, Shengyu Duan, Bofeng Zhang, Yanglan Gan, Yixin Chen 0001
IEEE Trans. Netw. Serv. Manag.3
2020 BTI Aging Monitoring based on SRAM Start-up Behavior
abstract
Bias Temperature Instability (BTI) is one of the dominant CMOS aging mechanisms. It causes time-dependent variation, threatening circuit lifetime reliability. BTI-induced circuit errors are not detectable at the fabrication stage. On-line monitoring schemes are therefore necessary to capture the degradations during the operational time. Traditional aging monitoring techniques exhibit high implementation complexity and low stability. In this paper, we propose a BTI monitoring approach by simply tracking the start-up behavior of SRAM cells. SRAM is a widely used on-chip device in many applications. We study the impact of BTI for SRAM start-up values and age some cells in a manipulated manner. The BTI degradation is evaluated based on the number of SRAM cells starting with a certain value. This technique can be used to estimate the degradation for on-chip logic circuits without introducing additional circuitry, and thus has very low implementation complexity. We use an SRAM array with 1024 cells to estimate the degradations for multiple logic circuits, and show the average mean absolute percentage error as 8.48%. In addition, this technique is robust considering process, voltage and temperature variations.
Shengyu Duan, Gaole Sai
ATS1
2019 A reliable PUF in a dual function SRAM
Mohd Syafiq Mispan, Shengyu Duan, Basel Halak, Mark Zwolinski
Integr.2
2018 Cell Flipping with Distributed Refresh for Cache Ageing Minimization
abstract
CMOS wear-out mechanisms, especially Bias Temperature Instability (BTI), have caused growing concerns about circuit reliability. For cache memories, BTI reduces the static noise margin (SNM), causing unreliable read operations. In practice, error-correction codes (ECCs) are often used to protect data from transient errors in caches, but the limited error correction capabilities are not always enough to overcome BTIinduced read failures. In this paper, we propose a cell flipping technique with distributed refresh phases (CFDR) to minimize cache degradations. The CFDR method flips and refreshes each cache block at different times, minimizing the interruption time and balancing the degradation rate, even for infrequently replaced cache blocks. We evaluate the CFDR technique on an instruction cache in a 32-bit ARM architecture and show our method reduces the number of error bits by 58.86% and 13.59%, compared with an ECC scheme and a traditional cell flipping technique. The cache lifetime can be improved by 125% by using CFDR with less than 1% area overhead, which is not only more effective but also more cost-efficient than the existing techniques.
Shengyu Duan, Basel Halak, Mark Zwolinski
ATS1
2018 A New Ageing-Aware Approach Via Path Isolation
abstract
NBTI is becoming one of the major circuit reliability issues in nano-scale technologies. BTI can cause a threshold voltage shift in CMOS devices and consequently increase circuit delay. This paper proposed a novel ageing aware approach to improve circuit's lifetime. The vulnerable circuit paths against ageing effects are isolated. In addition, minimum area overhead is consumed by adopting proposed synthesis algorithm. The simulation results show that the proposed approach can save up to 67.7% area compared with the conventional over-design technique.
Shengyu Duan, Tom J. Kazmierski
FDL2
2018 Lifetime Reliability-Aware Digital Synthesis
Shengyu Duan, Mark Zwolinski, Basel Halak
IEEE Trans. Very Large Scale Integr. Syst.1