EDBT 2026 Demo / reviewers in the wild / expert
Shao-I Chu
dblp:67/4550
· DBLP profile ↗
15ranked-venue papers
8as first author
6since 2021 · last 2026
0000-0003-2651-2357ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 4 first-author · 6 since 2021Computer networks · 5 · 3 first-authorDatabases, data management, data science and information retrieval · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hardware-efficient architecture of spiking neural networks based on sign-magnitude stochastic computing
Thai N. Nguyen, Jun-Xiang Shi, Shao-I Chu, Bing-Hong Liu |
Integr. | 4 |
| 2026 | Area-Time Efficient Formula-Based BCH Decoder With Trace Mechanism for WBAN ApplicationsabstractThis brief presents the area-time efficient decoding algorithm and architecture for the double error correcting$(63,51)$Bose–Chaudhuri–Hocquenghem (BCH) code with applications to wireless body area networks (WBANs). The formula for finding the roots of an error locator polynomial (ELP) is rederived to reduce the decoding complexity. The trace constraint for this formula is also taken into consideration to avoid the performance loss of bit error rate (BER). Hardware implementation results reveal that the proposed architecture surpasses the well-known Chien search-based and searchless decoders by the improvements of at least 50.89% and 29.35%, respectively, in terms of area-time complexity. Wei-Che Liang, Thai N. Nguyen, Shao-I Chu, Bing-Hong Liu, Chen-Yang Hong, Shao-Tong Chen |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2024 | Stochastic Circuits for Computing Weighted Ratio With Applications to Multiclass Bayesian Inference MachineabstractBayesian inference is one method of statistical inference in machine learning. It predicts the probability that a given test belongs to a certain class and is widely used in various applications such as medical diagnosis, spam classification and fraud detection. The conventional binary architecture of computing the posterior probability is inefficient in practical implementation, which is involved in multiplication, addition and division operations. Recently, it has been shown that simple Muller C-elements, the asynchronous logic units, can perform stochastic Bayesian inference motivated by its truth table when the data is encoded as the bit-stream. The Bayesian inference machine is therefore implemented with low hardware cost. However, such an architecture is employed to compute the posterior probability of two classes only. This brief presents two stochastic circuit designs for computing the weighted ratio with multiple weights for generalized multi-class Bayesian machines. The first design is mainly based on the JK flip flop and multiplexers. The second approach is to construct the finite state machine (FSM) by manipulating the correlation between the input bit-streams. The FSM-based design requires fewer random number sources (RNSs) as compared to the JK flip flop-based implementation. These facts lead to a reduction of hardware area and energy. Simulation results show that the accuracy of the proposed JK flip flop-based and FSM-based designs is almost the same in the tested data sets. As compared to the traditional binary design, the circuit area of the proposed stochastic design is improved by$96\%$at least in the cases of three and four classes. The consumed energy per operation is reduced by$58.1\%$at least in the cases of three and four classes. Shao-I Chu, Chi-Long Wu, Tzu-Heng Chien, Bing-Hong Liu, Tu N. Nguyen 0001 |
IEEE Trans. Computers | 1 |
| 2024 | Service Recovery in NFV-Enabled Networks: Algorithm Design and AnalysisabstractNetwork function virtualization (NFV), a novel network architecture, promises to offer a lot of convenience in network design, deployment, and management. This paradigm, although flexible, suffers from many risks engendering interruption of services, such as node and link failures. Thus, resiliency is one of the requirements in NFV-enabled network design for recovering network services once occurring failures. Therefore, in addition to a primary chain of virtual network functions (VNFs) for a service, one typically allocates the corresponding backup VNFs to satisfy the resiliency requirement. Nevertheless, this approach consumes network resources that can be inherently employed to deploy more services. Moreover, one can hardly recover all interrupted services due to the limitation of network backup resources. In this context, the importance of the services is one of the factors employed to judge the recovery priority. In this paper, we first assign each service a weight expressing its importance, then seek to retrieve interrupted services such that the total weight of the recovered services is maximum. Hence, we also call this issue the VNF restoration for recovering weighted services (VRRWS) problem. We next demonstrate the difficulty of the VRRWS problem is NP-hard and propose an effective technique, termed online recovery algorithm (ORA), to address the problem without necessitating the backup resources. Eventually, we conduct extensive simulations to evaluate the performance of the proposed algorithm as well as the factors affecting the recovery. The experiment shows that the available VNFs should be migrated to appropriate nodes during the recovery process to achieve better results. Dung H. P. Nguyen, Chih-Chieh Lin, Tu N. Nguyen 0001, Shao-I Chu, Bing-Hong Liu |
IEEE Trans. Cloud Comput. | 4 |
| 2023 | An Efficient Hard-Detection GRAND Decoder for Systematic Linear Block CodesabstractGuessing random additive noise decoding (GRAND) has been recently proposed as a code-agnostic decoding technique for linear block codes, which attempts to guess the possible error pattern applied on the received word to check if the result is a valid codeword. The GRAND with abandonment (GRANDAB) is a hard-detection decoder by limiting the number of generated test error patterns. This article presents efficient algorithms to reduce the number of queries for the GRANDAB when the codes are systematic and cyclic. These methods exploit the properties of the syndrome weight and cyclic codes to significantly improve the decoding latency as the GRANAB corrects up to the error-correcting capability of the code. The VLSI architecture of novel hard-detection GRANDAB is presented, which provides efficient decoding up to 128 bits and can correct up to 3 bit errors at or above the error-correcting capability of the code. The property of syndrome weight is integrated with the dial structure for parallelism, supporting the default mode and the mode of systematic encoding with the known error-correcting capability. When compared to the architecture developed by Abbas et al., the average decoding cycles of two and three errors for the (127, 106) Bose–Chaudhuri–Hocquenghem (BCH) code are improved by 30.90% and 48.63%, respectively. For the (128, 96) cyclic redundancy check (CRC) code, the average decoding cycles of two and three errors are reduced by 45.19% and 65.71%, respectively. At the signal-to-noise ratio (SNR) of 5 dB, the average latencies for decoding (128, 96) CRC and (127, 106) BCH codes are improved by 21.34% and 12.29%. The worst-case latency for decoding the (127, 113) BCH code in the presented design is shorter than that of the work implemented by Riaz et al. with the reduction of 98.02%. The developed hardware architecture is also superior to the original dial-based design by Abbas et al. in terms of area–time (AT) complexity. Shao-I Chu, Syuan-An Ke, Sheng-Jung Liu, Yan-Wei Lin |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2022 | Polynomial Computation Using Unipolar Stochastic Logic and Correlation TechniqueabstractThis paper addresses polynomial computation using unipolar stochastic logic by exploiting correlation between the bit-streams. The AND-OR, double-NAND, OR-AND and double-NOR circuits are presented for polynomials with all positive coefficients whose sum is less than or equal to one by mathematically analyzing the joint probability distribution of coefficient bit-streams. The NAND-AND expansion is also developed for polynomials with alternatively positive and negative coefficients whose absolute values are decreasing by applying the same idea. Unlike the original methods with multiple uncorrelated random number sources (RNSs) for coefficient bit-stream generation, the presented methods only require a single RNS. Since the RNSs take up huge hardware resource in stochastic circuits, the proposed RNS-sharing techniques for polynomial computation result in a significant reduction of hardware complexity. For the factorization technique in the general polynomials, this paper enhances the original stochastic designs for the second-order polynomial and further presents the simple correlation-dependent circuits. Results show that the proposed architectures are superior to the previous ones by reducing the total number of RNSs. Shao-I Chu, Chi-Long Wu, Tu N. Nguyen 0001, Bing-Hong Liu |
IEEE Trans. Computers | 1 |
| 2019 | Challenges, Designs, and Performances of a Distributed Algorithm for Minimum-Latency of Data-Aggregation in Multi-Channel WSNsabstractIn wireless sensor networks (WSNs), the sensed data by sensors need to be gathered, so that one very important application is periodical data collection. There is much effort which aimed at the data collection scheduling algorithm development to minimize the latency. Most of previous works investigating the minimum latency of data collection issue have an ideal assumption that the network is a centralized system, in which the entire network is completely synchronized with full knowledge of components. In addition, most of existing works often assume that any (or no) data in the network are allowed to be aggregated into one packet and the network models are often treated as tree structures. However, in practical, WSNs are more likely to be distributed systems, since each sensor's knowledge is disjointed to each other, and a fixed number of data are allowed to be aggregated into one packet. This is a formidable motivation for us to investigate the problem of minimum latency for the data aggregation without data collision in the distributed WSNs when the sensors are considered to be assigned the channels and the data are compressed with a flexible aggregation ratio, termed the minimum-latency collision-avoidance multiple-data-aggregation scheduling with multi-channel (MLCAMDAS-MC) problem. A new distributed algorithm, termed the distributed collision-avoidance scheduling (DCAS) algorithm, is proposed to address the MLCAMDAS-MC. Finally, we provide the theoretical analyses of DCAS and conduct extensive simulations to demonstrate the performance of DCAS. Tu N. Nguyen 0001, Bing-Hong Liu, Shao-I Chu, Hao-Zhe Weng |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2019 | Design of FSM-Based Function With Reduced Number of States in Integral Stochastic ComputingabstractStochastic computing (SC) is a promising computing paradigm with low power hardware circuitry. This brief proposes a new finite state machine (FSM)-based function implementation in the SC designs. The new architecture allows multiple input stochastic bitstreams to improve the processing latency and precision loss. As compared to previous integral SC design, the proposed algorithm reduces the total number of states needed in the FSM, while maintaining good accuracy performance. The proposed FSM-based construction is also verified by the mathematical analysis. Synthesized results of hardware implementation reveal that the proposed method has competitive advantages over the previous counterparts in terms of area and power consumption. The presented techniques can be applied in a lot of activation function implementations of the deep neural networks. Shao-I Chu, Chen-En Hsieh, Yu-Jung Huang |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2016 | Performance of switching-based partial relay selection scheme for amplify-and-forward cognitive relay networksabstractThe outage probabilities (OPs) of cognitive amplify‐and‐forward systems with the conventional and the switching‐based partial relay selection (PRS) schemes over independent but not identically distributed Rayleigh fading channels are evaluated. The conventional PRS scheme always depends on the instantaneous channel quality of the first hop, while the switching‐based PRS scheme is based on the estimated average channel state information (CSI) of the first and second hops. For the switching‐based PRS scheme, the instantaneous CSI for links with the smaller average channel power in the first and the second hops for each end‐to‐end path is used for relay selection. Thus, the switching‐based PRS scheme counts on the instantaneous CSI of either the first or second hop. The tight lower bounds and asymptotic expressions of the OP are derived. The feedback overheads of both schemes are discussed. Simulation results substantiate the theoretical analysis and also reveal that the switching‐based PRS outperforms the conventional one over the cognitive relay networks in terms of OP. Shao-I Chu, Chih-Yuan Lien, Chien-Liang Chiu |
IET Commun. | 1 |
| 2013 | Efficient decoding of the (23, 12, 7) Golay code up to five errors
Hung-Peng Lee, Shao-I Chu, Hsin-Chiu Chang |
Inf. Sci. | 2 |
| 2013 | Performance analysis and power allocation for decode-and-forward cooperative communications over Rician fading channelabstractABSTRACT This paper derives the asymptotic symbol error rate (SER) and outage probability of decode‐and‐forward (DF) cooperative communications over Rician fading channels. How to optimally allocate the total power is also addressed when the performance metric in terms of SER or outage probability is taken into consideration. Analysis reveals the insights that Rician factor has a great impact on the system performance as compared with the channel variance, and the relay–destination channel quality is of importance. In addition, the source–relay channel condition is irrelevant to the optimal power allocation design. Simulation and numerical evaluation substantiate the tightness of the asymptotic expressions in the high‐SNR regions and demonstrate the accuracy of our theoretical analysis. Copyright © 2011 John Wiley & Sons, Ltd. Shao-I Chu, Hung-Peng Lee, Hsin-Chiu Chang |
Wirel. Commun. Mob. Comput. | 1 |
| 2011 | Comments on "Performance Analysis of Amplify-and-Forward Opportunistic Relaying in Rician Fading"abstractIt has been pointed out that the asymptotic SER formula of the opportunistic relaying derived by B. Mahamis erroneous. The correct SER approximation is therefore presented. It is concluded that opportunistic relaying provides the coding gain in high SNR of${\rm R}!$times less as compared with the repetition-based cooperation. The simulations are also conducted to substantiate the corrected SER formula. Shao-I Chu |
IEEE Signal Process. Lett. | 1 |
| 2010 | High speed decoding of the binary (47, 24, 11) quadratic residue code
Tsung-Ching Lin, Hung-Peng Lee, Hsin-Chiu Chang, Shao-I Chu, Trieu-Kien Truong |
Inf. Sci. | 4 |
| 2007 | Time-of-day Internet-access management by combining empirical data-based pricing with quota-based priority controlabstractAn empirical data-based design methodology is proposed for Internet-access management to improve congestion, uneven usage and fairness, especially during peak hours, over a free-of-charge or flat-rate network. The design methodology combines time-of-day pricing (TDP) with quota-based priority control (QPC). Core to the design methodology are the innovations in characterising user demand and quota-allocation behaviour with respect to time and pricing. In-depth analyses of empirical data reveal distinctive behaviour patterns of myopic and prudent quota allocations over time and both patterns indicate high preference for peak-hour access. The user models adopt general utility functions and capture how pricing affects user behaviour as prudent or myopic. Preference parameters of users' utility over time are then estimated by collecting easily measurable user volumes. The TDP design problem is formulated and solved as a Stackelberg game. Tested on the empirical data of a 5000-user network, the TDP design leads to significant improvements in peak-hour usage and fairness, peak shaving and load balancing over pure QPC. The methodology requires only two simple and short-period data collections from an operational network and takes about 1 min of CPU time for TDP calculation. Results demonstrate the effectiveness of our design methodology when applied to Internet-access environments with frequent changes. Shao-I Chu, Shi-Chung Chang |
IET Commun. | 1 |
| 2004 | Management of abusive and unfair Internet access by quota-based priority control
Tsung-Ching Lin, Yeali S. Sun, Shi-Chung Chang, Shao-I Chu, Yi-Ting Chou, Mei-Wen Li |
Comput. Networks | 4 |