Sheng-De Wang

dblp:76/6381 · DBLP profile ↗
← Back
79ranked-venue papers
11as first author
9since 2021 · last 2025
0000-0001-8856-7850ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 5 first-author · 6 since 2021Systems, architecture and hardware · 20 · 2 first-authorSoftware engineering, systems software and programming languages · 9 · 1 first-authorSecurity and privacy · 8 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-authorDatabases, data management, data science and information retrieval · 6Human-computer interaction and ubiquitous computing · 5 · 1 first-authorComputer networks · 2Graphics, computer vision, multimedia, augmented reality and games · 2Theory of computation · 1
YearPublicationVenuePosition
2025 GWNet: A Lightweight Model for Low-Light Image Enhancement Using Gamma Correction and Wavelet Transform
Ming-Yu Kuo, Sheng-De Wang
ICAART (2)2
2025 HPE-DARTS: Hybrid Pruning and Proxy Evaluation in Differentiable Architecture Search
Hung-I Lin, Lin-Jing Kuo, Sheng-De Wang
ICAART (2)3
2024 VP-DARTS: Validated Pruning Differentiable Architecture Search
Tai-Che Feng, Sheng-De Wang
ICAART (2)2
2023 Anomaly Detection for Multivariate Industrial Sensor Data via Decoupled Generative Adversarial Network
Wei-Chin Chien, Sheng-De Wang
ICAART (3)2
2023 Soft Hybrid Filter Pruning using a Dual Ranking Approach
abstract
Conventional pruning techniques typically focus on evaluating a single structure in the network, such as the convolutional layer or batch normalization layer, to identify pruning targets. However, this approach fails to effectively leverage the potential of all structures within each layer of the network. In order to comprehensively consider the various structures in each layer, we propose a novel method called Soft Hybrid Filter Pruning using a Dual Ranking Approach (DR-SHFP), which builds upon Soft Filter Pruning (SFP) by introducing a dual-ranking approach. DR-SHFP incorporates a ranking system that assigns a rank to each filter in a collaborative manner, taking into account both convolutional layers and batch normalization layers. By simultaneously evaluating both types of layers, our method captures more information from the layer structures, overcoming the limitations of single-structure evaluation. Consequently, DR-SHFP can identify and select filters more effectively for pruning, leading to improved performance. Experimental results demonstrate the effectiveness of DR-SHFP on benchmark datasets such as CIFAR-10, CIFAR-100, and Tiny-ImageNet. The proposed method outperforms other soft pruning methods, showcasing its capability to achieve excellent performance in various settings.
Jen-Chieh Yang, Sheng-De Wang
TrustCom3
2023 A Hybrid Filter Pruning Method Based on Linear Region Analysis
abstract
This study proposes a hybrid filter pruning method based on linear region analysis. Our approach combines the advantages of cluster pruning and norm-based filter pruning by introducing thresholds based on Euclidean distance and norm distance. We also incorporate the linear region analysis approach in neural network architecture search to estimate the performance of trained model architectures. This enables us to efficiently search for the optimal pruned structure with corresponding thresholds for effective model compression.
Chang-Hsuan Hsieh, Jen-Chieh Yang, Hung-Yi Lin, Lin-Jing Kuo, Sheng-De Wang
TrustCom5
2023 MOFP: Multi-Objective Filter Pruning for Deep Learning Models
abstract
The paper proposes a new approach called Multi-Objective Filter Pruning (MOFP), which formulates the filter pruning of deep learning models as a multi-objective optimization problem. The proposed approach applies the Non-Dominated Sorting Genetic Algorithm II to solve the problem and the Asymmetric Gaussian Distribution (AGD) for population initialization. Compared with existing methods, MOFP shows competitiveness in terms of a balance of objectives between compression rates, computing power, and prediction accuracy. In addition, the search result of MOFP is a Pareto Front, which eliminates the need for multiple searches to obtain architectures with different compression rates, significantly improving overall search efficiency. The results show that the use of AGD for population initialization can enhance the search process by effectively exploring the search space, leading to higher quality results.
Jen-Chieh Yang, Hung-I Lin, Lin-Jing Kuo, Sheng-De Wang
TrustCom4
2022 Enhanced Local Gradient Smoothing: Approaches to Attacked-region Identification and Defense
You-Wei Cheng, Sheng-De Wang
ICAART (2)2
2022 Lagrange interpolation-driven access control mechanism: Towards secure and privacy-preserving fusion of personal health records
Yin-Tzu Huang, Dai-Lun Chiang, Tzer-Shyong Chen, Sheng-De Wang, Feipei Lai, Yu-Da Lin
Knowl. Based Syst.4
2020 Quality-Aware Streaming Network Embedding with Memory Refreshing
Hsi-Wen Chen, Hong-Han Shuai, Sheng-De Wang, De-Nian Yang
PAKDD (1)3
2020 A Novel Gaming Video Encoding Process Using In-Game Motion Vectors
abstract
In this paper, we propose a motion preprocessing method for the use in the game application pipeline, which is composed of two stages: the proposed preprocessing method and the High Efficiency Video Coding (HEVC) encoder. The method accepts the object information from the game application, preprocesses the motion vectors of objects, and pass the preprocessed motion data to the HEVC encoder. The HEVC encoder takes the motion data as the initial (or the dedicated) value of motion estimation. Therefore, the traditional diamond search can be skipped and hence increase the encoding performance of the HEVC encoder. In the motion preprocessing method, the following three steps are taken: a coordination system transformation, determining motion vectors for 4 × 4checkerboard blocks [Atomic Block (AB)], and the selection of proper motion vectors for all varieties of prediction units in the encoder. With a focus on the special issues, such as move-out zone and bi-directional prediction, we are able to further optimize the performance of the encoder. We examined two types of 2D gaming scenes in our experiments. The experimental results show that, as compared with the original diamond search method provided by the encoder, our algorithm is able to achieve up to 49.0% time reduction of video encoding. The Bjontegaard Delta bit rate can achieve up to -17.0% in the random_access mode while combining with the x265 encoder and up to -26.2% in the lowdelay mode while combining with the HM-16 encoder.
Chi-Wei Lu, Sheng-De Wang, Shao-Yi Chien
IEEE Trans. Circuits Syst. Video Technol.2
2019 Data Reduction for real-time bridge vibration data on Edge
abstract
In the Internet of Things (IoT) era, with the growing number of data sources, we need to face some challenges such as high cost of the cloud storage caused by large amounts of data. To minimize the communication time and enhance the performance, sending the entire large amount of data is not practical. Thus, it is appropriate to make use of edge computing, or data preprocessing on IoT gateways. In this paper, we propose a data reduction algorithm for the gateway of bridge vibration G-sensors. The data reduction algorithm is based on a pattern system, which is comprised of a pattern library and a pattern classifier. The pattern library is generated by using the K-means clustering method. The results show that the proposed approach is effective in data reduction and outlier detection for bridge vibration data collection on the IoT gateway.
Anthony Chen, Fu-Hsuan Liu, Sheng-De Wang
DSAA3
2018 Detach and Adapt: Learning Cross-Domain Disentangled Deep Representation
abstract
While representation learning aims to derive interpretable features for describing visual data, representation disentanglement further results in such features so that particular image attributes can be identified and manipulated. However, one cannot easily address this task without observing ground truth annotation for the training data. To address this problem, we propose a novel deep learning model of Cross-Domain Representation Disentangler (CDRD). By observing fully annotated source-domain data and unlabeled target-domain data of interest, our model bridges the information across data domains and transfers the attribute information accordingly. Thus, cross-domain feature disentanglement and adaptation can be jointly performed. In the experiments, we provide qualitative results to verify our disentanglement capability. Moreover, we further confirm that our model can be applied for solving classification tasks of unsupervised domain adaptation, and performs favorably against state-of-the-art image disentanglement and translation methods.
Yen-Cheng Liu, Yu-Ying Yeh, Tzu-Chien Fu, Sheng-De Wang, Walon Wei-Chen Chiu, Yu-Chiang Frank Wang
CVPR4
2018 Distributed Continuous Control with Meta Learning on Robotic Arms
abstract
Deep reinforcement learning has been proposed to train the control agent for robotic arms, such as Deep Q-Learning (DQN) and Policy Gradient (PG). The approach of Deterministic Deep Policy Gradient (DDPG) takes the advantage of deterministic policy instead of stochastic policy to further simplify the training process and improve the performance. Reinforcement Learning takes the reward from the environment and trains the underlying control agent to achieve the task. An appropriate reward will get better performance and shorter training time, but it requires the domain knowledge and the method of trial and error to define the appropriate reward function. In this paper, we proposed a method that is based on DDPG and makes use of Prioritized Experience Replay (PER), Asynchronous Agent Learning and Meta Learning. The proposed Meta Learning approach uses multiple distributed learners, called workers, to learn from consecutive previous states and rewards. Simulations are done on 6-DOF (IRB140) and 7-DOF (LBR iiwa 14 R820) robotic arms to train the control agents to reach random targets in the three dimension space. The experiments show that the algorithm we proposed is better than the algorithm using DDPG with specialized reward function on the task success rate and the training speed.
Kuan-Ting Chen, Sheng-De Wang
SMC2
2017 Streaming analytics processing in manufacturing performance monitoring and prediction
abstract
Having the capability to process live streaming data is now the fundamental requisite for the successful realization of Industrial Internet of Things (IoT) and poses huge benefits in terms of increased operational efficiency, lesser costs and diminished risk to the industrial world. The advent of IoT and big data analytics technology offers further opportunities in manufacturing business models and asset management. For industrial manufacturing processes that are typically fast-paced and ridden with sophisticated set of conditions, such on-the-fly, real-time, fine-tuning adjustment suggestions of a predictive nature are challenging to describe. However, when provided properly, streaming analytics is greatly useful in the pursuit of improved industrial performance. We developed a streaming analytics system that used to evaluate stable manufacturing efficiency of multiple production lines simultaneously. This paper illustrates an use case from semiconductor manufacturing industry in Taiwan to present the data-driven applicability of streaming analytics system that enables companies to collect a large number of real-time, heterogeneous plant data with steps of text extraction, causal correlation, statistical modeling, as well as real-time monitoring and anomaly detection, to improve overall equipment effectiveness (OEE) of industrial manufacturing.
Yi-Hsin Wu, Sheng-De Wang, Li-Jung Chen, Cheng-Juei Yu
IEEE BigData2
2015 OpenCL computing on FPGA using multiported shared memory
abstract
This paper focuses on memory access improvements for the OpenCL architecture for FPGAs with the goal of achieving trade-off between performance and required resources. In OpenCL compute units, there is usually a linear relation between computation time and local memory access latency. This latency is normally hidden by increasing the parallel workload. However, with such an approach, the target FPGA device could easily run out of resources. In this work, conflict-free multiported memories are used to minimize local memory access latency. Experiments show that multiported memories can successfully increase computation speed and reduce the required parallel workload for maximum throughput to practical amounts.
Tahsin Turker Mutlugun, Sheng-De Wang
FPL2
2015 Machine Learning Based Hybrid Behavior Models for Android Malware Analysis
abstract
Malware analysis on the Android platform has been an important issue as the platform became prevalent. The paper proposes a malware detection approach based on static analysis and machine learning techniques. By conducting SVM training on two different feature sets, malicious-preferred features and normal-preferred features, we built a hybrid-model classifier to improve the detection accuracy. With the consideration of normal behavior features, the ability of detecting unknown malwares can be improved. The experiments show that the accuracy is as high as 96.69% in predicting unknown applications. Further, the proposed approach can be applied to make confident decisions on labeling unknown applications. The experiment results show that the proposed hybrid model classifier can label 79.4% applications without false positive and false negative occurred in the labeling process.
Hsin-Yu Chuang, Sheng-De Wang
QRS2
2013 An efficient multicharacter transition string-matching engine based on the aho-corasick algorithm
abstract
A string-matching engine capable of inspecting multiple characters in parallel can multiply the throughput. However, the space required for implementing a matching engine that can process multiple characters in parallel generally grows exponentially with respect to the characters to be processed in parallel. Based on the Aho-Corasick algorithm (AC-algorithm), this work presents a novel multicharacter transition Nondeterministic Finite Automaton (NFA) approach, called multicharacter AC-NFA , to allow for the inspection of multiple characters in parallel. This approach first converts an AC-trie to an AC-NFA by allowing for the simultaneous activation of multiple states and then converts the AC-NFA to a k -character AC-NFA by an algorithm with concatenation operations and assistant transitions. Additionally, the alignment problem, which occurs while multiple characters are being inspected in parallel, is solved using assistant transitions. Moreover, a corresponding output is provided for each inspected character by introducing priority multiplexers to determine the final matching outputs during implementation of the multicharacter AC-NFA. Consequently, the number of derived k -character transitions grows linearly with respect to the number k . Furthermore, the derived multicharacter AC-NFA is implemented on FPGAs for evaluation. The resulting throughput grows approximately 14 times and the hardware cost grows about 18 times for 16-character AC-NFA implementation, as compared with that for 1-character AC-NFA implementation. The achievable throughput is 21.4Gbps for the 16-character AC-NFA implementation operating at a 167.36MHz clock.
Chien-Chi Chen, Sheng-De Wang
ACM Trans. Archit. Code Optim.2
2011 A Data Parallel Approach to XML Parsing and Query
abstract
Data-parallel XML parsing has a crucial problem in partitioning XML documents. Existing approaches need a pre-parse step to determine the partitions. In this paper, we propose a direct parallel method to solve this problem without pre-parsing. In the direct parallel method, we directly start the parallel parsing by finding the "light tower", which is a particular character with some exceptions, called clues. We handle the exceptions by watching the clues and reparsing the partition if it is required in the parsing stage. We also propose a non-synchronized splitter approach to the parallel XML querying using XPath expressions. In the non-synchronized splitter approach, we split an XPath expression into pieces to be executed by threads and we use a data structure, called the ancestor table, to help each thread handle its part of XPath expression independently without communications between threads. Our experiments show that our approach scales well from small sized files to huge sized files.
Cheng-Han You, Sheng-De Wang
HPCC2
2011 Algorithms and Hardware Architectures for Variable Block Size Motion Estimation
Sheng-De Wang, Chih-Hung Weng
UIC1
2010 An in-place search algorithm for the resource constrained scheduling problem during high-level synthesis
abstract
We propose an in-place search algorithm for computing the exact solutions to the resource constrained scheduling problem. This algorithm supports operation chaining, pipelining and multicycling in the underlying scheduling problem. Based on two lower-bound estimation mechanisms that are capable of predicting the criterion values of search nodes represented by partially scheduled data flow graphs, the proposed algorithm can effectively prune the nonpromising search space and finds the optimum usually several times faster than existing techniques. As opposed to existing search-based scheduling techniques whose space complexity is squared or exponential in the search depth, our approach requires only a constant storage space during the traversal of the search tree. The low space complexity is accomplished by using a combination-generating algorithm, which leads our approach to visit search nodes in such a way that each one is obtained by making only a small change to its sibling without keeping any parent nodes in memory. Experimental results on several well known benchmarks with varying resource constraints show the effectiveness of the proposed algorithm.
Cheng-Juei Yu, Yi-Hsin Wu, Sheng-De Wang
ACM Trans. Design Autom. Electr. Syst.3
2009 An enhanced uplink scheduling algorithm for video traffic transmission in IEEE 802.16 BWA systems
abstract
One of the major challenges in transmitting real time video data in broadband wireless access (BWA) systems is that the bit rate of the video traffic varies from time to time, and hence the bandwidth utilization is low. To cope with the difficulties, we propose an enhanced uplink scheduling algorithm for variable bit rate video (VBR) traffic transmission in IEEE 802.16 BWA systems. In the proposed algorithm, the base station (BS) assigns uplink bandwidth to the video user by considering the traffic state transitions. The uplink bandwidth is divided into several intervals and each interval represents a traffic state. Only when the traffic state is changed, the bandwidth request process is incurred. Also, by using two reserved bits in the generic MAC header of IEEE 802.16 BWA systems as piggyback bits, the information of video traffic state transition can be sent to the BS without extra overhead. Simulations conducted with QualNet show that the proposed algorithm provides better performance in that the bandwidth waste ratio of our proposed algorithm is less than that of the conventional algorithms.
Yeong-Sheng Chen, Der-Jiunn Deng, Yu-Ming Hsu, Sheng-De Wang
IWCMC4
2009 Choosing the kernel parameters for support vector machines by the inter-cluster distance in the feature space
Kuo-Ping Wu, Sheng-De Wang
Pattern Recognit.2
2008 Crossing Heterogeneous Grid Systems with a Single Sign-On Scheme Based on a P2P Layer
abstract
Grid systems provide a standard platform to share resources, such as computing facilities and backup services, between research communities and commercial organizations. Nowadays there are plenty of grids among different institutions, and many of them are probably heterogeneous to others because they are constructed by different types of middleware. Enabling all these grids to work as a single grid can provide users with even more computing power and services. Because of varieties of middleware and certification authority, cooperating across different grids is difficult. We will present a single sign-on scheme such that heterogeneous grids can cooperate through it on top of existing middleware. The proposed scheme focuses on achieving interoperability among different grids above a P2P layer. By means of the scheme and MyProxy X.509 credentials services, we demonstrate our solution to execute jobs on heterogeneous grids with single sign-on.
Yuan-Chin Wen, Sheng-De Wang
APSCC3
2008 Adaptive Neighbor Caching for Fast BSS Transition Using IEEE 802.11k Neighbor Report
abstract
Handoff latency is a severe bottleneck impacting the service continuity for voice and multimedia applications in WLAN. IEEE 802.11k neighbor report defines the neighbor APs which are potential transition candidates for the roaming target. But the selection method for the roaming target AP is left undefined. Several schemes have been proposed for fast handoff with neighbor APpsilas information. However, these schemes result in huge redundant transition messages overheads in the WLAN and require high computing power for the AAA (Authentication, Authorization, and Access control) server. In this paper, we propose an adaptive neighbor caching (ANC) method to achieve higher handoff prediction accuracy for selecting proper candidate APs in the Neighbor Report. An adaptive predictability index is introduced for selecting those potential roaming APs, which can mitigate the scanning latency and the pre-authentication key distribution message overhead in the WLAN as well as computing loading for the AAA server. Simulation results present up to 83.5% of transition messages are reduced in comparison to the Proactive Neighbor Caching (PNC), 56.4% of candidate AP selection accuracy is improved and 37.5% of transition messages are reduced in comparison to the Selective Neighbor Caching (SNC).
Ching-Hwa Yu, Michael Pan, Sheng-De Wang
ISPA3
2007 The Modified Grid Location Service for Mobile Ad-Hoc Networks
Hau-Han Wang, Sheng-De Wang
GPC2
2007 Energy Saving Based on CPU Voltage Scaling and Hardware Software Partitioning
abstract
We examine the possible energy savings by mapping critical software functions from a microprocessor to configurable logics. A system-on-a-chip containing configurable logic is now commercially available. The configurable logic is typically intended to implement peripherals and co-processors without increasing chip count. We show that reduced software energy is an extra significant benefit, making such chips even more useful. We identify critical software functions of an application and implement them in the configurable logic such that the application can complete sooner, allowing us to put the system in a low-power state for longer periods, thus reducing energy. We use estimation-based approach for a hypothetical device having a 32-bit MlPS-extension processor plus on-chip configurable logic, yielding energy savings of 40%, increasing to 54% assuming voltage scaling.
Chia Hsiang Hsu, Cheng-Juei Yu, Sheng-De Wang
PRDC3
2006 Choosing the Kernel parameters of Support Vector Machines According to the Inter-cluster Distance
abstract
This paper proposes using the inter-cluster distance between class means in the feature space to help choose parameters for a kernel function when training a support vector machine (SVM). With the proposed method, the square values of the distance between the two class means of the training data in different feature spaces are calculated. These values are used as the indexes of data separation in the feature space. The experiment results show that the proposed method can choose the parameters close to the best ones. As a result, fewer possible values of the kernel parameters are required to be tested when training an SVM, and thus the training time of total training process can be significantly shortened.
Kuo-Ping Wu, Sheng-De Wang
IJCNN2
2006 A Dependable Outbound Bandwidth Based Approach for Peer to Peer Media Streaming
abstract
A fundamental problem in peer-to-peer streaming is how to select peers with desired media data so that the best possible streaming quality can be maintained. In this paper, we propose an outbound bandwidth based streaming model in which peers are layered according to their offered outbound bandwidth and are permitted to request data of peers from upper layers peers only. Based on the layered approach, a media data assignment algorithm for the subset of media data is presented to select qualified sending peers to ensure that they are received before their scheduled playback time. We also present two resolutions for request conflicts, which arise when there are more than one peer simultaneously requesting data from the same sending peer that can't afford outbound bandwidth for all requests. We evaluated the proposed streaming model through simulations. Experimental results show that streaming quality of the proposed streaming model is excellent and the properties of scalability as well as robustness are obtained even in a highly dynamic environment where peers join and leave frequently
Zheng Yi Huang, Sheng-De Wang
PRDC2
2006 Choosing the Parameters of 2-norm Soft Margin Support Vector Machines According to the Cluster Validity
abstract
Determining the kernel and error penalty parameters for support vector machines (SVMs) is very problem-dependent in practice. This paper proposes using a cluster validation index in the feature space to help choose parameters for training 2-norm soft margin support vector machines. With the proposed method, the kernel parameters and the penalty parameter of the error term in the 2-norm soft margin SVM are considered to be the parameters of an alternative kernel for a hard margin SVM. Thus the values of cluster validation index can be calculated in the feature spaces which are defined by the kernels with the parameters. The cluster validation index shows whether the data are well-separated in a feature space, so it can be used to determine whether a combination of the kernel parameters leads to a feature space in which the data are easy to be classified. It guides the search of parameters toward a good testing accuracy, so the search range of the parameters is confined to a small region, and the parameters selecting time of the SVM training process can be shortened.
Kuo-Ping Wu, Sheng-De Wang
SMC2
2006 Transmission Range Designation Broadcasting Methods for Wireless Ad Hoc Networks
Jian-Feng Huang, Sheng-Yan Chuang, Sheng-De Wang
UIC3
2005 Local Repair Mechanisms for On-Demand Routing in Mobile Ad hoc Networks
abstract
With the dynamic and mobile nature of ad hoc wireless networks, links may fail due to topological changes by mobile nodes. As the degree of mobility increases, the wireless network would suffer more link errors. Ad hoc routing protocols that use broadcast to discover routes may become inefficient due to frequent failures of intermediate connections in an end-to-end communication. When an intermediate link breaks, it is beneficial to discover a new route locally without resorting to an end-to-end route discovery. Based on the concept of localizing the route request query, we propose an efficient approach to repair error links quickly. The approach can apply to the ad hoc on-demand distance vector (AODV) routing protocol. As an enhancement to AODV, the proposed approach leads to two routing protocols, called AODV-LRQ and AODV-LRT, which are aimed to efficiently repair the link errors. To evaluate the effects of the route repair, we define a factor, called bonus gain, as the ratio between the throughput increment to the routing overhead increment. Simulation results show that the proposed methods can get high bonus gain, that is, it can maintain the throughput as well as reduce the routing overheads.
Michael Pan, Sheng-Yan Chuang, Sheng-De Wang
PRDC3
2004 Japster: An Improved Peer-to-Peer Network Architecture
Sheng-De Wang, Hsuen-Ling Ko, YungYu Zhuang
EUC1
2004 Competitive algorithms for the clustering of noisy data
Tai-Ning Yang, Sheng-De Wang
Fuzzy Sets Syst.2
2004 Training algorithms for fuzzy support vector machines with noisy data
Chun-fu Lin, Sheng-De Wang
Pattern Recognit. Lett.2
2004 An adaptive H∞ controller design for bank-to-turn missiles using ridge Gaussian neural networks
abstract
A new autopilot design for bank-to-turn (BTT) missiles is presented. In the design of autopilot, a ridge Gaussian neural network with local learning capability and fewer tuning parameters than Gaussian neural networks is proposed to model the controlled nonlinear systems. We prove that the proposed ridge Gaussian neural network, which can be a universal approximator, equals the expansions of rotated and scaled Gaussian functions. Although ridge Gaussian neural networks can approximate the nonlinear and complex systems accurately, the small approximation errors may affect the tracking performance significantly. Therefore, by employing the Hinfinity control theory, it is easy to attenuate the effects of the approximation errors of the ridge Gaussian neural networks to a prescribed level. Computer simulation results confirm the effectiveness of the proposed ridge Gaussian neural networks-based autopilot with Hinfinity stabilization.
Chuan-Kai Lin, Sheng-De Wang
IEEE Trans. Neural Networks2
2002 A Packet-Based Caching Proxy with Loss Recovery for Video Streaming
abstract
With the popularity of broadband networks, video streaming is growing rapidly in the Internet. By deployment of caching proxies, backbone bandwidth can be saved significantly. In this paper, we propose a packet-based caching architecture for video streaming. The proposed caching scheme is based on streamed packets instead of video files and the consideration of packet loss recovery. We also propose an effective cache replacement algorithm, PLFU, and evaluate it through simulation.
Kuan-Sheng Hsueh, Sheng-De Wang
PRDC2
2002 Fuzzy support vector machines
abstract
A support vector machine (SVM) learns the decision surface from two distinct classes of the input points. In many applications, each input point may not be fully assigned to one of these two classes. In this paper, we apply a fuzzy membership to each input point and reformulate the SVMs such that different input points can make different contributions to the learning of decision surface. We call the proposed method fuzzy SVMs (FSVMs).
Chun-fu Lin, Sheng-De Wang
IEEE Trans. Neural Networks2
2001 Jato: A Compact Binary File Format for Java Class
abstract
Java has been a very important programming language, especially with its cross-platform characteristics, but the CLASS file format defined in the Java Virtual Machine (JVM) specification contains many redundancies and replications of information. These redundancies most come from the "constant pool" of a CLASS file. We propose a compact binary file format, called Jato, and its associated archive format, called Jatar, for the Java system. Using these two formats, many of the redundancies can be removed. We didn't utilize any text compression technique in the proposed formats, so they do not sacrifice the loading speed and are thus very suitable for use in embedded environments. We've also implemented a class loader that is capable of loading the Jato files into a regular JVM. Using this approach, we show that the Jato file format is effective and promising, while still keeping the cross-platform features of Java.
Sheng-De Wang, Yuhder Lin
ICPADS1
2001 Fault-Tolerant Routing in Two-Dimensional Mesh Networks with Less-Restricted Fault Patterns
abstract
Wormhole routing in networks is prone to deadlocks. Several techniques have been provided to solve the problem, including virtual channels and restriction on the fault patterns. We relax the fault patterns to one that does not contain the column-surrounded fault pattern. In our routing scheme, the concept of off-node is proposed to help messages leave the visited f-ring at an appropriate node such that no message encounters the same f-ring more than once, and therefore never gets trapped in faulty blocks. Virtual channels are simulated on physical channels to avoid cyclic dependence on channels.
Sheng-De Wang, Pao Hwa Sui
PRDC1
2001 Fingerprint feature reduction by principal Gabor basis function
Chih-Jen Lee, Sheng-De Wang
Pattern Recognit.2
2001 A rotation invariant printed Chinese character recognition system
Tai-Ning Yang, Sheng-De Wang
Pattern Recognit. Lett.2
2000 Adaptive tuning of the fuzzy controller for robots
Sheng-De Wang, Chuan-Kai Lin
Fuzzy Sets Syst.1
2000 A fault-tolerant routing algorithm for wormhole routed meshes
Pao Hwa Sui, Sheng-De Wang
Parallel Comput.2
2000 Fuzzy auto-associative neural networks for principal component extraction of noisy data
abstract
In this paper, we propose a fuzzy auto-associative neural network for principal component extraction. The objective function is based on reconstructing the inputs from the corresponding outputs of the auto-associative neural network. Unlike the traditional approaches, the proposed criterion is a fuzzy mean squared error.We prove that the proposed objective function is an appropriate fuzzy formulation of auto-associative neural network for principal component extraction. Simulations are given to show the performances of the proposed neural networks in comparison with the existing method.
Tai-Ning Yang, Sheng-De Wang
IEEE Trans. Neural Networks Learn. Syst.2
2000 Adaptive and Deadlock-Free Routing for Irregular Faulty Patterns in Mesh Multicomputers
abstract
Message routing achieves the internode communication in parallel computers. A reliable routing is supposed to be deadlock-free and fault-tolerant. While many routing algorithms are able to tolerate a large number of faults enclosed by rectangular faulty blocks, there is no existing algorithm that is capable of handling irregular faulty patterns for wormhole networks. In this paper, a two-staged adaptive and deadlock-free routing algorithm called "Routing for Irregular Faulty Patterns" (RIFP) is proposed. It can tolerate irregular faulty patterns by transmitting messages from sources or to destinations within faulty blocks via multiple "intermediate nodes." A method employed by RIFP is first introduced to generate intermediate nodes using the local failure information. By its aid, two communicating nodes can always exchange their data or intermediate results if there is at least one path between them. RIFP needs two virtual channels per physical link in meshes.
Ming-Jer Tsai, Sheng-De Wang
IEEE Trans. Parallel Distributed Syst.2
1999 A secure and practical electronic voting scheme
Wei-Chi Ku, Sheng-De Wang
Comput. Commun.2
1999 Fuzzy system identification using an adaptive learning rule with terminal attractors
Chuan-Kai Lin, Sheng-De Wang
Fuzzy Sets Syst.2
1999 Fuzzy system modeling using linear distance rules
Sheng-De Wang, Chien-Hui Lee
Fuzzy Sets Syst.1
1999 Fault-Tolerant Wormhole Routing in Two-Dimensional Mesh Networks with Convex Faults
Pao Hwa Sui, Sheng-De Wang
Inf. Sci.2
1999 Robust algorithms for principal component analysis
Tai-Ning Yang, Sheng-De Wang
Pattern Recognit. Lett.2
1998 Adaptive and Fault-Tolerant Routing with 100% Node Utilization for Mesh Multicomputer
abstract
We propose an adaptive and deadlock-free routing algorithm to tolerate irregular faulty patterns using two virtual channels per physical link. It can improve the node utilization up to 100%. When a node becomes faulty or recovered, the central control unit constructs a directed path graph which is used for generating the intermediate nodes of the message path. Thus a message can be transmitted from sources or to destinations within faulty blocks via a set of "intermediate nodes". Our method requires the global failure information if the central control unit is not available.
Sheng-De Wang, Ming-Jer Tsai
ICPADS1
1998 A self-organizing fuzzy control approach for bank-to-turn missiles
Chuan-Kai Lin, Sheng-De Wang
Fuzzy Sets Syst.2
1998 Perceptron-perceptron net
Sheng-De Wang, Tsong-Chih Hsu
Pattern Recognit. Lett.1
1998 A Fully Adaptive Routing Algorithm for Dynamically Injured Hypercubes, Meshes, and Tori
abstract
Unicast V is a progressive, misrouting algorithm for packet or virtual cut-through networks. A progressive protocol forwards a message at an intermediate node if a nonfaulty profitable link is available and waits, deroutes, or aborts otherwise. A misrouting protocol uses both profitable and nonprofitable links at each node; thus, a message can move farther away from its destination at some steps. Unicast V is simple for hardware implementation, requires a very small message overhead, and makes routing decisions by local failure information only. However, it is claimed to be partially adaptive and to be able to tolerate static faults in hypercubes only. In this paper, we uncover some new features of Unicast V: (1) it is fully-adaptive, (2) it also applies to meshes and tori, and (3) it can tolerate dynamic faults by careful implementation. In addition, we also provide bounds on the performance of the algorithm.
Ming-Jer Tsai, Sheng-De Wang
IEEE Trans. Parallel Distributed Syst.2
1997 The K1-Map Reduction for Pattern Classifications
abstract
A shortcut hand-reduction method known as the Karnaugh map (K map) is an efficient way of reducing Boolean functions to a minimum form for the purpose of minimizing hardware requirements. In this paper, by applying the prime group and the essential prime group concepts of the K maps to pattern classification problems, the K1-map reduction method is proposed. The K1-map reduction method can be used to design restricted Coulomb energy networks and to determine the number of hidden units problems in a systematic manner.
Tsong-Chih Hsu, Sheng-De Wang
IEEE Trans. Pattern Anal. Mach. Intell.2
1997 An Improved Algorithm for Fault-Tolerant Routing in Hypercubes
abstract
Boppana and Chalasani (1995) present simple methods to enhance wormhole routing algorithms for fault-tolerance in meshes, In this brief paper, we note that one of their algorithms, f-cube4, can further be improved. In particular, we show that only three virtual channels per physical channel are sufficient for tolerating multiple faulty regions. We also show that our scheme does not lead to deadlock with any combination of faults, while f-cube4 leads to deadlocks for some extreme combinations of fault regions.
Pao Hwa Sui, Sheng-De Wang
IEEE Trans. Computers2
1997 k-winners-take-all neural net with Θ(1) time complexity
abstract
In this article we present a k-winners-take-all (k-WTA) neural net that is established based on the concept of the constant time sorting machine by Hsu and Wang. It fits some specific applications, such as real-time processing, since its Theta(1) time complexity is independent to the problem size. The proposed k-WTA neural net produces the solution in constant time while the Hopfield network requires a relatively long transient to converge to the solution from some initial states.
Tsong-Chih Hsu, Sheng-De Wang
IEEE Trans. Neural Networks2
1996 High-Performance Low-Cost Non-Blocking Switch for ATM
abstract
The internal conflict and the output contention problem of the ATM switch are considered. We propose an ATM switch based on the Benes switch and the concept of input port controller. Following a simple scheduling algorithm we can obtain a set of cells for all input ports. When going through the Benes switch, no internal conflicts and output contentions will occur. Although its cost is very low, its maximum throughput is very high even for large switch size and full traffic load. After applying the scheduling algorithm, a series of routing bits for each cell can be found to support a self-routing function for the Benes switch.
Jeen-Fong Lin, Sheng-De Wang
INFOCOM2
1996 A self-organizing adaptive fuzzy controller
Chien-Hui Lee, Sheng-De Wang
Fuzzy Sets Syst.2
1996 Tiling Nested Loops into Maximal Rectangular Blocks
Yeong-Sheng Chen, Sheng-De Wang, Chien-Min Wang
J. Parallel Distributed Comput.2
1996 Transformations of star-delta and delta-star reliability networks
abstract
It is an open problem whether there exists an equivalent transformation between star reliability networks and delta reliability networks. We provide: criteria to check if there exists any equivalent transformation when a delta-network or a star-network is given, and a criterion about a special star-network which has an unreliable center node and give an example to explain how to use it.
Sheng-De Wang, Cha-Hon Sun
IEEE Trans. Reliab.1
1995 Comments on "On the design of feedforward neural networks for binary mapping [1]"
Tsong-Chih Hsu, Sheng-De Wang
Neurocomputing2
1995 Ring-Connected Networks and Their Relationship to Cubical Ring Connected Cycles and Dynamic Redundancy Networks
abstract
Reviews a 1-fault-tolerant (1-ft) hypercube model with degree 2r: the ring-connected network (RCN), which has the lowest degree among all 1-ft, one-spare node, r-dimensional hypercube architectures yet discovered. Then, we propose a constant-time reconfiguration algorithm via an add-and-modulo automorphism. Furthermore, by introducing the equivalence from hypercubes to cube-connected cycles (CCCs) and to butterflies (BFs), we find that there is also a corresponding equivalence from RCNs to cubical ring-connected cycles (CRCCs) and to dynamic redundancy networks (DRNs). From this fact, we find that once a symmetric fault-tolerant structure has been discovered for one of the three models, then it can be applied directly to the other hypercubic networks. Applying the technique, we find a degree-6, 1-ft Benes network. We think that more attention should be paid to the strong relationship between hypercubes, CCCs and BFs. Finally, from this equivalence relationship we propose three new bounded-degree k-ft models: k-ft CCCs, k-ft BFs and k-ft Benes networks.>
Isaac Yi-Yuan Lee, Sheng-De Wang
IEEE Trans. Parallel Distributed Syst.2
1994 Broadcasting on Faulty Hypercubes
abstract
In this paper we propose a method for constructing the maximum number of edge-disjoint spanning trees (in the directed sense) on a hypercube with arbitrary one faulty node. Each spanning tree is of optimal height. By taking the common neighbor of the roots of these edge-disjoint spanning trees as the new root and reversing the direction of the directed link from each root to the new root, a spanning graph, consisting of n-1 edge-disjoint spanning trees of optimal height is formed. Broadcasting based on the spanning graph has an optimal bandwidth utilization and an optimal latency.
Pao Hwa Sui, Sheng-De Wang, Isaac Yi-Yuan Lee
ICPADS2
1994 Compiler techniques for maximizing fine-grain and coarse-grain parallelism in loops with uniform dependences
abstract
In this paper, an approach to the problem of exploiting parallelism within nested loops is proposed. The proposed method first finds out all the initially independent computations, and then, based on them, identifies the valid partitioning bases to partition the entire iteration space of the loop nest. Because the shape of the iteration space is taken into account, pseudo-dependence relations are eliminated and hence more parallelism is exploited. Our approach provides a systematic method to maximize the degree of fine- or coarse-grain parallelism and is free from the open question of how to combine different loop transformations for the goal of maximizing parallelism. It is also shown that our approach can exploit more parallelism than other related work and have many advantages over them.
Yeong-Sheng Chen, Sheng-De Wang, Chien-Min Wang
International Conference on Supercomputing2
1994 Comments on "Distributed Algorithms for Network Recognition Problems"
abstract
K.V.S. Ramarao (1989) proposed distributed algorithms to recognize five network topologies. In these comments, we first use a counter example to comment that the approach by Ramarao of recognizing a star topology is incomplete. Then, we propose two modified approaches to do the work.>
Cha-Hon Sun, Sheng-De Wang
IEEE Trans. Computers2
1993 A new transformation method for nondominated coterie design
David Shou, Sheng-De Wang
Inf. Sci.2
1993 Performance modeling and analysis of load balancing policies with priority queueing
Rong-Chau Liu, Sheng-De Wang
J. Syst. Softw.2
1992 A Well-informed Approach to Distributed Task Assignment
Chiun-Chieh Hsu, Sheng-De Wang, Te-Son Kuo
Comput. J.2
1992 Heuristic task assignment for distributed computing systems
Chiun-Chieh Hsu, Sheng-De Wang
Inf. Sci.2
1992 A hybrid scheme for efficiently executing nested loops on multiprocessors
Chien-Min Wang, Sheng-De Wang
Parallel Comput.2
1992 Efficient Processor Assignment Algorithms and Loop Transformations for Executing Nested Parallel Loops on Multiprocessors
abstract
An important issue for the efficient use of multiprocessor systems is the assignment of parallel processors to nested parallel loops. It is desirable for a processor assignment algorithm to be fast and always generate an optimal processor assignment. The paper proposes two efficient algorithms to decide the optimal number of processors assigned to each individual loop. Efficient parallel counterparts of these two algorithms are also presented. These algorithms not only always generate an optimal processor assignment, but also are much faster than the exiting optimal algorithm in the literature. The paper discusses improving the performance of parallel execution by transforming a nested parallel loop into a semantically equivalent one. Three loop transformations are investigated. It is observed that, in most cases, the parallel execution time is improved after applying these transformations.>
Chien-Min Wang, Sheng-De Wang
IEEE Trans. Parallel Distributed Syst.2
1991 Compiler techniques to extract parallelism within a nested loop
abstract
By analyzing the dependences between instances, the authors propose a new compiler technique called cycle breaking for parallelizing nested loops. For a single dependence cycle, it extracts more parallelism than two similar techniques. Several versions of cycle braking are presented to extract parallelism within a nested loop by linearizing its multidimensional iteration space. It is observed that the order in which loops are linearized can dramatically affect the parallelism extracted by cycle breaking. Two loop reordering transformations are investigated. Methods to find the optimal linearization order of loops are proposed. These techniques can enhance the parallelism of a nested loop.>
Chien-Min Wang, Sheng-De Wang
COMPSAC2
1991 A Scheduling Scheme for Efficiently Executing Hybrid Nested Loops
Chien-Min Wang, Sheng-De Wang
ICPP (2)2
1990 A neural network approach for Chinese character recognition
abstract
A neural network model adapted from Fukushima's Neocognitron is applied to the pattern recognition of Chinese characters. Chinese characters are well known for their nonalphabetic and two-dimensional features. Chinese characters are viewed as two-dimensional patterns, and some simple subpatterns, called primitives, are identified for hierarchical processing in a neural network. Simulation results show the feasibility and the effectiveness of this approach
Sheng-De Wang, Chung-Chi Pan
IJCNN1
1990 Self-adaptive neural architectures for control applications
abstract
The potential use of the modeling capacity of neural networks for control applications is examined. A neuromorphic controller, called the self-adaptive neural controller (SANC), is designed by utilizing the neural modeling capacity. The results of this approach reveal at least two expected benefits: learning from example and dynamical adaptation. With the learning from example ability, SANC is essentially application-independent, even if the plant considered is too complex or too uncertain to be modeled by precise mathematical expressions. With the dynamical adaptation feature, SANC is shown to be robust, adaptive. and capable of learning, even if the environment varies too much to be controlled by traditional controllers
Sheng-De Wang, Hackerd M. S. Yeh
IJCNN1
1990 Structured partitioning of concurrent programs for execution on multiprocessors
Chien-Min Wang, Sheng-De Wang
Parallel Comput.2
1989 Minimization of task turnaround time for distributed systems
abstract
The problem of assigning a partitioned task to a distributed computing system is studied. Considering communication overhead and idle time, it is possible to develop a mathematical model to describe the cost function, which is defined to evaluate the task turnaround time, under a general model of distributed computing systems. Task assignment is formulated as a DU-mapping, which maps a directed acyclic task graph onto an undirected system graph. The search of optimal DU-mapping is NP-complete and is transformed into a state-space search problem. An approach called critical sink underestimate is developed to attain an optimal DU-mapping. This approach allows the most nodes in the state-space tree to be pruned. Experimental results reveal that this method performs very well due to its close evaluation to the real cost.>
Chiun-Chieh Hsu, Sheng-De Wang, Te-Son Kuo
COMPSAC2