Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Sheng-Wen Bai

dblp:03/1859 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
0since 2021 · last 2003
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Parallel and multicore computing · 66% Memory systems · 22% High-performance computing · 12%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing
multicomputer
0.022000
A Generalized Basic-Cycle Calculation Method for Efficient Array Redistribution · IEEE Trans. Parallel Distributed Syst. 2000
A Basic-Cycle Calculation Technique for Efficient Dynamic Data Redistribution · IEEE Trans. Parallel Distributed Syst. 1998
Parallel and multicore computing › data distribution
array redistribution
0.012000
A Generalized Basic-Cycle Calculation Method for Efficient Array Redistribution · IEEE Trans. Parallel Distributed Syst. 2000
Memory systems › memory access optimization
data movement reduction
0.012000
A Generalized Basic-Cycle Calculation Method for Efficient Array Redistribution · IEEE Trans. Parallel Distributed Syst. 2000
Parallel and multicore computing › data distribution
data redistribution
0.011998
A Basic-Cycle Calculation Technique for Efficient Dynamic Data Redistribution · IEEE Trans. Parallel Distributed Syst. 1998
High-performance computing
scientific computing
0.012000
A Generalized Basic-Cycle Calculation Method for Efficient Array Redistribution · IEEE Trans. Parallel Distributed Syst. 2000
High-performance computing
distributed memory systems
0.011998
A Basic-Cycle Calculation Technique for Efficient Dynamic Data Redistribution · IEEE Trans. Parallel Distributed Syst. 1998

Methods — techniques the papers use, named apart from their topics

cost modeling · 0.0basic-cycle calculation · 0.0performance measurement · 0.0
YearPublicationVenuePosition
2003 An Agent and Profile Management System for Mobile Users and Service Providers
abstract
With the development of mobile devices, people can execute many applications on their personal devices. However, many limitations of the mobile devices, e.g., CPU, memory, power supply, etc., make them impossible to be completely the same as the desktop PCs. In this paper we present an integrated management architecture for thin-client purpose called Agent and Profile Management System (APMS). The users only need to download a simple service agent to their mobile device and install it, and then the service agent will connect to the corresponding service provider send a request to the back-end server After receiving the request, the back-end server will execute appropriate processes and return a response to the mobile device. Most of the procedures are accomplished on the server side. Therefore, the workload of mobile devices is relatively low and the cost of mobile devices can be effectively reduced.
Tsung-Chuan Huang, Chu-Sing Yang, Sheng-Wen Bai, Szu-Hsuan Wang
AINA3
2003 Hierarchical Grown Bluetrees (HGB) - An Effective Topology for Bluetooth Scatternets
Tsung-Chuan Huang, Chu-Sing Yang, Chao-Chieh Huang, Sheng-Wen Bai
ISPA4
2002 Data Redistribution Using MPI User-Defined Types
abstract
In many parallel programs, run-time data redistribution is usually required to enhance data locality and reduce remote memory access on the distributed memory multicomputers. Recently researches in data redistribution algorithm have become very mature. The time required to generate data sets and processor sets is much lesser then before. That means packing/unpacking becomes a relatively heavy cost in the redistribution. In this paper we present methods to perform BLOCK-CYCLIC(s) to BLOCK-CYCLIC(t) redistribution using MPI user-defined types. In this approach, we can reduce the requirement of memory buffers and avoid unnecessary data-movement. The theoretical models are presented to determine the best method for redistribution. To evaluate the performance of the proposed methods, we have implemented our methods on an IBM SP2 parallel machine. The experimental results show that this approach can obviously improve the performance of redistribution in most cases.
Chu-Sing Yang, Sheng-Wen Bai
CW2
2000 A Generalized Basic-Cycle Calculation Method for Efficient Array Redistribution
abstract
In many scientific applications, dynamic array redistribution is usually required to enhance the performance of an algorithm. In this paper, we present a generalized basic-cycle calculation (GBCC) method to efficiently perform a BLOCK-CYCLIC(s) over P processors to BLOCK-CYCLIC(t) over Q processors array redistribution. In the GBCC method, a processor first computes the source/destination processor/data sets of array elements in the first generalized basic-cycle of the local array it owns. A generalized basic-cycle is defined as lcm(sP, tQ)/(gcd(s,t)/spl times/P) in the source distribution and lcm(sP, tQ)/(gcd(s,t)/spl times/Q) in the destination distribution. From the source/destination processor/data sets of array elements in the first generalized basic-cycle, we can construct packing/unpacking pattern tables to minimize the data-movement operations. Since each generalized basic-cycle has the same communication pattern, based on the packing/unpacking pattern tables, a processor can pack/unpack array elements efficiently. To evaluate the performance of the GBCC method, we have implemented this method on an IBM SP2 parallel machine, along with the PITFALLS method and the ScaLAPACK method. The cost models for these three methods are also presented. The experimental results show that the GBCC method outperforms the PITFALLS method and the ScaLAPACK method for all test samples. A brief description of the extension of the GBCC method to multidimensional array redistributions is also presented.
Ching-Hsien Hsu, Sheng-Wen Bai, Yeh-Ching Chung, Chu-Sing Yang
IEEE Trans. Parallel Distributed Syst.2
1998 A Generalized Basic Cycle Calculation Method for Efficient Array Redistribution
abstract
In many scientific applications, dynamic array redistribution is usually required to enhance the performance of an algorithm. We present a generalized basic cycle calculation (GBCC) method to efficiently perform a BLOCK-CYCLIC(s) over P processors to BLOCK-CYCLIC(t) over Q processors array redistribution. In the GBCC method, a processor first computes the source/destination processor/data sets of array elements in the first generalized basic cycle of the local array it owns. A generalized basic cycle is defined as lcm(sP,tQ)/(gcd(s,t)/spl times/P) in the source distribution and lcm(sP,tQ)/(gcd(s,t)/spl times/Q) in the destination distribution. From the source/destination processor/data sets of array elements in the first generalized basic cycle, we can construct packing/unpacking pattern tables. Based on the packing/unpacking pattern tables, a processor can pack/unpack array elements efficiently. To evaluate the performance of the GBCC method, we have implemented this method on an IBM SP2 parallel machine, along with the PITFALLS method and the ScaLAPACK method. The cost models for these three methods are also presented. The experimental results show that the GBCC method outperforms the PITFALLS method and the ScaLAPACK method for all test samples. A brief description of the extension of the GBCC method to multi dimensional array redistributions is also presented.
Yeh-Ching Chung, Sheng-Wen Bai, Ching-Hsien Hsu, Chu-Sing Yang
ICPADS2
1998 A Basic-Cycle Calculation Technique for Efficient Dynamic Data Redistribution
abstract
Array redistribution is usually required to enhance algorithm performance in many parallel programs on distributed memory multicomputers. Since it is performed at run-time, there is a performance trade-off between the efficiency of the new data decomposition for a subsequent phase of an algorithm and the cost of redistributing data among processors. In this paper, we present a basic-cycle calculation technique to efficiently perform BLOCK-CYCLIC(S) to BLOCK-CYCLIC(t) redistribution. The main idea of the basic-cycle calculation technique is, first, to develop closed forms for computing source/destination processors of some specific array elements in a basic-cycle, which is defined as icm(s,t)/gcd(s,t). These closed forms are then used to efficiently determine the communication sets of a basic-cycle. From the source/destination processor/data sets of a basic-cycle, we can efficiently perform a BLOCK-CYCLIC(s) to BLOCK-CYCLIC(t) redistribution. To evaluate the performance of the basic-cycle calculation technique, we have implemented this technique on an IBM SP2 parallel machine, along with the PITFALLS method and the multiphase method. The cost models for these three methods are also presented. The experimental results show that the basic-cycle calculation technique outperforms the PITFALLS method and the multiphase method for most test samples.
Yeh-Ching Chung, Ching-Hsien Hsu, Sheng-Wen Bai
IEEE Trans. Parallel Distributed Syst.3