Chenggong Charles Fan

dblp:37/797 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
0since 2021 · last 2001
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Interconnection networks and networks-on-chip · 32% Distributed systems · 32% Hardware reliability and fault tolerance · 23%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware reliability and fault tolerance
error control coding
0.012001
Computing in the RAIN: A Reliable Array of Independent Nodes · IEEE Trans. Parallel Distributed Syst. 2001
Distributed systems
fault tolerance
0.012001
Computing in the RAIN: A Reliable Array of Independent Nodes · IEEE Trans. Parallel Distributed Syst. 2001
Distributed systems › distributed system dependability
reliable distributed systems
0.012001
Computing in the RAIN: A Reliable Array of Independent Nodes · IEEE Trans. Parallel Distributed Syst. 2001
Storage systems
storage reliability
0.012001
Computing in the RAIN: A Reliable Array of Independent Nodes · IEEE Trans. Parallel Distributed Syst. 2001
Interconnection networks and networks-on-chip › nonblocking networks
extra-stage networks
0.012000
Tolerating Multiple Faults in Multistage Interconnection Networks with Minimal Extra Stages · IEEE Trans. Computers 2000
Interconnection networks and networks-on-chip › switching network › multistage interconnection network
fault-tolerant MIN
0.012000
Tolerating Multiple Faults in Multistage Interconnection Networks with Minimal Extra Stages · IEEE Trans. Computers 2000
Hardware reliability and fault tolerance › network fault tolerance
fault-tolerant network design
0.012000
Tolerating Multiple Faults in Multistage Interconnection Networks with Minimal Extra Stages · IEEE Trans. Computers 2000
Interconnection networks and networks-on-chip › switching network
multistage interconnection network
0.012000
Tolerating Multiple Faults in Multistage Interconnection Networks with Minimal Extra Stages · IEEE Trans. Computers 2000
Distributed systems
fault management
0.012001
Computing in the RAIN: A Reliable Array of Independent Nodes · IEEE Trans. Parallel Distributed Syst. 2001
Distributed systems › group communication
group membership
0.012001
Computing in the RAIN: A Reliable Array of Independent Nodes · IEEE Trans. Parallel Distributed Syst. 2001

Methods — techniques the papers use, named apart from their topics

software-implemented fault tolerance · 0.0error-control codes · 0.0
YearPublicationVenuePosition
2001 The Raincore Distributed Session Service for Networking Elements
abstract
Motivated by the explosive growth of the Internet, we study efficient and fault-tolerant distributed session layer \nprotocols for networking elements. These protocols are \ndesigned to enable a network cluster to share the state \ninformation necessary for balancing network traffic and \ncomputation load among a group of networking elements. \nIn addition, in the presence of failures, they allow \nnetwork traffic to fail-over from failed networking \nelements to healthy ones. To maximize the overall \nnetwork throughput of the networking cluster, we assume a unicast communication medium for these protocols. The Raincore Distributed Session Service is based on a fault-tolerant token protocol, and provides group membership, reliable multicast and mutual exclusion services in a networking environment. We show that this service provides atomic reliable multicast with consistent ordering. We also show that Raincore token protocol consumes less overhead than a broadcast-based protocol in this environment in terms of CPU task-switching. The Raincore technology was transferred to Rainfinity, a startup company that is focusing on software for Internet reliability and performance. Rainwall, Rainfinity’s first product, was developed using the Raincore Distributed Session Service. We present initial performance results of the Rainwall product that validates our design assumptions and goals.
Chenggong Charles Fan, Jehoshua Bruck
IPDPS1
2001 Computing in the RAIN: A Reliable Array of Independent Nodes
abstract
The RAIN project is a research collaboration between Caltech and NASA-JPL on distributed computing and data-storage systems for future spaceborne missions. The goal of the project is to identify and develop key building blocks for reliable distributed systems built with inexpensive off-the-shelf components. The RAIN platform consists of a heterogeneous cluster of computing and/or storage nodes connected via multiple interfaces to networks configured in fault-tolerant topologies. The RAIN software components run in conjunction with operating system services and standard network protocols. Through software-implemented fault tolerance, the system tolerates multiple node, link, and switch failures, with no single point of failure. The RAIN-technology has been transferred to Rainfinity, a start-up company focusing on creating clustered solutions for improving the performance and availability of Internet data centers. In this paper, we describe the following contributions: 1) fault-tolerant interconnect topologies and communication protocols providing consistent error reporting of link failures, 2) fault management techniques based on group membership, and 3) data storage schemes based on computationally efficient error-control codes. We present several proof-of-concept applications: a highly-available video server, a highly-available Web server, and a distributed checkpointing system. Also, we describe a commercial product, Rainwall, built with the RAIN technology.
Vasken Bohossian, Chenggong Charles Fan, Paul S. LeMahieu, Marc D. Riedel, Lihao Xu, Jehoshua Bruck
IEEE Trans. Parallel Distributed Syst.2
2000 Tolerating Multiple Faults in Multistage Interconnection Networks with Minimal Extra Stages
abstract
Adams and Siegel (1982) proposed an extra stage cube interconnection network that tolerates one switch failure with one extra stage. We extend their results and discover a class of extra stage interconnection networks that tolerate multiple switch failures with a minimal number of extra stages. Adopting the same fault model as Adams and Siegel, the faulty switches can be bypassed by a pair of demultiplexer/multiplexer combinations. It is easy to show that, to maintain point to point and broadcast connectivities, there must be at least S extra stages to tolerate I switch failures. We present the first known construction of an extra stage interconnection network that meets this lower-bound. This 12-dimensional multistage interconnection network has n+f stages and tolerates I switch failures. An n-bit label called mask is used for each stage that indicates the bit differences between the two inputs coming into a common switch. We designed the fault-tolerant construction such that it repeatedly uses the singleton basis of the n-dimensional vector space as the stage mask vectors. This construction is further generalized and we prove that an n-dimensional multistage interconnection network is optimally fault-tolerant if and only if the mask vectors of every n consecutive stages span the n-dimensional vector space.
Chenggong Charles Fan, Jehoshua Bruck
IEEE Trans. Computers1