Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jane-Ferng Chiu

dblp:95/3043 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
0since 2021 · last 2011
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 first-authorSecurity and privacy · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Distributed systems · 99% High-performance computing · 1%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems
fault tolerance
0.122011
A New Diskless Checkpointing Approach for Multiple Processor Failures · IEEE Trans. Dependable Secur. Comput. 2011
Process-Replication Technique for Fault Tolerance and Performance Improvement in Distributed Computing Systems · HPDC 1994
Distributed systems › fault tolerance › checkpointing
diskless checkpointing
0.112011
A New Diskless Checkpointing Approach for Multiple Processor Failures · IEEE Trans. Dependable Secur. Comput. 2011
Distributed systems
distributed coordination
0.011994
Process-Replication Technique for Fault Tolerance and Performance Improvement in Distributed Computing Systems · HPDC 1994
Distributed systems › distributed coordination
message ordering
0.011994
Process-Replication Technique for Fault Tolerance and Performance Improvement in Distributed Computing Systems · HPDC 1994
Distributed systems › replication
process replication
0.011994
Process-Replication Technique for Fault Tolerance and Performance Improvement in Distributed Computing Systems · HPDC 1994
High-performance computing
performance optimization
0.011994
Process-Replication Technique for Fault Tolerance and Performance Improvement in Distributed Computing Systems · HPDC 1994

Methods — techniques the papers use, named apart from their topics

XOR-based checkpoint encoding · 0.1simulation · 0.0
YearPublicationVenuePosition
2011 A New Diskless Checkpointing Approach for Multiple Processor Failures
abstract
Diskless checkpointing is an important technique for performing fault tolerance in distributed or parallel computing systems. This study proposes a new approach to enhance neighbor-based diskless checkpointing to tolerate multiple failures using simple checkpointing and failure recovery operations, without relying on dedicated checkpoint processors. In this scheme, each processor saves its checkpoints in a set of peer processors, called checkpoint storage nodes. In return, each processor uses simple XOR operations to store a collection of checkpoints for the processors for which it is a checkpoint storage node. This study defines the concept of safe recovery criterion, which specifies the requirement for ensuring that any failed processor can be recovered in a single step using the checkpoint data stored at one of the surviving processors, as long as no more than a given number of failures occur. This study further identifies the necessary and sufficient conditions for satisfying the safe recovery criterion and presents a method for designing checkpoint storage node sets that meet these requirements. The proposed scheme allows failure recovery to be performed in a distributed manner using XOR operations.
Ge-Ming Chiu, Jane-Ferng Chiu
IEEE Trans. Dependable Secur. Comput.2
2008 Mutual-Aid: Diskless Checkpointing Scheme for Tolerating Double Faults
abstract
Tolerating double faults is an important issue for diskless checkpointing due to the size and increase of the executing time. This is why Mutual-Aid checkpointing has become the first scheme to achieve the goal. Mutual-aid checkpointing combines the advantages of neighbor-based, with parity-based diskless approaches. This also tolerates all double processor faults by bitwising exclusive-or snapshots from its neighbor processors in its virtual assistant ring. In view of the fact that checkpointing and recovery of mutual-aid are so simple and efficient, this increases the performance, reduces application running time, and allows more frequent checkpoints. Moreover, it could be employed towards a very largescale and high performance computing field because of its distributed methods as well as localized operations. The degree of fault tolerance has achieved higher success than other schemes.
Jane-Ferng Chiu, Wei-Hua Hao
HPCC1
1994 Process-Replication Technique for Fault Tolerance and Performance Improvement in Distributed Computing Systems
abstract
The paper presents a process-replication protocol which aims at providing fault-tolerance as well as performance improvement to applications such as long-running and real-time tasks. Identical delivering order of messages are enforced on all replicas of a troupe using multicasts for inter- and intra-troupe communication. The detailed design of the protocol is given in the paper. The protocol is self-contained in the sense that crashes in a troupe are handled internally without affecting the operation of other troupes. The crash-handling procedure is simple and associated overhead during fail-free operation is small. The protocol takes advantages of the redundancy of processes to expedite the completion of a distributed task by speeding up the determination of message sequences and transmission of outgoing data messages at the expense of small control messages. Simulation is carried out to show the performance improvement.>
Jane-Ferng Chiu, Ge-Ming Chiu
HPDC1