EDBT 2026 Demo / reviewers in the wild / expert
Santosh K. Shrivastava
dblp:37/5945 · also Santosh Kumar Shrivastava
· DBLP profile ↗
66ranked-venue papers
16as first author
1since 2021 · last 2022
0000-0001-6105-2785ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 20 · 7 first-author · 1 since 2021Systems, architecture and hardware · 15 · 4 first-authorSecurity and privacy · 12 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 6 · 2 first-authorComputer networks · 1Databases, data management, data science and information retrieval · 1Theory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
13 papers |
Distributed systems · 76% Hardware reliability and fault tolerance · 11% Cloud and datacenter computing · 6% | |
| Software engineering, system software, and programming languages
6 papers |
Requirements engineering and software design · 76% Services computing and microservices · 22% Operating systems · 2% | |
| Databases, data mining, and information retrieval
1 paper |
Transaction processing and concurrency control · 100% | |
| Network and information security
1 paper |
Cryptographic protocols and secure computation · 87% Authentication and access control · 13% |
Topics — the 30 heaviest of 42, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Requirements engineering and software design
business process modeling |
0.1 | 1 | 2012 | A Model for Checking Contractual Compliance of Business Interactions · IEEE Trans. Serv. Comput. 2012 |
Distributed systems
fault tolerance |
0.1 | 10 | 2008 | Enhancing an Application Server to Support Available Components · IEEE Trans. Software Eng. 2008 Implementing Fail-Silent Nodes for Distributed Systems · IEEE Trans. Computers 1996 Structuring Fault-Tolerant Object Systems for Modularity in a Distributed Environment · IEEE Trans. Parallel Distributed Syst. 1994 |
Transaction processing and concurrency control
distributed transaction processing |
0.1 | 1 | 2008 | Enhancing an Application Server to Support Available Components · IEEE Trans. Software Eng. 2008 |
Transaction processing and concurrency control › distributed transaction management
multidatabase transaction management |
0.1 | 1 | 2008 | Enhancing an Application Server to Support Available Components · IEEE Trans. Software Eng. 2008 |
Cryptographic protocols and secure computation
fair exchange |
0.1 | 1 | 2005 | A Family of Trusted Third Party Based Fair-Exchange Protocols · IEEE Trans. Dependable Secur. Comput. 2005 |
Cryptographic protocols and secure computation › secure computation protocols › setup assumptions
trusted third party |
0.1 | 1 | 2005 | A Family of Trusted Third Party Based Fair-Exchange Protocols · IEEE Trans. Dependable Secur. Comput. 2005 |
Distributed systems
replication |
0.0 | 3 | 1996 | Implementing Fail-Silent Nodes for Distributed Systems · IEEE Trans. Computers 1996 Principal Features of the VOLTAN Family of Reliable Node Architectures for Distributed Systems · IEEE Trans. Computers 1992 Using Objects and Actions to Provide Fault Tolerance in Distributed, Real-Time Applications · RTSS 1991 |
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management |
0.0 | 1 | 2008 | Enhancing an Application Server to Support Available Components · IEEE Trans. Software Eng. 2008 |
Authentication and access control
smart card |
0.0 | 1 | 2005 | A Family of Trusted Third Party Based Fair-Exchange Protocols · IEEE Trans. Dependable Secur. Comput. 2005 |
Distributed systems › group communication
message ordering protocols |
0.0 | 1 | 1996 | Implementing Fail-Silent Nodes for Distributed Systems · IEEE Trans. Computers 1996 |
Hardware reliability and fault tolerance
self-checking systems |
0.0 | 1 | 1996 | Implementing Fail-Silent Nodes for Distributed Systems · IEEE Trans. Computers 1996 |
Distributed systems › fault tolerance
recoverability |
0.0 | 2 | 1994 | Structuring Fault-Tolerant Object Systems for Modularity in a Distributed Environment · IEEE Trans. Parallel Distributed Syst. 1994 Structuring Distributed Systems for Recoverability and Crash Resistance · IEEE Trans. Software Eng. 1981 |
Distributed systems › middleware
distributed programming systems |
0.0 | 1 | 1994 | Structuring Fault-Tolerant Object Systems for Modularity in a Distributed Environment · IEEE Trans. Parallel Distributed Syst. 1994 |
Distributed systems
persistent object systems |
0.0 | 1 | 1994 | Structuring Fault-Tolerant Object Systems for Modularity in a Distributed Environment · IEEE Trans. Parallel Distributed Syst. 1994 |
Hardware reliability and fault tolerance
fault masking |
0.0 | 1 | 1992 | Principal Features of the VOLTAN Family of Reliable Node Architectures for Distributed Systems · IEEE Trans. Computers 1992 |
Distributed systems › fault tolerance
fault-tolerant real-time systems |
0.0 | 1 | 1991 | Using Objects and Actions to Provide Fault Tolerance in Distributed, Real-Time Applications · RTSS 1991 |
Distributed systems
remote procedure call |
0.0 | 2 | 1988 | Rajdoot: A Remote Procedure Call Mechanism Supporting Orphan Detection and Killing · IEEE Trans. Software Eng. 1988 The Design of a Reliable Remote Procedure Call Mechanism · IEEE Trans. Computers 1982 |
Performance modeling and evaluation
analytical modeling |
0.0 | 1 | 1990 | A Performance Evaluation Study of Pipeline TMR Systems · IEEE Trans. Parallel Distributed Syst. 1990 |
Processor architecture and microarchitecture › pipelining
pipeline performance |
0.0 | 1 | 1990 | A Performance Evaluation Study of Pipeline TMR Systems · IEEE Trans. Parallel Distributed Syst. 1990 |
Hardware reliability and fault tolerance › redundancy › modular redundancy
triple modular redundancy |
0.0 | 1 | 1990 | A Performance Evaluation Study of Pipeline TMR Systems · IEEE Trans. Parallel Distributed Syst. 1990 |
Hardware reliability and fault tolerance › redundancy
replicated execution |
0.0 | 1 | 1989 | Constructing Replicated Systems Using Processors with Point-to-Point Communication Links · ISCA 1989 |
Processor architecture and microarchitecture
multicore design |
0.0 | 1 | 1992 | Principal Features of the VOLTAN Family of Reliable Node Architectures for Distributed Systems · IEEE Trans. Computers 1992 |
Embedded and real-time systems
distributed real-time systems |
0.0 | 1 | 1991 | Using Objects and Actions to Provide Fault Tolerance in Distributed, Real-Time Applications · RTSS 1991 |
Storage systems
crash resistance |
0.0 | 1 | 1981 | Structuring Distributed Systems for Recoverability and Crash Resistance · IEEE Trans. Software Eng. 1981 |
Performance modeling and evaluation
simulation |
0.0 | 1 | 1990 | A Performance Evaluation Study of Pipeline TMR Systems · IEEE Trans. Parallel Distributed Syst. 1990 |
Distributed systems
distributed coordination |
0.0 | 1 | 1989 | Constructing Replicated Systems Using Processors with Point-to-Point Communication Links · ISCA 1989 |
Hardware reliability and fault tolerance
majority voting |
0.0 | 1 | 1989 | Constructing Replicated Systems Using Processors with Point-to-Point Communication Links · ISCA 1989 |
Parallel and multicore computing › parallel programming models › message passing
point-to-point communication |
0.0 | 1 | 1989 | Constructing Replicated Systems Using Processors with Point-to-Point Communication Links · ISCA 1989 |
Distributed systems › fault tolerance
exactly-once processing |
0.0 | 1 | 1988 | Rajdoot: A Remote Procedure Call Mechanism Supporting Orphan Detection and Killing · IEEE Trans. Software Eng. 1988 |
Operating systems › resource management
resource allocation |
0.0 | 1 | 1978 | Reliable Resource Allocation Between Unreliable Processes · IEEE Trans. Software Eng. 1978 |
Methods — techniques the papers use, named apart from their topics
messaging middleware event processing · 0.3business rules · 0.3replication · 0.2nonblocking transaction processing · 0.2failover · 0.2smartcard-based protocols · 0.1protocol family · 0.1object-oriented design · 0.0message ordering · 0.0comparison protocol · 0.0fail-signal · 0.0application-level process replication · 0.0protocol design · 0.0object-oriented multilevel model · 0.0monitors · 0.0backward error recovery · 0.0SIMULA class and inner features · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | The evolution of the Arjuna transaction processing systemabstractAbstract TheArjunatransaction processing system began life in the mid‐1980s as an academic research project to examine the use of object‐oriented techniques in the development of fault‐tolerant distributed systems. Thirty‐five years later, and available in open source form, it forms an integral part of various middleware products from Red Hat where it is used to provide ACID as well as non‐ACID transaction services. This journey from an academic to a commercial environment has been neither easy nor smooth but it has been interesting from many different perspectives. This article gives an overview of this journey, discusses the features ofArjunathat enables it to support a wide variety of transaction services and concludes by presenting lessons learned. Mark C. Little, Santosh K. Shrivastava |
Softw. Pract. Exp. | 2 |
| 2012 | A Model for Checking Contractual Compliance of Business InteractionsabstractThe electronic representation of a contract for a business-to-business (B2B) partnership should be such that it can be used by a monitoring service for compliance checking of B2B interactions at runtime, ensuring that the interactions match the rights and obligations that each partner has promised to honor. With this view in mind, the paper develops a model for checking contractual compliance of business interactions. Specifically, the paper develops a novel way of representing contract clauses using business rules, that is specially suited to compliance checking and describes what events need to be captured from the underlying messaging middleware and how they can be processed in a careful manner to evaluate contractual compliance. Carlos Molina-Jiménez, Santosh K. Shrivastava, Massimo Strano |
IEEE Trans. Serv. Comput. | 2 |
| 2010 | A Case for Consumer - centric Resource Accounting ModelsabstractA pay-per-use cloud service should be made available to consumers with an unambiguous resource accounting model that precisely describes all the factors that are taken into account in calculating resource consumption charges. The paper proposes the notion of consumer-centric resource accounting model such that consumers can programmatically compute their consumption charges of a remotely used service. In particular, the notion of strongly consumer-centric accounting model is proposed that requires that all the data needed for calculating billing charges can be collected independently by the consumer (or a trusted third party, TTP); in effect, this means that a consumer (or a TTP) should be in a position to run their own measurement service. Strongly consumer-centric accounting models have the desirable property of openness and transparency, since service users are in a position to verify the charges billed to them. To illustrate the ideas, the accounting model of a given cloud infrastructure service (simple storage service, S3 from Amazon) is evaluated. The exercise reveals some shortcomings which can be fixed as indicated in this paper to make Amazon's model strongly consumer-centric. Service providers can learn from this evaluation study to re-examine their accounting models and perform any amendments. Ahmed Mihoob, Carlos Molina-Jiménez, Santosh K. Shrivastava |
IEEE CLOUD | 3 |
| 2009 | Proactive Fortification of Fault-Tolerant Services
Paul D. Ezhilchelvan, Dylan Clarke, Isi Mitrani, Santosh K. Shrivastava |
OPODIS | 4 |
| 2008 | Enhancing an Application Server to Support Available ComponentsabstractThree-tier middleware architecture is commonly used for hosting enterprise-distributed applications. Typically, the application is decomposed into three layers: front end, middle tier, and back end. Front end ("Web server") is responsible for handling user interactions and acts as a client of the middle tier, while back end provides storage facilities for applications. Middle tier ("application server") is usually the place where all computations are performed. One of the benefits of this architecture is that it allows flexible management of a cluster of computers for performance and scalability; further, availability measures, such as replication, can be introduced in each tier in an application-specific manner. However, incorporation of availability measures in a multitier system poses challenging system design problems of integrating open, nonproprietary solutions to transparent failover, exactly once execution of client requests, nonblocking transaction processing, and an ability to work with clusters. This paper describes how replication for availability can be incorporated within the middle and back-end tiers, meeting all these challenges. This paper develops an approach that requires enhancements to the middle tier only for supporting replication of both the middleware back-end tiers. The design, implementation, and performance evaluation of such a middle-tier-based replication scheme for multidatabase transactions on a widely deployed open source application server (JBoss) are presented. Achmad I. Kistijantoro, Graham Morgan, Santosh K. Shrivastava, Mark C. Little |
IEEE Trans. Software Eng. | 3 |
| 2007 | Implementing Business Conversations with Consistency Guarantees Using Message-Oriented MiddlewareabstractThe paper considers distributed applications where interactions between constituent services take place via messages in an asynchronous environment with unpredictable communication and processing delays; further, interacting parties are not required to be online at the same time. Message-oriented middleware (MoM) is commonly used for connecting such loosely coupled distributed applications. Despite loose coupling, many service interactions have temporal and message validation constraints. A failure to deliver a valid message within its time constraint could cause mutually conflicting views of an interaction (one party regarding it as timely whilst the other party regarding it as untimely) leading to application level inconsistencies. In a loosely coupled system, such inconsistencies could remain undetected for a long time, requiring costly application level recovery procedures. This paper describes how synchronisation support providing multilateral consistency guarantees can be provided using the underlying MoM to prevent inconsistencies from reaching application level. Carlos Molina-Jiménez, Santosh K. Shrivastava, Nick Cook |
EDOC | 2 |
| 2006 | Satem: Trusted Service Code Execution across TransactionsabstractInternet-scale distributed applications are frequently built as loosely coupled compositions of services. We would like that despite loose coupling, each constituent service has a mutually consistent view of the state of the application not withstanding software, hardware and network related problems (e.g., clock skews, unpredictable transmission delays, message loss, node crashes etc.). Some observations on how messaging middleware should be structured to accomplish this are presented. Santosh K. Shrivastava |
SRDS | 1 |
| 2006 | Design and Implementation of Web Services Middleware to Support Fair Non-Repudiable InteractionsabstractThe use of open, Internet-based communications for business-to-business (B2B) interactions requires accountability for and acknowledgment of the actions of participants. Accountability and acknowledgment can be achieved by the systematic maintenance of an irrefutable audit trail to render the interaction non-repudiable. To safeguard the interests of each party, the mechanisms used to meet this requirement should ensure fairness. That is, misbehavior should not disadvantage well-behaved parties. Despite the fact that Web services are increasingly used to enable B2B interactions, there is currently no systematic support to deliver such guarantees. This paper introduces a flexible framework to support fair non-repudiable B2B interactions based on a trusted delivery agent. A Web services implementation is presented. The role of the delivery agent can be adapted to different end user capabilities and to meet different application requirements. Nick Cook, Paul Robinson, Santosh K. Shrivastava |
Int. J. Cooperative Inf. Syst. | 3 |
| 2005 | A Method for Specifying Contract Mediated InteractionsabstractTo form and automatically manage partnerships within a virtual organisation, it is necessary to have an electronic representation of the contract governing business relationships that can be used to mediate the rights and obligations that each interacting entity promises to honour. The paper describes a general method of representing business interactions using a widely used modelling language Promela and discusses how to represent permissions, obligations, prohibitions, actors (agents), time constraints, and message type checking; that is, all the basic parameters that compose most typical business contracts. Two levels of contract representations are described: implementation neutral, and implementation specific, that is a refinement of the former to include technical details such as acknowledgements and synchronization messages that form an important part of any implementation. Carlos Molina-Jiménez, Santosh K. Shrivastava, John P. Warne |
EDOC | 2 |
| 2005 | Implementing Fair Non-repudiable Interactions with Web ServicesabstractThe use of open, Internet-based communications for business-to-business (B2B) interactions requires accountability for and acknowledgment of the actions of participants. Accountability and acknowledgment can be achieved by the systematic maintenance of an irrefutable audit trail to render the interaction non-repudiable. To safeguard the interests of each party, the mechanisms used to meet this requirement should ensure fairness. That is, misbehaviour should not disadvantage well-behaved parties. Despite the fact that Web services are increasingly used to enable B2B interactions, there is currently no systematic support to deliver such guarantees. This paper introduces a flexible framework to support fair non-repudiable B2B interactions based on a trusted delivery agent. A Web services implementation is presented. The role of the delivery agent can be adapted to different end user capabilities and to meet different application requirements. Paul Robinson, Nick Cook, Santosh K. Shrivastava |
EDOC | 3 |
| 2005 | A Family of Trusted Third Party Based Fair-Exchange ProtocolsabstractFair exchange protocols play an important role in application areas such as e-commerce where protocol participants require mutual guarantees that a transaction involving exchange of items has taken place in a specific manner. A protocol is fair if no protocol participant can gain any advantage over an honest participant by misbehaving. In addition, such a protocol is fault-tolerant if the protocol can ensure that an honest participant does not suffer any loss of fairness despite any failures of the participant's node. This paper presents a family of fair exchange protocols for two participants which make use of the presence of a trusted third party, under a variety of assumptions concerning participant misbehavior, message, delays, and node reliability. The development is systematic, beginning with the strongest set of the assumptions and gradually weakening the assumptions to the weakest set. The resulting protocol family exposes the impact of a given set of assumptions on solving the problem of fair exchange. Specifically, it highlights the relationships that exist between fairness and assumptions on the nature of participant misbehavior, communication delays, and node crashes. The paper also shows that the restrictions assumed on a dishonest participant's misbehavior can be realized through the use of smartcards and smartcard-based protocols. Paul D. Ezhilchelvan, Santosh K. Shrivastava |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2004 | Component Middleware to Support Non-repudiable Service InteractionsabstractThe wide variety of services and resources available over the Internet presents new opportunities to create value added, inter-organisational composite services (CSs)from multiple existing services. To preserve their autonomy and privacy, each organisation needs to regulate access both to their services and to shared information within the CS. Key mechanisms to facilitate such regulated interactions are the collection and verification of non-repudiable evidence of the actions of the parties to the CS. The paper describes how component-based middleware can be enhanced to support non-repudiable service invocation and information sharing. A generic implementation, based on a J2EE application server, is presented. Nick Cook, Paul Robinson, Santosh K. Shrivastava |
DSN | 3 |
| 2004 | Notations for the Specification and Verification of Composite Web Services
Simon J. Woodman, Doug J. Palmer, Santosh K. Shrivastava, Stuart M. Wheater |
EDOC | 3 |
| 2003 | Middleware Support for Non-repudiable Transactional Information Sharing between Enterprises
Nick Cook, Santosh K. Shrivastava, Stuart M. Wheater |
DAIS | 2 |
| 2003 | Systematic Development of a Family of Fair Exchange Protocols
Paul D. Ezhilchelvan, Santosh K. Shrivastava |
DBSec | 2 |
| 2003 | Model Checking Correctness Properties of Electronic Contracts
Ellis Solaiman, Carlos Molina-Jiménez, Santosh K. Shrivastava |
ICSOC | 3 |
| 2003 | Component Replication in Distributed Systems: A Case Study Using Enterprise Java BeansabstractA recent trend has seen the extension of object-oriented middleware. A major advantage components offer over objects is that only the business logic of an application needs to be addressed by a programmer with support services required incorporated into the application at deployment time. This is achieved via components (business logic of an application), containers that host components and are responsible for providing the underlying middleware services required by components and application servers that host containers. Well-known examples of component middleware architectures are Enterprise Java Beans (EJBs) and the CORBA Component model (CCM). Two of the many services available at deployment time in most component architectures are component persistence and atomic transactions. This paper examines, using EJBs, how replication for availability can be supported by containers so that components that are transparently using persistence and transactions can also be made highly available. Achmad I. Kistijantoro, Graham Morgan, Santosh K. Shrivastava, Mark C. Little |
SRDS | 3 |
| 2003 | The CORBA Activity Service Framework for supporting extended transactionsabstractAbstract Although it has long been realized that ACID (atomicity, consistency, isolation, durability) transactions by themselves are not adequate for structuring long‐lived applications and much research work has been done on developing specific extended transaction models, no middleware support for building extended transactions is currently available and the situation remains that a programmer often has to develop application specific mechanisms. The CORBA Activity Service Framework described in this paper is a way out of this situation. The design of the service is based on the insight that the various extended transaction models can be supported by providing a general purpose event signalling mechanism that can be programmed to enable activities—application specific units of computations—to coordinate each other in a manner prescribed by the model under consideration. The different extended transaction models can be mapped onto specific implementations of this framework, permitting such transactions to span a network of systems connected indirectly by some distribution infrastructure. The framework described in this paper is an overview of the OMG's (Object Management Group) Additional Structuring Mechanisms for the OTS standard. Through a number of examples the paper shows that the framework has the flexibility to support a wide variety of extended transaction models. Although the framework is presented here in CORBA specific terms, the main ideas are sufficiently general, so that it should be possible to use them in conjunction with other middleware. Copyright © 2003 John Wiley & Sons, Ltd. Iain Houston, Mark C. Little, Ian Robinson, Santosh K. Shrivastava, Stuart M. Wheater |
Softw. Pract. Exp. | 4 |
| 2002 | Distributed Object Middleware to Support Dependable Information Sharing between OrganisationsabstractOrganisations increasingly use the Internet to offer their own services and to utilise the services of others. This naturally leads to information sharing across organisational boundaries. However, despite the requirement to share information, the autonomy and privacy requirements of organisations must not be compromised. This demands the strict policing of inter-organisational interactions. Thus there is a requirement for dependable mechanisms for information sharing between organisations that do not necessarily trust each other. The paper describes the design of a novel distributed object middleware that guarantees both safety and liveness in this context. The safety property ensures that local policies are not compromised despite failures and/or misbehaviour by other parties. The liveness property ensures that, if no party misbehaves, agreed interactions will take place despite a bounded number of temporary network and computer related failures. The paper describes a prototype implementation with example applications. Nick Cook, Santosh K. Shrivastava, Stuart M. Wheater |
DSN | 2 |
| 2002 | Dependability and the Grid: Issues and Challenges
Richard D. Schlichting, Andrew A. Chien, Carl Kesselman, Keith Marzullo, James S. Plank, Santosh K. Shrivastava |
DSN | 6 |
| 2002 | Using Bloom Filters to Speed-up Name Lookup in Distributed SystemsabstractBloom filters make use of a ‘probabilistic’ hash-coding method to reduce the amount of space required to store a hash set. A Bloom filter offers a trade-off between its size and the probability that the filter returns the wrong result. It does this without storing the entire set, at the cost of occasionally incorrectly answering yes to the question ‘is $x$ a member of $s$?’. How Bloom filters can be used to speed up the name to location resolution process in large-scale distributed systems is discussed. The approach presented offers trade-offs between performance (the time taken to resolve an object's name to its location) and resource utilization (the amount of physical memory to store location information and the number of messages exchanged to obtain the object's address). Mark C. Little, Santosh K. Shrivastava, Neil A. Speirs |
Comput. J. | 2 |
| 2001 | Flexible Workflow Management in the OPENflow SystemabstractWorkflow management systems are required to provide flexible ways of managing workflows as the business processes they represent frequently require changes over time. Provision of flexibility features in workflow management systems is very much a research issue as workflow systems in use today have been found lacking in such features. Two approaches to achieving flexibility in workflows have been identified, namely flexibility by selection and flexibility by adaptation. Flexibility by selection is achieved by ensuring that there are a number of execution paths through the workflow process, such that key decision making points are well represented. This allows the appropriate path to be selected on a per-instance basis to take account of the prevailing circumstances. Flexibility by adaptation permits dynamic changes to workflows to include one or more new execution paths. The paper describes how flexibility is supported in the OPENflow distributed workflow system. In particular, it describes high level tool support for performing dynamic changes to a workflow. In OPENflow, dynamic reconfiguration mechanisms have been provided by making use of atomic transactions to add and remove one or more tasks and to allow the addition and removal of dependencies between tasks from a running workflow. Use of transactions ensures that changes are carried out atomically with respect to normal processing. An example application is described to illustrate flexible workflow management. J. J. Halliday, Santosh K. Shrivastava, Stuart M. Wheater |
EDOC | 2 |
| 2001 | A Model and Architecture for Conducting Hierarchically Structured AuctionsabstractThe paper develops a distributed systems architecture for dependable Internet based online auctions, meeting the requirements of data integrity, responsiveness, fairness and scalability. It is well known that a tree-based, recursive design approach caters well for scalability requirements. With this observation in mind, the paper develops an approach that permits an auction service to be mapped on to globally distributed auction servers. The paper selects a suitable auction model that treats sellers and buyers symmetrically. This symmetry enables a computational node to play at one level of the tree the role of a seller by dealing with a group of potential buyers as well as to play the role of a potential buyer at the next higher level. Such a symmetric auction (also known as a double auction) is used for supporting a standard auction to be carried out in a hierarchic manner. An architecture is developed and basic algorithms and protocols are presented. Paul D. Ezhilchelvan, Santosh K. Shrivastava, Mark C. Little |
ISORC | 2 |
| 2001 | The CORBA Activity Service Framework for Supporting Extended Transactions
Iain Houston, Mark C. Little, Ian Robinson, Santosh K. Shrivastava, Stuart M. Wheater |
Middleware | 4 |
| 2000 | Implementing Flexible Object Group Invocation in Networked SystemsabstractDistributed applications should be able to make use of an object group service in a number of application specific ways. Three main modes of interactions can be identified; (i) request-reply; a client issues a request to multiple servers and waits for their replies; this represents a commonly occurring scenario when a service is replicated; (ii) group-to-group request-reply; a generalisation of the previous case, where clients are themselves groups; and (iii) Peer Participation: here all the members are regularly multicasting messages (asynchronous invocation); this represents a commonly occurring scenario when the purpose of an application is to share information between members, (e.g., a teleconferencing application). Customisation within each class of interaction is frequently required for obtaining better performance. This paper describes the design and implementation of a flexible CORBA object group service that supports the three types of interactions and enables application specific customisation. Performance figures collected over low latency LAN and high latency WAN are presented to support the case for flexibility. Graham Morgan, Santosh K. Shrivastava |
DSN | 2 |
| 2000 | A Workflow and Agent Based Platform for Service ProvisioningabstractThe design and implementation of a dependable system that provides a composition and execution environment for distributed applications whose executions could span arbitrarily large durations is described. The objective is to create a framework for complex service provisioning. By complex service provisioning we primarily mean the ability to compose a given service out of existing ones as well as the ability to exercise dynamic control over the execution of the service. The approach taken is centred around building middleware services based on integration of workflow and agent technologies. The platform enables these two systems to interact via CORBA services. Service behaviour and service deployment are represented as workflow processes. Individual tasks that make up the workflow would be legacy applications, specially created tasks, and agent applications. Agents are able to create workflow instances, receive results from workflow and send inputs to workflows. This enables agents to act as user agents capable of managing workflows on behalf of users. Santosh K. Shrivastava, Luc Bellissard, David Féliot, Marc Herrmann, Noel De Palma, Stuart M. Wheater |
EDOC | 1 |
| 1999 | Design and implemantation of a CORBA fault-tolerant object group service
Graham Morgan, Santosh K. Shrivastava, Paul D. Ezhilchelvan, Mark C. Little |
DAIS | 2 |
| 1999 | Implementing support for work activity coordination within a distributed workflow systemabstractThere is growing interest in providing computer support for an organisation's business processes such as customer order processing, product support, stock taking and so forth. Workflow systems are normally used for this purpose to co-ordinate and monitor execution of multiple tasks arranged to form business processes. In this context, computer support for work activity co-ordination will be taken to mean tools and services made available to the employees of an organisation to enable them to carry out their tasks that form the part of these workflows. The paper illustrates how a transactional workflow system can be augmented with support for reliable management of worklists in an arbitrarily distributed environment. Ideally, the workflow enactment service should be neutral to the types of work activity coordination services to be made available to users. In this way, the needs of a wide variety of organisations can be supported by building organisation specific services. The approach taken here is to separate organisational aspects of workflow management from the workflow enactment service as much as possible. J. J. Halliday, Santosh K. Shrivastava, Stuart M. Wheater |
EDOC | 2 |
| 1999 | Enhancing Replica Management Services to Tolerate Group FailuresabstractIn a distributed system, replication of components, such as objects, is a well known way of achieving availability. For increased availability, crashed and disconnected components must be replaced by new components on available spare nodes. In this context, we address the problem of reconfiguring a group after the group as an entity has failed. Such a failure is termed a group failure which, for example, can be the crash of every component in the group or the group being partitioned into minority islands. The solution assumes crash-proof storage, and eventual recovery of crashed nodes and healing of partitions. It guarantees that: (i) the number of groups reconfigured after a group failure is never more than one, and (ii) the reconfigured group contains a majority of the components which were members just before the group failed, so that the loss of state information due to group failure is minimal. The protocol is efficient in terms of communication rounds and use of stable store, during both normal operations and reconfiguration after a group failure. Paul D. Ezhilchelvan, Santosh K. Shrivastava |
ISORC | 2 |
| 1999 | A Method for Combining Replication with CachingabstractObject replication and cacheing have been used individually in distributed systems for many years. There are benefits from being able to support both: replication for availability, and cacheing for performance. Although both involve handling multiple copies of objects, there are sufficient differences between the two to make the design of an integrated approach a challenging exercise. In this paper, we describe a design that provides the benefits of both replication and cacheing by allowing both types of protocols to co-exist within the same application. We do not propose a new replication or cacheing protocol, but rather a way in which existing implementations can be combined. Mark C. Little, Santosh K. Shrivastava |
SRDS | 2 |
| 1999 | On the Provision of Replicated Internet Auction ServicesabstractThe paper presents the design of a software infrastructure that can support negotiation and competition among buyers and sellers of goods (auctioning) over the Internet. The goals of data integrity, responsiveness, and scalability have been achieved by replicating the auction service across a number of auction servers. Fabio Panzieri, Santosh K. Shrivastava |
SRDS | 2 |
| 1998 | A Language for Specifying the Composition of Reliable Distributed ApplicationsabstractThis paper describes the design of a scripting language aimed at expressing task (unit of computation) composition and inter-task dependencies of distributed applications whose execution could span arbitrary large durations. This work is motivated by the observation that an increasingly large number of distributed applications are constructed by composing them out of existing applications and are executed in an heterogeneous environment. The resulting applications can be very complex in structure, containing many notification and dataflow dependencies between their constituent applications. The language enables applications to be structured with the properties of modularity, interoperability, dynamic reconfigurability and fault-tolerance. Frédéric Ranno, Santosh K. Shrivastava, Stuart M. Wheater |
ICDCS | 2 |
| 1998 | Checked Transactions in an Asynchronous Message Passing EnvironmentabstractTraditionally transactions have been single-threaded. In such an environment the thread terminating the transaction is, by definition, the thread which performed the work. Therefore, transaction termination is implicitly synchronised with the completion of the transactional work. With the increased availability of both software and hardware multi-threading, transaction services are being required to allow multiple threads to be active within a transaction. In these systems it is important to guarantee that all threads have completed when a transaction is terminated otherwise some work may not be performed transactionally. In this paper we present a protocol for the enforcement of checked transactional behaviour within an asynchronous environment. We illustrate the use of the protocol within a proposed implementation for a CORBA-compliant Object Transaction Service intended for a soft real-time application which makes extensive use of concurrency and asynchronous message passing. Steve J. Caughey, Mark C. Little, Santosh K. Shrivastava |
ISORC | 3 |
| 1998 | Inter-task Co-ordination in Long-Lived Distributed Applications
Santosh K. Shrivastava |
DISC | 1 |
| 1998 | Performance of Fault-Tolerant Data and Compute Intensive Programs over a Network of Workstations
J. A. Smith, Santosh K. Shrivastava |
Theor. Comput. Sci. | 2 |
| 1997 | A System for Specifying and Coordinating the Execution of Reliable Distributed Applications
Frédéric Ranno, Santosh K. Shrivastava, Stuart M. Wheater |
DAIS | 2 |
| 1997 | Fault-Tolerant Parallel Applications Using Queues and ActionsabstractThere are many techniques supporting execution of large computations over a network of workstations (NOW) but data intensive computations are usually run on high performance parallel machines. A NOW comprising individual user's machines typically has a low performance interconnect and suffers arbitrary changes of availability. Exploiting such resources to execute data intensive computations is difficult but even in a more constrained environment there is an unfulfilled need for fault-tolerance. The structuring approach presented fulfills this need. Performance exceeding 100 Mflop/s is demonstrated for large fault-tolerant out of core examples of matrix multiplication and Cholesky factorisation using five 133 MHz Pentium compute machines. J. A. Smith, Santosh K. Shrivastava |
ICPP | 2 |
| 1997 | Constructing Reliable Web Applications Using Atomic Actions
Mark C. Little, Santosh K. Shrivastava, Steve J. Caughey, David B. Ingham |
Comput. Networks | 2 |
| 1996 | Implementing Fail-Silent Nodes for Distributed SystemsabstractA fail-silent node is a self-checking node that either functions correctly or stops functioning after an internal failure is detected. Such a node can be constructed from a number of conventional processors. In a software-implemented fail-silent node, the nonfaulty processors of the node need to execute message order and comparison protocols to "keep in step" and check each other, respectively. In this paper, the design and implementation of efficient protocols for a two processor fail-silent node are described in detail. The performance figures obtained indicate that in a wide class of applications requiring a high degree of fault tolerance, software-implemented fail-silent nodes constructed simply by utilizing standard "off-the-shelf" components are an attractive alternative to their hardware-implemented counterparts that do require special-purpose hardware components, such as fault-tolerant clocks, comparator, and bus interface circuits. Francisco Vilar Brasileiro, Paul D. Ezhilchelvan, Santosh K. Shrivastava, Neil A. Speirs |
IEEE Trans. Computers | 3 |
| 1995 | Newtop: A Fault-Tolerant Group Communication ProtocolabstractA general purpose group communication protocol suite called Newtop is described. It is assumed that processes can simultaneously belong to many groups, group size could be large, and processes could be communicating over the Internet. Asynchronous communication environment is therefore assumed where message transmission times cannot be accurately estimated, and the underlying network may well get partitioned, preventing functioning processes from communicating with each other. Newtop can provide causality preserving total order delivery to members of a group, ensuring that total order delivery is preserved for multi-group processes. Both symmetric and asymmetric order protocols are supported, permitting a process to use say symmetric version in one group and asymmetric version in other. Paul D. Ezhilchelvan, Raimundo José de Araújo Macêdo, Santosh K. Shrivastava |
ICDCS | 3 |
| 1994 | Structuring Fault-Tolerant Object Systems for Modularity in a Distributed EnvironmentabstractThe object-oriented approach to system structuring has found widespread acceptance among designers and developers of robust computing systems. The authors propose a system structure for distributed programming systems that support persistent objects and describe how properties such as persistence and recoverability can be implemented. The proposed structure is modular, permitting easy exploitation of any distributed computing facilities provided by the underlying system. An existing system constructed according to the principles espoused here is examined to illustrate the practical utility of the proposed approach to system structuring.> Santosh K. Shrivastava, Daniel L. McCue |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 1993 | Maintaining Information about Persistent Replicated Objects in a Distributed SystemabstractPresents a general model for persistent replicated object management and identify what metainformation about objects needs to be maintained by a naming and binding service to ensure that objects named by application programs are bound to only those object replicas which are in a mutually consistent state. These ideas are developed within the framework of a distributed system in which application programs are composed of atomic actions (atomic transactions) manipulating persistent (long-lived) objects.> Mark C. Little, Daniel L. McCue, Santosh K. Shrivastava |
ICDCS | 3 |
| 1993 | The Duality of Fault-tolerant System StructuresabstractAbstract An examination of the structure of fault‐tolerant systems incorporating backward error recovery indicates a partitioning into two broad classes. Two canonical models, each representing a particular class of systems, have been constructed. The first model incorporates objects and actions as the entities for program construction whereas the second model employs communicating processes and conversations. Applications in areas such as office information and banking systems are typically described and built in terms of the first model whereas applications in the area of process control are usually described and built in terms of the second model. The paper claims that the two models are duals of each other and presents arguments and examples to substantiate this claim. It will be shown that the techniques that have been developed within the context of one model turn out to have interesting and hitherto unexplored duals in the other model. Santosh K. Shrivastava, Luigi V. Mancini, Brian Randell |
Softw. Pract. Exp. | 1 |
| 1992 | Principal Features of the VOLTAN Family of Reliable Node Architectures for Distributed SystemsabstractA VOLTAN node is composed of a number of conventional processors on which application-level processes are replicated to achieve fault-tolerance. The architecture of a family of such nodes with differing functionalities is presented. These include failure-masking, fail-signal, and fail-silent nodes. The software architectures of a three-processor failure-masking and a two-processor fail-silent node are discussed in detail. The suitability of VOLTAN nodes as building blocks of reliable distributed systems is also discussed.> Santosh K. Shrivastava, Paul D. Ezhilchelvan, Neil A. Speirs, Alan Tully |
IEEE Trans. Computers | 1 |
| 1991 | Using Objects and Actions to Provide Fault Tolerance in Distributed, Real-Time ApplicationsabstractAn object-oriented model is developed for structuring distributed real-time applications. Atomic atoms (atomic transactions) and exception handling techniques are used to introduce fault tolerance. Additional techniques are then developed to permit application and device specific commit and abort processing. Objects can be replicated to increase their availability. The authors examine the reasons why some of the previous real-time object models are not suitable for active replication and why the model proposed represents an improvement. Realistic examples are used to illustrate the practical utility of the approach.> Santosh K. Shrivastava, Adrian Waterworth |
RTSS | 1 |
| 1991 | Fault-Tolerant Reference Counting for Garbage Collection in Distributed SystemsabstractThe function of a garbage collector in a computer system is to reclaim storage that is no longer in use. Developing a garbage collector for a distributed system composed of autonomous computers (nodes) connected by a communication network poses a challenging problem: optimising performance whilst achieving fault-tolerance. The paper presents the design and implementation of a reference-count garbage collection scheme which is both efficient and fault-tolerant. A distributed object-based system is considered where operations on remote objects are invoked via remote procedure calls. The orphan treatment scheme associated with remote procedure calls has been enhanced to enable the collection of garbage arising from node crashes. Luigi V. Mancini, Santosh K. Shrivastava |
Comput. J. | 2 |
| 1990 | Implementing Fault-Tolerant Distributed ApplicationsabstractThe authors develop some control structures suitable for composing fault-tolerant distributed applications using atomic actions (atomic transactions) as building blocks. The authors describe how such structures may be implemented using the concept of multicolored actions. The reasons why other control structures, in addition to nested and concurrent atomic actions, are desirable and are identified, and three structures are proposed: serializing actions, glued actions, and top-level independent actions. A number of examples are used to illustrate their usefulness. A technique based on the concept of multicolored actions is presented as a uniform basis for implementing all of the three action structures presented.> Santosh K. Shrivastava, Stuart M. Wheater |
ICDCS | 1 |
| 1990 | Preventing State Divergence in Replicated Distributed ProgramsabstractReplicated execution of distributed programs, which provides a means of masking hardware (processor) failures in a distributed system, is discussed. Application-level entities (processes, objects) are replicated to execute on distinct processors. Such replica entities communicate by message passing. Nondeterminism within the replicas could cause messages to be processed in nonidentical order, producing a divergence of state. Possible sources of nondeterminism are identified, and a generic mechanism for ensuring that nonfaulty replicas process messages in identical order, thereby preventing state divergence among such replicate entities, is presented.> Alan Tully, Santosh K. Shrivastava |
SRDS | 2 |
| 1990 | A Performance Evaluation Study of Pipeline TMR SystemsabstractA distributed system in which a job can be broken into a number of subjobs which are processed sequentially at various processors is considered. The performance of such a system is then compared to the replicated (triple modular redundant, or TMR) version of the system in which each subjob will require concurrent replicated processing with majority voting. The effect of voting times and processor failure rates on the performance of the system is investigated with analytical approximations and computer simulations. The accuracy of the former is examined. The results indicate the possible existence of a threshold voting time, below which the TMR system performs better than the unreplicated one, and above which the situation is reversed. Such thresholds are observed, where possible, in systems with repairable servers, as well as in those with nonrepairable servers.> Paul D. Ezhilchelvan, Isi Mitrani, Santosh K. Shrivastava |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 1989 | The Treatment of Persistent Objects in Arjuna
Graeme N. Dixon, Graham D. Parrington, Santosh K. Shrivastava, Stuart M. Wheater |
ECOOP | 3 |
| 1989 | Constructing Replicated Systems Using Processors with Point-to-Point Communication LinksabstractReplicated processing with majority voting is a well known method of achieving fault tolerance. We consider the problem of constructing a distributed system composed of an arbitrarily large number of N-modular redundant (NMR) nodes, where each node itself is composed of N, N = 2m + 1 and m ≥ 1, processing and voting elements. Advanced microprocessors, such as Inmos Transputers, provide fast serial communication links for inter-processor communication, making it possible to construct large networks of processors. We describe how replicated processing with majority voting can be achieved for such processor networks. This paper will present the overall systems architecture, including voting and NMR synchronization algorithms specially developed to exploit fast point to point communication facilities. Paul D. Ezhilchelvan, Santosh K. Shrivastava, Alan Tully |
ISCA | 2 |
| 1989 | The Treatment of Persistent Objects in ArjunaabstractArjuna is a programming system which provides a set of tools for constructing fault-tolerant distributed applications. It supports an object-oriented model of computation in which atomic actions (atomic transactions) control sequences of operations invoked upon persistent objects. Persistent objects outlive the applications that create them and this paper concentrates on the mechanisms within Arjuna that are concerned with their management. The paper describes how these mechanisms are related to the other Arjuna mechanisms required by atomic actions for the distribution, concurrency control, recovery and commitment of persistent objects. Graeme N. Dixon, Graham D. Parrington, Santosh K. Shrivastava, Stuart M. Wheater |
Comput. J. | 3 |
| 1988 | Implementing Concurrency Control in Reliable Object-Oriented Systems
Graham D. Parrington, Santosh K. Shrivastava |
ECOOP | 2 |
| 1988 | Rajdoot: A Remote Procedure Call Mechanism Supporting Orphan Detection and KillingabstractRajdoot is a remote procedure call (RPC) mechanism with a number of fault tolerance capabilities. A discussion is presented of the reliability-related issues and how these issues have been dealt with in the RPC design. Rajdoot supports exactly-once semantics with call nesting capability, and incorporates effective measures for orphan detection and killing. Performance figures show that the reliability measures of Rajdoot impose little overhead.> Fabio Panzieri, Santosh K. Shrivastava |
IEEE Trans. Software Eng. | 2 |
| 1987 | Exploiting Type Inheritance Facilities to Implement Recoverability in Object Based Systems
Graeme N. Dixon, Santosh K. Shrivastava |
SRDS | 2 |
| 1982 | The Design of a Reliable Remote Procedure Call MechanismabstractIn this correspondence we describe the design of a reliable Remote Procedure Call mechanism intended for use in local area networks. Starting from the hardware level that provides primitive facilities for data transmission, we describe how such a mechanism can be constructed. We discuss various design issues involved, including the choice of a message passing system over which the remote call mechanism is to be constructed and the treatment of various abnormal situations such as lost messages and node crashes. We also investigate what the reliability requirements of the Remote Procedure Call mechanism should be with respect to both the application programs using it and the message passing system on which it itself is based. Santosh K. Shrivastava, Fabio Panzieri |
IEEE Trans. Computers | 1 |
| 1981 | Some Critical Comments on the Paper "An Optimal Approach to Fault Tolerant Software Systems Design" by Gannon and Shapiro
Peter A. Lee, J. L. Lloyd, Santosh K. Shrivastava |
IEEE Trans. Software Eng. | 3 |
| 1981 | Structuring Distributed Systems for Recoverability and Crash ResistanceabstractAn object-oriented multilevel model of computation is used to discuss recoverability and crash resistance issues in distributed systems. Of particular importance are the issues that are raised when recoverability and crash resistance properties are desired from objects whose concrete representations are distributed over several nodes. The execution of a program at a node of the system can give rise to a hierarchy of processes executing various parts of the program at different nodes. Recoverability and crash resistance properties are needed to ensure that such a group of processes leave the system state consistent despite faults in the system. Santosh K. Shrivastava |
IEEE Trans. Software Eng. | 1 |
| 1979 | Concurrent Pascal with Backward Error Recovery: Language Features and ExamplesabstractAbstract The programming language Concurrent Pascal has been extended to include some language features that facilitate the writing of fault‐tolerant software. As a result, it is possible now to (1) write operating systems with a measure of fault‐tolerance, and (2) for such an operating system to support fault‐tolerant user programs. The paper describes these language features and illustrates their use with the help of a few working examples. Santosh K. Shrivastava |
Softw. Pract. Exp. | 1 |
| 1979 | Concurrent Pascal with Backward Error Recovery: ImplementationabstractAbstract The implementation of backward error recovery features requires the support of a run time subsystem (called the recovery system) that is responsible for performing the task of state restoration. The recovery system implemented to support the recovery features of Concurrent Pascal includes, for each process, a recovery cache for recording appropriate recovery data. This paper describes the details of the recovery system that was implemented as a part of the interpreter of Brinch Hansen's Concurrent Pascal system. Santosh K. Shrivastava |
Softw. Pract. Exp. | 1 |
| 1978 | Reliable Software: A Selective Annotated BibliographyabstractAbstract A total of 64 references to papers, books and conference proceedings on the subject of software reliability have been selected. Each of these references is provided with an annotation consisting of a paragraph of commentary. Sections of the bibliography are devoted to requirements definition, programming methodology, certification, fault‐tolerance and reliability modelling. Thomas Anderson 0001, Santosh K. Shrivastava |
Softw. Pract. Exp. | 2 |
| 1978 | Sequential Pascal with Recovery BlocksabstractAbstract The programming language Sequential Pascal has been extended to include recovery blocks. This paper describes the modifications made to the kernel and interpreter of Brinch Hansen's Pascal system to support recovery blocks and the associated recovery caches needed for state restoration. Santosh K. Shrivastava |
Softw. Pract. Exp. | 1 |
| 1978 | A Model of Recoverability in Multilevel SystemsabstractBackward error recovery (that is, resetting an erroneous state of a system to a previous error-free state) is an important general technique for recovery from faults in a system, especially those faults which were not foreseen. However, the provision of backward error recovery can be complex, particularly if the implementation of the system is "multilever" and recovery is to be provided at a number of these levels. This paper discusses two distinct categories of multilevel system, and then examines in detail the issues involved in providing backward error recovery in both types of system. Thomas Anderson 0001, Peter A. Lee, Santosh K. Shrivastava |
IEEE Trans. Software Eng. | 3 |
| 1978 | Reliable Resource Allocation Between Unreliable ProcessesabstractBasic error recovery problems between interacting processes are first discussed and the desirability of having separate recovery mechanisms for cooperation and competition is demonstrated. The paper then concentrates on recovery mechanisms for processes competing for the use of the shared resources of a computer system. Appropriate programming language features are developed based on the class and inner features of SIMULA, and on the structuring concepts of recovery blocks and monitors. Santosh K. Shrivastava, Jean-Pierre Banâtre |
IEEE Trans. Software Eng. | 1 |
| 1976 | Systematic Programming of Scheduling AlgorithmsabstractAbstract This paper applies the technique of systematic (or structured) programming for programming scheduling algorithms as encountered in operating system design. Monitors are used for structuring scheduling algorithms and a synchronizing method is proposed for process scheduling. Some fairly difficult scheduling problems are solved systematically to illustrate the usefulness of the monitor concepts and the synchronizing method. Certain implementation aspects are also discussed. Santosh K. Shrivastava |
Softw. Pract. Exp. | 1 |
| 1975 | A View of Concurrent Process SynchronisationabstractThis paper reviews some important ideas which have emerged from the work of Dijkstra. In particular it examines his technique of structuring programs so that the process scheduling algorithms are easy to understand and verify. The paper then reviews some synchronisation techniques and presents the reasons for favouring the techniques that neatly exploit Dijkstra's work. The paper is expository rather than original. Santosh K. Shrivastava |
Comput. J. | 1 |