Alexander B. Romanovsky

dblp:r/AlexanderBRomanovsky · DBLP profile ↗
← Back
105ranked-venue papers
17as first author
7since 2021 · last 2025
0000-0002-4076-3331ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 55 · 6 first-author · 1 since 2021Systems, architecture and hardware · 16 · 6 first-authorSecurity and privacy · 12 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 3 first-author · 2 since 2021Theory of computation · 4 · 3 since 2021Databases, data management, data science and information retrieval · 3Computer networks · 1
YearPublicationVenuePosition
2025 Proof Semantics of Railway Interlocking
Linas Laibinis, Alexei Iliasov, Alexander B. Romanovsky
ABZ3
2024 Safety Invariant Engineering for Interlocking Verification
Alexei Iliasov, Dominic Taylor, Linas Laibinis, Alexander B. Romanovsky
SAFECOMP4
2023 A Refinement-based Formal Development of Cyber-physical Railway Signalling Systems
abstract
For years, formal methods have been successfully applied in the railway domain to formally demonstrate safety of railway systems. Despite that, little has been done in the field of formal methods to address the cyber-physical nature of modern railway signalling systems. In this article, we present an approach for a formal development of cyber-physical railway signalling systems that is based on a refinement-based modelling and proof-based verification. Our approach utilises the Event-B formal specification language together with a hybrid system and communication modelling patterns to developing a generic hybrid railway signalling system model that can be further refined to capture a specific railway signalling system. The main technical contribution of this article is the refinement of the hybrid train Event-B model with other railway signalling sub-systems. The complete model of the cyber-physical railway signalling system was formally proved to ensure a safe rolling stock separation and prevent their derailment. Furthermore, the article demonstrates the advantage of the refinement-based development approach of cyber-physical systems, which enables a problem decomposition and in turn reduction in the verification and modelling effort.
Yamine Aït-Ameur, Sergiy Bogomolov, Guillaume Dupont, Alexei Iliasov, Alexander B. Romanovsky, Paulius Stankaitis
Formal Aspects Comput.5
2023 Practical Verification of Railway Signalling Programs
abstract
SafeCap is a modern toolkit for modelling, simulation and formal verification of railway networks. This paper discusses the use of SafeCap for formal analysis and automated scalable safety verification of solid state interlocking (SSI) programs – a technology at the heart of many railway signalling solutions around the world. The main driving force behind SafeCap development was to make it easy for signalling engineers to use the technology and thus to ensure its smooth industrial deployment. The unique qualities and the novelty of SafeCap are in making the use of formal notations and proofs fully transparent for the engineers. In this paper we explain the formal foundations of the proposed method, its tool support, and its successful application by railway companies in developing industrial signalling projects.
Alexei Iliasov, Dominic Taylor, Linas Laibinis, Alexander B. Romanovsky
IEEE Trans. Dependable Secur. Comput.4
2021 A refinement-based development of a distributed signalling system
abstract
Abstract The decentralised railway signalling systems have a potential to increase capacity, availability and reduce maintenance costs of railway networks. However, given the safety-critical nature of railway signalling and the complexity of novel distributed signalling solutions, their safety should be guaranteed by using thorough system validation methods. To achieve such a high-level of safety assurance of these complex signalling systems, scenario-based testing methods are far from being sufficient despite that they are still widely used in the industry. Formal verification is an alternative approach which provides a rigorous approach to verifying complex systems and has been successfully used in the railway domain. Despite the successes, little work has been done in applying formal methods for distributed railway systems. In our research we are working towards a multifaceted formal development methodology of complex railway signalling systems. The methodology is based on the Event-B modelling language which provides an expressive modelling language, a stepwise development and a proof-based model verification. In this paper, we present the application of the methodology for the development and verification of a distributed protocol for reservation of railway sections. The main challenge of this work is developing a distributed protocol which ensures safety and liveness of the distributed railway system when message delays are allowed in the model.
Paulius Stankaitis, Alexei Iliasov, Tsutomu Kobayashi, Yamine Aït-Ameur, Fuyuki Ishikawa, Alexander B. Romanovsky
Formal Aspects Comput.6
2021 Mutation Testing for Rule-Based Verification of Railway Signaling Data
abstract
Industry applications of formal verification to signaling control tables require formulation of a large number of mathematical conjectures expressing verification rules. It is paramount to establish the validity and completeness of these conjectures. This article discusses a mutation-based validation technique that guides domain experts in the construction of such verification rules. Furthermore, we use genetic programming to quickly generate millions of well-formed data mutations of control tables and to synthesize mutation programs. The technique is illustrated by a synthetic running example and a discussion of our experience in using it in the industrial setting.
Linas Laibinis, Alexei Iliasov, Alexander B. Romanovsky
IEEE Trans. Reliab.3
2021 A customisable pipeline for the semi-automated discovery of online activists and social campaigns on Twitter
abstract
Abstract Substantial research is available on detectinginfluencerson social media platforms. In contrast, comparatively few studies exists on the role ofonline activists, defined informally as users who actively participate in socially-minded online campaigns. Automatically discovering activists who can potentially be approached by organisations that promote social campaigns is important, but not easy, as they are typically active only locally, and, unlike influencers, they are not central to large social media networks. We make the hypothesis that such interesting users can be found on Twitter within temporally and spatially localisedcontexts. We define these as small but topical fragments of the network, containing interactions about social events or campaigns with a significant online footprint. To explore this hypothesis, we have designed an iterative discovery pipeline consisting of two alternating phases of user discovery and context discovery. Multiple iterations of the pipeline result in a growing dataset of user profiles for activists, as well as growing set of online social contexts. This mode of exploration differs significantly from prior techniques that focus on influencers, and presents unique challenges because of the weak online signal available to detect activists. The paper describes the design and implementation of the pipeline as a customisable software framework, where user-defined operational definitions of online activism can be explored. We present an empirical evaluation on two extensive case studies, one concerning healthcare-related campaigns in the UK during 2018, the other related to online activism in Italy during the COVID-19 pandemic.
Flavio Primo, Alexander B. Romanovsky, Rafael Maiani de Mello, Alessandro F. Garcia 0001, Paolo Missier
World Wide Web2
2020 PARMA: Parallelization-Aware Run-Time Management for Energy-Efficient Many-Core Systems
abstract
Performance and energy efficiency considerations have shifted computing paradigms from single-core to many-core architectures. At the same time, traditional speedup models such as Amdahl's Law face challenges in the run-time reasoning for system performance and energy efficiency, because these models typically assume limited variations of the parallel fraction. Moreover, the parallel fraction, which varies dynamically in workloads, is generally unknown at run-time without application-level instrumentation. This article describes novel performance/energy trade-off models based on realistic architectural considerations, which describe the parallel fraction and speedup as functions of performance counter values available in modern processors, removing the need for application-level instrumentation. These are then used to develop a Parallelization-Aware Run-time Management (PARMA) approach. PARMA aims at controlling core allocations and operating voltage/frequency points for energy efficiency, according to the varying workload parallel fractions. The efficacy of our models and the PARMA approach is extensively validated using a number of PARSEC benchmark applications, involving two performance/energy trade-off metrics: energy-delay-product (EDP), typically used in high-performance applications and energy per instruction (EPI), suitable for energy-aware applications. Up to 48 and 68 percent improvements in EDP and EPI have been observed using the PARMA approach compared with parallelization-agnostic methods.
Mohammed A. Noaman Al-Hayanni, Ashur Rafiev, Fei Xia 0001, Rishad A. Shafik, Alexander B. Romanovsky, Alexandre Yakovlev
IEEE Trans. Computers5
2020 Dynamically Partitioning Workflow over Federated Clouds for Optimising the Monetary Cost and Handling Run-Time Failures
abstract
Several real-world problems in domain of healthcare, large scale scientific simulations, and manufacturing are organised as workflow applications. Efficiently managing workflow applications on the Cloud computing data-centres is challenging due to the following problems: (i) they need to perform computation over sensitive data (e.g., Healthcare workflows) hence leading to additional security and legal risks especially considering public cloud environments and (ii) the dynamism of the cloud environment can lead to several run-time problems such as data loss and abnormal termination of workflow task due to failures of computing, storage, and network services. To tackle above challenges, this paper proposes a novel workflow management framework call Deploy on Federated Cloud Framework (DoFCF) that can dynamically partition scientific workflows across federated cloud (public/private) data-centres for minimising the financial cost, adhering to security requirements, while gracefully handling run-time failures. The framework is validated in cloud simulation tool (CloudSim) as well as in a realistic workflow-based cloud platform (e-Science Central). The results showed that our approach is practical and is successful in meeting users security requirements and reduces overall cost, and dynamically adapts to the run-time failures.
Zhenyu Wen, Rawaa Qasha, Zequn Li 0002, Rajiv Ranjan 0001, Paul Watson 0001, Alexander B. Romanovsky
IEEE Trans. Cloud Comput.6
2020 GA-Par: Dependable Microservice Orchestration Framework for Geo-Distributed Clouds
abstract
Recent advances in composing Cloud applications have been driven by deployments of inter-networking heterogeneous microservices across multiple Cloud datacenters. System dependability has been of the upmost importance and criticality to both service vendors and customers. Security, a measurable attribute, is increasingly regarded as the representative example of dependability. Literally, with the increment of microservice types and dynamicity, applications are exposed to aggravated internal security threats and externally environmental uncertainties. Existing work mainly focuses on the QoS-aware composition of native VM-based Cloud application components, while ignoring uncertainties and security risks among interactive and interdependent container-based microservices. Still, orchestrating a set of microservices across datacenters under those constraints remains computationally intractable. This paper describes a new dependable microservice orchestration framework GA-Par to effectively select and deploy microservices whilst reducing the discrepancy between user security requirements and actual service provision. We adopt a hybrid (both whitebox and blackbox based) approach to measure the satisfaction of security requirement and the environmental impact of network QoS on system dependability. Due to the exponential grow of solution space, we develop a parallel Genetic Algorithm framework based on Spark to accelerate the operations for calculating the optimal or near-optimal solution. Large-scale real world datasets are utilized to validate models and orchestration approach. Experiments show that our solution outperforms the greedy-based security aware method with 42.34 percent improvement. GA-Par is roughly 4× faster than a Hadoop-based genetic algorithm solver and the effectiveness can be constantly guaranteed under different application scales.
Zhenyu Wen, Tao Lin 0004, Renyu Yang, Shouling Ji, Rajiv Ranjan 0001, Alexander B. Romanovsky, Chang-Ting Lin, Jie Xu 0007
IEEE Trans. Parallel Distributed Syst.6
2020 From Analyzing Operating System Vulnerabilities to Designing Multiversion Intrusion-Tolerant Architectures
abstract
This paper analyzes security problems of modern computer systems caused by vulnerabilities in their operating systems (OSs). Our scrutiny of widely used enterprise OSs focuses on their vulnerabilities by examining the statistical data available on how vulnerabilities in these systems are disclosed and eliminated, and by assessing their criticality. This is done by using statistics from both the National Vulnerabilities Database and the Common Vulnerabilities and Exposures System. The specific technical areas the paper covers are the quantitative assessment of forever-day vulnerabilities, estimation of days-of-grey-risk, the analysis of the vulnerabilities severity and their distributions by attack vector and impact on security properties. In addition, the study aims to explore those vulnerabilities that have been found across a diverse range of OSs. This leads us to analyzing how different intrusion-tolerant architectures deploying the OS diversity impact availability, integrity, and confidentiality.
Anatoliy Gorbenko, Alexander B. Romanovsky, Olga Tarasyuk, Oleksandr Biloborodov
IEEE Trans. Reliab.2
2019 Modelling Hybrid Train Speed Controller using Proof and Refinement
abstract
The modern radio-based railway signalling systems aim to increase network's capacity by enabling trains to run closer to each other. At the core of such systems is train's on-board computer (discrete) responsible for computing and controlling the speed (continuous) of the train. Such systems are best captured by hybrid models, which capture discrete and continuous system's aspects. Hybrid models are notoriously difficult to model and verify, in our research we address this problem by applying hybrid systems' modelling patterns and stepwise refinement for developing hybrid train speed controller model.
Paulius Stankaitis, Guillaume Dupont, Neeraj Kumar Singh 0001, Yamine Aït-Ameur, Alexei Iliasov, Alexander B. Romanovsky
ICECCS6
2019 A Customisable Pipeline for Continuously Harvesting Socially-Minded Twitter Users
Flavio Primo, Paolo Missier, Alexander B. Romanovsky, Mickael Figueredo, Nélio Cacho
ICWE3
2019 Fault tolerant internet computing: Benchmarking and modelling trade-offs between availability, latency and consistency
Anatoliy Gorbenko, Alexander B. Romanovsky, Olga Tarasyuk
J. Netw. Comput. Appl.2
2018 DroidEH: An Exception Handling Mechanism for Android Applications
abstract
App crashing is the most common cause of complaints about Android mobile phone apps according to recent studies. Since most Android applications are written in Java, exception handling is the primary mechanism they employ to report and handle errors, similar to standard Java applications. Unfortunately, the exception handling mechanism for the Android platform has two liabilities: (1) the "Terminate ALL" approach and (2) a lack of a holistic view on exceptional behavior. As a consequence, exceptions easily get "out of control" and, as system development progresses, exceptional control flows become less well-understood, with potentially negative effects on program reliability. This paper presents an innovative exception handling mechanism for the Android platform, named DroidEH, that provides abstractions to support systematic engineering of holistic fault tolerance by applying cross-cutting reasoning about systems and their components.
Juliana Oliveira, Hivana Macedo, Nélio Cacho, Alexander B. Romanovsky
ISSRE4
2018 Cost-aware Scheduling of Software Processes Execution in the Cloud
abstract
Using cloud computing to execute software processes brings several benefits to software development. In a previous work, we proposed a reference architecture, which treats software processes as wor ...
Sami Alajrami, Alexander B. Romanovsky, Barbara Gallina
MODELSWARD2
2018 Formal Verification of Signalling Programs with SafeCap
Alexei Iliasov, Dominic Taylor, Linas Laibinis, Alexander B. Romanovsky
SAFECOMP4
2018 VazaDengue: An information system for preventing and combating mosquito-borne diseases with social networks
Leonardo da Silva Sousa, Rafael Maiani de Mello, Diego Cedrim, Alessandro F. Garcia 0001, Paolo Missier, Anderson G. Uchôa, Anderson Oliveira, Alexander B. Romanovsky
Inf. Syst.8
2017 Recruiting from the Network: Discovering Twitter Users Who Can Help Combat Zika Epidemics
Paolo Missier, Callum McClean, Jonathan Carlton, Diego Cedrim, Leonardo da Silva Sousa, Alessandro F. Garcia 0001, Alexandre Plastino 0001, Alexander B. Romanovsky
ICWE8
2017 Experience Report: Study of Vulnerabilities of Enterprise Operating Systems
abstract
This experience report analyses security problems of modern computer systems caused by vulnerabilities in their operating systems. An aggregated vulnerability database has been developed by joining vulnerability records from two publicly available vulnerability databases: the Common Vulnerabilities and Exposures system (CVE) and the National Vulnerabilities database (NVD). The aggregated data allow us to investigate the stages of the vulnerability life cycle, vulnerability disclosure and the elimination statistics for different operating systems. The specific technical areas the paper covers are the quantitative assessment of vulnerabilities discovered and fixed in operating systems, the estimation of time that vendors spend on patch issuing, and the analysis of the vulnerability criticality and identification of vulnerabilities common for different operating systems.
Anatoliy Gorbenko, Alexander B. Romanovsky, Olga Tarasyuk, Oleksandr Biloborodov
ISSRE2
2017 Cost Effective, Reliable and Secure Workflow Deployment over Federated Clouds
abstract
The significant growth in cloud computing has led to increasing number of cloud providers, each offering their service under different conditions - one might be more secure whilst another might be less expensive or more reliable. At the same time user applications have become more and more complex. Often, they consist of a diverse collection of software components, and need to handle variable workloads, which poses different requirements on the infrastructure. Therefore, many organisations are considering using a combination of different clouds to satisfy these needs. It raises, however, a non-trivial issue of how to select the best combination of clouds to meet the application requirements. This paper presents a novel algorithm to deploy workflow applications on federated clouds. First, we introduce an entropy-based method to quantify the most reliable workflow deployments. Second, we apply an extension of the Bell-LaPadula Multi-Level security model to address application security requirements. Finally, we optimise deployment in terms of its entropy and also its monetary cost, taking into account the cost of computing power, data storage and inter-cloud communication. We implemented our new approach and compared it against two existing scheduling algorithms: Extended Dynamic Constraint Algorithm (EDCA) and Extended Biobjective dynamic level scheduling (EBDLS). We show that our algorithm can find deployments that are of equivalent reliability but are less expensive and meet security requirements. We have validated our solution through a set of realistic scientific workflows, using well-known cloud simulation tools (WorkflowSim and DynamicCloudSim) and a realistic cloud based data analysis system (e-Science Central).
Zhenyu Wen, Jacek Cala, Paul Watson 0001, Alexander B. Romanovsky
IEEE Trans. Serv. Comput.4
2016 Selective abstraction and stochastic methods for scalable power modelling of heterogeneous systems
abstract
With the increase of system complexity in both platforms and applications, power modelling of heterogeneous systems is facing grand challenges from the model scalability issue. To address these challenges, this paper studies two systematic methods: selective abstraction and stochastic techniques. The concept of selective abstraction via black-boxing is realised using hierarchical modelling and cross-layer cuts, respecting the concepts of boxability and error contamination. The stochastic aspect is formally underpinned by Stochastic Activity Networks (SANs). The proposed method is validated with experimental results from Odroid XU3 heterogeneous 8-core platform and is demonstrated to maintain high accuracy while improving scalability.
Ashur Rafiev, Fei Xia 0001, Alexei Iliasov, Rem Gensh, Ali Aalsaud, Alexander B. Romanovsky, Alexandre Yakovlev
FDL6
2016 Proving Event-B Models with Reusable Generic Lemmas
Alexei Iliasov, Paulius Stankaitis, Alexander B. Romanovsky
ICFEM3
2016 EXE-SPEM: Towards Cloud-based Executable Software Process Models
abstract
Executing software processes in the cloud can bring several benefits to software development. In this paper, we discuss the benefits and considerations of cloud-based software processes. EXE-SPEM is our extension of the Software and Systems Process Engineering (SPEM2.0) Meta-model to support creating cloud-based executable software process models. Since SPEM2.0 is a visual modelling language, we introduce an XML notation meta-model and mapping rules from EXE-SPEM to this notation which can be executed in a workflow engine. We demonstrate our approach by modelling an example software process using EXE-SPEM and mapping it to the XML notation.
Sami Alajrami, Barbara Gallina, Alexander B. Romanovsky
MODELSWARD3
2016 Software Development in the Post-PC Era: Towards Software Development as a Service
Sami Alajrami, Alexander B. Romanovsky, Barbara Gallina
PROFES2
2016 Towards Cloud-Based Enactment of Safety-Related Processes
Sami Alajrami, Barbara Gallina, Irfan Sljivo, Alexander B. Romanovsky, Petter Isberg
SAFECOMP4
2015 Cost Effective, Reliable, and Secure Workflow Deployment over Federated Clouds
abstract
The federation of clouds can provide benefits for cloud-based applications. Different clouds have different advantages - one might be more reliable whilst another might be more secure or less expensive. However, being able to select the best combination of clouds to meet the application requirements is not trivial. This paper presents a novel algorithm to deploy workflow applications on federated clouds. Firstly, we introduce an entropy-based method to quantify the most reliable workflow deployments. Secondly, we apply an extension of the Bell-LaPadula Multi-Level security model to meet application security requirements. Finally, we optimise deployment in terms of its entropy and also its monetary cost, taking into account the price of computing power, data storage and inter-cloud communication. To evaluate the new algorithm we compared it against two existing scheduling algorithms: Dynamic Constraint Algorithm (DCA) and Biobjective dynamic level scheduling (BDLS). We show that our algorithm can find deployments that are of equivalent reliability, but are less expensive and also meet security requirements. We have validated our solution using workflows implemented in the e-Science Central cloud-based data analysis system.
Zhenyu Wen, Jacek Cala, Paul Watson 0001, Alexander B. Romanovsky
CLOUD4
2015 The Impact of Consistency on System Latency in Fault Tolerant Internet Computing
Olga Tarasyuk, Anatoliy Gorbenko, Alexander B. Romanovsky, Vyacheslav S. Kharchenko, Vitalii Ruban
DAIS3
2015 A reactive architecture for cloud-based system engineering
abstract
The paper introduces an architecture to support system engineering on the cloud. It employs the main benefits of the cloud: scalability, parallelism, cost-effectiveness, multi-user access and flexibility. The architecture includes an open toolbox which provides tools as a service to support various phases of system engineering. The architecture uses the Open Services for Life-cycle Collaboration (OSLC) technology to create a reactive middleware that informs all stakeholders about any changes in the development artefacts. It facilitates the interoperability of tools and enables the workflow of tools to support complex engineering steps. Another component of the architecture is a shared repository of artefacts. All the artefacts generated during a system engineering process are stored in the repository, and can be accessed by relevant stakeholders. The shared repository also serves as a platform to support a protocol for formal model decomposition and group work on the decomposed models. Finally, the architecture includes components for ensuring the dependability of the system engineering process.
David Adjepon-Yamoah, Alexander B. Romanovsky, Alexei Iliasov
ICSSP2
2015 A Formal Specification and Prototyping Language for Multi-core System Management
abstract
We relate the experience of a defining a formal domain specific language (DSL) for the construction and reasoning about OS-level management logic of multi-core systems. The approach is based on a novel, iterative development principle where results of prototyping studies feed back into the next language revision. We illustrate the DSL with several examples of executable scripts.
Alexei Iliasov, Ashur Rafiev, Fei Xia 0001, Rem Gensh, Alexander B. Romanovsky, Alexandre Yakovlev
PDP5
2015 From Requirements Engineering to Safety Assurance: Refinement Approach
Linas Laibinis, Elena Troubitsyna, Yuliya Prokhorova, Alexei Iliasov, Alexander B. Romanovsky
SETTA5
2014 Special issue on Automated Verification of Critical Systems (AVoCS'11)
Cliff B. Jones, Alexander B. Romanovsky
Sci. Comput. Program.2
2014 Synthesis of Processor Instruction Sets from High-Level ISA Specifications
abstract
As processors continue to get exponentially cheaper for end users following Moore’s law, the costs involved in their design keep growing, also at an exponential rate. The reason is ever increasing complexity of processors, which modern EDA tools struggle to keep up with. This paper focuses on the design of Instruction Set Architecture (ISA), a significant part of the whole processor design flow. Optimal design of an instruction set for a particular combination of available hardware resources and software requirements is crucial for building processors with high performance and energy efficiency, and is a challenging task involving a lot of heuristics and high-level design decisions. This paper presents a new compositional approach to formal specification and synthesis of ISAs. The approach is based on a new formalism, called Conditional Partial Order Graphs, capable of capturing common behavioural patterns shared by processor instructions, and therefore providing a very compact and efficient way to represent and manipulate ISAs. The Event-B modelling framework is used as a formal specification and verification back-end to guarantee correctness of ISA specifications. We demonstrate benefits of the presented methodology on several examples, including Intel 8051 microcontroller.
Andrey Mokhov, Alexei Iliasov, Danil Sokolov, Maxim Rykunov, Alexandre Yakovlev, Alexander B. Romanovsky
IEEE Trans. Computers6
2013 Fault modelling for systems of systems
abstract
This paper proposes a systematic model-based approach to the architectural description of faults and fault tolerance mechanisms in systems of systems (SoSs). The challenges of engineering dependable SoSs motivate a proposal for the view elements that would be needed to support a fault tolerance profile for SoSs using the Systems Modelling Language (SysML). The effectiveness of the approach is evaluated on a case study based on a real emergency response SoS. Results suggest that this is a promising approach, and that a comprehensive solution to the engineering of dependable SoSs requires that such a profile is linked to methods and tools for requirements elicitation, safety analysis, architectural design and formal verification.
Zoe Andrews, John S. Fitzgerald, Richard John Payne, Alexander B. Romanovsky
ISADS4
2013 The SafeCap Platform for Modelling Railway Safety and Capacity
Alexei Iliasov, Ilya Lopatkin, Alexander B. Romanovsky
SAFECOMP3
2013 Developing mode-rich satellite software by refinement in Event-B
Alexei Iliasov, Elena Troubitsyna, Linas Laibinis, Alexander B. Romanovsky, Kimmo Varpaaniemi, Dubravka Ilic, Timo Latvala
Sci. Comput. Program.4
2011 Formal Derivation of a Distributed Program in Event B
Alexei Iliasov, Linas Laibinis, Elena Troubitsyna, Alexander B. Romanovsky
ICFEM4
2011 Rigorous Development of Dependable Systems Using Fault Tolerance Views
abstract
This paper introduces the Mode and Fault Tolerance Views approach to stepwise rigorous development of critical systems. It supports systematic, structured and recursive modelling of system fault tolerance, including error detection, error recovery and degraded modes. Built on our previous work extending the Event-B method with reasoning about fault tolerance, the paper focuses on a practical application and evaluation of the approach. The proposed modelling approach is backed by an integrated toolset. The paper is illustrated with a case study from the aerospace domain.
Ilya Lopatkin, Alexei Iliasov, Alexander B. Romanovsky
ISSRE3
2010 Developing Mode-Rich Satellite Software by Refinement in Event B
Alexei Iliasov, Elena Troubitsyna, Linas Laibinis, Alexander B. Romanovsky, Kimmo Varpaaniemi, Dubravka Ilic, Timo Latvala
FMICS4
2010 Patterns for Modelling Time and Consistency in Business Information Systems
abstract
Maintaining semantic consistency of data is a significant problem in distributed information systems, particularly those on which a business may depend. Our current work aims to use Event-B and the Rodin tools to support the specification and design of such systems in a way that integrates well into existing development processes. This paper presents Event-B patterns that may be used to represent recovery from time-bounded inconsistency and illustrates their use in a model derived from industrial applications.
Jeremy W. Bryans, John S. Fitzgerald, Alexander B. Romanovsky, Andreas Roth 0001
ICECCS3
2010 Verifying Mode Consistency for On-Board Satellite Software
Alexei Iliasov, Elena Troubitsyna, Linas Laibinis, Alexander B. Romanovsky, Kimmo Varpaaniemi, Pauli Väisänen, Dubravka Ilic, Timo Latvala
SAFECOMP4
2010 Real Distribution of Response Time Instability in Service-Oriented Architecture
abstract
This paper reports our practical experience of benchmarking a complex System Biology Web Service, and investigates the instability of its behaviour and the delays induced by the communication medium. We present the results of our statistical data analysis and distributions which fit and predict the response time instability typical of Service-Oriented Architectures (SOAs) built over the Internet. Our experiment has shown that the request processing time of the target e-science Web Service (WS) has a higher instability than the network round trip time. It has been found that by using a particular theoretical distribution, within short time intervals the request processing time can be represented better than the network round trip time. Moreover, certain characteristics of the probability distribution series of the round trip time make it particularly difficult to fit them theoretically. The experimental work reported in the paper supports our claim that dealing with the uncertainty inherent in the very nature of SOA and WSs is one of the main challenges in building dependable service-oriented systems. In particular, this uncertainty exhibits itself through very unstable web service response times and Internet data transfer delays that are hard to predict. Our findings indicate that the more experimental data is considered the less precise distributional approximations become. The paper concludes with a discussion of the lessons learnt about the analysis techniques to be used in such experiments, the validity of the data, the main causes of uncertainty and possible remedial actions.
Anatoliy Gorbenko, Vyacheslav S. Kharchenko, Seyran Mamutov, Olga Tarasyuk, Yuhui Chen, Alexander B. Romanovsky
SRDS6
2010 Guest Editors' Introduction to the Special Section on Exception Handling: From Requirements to Software Maintenance
abstract
The four papers in this special section focus on topics related to exception handling.
Alessandro F. Garcia 0001, Alexander B. Romanovsky, Valérie Issarny
IEEE Trans. Software Eng.2
2009 Formal Modelling and Analysis of Business Information Applications with Fault Tolerant Middleware
abstract
Distributed information systems are critical to the functioning of many businesses; designing them to be dependable is a challenging but important task. We report our experience in using formal methods to enhance processes and tools for development of business information software based on service-oriented architectures. In our work, which takes place in an industrial setting, we focus on the configuration of middleware, verifying application-level requirements in the presence of faults. In pilot studies provided by SAP, we used the Event-B formalism and the open Rodin tools platform to prove properties of models of business protocols and expose weaknesses of certain middleware configurations with respect to particular protocols. We then extended the approach to use models automatically generated from diagrammatic design tools, opening the possibility of seamless integration with current development environments. Increased automation in the verification process, through domain-specific models and theories, is a goal for future work.
Jeremy W. Bryans, John S. Fitzgerald, Alexander B. Romanovsky, Andreas Roth 0001
ICECCS3
2009 Benchmarking Dependability of a System Biology Application
abstract
In this paper we report our practical experience in benchmarking a System Biology Web Service, and investigate instability of its performance and the delays induced by the communication medium. We discuss the results of a statistical data analysis and discuss the causes affecting the Web Service performance. The uncertainty discovered in Web Services operations reduces the overall dependability of Service-Oriented Architecture and require specific resilience techniques.
Yuhui Chen, Alexander B. Romanovsky, Anatoliy Gorbenko, Vyacheslav S. Kharchenko, Seyran Mamutov, Olga Tarasyuk
ICECCS2
2009 Modal Systems: Specification, Refinement and Realisation
Fernando Luís Dotti, Alexei Iliasov, Leila Ribeiro 0001, Alexander B. Romanovsky
ICFEM4
2009 Frameworks for designing and implementing dependable systems using Coordinated Atomic Actions: A comparative study
Alfredo Capozucca, Nicolas Guelfi, Patrizio Pelliccione, Alexander B. Romanovsky, Avelino Francisco Zorzo
J. Syst. Softw.4
2009 Improving reliability of cooperative concurrent systems with exception flow analysis
Fernando Castor Filho, Alexander B. Romanovsky, Cecília M. F. Rubira
J. Syst. Softw.2
2008 How to Enhance UDDI with Dependability Capabilities
abstract
How dependability is to be assessed and ensured during Web service operation and how unbiased and trusted mechanisms supporting this are to be developed are still open issues. This paper addresses the following questions: who should publish dependability parameters, in which way they should be distributed, and who (and how) should monitor these parameters in the global service-oriented architecture. We discuss several techniques of on-line dependability monitoring and measurement, which extend the UDDI (Universal Description, Discovery and Integration) business registry with dependability metadata publishing and monitoring capabilities. The paper also proposes UDDI add-ons and light-weight user-side mechanisms for public operational and exceptional reporting.
Anatoliy Gorbenko, Alexander B. Romanovsky, Vyacheslav S. Kharchenko
COMPSAC2
2008 An aspect-oriented software architecture for code mobility
abstract
Abstract Mobile agents have come forward as a technique for tackling the complexity of open distributed applications. However, the pervasive nature of code mobility implies that it cannot be modularized using only object‐oriented (OO) concepts. In fact, developers frequently evidence the presence of mobility scattering in their system's modules. Despite these problems, they usually rely on OO application programming interfaces (APIs) offered by the mobility platforms. Such classical API‐oriented designs suffer a number of architectural restrictions, and there is a pressing need for empowering developers with an architectural framework supporting a flexible incorporation of code mobility in the agent applications. This work presents an aspect‐oriented software architecture, called ArchM, ensuring that code mobility has an enhanced modularization and variability in agent systems, and is straightforwardly introduced in otherwise stationary agents. It addresses OO APIs' restrictions and is independent of specific platforms and applications. An ArchM implementation also overcomes fine‐grained problems related to mobility tangling and scattering at the implementation level. The usefulness and usability of ArchM are assessed within the context of two case studies and through its composition with two mobility platforms. Copyright © 2008 John Wiley & Sons, Ltd.
Cidiane Lobato, Alessandro F. Garcia 0001, Alexander B. Romanovsky, Carlos José Pereira de Lucena
Softw. Pract. Exp.3
2007 A Framework for Open Distributed System Design
abstract
Building open distributed systems is an even more challenging task than building distributed systems, as their components are loosely synchronised, can move, become disconnected, and their behaviour may depend on the changing context. The approach we are putting forward relies on using a combination of formal methods applied for rigorous development of the critical parts of the system and a set of design abstractions proposed specifically for the open context-aware applications and supported by a special middleware. Our middleware provides system structuring through the concepts of roles, agents, locations and scopes, making it easier for application developers to achieve fault tolerance. We demonstrate our approach using a case study, in which we show the whole process of developing an ambient campus application - an example of open distributed systems - including its formal specification, refinement, and implementation.
Alexei Iliasov, Alexander B. Romanovsky, Budi Arief
COMPSAC (2)2
2007 On Rigorous Design and Implementation of Fault Tolerant Ambient Systems
abstract
Developing fault tolerant ambient systems requires many challenging factors to be considered due to the nature of such systems, which tend to contain a lot of mobile elements that change their behavior depending on the surrounding environment, as well as the possibility of their disconnection and reconnection. It is therefore necessary to construct the critical parts of fault tolerant ambient systems in a rigorous manner. This can be achieved by deploying formal approach at the design stage, coupled with sound framework and support at the implementation stage. In this paper, we briefly describe a middleware that we developed to provide system structuring through the concepts of roles, agents, locations and scopes, making it easier for the developers to achieve fault tolerance. We then outline our experience in developing an ambient lecture system using the combination of formal approach and our middleware
Alexei Iliasov, Alexander B. Romanovsky, Budi Arief, Linas Laibinis, Elena Troubitsyna
ISORC2
2007 Coordinated Atomic Actions for Dependable Distributed Systems: the Current State in Concepts, Semantics and Verification Means
abstract
Coordinated Atomic Actions (CAAs) have been introduced about ten years ago as a conceptual framework for developing fault-tolerant concurrent systems. All the work done since then extended the CAA framework with the capabilities to model, verify, and implement concurrent distributed systems following pre-defined development methodologies. As a result, CAAs, compared to other approaches available, offer a rich set of means for engineering dependable systems. Nevertheless, it is sometimes difficult to have a global and analytical view of all the features available as this concept provides a number of features which need to be applied in combination. The main contribution of this paper is in presenting a complete state-of-the-art overview of the work done around CAAs from the three perspectives: the definitions of the fundamental concepts, their various semantics and the means supporting formal verification. This paper is useful for the potential CAAs users in helping them to avoid misinterpretation when employing all the available features. Finally, our paper should contribute in better understanding of the likely directions in which the CAA framework may evolve in the near future.
Barbara Gallina, Nicolas Guelfi, Alexander B. Romanovsky
ISSRE3
2007 EFTS 2007: the 2nd international workshop on engineering fault tolerant systems
abstract
Fault tolerance engineering has been advocated as one of the main approaches to ensuring the overall system dependability. The 2nd International Workshop on Engineering Fault Tolerant Systems (EFTS 2007) aims to investigate how fault tolerance mechanisms can be taken into account when engineering complex software systems and to improve our understanding of where and how fault-tolerance should be integrated in the software life-cycle. The focus of the workshop is on developing novel models to be applied at different abstraction levels (requirements, architecture and design models for fault tolerance, together with new implementation schemes), innovative technologies (tools and frameworks for implementing distributed fault tolerant systems) and advanced verification environments (to assess the achieved level of fault tolerance and to evaluate the dependability properties of the systems). Recently there has been a growing interest in the areas directly related and overlapping with fault tolerance, such as self-healing, resilience, self-adaptation and self-management. The topics related to engineering of systems with such properties are in the scope of the workshop as the intention is to improve the current understanding of how fault tolerance engineering can benefit from research on these areas.
Nicolas Guelfi, Henry Muccini, Patrizio Pelliccione, Alexander B. Romanovsky
ESEC/SIGSOFT FSE4
2007 Architecting Fault Tolerant Systems
abstract
While typical solutions focus on fault tolerance (and specifically, exception handling) during the design and implementation phases of the software life-cycle (e.g., Java and Windows NT exception handling), more recently the need for explicit exception handling solutions during the entire life cycle has been advocated by some researchers. Several solutions have been proposed for fault tolerance via exception handling at the software architecture and component levels. This paper describes how the two concepts of fault tolerance and software architectures have been integrated so far. It is structured in two parts (overview on fault tolerance and exception handling, and integrating fault tolerance into software architecture) and is based on a survey study on architecting fault tolerant systems where more than fifteen approaches have been analyzed and classified. This paper concludes by identifying those issues that remain still open and require deeper investigation.
Henry Muccini, Patrizio Pelliccione, Alexander B. Romanovsky
WICSA3
2006 Workshop on Architecting Dependable Systems (WADS)
abstract
This workshop will continue the initiative, which started four years ago, of bringing together the international communities of dependability and software architectures. The first workshop on Architecting Dependable Systems was organised during the International Conference on Software Engineering 2002 (ICSE), and since then five workshops were organised and three books were published (http://www.cs.kent.ac.uk/wads). This activity has culminated in 2004 with the organisation of the Twin Workshops that were held at ICSE 2004 and DSN 2004. This series of workshops have shown to be a fertile ground for both communities for clarifying approaches that have been previously tried and succeeded, as well as those that have been tried but have not yet shown to be successful. This not only helps avoid the reinvention of the wheel, but also clarifies and promotes areas where the most promising research may lie.
Rogério de Lemos, Cristina Gacek, Alexander B. Romanovsky
DSN3
2006 Fifth workshop on software engineering for large-scale multi-agent systems (SELMAS)
abstract
Software is becoming present in every aspect of our lives, pushing us inevitably towards a world of ambient computing systems. Multi-agent systems (MAS) are a prominent technology which facilitates modeling and development of large-scale distributed systems. In recent years, software engineering research has focused on methodologies and techniques for improving MAS design and implementation. However, making large MAS dependable is still an open issue. The Fifth Workshop on Software Engineering for Large-Scale Multi-Agent Systems (SELMAS 2006) aims to bring together academic, industrial and commercial communities interested in agent-oriented software engineering topics to discuss the different technologies being defined and used in the development of dependable MAS.
Ricardo Choren, Ho-fung Leung, Alessandro F. Garcia 0001, Carlos José Pereira de Lucena, Holger Giese, Alexander B. Romanovsky
ICSE6
2006 Looking Ahead in Open Multithreaded Transactions
abstract
Open multithreaded transactions constitute building blocks that allow a developer to design and structure the execution of complex distributed systems featuring cooperative and competitive concurrency in a reliable way. In this paper we describe an optimization to the standard open multithreaded transaction model that does not impose any participant synchronization when committing a transaction, but still provides the same execution semantics. This optimization - letting participants "look ahead" and continue their execution on the outside of the transaction - makes it possible to speed up the execution of in individual transaction with multiple participants tremendously. The paper describes all technical issues that had to be solved, e.g. adapting concurrency control of transactional objects to be look-ahead aware, adapting joining rules for look-ahead participants, and re-defining exception handling in the presence of look-ahead
Maxime Monod, Jörg Kienzle, Alexander B. Romanovsky
ISORC3
2006 CAA-DRIP: a framework for implementing Coordinated Atomic Actions
abstract
This paper presents an implementation framework, called CAA-DRIP, that has been defined to allow a straightforward implementation of dependable distributed applications designed using the coordinated atomic action (CAA) paradigm. CAAs provide a coherent set of concepts adapted to the design of fault tolerant distributed systems that includes: structured transactions, distribution, cooperation, competition, and forward and backward error recovery mechanisms triggered by exceptions. DRIP (dependable remote interacting processes) is an efficient Java implementation framework, which provides support for implementing "dependable multiparty interactions (DMI)" which includes a general exception handling mechanism. As DMI has a softer exception handling semantics with respect to CAA semantics, a CAA design can be implemented by DRIP. The aim of the CAA-DRIP framework is to provide a set of Java classes that allows programmers to implement only the semantics of CAAs with the same terminology and concepts at the design and implementation levels. The new framework simplifies the implementation phase and at the same time reduces the size of the final system since it requires fewer number of instances for creating a CAA at runtime. Details of these improvements as well as a precise description of the CAAs behaviour in terms of state charts, which is used as a reference model to define the CAA-DRIP framework, are presented in this paper
Alfredo Capozucca, Nicolas Guelfi, Patrizio Pelliccione, Alexander B. Romanovsky, Avelino Francisco Zorzo
ISSRE4
2006 Architecting dependable systems
Rogério de Lemos, Cristina Gacek, Alexander B. Romanovsky
J. Syst. Softw.3
2005 Exception Handling in Coordination-Based Mobile Environments
abstract
Mobile agent systems have many attractive features including asynchrony, openness, dynamicity and anonymity, which makes them indispensable in designing complex modern applications that involve moving devices, human participants and software. To be comprehensive this list should include fault tolerance, yet as our analysis shows, this property is, unfortunately, often overlooked by middleware designers. A few existing solutions for fault tolerant mobile agents are developed mainly for tolerating hardware faults without providing any general support for application-specific recovery. In this paper we describe a novel exception handling model that allows application-specific recovery in coordination-based systems consisting of mobile agents. The proposed mechanism is general enough to be used in both loosely-and tightly-coupled communication models. The general ideas behind the mechanism are applied in the context of the Lime middleware.
Alexei Iliasov, Alexander B. Romanovsky
COMPSAC (1)2
2005 Software engineering for large-scale multi-agent systems - SELMAS'05
abstract
is becoming present in every aspect of our lives, pushing us inevitably towards a world of distributed, context-aware computing systems. SELMAS'05, Software Everywhere - Context-Aware Agents, builds on the success of precedent SELMAS workshops, but with a special emphasis on the impact of the agent technology in the development of large context-aware systems. SELMAS has a track record of bringing together researchers and practitioners with a variety of perspectives in order to engage in lively discussion and debate.
Alessandro F. Garcia 0001, Ricardo Choren, Carlos José Pereira de Lucena, Alexander B. Romanovsky, Tom Holvoet, Paolo Giorgini
ICSE4
2005 Workshop on architecting dependable systems (WADS 2005)
abstract
This workshop summary gives a brief overview of the workshop on "Architecting Dependable Systems" held in conjunction with the ICSE 2005. The main aim of this workshop is to promote cross-fertilization between the software architecture and dependability communities. We believe that both of them will benefit from clarifying approaches that have been previously tested and have succeeded as well as those that have been tried but have not yet been shown to be successful.
Rogério de Lemos, Alexander B. Romanovsky
ICSE2
2004 Twin Workshops on Architecting Dependable Systems (WADS 2004)
Rogério de Lemos, Cristina Gacek, Alexander B. Romanovsky
DSN3
2004 Software Engineering for Large-Scale Multi-agent Systems - SELMAS'04
abstract
The development of multiagent systems (MAS) is not a trivial task. In addition, with the advances in Internet technologies, MAS are undergoing a transition from closed to open architectures composed of a huge number of autonomous agents, which operate and move across different environments. In fact, openness introduces additional complexity to the system modeling, design and implementation. It also impacts on most quality attributes of MAS, including scalability, interoperability, reliability and adaptability. This paper brings together researchers and practitioners to discuss the current state and future direction of research in software engineering for open MAS. A particular interest is to understand those issues in the agent technology that make it difficult and/or improve the production of large open systems.
Ricardo Choren, Alessandro F. Garcia 0001, Carlos José Pereira de Lucena, Martin L. Griss, David Chenho Kung, Naftaly H. Minsky, Alexander B. Romanovsky
ICSE7
2004 Twin Workshops on Architecting Dependable Systems (WADS 2004)
Rogério de Lemos, Cristina Gacek, Alexander B. Romanovsky
ICSE3
2003 ICSE 2003 Workshop on Software Architectures for Dependable Systems
abstract
This workshop summary gives a brief overview of a one-day workshop on "Software Architectures for Dependable Systems" held in conjunctions with ICSE 2003.
Rogério de Lemos, Cristina Gacek, Alexander B. Romanovsky
ICSE3
2003 Software Engineering for Large-Scale Multi-Agent Systems - SELMAS'2003
abstract
Objects and agents are abstractions that exhibit Points of similarity, but the development of multi-agent systems (MASs) poses other challenges to Software Engineering since software agents are inherently more complex entities. In addition, a large MAS needs to satisfy multiple stringent requirements such as reliability, trustability, security, interoperability, scalability, reusability, and maintainability. This workshop brings together researchers and practitioners to discuss the current state of the art and the future research directions in software engineering for large-scale MASs. A particular interest is to understand those issues in the agent technology that make it difficult and/or improve the production of complex distributed systems.
Carlos José Pereira de Lucena, Alberto Sardinha, Alessandro F. Garcia 0001, Alexander B. Romanovsky, Jaelson Brelaz de Castro, Paulo S. C. Alencar, Donald D. Cowan
ICSE4
2003 Structuring Integrated Web Applications for Fault Tolerance
abstract
This paper shows how modern structuring techniques can be employed in integrating complex web applications such as travel agency systems. The main challenges the developers of such systems face are dealing with legacy web services and incorporating means for tolerating errors. Because of the very nature of such systems, exception handling is the main recovery technique to be applied in their development. We employ coordinated atomic actions to allow disciplined handling of such abnormal situations by recursively structuring the integrated system and by associating handlers with such actions. We use protective wrappers in such a way that each operation on legacy components is transformed into an atomic action with a well-defined interface. To accommodate a combined use of several ready-made environments (such as communication packages, services and run-time supports), we employ a multilevel exception handling. We believe that these techniques are generally applicable for both: structuring integrated web applications and providing their fault tolerance.
Alexander B. Romanovsky, Panos Periorellis, Avelino Francisco Zorzo
ISADS1
2003 Integrating COTS Software Components into Dependable Software Architectures
abstract
This paper considers the problem of integrating commercial off-the-shelf (COTS) software components into systems with high dependability requirements. These components, by their very nature, are built to be reused as black boxes that cannot be modified. Instead, the system architect has to rely on techniques external with respect to the component for resolving mismatches of the services required and provided that might arise in the interaction of the component and its environment. This paper proposes an architectural solution to turning COTS components into idealised fault-tolerant COTS components by adding protective wrappers to them.
Paulo Asterio de Castro Guerra, Alexander B. Romanovsky, Rogério de Lemos
ISORC2
2003 A fault-tolerant software architecture for COTS-based software systems
abstract
This paper considers the problem of integrating Commercial offthe-shelf (COTS) components into systems with high dependability requirements. Such components are built to be reused as black boxes that cannot be modified. The system architect has to rely on techniques that are external to the component for resolving mismatches between the services required and provided that might arise in the interaction of the component and its environment. The paper puts forward an approach that employs the layer-based C2 architectural style for structuring error detection and recovery mechanisms to be added to the component during system integration.
Paulo Asterio de Castro Guerra, Cecília M. F. Rubira, Alexander B. Romanovsky, Rogério de Lemos
ESEC / SIGSOFT FSE3
2003 Coordinated Forward Error Recovery for Composite Web Services
abstract
This paper proposes a solution based on forward error recovery, oriented towards providing dependability of composite Web services. While exploiting their possible support for fault tolerance (e.g., transactional support at the level of each service), the proposed solution has no impact on the autonomy of the individual Web services, our solution lies in system structuring in terms of co-operative atomic actions that have a well-defined behavior, both in the absence and in the presence of service failures. More specifically, we define the notion of Web Service Composition Action (WSCA), based on the Coordinated Atomic Action concept, which allows structuring composite Web services in terms of dependable actions. Fault tolerance can then be obtained as an emergent property of the aggregation of several potentially non-dependable services. We further introduce a framework enabling the development of composite Web services based on WSCAs, consisting of an XML-based language for the specification of WSCAs.
Ferda Tartanoglu, Valérie Issarny, Alexander B. Romanovsky, Nicole Lévy
SRDS3
2002 A Structured Approach to Handling On-Line Interface Upgrades
abstract
The integration of complex systems out of existing systems is an active area of research and development. There are many practical situations in which the interfaces of the component systems, for example belonging to separate organisations, are changed dynamically and without notification. In this paper we propose an approach to handling such upgrades in a structured and disciplined fashion. All interface changes are viewed as abnormal events and general fault tolerance mechanisms (exception handling, in particular) are applied to dealing with them. The paper outlines general ways of detecting such interface upgrades and recovering after them. An Internet Travel Agency is used as a case study.
Cliff B. Jones, Alexander B. Romanovsky, Ian Welch
COMPSAC2
2002 Dependable On-Line Upgrading of Distributed Systems
abstract
The main theme of the workshop is to develop approaches to dependable systematic on-line system upgrading. The workshop aims at: bringing together practitioners, researchers and system developers working on the issues related to online upgrading of distributed systems; developing a better understanding of the problems that developers face while dealing with on-line system upgrading; defining a research agenda for developing distributed upgradable systems. We sought submissions from both industry and academia on all topics related to on-line upgrading of distributed systems. This "Dependable On-line Upgrading of Distributed Systems" workshop broadens the discussion from the 'simple' upgrading of objects in a closed environment to the asynchronous integration of upgraded components (objects, tasks, interfaces, etc.) in a distributed, independently managed, heterogeneous environment in which dependability is key. Inevitable in this environment is the possibility of upgrade errors and mismatches.
Alexander B. Romanovsky, Iain Smith
COMPSAC1
2002 ICSE 2002 workshop on architecting dependable systems
abstract
This workshop summary gives a brief overview of a one day workshop on "Architecting Dependable Systems" held in conjunctions with ICSE 2002.
Rogério de Lemos, Cristina Gacek, Alexander B. Romanovsky
ICSE3
2002 Rigorous Development of an Embedded Fault-Tolerant System Based on Coordinated Atomic Actions
abstract
Describes our experience using coordinated atomic (CA) actions as a system structuring tool to design and validate a sophisticated and embedded control system for a complex industrial application that has high reliability and safety requirements. Our study is based on an extended production cell model, the specification and simulator for which were defined and developed by FZI (Forschungszentrum Informatik, Germany). This "fault-tolerant production cell" represents a manufacturing process involving redundant mechanical devices (provided in order to enable continued production in the presence of machine faults). The challenge posed by the model specification is to design a control system that maintains specified safety and liveness properties even in the presence of a large number and variety of device and sensor failures. Based on an analysis of such failures, we provide details of: (1) a design for a control program that uses CA actions to deal with both safety-related and fault tolerance concerns and (2) the formal verification of this design based on the use of model checking. We found that CA action structuring facilitated both the design and verification tasks by enabling the various safety problems (involving possible clashes of moving machinery) to be treated independently. Even complex situations involving the concurrent occurrence of any pairs of the many possible mechanical and sensor failures can be handled simply yet appropriately. The formal verification activity was performed in parallel with the design activity, and the interaction between them resulted in a combined exercise in "design for validation"; formal verification was very valuable in identifying some very subtle residual bugs in early versions of our design which would have been difficult to detect otherwise.
Jie Xu 0007, Brian Randell, Alexander B. Romanovsky, Robert J. Stroud, Avelino Francisco Zorzo, Ercument Canver, Friedrich W. von Henke
IEEE Trans. Computers3
2001 Exception Handling in Component-Based System Development
abstract
Designers of component-based software face two problems related to dealing with abnormal events: developing exception handling at the level of the integrated system and accommodating (and adjusting, if necessary) exceptions and exception handling provided by individual components. Our intention is to develop an exception handling framework suitable for component-based system development by applying general exception handling mechanisms which have been proposed and successfully used in concurrent/distributed systems and in programming languages. The framework is applied in three steps. Firstly, individual components are wrapped in such a way that the wrappers perform activity related to local error detection and exception handling, and signal, if necessary, external exceptions outside the component. At the second step the execution of the overall system is structured as a set of dynamic actions in which components take parts. Such actions have important properties which facilitate exception handling: they are atomic, contain erroneous information and serve as recovery regions. The last step is designing exception handling at the action level: each action (i.e. all components participating in it) handles exceptions signalled by individual wrapped components.
Alexander B. Romanovsky
COMPSAC1
2001 On Applying Coordinated Atomic Actions and Dependable Software Architectures for Developing Complex Systems
abstract
Modern concurrent and distributed applications are becoming increasingly complex; so, in order to provide fault tolerance, special structuring mechanisms are required to help reduce this complexity. Unfortunately, such structuring techniques are mostly introduced as design and implementation features, which complicates their employment. The approach we propose relies on introducing the appropriate software structuring together with associated fault tolerance measures at the earlier phases of software development and on supporting it with special software architectures and design patterns.
Delano M. Beder, Cecília M. F. Rubira, Brian Randell, Alexander B. Romanovsky
ISORC4
2001 Looking Ahead in Atomic Actions with Exception Handling
abstract
An approach to introducing exception handling into object-oriented N is presented. A novel atomic action scheme is developed that does not impose any participant synchronisation on action exit. In order to use cooperative exception handling at the action level as the main fault tolerance mechanism, we develop a distributed protocol that finds, for any exception raised, an action containing all potentially erroneous information, aborts all of its nested actions, resolves multiple concurrent exceptions and involves all the action participants into cooperative handling of the resolved exception. In the scheme, no service messages are sent and no service synchronisation is introduced if there are no exceptions raised. This flexible scheme can be applied in a number of emerging areas in which entities of a different nature (including software tasks, people, plants, documents, organisations, etc.) participate in cooperative activities.
Alexander B. Romanovsky
SRDS1
2001 Conversations with fixed and potential participants
Alexander B. Romanovsky, Paul D. Ezhilchelvan
J. Syst. Archit.1
2001 A comparative study of exception handling mechanisms for building dependable object-oriented software
Alessandro F. Garcia 0001, Cecília M. F. Rubira, Alexander B. Romanovsky, Jie Xu 0007
J. Syst. Softw.3
2000 An Exception Handling Framework for N-Version Programming in Object-Oriented Systems
abstract
An approach to introducing exception handling into object oriented N-version programming (NVP) is proposed. General principles of structuring systems with diversity are outlined. The importance of using exceptions while applying diversely developed software is shown. Internal and external exceptions are clearly separated in our framework: each version has its own internal exceptions but the external exceptions of all versions have to be the same and identical to the interface exceptions of the diversely designed class. This scheme requires an adjudicator of a special kind to allow signalling interface exceptions when a majority of versions have signalled the same exception. These ideas are demonstrated using a general class diversity framework developed recently. An Ada implementation is outlined.
Alexander B. Romanovsky
ISORC1
2000 Extending conventional languages by distributed/concurrent exception resolution
Alexander B. Romanovsky
J. Syst. Archit.1
2000 Concurrent Exception Handling and Resolution in Distributed Object Systems
abstract
We address the problem of how to handle exceptions in distributed object systems. In a distributed computing environment, exceptions may be raised simultaneously in different processing nodes and thus need to be treated in a coordinated manner. Mishandling concurrent exceptions can lead to catastrophic consequences. We take two kinds of concurrency into account: 1) Several objects are designed collectively and invoked concurrently to achieve a global goal and 2) multiple objects (or object groups) that are designed independently compete for the same system resources. We propose a new distributed algorithm for resolving concurrent exceptions and show that the algorithm works correctly even in complex nested situations, and is an improvement over previous proposals in that it requires only O(n/sub max/N/sup 2/) messages, thereby permitting quicker response to exceptions.
Jie Xu 0007, Alexander B. Romanovsky, Brian Randell
IEEE Trans. Parallel Distributed Syst.2
2000 Guest Editors' Introduction-Current Trends in Exception Handling
abstract
THE importance of exception handling is well-recognized by system designers and software engineers. Exception handing is very often the most important part of the system because it deals with abnormal situations. The goal of
Dewayne E. Perry, Alexander B. Romanovsky, Anand R. Tripathi
IEEE Trans. Software Eng.2
2000 Guest Editors' Introduction - Current Trends in Exception Handling
abstract
1 THE ARTICLES THE second part of this special issue on Current Trends in Exception Handling includes four papers which primarily deal with exception handling in human-centered systems such as workflow, requirements specification, and new interactive programming models such as spreadsheets. These research contributions demonstrate that exceptions are not restricted to programming languages, but occur in many, if not most, real-world situations. These papers also reflect that exceptions can be deviations from normal conditions and may not necessarily imply errors. This is similar to Goodenough's observations in his classic paper in the 1970s [1]. These papers lead us to observe that anything that has an algorithmic flow, whether it be workflow or a design process or a program, has a pervasive exception handling need. Programming may be the ultimate in an algorithm, so many of the problems encountered there have analogies in other areas. Moreover, the computation model presented by programming languages tends to be relatively more simple in regard to handling of exceptions in contrast to dealing with such problems in large-scale systems such as enterprise-wide workflow. In those environments, it is not simply exception handling language constructs that are needed, but a methodology on how to use exception handling. The programming model of spreadsheet systems raises many unique issues related to exception handling. This is a widely used model of programming by end users through the use of many commercial products. In the paper aException Handling in the Spreadsheet Paradigm,o Margaret Burnett, Anurag Agrawal, and Pieter van Zee discuss these issues and present an approach for handling exceptions in this programming paradigm. Many spreadsheet programs can be quite large and complex and, therefore, both reliability as well maintainability of such programs becomes an issue when exception handling is introduced. The authors present their experience with and analysis of the error value models for spreadsheet programs. Achieving a high level of fault tolerance is one of the main concerns in developing modern workflow systems due to many factors: Distributed environment, long duration of activities, and complexity of the software involved are among them. The paper aException Handling in Workflow Management Systemso by Clause Hagen and Gustavo Alonso describes an advanced fault tolerance mechanism for incorporating both transactions and exception handling into such systems. The approach is unique for workflow systems as it treats workflow support as a programming environment and relies on general research on developing fault-tolerant software. The authors use fundamental research on linguistic issues of exception handling and propose simple ways of applying these concepts to transactional workflow management. The modeling language incorporates special features for error detection and handling which are conceptually similar to exception handling features found in programming languages. Another important way in which this approach is new is how it combines transaction atomicity and exception handling. A validation technique is developed to make it possible to assess the correctness of workflow specification in situations when exceptions are raised and handled. In the paper aHandling of Irregularities in Human Centered Systems: A Unified Framework for Data Processes,o Takahiro Murata and Alex Borgida address exception handling problems in human-centered systems. In such systems, exception handling is required for dealing with errors, as well as deviations, in data as well as processes, from their normal constraints. The paper focuses on exception handling in enterprise workflow systems. Generally, in enterprise systems, process models are used for describing the dynamic nature of activities of humans and semi-automated system entities. Most often, such models do not capture many unanticipated deviations. Sometimes such deviations have to be corrected and other times they are to be tolerated. This paper presents a unified model for handling errors and deviations, which are treated as exception conditions resulting from violations of some specified constraints. When permitting deviations to persist, it relies on runtime checks for assessing their consequences. Axel van Lamsweerde and Emmanuel Letier address the issues of aHandling Obstacles in Goal-Oriented Requirements Engineering.o The requirements elicitation process often results in goals, requirements, and assumptions about the desired system that are too idealized and that do not take into account the various kinds of problems that can occur. Not anticipating exceptional behaviors results in unrealistic, unachievable, or incomplete requirements specifications. This, in turn, leads to systems that are not robust enough and which may fail at critical times, perhaps with critical consequences. The authors present formal techniques for IEEE TRANSACTIONS ON SOFTWARE ENGINEERING, VOL. 26, NO. 10, OCTOBER 2000 921
Dewayne E. Perry, Alexander B. Romanovsky, Anand R. Tripathi
IEEE Trans. Software Eng.2
1999 Formal Development and Validation of Java Dependable Distributed Systems
abstract
The rapid expansion of Java programs into the software market is often not supported by a proper development methodology. We present a formal development methodology, well suited for Java dependable distributed applications. It is based on the stepwise refinement of model oriented formal specifications, and enables validation of the obtained system wrt the client's requirements. Three refinement steps have been identified in the case of fault tolerant distributed applications: first, starting from informal requirements, an initial formal specification is derived. It does not depend on implementation constraints and provides a centralized solution; second, dependability and distribution constraints are integrated; third, the Java implementation is realised. The CO-OPN/2 language is used to express specifications formally; and the dependability and distribution design as based on the Coordinated Atomic action concept. The methodology and the three refinement steps are presented through a very simple fault tolerant distributed Java application.
Giovanna Di Marzo Serugendo, Nicolas Guelfi, Alexander B. Romanovsky, Avelino Francisco Zorzo
ICECCS3
1999 Engineering Look-ahead in Distributed Conversations
abstract
This paper investigates the effects of relaxing the synchronisation embedded in "classical" conversation schemes. Look-ahead conversation scheme enables the synchronisation mandated by conversations to be performed concurrently to other normal system activities, and thereby provides scope for enhancing system performance. In this paper, we take the view that permitting look-ahead must guarantee that executions with and without look-ahead be equivalent for identical inputs and run-time conditions. We identify and formulate the necessary condition for meeting this objective. We then present two schemes for realising this condition. The first scheme is based on piggybacking extra information onto ongoing messages, and the second one is a message passing protocol requiring each participant to send one message to every other participant of the conversation. These schemes do not require a conversation participant to know a priori all participants of the conversation, bur only those it is designed to interact with, during the conversation.
Paul D. Ezhilchelvan, Alexander B. Romanovsky
ISADS2
1999 Exception Handling in a Cooperative Object-Oriented Approach
abstract
A cooperative action (CO action) is a modelling abstraction for representing collaborative behaviour between objects at different phases of the software development. In this paper, the original definition of a cooperative object-oriented approach for software development is extended in order to include the description of exceptional behaviour. Unlike the traditional methods that usually deal with exceptions at the late design and implementation phases, the proposed approach emphasises the separation of treatments of application-related, design-related and implementation-related exceptions during the software life-cycle. The feasibility of the approach is demonstrated in terms of a benchmark case study.
Rogério de Lemos, Alexander B. Romanovsky
ISORC2
1999 Choosing Effective Methods for Design Diversity - How to Progress from Intuition to Science
Peter T. Popov, Lorenzo Strigini, Alexander B. Romanovsky
SAFECOMP3
1999 On Structuring Cooperative and Competitive Concurrent Systems
abstract
Developing advanced structuring techniques has always been of great importance for computer science and practice. Many structuring approaches are used to help capture certain characteristics of applications: group communications, replication features, file services, etc. Associating fault tolerance measures with structuring units allows us to benefit from state and behaviour encapsulation while designing dependable systems. Procedures were among the first general techniques intended for structuring application software. They reflect both the static and dynamic structures of sequential systems (a stack of procedure contexts of nested calls represents the state of program execution). The situation is much more complex in concurrent systems, in which the states of several concurrent components should be taken into consideration while describing system behaviour in general, and its fault tolerance in particular. The purpose of this survey is to outline recent trends in developing approaches to structuring competitive and cooperative concurrent systems, together with corresponding fault tolerance techniques, and to discuss different directions of research in this area and their interrelations.
Alexander B. Romanovsky
Comput. J.1
1999 Coordinated atomic actions as a technique for implementing distributed gamma computation
Alexander B. Romanovsky, Avelino Francisco Zorzo
J. Syst. Archit.1
1999 Class diversity support in object-oriented languages
Alexander B. Romanovsky
J. Syst. Softw.1
1999 Using Coordinated Atomic Actions to Design Safety-Critical Systems: a Production Cell Case Study
abstract
Coordinated Atomic actions (CA actions) are a unified approach to structuring complex concurrent activities and supporting error recovery between multiple interacting objects in object-oriented systems. This paper explains how we have used the CA action concept to design and implement a safety-critical application. We have used the Production Cell model that was developed in the Forschungszentrum Informatik (FZI), Karlsruhe, Germany, to present a realistic industry-oriented problem, where safety requirements play a significant role. Our design consists of two levels: the first level deals with the scheduling of CA actions, and the second level deals with the interactions between devices. Both the scheduling mechanism and the device interactions are enclosed by CA actions. Exception handling and error recovery are incorporated into CA actions in order to satisfy high safety and fault tolerance requirements. A controlling program based on our design was developed in the Java language and used to drive a graphical simulator provided by the FZI. Copyright © 1999 John Wiley & Sons, Ltd.
Avelino Francisco Zorzo, Alexander B. Romanovsky, Jie Xu 0007, Brian Randell, Robert J. Stroud, Ian Welch
Softw. Pract. Exp.2
1998 Coordinated Exception Handling in Distributed Object Systems: From Model to System Implementation
abstract
Exception handling in concurrent and distributed programs is a difficult task though it is often necessary. In many cases traditional exception mechanisms for sequential programs are no longer appropriate. One major difficulty is that the process of handling an exception may need to involve multiple concurrent components that are cooperating in pursuit of some global goal. Another complication is that several exceptions may be raised concurrently in different nodes of a distributed environment. Existing proposals and actual concurrent languages either ignore these difficulties or only cope with a limited form of them. The paper attempts a general solution, developed especially for distributed object systems, starting from a conceptual model, together with algorithms for coordinating concurrent components and resolving multiple exceptions, through to an actual system implementation. An industrial production cell is chosen as a case study to demonstrate the usefulness of the proposed model and algorithms. A system that supports coordinated atomic actions and exception resolution is implemented in distributed Ada 95 and examined through several performance-related experiments.
Jie Xu 0007, Alexander B. Romanovsky, Brian Randell
ICDCS2
1998 Coordinated Atomic Actions in Modelling Objects Cooperation
abstract
The approach described in this paper makes use of Coordinated Atomic Actions (CA actions)-a structural design mechanism, for representing the cooperation between objects, at different stages of the software development. The original concept of a CA action has been expanded for accommodating the modelling needs of the initial stages of software development. One of the motivations for using CA actions in an OO approach, is the ability of CA actions to extract from the specification of an object those issues which are related with its the collaborative activities, thus avoiding that a specification of a cooperation be scattered among the specifications of the objects. The feasibility of the approach is demonstrated in terms of a benchmark case study.
Rogério de Lemos, Alexander B. Romanovsky
ISORC2
1998 Exception Handling in Object-Oriented Real-Time Distributed Systems
abstract
Exception handling in a complex concurrent and distributed system (e.g. one involving cooperating rather than just competing activities) is often a necessary, but a very difficult, task. No widely accepted models or approaches exist in this area. The object-oriented paradigm, for all its structuring benefits, and real-time requirements each add further difficulties to the design and implementation of exception handling in such systems. In this paper, we develop a general structuring framework based on the coordinated atomic (CA) action concept for handling exceptions in an object-oriented distributed system, in which exceptions in both the value and the time domain are taken into account. In particular, we attempt to attack several difficult problems related to real-time system design and error recovery, including action-level timing constraints, time-triggered CA actions, and time-dependent exception handling. The proposed framework is then demonstrated and assessed using an industrial real-time application-the Production Cell III case study.
Alexander B. Romanovsky, Jie Xu 0007, Brian Randell
ISORC1
1998 Distributed Atomic Actions in Ada 95
abstract
This paper discusses the development of a distributed asynchronous atomic action scheme for Ada 95. The scheme makes use of many unique Ada 95 features including protected objects, asynchronous transfer of control and the distributed systems annex. We present the packages which implement the local and global action support and illustrate their use in a (partial) implementation of the FZI production cell problem. We also discuss a number of variations of the model and how these might be included. Finally, we discuss how the distribution model used in Ada 95 has influenced our design.
Stuart E. Mitchell, Andy J. Wellings, Alexander B. Romanovsky
Comput. J.3
1998 A study of atomic action schemes intended for standard Ada
Alexander B. Romanovsky
J. Syst. Softw.1
1997 Practical Exception Handling and Resolution in Concurrent Programs
Alexander B. Romanovsky
Comput. Lang.1
1997 Implementation of blocking coordinated atomic actions based on forward error recovery
Alexander B. Romanovsky, Brian Randell, Robert J. Stroud, Jie Xu 0007, Avelino Francisco Zorzo
J. Syst. Archit.1
1996 Exception Handling and Resolution in Distributed Object-oriented Systems
abstract
We address the problem of how to handle exceptions in distributed object-oriented systems. In a distributed computing environment exceptions may be raised simultaneously and thus need to be treated in a coordinated manner. We take two kinds of concurrency into account: 1) several objects are designed collectively and invoked concurrently to achieve a global goal, and 2) concurrent objects or object groups that are designed independently compete for the same system resources. We propose a new distributed algorithm for resolving concurrent exceptions and show that the algorithm works correctly even in complex nested situations, and is an improvement over previous proposals in that it requires only O(N/sup 2/) messages, and is fully object-oriented.
Jie Xu 0007, Alexander B. Romanovsky, Brian Randell
ICDCS2
1996 Application specific conversation schemes for ADA programs
Alexander B. Romanovsky
Microprocess. Microprogramming1
1995 Conversations of Objects
Alexander B. Romanovsky
Comput. Lang.1
1994 The problems of designing a conversation scheme for concurrent object oriented languages
Alexander B. Romanovsky
Microprocess. Microprogramming1