VLDB 2026 Research / reviewers in the wild / expert
Jacek Sroka
dblp:42/6606
· DBLP profile ↗
23ranked-venue papers
11as first author
3since 2021 · last 2025
0000-0002-1714-9667ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Theory of computation · 9 · 6 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 3Databases, data management, data science and information retrieval · 3 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FuzzyPPI: Large-Scale Interaction of Human Proteome at Fuzzy Semantic SpaceabstractLarge-scale protein-protein interaction (PPI) network of an organism provides key insights into its cellular and molecular functionalities, signaling pathways and underlying disease mechanisms. For any organism, the total unexplored protein interactions significantly outnumbers all known positive and negative interactions. For Human, all known PPI datasets contain only ∼ 5.61 million positive and ∼ 0.76 million negative interactions, which is ∼ 3.1% of potential interactions. We have implemented a distributed algorithm in Apache Spark that evaluates a Human PPI network of ∼ 180 million potential interactions resulting from 18 994 reviewed proteins for which Gene Ontology (GO) annotations are available. The computed scores have been validated against state-of-the-art methods on benchmark datasets.FuzzyPPI performed significantly better with an average F1 score of 0.62 compared to GOntoSim (0.39), GOGO (0.38), and Wang (0.38) when tested with the Gold Standard PPI Dataset. The resulting scores are published with a web server for non-commercial use athttp://fuzzyppi.mimuw.edu.pl/. Moreover, conventional PPI prediction methods produce binary results, but in fact this is just a simplification as PPIs have strengths or probabilities and recent studies show that protein binding affinities may prove to be effective in detecting protein complexes, disease association analysis, signaling network reconstruction, etc. Keeping these in mind, our algorithm is based on a fuzzy semantic scoring function and produces probabilities of interaction. Anup Kumar Halder, Soumyendu Sekhar Bandyopadhyay, Witold Jedrzejewski, Subhadip Basu, Jacek Sroka |
IEEE Trans. Big Data | 5 |
| 2023 | Aggregating over Dominated Points by Sorting, Scanning, Zip and Flat Maps
Jacek Sroka, Jerzy Tyszkiewicz |
ESA | 1 |
| 2021 | PartSeg: a tool for quantitative feature extraction from 3D microscopy images for dummiesabstractBACKGROUND: Bioimaging techniques offer a robust tool for studying molecular pathways and morphological phenotypes of cell populations subjected to various conditions. As modern high-resolution 3D microscopy provides access to an ever-increasing amount of high-quality images, there arises a need for their analysis in an automated, unbiased, and simple way. Segmentation of structures within the cell nucleus, which is the focus of this paper, presents a new layer of complexity in the form of dense packing and significant signal overlap. At the same time, the available segmentation tools provide a steep learning curve for new users with a limited technical background. This is especially apparent in the bulk processing of image sets, which requires the use of some form of programming notation. RESULTS: In this paper, we present PartSeg, a tool for segmentation and reconstruction of 3D microscopy images, optimised for the study of the cell nucleus. PartSeg integrates refined versions of several state-of-the-art algorithms, including a new multi-scale approach for segmentation and quantitative analysis of 3D microscopy images. The features and user-friendly interface of PartSeg were carefully planned with biologists in mind, based on analysis of multiple use cases and difficulties encountered with other tools, to offer an ergonomic interface with a minimal entry barrier. Bulk processing in an ad-hoc manner is possible without the need for programmer support. As the size of datasets of interest grows, such bulk processing solutions become essential for proper statistical analysis of results. Advanced users can use PartSeg components as a library within Python data processing and visualisation pipelines, for example within Jupyter notebooks. The tool is extensible so that new functionality and algorithms can be added by the use of plugins. For biologists, the utility of PartSeg is presented in several scenarios, showing the quantitative analysis of nuclear structures. CONCLUSIONS: In this paper, we have presented PartSeg which is a tool for precise and verifiable segmentation and reconstruction of 3D microscopy images. PartSeg is optimised for cell nucleus analysis and offers multi-scale segmentation algorithms best-suited for this task. PartSeg can also be used for the bulk processing of multiple images and its components can be reused in other systems or computational experiments. Grzegorz Bokota, Jacek Sroka, Subhadip Basu, Nirmal Das, Pawel Trzaskoma, Yana Yushkevich, Agnieszka Grabowska, Adriana Magalska, Dariusz Plewczynski |
BMC Bioinform. | 2 |
| 2018 | Verification of Dynamic Behaviour in Qualitative Molecular Networks Describing Gene Regulation, Signalling and Whole-cell MetabolismabstractWe present a tool for the verification of qualitative biological models. These models formalise observed behaviours and interrelations of molecular and cellular mechanisms. During its development a model is continuously verified. Predicted behaviours are compared with behaviours observed in experim ental data. Moreover, the model must not exhibit behaviours which contradict existing knowledge about capabilities of the biological system under investigation. Model development is an iterative process involving many rounds of prediction, verification and refinement. Due to the complexity of biological systems this process is laborious and error prone, which motivates the development of “model debugging” tools. The qualitative models we investigate represent large-scale molecular interaction networks describing gene regulation, signalling and whole-cell metabolism. We integrate a steady state model of whole-cell metabolism with a dynamic model of gene regulation and signalling represented as a Petri net. This Quasi-Steady State Petri Net (QSSPN) representation allows the generation of dynamic sequences of molecular events satisfying substrate, activator, inhibitor and metabolic flux requirements at every state transition. The reachability graph of the dynamic part of the model is examined and for every transition in this graph the satisfaction of metabolic flux requirements is verified by well-established linear programming techniques. Our approach is based on network connectivity alone and does not require any kinetic parameters. We demonstrate the applicability of our method by analysing a large-scale model of a nuclear receptor network regulating bile acid homeostasis in human hepatocyte. To date, simulation and verification of QSSPN models have been performed exclusively by Monte Carlo simulation. Random walks through the state space were used to find examples of behaviour satisfying properties of interest. Here, we provide for the first time for QSSPN models an exhaustive analysis of the state space up to a finite depth, which is possible due to several effective optimisations. Contrary to the Monte Carlo approach, we can prove that certain behaviour cannot be realised by the model within a given number of steps. This allows rejection of models which are not capable to reproduce experimentally observed behaviours, as well as verification that biologically unrealistic behaviours cannot occur in the simulation. We show an example of how these features improve identification of problems in large scale network models. Marek Grabowski, Grzegorz Bokota, Jacek Sroka, Andrzej M. Kierzek |
Fundam. Informaticae | 3 |
| 2018 | Simulation of multicellular populations with Petri nets and genome scale intracellular networks
Kamil Kedzia, Wojtek Ptak, Jacek Sroka, Andrzej M. Kierzek |
Sci. Comput. Program. | 3 |
| 2017 | On Determining the AND-OR Hierarchy in Workflow NetsabstractThis paper presents a notion of reduction where a WF net is transformed into a smaller net by iteratively contracting certain well-formed subnets into single nodes until no more of such contractions are possible. This reduction can reveal the hierarchical structure of a WF net, and since it preserves certain semantic properties such as soundness, can help with analysing and understanding why a WF net is sound or not. The reduction can also be used to verify if a WF net is an AND-OR net. This class of WF nets was introduced in earlier work, and arguably describes nets that follow good hierarchical design principles. It is shown that the reduction is confluent up to isomorphism, which means that despite the inherent non-determinism that comes from the choice of subnets that are contracted, the final result of the reduction is always the same up to the choice of the identity of the nodes. Based on this result, a polynomial-time algorithm is presented that computes this unique result of the reduction. Finally, it is shown how this algorithm can be used to verify if a WF net is an AND-OR net. Jacek Sroka, Jan Hidders |
Fundam. Informaticae | 1 |
| 2016 | AB-QSSPN: Integration of Agent-Based Simulation of Cellular Populations with Quasi-Steady State Simulation of Genome Scale Intracellular NetworksabstractWe present a tool for simulation of populations of living cells interacting in spatial structures. Each cell is modelled with the Quasi-Steady Petri Net that integrates dynamic regulatory network expressed with a Petri Net (PN) and Genome Scale Metabolic Networks (GSMN) where linear programming is used to explore the steady-state metabolic flux distributions in the whole-cell model. Similar simulations have already been conducted for single cells, but we present an architecture to simulate populations of millions of interacting cells organized in spatial structures which can be used to model tumour growth or formation of tuberculosis lesions. For that we use the Spark framework and organize the computation in an agent based “think like a vertex” fashion as in Pregel like systems. In the cluster we introduce a special kind of per node caching to speed up computation of the steady-state metabolic flux. Our tool can be used to provide a mechanistic link between genotype and behaviour of multicellular system. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves. Wojtek Ptak, Andrzej M. Kierzek, Jacek Sroka |
Petri Nets | 3 |
| 2015 | Recent advances in Scalable Workflow Enactment Engines and Technologies
Jan Hidders, Paolo Missier, Jacek Sroka |
Future Gener. Comput. Syst. | 3 |
| 2015 | On Generating Hierarchical Workflow Nets and their Extensions and Verifying HierarchicalityabstractFor designing and analyzing complex workflow nets the notion of hierarchical decomposition can be essential for keeping the structure of the workflow comprehensible. In this paper we study two classes of nets: hierarchical nets and extended hierarchical nets. The first have a simple hierarchical st ructure and can be defined in terms of five simple refinement rules. We show that for arbitrary nets it can be easily verified if they can be constructed this way, thus confirming their good design and the properties following from it. As we prove, this can be done by performing the refinements in reverse, i.e., by contracting subnets into single nodes. It is shown that the choice of the contracted subnet does not change the final result of the process, and therefore this procedure for checking the hierarchical structure requires no back-tracking. The second class, extended hierarchical nets, is an extension of the first class where two types of extra refinements are introduced that allow to indicate (1) the synchronization between two parallel running subworkflows or (2) the transfer of a thread from one subworkflow to another one. These refinements come with natural and necessary preconditions that ensure that result is still a sound workflow net. In case (1) where we want to synchronize two actions in two subworkflows, we should convince ourselves that the subworkflows represent parallel threads which always execute together, otherwise a deadlock could easily arise. Dually, in case (2), if after the moment that a choice was made between two subworkflows we at a later point in the workflow want to allow a transfer between them, this can be done safely provided that we did not enter any thread fork in the meantime. We show that the class of extended hierarchical nets, which is defined by adding these two additional types of refinement, is a proper superset of the hierarchical nets, but still all such nets exhibit the correctness property of *-soundness. We do this by showing that the class is a proper subset of the AND-OR nets which were in earlier work shown to have this property. Jacek Sroka, Piotr Chrzastowski-Wachtel, Jan Hidders |
Fundam. Informaticae | 1 |
| 2015 | Translating Relational Queries into SpreadsheetsabstractSpreadsheets are among the most commonly used applications for data management and analysis. They combine data processing with very diverse supplementary features: statistics, visualization, reporting, linear programming solvers, Web queries periodically downloading data from external sources, etc. However, the spreadsheet paradigm of computation still lacks sufficient analysis. In this article, we demonstrate that a spreadsheet can implement all data transformations definable in SQL, merely by utilizing spreadsheet formulas. We provide a query compiler, which translates any given SQL query into a worksheet of the same semantics, including NULL values. Thereby, database operations become available to the users who do not want to migrate to a database. They can define their queries using a high-level language and then get their execution plans in a plain vanilla spreadsheet. The functions available in spreadsheets impose limitations on the algorithms one can implement. In this paper, we offer O(n log2n) sorting spreadsheet, using a non-constant number of rows, and, surprisingly, Depth-First-Search and Breadth-First-Search on graphs. Jacek Sroka, Adrian Panasiuk, Krzysztof Stencel, Jerzy Tyszkiewicz |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2014 | On generating ⁎-sound nets with substitution
Jacek Sroka, Jan Hidders |
Inf. Syst. | 1 |
| 2013 | PrefaceabstractThis Special Issue originates from the First International Workshop on Scalable Workflow Enactment Engines and Technologies (SWEET), held in conjunction with the 2012 SIGMOD conference in Scottsdale, Arizona, USA on May 20th, 2012.The goal of the workshop was to bring together researchers and practitioners to explore the state of the art in workflow-based programming for scientific data-intensive applications, and the potential of cloud-based computing in this area.Amongst the main motivations for the adoption of workflow technology is the potential ability for computational scientists, who are the immediate beneficiaries of e-science infrastructure, to assemble data-centric science pipelines without the need for deep technical knowledge of the underlying data management infrastructure.In order to fulfill this potential, workflow middleware for science must offer ease of programming whilst providing portability across different computational environments, as well as transparent access to large pools of data storage and distributed computational resources.These user requirements translate into the need for a robust underlying data management infrastructure.The SWEET workshop was an initial exploration into the state of the art in such data-centric workflow technology, and its ability to exploit in particular the potential of elastic cloud infrastructure to achieve scalability over large size data problems.This special issue extends this initial exploration into this space.While it includes the extended version of just one of the SWEET papers, it also collects three additional high quality contributions.Collectively, these paint an exciting landscape of emerging cloud-aware workflow technology for science.The SWEET paper in this collection, entitled Turbine: A distributed-memory dataflow engine for high performance many-task applications, comes from the Argonne National Laboratory.Justin Wozniak, Timothy Armstrong and colleagues describe the Turbine workflow system, which is aimed at specifying programs on large-scale, high performance computing (HPC) systems using the Swift language.Turbine allows for distributed-memory evaluation of dataflow programs such that the overhead of program evaluation and task generation is spread throughout an extreme-scale computing system.This involves for example the introduction of futures, i.e., objects that act as proxies for results that are not yet available.Notable features are the detection of parallel loops and concurrent function invocations, which Special issue editors Jan Hidders, Paolo Missier, Jacek Sroka, Jan Van den Bussche |
Fundam. Informaticae | 3 |
| 2011 | Constrained Coalition FormationabstractThe conventional model of coalition formation considers every possible subset of agents as a potential coalition. However, in many real-world applications, there are inherent constraints on feasible coalitions: for instance, certain agents may be prohibited from being in the same coalition, or the coalition structure may be required to consist of coalitions of the same size. In this paper, we present the first systematic study of constrained coalition formation (CCF). We propose a general framework for this problem, and identify an important class of CCF settings, where the constraints specify which groups of agents should/should not work together. We describe a procedure that transforms such constraints into a structured input that allows coalition formation algorithms to identify, without any redundant computations, all the feasible coalitions. We then use this procedure to develop an algorithm for generating an optimal (welfare-maximizing) constrained coalition structure, and show that it outperforms existing state-of-the-art approaches by several orders of magnitude. Talal Rahwan, Tomasz P. Michalak, Edith Elkind, Piotr Faliszewski, Jacek Sroka, Michael J. Wooldridge, Nicholas R. Jennings |
AAAI | 5 |
| 2011 | CalcTav - integration of a spreadsheet and Taverna workbenchabstractMOTIVATION: Taverna workbench is an environment for construction, visualization and execution of bioinformatic workflows that integrates specialized tools available on the Internet. It already supports major bioinformatics services and is constantly gaining popularity. However, its user interface requires considerable effort to learn, and sometimes requires programming or scripting experience from its users. We have integrated Taverna with OpenOffice Calc, making the functions of the scientific workflow system available in the spreadsheet. In CalcTav, one can define workflows using the spreadsheet interface and analyze the results using the spreadsheet toolset. RESULTS: Technically, CalcTav is a plugin for OpenOffice Calc, which provides the functionality of Taverna available in the form of spreadsheet functions. Even basic familiarity with spreadsheets already suffices to define and use spreadsheet workflows with Taverna services. The data processed by the Taverna components is automatically transferred to and from spreadsheet cells, so all the visualization and data analysis tools of OpenOffice Calc are available to the workflow creator within one, consistent user interface. AVAILABILITY: CalcTav is available under GPLv2 from http://code.google.com/p/calctav/ CONTACT: [email protected]. Jacek Sroka, Lukasz Krupa, Andrzej M. Kierzek, Jerzy Tyszkiewicz |
Bioinform. | 1 |
| 2011 | Acorn: A grid computing system for constraint based modeling and visualization of the genome scale metabolic reaction networks via a web interfaceabstractBACKGROUND: Constraint-based approaches facilitate the prediction of cellular metabolic capabilities, based, in turn on predictions of the repertoire of enzymes encoded in the genome. Recently, genome annotations have been used to reconstruct genome scale metabolic reaction networks for numerous species, including Homo sapiens, which allow simulations that provide valuable insights into topics, including predictions of gene essentiality of pathogens, interpretation of genetic polymorphism in metabolic disease syndromes and suggestions for novel approaches to microbial metabolic engineering. These constraint-based simulations are being integrated with the functional genomics portals, an activity that requires efficient implementation of the constraint-based simulations in the web-based environment. RESULTS: Here, we present Acorn, an open source (GNU GPL) grid computing system for constraint-based simulations of genome scale metabolic reaction networks within an interactive web environment. The grid-based architecture allows efficient execution of computationally intensive, iterative protocols such as Flux Variability Analysis, which can be readily scaled up as the numbers of models (and users) increase. The web interface uses AJAX, which facilitates efficient model browsing and other search functions, and intuitive implementation of appropriate simulation conditions. Research groups can install Acorn locally and create user accounts. Users can also import models in the familiar SBML format and link reaction formulas to major functional genomics portals of choice. Selected models and simulation results can be shared between different users and made publically available. Users can construct pathway map layouts and import them into the server using a desktop editor integrated within the system. Pathway maps are then used to visualise numerical results within the web environment. To illustrate these features we have deployed Acorn and created a web server allowing constraint based simulations of the genome scale metabolic reaction networks of E. coli, S. cerevisiae and M. tuberculosis. CONCLUSIONS: Acorn is a free software package, which can be installed by research groups to create a web based environment for computer simulations of genome scale metabolic reaction networks. It facilitates shared access to models and creation of publicly available constraint based modelling resources. Jacek Sroka, Lukasz Bieniasz-Krzywiec, Szymon Gwozdz, Dariusz Leniowski, Jakub Lacki, Mateusz Markowski, Claudio Avignone-Rossa, Michael E. Bushell, Johnjoe McFadden, Andrzej M. Kierzek |
BMC Bioinform. | 1 |
| 2010 | A Network Flow Approach to Coalitional GamesabstractIn this paper we propose a novel approach to represent coalitional games, called a Coalition-Flow Network (CF-NET), that builds upon a generalization of the network flow literature. Specifically, this representation is based on our observation that the coalition formation process can be viewed as the problem of directing the flow through a network where every edge has certain capacity constraints. Talal Rahwan, Tomasz P. Michalak, Madalina Croitoru, Jacek Sroka, Nicholas R. Jennings |
ECAI | 4 |
| 2010 | JavaSpaces NetBeans: a linda workbench for distributed programming courseabstractIn this paper we introduce the JavaSpaces NetBeans IDE (JSN) which integrates the JavaSpaces technology, an implementation of Linda principles in Java, with the NetBeans IDE. JSN is a didactic tool for practical assignments during distributed programming courses. It hides advanced aspects of JavaSpaces configuration and lets students focus on interprocess coordination. An important component of JSN is a distributed debugger which can help to make concurrent programming classes easier to understand and more compelling. Magdalena Dukielska, Jacek Sroka |
ITiCSE | 2 |
| 2010 | A formal semantics for the Taverna 2 workflow model
Jacek Sroka, Jan Hidders, Paolo Missier, Carole A. Goble |
J. Comput. Syst. Sci. | 1 |
| 2009 | On representing coalitional games with externalitiesabstractWe consider the issue of representing coalitional games in multi-agent systems with externalities (i.e., in systems where the performance of one coalition may be affected by other co-existing coalitions). In addition to the conventional partition function game representation (PFG), we propose a number of new representations based on a new notion of externalities. In contrast to conventional game theory, our new concept is not related to the process by which the coalitions are formed, but rather to the effect that each coalition may have on the entire system and vice versa. We show that the new representations are fully expressive and, for many classes of games, more concise than the conventional PFG. Building upon these new representations, we propose a number of approaches to solve the coalition structure generation problem in systems with externalities. We show that, if externalities are characterised by various degrees of regularity, the new representations allow us to adapt coalition structure generation algorithms that were originally designed for domains with no externalities, so that they can be used when externalities are present. Finally, building upon Rahwan et al. [16] and Michalak et al. [9], we present a unified method to solve the coalition structure generation problem in any system, with or without externalities, provided sufficient information is available. Tomasz P. Michalak, Talal Rahwan, Jacek Sroka, Andrew James Dowell, Michael J. Wooldridge, Peter McBurney, Nicholas R. Jennings |
EC | 3 |
| 2009 | Towards a Formal Semantics for the Process Model of the Taverna Workbench. Part IabstractWorkflow development and enactment workbenches are becoming a standard tool for conducting in silico experiments. Their main advantages are easy to operate user interfaces, specialized and expressive graphical workflow specification languages and integration with a huge number of bioinformatic services. A popular example of such a workbench is Taverna, which has many additional useful features like service discovery, storing intermediate results and tracking data provenance. We discuss a detailed formal semantics for Scufl - the workflow definition language of the Taverna workbench. It has several interesting features that are notmet in other models including dynamic and transparent type coercion and implicit iteration, control edges, failure mechanisms, and incominglinks strategies. We study these features and investigate their usefulness separately as well as in combination, and discuss alternatives. The formal definition of such a detailed semantics not only allows to exactly understand what is being done in a given experiment, but is also the first step toward automatic correctness verification and allows the creation of auxiliary tools that would detect potential errors and suggest possible solutions to workflow creators, the same way as Integrated Development Environments aid modern programmers. A formal semantics is also essential for work on enactment optimization and in designing the means to effectively query workflow repositories. This paper is the first of two. It defines, explains and discusses fundamental notions for describing Scufl graphs and their semantics. Then, in the second part, we use these notions to define the semantics and show that our definition can be used to prove properties of Scufl graphs. Jacek Sroka, Jan Hidders |
Fundam. Informaticae | 1 |
| 2009 | Towards a Formal Semantics for the Process Model of the Taverna Workbench. Part II
Jacek Sroka, Jan Hidders |
Fundam. Informaticae | 1 |
| 2008 | DFL: A dataflow language based on Petri nets and nested relational calculus
Jan Hidders, Natalia Kwasnikowska, Jacek Sroka, Jerzy Tyszkiewicz, Jan Van den Bussche |
Inf. Syst. | 3 |
| 2006 | XQTav: an XQuery processor for Taverna environmentabstractUNLABELLED: Taverna workbench is an environment for construction, visualization and execution of bioinformatic workflows that integrate specialized tools available through the internet. It is gaining popularity fast, because of supporting the most important bioinformatic services and its simple, yet robust graphical notation. Here we present XQTav-an extension of Taverna that provides full integration with XQuery (the query language for XML) engine. XQTav allows execution of XQuery scripts in Taverna workflow diagrams. All existing Taverna processors can be accessed in the XQuery scripts. This provides an alternative way of specifying subworkflows in Taverna and is useful when one deals with query-like algorithms (e.g. filters and inner joins). Moreover, XQtav may be used to automatically generate an XQuery script that is equivalent to Taverna's workflow. This constitutes another way of creating and enacting bioinformatic workflows: overall structure of a diagram is drawn in Taverna environment, XQuery code is generated and possibly adjusted by hand. It can be executed by XQuery engines or incorporated into other software environments. AVAILABILITY: XQtav is an open source software. It may be downloaded from http://xqtav.sourceforge.net/. The page also contains various tutorials and examples, including the one described in this report. Jacek Sroka, Grzegorz Kaczor, Jerzy Tyszkiewicz, Andrzej M. Kierzek |
Bioinform. | 1 |