VLDB 2026 Research / reviewers in the wild / expert
Elzbieta Pustulka
dblp:318/5521 · also Ela Hunt, Ela Pustulka, Ela Pustulka-Hunt, Elzbieta Katarzyna Pustulka-Hunt
· DBLP profile ↗
15ranked-venue papers
2as first author
2since 2021 · last 2024
0000-0001-7379-847XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 5 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 4Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1Systems, architecture and hardware · 1Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | On the importance of CI/CD practices for database applicationsabstractSummary Continuous integration and continuous delivery (CI/CD) automate software integration and reduce repetitive engineering work. While the use of CI/CD presents efficiency gains, in database application development, this potential has not been fully exploited. We explore the state of the art in this area, with a focus on current practices, common software tools, challenges, and preconditions that apply to database applications. The work is grounded in a synoptic literature review and contributes a novel generic CI/CD pipeline for database system application development. Our generic pipeline was tailored to three industrial development use cases in which we measured the benefits of integration and deployment automation. The measurements demonstrate clearly that introducing CI/CD had significant benefits. It reduced the number of failed deployments, improved their stability, and increased the number of deployments. Interviews with the developers before and after the implementation of the CI/CD show that the pipeline brings clear benefits to the development team (i.e., a reduced cognitive load). These findings put current database release practices driven by business expectations, such as fixed release windows, in question. Jasmin Fluri, Fabrizio Fornari 0001, Elzbieta Pustulka |
J. Softw. Evol. Process. | 3 |
| 2023 | Measuring the Benefits of CI/CD Practices for Database Application DevelopmentabstractModern software development practices automate software integration and reduce repetitive software engineering work. Automation reduces the time it takes from defining software requirements to deploying the software in production. However, when it comes to database applications, the database integration and deployment are often executed manually, making it costly and error-prone. To mitigate this, we extended current software development methodologies by designing a CI/CD pipeline that takes into consideration the database setting. We report on two industrial case studies in which we implemented a newly designed pipeline and we measure the benefits of integration and deployment automation in database development projects. From a quantitative perspective, we found that introducing CI/CD pipelines reduces failed deployments, improves stability and increases the number of executed deployments. From a qualitative perspective, we interviewed the developers before and after the implementation of the CI/CD pipeline and the results show the CI/CD pipeline brings clear benefits to the development team (i.e., reduced cognitive load). This finding puts current database release practices driven by business expectations such as fixed release windows in question. Jasmin Fluri, Fabrizio Fornari 0001, Elzbieta Pustulka |
ICSSP | 3 |
| 2010 | VisGenome with CartoonPlus: Supporting large scale genomic analyses via physical space deformation
Joanna Jakubowska, Elzbieta Pustulka, Matthew Chalmers |
Future Gener. Comput. Syst. | 2 |
| 2009 | Mobile P2P Fast Similarity SearchabstractIn informal data sharing environments, misspellings cause problems for data indexing and retrieval. This is even more pronounced in mobile environments, in which devices with limited input devices are used. In a mobile environment, similarity search algorithms for finding misspelled data need to account for limited CPU and bandwidth. This demo shows P2P fast similarity search (P2PFastSS) running on mobile phones and laptops that is tailored to uncertain data entry and uses available resources efficiently. In this demo, users publish and search for textual content containing misspellings without relying on query logging, as done by Google, and with a minimum distributed indexing infrastructure. Similarity search is supported by using the concept of deletion neighborhood to evaluate the edit distance metric of string similarity. Thomas Bocek, Fabio Victora Hecht, David Hausheer, Elzbieta Pustulka, Burkhard Stiller |
CCNC | 4 |
| 2009 | Distributed Privilege Enforcement in PACS
Christoph Sturm, Elzbieta Pustulka, Marc H. Scholl |
DBSec | 2 |
| 2009 | Francisella tularensis novicida proteomic and transcriptomic data integration and annotation based on semantic web technologiesabstractBACKGROUND: This paper summarises the lessons and experiences gained from a case study of the application of semantic web technologies to the integration of data from the bacterial species Francisella tularensis novicida (Fn). Fn data sources are disparate and heterogeneous, as multiple laboratories across the world, using multiple technologies, perform experiments to understand the mechanism of virulence. It is hard to integrate these data sources in a flexible manner that allows new experimental data to be added and compared when required. RESULTS: Public domain data sources were combined in RDF. Using this connected graph of database cross references, we extended the annotations of an experimental data set by superimposing onto it the annotation graph. Identifiers used in the experimental data automatically resolved and the data acquired annotations in the rest of the RDF graph. This happened without the expensive manual annotation that would normally be required to produce these links. This graph of resolved identifiers was then used to combine two experimental data sets, a proteomics experiment and a transcriptomic experiment studying the mechanism of virulence through the comparison of wildtype Fn with an avirulent mutant strain. CONCLUSION: We produced a graph of Fn cross references which enabled the combination of two experimental datasets. Through combination of these data we are able to perform queries that compare the results of the two experiments. We found that data are easily combined in RDF and that experimental results are easily compared when the data are integrated. We conclude that semantic data integration offers a convenient, simple and flexible solution to the integration of published and unpublished experimental data. Nadia Anwar, Elzbieta Pustulka |
BMC Bioinform. | 2 |
| 2008 | Fast similarity search in peer-to-peer networksabstractPeer-to-peer (P2P) systems show numerous advantages over centralized systems, such as load balancing, scalability, and fault tolerance, and they require certain functionality, such as search, repair, and message and data transfer. In particular, structured P2P networks perform an exact search in logarithmic time proportional to the number of peers. However, keyword similarity search in a structured P2P network remains a challenge. Similarity search for service discovery can significantly improve service management in a distributed environment. As services are often described informally in text form, keyword similarity search can find the required services or data items more reliably. This paper presents a fast similarity search algorithm for structured P2P systems. The new algorithm, called P2P fast similarity search (P2PFastSS), finds similar keys in any distributed hash table (DHT) using the edit distance metric, and is independent of the underlying P2P routing algorithm. Performance analysis shows that P2PFastSS carries out a similarity search in time proportional to the logarithm of the number of peers. Simulations on PlanetLab confirm these results and show that a similarity search with 34,000 peers performs in less than three seconds on average. Thus, P2PFastSS is suitable for similarity search in large-scale network infrastructures, such as service description matching in service discovery or searching for similar terms in P2P storage networks. Thomas Bocek, Elzbieta Pustulka, David Hausheer, Burkhard Stiller |
NOMS | 2 |
| 2008 | PORSCHE: Performance ORiented SCHEma mediation
Khalid Saleem, Zohra Bellahsene, Elzbieta Pustulka |
Inf. Syst. | 3 |
| 2007 | Performance Oriented Schema Matching
Khalid Saleem, Zohra Bellahsene, Elzbieta Pustulka |
DEXA | 3 |
| 2007 | XBenchMatch: a Benchmark for XML Schema Matching Tools
Fabien Duchateau, Zohra Bellahsene, Elzbieta Pustulka |
VLDB | 3 |
| 2007 | VisGenome: visualization of single and comparative genome representationsabstractVisGenome visualizes single and comparative representations for the rat, the mouse and the human chromosomes at different levels of detail. The tool offers smooth zooming and panning which is more flexible than seen in other browsers. It presents information available in Ensembl for single chromosomes, as well as homologies (orthologue predictions including ortholog one2one, apparent ortholog one2one, ortholog many2many) for any two chromosomes from different species. The application can query supporting data from Ensembl by invoking a link in a browser. Joanna Jakubowska, Elzbieta Pustulka, Matthew Chalmers, Martin W. McBride, Anna F. Dominiczak |
Bioinform. | 2 |
| 2005 | System level visualization of eQTLs and pQTLs
Joanna Jakubowska, Elzbieta Pustulka, Matthew Chalmers, David Leader, Martin W. McBride, Anna F. Dominiczak |
BMC Bioinform. | 2 |
| 2004 | An object model and database for functional genomicsabstractMOTIVATION: Large-scale functional genomics analysis is now feasible and presents significant challenges in data analysis, storage and querying. Data standards are required to enable the development of public data repositories and to improve data sharing. There is an established data format for microarrays (microarray gene expression markup language, MAGE-ML) and a draft standard for proteomics (PEDRo). We believe that all types of functional genomics experiments should be annotated in a consistent manner, and we hope to open up new ways of comparing multiple datasets used in functional genomics. RESULTS: We have created a functional genomics experiment object model (FGE-OM), developed from the microarray model, MAGE-OM and two models for proteomics, PEDRo and our own model (Gla-PSI-Glasgow Proposal for the Proteomics Standards Initiative). FGE-OM comprises three namespaces representing (i) the parts of the model common to all functional genomics experiments; (ii) microarray-specific components; and (iii) proteomics-specific components. We believe that FGE-OM should initiate discussion about the contents and structure of the next version of MAGE and the future of proteomics standards. A prototype database called RNA And Protein Abundance Database (RAPAD), based on FGE-OM, has been implemented and populated with data from microbial pathogenesis. AVAILABILITY: FGE-OM and the RAPAD schema are available from http://www.gusdb.org/fge.html, along with a set of more detailed diagrams. RAPAD can be accessed by registration at the site. Andrew R. Jones, Elzbieta Pustulka, Jonathan M. Wastling, Angel D. Pizarro, Christian J. Stoeckert Jr. |
Bioinform. | 2 |
| 2002 | Database indexing for large DNA and protein sequence collections
Elzbieta Pustulka, Malcolm P. Atkinson 0001, Robert W. Irving |
VLDB J. | 1 |
| 2001 | A Database Index to Large Biological Sequences
Elzbieta Pustulka, Malcolm P. Atkinson 0001, Robert W. Irving |
VLDB | 1 |