Elzbieta Pustulka

dblp:318/5521 · also Ela Hunt, Ela Pustulka, Ela Pustulka-Hunt, Elzbieta Katarzyna Pustulka-Hunt · DBLP profile ↗
← Back
15ranked-venue papers
2as first author
2since 2021 · last 2024
0000-0001-7379-847XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 4Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1Systems, architecture and hardware · 1Security and privacy · 1
YearPublicationVenuePosition
2024 On the importance of CI/CD practices for database applications
abstract
Summary Continuous integration and continuous delivery (CI/CD) automate software integration and reduce repetitive engineering work. While the use of CI/CD presents efficiency gains, in database application development, this potential has not been fully exploited. We explore the state of the art in this area, with a focus on current practices, common software tools, challenges, and preconditions that apply to database applications. The work is grounded in a synoptic literature review and contributes a novel generic CI/CD pipeline for database system application development. Our generic pipeline was tailored to three industrial development use cases in which we measured the benefits of integration and deployment automation. The measurements demonstrate clearly that introducing CI/CD had significant benefits. It reduced the number of failed deployments, improved their stability, and increased the number of deployments. Interviews with the developers before and after the implementation of the CI/CD show that the pipeline brings clear benefits to the development team (i.e., a reduced cognitive load). These findings put current database release practices driven by business expectations, such as fixed release windows, in question.
Jasmin Fluri, Fabrizio Fornari 0001, Elzbieta Pustulka
J. Softw. Evol. Process.3
2023 Measuring the Benefits of CI/CD Practices for Database Application Development
abstract
Modern software development practices automate software integration and reduce repetitive software engineering work. Automation reduces the time it takes from defining software requirements to deploying the software in production. However, when it comes to database applications, the database integration and deployment are often executed manually, making it costly and error-prone. To mitigate this, we extended current software development methodologies by designing a CI/CD pipeline that takes into consideration the database setting. We report on two industrial case studies in which we implemented a newly designed pipeline and we measure the benefits of integration and deployment automation in database development projects. From a quantitative perspective, we found that introducing CI/CD pipelines reduces failed deployments, improves stability and increases the number of executed deployments. From a qualitative perspective, we interviewed the developers before and after the implementation of the CI/CD pipeline and the results show the CI/CD pipeline brings clear benefits to the development team (i.e., reduced cognitive load). This finding puts current database release practices driven by business expectations such as fixed release windows in question.
Jasmin Fluri, Fabrizio Fornari 0001, Elzbieta Pustulka
ICSSP3
2010 VisGenome with CartoonPlus: Supporting large scale genomic analyses via physical space deformation
Joanna Jakubowska, Elzbieta Pustulka, Matthew Chalmers
Future Gener. Comput. Syst.2
2009 Mobile P2P Fast Similarity Search
abstract
In informal data sharing environments, misspellings cause problems for data indexing and retrieval. This is even more pronounced in mobile environments, in which devices with limited input devices are used. In a mobile environment, similarity search algorithms for finding misspelled data need to account for limited CPU and bandwidth. This demo shows P2P fast similarity search (P2PFastSS) running on mobile phones and laptops that is tailored to uncertain data entry and uses available resources efficiently. In this demo, users publish and search for textual content containing misspellings without relying on query logging, as done by Google, and with a minimum distributed indexing infrastructure. Similarity search is supported by using the concept of deletion neighborhood to evaluate the edit distance metric of string similarity.
Thomas Bocek, Fabio Victora Hecht, David Hausheer, Elzbieta Pustulka, Burkhard Stiller
CCNC4
2009 Distributed Privilege Enforcement in PACS
Christoph Sturm, Elzbieta Pustulka, Marc H. Scholl
DBSec2
2009 Francisella tularensis novicida proteomic and transcriptomic data integration and annotation based on semantic web technologies
abstract
BACKGROUND: This paper summarises the lessons and experiences gained from a case study of the application of semantic web technologies to the integration of data from the bacterial species Francisella tularensis novicida (Fn). Fn data sources are disparate and heterogeneous, as multiple laboratories across the world, using multiple technologies, perform experiments to understand the mechanism of virulence. It is hard to integrate these data sources in a flexible manner that allows new experimental data to be added and compared when required. RESULTS: Public domain data sources were combined in RDF. Using this connected graph of database cross references, we extended the annotations of an experimental data set by superimposing onto it the annotation graph. Identifiers used in the experimental data automatically resolved and the data acquired annotations in the rest of the RDF graph. This happened without the expensive manual annotation that would normally be required to produce these links. This graph of resolved identifiers was then used to combine two experimental data sets, a proteomics experiment and a transcriptomic experiment studying the mechanism of virulence through the comparison of wildtype Fn with an avirulent mutant strain. CONCLUSION: We produced a graph of Fn cross references which enabled the combination of two experimental datasets. Through combination of these data we are able to perform queries that compare the results of the two experiments. We found that data are easily combined in RDF and that experimental results are easily compared when the data are integrated. We conclude that semantic data integration offers a convenient, simple and flexible solution to the integration of published and unpublished experimental data.
Nadia Anwar, Elzbieta Pustulka
BMC Bioinform.2
2008 Fast similarity search in peer-to-peer networks
abstract
Peer-to-peer (P2P) systems show numerous advantages over centralized systems, such as load balancing, scalability, and fault tolerance, and they require certain functionality, such as search, repair, and message and data transfer. In particular, structured P2P networks perform an exact search in logarithmic time proportional to the number of peers. However, keyword similarity search in a structured P2P network remains a challenge. Similarity search for service discovery can significantly improve service management in a distributed environment. As services are often described informally in text form, keyword similarity search can find the required services or data items more reliably. This paper presents a fast similarity search algorithm for structured P2P systems. The new algorithm, called P2P fast similarity search (P2PFastSS), finds similar keys in any distributed hash table (DHT) using the edit distance metric, and is independent of the underlying P2P routing algorithm. Performance analysis shows that P2PFastSS carries out a similarity search in time proportional to the logarithm of the number of peers. Simulations on PlanetLab confirm these results and show that a similarity search with 34,000 peers performs in less than three seconds on average. Thus, P2PFastSS is suitable for similarity search in large-scale network infrastructures, such as service description matching in service discovery or searching for similar terms in P2P storage networks.
Thomas Bocek, Elzbieta Pustulka, David Hausheer, Burkhard Stiller
NOMS2
2008 PORSCHE: Performance ORiented SCHEma mediation
Khalid Saleem, Zohra Bellahsene, Elzbieta Pustulka
Inf. Syst.3
2007 Performance Oriented Schema Matching
Khalid Saleem, Zohra Bellahsene, Elzbieta Pustulka
DEXA3
2007 XBenchMatch: a Benchmark for XML Schema Matching Tools
Fabien Duchateau, Zohra Bellahsene, Elzbieta Pustulka
VLDB3
2007 VisGenome: visualization of single and comparative genome representations
abstract
VisGenome visualizes single and comparative representations for the rat, the mouse and the human chromosomes at different levels of detail. The tool offers smooth zooming and panning which is more flexible than seen in other browsers. It presents information available in Ensembl for single chromosomes, as well as homologies (orthologue predictions including ortholog one2one, apparent ortholog one2one, ortholog many2many) for any two chromosomes from different species. The application can query supporting data from Ensembl by invoking a link in a browser.
Joanna Jakubowska, Elzbieta Pustulka, Matthew Chalmers, Martin W. McBride, Anna F. Dominiczak
Bioinform.2
2005 System level visualization of eQTLs and pQTLs
Joanna Jakubowska, Elzbieta Pustulka, Matthew Chalmers, David Leader, Martin W. McBride, Anna F. Dominiczak
BMC Bioinform.2
2004 An object model and database for functional genomics
abstract
MOTIVATION: Large-scale functional genomics analysis is now feasible and presents significant challenges in data analysis, storage and querying. Data standards are required to enable the development of public data repositories and to improve data sharing. There is an established data format for microarrays (microarray gene expression markup language, MAGE-ML) and a draft standard for proteomics (PEDRo). We believe that all types of functional genomics experiments should be annotated in a consistent manner, and we hope to open up new ways of comparing multiple datasets used in functional genomics. RESULTS: We have created a functional genomics experiment object model (FGE-OM), developed from the microarray model, MAGE-OM and two models for proteomics, PEDRo and our own model (Gla-PSI-Glasgow Proposal for the Proteomics Standards Initiative). FGE-OM comprises three namespaces representing (i) the parts of the model common to all functional genomics experiments; (ii) microarray-specific components; and (iii) proteomics-specific components. We believe that FGE-OM should initiate discussion about the contents and structure of the next version of MAGE and the future of proteomics standards. A prototype database called RNA And Protein Abundance Database (RAPAD), based on FGE-OM, has been implemented and populated with data from microbial pathogenesis. AVAILABILITY: FGE-OM and the RAPAD schema are available from http://www.gusdb.org/fge.html, along with a set of more detailed diagrams. RAPAD can be accessed by registration at the site.
Andrew R. Jones, Elzbieta Pustulka, Jonathan M. Wastling, Angel D. Pizarro, Christian J. Stoeckert Jr.
Bioinform.2
2002 Database indexing for large DNA and protein sequence collections
Elzbieta Pustulka, Malcolm P. Atkinson 0001, Robert W. Irving
VLDB J.1
2001 A Database Index to Large Biological Sequences
Elzbieta Pustulka, Malcolm P. Atkinson 0001, Robert W. Irving
VLDB1