VLDB 2026 Research / reviewers in the wild / expert
Bozena Malysiak-Mrozek
dblp:51/4657 · also Bozena Malysiak
· DBLP profile ↗
27ranked-venue papers
10as first author
8since 2021 · last 2025
0000-0003-4977-4915ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 10 first-author · 8 since 2021Databases, data management, data science and information retrieval · 10 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 3 since 2021Systems, architecture and hardware · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorTheory of computation · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | From Genetic Reads to Information Granules: Scalable Big NGS Data Cleaning with Apache Pig
Bozena Malysiak-Mrozek, Tomasz Sitek, Vaidy S. Sunderam, Boleslaw Pochopien, Krzysztof Tokarz, Che-Lun Hung, Dariusz Mrozek |
IEEE Big Data | 1 |
| 2025 | Effective Detection and Recognition of Traffic Signs with Light Convolutional Neural NetworksabstractDetection and recognition of traffic signs are two analytic processes in vehicular systems that contribute to increasing driver safety, warning, and preventing collisions by improving drivers’ focus and awareness. They are also crucial for developing self-driving vehicles that can sense the environment through a camera eye and understand the road restrictions. Convolutional Neural Networks (CNNs) play an essential role in both processes by finding the traffic sign objects on acquired images or video frames and recognizing their meaning. However, CNN architectures running on vehicles need minimization, ensuring satisfactory performance and efficient operation with decreased computational resources. In this paper, we investigate two CNN architectures for traffic sign detection and two architectures for traffic sign recognition. Our experiments confirm that light models for both processes can successfully perform the achieved tasks, reaching effectiveness close to the complex models reported in the scientific literature. Maciej Olszewski, Bozena Malysiak-Mrozek, Krzysztof Tokarz, Boleslaw Pochopien, Che-Lun Hung, Andrzej Pulka, Dariusz Mrozek |
KES | 2 |
| 2024 | Fuzzy Querying in the Cloud-based Environment for Data Stream-driven Predictive Maintenance in AGV-enabled Smart FactoriesabstractFuzzy data processing enables data enrichment and increases data interpretation in industrial environments. In the cloud-based IoT data ingestion pipelines, fuzzy data processing can be implemented in several locations, closer to the IoT events gateways, stream processors, or the persistence layer before the data is visualized. Since Automated Guided Vehicles (AGV)-enabled manufacturing can produce vast amounts of data, the decision on the placement of the fuzzy data processing can be important for secondary processes performed on the enriched data, like the predictive maintenance inferencing. In this paper, we analyze two locations of fuzzy data processing in the cloud-based environment built for monitoring AGVs in smart factories - by formulating fuzzy queries against data streams on stream processing units and data at rest in a database. The querying scenarios cover fuzzy filtering with simple and complex criteria, fuzzy filtering through assignment to a linguistic variable, and joining data streams by representing joining attributes as fuzzy numbers. The experimental results show that querying the data stream can be more efficient and profitable in the scalable environment of many AGVs. However, the enrichment provided for the data at rest is also beneficial when gathering data for building future predictive maintenance models. Bozena Malysiak-Mrozek, Dominik Romanów, Piotr Grzesik, Pawel Benecki, Alexandre Niyomugaba, Theodore Habimana, Daniel Kostrzewa, Krzysztof Tokarz, Che-Lun Hung, Dariusz Mrozek |
IEEE Big Data | 1 |
| 2024 | Decoding the Granular Puzzle of Macromolecules: Efficient 3D Protein Structure Alignment in the Age of Big Data with Apache SparkabstractProteins are complex biological information granules that play a crucial role in various cellular processes within living organisms. Processing 3D protein structures, which are the most informative from the biological point of view, is both intricate and time-consuming. In particular, performing 3D protein structure searches against large protein datasets involves identifying similarities and conducting structural alignments across numerous molecules (granules). This task demands advanced methods for matching identical and similar regions within protein structures and substantial computational resources to handle large collections of macromolecular data efficiently. In this paper, we present our parallel implementation of scalable 3D structural alignment on the Apache Spark big data platform. We describe a customized approach that leverages Spark data transformations within the data processing pipeline for the alignment process. Our experimental results demonstrate that this solution, tightly integrated with the Spark processing model, is both efficient and scalable, even with the increasing volume of protein structure data. Bozena Malysiak-Mrozek, Paulina Pawlowicz, Vaidy S. Sunderam, Che-Lun Hung, Andrzej Kwiecien, Dariusz Mrozek |
IEEE Big Data | 1 |
| 2024 | Interpreting Industrial IoT Data Streams Through Fuzzy Querying With Hysteretic Fuzzy Sets on Apache KafkaabstractIn industrial settings, querying data streams from Internet of Things (IoT) devices benefits from utilizing elastic criteria to enhance the interpretability of the current state of the monitored environment. Fuzzy sets provide this elasticity, enabling the aggregation and representation of similar values in a human-comprehensible manner. However, many sensor signals exhibit temporal oscillations, leading to varying interpretations of the signal based on its current trend (rising or falling). This hysteresis in signal (and subsequently of the production device) interpretation inspired us to introduce this phenomenon into data stream processing, resulting in the novel concept of hysteretic fuzzy sets. This paper demonstrates how fuzzy searching and grouping can be applied to IoT sensor signals in flexible Big Data stream processing on Apache Kafka. We illustrate the impact of data stream querying with KSQL queries involving fuzzy sets (encompassing fuzzy filtering of data stream events, fuzzy transformation of data stream attributes, fuzzy grouping, and joining) on the flexibility of executed operations and computational resources utilized by the Kafka processing engine. Finally, our experiments with hysteretic fuzzy sets while analyzing sensor signals in power plants demonstrate that this novel approach effectively reduces the number of alarms while monitoring the state of the production machine. Bozena Malysiak-Mrozek, Bartlomiej Ryba, Marek Moleda, Che-Lun Hung, Witold Pedrycz, Weiping Ding 0001, Dariusz Mrozek |
IEEE Trans. Fuzzy Syst. | 1 |
| 2023 | Predicting Conflict Zones on Terrestrial Routes of Automated Guided Vehicles with Fuzzy Querying on Apache KafkaabstractIn today’s world, smart factories are a coexisting element of smarticizing cities. Smart manufacturing of today relies on the automation of many component tasks of the production process. Automated guided vehicles (AGVs) that transport materials on the production lines are important elements of this automation. Appropriate management of a fleet of AGVs requires avoiding collisions. However, prediction and early detection of approaching collision points on the transportation routes not only prevent collisions but also enables adjusting the AGV operation and improving its flow. In this paper, we demonstrate the use of fuzzy sets and linguistic variables in collision prevention by processing AGV data streams with Apache Kafka. We extend the capabilities of Apache Kafka and ksqlDB towards fuzzy stream processing and use fuzzy KSQL queries to predict collisions. Our experiments prove that fuzzy querying against AGV data streams does not consume much time and computational resources, and we can successfully avoid collisions by predicting future positions of the AGV for various densities of data streams and widths of time windows. Bozena Malysiak-Mrozek, Mario Bas, Vaidy S. Sunderam, Stanislaw Kozielski, Dariusz Mrozek |
DSAA | 1 |
| 2022 | High-Efficient Fuzzy Querying With HiveQL for Big Data WarehousingabstractQuerying and reporting from large volumes of structured, semistructured, and unstructured data often requires some flexibility. This flexibility provided by fuzzy sets allows for categorization of the surrounding world in a flexible, human-mind-like manner. Apache Hive is a data warehousing framework working on top of the Hadoop platform for big data processing. Hive allows executing queries and aggregating and analyzing data stored in Hadoop distributed file system and other repositories. Hive responds to the current needs for efficient big data warehousing, which is impossible with traditional data warehouses due to their rigid nature. This article presents the FuzzyHive library that extends the Hive framework with fuzzy sets based techniques for querying, analyzing, and reporting on big data warehouses. We formalize the fuzzy techniques used while operating on Hive-based data warehouses (including fuzzy filtering on dimensional attributes, projection with fuzzy transformation, fuzzy grouping, and joining). We also show how we embedded these operations in Hive query language, which was not studied so far. Such extensions make big data warehousing more flexible and contribute to the portfolio of tools used by the community of people working with fuzzy sets and data analysis. The FuzzyHive library complements the spectrum of available solutions for fuzzy data processing and querying in large datasets. We investigate Hive fuzzy querying performance, effectiveness, and scalability for various data storage formats (text, Avro, and Parquet). Our experiments demonstrate that the proposed extensions introduce more elasticity and are also efficient for big data warehousing, which is the first such kind of solution for this environment. Bozena Malysiak-Mrozek, Jadwiga Wieszok, Witold Pedrycz, Weiping Ding 0001, Dariusz Mrozek |
IEEE Trans. Fuzzy Syst. | 1 |
| 2021 | Fuzzy Filtering in Large-Scale Prediction of Intrinsically Disordered Regions of Proteins on Apache SparkabstractIntrinsically disordered proteins (IDPs) participate in many cellular processes. They are also studied for their participation in the course and formation of many diseases. Experimental determination of disordered regions (IDRs) is costly and not always possible. Due to the exponential growth of protein sequences, for which the 3D structure cannot be experimentally determined, computational prediction becomes an important alternative. Spark-IDPP is the large-scale meta-predictor for IDRs and IDPs designed to run on the Apache Spark cluster. The meta-prediction with Spark-IDPP includes fuzzy filtering of produced prediction output. Here, we experimentally validate various fuzzy filters and show that a properly designed characteristic function for fuzzy filtering may improve the prediction quality in all modes of the Spark-IDPP execution. Bozena Malysiak-Mrozek, Lukasz Bozek, Dariusz Mrozek |
CEC | 1 |
| 2020 | Fall detection in older adults with mobile IoT devices and machine learning in the cloud and on the edgeabstractRemote monitoring of older adults and detecting dangers in the state of human health have become essential elements of modern telemedicine. Falls are a frequent reason for deaths or post-traumatic complications in the elderly. Therefore, the early detection of falls can be crucial for the survival of a person or for providing necessary support. However, telemedicine data centers require scalable computing and storage resources for the growing number of monitored people. Dedicated approaches that allow for minimal data transmission of strictly interesting cases are also required. In this paper, we show a scalable architecture of a system that can monitor thousands of older adults, detect falls, and notify caregivers. Scalability tests that disclose requirements to enable large scale system operations were also performed. Moreover, we validated several Machine Learning models to evaluate their suitability in the detection process. Among the tested models, Boosted Decisions Trees resulted in the best classification performance. We also experimentally tested the detection of falls inside a Cloud-based data center and on an Edge IoT device. Results of tests on the device-to-cloud data transmission confirmed that significant reduction in the size of stored and transmitted data can be achieved while performing fall detection on the Edge. Dariusz Mrozek, Anna Koczur, Bozena Malysiak-Mrozek |
Inf. Sci. | 3 |
| 2020 | A Hopping Umbrella for Fuzzy Joining Data Streams From IoT Devices in the Cloud and on the EdgeabstractInternet of Things (IoT) is a new technology that changes the image of the current world, yielding new possibilities, but also proliferating data. IoT devices may constantly produce enormous amounts of data as data streams that can be analyzed in real time and also collected for further exploration in data lakes in huge data centers. Due to their scaling capabilities, these data centers are frequently located in the Cloud. However, recent rapid growth in the number of IoT devices and their applications in manufacturing, transport, and health care motivates moving the burden of data processing and analysis to the Edge. One of the phases of data processing is combining data streams from two (or more) IoT devices that monitor the same object and work asynchronously. Since they generate sensor readings at various moments of time, their data streams must be properly combined in order to obtain a complete image of the monitored object or process. In this article, we present the idea of a hopping umbrella which fuzzifies timestamps from sensor readings while joining data streams from asynchronous IoT devices in a flexible way. In contrast to processing data at rest, the hopping umbrella implements the fuzzy join operation in time windows for data streams (data in motion). By using fuzzy sets, the hopping umbrella not only allows combining asynchronous events from multiple sensors, but also facilitates evaluation of the degree of matching of the combined sensor readings, and consequently allows for reduction of the output stream size. Our experiments performed in Cloud and on Edge devices proved that with the use of this idea, we are able to properly join the best matching sensor readings and in some scenarios, reduce the number of data transferred to the Cloud data center without significant overhead in resource utilization of stream processing units. Dariusz Mrozek, Krzysztof Tokarz, Daniel Pankowski, Bozena Malysiak-Mrozek |
IEEE Trans. Fuzzy Syst. | 4 |
| 2019 | High-throughput and scalable protein function identification with Hadoop and Map-only pattern of the MapReduce processing modelabstractEfficient computational solutions for identification of protein functions or finding structural homologs of proteins gain importance in the era of structural genomics and in the face of growing volumes of biological data. Structural alignments, which underlie these two processes, take a lot of time to complete, especially when performed for large collections of 3D protein structures. Fortunately, structural alignments can be carried out on well-separable and independent subsets of the whole macromolecular data repository, which perfectly fits the MapReduce processing paradigm of bringing computations to data. In this paper, we show how the protein function identification and finding structural homologs can be efficiently accelerated with the use of the MapReduce procedure executed on Hadoop cluster established in a virtualized compute environment or a private cloud. For this purpose, we propose Map-only processing pattern of the MapReduce procedure, which is formally defined in this paper. The solution that we show joins advantages of performing computations in small virtualized compute environments with large-scale computations in public clouds, thus allowing to perform structural alignments for a number of usage scenarios, including comparison of pairs of 3D protein structures during evaluation of predicted protein models, one-to-many comparisons while identifying possible functions of the given structure, or all-to-all alignments while investigating the divergence between known protein structures and classifying proteins by their fold. In this paper, we also present results of performance tests when scaling up nodes of the Hadoop cluster and increasing the degree of parallelism with the intention of improving efficiency of the computations. Dariusz Mrozek, Marek Suwala, Bozena Malysiak-Mrozek |
Knowl. Inf. Syst. | 3 |
| 2018 | Soft and Declarative Fishing of Information in Big Data LakeabstractIn recent years, many fields that experience a sudden proliferation of data, which increases the volume of data that must be processed and the variety of formats the data is stored in have been identified. This causes pressure on existing compute infrastructures and data analysis methods, as more and more data are considered as a useful source of information for making critical decisions in particular fields. Among these fields exist several areas related to human life, e.g., various branches of medicine, where the uncertainty of data complicates the data analysis, and where the inclusion of fuzzy expert knowledge in data processing brings many advantages. In this paper, we show how fuzzy techniques can be incorporated in big data analytics carried out with the declarative U-SQL language over a big data lake located on the cloud. We define the concept of big data lake together with the Extract, Process, and Store process performed while schematizing and processing data from the Data Lake, and while storing results of the processing. Our solution, developed as a Fuzzy Search Library for Data Lake, introduces the possibility of massively parallel, declarative querying of big data lake with simple and complex fuzzy search criteria, using fuzzy linguistic terms in various data transformations, and fuzzy grouping. Presented ideas are exemplified by a distributed analysis of large volumes of biomedical data on Microsoft Azure cloud. Results of performed tests confirm that the presented solution is highly scalable on the Cloud and is a successful step toward soft and declarative processing of data on a large scale. The solution presented in this paper directly addresses three characteristics of big data, i.e., volume, variety, and velocity, and indirectly addresses, veracity and value. Bozena Malysiak-Mrozek, Marek Stabla, Dariusz Mrozek |
IEEE Trans. Fuzzy Syst. | 1 |
| 2017 | Orchestrating Task Execution in Cloud4PSi for Scalable Processing of Macromolecular Data of 3D Protein Structures
Dariusz Mrozek, Artur Klapcinski, Bozena Malysiak-Mrozek |
ACIIDS (2) | 3 |
| 2017 | Life Sciences Data Analysis
Dariusz Mrozek, Pawel Kasprowski, Bozena Malysiak-Mrozek, Stanislaw Kozielski |
Inf. Sci. | 3 |
| 2016 | HDInsight4PSi: Boosting performance of 3D protein structure similarity searching with HDInsight clusters in Microsoft Azure cloud
Dariusz Mrozek, Pawel Danilowicz, Bozena Malysiak-Mrozek |
Inf. Sci. | 3 |
| 2016 | An efficient and flexible scanning of databases of protein secondary structures - with the segment index and multithreaded alignmentabstractProtein secondary structure describe protein construction in terms of regular spatial shapes, including alpha-helices, beta-strands, and loops, which protein amino acid chain can adopt in some of its regions. This information is supportive for protein classification, functional annotation, and 3D structure prediction. The relevance of this information and the scope of its practical applications cause the requirement for its effective storage and processing. Relational databases, widely-used in commercial systems in recent years, are one of the serious alternatives honed by years of experience, enriched with developed technologies, equipped with the declarative SQL query language, and accepted by the large community of programmers. Unfortunately, relational database management systems are not designed for efficient storage and processing of biological data, such as protein secondary structures. In this paper, we present a new search method implemented in the search engine of the PSS-SQL language. The PSS-SQL allows formulation of queries against a relational database in order to find proteins having secondary structures similar to the structural pattern specified by a user. In the paper, we will show how the search process can be accelerated by multiple scanning of the Segment Index and parallel implementation of the alignment procedure using multiple threads working on multiple-core CPUs. Dariusz Mrozek, Bartek Socha, Stanislaw Kozielski, Bozena Malysiak-Mrozek |
J. Intell. Inf. Syst. | 4 |
| 2015 | Scaling Ab Initio Predictions of 3D Protein Structures in Microsoft Azure CloudabstractComputational methods for protein structure prediction allow us to determine a three-dimensional structure of a protein based on its pure amino acid sequence. These methods are a very important alternative to costly and slow experimental methods, like X-ray crystallography or Nuclear Magnetic Resonance. However, conventional calculations of protein structure are time-consuming and require ample computational resources, especially when carried out with the use of ab initio methods that rely on physical forces and interactions between atoms in a protein. Fortunately, at the present stage of the development of computer science, such huge computational resources are available from public cloud providers on a pay-as-you-go basis. We have designed and developed a scalable and extensible system, called Cloud4PSP, which enables predictions of 3D protein structures in the Microsoft Azure commercial cloud. The system makes use of the Warecki-Znamirowski method as a sample procedure for protein structure prediction, and this prediction method was used to test the scalability of the system. The results of the efficiency tests performed proved good acceleration of predictions when scaling the system vertically and horizontally. In the paper, we show the system architecture that allowed us to achieve such good results, the Cloud4PSP processing model, and the results of the scalability tests. At the end of the paper, we try to answer which of the scaling techniques, scaling out or scaling up, is better for solving such computational problems with the use of Cloud computing. Dariusz Mrozek, Pawel Gosk, Bozena Malysiak-Mrozek |
J. Grid Comput. | 3 |
| 2014 | Cloud4Psi: cloud computing for 3D protein structure similarity searchingabstractSUMMARY: Popular methods for 3D protein structure similarity searching, especially those that generate high-quality alignments such as Combinatorial Extension (CE) and Flexible structure Alignment by Chaining Aligned fragment pairs allowing Twists (FATCAT) are still time consuming. As a consequence, performing similarity searching against large repositories of structural data requires increased computational resources that are not always available. Cloud computing provides huge amounts of computational power that can be provisioned on a pay-as-you-go basis. We have developed the cloud-based system that allows scaling of the similarity searching process vertically and horizontally. Cloud4Psi (Cloud for Protein Similarity) was tested in the Microsoft Azure cloud environment and provided good, almost linearly proportional acceleration when scaled out onto many computational units. AVAILABILITY AND IMPLEMENTATION: Cloud4Psi is available as Software as a Service for testing purposes at: http://cloud4psi.cloudapp.net/. For source code and software availability, please visit the Cloud4Psi project home page at http://zti.polsl.pl/dmrozek/science/cloud4psi.htm. Dariusz Mrozek, Bozena Malysiak-Mrozek, Artur Klapcinski |
Bioinform. | 2 |
| 2013 | search GenBank: interactive orchestration and ad-hoc choreography of Web services in the exploration of the biomedical resources of the National Center For Biotechnology InformationabstractBACKGROUND: Due to the growing number of biomedical entries in data repositories of the National Center for Biotechnology Information (NCBI), it is difficult to collect, manage and process all of these entries in one place by third-party software developers without significant investment in hardware and software infrastructure, its maintenance and administration. Web services allow development of software applications that integrate in one place the functionality and processing logic of distributed software components, without integrating the components themselves and without integrating the resources to which they have access. This is achieved by appropriate orchestration or choreography of available Web services and their shared functions. After the successful application of Web services in the business sector, this technology can now be used to build composite software tools that are oriented towards biomedical data processing. RESULTS: We have developed a new tool for efficient and dynamic data exploration in GenBank and other NCBI databases. A dedicated search GenBank system makes use of NCBI Web services and a package of Entrez Programming Utilities (eUtils) in order to provide extended searching capabilities in NCBI data repositories. In search GenBank users can use one of the three exploration paths: simple data searching based on the specified user's query, advanced data searching based on the specified user's query, and advanced data exploration with the use of macros. search GenBank orchestrates calls of particular tools available through the NCBI Web service providing requested functionality, while users interactively browse selected records in search GenBank and traverse between NCBI databases using available links. On the other hand, by building macros in the advanced data exploration mode, users create choreographies of eUtils calls, which can lead to the automatic discovery of related data in the specified databases. CONCLUSIONS: search GenBank extends standard capabilities of the NCBI Entrez search engine in querying biomedical databases. The possibility of creating and saving macros in the search GenBank is a unique feature and has a great potential. The potential will further grow in the future with the increasing density of networks of relationships between data stored in particular databases. search GenBank is available for public use at http://sgb.biotools.pl/. Dariusz Mrozek, Bozena Malysiak-Mrozek, Artur Siaznik |
BMC Bioinform. | 2 |
| 2011 | Fast and Accurate Similarity Searching of Biopolymer Sequences with GPU and CUDA
Robert Pawlowski, Bozena Malysiak-Mrozek, Stanislaw Kozielski, Dariusz Mrozek |
ICA3PP (1) | 2 |
| 2011 | Scalable System for Protein Structure Similarity Searching
Bozena Malysiak-Mrozek, Alina Momot, Dariusz Mrozek, Lukasz Hera, Stanislaw Kozielski, Michal Momot |
ICCCI (2) | 1 |
| 2010 | Improving Performance of Protein Structure Similarity Searching by Distributing Computations in Hierarchical Multi-Agent System
Alina Momot, Bozena Malysiak-Mrozek, Stanislaw Kozielski, Dariusz Mrozek, Lukasz Hera, Sylwia Górczynska-Kosiorz, Michal Momot |
ICCCI (1) | 2 |
| 2010 | Processing of Crisp and Fuzzy Measures in the Fuzzy Data Warehouse for Global Natural Resources
Bozena Malysiak-Mrozek, Dariusz Mrozek, Stanislaw Kozielski |
IEA/AIE (3) | 1 |
| 2009 | The Energy Distribution Data Bank: Collecting Energy Features of Protein Molecular StructuresabstractThe analysis of structural and energy features of proteins can be a key to understand how proteins work and interact to each other in cellular reactions. Potential energy is a function of atomic positions in a protein structure. The distributions of energy over each atom in protein structures can be very supportive for the studies of the complex processes proteins are involved in. Energy profiles contain distributions of different potential energies in protein molecular structures. Therefore, they constitute a full descriptor of energy properties for protein structures. The Energy Distribution Data Bank (EDB, http://edb.aei.polsl.pl) stores energy profiles for protein molecular structures retrieved from the well-known Protein Data Bank. In the paper, we describe the purpose of the EDB, a possible use of the information stored in it, query possibilities, and plans for future development. Dariusz Mrozek, Bozena Malysiak-Mrozek, Stanislaw Kozielski, Andrzej Swierniak |
BIBE | 2 |
| 2009 | The EDML Format to Exchange Energy Profiles of Protein Molecular Structures
Dariusz Mrozek, Bozena Malysiak-Mrozek, Stanislaw Kozielski, Sylwia Górczynska-Kosiorz |
ICIC (1) | 2 |
| 2007 | An Optimal Alignment of Proteins Energy Characteristics with Crisp and Fuzzy Similarity AwardsabstractWe discuss the usage of constant and fuzzy similarity awards while establishing an optimal alignment between energy characteristics of two compared protein energy profiles. Single protein energy profile is a set of energy characteristics of various types of energy. The energy profile is determined for a given protein structure. We use these profiles to find protein molecules of the same structural protein family and inspect conformational modifications in their molecular structures as an effect of biochemical reactions or environmental influences. Energy profiles are received in the computational process based on the molecular mechanics theory. Afterwards, these profiles can be stored in the special purpose database (EDB) and used by the search engine to find similar fragments of protein structures on the energy level. To optimize the alignment path we use modified, energy-adapted Smith-Waterman method with one of the tested similarity awards. Dariusz Mrozek, Bozena Malysiak-Mrozek, Stanislaw Kozielski |
FUZZ-IEEE | 2 |
| 2007 | Agent-Supported Protein Structure Similarity Searching
Dariusz Mrozek, Bozena Malysiak-Mrozek, Wojciech Augustyn |
PRIMA | 2 |