Robert Wrembel

dblp:41/3391 · DBLP profile ↗
← Back
50ranked-venue papers in the field
9as first author
21since 2021 · last 2026
0000-0001-6037-5718ORCID · verified

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 39 (6 first)Data Mining & Knowledge Discovery · 9 (2 first)Information Retrieval & Web Search · 2 (1 first)
YearPublicationVenuePosition
2026 Relational Database Data Lineage Ontology
Jakub Dutkiewicz, Pawel Misiorek, Robert Wrembel
DaWaK3
2025 Data Integration in the AI Era: Research Trends and Still Open Issues
Robert Wrembel
DaWaK1
2025 Leveraging Machine Learning Techniques for Customer Data Deduplication - Hard-Won Lessons from a Real-World Project in the Financial Industry
Robert Wrembel, Witold Andrzejewski, Pawel Boinski, Bartosz Bebel
DaWaK1
2025 Advances in databases and information systems - Selected papers from ADBIS 2023
Alberto Abelló, Ladjel Bellatreche, Oscar Romero 0001, Panos Vassiliadis, Robert Wrembel
Inf. Syst.5
2025 Advances on data management systems
Ladjel Bellatreche, Marlon Dumas, Panagiotis Karras, Raimundas Matulevicius, Silvia Chiusano, Tania Cerquitelli, Robert Wrembel
Inf. Syst.7
2024 Optimizing Data Integration Processes with the Support of Machine Learning - Is it really possible?
Robert Wrembel
DOLAP1
2024 Integrating the Biological Knowledge from Protein Databases Into Spatial RNA Sequencing Analyses
Anna Lesniewska, Szymon Dziegielewski, Elena Melnyk, Michal J. Okoniewski, Robert Wrembel
iiWAS (2)5
2024 Data analytics and knowledge discovery on big data: Algorithms, architectures, and applications
Robert Wrembel, Johann Gamper
Data Knowl. Eng.1
2024 On tuning parameters guiding similarity computations in a data deduplication pipeline for customers records: Experience from a R&D project
Witold Andrzejewski, Bartosz Bebel, Pawel Boinski, Robert Wrembel
Inf. Syst.4
2023 On Tuning the Sorted Neighborhood Method for Record Comparisons in a Data Deduplication Pipeline - Industrial Experience Report
Pawel Boinski, Witold Andrzejewski, Bartosz Bebel, Robert Wrembel
DEXA (1)4
2023 Data Integration Revitalized: From Data Warehouse Through Data Lake to Data Mesh
Robert Wrembel
DEXA (1)1
2023 Text Similarity Measures in a Data Deduplication Pipeline for Customers Records
Witold Andrzejewski, Bartosz Bebel, Pawel Boinski, Mariusz Sienkiewicz, Robert Wrembel
DOLAP5
2023 Data Source Connectors Layer as a Service - Design Patterns
Michal Bodziony, Robert Wrembel
DOLAP2
2023 Big Data Analytics and Knowledge Discovery
Matteo Golfarelli, Robert Wrembel
Data Knowl. Eng.2
2023 A large reproducible benchmark on text classification for the legal domain based on the ECHR-OD repository
Alexandre Quemy, Robert Wrembel, Natalia Lopuszynska, George Papadakis 0001, Agustín D. Delgado
Inf. Syst.2
2022 Quality Versus Speed in Energy Demand Prediction - Experience Report from an R &D project
Witold Andrzejewski, Jedrzej Potoniec, Maciej Drozdowski, Jerzy Stefanowski, Robert Wrembel, Pawel Stapf
DEXA (1)5
2022 What Logical Model Is Suitable for Relational Trajectory Data Warehouses? - Application to Agricultural Autonomous Robots
Georgia Garani, Sandro Bimonte, Robert Wrembel
DEXA (1)4
2022 Data Integration, Cleaning, and Deduplication: Research Versus Industrial Projects
Robert Wrembel
iiWAS1
2022 Data processing in modern distributed architectures
Jérôme Darmont, Boris Novikov 0001, Robert Wrembel, Ladjel Bellatreche
Inf. Syst.3
2022 ECHR-OD: On building an integrated open repository of legal documents for machine learning applications
Alexandre Quemy, Robert Wrembel
Inf. Syst.2
2021 Reference Architecture for Running Large Scale Data Integration Experiments
Michal Bodziony, Robert Wrembel
DEXA (1)2
2020 Framework to Optimize Data Processing Pipelines Using Performance Metrics
Syed Muhammad Fawad Ali, Robert Wrembel
DaWaK2
2020 Data Engineering for Data Science: Two Sides of the Same Coin
Oscar Romero 0001, Robert Wrembel
DaWaK2
2020 On Integrating and Classifying Legal Text Documents
Alexandre Quemy, Robert Wrembel
DEXA (1)2
2020 On Evaluating Performance of Balanced Optimization of ETL Processes for Streaming Data Sources
Michal Bodziony, Szymon Roszyk, Robert Wrembel
DOLAP3
2020 An Alternative View on Data Processing Pipelines from the DOLAP 2019 Perspective
Oscar Romero 0001, Robert Wrembel, Il-Yeol Song
Inf. Syst.2
2019 Towards a Cost Model to Optimize User-Defined Functions in an ETL Workflow Based on User-Defined Performance Metrics
Syed Muhammad Fawad Ali, Robert Wrembel
ADBIS2
2019 Metadata Discovery Using Data Sampling and Exploratory Data Analysis
Hiba Khalid, Robert Wrembel, Esteban Zimányi
MEDI2
2019 PRESISTANT: Learning based assistant for data pre-processing
Besim Bilalli, Alberto Abelló, Tomàs Aluja-Banet, Robert Wrembel
Data Knowl. Eng.4
2019 DOLAP data warehouse research over two decades: Trends and challenges
Robert Wrembel, Alberto Abelló, Il-Yeol Song
Inf. Syst.1
2018 Fuzzy Metadata Strategies for Enhanced Data Integration
Hiba Khalid, Esteban Zimányi, Robert Wrembel
DATA3
2017 Context Similarity for Retrieval-Based Imputation
abstract
Completeness as one of the four major dimensions of data quality is a pervasive issue in modern databases. Although data imputation has been studied extensively in the literature, most of the research is focused on inference-based approach. We propose to harness Web tables as an external data source to effectively and efficiently retrieve missing data while taking into account the inherent uncertainty and lack of veracity that they contain.
Ahmad Ahmadov, Maik Thiele, Wolfgang Lehner, Robert Wrembel
ASONAM4
2017 From conceptual design to performance optimization of ETL workflows: current state of research and open problems
abstract
In this paper, we discuss the state of the art and current trends in designing and optimizing ETL workflows. We explain the existing techniques for: (1) constructing a conceptual and a logical model of an ETL workflow, (2) its corresponding physical implementation, and (3) its optimization, illustrated by examples. The discussed techniques are analyzed w.r.t. their advantages, disadvantages, and challenges in the context of metrics such as autonomous behavior, support for quality metrics, and support for ETL activities as user-defined functions. We draw conclusions on still open research and technological issues in the field of ETL. Finally, we propose a theoretical ETL framework for ETL optimization.
Syed Muhammad Fawad Ali, Robert Wrembel
VLDB J.2
2016 On Representing Interval Measures by Means of Functions
Gastón Bakkalian, Christian Koncilia, Robert Wrembel
MEDI3
2016 Automated Data Pre-processing via Meta-learning
Besim Bilalli, Alberto Abelló, Tomàs Aluja-Banet, Robert Wrembel
MEDI4
2015 A Generic Data Warehouse Architecture for Analyzing Workflow Logs
Christian Koncilia, Horst Pichler, Robert Wrembel
ADBIS3
2015 Sequential Data Analytics by Means of Seq-SQL Language
Bartosz Bebel, Tomasz Cichowicz, Tadeusz Morzy, Filip Rytwinski, Robert Wrembel, Christian Koncilia
DEXA (1)5
2014 A Logical Model for Multiversion Data Warehouses
Waqas Ahmed 0003, Esteban Zimányi, Robert Wrembel
DaWaK3
2014 Interval OLAP: Analyzing Interval Data
Christian Koncilia, Tadeusz Morzy, Robert Wrembel, Johann Eder
DaWaK3
2012 RTDW-bench: Benchmark for Testing Refreshing Performance of Real-Time Data Warehouse
Jacek Jedrzejczak, Tomasz Koszlajda, Robert Wrembel
DEXA (2)3
2012 Time-HOBI: Index for optimizing star queries
Tadeusz Morzy, Robert Wrembel, Jan Chmiel, Artur Wojciechowski
Inf. Syst.2
2010 GPU-WAH: Applying GPUs to Compressing Bitmap Indexes with Word Aligned Hybrid
Witold Andrzejewski, Robert Wrembel
DEXA (2)2
2010 Time-HOBI: indexing dimension hierarchies by means of hierarchically organized bitmaps
abstract
One of the important research and technological problems in data warehousing is the optimization of star queries. So far, most of the research focused on optimizing such queries by means of join indexes and bitmap join indexes. In this paper we propose an index, called Time-HOBI, for optimizing star queries and computing aggregates along dimension hierarchies. Time-HOBI, created on a dimension hierarchy, is composed of: (1) a hierarchically organized bitmap index (HOBI), one bitmap index for one dimension level and (2) a time index (TI) that implicitly encodes time in every dimension. HOBI allows to quickly search for fact rows that fulfill selection criteria. With the support of TI joining a fact table with the Time dimension is avoided. Time-HOBI was implemented and evaluated experimentally on a real dataset, coming from the biggest East-European Internet auction platform Allegro.pl. The experiments show that Time-HOBI offers a promising star query performance.
Jan Chmiel, Tadeusz Morzy, Robert Wrembel
DOLAP3
2009 HOBI: Hierarchically Organized Bitmap Index for Indexing Dimensional Data
Jan Chmiel, Tadeusz Morzy, Robert Wrembel
DaWaK3
2009 RLH: Bitmap compression technique based on run-length and Huffman encoding
Michal Stabno, Robert Wrembel
Inf. Syst.2
2007 RLH: bitmap compression technique based on run-length and huffman encoding
abstract
In this paper we present a technique of compressing bitmap indexes for application in data warehouses. The developed compression technique, called Run-Length Huffman (RLH), is based on the run-length encoding and on the Huffman encoding. RLH was implemented and experimentally compared to the well known Word Aligned Hybrid bitmap compression technique that has been reported to provide the shortest query execution time. The experiments discussed in this paper show that RLH offers shorter query response times than WAH, for certain cardinalities of indexed attributes. Moreover, bitmaps compressed with RLH are smaller than corresponding bitmaps compressed with WAH. Additionally, we propose a modified RLH, called RLH-1024, which is designed to better support bitmap updates.
Michal Stabno, Robert Wrembel
DOLAP2
2006 Dynamic Method Materialization: A Framework for Optimizing Data Access Via Methods
Robert Wrembel, Mariusz Masewicz, Krzysztof Jankiewicz
DEXA1
2006 Managing and Querying Versions of Multiversion Data Warehouse
Robert Wrembel, Tadeusz Morzy
EDBT1
2004 On querying versions of multiversion data warehouse
abstract
A data warehouse (DW) is fed with data that come from external data sources that are production systems. External data sources, which are usually autonomous, often change not only their content but also their structure. The evolution of external data sources has to be reflected in a DW, that uses the sources. Traditional DW systems offer a limited support for handling dynamics in their structure and content. A promising approach to handling changes in DW structure and content is based on a multiversion data warehouse. In such a DW, each DW version describes a schema and data at certain period of time or a given business scenario, created for simulation purposes. In order to appropriately analyze multiversion data, an extension to a traditional SQL language is required. In this paper we propose an approach to querying a multiversion DW. To this end, we extended a SQL language and built a multiversion query language interface with functionality that allows: (1) expressing queries that address several DW versions and (2) presenting their results annotated with metadata information.
Tadeusz Morzy, Robert Wrembel
DOLAP2
2001 Hierarchical Materialisation of Methods in OO Views: Design, Maintenance, and Experimental Evaluation
abstract
The application of materialised object-oriented views in object-relational data warehousing systems is promising. In this paper we propose a novel technique for the materialisation of method results in object-oriented views, called hierarchical materialisation. When an object used to materialise the result of method m is updated, then m has to be recomputed. This recomputation can use unaffected intermediate materialised results of methods called from m, thus reducing a recomputation time. The hierarchical materialisation technique was implemented and evaluated by a number of experiments concerning methods without input arguments as well as methods with input arguments. The results showed that hierarchical materialisation reduces method recomputation time. Moreover, materialising methods with input arguments of narrow discrete domains introduces only a small time overhead.
Bartosz Bebel, Robert Wrembel
DOLAP2