EDBT 2026 Demo / reviewers in the wild / expert
Christos Tjortjis
dblp:65/6096
· DBLP profile ↗
11ranked-venue papers in the field
0as first author
8since 2021 · last 2026
0000-0001-8263-9024ORCID · corroborated
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 9Data Mining & Knowledge Discovery · 1Knowledge Engineering, Semantic Web & Information Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DynaHash: An efficient blocking structure for streaming record linkageabstractRecord linkage holds a crucial position in data management and analysis by identifying and merging records from disparate data sets that pertain to the same real-world entity. As data volumes grow, the intricacies of record linkage amplify, presenting challenges, such as potential redundancies and computational complexities. This paper introduces DynaHash, a novel randomized record linkage mechanism that utilizes (a) the MinHash technique to generate compact representations of blocking keys and (b) Hamming Locality-Sensitive Hashing (LSH) to construct the blocking structure from these vectors. By employing these methods, DynaHash offers theoretical guarantees of accuracy and achieves sublinear runtime complexities, with appropriate parameter tuning. It comprises two key components: a persistent storage system for permanently storing the blocking structure to ensure complete results, and an in-memory component for generating very fast partial results by summarizing the persisted blocking structure. Additionally, DynaHash leverages Multi-Probe matching to scan multiple neighboring blocks, in terms of their Hamming distances, in order to find matches. Our theoretical work derives a decrease factor in the space requirements, which depends on the Hamming threshold, compared with the baseline LSH. Our experimental evaluation against three state-of-the-art methods on six real-world data sets demonstrates DynaHash’s exceptional recall rates and query times, which are at least 2 × faster than its competitors and do not depend on the size of the underlying data sets. Dimitrios Karapiperis, Christos Tjortjis, Vassilios S. Verykios |
Inf. Syst. | 2 |
| 2025 | LSBlock: A Hybrid Blocking System Combining Lexical and Semantic Similarity Search for Record Linkage
Dimitrios Karapiperis, Christos Tjortjis, Vassilios S. Verykios |
ADBIS | 2 |
| 2025 | Adaptive sliding window normalization
George Papageorgiou 0005, Christos Tjortjis |
Inf. Syst. | 2 |
| 2024 | Evaluating the effectiveness of machine learning models for performance forecasting in basketball: a comparative studyabstractAbstract Sports analytics (SA) incorporate machine learning (ML) techniques and models for performance prediction. Researchers have previously evaluated ML models applied on a variety of basketball statistics. This paper aims to benchmark the forecasting performance of 14 ML models, based on 18 advanced basketball statistics and key performance indicators (KPIs). The models were applied on a filtered pool of 90 high-performance players. This study developed individual forecasting scenarios per player and experimented using all 14 models. The models’ performance ranking was developed using a bespoke evaluation metric, called weighted average percentage error (WAPE), formulated from the weighted mean absolute percentage error (MAPE) evaluation results of each forecasted statistic and model. Moreover, we employed a comprehensive forecasting approach to improve KPI's results. Results showed that Tree-based models, namely Extra Trees, Random Forest, and Decision Tree, are the best performers in most of the forecasted performance indicators, with the best performance achieved by Extra Trees with a WAPE of 34.14%. In conclusion, we achieved a 3.6% MAPE improvement for the selected KPI with our approach on unseen data. George Papageorgiou 0005, Vangelis Sarlis, Christos Tjortjis |
Knowl. Inf. Syst. | 3 |
| 2024 | A Suite of Efficient Randomized Algorithms for Streaming Record LinkageabstractOrganizations leverage massive volumes of information and new types of data to generate unprecedented insights and improve their outcomes. Correctly identifying duplicate records that represent the same entity, such as user, customer, patient and so on, a process commonly known as record linkage, can improve service levels, accelerate sales, or elevate healthcare decision support. Towards this direction, blocking methods are used with the aim to group matching records in the same block using a combination of their attributes as blocking keys. This paper introduces a suite of randomized algorithms specifically crafted for streaming record linkage settings. Using a bounded in-memory data structure, in terms of the number of blocks and positions within each block, our algorithms guarantee that the most frequently accessed and the most recently used blocks remain in main memory and, additionally, the records within a block are renewed on a rolling basis. The operation of our algorithms rely on simple random choices, instead of utilizing cumbersome sorting data structures, which ensure that the probability of inactive blocks and older records to remain in main memory decays in order to free space for more promising blocks and fresher records, respectively. We also introduce an algorithm that performs approximate blocking to tackle the problem of misspellings and typos present in the blocking keys. The experimental evaluation showcases that our proposed algorithms scale efficiently to data streams by providing certain accuracy guarantees. Dimitrios Karapiperis, Christos Tjortjis, Vassilios S. Verykios |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | A Randomized Blocking Structure for Streaming Record LinkageabstractA huge amount of data, in terms of streams, are collected nowadays via a variety of sources, such as sensors, mobile devices, or even raw log files. The unprecedented rate at which these data are generated and collected calls for novel record linkage methods to identify matching records pairs, which refer to the same real-world entity. Towards this direction, blocking methods are used in order to reduce the number of candidate record pairs while still maintaining high levels of accuracy. This paper introduces ExpBlock, a randomized record linkage structure, which guarantees that both the most frequently accessed and recently used blocks remain in main memory and, additionally, the records within a block are renewed on a rolling basis. Specifically, the probability of inactive blocks and older records to remain in main memory decays in order to make room for more promising blocks and fresher records, respectively. We implement these features using random choices instead of utilizing cumbersome sorting data structures in order to favour simplicity of implementation and efficiency. We showcase, through the experimental evaluation, that ExplBlock scales efficiently to data streams by providing accurate results in a timely fashion. Dimitrios Karapiperis, Christos Tjortjis, Vassilios S. Verykios |
Proc. VLDB Endow. | 2 |
| 2022 | Mining association rules from COVID-19 related twitter data to discover word patterns, topics and inferences
Paraskevas Koukaras, Christos Tjortjis, Dimitris Rousidis |
Inf. Syst. | 2 |
| 2021 | A Data Science approach analysing the Impact of Injuries on Basketball Player and Team Performance
Vangelis Sarlis, Vasilis Chatziilias, Christos Tjortjis, Dimitris Mandalidis |
Inf. Syst. | 3 |
| 2020 | Sports analytics - Evaluation of basketball players and team performance
Vangelis Sarlis, Christos Tjortjis |
Inf. Syst. | 2 |
| 2017 | ARMICA-Improved: A New Approach for Association Rule Mining
Shahpar Yakhchi, Seyed Mohssen Ghafari, Christos Tjortjis, Mahdi Fazeli |
KSEM | 3 |
| 2007 | An improved methodology on information distillation by mining program source code
Yiannis Kanellopoulos, Christos Makris 0001, Christos Tjortjis |
Data Knowl. Eng. | 3 |