Nikolaos Vasiloglou

dblp:96/4527 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
4since 2021 · last 2025
0000-0001-8499-7781ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 7 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Security and privacy · 3Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Scalable Tabular Dataset Labeling with Language Models
Yaojie Hu, Ilias Fountalis, Jin Tian 0001, Nikolaos Vasiloglou
SSDBM4
2024 CHORUS: Foundation Models for Unified Data Discovery and Exploration
abstract
We apply foundation models to data discovery and exploration tasks. Foundation models are large language models (LLMS) that show promising performance on a range of diverse tasks unrelated to their training. We show that these models are highly applicable to the data discovery and data exploration domain. When carefully used, they have superior capability on three representative tasks: table-class detection, column-type annotation and join-column prediction. On all three tasks, we show that a foundation-model-based approach outperforms the task-specific models and so the state of the art. Further, our approach often surpasses human-expert task performance. We investigate the fundamental characteristics of this approach including generalizability to several foundation models and the impact of non-determinism on the outputs. All in all, this suggests a future direction in which disparate data management tasks can be unified under foundation models.
Moe Kayali, Anton Lykov, Ilias Fountalis, Nikolaos Vasiloglou, Dan Olteanu, Dan Suciu
Proc. VLDB Endow.4
2022 Named Entity Recognition in Long Documents: An End-to-end Case Study in the Legal Domain
abstract
Named entity recognition (NER) is a fundamental task for several important applications such as knowledge base construction and semantic search. So far, the focus has been on building machine learning models, which identify generic named entities (e.g., person, date). Such models can be used off-the-shelf without requiring ground truth labels for training. However, such models cannot generalize to specialized domains that have domain-specific named entities (e.g., the legal domain). In these cases, it is inevitable to generate ground truth data and experiment with a variety of models in order to achieve good performance. Motivated by a real use case from the financial sector, we discuss the approach and lessons learned when solving the NER problem in the legal domain. This task is particularly challenging because it requires extensive human expertise to produce high quality ground-truth labels. For solving the legal domain NER problem, we first crawl a large dataset of legal documents and then introduce a semi-automated process to generate high-quality labels for a set of eleven predefined named entities. We validate that the proposed approach achieves high quality labels that outperform popular out-of-the-box NER methods. On top of that, our method once followed, can generate ground truth labels for the pre-defined named entities for an unbounded number of documents. Next, we experiment with a set of models and training procedures and report their performance on the NER task. Our experimental evaluation confirms that most of the models can generalize very well, achieving F1-score between 86% and 98.9%. The dataset, the labels produced by human annotators and our semi-supervised approach, as well as our code are made available to the research community.
Hossein Keshavarz, Zografoula Vagena, Pigi Kouki, Ilias Fountalis, Mehdi Mabrouki, Aziz Belaweid, Nikolaos Vasiloglou
IEEE Big Data7
2022 Simulating Complex Problems Inside a Database
abstract
The standard way to store and interact with the large amount of data that are central to the functioning of any modern business is through the use of a relational Knowledge Graph Management System (KGMS). In this paper we show how the relational model can be successfully exploited to model complex analytic scenarios while enjoying the same characteristics of clarity and flexibility as when modeling the data themselves. Using the Rel language, we simulate the daily schedule of an airline company as an agentbased system, and we will show how modeling this system through a set of relationships and logical rules will let us focus directly on the inherent complexity of our model, taking away most of the incidental effort in actually implementing our simulation.
Giancarlo Fissore, Nikolaos Vasiloglou
CIKM2
2020 From the lab to production: A case study of session-based recommendations in the home-improvement domain
abstract
E-commerce applications rely heavily on session-based recommendation algorithms to improve the shopping experience of their customers. Recent progress in session-based recommendation algorithms shows great promise. However, translating that promise to real-world outcomes is a challenging task for several reasons, but mostly due to the large number and varying characteristics of the available models. In this paper, we discuss the approach and lessons learned from the process of identifying and deploying a successful session-based recommendation algorithm for a leading e-commerce application in the home-improvement domain. To this end, we initially evaluate fourteen session-based recommendation algorithms in an offline setting using eight different popular evaluation metrics on three datasets. The results indicate that offline evaluation does not provide enough insight to make an informed decision since there is no clear winning method on all metrics. Additionally, we observe that standard offline evaluation metrics fall short for this application. Specifically, they reward an algorithm only when it predicts the exact same item that the user clicked next or eventually purchased. In a practical scenario, however, there are near-identical products which, although they are assigned different identifiers, they should be considered as equally-good recommendations. To overcome these limitations, we perform an additional round of evaluation, where human experts provide both objective and subjective feedback for the recommendations of five algorithms that performed the best in the offline evaluation. We find that the experts’ opinion is oftentimes different from the offline evaluation results. Analysis of the feedback confirms that the performance of all models is significantly higher when we evaluate near-identical product recommendations as relevant. Finally, we run an A/B test with one of the models that performed the best in the human evaluation phase. The treatment model increased conversion rate by 15.6% and revenue per visit by 18.5% when compared with a leading third-party solution.
Pigi Kouki, Ilias Fountalis, Nikolaos Vasiloglou, Xiquan Cui, Edo Liberty, Khalifeh Al Jadda
RecSys3
2019 Product collection recommendation in online retail
abstract
Recommender systems are an integral part of eCommerce services, helping to optimize revenue and user satisfaction. Bundle recommendation has recently gained attention by the research community since behavioral data supports that users often buy more than one product in a single transaction. In most cases, bundle recommendations are of the form "users who bought product A also bought products B, C, and D". Although such recommendations can be useful, there is no guarantee that products A, B, C, and D may actually be related to each other. In this paper, we address the problem of collection recommendation, i.e., recommending a collection of products that share a common theme and can potentially be purchased together in a single transaction. We extend on traditional approaches that use mostly transactional data by incorporating both domain knowledge from product suppliers in the form of hierarchies, as well as textual attributes from the products. Our approach starts by combining product hierarchies together with transactional data or domain knowledge to identify candidate sets of product collections. Then, it generates the product collection recommendations from these candidate sets by learning a deep similarity model that leverages textual attributes. Experimental evaluation on real data from the Home Depot online retailer shows that the proposed solution can recommend collections of products with increased accuracy when compared to expert-crafted collections.
Pigi Kouki, Ilias Fountalis, Nikolaos Vasiloglou, Nian Yan, Unaiza Ahsan, Khalifeh Al Jadda, Huiming Qu
RecSys3
2017 Practical Attacks Against Graph-based Clustering
abstract
Graph modeling allows numerous security problems to be tackled in a general way, however, little work has been done to understand their ability to withstand adversarial attacks. We design and evaluate two novel graph attacks against a state-of-the-art network-level, graph-based detection system. Our work highlights areas in adversarial machine learning that have not yet been addressed, specifically: graph-based clustering techniques, and a global feature space where realistic attackers without perfect knowledge must be accounted for (by the defenders) in order to be practical. Even though less informed attackers can evade graph clustering with low cost, we show that some practical defenses are possible.
Yizheng Chen 0001, Yacin Nadji, Athanasios Kountouras, Fabian Monrose, Roberto Perdisci, Manos Antonakakis, Nikolaos Vasiloglou
CCS7
2017 Multi-way Interacting Regression via Factorization Machines
abstract
We propose a Bayesian regression method that accounts for multi-way interactions of arbitrary orders among the predictor variables. Our model makes use of a factorization mechanism for representing the regression coefficients of interactions among the predictors, while the interaction selection is guided by a prior distribution on random hypergraphs, a construction which generalizes the Finite Feature Model. We present a posterior inference algorithm based on Gibbs sampling, and establish posterior consistency of our regression model. Our method is evaluated with extensive experiments on simulated data and demonstrated to be able to identify meaningful interactions in applications in genetics and retail demand forecasting.
Mikhail Yurochkin, XuanLong Nguyen, Nikolaos Vasiloglou
NIPS3
2017 ERBlox: Combining matching dependencies with machine learning for entity resolution
Zeinab Bahmani, Leo Bertossi, Nikolaos Vasiloglou
Int. J. Approx. Reason.3
2012 From Throw-Away Traffic to Bots: Detecting the Rise of DGA-Based Malware
Manos Antonakakis, Roberto Perdisci, Yacin Nadji, Nikolaos Vasiloglou, Saeed Abu-Nimeh, Wenke Lee, David Dagon
USENIX Security Symposium4
2011 Detecting Malware Domains at the Upper DNS Hierarchy
Manos Antonakakis, Roberto Perdisci, Wenke Lee, Nikolaos Vasiloglou, David Dagon
USENIX Security Symposium4
2009 Non-negative Matrix Factorization, Convexity and Isometry
abstract
In this paper we explore avenues for improving the reliability of dimensionality reduction methods such as Non-Negative Matrix Factorization (NMF) as interpretive exploratory data analysis tools. We first explore the difficulties of the optimization problem underlying NMF, showing for the first time that non-trivial NMF solutions always exist and that the optimization problem is actually convex, by using the theory of Completely Positive Factorization. We subsequently explore four novel approaches to finding globally-optimal NMF solutions using various ideas from convex optimization. We then develop a new method, isometric NMF (isoNMF), which preserves non-negativity while also providing an isometric embedding, simultaneously achieving two properties which are helpful for interpretation. Though it results in a more difficult optimization problem, we show experimentally that the resulting method is scalable and even achieves more compact spectra than standard NMF.
Nikolaos Vasiloglou, Alexander G. Gray, David V. Anderson
SDM1