Petar Ristoski

dblp:118/3466 · DBLP profile ↗
← Back
25ranked-venue papers
12as first author
4since 2021 · last 2022
0000-0002-1890-1507ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 23 · 11 first-author · 4 since 2021Artificial intelligence and machine learning · 9 · 5 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2022 Fifty Shades of Pink: Understanding Color in e-commerce using Knowledge Graphs
abstract
The color of the products is one of the most prevalent aspects in many e-commerce domains, and it is one of the decisive purchasing factors. Besides having thousands of color variations and shades, many brands continuously develop proprietary colors and color names to attract more customers. This often leads to color ambiguity (textual and visual), and vocabulary mismatch between buyers and sellers. Therefore, it is crucial for any e-commerce search engine to correctly identify the buyer's color intent and match it to the corresponding product listings. To address this challenge, in this work, we introduce a color query expansion approach using color Knowledge Graphs. We use Knowledge Graphs to unambiguously identify all the colors based on their properties, and the relationships to other colors, which allows us to perform semantic query expansion. Similar expansion concepts could be applied to domains outside of color.
Lizzie Liang, Sneha Kamath, Petar Ristoski, Qunzhi Zhou
CIKM3
2022 Shoe Size Resolution in Search Queries and Product Listings using Knowledge Graphs
abstract
The Fashion domain is one of the most profitable domains in most of the e-commerce shops, shoes being one of the top-selling categories within this domain. When shopping for shoes, one of the most important aspects for the buyers is the shoe size. Shoe size charts differ between different brands, geographical regions, genders and age groups. Not providing some of these details, as a buyer or a seller, could lead to a query intent to inventory mismatch and reduced or wrong search results. Furthermore, buying the wrong shoe size is one of the top reasons for product returns, which causes shipping delays and loss in revenue. To address this issue, we propose an approach for shoe size resolution and normalization in search queries and product listings using Knowledge Graphs.
Petar Ristoski, Aritra Mandal, Simon Becker, Anu Mandalam, Ethan Hart, Sanjika Hewavitharana, Qunzhi Zhou
CIKM1
2021 SAUCE: Truncated Sparse Document Signature Bit-Vectors for Fast Web-Scale Corpus Expansion
abstract
Recent advances in text representation have shown that training on large amounts of text is crucial for natural language understanding. However, models trained without predefined notions of topical interest typically require careful fine-tuning when transferred to specialized domains. When a sufficient amount of within-domain text may not be available, expanding a seed corpus of relevant documents from large-scale web data poses several challenges. First, corpus expansion requires scoring and ranking each document in the collection, an operation that can quickly become computationally expensive as the web corpora size grows. Relying on dense vector spaces and pairwise similarity adds to the computational expense. Secondly, as the domain concept becomes more nuanced, capturing the long tail of domain-specific rare terms becomes non-trivial, especially under limited seed corpora scenarios.
Muntasir Wahed, Daniel Gruhl, Alfredo Alba, Anna Lisa Gentile, Petar Ristoski, Chad DeLuca, Steve Welch, Ismini Lourentzou
CIKM5
2021 KG-ZESHEL: Knowledge Graph-Enhanced Zero-Shot Entity Linking
abstract
Entity linking is a fundamental task for a successful use of knowledge graphs in many information systems. It maps textual mentions to their corresponding entities in a given knowledge graph. However, with the rapid evolution of knowledge graphs, a large number of entities is continuously added over time. Performing entity linking on new, or unseen, entities poses a great challenge, as standard entity linking approaches require large amounts of labeled data for all new entities, and the underlying model must be regularly updated. To address this challenge, several zero-shot entity linking approaches have been proposed, which don't require additional labeled data to perform entity linking over unseen entities and new domains. Most of these approaches use large language models, such as BERT, to encode the textual description of the mentions and entities in a common embedding space, which allows linking mentions to unseen entities. While such approaches have shown good performance, one big drawback is that they are not able to exploit the entity symbolic information from the knowledge graph, such as entity types, relations, popularity scores and graph embeddings. In this paper, we present KG-ZESHEL, a knowledge graph-enhanced zero-shot entity linking approach, which extends an existing BERT-based zero-shot entity linking approach with mention and entity auxiliary information. Experiments on two benchmark entity linking datasets, show that our proposed approach outperforms the related BERT-based state-of-the-art entity linking models.
Petar Ristoski, Zhizhong Lin, Qunzhi Zhou
K-CAP1
2020 Expert-in-the-loop AI for Polymer Discovery
abstract
The use of AI in knowledge dense domains, e.g., chemistry, medicine, biology, etc. - is extremely promising, but often suffers from slow deployment and adaptation to different tasks. We propose a methodology to quickly capture the intent and expertise of a domain expert in order to train personalized AI models for specific tasks. Specifically we focus on the domain of polymer materials design and discovery: it often takes 10 years or more to design, synthesize, test, and introduce a new polymer material into the market. One way to accelerate up the design of polymer materials is through the use of computational methods to design the material, such as combinatorial screening, generative models, inverse design, etc. The drawback of these methods is that they generate a large number of candidates for new molecules, which then need to be manually reviewed by subject matter experts who select only a dozen for further investigation. Our solution is a human-in-the-loop methodology where we rank the candidates according to a utility function that is learned via the continued interaction with the subject matter experts, but which is also constrained by specific chemical knowledge. We prove the viability of our proposed methodology in a polymer production lab and we (i) evaluate against datasets of polymers previously produced in the lab as well as (ii) producing several novel materials that are undergoing experimental development, and (iii) quantitatively show that standard synthetic accessibility scores do not inform about patterns of SME decisions.
Petar Ristoski, Dmitry Zubarev, Anna Lisa Gentile, Nathaniel Park, Dan Sanders 0002, Daniel Gruhl, Linda Kato, Steve Welch
CIKM1
2020 GEval: A Modular and Extensible Evaluation Framework for Graph Embedding Techniques
Maria Angela Pellegrino, Abdulrahman Altabba, Martina Garofalo, Petar Ristoski, Michael Cochez
ESWC4
2020 Understanding Data Centers from Logs: Leveraging External Knowledge for Distant Supervision
Chad DeLuca, Anna Lisa Gentile, Petar Ristoski, Steve Welch
ISWC (2)3
2020 Large-scale relation extraction from web documents and knowledge graphs with human-in-the-loop
Petar Ristoski, Anna Lisa Gentile, Alfredo Alba, Daniel Gruhl, Steve Welch
J. Web Semant.1
2019 Explore and Exploit. Dictionary Expansion with Human-in-the-Loop
abstract
Many Knowledge Extraction systems rely on semantic resources - dictionaries, ontologies, lexical resources - to extract information from unstructured text. A key for successful information extraction is to consider such resources as evolving artifacts and keep them up-to-date. In this paper, we tackle the problem of dictionary expansion and we propose a human-in-the-loop approach: we couple neural language models with tight human supervision to assist the user in building and maintaining domain-specific dictionaries. The approach works on any given input text corpus and is based on the explore and exploit paradigm: starting from a few seeds (or an existing dictionary) it effectively discovers new instances (explore) from the text corpus as well as predicts new potential instances which are not in the corpus, i.e. “unseen” , using the current dictionary entries (exploit). We evaluate our approach on five real-world dictionaries, achieving high accuracy with a rapid expansion rate.
Anna Lisa Gentile, Daniel Gruhl, Petar Ristoski, Steve Welch
ESWC3
2019 Personalized Knowledge Graphs for the Pharmaceutical Domain
Anna Lisa Gentile, Daniel Gruhl, Petar Ristoski, Steve Welch
ISWC (2)3
2018 User-Centric Ontology Population
Kenneth L. Clarkson, Anna Lisa Gentile, Daniel Gruhl, Petar Ristoski, Joseph Terdiman, Steve Welch
ESWC4
2017 Entity Matching on Web Tables: a Table Embeddings approach for Blocking
abstract
Entity matching, or record linkage, is the task of identifying records that refer to the same entity.Naive entity matching techniques (i.e., brute-force pairwise comparisons) have quadratic complexity.A typical shortcut to the problem is to employ blocking techniques to reduce the number of comparisons, i.e. to partition the data in several blocks and only compare records within the same block.While classic blocking methods are designed for data from relational databases with clearly defined schemas, they are not applicable to data from Web tables, which are more prone to noise and do not come with an explicit schema.At the same time, Web tables are an interesting data source for many knowledge intensive tasks, which makes record linkage on Web Tables an important challenge.In this work, we propose an unsupervised approach to partition the data, that does not exploit any external knowledge, but only relies on heuristics to select the blocking attributes.We compare different partitioning methods: we use (i) clustering on bagof-words, (ii) binning via Locality-Sensitive Hashing and (iii) clustering using word embeddings.In particular, the clustering methods show good results on a standard dataset of Web Tables, and, when combined with word embeddings, are a robust solution which allows for computing the clusters in a dense, low-dimensional space.
Anna Lisa Gentile, Petar Ristoski, Steffen Eckel, Dominique Ritze, Heiko Paulheim
EDBT2
2017 Multi-lingual Concept Extraction with Linked Data and Human-in-the-Loop
abstract
Ontologies are dynamic artifacts that evolve both in structure and content. Keeping them up-to-date is a very expensive and critical operation for any application relying on semantic Web technologies. In this paper we focus on evolving the content of an ontology by extracting relevant instances of ontological concepts from text. We propose a novel technique which is (i) completely language independent, (ii) combines statistical methods with human-in-the-loop and (iii) exploits Linked Data as bootstrapping source. Our experiments on a publicly available medical corpus and on a Twitter dataset show that the proposed solution achieves comparable performances regardless of language, domain and style of text. Given that the method relies on a human-in-the-loop, our results can be safely fed directly back into Linked Data resources.
Alfredo Alba, Anni Coden, Anna Lisa Gentile, Daniel Gruhl, Petar Ristoski, Steve Welch
K-CAP5
2017 Global RDF Vector Space Embeddings
Michael Cochez, Petar Ristoski, Simone Paolo Ponzetto, Heiko Paulheim
ISWC (1)2
2017 Large-scale taxonomy induction using entity and word embeddings
abstract
Taxonomies are an important ingredient of knowledge organization, and serve as a backbone for more sophisticated knowledge representations in intelligent systems, such as formal ontologies. However, building taxonomies manually is a costly endeavor, and hence, automatic methods for taxonomy induction are a good alternative to build large-scale taxonomies. In this paper, we propose TIEmb, an approach for automatic unsupervised class subsumption axiom extraction from knowledge bases using entity and text embeddings. We apply the approach on the WebIsA database, a database of subsumption relations extracted from the large portion of the World Wide Web, to extract class hierarchies in the Person and Place domain.
Petar Ristoski, Stefano Faralli 0001, Simone Paolo Ponzetto, Heiko Paulheim
WI1
2017 Event-based knowledge reconciliation using frame embeddings and frame similarity
Mehwish Alam, Diego Reforgiato Recupero, Misael Mongiovì, Aldo Gangemi, Petar Ristoski
Knowl. Based Syst.5
2016 Enriching Product Ads with Metadata from HTML Annotations
Petar Ristoski, Peter Mika
ESWC1
2016 RDF2Vec: RDF Graph Embeddings for Data Mining
Petar Ristoski, Heiko Paulheim
ISWC (1)1
2016 A Collection of Benchmark Datasets for Systematic Evaluations of Machine Learning on the Semantic Web
Petar Ristoski, Gerben de Vries, Heiko Paulheim
ISWC (2)1
2016 Semantic Web in data mining and knowledge discovery: A comprehensive survey
Petar Ristoski, Heiko Paulheim
J. Web Semant.1
2015 Towards Linked Open Data Enabled Data Mining - Strategies for Feature Generation, Propositionalization, Selection, and Consolidation
Petar Ristoski
ESWC1
2015 Event-Based Clustering for Reducing Labeling Costs of Event-related Microposts
Axel Schulz 0001, Frederik Janssen, Petar Ristoski, Johannes Fürnkranz
ICWSM3
2015 The Mannheim Search Join Engine
Oliver Lehmberg, Dominique Ritze, Petar Ristoski, Robert Meusel, Heiko Paulheim, Christian Bizer
J. Web Semant.3
2015 Mining the Web of Linked Data with RapidMiner
Petar Ristoski, Christian Bizer, Heiko Paulheim
J. Web Semant.1
2014 Feature Selection in Hierarchical Feature Spaces
Petar Ristoski, Heiko Paulheim
Discovery Science1