Pedro A. Szekely

dblp:73/4919 · DBLP profile ↗
← Back
35ranked-venue papers in the field
3as first author
9since 2021 · last 2022
0000-0002-4621-2266ORCID · verified

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 25 (3 first)Information Retrieval & Web Search · 6Data Mining & Knowledge Discovery · 4
YearPublicationVenuePosition
2022 A study of the quality of Wikidata
Kartik Shenoy, Filip Ilievski, Daniel Garijo, Daniel Schwabe 0001, Pedro A. Szekely
J. Web Semant.5
2021 AMPPERE: A Universal Abstract Machine for Privacy-Preserving Entity Resolution Evaluation
abstract
Entity resolution is the task of identifying records in different datasets that refer to the same entity in the real world. In sensitive domains (e.g. financial accounts, hospital health records), entity resolution must meet privacy requirements to avoid revealing sensitive information such as personal identifiable information to untrusted parties. Existing solutions are either too algorithmically-specific or come with an implicit trade-off between accuracy of the computation, privacy, and run-time efficiency. We propose AMMPERE, an abstract computation model for performing universal privacy-preserving entity resolution. AMMPERE offers abstractions that encapsulate multiple algorithmic and platform-agnostic approaches using variants of Jaccard similarity to perform private data matching and entity resolution. Specifically, we show that two parties can perform entity resolution over their data, without leaking sensitive information. We rigorously compare and analyze the feasibility, performance overhead and privacy-preserving properties of these approaches on the Sharemind multi-party computation (MPC) platform as well as on PALISADE, a lattice-based homomorphic encryption library. The AMMPERE system demonstrates the efficacy of privacy-preserving entity resolution for real-world data while providing a precise characterization of the induced cost of preventing information leakage.
Yixiang Yao, Tanmay Ghai, Srivatsan Ravi, Pedro A. Szekely
CIKM4
2021 CSKG: The CommonSense Knowledge Graph
Filip Ilievski, Pedro A. Szekely
ESWC2
2021 Generating Explainable Abstractions for Wikidata Entities
abstract
The large coverage and quality of the Wikidata knowledge graph make it suitable for usage in downstream applications, such as entity summarization, entity linking, and question answering. Yet, most retrieval and similarity-based methods for Wikidata make limited use of its semantics, and lose the link between the rich structure in Wikidata and the decision-making algorithm. In this paper, we investigate how to define abstractive representations (profiles) of Wikidata entities. We propose a scalable method that can produce profiles for Wikidata entities based on salient labels associated with their types. We represent the resulting profiles as a graph, and compute profile embeddings. Our empirical analysis shows that the profiles can capture similarity competitively to baselines, but excel in terms of explainability. On the task of neural entity linking in tables, the profiles outperform all baselines in terms of accuracy, whereas their human-readable representation clearly explains the source of improvement. We make our code and data available to facilitate novel use cases based on the Wikidata profiles.
Nicholas Klein, Filip Ilievski, Pedro A. Szekely
K-CAP3
2021 From Tables to Knowledge: Recent Advances in Table Understanding
abstract
A wealth of human knowledge is expressed in structured tables, across web pages, scientific articles, spreadsheets, and databases. This wealth of knowledge is mirrored by diversity in the vast number of layout structures, content types, formats, and surface forms used to express tables. Recent advances in representation learning and knowledge representation have made progress in exploiting structural regularities in tabular data to unlock this knowledge. In this tutorial, we provide a survey of these advances for a host of table understanding tasks, including table segmentation, semantic typing of cells, transforming tables to knowledge graphs, entity linking, and table retrieval tasks for question answering.
Jay Pujara, Pedro A. Szekely, Huan Sun 0001, Muhao Chen 0001
KDD2
2021 A Graph-Based Approach for Inferring Semantic Descriptions of Wikipedia Tables
Craig A. Knoblock, Pedro A. Szekely, Minh Pham 0004, Jay Pujara
ISWC3
2021 Retrieving Complex Tables with Multi-Granular Graph Representation Learning
abstract
The task of natural language table retrieval (NLTR) seeks to retrieve semantically relevant tables based on natural language queries. Existing learning systems for this task often treat tables as plain text based on the assumption that tables are structured as dataframes. However, tables can have complex layouts which indicate diverse dependencies between subtable structures, such as nested headers. As a result, queries may refer to different spans of relevant content that is distributed across these structures. Moreover, such systems fail to generalize to novel scenarios beyond those seen in the training set. Prior methods are still distant from a generalizable solution to the NLTR problem, as they fall short in handling complex table layouts or queries over multiple granularities. To address these issues, we propose Graph-based Table Retrieval (GTR), a generalizable NLTR framework with multi-granular graph representation learning. In our framework, a table is first converted into a tabular graph, with cell nodes, row nodes and column nodes to capture content at different granularities. Then the tabular graph is input to a Graph Transformer model that can capture both table cell content and the layout structures. To enhance the robustness and generalizability of the model, we further incorporate a self-supervised pre-training task based on graph-context matching. Experimental results on two benchmarks show that our method leads to significant improvements over the current state-of-the-art systems. Further experiments demonstrate promising performance of our method on cross-dataset generalization, and enhanced capability of handling complex tables and fulfilling diverse query intents.
Fei Wang 0060, Kexuan Sun 0002, Muhao Chen 0001, Jay Pujara, Pedro A. Szekely
SIGIR5
2021 TOMATE: A heuristic-based approach to extract data from HTML tables
Juan C. Roldán, Patricia Jiménez, Pedro A. Szekely, Rafael Corchuelo
Inf. Sci.3
2021 Learning cell embeddings for understanding table layouts
Majid Ghasemi-Gol, Jay Pujara, Pedro A. Szekely
Knowl. Inf. Syst.3
2020 KGTK: A Toolkit for Large Knowledge Graph Manipulation and Analysis
Filip Ilievski, Daniel Garijo, Hans Chalupsky, Naren Teja Divvala, Yixiang Yao, Craig Milo Rogers, Ronpeng Li, Daniel Schwabe 0001, Pedro A. Szekely
ISWC (2)11
2019 Tabular Cell Classification Using Pre-Trained Cell Embeddings
abstract
There is a large amount of data on the web in tabular form, such as excel sheets, CSVs, and web tables. Often, tabular data is meant for human consumption, using data layouts that are difficult for machines to interpret automatically. Previous work uses the stylistic features of tabular cells (e.g. font size, border type, background color) to classify tabular cells by their role in the data layout of the document (top attribute, data, metadata, etc.). In this paper, we propose a method to embed the semantic and contextual information about tabular cells in a low dimension cell embedding space. We then propose an RNN-based classification technique to use these cell vector representations, combining them with stylistic features introduced in previous work, in order to improve the performance of cell type classification in complex documents. We evaluate the performance of our system on three datasets containing documents with various data layouts, in two settings, in-domain, and cross-domain training. Our evaluation result shows that our proposed cell vector representations in combination with our RNN-based classification technique significantly improves cell type classification performance.
Majid Ghasemi-Gol, Jay Pujara, Pedro A. Szekely
ICDM3
2019 T2WML: Table To Wikidata Mapping Language
abstract
The web contains millions of useful spreadsheets and CSV files, but these files are difficult to use in applications because they use a wide variety of data layouts and terminology. We present Table To Wikidata Mapping Language (T2WML), a language that makes it easy to map and link arbitrary spreadsheets and CSV files to the Wikidata data model. The output of T2WML consists of Wikidata statements that can be loaded in the public Wikidata knowledge base or in a Wikidata clone repository, creating an augmented Wikidata knowledge graph that application developers can query using SPARQL.
Pedro A. Szekely, Daniel Garijo, Divij Bhatia, Yixiang Yao, Jay Pujara
K-CAP1
2019 Coupled Clustering of Time-Series and Networks
abstract
Motivated by the problem of human-trafficking, where it is often observed that criminal organizations are linked and behave similarly over time, we introduce the problem of Coupled Clustering of Time-series and their underlying Network. The goal is to find tightly connected subgroups of nodes that also have similar node-specific time series (temporal—not necessarily structural—behavior). We formulate the problem as a coupled matrix factorization for the time series, combined with regularization for network smoothness. We propose CCTN, and an incrementally-updated counterpart, CCTN-inc, which efficiently handles network updates. Extensive experiments show that CCTN is up to 4x more accurate than baselines that consider graph structure or time series alone, and CCTN-inc is up to 55x faster than CCTN. As an application, we explore an exclusive database with millions of online ads on human trafficking, and successfully deploy our technique to detect criminal organizations.
Linhong Zhu, Pedro A. Szekely, Aram Galstyan, Danai Koutra
SDM3
2019 Expert-Guided Entity Extraction using Expressive Rules
abstract
Knowledge Graph Construction (KGC) is an important problem that has many domain-specific applications, including semantic search and predictive analytics. As sophisticated KGC algorithms continue to be proposed, an important, neglected use case is to empower domain experts who do not have much technical background to construct high-fidelity, interpretable knowledge graphs. Such domain experts are a valuable source of input because of their (both formal and learned) knowledge of the domain. In this demonstration paper, we present a system that allows domain experts to construct knowledge graphs by writing sophisticated rule-based entity extractors with minimal training, using a GUI-based editor that offers a range of complex facilities.
Mayank Kejriwal, Runqi Shao, Pedro A. Szekely
SIGIR3
2018 Structured Event Entity Resolution in Humanitarian Domains
Mayank Kejriwal, Pedro A. Szekely
ISWC (1)4
2017 Neural Embeddings for Populated Geonames Locations
Mayank Kejriwal, Pedro A. Szekely
ISWC (2)2
2017 An Investigative Search Engine for the Human Trafficking Domain
Mayank Kejriwal, Pedro A. Szekely
ISWC (2)2
2017 Lessons Learned in Building Linked Data for the American Art Collaborative
Craig A. Knoblock, Pedro A. Szekely, Eleanor E. Fink, Duane Degler, David Newbury, Robert Sanderson, Kate Blanch, Sara Snyder, Nilay Chheda, Nimesh Jain, Ravi Raju Krishna, Nikhila Begur Sreekanth, Yixiang Yao
ISWC (2)2
2017 Information Extraction in Illicit Web Domains
abstract
Extracting useful entities and attribute values from illicit domains such as human trafficking is a challenging problem with the potential for widespread social impact. Such domains employ atypical language models, have 'long tails' and suffer from the problem of concept drift. In this paper, we propose a lightweight, feature-agnostic Information Extraction (IE) paradigm specifically designed for such domains. Our approach uses raw, unlabeled text from an initial corpus, and a few (12-120) seed annotations per domain-specific attribute, to learn robust IE models for unobserved pages and websites. Empirically, we demonstrate that our approach can outperform feature-centric Conditional Random Field baselines by over 18% F-Measure on five annotated sets of real-world human trafficking datasets in both low-supervision and high-supervision settings. We also show that our approach is demonstrably robust to concept drift, and can be efficiently bootstrapped even in a serial computing environment.
Mayank Kejriwal, Pedro A. Szekely
WWW2
2016 A Scalable Approach to Incrementally Building Knowledge Graphs
Gleb Gawriljuk, Andreas Harth, Craig A. Knoblock, Pedro A. Szekely
TPDL4
2016 Efficient Graph-Based Document Similarity
Christian Paul, Achim Rettinger, Aditya Mogadala, Craig A. Knoblock, Pedro A. Szekely
ESWC5
2016 Comparing Vocabulary Term Recommendations Using Association Rules and Learning to Rank: A User Study
Johann Schaible, Pedro A. Szekely, Ansgar Scherp
ESWC2
2016 Semantic Labeling: A Domain-Independent Approach
Minh Pham 0004, Suresh Alse, Craig A. Knoblock, Pedro A. Szekely
ISWC (1)4
2016 Leveraging Linked Data to Discover Semantic Relations Within Data Sources
Mohsen Taheriyan, Craig A. Knoblock, Pedro A. Szekely, José Luis Ambite
ISWC (1)3
2016 Unsupervised Entity Resolution on Multi-type Graphs
Linhong Zhu, Majid Ghasemi-Gol, Pedro A. Szekely, Aram Galstyan, Craig A. Knoblock
ISWC (1)3
2016 Learning the semantics of structured data sources
Mohsen Taheriyan, Craig A. Knoblock, Pedro A. Szekely, José Luis Ambite
J. Web Semant.3
2015 Assigning Semantic Labels to Data Sources
S. K. Ramnandan, Amol Mittal, Craig A. Knoblock, Pedro A. Szekely
ESWC4
2015 Building and Using a Knowledge Graph to Combat Human Trafficking
Pedro A. Szekely, Craig A. Knoblock, Jason Slepicka, Andrew Philpot, Chengye Yin, Dipsy Kapoor, Premkumar Natarajan, Daniel Marcu, Kevin Knight, David Stallard, Subessware S. Karunamoorthy, Rajagopal Bojanapalli, Steven Minton, Brian Amanatullah, Todd Hughes, Mike Tamayo, David Flynt, Rachel Artiss, Shih-Fu Chang, Tao Chen 0015, Gerald Hiebel, Lidia Silva Ferreira
ISWC (2)1
2013 Connecting the Smithsonian American Art Museum to the Linked Data Cloud
Pedro A. Szekely, Craig A. Knoblock, Xuming Zhu, Eleanor E. Fink, Rachel Allen, Georgina Goodlander
ESWC1
2013 A Graph-Based Approach to Learn Semantic Descriptions of Data Sources
Mohsen Taheriyan, Craig A. Knoblock, Pedro A. Szekely, José Luis Ambite
ISWC (1)3
2012 Semi-automatically Mapping Structured Sources into the Semantic Web
Craig A. Knoblock, Pedro A. Szekely, José Luis Ambite, Aman Goel, Kristina Lerman, Maria Muslea, Mohsen Taheriyan, Parag Mallick
ESWC2
2012 Rapidly Integrating Services into the Linked Data Cloud
Mohsen Taheriyan, Craig A. Knoblock, Pedro A. Szekely, José Luis Ambite
ISWC (1)3
2011 Mind Your Metadata: Exploiting Semantics for Configuration, Adaptation, and Provenance in Scientific Workflows
Yolanda Gil, Pedro A. Szekely, Sandra R. Villamizar, Thomas C. Harmon, Varun Ratnakar, Maria Muslea, Fabio Silva, Craig A. Knoblock
ISWC (2)2
2011 Building Mashups by Demonstration
abstract
The latest generation of WWW tools and services enables Web users to generate applications that combine content from multiple sources. This type of Web application is referred to as a mashup. Many of the tools for constructing mashups rely on a widget paradigm, where users must select, customize, and connect widgets to build the desired application. While this approach does not require programming, the users must still understand programming concepts to successfully create a mashup. As a result, they are put off by the time, effort, and expertise needed to build a mashup. In this article, we describe our programming-by-demonstration approach to building mashup by example. Instead of requiring a user to select and customize a set of widgets, the user simply demonstrates the integration task by example. Our approach addresses the problems of extracting data from Web sources, cleaning and modeling the extracted data, and integrating the data across sources. We implemented these ideas in a system called Karma, and evaluated Karma on a set of 23 users. The results show that, compared to other mashup construction tools, Karma allows more of the users to successfully build mashups and makes it possible to build these mashups significantly faster compared to using a widget-based approach.
Rattapoom Tuchinda, Craig A. Knoblock, Pedro A. Szekely
ACM Trans. Web3
2003 WebScripter: Grass-Roots Ontology Alignment via End-User Report Creation
Baoshi Yan, Martin R. Frank, Pedro A. Szekely, Robert Neches, Juan Lopez
ISWC3