Patricia Jiménez

dblp:83/9015 · DBLP profile ↗
← Back
12ranked-venue papers
10as first author
7since 2021 · last 2022
0000-0001-5070-5904ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 6 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2022 On validating web information extraction proposals
Patricia Jiménez, Rafael Corchuelo
Expert Syst. Appl.1
2022 An Experimental Study of Neural Approaches to Multi-Hop Inference in Question Answering
abstract
Question answering aims at computing the answer to a question given a context with facts. Many proposals focus on questions whose answer is explicit in the context; lately, there has been an increasing interest in questions whose answer is not explicit and requires multi-hop inference to be computed. Our analysis of the literature reveals that there is a seminal proposal with increasingly complex follow-ups. Unfortunately, they were presented without an extensive study of their hyper-parameters, the experimental studies focused exclusively on English, and no statistical analysis to sustain the conclusions was ever performed. In this paper, we report on our experience devising a very simple neural approach to address the problem, on our extensive grid search over the space of hyper-parameters, on the results attained with English, Spanish, Hindi, and Portuguese, and sustain our conclusions with statistically sound analyses. Our findings prove that it is possible to beat many of the proposals in the literature with a very simple approach that was likely overlooked due to the difficulty to perform an extensive grid search, that the language does not have a statistically significant impact on the results, and that the empirical differences found among some existing proposals are not statistically significant.
Patricia Jiménez, Rafael Corchuelo
Int. J. Neural Syst.1
2022 On exploring data lakes by finding compact, isolated clusters
Patricia Jiménez, Juan C. Roldán, Rafael Corchuelo
Inf. Sci.1
2022 A hybrid quantum approach to leveraging data from HTML tables
Patricia Jiménez, Juan C. Roldán, Rafael Corchuelo
Knowl. Inf. Syst.1
2022 On the design of an advanced business rule engine
abstract
Abstract Business rules govern how well‐managed companies perform every day. They are expected to be written in natural language because they are devised by business people. That makes it difficult to translate them automatically into executable rules that can be integrated into typical information systems. Our industrial and academic research concludes that the current tools have a number of problems, namely: many of them do not pay any attention to the SBVR standard; some of them are not multi‐language; most of them cannot achieve perfect parsing accuracy; some of them are not domain agnostic; some of them use proprietary technologies; some of them do not produce executable rules; and none supports exploratory what‐if analyses natively. In this article, we present Tier‐Rules, which is a system that overcomes the previous problems. We also report on an industrial case study that helps illustrate it in practice.
Patricia Jiménez, Rafael Corchuelo
Softw. Pract. Exp.1
2021 A clustering approach to extract data from HTML tables
Patricia Jiménez, Juan C. Roldán, Rafael Corchuelo
Inf. Process. Manag.1
2021 TOMATE: A heuristic-based approach to extract data from HTML tables
Juan C. Roldán, Patricia Jiménez, Pedro A. Szekely, Rafael Corchuelo
Inf. Sci.2
2020 On extracting data from tables that are encoded using HTML
Juan C. Roldán, Patricia Jiménez, Rafael Corchuelo
Knowl. Based Syst.2
2020 On the synthesis of metadata tags for HTML files
abstract
Summary RDFa, JSON‐LD, Microdata, and Microformats allow to endow the data in HTML files with metadata tags that help software agents understand them. Unluckily, there are many HTML files that do not have any metadata tags, which has motivated many authors to work on proposals to synthesize them. But they have some problems: the authors either provide an overall picture of their designs without too many details on the techniques behind the scenes or focus on the techniques but do not describe the design of the software systems that support them; many of them cannot deal with data that are encoded using semistructured formats like forms, listings, or tables; and the few proposals that can work on tables can deal with horizontal listings only. In this article, we describe the design of a system that overcomes the previous limitations using a novel embedding approach that has proven to outperform four state‐of‐the‐art techniques on a repository with randomly selected HTML files from 40 different sites. According to our experimental analysis, our proposal can achieve an F1 score that outperforms the others by 10.14%; this difference was confirmed to be statistically significant at the standard confidence level.
Patricia Jiménez, Juan C. Roldán, Fernando O. Gallego, Rafael Corchuelo
Softw. Pract. Exp.1
2016 On learning web information extraction rules with TANGO
Patricia Jiménez, Rafael Corchuelo
Inf. Syst.1
2016 Roller: a novel approach to Web information extraction
Patricia Jiménez, Rafael Corchuelo
Knowl. Inf. Syst.1
2016 ARIEX: Automated ranking of information extractors
Patricia Jiménez, Rafael Corchuelo, Hassan A. Sleiman
Knowl. Based Syst.1