Jorge Martinez-Gil

dblp:67/2895 · DBLP profile ↗
← Back
33ranked-venue papers
24as first author
15since 2021 · last 2026
0000-0002-5730-7965ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 17 · 11 first-author · 7 since 2021Artificial intelligence and machine learning · 16 · 12 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 6 first-author · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Drift Adaptation as Supervision Routing Under Heterogeneous Costs
Jorge Martinez-Gil, Florian Bachinger, Rudolf Ramler, Francois Picard, Leïla Belmerhnia, Georgios P. Spathoulas
DEXA (2)1
2025 Framework to automatically determine the quality of open data catalogs
Jorge Martinez-Gil
Expert Syst. Appl.1
2025 Augmenting the Interpretability of GraphCodeBERT for Code Similarity Tasks
abstract
Assessing the degree of similarity of code fragments is crucial for ensuring software quality, but it remains challenging due to the need to capture the deeper semantic aspects of code. Traditional syntactic methods often fail to identify these connections. Recent advancements have addressed this challenge, though they frequently sacrifice interpretability. To improve this, we present an approach aiming to augment the transparency of the similarity assessment by using GraphCodeBERT, which enables the identification of semantic relationships between code fragments. This approach identifies similar code fragments and clarifies the reasons behind that identification, helping developers better understand and trust the results. The source code for our implementation is available at https://www.github.com/jorge-martinez-gil/graphcodebert-interpretability .
Jorge Martinez-Gil
Int. J. Softw. Eng. Knowl. Eng.1
2025 MTD-DS: An SLA-Aware Decision Support Benchmark for Multi-Tenant Parallel DBMSs
abstract
Multi-tenant DBMSs are used by cloud providers for their Database-as-a-Service products. They could be single-node DBMSs installed in virtual machines, SQL-on-Hadoop systems or classic parallel relational DBMSs running on top of a shared-nothing or shared-disk architecture. For a cloud provider, it is interesting to measure these systems’ capability of dealing with multi-tenant workloads, i.e., taking advantage of the statistical multiplexing to obtain economic gain while being attractive by providing a good quality of service and a low bill to the tenants. In this paper, we present MTD-DS benchmark (with MTD for Multi-Tenant parallel DBMSs and DS for Decision Support). MTD-DS extends TPC-DS by adding a multi-tenant query workload generator, a performance Service Level Objectives generator, configurable Database-as-a-Service pricing models, and new metrics to measure the potential capability of a multi-tenant parallel DBMS in obtaining the best trade-off between the provider's benefit and the tenants’ satisfaction. Example experimental results have been produced to show the relevance and the feasibility of the MTD-DS benchmark.
Shaoyi Yin, Franck Morvan, Jorge Martinez-Gil, Abdelkader Hameurlain
IEEE Trans. Knowl. Data Eng.3
2024 Optimizing readability using genetic algorithms
Jorge Martinez-Gil
Knowl. Based Syst.1
2023 Neurofuzzy semantic similarity measurement
Jorge Martinez-Gil, Riad Mokadem, Josef Küng, Abdelkader Hameurlain
Data Knowl. Eng.1
2023 A Comparative Study of Ensemble Techniques Based on Genetic Programming: A Case Study in Semantic Similarity Assessment
abstract
The challenge of assessing semantic similarity between pieces of text through computers has attracted considerable attention from industry and academia. New advances in neural computation have developed very sophisticated concepts, establishing a new state of the art in this respect. In this paper, we go one step further by proposing new techniques built on the existing methods. To do so, we bring to the table the stacking concept that has given such good results and propose a new architecture for ensemble learning based on genetic programming. As there are several possible variants, we compare them all and try to establish which one is the most appropriate to achieve successful results in this context. Analysis of the experiments indicates that Cartesian Genetic Programming seems to give better average results.
Jorge Martinez-Gil
Int. J. Softw. Eng. Knowl. Eng.1
2022 Multi-Cloud Query Optimisation with Accurate and Efficient Quoting
abstract
A recent trend among major organisations is to release their datasets in the cloud over various Database-as-a-Service (DBaaS) providers’ premises, creating a use case for multi-cloud querying. As identified in the literature, middlewares with such capabilities should quote the monetary cost and the response time of the queries in order to gain the trust of their users, and also optimise the queries so as to avoid cost overruns and meet the quotations. Considering those requirements, this paper introduces an accurate cost model and an efficient execution plan search strategy for dealing with large-scale multi-cloud queries. The former is an ensemble learning stack leveraging online machine learning models, and the latter is a randomised method inspired by iterative improvement. We evaluated our middleware over simulated providers by using the Join Order Benchmark. Experiments showed that the cost model manages to correct the estimations from the providers. The randomised strategy can produce more efficiently execution plans that yield better performances and a lower monetary cost compared to an exhaustive approach from previous work.
Damien T. Wojtowicz, Shaoyi Yin, Jorge Martinez-Gil, Franck Morvan, Abdelkader Hameurlain
IEEE Big Data3
2022 Interpretable ontology meta-matching in the biomedical domain using Mamdani fuzzy inference
Jorge Martinez-Gil, José Manuel Chaves-González
Expert Syst. Appl.1
2022 Special Issue on Machine Learning and Knowledge Graphs
Mehwish Alam, Anna Fensel, Jorge Martinez-Gil, Bernhard Moser 0001, Diego Reforgiato Recupero, Harald Sack
Future Gener. Comput. Syst.3
2022 A reproducible POI recommendation framework: Works mapping and benchmark evaluation
abstract
This work is a companion reproducibility paper that presents a framework to reproduce our previous experiments and results reported in Werneck et al. (2021). In that previous paper, we introduced a systematic mapping process of points-of-interest (POI) recommendation methods and provided a uniform evaluation methodology based on metrics covering different aspects besides accuracy. Due to the lack of reproducible and extensible benchmarks, our work introduces a reproducibility framework for POI methods based on a collection of Python software libraries and a Docker image. Our proposal is composed of: (1) a package to perform a protocol that reproduces our systematic mapping process Werneck et al. (2021), containing all collected data, insightful views on current advances and opened challenges; and (2) an extensible benchmark to perform a protocol to reproduce experimental evaluations on POI recommendation, considering different datasets, metrics, and the strongest baselines in the literature. This work also demonstrates all processes required to instantiate its framework. Moreover, our work can be considered at least weakly reproducible, since we were able to reproduce the results of the previous paper, leading us to the same conclusions.
Heitor Werneck, Nícollas Silva, Adriano C. M. Pereira, Matheus Carvalho Viana, Alejandro Bellogín, Jorge Martinez-Gil, Fernando Mourão, Leonardo Rocha 0001
Inf. Syst.6
2021 A Novel Neurofuzzy Approach for Semantic Similarity Measurement
Jorge Martinez-Gil, Riad Mokadem, Josef Küng, Abdelkader Hameurlain
DaWaK1
2021 Matching Large Biomedical Ontologies Using Symbolic Regression
abstract
The problem of ontology matching consists of finding the semantic correspondences between two ontologies that, although belonging to the same domain, have been developed separately. Matching methods are of great importance since they allow us to find the pivot points from which an automatic data integration process can be established. Unlike the most recent developments based on deep learning, this study presents our research on the development of new methods for ontology matching that are accurate and interpretable at the same time. For this purpose, we rely on a symbolic regression model specifically trained to find the mathematical expression that can solve the ground truth accurately, with the possibility of being understood by a human operator and forcing the processor to consume as little energy as possible. The experimental evaluation results show that our approach seems to be promising.
Jorge Martinez-Gil, Shaoyi Yin, Josef Küng, Franck Morvan
iiWAS1
2021 Indirect Mass Flow Estimation based on Power Measurements of Conveyor Belts in Mineral Processing Applications
abstract
This study presents our most recent advances in the design of a data-driven method for mass flow rate estimation of conveyor belts. Our proposal is focused on obtaining an indirect method that uses power measurement from the conveyor belt. The aim is to replace traditional expensive measurement hardware, which results in benefits such as lowering overall costs as well as the possibility of working in hostile environments such those with adverse weather conditions and the presence of dust and vibration. The mass estimation is based on data-driven estimations of idle power and net energy consumption. We discuss different models describing the relationship between energy input and transported mass: a constant proportionality factor, a time-dependent factor and a regression model depending on the idle power. We illustrate our approach on a case study where the state-dependent model yields the most promising results across multiple working periods.
Bernhard Heinzl, Jorge Martinez-Gil, Johannes Himmelbauer, Michael Rossbory, Christian Hinterdorfer, Christian Hinterreiter
INDIN2
2021 Semantic similarity controllers: On the trade-off between accuracy and interpretability
Jorge Martinez-Gil, José Manuel Chaves-González
Knowl. Based Syst.1
2020 A novel method based on symbolic regression for interpretable semantic similarity measurement
Jorge Martinez-Gil, José Manuel Chaves-González
Expert Syst. Appl.1
2019 Multiple Choice Question Answering in the Legal Domain Using Reinforced Co-occurrence
Jorge Martinez-Gil, Bernhard Freudenthaler, A Min Tjoa
DEXA (1)1
2019 Optimal Selection of Training Courses for Unemployed People based on Stable Marriage Model
abstract
The problem that we address here is given n job seekers and n job offers, where each job seeker has ranked all job offers in order of preference given by a suitability function, and vice versa; the goal is to compute the minimum set of skills to be offered to the job seekers, so that a) a global stable marriage between job seekers and potential employers can be reached, and b) the degree of satisfaction for that stable marriage might be maximum. To achieve this goal, we have designed an iterative algorithmic solution that can be solved in polynomial time. Additionally, we illustrate our solution with an use case based on a numerical example.
Jorge Martinez-Gil, Bernhard Freudenthaler
iiWAS1
2019 Automatic design of semantic similarity controllers based on fuzzy logics
Jorge Martinez-Gil, José Manuel Chaves-González
Expert Syst. Appl.1
2019 Semantic similarity aggregators for very short textual expressions: a case study on landmarks and points of interest
Jorge Martinez-Gil
J. Intell. Inf. Syst.1
2018 Accurate and efficient profile matching in knowledge bases
Jorge Martinez-Gil, Alejandra Lorena Paoletti, Gábor Rácz, Attila Sali, Klaus-Dieter Schewe
Data Knowl. Eng.1
2017 Management of accurate profile matching using multi-cloud service interaction
abstract
The current paper describes our research towards a cloud infrastructure for the universal access and interaction with a number of services implementing methods for enriching, matching and querying information about job offers and applicant profiles in the cloud. These methods exploit well-known recruitment knowledge bases in order to deliver valuable information to such organizations as public and private employment agencies that we assume to be geographically distributed. The rationale behind our approach is to offer an universal, yet inexpensive, distribution model able to reduce the cost of installing and maintaining the recruitment technology within the client's businesses.
Andreea Buga, Bernhard Freudenthaler, Jorge Martinez-Gil, Sorana Tania Nemes, Alejandra Lorena Paoletti
iiWAS3
2017 Automatic recommendation of prognosis measures for mechanical components based on massive text mining
abstract
Automatically providing suggestions for predicting the likely status of a mechanical component is a key challenge in a wide variety of industrial domains. Existing solutions based on ontological models have proven to be appropriate for fault diagnosis, but they fail when suggesting activities leading to a successful prognosis of mechanical components. The major reason is that fault prognosis is an activity that, unlike fault diagnosis, involves a lot of uncertainty and it is not always possible to envision a model for predicting possible faults. In this work, we propose a solution based on massive text mining for automatically suggesting prognosis activities concerning mechanical components. The great advantage of text mining is that it is possible to automatically analyze vast amounts of unstructured information in order to find strategies that have been successfully exploited, and formally or informally documented, in the past in any part of the world.
Jorge Martinez-Gil, Bernhard Freudenthaler, Thomas Natschläger
iiWAS1
2016 Top-k Matching Queries for Filter-Based Profile Matching in Knowledge Bases
Alejandra Lorena Paoletti, Jorge Martinez-Gil, Klaus-Dieter Schewe
DEXA (2)2
2016 Maintenance of Profile Matchings in Knowledge Bases
Jorge Martinez-Gil, Alejandra Lorena Paoletti, Gábor Rácz, Attila Sali, Klaus-Dieter Schewe
MEDI1
2016 Accurate Semantic Similarity Measurement of Biomedical Nomenclature by Means of Fuzzy Logic
abstract
Semantic similarity measurement of biomedical nomenclature aims to determine the likeness between two biomedical expressions that use different lexicographies for representing the same real biomedical concept. There are many semantic similarity measures for trying to address this issue, many of them have represented an incremental improvement over the previous ones. In this work, we present yet another incremental solution that is able to outperform existing approaches by using a sophisticated aggregation method based on fuzzy logic. Results show us that our strategy is able to consistently beat existing approaches when solving well-known biomedical benchmark data sets.
Jorge Martinez-Gil
Int. J. Uncertain. Fuzziness Knowl. Based Syst.1
2015 Extending Knowledge-Based Profile Matching in the Human Resources Domain
Alejandra Lorena Paoletti, Jorge Martinez-Gil, Klaus-Dieter Schewe
DEXA (2)2
2013 Evolutionary algorithm based on different semantic similarity functions for synonym recognition in the biomedical domain
José Manuel Chaves-González, Jorge Martinez-Gil
Knowl. Based Syst.2
2012 KnoE: A Web Mining Tool to Validate Previously Discovered Semantic Correspondences
Jorge Martinez-Gil, José Francisco Aldana-Montes
J. Comput. Sci. Technol.1
2011 Evaluation of two heuristic approaches to solve the ontology meta-matching problem
Jorge Martinez-Gil, José Francisco Aldana-Montes
Knowl. Inf. Syst.1
2010 Statistical Study about Existing OWL Ontologies from a Significant Sample as Previous Step for their Alignment
abstract
In this work, we present a proposal for characterizing the OWL ontologies available on the Web from a significant sample. We have conducted a study to review the specific characteristics of these ontologies paying attention to features which can be important from the point of view of the ontology alignment: language, sizes, number, and kind of entities that are represented in them. As a result, we offer some statistical data that can be helpful in order to understand the current situation of OWL ontologies in the Web and, therefore to guide the process of taking decisions when developing applications for aligning them.
Jorge Martinez-Gil, Enrique Alba 0001, José Francisco Aldana-Montes
CISIS1
2008 Comparison of Textual Renderings of Ontologies for Improving Their Alignment
abstract
This work is about an experiment in which we have compared the textual rendering of ontologies in order to get more accurate alignments between them. The experiments we have performed consist on three main steps: rendering in a textual way two ontologies, comparing the obtained text with several algorithms for text comparing and, using the obtained result as a factor to improve the alignments between them. As result, we got some evidences that this technique gives us a good measure of the similarity of ontologies and, therefore can allow us to improve the effectiveness of the alignment process.
Jorge Martinez-Gil, Ismael Navas-Delgado, Antonio Polo Márquez, José Francisco Aldana-Montes
CISIS1
2007 Thinking on the Web: Berners-Lee, Gödel and Turing
abstract
This book is a great compilation of the foundations of the World Wide Web. The authors review, in depth, most of the computation principles with a rigorous, nice and understandable style. The book is divided into two parts, the first one presents the web from a philosophical point of view and the second part offers us all technological issues which support it. The amazing jobs, which have been done by three geniuses: Berners-Lee, Gödel and Turing, are highlighted in the first part of the book. Their contributions to computer science allow us to have a firm, solid and mature theory which inspires the Information Revolution we have been living to. Semantic technologies are shown in the second part of the book, which is written so simply that it is the most understandable work related to this subject I have ever read. When people read something about Semantic Web they usually...
Jorge Martinez-Gil
Comput. J.1