VLDB 2026 Research / reviewers in the wild / expert
Manuel Vilares Ferro
dblp:v/MVilaresF
· DBLP profile ↗
35ranked-venue papers
21as first author
2since 2021 · last 2023
0000-0003-3414-6211ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 13 first-author · 1 since 2021Databases, data management, data science and information retrieval · 11 · 5 first-authorTheory of computation · 9 · 7 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Early stopping by correlating online indicators in neural networksabstractIn order to minimize the generalization error in neural networks, a novel technique to identify overfitting phenomena when training the learner is formally introduced. This enables support of a reliable and trustworthy early stopping condition, thus improving the predictive power of that type of modeling. Our proposal exploits the correlation over time in a collection of online indicators, namely characteristic functions for indicating if a set of hypotheses are met, associated with a range of independent stopping conditions built from a canary judgment to evaluate the presence of overfitting. That way, we provide a formal basis for decision making in terms of interrupting the learning process. As opposed to previous approaches focused on a single criterion, we take advantage of subsidiarities between independent assessments, thus seeking both a wider operating range and greater diagnostic reliability. With a view to illustrating the effectiveness of the halting condition described, we choose to work in the sphere of natural language processing, an operational continuum increasingly based on machine learning. As a case study, we focus on parser generation, one of the most demanding and complex tasks in the domain. The selection of cross-validation as a canary function enables an actual comparison with the most representative early stopping conditions based on overfitting identification, pointing to a promising start toward an optimal bias and variance control. Manuel Vilares Ferro, Yerai Doval Mosquera, Francisco J. Ribadas, Victor M. Darriba |
Neural Networks | 1 |
| 2022 | Absolute convergence and error thresholds in non-active adaptive samplingabstractIn order to minimize the generalization error in neural networks, a novel technique to identify overfitting phenomena when training the learner is formally introduced. This enables support of a reliable and trustworthy early stopping condition, thus improving the predictive power of that type of modeling. Our proposal exploits the correlation over time in a collection of online indicators, namely characteristic functions for indicating if a set of hypotheses are met, associated with a range of independent stopping conditions built from a canary judgment to evaluate the presence of overfitting. That way, we provide a formal basis for decision making in terms of interrupting the learning process. As opposed to previous approaches focused on a single criterion, we take advantage of subsidiarities between independent assessments, thus seeking both a wider operating range and greater diagnostic reliability. With a view to illustrating the effectiveness of the halting condition described, we choose to work in the sphere of natural language processing, an operational continuum increasingly based on machine learning. As a case study, we focus on parser generation, one of the most demanding and complex tasks in the domain. The selection of cross-validation as a canary function enables an actual comparison with the most representative early stopping conditions based on overfitting identification, pointing to a promising start toward an optimal bias and variance control. Manuel Vilares Ferro, Victor M. Darriba, Jesús Vilares |
J. Comput. Syst. Sci. | 1 |
| 2020 | Adaptive scheduling for adaptive sampling in pos taggers constructionabstractWe introduce an adaptive scheduling for adaptive sampling as a novel way of machine learning in the construction of part-of-speech taggers. The goal is to speed up the training on large data sets, without significant loss of performance with regard to an optimal configuration. In contrast to previous methods using a random, fixed or regularly rising spacing between the instances, ours analyzes the shape of the learning curve geometrically in conjunction with a functional model to increase or decrease it at any time. The algorithm proves to be formally correct regarding our working hypotheses. Namely, given a case, the following one is the nearest ensuring a net gain of learning ability from the former, it being possible to modulate the level of requirement for this condition. We also improve the robustness of sampling by paying greater attention to those regions of the training data base subject to a temporary inflation in performance, thus preventing the learning from stopping prematurely. The proposal has been evaluated on the basis of its reliability to identify the convergence of models, corroborating our expectations. While a concrete halting condition is used for testing, users can choose any condition whatsoever to suit their own specific needs. Manuel Vilares Ferro, Victor M. Darriba, Jesús Vilares |
Comput. Speech Lang. | 1 |
| 2018 | On the performance of phonetic algorithms in microtext normalization
Yerai Doval, Manuel Vilares Ferro, Jesús Vilares |
Expert Syst. Appl. | 2 |
| 2018 | Wikipedia-based hybrid document representation for textual news classification
Marcos Mouriño-García, Roberto Pérez-Rodríguez, Luis E. Anido-Rifón, Manuel Vilares Ferro |
Soft Comput. | 4 |
| 2017 | Modeling of learning curves with applications to POS taggingabstractAn algorithm to estimate the evolution of learning curves on the whole of a training data base, based on the results obtained from a portion and using a functional strategy, is introduced.We approximate iteratively the sought value at the desired time, independently of the learning technique used and once a point in the process, called prediction level, has been passed.The proposal proves to be formally correct with respect to our working hypotheses and includes a reliable proximity condition.This allows the user to fix a convergence threshold with respect to the accuracy finally achievable, which extends the concept of stopping criterion and seems to be effective even in the presence of distorting observations.Our aim is to evaluate the training effort, supporting decision making in order to reduce the need for both human and computational resources during the learning process.The proposal is of interest in at least three operational procedures.The first is the anticipation of accuracy gain, with the purpose of measuring how much work is needed to achieve a certain degree of performance.The second relates the comparison of efficiency between systems at training time, with the objective of completing this task only for the one that best suits our requirements.The prediction of accuracy is also a valuable item of information for customizing systems, since we can estimate in advance the impact of settings on both the performance and the development costs.Using the generation of part-of-speech taggers as an example application, the experimental results are consistent with our expectations. Manuel Vilares Ferro, Victor M. Darriba, Francisco J. Ribadas |
Comput. Speech Lang. | 1 |
| 2016 | On the feasibility of character n-grams pseudo-translation for Cross-Language Information Retrieval tasks
Jesús Vilares, Manuel Vilares Ferro, Miguel A. Alonso 0001, Michael P. Oakes |
Comput. Speech Lang. | 2 |
| 2016 | Studying the effect and treatment of misspelled queries in Cross-Language Information Retrieval
Jesús Vilares, Miguel A. Alonso 0001, Yerai Doval, Manuel Vilares Ferro |
Inf. Process. Manag. | 4 |
| 2015 | Towards a Multi-label Classification of Open Educational ResourcesabstractNowadays, there are a lot of online repositories containing thousands of very useful educational resources for the educational community. To take full advantage of these resources requires a simple, direct and effective access to those resources that are of interest, therefore, it is necessary that those resources are ordered or ranked based on some criteria? -- that is to say, they have to be classified. Classification is usually done manually by the resource provider, which directly implies a main problem: the time spent categorising resources. In this paper we propose a solution to this problem through the design and implementation of a multi-label classifier that enables the automatic classification of a set of educational resources in their most suitable category or categories, thus eliminating the need to manually perform this classification. Evaluation results show that the performance of the OER classifier is comparable to classification of a de-facto standard corpus: OHSUMED. Marcos Mouriño-García, Roberto Pérez-Rodríguez, Luis E. Anido-Rifón, Manuel Vilares Ferro |
ICALT | 4 |
| 2015 | Supporting knowledge discovery for biodiversity
Manuel Vilares Ferro, Milagros Fernández Gavilanes, Adrián Blanco González |
Data Knowl. Eng. | 1 |
| 2011 | Managing misspelled queries in IR applications
Jesús Vilares, Manuel Vilares Ferro, Juan Otero Pombo |
Inf. Process. Manag. | 2 |
| 2010 | Error-repair parsing schemata
Carlos Gómez-Rodríguez, Miguel A. Alonso 0001, Manuel Vilares Ferro |
Theor. Comput. Sci. | 3 |
| 2009 | A General Method for Transforming Standard Parsers into Error-Repair Parsers
Carlos Gómez-Rodríguez, Miguel A. Alonso 0001, Manuel Vilares Ferro |
CICLing | 3 |
| 2008 | Extraction of complex index terms in non-English IR: A shallow parsing based approach
Jesús Vilares, Miguel A. Alonso 0001, Manuel Vilares Ferro |
Inf. Process. Manag. | 3 |
| 2007 | Character N-Grams Translation in Cross-Language Information Retrieval
Jesús Vilares, Michael P. Oakes, Manuel Vilares Ferro |
NLDB | 3 |
| 2006 | Regional vs. Global Robust Spelling Correction
Manuel Vilares Ferro, Juan Otero Pombo, Victor M. Darriba |
CICLing | 1 |
| 2005 | Regional Versus Global Finite-State Error Repair
Manuel Vilares Ferro, Juan Otero Pombo, Jorge Graña Gil |
CICLing | 1 |
| 2005 | Robust Spelling Correction
Manuel Vilares Ferro, Juan Otero Pombo, Jesús Vilares |
CIAA | 1 |
| 2004 | Parsing Incomplete Sentences Revisited
Manuel Vilares Ferro, Victor M. Darriba, Jesús Vilares |
CICLing | 1 |
| 2004 | Phrase Similarity through the Edit Distance
Manuel Vilares Ferro, Francisco J. Ribadas, Jesús Vilares |
DEXA | 1 |
| 2004 | Morphological and Syntactic Processing for Text Retrieval
Jesús Vilares, Miguel A. Alonso 0001, Manuel Vilares Ferro |
DEXA | 3 |
| 2004 | On Asymptotic Finite-State Error Repair
Manuel Vilares Ferro, Juan Otero Pombo, Jorge Graña Gil |
SPIRE | 1 |
| 2004 | Regional Finite-State Error Repair
Manuel Vilares Ferro, Juan Otero Pombo, Jorge Graña Gil |
CIAA | 1 |
| 2004 | A formal frame for robust parsing
Manuel Vilares Ferro, Victor M. Darriba, Jesús Vilares, Francisco J. Ribadas |
Theor. Comput. Sci. | 1 |
| 2003 | Generation of Incremental Parsers
Manuel Vilares Ferro, Miguel A. Alonso 0001, Victor M. Darriba |
CICLing | 1 |
| 2003 | Robust Parsing Using Dynamic Programming
Manuel Vilares Ferro, Victor M. Darriba, Jesús Vilares, Leandro Rodríguez-Liñares |
CIAA | 1 |
| 2002 | Tabulation of Bidirectional Push Down Automata
Miguel A. Alonso 0001, Víctor J. Díaz, Manuel Vilares Ferro |
CIAA | 3 |
| 2002 | Searching for Asymptotic Error Repair
Manuel Vilares Ferro, Victor M. Darriba, Miguel A. Alonso 0001 |
CIAA | 1 |
| 2001 | Approximate VLDC Pattern Matching in Shared-Forest
Manuel Vilares Ferro, Francisco J. Ribadas, Victor M. Darriba |
CICLing | 1 |
| 2001 | Approximately Common Patterns in Shared-ForestsabstractWe present a proposal intended to demonstrate the applicability of tabulation techniques for detecting approximately common patterns when dealing with structures sharing some common parts. This sharing saves on the space needed to represent the structures and also on their later processing, by factorizing the filtering of substructure matching. As a consequence, preliminary experimental tests indicate a reduction of the running time. Manuel Vilares Ferro, Francisco J. Ribadas, Jorge Graña Gil |
CIKM | 1 |
| 2001 | Towards the Development of Heuristics for Automatic Query Expansion
Jesús Vilares, Manuel Vilares Ferro, Miguel A. Alonso 0001 |
DEXA | 2 |
| 2000 | Approximate Pattern Matching in Shared-Forest
Manuel Vilares Ferro, Francisco J. Ribadas, Victor M. Darriba |
DEXA | 1 |
| 2000 | Regional Least-Cost Error Repair
Manuel Vilares Ferro, Victor M. Darriba, Francisco J. Ribadas |
CIAA | 1 |
| 1999 | Tabular Algorithms for TAG Parsing
Miguel A. Alonso 0001, David Cabrero Souto, Éric Villemonte de la Clergerie, Manuel Vilares Ferro |
EACL | 4 |
| 1996 | Finite state morphology and formal verificationabstractThe full paper describes an environment for the generation of non-deterministic taggers, currently used for the development of a Spanish lexicon. In relation to previous approaches, our system includes the use of verification tools in order to assure the robustness of the generated taggers. A wide variety of user defined criteria can be applied for checking the exact properties of the system. Manuel Vilares Ferro, Jorge Graña Gil, Pilar Alvariño |
Nat. Lang. Eng. | 1 |