VLDB 2026 Research / reviewers in the wild / expert
Victor M. Darriba
dblp:75/5475 · also Víctor Manuel Darriba Bilbao
· DBLP profile ↗
14ranked-venue papers
0as first author
2since 2021 · last 2023
0000-0001-5566-1699ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 1 since 2021Theory of computation · 5 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Early stopping by correlating online indicators in neural networksabstractIn order to minimize the generalization error in neural networks, a novel technique to identify overfitting phenomena when training the learner is formally introduced. This enables support of a reliable and trustworthy early stopping condition, thus improving the predictive power of that type of modeling. Our proposal exploits the correlation over time in a collection of online indicators, namely characteristic functions for indicating if a set of hypotheses are met, associated with a range of independent stopping conditions built from a canary judgment to evaluate the presence of overfitting. That way, we provide a formal basis for decision making in terms of interrupting the learning process. As opposed to previous approaches focused on a single criterion, we take advantage of subsidiarities between independent assessments, thus seeking both a wider operating range and greater diagnostic reliability. With a view to illustrating the effectiveness of the halting condition described, we choose to work in the sphere of natural language processing, an operational continuum increasingly based on machine learning. As a case study, we focus on parser generation, one of the most demanding and complex tasks in the domain. The selection of cross-validation as a canary function enables an actual comparison with the most representative early stopping conditions based on overfitting identification, pointing to a promising start toward an optimal bias and variance control. Manuel Vilares Ferro, Yerai Doval Mosquera, Francisco J. Ribadas, Victor M. Darriba |
Neural Networks | 4 |
| 2022 | Absolute convergence and error thresholds in non-active adaptive samplingabstractIn order to minimize the generalization error in neural networks, a novel technique to identify overfitting phenomena when training the learner is formally introduced. This enables support of a reliable and trustworthy early stopping condition, thus improving the predictive power of that type of modeling. Our proposal exploits the correlation over time in a collection of online indicators, namely characteristic functions for indicating if a set of hypotheses are met, associated with a range of independent stopping conditions built from a canary judgment to evaluate the presence of overfitting. That way, we provide a formal basis for decision making in terms of interrupting the learning process. As opposed to previous approaches focused on a single criterion, we take advantage of subsidiarities between independent assessments, thus seeking both a wider operating range and greater diagnostic reliability. With a view to illustrating the effectiveness of the halting condition described, we choose to work in the sphere of natural language processing, an operational continuum increasingly based on machine learning. As a case study, we focus on parser generation, one of the most demanding and complex tasks in the domain. The selection of cross-validation as a canary function enables an actual comparison with the most representative early stopping conditions based on overfitting identification, pointing to a promising start toward an optimal bias and variance control. Manuel Vilares Ferro, Victor M. Darriba, Jesús Vilares |
J. Comput. Syst. Sci. | 2 |
| 2020 | Adaptive scheduling for adaptive sampling in pos taggers constructionabstractWe introduce an adaptive scheduling for adaptive sampling as a novel way of machine learning in the construction of part-of-speech taggers. The goal is to speed up the training on large data sets, without significant loss of performance with regard to an optimal configuration. In contrast to previous methods using a random, fixed or regularly rising spacing between the instances, ours analyzes the shape of the learning curve geometrically in conjunction with a functional model to increase or decrease it at any time. The algorithm proves to be formally correct regarding our working hypotheses. Namely, given a case, the following one is the nearest ensuring a net gain of learning ability from the former, it being possible to modulate the level of requirement for this condition. We also improve the robustness of sampling by paying greater attention to those regions of the training data base subject to a temporary inflation in performance, thus preventing the learning from stopping prematurely. The proposal has been evaluated on the basis of its reliability to identify the convergence of models, corroborating our expectations. While a concrete halting condition is used for testing, users can choose any condition whatsoever to suit their own specific needs. Manuel Vilares Ferro, Victor M. Darriba, Jesús Vilares |
Comput. Speech Lang. | 2 |
| 2017 | Modeling of learning curves with applications to POS taggingabstractAn algorithm to estimate the evolution of learning curves on the whole of a training data base, based on the results obtained from a portion and using a functional strategy, is introduced.We approximate iteratively the sought value at the desired time, independently of the learning technique used and once a point in the process, called prediction level, has been passed.The proposal proves to be formally correct with respect to our working hypotheses and includes a reliable proximity condition.This allows the user to fix a convergence threshold with respect to the accuracy finally achievable, which extends the concept of stopping criterion and seems to be effective even in the presence of distorting observations.Our aim is to evaluate the training effort, supporting decision making in order to reduce the need for both human and computational resources during the learning process.The proposal is of interest in at least three operational procedures.The first is the anticipation of accuracy gain, with the purpose of measuring how much work is needed to achieve a certain degree of performance.The second relates the comparison of efficiency between systems at training time, with the objective of completing this task only for the one that best suits our requirements.The prediction of accuracy is also a valuable item of information for customizing systems, since we can estimate in advance the impact of settings on both the performance and the development costs.Using the generation of part-of-speech taggers as an example application, the experimental results are consistent with our expectations. Manuel Vilares Ferro, Victor M. Darriba, Francisco J. Ribadas |
Comput. Speech Lang. | 2 |
| 2015 | Undirected Dependency ParsingabstractDependency parsers, which are widely used in natural language processing tasks, employ a representation of syntax in which the structure of sentences is expressed in the form of directed links (dependencies) between their words. In this article, we introduce a new approach to transition‐based dependency parsing in which the parsing algorithm does not directly construct dependencies, but rather undirected links, which are then assigned a direction in a postprocessing step. We show that this alleviates error propagation, because undirected parsers do not need to observe the single‐head constraint, resulting in better accuracy. Undirected parsers can be obtained by transforming existing directed transition‐based parsers as long as they satisfy certain conditions. We apply this approach to obtain undirected variants of three different parsers (the Planar, 2‐Planar, and Covington algorithms) and perform experiments on several data sets from the CoNLL‐X shared tasks and on the Wall Street Journal portion of the Penn Treebank, showing that our approach is successful in reducing error propagation and produces improvements in parsing accuracy in most of the cases and achieving results competitive with state‐of‐the‐art transition‐based parsers. Carlos Gómez-Rodríguez, Daniel Fernández-González, Victor M. Darriba |
Comput. Intell. | 3 |
| 2006 | Regional vs. Global Robust Spelling Correction
Manuel Vilares Ferro, Juan Otero Pombo, Victor M. Darriba |
CICLing | 3 |
| 2004 | Parsing Incomplete Sentences Revisited
Manuel Vilares Ferro, Victor M. Darriba, Jesús Vilares |
CICLing | 2 |
| 2004 | A formal frame for robust parsing
Manuel Vilares Ferro, Victor M. Darriba, Jesús Vilares, Francisco J. Ribadas |
Theor. Comput. Sci. | 2 |
| 2003 | Generation of Incremental Parsers
Manuel Vilares Ferro, Miguel A. Alonso 0001, Victor M. Darriba |
CICLing | 3 |
| 2003 | Robust Parsing Using Dynamic Programming
Manuel Vilares Ferro, Victor M. Darriba, Jesús Vilares, Leandro Rodríguez-Liñares |
CIAA | 2 |
| 2002 | Searching for Asymptotic Error Repair
Manuel Vilares Ferro, Victor M. Darriba, Miguel A. Alonso 0001 |
CIAA | 2 |
| 2001 | Approximate VLDC Pattern Matching in Shared-Forest
Manuel Vilares Ferro, Francisco J. Ribadas, Victor M. Darriba |
CICLing | 3 |
| 2000 | Approximate Pattern Matching in Shared-Forest
Manuel Vilares Ferro, Francisco J. Ribadas, Victor M. Darriba |
DEXA | 3 |
| 2000 | Regional Least-Cost Error Repair
Manuel Vilares Ferro, Victor M. Darriba, Francisco J. Ribadas |
CIAA | 2 |