VLDB 2026 Research / reviewers in the wild / expert
José-Miguel Benedí
dblp:04/2129 · also José-Miguel Benedí Ruíz
· DBLP profile ↗
44ranked-venue papers
3as first author
5since 2021 · last 2024
0000-0001-6516-2746ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 38 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13Databases, data management, data science and information retrieval · 5 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Speed-Up Pre-trained Vision Encoder-Decoder Transformers by Leveraging Lightweight Mixer Layers for Text Recognition
Daniel Parres, Dan Anitei, Roberto Paredes, Joan-Andreu Sánchez, José-Miguel Benedí |
DAS | 5 |
| 2024 | Improving Efficiency and Performance Through CTC-Based Transformers for Mathematical Expression Recognition
Dan Anitei, Daniel Parres, Joan-Andreu Sánchez, José-Miguel Benedí |
ICDAR (5) | 4 |
| 2023 | Discriminative estimation of probabilistic context-free grammars for mathematical expression recognition and retrievalabstractAbstract We present a discriminative learning algorithm for the probabilistic estimation of two-dimensional probabilistic context-free grammars (2D-PCFG) for mathematical expressions recognition and retrieval. This algorithm is based on a generalization of the H-criterion as the objective function and the growth transformations as the optimization method. For the development of the discriminative estimation algorithm, the N-best interpretations provided by the 2D-PCFG have been considered. Experimental results are reported on two available datasets: Im2Latex and IBEM. The first experiment compares the proposed discriminative estimation method with the classic Viterbi-based estimation method. The second one studies the performance of the estimated models depending on the length of the mathematical expressions and the number of admissible errors in the metric used. Ernesto Noya, José-Miguel Benedí, Joan-Andreu Sánchez, Dan Anitei |
Pattern Anal. Appl. | 2 |
| 2023 | The IBEM dataset: A large printed scientific image dataset for indexing and searching mathematical expressionsabstractSearching for information in printed scientific documents is a challenging problem that has recently received special attention from the Pattern Recognition research community. Mathematical expressions are complex elements that appear in scientific documents, and developing techniques for locating and recognizing them requires the preparation of datasets that can be used as benchmarks. Most current techniques for dealing with mathematical expressions are based on Machine Learning techniques which require a large amount of annotated data. These datasets must be prepared with ground-truth information for automatic training and testing. However, preparing large datasets with ground-truth is a very expensive and time-consuming task. This paper introduces the IBEM dataset, consisting of scientific documents that have been prepared for mathematical expression recognition and searching. This dataset consists of 600 documents, more than 8200 page images with more than 160000 mathematical expressions. It has been automatically generated from the version of the documents and can be enlarged easily. The ground-truth includes the position at the page level and the transcript for mathematical expressions both embedded in the text and displayed. This paper also reports a baseline classification experiment with mathematical symbols and a baseline experiment of Mathematical Expression Recognition performed on the IBEM dataset. These experiments aim to provide some benchmarks for comparison purposes so that future users of the IBEM dataset can have a baseline framework. Dan Anitei, Joan-Andreu Sánchez, José-Miguel Benedí, Ernesto Noya |
Pattern Recognit. Lett. | 3 |
| 2021 | ICDAR 2021 Competition on Mathematical Formula Detection
Dan Anitei, Joan-Andreu Sánchez, José Manuel Fuentes, Roberto Paredes, José-Miguel Benedí |
ICDAR (4) | 5 |
| 2020 | The Carabela Project and Manuscript Collection: Large-Scale Probabilistic Indexing and Content-based ClassificationabstractThe main aim of the Carabela project was to develop and apply techniques that allow textual searching on massive Spanish collections of 15th-19th century manuscripts. The project focused on a relatively small subset of 125 000 images of collections of interest to underwater archaeology. For this type of manuscripts, state-of-the-art automatic transcription techniques, generally fail to achieve usable transcription accuracy. Therefore, rather than insisting in actual transcription, methodologies for probabilistic indexing of handwritten text images have been adopted. This has allowed us to effectively cope with the intrinsically high degree of uncertainty of the text contained in most historical manuscripts, leading to highly effective systems for textual search and retrieval. Carabela has gone one step further by developing new techniques to classify probabilistically indexed, but otherwise untranscribed, text images according to their textual content. These techniques have been successfully used to automatically classify Carabela bundels (each containing hundreds or thousands of pages) according to their “level of risk” of public exposure, in order to control their access and avoid as much as possible the plundering of Spanish underwater heritage. Enrique Vidal 0001, Verónica Romero 0001, Alejandro H. Toselli, Joan-Andreu Sánchez, Vicente Bosch, Lorenzo Quirós, José-Miguel Benedí, José Ramón Prieto, Moisés Pastor, Francisco Casacuberta, Carlos Alonso, Carmen García, Lourdes Márquez, Carmen Orcero |
ICFHR | 7 |
| 2020 | Generation of Hypergraphs from the N-Best Parsing of 2D-Probabilistic Context-Free Grammars for Mathematical Expression RecognitionabstractWe consider hypergraphs as a tool obtained with bidimensional Probabilistic Context-Free Grammars to compactly represent the result of the n-best parse trees for an input image that represents a mathematical expression. More specifically, in this paper we propose: i) an algorithm to compute the N-best parse trees from a 2D-PCFGs, ii) an algorithm to represent the n-best parse trees using a compact representation in the form of hypergraphs, and iii) a formal framework for the development of inference algorithms (inside and outside) and normalization strategies of hypergraphs. Ernesto Noya, Joan-Andreu Sánchez, José-Miguel Benedí |
ICPR | 3 |
| 2016 | Beyond Prefix-Based Interactive Translation PredictionabstractCurrent automatic machine translation systems require heavy human proofreading to produce high-quality translations.We present a new interactive machine translation approach aimed at providing a natural collaboration between humans and translation systems.As such, we grant the user complete freedom to validate and correct any part of the translations suggested by the system.Our approach is then designed according to the requirements placed by this unrestricted proofreading protocol.In particular, the ability of the system to suggest new translations coherent with the set of potentially disjoint translation segments validated by the user.We evaluate our approach in a usersimulated setting where reference translations are considered the output desired by a human expert.Results show important reductions in the number of edits in comparison to decoupled post-editing and conventional prefix-based interactive translation prediction.Additionally, we provide evidence that it can also reduce the cognitive overload reported for interactive translation systems in previous user studies. Jesús González-Rubio, Daniel Ortiz-Martínez, Francisco Casacuberta, José-Miguel Benedí |
CoNLL | 4 |
| 2016 | An integrated grammar-based approach for mathematical expression recognition
Francisco Alvaro, Joan-Andreu Sánchez, José-Miguel Benedí |
Pattern Recognit. | 3 |
| 2015 | Structure detection and segmentation of documents using 2D stochastic context-free grammars
Francisco Alvaro, Francisco Cruz 0003, Joan-Andreu Sánchez, Oriol Ramos Terrades, José-Miguel Benedí |
Neurocomputing | 5 |
| 2015 | Unsegmented Dialogue Act Annotation and Decoding With N-Gram TransducersabstractMost studies on dialogue corpora, as well as most dialogue systems, employ dialogue acts as the basic units for interpreting discourse structure, user input and system actions. The definition of the discourse structure and the dialogue strategy consequently require the tagging of dialogue corpora in terms of dialogue acts. The tagging problem presents two basic variants: a batch variant (annotation of whole dialogues, in order to define dialogue strategy or study discourse structure) and an online variant (decoding of the dialogue act sequence of a given turn, in order to interpret user intentions). In the two variants is unusual having the segmentation of each turn into the dialogue meaningful units (segments) to which a dialogue act is assigned. In this paper we present the use of the N-Gram Transducer technique for tagging dialogues, without needing to provide a prior segmentation, in these two different variants (dialogue annotation and turn decoding). Experiments were performed in two corpora of different nature and results show that N-Gram Transducer models are suitable for these tasks and provide good performance. Carlos D. Martínez-Hinarejos, José-Miguel Benedí, Vicent Tamarit |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Offline Features for Classifying Handwritten Math Symbols with Recurrent Neural NetworksabstractIn mathematical expression recognition, symbol classification is a crucial step. Numerous approaches for recognizing handwritten math symbols have been published, but most of them are either an online approach or a hybrid approach. There is an absence of a study focused on offline features for handwritten math symbol recognition. Furthermore, many papers provide results difficult to compare. In this paper we assess the performance of several well-known offline features for this task. We also test a novel set of features based on polar histograms and the vertical repositioning method for feature extraction. Finally, we report and analyze the results of several experiments using recurrent neural networks on a large public database of online handwritten math expressions. The combination of online and offline features significantly improved the recognition rate. Francisco Alvaro, Joan-Andreu Sánchez, José-Miguel Benedí |
ICPR | 3 |
| 2014 | Recognition of on-line handwritten mathematical expressions using 2D stochastic context-free grammars and hidden Markov models
Francisco Alvaro, Joan-Andreu Sánchez, José-Miguel Benedí |
Pattern Recognit. Lett. | 3 |
| 2013 | Interactive Machine Translation using Hierarchical Translation ModelsabstractCurrent automatic machine translation systems are not able to generate error-free translations and human intervention is often required to correct their output.Alternatively, an interactive framework that integrates the human knowledge into the translation process has been presented in previous works.Here, we describe a new interactive machine translation approach that is able to work with phrase-based and hierarchical translation models, and integrates error-correction all in a unified statistical framework.In our experiments, our approach outperforms previous interactive translation systems, and achieves estimated effort reductions of as much as 48% relative over a traditional post-edition system. Jesús González-Rubio, Daniel Ortiz-Martínez, José-Miguel Benedí, Francisco Casacuberta |
EMNLP | 3 |
| 2013 | Classification of On-Line Mathematical Symbols with Hybrid Features and Recurrent Neural NetworksabstractRecognition of on-line handwritten mathematical symbols has been tackled using different methods, but the recognition rates achieved until now still leave room for improvement. Many of the published approaches are based on hidden Markov models, and some of them use off-line information extracted from the on-line data. In this paper, we present a set of hybrid features that combine both on-line and off-line information. Lately, recurrent neural networks have demonstrated to obtain good results and they have outperformed hidden Markov models in several sequence learning tasks, including handwritten text recognition. Hence, we also studied a state-of-the-art recurrent neural network classifier and we compared its performance with a classifier based on hidden Markov models. Experiments using a large public database showed that both the new proposed features and recurrent neural network classifier improved significantly the classification results. Francisco Alvaro, Joan-Andreu Sánchez, José-Miguel Benedí |
ICDAR | 3 |
| 2013 | Evaluating spoken dialogue models under the interactive pattern recognition frameworkabstract[EN] The new Interactive Pattern Recognition (IPR) framework has \nbeen proposed to deal with human-machine interaction. In this \ncontext a new formulation has been recently defined to represent \na Spoken Dialogue System as an IPR problem. In this work this \nformulation is applied to define graphical models that deal with \nSpoken Dialogue Systems. The definition of both a Dialogue \nManager and a User Model are shown and the estimation of \nthe parameters and smoothing techniques are presented in the \npaper. These models were evaluated in a dialogue generation \ntask on two very different corpora: Dihana corpus consisting \nof Spanish spoken dialogues acquired with the Wizard of Oz \ntechnique and Let’s Go corpus consisting of spoken dialogues \nin English between real users and the Ravenclaw dialogue manager \ndeveloped by CMU. The results obtained show that original \nand simulated dialogues exhibited very similar behaviours, \nthus demonstrating the learning capacity of the proposed models \nin both a controlled Wizard of Oz task and a spoken dialogue \nsystem that interacts with real users. This formulation can then \nbe considered as a promising framework to deal with Spoken \nDialogue Systems. Fabrizio Ghigi, M. Inés Torres, Raquel Justo, José-Miguel Benedí |
INTERSPEECH | 4 |
| 2012 | Unbiased Evaluation of Handwritten Mathematical Expression RecognitionabstractSeveral approaches have been proposed to tackle the problem of mathematical expression recognition, and automatic methods for performance evaluation are required. Mathematical expressions are usually encoded as a LaTeX string or a tree (MathML) for evaluation purpose, but these formats do not enforce uniqueness. Consequently, given that there can be several representations syntactically different but semantically equivalent, the automatic performance evaluation of mathematical expressions can be biased. Given a mathematical expression recognition tree and its ground-truth tree, the error is usually computed by comparing them. In this paper we propose to obtain a new tree, equivalent to the ground-truth tree, according to the model representation criteria. Then, we can compute an error by comparing the recognized tree with the obtained by using the model, both with the same bias. Several experiments were carried out in order to evaluate this approach and results showed that representation criteria had a significative effect in the evaluation results. Francisco Alvaro, Joan-Andreu Sánchez, José-Miguel Benedí |
ICFHR | 3 |
| 2012 | Estimating the number of segments for improving dialogue act labellingabstractAbstract In dialogue systems it is important to label the dialogue turns with dialogue-related meaning. Each turn is usually divided into segments and these segments are labelled with dialogue acts (DAs). A DA is a representation of the functional role of the segment. Each segment is labelled with one DA, representing its role in the ongoing discourse. The sequence of DAs given a dialogue turn is used by the dialogue manager to understand the turn. Probabilistic models that perform DA labelling can be used on segmented or unsegmented turns. The last option is more likely for a practical dialogue system, but it provides poorer results. In that case, a hypothesis for the number of segments can be provided to improve the results. We propose some methods to estimate the probability of the number of segments based on the transcription of the turn. The new labelling model includes the estimation of the probability of the number of segments in the turn. We tested this new approach with two different dialogue corpora:SwitchBoardandDihana. The results show that this inclusion significantly improves the labelling accuracy. Vicent Tamarit, Carlos D. Martínez-Hinarejos, José-Miguel Benedí |
Nat. Lang. Eng. | 3 |
| 2011 | Recognition of Printed Mathematical Expressions Using Two-Dimensional Stochastic Context-Free GrammarsabstractIn this work, a system for recognition of printed mathematical expressions has been developed. Hence, a statistical framework based on two-dimensional stochastic context-free grammars has been defined. This formal framework allows to jointly tackle the segmentation, symbol recognition and structural analysis of a mathematical expression by computing its most probable parsing. In order to test this approach a reproducible and comparable experiment has been carried out over a large publicly available (InftyCDB-1) database. Results are reported using a well-defined global dissimilitude measure. Experimental results show that this technique is able to properly recognize mathematical expressions, and that the structural information improves the symbol recognition step. Francisco Alvaro, Joan-Andreu Sánchez, José-Miguel Benedí |
ICDAR | 3 |
| 2010 | Dialogue act tagging and segmentation with a single perceptron
Ramón Granell, Stephen G. Pulman, Carlos D. Martínez-Hinarejos, José-Miguel Benedí |
INTERSPEECH | 4 |
| 2010 | Evaluation of HMM-based Models for the Annotation of Unsegmented Dialogue Turns
Carlos D. Martínez-Hinarejos, Vicent Tamarit, José-Miguel Benedí |
LREC | 3 |
| 2010 | Enlarged Search Space for SITG Parsing
Guillem Gascó i Mora, Joan-Andreu Sánchez, José-Miguel Benedí |
HLT-NAACL | 3 |
| 2009 | Reducing the Plagiarism Detection Search Space on the Basis of the Kullback-Leibler Distance
Alberto Barrón-Cedeño, Paolo Rosso, José-Miguel Benedí |
CICLing | 3 |
| 2009 | Improving Unsegmented Dialogue Turns Annotation with N-gram Transducers
Carlos D. Martínez-Hinarejos, Vicent Tamarit, José-Miguel Benedí |
PACLIC | 3 |
| 2008 | Statistical framework for a Spanish spoken dialogue corpus
Carlos D. Martínez-Hinarejos, José-Miguel Benedí, Ramón Granell |
Speech Commun. | 2 |
| 2007 | ANERsys: An Arabic Named Entity Recognition System Based on Maximum Entropy
Yassine Benajiba, Paolo Rosso, José-Miguel Benedí |
CICLing | 3 |
| 2007 | Clustering Narrow-Domain Short Texts by Using the Kullback-Leibler Distance
David Pinto 0001, José-Miguel Benedí, Paolo Rosso |
CICLing | 2 |
| 2006 | Segmented and Unsegmented Dialogue-Act Annotation with Statistical Dialogue Models
Carlos D. Martínez-Hinarejos, Ramón Granell, José-Miguel Benedí |
ACL | 3 |
| 2006 | Obtaining Word Phrases with Stochastic Inversion Translation Grammars for Phrase-based Statistical Machine Translation
Joan-Andreu Sánchez, José-Miguel Benedí |
EAMT | 2 |
| 2006 | Design and acquisition of a telephone spontaneous speech dialogue corpus in Spanish: DIHANA
José-Miguel Benedí, Eduardo Lleida, Amparo Varona, María José Castro Bleda, Isabel Galiano, Raquel Justo, Iñigo López de Letona, Antonio Miguel |
LREC | 1 |
| 2005 | Estimation of stochastic context-free grammars and their use as language models
José-Miguel Benedí, Joan-Andreu Sánchez |
Comput. Speech Lang. | 1 |
| 2004 | A hybrid language model based on a combination of N-grams and stochastic context-free grammarsabstractIn this paper, a hybrid language model is defined as a combination of a word-based n-gram, which is used to capture the local relations between words, and a category-based stochastic context-free grammar (SCFG) with a word distribution into categories, which is defined to represent the long-term relations between these categories. The problem of unsupervised learning of a SCFG in General Format and in Chomsky Normal Form by means of estimation algorithms is studied. Moreover, a bracketed version of the classical estimation algorithm based on the Earley algorithm is proposed. This paper also explores the use of SCFGs obtained from a treebank corpus as initial models for the estimation algorithms. Experiments on the UPenn Treebank corpus are reported. These experiments have been carried out in terms of the test set perplexity and the word error rate in a speech recognition experiment. Diego Linares, José-Miguel Benedí, Joan-Andreu Sánchez |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2002 | RNA Modeling by Combining Stochastic Context-Free Grammars and n-Gram ModelsabstractThe RNA sentences present structured regions caused by pairwise correlations, and nonstructured regions where any global relation can be found. In this paper, we present a combination of stochastic context-free grammars (SCFG) and bigram models. The SCFGs are used to represent the long-term relations of the structured part of RNA sequences, while the bigram models are used to capture the local relations of the nonstructured part. A stochastic version of Sakakibara's algorithm is used to study the SCFGs. Finally, experiments to evaluate the behavior of this proposal were carried out. Ismael Salvador, José-Miguel Benedí |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2001 | Improvement of a Whole Sentence Maximum Entropy Language Model Using Grammatical FeaturesabstractIn this paper, we propose adding long-term grammatical information in a Whole Sentence Maximun Entropy Language Model (WSME) in order to improve the performance of the model. The grammatical information was added to the WSME model as features and were obtained from a Stochastic Context-Free grammar. Finally, experiments using a part of the Penn Treebank corpus were carried out and significant improvements were acheived. Fredy A. Amaya, José-Miguel Benedí |
ACL | 2 |
| 2001 | Language Simplification through Error-Correcting and Grammatical Inference Techniques
Juan-Carlos Amengual, Alberto Sanchís, Enrique Vidal 0001, José-Miguel Benedí |
Mach. Learn. | 4 |
| 2000 | Combination Of N-Grams And Stochastic Context-Free Grammars For Language Modeling
José-Miguel Benedí, Joan-Andreu Sánchez |
COLING | 1 |
| 2000 | The EuTrans Spoken Language Translation System
Juan-Carlos Amengual, M. Asunción Castaño, Antonio Castellanos, Víctor M. Jiménez, David Llorens, Andrés Marzal, Federico Prat, Juan Miguel Vilar, José-Miguel Benedí, Francisco Casacuberta, Moisés Pastor, Enrique Vidal 0001 |
Mach. Transl. | 9 |
| 1999 | Learning of stochastic context-free grammars by means of estimation algorithms
Joan-Andreu Sánchez, José-Miguel Benedí |
EUROSPEECH | 2 |
| 1998 | Estimation of the probability distributions of stochastic context-free grammars from the k-best derivationsabstractThe use of the Inside-Outside (IO) algorithm for the estimation of the probability distributions of Stochastic Context-Free Grammars (SCFGs) in Natural-Language processing is restricted due to the time complexity per iteration and the large number of iterations that it needs to converge. Alternatively, an algorithm based on the Viterbi score (VS) is used. This VS algorithm converges more rapidly, but obtains less competitive models. We describe here a new algorithm that only considers the k-best derivations in the estimation process. The experimental results show that this algorithm achieves faster convergence than the IO and better models than the VS algorithm. Joan-Andreu Sánchez, José-Miguel Benedí |
ICSLP | 2 |
| 1997 | Speech translation based on automatically trainable finite-state modelsabstractThis paper extends previous work exploring the use of Subsequential Transducers to perform speech-input translation in limited-domain tasks. This is done following an integrated approach in which a Subsequential Transducer replaces the input-language model of a conventional speech recognition system, and is used both as language and translation model. This way, the search for the recognised sentence also produces the corresponding translation. A corpus-based approach is adopted in order to build the required models from training data. Experimental results are presented for the translation task considered in the EUTRANS project: one in the hotel domain with more than 500 words per language and language perplexities near to 10. Juan-Carlos Amengual, José-Miguel Benedí, Klaus Beulen, Francisco Casacuberta, M. Asunción Castaño, Antonio Castellanos, Víctor M. Jiménez, David Llorens, Andrés Marzal, Hermann Ney, Federico Prat, Enrique Vidal 0001, Juan Miguel Vilar |
EUROSPEECH | 2 |
| 1997 | Consistency of Stochastic Context-Free Grammars From Probabilistic Estimation Based on Growth TransformationsabstractAn important problem related to the probabilistic estimation of stochastic context-free grammars (SCFGs) is guaranteeing the consistency of the estimated model. This problem was considered by Booth-Thompson (1973) and Wetherell (1980) and studied by Maryanski (1974) and Chaudhuri et al. (1983) for unambiguous SCFGs only, when the probability distributions were estimated by the relative frequencies in a training sample. In this work, we extend this result by proving that the property of consistency is guaranteed for all SCFGs without restrictions, when the probability distributions are learned from the classical inside-outside and Viterbi algorithms, both of which are based on growth transformations. Other important probabilistic properties which are related to these results are also proven. Joan-Andreu Sánchez, José-Miguel Benedí |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1996 | Simplifying language through error-correcting decodingabstractIn many speech processing tasks, most of the sentences generally convey rather simple meanings.In these tasks, the "wordrecognition" problem is much more difficult than the underlying "speech understanding" problem would be.Accordingly we try to develop an adequate framework to focus on a properly defined "understanding" of the sentences rather than "recognizing" the (possibly) superfluous words.This can be seen as closely related with Spontaneous Language Understanding and Disfluence Modeling.In our approach, these problems are placed under the framework of Error-Correcting Decoding (ECD).A complex task is modeled in terms of a basic stochastic grammar, G, and an Error Model, E (taking insertions, substitutions and deletions into account).G should account for the basic (syntactic) structures underlying this task which would convey the semantics.E should account for general vocabulary variations, speech disfluencies, word disappearance, superfluous words, and so on.Each "complex" user sentence, x, will thus be considered as a corrupted version (according to E) of some "simple" sentence y of LG.Recognition can then be seen as an ECD process: given x, find a sentence y of LG with maximum posterior probability.We introduce fast ECD techniques and adequate procedures for simultaneously training G and E and apply these ideas to a simple task with results showing the potential of the proposed approach. Juan-Carlos Amengual, Enrique Vidal 0001, José-Miguel Benedí |
ICSLP | 3 |
| 1996 | An analysis of general acoustic-phonetic features for Spanish speech produced with the Lombard effect
Antonio Castellanos, José-Miguel Benedí, Francisco Casacuberta |
Speech Commun. | 2 |
| 1988 | On the verification of triangle inequality by dynamic time-warping dissimilarity measures
Enrique Vidal 0001, Francisco Casacuberta, José-Miguel Benedí, Maria-José Lloret, Hector Rulot Segovia |
Speech Commun. | 3 |