VLDB 2026 Research / reviewers in the wild / expert
Carlos D. Martínez-Hinarejos
dblp:04/1927 · also Carlos David Martínez, Carlos David Martínez-Hinarejos
· DBLP profile ↗
46ranked-venue papers
11as first author
14since 2021 · last 2026
0000-0002-6139-2891ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 9 first-author · 9 since 2021Databases, data management, data science and information retrieval · 11 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Explainable Children Autism Detection using Gaze Features in Audio-Visual Speech Comprehension ETRA012abstractAutism Spectrum Disorder (ASD) is a neurodevelopmental condition marked by impairments in social interaction and delayed language acquisition. Early and accurate identification is crucial for timely interventions that support cognitive and social development. Motivated by the subjectivity of traditional behavior-based assessments, computational methodologies offer more objective and cost-effective alternatives. Among these, eye-tracking stands out for capturing subtle attentional and perceptual patterns. This paper investigates the use of eye-tracking data for automatic ASD detection in children during audio-visual storytelling interactions, emphasizing traditional yet explainable machine learning methods. Although performance remains modest, our analyses reveal that fixation duration and revisit patterns to facial regions may serve as potential biomarkers. Further analyses highlight the impact of stimulus modality, suggesting that the inclusion of visual speech cues provides valuable discriminative information. These findings have the potential to support and guide the work of psychologists in the assessment of ASD within speech comprehension contexts. Miguel Zaragozá-Portolés, David Gimeno-Gómez, Vicenta Ávila, Inmaculada Fajardo, Antonio Ferrer, Nadina Gómez-Merino, Noemi Skrobiszewska, Carlos D. Martínez-Hinarejos |
Proc. ACM Hum. Comput. Interact. | 8 |
| 2025 | Improving Lightweight Named Entity Recognition in Handwritten Documents by Predicting Pyramidal Histograms of CharactersabstractNamed Entity Recogniton (NER) consists of tagging parts of an unstructured text containing particular semantic information. When applied to handwritten documents, it is possible to do it as a two-step approach in which Handwritten Text Recognition (HTR) is performed prior to tagging the automatic transcription. However, it is also possible to do both tasks simultaneously by using an HTR model that learns to output the transcription and the tagging symbols. In this paper, we focus on improving the one-step approach by introducing the auxiliary task of predicting Pyramidal Histograms of Characters (PHOC) in a Convolutional Recurrent Neural Network (CRNN) model. Moreover, given the recent rise of models that digest large amounts of data, we also study the usage of synthetic data to pretrain the proposed architecture. Our experiments show that by pretraining the PHOC-based architecture on synthetic data substantial improvements can be made in both transcription and tagging quality without compromising the computational cost of the decoding step. The resulting model matches the NER performance of the state-of-the-art while keeping its lightweight nature. David Villanova-Aparisi, Carlos D. Martínez-Hinarejos, Verónica Romero 0001, Moisés Pastor |
DocEng | 2 |
| 2025 | Acoustic and Linguistic Biomarkers for Cognitive Impairment Detection from Speech
Catarina Botelho, David Gimeno-Gómez, Francisco Teixeira, John Mendonça, Patrícia Pereira, Diogo A. P. Nunes, Thomas Rolland, Anna Pompili, Rubén Solera-Ureña, Maria Ponte, David Martins de Matos, Carlos D. Martínez-Hinarejos, Isabel Trancoso, Alberto Abad |
INTERSPEECH | 12 |
| 2025 | On the Relevance of Clinical Assessment Tasks for the Automatic Detection of Parkinson's Disease Medication State from Speech
David Gimeno-Gómez, Rubén Solera-Ureña, Anna Pompili, Carlos D. Martínez-Hinarejos, Rita Cardoso, Isabel Guimarães, Joaquim Ferreira 0002, Alberto Abad |
INTERSPEECH | 4 |
| 2025 | Tailored design of Audio-Visual Speech Recognition models using BranchformersabstractRecent advances in Audio-Visual Speech Recognition (AVSR) have led to unprecedented achievements in the field, improving the robustness of this type of system in adverse, noisy environments. In most cases, this task has been addressed through the design of models composed of two independent encoders, each dedicated to a specific modality. However, while recent works have explored unified audio-visual encoders, determining the optimal cross-modal architecture remains an ongoing challenge. Furthermore, such approaches often rely on models comprising vast amounts of parameters and high computational cost training processes. In this paper, we aim to bridge this research gap by introducing a novel audio-visual framework. Our proposed method constitutes, to the best of our knowledge, the first attempt to harness the flexibility and interpretability offered by encoder architectures, such as the Branchformer, in the design of parameter-efficient AVSR systems. To be more precise, the proposed framework consists of two steps: first, estimating audio- and video-only systems, and then designing a tailored audio-visual unified encoder based on the layer-level branch scores provided by the modality-specific models. Extensive experiments on English and Spanish AVSR benchmarks covering multiple data conditions and scenarios demonstrated the effectiveness of our proposed method. Even when trained on a moderate scale of data, our models achieve competitive word error rates (WER) of approximately 2.5% for English and surpass existing approaches for Spanish, establishing a new benchmark with an average WER of around 9.1%. These results reflect how our tailored AVSR system is able to reach state-of-the-art recognition rates while significantly reducing the model complexity w.r.t. the prevalent approach in the field. Code and pre-trained models are available at https://github.com/david-gimeno/tailored-avsr . David Gimeno-Gómez, Carlos D. Martínez-Hinarejos |
Comput. Speech Lang. | 2 |
| 2024 | Implementing and Evaluating Trustworthy Conversational Agents for Children
Marina Escobar-Planas, Roberto Ruiz Sánchez, Pedro Frau, Vicky Charisi, Carlos D. Martínez-Hinarejos, Emilia Gómez, Luis Merino |
CHIRA (1) | 5 |
| 2024 | AnnoTheia: A Semi-Automatic Annotation Toolkit for Audio-Visual Speech TechnologiesabstractMore than 7,000 known languages are spoken around the world. However, due to the lack of annotated resources, only a small fraction of them are currently covered by speech technologies. Albeit self-supervised speech representations, recent massive speech corpora collections, as well as the organization of challenges, have alleviated this inequality, most studies are mainly benchmarked on English. This situation is aggravated when tasks involving both acoustic and visual speech modalities are addressed. In order to promote research on low-resource languages for audio-visual speech technologies, we present AnnoTheia, a semi-automatic annotation toolkit that detects when a person speaks on the scene and the corresponding transcription. In addition, to show the complete process of preparing AnnoTheia for a language of interest, we also describe the adaptation of a pre-trained model for active speaker detection to Spanish, using a database not initially conceived for this type of task. Prior evaluations show that the toolkit is able to speed up to four times the annotation process. The AnnoTheia toolkit, tutorials, and pre-trained models are available at https://github.com/joactr/AnnoTheia/. José-M. Acosta-Triana, David Gimeno-Gómez, Carlos D. Martínez-Hinarejos |
LREC/COLING | 3 |
| 2024 | Comparison of Conventional Hybrid and CTC/Attention Decoders for Continuous Visual Speech RecognitionabstractThanks to the rise of deep learning and the availability of large-scale audio-visual databases, recent advances have been achieved in Visual Speech Recognition (VSR). Similar to other speech processing tasks, these end-to-end VSR systems are usually based on encoder-decoder architectures. While encoders are somewhat general, multiple decoding approaches have been explored, such as the conventional hybrid model based on Deep Neural Networks combined with Hidden Markov Models (DNN-HMM) or the Connectionist Temporal Classification (CTC) paradigm. However, there are languages and tasks in which data is scarce, and in this situation, there is not a clear comparison between different types of decoders. Therefore, we focused our study on how the conventional DNN-HMM decoder and its state-of-the-art CTC/Attention counterpart behave depending on the amount of data used for their estimation. We also analyzed to what extent our visual speech features were able to adapt to scenarios for which they were not explicitly trained, either considering a similar dataset or another collected for a different language. Results showed that the conventional paradigm reached recognition rates that improve the CTC/Attention model in data-scarcity scenarios along with a reduced training time and fewer parameters. David Gimeno-Gómez, Carlos D. Martínez-Hinarejos |
LREC/COLING | 2 |
| 2024 | Reading Between the Frames: Multi-modal Depression Detection in Videos from Non-verbal Cues
David Gimeno-Gómez, Ana-Maria Bucur, Adrian Cosma, Carlos D. Martínez-Hinarejos, Paolo Rosso |
ECIR (1) | 4 |
| 2024 | Reading Order Independent Metrics for Information Extraction in Handwritten Documents
David Villanova-Aparisi, Solène Tarride, Carlos D. Martínez-Hinarejos, Verónica Romero 0001, Christopher Kermorvant, Moisés Pastor |
ICDAR (2) | 3 |
| 2023 | Consistent Nested Named Entity Recognition in Handwritten Documents via Lattice Rescoring
David Villanova-Aparisi, Carlos D. Martínez-Hinarejos, Verónica Romero 0001, Moisés Pastor |
ICDAR (1) | 2 |
| 2023 | Evaluation of Different Tagging Schemes for Named Entity Recognition in Handwritten Documents
David Villanova-Aparisi, Carlos D. Martínez-Hinarejos, Verónica Romero 0001, Moisés Pastor |
ICDAR (3) | 2 |
| 2022 | Evaluation of Named Entity Recognition in Handwritten Documents
David Villanova-Aparisi, Carlos D. Martínez-Hinarejos, Verónica Romero 0001, Moisés Pastor |
DAS | 2 |
| 2022 | LIP-RTVE: An Audiovisual Database for Continuous Spanish in the WildabstractSpeech is considered as a multi-modal process where hearing and vision are two fundamentals pillars. In fact, several studies have demonstrated that the robustness of Automatic Speech Recognition systems can be improved when audio and visual cues are combined to represent the nature of speech. In addition, Visual Speech Recognition, an open research problem whose purpose is to interpret speech by reading the lips of the speaker, has been a focus of interest in the last decades. Nevertheless, in order to estimate these systems in the currently Deep Learning era, large-scale databases are required. On the other hand, while most of these databases are dedicated to English, other languages lack sufficient resources. Thus, this paper presents a semi-automatically annotated audiovisual database to deal with unconstrained natural Spanish, providing 13 hours of data extracted from Spanish television. Furthermore, baseline results for both speaker-dependent and speaker-independent scenarios are reported using Hidden Markov Models, a traditional paradigm that has been widely used in the field of Speech Technologies. David Gimeno-Gómez, Carlos D. Martínez-Hinarejos |
LREC | 2 |
| 2020 | Study of the influence of lexicon and language restrictions on computer assisted transcription of historical manuscripts
Emilio Granell, Verónica Romero 0001, Carlos D. Martínez-Hinarejos |
Neurocomputing | 3 |
| 2019 | Image-speech combination for interactive computer assisted transcription of handwritten documents
Emilio Granell, Verónica Romero 0001, Carlos D. Martínez-Hinarejos |
Comput. Vis. Image Underst. | 3 |
| 2018 | Comparing Different Feedback Modalities in Assisted Transcription of ManuscriptsabstractTranscription of handwritten text can be speed-up by using off-line Handwritten Text Recognition techniques, that allow the obtention of an initial draft transcription of an image with handwritten text. However, this draft transcription usually contains errors that must be amended by the transcriber by providing a feedback signal. The usual approach is post-edition, where each error is corrected without modifying the rest of the current transcription. A more sophisticated approach can employ the current modification to provide a new whole transcription, hopefully with less errors. Apart from that, feedback can be provided in different modalities: keyboard input, on-line handwritten text, or speech. Each of these modalities presents different features with respect to ambiguity, derived errors, and final transcription time. In this work we study how the different modalities behave in the assisted transcription of a historical handwritten text document in Spanish and we evaluate their transcription productivity. Carlos D. Martínez-Hinarejos, Emilio Granell, Verónica Romero 0001 |
DAS | 1 |
| 2018 | Multimodality, interactivity, and crowdsourcing for document transcriptionabstractAbstract Knowledge mining from documents usually use document engineering techniques that allow the user to access the information contained in documents of interest. In this framework, transcription may provide efficient access to the contents of handwritten documents. Manual transcription is a time‐consuming task that can be sped up by using different mechanisms. A first possibility is employing state‐of‐the‐art handwritten text recognition systems to obtain an initial draft transcription that can be manually amended. A second option is employing crowdsourcing to obtain a massive but not error‐free draft transcription. In this case, when collaborators employ mobile devices, speech dictation can be used as a transcription source, and speech and handwritten text recognition can be fused to provide a better draft transcription, which can be amended with even less effort. A final option is using interactive assistive frameworks, where the automatic system that provides the draft transcription and the transcriber cooperate to generate the final transcription. The novel contributions presented in this work include the study of the data fusion on a multimodal crowdsourcing framework and its integration with an interactive system. The use of the proposed solutions reduces the required transcription effort and optimizes the overall performance and usability, allowing for a better transcription process. Emilio Granell, Verónica Romero 0001, Carlos D. Martínez-Hinarejos |
Comput. Intell. | 3 |
| 2017 | Baseline Detection on Arabic Handwritten DocumentsabstractDocument processing comprises different steps depending on the nature of the documents. For text documents, specially for handwritten documents, transcription of their contents is one of the main tasks. Handwritten Text Recognition (HTR) is the process of automatically obtaining the transcription of the content of a handwritten text document. In document processing, the basic unit for the acquisition process is the page image, whilst line image is the basic form for the HTR process. This is a bottle-neck which is holding back the massive industrial document processing. Baseline detection can be used not only to segment page images into line images but also for many other document processing steps. Baseline detection problem can be formulated as a clustering problem over a set of interest points. In this work, we study the use of an automatic baseline detection technique, based on interest point clustering, in Arabic handwritten documents. The experiments reveal that this technique provides promising results for this task. Ahmed Fawzi, Moisés Pastor, Carlos D. Martínez-Hinarejos |
DocEng | 3 |
| 2017 | Spanish Sign Language Recognition with Different Topology Hidden Markov Models
Carlos D. Martínez-Hinarejos, Zuzanna Parcheta |
INTERSPEECH | 1 |
| 2017 | Improving the automatic segmentation of subtitles through conditional random field
Aitor Álvarez 0001, Carlos D. Martínez-Hinarejos, Haritz Arzelus, Marina Balenciaga, Arantza del Pozo |
Speech Commun. | 2 |
| 2017 | Multimodal Crowdsourcing for Transcribing Handwritten DocumentsabstractTranscription of handwritten documents is an important research topic for multiple applications, such as document classification or information extraction. In the case of historical documents, their transcription allows to preserve cultural heritage because of the amount of historical data contained in those documents. The transcription process can employ state-of-the-art handwritten text recognition systems in order to obtain an initial transcription. This transcription is usually not good enough for the quality standards, but that may speed up the final transcription of the expert. In this framework, the use of collaborative transcription applications (crowdsourcing) has risen in the recent years, but these platforms are mainly limited by the use of non-mobile devices. Thus, the recruiting initiatives get reduced to a smaller set of potential volunteers. In this paper, an alternative that allows the use of mobile devices is presented. The proposal consists of using speech dictation of handwritten text lines. Then, by using multimodal combination of speech and handwritten text images, a draft transcription can be obtained, presenting more quality than that obtained by only using handwritten text recognition. The speech dictation platform is implemented as a mobile device application, which allows for a wider range of population for recruiting volunteers. A real acquisition on the contents of a Spanish historical handwritten book was obtained with the platform. This data was used to perform experiments on the behaviour of the proposed framework. Some experiments were performed to study how to optimise the collaborators effort in terms of number of collaborations, including how many lines and which lines should be selected for the speech dictation. Emilio Granell, Carlos D. Martínez-Hinarejos |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | An Interactive Approach with Off-Line and On-Line Handwritten Text Recognition Combination for Transcribing Historical DocumentsabstractAutomatic transcription of historical documents is becoming an important research topic, specially because of the increasing number of digitised historical documents that libraries and archives are publishing. However, state-of-the-art handwritten text recognition systems are far from being perfect. Therefore, to have perfect transcriptions, human expert revision is required to really produce a transcription of standard quality. In this context, an interactive assistive scenario, where the automatic system and the human transcriber cooperate to generate the perfect transcription, would allow for a more effective approach. In this paper we present a multimodal interactive transcription system where user feedback is provided by means of touchscreen pen strokes, traditional keyboard and mouse operations. The combination of both the main and the feedback data stream is based on the use of Confusion Networks derived from the output of the on-line and off-line handwritten text recognition systems. The use of the proposed combination help to optimise overall performance and usability. Emilio Granell, Verónica Romero 0001, Carlos D. Martínez-Hinarejos |
DAS | 3 |
| 2016 | A Multimodal Crowdsourcing Framework for Transcribing Historical Handwritten DocumentsabstractTranscription of handwritten historical documents is one of the main topics in document analysis systems, due to cultural reasons. State-of-the-art handwritten text recognition systems allow to speed up the transcription task. Currently, this automatic transcription is far from perfect, and human expert revision is required in order to obtain the actual transcription. In this context, crowdsourcing emerged as a powerful tool for massive transcription at a relatively low cost, since the supervision effort of professional transcribers may be dramatically reduced. However, current transcription crowdsourcing platforms are mainly limited to the use of non-mobile devices, since the use of keyboards in mobile devices is not friendly enough for most users. This work presents the alternative of using speech dictation of handwritten text lines as transcription source in a crowdsourcing platform. The experiments explore how an initial handwritten text recognition hypothesis can be improved by using the contribution of speech recognition from several speakers, providing as a final result a better hypothesis to be amended by a professional transcriber with less effort. Emilio Granell, Carlos D. Martínez-Hinarejos |
DocEng | 2 |
| 2016 | Impact of Automatic Segmentation on the Quality, Productivity and Self-reported Post-editing Effort of Intralingual Subtitles
Aitor Álvarez 0001, Marina Balenciaga, Arantza del Pozo, Haritz Arzelus, Anna Matamala, Carlos D. Martínez-Hinarejos |
LREC | 6 |
| 2015 | Multimodal Output Combination for Transcribing Historical Handwritten Documents
Emilio Granell, Carlos D. Martínez-Hinarejos |
CAIP (1) | 2 |
| 2015 | Combining handwriting and speech recognition for transcribing historical handwritten documentsabstractTranscription of historical documents is an interesting task for libraries in order to make available their funds. In the lasts years, the use of Handwritten Text Recognition allowed paleographs to speed up the manual transcription process, since they are able to correct on a draft transcription. Another alternative is obtaining the draft transcription by dictating the contents to an Automatic Speech Recognition system. When both sources (image and speech) are available, a multimodal combination is possible, and an iterative process can be used in order to refine the final hypothesis. In this work, a multimodal combination based on confusion networks is presented. Results on two different sets of data, with different difficulty level, show that the proposed technique provides similar or better draft transcriptions than a previously proposed approach, allowing for a faster transcription process. Emilio Granell, Carlos D. Martínez-Hinarejos |
ICDAR | 2 |
| 2015 | Unsegmented Dialogue Act Annotation and Decoding With N-Gram TransducersabstractMost studies on dialogue corpora, as well as most dialogue systems, employ dialogue acts as the basic units for interpreting discourse structure, user input and system actions. The definition of the discourse structure and the dialogue strategy consequently require the tagging of dialogue corpora in terms of dialogue acts. The tagging problem presents two basic variants: a batch variant (annotation of whole dialogues, in order to define dialogue strategy or study discourse structure) and an online variant (decoding of the dialogue act sequence of a given turn, in order to interpret user intentions). In the two variants is unusual having the segmentation of each turn into the dialogue meaningful units (segments) to which a dialogue act is assigned. In this paper we present the use of the N-Gram Transducer technique for tagging dialogues, without needing to provide a prior segmentation, in these two different variants (dialogue annotation and turn decoding). Experiments were performed in two corpora of different nature and results show that N-Gram Transducer models are suitable for these tasks and provide good performance. Carlos D. Martínez-Hinarejos, José-Miguel Benedí, Vicent Tamarit |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2014 | An iterative multimodal framework for the transcription of handwritten historical documents
Vicente Alabau, Carlos D. Martínez-Hinarejos, Verónica Romero 0001, Antonio L. Lagarda |
Pattern Recognit. Lett. | 2 |
| 2012 | Estimating the number of segments for improving dialogue act labellingabstractAbstract In dialogue systems it is important to label the dialogue turns with dialogue-related meaning. Each turn is usually divided into segments and these segments are labelled with dialogue acts (DAs). A DA is a representation of the functional role of the segment. Each segment is labelled with one DA, representing its role in the ongoing discourse. The sequence of DAs given a dialogue turn is used by the dialogue manager to understand the turn. Probabilistic models that perform DA labelling can be used on segmented or unsegmented turns. The last option is more likely for a practical dialogue system, but it provides poorer results. In that case, a hypothesis for the number of segments can be provided to improve the results. We propose some methods to estimate the probability of the number of segments based on the transcription of the turn. The new labelling model includes the estimation of the probability of the number of segments in the turn. We tested this new approach with two different dialogue corpora:SwitchBoardandDihana. The results show that this inclusion significantly improves the labelling accuracy. Vicent Tamarit, Carlos D. Martínez-Hinarejos, José-Miguel Benedí |
Nat. Lang. Eng. | 2 |
| 2011 | A Multimodal Approach to Dictation of Handwritten Historical Documents
Vicente Alabau, Verónica Romero 0001, Antonio L. Lagarda, Carlos D. Martínez-Hinarejos |
INTERSPEECH | 4 |
| 2010 | Dialogue act tagging and segmentation with a single perceptron
Ramón Granell, Stephen G. Pulman, Carlos D. Martínez-Hinarejos, José-Miguel Benedí |
INTERSPEECH | 3 |
| 2010 | Evaluation of HMM-based Models for the Annotation of Unsegmented Dialogue Turns
Carlos D. Martínez-Hinarejos, Vicent Tamarit, José-Miguel Benedí |
LREC | 1 |
| 2009 | Improving Unsegmented Dialogue Turns Annotation with N-gram Transducers
Carlos D. Martínez-Hinarejos, Vicent Tamarit, José-Miguel Benedí |
PACLIC | 1 |
| 2009 | Simultaneous Dialogue Act Segmentation and Labelling using Lexical and Syntactic Features
Ramón Granell, Stephen G. Pulman, Carlos D. Martínez-Hinarejos |
SIGDIAL Conference | 3 |
| 2008 | Evaluation of several Maximum Likelihood Linear Regression Variants for Language Adaptation
Míriam Luján-Mares, Carlos D. Martínez-Hinarejos, Vicent Alabau Gonzalvo |
LREC | 2 |
| 2008 | Evaluation of Different Segmentation Techniques for Dialogue Turns
Carlos D. Martínez-Hinarejos, Vicent Tamarit |
LREC | 1 |
| 2008 | Statistical framework for a Spanish spoken dialogue corpus
Carlos D. Martínez-Hinarejos, José-Miguel Benedí, Ramón Granell |
Speech Commun. | 1 |
| 2006 | Segmented and Unsegmented Dialogue-Act Annotation with Statistical Dialogue Models
Carlos D. Martínez-Hinarejos, Ramón Granell, José-Miguel Benedí |
ACL | 1 |
| 2006 | Bilingual speech corpus in two phonetically similar languages
Vicente Alabau, Carlos D. Martínez-Hinarejos |
LREC | 2 |
| 2006 | Computer-assisted translation using speech recognitionabstractCurrent machine translation systems are far from being perfect. However, such systems can be used in computer-assisted translation to increase the productivity of the (human) translation process. The idea is to use a text-to-text translation system to produce portions of target language text that can be accepted or amended by a human translator using text or speech. These user-validated portions are then used by the text-to-text translation system to produce further, hopefully improved suggestions. There are different alternatives of using speech in a computer-assisted translation system: From pure dictated translation to simple determination of acceptable partial translations by reading parts of the suggestions made by the system. In all the cases, information from the text to be translated can be used to constrain the speech decoding search space. While pure dictation seems to be among the most attractive settings, unfortunately perfect speech decoding does not seem possible with the current speech processing technology and human error-correcting would still be required. Therefore, approaches that allow for higher speech recognition accuracy by using increasingly constrained models in the speech recognition process are explored here. All these approaches are presented under the statistical framework. Empirical results support the potential usefulness of using speech within the computer-assisted translation paradigm. Enrique Vidal 0001, Francisco Casacuberta, Luis Rodríguez, Jorge Civera, Carlos D. Martínez-Hinarejos |
IEEE Trans. Speech Audio Process. | 5 |
| 2004 | Some approaches to statistical and finite-state speech-to-speech translation
Francisco Casacuberta, Hermann Ney, Franz Josef Och, Enrique Vidal 0001, Juan Miguel Vilar, Sergio Barrachina 0001, Ismael García-Varea, David Llorens, Carlos D. Martínez-Hinarejos, Sirko Molau |
Comput. Speech Lang. | 9 |
| 2003 | Median strings for k-nearest neighbour classification
Carlos D. Martínez-Hinarejos, Alfons Juan-Císcar, Francisco Casacuberta |
Pattern Recognit. Lett. | 1 |
| 2002 | A Labelling Proposal to Annotate Dialogues
Carlos D. Martínez-Hinarejos, Emilio Sanchis Arnal, Fernando García 0001, Pablo Aibar |
LREC | 1 |
| 2001 | Speech-to-speech translation based on finite-state transducersabstractNowadays, the most successful speech recognition systems are based on stochastic finite-state networks (hidden Markov models and n-grams). Speech translation can be accomplished in a similar way as speech recognition. Stochastic finite-state transducers, which are specific stochastic finite-state networks, have proved very adequate for translation modeling. In this work a speech-to-speech translation system, the EuTRANS system, is presented. The acoustic, language and translation models are finite-state networks that are automatically learnt from training samples. This system was assessed in a series of translation experiments from Spanish to English and from Italian to English in an application involving the interaction (by telephone) of a customer with a receptionist at the front-desk of a hotel. Francisco Casacuberta, David Llorens, Carlos D. Martínez-Hinarejos, Sirko Molau, Francisco Nevado, Hermann Ney, Moisés Pastor, David Picó, Alberto Sanchís, Enrique Vidal 0001, Juan Miguel Vilar |
ICASSP | 3 |
| 2000 | Use of Median String for Classification abstractA string that minimizes the sum of distances to the strings of a given set is known as (generalized) median string of the set. This concept is important in pattern recognition for modelling a (large) set of garbled strings or patterns. The search of such a string is an NP-Hard problem and, therefore, no efficient algorithms to compute the median strings can be designed. A greedy approach has been proposed to compute an approximate median string of a set of strings. In this work an algorithm is proposed that iteratively improves the approximate solution given above. Experiments have been carried out on synthetic and real data to compare the performances of the approximate median string with the conventional set median. These experiments showed that the proposed median string is a better representation of a given set than the corresponding set median. Carlos D. Martínez-Hinarejos, Alfons Juan-Císcar, Francisco Casacuberta |
ICPR | 1 |