Alfons Juan-Císcar

dblp:j/AlfonsJuan · also Alfons Juan · DBLP profile ↗
← Back
63ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0002-9984-4072ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 47 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-authorDatabases, data management, data science and information retrieval · 4Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author
YearPublicationVenuePosition
2025 LHCP-ASR: An English Speech Corpus of High-Energy Particle Physics Talks for Narrow-Domain ASR Benchmarking
Jaume Santamaria-Jorda, Pablo Segovia-Martínez, Gonçal V. Garcés Díaz-Munío, Joan Albert Silvestre-Cerdà, Adrià Giménez, Rubén Gaspar Aparicio, René Fernández Sánchez, Jorge Civera, Alberto Sanchís, Alfons Juan-Císcar
INTERSPEECH10
2025 Speech translation for multilingual medical education leveraged by large language models
abstract
The application of large language models (LLMs) to speech translation (ST) or, in general, to machine translation (MT) has recently provided excellent results, superseding conventional encoder-decoder MT systems in the general domain. However, this is not clearly the case when LLMs as MT systems are translating medical-related materials. In this respect, the provision of multilingual training materials for oncology professionals is a goal of the EU project Interact-Europe in which this work was framed. To this end, cross-language technology adapted to the oncology domain was developed, evaluated and deployed for multilingual interspecialty medical education. More precisely, automatic speech recognition (ASR) and MT models were adapted to the oncology domain to translate English pre-recorded training videos, kindly provided by the European School of Oncology (ESO), into French, Spanish, German and Slovene. In this work, three categories of MT models adapted to the medical domain were assessed: bilingual encoder-decoder MT models trained from scratch, pre-trained large multilingual encoder-decoder MT models, and multilingual decoder-only LLMs. The experimental results underline the competitiveness in translation quality of LLMs compared to encoder-decoder MT models. Finally, the ESO speech dataset, comprising roughly 1000 videos and 745 h for the training and evaluation of ASR, MT and ST models, was publicly released for the scientific community.
Jorge Iranzo-Sánchez, Jaume Santamaria-Jorda, Gerard Mas-Mollà, Gonçal V. Garcés Díaz-Munío, Javier Iranzo-Sánchez, Javier Jorge, Joan Albert Silvestre-Cerdà, Adrià Giménez, Jorge Civera, Alberto Sanchís, Alfons Juan-Císcar
Artif. Intell. Medicine11
2024 Segmentation-Free Streaming Machine Translation
abstract
Abstract Streaming Machine Translation (MT) is the task of translating an unbounded input text stream in real-time. The traditional cascade approach, which combines an Automatic Speech Recognition (ASR) and an MT system, relies on an intermediate segmentation step which splits the transcription stream into sentence-like units. However, the incorporation of a hard segmentation constrains the MT system and is a source of errors. This paper proposes a Segmentation-Free framework that enables the model to translate an unsegmented source stream by delaying the segmentation decision until after the translation has been generated. Extensive experiments show how the proposed Segmentation-Free framework has better quality-latency trade-off than competing approaches that use an independent segmentation model.1
Javier Iranzo-Sánchez, Jorge Iranzo-Sánchez, Adrià Giménez, Jorge Civera, Alfons Juan-Císcar
Trans. Assoc. Comput. Linguistics5
2022 From Simultaneous to Streaming Machine Translation by Leveraging Streaming History
abstract
Simultaneous Machine Translation is the task of incrementally translating an input sentence before it is fully available.Currently, simultaneous translation is carried out by translating each sentence independently of the previously translated text.More generally, Streaming MT can be understood as an extension of Simultaneous MT to the incremental translation of a continuous input text stream.In this work, a state-of-the-art simultaneous sentencelevel MT system is extended to the streaming setup by leveraging the streaming history.Extensive empirical results are reported on IWSLT Translation Tasks, showing that leveraging the streaming history leads to significant quality gains.In particular, the proposed system proves to compare favorably to the best performing systems.
Javier Iranzo-Sánchez, Jorge Civera, Alfons Juan-Císcar
ACL (1)3
2022 Live Streaming Speech Recognition Using Deep Bidirectional LSTM Acoustic Models and Interpolated Language Models
abstract
Although Long-Short Term Memory (LSTM) networks and deep Transformers are now extensively used in offline ASR, it is unclear how best offline systems can be adapted to work with them under the streaming setup. After gaining considerable experience on this regard in recent years, in this paper we show how an optimized, low-latency streaming decoder can be built in which bidirectional LSTM acoustic models, together with general interpolated language models, can be nicely integrated with minimal perfomance degradation. In brief, our streaming decoder consists of a one-pass, real-time search engine relying on a limited-duration window sliding over time and a number of ad hoc acoustic and language model pruning techniques. Extensive empirical assessment is provided on truly streaming tasks derived from the well-known LibriSpeech and TED talks datasets, as well as from TV shows on a main Spanish broadcasting station.
Javier Jorge, Adrià Giménez, Joan Albert Silvestre-Cerdà, Jorge Civera, Alberto Sanchís, Alfons Juan-Císcar
IEEE ACM Trans. Audio Speech Lang. Process.6
2021 Europarl-ASR: A Large Corpus of Parliamentary Debates for Streaming ASR Benchmarking and Speech Data Filtering/Verbatimization
Gonçal V. Garcés Díaz-Munío, Joan Albert Silvestre-Cerdà, Javier Jorge, Adrià Giménez-Pastor, Javier Iranzo-Sánchez, Pau Baquero-Arnal, Nahuel Roselló, Alejandro Pérez González de Martos, Jorge Civera, Alberto Sanchís, Alfons Juan-Císcar
Interspeech11
2021 Towards Simultaneous Machine Interpretation
abstract
[EN] Automatic speech-to-speech translation (S2S) is one of the most challenging speech and language processing tasks, especially when considering its application to real-time settings. Recent advances on streaming Automatic Speech Recognition (ASR), simultaneous Machine Translation (MT) and incremental neural Text-To-Speech (TTS) make it possible to develop real-time cascade S2S systems with greatly improved accuracy. On the way to simultaneous machine interpretation, a state-of-the-art cascade streaming S2S system is described and empirically assessed in the simultaneous interpretation of European Parliament debates. We pay particular attention to the TTS component, particularly in terms of speech naturalness under a variety of response-time settings, as well as in terms of speaker similarity for its cross-lingual voice cloning capabilities.
Alejandro Pérez González de Martos, Javier Iranzo-Sánchez, Adrià Giménez-Pastor, Javier Jorge, Joan Albert Silvestre-Cerdà, Jorge Civera, Alberto Sanchís, Alfons Juan-Císcar
Interspeech8
2021 Towards cross-lingual voice cloning in higher education
abstract
The rapid progress of modern AI tools for automatic speech recognition and machine translation is leading to a progressive cost reduction to produce publishable subtitles for educational videos in multiple languages. Similarly, text-to-speech technology is experiencing large improvements in terms of quality, flexibility and capabilities. In particular, state-of-the-art systems are now capable of seamlessly dealing with multiple languages and speakers in an integrated manner, thus enabling lecturer’s voice cloning in languages she/he might not even speak. This work is to report the experience gained on using such systems at the Universitat Politècnica de València (UPV), mainly as a guidance for other educational organizations willing to conduct similar studies. It builds on previous work on the UPV’s main repository of educational videos, MediaUPV, to produce multilingual subtitles at scale and low cost. Here, a detailed account is given on how this work has been extended to also allow for massive machine dubbing of MediaUPV. This includes collecting 59 h of clean speech data from UPV’s academic staff, and extending our production pipeline of subtitles with a state-of-the-art multilingual and multi-speaker text-to-speech system trained from the collected data. Our main result comes from an extensive, subjective evaluation of this system by lecturers contributing to data collection. In brief, it is shown that text-to-speech technology is not only mature enough for its application to MediaUPV, but also needed as soon as possible by students to improve its accessibility and bridge language barriers.
Gonçal V. Garcés Díaz-Munío, Adrià Giménez, Joan Albert Silvestre-Cerdà, Alberto Sanchís, Jorge Civera, Manuel Jiménez, Carlos Turro, Alfons Juan-Císcar
Eng. Appl. Artif. Intell.9
2021 Streaming cascade-based speech translation leveraged by a direct segmentation model
abstract
The cascade approach to Speech Translation (ST) is based on a pipeline that concatenates an Automatic Speech Recognition (ASR) system followed by a Machine Translation (MT) system. Nowadays, state-of-the-art ST systems are populated with deep neural networks that are conceived to work in an offline setup in which the audio input to be translated is fully available in advance. However, a streaming setup defines a completely different picture, in which an unbounded audio input gradually becomes available and at the same time the translation needs to be generated under real-time constraints. In this work, we present a state-of-the-art streaming ST system in which neural-based models integrated in the ASR and MT components are carefully adapted in terms of their training and decoding procedures in order to run under a streaming setup. In addition, a direct segmentation model that adapts the continuous ASR output to the capacity of simultaneous MT systems trained at the sentence level is introduced to guarantee low latency while preserving the translation quality of the complete ST system. The resulting ST system is thoroughly evaluated on the real-life streaming Europarl-ST benchmark to gauge the trade-off between quality and latency for each component individually as well as for the complete ST system.
Javier Iranzo-Sánchez, Javier Jorge, Pau Baquero-Arnal, Joan Albert Silvestre-Cerdà, Adrià Giménez, Jorge Civera, Alberto Sanchís, Alfons Juan-Císcar
Neural Networks8
2020 Direct Segmentation Models for Streaming Speech Translation
abstract
Javier Iranzo-Sánchez, Adrià Giménez Pastor, Joan Albert Silvestre-Cerdà, Pau Baquero-Arnal, Jorge Civera Saiz, Alfons Juan. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Javier Iranzo-Sánchez, Adrià Giménez-Pastor, Joan Albert Silvestre-Cerdà, Pau Baquero-Arnal, Jorge Civera Saiz, Alfons Juan-Císcar
EMNLP (1)6
2020 Europarl-ST: A Multilingual Corpus for Speech Translation of Parliamentary Debates
abstract
Current research into spoken language translation (SLT), or speech-to-text translation, is often hampered by the lack of specific data resources for this task, as currently available SLT datasets are restricted to a limited set of language pairs. In this paper we present Europarl-ST, a novel multilingual SLT corpus containing paired audio-text samples for SLT from and into 6 European languages, for a total of 30 different translation directions. This corpus has been compiled using the debates held in the European Parliament in the period between 2008 and 2012. This paper describes the corpus creation process and presents a series of automatic speech recognition, machine translation and spoken language translation experiments that highlight the potential of this new resource. The corpus is released under a Creative Commons license and is freely accessible and downloadable.
Javier Iranzo-Sánchez, Joan Albert Silvestre-Cerdà, Javier Jorge, Nahuel Roselló, Adrià Giménez, Alberto Sanchís, Jorge Civera, Alfons Juan-Císcar
ICASSP8
2020 LSTM-Based One-Pass Decoder for Low-Latency Streaming
abstract
Current state-of-the-art models based on Long-Short Term Memory (LSTM) networks have been extensively used in ASR to improve performance. However, using LSTMs under a streaming setup is not straightforward due to real-time constraints. In this paper we present a novel streaming decoder that includes a bidirectional LSTM acoustic model as well as an unidirectional LSTM language model to perform the decoding efficiently while keeping the performance comparable to that of an off-line setup. We perform a one-pass decoding using a sliding window scheme for a bidirectional LSTM acoustic model and an LSTM language model. This has been implemented and assessed under a pure streaming setup, and deployed into our production systems. We report WER and latency figures for the well-known LibriSpeech and TED-LIUM tasks, obtaining competitive WER results with low-latency responses.
Javier Jorge, Adrià Giménez, Javier Iranzo-Sánchez, Joan Albert Silvestre-Cerdà, Jorge Civera, Alberto Sanchís, Alfons Juan-Císcar
ICASSP7
2020 Improved Hybrid Streaming ASR with Transformer Language Models
Pau Baquero-Arnal, Javier Jorge, Adrià Giménez, Joan Albert Silvestre-Cerdà, Javier Iranzo-Sánchez, Alberto Sanchís, Jorge Civera, Alfons Juan-Císcar
INTERSPEECH8
2019 Real-Time One-Pass Decoder for Speech Recognition Using LSTM Language Models
abstract
Recurrent Neural Networks, in particular Long-Short TermMemory (LSTM) networks, are widely used in Automatic Speech Recognition for language modelling during decoding,usually as a mechanism for rescoring hypothesis. This paperproposes a new architecture to perform real-time one-pass de-coding using LSTM language models. To make decoding ef-ficient, the estimation of look-ahead scores was accelerated byprecomputing static look-ahead tables. These static tables wereprecomputed from a prunedn-gram model, reducing drasti-cally the computational cost during decoding. Additionally,the LSTM language model evaluation was efficiently performedusing Variance Regularization along with a strategy of lazyevaluation. The proposed one-pass decoder architecture wasevaluated on the well-known LibriSpeech and TED-LIUMv3datasets. Results showed that the proposed algorithm obtainsvery competitive WERs with 0.6 RTFs. Finally, our one-passdecoder is compared with a decoupled two-pass decoder.
Javier Jorge, Adrià Giménez, Javier Iranzo-Sánchez, Jorge Civera, Alberto Sanchís, Alfons Juan-Císcar
INTERSPEECH6
2018 Speaker-Adapted Confidence Measures for ASR Using Deep Bidirectional Recurrent Neural Networks
abstract
In the last years, deep bidirectional recurrent neural networks (DBRNN) and DBRNN with long short-term memory cells (DBLSTM) have outperformed the most accurate classifiers for confidence estimation in automatic speech recognition. At the same time, we have recently shown that speaker adaptation of confidence measures using DBLSTM yields significant improvements over non-adapted confidence measures. In accordance with these two recent contributions to the state of the art in confidence estimation, this paper presents a comprehensive study of speaker-adapted confidence measures using DBRNN and DBLSTM models. First, we present new empirical evidences of the superiority of recurrent neural networks (RNN)-based confidence classifiers evaluated over a large speech corpus consisting of the English LibriSpeech and the Spanish poliMedia tasks. Second, we show new results on speaker-adapted confidence measures considering a multitask framework in which RNN-based confidence classifiers trained with LibriSpeech are adapted to speakers of the TED-LIUM corpus. These experiments confirm that speaker-adapted confidence measures outperform their non-adapted counterparts. Last, we describe an unsupervised adaptation method of the acoustic DBLSTM model based on confidence measures that results in better automatic speech recognition performance.
Miguel A. del Agua, Adrià Giménez, Alberto Sanchís, Jorge Civera, Alfons Juan-Císcar
IEEE ACM Trans. Audio Speech Lang. Process.5
2016 ASR Confidence Estimation with Speaker-Adapted Recurrent Neural Networks
Miguel A. del Agua, Santiago Piqueras, Adrià Giménez, Alberto Sanchís, Jorge Civera, Alfons Juan-Císcar
INTERSPEECH6
2016 Speaker-adapted confidence measures for speech recognition of video lectures
abstract
Automatic speech recognition applications can benefit from a confidence measure (CM) to predict the reliability of the output. Previous works showed that a word-dependent naïve Bayes (NB) classifier outperforms the conventional word posterior probability as a CM. However, a discriminative formulation usually renders improved performance due to the available training techniques. Taking this into account, we propose a logistic regression (LR) classifier defined with simple input functions to approximate to the NB behaviour. Additionally, as a main contribution, we propose to adapt the CM to the speaker in cases in which it is possible to identify the speakers, such as online lecture repositories. The experiments have shown that speaker-adapted models outperform their non-adapted counterparts on two difficult tasks from English (videoLectures.net) and Spanish (poliMedia) educational lectures. They have also shown that the NB model is clearly superseded by the proposed LR classifier.
Isaias Sanchez-Cortina, Jesús Andrés-Ferrer, Alberto Sanchís, Alfons Juan-Císcar
Comput. Speech Lang.4
2015 Efficient Generation of High-Quality Multilingual Subtitles for Video Lecture Repositories
abstract
Video lectures are a valuable educational tool in higher education to support or replace face-to-face lectures in active learning strategies. In 2007 the Universitat Politècnica de València (UPV) implemented its video lecture capture system, resulting in a high quality educational video repository, called poliMedia, with more than 10.000 mini lectures created by 1.373 lecturers. Also, in the framework of the European project transLectures, UPV has automatically generated transcriptions and translations in Spanish, Catalan and English for all videos included in the poliMedia video repository. transLectures’s objective responds to the widely-recognised need for subtitles to be provided with video lectures, as an essential service for non-native speakers and hearing impaired persons, and to allow advanced repository functionalities. Although high-quality automatic transcriptions and translations were generated in transLectures, they were not error-free. For this reason, lecturers need to manually review video subtitles to guarantee the absence of errors. The aim of this study is to evaluate the efficiency of the manual review process from automatic subtitles in comparison with the conventional generation of video subtitles from scratch. The reported results clearly indicate the convenience of providing automatic subtitles as a first step in the generation of video subtitles and the significant savings in time of up to almost 75 % involved in reviewing subtitles.
Juan Daniel Valor Miró, Joan Albert Silvestre-Cerdà, Jorge Civera, Carlos Turro, Alfons Juan-Císcar
EC-TEL5
2015 Window repositioning for printed Arabic recognition
Ihab Khoury, Adrià Giménez, Alfons Juan-Císcar, Jesús Andrés-Ferrer
Pattern Recognit. Lett.3
2015 Efficiency and usability study of innovative computer-aided transcription strategies for video lecture repositories
Juan Daniel Valor Miró, Joan Albert Silvestre-Cerdà, Jorge Civera, Carlos Turro, Alfons Juan-Císcar
Speech Commun.5
2014 Interactive handwriting recognition with limited user effort
Adrià Giménez, Jorge Civera, Alberto Sanchís, Alfons Juan-Císcar
Int. J. Document Anal. Recognit.5
2014 Discriminative Bernoulli HMMs for isolated handwritten word recognition
Adrià Giménez, Jesús Andrés-Ferrer, Alfons Juan-Císcar
Pattern Recognit. Lett.3
2014 Handwriting word recognition using windowed Bernoulli HMMs
Adrià Giménez, Ihab Khoury, Jesús Andrés-Ferrer, Alfons Juan-Císcar
Pattern Recognit. Lett.4
2014 Effective balancing error and user effort in interactive handwriting recognition
Jorge Civera, Alberto Sanchís, Alfons Juan-Císcar
Pattern Recognit. Lett.4
2013 Language model adaptation for video lectures transcription
abstract
Videolectures are currently being digitised all over the world for its enormous value as reference resource. Many of these lectures are accompanied with slides. The slides offer a great opportunity for improving ASR systems performance. We propose a simple yet powerful extension to the linear interpolation of language models for adapting language models with slide information. Two types of slides are considered, correct slides, and slides automatic extracted from the videos with OCR. Furthermore, we compare both time aligned and unaligned slides. Results report an improvement of up to 3.8 % absolute WER points when using correct slides. Surprisingly, when using automatic slides obtained with poor OCR quality, the ASR system still improves up to 2.2 absolute WER points.
Adria A. Martinez-Villaronga, Miguel A. del Agua, Jesús Andrés-Ferrer, Alfons Juan-Císcar
ICASSP4
2013 A System Architecture to Support Cost-Effective Transcription and Translation of Large Video Lecture Repositories
abstract
Online video lecture repositories are rapidly growing and becoming established as fundamental knowledge assets. However, most lectures are neither transcribed nor translated because of the lack of cost-effective solutions that can give accurate enough results. In this paper, we describe a system architecture that supports the cost-effective transcription and translation of large video lecture repositories. This architecture has been adopted in the EU project transLectures and is now being tested on a repository of more than 9000 video lectures at the Universitat Politecnica de Valencia. Following a brief description of this repository and of the transLectures project, we describe the proposed system architecture in detail. We also report empirical results on the quality of the transcriptions and translations currently being maintained and steadily improved.
Joan Albert Silvestre-Cerdà, Alejandro Pérez González de Martos, Manuel Jiménez, Carlos Turro, Alfons Juan-Císcar, Jorge Civera
SMC5
2012 Comparison of Bernoulli and Gaussian HMMs Using a Vertical Repositioning Technique for Off-Line Handwriting Recognition
abstract
In this paper a vertical repositioning method based on the center of gravity is investigated for handwriting recognition systems and evaluated on databases containing Arabic and French handwriting. Experiments show that vertical distortion in images has a large impact on the performance of HMM based handwriting recognition systems. Recently good results were obtained with Bernoulli HMMs (BHMMs) using a preprocessing with vertical repositioning of binarized images. In order to isolate the effect of the preprocessing from the BHMM model, experiments were conducted with Gaussian HMMs and the LSTM-RNN tandem HMM approach with relative improvements of 33% WER on the Arabic and up to 62% on the French database.
Patrick Doetsch, Mahdi Hamdani, Hermann Ney, Adrià Giménez, Jesús Andrés-Ferrer, Alfons Juan-Císcar
ICFHR6
2012 A prototype for interactive speech transcription balancing error and supervision effort
abstract
A system to transcribe speech data is presented following an interactive paradigm in which both, the system produces automatically speech transcriptions and the user is assisted by the system to amend output errors as efficiently as possible. Partially supervised transcriptions with a tolerance error fixed by the user are used to incrementally adapt the underlying system models. The prototype uses a simple yet effective method to find an optimal balance between recognition error and supervision effort.
Isaias Sanchez-Cortina, Alberto Sanchís, Alfons Juan-Císcar
IUI4
2012 A Word-Based Naïve Bayes Classifier for Confidence Estimation in Speech Recognition
abstract
Confidence estimation has been largely used in speech recognition to detect words in the recognized sentence that have been likely misrecognized. Confidence estimation can be seen as a conventional pattern classification problem in which a set of features is obtained for each hypothesized word in order to classify it as either correct or incorrect. We propose a smoothed naïve Bayes classification model to profitably combine these features. The model itself is a combination of word-dependent (specific) and word-independent (generalized) naïve Bayes models. As in statistical language modeling, the purpose of the generalized model is to smooth the (class posterior) estimates given by the specific models. Our classification model is empirically compared with confidence estimation based on posterior probabilities computed on word graphs. Empirical results clearly show that the good performance of word graph-based posterior probabilities can be improved by using the naïve Bayes combination of features.
Alberto Sanchís, Alfons Juan-Císcar, Enrique Vidal 0001
IEEE Trans. Speech Audio Process.2
2011 Minimum Bayes-risk System Combination
Jesús González-Rubio, Alfons Juan-Císcar, Francisco Casacuberta
ACL2
2011 Discriminative Bernoulli Mixture Models for Handwritten Digit Recognition
abstract
Bernoulli-based models such as Bernoulli mixtures or Bernoulli HMMs (BHMMs), have been successfully applied to several handwritten text recognition (HTR) tasks which range from character recognition to continuous and isolated handwritten words. All these models belong to the generative model family and, hence, are usually trained by (joint) maximum likelihood estimation (MLE). Despite the good properties of the MLE criterion, there are better training criteria such as maximum mutual information (MMI). The MMI is a widespread criterion that is mainly employed to train discriminative models such as log-linear (or maximum entropy) models. Inspired by the Bernoulli mixture classifier, in this work a log-linear model for binary data is proposed, the so-called mixture of multi-class logistic regression. The proposed model is proved to be equivalent to the Bernoulli mixture classifier. In this way, we give a discriminative training framework for Bernoulli mixture models. The proposed discriminative training framework is applied to a well-known Indian digit recognition task.
Adrià Giménez, Jesús Andrés-Ferrer, Alfons Juan-Císcar
ICDAR3
2010 The APP Oracle - An Interactive Student Competition on Pattern Recognition
Alfons Juan-Císcar, Jesús Andrés-Ferrer, Adrià Giménez, Jorge Civera, Roberto Paredes, Enrique Vidal 0001
CSEDU (2)1
2010 Interactive layout analysis and transcription systems for historic handwritten documents
abstract
The amount of digitized legacy documents has been rising dramatically over the last years due mainly to the increasing number of on-line digital libraries publishing this kind of documents, waiting to be classified and finally transcribed into a textual electronic format (such as ASCII or PDF). Nevertheless, most of the available fully-automatic applications addressing this task are far from being perfect and heavy and inefficient human intervention is often required to check and correct the results of such systems. In contrast, multimodal interactive-predictive approaches may allow the users to participate in the process helping the system to improve the overall performance. With this in mind, two sets of recent advances are introduced in this work: a novel interactive method for text block detection and two multimodal interactive handwritten text transcription systems which use active learning and interactive-predictive technologies in the recognition process.
Oriol Ramos Terrades, Alejandro H. Toselli, Verónica Romero 0001, Enrique Vidal 0001, Alfons Juan-Císcar
ACM Symposium on Document Engineering6
2010 Windowed Bernoulli Mixture HMMs for Arabic Handwritten Word Recognition
abstract
Hidden Markov Models (HMMs) are now widely used in off-line handwriting recognition and, in particular, in Arabic handwritten word recognition. In contrast to the conventional approach, based on Gaussian mixture HMMs, we have recently proposed to directly fed columns of raw, binary pixels into Bernoulli mixture HMMs. In this work, column bit vectors are extended by means of a sliding window of adequate width to better capture image context at each horizontal position of the word image. Using these windowed Bernoulli mixture HMMs, very good results are reported on the well-known IfN/ENIT database of Arabic handwritten Tunisian town names.
Adrià Giménez, Ihab Khoury, Alfons Juan-Císcar
ICFHR3
2010 Balancing error and supervision effort in interactive-predictive handwriting recognition
abstract
An effective approach to transcribe handwritten text documents is to follow an interactive-predictive paradigm in which both, the system is guided by the user, and the user is assisted by the system to complete the transcription task as efficiently as possible. This approach has been recently implemented in a system prototype called GIDOC, in which standard speech technology is adapted to handwritten text (line) images: HMM-based text image modeling, n-gram language modeling, and also confidence measures on recognized words. Confidence measures are used to assist the user in locating possible transcription errors, and thus validate system output after only supervising those (few) words for which the system is not highly confident. However, a certain degree of supervision is required for proper model adaptation from partially supervised transcriptions. Here, we propose a simple yet effective method to find an optimal balance between recognition error and supervision effort.
Alberto Sanchís, Alfons Juan-Císcar
IUI3
2010 Saturnalia: A Latin-Catalan Parallel Corpus for Statistical MT
Jesús González-Rubio, Jorge Civera, Alfons Juan-Císcar, Francisco Casacuberta
LREC3
2010 The RODRIGO Database
Alfons Juan-Císcar
LREC3
2010 Constrained domain maximum likelihood estimation for naive Bayes text classification
Jesús Andrés-Ferrer, Alfons Juan-Císcar
Pattern Anal. Appl.2
2009 Embedded Bernoulli Mixture HMMs for Continuous Handwritten Text Recognition
Adrià Giménez, Alfons Juan-Císcar
CAIP2
2009 A Phrase-Based Hidden Semi-Markov Approach to Machine Translation
Jesús Andrés-Ferrer, Alfons Juan-Císcar
EAMT2
2009 Embedded Bernoulli Mixture HMMs for Handwritten Word Recognition
abstract
Hidden Markov Models (HMMs) are now widely used in off-line handwritten word recognition. As in speech recognition, they are usually built from shared, embedded HMMs at symbol level, in which state-conditional probability density functions are modelled with Gaussian mixtures. In contrast to speech recognition, however, it is unclear which kind of real-valued features should be used and, indeed, very different features sets are in use today. In this paper, we propose to by-pass feature extraction and directly fed columns of raw, binary image pixels into embedded Bernoulli mixture HMMs, that is, embedded HMMs in which the emission probabilities are modelled with Bernoulli mixtures. The idea is to ensure that no discriminative information is filtered out during feature extraction, which in some sense is integrated into the recognition model. Empirical results are reported in which similar results are obtained with both Bernoulli and Gaussian mixtures, though Bernoulli mixtures are much simpler.
Adrià Giménez, Alfons Juan-Císcar
ICDAR2
2009 The GERMANA Database
abstract
A new handwritten text database, GERMANA, is presented to facilitate empirical comparison of different approaches to text line extraction and off-line handwriting recognition. GERMANA is the result of digitising and annotating a 764-page Spanish manuscript from 1891, in which most pages only contain nearly calligraphed text written on ruled sheets of well-separated lines. To our knowledge, it is the first publicly available database for handwriting research, mostly written in Spanish and comparable in size to standard databases. Due to its sequential book structure, it is also well-suited for realistic assessment of interactive handwriting recognition systems. To provide baseline results for reference in future studies, empirical results are also reported, using standard techniques and tools for preprocessing, feature extraction, HMM-based image modelling, and language modelling.
Daniel Gracia Pérez, Lionel Tarazón, Oriol Ramos Terrades, Alfons Juan-Císcar
ICDAR6
2009 Adaptation from partially supervised handwritten text transcriptions
abstract
An effective approach to transcribe handwritten text documents is to follow an interactive-predictive paradigm in which both, the system is guided by the user, and the user is assisted by the system to complete the transcription task as efficiently as possible. This approach has been recently implemented in a system prototype called GIDOC, in which standard speech technology is adapted to handwritten text (line) images: HMM-based text image modelling, n-gram language modelling, and also confidence measures on recognized words. Confidence measures are used to assist the user in locating possible transcription errors, and thus validate system output after only supervising those (few) words for which the system is not highly confident. Here, we study the effect of using these partially supervised transcriptions on the adaptation of image and language models to the task.
Daniel Gracia Pérez, Alberto Sanchís, Alfons Juan-Císcar
ICMI4
2008 A novel alignment model inspired on IBM Model 1
Jesús González-Rubio, Germán Sanchis-Trilles, Alfons Juan-Císcar, Francisco Casacuberta
EAMT3
2008 Maximum entropy models for speech confidence estimation
abstract
In this work we implement a confidence estimation system based on a Naive Bayes classifier, by using the maximum entropy paradigm. The model takes information from various sources including a set of scores which have proved to be useful in confidence estimation tasks. Two different approaches are modeled. First a basic model which takes advantages of smoothing techniques used in a previous work, and second an optimized model, which is designed to hold a set of very few but essential characteristics of the model, without decrease in the performance. A considerably reduction in the number of parameters is obtained compared to the basic model. Both models are evaluated with two different corpora and compared to a model previously developed.
Claudio Estienne, Alberto Sanchís, Alfons Juan-Císcar, Enrique Vidal 0001
ICASSP3
2008 Bilingual Text Classification using the IBM 1 Translation Model
Jorge Civera, Alfons Juan-Císcar
LREC2
2007 Estimation of confidence measures for machine translation
Alberto Sanchís, Alfons Juan-Císcar, Enrique Vidal 0001
MTSummit2
2007 Iterative Contextual Recurrent Classification of Chromosomes
César Ernesto Martínez, Alfons Juan-Císcar, Francisco Casacuberta
Neural Process. Lett.2
2006 Sense Cluster Based Categorization and Clustering of Abstracts
Davide Buscaldi, Paolo Rosso, Mikhail Alexandrov, Alfons Juan-Císcar
CICLing4
2006 Mixtures of IBM Model 2
Jorge Civera, Alfons Juan-Císcar
EAMT2
2006 Local transformation models for speech recognition
abstract
This paper presents a novel acoustic modeling framework that naturally extends the Hidden Markov Model (HMM) approach. The novel models reduce the errors caused by speaker variability by means of a local spectral mismatch reduction. A more complex and flexible speech production scheme can be assumed, in which the local temporal and frequency elastic deformations of the speech are captured by the model. In the new framework the states of a standard HMM, which are usually associated with temporal transitions, are expanded so that a new degree of freedom for the model is provided and it is then possible to estimate an optimum frequency warping factor at the same time as the decoder finds the best state sequence. In the local spectral warping based models the states become time-frequency related states and the number of parameters of the model is comparable to the standard HMM since they share a certain amount of parameters as it will be shown. The novel models are evaluated in the noise-free TIDIGITS corpus, which includes connected digits uttered by male, female and children. It has been found that, under speaker group (age-gender) mismatch conditions, the local frequency warping reduced Word Error Rate (WER) in mean by a 70%, using the initial models. When matched speaker group conditions were tested the error was reduced in mean in a 9.7% after reestimating the models. Index Terms: speaker variability, local frequency warping.
Antonio Miguel, Eduardo Lleida, Alfons Juan-Císcar, Luis Buera, Alfonso Ortega Giménez, Oscar Saz-Torralba
INTERSPEECH3
2006 Bilingual Machine-Aided Indexing
Jorge Civera, Alfons Juan-Císcar
LREC2
2004 New features based on multiple word graphs for utterance verification
abstract
The goal of Utterance Verification is to estimate a confidence measure which helps detecting words in the hypothesized sentence that are likely to have been missrecognized. Word graphs have been extensively employed for directly estimating the confidence measure and for extracting important predictor features. In all the cases, a single word graph which is obtained through the recognition process. In this paper we propose the use of multiple word graphs to compute new features. The experimental study shows that these proposed features outperform those computed on a single word graph and other well-known predictor features. Moreover, the combination of the proposed features along with other kind of features provides improvements in the verification accuracy.
Alberto Sanchís, Alfons Juan-Císcar, Enrique Vidal 0001
INTERSPEECH2
2004 Integrated Handwriting Recognition And Interpretation Using Finite-State Models
abstract
The interpretation of handwritten sentences is carried out using a holistic approach in which both text image recognition and the interpretation itself are tightly integrated. Conventional approaches follow a serial, first-recognition then-interpretation scheme which cannot adequately use semantic–pragmatic knowledge to recover from recognition errors. Stochastic finite-sate transducers are shown to be suitable models for this integration, permitting a full exploitation of the final interpretation constraints. Continuous-density hidden Markov models are embedded in the edges of the transducer to account for lexical and morphological constraints. Robustness with respect to stroke vertical variability is achieved by integrating tangent vectors into the emission densities of these models. Experimental results are reported on a syntax-constrained interpretation task which show the effectiveness of the proposed approaches. These results are also shown to be comparatively better than those achieved with other conventional, N-gram-based techniques which do not take advantage of full integration.
Alejandro H. Toselli, Alfons Juan-Císcar, Ismael Salvador, Enrique Vidal 0001, Francisco Casacuberta, Daniel Keysers, Hermann Ney
Int. J. Pattern Recognit. Artif. Intell.2
2003 Improving utterance verification using a smoothed naive Bayes model
abstract
Utterance verification can be seen as a conventional pattern classification problem in which a feature vector is obtained for each hypothesized word in order to classify it as either correct or incorrect. It is unclear, however, which predictor (pattern) features and classification model should be used. Regarding the features, we have proposed a new feature, called word trellis stability (WTS), that can be profitably used in conjunction with more or less standard features such as acoustic stability. This is confirmed in this paper, where a smoothed naive Bayes classification model is proposed to adequately combine predictor features. On a series of experiments with this classification model and several features, we have found that the results provided by each feature alone are outperformed by certain combinations. In particular, the combination of the two above-mentioned features has been consistently found to give the most accurate result in two verification tasks.
Alberto Sanchís, Alfons Juan-Císcar, Enrique Vidal 0001
ICASSP (1)2
2003 Utterance verification using an optimized k-nearest neighbour classifier
Roberto Paredes, Alberto Sanchís, Enrique Vidal 0001, Alfons Juan-Císcar
INTERSPEECH4
2003 Median strings for k-nearest neighbour classification
Carlos D. Martínez-Hinarejos, Alfons Juan-Císcar, Francisco Casacuberta
Pattern Recognit. Lett.2
2002 Using Recurrent Neural Networks for Automatic Chromosome Classification
César Ernesto Martínez, Alfons Juan-Císcar, Francisco Casacuberta
ICANN2
2002 On the use of Bernoulli mixture models for text classification
Alfons Juan-Císcar, Enrique Vidal 0001
Pattern Recognit.1
2000 On the Use of Normalized Edit Distances and an Efficient k-NN Search Technique (k-AESA) for Fast and Accurate String Classification
abstract
Classification based on nearest neighbours (NN) is a uniformly good approach to many pattern recognition tasks. However, two important aspects need to be taken into account to actually achieve good performance in practice: 1) the metric or dissimilarity measure adopted to compare the considered patterns; and 2) the computational cost incurred by the NN searching operation. As it is shown in this paper, by using adequate techniques to cope with these two issues, the NN-based classification leads to better results than those obtained by other approaches that have been applied to a task of human banded chromosomes classification.
Alfons Juan-Císcar, Enrique Vidal 0001
ICPR1
2000 Use of Median String for Classification
abstract
A string that minimizes the sum of distances to the strings of a given set is known as (generalized) median string of the set. This concept is important in pattern recognition for modelling a (large) set of garbled strings or patterns. The search of such a string is an NP-Hard problem and, therefore, no efficient algorithms to compute the median strings can be designed. A greedy approach has been proposed to compute an approximate median string of a set of strings. In this work an algorithm is proposed that iteratively improves the approximate solution given above. Experiments have been carried out on synthetic and real data to compare the performances of the approximate median string with the conventional set median. These experiments showed that the proposed median string is a better representation of a given set than the corresponding set median.
Carlos D. Martínez-Hinarejos, Alfons Juan-Císcar, Francisco Casacuberta
ICPR2
1998 Fast k-nearest-neighbours searching through extended versions of the approximating and eliminating search algorithm (AESA)
abstract
The approximating and eliminating search algorithm (AESA) is probably the technique requiring the fewest distance computations for nearest-neighbour searching in general metric spaces. In this paper we propose direct and refined extensions to the AESA for finding k-nearest-neighbours. Results of a number of experiments involving synthetic data are reported, showing that both extensions, and especially the last one, lead to computational savings similar to that of the original (1-NN) AESA.
Alfons Juan-Císcar, Enrique Vidal 0001, Pablo Aibar
ICPR1
1994 Fast K-means-like clustering in metric spaces
Alfons Juan-Císcar, Enrique Vidal 0001
Pattern Recognit. Lett.1