Jorge Civera

dblp:24/4840 · DBLP profile ↗
← Back
34ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0002-0963-0143ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4Human-computer interaction and ubiquitous computing · 3
YearPublicationVenuePosition
2025 LHCP-ASR: An English Speech Corpus of High-Energy Particle Physics Talks for Narrow-Domain ASR Benchmarking
Jaume Santamaria-Jorda, Pablo Segovia-Martínez, Gonçal V. Garcés Díaz-Munío, Joan Albert Silvestre-Cerdà, Adrià Giménez, Rubén Gaspar Aparicio, René Fernández Sánchez, Jorge Civera, Alberto Sanchís, Alfons Juan-Císcar
INTERSPEECH8
2025 Speech translation for multilingual medical education leveraged by large language models
abstract
The application of large language models (LLMs) to speech translation (ST) or, in general, to machine translation (MT) has recently provided excellent results, superseding conventional encoder-decoder MT systems in the general domain. However, this is not clearly the case when LLMs as MT systems are translating medical-related materials. In this respect, the provision of multilingual training materials for oncology professionals is a goal of the EU project Interact-Europe in which this work was framed. To this end, cross-language technology adapted to the oncology domain was developed, evaluated and deployed for multilingual interspecialty medical education. More precisely, automatic speech recognition (ASR) and MT models were adapted to the oncology domain to translate English pre-recorded training videos, kindly provided by the European School of Oncology (ESO), into French, Spanish, German and Slovene. In this work, three categories of MT models adapted to the medical domain were assessed: bilingual encoder-decoder MT models trained from scratch, pre-trained large multilingual encoder-decoder MT models, and multilingual decoder-only LLMs. The experimental results underline the competitiveness in translation quality of LLMs compared to encoder-decoder MT models. Finally, the ESO speech dataset, comprising roughly 1000 videos and 745 h for the training and evaluation of ASR, MT and ST models, was publicly released for the scientific community.
Jorge Iranzo-Sánchez, Jaume Santamaria-Jorda, Gerard Mas-Mollà, Gonçal V. Garcés Díaz-Munío, Javier Iranzo-Sánchez, Javier Jorge, Joan Albert Silvestre-Cerdà, Adrià Giménez, Jorge Civera, Alberto Sanchís, Alfons Juan-Císcar
Artif. Intell. Medicine9
2024 Segmentation-Free Streaming Machine Translation
abstract
Abstract Streaming Machine Translation (MT) is the task of translating an unbounded input text stream in real-time. The traditional cascade approach, which combines an Automatic Speech Recognition (ASR) and an MT system, relies on an intermediate segmentation step which splits the transcription stream into sentence-like units. However, the incorporation of a hard segmentation constrains the MT system and is a source of errors. This paper proposes a Segmentation-Free framework that enables the model to translate an unsegmented source stream by delaying the segmentation decision until after the translation has been generated. Extensive experiments show how the proposed Segmentation-Free framework has better quality-latency trade-off than competing approaches that use an independent segmentation model.1
Javier Iranzo-Sánchez, Jorge Iranzo-Sánchez, Adrià Giménez, Jorge Civera, Alfons Juan-Císcar
Trans. Assoc. Comput. Linguistics4
2022 From Simultaneous to Streaming Machine Translation by Leveraging Streaming History
abstract
Simultaneous Machine Translation is the task of incrementally translating an input sentence before it is fully available.Currently, simultaneous translation is carried out by translating each sentence independently of the previously translated text.More generally, Streaming MT can be understood as an extension of Simultaneous MT to the incremental translation of a continuous input text stream.In this work, a state-of-the-art simultaneous sentencelevel MT system is extended to the streaming setup by leveraging the streaming history.Extensive empirical results are reported on IWSLT Translation Tasks, showing that leveraging the streaming history leads to significant quality gains.In particular, the proposed system proves to compare favorably to the best performing systems.
Javier Iranzo-Sánchez, Jorge Civera, Alfons Juan-Císcar
ACL (1)2
2022 Live Streaming Speech Recognition Using Deep Bidirectional LSTM Acoustic Models and Interpolated Language Models
abstract
Although Long-Short Term Memory (LSTM) networks and deep Transformers are now extensively used in offline ASR, it is unclear how best offline systems can be adapted to work with them under the streaming setup. After gaining considerable experience on this regard in recent years, in this paper we show how an optimized, low-latency streaming decoder can be built in which bidirectional LSTM acoustic models, together with general interpolated language models, can be nicely integrated with minimal perfomance degradation. In brief, our streaming decoder consists of a one-pass, real-time search engine relying on a limited-duration window sliding over time and a number of ad hoc acoustic and language model pruning techniques. Extensive empirical assessment is provided on truly streaming tasks derived from the well-known LibriSpeech and TED talks datasets, as well as from TV shows on a main Spanish broadcasting station.
Javier Jorge, Adrià Giménez, Joan Albert Silvestre-Cerdà, Jorge Civera, Alberto Sanchís, Alfons Juan-Císcar
IEEE ACM Trans. Audio Speech Lang. Process.4
2021 Europarl-ASR: A Large Corpus of Parliamentary Debates for Streaming ASR Benchmarking and Speech Data Filtering/Verbatimization
Gonçal V. Garcés Díaz-Munío, Joan Albert Silvestre-Cerdà, Javier Jorge, Adrià Giménez-Pastor, Javier Iranzo-Sánchez, Pau Baquero-Arnal, Nahuel Roselló, Alejandro Pérez González de Martos, Jorge Civera, Alberto Sanchís, Alfons Juan-Císcar
Interspeech9
2021 Towards Simultaneous Machine Interpretation
abstract
[EN] Automatic speech-to-speech translation (S2S) is one of the most challenging speech and language processing tasks, especially when considering its application to real-time settings. Recent advances on streaming Automatic Speech Recognition (ASR), simultaneous Machine Translation (MT) and incremental neural Text-To-Speech (TTS) make it possible to develop real-time cascade S2S systems with greatly improved accuracy. On the way to simultaneous machine interpretation, a state-of-the-art cascade streaming S2S system is described and empirically assessed in the simultaneous interpretation of European Parliament debates. We pay particular attention to the TTS component, particularly in terms of speech naturalness under a variety of response-time settings, as well as in terms of speaker similarity for its cross-lingual voice cloning capabilities.
Alejandro Pérez González de Martos, Javier Iranzo-Sánchez, Adrià Giménez-Pastor, Javier Jorge, Joan Albert Silvestre-Cerdà, Jorge Civera, Alberto Sanchís, Alfons Juan-Císcar
Interspeech6
2021 Towards cross-lingual voice cloning in higher education
abstract
The rapid progress of modern AI tools for automatic speech recognition and machine translation is leading to a progressive cost reduction to produce publishable subtitles for educational videos in multiple languages. Similarly, text-to-speech technology is experiencing large improvements in terms of quality, flexibility and capabilities. In particular, state-of-the-art systems are now capable of seamlessly dealing with multiple languages and speakers in an integrated manner, thus enabling lecturer’s voice cloning in languages she/he might not even speak. This work is to report the experience gained on using such systems at the Universitat Politècnica de València (UPV), mainly as a guidance for other educational organizations willing to conduct similar studies. It builds on previous work on the UPV’s main repository of educational videos, MediaUPV, to produce multilingual subtitles at scale and low cost. Here, a detailed account is given on how this work has been extended to also allow for massive machine dubbing of MediaUPV. This includes collecting 59 h of clean speech data from UPV’s academic staff, and extending our production pipeline of subtitles with a state-of-the-art multilingual and multi-speaker text-to-speech system trained from the collected data. Our main result comes from an extensive, subjective evaluation of this system by lecturers contributing to data collection. In brief, it is shown that text-to-speech technology is not only mature enough for its application to MediaUPV, but also needed as soon as possible by students to improve its accessibility and bridge language barriers.
Gonçal V. Garcés Díaz-Munío, Adrià Giménez, Joan Albert Silvestre-Cerdà, Alberto Sanchís, Jorge Civera, Manuel Jiménez, Carlos Turro, Alfons Juan-Císcar
Eng. Appl. Artif. Intell.6
2021 Streaming cascade-based speech translation leveraged by a direct segmentation model
abstract
The cascade approach to Speech Translation (ST) is based on a pipeline that concatenates an Automatic Speech Recognition (ASR) system followed by a Machine Translation (MT) system. Nowadays, state-of-the-art ST systems are populated with deep neural networks that are conceived to work in an offline setup in which the audio input to be translated is fully available in advance. However, a streaming setup defines a completely different picture, in which an unbounded audio input gradually becomes available and at the same time the translation needs to be generated under real-time constraints. In this work, we present a state-of-the-art streaming ST system in which neural-based models integrated in the ASR and MT components are carefully adapted in terms of their training and decoding procedures in order to run under a streaming setup. In addition, a direct segmentation model that adapts the continuous ASR output to the capacity of simultaneous MT systems trained at the sentence level is introduced to guarantee low latency while preserving the translation quality of the complete ST system. The resulting ST system is thoroughly evaluated on the real-life streaming Europarl-ST benchmark to gauge the trade-off between quality and latency for each component individually as well as for the complete ST system.
Javier Iranzo-Sánchez, Javier Jorge, Pau Baquero-Arnal, Joan Albert Silvestre-Cerdà, Adrià Giménez, Jorge Civera, Alberto Sanchís, Alfons Juan-Císcar
Neural Networks6
2020 Europarl-ST: A Multilingual Corpus for Speech Translation of Parliamentary Debates
abstract
Current research into spoken language translation (SLT), or speech-to-text translation, is often hampered by the lack of specific data resources for this task, as currently available SLT datasets are restricted to a limited set of language pairs. In this paper we present Europarl-ST, a novel multilingual SLT corpus containing paired audio-text samples for SLT from and into 6 European languages, for a total of 30 different translation directions. This corpus has been compiled using the debates held in the European Parliament in the period between 2008 and 2012. This paper describes the corpus creation process and presents a series of automatic speech recognition, machine translation and spoken language translation experiments that highlight the potential of this new resource. The corpus is released under a Creative Commons license and is freely accessible and downloadable.
Javier Iranzo-Sánchez, Joan Albert Silvestre-Cerdà, Javier Jorge, Nahuel Roselló, Adrià Giménez, Alberto Sanchís, Jorge Civera, Alfons Juan-Císcar
ICASSP7
2020 LSTM-Based One-Pass Decoder for Low-Latency Streaming
abstract
Current state-of-the-art models based on Long-Short Term Memory (LSTM) networks have been extensively used in ASR to improve performance. However, using LSTMs under a streaming setup is not straightforward due to real-time constraints. In this paper we present a novel streaming decoder that includes a bidirectional LSTM acoustic model as well as an unidirectional LSTM language model to perform the decoding efficiently while keeping the performance comparable to that of an off-line setup. We perform a one-pass decoding using a sliding window scheme for a bidirectional LSTM acoustic model and an LSTM language model. This has been implemented and assessed under a pure streaming setup, and deployed into our production systems. We report WER and latency figures for the well-known LibriSpeech and TED-LIUM tasks, obtaining competitive WER results with low-latency responses.
Javier Jorge, Adrià Giménez, Javier Iranzo-Sánchez, Joan Albert Silvestre-Cerdà, Jorge Civera, Alberto Sanchís, Alfons Juan-Císcar
ICASSP5
2020 Improved Hybrid Streaming ASR with Transformer Language Models
Pau Baquero-Arnal, Javier Jorge, Adrià Giménez, Joan Albert Silvestre-Cerdà, Javier Iranzo-Sánchez, Alberto Sanchís, Jorge Civera, Alfons Juan-Císcar
INTERSPEECH7
2019 Real-Time One-Pass Decoder for Speech Recognition Using LSTM Language Models
abstract
Recurrent Neural Networks, in particular Long-Short TermMemory (LSTM) networks, are widely used in Automatic Speech Recognition for language modelling during decoding,usually as a mechanism for rescoring hypothesis. This paperproposes a new architecture to perform real-time one-pass de-coding using LSTM language models. To make decoding ef-ficient, the estimation of look-ahead scores was accelerated byprecomputing static look-ahead tables. These static tables wereprecomputed from a prunedn-gram model, reducing drasti-cally the computational cost during decoding. Additionally,the LSTM language model evaluation was efficiently performedusing Variance Regularization along with a strategy of lazyevaluation. The proposed one-pass decoder architecture wasevaluated on the well-known LibriSpeech and TED-LIUMv3datasets. Results showed that the proposed algorithm obtainsvery competitive WERs with 0.6 RTFs. Finally, our one-passdecoder is compared with a decoupled two-pass decoder.
Javier Jorge, Adrià Giménez, Javier Iranzo-Sánchez, Jorge Civera, Alberto Sanchís, Alfons Juan-Císcar
INTERSPEECH4
2018 Speaker-Adapted Confidence Measures for ASR Using Deep Bidirectional Recurrent Neural Networks
abstract
In the last years, deep bidirectional recurrent neural networks (DBRNN) and DBRNN with long short-term memory cells (DBLSTM) have outperformed the most accurate classifiers for confidence estimation in automatic speech recognition. At the same time, we have recently shown that speaker adaptation of confidence measures using DBLSTM yields significant improvements over non-adapted confidence measures. In accordance with these two recent contributions to the state of the art in confidence estimation, this paper presents a comprehensive study of speaker-adapted confidence measures using DBRNN and DBLSTM models. First, we present new empirical evidences of the superiority of recurrent neural networks (RNN)-based confidence classifiers evaluated over a large speech corpus consisting of the English LibriSpeech and the Spanish poliMedia tasks. Second, we show new results on speaker-adapted confidence measures considering a multitask framework in which RNN-based confidence classifiers trained with LibriSpeech are adapted to speakers of the TED-LIUM corpus. These experiments confirm that speaker-adapted confidence measures outperform their non-adapted counterparts. Last, we describe an unsupervised adaptation method of the acoustic DBLSTM model based on confidence measures that results in better automatic speech recognition performance.
Miguel A. del Agua, Adrià Giménez, Alberto Sanchís, Jorge Civera, Alfons Juan-Císcar
IEEE ACM Trans. Audio Speech Lang. Process.4
2016 ASR Confidence Estimation with Speaker-Adapted Recurrent Neural Networks
Miguel A. del Agua, Santiago Piqueras, Adrià Giménez, Alberto Sanchís, Jorge Civera, Alfons Juan-Císcar
INTERSPEECH5
2015 Efficient Generation of High-Quality Multilingual Subtitles for Video Lecture Repositories
abstract
Video lectures are a valuable educational tool in higher education to support or replace face-to-face lectures in active learning strategies. In 2007 the Universitat Politècnica de València (UPV) implemented its video lecture capture system, resulting in a high quality educational video repository, called poliMedia, with more than 10.000 mini lectures created by 1.373 lecturers. Also, in the framework of the European project transLectures, UPV has automatically generated transcriptions and translations in Spanish, Catalan and English for all videos included in the poliMedia video repository. transLectures’s objective responds to the widely-recognised need for subtitles to be provided with video lectures, as an essential service for non-native speakers and hearing impaired persons, and to allow advanced repository functionalities. Although high-quality automatic transcriptions and translations were generated in transLectures, they were not error-free. For this reason, lecturers need to manually review video subtitles to guarantee the absence of errors. The aim of this study is to evaluate the efficiency of the manual review process from automatic subtitles in comparison with the conventional generation of video subtitles from scratch. The reported results clearly indicate the convenience of providing automatic subtitles as a first step in the generation of video subtitles and the significant savings in time of up to almost 75 % involved in reviewing subtitles.
Juan Daniel Valor Miró, Joan Albert Silvestre-Cerdà, Jorge Civera, Carlos Turro, Alfons Juan-Císcar
EC-TEL3
2015 Efficiency and usability study of innovative computer-aided transcription strategies for video lecture repositories
Juan Daniel Valor Miró, Joan Albert Silvestre-Cerdà, Jorge Civera, Carlos Turro, Alfons Juan-Císcar
Speech Commun.3
2014 Interactive handwriting recognition with limited user effort
Adrià Giménez, Jorge Civera, Alberto Sanchís, Alfons Juan-Císcar
Int. J. Document Anal. Recognit.3
2014 Effective balancing error and user effort in interactive handwriting recognition
Jorge Civera, Alberto Sanchís, Alfons Juan-Císcar
Pattern Recognit. Lett.2
2013 A System Architecture to Support Cost-Effective Transcription and Translation of Large Video Lecture Repositories
abstract
Online video lecture repositories are rapidly growing and becoming established as fundamental knowledge assets. However, most lectures are neither transcribed nor translated because of the lack of cost-effective solutions that can give accurate enough results. In this paper, we describe a system architecture that supports the cost-effective transcription and translation of large video lecture repositories. This architecture has been adopted in the EU project transLectures and is now being tested on a repository of more than 9000 video lectures at the Universitat Politecnica de Valencia. Following a brief description of this repository and of the transLectures project, we describe the proposed system architecture in detail. We also report empirical results on the quality of the transcriptions and translations currently being maintained and steadily improved.
Joan Albert Silvestre-Cerdà, Alejandro Pérez González de Martos, Manuel Jiménez, Carlos Turro, Alfons Juan-Císcar, Jorge Civera
SMC6
2012 Explicit length modelling for statistical machine translation
Joan Albert Silvestre-Cerdà, Jesús Andrés-Ferrer, Jorge Civera
Pattern Recognit.3
2010 The APP Oracle - An Interactive Student Competition on Pattern Recognition
Alfons Juan-Císcar, Jesús Andrés-Ferrer, Adrià Giménez, Jorge Civera, Roberto Paredes, Enrique Vidal 0001
CSEDU (2)4
2010 Saturnalia: A Latin-Catalan Parallel Corpus for Statistical MT
Jesús González-Rubio, Jorge Civera, Alfons Juan-Císcar, Francisco Casacuberta
LREC2
2009 Statistical Approaches to Computer-Assisted Translation
abstract
Current machine translation (MT) systems are still not perfect. In practice, the output from these systems needs to be edited to correct errors. A way of increasing the productivity of the whole translation process (MT plus human work) is to incorporate the human correction activities within the translation process itself, thereby shifting the MT paradigm to that of computer-assisted translation. This model entails an iterative process in which the human translator activity is included in the loop: In each iteration, a prefix of the translation is validated (accepted or amended) by the human and the system computes its best (or n-best) translation suffix hypothesis to complete this prefix. A successful framework for MT is the so-called statistical (or pattern recognition) framework. Interestingly, within this framework, the adaptation of MT systems to the interactive scenario affects mainly the search process, allowing a great reuse of successful techniques and models. In this article, alignment templates, phrase-based models, and stochastic finite-state transducers are used to develop computer-assisted translation systems. These systems were assessed in a European project (TransType2) in two real tasks: The translation of printer manuals; manuals and the translation of the Bulletin of the European Union. In each task, the following three pairs of languages were involved (in both translation directions): English-Spanish, English-German, and English-French.
Sergio Barrachina 0001, Oliver Bender, Francisco Casacuberta, Jorge Civera, Elsa Cubel, Shahram Khadivi, Antonio L. Lagarda, Hermann Ney, Jesús Tomás, Enrique Vidal 0001, Juan Miguel Vilar
Comput. Linguistics4
2008 Improving Interactive Machine Translation via Mouse Actions
Germán Sanchis-Trilles, Daniel Ortiz-Martínez, Jorge Civera, Francisco Casacuberta, Enrique Vidal 0001, Hieu Hoang
EMNLP3
2008 Bilingual Text Classification using the IBM 1 Translation Model
Jorge Civera, Alfons Juan-Císcar
LREC1
2006 Mixtures of IBM Model 2
Jorge Civera, Alfons Juan-Císcar
EAMT1
2006 A Computer-Assisted Translation Tool based on Finite-State Technology
Jorge Civera, Antonio L. Lagarda, Elsa Cubel, Francisco Casacuberta, Enrique Vidal 0001, Juan Miguel Vilar, Sergio Barrachina 0001
EAMT1
2006 Bilingual Machine-Aided Indexing
Jorge Civera, Alfons Juan-Císcar
LREC1
2006 Computer-assisted translation using speech recognition
abstract
Current machine translation systems are far from being perfect. However, such systems can be used in computer-assisted translation to increase the productivity of the (human) translation process. The idea is to use a text-to-text translation system to produce portions of target language text that can be accepted or amended by a human translator using text or speech. These user-validated portions are then used by the text-to-text translation system to produce further, hopefully improved suggestions. There are different alternatives of using speech in a computer-assisted translation system: From pure dictated translation to simple determination of acceptable partial translations by reading parts of the suggestions made by the system. In all the cases, information from the text to be translated can be used to constrain the speech decoding search space. While pure dictation seems to be among the most attractive settings, unfortunately perfect speech decoding does not seem possible with the current speech processing technology and human error-correcting would still be required. Therefore, approaches that allow for higher speech recognition accuracy by using increasingly constrained models in the speech recognition process are explored here. All these approaches are presented under the statistical framework. Empirical results support the potential usefulness of using speech within the computer-assisted translation paradigm.
Enrique Vidal 0001, Francisco Casacuberta, Luis Rodríguez, Jorge Civera, Carlos D. Martínez-Hinarejos
IEEE Trans. Speech Audio Process.4
2005 On the use of speech recognition in computer assisted translation
Luis Rodríguez, Jorge Civera, Enrique Vidal 0001, Francisco Casacuberta, César Ernesto Martínez
INTERSPEECH2
2005 Text Mining Biomedical Literature for Discovering Gene-to-Gene Relationships: A Comparative Study of Algorithms
abstract
Partitioning closely related genes into clusters has become an important element of practically all statistical analyses of microarray data. A number of computer algorithms have been developed for this task. Although these algorithms have demonstrated their usefulness for gene clustering, some basic problems remain. This paper describes our work on extracting functional keywords from MEDLINE for a set of genes that are isolated for further study from microarray experiments based on their differential expression patterns. The sharing of functional keywords among genes is used as a basis for clustering in a new approach called BEA-PARTITION in this paper. Functional keywords associated with genes were extracted from MEDLINE abstracts. We modified the Bond Energy Algorithm (BEA), which is widely accepted in psychology and database design but is virtually unknown in bioinformatics, to cluster genes by functional keyword associations. The results showed that BEA-PARTITION and hierarchical clustering algorithm outperformed k-means clustering and self-organizing map by correctly assigning 25 of 26 genes in a test set of four known gene groups. To evaluate the effectiveness of BEA-PARTITION for clustering genes identified by microarray profiles, 44 yeast genes that are differentially expressed during the cell cycle and have been widely studied in the literature were used as a second test set. Using established measures of cluster quality, the results produced by BEA-PARTITION had higher purity, lower entropy, and higher mutual information than those produced by k-means and self-organizing map. Whereas BEA-PARTITION and the hierarchical clustering produced similar quality of clusters, BEA-PARTITION provides clear cluster boundaries compared to the hierarchical clustering. BEA-PARTITION is simple to implement and provides a powerful approach to clustering genes or to any clustering problem where starting matrices are available from experimental observations.
Ying Liu 0007, Shamkant B. Navathe, Jorge Civera, Venu Dasigi, Ashwin Ram 0001, Brian J. Ciliax, Ray Dingledine
IEEE ACM Trans. Comput. Biol. Bioinform.3
2004 Finite-State Models for Computer Assisted Translation
Elsa Cubel, Jorge Civera, Juan Miguel Vilar, Antonio L. Lagarda, Francisco Casacuberta, Enrique Vidal 0001, David Picó, Luis Rodríguez
ECAI2
2004 From Machine Translation to Computer Assisted Translation using Finite-State Models
Jorge Civera, Elsa Cubel, Antonio L. Lagarda, David Picó, Enrique Vidal 0001, Francisco Casacuberta, Juan Miguel Vilar, Sergio Barrachina 0001
EMNLP1