Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Sérgio Paulo

dblp:41/6096 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
2since 2021 · last 2026
0000-0003-0840-4985ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 4 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 1 since 2021Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Robot manipulation · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Embedded and real-time systems · 50% Distributed systems · 50%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Robot manipulation
industrial robot
0.012003
High efficient robotic de-palletizing system for the non-flat ceramic industry · ICRA 2003
Distributed systems › distributed system architecture
distributed software architecture
0.012003
High efficient robotic de-palletizing system for the non-flat ceramic industry · ICRA 2003
Embedded and real-time systems
industrial control systems
0.012003
High efficient robotic de-palletizing system for the non-flat ceramic industry · ICRA 2003

Methods — techniques the papers use, named apart from their topics

industrial networks · 0.1PLC control · 0.1
YearPublicationVenuePosition
2026 FalAR: A Large-scale Speaker-Annotated European Portuguese Speech Corpus of Parliamentary Sessions
Francisco Teixeira, Carlos Carvalho 0003, Mariana Julião, Catarina Botelho, Rubén Solera-Ureña, Sérgio Paulo, Thomas Rolland, Ben Peters, Isabel Trancoso, Alberto Abad
LREC6
2025 CAMÕES: A Comprehensive Automatic Speech Recognition Benchmark for European Portuguese
abstract
Existing resources for Automatic Speech Recognition in Portuguese are mostly focused on Brazilian Portuguese, leaving European Portuguese (EP) and other varieties underexplored. To bridge this gap, we introduce CAMÕES, the first open framework for EP and other Portuguese varieties. It consists of (1) a comprehensive evaluation benchmark, including 46 h of EP test data spanning multiple domains; and (2) a collection of state-of-the-art models. For the latter, we consider multiple foundation models, evaluating their zero-shot and fine-tuned performances, as well as E-Branchformer models trained from scratch. A curated set of $\mathbf{4 2 5 h}$ of EP was used for both fine-tuning and training. Our results show comparable performance for EP between fine-tuned foundation models and the E-Branchformer. Furthermore, the best-performing models achieve relative improvements above 35% WER, compared to the strongest zero-shot foundation model, establishing a new state-of-the-art for EP and other varieties.
Carlos Carvalho 0003, Francisco Teixeira, Catarina Botelho, Anna Pompili, Rubén Solera-Ureña, Sérgio Paulo, Mariana Julião, Thomas Rolland, John Mendonça, Diogo A. P. Nunes, Isabel Trancoso, Alberto Abad
ASRU6
2016 Automating live and batch subtitling of multimedia contents for several European languages
Aitor Álvarez 0001, Carlos Mendes, Matteo Raffaelli, Tiago Luís, Sérgio Paulo, Nicola Piccinini, Haritz Arzelus, Joao P. Neto, Carlo Aliprandi, Arantza del Pozo
Multim. Tools Appl.5
2014 SAVAS: Collecting, Annotating and Sharing Audiovisual Language Resources for Automatic Subtitling
Arantza del Pozo, Carlo Aliprandi, Aitor Álvarez 0001, Carlos Mendes, Joao P. Neto, Sérgio Paulo, Nicola Piccinini, Matteo Raffaelli
LREC6
2008 Methodologies for Designing and Recording Speech Databases for Corpus Based Synthesis
Luís C. Oliveira, Sérgio Paulo, Luís Figueira, Carlos Mendes, Joaquim Godinho
LREC2
2007 MuLAS: a framework for automatically building multi-tier corpora
abstract
The Multi-Level Alignment System (MuLAS) is the L2F tool for building multi-tier speech corpora with reduced or no human intervention at all. MuLAS automatically combines information coming from external speech annotations, human or machine-generated, with the text-based utterance descriptions that it creates, in order to build more reliable and complete descriptions of the spoken utterances. This paper presents our methods for multi-tier annotation synchronization, which lie behind the MuLAS operation. Such methods have allowed us to expand the building of multi-tier corpora to new languages without spending too much effort. MuLAS has been successfully applied to the building of multitier corpora for speech synthesis in American and British English, European Portuguese and German. Natural prosody generation has benefited from MuLAS, too, since prosodic models can be derived from corpora built by MuLAS.
Sérgio Paulo, Luís C. Oliveira
INTERSPEECH1
2005 Reducing the corpus-based TTS signal degradation due to speaker's word pronunciations
Sérgio Paulo, Luís C. Oliveira
INTERSPEECH1
2005 Generation of word alternative pronunciations using weighted finite state transducers
abstract
This paper describes a speech segmentation tool allowing alternative word pronunciations within a WFST framework. Two approaches to word pronunciation graph generation were developed and evaluated. The first approach is grapheme-based where each grapheme is converted into all the phones it can give rise to, in the form of a WFST. Word graphs are obtained by concatenating all grapheme WFSTs. In the second approach, a training corpus is used to find the different realizations of the syllable. This information is used to generate alternative syllable-level pronunciations, represented as WFSTs, that are concatenated to produce the word graphs. Both approaches were evaluated by aligning the phone sequence generated by each approach with the manually labelled phone sequence for all utterance in the corpus. This alignment was used for computing F-rate values for each phone. The syllable-based approach produced the best results.
Sérgio Paulo, Luís C. Oliveira
INTERSPEECH1
2003 High efficient robotic de-palletizing system for the non-flat ceramic industry
abstract
In this paper we describe a system designed to be used on a shop floor of a factory that produces non-flat ceramic products, where an extensive mixture of human and automatic labour is present. Today manufacturing setups rely increasingly on technology. It is common to have all sources of equipment on the shop floor, commanded by industrial PCs or PLCs connected by an industrial network to other factory resources. Also, the production systems are becoming more and more autonomous requiring less operator intervention in every day normal operation. That means using computers for controlling and supervision of the production systems, industrial networks and distributed software architectures. It means also designing application software that is really distributed in the shop floor, taking advantage of the flexibility installed by using programmable equipment.
J. Norberto Pires, Sérgio Paulo
ICRA2
2003 DTW-based phonetic alignment using multiple acoustic features
abstract
This paper presents the results of our effort in improving the accuracy of a DTW-based automatic phonetic aligner. The adopted model assumes that the phonetic segment sequence is already known and so the goal is only to align the spoken utterance with a reference synthetic signal produced by waveform concatenation without prosodic modifications. Instead of using a single acoustic measure to compute the alignment cost function, our strategy uses a combination of acoustic features depending on the pair of phonetic segment classes being aligned. The results show that this strategy considerably reduces the segment boundary location errors, even when aligning synthetic and natural speech signals of different gender speakers.
Sérgio Paulo, Luís C. Oliveira
INTERSPEECH1