Fernando Batista

dblp:07/6052 · DBLP profile ↗
← Back
37ranked-venue papers
9as first author
6since 2021 · last 2024
0000-0002-1075-0177ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 33 · 7 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 6 first-authorDatabases, data management, data science and information retrieval · 7 · 3 since 2021
YearPublicationVenuePosition
2024 Counter Hate Speech Detection in Youtube Conversations
Pedro Fialho, Ricardo Ribeiro 0001, Fernando Batista, Gil Ramos, António Fonseca, Sérgio Moro, Rita Guerra, Paula Carvalho 0001, Catarina Marques, Cláudia Silva 0002
IPMU (3)3
2024 Unveiling Patterns of Hate Speech in the Portuguese Sphere: A Social Network Analysis Approach
Catarina Pontes, António Fonseca, Sérgio Moro, Fernando Batista, Ricardo Ribeiro 0001, Catarina Marques, Paula Carvalho 0001, Cláudia Silva 0002, Rita Guerra
IPMU (3)4
2023 Semantic similarity for mobile application recommendation under scarce user data
abstract
The More Like This recommendation approach is ubiquitous in multiple domains and consists in recommending items similar to the one currently selected by the user, being particularly relevant when user data is scarce. We studied the impact of using semantic similarity in the context of the More Like This recommendation for mobile applications, by leveraging dense representations in order to infer the similarity between applications, based on their textual fields. Our approach was validated by comparing it to the solution currently in use by Aptoide, a mobile application store, since no benchmarks are available for this specific task. To further evaluate the proposed model, we asked 1262 users to compare the results achieved by both approaches, also allowing us to build an annotated dataset of similar applications. Results show that the semantic representations are able to capture the context of the applications, with more useful recommendations being presented to users, when compared to Aptoide’s current solution. For replication and future research, all the code and data used in this study was made publicly available, including two novel datasets (installed applications for more than one million users, and app user-labeled similarity), the fine-tuned model, and the test platform.
João Coelho, Diogo Mano, Beatriz Paula, Carlos Coutinho, João Oliveira 0001, Ricardo Ribeiro 0001, Fernando Batista
Eng. Appl. Artif. Intell.7
2022 Early Experiments on Automatic Annotation of Portuguese Medieval Texts
Maria Inês Bico, Jorge Baptista, Fernando Batista, Esperança Cardeira
TPDL3
2022 Hate Speech Dynamics Against African descent, Roma and LGBTQI Communities in Portugal
abstract
This paper introduces FIGHT, a dataset containing 63,450 tweets, posted before and after the official declaration of Covid-19 as a pandemic by online users in Portugal. This resource aims at contributing to the analysis of online hate speech targeting the most representative minorities in Portugal, namely the African descent and the Roma communities, and the LGBTQI community, the most commonly reported target of hate speech in social media at the European context. We present the methods for collecting the data, and provide insightful statistics on the distribution of tweets included in FIGHT, considering both the temporal and spatial dimensions. We also analyze the availability over time of tweets targeting the above-mentioned communities, distinguishing public, private and deleted tweets. We believe this study will contribute to better understand the dynamics of online hate speech in Portugal, particularly in adverse contexts, such as a pandemic outbreak, allowing the development of more informed and accurate hate speech resources for Portuguese.
Paula Carvalho 0001, Bernardo Cunha Matos, Raquel Bento Santos, Fernando Batista, Ricardo Ribeiro 0001
LREC4
2021 Towards better subtitles: A multilingual approach for punctuation restoration of speech transcripts
Nuno Miguel Guerreiro, Ricardo Rei, Fernando Batista
Expert Syst. Appl.3
2020 Using Topic Information to Improve Non-exact Keyword-Based Search for Mobile Applications
Eugénio Ribeiro, Ricardo Ribeiro 0001, Fernando Batista, João Oliveira 0001
IPMU (1)3
2020 Creating Classification Models from Textual Descriptions of Companies Using Crunchbase
Marco Felgueiras, Fernando Batista, João Paulo Carvalho 0001
IPMU (1)2
2020 Automatic Truecasing of Video Subtitles Using BERT: A Multilingual Adaptable Approach
Ricardo Rei, Nuno Miguel Guerreiro, Fernando Batista
IPMU (1)3
2018 Acoustic-prosodic Entrainment in Structural Metadata Events
abstract
This paper presents an acoustic-prosodic analysis of entrain- ment in a Portuguese map-task corpus. Our aim is to ana- lyze how turn-by-turn entrainment varies with distinct structural metadata events: types of sentence-like units (SU) in consecu- tive turns (e.g. interrogatives followed by declaratives, or both declaratives), and with the presence of discourse markers, affir- mative cue words, and disfluencies in the beginning of turns. Entrainment at turn-exchanges may be observed in terms of pitch, energy, duration, and voice quality. Regarding SU types, question-answer turns are the ones with stronger similarity, and declarative-interrogative pairs are the ones where less entrain- ment occurs, as expected. Moreover, in question-answer pairs, there is also stronger evidence of entrainment with Yes/No and Tag questions than with Wh- questions. In fact, these subtypes are coded in distinctive prosodic ways (moreover, the first sub- type has no associated lexical-syntactic cues in Portuguese, only prosodic). As for turn-initial structures, entrainment is stronger when the second turn begins with an affirmative cue word; less strong with ambiguous structures (such as ‘OK’), emphatic af- firmative answers, and negative answers; and scarce with dis- fluencies and discourse markers. The different degrees of local entrainment may be related with the informative structure of distinct structural metadata events.
Vera Cabarrão, Fernando Batista, Helena Moniz, Isabel Trancoso, Ana Isabel Mata
INTERSPEECH2
2017 Detecting relevant tweets in very large tweet collections: The London Riots case study
abstract
In this paper we propose to approach the subject of detecting relevant tweets when in the presence of very large tweet collections containing a large number of different trending topics. We use a large database of tweets collected during the 2011 London Riots as a case study to demonstrate the application of the proposed techniques. In order to extract relevant content, we extend, formalize and apply a recent technique, called Twitter Topic Fuzzy Fingerprints, which, in the scope of social media, outperforms other well known text based classification methods, while being less computationally demanding, an essential feature when processing large volumes of streaming data. Using this technique we were able to detect 45% additional relevant tweets within the database.
João Paulo Carvalho 0001, Hugo Rosa, Fernando Batista
FUZZ-IEEE3
2017 A Semi-Supervised Learning Approach for Acoustic-Prosodic Personality Perception in Under-Resourced Domains
abstract
Automatic personality analysis has gained attention in the last years as a fundamental dimension in human-To-human and human-To-machine interaction. However, it still suffers from limited number and size of speech corpora for specific domains, such as the assessment of children's personality. This paper investigates a semi-supervised training approach to tackle this scenario. We devise an experimental setup with age and language mismatch and two training sets: A small labeled training set from the Interspeech 2012 Personality Sub-challenge, containing French adult speech labeled with personality OCEAN traits, and a large unlabeled training set of Portuguese children's speech. As test set, a corpus of Portuguese children's speech labeled with OCEAN traits is used. Based on this setting, we investigate a weak supervision approach that iteratively refines an initial model trained with the labeled data-set using the unlabeled data-set. We also investigate knowledge-based features, which leverage expert knowledge in acoustic-prosodic cues and thus need no extra data. Results show that, despite the large mismatch imposed by language and age differences, it is possible to attain improvements with these techniques, pointing both to the benefits of using a weak supervision and expert-based acoustic-prosodic features across age and language.
Rubén Solera-Ureña, Helena Moniz, Fernando Batista, Vera Cabarrão, Anna Pompili, Ramón Fernandez Astudillo, Joana Campos 0001, Ana Paiva 0001, Isabel Trancoso
INTERSPEECH3
2017 MISNIS: An intelligent platform for twitter topic mining
João Paulo Carvalho 0001, Hugo Rosa, Gaspar Brogueira, Fernando Batista
Expert Syst. Appl.4
2016 Creating Extended Gender Labelled Datasets of Twitter Users
Marco Vicente, Fernando Batista, João Paulo Carvalho 0001
IPMU (2)2
2016 SPA: Web-based Platform for easy Access to Speech Processing Modules
Fernando Batista, Pedro Curto, Isabel Trancoso, Alberto Abad, Jaime Ferreira, Eugénio Ribeiro, Helena Moniz, David Martins de Matos, Ricardo Ribeiro 0001
LREC1
2015 Text based classification of companies in CrunchBase
abstract
This paper introduces two fuzzy fingerprint based text classification techniques that were successfully applied to automatically label companies from CrunchBase, based purely on their unstructured textual description. This is a real and very challenging problem due to the large set of possible labels (more than 40) and also to the fact that the textual descriptions do not have to abide by any criteria and are, therefore, extremely heterogeneous. Fuzzy fingerprints are a recently introduced technique that can be used for performing fast classification. They perform well in the presence of unbalanced datasets and can cope with a very large number of classes. In the paper, a comparison is performed against some of the best text classification techniques commonly used to address similar problems. When applied to the CrunchBase dataset, the fuzzy fingerprint based approach outperformed the other techniques.
Fernando Batista, João Paulo Carvalho 0001
FUZZ-IEEE1
2015 Twitter gender classification using user unstructured information
abstract
This paper describes an approach to automatically detect the gender of Twitter users, based only on clues provided by their profile information in an unstructured form. A number of features that capture phenomena specific of Twitter users is proposed and evaluated on a dataset of about 242K English language users. Different supervised and unsupervised approaches are used to assess the performance of the proposed features, including Naive Bayes variants, Logistic Regression, Support Vector Machines, Fuzzy c-Means clustering, and K-means. An unsupervised approach based on Fuzzy c-Means proved to be very suitable for this task, returning the correct gender for about 96% of the users.
Marco Vicente, Fernando Batista, João Paulo Carvalho 0001
FUZZ-IEEE2
2015 Detecting repetitions in spoken dialogue systems using phonetic distances
abstract
This paper addresses the problem of automatic detection of re-peated turns in Spoken Dialogue Systems. Repetitions can be a symptom of problematic communication between users and systems. Such repetitions are often due to speech recognition errors, which in turn makes it hard to use speech recognition to detect repetitions. We present an approach to detect rep-etition using the phonetic distance to find the best alignment between turns in the same dialogue. The alignment score ob-tained is combined with different features to improve repeti-tion detection. To evaluate the method proposed we compare several alignment techniques from edit distance to DTW-based distance, previously used in Spoken-Term detection tasks. We also compare two different methods to compute the phonetic distance: the first one using the phoneme sequence, and the second one using the distance between the phone posterior vec-tors. Two different datasets were used in this evaluation: a bus-schedule information system (in English) and a call routing system (in Swedish). The results show that approaches using phoneme distances over-perform approaches using Levenshtein distances between ASR outputs for repetition detection. Index Terms: spoken dialogue systems, repetition detection, phonetic distance
José Lopes 0001, Giampiero Salvi, Gabriel Skantze, Alberto Abad, Joakim Gustafson, Fernando Batista, Raveesh Meena, Isabel Trancoso
INTERSPEECH6
2015 Combining multiple approaches to predict the degree of nativeness
abstract
Automatic speaker nativeness assessment has multiple applications, such as second language learning and IVR systems. In this paper we view this as a regression problem, since the available labels are on a continuous scale. Multiple approaches were applied, such as phonotactic models, i-vectors, and goodness of pronunciation, covering both segmental and suprasegmental features. Different phonotactic models were adopted, either trained with the challenge data, or using additional multilingual data from other domains. The obtained values were later combined in multiple ways and fed to a support vector machine regressor. Results on the test set surpass the provided baseline and are in line with the results obtained on the remaining sets. This suggests that our models generalize well to other datasets
Eugénio Ribeiro, Jaime Ferreira, Julia Olcoz, Alberto Abad, Helena Moniz, Fernando Batista, Isabel Trancoso
INTERSPEECH6
2014 Twitter Topic Fuzzy Fingerprints
abstract
In this paper we propose to approach the subject of Twitter Topic Detection using a new technique called Topic Fuzzy Fingerprints. A comparison is made with two popular text classification techniques, Support Vector Machines (SVM) and fc-Nearest Neighbours (fcNN). Preliminary results show that Twitter Topic Fuzzy Fingerprints outperforms the other two techniques achieving better Precision and Recall, while still being much faster, which is an essential feature when processing large volumes of streaming data.
Hugo Rosa, Fernando Batista, João Paulo Carvalho 0001
FUZZ-IEEE2
2014 OpenLogos Semantico-Syntactic Knowledge-Rich Bilingual Dictionaries
Anabela Barreiro, Fernando Batista, Ricardo Ribeiro 0001, Helena Moniz, Isabel Trancoso
LREC2
2014 Linguistic Evaluation of Support Verb Constructions by OpenLogos and Google Translate
Anabela Barreiro, Johanna Monti, Brigitte Orliac, Susanne Preuß, Kutz Arrieta, Wang Ling, Fernando Batista, Isabel Trancoso
LREC7
2014 Revising the annotation of a Broadcast News corpus: a linguistic approach
Vera Cabarrão, Helena Moniz, Fernando Batista, Ricardo Ribeiro 0001, Nuno J. Mamede, Hugo Meinedo, Isabel Trancoso, Ana Isabel Mata, David Martins de Matos
LREC3
2014 Teenage and adult speech in school context: building and processing a corpus of European Portuguese
Ana Isabel Mata, Helena Moniz, Fernando Batista, Julia Hirschberg
LREC3
2014 Prosodic, syntactic, semantic guidelines for topic structures across domains and corpora
Ana Isabel Mata, Helena Moniz, Telmo Móia, Anabela Gonçalves, Fátima Silva, Fernando Batista, Inês Duarte, Fátima De Cassia E. Oliveira, Isabel Falé
LREC6
2014 Speaking style effects in the production of disfluencies
Helena Moniz, Fernando Batista, Ana Isabel Mata, Isabel Trancoso
Speech Commun.2
2013 Disfluency detection based on prosodic features for university lectures
abstract
This paper focuses on the identification of disfluent sequences and their distinct structural regions, based on acoustic and prosodic features. Reported experiments are based on a corpus of university lectures in European Portuguese, with roughly 32h, and a relatively high percentage of disfluencies (7.6%). The set of features automatically extracted from the corpus proved to be discriminant of the regions contained in the production of a disfluency. Several machine learning methods have been applied, but the best results were achieved using Classification and Regression Trees (CART). The set of features which was most informative for cross-region identification encompasses word duration ratios, word confidence score, silent ratios, and pitch and energy slopes. Features such as the number of phones and syllables per word proved to be more useful for the identification of the interregnum, whereas energy slopes were most suited for identifying the interruption point.
Henrique Medeiros, Helena Moniz, Fernando Batista, Isabel Trancoso, Luís Nunes
INTERSPEECH3
2012 A critical survey on the use of Fuzzy Sets in Speech and Natural Language Processing
abstract
This paper shows how the use and applications of Fuzzy Sets (FS) in Speech and Natural Language Processing (SNLP) have seen a steady decline to a point where FS are virtually unknown or unappealing for most of the researchers currently working in the SNLP field, tries to find the reasons behind this decline, and proposes some guidelines on what could be done to reverse it and make FS assume a relevant role in SNLP.
João Paulo Carvalho 0001, Fernando Batista, Luísa Coheur
FUZZ-IEEE2
2012 Prosodic contex-based analysis of disfluencies
abstract
This work explores prosodic cues of disfluencies in a corpus of university lectures. Results show three significant (p < 0.001) trends: pitch and energy slopes are significantly different between the disfluency and the onset of fluency; those features are also relevant to disfluency type differentiation; and they do not seem to be a speakereffect. The best combination of linguistic features one can use to better predict the onset of fluency are pitch and energy resets as well as the presence of a silent pause immediately before a repair. Our results, thus, point out to a strategy of prosodic contrast rather than of parallelism. With this work we hope to contribute to the analysis of the prosodic behaviors in the production of the so called disfluencies and in the fluency repair in European Portuguese.
Helena Moniz, Fernando Batista, Isabel Trancoso, Ana Isabel Mata
INTERSPEECH2
2012 Bilingual Experiments on Automatic Recovery of Capitalization and Punctuation of Automatic Speech Transcripts
abstract
This paper focuses on the tasks of recovering capitalization and punctuation marks from texts without that information, such as spoken transcripts, produced by automatic speech recognition systems. These two practical rich transcription tasks were performed using the same discriminative approach, based on maximum entropy, suitable for on-the-fly usage. Reported experiments were conducted both over Portuguese and English broadcast news data. Both force aligned and automatic transcripts were used, allowing to measure the impact of the speech recognition errors. Capitalized words and named entities are intrinsically related, and are influenced by time variation effects. For that reason, the so-called language dynamics have been addressed for the capitalization task. Language adaptation results indicate, for both languages, that the capitalization performance is affected by the temporal distance between the training and testing data. In what regards the punctuation task, this paper covers the three most frequent punctuation marks: full stop, comma, and question marks. Different methods were explored for improving the baseline results for full stop and comma. The first uses punctuation information extracted from large written corpora. The second applies different levels of linguistic structure, including lexical, prosodic, and speaker related features. The comma detection improved significantly in the first method, thus indicating that it depends more on lexical features. The second method provided even better results, for both languages and both punctuation marks, best results being achieved mainly for full stop. As for question marks, there is a small gain, but differences are not very significant, due to the relatively small number of question marks in the corpora.
Fernando Batista, Helena Moniz, Isabel Trancoso, Nuno J. Mamede
IEEE Trans. Speech Audio Process.1
2010 Extending the punctuation module for european portuguese
abstract
This paper describes our recent work on extending the punctuation module of automatic subtitles for Portuguese Broadcast News. The main improvement was achieved by the use of prosodic information. This enabled the extension of the previous module which covered only full stops and commas, to cover question marks as well. The approach uses lexical, acoustic and prosodic information. Our results show that the latter is relevant for all types of punctuation. An analysis of the results also shows what type of interrogative is better dealt with by our method, taking into account the specificities of Portuguese. This may lead to different results for different types of corpora, depending on the types of interrogatives that are more frequent.
Fernando Batista, Helena Moniz, Isabel Trancoso, Hugo Meinedo, Ana Isabel Mata, Nuno J. Mamede
INTERSPEECH1
2009 Comparing automatic rich transcription for Portuguese, Spanish and English Broadcast News
abstract
This paper describes and evaluates a language independent approach for automatically enriching the speech recognition output with punctuation marks and capitalization information. The two tasks are treated as two classification problems, using a maximum entropy modeling approach, which achieves results within state-of-the-art. The language independence of the approach is attested with experiments conducted on Portuguese, Spanish and English broadcast news corpora. This paper provides the first comparative study between the three languages, concerning these tasks.
Fernando Batista, Isabel Trancoso, Nuno J. Mamede
ASRU1
2008 The impact of language dynamics on the capitalization of broadcast news
abstract
This paper investigates the impact of language dynamics on the capitalization of transcriptions of broadcast news. Most of the capitalization information is provided by a large newspaper corpus. Three different speech corpora subsets, from different time periods, are used for evaluation, assessing the importance of available training data in nearby time periods. Results are provided both for manual and automatic transcriptions, showing also the impact of the recognition errors in the capitalization task. Our approach is based on maximum entropy models, uses unlimited vocabulary, and is suitable for language adaptation. The language model for a given language period is produced by retraining a previous language model with data from that time period. The language model produced with this approach can be sorted and then pruned, in order to reduce computational resources, without much impact in the final results.
Fernando Batista, Nuno J. Mamede, Isabel Trancoso
INTERSPEECH1
2008 Impact of dynamic model adaptation beyond speech recognition
abstract
The application of speech recognition to live subtitling of Broadcast News has motivated the adaptation of the lexical and language models of the recognizer on a daily basis with text material retrieved from online newspapers. This paper studies the impact of this adaptation on two of the blocks following the speech recognition module: capitalization and topic indexation. We describe and evaluate different adaptation approaches that try to explore the language dynamics.
Fernando Batista, Rui Amaral, Isabel Trancoso, Nuno J. Mamede
SLT1
2008 Recovering capitalization and punctuation marks for automatic speech recognition: Case study for Portuguese broadcast news
Fernando Batista, Diamantino Caseiro, Nuno J. Mamede, Isabel Trancoso
Speech Commun.1
2007 Recovering punctuation marks for automatic speech recognition
abstract
This paper shows results of recovering punctuation over speech transcriptions for a Portuguese broadcast news corpus. The approach is based on maximum entropy models and uses word, part-of-speech, time and speaker information. The contribution of each type of feature is analyzed individually. Separate results for each focus condition are given, making it possible to analyze the differences of performance between planned and spontaneous speech. Index Terms: rich transcription, punctuation recovery, sentence boundary detection, maximum entropy.
Fernando Batista, Diamantino Caseiro, Nuno J. Mamede, Isabel Trancoso
INTERSPEECH1
2000 Some Language Resources and Tools for Computational Processing of Portuguese at INESC
Luzia Wittmann, Ricardo Ribeiro 0001, Tânia Pêgo, Fernando Batista
LREC4