VLDB 2026 Research / reviewers in the wild / expert
João Paulo Carvalho 0001
dblp:98/4980 · also João P. Carvalho 0001
· DBLP profile ↗
51ranked-venue papers
11as first author
9since 2021 · last 2026
0000-0003-0005-8299ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 47 · 11 first-author · 9 since 2021Databases, data management, data science and information retrieval · 9 · 2 since 2021Systems, architecture and hardware · 2Computer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Genetic algorithms as an optimization strategy to reduce the complexity of fuzzy fingerprints from large language modelsabstractLarge Language Models currently represent the forefront of Natural Language Processing for classification tasks. Recent studies have integrated Fuzzy Fingerprints as a novel classification layer within these models to enhance result interpretability and reduce overall model complexity, without substantially compromising performance. However, in more challenging settings, this framework requires a larger fingerprint size to achieve competitive results when compared to using the full classification model. In this work, we employ Genetic Algorithms to the Fuzzy Fingerprint framework to further optimize these fingerprints to surpass the performance of conventional large pre-trained classifiers, while also offering improvements in interpretability and reduced complexity. Our empirical analysis demonstrates that optimizing to a smaller fingerprint size not only improves interpretability but also, when compared to a baseline with a larger fingerprint size, delivers higher performance relative to leading-edge methodologies. João Paulo Carvalho 0001, Luísa Coheur |
Fuzzy Sets Syst. | 2 |
| 2025 | Towards Cyberbullying Detection: Building, Benchmarking and Longitudinal Analysis of Aggressiveness and Conflicts/Attacks Datasets From TwitterabstractOffense and hate speech are a source of online conflicts which have become common in social media and, as such, their study is a growing topic of research in machine learning and natural language processing. This article presents two Portuguese language offense-related datasets that deepen the study of the subject: an Aggressiveness dataset and a Conflicts/Attacks dataset. While the former is similar to other offense detection related datasets, the latter constitutes a novelty due to the use of the history of the interaction between users. Several studies were carried out to construct and analyze the data in the datasets. The first study included gathering expressions of verbal aggression witnessed by adolescents to guide data extraction for the datasets. The second study included extracting data from Twitter (in Portuguese) that matched the most frequent expressions/words/sentences that were identified in the previous study. The third study consisted in the development of the Aggressiveness dataset, the Conflicts/Attacks dataset, and classification models. In our fourth study, we proposed to examine whether online aggression and conflicts/attacks revealed any trend changes over time with a sample of 86 adolescents. With this study, we also proposed to investigate whether the amount of tweets sent over a period of 273 days was related to online aggression and conflicts/attacks. Finally, we analyzed the percentage of participants who participated in the aggressions and/or attacks/conflicts. Paula Ferreira 0004, Nádia Salgado Pereira, Hugo Rosa, Sofia Oliveira, Luísa Coheur, Sofia Mateus Francisco, Sidclay Bezerra de Souza, Ricardo Ribeiro 0001, João Paulo Carvalho 0001, Paula Paulino, Isabel Trancoso, Ana Margarida Veiga Simão |
IEEE Trans. Affect. Comput. | 9 |
| 2024 | Information Retrieval Using Fuzzy Fingerprints
Gonçalo Raposo, João Paulo Carvalho 0001, Luísa Coheur, Bruno Martins 0001 |
IPMU (1) | 2 |
| 2023 | Who Said That?: Selecting the Correct Persona from Conversational TextabstractIn this paper, we explore the ability to detect personas from conversational data. First, we adapt the Persona-Chat dataset, a well-known dialogue dataset, to support the task of selecting the correct persona out of various candidates. Then, we introduce persona perturbations to create additional identical personas that act as more challenging distractors. We train three different BERT-based models in a multiple-choice fashion to select the correct persona from a group of distractor personas. We show that this approach is able to discern between the group of original persona candidates, however, these models struggle to maintain high performance when we employ very identical distractors obtained from the proposed perturbations. João Paulo Carvalho 0001, Luísa Coheur |
IVA | 2 |
| 2023 | PGTask: Introducing the Task of Profile Generation from DialoguesabstractRecent approaches have attempted to personalize dialogue systems by leveraging profile information into models.However, this knowledge is scarce and difficult to obtain, which makes the extraction/generation of profile information from dialogues a fundamental asset.To surpass this limitation, we introduce the Profile Generation Task (PGTask).We contribute with a new dataset for this problem, comprising profile sentences aligned with related utterances, extracted from a corpus of dialogues.Furthermore, using state-of-the-art methods, we provide a benchmark for profile generation on this novel dataset.Our experiments disclose the challenges of profile generation, and we hope that this introduces a new research direction. João Paulo Carvalho 0001, Luísa Coheur |
SIGDIAL | 2 |
| 2023 | EmulART: Emulating radiative transfer - a pilot study on autoencoder-based dimensionality reduction for radiative transfer modelsabstractAbstract Dust is a major component of the interstellar medium. Through scattering, absorption and thermal re-emission, it can profoundly alter astrophysical observations. Models for dust composition and distribution are necessary to better understand and curb their impact on observations. A new approach for serial and computationally inexpensive production of such models is here presented. Traditionally these models are studied with the help of radiative transfer modelling, a critical tool to understand the impact of dust attenuation and reddening on the observed properties of galaxies and active galactic nuclei. Such simulations present, however, an approximately linear computational cost increase with the desired information resolution. Our new efficient model generator proposes a denoising variational autoencoder (or alternatively PCA), for spectral compression, combined with an approximate Bayesian method for spatial inference, to emulate high information radiative transfer models from low information models. For a simple spherical dust shell model with anisotropic illumination, our proposed approach successfully emulates the reference simulation starting from less than 1% of the information. Our emulations of the model at different viewing angles present median residuals below 15% across the spectral dimension and below 48% across spatial and spectral dimensions. EmulART infers estimates for $$\sim $$ ∼ 85% of information missing from the input, all within a total running time of around 20 minutes, estimated to be 6 $$\times $$ × faster than the present target high information resolution simulations, and up to 50 $$\times $$ × faster when applied to more complicated simulations. João Rino-Silvestre, Santiago Gonzalez-Gaitan, Marko Stalevski, Majda Smole, Pedro Guilherme-Garcia, João Paulo Carvalho 0001, Ana Maria Mourão |
Neural Comput. Appl. | 6 |
| 2022 | Fast Text Based Classification of News Snippets for Telecom Assurance
Artur Simões, João Paulo Carvalho 0001 |
IPMU (2) | 2 |
| 2021 | Fuzzy Influence in Fuzzy Semantic Similarity MeasuresabstractThe field of Computing with Words has been pivotal in the development of fuzzy semantic similarity measures. Fuzzy semantic similarity measures allow the modelling of words in a given context with a tolerance for the imprecise nature of human perceptions. In this work, we look at how this imprecision can be addressed with the use of fuzzy semantic similarity measures in the field of natural language processing. A fuzzy influence factor is introduced into an existing measure known as FUSE. FUSE computes the similarity between two short texts based on weighted syntactic and semantic components in order to address the issue of comparing fuzzy words that exist in different word categories. A series of empirical experiments investigates the effect of introducing a fuzzy influence factor into FUSE across a number of short text datasets. Comparisons with other similarity measures demonstrates that the fuzzy influence factor has a positive effect in improving the correlation of machine similarity judgments with similarity judgments of humans. Naeemeh Adel, Keeley A. Crockett, João Paulo Carvalho 0001, Valerie V. Cross |
FUZZ-IEEE | 3 |
| 2021 | Retrieval Augmentation for Deep Neural NetworksabstractDeep neural networks have achieved state-of-the-art results in various vision and/or language tasks. Despite the use of large training datasets, most models are trained by iterating over single input-output pairs, discarding the remaining examples for the current prediction. In this work, we actively exploit the training data, using the information from nearest training examples to aid the prediction both during training and testing. Specifically, our approach uses the target of the most similar training example to initialize the memory state of an LSTM model, or to guide attention mechanisms. We apply this approach to image captioning and sentiment analysis, respectively through image and text retrieval. Results confirm the effectiveness of the proposed approach for the two tasks, on the widely used Flickr8 and IMDB datasets. Our code is publicly available11htttp://github.com/RitaRamo/retrieval-augmentation-nn. Rita Ramos, Patrícia Pereira, Helena Moniz, João Paulo Carvalho 0001, Bruno Martins 0001 |
IJCNN | 4 |
| 2020 | Interpreting Human Responses in Dialogue Systems using Fuzzy Semantic Similarity MeasuresabstractDialogue systems are automated systems that interact with humans using natural language. Much work has been done on dialogue management and learning using a range of computational intelligence based approaches, however the complexity of human dialogue in different contexts still presents many challenges. The key impact of work presented in this paper is to use fuzzy semantic similarity measures embedded within a dialogue system to allow a machine to semantically comprehend human utterances in a given context and thus communicate more effectively with a human in a specific domain using natural language. To achieve this, perception based words should be understood by a machine in context of the dialogue. In this work, a simple question and answer dialogue system is implemented for a café customer satisfaction feedback survey. Both fuzzy and crisp semantic similarity measures are used within the dialogue engine to assess the accuracy and robustness of rule firing. Results from a 32 participant study, show that the fuzzy measure improves rule matching within the dialogue system by 21.88% compared with the crisp measure known as STASIS, thus providing a more natural and fluid dialogue exchange. Naeemeh Adel, Keeley A. Crockett, David Chandran, João Paulo Carvalho 0001 |
FUZZ-IEEE | 4 |
| 2020 | Relevance Ranking for Web SearchabstractRelevance ranking is a core problem of Information Retrieval which plays a fundamental role in various real world applications, such as search engines. Given a query and a set of candidate text documents, relevance ranking algorithms determine how relevant each text document is for the given query. This degree of relevance allows them to rank the text documents and perform actions such as returning the best matching documents for the query. As in other machine learning and computational intelligence disciplines, deep learning techniques have recently achieved state of the art results by successfully capturing relevance matching signals between query-textual document pairs. This paper focuses on the PositionAware Convolutional-Recurrent Relevance Matching approach. On a first phase, it reimplements the original work, reproduces the published results and performs a number of additional experiments that identify potential model limitations. On a second phase, it explores possible model improvements based on deep learning techniques such as soft self-attention and deep transfer learning. Experiments on the well-known TREC Web Track data show that it is possible to obtain small improvements over the original model and point to a number of limitations of the general approach due to the information bottlenecks involved. João Lages, João Paulo Carvalho 0001 |
FUZZ-IEEE | 2 |
| 2020 | Creating Classification Models from Textual Descriptions of Companies Using Crunchbase
Marco Felgueiras, Fernando Batista, João Paulo Carvalho 0001 |
IPMU (1) | 3 |
| 2020 | Electrical Power Grid Frequency Estimation with Fuzzy Boolean Nets
Nuno M. Rodrigues, João Paulo Carvalho 0001, Fernando M. Janeiro, Pedro M. Ramos |
IPMU (1) | 2 |
| 2020 | Comparing different solutions for forecasting the energy production of a wind farm
Darío Baptista, João Paulo Carvalho 0001, Fernando Morgado Dias |
Neural Comput. Appl. | 2 |
| 2020 | FUZYE: A Fuzzy c-Means Analog IC Yield Optimization Using Evolutionary-Based AlgorithmsabstractThis paper presents fuzzy c-means-based yield estimation (FUZYE), a methodology that reduces the time impact caused by Monte Carlo (MC) simulations in the context of analog integrated circuits (ICs) yield estimation, enabling it for yield optimization with population-based algorithms, e.g., the genetic algorithm (GA). MC analysis is the most general and reliable technique for yield estimation, yet the considerable amount of time it requires has discouraged its adoption in population-based optimization tools. The proposed methodology reduces the total number of MC simulations that are required, since, at each GA generation, the population is clustered using a fuzzy c-means (FCMs) technique, and, only the representative individual (RI) from each cluster is subject to MC simulations. This paper shows that the yield for the rest of the population can be estimated based on the membership degree of FCM and RIs yield values alone. This new method was applied on two real circuit-sizing optimization problems and the obtained results were compared to the exhaustive approach, where all individuals of the population are subject to MC analysis. The FCM approach presents a reduction of 89% in the total number of MC simulations, when compared to the exhaustive MC analysis over the full population. Moreover, a k-means-based clustering algorithm was also tested and compared with the proposed FUZYE, with the latest showing an improvement up to 13% in yield estimation accuracy. António Canelas, Ricardo Povoa, Ricardo Martins 0003, Nuno Lourenço 0003, Jorge Guilherme, João Paulo Carvalho 0001, Nuno Horta |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2019 | Human Hedge Perception - and its Application in Fuzzy Semantic Similarity MeasuresabstractFuzzy Semantic Similarity Measures are algorithms that are able to compare two or more short texts that contain human perception based words and return a numeric measure of similarity of meaning between them. Such similarity is computed using a weighting, comprised of the semantic and the syntactic composition of the short text. Similarities of individual words are computed through the use of a corpus, and ontological structures based on both WordNet - a well-known lexical database of English, and on category specific fuzzy ontologies created from the derivation of Type-I or Type-II interval fuzzy sets from human perceptions of fuzzy words. Currently, linguistic hedges are not utilized in the similarity calculation within fuzzy semantic similarity measures and are ignored. This paper describes a study, which aims to capture human perceptions for linguistic hedges typically used in natural language. Twelve linguistic hedges used within natural language are selected and an experiment is conducted to capture human perceptions of the impact of hedges on fuzzy category words. A dataset of hedge sentence pairs is created and rated in terms of similarity by human participants. Excellent inter-rater correlations and inter-class correlations are established between the average human ratings and an established fuzzy semantic similarity measure. Naeemeh Adel, Keeley A. Crockett, Alan Crispin, João Paulo Carvalho 0001, David Chandran |
FUZZ-IEEE | 4 |
| 2019 | 2Gather4Health: Automatic Web Identification of Solutions in Patient InnovationabstractPatient Innovation is an online open platform, with a community of over 60.000 users and more than 800 innovative solutions developed by patients and informal caregivers from all over the world. These solutions and/or creators were found by manually searching the Web through a combination of appropriate keywords and using experts to curate the results. In this paper we present a dedicated web-crawler architecture that includes a text classifier able to automatically identify Patient Innovation solutions from the web. The classifier is composed by a 2-layer hybrid MNB and Fuzzy Fingerprint classifier. João N. Almeida, João Paulo Carvalho 0001, Salomé Azevedo |
FUZZ-IEEE | 2 |
| 2019 | An Architecture Based on Fuzzy Systems for Personalized Medicine in ICUsabstractThis paper proposes a decision support system based on fuzzy clustering, fuzzy modeling and fuzzy fingerprints, to provide personalized therapy for critically ill patients. It is hypothesized that the ‘collective experience’ from large clinical databases, where clinical decisions are linked with patient outcomes, can be used to identify specific patient sub-groups and build personalized therapy models towards a new era of personalized medicine, allowing the improvement of patient outcomes in the Intensive Care Unit (ICU). The validity of the proposed systems will be tested using the case study of patients admitted to the ICU who then develop acute kidney injury (AKI); Two-thirds of patients with AKI require renal support therapy. Generalized severity scoring systems have consistently performed poorly for patients with AKI. João Miguel da Costa Sousa, Susana M. Vieira, João Paulo Carvalho 0001, Sara C. Madeira, Leo A. Celi, Stan N. Finkelstein |
FUZZ-IEEE | 3 |
| 2018 | FUSE (Fuzzy Similarity Measure) - A measure for determining fuzzy short text similarity using Interval Type-2 fuzzy setsabstractMeasurement of the semantic and syntactic similarity of human utterances is essential in developing language that is understandable when machines engage in dialogue with users. However, human language is complex and the semantic meaning of an utterance is usually dependent on context at a given time and also based on learnt experience of the meaning of the perception based words that are used. Limited work in terms of the representation and coverage has been done on the development of fuzzy semantic similarity measures. This paper proposes a new measure known as FUSE (FUzzy Similarity mEasure) which determines similarity using expanded categories of perception based words that have been modelled using Interval Type-2 fuzzy sets. The paper describes the method of obtaining the human ratings of these words based on Mendel's methodology and applies them within the FUSE algorithm. FUSE is then evaluated on three established datasets and is compared with two known semantic similarity algorithms. Results indicate FUSE provides higher correlations to human ratings. Naeemeh Adel, Keeley A. Crockett, Alan Crispin, David Chandran, João Paulo Carvalho 0001 |
FUZZ-IEEE | 5 |
| 2018 | Towards Computational Fact-Checking: Is the information checkable?abstractFact-checking has recently become a real world hot topic, especially in what concerns political claims. Several big players, such as, for example, Google or Facebook, have started addressing/making contributions to make "fact-checking" possible/available to the general public. However, most, if not all fact-checking platforms are largely manual, in the sense that most of the contributions and of the actual checking is performed by humans. Automatic computational fact-checking is still very far from being reliable and available on a large scale.The current work is a contribution to the goal of automatic fact-checking by presenting features to distinguish checkable from uncheckable sentences and a fuzzy approach to computing sentence checkability, i.e., to answer the question: "is it possible to know if a sentence is worth to be checked?". Even though, this is a hot topic, few to none solutions have been presented to automatically assess the worthiness and liability of the verification of a sentence. The solution that is proposed is mainly based on Natural Language Processing methods, linguistics and fuzzy logic.Assessing the checkability of a sentence can have many applications besides automatic fact-checking, like, for example: a broader fact-checking view (automatic checking of webpages and articles), the summarizing of information and the evaluation of factual information on texts. The main goal, however, was to focus on analysis of individual sentences, to provide an important tool to automatic fact-checking, by finding if a sentence is worth to, and can, be checked. Hugo Farinha, João Paulo Carvalho 0001 |
FUZZ-IEEE | 2 |
| 2018 | Using Fuzzy Fingerprints for Cyberbullying Detection in Social NetworksabstractAs cyberbullying becomes more and more frequent in social networks, automatically detecting it and pro-actively acting upon it becomes of the utmost importance. In this work, we study how a recent technique with proven success in similar tasks, Fuzzy Fingerprints, performs when detecting textual cyberbullying in social networks. Despite being commonly treated as binary classification task, we argue that this is in fact a retrieval problem where the only relevant performance is that of retrieving cyberbullying interactions. Experiments show that the Fuzzy Fingerprints slightly outperforms baseline classifiers when tested in a close to real life scenario, where cyberbullying instances are rarer than those without cyberbullying. Hugo Rosa, João Paulo Carvalho 0001, Pável Calado, Bruno Martins 0001, Ricardo Ribeiro 0001, Luísa Coheur |
FUZZ-IEEE | 2 |
| 2018 | A "Deeper" Look at Detecting Cyberbullying in Social NetworksabstractAs cyberbullying becomes more and more frequent in social networks, automatically detecting it and pro-actively acting upon it becomes of the utmost importance. In this work, a detailed look at the current state-of-the-art in cyberbullying detection reveals that deep learning techniques have seldom been used to tackle this problem, despite growing reputation in other text-based classification tasks. Motivated by neural networks' documented success, three architectures are implemented from similar works: a simple CNN, a hybrid CNN-LSTM and a mixed CNN-LSTM-DNN. In addition, three text representations are trained from three different sources, via the word2vec model: Google-News, Twitter and Formspring. The experiment shows that these models with one of the above embeddings beat other benchmark classifiers (Support Vector Machines and Logistic Regression) both in an unbalanced and balanced version of the same dataset. Hugo Rosa, David Martins de Matos, Ricardo Ribeiro 0001, Luísa Coheur, João Paulo Carvalho 0001 |
IJCNN | 5 |
| 2018 | Tag-Based User Fuzzy Fingerprints for Recommender Systems
André Luiz da Costa Carvalho, Pável Calado, João Paulo Carvalho 0001 |
IPMU (3) | 3 |
| 2017 | Combining ratings and item descriptions in recommendation systems using fuzzy fingerprintsabstractMemory-based Collaborative filtering solutions are dominant in the Recommendation Systems domain, due to their low implementation effort and service maintenance, when compared to Model-based approaches. Memory-based systems often rely on similarity metrics to compute similarities between items (or users) using ratings, in what is often named neighbor-based Collaborative filtering. This paper applies Fuzzy Fingerprints to create a novel similarity metric. In it, the Fuzzy Fingerprint of each item is described with a ranking of users ratings, combined with words obtained from the items' description. This allows the presented similarity metric to use fewer neighbors than other well-known metrics such as Cosine similarity or Pearson Correlation. Our proposal is able to reduce RMSE by at least 0.030 and improve NDCG@10 by at least 0.017 when compared with the best baseline here presented. André Luiz da Costa Carvalho, Pável Calado, João Paulo Carvalho 0001 |
FUZZ-IEEE | 3 |
| 2017 | Detecting relevant tweets in very large tweet collections: The London Riots case studyabstractIn this paper we propose to approach the subject of detecting relevant tweets when in the presence of very large tweet collections containing a large number of different trending topics. We use a large database of tweets collected during the 2011 London Riots as a case study to demonstrate the application of the proposed techniques. In order to extract relevant content, we extend, formalize and apply a recent technique, called Twitter Topic Fuzzy Fingerprints, which, in the scope of social media, outperforms other well known text based classification methods, while being less computationally demanding, an essential feature when processing large volumes of streaming data. Using this technique we were able to detect 45% additional relevant tweets within the database. João Paulo Carvalho 0001, Hugo Rosa, Fernando Batista |
FUZZ-IEEE | 1 |
| 2017 | Application of fuzzy semantic similarity measures to event detection within tweetsabstractThis paper examines the suitability of applying fuzzy semantic similarity measures (FSSM) to the task of detecting potential future events through the use of a group of prototypical event tweets. FSSM are ideal measures to be used to analyse the semantic textual content of tweets due to the ability to deal equally with not only nouns, verbs, adjectives and adverbs, but also perception based fuzzy words. The proposed methodology first creates a set of prototypical event related tweets and a control group of tweets from a data source, then calculates the semantic similarity against an event dataset compiled from tweets issued during the 2011 London riots. The dataset of tweets contained a proportion of tweets that the Guardian Newspaper publically released that were attributed to 200 influential Twitter users during the actual riot. The effects of changing the semantic similarity threshold are investigated in order to evaluate if Twitter tweets can be used in conjunction with fuzzy short text similarity measures and prototypical event related tweets to determine if an event is more likely to occur. By looking at the increase in frequency of tweets in the dataset, over a certain similarity threshold when matched with prototypical event tweets about riots, the results have shown that a potential future event can be detected. Keeley A. Crockett, Naeemeh Adel, James O'Shea, Alan Crispin, David Chandran, João Paulo Carvalho 0001 |
FUZZ-IEEE | 6 |
| 2017 | MISNIS: An intelligent platform for twitter topic mining
João Paulo Carvalho 0001, Hugo Rosa, Gaspar Brogueira, Fernando Batista |
Expert Syst. Appl. | 1 |
| 2017 | Daily prediction of ICU readmissions using feature engineering and ensemble fuzzy modeling
Rita Viegas, Cátia M. Salgado, Sérgio Curto, João Paulo Carvalho 0001, Susana M. Vieira, Stan N. Finkelstein |
Expert Syst. Appl. | 4 |
| 2016 | Predicting ICU readmissions based on bedside medical text notesabstractPatients are often discharged prematurely from Intensive Care Units (ICU) due to clinical resource limitations, economic pressure or poor discharge planning. The readmission of such patients is associated with an increased risk of death and is currently viewed as a marker for poor quality care. Several studies have focused on predicting which patients are likely to be readmitted, using techniques such as logistic regression or machine learning algorithms, and based on physiological data measured during the patients' stay at the ICU. So far, no published algorithms have been able to predict readmissions to a satisfactory degree. In this work we hypothesize that physicians' and nurses' notes could give a better explanation of both ICU discharges and readmissions, and propose using the text notes in an ICU database in order to build classification models for the prediction of readmissions. We tested the use of Fuzzy Fingerprints and other traditional text classifiers and compared them to a previously proposed model based on numerical data, obtaining very relevant improvements in the classification results, namely an AUC=0.8. Sérgio Curto, João Paulo Carvalho 0001, Cátia M. Salgado, Susana M. Vieira, João Miguel da Costa Sousa |
FUZZ-IEEE | 2 |
| 2016 | Creating Extended Gender Labelled Datasets of Twitter Users
Marco Vicente, Fernando Batista, João Paulo Carvalho 0001 |
IPMU (2) | 3 |
| 2015 | Text based classification of companies in CrunchBaseabstractThis paper introduces two fuzzy fingerprint based text classification techniques that were successfully applied to automatically label companies from CrunchBase, based purely on their unstructured textual description. This is a real and very challenging problem due to the large set of possible labels (more than 40) and also to the fact that the textual descriptions do not have to abide by any criteria and are, therefore, extremely heterogeneous. Fuzzy fingerprints are a recently introduced technique that can be used for performing fast classification. They perform well in the presence of unbalanced datasets and can cope with a very large number of classes. In the paper, a comparison is performed against some of the best text classification techniques commonly used to address similar problems. When applied to the CrunchBase dataset, the fuzzy fingerprint based approach outperformed the other techniques. Fernando Batista, João Paulo Carvalho 0001 |
FUZZ-IEEE | 2 |
| 2015 | Twitter gender classification using user unstructured informationabstractThis paper describes an approach to automatically detect the gender of Twitter users, based only on clues provided by their profile information in an unstructured form. A number of features that capture phenomena specific of Twitter users is proposed and evaluated on a dataset of about 242K English language users. Different supervised and unsupervised approaches are used to assess the performance of the proposed features, including Naive Bayes variants, Logistic Regression, Support Vector Machines, Fuzzy c-Means clustering, and K-means. An unsupervised approach based on Fuzzy c-Means proved to be very suitable for this task, returning the correct gender for about 96% of the users. Marco Vicente, Fernando Batista, João Paulo Carvalho 0001 |
FUZZ-IEEE | 3 |
| 2015 | Incremental dataflow execution, resource efficiency and probabilistic guarantees with Fuzzy Boolean nets
Sérgio Esteves, João Nuno de Oliveira e Silva, João Paulo Carvalho 0001, Luís Veiga |
J. Parallel Distributed Comput. | 3 |
| 2014 | Twitter Topic Fuzzy FingerprintsabstractIn this paper we propose to approach the subject of Twitter Topic Detection using a new technique called Topic Fuzzy Fingerprints. A comparison is made with two popular text classification techniques, Support Vector Machines (SVM) and fc-Nearest Neighbours (fcNN). Preliminary results show that Twitter Topic Fuzzy Fingerprints outperforms the other two techniques achieving better Precision and Recall, while still being much faster, which is an essential feature when processing large volumes of streaming data. Hugo Rosa, Fernando Batista, João Paulo Carvalho 0001 |
FUZZ-IEEE | 3 |
| 2014 | Fuzzy Boolean Nets - a nature inspired model for learning and reasoning
José Alberto Batista Tomé, João Paulo Carvalho 0001 |
Fuzzy Sets Syst. | 2 |
| 2013 | Introducing UWS - A fuzzy based word similarity function with good discrimination capability: Preliminary resultsabstractThis paper introduces a novel word similarity function, the Uke Similarity Function (UWS), that fuses the most interesting characteristics of the two main philosophies in word and string matching: the edit distance and the n-gram similarity approach. It also uses fuzzy sets to integrate expert knowledge about typographical errors and to easily include phonetic and token related errors. The UWS was developed with the goal of automatic detection and correction of typographical and other word errors in unedited corpus data when creating word lists. João Paulo Carvalho 0001, Luísa Coheur |
FUZZ-IEEE | 1 |
| 2013 | On the semantics and the use of fuzzy cognitive maps and dynamic cognitive maps in social sciences
João Paulo Carvalho 0001 |
Fuzzy Sets Syst. | 1 |
| 2012 | A critical survey on the use of Fuzzy Sets in Speech and Natural Language ProcessingabstractThis paper shows how the use and applications of Fuzzy Sets (FS) in Speech and Natural Language Processing (SNLP) have seen a steady decline to a point where FS are virtually unknown or unappealing for most of the researchers currently working in the SNLP field, tries to find the reasons behind this decline, and proposes some guidelines on what could be done to reverse it and make FS assume a relevant role in SNLP. João Paulo Carvalho 0001, Fernando Batista, Luísa Coheur |
FUZZ-IEEE | 1 |
| 2011 | Web user identification with fuzzy fingerprintsabstractFingerprint identification is a well-known technique in forensic sciences. The basic idea of identifying a subject based on a set of features left by the subject actions or behavior can be applied to other domains. Identifying a web user based on a user fingerprint is one such application. This paper considers the problem of extracting fingerprints from web usage logs and matching them with those obtained from a set of known users. It presents an innovative fuzzy fingerprint algorithm based on vector valued fuzzy sets. Used sites are used as base features to create the fingerprint. The assumption is that sites accessed by each user remain approximately stable and are distinctive. The paper presents the proposed algorithm and shows some experimental results that validate the approach. The use of fast and compact algorithms is critical due to the possible huge number of users, and allows this method to be used on near real time. Nuno Homem, João Paulo Carvalho 0001 |
FUZZ-IEEE | 2 |
| 2010 | On the semantics and the use of Fuzzy Cognitive Maps in social sciencesabstractFuzzy Cognitive Maps (FCM) have been around for more than twenty years, but the way how they have been used and the interpretation of their results are nowadays far from their original intended goal. This paper focus on discussing the structure, the semantics and the possible use of FCM as tools to model and simulate complex social, economic and political systems, while clarifying some issues that have been recurrent in published FCM papers. João Paulo Carvalho 0001 |
FUZZ-IEEE | 1 |
| 2010 | Estimating Top-k Destinations in Data Streams
Nuno Homem, João Paulo Carvalho 0001 |
IPMU | 2 |
| 2010 | Dispersion Estimates for Telecommunications Fraud
Nuno Homem, João Paulo Carvalho 0001 |
IPMU | 2 |
| 2010 | Finding top-k elements in data streams
Nuno Homem, João Paulo Carvalho 0001 |
Inf. Sci. | 2 |
| 2009 | Distributed Routing Path Optimization for OBS Networks Based on Ant Colony OptimizationabstractThis work proposes a distributed framework for routing path optimization in Optical Burst-Switched (OBS) networks loosely mimicking the foraging behavior of ants, which in the past has originated the Ant Colony Optimization (ACO) metaheuristic. The distributed framework consists of additional data structures stored at the nodes and special control packets used to estimate the goodness of the routing paths and update the routing tables of the nodes. The performance of the ACO-based framework is evaluated, through network simulation, using two reference network topologies and compared with that obtained with shortest path routing and centralized routing path optimization. The simulation results show that the distributed framework significantly improves the performance of OBS networks, when compared to that of using shortest path routing, and attains a comparable performance to that of the centralized strategy. Moreover, the results also suggest that the framework is robust, as it does not require fine tuning its main parameters. João Pedro 0001, João Pires 0001, João Paulo Carvalho 0001 |
GLOBECOM | 3 |
| 2008 | Issues on Dynamic Cognitive Map modelling of purse-seine fishing skippers behaviorabstractThis paper focus on obtaining a qualitative dynamic model based on real world data taken from a real world qualitative system: the day to day behavior of purse seine fishing fleet skippers. The model is based on a dynamic cognitive mapping approach (rule based fuzzy cognitive maps - RB-FCM) where several developments had to be made in order to obtain a workable system. Most changes were due to timing issues, which are essential in the study of system dynamics but have traditionally been avoided in most dynamic cognitive maps modelling approaches. João Paulo Carvalho 0001, Laura Wise, Alberto Murta, Marta Mesquita |
FUZZ-IEEE | 1 |
| 2007 | Fuzzy Boolean Networks Learning BehaviourabstractIn this paper one studies the learning behaviour of an entire rule base in fuzzy Boolean networks. It is analyzed the influence of a set of factors such as number of inputs per neuron, granularity of antecedent spaces and number of teaching experiments on learning effectiveness without cross influence between rules and on interpolation capabilities of the network. Both one dimensional problems and two dimensional problems are tested and results interpreted using theoretical results also presented. José Alberto Batista Tomé, João Paulo Carvalho 0001 |
ISDA | 2 |
| 2007 | Qualitative optimization of Fuzzy Causal Rule Bases using Fuzzy Boolean Nets
João Paulo Carvalho 0001, José Alberto Batista Tomé |
Fuzzy Sets Syst. | 1 |
| 2006 | Using Rule-based Fuzzy Cognitive Maps to Model Dynamic Cell Behavior in Voronoi Based Cellular AutomataabstractThis paper focus on the use of Rule Based Fuzzy Cognitive Maps to represent cell behaviour in Voronoi Based Cellular Automata in order to model the dynamics of temporal and spatial propagation processes. As an application example, the proposed approach is applied to modelling and simulation of forest fire propagation. João Paulo Carvalho 0001, Marco Carola, José Alberto Batista Tomé |
FUZZ-IEEE | 1 |
| 2005 | Market Index Prediction using Fuzzy Boolean NetsabstractA wide range of applications can be identified for time series prediction, including energy systems planning, currency forecasting, or traffic prediction. Specifically, stock exchange operations can greatly benefit from efficient forecast techniques. Therefore, a number of different prediction approaches have been proposed such as linear models, feedforward neural network models, recurrent neural networks or fuzzy neural models. In this paper one presents a prediction model based on fuzzy rules that relate past data values with the next unknown value to be estimated. A fuzzy Boolean neural network has been used for this purpose, which has been applied to the Nasdaq index prediction. The results turned to be encouraging, namely on the percentage of correct up/down trend prediction. José Alberto Batista Tomé, João Paulo Carvalho 0001 |
HIS | 2 |
| 2004 | Qualitative modelling of an economic system using rule-based fuzzy cognitive mapsabstractTruly qualitative modelling of qualitative dynamic systems is a delicate issue in a sense that even when one uses a "qualitative modelling tool" one often end up adopting quantitative model approaches in disguise. In this paper, we use rule based fuzzy cognitive maps to obtain and simulate a qualitative model of an economic system. João Paulo Carvalho 0001, José Alberto Batista Tomé |
FUZZ-IEEE | 1 |
| 2001 | Rule Based Fuzzy Cognittive Maps-Expressing Time in Qualitative System DynamicsabstractTime is essential in the study of system dynamics. When representing and analyzing the dynamics of complex quantitative systems, the problem of expressing the "effect" of time flow is naturally solved since the mathematical equations that describe the relations between the entities of the system are a function of time. However, if we are dealing with real world qualitative systems that are impossible or difficult to model using mathematical equations, then the use of natural language becomes the best tool to represent the system and expressing time influence becomes a real issue that has not been addressed before. This paper introduces a coherent procedure to implicitly represent time in rule based fuzzy cognitive maps which are a previously introduced methodology and tool to represent and simulate the dynamics of qualitative systems. João Paulo Carvalho 0001, José Alberto Batista Tomé |
FUZZ-IEEE | 1 |