Carlos Gonçalves

dblp:22/7871 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
6since 2021 · last 2024
0000-0001-9113-6269ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2024 Sovereign Citizen on Digital Regulated Services Ecosystem
A. Luís Osório, Luis M. Camarinha-Matos, Carlos Gonçalves, Tiago M. Dias
PRO-VE (1)3
2023 A Business Technology Alignment Strategy for Digital Collaborative Networks
A. Luís Osório, Ricardo J. Rabelo, Carlos Gonçalves
PRO-VE3
2022 A Multi-supplier Collaborative Monitoring Framework for Informatics System of Systems
Carlos Gonçalves, Tiago Dias 0001, A. Luís Osório, Luis M. Camarinha-Matos
PRO-VE1
2022 Collaborative Management of Traffic Accidents Data for Social Impact Analytics
A. Luís Osório, Cláudia Antunes, Luis M. Camarinha-Matos, Carlos Gonçalves
PRO-VE4
2021 A Collaborative Cyber-Physical Microservices Platform - the SITL-IoT Case
Carlos Gonçalves, A. Luís Osório, Luis M. Camarinha-Matos, Tiago Dias 0001, José Tavares
PRO-VE1
2021 Open and Collaborative Micro Services in Digital Transformation
A. Luís Osório, Luis M. Camarinha-Matos, Tiago Dias 0001, Carlos Gonçalves, José Tavares
PRO-VE4
2016 A theoretical model for n-gram distribution in big data corpora
abstract
There is a wide diversity of applications relying on the identification of the sequences of n consecutive words (n-grams) occurring in corpora. Many studies follow an empirical approach for determining the statistical distribution of the n-grams but are usually constrained by the corpora sizes, which for practical reasons stay far away from Big Data. However, Big Data sizes imply hidden behaviors to the applications, such as extraction of relevant information from Web scale sources. In this paper we propose a theoretical approach for estimating the number of distinct n-grams in each corpus. It is based on the Zipf-Mandelbrot Law and the Poisson distribution, and it allows an efficient estimation of the number of distinct 1-grams, 2-grams,..., 6-grams, for any corpus size. The proposed model was validated for English and French corpora. We illustrate a practical application of this approach to the extraction of relevant expressions from natural language corpora, and predict its asymptotic behaviour for increasingly large sizes.
Joaquim Ferreira da Silva, Carlos Gonçalves, José C. Cunha
IEEE BigData2
2016 An n-gram cache for large-scale parallel extraction of multiword relevant expressions with LocalMaxs
abstract
LocalMaxs extracts relevant multiword terms based on their cohesion but is computationally intensive, a critical issue for very large natural language corpora. The corpus properties concerning n-gram distribution determine the algorithm complexity and were empirically analyzed for corpora up to 982 million words. A parallel LocalMaxs implementation exhibits almost linear relative efficiency, speedup, and sizeup, when executed with up to 48 cloud virtual machines and a distributed key-value store. To reduce the remote data communication, we present a novel n-gram cache with cooperative-based warm-up, leading to reduced miss ratio and time penalty. A cache analytical model is used to estimate the performance of cohesion calculation of n-gram expressions, based on corpus empirical data. The model estimates agree with the real execution results.
Carlos Gonçalves, Joaquim Ferreira da Silva, José C. Cunha
eScience1
2012 Data analytics in the cloud with flexible MapReduce workflows
abstract
Data analytic applications are characterized by large data sets that are subject to a series of processing phases. Some of these phases are executed sequentially but others can be executed concurrently or in parallel on clusters, grids or clouds. The MapReduce programming model has been applied to process large data sets in cluster and cloud environments. For developing an application using MapReduce there is a need to install/configure/access specific frameworks such as Apache Hadoop or Elastic MapReduce in Amazon Cloud. It would be desirable to provide more flexibility in adjusting such configurations according to the application characteristics. Furthermore the composition of the multiple phases of a data analytic application requires the specification of all the phases and their orchestration. The original MapReduce model and environment lacks flexible support for such configuration and composition. Recognizing that scientific workflows have been successfully applied to modeling complex applications, this paper describes our experiments on implementing MapReduce as sub-workflows in the AWARD framework (Autonomic Workflow Activities Reconfigurable and Dynamic). A text mining data analytic application is modeled as a complex workflow with multiple phases, where individual workflow nodes support MapReduce computations. As in typical MapReduce environments, the end user only needs to define the application algorithms for input data processing and for the map and reduce functions. In the paper we present experimental results when using the AWARD framework to execute MapReduce workflows deployed over multiple Amazon EC2 (Elastic Compute Cloud) instances.
Carlos Gonçalves, Luís Assunção, José C. Cunha
CloudCom1
2005 Open Multi-Technology Service Oriented Architectur for "Its" Business Models: The ITSIBus Etoll Services
A. Luís Osório, Carlos Gonçalves, Paul Araújo, Manuel Barata, J. Sales Gomes, Gastão C. Jacquet, Bui M. Dias
PRO-VE2