Jorge Bernardino

dblp:43/3799 · DBLP profile ↗
← Back
81ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0001-9660-2011ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 37 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 23 · 1 since 2021Artificial intelligence and machine learning · 17 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 17 · 2 since 2021Security and privacy · 3Systems, architecture and hardware · 1Theory of computation · 1
YearPublicationVenuePosition
2026 Automating the generation of database artifacts: From ER+ to SQL
abstract
Data engineers often need to transform a conceptual understanding of an application into deployable database artifacts spanning operational and analytical layers, heterogeneous locations, and explicit data-transformation pipelines. In this setting, conceptual specifications are valuable not only for the initial authoring of schemas and queries but also for coherently propagating subsequent conceptual changes into implementation artifacts. ER+ offers constructs for grouping, aggregation, line functions, and data transport, but translating these constructs into consistent executable artifacts remains a demanding task. This paper presents an extension of Online Database Architect (ONDA) that supports ER+ and automatically generates relational schemas and executable Structured Query Language (SQL) for both operational and analytical layers. The approach is defined by explicit mapping rules grounded in relational-algebraic semantics and constrained by invariants that ensure deterministic naming, key preservation, referential integrity, and sound compilation of grouping, aggregation, and line-function semantics. In particular, summaries involving line functions and aggregates are compiled through Common Table Expression (CTE)-based SQL patterns that preserve the intended grouping grain and avoid mixed-granularity expressions. We evaluate the approach using two case studies: a handcrafted ER+ model illustrating the end-to-end workflow and a TPC-H Q11-based scenario that represents a realistic analytical pattern. The generated artifacts show that ER+ specifications can be translated systematically into executable relational and analytical structures. A lightweight comparative evaluation further suggests reduced manual effort in the initial production of artifacts, improved consistency of generated analytical SQL, and maintainability when conceptual changes must be propagated into dependent implementations. These results indicate that conceptual models can serve as a practical basis for producing and evolving deployment-ready database artifacts.
Gonçalo Carvalho, Deolinda Rasteiro, Nour Dorgham, Bruno Cabral 0001, Jorge Bernardino, Vasco Pereira
Inf. Syst.5
2022 NoSQL Document Databases Assessment: Couchbase, CouchDB, and MongoDB
Inês Carvalho, Filipe Sá, Jorge Bernardino
DATA3
2022 Experimental Evaluation of Low Code development, Java Swing and JavaScript programming
abstract
Low Code is a technology that has been gaining popularity over the years, due to its potential and simplicity. But so far there has not been an experimental evaluation with other programming methods. This paper aims to introduce Low Code development technology and compare it with Java Swing programming and manual development with HTML (HyperText Markup Language), CSS (Cascading Style Sheet), and JavaScript. These technologies are compared using the following metrics: development time, execution time, and the number of written code lines. In this evaluation, two applications are implemented, a simple calculator, and a text editor, developed in all technologies. It is concluded that it is faster to develop applications in Low Code but in terms of execution time, these are usually slower. Although the Low Code development is still at a somewhat embryonic stage which leads to some bugs and errors, Low Code development is better in general than Java Swing programming, and somewhat similar to manual programming with HTML, CSS, and JavaScript. Another benefit is that Low Code generates HTML and CSS automatically.
André Calçada, Jorge Bernardino
IDEAS2
2022 Injecting software faults in Python applications
Henrique Marques, Nuno Laranjeiro, Jorge Bernardino
Empir. Softw. Eng.3
2021 A holistic data modeling approach for multi-database systems
abstract
IoT, edge-oriented systems, and the growing ubiquity of access to the Internet have driven the development of the most complex software systems to date. Designing such systems is demanding due to their distributed nature, different technologies, multi-layer, hard-to-meet quality attributes, and the integration of several databases with diverse technologies. This work proposes a data modeling method able to represent holistically these systems’ data structure, data transport, and transformation.
Gonçalo Carvalho, Jorge Bernardino, Vasco Pereira, Bruno Cabral 0001
IEEE BigData2
2021 An Analysis of Public REST Web Service APIs
abstract
Businesses are increasingly deploying their services on the web, in the form of web applications, SOAP services, message-based services, and, more recently, REST services. Although the movement towards REST is widely recognized, there is not much concrete information regarding the technical features being used in the field, such as typical data formats, how HTTP verbs are being used, or typical URI structures, just to name a few. In this paper, we go through the Alexa.com top 4000 most popular sites to identify precisely 500 websites claiming to provide a REST web service API. We analyze these 500 APIs for key technical features, degree of compliance with REST architectural principles (e.g., resource addressability), and for adherence to best practices (e.g., API versioning). We observed several trends (e.g., widespread JSON support, software-generated documentation), but, at the same time, high diversity in services, including differences in adherence to best practices, with only 0.8 percent of services strictly complying with all REST principles. Our results can help practitioners evolve guidelines and standards for designing higher quality services and also understand deficiencies in currently deployed services. Researchers may also benefit from the identification of key research areas, contributing to the deployment of more reliable services.
Andy Neumann, Nuno Laranjeiro, Jorge Bernardino
IEEE Trans. Serv. Comput.3
2020 Analysis of Data Anonymization Techniques
Joana Ferreira Marques, Jorge Bernardino
KEOD2
2020 Big Data Streaming Platforms to Support Real-time Analytics
Eliana Fernandes, Ana Carolina Salgado, Jorge Bernardino
ICSOFT3
2020 Computation offloading in Edge Computing environments using Artificial Intelligence techniques
Gonçalo Carvalho, Bruno Cabral 0001, Vasco Pereira, Jorge Bernardino
Eng. Appl. Artif. Intell.4
2020 Automating orthogonal defect classification using machine learning algorithms
Fábio Lopes 0002, João Agnelo, César Alexandre Teixeira, Nuno Laranjeiro, Jorge Bernardino
Future Gener. Comput. Syst.5
2020 Intrusion Detection Systems for Mitigating SQL Injection Attacks: Review and State-of-Practice
abstract
Databases are widely used by organizations to store business-critical information, which makes them one of the most attractive targets for security attacks. SQL Injection is the most common attack to webpages with dynamic content. To mitigate it, organizations use Intrusion Detection Systems (IDS) as part of the security infrastructure, to detect this type of attack. However, the authors observe a gap between the comprehensive state-of-the-art in detecting SQL Injection attacks and the state-of-practice regarding existing tools capable of detecting such attacks. The majority of IDS implementations provide little or no protection against SQL Injection attacks, with exceptions like the tools Bro and ModSecurity. In this article, the authors compare these tools using the CSIC dataset in order to examine the state-of-practice in database protection from SQL Injection attacks, identifying the main characteristics and implementation details needed for IDSs to successfully detect such attacks. The experiments indicate that signature-based IDS provide the greatest coverage against SQL Injection.
Raul Barbosa, Jorge Bernardino
Int. J. Inf. Secur. Priv.3
2020 Using Orthogonal Defect Classification to characterize NoSQL database defects
abstract
NoSQL databases are increasingly used for storing and managing data in business-critical Big Data systems. The presence of software defects (i.e., bugs) in these databases can bring in severe consequences to the NoSQL services being offered, such as data loss or service unavailability. Thus, it is essential to understand the types of defects that frequently affect these databases, allowing developers take action in an informed manner (e.g., redirect testing efforts). In this paper, we use Orthogonal Defect Classification (ODC) to classify a total of 4096 software defects from three of the most popular NoSQL databases: MongoDB, Cassandra, and HBase. The results show great similarity for the defects across the three different NoSQL systems and, at the same time, show the differences and heterogeneity regarding research carried out in other domains and types of applications, emphasizing the need for possessing such information. Our results expose the defect distributions in NoSQL databases, provide a foundation for selecting representative defects for NoSQL systems, and, overall, can be useful for developers for verifying and building more reliable NoSQL database systems.
João Agnelo, Nuno Laranjeiro, Jorge Bernardino
J. Syst. Softw.3
2019 A Case for Machine Learning in Edge-Oriented Computing to Enhance Mobility as a Service
abstract
The study of human mobility tries to understand human flows and synergies with the geographical environment. Mobility as a Service (MaaS) is a new mobility concept that promises to revolutionize commuting by merging public and private transport providers around a common platform that travelers will use as a service, thus providing new research opportunities for human mobility. One of MaaS main offerings is the ability to calculate both routes and commuting strategies based on the availability of transports and specific user constraints. In this work, we discuss how Edge-Oriented Computing (EOC) and Machine Learning (ML) can contribute to extending the reach of MaaS in the upcoming years. EOC enables technologies to perform computation at the edge of the network, reducing latency and communication overheads, which 5G technologies are committed to further diminish, thus benefitting the proliferation of MaaS. Also, ML techniques are one of the most robust approaches for planning routes and predicting future movements. Finally, we present open research topics that will promote the attractiveness of MaaS.
Gonçalo Carvalho, Bruno Cabral 0001, Vasco Pereira, Jorge Bernardino
DCOSS4
2019 Open Source Project Management Tools Assessment using QSOS Methodology
Anabela Carreira, Jorge Bernardino
KEOD2
2019 Project Management Tools Assessment with OSSpal
Samuel Cruz, Jorge Bernardino
KEOD2
2019 Evaluation of Asana, Odoo, and ProjectLibre Project Management Tools using the OSSpal Methodology
abstract
In order to successfully complete projects it is essential for companies to acquire a project management tool that assists in their planning, cost and resource management. Currently, there are several open source project management tools on the market that have much of the functionality required. Thus, it is important to choose the most appropriate one according to the needs of the user. To help with this choice, open source software evaluation methodologies can be used. In this paper, we use the OSSpal methodology to evaluate three popular open source project management tools: Asana, Odoo, and ProjectLibre.
Joana Ferreira Marques, Jorge Bernardino
KEOD2
2019 Evaluating Gant Project, Orange Scrum, and ProjeQtOr Open Source Project Management Tools using QSOS
abstract
The task of managing a software project is an extremely complex job, drawing on many personal, team, and organizational resources. We realized that exist many project management tools and software being developed every day to help managers to automate the administration of individual projects or groups of projects during their life-cycle. Therefore, it is important to identify the software functionalities and compare them to the intended requirements of the company, to select which software complements the expectations. This paper presents a comparison between three of the most popular open source project management tools: Gantt Project, OrangeScrum, ProjeQtOr. To assess these project management tools is used the Qualification and Selection Open Source (QSOS) methodology.
Catarina Romão Proença, Jorge Bernardino
ICSOFT2
2019 Java Web Services: A Performance Analysis
abstract
Service-oriented architecture (SOA) is being increasingly used by developers both in web applications and in mobile applications. Within web services there are two main implementations: SOAP communication protocol and REST. This work presents a comparative study of performance between these two types of web services, SOAP versus REST, as well as analyses factors that may affect the efficiency of applications that are based on this architecture. In this experimental evaluation we used an application deployed in a Wildfly server and then used the JMeter test tool to launch requests in different numbers of threads and calls. Contrary to the more general idea that REST web services are significantly faster than SOAP, our results show that REST web services are 1% faster than SOAP. As this programming paradigm is increasingly used in a growing number of client and server applications, we conclude that the REST implementation is more efficient for systems which have to respond to less calls but have more requests in a connection.
Pedro Costa e Silva, Jorge Bernardino
ICSOFT2
2019 Evaluating Open Source Project Management Tools using OSSPal Methodology
abstract
One of the major differences between a successful project and a failed one is the project management abilities. Project management leads to better alignment of projects within the business’ strategy, so that companies can reduce their costs, accelerate product development, and focus on meeting their customers’ needs. To help with that, project management tools are highly recommended, once they ease planning, scheduling, resource allocation, communication and documentation tasks. In this paper, we assess three popular open source project management tools: OpenProject, Orangescrum and ProjectLibre with the help of the OSSPal methodology. This study can help project managers and programmers on choosing an adequate, current, high quality and affordable tool to perform their projects.
Antonio Oliveira, Jorge Bernardino
WEBIST2
2019 An Application of OSSpal for the Assessment of Open Source Project Management Tools
abstract
Projects are a necessity within any competitive business, and as the execution of complex projects becomes the norm, so grows the need for advances in project management. The use of project management tools is key towards taming said complexity. There are many such tools available; the current challenge resides in picking the right one. In this paper, we evaluate three different tools - OpenProject, dotProject, and Odoo - using the OSSpal methodology.
Hugo Carvalho de Paula, Jorge Bernardino
WEBIST2
2019 Evaluating GitLab, OpenProject, and Redmine using QSOS Methodology
abstract
Measure and planning all aspects and variables of a project is extremely important to have success. Therefore, having the right tool is extremely important. To evaluate project management tools, we can use several methodologies, that allow to choose the best tools according to our criteria. QSOS is one of the methodologies that allows to make a more weighted choose. In this paper, we evaluate popular open source project management tools GitLab, OpenProject, and Redmine, using QSOS methodology.
André Vicente, Jorge Bernardino
WEBIST2
2018 Graph Databases Comparison: AllegroGraph, ArangoDB, InfiniteGraph, Neo4J, and OrientDB
Diogo Fernandes, Jorge Bernardino
DATA2
2018 Benchmarking Auto-WEKA on a Commodity Machine
João Freitas, Nuno Lavado, Jorge Bernardino
DATA3
2018 System to Predict Diseases in Vineyards and Olive Groves using Data Mining and Geolocation
Rodrigo Rocha Silva, Jorge Bernardino
ICSOFT3
2018 Open Source Data Mining Tools Evaluation using OSSpal Methodology
Any Keila Pereira, Ana Paula Sousa, João Ramalho Santos, Jorge Bernardino
ICSOFT4
2018 Virtualization: Past and Present Challenges
Frederico Cerveira, Raul Barbosa, Jorge Bernardino
ICSOFT4
2018 Freemium Project Management Tools: Asana, Freedcamp and Ace Project
Tânia Ferreira, Juncal Gutiérrez-Artacho, Jorge Bernardino
WorldCIST (1)3
2018 PreX: A predictive model to prevent exceptions
João Ricardo Lourenço, Bruno Cabral 0001, Jorge Bernardino
J. Syst. Softw.3
2017 On the Use of CEP in Safety-critical Systems
abstract
Nos dias de hoje, a informação é um dos principais recursos de qualquer empresa e desempenha um papel importante na tomada de decisão. Para as equipas de gestão, a obtenção de informações importantes o mais rápido possível é uma prioridade, que pode se tornar desafiadora quando existem muitos dados a serem processados. Os mecanismos complexos de processamento de eventos (CEP) são capazes de processar o contínuo fluxo de dados, separando a informação e filtrando os dados de pouca relevância. Assim, os sistemas da CEP são capazes de analisar milhares de registros de dados de forma muito rápida, reduzindo o atraso entre a receção e processamento de dados, tornando-se conveniente para a tomada de decisão. Acreditamos que esta capacidade operacional, poderia beneficiar qualquer sistema atual de gestão de dados. Mas, existem diversos tipos de sistemas de informação, aplicados a uma variedade de áreas empresariais, que operam em diferentes ambientes, além de exigir inúmeros métodos para garantir sua correta operacionalidade. Um tipo desses sistemas, são sistemas críticos, responsáveis por infraestruturas com grande impacto e que podem causar danos elevados às pessoas, à sociedade ou ao meio ambiente. Os sistemas críticos são responsáveis pela realização de operações em ambientes críticos, tais como armazenamento de água, estações de petróleo e atómicas, sistemas de veículos e aviários, dispositivos médicos, etc. Uma vez que esses sistemas são necessários para gerar resposta e alertas em tempo real, é possível que os motores CEP podem ser uma solução para melhorar o desempenho desses sistemas. Mas, os sistemas críticos possuem outros atributos de qualidade, como a proteção, a confiabilidade e a segurança. Neste trabalho, investigamos se os motores CEP podem ser usados em sistemas críticos e se são capazes de lidar com os atributos de qualidade desses sistemas. Depois de descrever os motores CEP, os sistemas críticos e seus atributos de qualidade, nos concentramos em segurança e proteção e fornecemos uma solução para a autenticidade de dados como mecanismo adicionado a um dos motores CEP mais populares, o ESPER. Concluímos que a solução proposta fornece autenticidade de dados, mas também tem um impacto considerável no desempenho
Veronika Abramova, Jorge Bernardino, Bruno Cabral 0001
COMPLEXIS2
2017 Using DataMining to Predict Diseases in Vineyards and Olive Groves
Rodrigo Rocha Silva, Jorge Bernardino
KEOD3
2017 NewSQL Databases - MemSQL and VoltDB Experimental Evaluation
Jorge Bernardino
KEOD2
2017 An Ontology for Clinical Decision Support System to Predict Female's Fertile Period
Francisco Vaz, Rodrigo Rocha Silva, Jorge Bernardino
KEOD3
2017 The Ability of Cloud Computing Performance Benchmarks to Measure Dependability
Eduardo Carvalho, Raul Barbosa, Jorge Bernardino
ICSOFT3
2017 Big Data Analytics: A Preliminary Study of Open Source Platforms
Jorge Nereu, Ana de Almeida 0001, Jorge Bernardino
ICSOFT3
2017 An Event Search Platform Using Machine Learning
abstract
Currently, the evolution of technology allows to find which events occurs around us at any given location.Social networks are one of the reasons of this trend and new applications are emerging aiming at finding and disclosing events.This paper proposes a platform of event searching.In particular, we propose a new architecture that uses machine learning to classify events with tags.An experimental evaluation with different types of algorithms was done using Facebook as a source of dataset events.
Marcelo Rodrigues, Rodrigo Rocha Silva, Jorge Bernardino
SEKE3
2017 Custom Process to Small Business
abstract
In this work we customized the RUP process with Scrum practices, and proposed a differentiate traceability matrix, applying in a small company.The experimental results show that our customization can be adopted as an alternative to a systematic and less-intrusive process.
Rodrigo Rocha Silva, Fernanda Yuri Kimura, Jorge Bernardino, Joubert de Castro Lima
SEKE3
2017 Comparative Analysis of Web Platform Assessment Tools
Solange Paz, Jorge Bernardino
WEBIST2
2017 Evaluation of Firewall Open Source Software
Diogo Sampaio, Jorge Bernardino
WEBIST2
2017 A Web Integration Framework for Cheap Flight Fares
abstract
Travel agencies offer their services via the Internet, which creates new methods of communication and connection between customers and third companies. Due to the difficulty that the management of large volume of flight routes represents, it is necessary to capture the information provided by airlines through a variety of services, providing end customers with competitive fares. In this paper, we analyse the information source of flight fares offered by airlines, studying the difficulties, limitations and costs involved in accessing these data. We also review the storage systems available to undertake a study on flight information obtained, serving mainly to show differences in fares for each possible route. A framework that explores the possibilities of finding "hidden" flight fares that result in much cheaper options in comparison to the average price of each flight route will be presented.
Manuel Sánchez, Juncal Gutiérrez-Artacho, Jorge Bernardino
WEBIST3
2017 Insider Attacks in a Non-secure Hadoop Environment
Pedro Camacho, Bruno Cabral 0001, Jorge Bernardino
WorldCIST (2)3
2017 Open Source CRM Tools for Small Companies
Telma Marques Cruz, Juncal Gutiérrez-Artacho, Jorge Bernardino
WorldCIST (1)3
2017 Business Intelligence for E-commerce: Survey and Research Directions
Tânia Ferreira, Isabel Pedrosa, Jorge Bernardino
WorldCIST (1)3
2017 Dashboards and Indicators for a BI Healthcare System
Sónia Rocha, Jorge Bernardino, Isabel Pedrosa, Ilda Ferreira
WorldCIST (1)2
2017 Describing and Comparing Big Data Querying Tools
Mário Rodrigues, Maribel Yasmina Santos, Jorge Bernardino
WorldCIST (1)3
2016 Evaluating Open Source Data Mining Tools for Business
abstract
Businesses are struggling to stay ahead of competition in a globalized economy where there are more and stronger competitors. Managers are constantly looking for advantages that can generate benefits at low costs. One way to have such advantage is using the data about customers, demographic data, purchase history, customer behavior and preferences that can help to take better business decisions. Data Mining addresses the challenges of collecting value inside data and the ways to put that value to use for virtually any area of our lives, including business. In this paper, we address the interest of Data Mining for business and analyze three popular Open Source Data Mining Tools – KNIME, Orange and RapidMiner – considered as a good starting point for enterprises to begin exploring the power of Data Mining and its benefits.
Pedro Daniel Coimbra de Almeida, Le Gruenwald, Jorge Bernardino
DATA3
2016 Cassandra's Performance and Scalability Evaluation
abstract
In the past, relational databases were the most commonly used technology for storing and retrieving data, allowing easier management and retrieval of any stored information organized as a set of tables. However, today databases are larger in size and the query execution time can become very long, requiring servers with bigger capacities. The purpose of this paper is to describe and analyze the Cassandra NoSQL database using the Yahoo! Cloud Serving Benchmark in order to better understand the execution capabilities for various types of applications in environments with different amounts of stored data. The experiments with Cassandra show good scalability and performance results and how the database size and number of nodes affect it.
Melyssa Barata, Jorge Bernardino
DATA2
2016 XSX: Lightweight Encryption for Data Warehousing Environments
Ricardo Jorge Santos, Marco Vieira, Jorge Bernardino
DaWaK3
2016 Open Source vs Proprietary Project Management Tools
Veronika Abramova, Francisco Pires, Jorge Bernardino
WorldCIST (1)3
2016 A Survey on Open Source Data Mining Tools for SMEs
Pedro Daniel Coimbra de Almeida, Jorge Bernardino
WorldCIST (1)2
2016 A Predictive Model for Exception Handling
João Ricardo Lourenço, Bruno Cabral 0001, Jorge Bernardino
WorldCIST (1)3
2015 What is BigQuery?
abstract
Big Data information is continuously increasing in volume and variety. This is the information that companies would like to quickly explore to identify strategic answers to the business. To overcome the problem of traditional database management systems to support large volumes of data arises Google BigQuery platform. This solution runs in the Cloud, SQL-like queries against massive quantities of data, providing real-time insights about the data. In this paper, we will analyze the main features of BigQuery that Google offers to manage large-scale data.
Sérgio Fernandes 0004, Jorge Bernardino
IDEAS2
2015 Big Data Issues
abstract
Big Data is a new trend regarded by both academics and business areas as an interesting concept. The paradigm includes the storage and processing of petabyte-size datasets, boosting knowledge discovery over data and providing organizations with competitive advantage over their contenders. This paper comes to provide an overview of the concept, answering questions that one has when faced with the term for the first time: What is Big Data and what are its advantages? How does it work? How is it accepted among enterprises? How to deploy a Big Data solution?
Pedro Caldeira Neves, Jorge Bernardino
IDEAS2
2015 A Survey on Data Quality: Classifying Poor Data
abstract
Data is part of our everyday life and an essential asset in numerous businesses and organizations. The quality of the data, i.e., the degree to which the data characteristics fulfill requirements, can have a tremendous impact on the businesses themselves, the companies, or even in human lives. In fact, research and industry reports show that huge amounts of capital are spent to improve the quality of the data being used in many systems, sometimes even only to understand the quality of the information in use. Considering the variety of dimensions, characteristics, business views, or simply the specificities of the systems being evaluated, understanding how to measure data quality can be an extremely difficult task. In this paper we survey the state of the art in classification of poor data, including the definition of dimensions and specific data problems, we identify frequently used dimensions and map data quality problems to the identified dimensions. The huge variety of terms and definitions found suggests that further standardization efforts are required. Also, data quality research on Big Data appears to be in its initial steps, leaving open space for further research.
Nuno Laranjeiro, Seyma Nur Soydemir, Jorge Bernardino
PRDC3
2015 An Overview of Decision Support Benchmarks: TPC-DS, TPC-H and SSB
Melyssa Barata, Jorge Bernardino, Pedro Furtado 0001
WorldCIST (1)2
2015 Scalability of Facebook Architecture
Hugo Barrigas, Daniel Barrigas, Melyssa Barata, Jorge Bernardino, Pedro Furtado 0001
WorldCIST (1)4
2015 Commercial Business Intelligence Suites Comparison
Joaquim Lapa, Jorge Bernardino, Ana de Almeida 0001
WorldCIST (1)2
2015 NoSQL Databases: A Software Engineering Perspective
João Ricardo Lourenço, Veronika Abramova, Marco Vieira, Bruno Cabral 0001, Jorge Bernardino
WorldCIST (1)5
2015 Open Source Backup Systems for SMEs
Diogo Sampaio, Jorge Bernardino
WorldCIST (1)2
2014 Evaluating Cassandra Scalability with YCSB
Veronika Abramova, Jorge Bernardino, Pedro Furtado 0001
DEXA (2)2
2014 Survey on Big Data and Decision Support Benchmarks
Melyssa Barata, Jorge Bernardino, Pedro Furtado 0001
DEXA (2)2
2014 Open source business intelligence in manufacturing
abstract
In an increasingly competitive environment the manufacturing industry faces many challenges, like global competition and ever more demanding customers. This leads companies to be proactive in managing and to take advantage of corporate data if they want to keep up or stay ahead of the competition. That's where Business Intelligence can improve company's competitive edge. In this paper we show the suitability of an open source business intelligence tool to manufacturing industry.
Eduardo Jesus, Jorge Bernardino
IDEAS2
2014 An overview of openstack architecture
abstract
Cloud Computing concept refers to both the applications delivered as services over the Internet and the servers and system software in the datacenters that provide those services. These solutions offer pools of virtualized computing resources, paid on a pay-per-use basis, and drastically reduce the initial investment and maintenance costs. Efficient and flexible resource management is the main focus for the cloud solutions on the market as well as scalability and adaptability to new environments. Openstack exceeded the market as a scalable, performant and highly adaptive open source architecture for both public and private cloud solutions as well as leveraging from hardware resources either they be professional or entry level. This paper gives an overview of Openstack software components functionalities in order to design and implement unique cloud computing solutions to fit enterprises purposes.
Tiago Rosado, Jorge Bernardino
IDEAS2
2014 Testing Cloud Benchmark Scalability with Cassandra
abstract
NoSQL databases were developed as highly scalable databases that allow easy data distribution over a number of servers. With the increased interest of researchers and companies in non-relational technology, NoSQL databases became widely used and a common belief emerged defending that those engines scale well. This means that the use of more nodes would result in reduced execution time of requests and the system would scale adequately, by adding nodes proportionally to data size and load. However, sometimes, adding nodes may not result in improvement of request-serving time. Therefore, it is useful to investigate how different factors, such as workload, data size and number of simultaneous sessions influence scaling capabilities. We will review the architecture of Cassandra, which is known for being one of the most efficient NoSQL engines, and analyze its scalability, using the Yahoo Cloud Serving Benchmark. The results will allow a better understanding of scalability and scalability limitations in that type of environment.
Veronika Abramova, Jorge Bernardino, Pedro Furtado 0001
SERVICES2
2013 A Specific Encryption Solution for Data Warehouses
Ricardo Jorge Santos, Deolinda Dias Rasteiro, Jorge Bernardino, Marco Vieira
DASFAA (2)3
2012 Leveraging 24/7 Availability and Performance for Distributed Real-Time Data Warehouses
abstract
Real-time Data Warehouses (DWs) must be able to deal with continuous updates while ensuring 24/7 availability. To improve their performance, distributing data using round-robin algorithms on clusters of shared-nothing machines is normally used. This paper proposes a solution for distributed DW databases that ensures its continuous availability and deals with frequent data loading requirements, while adding small performance overhead. We use a data striping and replication architecture to distribute portions of each fact table among pairs of slave nodes, where each slave node is an exact replica of its partner. This allows balancing query execution and replacing any defective node, ensuring the system's continuous availability. The size of each portion in a given node depends on its individual features, namely performance benchmark measures and dedicated database RAM. The estimated cost for executing each query workload in each slave node is also used for balancing query performance. We include experiments using the TPC-H decision support benchmark to evaluate the scalability of the proposed solution and show that it outperforms standard round-robin distributed DW setups.
Ricardo Jorge Santos, Jorge Bernardino, Marco Vieira
COMPSAC2
2012 Evaluating the Feasibility Issues of Data Confidentiality Solutions from a Data Warehousing Perspective
Ricardo Jorge Santos, Jorge Bernardino, Marco Vieira
DaWaK2
2012 Securing Data Warehouses from Web-Based Intrusions
Ricardo Jorge Santos, Jorge Bernardino, Marco Vieira, Deolinda Dias Rasteiro
WISE2
2011 24/7 Real-Time Data Warehousing: A Tool for Continuous Actionable Knowledge
abstract
Technological evolution has redefined many business models. Many decision makers are now required to act near real-time, instead of periodically, given the latest transactional information. Decision-making occurs much more frequently and considers the latest business data. Since data warehouses (DWs) are the core of business intelligence, decision support systems need to deal with 24/7 real-time requirements. Thus, the ability to deal with continuous data loading and decision support availability simultaneously is critical, for producing continuous actionable knowledge. The main challenge in this context is to efficiently manage the DW's refreshment, when data sources change, to recapture consistency and accuracy with those sources, while maintaining OLAP availability and database performance. This paper proposes a simple, fast and efficient solution based on database replication and temporary tables to change a traditional enterprise DW into a real-time DW, enabling continuous data loading and OLAP availability on a 24/7 schedule. Experimental evaluations using a real-world DW and the TPC-H decision support benchmark show its advantages and analyze its impact in OLAP performance.
Ricardo Jorge Santos, Jorge Bernardino, Marco Vieira
COMPSAC2
2011 A data masking technique for data warehouses
abstract
Data Warehouses (DWs) are the enterprise's most valuable asset in what concerns critical business information, making them an appealing target for attackers. Packaged database encryption solutions are considered the best solution to protect sensitive data. However, given the volume of data typically processed by DW queries, the existing encryption solutions heavily increase storage space and introduce very large overheads in query response time, due to decryption costs. In many cases, this performance degradation makes encryption unfeasible for use in DWs. In this paper we propose a transparent data masking solution for numerical values in DWs based on the mathematical modulus operator, which can be used without changing user application and DBMS source code. Our solution provides strong data security while introducing small overheads in both storage space and database performance. Several experimental evaluations using the TPC-H decision support benchmark and a real-world DW are included. The results show the overall efficiency of our proposal, demonstrating that it is a valid alternative to existing standard encryption routines for enforcing data confidentiality in DWs.
Ricardo Jorge Santos, Jorge Bernardino, Marco Vieira
IDEAS2
2011 Balancing Security and Performance for Enhancing Data Privacy in Data Warehouses
abstract
Data Warehouses (DWs) store the golden nuggets of the business, which makes them an appealing target. To ensure data privacy, encryption solutions have been used and proven efficient in their security purpose. However, they introduce massive storage space and performance overheads, making them unfeasible for DWs. We propose a data masking technique for protecting sensitive business data in DWs that balances security strength with database performance, using a formula based on the mathematical modular operator. Our solution manages apparent randomness and distribution of the masked values, while introducing small storage space and query execution time overheads. It also enables a false data injection method for misleading attackers and increasing the overall security strength. It can be easily implemented in any DataBase Management System (DBMS) and transparently used, without changes to application source code. Experimental evaluations using a real-world DW and TPC-H decision support benchmark implemented in leading commercial DBMS Oracle llg and Microsoft SQL Server 2008 demonstrate its overall effectiveness. Results show substantial savings of its implementation costs when compared with state of the art data privacy solutions provided by those DBMS and that it outperforms those solutions in both data querying and insertion of new data.
Ricardo Jorge Santos, Jorge Bernardino, Marco Vieira
TrustCom2
2010 A 24/7 monitorization tool for avoiding hypotensive episodes in critical care
abstract
The sudden fall of blood pressure (hypotension) is a common complication in medical care. In critical care patients, hypotension (HT) may cause serious heart, endocrine or neurological disorders, inducing severe or even lethal events. Moreover, recent studies report an increase of mortality in HT prone hemodialysis patients in need of critical care. If HT could be predicted in advance, medical staff could take action to minimize its effects, or even avoid its occurrence. Typically, most medical systems have focused on monitoring and detecting current patient status, rather than determining biosignal trends or predicting a patient's future status. Therefore, predicting HT episodes in advance remains a challenge. Furthermore, since critical care actions such as hemodialysis are oftenly inconvenient and uncomfortable procedures, HT prediction or detection methods should be non-invasive, whenever possible. In this paper, we present a solution for continuous monitorization and prediction of HT episodes, using heart rate (HR) and mean blood pressure (BP) non-invasive measured biosignals. We propose an architecture for a HT Predictor (HTP) Tool, presenting a set of tools and a real-time database capable of continuously storing and real-time monitoring all patient's historical HR and BP biosignal data, and efficiently alerting both probable and detected occurrences of HT episodes for each patient for the following 60 minutes. Additionally, the system promotes medical staff mobility, by taking advantage of using mobile personal devices such as mobile phones and PDA's, optimizing human resources. Finally, an experimental evaluation on real-life data from the well known Physionet database shows the efficiency of the tool, outperforming the winning proposal of the Physionet 2009 Challenge.
Ricardo Jorge Santos, Jorge Bernardino, Jorge Henriques
IDEAS2
2009 A Query Cache Tool for Optimizing Repeatable and Parallel OLAP Queries
Ricardo Jorge Santos, Jorge Bernardino
DEXA2
2009 Optimizing data warehouse loading procedures for enabling useful-time data warehousing
abstract
The purpose of a data warehouse is to aid decision making. As the real-time enterprise evolves, synchronism between transactional data and data warehouses is redefined. To cope with real-time requirements, the data warehouses must be able to enable continuous data integration, in order to deal with the most recent business data. Traditional data warehouses are unable to support any dynamics in structure and content while they are available for OLAP. Their data is periodically updated because they are unprepared for continuous data integration. For real-time enterprises with needs in decision support while the transactions are occurring, (near) real-time data warehousing seem very promising. In this paper we present a survey on testing today's most used loading techniques and analyze which are the best data loading methods, presenting a methodology for efficiently supporting continuous data integration for data warehouses. To accomplish this, we use techniques such as table structure replication with minimum content and query predicate restrictions for selecting data, to enable loading data in the data warehouse continuously, with minimum impact in query execution time. We demonstrate the efficiency of the method using benchmark TPC-H and executing query workloads while simultaneously performing continuous data integration.
Ricardo Jorge Santos, Jorge Bernardino
IDEAS2
2008 Efficient Data Distribution for DWS
Raquel Almeida 0002, Jorge Vieira, Marco Vieira, Henrique Madeira, Jorge Bernardino
DaWaK5
2008 Real-time data warehouse loading methodology
abstract
A data warehouse provides information for analytical processing, decision making and data mining tools. As the concept of real-time enterprise evolves, the synchronism between transactional data and data warehouses, statically implemented, has been redefined. Traditional data warehouse systems have static structures of their schemas and relationships between data, and therefore are not able to support any dynamics in their structure and content. Their data is only periodically updated because they are not prepared for continuous data integration. For real-time enterprises with needs in decision support purposes, real-time data warehouses seem to be very promising. In this paper we present a methodology on how to adapt data warehouse schemas and user-end OLAP queries for efficiently supporting real-time data integration. To accomplish this, we use techniques such as table structure replication and query predicate restrictions for selecting data, to enable continuously loading data in the data warehouse with minimum impact in query execution time. We demonstrate the efficiency of the method by analyzing its impact in query performance using benchmark TPC-H executing query workloads while simultaneously performing continuous data integration at various insertion time rates.
Ricardo Jorge Santos, Jorge Bernardino
IDEAS2
2006 Global Epidemiological Outbreak Surveillance System Architecture
abstract
Diseases such as avian influenza, severe acute respiratory syndrome (SARS) and Creutzfeldt-Jacob syndrome represent a new era of biological threats. Nowadays, these hazards breed, mutate and evolve at tremendous speed. Furthermore, they may spread out at the same speed as which we travel. This reveals an urgent need for an agent capable of dealing with such threats. Data warehouses are databases which provide decision support by on-line analytical processing (OLAP) techniques. We present the architecture for an effective information system infrastructure enabling the prediction and near real-time detection of disease outbreaks, using knowledge extraction algorithms to explore a symptoms/diseases data warehouse in a continuous and active form. To collect such data, we take advantage of the Internet and features existing in today?s common communication devices such as personal computers, portable digital assistants and cellular phones. We present a case-simulation based on a small country, showing the system can detect an outbreak within hours or even minutes after its physical occurrence, alerting health decision makers and providing quick interaction and feedback between all users. The architecture is also functionally independent from its geographical dimension.
Ricardo Jorge Santos, Jorge Bernardino
IDEAS2
2005 Efficient Compression of Text Attributes of Data Warehouse Dimensions
Jorge Vieira, Jorge Bernardino, Henrique Madeira
DaWaK2
2002 DWS-AQA: A Cost Effective Approach for Very Large Data Warehouses
abstract
Data warehousing applications typically involve massive amounts of data that push database management technology to the limit. A scalable architecture is crucial, not only to handle very large amount of data but also to assure interactive response time to the users. Large data warehouses require a very expensive setup, typically based on high-end servers or high-performance clusters. In this paper we propose and evaluate a simple but very effective method to implement a data warehouse using the computers and workstations typically available in large organizations. The proposed approach is called data warehouse striping with approximate query answering (DWS-AQA). The goal is to use the processing and disk capacity normally available in large workstation networks to implement a data warehouse with a very reduced infrastructure cost. As the data warehouse shares computers that are also being used for other purposes, most of the times only a fraction of the computers will be able to execute the partial queries in time. However, as we show in the paper, the approximated answers estimated from partial results have a very small error for most of the plausible scenarios. Moreover, as the data warehouse facts are partitioned in a strict uniform way, it is possible to calculate tight confidence intervals for the approximated answers, providing the user with a measure of the accuracy of the query results. A set of experiments on the TPC-H benchmark database is presented to show the accuracy of DWS-AQA for a large number of scenarios.
Jorge Bernardino, Pedro Furtado 0001, Henrique Madeira
IDEAS1
2002 Approximate Query Answering Using Data Warehouse Striping
Jorge Bernardino, Pedro Furtado 0001, Henrique Madeira
J. Intell. Inf. Syst.1
2001 Approximate Query Answering Using Data Warehouse Striping
Jorge Bernardino, Pedro Furtado 0001, Henrique Madeira
DaWaK1
2001 Experimental Evaluation of a New Distributed Partitioning Technique for Data Warehouses
abstract
Since data warehousing has become a major field of research there has been a lot of interest in reducing the response time of complex queries posed over the very large databases. The problem is that data warehouses store large amounts of data for decision support, requiring a high level of query performance and scalability to the database engines. A novel round-robin data partitioning approach especially designed for relational data warehouse environments is proposed and experimentally evaluated. This approach is specific to data warehouses implemented over relational repositories using the star schema, as it takes advantage of the specific characteristics of star schemas and typical data warehouse query profiles. The proposed approach guarantees optimal load balancing of query execution and assures high scalability. The experimental evaluation presented in the paper, using a comprehensive set of typical queries from the APB-I benchmark running over Oracle 8, shows that an optimal speedup can be obtained with this technique. The proposed technique constitutes an effective and practical way of coping with very large data warehouses and can be applied to existing database technology.
Jorge Bernardino, Henrique Madeira
IDEAS1