Moditha Hewasinghage

dblp:206/6189 · DBLP profile ↗
← Back
9ranked-venue papers
7as first author
4since 2021 · last 2023
0000-0001-6288-2043ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 9 · 7 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2023 Automated database design for document stores with multicriteria optimization
abstract
Abstract Document stores have gained popularity among NoSQL systems mainly due to the semi-structured data storage structure and the enhanced query capabilities. The database design in document stores expands beyond the first normal form by encouraging de-normalization through nesting. This hinders the process, as the number of alternatives grows exponentially with multiple choices in nesting (including different levels) and referencing (including the direction of the reference). Due to this complexity, document store data design is mostly carried out in trial-and-error or ad-hoc rule-based approaches. However, the choices affect multiple, often conflicting, aspects such as query performance, storage space, and complexity of the documents. To overcome these issues, in this paper, we apply multicriteria optimization. Our approach is driven by a query workload and a set of optimization objectives. First, we formalize a canonical model to represent alternative designs and introduce an algebra of transformations that can systematically modify a design. Then, using these transformations, we implement a local search algorithm driven by a loss function that can propose near-optimal designs with high probability. Finally, we compare our prototype against an existing document store data design solution purely driven by query cost, where our proposed designs have better performance and are more compact with less redundancy.
Moditha Hewasinghage, Sergi Nadal, Alberto Abelló, Esteban Zimányi
Knowl. Inf. Syst.1
2021 DocDesign 2.0: Automated Database Design for Document Stores with Multi-criteria Optimization
abstract
We present DocDesign 2.0, a novel system that supports database design for document stores. DocDesign 2.0 automatically generates a document store design driven by a query workload and a set of optimization objectives. In the presence of a massive search space, DocDesign 2.0 adopts multi-objective optimization techniques that, with high probability, guarantee to yield the optimal design based on the preferences (i.e., weights) provided by the end-user. In this paper, we demonstrate how DocDesign 2.0 improves the productivity on the task of designing a document store, as well as how the quality of the results is improved with respect to those obtained by manually generating the design.
Moditha Hewasinghage, Sergi Nadal, Alberto Abelló
EDBT1
2021 Managing polyglot systems metadata with hypergraphs
abstract
A single type of data store can hardly fulfill every end-user requirements in the NoSQL world. Therefore, polyglot systems use different types of NoSQL datastores in combination. However, the heterogeneity of the data storage models makes managing the metadata a complex task in such systems, with only a handful of research carried out to address this. In this paper, we propose a hypergraph-based approach for representing the catalog of metadata in a polyglot system. Taking an existing common programming interface to NoSQL systems, we extend and formalize it as hypergraphs. Then, we define design constraints and query transformation rules for three representative data store types. Next, we propose a simple query rewriting algorithm from the metadata of the catalog to underlying data store specific ones and provide a prototype implementation. Furthermore, we introduce a storage statistics estimator on the underlying data stores. Finally, we show the feasibility of our approach on a use case of an existing polyglot system, and its usefulness in metadata and physical query path calculations.
Moditha Hewasinghage, Alberto Abelló, Jovan Varga, Esteban Zimányi
Data Knowl. Eng.1
2021 A cost model for random access queries in document stores
Moditha Hewasinghage, Alberto Abelló, Jovan Varga, Esteban Zimányi
VLDB J.1
2020 DocDesign: Cost-Based Database Design for Document Stores
abstract
Document stores have become one of the most popular NoSQL systems, mainly due to their semi-structured data storage structure and well-developed query capabilities. The semi-structured nature allows them to have database designs beyond traditional normalization theories. This makes the database design decisions more complicated with a myriad of possibilities. Thus, the database design process for them has resorted to ad-hoc trial and error methods. However, having a good database design is essential for any data storage system’s performance, and bad design decisions cannot always be compensated by adding more powerful hardware. Thus, in this work, we propose DocDesign, a decision aid tool for document store database design. DocDesign allows its users to evaluate different database designs for data storage requirements under a particular workload. Through DocDesign, users can make informed decisions for a design by evaluating the estimated storage statistics and query runtimes without testing it on an actual document store. DocDesign also generates design specific queries for the input workload. This not only cuts down the time and the effort taken in design decision making and development but also save money spent on fixing poor designs in the long run. On-site, we will showcase how DocDesign facilitates the design decision-making process for MongoDB with both synthetic and real-world examples.
Moditha Hewasinghage, Alberto Abelló, Jovan Varga, Esteban Zimányi
SSDBM1
2018 Modeling Strategies for Storing Data in Distributed Heterogeneous NoSQL Databases
Moditha Hewasinghage, Nacéra Bennacer Seghouani, Francesca Bugiotti
ER1
2018 Managing Polyglot Systems Metadata with Hypergraphs
Moditha Hewasinghage, Jovan Varga, Alberto Abelló, Esteban Zimányi
ER1
2018 A Frequent Named Entities-Based Approach for Interpreting Reputation in Twitter
abstract
Twitter is a social network that provides a powerful source of data. The analysis of those data offers many challenges among those stands out the opportunity to find reputation of a product, a person or any other entity of interest. Several approaches for sentiment analysis have been proposed in the literature to assess the general opinion expressed in tweets on an entity. Nevertheless, these methods aggregate sentiment scores retrieved from tweets, which is a static view to evaluate the overall reputation of an entity. The reputation of an entity is not static; entities collaborate with each other, and they get involved in different events over time. A simple aggregation of sentiment scores is then not sufficient to represent this dynamism. In this paper, we present a new approach to determine the reputation of an entity on the basis of the set of events in which it is involved. To achieve this, we propose a new sampling method driven by a tweet weighting measure to give a better quality and summary of the target entity. We introduce the concept of Frequent Named Entities to determine the events involving the target entity. Our evaluation achieved for different entities shows that 90% of the reputation of an entity originates from the events it is involved in and the breakdown into events allows interpreting the reputation in a transparent and self-explanatory way.
Nacéra Bennacer Seghouani, Francesca Bugiotti, Moditha Hewasinghage, Suela Isaj, Gianluca Quercini
Data Sci. Eng.3
2017 Interpreting Reputation Through Frequent Named Entities in Twitter
Nacéra Bennacer Seghouani, Francesca Bugiotti, Moditha Hewasinghage, Suela Isaj, Gianluca Quercini
WISE (1)3