Vincent-nam Dang

dblp:301/3675 · also Vincent-Nam Dang · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
5since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Unified access to interdisciplinary open data platforms: Open Science Data Network
abstract
Open Science is based on a collaborative network to develop transparent, accessible, and shared knowledge. Open Research Data Platforms (ORDPs) are deployed to fulfill the needs for data sharing of a specific community and/or scientific discipline. The high variety of research areas creates a barrier to data sharing between research entities. To enable this research data to be found by the research entities that need it, it is necessary to establish access to different ORDPs that are unknown to these research entities. The goal of this article is to provide a quantitative analysis showing the current limitations of data sharing between ORDPs in Open Science. We then propose a solution to improve data access and sharing based on theoretical foundations and an experimental approach. We propose to extend our theoretical interoperability model, which helps us to define the necessary steps to interoperate ORDPs. We present and discuss a quantitative evaluation of ORDPs’ interoperability. Based on this exploratory study, we propose a solution that enables research entities to discover unknown ORDPs, thereby facilitating access to relevant data. This solution is the Open Science Data Network (OSDN), a decentralized, distributed, and federated network of ORDPs that integrates a query propagation process and robustness features. To enable the deployment of OSDN at an Open Science scale, we designed our solution by considering its adoption cost relative to a non-organized interoperability approach. With two ORDPs integrated into the OSDN, the adoption cost is estimated to be reduced by at least 17%. This reduction approaches 100% as the number of integrated ORDPs increases. To demonstrate the feasibility of the solution, we developed a Proof of Concept (POC) and applied it to two research projects from different domains and involving distinct research communities. For the first research project, we measured a 7% increase in the volume of accessed data and an 80% reduction in the time needed to find this data. In addition, researcher from this experiment was able to formulate new intra- and interdisciplinary research questions thanks to the newly accessed data. In the second research project, we observed an increase in data volume of up to a factor of 3968. More importantly, this process led to the discovery of new essential data that was previously missing.
Vincent-nam Dang, Nathalie Aussenac-Gilles, Imen Megdiche, Franck Ravat
Data Knowl. Eng.1
2024 OSDN: An Open Science Data Network for Interdisciplinary Research
Vincent-nam Dang, Nathalie Aussenac-Gilles, Imen Megdiche, Franck Ravat
DASFAA (7)1
2024 Enabling Interdisciplinary Research in Open Science: Open Science Data Network
Vincent-nam Dang, Nathalie Aussenac-Gilles, Imen Megdiche, Franck Ravat
RCIS (1)1
2023 Interoperability of Open Science Metadata: What About the Reality?
Vincent-nam Dang, Nathalie Aussenac-Gilles, Imen Megdiche, Franck Ravat
RCIS1
2021 A Zone-Based Data Lake Architecture for IoT, Small and Big Data
abstract
Data lakes are supposed to enable analysts to perform more efficient and efficacious data analysis by crossing multiple existing data sources, processes and analyses. However, it is impossible to achieve that when a data lake does not have a metadata governance system that progressively capitalizes on all the performed analysis experiments. The objective of this paper is to have an easily accessible, reusable data lake that capitalizes on all user experiences. To meet this need, we propose an analysis-oriented metadata model for data lakes. This model includes the descriptive information of datasets and their attributes, as well as all metadata related to the machine learning analyzes performed on these datasets. To illustrate our metadata solution, we implemented a web application of data lake metadata management. This application allows users to find and use existing data, processes and analyses by searching relevant metadata stored in a NoSQL data store within the data lake. To demonstrate how to easily discover metadata with the application, we present two use cases, with real data, including datasets similarity detection and machine learning guidance.
Yan Zhao 0022, Imen Megdiche, Franck Ravat, Vincent-nam Dang
IDEAS4