Shaoyi Yin

dblp:99/3212 · DBLP profile ↗
← Back
18ranked-venue papers
7as first author
4since 2021 · last 2025
0000-0002-5335-2443ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 14 · 7 first-author · 3 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 MTD-DS: An SLA-Aware Decision Support Benchmark for Multi-Tenant Parallel DBMSs
abstract
Multi-tenant DBMSs are used by cloud providers for their Database-as-a-Service products. They could be single-node DBMSs installed in virtual machines, SQL-on-Hadoop systems or classic parallel relational DBMSs running on top of a shared-nothing or shared-disk architecture. For a cloud provider, it is interesting to measure these systems’ capability of dealing with multi-tenant workloads, i.e., taking advantage of the statistical multiplexing to obtain economic gain while being attractive by providing a good quality of service and a low bill to the tenants. In this paper, we present MTD-DS benchmark (with MTD for Multi-Tenant parallel DBMSs and DS for Decision Support). MTD-DS extends TPC-DS by adding a multi-tenant query workload generator, a performance Service Level Objectives generator, configurable Database-as-a-Service pricing models, and new metrics to measure the potential capability of a multi-tenant parallel DBMS in obtaining the best trade-off between the provider's benefit and the tenants’ satisfaction. Example experimental results have been produced to show the relevance and the feasibility of the MTD-DS benchmark.
Shaoyi Yin, Franck Morvan, Jorge Martinez-Gil, Abdelkader Hameurlain
IEEE Trans. Knowl. Data Eng.1
2022 Multi-Cloud Query Optimisation with Accurate and Efficient Quoting
abstract
A recent trend among major organisations is to release their datasets in the cloud over various Database-as-a-Service (DBaaS) providers’ premises, creating a use case for multi-cloud querying. As identified in the literature, middlewares with such capabilities should quote the monetary cost and the response time of the queries in order to gain the trust of their users, and also optimise the queries so as to avoid cost overruns and meet the quotations. Considering those requirements, this paper introduces an accurate cost model and an efficient execution plan search strategy for dealing with large-scale multi-cloud queries. The former is an ensemble learning stack leveraging online machine learning models, and the latter is a randomised method inspired by iterative improvement. We evaluated our middleware over simulated providers by using the Join Order Benchmark. Experiments showed that the cost model manages to correct the estimations from the providers. The randomised strategy can produce more efficiently execution plans that yield better performances and a lower monetary cost compared to an exhaustive approach from previous work.
Damien T. Wojtowicz, Shaoyi Yin, Jorge Martinez-Gil, Franck Morvan, Abdelkader Hameurlain
IEEE Big Data2
2021 Cost-Effective Dynamic Optimisation for Multi-Cloud Queries
abstract
The provision of public data through various Database-as-a-Service (DBaaS) providers has recently emerged as a significant trend, backed by major organisations. This paper introduces Nebula, a non-profit middleware providing multi-cloud querying capabilities by fully outsourcing its users' queries to the involved DBaaS providers. First, we propose a quoting procedure for those queries, whose need stems from the pay-per-query policy of the providers. Those quotations contain monetary cost and response time estimations, and are computed using provider-generated tenders. Then, we present an agent-based dynamic optimisation engine that orchestrates the outsourced execution of the queries. Agents within this engine cooperate in order to meet the quoted values. We evaluated Nebula over simulated providers by using the Join Order Benchmark (JOB). Experimental results showed Nebula's approach is, in most cases, more competitive in terms of monetary cost and response time than existing work in the multi-cloud DBMS literature.
Damien T. Wojtowicz, Shaoyi Yin, Franck Morvan, Abdelkader Hameurlain
CLOUD2
2021 Matching Large Biomedical Ontologies Using Symbolic Regression
abstract
The problem of ontology matching consists of finding the semantic correspondences between two ontologies that, although belonging to the same domain, have been developed separately. Matching methods are of great importance since they allow us to find the pivot points from which an automatic data integration process can be established. Unlike the most recent developments based on deep learning, this study presents our research on the development of new methods for ontology matching that are accurate and interpretable at the same time. For this purpose, we rely on a symbolic regression model specifically trained to find the mathematical expression that can solve the ground truth accurately, with the possibility of being understood by a human operator and forcing the processor to consume as little energy as possible. The experimental evaluation results show that our approach seems to be promising.
Jorge Martinez-Gil, Shaoyi Yin, Josef Küng, Franck Morvan
iiWAS2
2020 SLA-driven resource re-allocation for SQL-like queries in the cloud
Mohamed Mehdi Kandi, Shaoyi Yin, Abdelkader Hameurlain
Knowl. Inf. Syst.2
2018 SLA Definition for Multi-Tenant DBMS and its Impact on Query Optimization
abstract
In the cloud context, users are often called tenants. A cloud DBMS shared by many tenants is called a multi-tenant DBMS. The resource consolidation in such a DBMS allows the tenants to only pay for the resources that they consume, while providing the opportunity for the provider to increase its economic gain. For this, a Service Level Agreement (SLA) is usually established between the provider and a tenant. However, in the current systems, the SLA is often defined by the provider, while the tenant should agree with it before using the service. In addition, only the availability objective is described in the SLA, but not the performance objective. In this paper, an SLA negotiation framework is proposed, in which the provider and the tenant define the performance objective together in a fair way. To demonstrate the feasibility and the advantage of this framework, we evaluate its impact on query optimization. We formally define the problem by including the cost-efficiency aspect, we design a cost model and study the plan search space for this problem, we revise two search methods to adapt to the new context, and we propose a heuristic to solve the resource contention problem caused by concurrent queries of multiple tenants. We also conduct a performance evaluation to show that, our optimization approach (i.e., driven by the SLA) can be much more cost-effective than the traditional approach which always minimizes the query completion time.
Shaoyi Yin, Abdelkader Hameurlain, Franck Morvan
IEEE Trans. Knowl. Data Eng.1
2016 Adaptive Join Operator for Federated Queries over Linked Data Endpoints
Damla Oguz, Shaoyi Yin, Abdelkader Hameurlain, Belgin Ergenç, Oguz Dikenelli
ADBIS2
2014 MILo-DB: a personal, secure and portable database machine
Nicolas Anciaux, Luc Bouganim, Philippe Pucheral, Yanli Guo, Lionel Le Folgoc, Shaoyi Yin
Distributed Parallel Databases6
2013 Resource Allocation for Query Optimization in Data Grid Systems: Static Load Balancing Strategies
Shaoyi Yin, Igor Epimakhov, Franck Morvan, Abdelkader Hameurlain
ADBIS1
2013 Dynamic Multi-probe LSH: An I/O Efficient Index Structure for Approximate Nearest Neighbor Search
Shaoyi Yin, Mehdi Badr, Dan Vodislav
DEXA (1)1
2013 Multi-criteria search algorithm: An efficient approximate k-NN algorithm for image retrieval
abstract
We propose a new method for approximate k-NN search in large scale image databases, based on top-k multi-criteria search techniques. The method defines a simple index structure based on sorted lists, which provides a good compromise between fast retrieval, storage requirements and update cost. The search algorithm delivers approximate results with guarantees about false negatives, with fast emergence of good approximations, monotonically improved and leading if necessary to an exact result. Experiments with the on-disk implementation show that our method produces very good approximate results several times faster than the Baseline method.
Mehdi Badr, Dan Vodislav, David Picard, Shaoyi Yin, Philippe Henri Gosselin
ICIP4
2013 Mobile Agent-based Dynamic Resource Allocation Method for Query Optimization in Data Grid Systems
abstract
Resource allocation is one of the principal stages of query processing in relational data grid systems. Specific characteristics of the data grid environment, such as dynamicity, heterogeneity and large scale, impose serious restrictions to the resource allocation process. Static resource allocation before the query execution may be far from optimal due to the dynamic changes of the system. One possible optimization is to adjust dynamically the allocation of resources during the query execution. Some methods of dynamic resource allocation have been proposed, however, most of them use centralized control mechanisms. In this study we argue that the decentralized approach meets better the requirements of the data grid systems. In this study we propose a decentralized method of dynamic resource allocation that is based on the mobile agent paradigm. We consider the participating nodes as autonomous and independent elements of the system, each of which can detect if it is overloaded and make the decision to react. Then we consider each relational operation as a mobile agent running on the allocated node, meaning that, it keeps track of its own status and can migrate to another node at any time. A two-level cooperation mechanism between such autonomous nodes and autonomous operations is described in detail. Performance evaluation proves the efficiency of the proposed method.
Igor Epimakhov, Abdelkader Hameurlain, Franck Morvan, Shaoyi Yin
KES-AMSTA4
2012 PBFilter: A flash-based indexing scheme for embedded systems
Shaoyi Yin, Philippe Pucheral
Inf. Syst.1
2010 Pluggable personal data servers
abstract
An increasing amount of personal data is automatically gathered on servers by administrations, hospitals and private companies while several security surveys highlight the failure of database servers to keep confidential data really private. The advent of powerful secure tokens, combining the security of smart card microcontrollers with the storage capacity of NAND Flash chips, introduces a credible alternative to the systematic centralization of personal data. By embedding a full-fledged database server in such device, an individual can now store her personal data in her own secure token, kept under her control, and never disclose in clear her private data to the outside untrusted world. This demonstration shows the benefit of the proposed approach in terms of privacy protection and pervasiveness through a healthcare scenario. This scenario is extracted from a field experiment where medical folders embedded in secure tokens are used to improve the coordination of medical care at home for elderly people. The demonstration also highlights interesting features of the embedded DBMS engine introduced to tackle the secure token's strong hardware constraints.
Nicolas Anciaux, Luc Bouganim, Yanli Guo, Philippe Pucheral, Jean-Jacques Vandewalle, Shaoyi Yin
SIGMOD Conference6
2010 Secure Personal Data Servers: a Vision Paper
abstract
An increasing amount of personal data is automatically gathered and stored on servers by administrations, hospitals, insurance companies, etc. Citizen themselves often count on internet companies to store their data and make them reliable and highly available through the internet. However, these benefits must be weighed against privacy risks incurred by centralization. This paper suggests a radically different way of considering the management of personal data. It builds upon the emergence of new portable and secure devices combining the security of smart cards and the storage capacity of NAND Flash chips. By embedding a full-fledged Personal Data Server in such devices, user control of how her sensitive data is shared by others (by whom, for how long, according to which rule, for which purpose) can be fully reestablished and convincingly enforced. To give sense to this vision, Personal Data Servers must be able to interoperate with external servers and must provide traditional database services like durability, availability, query facilities, transactions. This paper proposes an initial design for the Personal Data Server approach, identifies the main technical challenges associated with it and sketches preliminary solutions. We expect that this paper will open exciting perspectives for future database research.
Tristan Allard, Nicolas Anciaux, Luc Bouganim, Yanli Guo, Lionel Le Folgoc, Benjamin Nguyen, Philippe Pucheral, Indrajit Ray, Indrakshi Ray, Shaoyi Yin
Proc. VLDB Endow.10
2009 A sequential indexing scheme for flash-based embedded systems
abstract
NAND Flash has become the most popular stable storage medium for embedded systems. As on-board storage capacity increases, the need for efficient indexing techniques arises. Such techniques are very challenging to design due to a combination of NAND Flash constraints (for example the block-erase-before-page-rewrite constraint and limited number of erase cycles) and embedded system constraints (for example tiny RAM and resource consumption predictability). Previous work adapted traditional indexing methods to cope with Flash constraints by deferring index updates using a log and batching them to decrease the number of rewrite operations in Flash memory. However, these methods were not designed with embedded system constraints in mind and do not address them. In this paper, we propose a new alternative for indexing Flash-resident data that specifically addresses the embedded context. This approach, called PBFilter, organizes the index structure in a purely sequential way. Key lookups are sped up thanks to two principles called Summarization and Partitioning. We instantiate these principles with data structures and algorithms based on Bloom Filters and show the effectiveness of this approach through a comprehensive performance study.
Shaoyi Yin, Philippe Pucheral, Xiaofeng Meng 0001
EDBT1
2008 Restoring the Patient Control over Her Medical History
abstract
Paper-based folders have been widely used to coordinate cares in medical-social networks, but they introduce some burning issues (e.g. privacy protection, remote access to the folder). Replacing the paper-based folder system by a traditional Electronic Healthcare Record (EHR) introduces new drawbacks: forcing of the patient consent, unbounded data retention, no security guarantee outside the server domain and no disconnected access to the folder. To solve these problems, this paper proposes an experimental platform which combines an EHR system with medical-social folders embedded in a new hardware portable device. The objectives pursued are (1) to re-establish a natural and powerful way of protecting and sharing highly sensitive information among trusted parties and (2) to build a shared medical-social folder providing the highest degree of availability, whatever the mode of operation (disconnected or not).
Nicolas Anciaux, Mehdi Benzine, Luc Bouganim, Kévin Jacquemin, Philippe Pucheral, Shaoyi Yin
CBMS6
2008 PBFilter: indexing flash-resident data through partitioned summaries
abstract
NAND Flash has become the most popular persistent data storage medium for mobile and embedded devices. The hardware characteristics of NAND Flash (e.g. page granularity for read/write with a block-erase-before-rewrite constraint, limited number of erase cycles) preclude in-place updates. In this paper, we propose a new indexing scheme, called PBFilter, designed from the outset to exploit the peculiarities of NAND Flash.
Shaoyi Yin, Philippe Pucheral, Xiaofeng Meng 0001
CIKM1