EDBT 2026 Demo / reviewers in the wild / expert
Omar Boussaïd
dblp:63/3490 · also Omar Boussaid
· DBLP profile ↗
53ranked-venue papers
2as first author
4since 2021 · last 2023
0000-0001-6388-3152ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 34 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 20 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 12Systems, architecture and hardware · 2 · 1 since 2021Security and privacy · 2Human-computer interaction and ubiquitous computing · 2Theory of computation · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Temporal Multidimensional Model for Evolving Graph-Based Data WarehousesabstractInternational audience Redha Benhissen, Fadila Bentayeb, Omar Boussaïd |
DATA | 3 |
| 2023 | GAMM: Graph-Based Agile Multidimensional Model
Redha Benhissen, Fadila Bentayeb, Omar Boussaïd |
DOLAP | 3 |
| 2022 | Building a novel physical design of a distributed big data warehouse over a Hadoop cluster to enhance OLAP cube query performance
Yassine Ramdane, Omar Boussaïd, Doulkifli Boukraâ, Nadia Kabachi, Fadila Bentayeb |
Parallel Comput. | 2 |
| 2021 | Medical-Based Text Classification Using FastText Features and CNN-LSTM Model
Mohamed Walid Zeghdaoui, Omar Boussaïd, Fadila Bentayeb, Frederik Joly |
DEXA (1) | 2 |
| 2020 | Tag's Depth-Based Expert Profiling Using a Topic Modeling TechniqueabstractExpert finding and expert profiling are two important tasks for organizations, researchers, and work seekers. This importance can also be seen in online communities especially with the explosion of social networks. Expert finding on one hand addresses the task of finding the right person with the appropriate knowledge or skills. Expert profiling on the other hand gives a concise and meaningful description of a candidate expert. This paper focuses on what social tagging can bring to improve expert finding and profiling. A novel expertise indicator that models and assesses an expert based on the expert's tagging activities is proposed. First, tags are used as interest indicator to build candidate's profiles; then, Latent Dirichlet Allocation algorithm (LDA) is used to construct the tags distribution over topics by exploiting the tag's semantic characteristics. Topics of interest are then filtered using tag's depth. The latter is finally used as the expertise indicator. Experiments performed on the stack overflow dataset show the accuracy of the proposed approach. Saida Kichou, Omar Boussaïd, Abdelkrim Meziane |
Int. J. Semantic Web Inf. Syst. | 2 |
| 2019 | SDWP: A New Data Placement Strategy for Distributed Big Data Warehouses in Hadoop
Yassine Ramdane, Nadia Kabachi, Omar Boussaïd, Fadila Bentayeb |
DaWaK | 3 |
| 2019 | SkipSJoin: A New Physical Design for Distributed Big Data Warehouses in Hadoop
Yassine Ramdane, Nadia Kabachi, Omar Boussaïd, Fadila Bentayeb |
ER | 3 |
| 2019 | AuMixDw: Towards an automated hybrid approach for building XML data warehouses
Zoubir Ouaret, Doulkifli Boukraâ, Omar Boussaïd, Rachid Chalal |
Data Knowl. Eng. | 3 |
| 2018 | Partitioning and Bucketing Techniques to Speed up Query Processing in Spark-SQLabstractHorizontal partitioning is an optimization technique applied to improve query processing time in distributed data warehouses. Scanning a large number of HDFS data blocks, to respond to the ad-hoc or OLAP queries, is a heavy operation. We can skip loading unnecessary data blocks if we partition or index some tables by the appropriate predicate attributes. However, the way of selecting the candidate's attributes remains a challenging task to handle. In this paper, we propose a technique based on frequent itemset mining, to Partition, Bucket and Sort the Tables (PBSTs) of a big data warehouse with the more frequent predicate attributes in the queries. We take into account the density of the attributes of the tables, data skew, and the physical characteristics of the cluster nodes. To evaluate our approach, we conducted some experiments in a cluster of 15 slave nodes. Experimental results show that with our method, we improve the query response time by 50 % over existing skipping techniques. Yassine Ramdane, Omar Boussaïd, Nadia Kabachi, Fadila Bentayeb |
ICPADS | 2 |
| 2017 | S2D: Shared Distributed Datasets, Storing Shared Data for Multiple and Massive Queries Optimization in a Distributed Data Warehouse
Rado Ratsimbazafy, Omar Boussaïd, Fadila Bentayeb |
DaWaK | 2 |
| 2017 | Logical Schema for Data Warehouse on Column-Oriented NoSQL Databases
Mohamed Boussahoua, Omar Boussaïd, Fadila Bentayeb |
DEXA (2) | 2 |
| 2017 | Minimizing Negative Influence in Social Networks: A Graph OLAP Based Approach
Zakia Challal, Omar Boussaïd, Kamel Boukhalfa |
DEXA (2) | 2 |
| 2017 | A Fine‐Grained Distribution Approach for ETL Processes in Big Data Environments
Mahfoud Bala, Omar Boussaïd, Zaia Alimazighi |
Data Knowl. Eng. | 2 |
| 2016 | Handicraft Women Recommendation Approach Based on User's Social Tagging OperationsabstractRecommendation predicts which items the user might be interested in, and aims to help users finding the adequate element. Such as movies, music and commercial products, persons may be also recommended. In the case of handicraft women, we propose a recommendation approach based on extracted user's interest using his/her social tagging operations to improve business activities of the handicrafts women, the approach is applied with preliminary tests. Saida Kichou, Hakima Mellah, Omar Boussaïd, Abdelkrim Meziane |
WI | 3 |
| 2016 | An MDA approach to secure access to data on cloud using implicit securityabstractCloud computing has been developed to deliver information technologies services on demand for organisations or as individual users. In this paper, we describe an approach which proposes a security model bases on organisation role-based access control (ORBAC) and encryption system. This model aims to give users a possibility to control security of their data. In this scheme, a secret key is partitioned using a Galois field GF(22). Our proposal has been aligned with model driven architecture (MDA). Yasmina Ghebghoub, Saliha Oukid, Omar Boussaïd |
Int. J. Inf. Comput. Secur. | 3 |
| 2015 | A Data Mining-based Blocks Placement Optimization for Distributed Data WarehousesabstractThe amount of data that is captured and generated by modern computing devices has augmented exponentially over the last years. The Hadoop framework - an open source project based on the MapReduce paradigm - is a popular choice for processing these large volumes of data or big data. However, the performance gained from Hadoop's features is currently limited by its default block placement policy, which does not take any data characteristics into account. This is particularly true for relational data bases and data warehouses. Indeed, the efficiency of many operations can be improved by a careful data placement, including indexing, grouping, aggregation and joins. In this paper we propose a data warehouse distribution strategy to improve query gain performances on multi-nodes clusters, especially Hadoop clusters. Based on k-means clustering method that allows to master the number of clusters through its k parameter, we investigate the performance gain for OLAP cube construction with and without data organization. And this, by varying the number of clusters and data warehouse size. Our experiments suggest that a good data placement on a cluster during the implementation of the data warehouse increase significantly the OLAP cube construction and querying performances. Billel Arres, Nadia Kabachi, Omar Boussaïd, Fadila Bentayeb |
IDEAS | 3 |
| 2015 | Optimizing OLAP Cubes Construction by Improving Data Placement on Multi-nodes ClustersabstractThe increasing volumes of relational data let us find an alternative to cope with them. The Hadoop framework - which is an open source project based on the MapReduce paradigm - is a popular choice for big data analytics. However, the performance gained from Hadoop's features is currently limited by its default block placement policy, which does not take any data characteristics into account. Indeed, the efficiency of many operations can be improved by a careful data placement, including indexing, grouping, aggregation and joins. In this paper we propose a data warehouse placement policy to improve query gain performances on multi nodes clusters, especially Hadoop clusters. We investigate the performance gain for OLAP cube construction query with and without data organization. And this, by varying the number of nodes and data warehouse size. It has been found that, the proposed data placement policy has lowered global execution time for building OLAP data cubes up to 20 percent compared to default data placement. Billel Arres, Nadia Kabachi, Omar Boussaïd |
PDP | 3 |
| 2015 | Intentional Data Placement Optimization for Distributed Data WarehousesabstractParallel computing is a fundamental technique in the management of large quantities of data as it leverages on the concurrent utilization of multiple computing resources. One of the technologies that made big data analytics popular and accessible to enterprises of all sizes is MapReduce (and its open-source Hadoop implementation). With the ability to automatically parallelize the application on a cluster of commodity hardware, MapReduce allows enterprises to analyze terabytes and petabytes of data more conveniently than ever. However, the performance gained from Hadoop's features is currently limited by its default block placement policy, which does not take any data characteristics into account. Indeed, the efficiency of many operations can be improved by a careful data placement, including indexing, grouping, aggregation and joins. In this paper, we present a MapReduce data blocks allocation approach to improve MapReduce jobs execution and query performances on multi-nodes clusters, especially Hadoop clusters. Based on k-means clustering method that allows to master the number of clusters through its k parameter, we study the influence of number of clusters on queries execution instead of queries performances with and without data organization. For this, we used well-known, large-scale data analysis benchmark: TPC-H. Our experiments suggest that defining a good data placement on a cluster during the implementation of a data warehouse increase significantly the OLAP cube construction and querying performances. Billel Arres, Nadia Kabachi, Omar Boussaïd, Fadila Bentayeb |
SMC | 3 |
| 2014 | P-TRIAR: Personalization Based on TRIadic Association Rules
Sid-Ali Selmane, Omar Boussaïd, Fadila Bentayeb |
ADBIS | 2 |
| 2014 | An original approach for processing public open data with MapReduce: A case studyabstractNowadays, many governments and states are involved in an opening strategy of their public data. However, the volume of these opened data is constantly increasing, and will reach in the near future limitations of current treatment and storage capacity. On the other hand, the MapReduce paradigm is one of the most used parallel programming models. With a master-slave architecture, it allows parallel processing of very large data sets. In this paper, we propose a parallel approach based on Mapreduce to process public open data. Applied, as a case study, to the official data sets from the French Ministry of Communication. We implement a parallel algorithm as a solution to define a ranking of national museums and galleries according to the accessibility degrees for people with disabilities. We studied the feasibility of our approach in two main parts: The performance in terms of execution time, and, the visualization of the obtained results in order to integrate them into solutions such as geographic BI. This work can be applied to other cases with very large data sets. Billel Arres, Nadia Kabachi, Fadila Bentayeb, Omar Boussaïd |
AICCSA | 4 |
| 2014 | P-ETL: Parallel-ETL based on the MapReduce paradigmabstractBig data is an opportunity in the emergence of novel business applications such as “Big Data Analytics” (BDA). However, these data with non-traditional volumes create a real problem given the capacity constraints of traditional systems. The aim of this paper is to deal with the impact of big data in a decision-support environment and more particularly in the data integration phase. In this context, we developed a platform, called P-ETL (Parallel-ETL) for extracting (E), transforming (T) and loading (L) very large data in a data warehouse (DW). To cope with very large data, ETL processes under our P-ETL platform run on a cluster of computers in parallel way with MapReduce paradigm. The conducted experiment shows mainly that increasing tasks dealing with large data speeds-up the ETL process. Mahfoud Bala, Omar Boussaïd, Zaia Alimazighi |
AICCSA | 2 |
| 2014 | A global and comprehensive approach for XML data warehouse designabstractThe increasing amounts of interesting data stored in the XML format is the most challenging issue for BI community, thus it is desirable to successfully extract, store and integrate this large sources of information special purpose systems called “data warehouse” for further analysis and decision-making. However, compared with the well structured relational databases of a company, XML data presents a complex hierarchical structure, which renders inappropriate, existing traditional data warehouse approaches and techniques. In this paper, we propose a semi-automatic approach for XML data warehouse design starting from XML schemas as data sources. The first step consists in automatically generating the UML Class diagram from W3C XML Schema (XSD). However, the obtained diagram can be very large and hard to understand. To overcome this situation, we use a set of rules based on basic techniques for object oriented design quality to develop a simplification algorithm that efficiently generates high-quality diagrams with limited number of classes. Then, we propose a multi-dimensional (MD) element extraction algorithm to automatically identify facts, measures and their corresponding dimensions. We also present a new metric for ranking obtained MD schemas according to their relevance. The final step consists in automatically generating the star XML schema that corresponds to the XML Data warehouse schema. Finally, we have implemented our approach using JAVA and we have evaluated this tool on several XML schemas. Zoubir Ouaret, Omar Boussaïd, Rachid Chalal |
AICCSA | 2 |
| 2014 | Towards an OLAP Environment for Column-Oriented Data Warehouses
Khaled Dehdouh, Fadila Bentayeb, Omar Boussaïd, Nadia Kabachi |
DaWaK | 3 |
| 2014 | An Efficient Method for Community Detection Based on Formal Concept Analysis
Sid-Ali Selmane, Fadila Bentayeb, Rokia Missaoui, Omar Boussaïd |
ISMIS | 4 |
| 2014 | The Multidimensional Semantic Model of Text Objects(MSMTO): A Framework for Text Data Analysis
Sarah Attaf, Nadjia Benblidia, Omar Boussaïd |
MEDI | 3 |
| 2014 | Columnar NoSQL Star Schema Benchmark
Khaled Dehdouh, Omar Boussaïd, Fadila Bentayeb |
MEDI | 2 |
| 2014 | Columnar NoSQL CUBE: Agregation operator for columnar NoSQL data warehouseabstractThe emergence of large volumes of data imposed by the major players of the web requires new management models and new data storage architectures and treatment able to find information quickly in a large volume of data. The column-oriented NoSQL (Not Only SQL) database provide for big data the most suitable model to the data warehouse and the structure of multidimensional data in OLAP cube form. However, in the absence of OLAP cube computation operators, we propose in this paper, a new aggregation operator called CN-CUBE (Columnar NoSQL CUBE), which allows data cubes to be computed from data warehouses stored in column-oriented NoSQL database management system. We implemented the CNCUBE operator using the SQL Phoenix interface of HBase DBMS and conducted experiments on a public data warehouse in a distributed environment produced using the Hadoop platform. Thus we have shown that our CN-CUBE operator has OLAP cubes computation times very suitable for NoSQL warehouses. Khaled Dehdouh, Fadila Bentayeb, Omar Boussaïd, Nadia Kabachi |
SMC | 3 |
| 2014 | Complex Object-Based Multidimensional Modeling and Cube ConstructionabstractThis paper presents a multidimensional model and a language to construct cubes for the purpose of on-line analytical processing. Both the multidimensional model and the cube model are based on the concept of complex object which models complex entiti Doulkifli Boukraâ, Omar Boussaïd, Fadila Bentayeb |
Fundam. Informaticae | 2 |
| 2013 | Building OLAP cubes on a Cloud Computing environment with MapReduceabstractLarge-scale data analysis has become increasingly important for many enterprises, and Cloud Computing, under the impulse of large companies, has recently endowed a special attention both in industry and academic researches. Hadoop, based on a new distributed computing paradigm, called MapReduce, has allowed to facilitate access to such environments, due to its impressive scalability and flexibility to handle structured as well as unstructured data. The goal of our work is to develop a Cloud Computing environment for exploiting data warehouses and perform online analysis. It consists of handling large nonrelational databases and supporting data warehouse with a new generation of database management systems (DBMS) such as Hive. Thus, to set up such an environment, we implemented a data warehouse under Hadoop and Hive and we used the Map and Reduce functions of this environment, then we compared the cost of loading the warehoused data and constructing OLAP cubes between a virtual and a physical cluster, as well as the rise in data loading on a physical cluster. Obtained results allows MapReduce developers to fully compare the performance, help in the choice of platform, in which a customer application can be developed to translate SQL requests to HQL (Hive-QL) requests, and check if a not-relational model is adequate or not. Billel Arres, Nadia Kabachi, Omar Boussaïd |
AICCSA | 3 |
| 2013 | Semantic metadata mediation: XML, RDF and RuleMLabstractThis work is situated in the general context of stored information heterogeneity in a decisional system such as data, metadata and knowledge, for cohabitation and reconciliation of these information by mediation. In this paper we focus on the heterogeneous metadata integration, with the definition of a structural and semantic mediation model. Our aim is to propose a mediation architecture for the heterogeneous sources metadata, represented by XML, RDF and RuleML model, providing to user the metadata transparency. This, by including data structures, of natures fundamentally different, and allowing the decomposition of a query involving multiple sources, to specific queries to these sources, then recompose the result. We use ontology for managing structural and semantic heterogeneity. Messaouda Fareh, Omar Boussaïd, Rachid Chalal |
AICCSA | 2 |
| 2013 | A Layered Multidimensional Model of Complex Objects
Doulkifli Boukraâ, Omar Boussaïd, Fadila Bentayeb, Djamel Eddine Zegour |
CAiSE | 2 |
| 2013 | Social microblogging cubeabstractMicroblogging sites have become a staple in our modern world. They provide the users with the ability to keep in touch with their contacts, using up of 140 characters in the case of Twitter sites. Responding to this emerging trend, it becomes critically important to interactively view and analyze the massive amount of microblogging data from different perspectives and with multiple granularities. In the area of Business intelligence, On-line analytical processing (OLAP) is a powerful primitive for data analysis. However, OLAP tools face major challenges in manipulating unstructured text such as microblogging data. Lilia Hannachi, Nadjia Benblidia, Fadila Bentayeb, Omar Boussaïd |
DOLAP | 4 |
| 2013 | CXT-cube: contextual text cube model and aggregation operator for text OLAPabstractTraditional data warehousing technologies and On-Line Analytical Processing (OLAP) are unable to analyze textual data. Moreover, as OLAP queries of a decision-maker are generally related to a context, contextual information must be taken into account during the exploitation of data warehouses. Thus, we propose a contextual text cube model denoted CXT-Cube which considers several contextual factors during the OLAP analysis in order to better consider the contextual information associated with textual data. CXT-Cube is characterized by several contextual dimensions, each one related to a contextual factor. In addition, we extend our aggregation OLAP operator for textual data ORank (OLAP-Rank) to consider all the contextual factors defined in our CXT-Cube model. To validate our model, we perform an experimental study and the preliminary results show the importance of our approach for integrating textual data into a data warehouse and improving the decision-making. Lamia Oukid, Ounas Asfari, Fadila Bentayeb, Nadjia Benblidia, Omar Boussaïd |
DOLAP | 5 |
| 2013 | Semi-structured Documents Mining: A Review and ComparisonabstractThe number of semi-structured documents that is produced is steadily increasing. Thus, it will be essential for discovering new knowledge from them. In this survey paper, we review popular semi-structured documents mining approaches (structure alone and both structure and content). We provide a brief description of each technique as well as efficient algorithms for implementing the technique and comparing them using different comparison criteria. Amina Madani, Omar Boussaïd, Djamel Eddine Zegour |
KES | 2 |
| 2012 | Community Extraction Based on Topic-Driven-Model for Clustering Users Tweets
Lilia Hannachi, Ounas Asfari, Nadjia Benblidia, Fadila Bentayeb, Nadia Kabachi, Omar Boussaïd |
ADMA | 6 |
| 2012 | Managing a fragmented XML data cube with oracle and timestenabstractIn this paper, we cross two techniques for performance tuning of an XML cube. We analyze six configurations for managing the cube. The configurations result from storing two variants of the cube (unfragmented and fragmented) in different ways. First, we consider a disk-resident database. Then, we consider caching the frequent properties of the unfragmented cube and the frequent fragments of the fragmented cube. Finally, we load and manage the entire cube into the main memory. We show the benefits of vertical fragmentation and in-memory management of the XML cube through a set of experiments. Doulkifli Boukraâ, Omar Boussaïd, Fadila Bentayeb, Djamel Eddine Zegour |
DOLAP | 2 |
| 2012 | Verification of Security Coherence in Data Warehouse Designs
Ali Salem, Salah Triki, Hanêne Ben-Abdallah, Nouria Harbi, Omar Boussaïd |
TrustBus | 5 |
| 2011 | Vertical Fragmentation of XML Data Warehouses Using Frequent Path Sets
Doulkifli Boukraâ, Omar Boussaïd, Fadila Bentayeb |
DaWaK | 2 |
| 2011 | Efficient incremental breadth-depth XML event miningabstractMany applications log a large amount of events continuously. Extracting interesting knowledge from logged events is an emerging active research area in data mining. In this context, we propose an approach for mining frequent events and association rules from logged events in XML format. This approach is composed of two-main phases: I) constructing a novel tree structure called Frequency XML-based Tree (FXT), which contains the frequency of events to be mined; II) querying the constructed FXT using XQuery to discover frequent itemsets and association rules. The FXT is constructed with a single-pass over logged data. We implement the proposed algorithm and study various performance issues. The performance study shows that the algorithm is efficient, for both constructing the FXT and discovering association rules. Rashed K. Salem, Jérôme Darmont, Omar Boussaïd |
IDEAS | 3 |
| 2011 | Securing Data Warehouses: A Semi-automatic Approach for Inference Prevention at the Design Level
Salah Triki, Hanêne Ben-Abdallah, Nouria Harbi, Omar Boussaïd |
MEDI | 4 |
| 2010 | OLAP Operators for Complex Object Data Cubes
Doulkifli Boukraâ, Omar Boussaïd, Fadila Bentayeb |
ADBIS | 2 |
| 2010 | UNICOMP: Identification of Enterprise Competencies to Build Collaborative Networks
Kafil Hajlaoui, Xavier Boucher, Omar Boussaïd |
PRO-VE | 3 |
| 2007 | Evolution of Data Warehouses' Optimization: A Workload Perspective
Cécile Favre, Fadila Bentayeb, Omar Boussaïd |
DaWaK | 3 |
| 2007 | Ontology-Based Object Recognition for Remote Sensing Image InterpretationabstractThe multiplication of Very High Resolution (spatial or spectral) remote sensing images appears to be an opportu- nity to identify objects in urban and periurban areas. The classification methods applied in the object-oriented image analysis approach could be based on the use of domain knowledge. A major issue in these approaches is domain knowledge formalization and exploitation. In this paper, we propose a recognition method based on an ontology which has been developed by experts of the domain. In order to give objects a semantic meaning, we have developed a matching process between an object and the concepts of the ontology. Experiments are made on a Quickbird image. The quality of the results shows the effectiveness of the proposed method. Nicolas Durand 0001, Sébastien Derivaux, Germain Forestier, Cédric Wemmert, Pierre Gançarski, Omar Boussaïd, Anne Puissant |
ICTAI (1) | 6 |
| 2007 | Integration and dimensional modeling approaches for complex data warehousing
Omar Boussaïd, Adrian Tanasescu, Fadila Bentayeb, Jérôme Darmont |
J. Glob. Optim. | 1 |
| 2006 | X-Warehousing: An XML-Based Approach for Warehousing Complex Data
Omar Boussaïd, Riadh Ben Messaoud, Rémy Choquet, Stéphane Anthoard |
ADBIS | 1 |
| 2006 | Enhanced mining of association rules from data cubesabstractOn-line analytical processing (OLAP) provides tools to explore and navigate into data cubes in order to extract interesting information. Nevertheless, OLAP is not capable of explaining relationships that could exist in a data cube. Association rules are one kind of data mining techniques which finds associations among data. In this paper, we propose a framework for mining inter-dimensional association rules from data cubes according to a sum-based aggregate measure more general than simple frequencies provided by the traditional COUNT measure. Our mining process is guided by a meta-rule context driven by analysis objectives and exploits aggregate measures to revisit the definition of support and confidence. We also evaluate the interestingness of mined association rules according to Lift and Loevinger criteria and propose an efficient algorithm for mining inter-dimensional association rules directly from a multidimensional data. Riadh Ben Messaoud, Sabine Loudcher, Omar Boussaïd, Rokia Missaoui |
DOLAP | 3 |
| 2006 | Efficient multidimensional data representations based on multiple correspondence analysisabstractIn the On Line Analytical Processing (OLAP) context, exploration of huge and sparse data cubes is a tedious task which does not always lead to efficient results. In this paper, we couple OLAP with the Multiple Correspondence Analysis (MCA) in order to enhance visual representations of data cubes and thus, facilitate their interpretations and analysis. We also provide a quality criterion to measure the relevance of obtained representations. The criterion is based on a geometric neighborhood concept and a similarity metric between cells of a data cube. Experimental results on real data proved the interest and the efficiency of our approach. Riadh Ben Messaoud, Omar Boussaïd, Sabine Loudcher |
KDD | 2 |
| 2005 | Preparing complex data for warehousingabstractSummary form only given. In order to prepare complex data for relevant analysis, a data warehousing-based approach is needed. However, a good multidimensional modeling requires efficient preparation of data starting with a data integration phase. We present in this paper two principal steps of the complex data warehousing process. The data integration is the first one. To do that, we define a generic UML data model capable of representing a wide range of complex data including their possible semantic properties. Furthermore, complex data are represented as XML documents generated through an implemented prototype. The second important phase is the preparation of data for the multidimensional modeling. We demonstrate that we can use data mining techniques to help the user in building a better multidimensional model. Adrian Tanasescu, Omar Boussaïd, Fadila Bentayeb |
AICCSA | 2 |
| 2005 | Evaluation of a MCA-based approach to organize data cubesabstractIn the OLAP context, exploration of huge and sparse data cubes is a tedious task that does not always lead to efficient results. We propose to use a Multiple Correspondence Analysis (MCA) in order to enhance data cube representations and make them more suitable for visualization and thus, easier to analyze. We also provide an original quality criterion to measure the relevance of the obtained data representations. Experimental results we led on real data samples have shown the interest and the efficiency of our approach. Riadh Ben Messaoud, Omar Boussaïd, Sabine Loudcher |
CIKM | 2 |
| 2005 | Automatic Selection of Bitmap Join Indexes in Data Warehouses
Kamel Aouiche, Jérôme Darmont, Omar Boussaïd, Fadila Bentayeb |
DaWaK | 3 |
| 2005 | DWEB: A Data Warehouse Engineering Benchmark
Jérôme Darmont, Omar Boussaïd, Fadila Bentayeb |
DaWaK | 2 |
| 2004 | A new OLAP aggregation based on the AHC techniqueabstractNowadays, decision support systems are evolving in order to handle complex data. Some recent works have shown the interest of combining on-line analysis processing (OLAP) and data mining. We think that coupling OLAP and data mining would provide excellent solutions to treat complex data. To do that, we propose an enhanced OLAP operator based on the agglomerative hierarchical clustering (AHC). The here proposed operator, called OpAC (Operator for Aggregation by Clustering) is able to provide significant aggregates of facts refereed to complex objects. We complete this operator with a tool allowing the user to evaluate the best partition from the AHC results corresponding to the most interesting aggregates of facts. Riadh Ben Messaoud, Omar Boussaïd, Sabine Loudcher |
DOLAP | 2 |