VLDB 2026 Research / reviewers in the wild / expert
Raja Chiky
dblp:48/2215
· DBLP profile ↗
28ranked-venue papers
2as first author
6since 2021 · last 2024
0000-0001-8346-318XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 13 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 since 2021Systems, architecture and hardware · 1Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Online log parsing using evolving research tree
Arthur Vervaet, Mar Callau-Zori, Yousra Chabchoub, Raja Chiky |
Knowl. Inf. Syst. | 4 |
| 2023 | StreamMLOps: Operationalizing Online Learning for Big Data Streaming & Real-Time ApplicationsabstractContinuously learning and serving from evolving streaming data and serving in real-time is a challenging problem. Traditionally, data is partitioned and processed in batches to train machine learning (ML) models. In industrial applications, static models’ performance drops over time (model degradation, concept drift), requiring new models to be trained with recent data and redeployed in production. The scientific community has been studying online and adaptive methods to address batch-learning limitations and continuously train AI tasks for industrial applications such as cyber-security, AIOps, anomaly scoring, and drift detection in stock markets. This paper deals with the MLOps aspects of deploying such online and dynamic models to address the requirements in the production systems for real-time applications. Our architectures - based on open-source tools such as Kafka and River - demonstrated how online learning methods could be scaled horizontally in production to meet the demands of a high-velocity streaming pipeline. We demonstrate an MLOps strategy to perform incremental learning from streaming data and continuously deploy the online learning model without pausing the inference pipeline. Indeed, the design satisfies requirements such as model versioning, monitoring, audibility and reproducibility of prediction in both a supervised and semi-supervised setting. Our experiments - for malicious URLs detection task - performed on high-dimensional and feature-evolving streaming data (more than 3 million features) establish the effectiveness and efficiency of online learning models compared to batch (static) machine learning regarding both time and space complexity. Finally, we provide some best practices on data engineering for deploying online models to process a real-time feature stream in production environments. Code is publicly available for reproducibility. Mariam Barry, Jacob Montiel, Albert Bifet, Sameer Wadkar, Nikolay Manchev, Max Halford, Raja Chiky, Saad El Jaouhari, Katherine B. Shakman, Joudi Al Fehaily, Fabrice Le Deit, Vinh-Thuy Tran, Eric Guerizec |
ICDE | 7 |
| 2022 | SDG-Meter: A Deep Learning Based Tool for Automatic Text Classification of the Sustainable Development Goals
Jade Eva Guisiano, Raja Chiky, Jonathas De Mello |
ACIIDS (1) | 2 |
| 2022 | Stream2Graph: Dynamic Knowledge Graph for Online Learning Applied in Large-scale NetworkabstractKnowledge Graphs (KG) are valuable information sources that store knowledge in a domain (healthcare, finance, e-commerce, cyber-security.). Most industrial KGs are dynamic by nature as they are updated regularly with streaming data (customer activity, network traffic, application logs, IT process). However, extracting insights from continuously updated data comes with major challenges, particularly in big data settings. In this paper, we address the following challenges: 1) ingesting heterogeneous data, 2) training and deployment of predictive models on continuously evolving data, and 3) implementation of data pipelines for updating and maintaining the KG in production. We cover multiple aspects of this process, from knowledge collection to its operationalization. We propose Stream2Graph, a stream-based system for building and updating the knowledge base dynamically in real time. Then we show how graph features can be used in downstream online machine learning models. The solution speeds up big data stream learning and knowledge extraction to enhance Graph-based AI applications. Experimental results show the effectiveness of our solution for knowledge base construction and improvement of big data learning capabilities. Using data from Stream2Graph resulted in speedups for training and inference time in the range from 547x to 2000x in downstream ML models. Finally, we provide the lessons learned from applying graph-based online learning on large-scale network processing high-velocity streaming data. Mariam Barry, Albert Bifet, Raja Chiky, Saad El Jaouhari, Jacob Montiel, Aissa El Ouafi, Eric Guerizec |
IEEE Big Data | 3 |
| 2022 | StreamFlow: A System for Summarizing and Learning Over Industrial Big Data StreamsabstractThe growing need for predictive analytics over streaming data in the industry requires a flexible and continuously scalable big data system. In real-time big data applications (cybersecurity, AIOps, anomaly detection, predictive maintenance, IoT etc.), efficient machine learning models must be trained and industrialized within existing data processing plat-forms and industrial tools. This requires interoperability between various components: data collection, processing, summarization, modelling and analytics. Existing works focus on building AI models for big data, neglecting real-world challenges when integrating such models into an existing industrial production framework. In this paper, we propose StreamFlow, an operational data pipeline to address industrial challenges for continuous learning over big data streams. We also propose an online method using sliding windows to summarize high-velocity data. The final result of the framework is a feature vector that describes the underlying processes and is ready to use in machine learning tasks. Moreover, we showcase real-world applications such as automated feature engineering for real-time monitoring and online machine learning for event classification. The proposed system has been deployed within production in a banking system, processing billions of daily traffic operations. Our experiments demonstrate the effectiveness and performance of our approach by evaluating it at different levels: processing, summarization, improvement of machine learning performance and effectiveness in an industrial setting. In the case of downstream machine learning tasks, using summarized data generated by StreamFlow results in up to 2 orders of magnitude speedups in training time without compromising predictive performance. Mariam Barry, Saad El Jaouhari, Albert Bifet, Jacob Montiel, Eric Guerizec, Raja Chiky |
IEEE Big Data | 6 |
| 2021 | USTEP: Unfixed Search Tree for Efficient Log ParsingabstractLogs record valuable system information at runtime. They are widely used by data-driven approaches for development and monitoring purposes. Parsing log messages to structure their format is a classic preliminary step for log-mining tasks. As they appear upstream, parsing operations can become a processing time bottleneck for downstream applications. The quality of parsing also has a direct influence on their efficiency. Previous approaches toward online log parsing focused on stateful methods. But an increasing number of tasks ask for real time monitoring. Regarding this problem, we propose USTEP, an online log parsing method based on an evolving tree structure. Evaluation results on a panel of 13 datasets coming from different real-world systems demonstrate USTEP superiority in terms of both effectiveness and robustness when compared to other online methods. We also introduce USTEP-UP, a way of running multiple decentralized instances of USTEP in parallel. Arthur Vervaet, Raja Chiky, Mar Callau-Zori |
ICDM | 2 |
| 2020 | Anomaly Detection for Data Streams Based on Isolation Forest Using Scikit-Multiflow
Maurras Togbe, Mariam Barry, Aliou Boly, Yousra Chabchoub, Raja Chiky, Jacob Montiel, Vinh-Thuy Tran |
ICCSA (4) | 5 |
| 2020 | Movies Emotional Analysis Using Textual Contents
Amir Kazem Kayhani, Farid Meziane, Raja Chiky |
NLDB | 3 |
| 2019 | Workload Characterization for a Non-Hyperscale Public Cloud PlatformabstractThe improvement of automated resource management techniques for cloud computing platforms requires a deep understanding of the workload. Previous works focused on virtual machines (VMs), and neglected complementary virtual resources such as images, volumes, snapshots and security groups. Besides, most attention went to public hyperscale platforms with more than ten thousand servers, or small on-premise platforms. To fill the gap, we perform a holistic workload characterization of a non-hyperscale platform. We have collected a three-month-long trace allowing us to characterize the correlated utilization of virtual resources; the consumption of CPU, memory and disk by VMs; and the CPU interferences between VMs. Loïc Pérennou, Mar Callau-Zori, Sylvain Lefebvre 0002, Raja Chiky |
CLOUD | 4 |
| 2019 | Applying Supervised Machine Learning to Predict Virtual Machine Runtime for a Non-hyperscale Cloud Provider
Loïc Pérennou, Raja Chiky |
ICCCI (2) | 2 |
| 2019 | A Distributed Pollution Monitoring System: The Application of Blockchain to Air Quality Monitoring
Cameron Thouati de Tazoult, Raja Chiky, Valentin Foltescu |
ICCCI (2) | 2 |
| 2019 | Improving the Attribute-Based Active Learning by Clustering the New ItemsabstractThe issue that recommender system often meets is cold-start problem, where the system does not have any ratings of new items or new users. Thus, it can not provide relevant recommendation for the new users or new items. In previous research, when dealing with item cold-start problem, some scientists combined content information and active learning method, and used factorization machine to model the prediction task. However, a shortcoming in this method is that when using factorization machine model to select users to give ratings to new items, the active users may be selected for too many times, leading to a result that they refuse to give ratings for new items, or randomly give their ratings, which does not exactly show their preferences. In this paper, to solve this issue, we use clustering algorithm to divide new items into different groups and choose one item to represent the group, and only request users giving ratings for the representative items. Junxin Zhou, Raja Chiky |
SERVICES | 2 |
| 2018 | An In-depth Analysis of CUSUM Algorithm for the Detection of Mean and Variability Deviation in Time Series
Rayane El Sibai, Yousra Chabchoub, Raja Chiky, Jacques Demerjian, Kablan Barbar |
W2GIS | 3 |
| 2017 | Enhancing New User Cold-Start Based on Decision Trees Active Learning by Using Past Warm-Users Predictions
Manuel Pozo, Raja Chiky, Farid Meziane, Elisabeth Métais |
ICCCI (1) | 2 |
| 2017 | Assessing and Improving Sensors Data Quality in Streaming Context
Rayane El Sibai, Yousra Chabchoub, Raja Chiky, Jacques Demerjian, Kablan Barbar |
ICCCI (2) | 3 |
| 2017 | Evaluating Non-personalized Single-Heuristic Active Learning Strategies for Collaborative Filtering Recommender SystemsabstractIn collaborative filtering recommender systems, the users rate items, and this process helps in understanding their preferences. The systems can suffer from the cold-start problem, which refers to the absence or insufficiency of ratings for new users. This can be solved by using active learning strategies, which can be non-personalized or personalized, and which were evaluated and tested previously using different datasets and metrics. In this paper, we present a clearer study by implementing the main non-personalized single-heuristic strategies (random, popularity, co—coverage, variance, entropy, entropy0) on the same dataset, and by evaluating them using the same metrics, in order to have a better comparison. We use the public MovieLens dataset in the experimentations and the results show that the random strategy performs the worst, whereas the entropy0 leads to the best results. All strategies except the random strategy lead to very close results at a certain point, where ratings for almost the same items will have been elicited. Georges Chaaya, Elisabeth Métais, Jacques Bou Abdo, Raja Chiky, Jacques Demerjian, Kablan Barbar |
ICMLA | 4 |
| 2016 | FreGraPaD: Frequent RDF graph patterns detection for semantic data streamsabstractNowadays, high volumes of data are generated and published at a very high velocity by real-time systems, such as social networks, e-commerce, weather stations and sensors, producing heterogeneous data streams. To take advantage of linked data and offer interoperable solutions, semantic Web technologies have been used. To analyze these huge volumes of data, different stream mining algorithms exist such as compression or load-shedding. Nevertheless, most of them need many passes through the data and often store part of it on disk. If we want to apply efficient compression on semantic data streams, we need to first detect frequent graph patterns in RDF streams. In this article, we present FreGraPaD, an algorithm that detects those patterns in a single pass, using exclusively internal memory and following a data structure oriented approach. Experimental results clearly confirm the good accuracy of FreGraPaD in detecting frequent graph patterns from semantic data streams. Fethi Belghaouti, Amel Bouzeghoub, Zakia Kazi-Aoul, Raja Chiky |
RCIS | 4 |
| 2016 | An item/user representation for recommender systems based on bloom filtersabstractThis paper focuses on the items/users representation in the domain of recommender systems. These systems compute similarities between items (and/or users) to recommend new items to users based on their previous preferences. It is often useful to consider the characteristics (a.k.a features or attributes) of the items and/or users. This represents items/users by vectors that can be very large, sparse and space-consuming. In this paper, we propose a new accurate method for representing items/users with low size data structures that relies on two concepts: (1) item/user representation is based on bloom filter vectors, and (2) the usage of these filters to compute bitwise AND similarities and bitwise XNOR similarities. This work is motivated by three ideas: (1) detailed vector representations are large and sparse, (2) comparing more features of items/users may achieve better accuracy for items similarities, and (3) similarities are not only in common existing aspects, but also in common missing aspects. We have experimented this approach on the publicly available MovieLens dataset. The results show a good performance in comparison with existing approaches such as standard vector representation and Singular Value Decomposition (SVD). Manuel Pozo, Raja Chiky, Farid Meziane, Elisabeth Métais |
RCIS | 2 |
| 2016 | POL: A Pattern Oriented Load-Shedding for Semantic Data Stream Processing
Fethi Belghaouti, Amel Bouzeghoub, Zakia Kazi-Aoul, Raja Chiky |
WISE (2) | 4 |
| 2016 | From Business Intelligence to semantic data stream management
Marie-Aude Aufaure, Raja Chiky, Olivier Curé, Houda Khrouf, Gabriel Képéklian |
Future Gener. Comput. Syst. | 2 |
| 2015 | CaLibRe: A Better Consistency-Latency Tradeoff for Quorum Based Replication Systems
Sathiya Prabhu Kumar, Sylvain Lefebvre 0002, Raja Chiky, Eric Gressier-Soudan |
DEXA (2) | 3 |
| 2015 | Using collaborative filtering to enhance domain-independent CBR recommender's personalizationabstractCase-Based Reasoning (CBR) is a problem solving methodology that reuses the knowledge of past experiences to solve new problems. It's a knowledge-based technique that has been introduced to the recommendation field to allow reasoning on domain knowledge and to generate more accurate recommendations. If CBR helps suggesting items that meet the users' search criteria, it has the disadvantage of being domain-dependent (all the reasoning process is generally based on hard-coded domain knowledge) and generating less personalized recommendations. In this paper, we propose an approach for a generic and personalized CBR-based recommender system. First, we use a generic ontology to formalize all the knowledge required during the reasoning process. The ontology represents an intermediate layer between the recommender engine and the application domain to ensure the domain-independence criteria. Second, we propose a hybridization strategy that combines CBR and collaborative filtering to alleviate the limitations of CBR and improve the personalized character of the recommendations. Finally, preliminary validation is performed using a publicly available data set of restaurants. Jihane Karim, Matthieu Manceny, Raja Chiky, Michel Manago, Marie-Aude Aufaure |
RCIS | 3 |
| 2014 | Enhancing Collaborative Filtering Using Semantic Relations in Data
Manuel Pozo, Raja Chiky, Zakia Kazi-Aoul |
ICCCI | 2 |
| 2014 | How can sliding HyperLogLog and EWMA detect port scan attacks in IP traffic?abstractIP networks are constantly targeted by new techniques of denial of service attacks (SYN flooding, port scan, UDP flooding, etc), causing service disruption and considerable financial damage. The on-line detection of DoS attacks in the current high-bit rate IP traffic is a big challenge. We propose in this paper an on-line algorithm for port scan detection. It is composed of two complementary parts: First, a probabilistic counting part, where the number of distinct destination ports is estimated by adapting a method called ‘sliding HyperLogLog’ to the context of port scan in IP traffic. Second, a decisional mechanism is performed on the estimated number of destination ports in order to detect in real time any behavior that could be related to a malicious traffic. This latter part is mainly based on the exponentially weighted moving average algorithm (EWMA) that we adapted to the context of on-line analysis by adding a learning step (supposed without attacks) and improving its update mechanism. The obtained port scan detecting method is tested against real IP traffic containing some attacks. It detects all the port scan attacks within a very short time response (of about 30 s) and without any false positive. The algorithm uses a very small total memory of less than 22 kb and has a very good accuracy on the estimation of the number of destination ports (a relative error of about 3.25 % ), which is in agreement with the theoretical bounds provided by the sliding HyperLogLog algorithm. Yousra Chabchoub, Raja Chiky, Betul Dogan |
EURASIP J. Inf. Secur. | 2 |
| 2012 | A clustering approach for sampling data streams in sensor networks
Alzennyr Da Silva, Raja Chiky, Georges Hébrail |
Knowl. Inf. Syst. | 2 |
| 2010 | Aggregation of asynchronous electric power consumption time series knowing the integralabstractMore and more data mining algorithms are applied to a large number of long time series issued by many distributed sensors. The consequence of the huge volume of data is that data warehouses often contain asynchronous time series, i.e. the values have been sampled and are not anymore observed at the same instants. This is a problem when applying data mining algorithms to such asynchronous time series. The standard way to solve this problem is to interpolate intermediate points. We present here two new interpolation approaches which take into account the knowledge of the integral of the time series between two points. The first approach is naive and uses the history of slope values. The second approach is stochastic and provides a confidence interval of interpolated values. The two methods have been assessed experimentally on a real dataset of electric power consumption time series issued from smart meters. Raja Chiky, Laurent Decreusefond, Georges Hébrail |
EDBT | 1 |
| 2010 | CLUSMASTER: A Clustering Approach for Sampling Data Streams in Sensor NetworksabstractThe growing usage of embedded devices and sensors in our daily lives has been profoundly reshaping the way we interact with our environment and our peers. As more and more sensors will pervade our future cities, increasingly efficient infrastructures to collect, process, and store massive amounts of data streams from a wide variety of sources will be required. Despite the different application-specific features and hardware platforms, sensor network applications share a common goal: periodically sample and store data collected from different sensors in a common persistent memory. In this article we present a clustering approach for rapidly and efficiently computing the best sampling rate which minimizes the SSE (Sum of Square Errors) for each particular sensor in a network. In order to evaluate the efficiency of the proposed approach, we carried out experiments on real electric power consumption data streams produced by a 1-thousand sensor network provided by the French energy group-EDF (Electricite de France). Alzennyr Da Silva, Raja Chiky, Georges Hébrail |
ICDM | 2 |
| 2008 | Summarizing Distributed Data Streams for Storage in Data Warehouses
Raja Chiky, Georges Hébrail |
DaWaK | 1 |