Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Max Halford

dblp:239/4486 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
2since 2021 · last 2023
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Learning paradigms · 41% Learning theory · 29% Time series and sequential data · 29%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 77% Data stream processing · 23%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning and data management
machine learning lifecycle management
0.712023
StreamMLOps: Operationalizing Online Learning for Big Data Streaming & Real-Time Applications · ICDE 2023
Machine learning › Learning paradigms
continual learning
0.512021
River: machine learning for streaming data in Python · J. Mach. Learn. Res. 2021
Machine learning › Learning theory
online learning
0.512021
River: machine learning for streaming data in Python · J. Mach. Learn. Res. 2021
Machine learning › Time series and sequential data
streaming data
0.512021
River: machine learning for streaming data in Python · J. Mach. Learn. Res. 2021
Machine learning › Learning paradigms
incremental learning
0.212023
StreamMLOps: Operationalizing Online Learning for Big Data Streaming & Real-Time Applications · ICDE 2023
Data stream processing › stream mining
streaming machine learning
0.212023
StreamMLOps: Operationalizing Online Learning for Big Data Streaming & Real-Time Applications · ICDE 2023

Methods — techniques the papers use, named apart from their topics

river · 1.3online learning · 1.3kafka · 1.3
YearPublicationVenuePosition
2023 StreamMLOps: Operationalizing Online Learning for Big Data Streaming & Real-Time Applications
abstract
Continuously learning and serving from evolving streaming data and serving in real-time is a challenging problem. Traditionally, data is partitioned and processed in batches to train machine learning (ML) models. In industrial applications, static models’ performance drops over time (model degradation, concept drift), requiring new models to be trained with recent data and redeployed in production. The scientific community has been studying online and adaptive methods to address batch-learning limitations and continuously train AI tasks for industrial applications such as cyber-security, AIOps, anomaly scoring, and drift detection in stock markets. This paper deals with the MLOps aspects of deploying such online and dynamic models to address the requirements in the production systems for real-time applications. Our architectures - based on open-source tools such as Kafka and River - demonstrated how online learning methods could be scaled horizontally in production to meet the demands of a high-velocity streaming pipeline. We demonstrate an MLOps strategy to perform incremental learning from streaming data and continuously deploy the online learning model without pausing the inference pipeline. Indeed, the design satisfies requirements such as model versioning, monitoring, audibility and reproducibility of prediction in both a supervised and semi-supervised setting. Our experiments - for malicious URLs detection task - performed on high-dimensional and feature-evolving streaming data (more than 3 million features) establish the effectiveness and efficiency of online learning models compared to batch (static) machine learning regarding both time and space complexity. Finally, we provide some best practices on data engineering for deploying online models to process a real-time feature stream in production environments. Code is publicly available for reproducibility.
Mariam Barry, Jacob Montiel, Albert Bifet, Sameer Wadkar, Nikolay Manchev, Max Halford, Raja Chiky, Saad El Jaouhari, Katherine B. Shakman, Joudi Al Fehaily, Fabrice Le Deit, Vinh-Thuy Tran, Eric Guerizec
ICDE6
2021 River: machine learning for streaming data in Python
abstract
River is a machine learning library for dynamic data streams and continual learning. It provides multiple state-of-the-art learning methods, data generators/transformers, performance metrics and evaluators for different stream learning problems. It is the result from the merger of two popular packages for stream learning in Python: Creme and scikit-multiflow. River introduces a revamped architecture based on the lessons learnt from the seminal packages. River's ambition is to be the go-to library for doing machine learning on streaming data. Additionally, this open source package brings under the same umbrella a large community of practitioners and researchers. The source code is available at https://github.com/online-ml/river.
Jacob Montiel, Max Halford, Saulo Martiello Mastelini, Geoffrey Bolmier, Raphaël Sourty, Robin Vaysse, Adil Zouitine, Heitor Murilo Gomes, Jesse Read, Talel Abdessalem, Albert Bifet
J. Mach. Learn. Res.2
2019 An Approach Based on Bayesian Networks for Query Selectivity Estimation
Max Halford, Philippe Saint-Pierre, Franck Morvan
DASFAA (2)1