EDBT 2026 Demo / reviewers in the wild / expert
Guillaume Habault
dblp:153/0361
· DBLP profile ↗
14ranked-venue papers
6as first author
9since 2021 · last 2026
0000-0002-3364-5863ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 7 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sample-Level Prototypical Federated LearningabstractWith the increasing concerns about privacy and data regulations, federated learning (FL) has been emerging as a solution to train machine learning models collaboratively with non-exchangeable data from multiple clients. As a result of data locality, data is usually not identically or independently (non-IID) distributed across clients, and the non-IID property has long been the key challenge in FL. Furthermore, in real-world cross-silo scenarios, it is ubiquitous that clients are organizations owning private data from multiple domains internally, which exacerbates the non-IID issue. For example, in healthcare applications, each client (hospital) gathers data from patients with heterogeneous demographics. While previous works have made efforts to address the non-IID challenge across clients by assuming various relations among client-level data distributions and enabling personalized models at the client level, they ignore the internal data heterogeneity within each client or require explicit data domain indicators, which are hardly accessible in real-world data. Here, we propose Sample-Level Prototypical Federated Learning (SL-PFL) to bridge the gap. SL-PFL incorporates prototypical learning under the FL framework and provides a fine-grained personalized model for each data sample instead of learning one uniform model for all samples of each client. Meanwhile, it can be trained using data without ground-truth domain indicators. Experimental results demonstrate that our proposed method with sample-level personalized models outperforms existing FL methods with a global model or client-level personalized models on various real-world regression and classification tasks from weather, computer vision, and healthcare applications. Chuizheng Meng, Jianke Yang, Hao Niu 0001, Guillaume Habault, Roberto Legaspi, Shinya Wada, Chihiro Ono, Yan Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | CartoMapQA: A Fundamental Benchmark Dataset Evaluating Vision-Language Models on Cartographic Map UnderstandingabstractThe rise of Large Visual-Language Models (LVLMs) has unlocked new possibilities for seamlessly integrating visual and textual information. However, their ability to interpret cartographic maps remains largely unexplored. In this paper, we introduce CartoMapQA, a benchmark specifically designed to evaluate LVLMs' understanding of cartographic maps through question-answering tasks. The dataset includes over 2000 samples, each composed of a cartographic map, a question (with open-ended or multiple-choice answers), and a ground-truth answer. These tasks span key low-, mid- and high-level map interpretation skills, including symbol recognition, embedded information extraction, scale interpretation, and route-based reasoning. Our evaluation of both open-source and proprietary LVLMs reveals persistent challenges: models frequently struggle with map-specific semantics, exhibit limited geospatial reasoning, and are prone to Optical Character Recognition (OCR)-related errors. By isolating these weaknesses, CartoMapQA offers a valuable tool for guiding future improvements in LVLM architectures. Ultimately, it supports the development of models better equipped for real-world applications that depend on robust and reliable map understanding, such as navigation, geographic search, and urban planning. Our source code and data are openly available to the research community at: https://github.com/ungquanghuy-kddi/CartoMapQA.git Huy Quang Ung, Guillaume Habault, Yasutaka Nishimura, Hao Niu 0001, Roberto Legaspi, Tomoki Oya, Ryoichi Kojima, Masato Taya, Chihiro Ono, Atsunori Minamikawa, Yan Liu 0002 |
SIGSPATIAL/GIS | 2 |
| 2024 | Mixture of Projection Experts for Multivariate Long-Term Time Series ForecastingabstractMultivariate long-term time series forecasting (MLTSF), applicable across various domains, has gained increasing research attention. Channel-independent (CI) models, including Linear and Transformer-based architectures, have recently achieved state-of-the-art (SOTA) performance for MLTSF. Notably, Linear models can deliver satisfactory forecasting performance even with just a single linear projection layer. However, we identify a limitation in this architecture: a single linear projection struggles to adequately capture the inter- and intra-variate heterogeneity in temporal patterns. Similarly, any complex models like Transformer-based models that use a single projection layer to generate final predictions, may face capacity bottlenecks. To overcome this, we propose the Mixture of Projection Experts (MoPE), which replaces the single linear projection with multiple projection branches, and employs a gate network to dynamically assign weights to each branch based on the input data. We applied MoPE to multiple SOTA models and evaluated it on nine real-world datasets. Results show that MoPE boosts forecasting accuracy by an average of 9.59%, demonstrating its effectiveness in mitigating the limitations of a single projection layer. Additionally, our experiments demonstrate that integrating our proposal into CI models enhances their generalization to unseen variates. Interpretability analysis also reveals MoPE's ability to disentangle different temporal patterns. Overall, our paper establishes MoPE as an effective solution for MLTSF tasks. Hao Niu 0001, Guillaume Habault, Defu Cao, Roberto Legaspi, Huy Quang Ung, James Enouen, Shinya Wada, Chihiro Ono, Atsunori Minamikawa, Yan Liu 0002 |
ICMLA | 2 |
| 2023 | ICDAR'23: Intelligent Cross-Data Analysis and RetrievalabstractRecently, there has been an increased interest in cross-data research problems, such as predicting air quality using life logging images, predicting congestion using weather and tweets data, and predicting sleep quality using daily exercises and meals. Although several research focusing on multimodal data analytics have been performed, few studies have been conducted on cross-data research (e.g., cross-modal data, cross-domain, cross-platform). The article collection “Intelligent Cross-Data Analysis and Retrieval” aims to encourage research in intelligent cross-data analytics and retrieval and contribute to the creation of a sustainable society. Researchers from diverse domains such as well-being, disaster prevention and mitigation, mobility, climate, tourism and healthcare are welcome to contribute to this Research Topic. Guillaume Habault, Minh-Son Dao, Michael Riegler 0001, Duc-Tien Dang-Nguyen, Yuta Nakashima, Cathal Gurrin |
ICMR | 1 |
| 2023 | Time-delayed Multivariate Time Series PredictionsabstractA major issue with real-time monitoring is to collect complete data. Hardware or software failures, network issues or, more frequently, time delays can disrupt such a collection. This results in having two versions of the same information: one in real-time but with potentially missing data, and the another, albeit complete, is delayed. Many works have studied how to handle missing data for classification and prediction. However, to the best of our knowledge, they do not consider how to leverage the delayed complete data to assist in learning the representation of real-time available data with missing values. This is despite the fact that the delayed complete data contain all the information (e.g., periodicities and trends). In this paper, we propose a framework to enhance the representation learning of the real-time available data by aligning the representation of past real-time but with missing data to that of past delayed but complete data. We test both a distance metric and contrastive learning to achieve this alignment. We implement our framework on a Transformer-based model and experiment it on three datasets. The efficiency of our solution is evaluated against seven baselines and considering four distinct patterns of missing data. Our experiments show that this proposal has a significant improvement in prediction accuracy (5.21% on average) over the baselines. Hao Niu 0001, Guillaume Habault, Roberto Legaspi, Chuizheng Meng, Defu Cao, Shinya Wada, Chihiro Ono, Yan Liu 0002 |
SDM | 2 |
| 2022 | Physics-Informed Long-Sequence Forecasting From Multi-Resolution Spatiotemporal DataabstractSpatiotemporal data aggregated over regions or time windows at various resolutions demonstrate heterogeneous patterns and dynamics in each resolution. Meanwhile, the multi-resolution characteristic provides rich contextual information, which is critical for effective long-sequence forecasting. The importance of such inter-resolution information is more significant in practical cases, where fine-grained data is usually collected via approaches with lower costs but also lower qualities compared to those for coarse-grained data. However, existing works focus on uni-resolution data and cannot be directly applied to fully utilize the aforementioned extra information in multi-resolution data. In this work, we propose Spatiotemporal Koopman Multi-Resolution Network (ST-KMRN), a physics-informed learning framework for long-sequence forecasting from multi-resolution spatiotemporal data. Our method jointly models data aggregated in multiple resolutions and captures the inter-resolution dynamics with the self-attention mechanism. We also propose downsampling and upsampling modules among resolutions to further strengthen the connections among data of multiple resolutions. Moreover, we enhance the modeling of intra-resolution dynamics with physics-informed modules based on Koopman theory. Experimental results demonstrate that our proposed approach achieves the best performance on the long-sequence forecasting tasks compared to baselines without a specific design for multi-resolution data. Chuizheng Meng, Hao Niu 0001, Guillaume Habault, Roberto Legaspi, Shinya Wada, Chihiro Ono, Yan Liu 0002 |
IJCAI | 3 |
| 2022 | Mu2ReST: Multi-resolution Recursive Spatio-Temporal Transformer for Long-Term Prediction
Hao Niu 0001, Chuizheng Meng, Defu Cao, Guillaume Habault, Roberto Legaspi, Shinya Wada, Chihiro Ono, Yan Liu 0002 |
PAKDD (1) | 4 |
| 2021 | Influence of Land Use information over performance when predicting spatiotemporal electricity load demandabstractIn a world where the growing concern about climate change becomes more apparent, electricity industry has its role to play in slowing down such an evolution. In fact, tools and services could be provided to consumers in order to better manage their ever-increasing consumption. But, while waiting for such a complex solution to be deployed, a first step would be to better forecast current consumption. Such an improvement would enable to generate electricity accordingly, while renewable solutions are not presently optimal – i.e., intermittent, long-term storage issue, etc.Electricity Load Demand (ELD) fluctuations do not solely depend on temporal factors, they also have a spatial contribution that is often underestimated. Indeed, it is possible to map different ELD profiles with their Land Use (LU) information. For example, ELD in residential areas fits with residents’ commuting style, while in industrial areas, it mostly correlates with working hours. This paper aims to investigate such a relation. Considering ELD data on a 500m square grid and LU data on a 100m square grid, different 500m cell labeling has been defined. In order to assess if cells with similar labels have similar ELD, they have been grouped under the same forecast model. Several types of models (Multilayer Perceptron, Recurrent Neural Network) have been used to compare their performance and efficiency. This study confirms that, all models considered, some labels are more difficult to forecast than others. Such associations can reduce by over 30% the prediction error compare to a per cell scenario. Additional investigations would be needed to further reduce prediction error and to help models better seize the land specificity of each grid-cell. External data also affects ELD, and pairing them with an optimal LU labeling could be a promising solution. Guillaume Habault, Shinya Wada, Chihiro Ono |
IEEE BigData | 1 |
| 2021 | Is Adding More Modalities Better in a Multimodal Spatio-temporal Prediction Scenario? A Case Study on Japan Air Quality
Yutaro Mishima, Guillaume Habault, Shinya Wada |
MobiQuitous | 2 |
| 2020 | Elucidating the extent by which population staying patterns help improve electricity load demand predictionsabstractThe need for electricity has never been more important these days. In order to achieve balance between generation and distribution - as well as schedule operations accordingly-high-accuracy load demand predictions are mandatory. But our society is currently undergoing modifications in electricity consumption allocation. We are witnessing a fast shift from office-to home- based working style. As a consequence, Electricity load demand prediction models are in need for additional data in order to quickly adapt to these modifications and maintain efficient predictions accuracy. The rising popularity of "tracking" devices and alike-applications opens up to a new type of multi-modal investigations. The availability of associated location data enables researcher to study mobility routine and patterns in order to cross it with other data. Electricity consumption is one domain impacted by people's mobility behavior (commuting, telework, etc.) as people are not "plugged" onto the power grid while moving. This paper presents a study on population staying patterns and how it can relate to electricity load demand. Time-series data providing the number of people staying in a given area has been used within a Deep Learning model in order to enhance electricity load demand predictions at the provider level. It unveils the potential usage of such dynamics data, while setting the foundations for more complex studies. Guillaume Habault, Shinya Wada, Rui Kimura, Chihiro Ono |
IEEE BigData | 1 |
| 2019 | Detecting errors in short-term electricity demand forecast using people dynamicsabstractThe landscape of power grids is gradually changing. The growing number of electrical appliances as well as the outbreak of Electric Vehicles (EVs) is increasing the need for electricity. As a consequence, high-accuracy consumption predictions are necessary in order to both schedule and plan production and operations accordingly. The emergence of connected “tracking” devices opens up new data-sets into both Internet-of-Things and Big Data worlds. It provides information on human dynamics (people mobility behavior) and with it several opportunities. Electricity consumption is impacted by people movements as while moving they are not “connected” to the power grid. Therefore, predicting such movement patterns and volume could help electricity providers improve their own consumption predictions. This paper presents a system and methods used to predict people movement behaviors as well as detect any anomaly. A scoring system is used to both evaluate the dynamics predictions and raise alerts when the computed score surpasses established thresholds. This proposal is tested over a scenario using available datasets and demonstrates that modifications in people movement behavior is affecting the consumption profile. Guillaume Habault, Yasutaka Nishimura, Kiyohito Yoshihara, Chihiro Ono |
IEEE BigData | 1 |
| 2018 | Delivery Management System Based on Vehicles Monitoring and a Machine-Learning MechanismabstractThe continuously growing online shopping is increasing the number of attended home deliveries. The last-mile delivery plays an important role in online shopping satisfaction and especially for food deliveries. This paper focuses on food delivery retailers and particularly investigates the possibility to enhance deliveries using information and data knowledge. In fact, in addition to optimize and to share delivery routes, delivery vehicles could be monitored in order to always maintain shortest delivery delays. We propose in this paper a delivery management architecture targeting these principles. This system is composed of several core mechanisms that should keep delivery delays to a minimum while maintaining low service times. A proof-of-concept of this delivery management system has been developed using Electric Scooters, smartphones and several algorithms. It demonstrates how this architecture could work in a food delivery scenario. Guillaume Habault, Yuya Taniguchi, Naoaki Yamanaka |
VTC Fall | 1 |
| 2017 | Monitoring Traffic Optimization in a Smart GridabstractThe emergence of microgeneration systems steadily increases, and it raises concerns regarding their impact on the power grid. It is, therefore, crucial to efficiently integrate them into future smart grid architectures, as there is not any standard way to monitor production units. Moreover, current data collection systems are simple and do not consider their impact on local area networks. This paper presents a set of proposed mechanisms that reduces the monitoring traffic, while offering management flexibility on large-scale systems. This study is illustrated with measurements performed on a small grid, and it shows that, for monitoring a photovoltaic production, both 1-min and 1-s intervals provide the same production estimation, while significantly decreasing the associated traffic. It can be reduced even more by aggregating several measurements during a given period before sending them and by using specific mechanisms to ensure reliability. This experiment also helps authors identify best practices for monitoring different equipment based on their behaviors. Guillaume Habault, Maxime Lefrançois, François Lemercier, Nicolas Montavont, Periklis Chatzimisios, Georgios Z. Papadopoulos |
IEEE Trans. Ind. Informatics | 1 |
| 2016 | Effective communicating optimization for V2G with electric busabstractThe number of connected devices - also known as Internet of Things (IoT) - is exponentially increasing. Such sensors and devices also appear in transportation systems giving some intelligence to roads, equipment and vehicles. Nowadays, it is possible to communicate with the environment in order to have better everyday services. Furthermore, the number of registered - public or private - Electric Vehicle (EVs) is continuously increasing. These vehicles, equipped with large battery, need to be charged and so, have a significant impact on power grids. However, these EVs can also be seen as energy sources. It is therefore important to be able to plan both the charge and discharge of EVs. Including these vehicles into Vehicle-to-Grid technology is a way to efficiently manage such pools of batteries. But, as a consequence, grid requires to have almost real-time data on these vehicles and especially their battery status. This paper studies an optimized data aggregation method for a fleet of electric buses. Each bus provides different type of information with different priority level. The efficiency of the studied method was evaluated with a simulation platform developed with ns-3. Simulation results - based on real route and bus stop positions - show that an optimal buffer size has been found to both satisfy transmission delays and optimize communications. Toshichika Shiobara, Guillaume Habault, Jean-Marie Bonnin, Hiroaki Nishi |
INDIN | 2 |