VLDB 2026 Research / reviewers in the wild / expert
Daniel Aloise
dblp:29/6167 · also Daniel J. Aloise
· DBLP profile ↗
29ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0002-9876-2921ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 8 · 4 since 2021Theory of computation · 7 · 5 first-author · 2 since 2021Software engineering, systems software and programming languages · 6 · 5 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Extracting Causal Relations from Log Sequences Using Causal Language Models
Vithor Gomes Bertalan, Fateme Faraji Daneshgar, Daniel Aloise |
SANER | 3 |
| 2026 | A Column Generation Algorithm with Dynamic Constraint Aggregation for Minimum Sum-of-Squares ClusteringabstractThe minimum sum-of-squares clustering problem (MSSC), also known as k-means clustering, refers to the problem of partitioning n data points into k clusters, with the objective of minimizing the total sum of squared Euclidean distances between each point and the center of its assigned cluster. We propose an efficient algorithm for solving large-scale MSSC instances, which combines column generation (CG) with dynamic constraint aggregation (DCA) to effectively reduce the number of constraints considered in the CG master problem. DCA was originally conceived to reduce degeneracy in set partitioning problems by utilizing an aggregated restricted master problem obtained from a partition of the set partitioning constraints into disjoint clusters. In this work, we explore the use of DCA within a CG algorithm for MSSC exact solution. Our method is fine-tuned by a series of ablation studies on DCA design choices, and is demonstrated to significantly outperform existing state-of-the-art exact approaches available in the literature. History: Accepted by Andrea Lodi, Area Editor for Design & Analysis of Algorithms–Discrete. Funding: This work was supported by Natural Sciences and Engineering Research Council of Canada [Grant 2023-04466]. Supplemental Material: The software that supports the findings of this study is available within the paper and its Supplemental Information ( https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2024.0938 ) as well as from the IJOC GitHub software repository ( https://github.com/INFORMSJoC/2024.0938 ). The complete IJOC Software and Data Repository is available at https://informsjoc.github.io/ . Antonio Maria Sudoso, Daniel Aloise |
INFORMS J. Comput. | 2 |
| 2025 | Data Mining-Driven Shift Enumeration for Accelerating the Solution of Large-Scale Personnel Scheduling ProblemsabstractThis study addresses large-scale personnel scheduling problems in the service industry by combining mathematical programming with data mining techniques to enhance efficiency. The studied problem aims at efficiently scheduling skilled employees over a one-week planning horizon, minimizing costs while meeting diverse job demands. In service industries, shift planning is intricately tied to customer presence, leading to a multitude of potential shifts and a difficult optimization problem that cannot be easily solved using a commercial mixed-integer programming solver. Nevertheless, these problems are categorized as recurrent problems, where distinct instances share common characteristics and solution structures that differ only in a few parameters over time. We propose to use a data mining technique, namely, the \(k\) -nearest neighbors algorithm, to expedite the solution process while upholding solution quality. We suggest using schedules of past solutions to reduce the problem size. Thus, for an upcoming instance, we identify similar historical instances and streamline the enumeration of shifts to align with the comparable historical instances’ schedules. This approach allows us to solve the problem using a commercial solver within a reasonable timeframe while preserving solution quality. Moreover, our methodology offers decision-makers the flexibility to determine the extent to which they wish to scale down the problem. Our experiments conducted on instances generated from real historical data with up to 12 jobs and 252 employees, yield an average removal of up to 85.5% of decision variables. This resulted in an average speedup factor of up to 15.5, with a marginal average cost increase of approximately 1.2%. Farin Rastgar-Amini, Daniel Aloise, Claudio Contardo, Guy Desaulniers |
ACM Trans. Evol. Learn. Optim. | 2 |
| 2023 | Using Transformer Models and Textual Analysis for Log ParsingabstractLog parsing has become an essential tool for extracting valuable information from a vast volume of log lines. It involves identifying standard patterns and extracting templates from these lines, enabling researchers and companies to employ advanced mining techniques like log deduplication and log anomaly detection. However, existing log parsing approaches have limitations. They often operate on small batches of log text and lack consideration for the entire context, necessitating prior knowledge of the log dataset. Furthermore, there is a scarcity of practical experience reports on the utilization of these log parsing approaches in the literature. In our paper, we address these challenges by proposing a novel log parsing approach that combines Transformers with a customized textual analysis. This textual analysis balances the density clustering of similar log lines, the calculation of word frequencies inside each cluster and the presence of the words inside an English vocabulary to parse new log lines. Our method outperforms existing parsing methods in terms of accuracy and operates as an unsupervised learning approach, eliminating the need for prior knowledge. Additionally, we present a proof-of-concept application of our parsing method in an industrial setting, showcasing its practical implementation. Vithor Gomes Bertalan, Daniel Aloise |
ISSRE | 2 |
| 2023 | A Lagrangian-based approach to learn distance metrics for clustering with minimal data transformationabstractDistance metric learning algorithms aim to learn how to measure similarities between data objects in a metric space. In the context of clustering, metric learning typically relies on side-information provided by experts, most commonly expressed in the form of pairwise constraints. In this setting, algorithms for metric learning execute data transformations that bring pairs of data points involved in must-link constraints close together, whereas pair of points involved in cannot-link constraints are moved away from each other. One caveat to such methods is that they can considerably change the original data distribution properties. With that in mind, we propose a Lagrangian-based approach to assist distance metric learning algorithms for clustering. Our method is developed to identify the least impactful transformations to the original data space, while still learning a more suitable metric space for grouping the data using the provided side information. Our results demonstrate that the proposed methodology is able to achieve a competitive clustering performance with respect to truth classification. Furthermore, the method is able to provide more accurate views of the transformed datasets, which can lead to more reliable clustering interpretations. Rodrigo Randel, Daniel Aloise, Alain Hertz |
SDM | 2 |
| 2022 | OptiMaP: swarm-powered Optimized 3D Mapping Pipeline for emergency response operationsabstractA smart application in sensing is mainly powered by a two-stage process comprising sensing (collect data) and computing (process data). While the sensing stage is typically performed locally through a dedicated Internet of Things infrastructure, the computing stage may require a powerful infrastructure in the cloud. However, when connectivity is poor and low latency becomes a requirement — as in emergency response and disaster relief operations — edge computing and ad hoc cloud paradigms come in support to keep the computing stage locally. Being local network connectivity and data processing limited, it is vital to properly optimize how the computing workload will be consumed by the local ad hoc cloud. For this purpose, we present and evaluate the swarm-powered Optimized 3D Mapping Pipeline (OptiMaP) for emergency response 3D mapping missions, which is implemented as a collaborative embedded Robot Operating System (ROS) application integrating an ad hoc telecommunication middleware.We simulate — with Software-In-The-Loop — realistic 3D mapping missions comprising up to 5 drones and 363 images covering 0.293km2. We show how the completion times of mapping missions carried out in a typical centralized manner can be dramatically reduced by two versions of the OptiMaP framework powered, respectively, by a variable neighborhood search heuristic and a greedy method. Leandro Rincon Costa, Daniel Aloise, Luca Giovanni Gianoli, Andrea Lodi 0001 |
DCOSS | 2 |
| 2022 | HurricaneLog: A humanitarian logistics game on hurricane preparedness and response operations
Thiago Pereira, Daniel Aloise, Marie-Ève Rancourt, Danny Godin |
DiGRA | 2 |
| 2022 | FaST: A linear time stack trace alignment heuristic for crash report deduplicationabstractIn software projects, applications are often monitored by systems that automatically identify crashes, collect their information into reports, and submit them to developers. Especially in popular applications, such systems tend to generate a large number of crash reports in which a significant portion of them are duplicate. Due to this high submission volume, in practice, the crash report deduplication is supported by devising automatic systems whose efficiency is a critical constraint. In this paper, we focus on improving deduplication system throughput by speeding up the stack trace comparison. In contrast to the state-of-the-art techniques, we propose FaST, a novel sequence alignment method that computes the similarity score between two stack traces in linear time. Our method independently aligns identical frames in two stack traces by means of a simple alignment heuristic. We evaluate FaST and five competing methods on four datasets from open-source projects using ranking and binary metrics. Despite its simplicity, FaST consistently achieves state-of-the-art performance regarding all metrics considered. Moreover, our experiments confirm that FaST is substantially more efficient than methods based on optimal sequence alignment. Irving Muller Rodrigues, Daniel Aloise, Eraldo Rezende Fernandes |
MSR | 2 |
| 2022 | TraceSim: An Alignment Method for Computing Stack Trace Similarity
Irving Muller Rodrigues, Aleksandr Khvorov, Daniel Aloise, Roman Vasiliev, Dmitrij V. Koznov, Eraldo Rezende Fernandes, George A. Chernishev, Dmitry V. Luciv, Nikita Povarov |
Empir. Softw. Eng. | 3 |
| 2021 | On Improving Deep Learning Trace Analysis with System Call ArgumentsabstractKernel traces are sequences of low-level events comprising a name and multiple arguments, including a timestamp, a process id, and a return value, depending on the event. Their analysis helps uncover intrusions, identify bugs, and find latency causes. However, their effectiveness is hindered by omitting the event arguments. To remedy this limitation, we introduce a general approach to learning a representation of the event names along with their arguments using both embedding and encoding. The proposed method is readily applicable to most neural networks and is task-agnostic. The benefit is quantified by conducting an ablation study on three groups of arguments: call-related, process-related, and time-related. Experiments were conducted on a novel web request dataset and validated on a second dataset collected on pre-production servers by Ciena, our partnering company. By leveraging additional information, we were able to increase the performance of two widely-used neural networks, an LSTM and a Transformer, by up to 11.3% on two unsupervised language modelling tasks. Such tasks may be used to detect anomalies, pre-train neural networks to improve their performance, and extract a contextual representation of the events. Quentin Fournier, Daniel Aloise, Seyed Vahid Azhari, François Tetreault |
MSR | 2 |
| 2021 | A Lagrangian-based score for assessing the quality of pairwise constraints in semi-supervised clustering
Rodrigo Randel, Daniel Aloise, Simon J. Blanchard, Alain Hertz |
Data Min. Knowl. Discov. | 2 |
| 2021 | The Covering-Assignment Problem for Swarm-Powered Ad Hoc Clouds: A Distributed 3-D Mapping UsecaseabstractThe popularity of drones is rapidly increasing across the different sectors of the economy. Aerial capabilities and relatively low costs make drones the perfect solution to improve the efficiency of operations that are typically carried out by humans. Besides automating field operations, drones acting de facto as a swarm can serve as an ad hoc cloud infrastructure built on top of computing and storage resources available across the swarm members and other elements. Even in the absence of Internet connectivity, this cloud can serve the workloads generated by the swarm members and the field agents. By considering the practical example of a swarm-powered 3-D reconstruction application on top of such cloud infrastructure, we present a new optimization problem for the efficient generation and execution of multinode computing workloads subject to data geolocation and clustering constraints. The objective is the minimization of the overall computing times, including both networking delays caused by the interdrone data transmission and computation delays. We prove that the problem is NP-hard and present two combinatorial formulations to model it. Computational results on the solution of the formulations show that one of them can be used to solve, within the configured time-limit, more than 50% of the considered real-world instances involving up to two hundred images and six drones. Leandro Rincon Costa, Daniel Aloise, Luca Giovanni Gianoli, Andrea Lodi 0001 |
IEEE Internet Things J. | 2 |
| 2021 | Preface to the special issue of JOGO on the occasion of the 40th anniversary of the Group for Research in Decision Analysis (GERAD)
Daniel Aloise, Gilles Caporossi, Sébastien Le Digabel |
J. Glob. Optim. | 1 |
| 2020 | An Exact CP Approach for the Cardinality-Constrained Euclidean Minimum Sum-of-Squares Clustering Problem
Mohammed Najib Haouas, Daniel Aloise, Gilles Pesant |
CPAIOR | 2 |
| 2020 | A Soft Alignment Model for Bug DeduplicationabstractBug tracking systems (BTS) are widely used in software projects. An important task in such systems consists of identifying duplicate bug reports, i.e., distinct reports related to the same software issue. For several reasons, reporting bugs that have already been reported is quite frequent, making their manual triage impractical in large BTSs. In this paper, we present a novel deep learning network based on soft-attention alignment to improve duplicate bug report detection. For a given pair of possibly duplicate reports, the attention mechanism computes interdependent representations for each report, which is more powerful than previous approaches. We evaluate our model on four well-known datasets derived from BTSs of four popular open-source projects. Our evaluation is based on a ranking-based metric, which is more realistic than decision-making metrics used in many previous works. Achieved results demonstrate that our model outperforms state-of-the-art systems and strong baselines in different scenarios. Finally, an ablation study is performed to confirm that the proposed architecture improves the duplicate bug reports detection. Irving Muller Rodrigues, Daniel Aloise, Eraldo Rezende Fernandes, Michel R. Dagenais |
MSR | 2 |
| 2020 | RecSeats: A Hybrid Convolutional Neural Network Choice Model for Seat Recommendations at Reserved Seating VenuesabstractPredicting locational choices (i.e., where one chooses to sit) is a challenging task because preferences are highly heterogeneous and depend not only on the location of the seats in the environment but also on the location of others. In the present research, we propose RecSeats - a framework to predict locational choices. The framework augments individual-level discrete choice models with a convolutional neural network (CNN) which can capture higher order interactions between features of available seats. The framework is flexible and can accommodate complexity in real-world locational choice data such as variability in the number of tickets purchased and the number and locations from past purchases. Applied to both locational choice experiment data and to ticketing data from a large North-American concert hall, we show that augmenting individual-level discrete choice models with a CNN consistently provides strong predictive accuracy. Théo Moins, Daniel Aloise, Simon J. Blanchard |
RecSys | 2 |
| 2020 | Convex fuzzy k-medoids clustering
Daniel Nobre Pinheiro, Daniel Aloise, Simon J. Blanchard |
Fuzzy Sets Syst. | 2 |
| 2018 | Towards Station-Level Demand Prediction for Effective Rebalancing in Bike-Sharing SystemsabstractBike sharing systems continue gaining worldwide popularity as they offer benefits on various levels, from society to environment. Given that those systems tend to be unbalanced along time, bikes are typically redistributed throughout the day to better meet the demand. Reasonably accurate demand prediction is key to effective redistribution; however, it is has received only little attention in the literature. In this paper, we focus on predicting the hourly demand for demand rentals and returns at each station of the system. The proposed model uses temporal and weather features to predict demand mean and variance. It first extracts the main traffic behaviors from the stations. These simplified behaviors are then predicted and used to perform station-level predictions based on machine learning and statistical inference techniques. We then focus on determining decision intervals, which are often used by bike sharing companies for their online rebalancing operations. Our models are validated on a two-year period of real data from BIXI Montréal. A worst-case analysis suggests that the intervals generated by our models may decrease unsatisfied demands by 30% when compared to the current methodology employed in practice. Pierre Hulot, Daniel Aloise, Sanjay Dominik Jena |
KDD | 2 |
| 2018 | A sampling-based exact algorithm for the solution of the minimax diameter clustering problem
Daniel Aloise, Claudio Contardo |
J. Glob. Optim. | 1 |
| 2018 | Parallel synchronous and asynchronous coupled simulated annealing
Kayo Gonçalves-e-Silva, Daniel Aloise, Samuel Xavier de Souza |
J. Supercomput. | 2 |
| 2017 | Less is more: basic variable neighborhood search heuristic for balanced minimum sum-of-squares clustering
Leandro Rincon Costa, Daniel Aloise, Nenad Mladenovic |
Inf. Sci. | 2 |
| 2017 | NP-Hardness of balanced minimum sum-of-squares clustering
Artem V. Pyatkin, Daniel Aloise, Nenad Mladenovic |
Pattern Recognit. Lett. | 2 |
| 2014 | Reactive Search strategies using Reinforcement Learning, local search algorithms and Variable Neighborhood Search
João Paulo Queiroz dos Santos, Jorge Dantas de Melo, Adrião Duarte Dória Neto, Daniel Aloise |
Expert Syst. Appl. | 4 |
| 2014 | Global optimization workshop 2012
Daniel Aloise, Pierre Hansen, Caroline T. M. Rocha |
J. Glob. Optim. | 1 |
| 2014 | Column generation bounds for numerical microaggregation
Daniel Aloise, Pierre Hansen, Caroline T. M. Rocha, Éverton Santi |
J. Glob. Optim. | 1 |
| 2012 | A VNS heuristic for escaping local extrema entrapment in normalized cut clustering
Pierre Hansen, Daniel Aloise |
Pattern Recognit. | 3 |
| 2011 | Evaluating a branch-and-bound RLT-based algorithm for minimum sum-of-squares clustering
Daniel Aloise, Pierre Hansen |
J. Glob. Optim. | 1 |
| 2009 | NP-hardness of Euclidean sum-of-squares clustering
Daniel Aloise, Amit Deshpande 0001, Pierre Hansen, Preyas Popat |
Mach. Learn. | 1 |
| 2006 | Scheduling workover rigs for onshore oil production
Dario J. Aloise, Daniel Aloise, Caroline T. M. Rocha, Celso C. Ribeiro, José C. Ribeiro Filho, Luiz S. S. Moura |
Discret. Appl. Math. | 2 |