VLDB 2026 Research / reviewers in the wild / expert
Michael Dubé
dblp:222/8017
· DBLP profile ↗
19ranked-venue papers
11as first author
13since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 7 first-author · 9 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Let's Talk 'Bout Mutation: Evolutionary Programming for DNA SequencesabstractSelf-driving automata (SDAs) are extensions of finite state automata that both read and output symbols. Previously, genetic algorithms were used to evolve SDAs to generate sequences that closely matched given DNA sequences, with the eventual goal of finding patterns in those sequences that were not achievable using traditional biological methods. Previous work demonstrated that improvements in fitness were almost exclusively due to mutation not crossover. This paper evaluates the use of evolutionary programming (EP) using SDAs as the representation. EP allows for easy handling of multiple types of mutation, including changes in the number of states, which was not available in earlier approaches. This work uses three fitness metrics: primary sequence matching fitness, secondary sequence similarity fitness, and a relative fitness function known as bout score. Tested on a set of six target DNA sequences, the approach matched 84.4–98.2% of each sequence, and discovered some features of the sequences for future exploration. James Sargant, Michael Dubé, Sheridan K. Houghten, Steffen Graether |
CIBCB | 2 |
| 2024 | Promoting Diversity in the Evolution of Biological Sequence DataabstractA population of Self-Driving Automata (SDAs) is evolved using a steady-state evolutionary algorithm tasked with matching six DNA sequences. This is a step towards using SDAs to assist in identifying patterns within groups of DNA sequences for which conventional methods for identifying patterns fail. The fitness function uses both a primary and a secondary fitness metric to determine the overall fitness of an SDA. The primary metric is sequence matching fitness, evaluating how well the evolved sequence matches the target sequence. The secondary metric is sequence diversity fitness: considering pairs of sequences with the same primary fitness, this counts the number of differences that exist between them. It is found that promoting diversity in this manner by using a secondary fitness metric dramatically improves results for the primary fitness metric. Michael Dubé, Sheridan K. Houghten, Steffen Graether |
CEC | 1 |
| 2024 | Driving Evolution Towards Discovery of Patterns in Sets of Weakly-Conserved DNA SequencesabstractAn evolutionary algorithm is used to evolve a population of self-driving automata (SDAs), modified state machines, that are used to produce output in the form of a DNA sequence. The SDAs are evaluated based on their ability to create sequences that closely match all DNA sequences within a given set. This is evaluated in a pairwise fashion, but attempts to match all sequences concurrently, using two fitness functions that are differentiated by whether they allow for gaps. Additionally, a secondary fitness metric, Sequence Diversity Fitness, encourages diversity among the output of the SDAs within the population throughout evolution. The target sequences are Φ-segments of dehydrin proteins, which are weakly-conserved, can vary considerably in length, and for which traditional methods fail when used to find patterns within them. The ultimate goal is to use SDAs to assist in identifying patterns within Φ-segments. Locating such a pattern could prove fruitful for understanding the functions of dehydrins and how they contribute to the protection of plants from abiotic stresses. Several sets of target sequences are used for analysis, with some sets being more closely-related than others. The evolutionary algorithm was found to produce sequences that matched (according to one of the fitness functions) up to 100% of a given set of target sequences under certain conditions, with closely-related sequences being more accurately matched. Michael Dubé, Sheridan K. Houghten, Steffen Graether |
CIBCB | 1 |
| 2024 | Immunity Vanishing Act: Epidemic Variant and Immunity Analysis via Evolutionary ComputationabstractAn evolutionary algorithm is employed to evolve contact networks representing interactions between individuals in a population. These networks play a crucial role in providing insights into epidemic behaviour and how viruses propagate through the population. The networks are evolved using two fitness functions: one focused on maximizing infection severity and one focused on maximizing the total number of infections (spread). Both evaluate potential networks by simulating SIR epidemics in the context of epidemic variants. The impact of different types of immunity and different probabilities of new variants being generated are evaluated with respect to the number and severity of infections, the characteristics of the evolved contact networks, and the length of the epidemic. The evolutionary algorithm successfully created networks likely to result in more severe infections when focusing on epidemic severity and networks likely to result in a higher number of infections of any severity when focusing on spread, although the immunity type and variant probability both had significant impact. Farhana Yasmeen, Michael Dubé, Sheridan K. Houghten |
CIBCB | 2 |
| 2023 | Comparison of Representations to Evolve Weighted Contact Networks with Epidemic PropertiesabstractTwo evolutionary algorithms are presented for the construction of weighted graphs: one based on self-driving automata (SDA), and one based on “editing” the edges of a graph. The algorithms are evaluated for their success at generating weighted contact networks likely to exhibit specified epidemic behaviour. Two main problems are considered: maximizing the length of the epidemic, and matching the profile (“curve”) of an epidemic, including one based on real-life data. Both algorithms significantly improve upon previous results using unweighted graphs. In most experiments the best overall results are obtained by the SDA algorithm, while the edge-editing algorithm usually has better mean fitness. This result is in part due to the SDA algorithm being more exploratory when compared to the one based on edge editing, which is more exploitative. Michael Dubé, James Sargant, Sheridan K. Houghten |
CIBCB | 1 |
| 2023 | From Bits to Bases: Evolving a Versatile Construct for Biological Sequence and Network DataabstractEvolutionary algorithms are used to evolve Self-Driving Automata (SDAs), finite automata that both read and output symbols. The output of the SDA can be used to generate biological data in the form of sequences or networks. The fitness of an SDA is assessed based on its ability to match real target data: DNA sequences and weighted contact networks. In sequence matching, the SDA method achieves 96.5% accuracy for one of the sequences, and over 90% for half of the target sequences, which range in length from 57 to 102 bases. In network matching, the SDA method is compared to another well-known method using three fitness functions. While the SDA method successfully reproduces multiple clusters in the target network, in general the results lag behind the comparator method. Several avenues for future work are identified, with the eventual goal of using SDAs to identify patterns in biological sequence data. Michael Dubé, James Sargant, Sheridan K. Houghten, Steffen Graether |
CIBCB | 1 |
| 2022 | Now I Know My Alpha, Beta, Gammas: Variants in an Epidemic SchemeabstractPersonal contact networks are used to represent the social connections that exist between individuals within a population. Producing accurate networks that represent the actual vectors of infection that exist within a network can be useful for modelling epidemic trajectory and outcomes, which is significantly impacted by a network's structure. An evolutionary algorithm is used to evolve these networks subject to two fitness measures: epidemic duration and epidemic spread through a population. With each infection there is a small probability of a new variant being generated. Being infected with one variant provides partial immunity to future variants. This allows us to evaluate the impact of each variant, a significant innovation in comparison to other work. The amount by which each variant was allowed to change had a significant impact upon epidemic spread. For epidemic duration, the probability of new variants was the primary cause of increased epidemic duration. Michael Dubé, Sheridan K. Houghten |
CEC | 1 |
| 2022 | Evolving Weighted Contact Networks for Epidemic Modeling: the Ring and the PowerabstractA generative evolutionary algorithm is used to evolve weighted personal contact networks that represent physical contact between individuals, and thus possible paths of infection during an epidemic. The evolutionary algorithm evolves a list of edge-editing operations applied to an initial graph. Two initial graphs are considered, a ring graph and a power-law graph. Different probabilities of infection and a wide range of weights are considered, which improve performance over other work. Modified edge operations are introduced, which also improve performance. It is shown that when trying to maximize epidemic duration, the best results are obtained when using the ring graph as the initial graph. When attempting to match a given epidemic profile, similar results are obtained when using either initial graph, but both improve performance over other work. James Sargant, Sheridan K. Houghten, Michael Dubé |
CEC | 3 |
| 2022 | Evaluation of Frameworks for Epidemic Variants and Infectivity using an Evolutionary AlgorithmabstractAn evolutionary algorithm is used to evolve personal contact networks representing the individuals in a population and the interactions between them. Such networks can be used to track the progress of an epidemic, as it passes from infected individuals to others. Two fitness functions are used: epidemic duration and epidemic spread. Each of these is evaluated in the context of new variants being introduced during the course of the epidemic. Individuals infected with one variant obtain immunity to that variant and possible partial immunity to future variants. Two frameworks for epidemic variants are presented. In the first, infectivity is coupled directly to how well an individual's immunity covers the variant. In the second, infectivity is decoupled, causing a much higher number of infections but with many of lessened severity due to immunity. Michael Dubé, Sheridan K. Houghten |
CIBCB | 1 |
| 2022 | Evolving Lockdown Strategies to Minimize Infections in an EpidemicabstractIn this paper we evaluate the impact of different lockdown strategies upon the total number of infections during an epidemic. The strategies are based upon the percentage of the population infected during a given time step, as well as upon the amount by which interactions must be reduced during lockdown. We use a weighted personal contact network to represent the population, its interactions, and the relative strengths of those interactions. During lockdown edges from this network are removed. We use an evolutionary algorithm to choose the set of edges to be removed so as to minimize infections, comparing different strategies. We show that allowing the evolutionary algorithm to choose which edges to remove significantly reduces the overall number of infections in comparison to random selection. In fact, the EA results for the least stringent conditions were similar or better to the random results for the most stringent conditions, showing that a judicious choice of restrictions during lockdown has the greatest effect on reducing infections. The evolutionary algorithm tends to favour a situation in which during lockdown individuals would reduce their number of contacts, as opposed to lessening the strength of their connections. James Sargant, Michael Dubé, Sheridan K. Houghten |
CIBCB | 2 |
| 2021 | Weighting on the World to Change... an EpidemicabstractA generative evolutionary algorithm is used to create personal contact networks representing which individuals can infect others during an infectious disease scenario. Two problems are considered: (i) finding networks that maximize the length of a simulated epidemic, and (ii) finding networks that match given epidemic profiles. A significant innovation is the introduction of weighted edges to represent the strength of the contact between individuals. Different weight initialization conditions are investigated and evaluated for their performance using a parameter selection mechanism designed to explore the parameter space. Results show that weighted edges were able to increase the overall performance achieved by the evolved networks for both problems considered. Furthermore, it is shown that initializing the weights with a value greater than one further improves performance. The results of the parameter selection mechanism were used to test additional parameter settings thoroughly which further maximize the length of the simulated epidemic for the evolved graphs. Rodrigo Vega Jimenez, Michael Dubé, Sheridan K. Houghten, James Alexander Hughes |
CEC | 2 |
| 2021 | Ring Optimization of Epidemic Contact NetworksabstractThis study compares a current representation for evolving networks to model epidemic spread with a novel representation also studied in a companion paper. This study applies a powerful diversity-friendly algorithm called ring optimization to this novel representation. The problem addressed is that the baseline method is found to optimize only locally; use of the novel representation improves the situation, but not much. The use of ring optimization yields similar or better performance for the ability of the evolved networks to model epidemics while substantially increasing the diversity of those networks. Dan Ashlock, Joseph Alexander Brown, Wendy Ashlock, Michael Dubé |
CIBCB | 4 |
| 2021 | A Comparison of Novel Representations for Evolving Epidemic NetworksabstractRecent work in representation has developed small, evolvable structures called a complex string generator that generate infinite, aperiodic strings of characters. Such a string can be sectioned to provide an arbitrary list of parameters of indefinite length. Other work in evolving networks to model disease transmission has an issue common in many high-dimensional problems, evolution is less efficient when it must get a large number of parameter values correct. Specifying many parameters with a small evolvable object is a potential solution to this problem. In this study we compare three different implementations of representations, two of which employ complex string generators, to specify social contact graphs that plausibly explain the pattern of infection in a small epidemic. Representations that edit a starting network are found to have results that clump in network space while evolving the adjacency matrix provides increased diversity: none of the representations overlap in their results. The adjacency matrix based representation also generated outliers that outperform a baseline representation, probably because of its enhance diversity of solutions. Dan Ashlock, Michael Dubé |
CIBCB | 2 |
| 2020 | Modelling of Vaccination Strategies for Epidemics using Evolutionary ComputationabstractPersonal contact networks that represent social interactions can be used to identify who can infect whom during the spread of an epidemic. The structure of a personal contact network has great impact upon both epidemic duration and the total number of infected individuals. A vaccine, with varying degrees of success, can reduce both the length and spread of an epidemic, but in the case of a limited supply of vaccine a vaccination strategy must be chosen, and this has a significant effect on epidemic behaviour.In this study we consider four different vaccination strategies and compare their effects upon epidemic duration and spread. These are random vaccination, high degree vaccination, ring vaccination, and the base case of no vaccination. All vaccinations are applied as the epidemic progresses, as opposed to in advance. The strategies are initially applied to static personal contact networks that are known ahead of time. They are then applied to personal contact networks that are evolved as the vaccination strategy is applied. When any form of vaccination is applied, all strategies reduce both duration and spread of the epidemic. When applied to a static network, random vaccination performs poorly in terms of reducing epidemic duration in comparison to strategies that take into account connectivity of the network. However, it performs surprisingly well when applied on the evolved networks, possibly because the evolutionary algorithm is unable to take advantage of a fixed strategy. Michael Dubé, Sheridan K. Houghten, Dan Ashlock |
CEC | 1 |
| 2020 | Evolving the CurveabstractEvolutionary algorithms are used to generate personal contact networks, modelling human populations, that are most likely to match a given epidemic profile. The Susceptible-Infected-Removed (SIR) model is used and also expanded upon to allow for an extended period of infection, termed the SIIR model. The networks generated for each of these models are thoroughly evaluated for their ability to match nine different epidemic profiles. The addition of the SIIR model showed that the model of infection has an impact on the networks generated. For the SIR and SIIR models, these differences were relatively minor in most cases. Michael Dubé, Sheridan K. Houghten, Dan Ashlock, James Alexander Hughes |
CIBCB | 1 |
| 2020 | Vaccinating a Population is a Programming ProblemabstractIt is important to understand how best to apply a limited number of vaccines to a population such that the spread of a disease, like SARS-CoV-2, is minimized. Although intuition provides a number of mitigation strategies that may be effective, they remain largely untested.A system was developed to test a given disease mitigation strategy. It is designed to work with a graph representing a real social network. A Genetic Programming system was used to discover novel mitigation strategies that are easily interpretable by a public health decision maker.Effective strategies were developed by the GP system. The strategies are easily explainable and intuitive. Novel mitigation strategies were compared to simple baseline strategies with varying success using a number of different metrics. Many of these strategies proved effective in general, however the topology of the graph influences the effectiveness of a strategy.The system has been made publicly available and the authors call on the research community to contribute their own mitigation strategies and measure their efficacy. James Alexander Hughes, Michael Dubé, Sheridan K. Houghten, Dan Ashlock |
CIBCB | 2 |
| 2019 | Representation for Evolution of Epidemic ModelsabstractCreating a representation capable of generating personal contact networks that are most likely to exhibit specific epidemic behavior is difficult due to the inherit volatility of an epidemic and the numerous parameters accompanying the problem. To surpass these hurdles, evolutionary algorithms are used to create a generative solution which generates personal contact networks, modeling human populations, to satisfy the epidemic duration and epidemic profile matching problems. This representation is entitled the Local THADS-N representation. Two new operators are added to the original THADS-N system, and tested with a traditional parameter sweep and a parameter selection method known as point packing on nine epidemic profiles. Additionally, a new epidemic model is implemented in order to allow for lost immunity within a population thus increasing the length of an epidemic. Michael Dubé, Sheridan K. Houghten, Dan Ashlock |
CEC | 1 |
| 2019 | Pandemic: A Graph Evolution StoryabstractThe Graph Evolution Tool (GET)was built to generate personal contact networks representing who can infect whom within a community. The tool is expanded in order to permit an infection scheme which divides the community into different districts, thus permitting within-district and between-district infections. The evolutionary algorithm comprising GET is expanded upon to simulate communities which include 512 individuals in up to eight districts, initially infecting one person in one district and spreading through a community. The overall goal is to generate communities that will maximize the length of an epidemic. The problem associated with adequately exploring the numerous parameters accompanying evolutionary algorithms is addressed using a point packing and insight from previous work. The Susceptible-Infected-Removed (SIR)model of infection was chosen as it provides a sufficient balance of simplicity and complexity for the problem. Michael Dubé, Sheridan K. Houghten, Dan Ashlock |
CIBCB | 1 |
| 2018 | Parameter selection for modeling of epidemic networksabstractThe accurate modeling of epidemics on social contact networks is difficult due to the variation between different epidemics and the large number of parameters inherent to the problem. To reduce complexity, evolutionary computation is used to create a generative representation of the epidemic model. Previous gains from the use of local, verses global, operators are further explored to better balance exploration and exploitation of the genetic algorithm. A typical parameter study is conducted to test this new local operator and the new method of point packing is utilized as a proof of concept to perform a better search of the parameter space. All experiments from both approaches are tested against nine epidemic profiles. The point-packing driven parameter search demonstrates that the algorithm parameters interact substantially and in a non-linear fashion, and also shows that the good parameter settings are problem specific. Michael Dubé, Sheridan K. Houghten, Dan Ashlock |
CIBCB | 1 |