EDBT 2026 Demo / reviewers in the wild / expert
Erik Hemberg
dblp:56/3697
· DBLP profile ↗
55ranked-venue papers
10as first author
22since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 50 · 10 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 1 first-author · 3 since 2021Systems, architecture and hardware · 7 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Hybrid LLM-Coevolution Framework to Generate Abusive Tax Strategies
Joy Sera Bhattacharya, Erik Hemberg, Una-May O'Reilly |
EuroGP | 2 |
| 2026 | Discrete Self-Adaptation in Competitive Coevolution for Constrained HardwareabstractSelf-adaptive competitive coevolutionary algorithms (CCEAs) typically rely on real-valued evolutionary parameters such as mutation rate. Yet emerging deployment targets, including specialized and resource-constrained hardware, often require integer-only or low-precision arithmetic which raises questions about how self-adaptation behaves under discrete constraints. Steven Jorgensen, Jennifer Fairhurst, Erik Hemberg, Una-May O'Reilly |
GECCO | 3 |
| 2026 | Evolutionary Computation for Tax-Minimizing Strategies in Special Economic ZonesabstractGovernments establish Special Tax Zones to promote specific economic objectives through tax incentives. However, firms often exploit these incentives by complying with the letter of the law while ignoring its purpose. Because economic transactions involve many interacting factors, tax authorities typically rely on reactive enforcement to identify abusive schemes arising from unforeseen legal loopholes. We address this problem by proposing an anticipatory modeling approach based on evolutionary computation for settings in which tax incentives affect different stages of a production chain. We formalize the decision space of profit-shifting strategies adopted by entities acting jointly facing different tax rates along production chains as a combinatorial optimization problem through the construction of a Production and Tax Zone Model, which is solved using a genetic algorithm. Our computational experiments in settings inspired by real tax laws show that this framework can generate production schemes that exploit tax incentives in varied ways. In addition, we conduct a case study that provides insights into how changes in economic and tax environments interact with the emergence of different abusive schemes, highlighting the model's potential as a policy-oriented tool for anticipating law-induced loopholes in complex production settings. Andres Leguizamon, Carlos David Sanchez, Sofia Ocampo, Una-May O'Reilly, Erik Hemberg |
GECCO | 5 |
| 2026 | Policy Search through Genetic Programming and LLM-assisted Curriculum LearningabstractCurriculum learning (CL) consists in using a diverse set of user-provided test cases, with varying levels of difficulty and organized in a suitable progression, for learning a policy. The quality of test cases is important to allow optimization techniques as genetic programming (GP) to solve policy search problems. In this work, we evaluate large language models (LLMs) as providers of test cases for GP-based policy search. We consider two policy search tasks, a single-player and a multi-player game, and four LLMs differing in complexity and specialization, which we prompt in order to generate suitable test cases for the two games. We experimentally assess the intrinsic quality of LLM-generated test cases and their utility when inserted in a curriculum consumed by a GP optimization. We evaluate the robustness of the approach with respect to the way cases are scheduled in curricula and with respect to the policy representation, for which we use both graphs and linear programs evolved by GP. We observe that the effectiveness of LLM-assisted CL depends on both the choice of LLM and the design of the prompting and scheduling strategies. These findings highlight important considerations for leveraging LLMs in automated curriculum design for GP-based optimization. Steven Jorgensen, Giorgia Nadizar, Gloria Pietropolli, Luca Manzoni, Eric Medvet, Una-May O'Reilly, Erik Hemberg |
ACM Trans. Evol. Learn. Optim. | 7 |
| 2025 | Runtime Bounds for a Coevolutionary Algorithm on Classes of Potential GamesabstractCoevolutionary algorithms are a family of black-box optimisation algorithms with many applications in game theory. We study a coevolutionary algorithm on an important class of games in game theory: potential games. In these games, a real-valued function defined over the entire strategy space encapsulates the strategic choices of all players collectively. We present the first theoretical analysis of a coevolutionary algorithm on potential games, showing a runtime guarantee that holds for all exact potential games, some weighted and ordinal potential games, and certain non-potential games. Using this result, we show a polynomial runtime on singleton congestion games. Furthermore, we show that there exist games for which coevolutionary algorithms find Nash equilibria exponentially faster than best or better response dynamics, and games for which coevolutionary algorithms find better Nash equilibria as well. Finally, we conduct experimental evaluations showing that our algorithm can outperform widely used algorithms, such as better response on random instances of singleton congestion games, as well as fictitious play, counterfactual regret minimisation (CFR), and external sampling CFR on dynamic routing games. Mario Alejandro Hevia Fajardo, Jamal Toutouh, Erik Hemberg, Una-May O'Reilly, Per Kristian Lehre |
FOGA | 3 |
| 2025 | Guiding Evolutionary AutoEncoder Training with Activation-Based Pruning OperatorsabstractThis study explores a novel approach to neural network pruning using evolutionary computation, focusing on simultaneously pruning the encoder and decoder of an autoencoder. We introduce two new mutation operators that use layer activations to guide weight pruning. Our findings reveal that one of these activation-informed operators outperforms random pruning, resulting in more efficient autoencoders with comparable performance to canonically trained models. Prior work has established that autoencoder training is effective and scalable with a spatial coevolutionary algorithm that cooperatively coevolves a population of encoders with a population of decoders, rather than one autoencoder. We evaluate how the same activity-guided mutation operators transfer to this context. We find that random pruning is better than guided pruning, in the coevolutionary setting. This suggests activation-based guidance proves more effective in low-dimensional pruning environments, where constrained sample spaces can lead to deviations from true uniformity in randomization. Conversely, population-driven strategies enhance robustness by expanding the total pruning dimensionality, achieving statistically uniform randomness that better preserves system dynamics. We experiment with pruning according to different schedules and present best combinations of operator and schedule for the canonical and coevolving populations cases. Steven Jorgensen, Erik Hemberg, Jamal Toutouh, Una-May O'Reilly |
GECCO | 2 |
| 2025 | LLM-Supported Natural Language to Bash TranslationabstractFinnian Westenfelder, Erik Hemberg, Stephen Moskal, Una-May O’Reilly, Silviu Chiricescu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Finnian Westenfelder, Erik Hemberg, Stephen Moskal, Una-May O'Reilly, Silviu Chiricescu |
NAACL (Long Papers) | 2 |
| 2024 | A Self-adaptive Coevolutionary AlgorithmabstractCoevolutionary algorithms are helpful computational abstractions of adversarial behavior and they demonstrate multiple ways that populations of competing adversaries influence one another. We introduce the ability for each competitor's mutation rate to evolve through self-adaptation. Because dynamic environments are frequently addressed with self-adaptation, we set up dynamic problem environments to investigate the impact of this ability. For a simple bilinear problem, a sensitivity analysis of the adaptive method's parameters reveals that it is robust over a range of multiplicative rate factors, when the rate is changed up or down with equal probability. An empirical study determines that each population's mutation rates converge to values close to the error threshold. Mutation rate dynamics are complex when both populations adapt their rates. Large scale empirical self-adaptation results reveal that both reasonable solutions and rates can be found. This addresses the challenge of selecting ideal static mutation rates in coevolutionary algorithms. The algorithm's payoffs are also robust. They are rarely poor and frequently they are as high as the payoff of the static rate to which they converge. On rare runs, they are higher. Mario Alejandro Hevia Fajardo, Erik Hemberg, Jamal Toutouh, Una-May O'Reilly, Per Kristian Lehre |
GECCO | 2 |
| 2024 | Cooperative Coevolutionary Spatial Topologies for Autoencoder TrainingabstractTraining autoencoders is non-trivial. Convergence to the identity function or overfitting are common pitfalls. Population based algorithms like coevolutionary algorithms can provide diversity. To more robustly train autoencoders, we introduce a novel cooperative coevolutionary algorithm that exploits a spatial topology. We investigate the impact of algorithm parameters and design choices on the performance. On a simple tunable benchmark problem we observe that the performance can be improved over that of an conventionally trained autoencoder. However, the training convergence can be slow, despite the final model performance being competitive with a conventional autoencoder. Erik Hemberg, Una-May O'Reilly, Jamal Toutouh |
GECCO | 1 |
| 2024 | Large Language Model-based Test Case Generation for GP AgentsabstractGenetic programming (GP) is a popular problem-solving and optimization technique. However, generating effective test cases for training and evaluating GP programs requires strong domain knowledge. Furthermore, GP programs often prematurely converge on local optima when given excessively difficult problems early in their training. Curriculum learning (CL) has been effective in addressing similar issues across different reinforcement learning (RL) domains, but it requires the manual generation of progressively difficult test cases as well as their careful scheduling. In this work, we leverage the domain knowledge and the strong generative abilities of large language models (LLMs) to generate effective test cases of increasing difficulties and schedule them according to various curricula. We show that by integrating a curriculum scheduler with LLM-generated test cases we can effectively train a GP agent player with environments-based curricula for a single-player game and opponent-based curricula for a multi-player game. Finally, we discuss the benefits and challenges of implementing this method for other problem domains. Steven Jorgensen, Giorgia Nadizar, Gloria Pietropolli, Luca Manzoni, Eric Medvet, Una-May O'Reilly, Erik Hemberg |
GECCO | 7 |
| 2023 | Genetic Programming and Coevolution to Play the Bomberman™ Video Game
Robert Gold, Henrique Branquinho, Erik Hemberg, Una-May O'Reilly, Pablo García-Sánchez |
EvoApplications@EvoStar | 3 |
| 2023 | Analysis of a Pairwise Dominance Coevolutionary Algorithm And DefendItabstractWhile competitive coevolutionary algorithms are ideally suited to model adversarial dynamics, their complexity makes it difficult to understand what is happening when they execute. To achieve better clarity, we introduce a game named DefendIt and explore a previously developed pairwise dominance coevolutionary algorithm named PDCoEA. We devise a methodology for consistent algorithm comparison, then use it to empirically study the impact of population size, the impact of relative budget limits between the defender and attacker, and the impact of mutation rates on the dynamics and payoffs. Our methodology provides reliable comparisons and records of run and multi-run dynamics. Our supplementary material also offers enticing and detailed animations of a pair of players' game moves over the course of a game of millions of moves matched to the same run's populations' payoffs. Per Kristian Lehre, Mario Alejandro Hevia Fajardo, Jamal Toutouh, Erik Hemberg, Una-May O'Reilly |
GECCO | 4 |
| 2023 | Semi-Supervised Learning with Coevolutionary Generative Adversarial NetworksabstractIt can be expensive to label images for classification. Good classifiers or high-quality images can be trained on unlabeled data with Generative Adversarial Network (GAN) methods. We use coevolutionary algorithms with Semi-Supervised GANs (SSL-GANs) that work with a few labeled and some more unlabeled images to train both a good classifier and a high-quality image generator. A spatial coevolutionary algorithm introduces diversity into the GAN training. We use a two-dimensional grid of GANs to gain discriminator loss diversity with a distributed cell-level coevolutionary algorithm. The GAN components are exchanged between neighboring cells based on performance and population-based hyperparameter tuning. The approach is demonstrated on two separate benchmark datasets, and with only a few labels, we simultaneously achieve good classification accuracy and high generated image quality score. In addition, the generated image quality and classification accuracy are competitive to state-of-the-art methods. Jamal Toutouh, Subhash Nalluru, Erik Hemberg, Una-May O'Reilly |
GECCO | 3 |
| 2023 | Investigating Student's Problem-solving Approaches in MOOCs using Natural Language ProcessingabstractProblem-solving approaches are an essential part of learning. Knowing how students approach solving problems can help instructors improve their instructional designs and effectively guide the learning process of students. We propose a natural language processing (NLP) driven method to capture online learners’ problem-solving approaches at scale while using Massive Open Online Courses (MOOCs) as a learning platform. We employ an online survey to gather data, NLP techniques, and existing educational theories to investigate this in the lens of both computer science and education. The paper shows how NLP techniques, i.e. preprocessing, topic modeling, and text summarization, must be tuned to extract information from a large-scale text corpus. The proposed method discovered 18 problem-solving approaches from the text data, such as using pen and paper, peer learning, trial and error, etc. We also observed topics that appear over the years, such as clarifying code logic, watching videos, etc. We observed that students heavily rely on "tools" for solving programming problems and can expect that such selection of methods can vary depending on the type of task. ByeongJo Kong, Erik Hemberg, Ana Bell, Una-May O'Reilly |
LAK | 2 |
| 2022 | Exploiting Knowledge from Code to Guide Program Search
Dirk Schweim, Erik Hemberg, Dominik Sobania, Una-May O'Reilly |
EuroGP | 2 |
| 2022 | Synthesizing Programs from Program Pieces Using Genetic Programming and Refinement Type Checking
Sabrina Tseng, Erik Hemberg, Una-May O'Reilly |
EuroGP | 2 |
| 2022 | Coevolutionary generative adversarial networks for medical image augumentation at scaleabstractMedical image processing can lack images for diagnosis. Generative Adversarial Networks (GANs) provide a method to train generative models for data augmentation. Synthesized images can be used to improve the robustness of computer-aided diagnosis systems. However, GANs are difficult to train due to unstable training dynamics that may arise during the learning process, e.g., mode collapse and vanishing gradients. This paper focuses on Lipizzaner, a GAN training framework that combines spatial coevolution with gradient-based learning, which has been used to mitigate GAN training pathologies. Lipizzaner improves performance by taking advantage of its distributed nature and running at scale. Thus, the Lipizzaner algorithm and implementation robustness can be scaled to high-performance computing (HPC) systems to provide more accurate generative models. We address medical imaging data augmentation to create chest X-Ray images by using Lipizzaner on the HPC infrastructure provided by Oak Ridge National Labs' Summit Supercomputer. The experimental analysis shows improved performance by increasing the scale of the Lipizzaner GAN training. We also demonstrate that distributed coevolutionary learning improves performance even when using suboptimal neural network architectures due to hardware constraints. Diana Flores, Erik Hemberg, Jamal Toutouh, Una-May O'Reilly |
GECCO | 2 |
| 2022 | Analyzing multi-agent reinforcement learning and coevolution in cybersecurityabstractCybersecurity simulations can offer deep insights into the behavior of agents in the battle to secure computer systems. We build on existing work modeling the competition between an attacker and defender on a network architecture in a zero-sum game using a graph database linking cybersecurity attack patterns, vulnerabilities, and software. We apply coevolution to this challenging environment, and in a novel modeling approach for this problem, interpret each population as a distribution over fixed strategies to form a mixed strategy Nash equilibrium. We compare the results to solutions generated by multi-agent reinforcement learning and show that evolutionary methods demonstrate a considerable degree of robustness to parameter misspecification in this environment. Matthew J. Turner 0001, Erik Hemberg, Una-May O'Reilly |
GECCO | 2 |
| 2021 | Getting a Head Start on Program Synthesis with Genetic Programming
Jordan Wick, Erik Hemberg, Una-May O'Reilly |
EuroGP | 2 |
| 2021 | Coevolutionary modeling of cyber attack patterns and mitigations using public datasetsabstractThe evolution of advanced persistent threats (APTs) spurs us to explore computational models of coevolutionary dynamics arising from efforts to secure cyber systems from them. In a first for evolutionary algorithms, we incorporate known threats and vulnerabilities into a stylized "competition" that pits cyber attack patterns against mitigations. Variations of attack patterns that are drawn from the public CAPEC catalog offering Common Attack Pattern Enumeration and Classifications. Mitigations take two forms: software updates or monitoring, and the software that is mitigated is identified by drawing from the public CVE dictionary of Common Vulnerabilities and Exposures. In another first, we quantify the outcome of a competition by incorporating the public Common Vulnerability Scoring System - CVSS. We align three abstract models of population-level dynamics where APTs interact with defenses with three competitive, coevolutionary algorithm variants that use the competition. A comparative study shows that the way a defensive population preferentially acts, e.g. shifting to mitigating recent attack patterns, results in different evolutionary outcomes, expressed as different dominant attack patterns and mitigations. Michal Shlapentokh-Rothman, Jonathan Kelly, Avital Baral, Erik Hemberg, Una-May O'Reilly |
GECCO | 4 |
| 2021 | Analyzing Student Reflection Sentiments and Problem-Solving Procedures in MOOCsabstractStudent reflection is thought to be an important part of retaining and understanding knowledge gained in a course. Using natural language processing, we analyze and interpret student reflections from Massive Open Online Courses (MOOCs) to understand the students' sentiments and problem-solving procedures. The reflections are free text responses to questions from MIT 6.00.1x, an introductory programming MOOC. We compare different sentiment analysis methods, and conclude that the best-performing methods can robustly classify sentiment of student responses. In addition, we develop methods to analyze student problem-solving procedures using sentence parsing and topic modeling. We find our method can distinguish some common problem-solving procedures such as utilizing course resources. Alexander Shashkov, Robert Gold, Erik Hemberg, ByeongJo Kong, Ana Bell, Una-May O'Reilly |
L@S | 3 |
| 2021 | Spatial Coevolution for Generative Adversarial Network TrainingabstractGenerative Adversarial Networks (GANs) are difficult to train because of pathologies such as mode and discriminator collapse. Similar pathologies have been studied and addressed in competitive evolutionary computation by increased diversity. We study a system, Lipizzaner, that combines spatial coevolution with gradient-based learning to improve the robustness and scalability of GAN training. We study different features of Lipizzaner’s evolutionary computation methodology. Our ablation experiments determine that communication, selection, parameter optimization, and ensemble optimization each, as well as in combination, play critical roles. Lipizzaner succumbs less frequently to critical collapses and, as a side benefit, demonstrates improved performance. In addition, we show a GAN-training feature of Lipizzaner: the ability to train simultaneously with different loss functions in the gradient descent parameter learning framework of each GAN at each cell. We use an image generation problem to show that different loss function combinations result in models with better accuracy and more diversity in comparison to other existing evolutionary GAN models. Finally, Lipizzaner with multiple loss function options promotes the best model diversity while requiring a large grid size for adequate accuracy. Erik Hemberg, Jamal Toutouh, Abdullah Al-Dujaili, Tom Schmiedlechner, Una-May O'Reilly |
ACM Trans. Evol. Learn. Optim. | 1 |
| 2020 | Re-purposing heterogeneous generative ensembles with evolutionary computationabstractGenerative Adversarial Networks (GANs) are popular tools for generative modeling. The dynamics of their adversarial learning give rise to convergence pathologies during training such as mode and discriminator collapse. In machine learning, ensembles of predictors demonstrate better results than a single predictor for many tasks. In this study, we apply two evolutionary algorithms (EAs) to create ensembles to re-purpose generative models, i.e., given a set of heterogeneous generators that were optimized for one objective (e.g., minimize Fréchet Inception Distance), create ensembles of them for optimizing a different objective (e.g., maximize the diversity of the generated samples). The first method is restricted by the exact size of the ensemble and the second method only restricts the upper bound of the ensemble size. Experimental analysis on the MNIST image benchmark demonstrates that both EA ensembles creation methods can re-purpose the models, without reducing their original functionality. The EA-based demonstrate significantly better performance compared to other heuristic-based methods. When comparing both evolutionary, the one with only an upper size bound on the ensemble size is the best. Jamal Toutouh, Erik Hemberg, Una-May O'Reilly |
GECCO | 2 |
| 2020 | Analyzing Pre-Existing Knowledge and Performance in a Programming MOOCabstractMassive Open Online Courses (MOOCs) are accessible to anyone with a device that can connect to the internet. MOOCs aim to increase the accessibility of higher-level knowledge and skills, such as programming. To understand how students are performing and struggling in the course, we investigate a popular MITx MOOC that teaches introductory programming. We look at problem set questions and examine students with different levels of pre-existing knowledge. Specifically, we study the number of attempts of each group per question and the mean final accuracy of each group per question. We find that for nearly all questions, students with no programming experience struggle more than students with prior programming experience. Moreover, we observe a potential turning point in the course where students of all experience levels begin to struggle. Our findings both show that two groups of MOOC students perform differently and inform question design in MOOCs by demonstrating which question types are particularly arduous. Hannah Burd, Ana Bell, Erik Hemberg, Una-May O'Reilly |
L@S | 3 |
| 2020 | Analyzing K-12 Blended MOOC Learning BehaviorsabstractWe investigate student learning behaviors in a Massive Open Online Course with in-person components. Our goal is to improve the design of the course through learning analytics. The programming language taught, App Inventor, is a drag-and-drop language to create Android applications. We visualize and quantify student behaviors such as automatic and manual saving of code, video sections viewed, and the various forms of knowledge required to understand the course material. It appears students are less likely to go from course material that teaches procedures to other material that teaches procedures than we would expect, and rarely review previous topics covered in the course. We also find students tend to save marginally less at the beginning and end of sessions. However, since the data set is small, our conclusions are limited. Robert Gold, Erik Hemberg, Una-May O'Reilly |
L@S | 2 |
| 2020 | Understanding Learner Behavior Through Learning Design Informed Learning AnalyticsabstractA goal of learning analytics is to inform and improve learning design. Previous studies have attempted to interpret learners' clickstream data based on learning science theories. Many of these interpretations are made without reference to the specific learning designs of the courses being analyzed. Here, we report on a learning design informed analytics exploration of an introductory MOOC on Computer Science and Python programming. The learning resources (videos) and practice resources (short exercises and problem sets) are analyzed according to the knowledge types and cognitive process levels respectively, both based on a revised Bloom's Taxonomy. A heat map visualization of the access intensity on a learner resource access transition matrix and social network analysis are used to analyze learners' behavior with respect to the different resource categories. The results show distinctively different patterns of access between groups of students with different course performance and different academic backgrounds. Leming Liang, Nancy Law, Erik Hemberg, Una-May O'Reilly |
L@S | 4 |
| 2020 | Analyzing the Components of Distributed Coevolutionary GAN Training
Jamal Toutouh, Erik Hemberg, Una-May O'Reilly |
PPSN (1) | 2 |
| 2019 | Improving Genetic Programming with Novel Exploration - Exploitation Control
Jonathan Kelly, Erik Hemberg, Una-May O'Reilly |
EuroGP | 2 |
| 2019 | On domain knowledge and novelty to improve program synthesis performance with grammatical evolutionabstractProgrammers solve coding problems with the support of both programming and problem specific knowledge. They integrate this domain knowledge to reason by computational abstraction. Correct and readable code arises from sound abstractions and problem solving. We attempt to transfer insights from such human expertise to genetic programming (GP) for solving automatic program synthesis. We draw upon manual and non-GP Artificial Intelligence methods to extract knowledge from synthesis problem definitions to guide the construction of the grammar that Grammatical Evolution uses and to supplement its fitness function. We examine the impact of using such knowledge on 21 problems from the GP program synthesis benchmark suite. Additionally, we investigate the compounding impact of this knowledge and novelty search. The resulting approaches exhibit improvements in accuracy on a majority of problems in the field's benchmark suite of program synthesis problems. Erik Hemberg, Jonathan Kelly, Una-May O'Reilly |
GECCO | 1 |
| 2019 | Spatial evolutionary generative adversarial networksabstractGenerative adversary networks (GANs) suffer from training pathologies such as instability and mode collapse. These pathologies mainly arise from a lack of diversity in their adversarial interactions. Evolutionary generative adversarial networks apply the principles of evolutionary computation to mitigate these problems. We hybridize two of these approaches that promote training diversity. One, E-GAN, at each batch, injects mutation diversity by training the (replicated) generator with three independent objective functions then selecting the resulting best performing generator for the next batch. The other, Lipizzaner, injects population diversity by training a two-dimensional grid of GANs with a distributed evolutionary algorithm that includes neighbor exchanges of additional training adversaries, performance based selection and population-based hyper-parameter tuning. We propose to combine mutation and population approaches to diversity improvement. We contribute a superior evolutionary GANs training method, Mustangs, that eliminates the single loss function used across Lipizzaner's grid. Instead, each training round, a loss function is selected with equal probability, from among the three E-GAN uses. Experimental analyses on standard benchmarks, MNIST and CelebA, demonstrate that Mustangs provides a statistically faster training method resulting in more accurate networks. Jamal Toutouh, Erik Hemberg, Una-May O'Reilly |
GECCO | 2 |
| 2019 | Adversarially Adapting Deceptive Views and Reconnaissance Scans on a Software Defined Network
Jonathan Kelly, Michael DeLaus, Erik Hemberg, Una-May O'Reilly |
IM | 3 |
| 2019 | Transfer Learning using Representation Learning in Massive Open Online CoursesabstractIn a Massive Open Online Course (MOOC), predictive models of student behavior can support multiple aspects of learning, including instructor feedback and timely intervention. Ongoing courses, when the student outcomes are yet unknown, must rely on models trained from the historical data of previously offered courses. It is possible to transfer models, but they often have poor prediction performance. One reason is features that inadequately represent predictive attributes common to both courses. We present an automated transductive transfer learning approach that addresses this issue. It relies on problem-agnostic, temporal organization of the MOOC clickstream data, where, for each student, for multiple courses, a set of specific MOOC event types is expressed for each time unit. It consists of two alternative transfer methods based on representation learning with auto-encoders: a passive approach using transductive principal component analysis and an active approach that uses a correlation alignment loss term. With these methods, we investigate the transferability of dropout prediction across similar and dissimilar MOOCs and compare with known methods. Results show improved model transferability and suggest that the methods are capable of automatically learning a feature representation that expresses common predictive characteristics of MOOCs. Mucong Ding, Yanbang Wang, Erik Hemberg, Una-May O'Reilly |
LAK | 3 |
| 2019 | Using Detailed Access Trajectories for Learning Behavior AnalysisabstractStudent learning activity in MOOCs can be viewed from multiple perspectives. We present a new organization of MOOC learner activity data at a resolution that is in between the fine granularity of the clickstream and coarse organizations that count activities, aggregate students or use long duration time units. A detailed access trajectory (DAT) consists of binary values and is two dimensional with one axis that is a time series, and the other that is a chronologically ordered list of a MOOC component type's instances, videos in instructional order, for example. Most popular MOOC platforms generate data that can be organized as detailed access trajectories (DATs). We explore the value of DATs by conducting four empirical mini-studies. Our studies suggest DATs contain rich information about students' learning behaviors and facilitate MOOC learning analyses. Yanbang Wang, Nancy Law, Erik Hemberg, Una-May O'Reilly |
LAK | 3 |
| 2019 | Student Code Trajectories in an Introductory Programming MOOCabstractIn classrooms, instructors teaching students how to code have the ability to monitor progress and provide feedback through regular interaction. There is generally no analogous tracing of learning progression in programming MOOCs, hindering the ability of MOOC platforms to provide automated feedback at scale. We explore features for every certified student's history of code submissions to specific problems in a programming MOOC and measure similarity to sample solutions. We seek to understand whether students who succeed in the course reach solutions similar to these instructor-intended sample solutions, in terms of the concepts and mechanisms they contain. Furthermore, do students learn to conform to instructor expectations as the course progresses, and does prior experience have correlations with student behavior? We also explore what feature representations are sufficient for code submission history, since they are directly applicable to the development of automated tutors for progress tracking. Ayesha Bajwa, Erik Hemberg, Ana Bell, Una-May O'Reilly |
L@S | 2 |
| 2019 | Investigating Learning Design Categorization and Learning Behaviour in Computational MOOCSabstractWe investigate learner efficiency by categorizing a computational MOOC and analyzing user behavior data from a learning design point of view. Learning design is important both when designing courses as well as studying them. Learning behavior can be observed from the MOOC platform data. For this study we ask two learning designer experts to categorize a course on MITx: "6.00.1x Introduction to Computer Science and Programming Using Python". We use these categorizations to investigate relationships with learning behavior by analyzing the MOOC platform data. Our study verifies that learning design can be correlated to learning behavior, e.g. students exhibit a pattern of behavior associated to a component's difficulty and category. Sagar Biswas, Nancy Law, Erik Hemberg, Una-May O'Reilly |
L@S | 3 |
| 2019 | On the Influence of Grades on Learning Behavior of Students in MOOCsabstractMOOCs (Massive Open Online Courses) frequently use grades to calculate whether a student passes the course. To better understand how student behavior is influenced by grade feedback, we conduct a study on the changes of certified students' behavior before and after they have received their grade. We use observational student data from two MITx MOOCs to examine student behavior before and after a grade is released and calculate the difference (the delta-activity). We then analyze the changes in the delta-activity distributions across all graded assignments a we observe that the variation in delta-activity decreases as grade decreases, with students who have the lowest grade exhibiting little or no change in weekly activity. This trend persists throughout each course, in all course offerings, suggesting that a change in grade does not correlate with a change in the behavior of certified MOOC students. Erik Hemberg, Una-May O'Reilly |
L@S | 2 |
| 2016 | Discrete Planar Truss Optimization by Node Position Variation Using Grammatical EvolutionabstractThe majority of existing discrete truss optimization methods focus primarily on optimizing global truss topology using a ground structure approach, in which all possible node and beam locations are specifieda priori. The ground structure discrete optimization method has been shown to be restrictive as it limits derivable solutions to what is explicitly defined. Greater representational freedom can improve performance. In this paper, grammatical evolution is applied. It can represent a variable number of nodes and their locations on a continuum. A novel method of connecting evolved nodes using a Delaunay triangulation algorithm shows that fully triangulated, kinematically stable structures can be generated. Discrete beam-truss structures can be optimized without the need for any information about the desired form of the solution other than the design envelope. Our technique is compared to existing discrete optimization techniques, and notable savings in structure self-weight are demonstrated. In particular, our new method can produce results superior to those reported in the literature in cases in which the problem is ill-defined and the structure of the solution is not knowna priori. Michael Fenton, Ciaran McNally, Jonathan Byrne, Erik Hemberg, James McDermott, Michael O'Neill 0001 |
IEEE Trans. Evol. Comput. | 4 |
| 2015 | Tax non-compliance detection using co-evolution of tax evasion risk and audit likelihoodabstractWe detect tax law abuse by simulating the co-evolution of tax evasion schemes and their discovery through audits. Tax evasion accounts for billions of dollars of lost income each year. When the IRS pursues a tax evasion scheme and changes the tax law or audit procedures, the tax evasion schemes evolve and change into undetectable forms. The arms race between tax evasion schemes and tax authorities presents a serious compliance challenge. Tax evasion schemes are sequences of transactions where each transaction is individually compliant. However, when all transactions are combined they have no other purpose than to evade tax and are thus non-compliant. Our method consists of an ownership network and a sequence of transactions, which outputs the likelihood of conducting an audit, and requires no prior tax return or audit data. We adjust audit procedures for a new generation of evolved tax evasion schemes by simulating the gradual change of tax evasion schemes and audit points, i.e. methods used for detecting non-compliance. Additionally, we identify, for a given audit scoring procedure, which tax evasion schemes will likely escape auditing. The approach is demonstrated in the context of partnership tax law and the Installment Bogus Optional Basis tax evasion scheme. The experiments show the oscillatory behavior of a co-adapting system and that it can model the co-evolution of tax evasion schemes and their detection. Erik Hemberg, Jacob B. Rosen, Geoff Warner, Sanith Wijesinghe, Una-May O'Reilly |
ICAIL | 1 |
| 2015 | Optimising complex pylon structures with grammatical evolution
Jonathan Byrne, Michael Fenton, Erik Hemberg, James McDermott, Michael O'Neill 0001 |
Inf. Sci. | 3 |
| 2013 | Understanding Expansion Order and Phenotypic Connectivity in πGE
David Fagan, Erik Hemberg, Michael O'Neill 0001, Seán McGarraghy |
EuroGP | 2 |
| 2013 | Introducing graphical models to analyze genetic programming dynamicsabstractWe propose graphical models as a new means of understanding genetic programming dynamics. Herein, we describe how to build an unbiased graphical model from a population of genetic programming trees. Graphical models both express information about the conditional dependency relations among a set of random variables and they support probabilistic inference regarding the likelihood of a random variable's outcome. We focus on the former information: by their structure, graphical models reveal structural dependencies between the nodes of genetic programming trees. We identify graphical model properties of potential interest in this regard - edge quantity and dependency among nodes expressed in terms of family relations. Using a simple symbolic regression problem we generate a graphical model of the population each generation. Then we interpret the graphical models with respect to conventional knowledge about the influence of subtree crossover and mutation upon tree structure. Erik Hemberg, Constantin Berzan, Kalyan Veeramachaneni, Una-May O'Reilly |
FOGA | 1 |
| 2012 | Grammar Bias and Initialisation in Grammar Based Genetic Programming
Eoin Murphy, Erik Hemberg, Miguel Nicolau, Michael O'Neill 0001, Anthony Brabazon |
EuroGP | 2 |
| 2012 | An investigation of local patterns for estimation of distribution genetic programmingabstractWe present an improved estimation of distribution (EDA) genetic programming (GP) algorithm which does not rely upon a prototype tree. Instead of using a prototype tree, Operator-Free Genetic Programming learns the distribution of ancestor node chains, "n-grams", in a fit fraction of each generation's population. It then uses this information, via sampling, to create trees for the next generation. Ancestral n-grams are used because an analysis of a GP run conducted by learning depth first graphical models for each generation indicated their emergence as substructures of conditional dependence. We are able to show that our algorithm, without an operator and a prototype tree, achieves, on average, performance close to conventional tree based crossover GP on the problem we study. Our approach sets a direction for pattern-based EDA GP which off ers better tractability and improvements over GP with operators or EDAs using prototype trees. Erik Hemberg, Kalyan Veeramachaneni, James McDermott, Constantin Berzan, Una-May O'Reilly |
GECCO | 1 |
| 2012 | Comparing methods for module identification in grammatical evolutionabstractModularity has been an important vein of research in evolutionary algorithms. Past research in evolutionary computation has shown that techniques able to decompose the benchmark problems examined in this work into smaller, more easily solved, sub-problems have an advantage over those which do not. This work describes and analyzes a number of approaches to discover sub-solutions (modules) in the grammatical evolution algorithm. Data from the experiments carried out show that particular approaches to identifying modules are better suited to certain problem types, at varying levels of difficulty. The results presented here show that some of these approaches are able to significantly outperform standard grammatical evolution and grammatical evolution using automatically defined functions on a subset of the problems tested. The results also point to a number of possibilities for extending this work to further enhance approaches to modularity. John Mark Swafford, Miguel Nicolau, Erik Hemberg, Michael O'Neill 0001, Anthony Brabazon |
GECCO | 3 |
| 2012 | Evolving Femtocell Algorithms with Dynamic and Stationary Training Scenarios
Erik Hemberg, Lester T. W. Ho, Michael O'Neill 0001, Holger Claussen 0001 |
PPSN (2) | 1 |
| 2012 | Differential Gene Expression with Tree-Adjunct Grammars
Eoin Murphy, Miguel Nicolau, Erik Hemberg, Michael O'Neill 0001, Anthony Brabazon |
PPSN (1) | 3 |
| 2012 | Analyzing Module Usage in Grammatical Evolution
John Mark Swafford, Erik Hemberg, Michael O'Neill 0001, Anthony Brabazon |
PPSN (1) | 2 |
| 2011 | Investigation of the Performance of Different Mapping Orders for GE on the Max Problem
David Fagan, Miguel Nicolau, Erik Hemberg, Michael O'Neill 0001, Anthony Brabazon, Seán McGarraghy |
EuroGP | 3 |
| 2011 | Combining Structural Analysis and Multi-Objective Criteria for Evolutionary Architectural Design
Jonathan Byrne, Michael Fenton, Erik Hemberg, James McDermott, Michael O'Neill 0001, Elizabeth Shotton, Ciaran McNally |
EvoApplications (2) | 3 |
| 2011 | A non-destructive grammar modification approach to modularity in grammatical evolutionabstractModularity has proven to be an important aspect of evolutionary computation. This work is concerned with discovering and using modules in one form of grammar-based genetic programming, grammatical evolution (GE). Previous work has shown that simply adding modules to GE's grammar has the potential to disrupt fit individuals developed by evolution up to that point. This paper presents a solution to prevent the disturbance in fitness that can come with modifying GE's grammar with previously discovered modules. The results show an increase in performance from a previously examined grammar modification approach and also an increase in performance when compared to standard GE. John Mark Swafford, Erik Hemberg, Michael O'Neill 0001, Miguel Nicolau, Anthony Brabazon |
GECCO | 2 |
| 2009 | Analysis of constant creation techniques on the binomial-3 problem with grammatical evolutionabstractThis paper studies the difference between Persistent Random Constants (PRC) and digit concatenation as methods for generating constants. It has been shown that certain problems have different fitness landscapes depending on how they are represented, independent of changes to the combinatorial search space, thus changing problem difficulty. In this case we show that the method for generating the constants can also influence how hard the problem is for genetic programming. Jonathan Byrne, Michael O'Neill 0001, Erik Hemberg, Anthony Brabazon |
IEEE Congress on Evolutionary Computation | 3 |
| 2008 | Grammatical bias and building blocks in meta-grammar Grammatical EvolutionabstractThis paper describes and tests the utility of a meta grammar approach to grammatical evolution (GE). Rather than employing a fixed grammar as is the case with canonical GE, under a meta grammar approach the grammar that is used to specify the construction of a syntactically correct solution is itself allowed to evolve. The ability to evolve a grammar in the context of GE means that useful bias towards specific structures and solutions can be evolved and directly incorporated into the grammar during a run. This approach facilitates the evolution of modularity and reuse both on structural and symbol levels and consequently could enhance both the scalability of GE and its adaptive potential in dynamic environments. In this paper an analysis of the extent that building block structures created in the grammars are used in the solution is undertaken. It is demonstrated that building block structures are incorporated into the evolving grammars and solutions at a rate higher than would be expected by random search. Furthermore, the results indicate that grammar design can be an important factor in performance. Erik Hemberg, Michael O'Neill 0001, Anthony Brabazon |
IEEE Congress on Evolutionary Computation | 1 |
| 2008 | Subtree deactivation control with grammatical Genetic Programming in dynamic environmentsabstractWe investigate the usefulness of a subtree deactivation control mechanism which is open to evolutionary learning. It is hypothesised that this representation confers an adaptive advantage in dynamic environments over the standard sub-tree representation adopted in Genetic Programming. Results presented on benchmark dynamic problem instances provides evidence to support that such an adaptive advantage exists. Michael O'Neill 0001, Anthony Brabazon, Erik Hemberg |
IEEE Congress on Evolutionary Computation | 3 |
| 2008 | Altering Search Rates of the Meta and Solution Grammars in the mGGA
Erik Hemberg, Michael O'Neill 0001, Anthony Brabazon |
EuroGP | 1 |
| 2007 | A Grammatical Genetic Programming Approach to Modularity in Genetic Algorithms
Erik Hemberg, Conor Gilligan, Michael O'Neill 0001, Anthony Brabazon |
EuroGP | 1 |