VLDB 2026 Research / reviewers in the wild / expert
Herna L. Viktor
dblp:89/3149 · also Herna Lydia Viktor
· DBLP profile ↗
51ranked-venue papers
4as first author
16since 2021 · last 2026
0000-0003-1914-5077ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 35 · 4 first-author · 12 since 2021Databases, data management, data science and information retrieval · 17 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Computer networks · 1Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ReMemDiff: Multi-label Lifelong Machine Learning Using Deep Generative Replay
Mohammed Awal Kassim, Herna L. Viktor, Wojtek Michalowski |
ISMIS | 2 |
| 2025 | CUE-X: A Framework for the Automatic Evaluation of Clinical Usefulness of Explanations for the Multimorbidity Problem
Martin Michalowski, Szymon Wilk, Jenny M. Bauer, Marc Carrier, Herna L. Viktor, Wojtek Michalowski |
AIME (1) | 5 |
| 2025 | Protein structure generation using a variational autoencoder with Lévy noise and quantum graph transformerabstractProtein generation has a wide range of applications in the design of therapeutic antibodies and the creation of new drugs. Nevertheless, it is a challenging endeavour, largely due to the complexities intrinsic to protein structures and the constraints of current generative models. The complex three-dimensional structure of proteins and the vast number of potential conformations that they can adopt present significant challenges for sampling. This paper introduces a novel variational autoencoder based on Lévy noise and a quantum graph transformer attention mechanism, which enables a more effective exploration of the conformational space. The method was applied to two protein datasets, resulting in enhanced outcomes in terms of Fréchet distance by a factor of up to 168 in comparison to a variational autoencoder using Gaussian noise and a bilinear attention mechanism. Eric Paquet, Herna L. Viktor, Wojtek Michalowski |
IJCNN | 2 |
| 2025 | X-HEART: eXplainable heterogeneous log anomaly detection using robust transformers
Paul Kiyambu Mvula, Paula Branco, Guy-Vincent Jourdan, Herna L. Viktor |
Knowl. Inf. Syst. | 4 |
| 2024 | Enhancing Temporal Transformers for Financial Time Series via Local Surrogate Interpretability
Muhammed K. Olorunnimbe, Herna L. Viktor |
ISMIS | 2 |
| 2024 | Ensemble of temporal Transformers for financial time series
Muhammed K. Olorunnimbe, Herna L. Viktor |
J. Intell. Inf. Syst. | 2 |
| 2023 | HEART: Heterogeneous Log Anomaly Detection Using Robust Transformers
Paul Kiyambu Mvula, Paula Branco, Guy-Vincent Jourdan, Herna L. Viktor |
DS | 4 |
| 2023 | Measuring Improvement of F1-Scores in Detection of Self-Admitted Technical DebtabstractArtificial Intelligence and Machine Learning have witnessed rapid, significant improvements in Natural Language Processing (NLP) tasks. Utilizing Deep Learning, researchers have taken advantage of repository comments in Software Engineering to produce accurate methods for detecting Self-Admitted Technical Debt (SATD) from 20 open-source Java projects’ code. In this work, we improve SATD detection with a novel approach that leverages the Bidirectional Encoder Representations from Transformers (BERT) architecture. For comparison, we re-evaluated previous deep learning methods and applied stratified 10-fold cross-validation to report reliable F1-scores. We examine our model in both cross-project and intra-project contexts. For each context, we use re-sampling and duplication as augmentation strategies to account for data imbalance. We find that our trained BERT model improves over the best performance of all previous methods in 19 of the 20 projects in cross-project scenarios. However, the data augmentation techniques were not sufficient to overcome the lack of data present in the intra-project scenarios, and existing methods still perform better. Future research will look into ways to diversify SATD datasets in order to maximize the latent power in large BERT models. William Aiken, Paul Kiyambu Mvula, Paula Branco, Guy-Vincent Jourdan, Mehrdad Sabetzadeh, Herna L. Viktor |
TechDebt@ICSE | 6 |
| 2023 | DynaQ: online learning from imbalanced multi-class streams through dynamic samplingabstractAbstract Online supervised learning from fast-evolving data streams, particularly in domains such as health, the environment, and manufacturing, is a crucial research area. However, these domains often experience class imbalance, which can skew class distributions. It is essential for online learning algorithms to analyze large datasets in real-time while accurately modeling rare or infrequent classes that may appear in bursts. While methods have been proposed to handle binary class imbalance, there is a lack of attention to multi-class imbalanced settings with varying degrees of imbalance in evolving streams. In this paper, we present the Dynamic Queues (DynaQ) algorithm for online learning in multi-class imbalanced settings to fill this knowledge gap. Our approach utilizes a batch-based resampling method that creates an instance queue for each class to balance the number of instances. We maintain a queue threshold and remove older samples during training. Additionally, we dynamically oversample minority classes based on one of four rate parameters: recall, F1-score, $$\kappa _m$$ κ m , and Euclidean distance. Our learning algorithm consists of an ensemble that uses sliding windows and a soft voting schema while incorporating a drift detection mechanism. Our experimental results demonstrate the superiority of the DynaQ approach over state-of-the-art methods. Farnaz Sadeghi, Herna L. Viktor, Parsa Vafaie |
Appl. Intell. | 2 |
| 2023 | Deformable Protein Shape Classification Based on Deep Learning, and the Fractional Fokker-Planck and Kähler-Dirac EquationsabstractThe classification of deformable protein shapes, based solely on their macromolecular surfaces, is a challenging problem in protein-protein interaction prediction and protein design. Shape classification is made difficult by the fact that proteins are dynamic, flexible entities with high geometrical complexity. In this paper, we introduce a novel description for such deformable shapes. This description is based on the bifractional Fokker-Planck and Dirac-Kähler equations. These equations analyse and probe protein shapes in terms of a scalar, vectorial and non-commuting quaternionic field, allowing for a more comprehensive description of the protein shapes. An underlying non-Markovian Lévy random walk establishes geometrical relationships between distant regions while recalling previous analyses. Classification is performed with a multiobjective deep hierarchical pyramidal neural network, thus performing a multilevel analysis of the description. Our approach is applied to the SHREC'19 dataset for deformable protein shapes classification and to the SHREC'16 dataset for deformable partial shapes classification, demonstrating the effectiveness and generality of our approach. Eric Paquet, Herna L. Viktor, Kamel Madi, Junzheng Wu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | MQ-OFL: Multi-sensitive Queue-based Online Fair Learning
Farnaz Sadeghi, Herna L. Viktor |
DS | 2 |
| 2022 | Adversarial Robustness of Neural-Statistical Features in Detection of Generative TransformersabstractThe detection of computer-generated text is an area of rapidly increasing significance as nascent generative models allow for efficient creation of compelling human-like text, which may be abused for the purposes of spam, disinformation, phishing, or online influence campaigns. Past work has studied detection of current state-of-the-art models, but despite a developing threat landscape, there has been minimal analysis of the robustness of detection methods to adversarial attacks. To this end, we evaluate neural and non-neural approaches on their ability to detect computer-generated text, their robustness against text adversarial attacks, and the impact that successful adversarial attacks have on human judgement of text quality. We find that while statistical features underperform neural features, statistical features provide additional adversarial robustness that can be leveraged in ensemble detection models. In the process, we find that previously effective complex phrasal features for detection of computer-generated text hold little predictive power against contemporary generative models, and identify promising statistical features to use instead. Finally, we pioneer the usage of $\Delta$MAUVE as a proxy measure for human judgement of adversarial text quality. Evan Crothers, Nathalie Japkowicz, Herna L. Viktor, Paula Branco |
IJCNN | 3 |
| 2022 | Similarity Embedded Temporal Transformers: Enhancing Stock Predictions with Historically Similar Trends
Muhammed K. Olorunnimbe, Herna L. Viktor |
ISMIS | 2 |
| 2022 | COVID-19 malicious domain names classificationabstractDue to the rapid technological advances that have been made over the years, more people are changing their way of living from traditional ways of doing business to those featuring greater use of electronic resources. This transition has attracted (and continues to attract) the attention of cybercriminals, referred to in this article as "attackers", who make use of the structure of the Internet to commit cybercrimes, such as phishing, in order to trick users into revealing sensitive data, including personal information, banking and credit card details, IDs, passwords, and more important information via replicas of legitimate websites of trusted organizations. In our digital society, the COVID-19 pandemic represents an unprecedented situation. As a result, many individuals were left vulnerable to cyberattacks while attempting to gather credible information about this alarming situation. Unfortunately, by taking advantage of this situation, specific attacks associated with the pandemic dramatically increased. Regrettably, cyberattacks do not appear to be abating. For this reason, cyber-security corporations and researchers must constantly develop effective and innovative solutions to tackle this growing issue. Although several anti-phishing approaches are already in use, such as the use of blacklists, visuals, heuristics, and other protective solutions, they cannot efficiently prevent imminent phishing attacks. In this paper, we propose machine learning models that use a limited number of features to classify COVID-19-related domain names as either malicious or legitimate. Our primary results show that a small set of carefully extracted lexical features, from domain names, can allow models to yield high scores; additionally, the number of subdomain levels as a feature can have a large influence on the predictions. Paul Kiyambu Mvula, Paula Branco, Guy-Vincent Jourdan, Herna L. Viktor |
Expert Syst. Appl. | 4 |
| 2021 | Paying Attention: Using a Siamese Pyramid Network for the Prediction of Protein-Protein Interactions with Folding and Self-Binding Primary SequencesabstractProtein-protein interactions play a fundamental role in drug design, gene therapy and vaccine development. The study of protein-protein interactions relies heavily on complex and time-consuming experiments, which has a severe impact on research throughputs. Thus, it is important to provide the experimentalist with the most promising cases by screening rapidly through a very large number of potential candidates. We propose a new deep neural network architecture that allows the binding probability for two proteins to be predicted instantly based solely on their amino acid sequences. Subsequently, screenings are performed based on the binding probabilities. The novelty of our approach lies in the fact that we consider self-binding and folding amino acid sequences, rather than just looking at these sequences per se. Our novel Siamese Pyramid Network (SPNet) architecture is inspired by Feature Pyramid Networks and consists of a multi-level Siamese neural network with an attention mechanism and a multilevel, trainable binding probability prediction network. Our experimental evaluation is performed on a strict dataset and shows that SPNet outperforms the state-of-the-art architectures. In addition, we employ SPNet to find the proteins that are most likely to bind with the Covid-2019 spike, thus providing a small and potentially valuable set of candidates for a future therapeutic vaccine. Junzheng Wu, Eric Paquet, Herna L. Viktor, Wojtek Michalowski |
IJCNN | 3 |
| 2021 | Catering for unique tastes: Targeting grey-sheep users recommender systems through one-class machine learning
Rabaa Alabdulrahman, Herna L. Viktor |
Expert Syst. Appl. | 2 |
| 2019 | Improved Programming-Language Independent MapReduce on Shared-Memory Systems
Erik G. Selin, Herna L. Viktor |
DaWaK | 2 |
| 2019 | Sensory Data-Driven Modeling of Adversaries in Mobile Crowdsensing PlatformsabstractThe advent of Mobile crowdsensing (MCS) facilitates the adoption of ubiquitous sensing solutions in smart environments. Despite its benefits, MCS calls for proper security and trust solutions. Various threatening attacks, such as injection attacks, can compromise both the veracity and integrity of crowdsensed data. This work leverages adversarial machine learning to introduce a smart injection attacker model (SINAM) that may be used in the design of security solutions against injection attacks in MCS. SINAM has been validated during an authentic MCS campaign. Unlike most random data injection models, SINAM monitors data traffic in an online- learning manner, successfully injecting malicious data across multiple victims with near-perfect accuracy rates of 99%. SINAM uses accomplices within the sensing campaign to predict accurate injections based on both behavioral analysis and context similarities. Kyle Quintal, Ertugrul Kara, Murat Simsek, Burak Kantarci, Herna L. Viktor |
GLOBECOM | 5 |
| 2019 | Active Learning and Deep Learning for the Cold-Start Problem in Recommendation System: A Comparative Study
Rabaa Alabdulrahman, Herna L. Viktor, Eric Paquet |
IC3K | 2 |
| 2018 | HCC-Learn Framework for Hybrid Learning in Recommender Systems
Rabaa Alabdulrahman, Herna L. Viktor, Eric Paquet |
IC3K | 2 |
| 2018 | McDiarmid Drift Detection Methods for Evolving Data StreamsabstractIncreasingly, Internet of Things (IoT) domains, such as sensor networks, smart cities, and social networks, generate vast amounts of data. Such data are not only unbounded and rapidly evolving. Rather, the content thereof dynamically evolves over time, often in unforeseen ways. These variations are due to so-called concept drifts, caused by changes in the underlying data generation mechanisms. In a classification setting, concept drift causes the previously learned models to become inaccurate, unsafe and even unusable. Accordingly, concept drifts need to be detected, and handled, as soon as possible. In medical applications and emergency response settings, for example, change in behaviours should be detected in near real-time, to avoid potential loss of life. To this end, we introduce the McDiarmid Drift Detection Method (MDDM), which utilizes McDiarmid's inequality [1] in order to detect concept drift. The MDDM approach proceeds by sliding a window over prediction results, and associate window entries with weights. Higher weights are assigned to the most recent entries, in order to emphasize their importance. As instances are processed, the detection algorithm compares a weighted mean of elements inside the sliding window with the maximum weighted mean observed so far. A significant difference between the two weighted means, upper-bounded by the McDiarmid inequality, implies a concept drift. Our extensive experimentation against synthetic and real-world data streams show that our novel method outperforms the state-of-the-art. Specifically, MDDM yields shorter detection delays as well as lower false negative rates, while maintaining high classification accuracies. Ali Pesaranghader, Herna L. Viktor, Eric Paquet |
IJCNN | 2 |
| 2018 | SCUT-DS: Learning from Multi-class Imbalanced Canadian Weather Data
Olubukola M. Olaitan, Herna L. Viktor |
ISMIS | 2 |
| 2018 | Clustering in the Presence of Concept Drift
Richard Hugh Moulton, Herna L. Viktor, Nathalie Japkowicz, João Gama 0001 |
ECML/PKDD (1) | 2 |
| 2018 | Dynamic adaptation of online ensembles for drifting data streams
Muhammed K. Olorunnimbe, Herna L. Viktor, Eric Paquet |
J. Intell. Inf. Syst. | 2 |
| 2018 | Reservoir of diverse adaptive learners and stacking fast hoeffding drift detection methods for evolving data streams
Ali Pesaranghader, Herna L. Viktor, Eric Paquet |
Mach. Learn. | 2 |
| 2017 | Context-Based Abrupt Change Detection and Adaptation for Categorical Data Streams
Sarah D'Ettorre, Herna L. Viktor, Eric Paquet |
DS | 2 |
| 2017 | An Exploratory Study of Oral and Dental Health in CanadaabstractHealthcare practitioners agree that good oral health is a critical indicator of general health and wellness of a population. The lack of access to mandatory coverage for common issues such as cavities and non-surgical periodontal care often lead not only to medical problems, but also to loss of productivity. This trend is especially evident for older individuals and lower-income families. This paper discusses the results of our exploration of the annual Canadian Community Health Survey (CCHS), in order to further study the interplay between socio-economic factors and oral and dental health. To this end, we present the results when applying a number of machine learning algorithms to a CCHS data mart. Our results reaffirm that individuals' levels and sources of income are strong indicators of the number of dental visits per year. In addition, we found that younger adults and youth, who usually live in larger households, visit the dentist less frequently than all other survey respondents. Andrei Belcin, Sean Floyd, Areej Asiri, Herna L. Viktor |
ICMLA | 4 |
| 2016 | A Framework for Classification in Data Streams Using Multi-strategy Learning
Ali Pesaranghader, Herna L. Viktor, Eric Paquet |
DS | 2 |
| 2016 | Fast Hoeffding Drift Detection Method for Evolving Data Streams
Ali Pesaranghader, Herna L. Viktor |
ECML/PKDD (2) | 2 |
| 2015 | Multi-label Classification of Anemia PatientsabstractThis work examines the application of machine learning to an important area of medicine which aims to diagnose paediatric patients with β-thalassemia minor, iron deficiency anemia or the co-occurrence of these ailments. Iron deficiency anemia is a major cause of microcytic anemia and is considered an important task in global health. Whilst existing methods, based on linear equations, are proficient at distinguishing between the two classes of anemia, they fail to identify the co-occurrence of this issues. Machine learning algorithms, however, can induce non-linear decision boundaries that enable accurate classification within complex domains. Through a multi-label classification technique, known as problem transformations, we convert the learning task to one that is appropriate for machine learning and examine the effectiveness of machine learning algorithms on this domain. Our results show that machine learning classifiers produce good overall accuracy and are able to identify instances of the co-occurrence class unlike the existing methods. Colin Bellinger, Ali Amid, Nathalie Japkowicz, Herna L. Viktor |
ICMLA | 4 |
| 2015 | Tweets as a Vote: Exploring Political Sentiments on Twitter for Opinion Mining
Muhammed K. Olorunnimbe, Herna L. Viktor |
ISMIS | 2 |
| 2014 | Building Smart Cubes for Reliable and Faster Access to Data
Daniel K. Antwi, Herna L. Viktor |
DaWaK | 2 |
| 2014 | A Quantum Particle Swarm Optimization and Genetic Algorithm approach to the correspondence problemabstractFinding correspondences between deformable objects has wide application in many domains. In information retrieval, researchers may be interested in finding similar objects, while computer animation experts may be considering ways to morph shapes. The correspondence problem is especially challenging when the objects under consideration are suspect to non-rigid deformations, noise and/or distortions. In this paper, a novel method using Quantum Particle Swarm Optimization (QPSO) and Genetic Algorithms (GA) is presented to address this issue. In our QPSO-GA algorithm we formulate the problem of correspondence detection as an optimization problem over all possible mapping in between the geodesic distance matrices associated with two sets of point clouds. We proceed to identify the optimal mapping, by first applying Quantum Particle Swarm Optimization to the permutation matrices associated with their geodesic distance matrices and then employing Genetic Algorithms in order to guide the search. Experimental results suggest that our QPSO-GA algorithm is fast, scalable, and robust. Our method accurately identifies the correspondences between objects, even in the presence of noise and distortion. Hamid Hadavi, Herna L. Viktor, Eric Paquet |
INISTA | 2 |
| 2013 | An isometry-invariant spectral approach for protein-protein dockingabstractThe protein docking problem refers to the task of predicting the appropriate matching of one protein molecule (the receptor) to another (the ligand), when attempting to bind them to form a stable complex. Research shows that matching the three-dimensional geometric structures of proteins plays a key role in determining a so-called docking pair. However, the active sites which are responsible for the binding do not always present a rigid-body shape matching problem. Rather, they may undergo deformations when docking occurs, which complicates the process. To address this issue, we present an isometry-invariant and topologically robust partial shape descriptor method for finding complementary protein sites. Our method employs Heat Kernel Signature shape descriptors which are based on the diffusion of heat on surfaces. Our experimental results against the Protein-Protein Benchmark 4.0 demonstrate the viability of our approach. Dela De Youngster, Eric Paquet, Herna L. Viktor, Emil M. Petriu |
BIBE | 3 |
| 2013 | Isometrically Invariant Description of Deformable Objects Based on the Fractional Heat Equation
Eric Paquet, Herna L. Viktor |
CAIP (2) | 2 |
| 2013 | Artificial neural networks for predicting 3D protein shapes from amino acid sequencesabstractResearch has shown that the functionalities of proteins are largely influenced by their three dimensional (3D) shapes. This observation is especially relevant in drug design, where the knowledge of the 3D structure of a protein enables pharmacologists to select the best binding proteins when aiming to moderate functions. However, a relatively small number of 3D shapes are known. In contrast, amino acid sequences may be acquired through very efficient automated, high throughput experimental methods and the amino acid sequences of a vast number of proteins have therefore been identified. It follows that it is important to address this knowledge gap. To this end, this paper introduces an approach to predict the 3D shapes of proteins, utilizing feed-forward artificial neural networks. Our novel solution allows one to learn the representations of the 3D shape associated with a protein by starting directly from its amino acid sequence descriptors. Once a neural network is trained, our search engine enables one to retrieve the closest known 3D shape associated with an unknown, so-called query protein. We evaluate the performance of our approach against the Protein Data Bank (PDB), by considering proteins from a diverse set of families. Our results indicate that our system is able to accurately find the most similar protein structures for a wide variety of protein 3D shapes and diverse protein family sizes. Herna L. Viktor, Eric Paquet |
IJCNN | 1 |
| 2013 | Reducing the size of databases for multirelational classification: a subgraph-based approach
Herna L. Viktor, Eric Paquet |
J. Intell. Inf. Syst. | 2 |
| 2012 | Aggregation and privacy in multi-relational databasesabstractThe aim of privacy-preserving data mining is to construct highly accurate predictive models while not disclosing privacy information. Aggregation functions, such as sum and count are often used to pre-process the data prior to applying data mining techniques to relational databases. Often, it is implicitly assumed that the aggregated (or summarized) data are less likely to lead to privacy violations during data mining. This paper investigates this claim, within the relational database domain. We introduce the PBIRD (Privacy Breach Investigation in Relational Databases) methodology. Our experimental results show that aggregation potentially introduces new privacy violations. That is, potentially harmful attributes obtained with aggregation are often different from the ones obtained from non-aggregated databases. This indicates that, even when privacy is enforced on non-aggregated data, it is not automatically enforced on the corresponding aggregated data. Consequently, special care should be taken during model building in order to fully enforce privacy when the data are aggregated. Yasser Jafer, Herna L. Viktor, Eric Paquet |
PST | 2 |
| 2011 | Spectral Clustering: An Explorative Study of Proximity Measures
Nadia Farhanaz Azam, Herna L. Viktor |
IC3K | 2 |
| 2011 | Multimodal Representations, Indexing, Unexpectedness and Proteins
Eric Paquet, Herna L. Viktor |
IEA/AIE (1) | 2 |
| 2009 | Finding Clothing That Fit through Cluster Analysis and Objective Interestingness Measures
Isis Peña, Herna L. Viktor, Eric Paquet |
DaWaK | 2 |
| 2008 | Learning from Skewed Class Multi-relational Databases
Herna L. Viktor |
Fundam. Informaticae | 2 |
| 2008 | Multirelational classification: a multiple view approach
Herna L. Viktor |
Knowl. Inf. Syst. | 2 |
| 2008 | Capri/MR: exploring protein databases from a structural and physicochemical point of viewabstractWith the advent of high throughput systems to experimentally determine the three-dimensional (3-D) structure of proteins, molecular biologists are in urgent need of systems to automatically store, maintain and explore the vast structural databases that are thus being created. We have designed and implemented the Capri/MR system which makes it possible to identify families of protein structures, as contained in such very large 3-D protein structure databases. Our system is able to automatically index and search a database of proteins by three-dimensional shape, structural and/or physicochemical properties. For each of these diverse protein structure representations, we create a compact rotation and translation invariant index (or signature) which is placed in a database for future querying. A similarity search algorithm performs an exhaustive search against the entire database. Our search algorithm takes advantage of the compact signatures to rapidly find protein structures that are similar in 3-D shape and/or two-dimensional (2-D) properties. As a result, queries in our Capri/MR system run within a fraction of a second, and we are able to accurately group protein structures into the correct families, with very high precision and recall. In addition, our system dynamically processes new protein structures as they become available. We demonstrate the power of Capri/MR against the Protein Data Bank, which contains all known, experimentally determined, 3-D protein structures (48.000 as of January 2008). The main applications of our Capri/MR system lie in structural proteomics, protein evolution and mutation, as well as in drug design, in particular for studying the docking problem and the computer aided design of non-toxic drugs. Eric Paquet, Herna L. Viktor |
Proc. VLDB Endow. | 2 |
| 2007 | Pruning Relations for Substructure Discovery of Multi-relational Databases
Herna L. Viktor, Eric Paquet |
PKDD | 2 |
| 2006 | Multi-view ANNs for Multi-relational ClassificationabstractArtificial neural networks (ANNs) provide a general, effective and practical approach for learning complex target functions. However, ANNs are not suitable for handling relational data, where information about the target concept is distributed over multiple related relations. ANNs algorithms usually only explore one relation, the so-called target relation, thus excluding crucial knowledge embedded in the related so-called background relations. This paper introduces a new approach, the multiple view artificial neural networks (MVNNs) method, to address the need for bridging the gap between ANNs and relational databases. The MVNNs strategy, firstly, propagates essential information held in the target relation to all background relations. Subsequently, it exploits multiple ANNs, which explore the target concepts against the separate back-ground relations. Thirdly, it incorporates crucial background knowledge, as obtained by the ANNs, into a meta-learning mechanism to construct the final model. Our experiments on eight data sets show that the MVNNs method achieves promising results in terms of overall accuracy obtained, when compared with two other relational data mining algorithms. Herna L. Viktor |
IJCNN | 2 |
| 2006 | Mining relational data through correlation-based multiple view validationabstractCommercial relational databases currently store vast amounts of real-world data. The data within these relational repositories are represented by multiple relations, which are inter-connected by means of foreign key joins. The mining of such interrelated data poses a major challenge to the data mining community. Unfortunately, traditional data mining algorithms usually only explore one relation, the so-called target relation, thus excluding crucial knowledge embedded in the related so-called background relations. In this paper, we propose a novel approach for classifying relational such domains. This strategy employs multiple views to capture crucial information not only from the target relation, but also from related relations. This information is integrated into the relational mining process. The framework presented here, firstly, explore the relational domain to partition its features space into multiple subsets. Subsequently, these subsets are used to construct multiple uncorrelated views, based on a novel correlation-based view validation method, against the target concept. Finally, the knowledge possessed by multiple views are incorporated into a meta-learning mechanism to augment one another. Based on this framework, a wide range of conventional data mining methods can be applied to mine relational databases. Our experiments on benchmark real-world data sets show that the proposed method achieves promising results both in terms of overall accuracy obtained and run time, when compared with two other relational data mining approaches. Herna L. Viktor |
KDD | 2 |
| 2006 | Measuring to Fit: Virtual Tailoring Through Cluster Analysis and Classification
Herna L. Viktor, Eric Paquet |
PKDD | 1 |
| 2004 | Boosting with Data Generation: Improving the Classification of Hard to Learn Examples
Herna L. Viktor |
IEA/AIE | 2 |
| 2000 | Generating New Patterns for Information Gain and Improved Neural Network LearningabstractThis paper introduces an approach to generate new patterns for improved neural network training. The patterns are based on the information obtained by means of a rule extraction approach. In this way, the training process is re-iterated using the most informative patterns. The data generation process is further enhanced by incorporating the high quality rules obtained from a decision tree. Results indicate that the approach results in improved generalization, especially in difficult to learn domains. Herna L. Viktor |
IJCNN (4) | 1 |
| 1999 | Improved generalisation using cooperative learning and rule extractionabstractRule extraction from artificial neural networks (ANN) addresses the need for symbolic representations to explain an ANN's decisions. The ANNSER rule extraction approach extracts rules from feedforward ANN that are used for classification. The quality of the ANNSER rules reflects the strengths and weaknesses of the ANN and indicates those classes over which the ANN did not generalize well. This paper introduces a new approach to ANN training in which two or more ANNSER learners co-exist in a cooperative multiagent learning environment. The ANNSER learners cooperate by using one another's high quality rules to generate new training instances. In this way, the generalization of the ANN are improved, leading to a set of high quality rules that describe the knowledge embedded in the trained ANN. Herna L. Viktor, Ian Cloete |
IJCNN | 1 |