Herna L. Viktor

dblp:89/3149 · also Herna Lydia Viktor · DBLP profile ↗
← Back
51ranked-venue papers
4as first author
16since 2021 · last 2026
0000-0003-1914-5077ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 35 · 4 first-author · 12 since 2021Databases, data management, data science and information retrieval · 17 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Computer networks · 1Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 ReMemDiff: Multi-label Lifelong Machine Learning Using Deep Generative Replay
Mohammed Awal Kassim, Herna L. Viktor, Wojtek Michalowski
ISMIS2
2025 CUE-X: A Framework for the Automatic Evaluation of Clinical Usefulness of Explanations for the Multimorbidity Problem
Martin Michalowski, Szymon Wilk, Jenny M. Bauer, Marc Carrier, Herna L. Viktor, Wojtek Michalowski
AIME (1)5
2025 Protein structure generation using a variational autoencoder with Lévy noise and quantum graph transformer
abstract
Protein generation has a wide range of applications in the design of therapeutic antibodies and the creation of new drugs. Nevertheless, it is a challenging endeavour, largely due to the complexities intrinsic to protein structures and the constraints of current generative models. The complex three-dimensional structure of proteins and the vast number of potential conformations that they can adopt present significant challenges for sampling. This paper introduces a novel variational autoencoder based on Lévy noise and a quantum graph transformer attention mechanism, which enables a more effective exploration of the conformational space. The method was applied to two protein datasets, resulting in enhanced outcomes in terms of Fréchet distance by a factor of up to 168 in comparison to a variational autoencoder using Gaussian noise and a bilinear attention mechanism.
Eric Paquet, Herna L. Viktor, Wojtek Michalowski
IJCNN2
2025 X-HEART: eXplainable heterogeneous log anomaly detection using robust transformers
Paul Kiyambu Mvula, Paula Branco, Guy-Vincent Jourdan, Herna L. Viktor
Knowl. Inf. Syst.4
2024 Enhancing Temporal Transformers for Financial Time Series via Local Surrogate Interpretability
Muhammed K. Olorunnimbe, Herna L. Viktor
ISMIS2
2024 Ensemble of temporal Transformers for financial time series
Muhammed K. Olorunnimbe, Herna L. Viktor
J. Intell. Inf. Syst.2
2023 HEART: Heterogeneous Log Anomaly Detection Using Robust Transformers
Paul Kiyambu Mvula, Paula Branco, Guy-Vincent Jourdan, Herna L. Viktor
DS4
2023 Measuring Improvement of F1-Scores in Detection of Self-Admitted Technical Debt
abstract
Artificial Intelligence and Machine Learning have witnessed rapid, significant improvements in Natural Language Processing (NLP) tasks. Utilizing Deep Learning, researchers have taken advantage of repository comments in Software Engineering to produce accurate methods for detecting Self-Admitted Technical Debt (SATD) from 20 open-source Java projects’ code. In this work, we improve SATD detection with a novel approach that leverages the Bidirectional Encoder Representations from Transformers (BERT) architecture. For comparison, we re-evaluated previous deep learning methods and applied stratified 10-fold cross-validation to report reliable F1-scores. We examine our model in both cross-project and intra-project contexts. For each context, we use re-sampling and duplication as augmentation strategies to account for data imbalance. We find that our trained BERT model improves over the best performance of all previous methods in 19 of the 20 projects in cross-project scenarios. However, the data augmentation techniques were not sufficient to overcome the lack of data present in the intra-project scenarios, and existing methods still perform better. Future research will look into ways to diversify SATD datasets in order to maximize the latent power in large BERT models.
William Aiken, Paul Kiyambu Mvula, Paula Branco, Guy-Vincent Jourdan, Mehrdad Sabetzadeh, Herna L. Viktor
TechDebt@ICSE6
2023 DynaQ: online learning from imbalanced multi-class streams through dynamic sampling
abstract
Abstract Online supervised learning from fast-evolving data streams, particularly in domains such as health, the environment, and manufacturing, is a crucial research area. However, these domains often experience class imbalance, which can skew class distributions. It is essential for online learning algorithms to analyze large datasets in real-time while accurately modeling rare or infrequent classes that may appear in bursts. While methods have been proposed to handle binary class imbalance, there is a lack of attention to multi-class imbalanced settings with varying degrees of imbalance in evolving streams. In this paper, we present the Dynamic Queues (DynaQ) algorithm for online learning in multi-class imbalanced settings to fill this knowledge gap. Our approach utilizes a batch-based resampling method that creates an instance queue for each class to balance the number of instances. We maintain a queue threshold and remove older samples during training. Additionally, we dynamically oversample minority classes based on one of four rate parameters: recall, F1-score, $$\kappa _m$$ κ m , and Euclidean distance. Our learning algorithm consists of an ensemble that uses sliding windows and a soft voting schema while incorporating a drift detection mechanism. Our experimental results demonstrate the superiority of the DynaQ approach over state-of-the-art methods.
Farnaz Sadeghi, Herna L. Viktor, Parsa Vafaie
Appl. Intell.2
2023 Deformable Protein Shape Classification Based on Deep Learning, and the Fractional Fokker-Planck and Kähler-Dirac Equations
abstract
The classification of deformable protein shapes, based solely on their macromolecular surfaces, is a challenging problem in protein-protein interaction prediction and protein design. Shape classification is made difficult by the fact that proteins are dynamic, flexible entities with high geometrical complexity. In this paper, we introduce a novel description for such deformable shapes. This description is based on the bifractional Fokker-Planck and Dirac-Kähler equations. These equations analyse and probe protein shapes in terms of a scalar, vectorial and non-commuting quaternionic field, allowing for a more comprehensive description of the protein shapes. An underlying non-Markovian Lévy random walk establishes geometrical relationships between distant regions while recalling previous analyses. Classification is performed with a multiobjective deep hierarchical pyramidal neural network, thus performing a multilevel analysis of the description. Our approach is applied to the SHREC'19 dataset for deformable protein shapes classification and to the SHREC'16 dataset for deformable partial shapes classification, demonstrating the effectiveness and generality of our approach.
Eric Paquet, Herna L. Viktor, Kamel Madi, Junzheng Wu
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 MQ-OFL: Multi-sensitive Queue-based Online Fair Learning
Farnaz Sadeghi, Herna L. Viktor
DS2
2022 Adversarial Robustness of Neural-Statistical Features in Detection of Generative Transformers
abstract
The detection of computer-generated text is an area of rapidly increasing significance as nascent generative models allow for efficient creation of compelling human-like text, which may be abused for the purposes of spam, disinformation, phishing, or online influence campaigns. Past work has studied detection of current state-of-the-art models, but despite a developing threat landscape, there has been minimal analysis of the robustness of detection methods to adversarial attacks. To this end, we evaluate neural and non-neural approaches on their ability to detect computer-generated text, their robustness against text adversarial attacks, and the impact that successful adversarial attacks have on human judgement of text quality. We find that while statistical features underperform neural features, statistical features provide additional adversarial robustness that can be leveraged in ensemble detection models. In the process, we find that previously effective complex phrasal features for detection of computer-generated text hold little predictive power against contemporary generative models, and identify promising statistical features to use instead. Finally, we pioneer the usage of $\Delta$MAUVE as a proxy measure for human judgement of adversarial text quality.
Evan Crothers, Nathalie Japkowicz, Herna L. Viktor, Paula Branco
IJCNN3
2022 Similarity Embedded Temporal Transformers: Enhancing Stock Predictions with Historically Similar Trends
Muhammed K. Olorunnimbe, Herna L. Viktor
ISMIS2
2022 COVID-19 malicious domain names classification
abstract
Due to the rapid technological advances that have been made over the years, more people are changing their way of living from traditional ways of doing business to those featuring greater use of electronic resources. This transition has attracted (and continues to attract) the attention of cybercriminals, referred to in this article as "attackers", who make use of the structure of the Internet to commit cybercrimes, such as phishing, in order to trick users into revealing sensitive data, including personal information, banking and credit card details, IDs, passwords, and more important information via replicas of legitimate websites of trusted organizations. In our digital society, the COVID-19 pandemic represents an unprecedented situation. As a result, many individuals were left vulnerable to cyberattacks while attempting to gather credible information about this alarming situation. Unfortunately, by taking advantage of this situation, specific attacks associated with the pandemic dramatically increased. Regrettably, cyberattacks do not appear to be abating. For this reason, cyber-security corporations and researchers must constantly develop effective and innovative solutions to tackle this growing issue. Although several anti-phishing approaches are already in use, such as the use of blacklists, visuals, heuristics, and other protective solutions, they cannot efficiently prevent imminent phishing attacks. In this paper, we propose machine learning models that use a limited number of features to classify COVID-19-related domain names as either malicious or legitimate. Our primary results show that a small set of carefully extracted lexical features, from domain names, can allow models to yield high scores; additionally, the number of subdomain levels as a feature can have a large influence on the predictions.
Paul Kiyambu Mvula, Paula Branco, Guy-Vincent Jourdan, Herna L. Viktor
Expert Syst. Appl.4
2021 Paying Attention: Using a Siamese Pyramid Network for the Prediction of Protein-Protein Interactions with Folding and Self-Binding Primary Sequences
abstract
Protein-protein interactions play a fundamental role in drug design, gene therapy and vaccine development. The study of protein-protein interactions relies heavily on complex and time-consuming experiments, which has a severe impact on research throughputs. Thus, it is important to provide the experimentalist with the most promising cases by screening rapidly through a very large number of potential candidates. We propose a new deep neural network architecture that allows the binding probability for two proteins to be predicted instantly based solely on their amino acid sequences. Subsequently, screenings are performed based on the binding probabilities. The novelty of our approach lies in the fact that we consider self-binding and folding amino acid sequences, rather than just looking at these sequences per se. Our novel Siamese Pyramid Network (SPNet) architecture is inspired by Feature Pyramid Networks and consists of a multi-level Siamese neural network with an attention mechanism and a multilevel, trainable binding probability prediction network. Our experimental evaluation is performed on a strict dataset and shows that SPNet outperforms the state-of-the-art architectures. In addition, we employ SPNet to find the proteins that are most likely to bind with the Covid-2019 spike, thus providing a small and potentially valuable set of candidates for a future therapeutic vaccine.
Junzheng Wu, Eric Paquet, Herna L. Viktor, Wojtek Michalowski
IJCNN3
2021 Catering for unique tastes: Targeting grey-sheep users recommender systems through one-class machine learning
Rabaa Alabdulrahman, Herna L. Viktor
Expert Syst. Appl.2
2019 Improved Programming-Language Independent MapReduce on Shared-Memory Systems
Erik G. Selin, Herna L. Viktor
DaWaK2
2019 Sensory Data-Driven Modeling of Adversaries in Mobile Crowdsensing Platforms
abstract
The advent of Mobile crowdsensing (MCS) facilitates the adoption of ubiquitous sensing solutions in smart environments. Despite its benefits, MCS calls for proper security and trust solutions. Various threatening attacks, such as injection attacks, can compromise both the veracity and integrity of crowdsensed data. This work leverages adversarial machine learning to introduce a smart injection attacker model (SINAM) that may be used in the design of security solutions against injection attacks in MCS. SINAM has been validated during an authentic MCS campaign. Unlike most random data injection models, SINAM monitors data traffic in an online- learning manner, successfully injecting malicious data across multiple victims with near-perfect accuracy rates of 99%. SINAM uses accomplices within the sensing campaign to predict accurate injections based on both behavioral analysis and context similarities.
Kyle Quintal, Ertugrul Kara, Murat Simsek, Burak Kantarci, Herna L. Viktor
GLOBECOM5
2019 Active Learning and Deep Learning for the Cold-Start Problem in Recommendation System: A Comparative Study
Rabaa Alabdulrahman, Herna L. Viktor, Eric Paquet
IC3K2
2018 HCC-Learn Framework for Hybrid Learning in Recommender Systems
Rabaa Alabdulrahman, Herna L. Viktor, Eric Paquet
IC3K2
2018 McDiarmid Drift Detection Methods for Evolving Data Streams
abstract
Increasingly, Internet of Things (IoT) domains, such as sensor networks, smart cities, and social networks, generate vast amounts of data. Such data are not only unbounded and rapidly evolving. Rather, the content thereof dynamically evolves over time, often in unforeseen ways. These variations are due to so-called concept drifts, caused by changes in the underlying data generation mechanisms. In a classification setting, concept drift causes the previously learned models to become inaccurate, unsafe and even unusable. Accordingly, concept drifts need to be detected, and handled, as soon as possible. In medical applications and emergency response settings, for example, change in behaviours should be detected in near real-time, to avoid potential loss of life. To this end, we introduce the McDiarmid Drift Detection Method (MDDM), which utilizes McDiarmid's inequality [1] in order to detect concept drift. The MDDM approach proceeds by sliding a window over prediction results, and associate window entries with weights. Higher weights are assigned to the most recent entries, in order to emphasize their importance. As instances are processed, the detection algorithm compares a weighted mean of elements inside the sliding window with the maximum weighted mean observed so far. A significant difference between the two weighted means, upper-bounded by the McDiarmid inequality, implies a concept drift. Our extensive experimentation against synthetic and real-world data streams show that our novel method outperforms the state-of-the-art. Specifically, MDDM yields shorter detection delays as well as lower false negative rates, while maintaining high classification accuracies.
Ali Pesaranghader, Herna L. Viktor, Eric Paquet
IJCNN2
2018 SCUT-DS: Learning from Multi-class Imbalanced Canadian Weather Data
Olubukola M. Olaitan, Herna L. Viktor
ISMIS2
2018 Clustering in the Presence of Concept Drift
Richard Hugh Moulton, Herna L. Viktor, Nathalie Japkowicz, João Gama 0001
ECML/PKDD (1)2
2018 Dynamic adaptation of online ensembles for drifting data streams
Muhammed K. Olorunnimbe, Herna L. Viktor, Eric Paquet
J. Intell. Inf. Syst.2
2018 Reservoir of diverse adaptive learners and stacking fast hoeffding drift detection methods for evolving data streams
Ali Pesaranghader, Herna L. Viktor, Eric Paquet
Mach. Learn.2
2017 Context-Based Abrupt Change Detection and Adaptation for Categorical Data Streams
Sarah D'Ettorre, Herna L. Viktor, Eric Paquet
DS2
2017 An Exploratory Study of Oral and Dental Health in Canada
abstract
Healthcare practitioners agree that good oral health is a critical indicator of general health and wellness of a population. The lack of access to mandatory coverage for common issues such as cavities and non-surgical periodontal care often lead not only to medical problems, but also to loss of productivity. This trend is especially evident for older individuals and lower-income families. This paper discusses the results of our exploration of the annual Canadian Community Health Survey (CCHS), in order to further study the interplay between socio-economic factors and oral and dental health. To this end, we present the results when applying a number of machine learning algorithms to a CCHS data mart. Our results reaffirm that individuals' levels and sources of income are strong indicators of the number of dental visits per year. In addition, we found that younger adults and youth, who usually live in larger households, visit the dentist less frequently than all other survey respondents.
Andrei Belcin, Sean Floyd, Areej Asiri, Herna L. Viktor
ICMLA4
2016 A Framework for Classification in Data Streams Using Multi-strategy Learning
Ali Pesaranghader, Herna L. Viktor, Eric Paquet
DS2
2016 Fast Hoeffding Drift Detection Method for Evolving Data Streams
Ali Pesaranghader, Herna L. Viktor
ECML/PKDD (2)2
2015 Multi-label Classification of Anemia Patients
abstract
This work examines the application of machine learning to an important area of medicine which aims to diagnose paediatric patients with β-thalassemia minor, iron deficiency anemia or the co-occurrence of these ailments. Iron deficiency anemia is a major cause of microcytic anemia and is considered an important task in global health. Whilst existing methods, based on linear equations, are proficient at distinguishing between the two classes of anemia, they fail to identify the co-occurrence of this issues. Machine learning algorithms, however, can induce non-linear decision boundaries that enable accurate classification within complex domains. Through a multi-label classification technique, known as problem transformations, we convert the learning task to one that is appropriate for machine learning and examine the effectiveness of machine learning algorithms on this domain. Our results show that machine learning classifiers produce good overall accuracy and are able to identify instances of the co-occurrence class unlike the existing methods.
Colin Bellinger, Ali Amid, Nathalie Japkowicz, Herna L. Viktor
ICMLA4
2015 Tweets as a Vote: Exploring Political Sentiments on Twitter for Opinion Mining
Muhammed K. Olorunnimbe, Herna L. Viktor
ISMIS2
2014 Building Smart Cubes for Reliable and Faster Access to Data
Daniel K. Antwi, Herna L. Viktor
DaWaK2
2014 A Quantum Particle Swarm Optimization and Genetic Algorithm approach to the correspondence problem
abstract
Finding correspondences between deformable objects has wide application in many domains. In information retrieval, researchers may be interested in finding similar objects, while computer animation experts may be considering ways to morph shapes. The correspondence problem is especially challenging when the objects under consideration are suspect to non-rigid deformations, noise and/or distortions. In this paper, a novel method using Quantum Particle Swarm Optimization (QPSO) and Genetic Algorithms (GA) is presented to address this issue. In our QPSO-GA algorithm we formulate the problem of correspondence detection as an optimization problem over all possible mapping in between the geodesic distance matrices associated with two sets of point clouds. We proceed to identify the optimal mapping, by first applying Quantum Particle Swarm Optimization to the permutation matrices associated with their geodesic distance matrices and then employing Genetic Algorithms in order to guide the search. Experimental results suggest that our QPSO-GA algorithm is fast, scalable, and robust. Our method accurately identifies the correspondences between objects, even in the presence of noise and distortion.
Hamid Hadavi, Herna L. Viktor, Eric Paquet
INISTA2
2013 An isometry-invariant spectral approach for protein-protein docking
abstract
The protein docking problem refers to the task of predicting the appropriate matching of one protein molecule (the receptor) to another (the ligand), when attempting to bind them to form a stable complex. Research shows that matching the three-dimensional geometric structures of proteins plays a key role in determining a so-called docking pair. However, the active sites which are responsible for the binding do not always present a rigid-body shape matching problem. Rather, they may undergo deformations when docking occurs, which complicates the process. To address this issue, we present an isometry-invariant and topologically robust partial shape descriptor method for finding complementary protein sites. Our method employs Heat Kernel Signature shape descriptors which are based on the diffusion of heat on surfaces. Our experimental results against the Protein-Protein Benchmark 4.0 demonstrate the viability of our approach.
Dela De Youngster, Eric Paquet, Herna L. Viktor, Emil M. Petriu
BIBE3
2013 Isometrically Invariant Description of Deformable Objects Based on the Fractional Heat Equation
Eric Paquet, Herna L. Viktor
CAIP (2)2
2013 Artificial neural networks for predicting 3D protein shapes from amino acid sequences
abstract
Research has shown that the functionalities of proteins are largely influenced by their three dimensional (3D) shapes. This observation is especially relevant in drug design, where the knowledge of the 3D structure of a protein enables pharmacologists to select the best binding proteins when aiming to moderate functions. However, a relatively small number of 3D shapes are known. In contrast, amino acid sequences may be acquired through very efficient automated, high throughput experimental methods and the amino acid sequences of a vast number of proteins have therefore been identified. It follows that it is important to address this knowledge gap. To this end, this paper introduces an approach to predict the 3D shapes of proteins, utilizing feed-forward artificial neural networks. Our novel solution allows one to learn the representations of the 3D shape associated with a protein by starting directly from its amino acid sequence descriptors. Once a neural network is trained, our search engine enables one to retrieve the closest known 3D shape associated with an unknown, so-called query protein. We evaluate the performance of our approach against the Protein Data Bank (PDB), by considering proteins from a diverse set of families. Our results indicate that our system is able to accurately find the most similar protein structures for a wide variety of protein 3D shapes and diverse protein family sizes.
Herna L. Viktor, Eric Paquet
IJCNN1
2013 Reducing the size of databases for multirelational classification: a subgraph-based approach
Herna L. Viktor, Eric Paquet
J. Intell. Inf. Syst.2
2012 Aggregation and privacy in multi-relational databases
abstract
The aim of privacy-preserving data mining is to construct highly accurate predictive models while not disclosing privacy information. Aggregation functions, such as sum and count are often used to pre-process the data prior to applying data mining techniques to relational databases. Often, it is implicitly assumed that the aggregated (or summarized) data are less likely to lead to privacy violations during data mining. This paper investigates this claim, within the relational database domain. We introduce the PBIRD (Privacy Breach Investigation in Relational Databases) methodology. Our experimental results show that aggregation potentially introduces new privacy violations. That is, potentially harmful attributes obtained with aggregation are often different from the ones obtained from non-aggregated databases. This indicates that, even when privacy is enforced on non-aggregated data, it is not automatically enforced on the corresponding aggregated data. Consequently, special care should be taken during model building in order to fully enforce privacy when the data are aggregated.
Yasser Jafer, Herna L. Viktor, Eric Paquet
PST2
2011 Spectral Clustering: An Explorative Study of Proximity Measures
Nadia Farhanaz Azam, Herna L. Viktor
IC3K2
2011 Multimodal Representations, Indexing, Unexpectedness and Proteins
Eric Paquet, Herna L. Viktor
IEA/AIE (1)2
2009 Finding Clothing That Fit through Cluster Analysis and Objective Interestingness Measures
Isis Peña, Herna L. Viktor, Eric Paquet
DaWaK2
2008 Learning from Skewed Class Multi-relational Databases
Herna L. Viktor
Fundam. Informaticae2
2008 Multirelational classification: a multiple view approach
Herna L. Viktor
Knowl. Inf. Syst.2
2008 Capri/MR: exploring protein databases from a structural and physicochemical point of view
abstract
With the advent of high throughput systems to experimentally determine the three-dimensional (3-D) structure of proteins, molecular biologists are in urgent need of systems to automatically store, maintain and explore the vast structural databases that are thus being created. We have designed and implemented the Capri/MR system which makes it possible to identify families of protein structures, as contained in such very large 3-D protein structure databases. Our system is able to automatically index and search a database of proteins by three-dimensional shape, structural and/or physicochemical properties. For each of these diverse protein structure representations, we create a compact rotation and translation invariant index (or signature) which is placed in a database for future querying. A similarity search algorithm performs an exhaustive search against the entire database. Our search algorithm takes advantage of the compact signatures to rapidly find protein structures that are similar in 3-D shape and/or two-dimensional (2-D) properties. As a result, queries in our Capri/MR system run within a fraction of a second, and we are able to accurately group protein structures into the correct families, with very high precision and recall. In addition, our system dynamically processes new protein structures as they become available. We demonstrate the power of Capri/MR against the Protein Data Bank, which contains all known, experimentally determined, 3-D protein structures (48.000 as of January 2008). The main applications of our Capri/MR system lie in structural proteomics, protein evolution and mutation, as well as in drug design, in particular for studying the docking problem and the computer aided design of non-toxic drugs.
Eric Paquet, Herna L. Viktor
Proc. VLDB Endow.2
2007 Pruning Relations for Substructure Discovery of Multi-relational Databases
Herna L. Viktor, Eric Paquet
PKDD2
2006 Multi-view ANNs for Multi-relational Classification
abstract
Artificial neural networks (ANNs) provide a general, effective and practical approach for learning complex target functions. However, ANNs are not suitable for handling relational data, where information about the target concept is distributed over multiple related relations. ANNs algorithms usually only explore one relation, the so-called target relation, thus excluding crucial knowledge embedded in the related so-called background relations. This paper introduces a new approach, the multiple view artificial neural networks (MVNNs) method, to address the need for bridging the gap between ANNs and relational databases. The MVNNs strategy, firstly, propagates essential information held in the target relation to all background relations. Subsequently, it exploits multiple ANNs, which explore the target concepts against the separate back-ground relations. Thirdly, it incorporates crucial background knowledge, as obtained by the ANNs, into a meta-learning mechanism to construct the final model. Our experiments on eight data sets show that the MVNNs method achieves promising results in terms of overall accuracy obtained, when compared with two other relational data mining algorithms.
Herna L. Viktor
IJCNN2
2006 Mining relational data through correlation-based multiple view validation
abstract
Commercial relational databases currently store vast amounts of real-world data. The data within these relational repositories are represented by multiple relations, which are inter-connected by means of foreign key joins. The mining of such interrelated data poses a major challenge to the data mining community. Unfortunately, traditional data mining algorithms usually only explore one relation, the so-called target relation, thus excluding crucial knowledge embedded in the related so-called background relations. In this paper, we propose a novel approach for classifying relational such domains. This strategy employs multiple views to capture crucial information not only from the target relation, but also from related relations. This information is integrated into the relational mining process. The framework presented here, firstly, explore the relational domain to partition its features space into multiple subsets. Subsequently, these subsets are used to construct multiple uncorrelated views, based on a novel correlation-based view validation method, against the target concept. Finally, the knowledge possessed by multiple views are incorporated into a meta-learning mechanism to augment one another. Based on this framework, a wide range of conventional data mining methods can be applied to mine relational databases. Our experiments on benchmark real-world data sets show that the proposed method achieves promising results both in terms of overall accuracy obtained and run time, when compared with two other relational data mining approaches.
Herna L. Viktor
KDD2
2006 Measuring to Fit: Virtual Tailoring Through Cluster Analysis and Classification
Herna L. Viktor, Eric Paquet
PKDD1
2004 Boosting with Data Generation: Improving the Classification of Hard to Learn Examples
Herna L. Viktor
IEA/AIE2
2000 Generating New Patterns for Information Gain and Improved Neural Network Learning
abstract
This paper introduces an approach to generate new patterns for improved neural network training. The patterns are based on the information obtained by means of a rule extraction approach. In this way, the training process is re-iterated using the most informative patterns. The data generation process is further enhanced by incorporating the high quality rules obtained from a decision tree. Results indicate that the approach results in improved generalization, especially in difficult to learn domains.
Herna L. Viktor
IJCNN (4)1
1999 Improved generalisation using cooperative learning and rule extraction
abstract
Rule extraction from artificial neural networks (ANN) addresses the need for symbolic representations to explain an ANN's decisions. The ANNSER rule extraction approach extracts rules from feedforward ANN that are used for classification. The quality of the ANNSER rules reflects the strengths and weaknesses of the ANN and indicates those classes over which the ANN did not generalize well. This paper introduces a new approach to ANN training in which two or more ANNSER learners co-exist in a cooperative multiagent learning environment. The ANNSER learners cooperate by using one another's high quality rules to generate new training instances. In this way, the generalization of the ANN are improved, leading to a set of high quality rules that describe the knowledge embedded in the trained ANN.
Herna L. Viktor, Ian Cloete
IJCNN1