Georgina Cosma

dblp:40/9504 · DBLP profile ↗
← Back
27ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0002-4663-6907ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 7 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorSystems, architecture and hardware · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Image Complexity-Aware Adaptive Retrieval for Efficient Vision-Language Models
Mikel Williams-Lekuona, Georgina Cosma
ECIR (1)2
2026 Neural corrective machine unranking
abstract
Machine unlearning in neural information retrieval (IR) systems requires removing specific data while maintaining model performance. Applying existing machine unlearning methods to IR may compromise retrieval effectiveness or inadvertently expose unlearning actions due to the removal of particular items from the retrieved results presented to users. We formalise corrective unranking , which extends machine unlearning in the (neural) IR context by integrating substitute documents to preserve ranking integrity, and propose a novel teacher–student framework, Corrective unRanking Distillation (CuRD), for this task. CuRD (1) facilitates forgetting by adjusting the (trained) neural IR model such that its output relevance scores of to-be-forgotten samples mimic those of low-ranking, non-retrievable samples; (2) enables correction by fine-tuning the relevance scores for the substitute samples to match those of corresponding to-be-forgotten samples closely; (3) seeks to preserve performance on samples that are not targeted for forgetting. We evaluate CuRD on four neural IR models (BERTcat, BERTdot, ColBERT, PARADE) using MS MARCO and TREC CAR datasets. Experiments with forget set sizes from 1% to 20% of the training dataset demonstrate that CuRD outperforms seven state-of-the-art baselines in terms of forgetting and correction while maintaining model retention and generalisation capabilities.
Jingrui Hou, Axel Finke, Georgina Cosma
Inf. Sci.3
2026 Neural Machine Unranking
abstract
We address the problem of machine unlearning in neural information retrieval (IR), introducing a novel task termed neural machine unranking (NuMuR). This problem is motivated by growing demands for data privacy compliance and selective information removal in neural IR systems. Existing task-agnostic or model-agnostic unlearning approaches, primarily designed for classification tasks, are suboptimal for NuMuR due to two core challenges: 1) neural rankers output unnormalised relevance scores rather than probability distributions, limiting the effectiveness of traditional teacher-student distillation frameworks and 2) entangled data scenarios, where queries and documents appear simultaneously across both forget and retain sets, may degrade retention performance in existing methods. To address these issues, we propose contrastive and consistent loss (CoCoL), a dual-objective framework. CoCoL comprises 1) a contrastive loss that reduces relevance scores on forget sets while maintaining performance on entangled samples and 2) a consistent loss that preserves accuracy on the retain set. Extensive experiments on two datasets, across four neural IR models, demonstrate that CoCoL achieves substantial forgetting with minimal retention and generalization performance loss. CoCoL facilitates more effective and controllable data removal than existing techniques.
Jingrui Hou, Axel Finke, Georgina Cosma
IEEE Trans. Neural Networks Learn. Syst.3
2025 A meta-heuristic approach to estimate and explain classifier uncertainty
abstract
Abstract Trust is a crucial factor affecting the adoption of machine learning (ML) models. Qualitative studies have revealed that end-users, particularly in the medical domain, need models that can express their uncertainty in decision-making allowing users to know when to ignore the model’s recommendations. However, existing approaches for quantifying decision-making uncertainty are not model-agnostic, or they rely on complex mathematical derivations that are not easily understood by laypersons or end-users, making them less useful for explaining the model’s decision-making process. This work proposes a set of class-independent meta-heuristics that can characterise the complexity of an instance in terms of factors that are mutually relevant to both human and ML decision-making. The measures are integrated into a meta-learning framework that estimates the risk of misclassification. The proposed framework outperformed predicted probabilities and entropy-based methods of identifying instances at risk of being misclassified. Furthermore, the proposed approach resulted in uncertainty estimates that proves more independent of model accuracy and calibration than existing approaches. The proposed measures and framework demonstrate promise for improving model development for more complex instances and provides a new means of model abstention and explanation.
Andrew Houston, Georgina Cosma
Appl. Intell.2
2025 Advancing continual lifelong learning in neural information retrieval: Definition, dataset, framework, and empirical evaluation
Jingrui Hou, Georgina Cosma, Axel Finke
Inf. Sci.2
2024 VITR: Augmenting Vision Transformers with Relation-Focused Learning for Cross-modal Information Retrieval
abstract
The relations expressed in user queries are vital for cross-modal information retrieval. Relation-focused cross-modal retrieval aims to retrieve information that corresponds to these relations, enabling effective retrieval across different modalities. Pre-trained networks, such as Contrastive Language-Image Pre-training networks, have gained significant attention and acclaim for their exceptional performance in various cross-modal learning tasks. However, the Vision Transformer (ViT) used in these networks is limited in its ability to focus on image region relations. Specifically, ViT is trained to match images with relevant descriptions at the global level, without considering the alignment between image regions and descriptions. This article introduces VITR, a novel network that enhances ViT by extracting and reasoning about image region relations based on a local encoder. VITR is comprised of two key components. Firstly, it extends the capabilities of ViT-based cross-modal networks by enabling them to extract and reason with region relations present in images. Secondly, VITR incorporates a fusion module that combines the reasoned results with global knowledge to predict similarity scores between images and descriptions. The proposed VITR network was evaluated through experiments on the tasks of relation-focused cross-modal information retrieval. The results derived from the analysis of the Flickr30K, MS-COCO, RefCOCOg, and CLEVR datasets demonstrated that the proposed VITR network consistently outperforms state-of-the-art networks in image-to-text and text-to-image retrieval.
Georgina Cosma, Axel Finke
ACM Trans. Knowl. Discov. Data2
2023 A genetically-optimised artificial life algorithm for complexity-based synthetic dataset generation
abstract
Algorithmic evaluation is a vital step in developing new approaches to machine learning and relies on the availability of existing datasets. However, real-world datasets often do not cover the necessary complexity space required to understand an algorithm’s domains of competence. As such, the generation of synthetic datasets to fill gaps in the complexity space has gained attention, offering a means of evaluating algorithms when data is unavailable. Existing approaches to complexity-focused data generation are limited in their ability to generate solutions that invoke similar classification behaviour to real data. The present work proposes a novel method (Sy:Boid) for complexity-based synthetic data generation, adapting and extending the Boid algorithm that was originally intended for computer graphics simulations. Sy:Boid embeds the modified Boid algorithm within an evolutionary multi-objective optimisation algorithm to generate synthetic datasets which satisfy predefined magnitudes of complexity measures. Sy:Boid is evaluated and compared to labelling-based and sampling-based approaches to data generation to understand its ability to generate a wide variety of realistic datasets. Results demonstrate Sy:Boid is capable of generating datasets across a greater portion of the complexity space than existing approaches. Furthermore, the produced datasets were observed to invoke very similar classification behaviours to that of real data.
Andrew Houston, Georgina Cosma
Inf. Sci.2
2023 Improving visual-semantic embeddings by learning semantically-enhanced hard negatives for cross-modal information retrieval
abstract
Visual Semantic Embedding (VSE) networks aim to extract the semantics of images and their descriptions and embed them into the same latent space for cross-modal information retrieval. Most existing VSE networks are trained by adopting a hard negatives loss function which learns an objective margin between the similarity of relevant and irrelevant image–description embedding pairs. However, the objective margin in the hard negatives loss function is set as a fixed hyperparameter that ignores the semantic differences of the irrelevant image–description pairs. To address the challenge of measuring the optimal similarities between image–description pairs before obtaining the trained VSE networks, this paper presents a novel approach that comprises two main parts: (1) finds the underlying semantics of image descriptions; and (2) proposes a novel semantically-enhanced hard negatives loss function, where the learning objective is dynamically determined based on the optimal similarity scores between irrelevant image–description pairs. Extensive experiments were carried out by integrating the proposed methods into five state-of-the-art VSE networks that were applied to three benchmark datasets for cross-modal information retrieval tasks. The results revealed that the proposed methods achieved the best performance and can also be adopted by existing and future VSE networks.
Georgina Cosma
Pattern Recognit.2
2022 Multifaceted Hierarchical Report Identification for Non-Functional Bugs in Deep Learning Frameworks
abstract
Non-functional bugs (e.g., performance-or accuracy-related bugs) in Deep Learning (DL) frameworks can lead to some of the most devastating consequences. Reporting those bugs on a repository such as GitHub is a standard route to fix them. Yet, given the growing number of new GitHub reports for DL frameworks, it is intrinsically difficult for developers to distinguish those that reveal non-functional bugs among the others, and assign them to the right contributor for investigation in a timely manner. In this paper, we propose MHNurf -an end-to-end tool for automatically identifying non-functional bug related reports in DL frameworks. The core of MHNurf is a Multifaceted Hierarchical Attention Network (MHAN) that tackles three unaddressed challenges: (1) learning the semantic knowledge, but doing so by (2) considering the hierarchy (e.g., words/tokens in sentences/statements) and focusing on the important parts (i.e., words, tokens, sentences, and statements) of a GitHub report, while (3) independently extracting information from different types of features, i.e., content, comment, code, command, and label.To evaluate MHNurf, we leverage 3,721 GitHub reports from five DL frameworks for conducting experiments. The results show that MHNurf works the best with a combination of content, comment, and code, which considerably outperforms the classic HAN where only the content is used. MHNurf also produces significantly more accurate results than nine other state-of-the-art classifiers with strong statistical significance, i.e., up to 71% AUC improvement and has the best Scott-Knott rank on four frameworks while 2nd on the remaining one. To facilitate reproduction and promote future research, we have made our dataset, code, and detailed supplementary results publicly available at: https://github.com/ideas-labo/APSEC2022-MHNurf.
Guoming Long, Tao Chen 0001, Georgina Cosma
APSEC3
2021 TAGA: Tabu Asexual Genetic Algorithm embedded in a filter/filter feature selection approach for high-dimensional data
Sadegh Salesi, Georgina Cosma, Michalis Mavrovouniotis
Inf. Sci.2
2021 Generalisation Power Analysis for finding a stable set of features using evolutionary algorithms for feature selection
Sadegh Salesi, Georgina Cosma
Knowl. Based Syst.2
2020 Enhancing Prediction in Cyclone Separators through Computational Intelligence
abstract
Pressure drop prediction is critical to the design and performance of cyclone separators as industrial gas cleaning devices. The complex non-linear relationship between cyclone Pressure Drop Coefficient (PDC) and geometrical dimensions suffice the need for state-of-the-art predictive modelling methods. Existing solutions have applied theoretical/semi-empirical techniques which fail to generalise well, and the suitability of intelligent techniques has not been widely explored for the task of pressure drop prediction in cyclone separators. To this end, this paper firstly introduces a fuzzy modelling methodology, then presents an alternative version of the Extended Kalman Filter (EKF) to train a Multi-Layer Neural Network (MLNN). The Lagrange dual formulation of Support Vector Machine (SVM) regression model is also deployed for comparison purposes. For optimal design of these models, manual and grid search techniques are used in a cross-validation setting subsequent to training. Based on the prediction accuracy of PDC, results show that the Fuzzy System (FS) is highly performing with testing mean squared error (MSE) of 3.97e-04 and correlation coefficient (R) of 99.70%. Furthermore, a significant improvement of EKF-trained network (MSE = 1.62e-04, R = 99.82%) over the traditional Back-Propagation Neural Network (BPNN) (MSE = 4.87e-04, R = 99.53%) is observed. SVM gives better prediction with radial basis kernel (MSE = 2.22e-04, R = 99.75%) and provides comparable performance to universal approximators. Of the conventional models considered, the model of Shepherd and Lapple (MSE = 7.3e-03, R = 97.88%) gives the best result which is still inferior to the intelligent models.
Oluwaseyi Ogun, Mbetobong Enoh, Georgina Cosma, Aboozar Taherkhani, Vincenzo Madonna
CEC3
2020 Classifying Imbalanced Multi-modal Sensor Data for Human Activity Recognition in a Smart Home using Deep Learning
abstract
In smart homes, data generated from real-time sensors for human activity recognition is complex, noisy and imbalanced. It is a significant challenge to create machine learning models that can classify activities which are not as commonly occurring as other activities. Machine learning models designed to classify imbalanced data are biased towards learning the more commonly occurring classes. Such learning bias occurs naturally, since the models better learn classes which contain more records. This paper examines whether fusing real-world imbalanced multi-modal sensor data improves classification results as opposed to using unimodal data; and compares deep learning approaches to dealing with imbalanced multi-modal sensor data when using various resampling methods and deep learning models. Experiments were carried out using a large multi-modal sensor dataset generated from the Sensor Platform for HEalthcare in a Residential Environment (SPHERE). The data comprises 16104 samples, where each sample comprises 5608 features and belongs to one of 20 activities (classes). Experimental results using SPHERE demonstrate the challenges of dealing with imbalanced multi-modal data and highlight the importance of having a suitable number of samples within each class for sufficiently training and testing deep learning models. Furthermore, the results revealed that when fusing the data and using the Synthetic Minority Oversampling Technique (SMOTE) to correct class imbalance, CNN-LSTM achieved the highest classification accuracy of 93.67% followed by CNN, 93.55%, and LSTM, i.e. 92.98%.
Ali A. Alani, Georgina Cosma, Aboozar Taherkhani
IJCNN2
2020 Predicting Insulation Resistance of Enamelled Wire using Neural Network and Curve Fit Methods Under Thermal Aging
abstract
Health monitoring has gained a massive interest in power systems engineering, as it has the advantage to reduce operating costs, improve reliability of power supply and provide a better service to customers. This paper presents surrogate methods to predict the electrical insulation lifetime using the neural network approach and three curve fitting models. These can be used for the health monitoring of insulating systems in electrical equipment, such as motors, generators, and transformers. The curve fit models and the supervised backpropagation neural network are employed to predict the insulation resistance trend of enameled copper wires, when stressed with a temperature of 290 °C. After selecting a suitable end of life criterion, the specimens' mean time-to-failure is estimated, and the performance of each of the analyzed models is apprised through a comparison with the standard method for thermal life evaluation of enameled wires. Amongst all, the best prediction accuracy is achieved by a Backpropagation neural network approach, which gives an error of just 3.29% when compared with the conventional life evaluation method, whereas, the error is above 10% for all the three investigated curve fit models.
Gulrukh Turabee, Georgina Cosma, Vincenzo Madonna, Paolo Giangrande, Muhammad Raza Khowja, Gaurang Vakil, Chris Gerada, Michael Galea
IJCNN2
2020 AdaBoost-CNN: An adaptive boosting algorithm for convolutional neural networks to classify multi-class imbalanced datasets using transfer learning
Aboozar Taherkhani, Georgina Cosma, T. Martin McGinnity
Neurocomputing2
2020 A review of learning in biologically plausible spiking neural networks
Aboozar Taherkhani, Ammar Belatreche, Yuhua Li 0001, Georgina Cosma, Liam P. Maguire, T. Martin McGinnity
Neural Networks4
2019 NatCSNN: A Convolutional Spiking Neural Network for Recognition of Objects Extracted from Natural Images
Georgina Cosma, T. Martin McGinnity
ICANN (1)2
2019 A Hybrid Semantic Knowledgebase-Machine Learning Approach for Opinion Mining
abstract
Opinion mining tools enable users to efficiently process a large number of online reviews in order to determine the underlying opinions. This paper presents a Hybrid Semantic Knowledgebase-Machine Learning approach for mining opinions at the domain feature level and classifying the overall opinion on a multi-point scale. The proposed approach benefits from the advantages of deploying a novel Semantic Knowledgebase approach to analyse a collection of reviews at the domain feature level and produce a set of structured information that associates the expressed opinions with specific domain features. The information in the knowledgebase is further supplemented with domain-relevant facts sourced from public Semantic datasets, and the enriched semantically-tagged information is then used to infer valuable semantic information about the domain as well as the expressed opinions on the domain features by summarising the overall opinions about the domain across multiple reviews, and by averaging the overall opinions about other cinematic features. The retrieved semantic information represents a valuable resource for modelling a machine learning classifier to predict the numerical rating of each review. Experimental evaluation revealed that the proposed Hybrid Semantic Knowledgebase-Machine Learning approach improved the precision and recall of the extracted domain features, and hence proved suitable for producing an enriched dataset of semantic features that resulted in higher classification accuracy.
Rowida Alfrjani, Taha Osman, Georgina Cosma
Data Knowl. Eng.3
2018 Bio-Inspired Ganglion Cell Models for Detecting Horizontal and Vertical Movements
abstract
The retina performs the earlier stages of image processing in living beings and is composed of six different groups of cells, namely, the rods, cones, horizontal, bipolar, amacrine and ganglion cells. Each of those group of cells can be sub-divided into other types of cells that vary in shape, size, connectivity and functionality. Each cell is responsible for performing specific tasks in these early stages of biological image processing. Some of those cells are sensitive to horizontal and vertical movements. This paper proposes a multi-hierarchical spiking neural network architecture for detecting horizontal and vertical movements using a custom dataset which was generated in laboratory settings. The proposed architecture was designed to reflect the connectivity, behaviour and the number of layers found in the majority of vertebrates retinas, including humans. The architecture was trained using 2303 images and tested using 816 images. Simulation results revealed that each cell model is sensitive to vertical and horizontal movements with a detection error of 6.75 percent.
Andreas Oikonomou, Georgina Cosma, T. Martin McGinnity
IJCNN3
2018 Deep-FS: A feature selection algorithm for Deep Boltzmann Machines
abstract
A Deep Boltzmann Machine is a model of a Deep Neural Network formed from multiple layers of neurons with nonlinear activation functions. The structure of a Deep Boltzmann Machine enables it to learn very complex relationships between features and facilitates advanced performance in learning of high-level representation of features, compared to conventional Artificial Neural Networks. Feature selection at the input level of Deep Neural Networks has not been well studied, despite its importance in reducing the input features processed by the deep learning model, which facilitates understanding of the data. This paper proposes a novel algorithm, Deep Feature Selection (Deep-FS), which is capable of removing irrelevant features from large datasets in order to reduce the number of inputs which are modelled during the learning process. The proposed Deep-FS algorithm utilizes a Deep Boltzmann Machine, and uses knowledge which is acquired during training to remove features at the beginning of the learning process. Reducing inputs is important because it prevents the network from learning the associations between the irrelevant features which negatively impact on the acquired knowledge of the network about the overall distribution of the data. The Deep-FS method embeds feature selection in a Restricted Boltzmann Machine which is used for training a Deep Boltzmann Machine. The generative property of the Restricted Boltzmann Machine is used to reconstruct eliminated features and calculate reconstructed errors, in order to evaluate the impact of eliminating features. The performance of the proposed approach was evaluated with experiments conducted using the MNIST, MIR-Flickr, GISETTE, MADELON and PANCAN datasets. The results revealed that the proposed Deep-FS method enables improved feature selection without loss of accuracy on the MIR-Flickr dataset, where Deep-FS reduced the number of input features by removing 775 features without reduction in performance. With regards to the MNIST dataset, Deep-FS reduced the number of input features by more than 45%; it reduced the network error from 0.97% to 0.90%, and also reduced processing and classification time by more than 5.5%. Additionally, when compared to classical feature selection methods, Deep-FS returned higher accuracy. The experimental results on GISETTE, MADELON and PANCAN showed that Deep-FS reduced 81%, 57% and 77% of the number of input features, respectively. Moreover, the proposed feature selection method reduced the classifier training time by 82%, 70% and 85% on GISETTE, MADELON and PANCAN datasets, respectively. Experiments with various datasets, comprising a large number of features and samples, revealed that the proposed Deep-FS algorithm overcomes the main limitations of classical feature selection algorithms. More specifically, most classical methods require, as a prerequisite, a pre-specified number of features to retain, however in Deep-FS this number is identified automatically. Deep-FS performs the feature selection task faster than classical feature selection algorithms which makes it suitable for deep learning tasks. In addition, Deep-FS is suitable for finding features in large and big datasets which are normally stored in data batches for faster and more efficient processing.
Aboozar Taherkhani, Georgina Cosma, T. Martin McGinnity
Neurocomputing2
2017 A multivariate feature selection framework for high dimensional biomedical data classification
abstract
High dimensional biomedical data are becoming common in various predictive models developed for disease diagnosis and prognosis. Extracting knowledge from high dimensional data which contain a large number of features and a small sample size presents intrinsic challenges for classification models. Genetic Algorithms can be successfully adopted to efficiently search through high dimensional spaces, and multivariate classification methods can be utilized to evaluate combinations of features for constructing optimized predictive models. This paper proposes a framework which can be adopted for building prediction models for high dimensional biomedical data. The proposed framework comprises of three main phases. The feature filtering phase which filters out the noisy features; the feature selection phase which is based on multivariate machine learning techniques and the Genetic Algorithm to evaluate the filtered features and select the most informative subsets of features for achieving maximum classification performance; and the predictive modeling phase during which machine learning algorithms are trained on the selected features to construct a reliable prediction model. Experiments were conducted using four high dimensional biomedical datasets including protein and geneexpression data. The results revealed optimistic performances for the multivariate selection approaches which utilize classification measurements based on implicit assumptions.
Abeer Alzubaidi, Georgina Cosma
CIBCB2
2017 Style Analysis for Source Code Plagiarism Detection - An Analysis of a Dataset of Student Coursework
abstract
Plagiarism has become an increasing problem in higher education in recent years. Coding style can be used to detect source code plagiarism that involves writing and deciding the structure of the code which does not affect the logic of a program, thus offering a way to differentiate between different code authors. This paper focuses to identify whether a data set consisting of student programming assignments is rich enough to apply coding style metrics to detect similarities between code sequences, and we use the BlackBox dataset as a case study.
Olfat M. Mirza, Mike Joy, Georgina Cosma
ICALT3
2017 A survey on computational intelligence approaches for predictive modeling in prostate cancer
Georgina Cosma, David J. Brown 0001, Matthew Archer, Masood Khan, A. Graham Pockley
Expert Syst. Appl.1
2017 Perceptual Comparison of Source-Code Plagiarism within Students from UK, China, and South Cyprus Higher Education Institutions
abstract
Perspectives of students on what constitutes source-code plagiarism may differ based on their educational background. Surveys have been conducted with home students undertaking computing and joint computing subject degrees at higher education institutions throughout the UK, China, and South Cyprus, and a total of 984 responses have been statistically analysed to determine the common areas of understanding and misunderstanding among students on various topics related to source-code plagiarism. The study identifies those topics which are well understood, and those topics which are not properly understood across the different groups of students, and is the first study which specifically discusses Cypriot student perceptions on source-code plagiarism. This study provides useful information to educators (teaching home and international students) who wish to better inform their students on the issues of plagiarism and source-code plagiarism. Finally, the survey results revealed that although students who were informed about plagiarism better understood what actions constitute plagiarism, some topics were still unclear among students regardless of the students’ educational background and whether they had been previously informed about plagiarism.
Georgina Cosma, Mike Joy, Jane E. Sinclair, Margarita Andreou, Dongyong Zhang, Beverley Cook, Russell Boyatt
ACM Trans. Comput. Educ.1
2015 A Fuzzy-based approach to programming language independent source-code plagiarism detection
abstract
Source-code plagiarism detection in programming, concerns the identification of source-code files that contain similar and/or identical source-code fragments. Fuzzy clustering approaches are a suitable solution to detecting source-code plagiarism due to their capability to capture the qualitative and semantic elements of similarity. This paper proposes a novel Fuzzy-based approach to source-code plagiarism detection, based on Fuzzy C-Means and the Adaptive-Neuro Fuzzy Inference System (ANFIS). In addition, performance of the proposed approach is compared to the Self- Organising Map (SOM) and the state-of-the-art plagiarism detection Running Karp-Rabin Greedy-String-Tiling (RKR-GST) algorithms. The advantages of the proposed approach are that it is programming language independent, and hence there is no need to develop any parsers or compilers in order for the fuzzy-based predictor to provide detection in different programming languages. The results demonstrate that the performance of the proposed fuzzy-based approach overcomes all other approaches on well-known source code datasets, and reveals promising results as an efficient and reliable approach to source-code plagiarism detection.
Giovanni Acampora, Georgina Cosma
FUZZ-IEEE2
2014 An extended neuro-fuzzy approach for efficiently predicting review ratings in E-markets
abstract
Internet has opened new interesting scenarios in the fields of commerce and marketing. In particular, the idea of e-commerce has enabled customers to perform their transactions in a faster and cheaper way than conventional markets, and it has allowed companies to increase their sales volume thanks to a world-wide visibility. However, one of the problems that can strongly affect the performance of any e-commerce portal is related to the quality and validity of ratings provided by customers in their past transactions. Indeed, these reviews are used to determine the extent of customers acceptance and satisfaction of a product or service and they can affect the future selling performance and market share of a company. As a consequence, an efficient analysis of customer feedback could allow e-commerce portals to improve their selling capabilities and revenue. This paper introduces an innovative computational intelligence framework for efficiently learning review ratings in e-commerce by addressing different issues involved in this significant task: the dimension and imprecision of ratings data. In particular, we integrate the techniques of Singular Value Decomposition (SVD), Fuzzy C-Means (FCM) and ANFIS and, as shown in experimental results, this synergetic approach yields better learning performance than other rating predictors based on a conventional artificial neural network and FCM algorithm.
Giovanni Acampora, Georgina Cosma, Taha Osman
FUZZ-IEEE2
2012 An Approach to Source-Code Plagiarism Detection and Investigation Using Latent Semantic Analysis
abstract
Plagiarism is a growing problem in academia. Academics often use plagiarism detection tools to detect similar source-code files. Once similar files are detected, the academic proceeds with the investigation process which involves identifying the similar source-code fragments within them that could be used as evidence for proving plagiarism. This paper describes PlaGate, a novel tool that can be integrated with existing plagiarism detection tools to improve plagiarism detection performance. The tool also implements a new approach for investigating the similarity between source-code files with a view to gathering evidence for proving plagiarism. Graphical evidence is presented that allows for the investigation of source-code fragments with regards to their contribution toward evidence for proving plagiarism. The graphical evidence indicates the relative importance of the given source-code fragments across files in a corpus. This is done by using the Latent Semantic Analysis information retrieval technique to detect how important they are within the specific files under investigation in relation to other files in the corpus.
Georgina Cosma, Mike Joy
IEEE Trans. Computers1