VLDB 2026 Research / reviewers in the wild / expert
Héctor D. Menéndez 0001
dblp:45/9854 · also Héctor David Menéndez 0001, Héctor David Menéndez-Benito, Héctor Menéndez 0001
· DBLP profile ↗
43ranked-venue papers
19as first author
21since 2021 · last 2026
0000-0002-6314-3725ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 12 first-author · 10 since 2021Software engineering, systems software and programming languages · 16 · 5 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 4 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MuSEvo: Assessing the Robustness of Multilingual Summarization Metrics Through Evolutionary Adversarial Natural Language Processing-Based Attacks
Gema Bello Orgaz, Aidan Dakhama, Cristian Ramírez-Atencia, Héctor D. Menéndez 0001 |
ICCSA (2) | 4 |
| 2026 | StableZero: Improving General Image Generation Through Visualisation Feedback
Kazim Laos, Héctor D. Menéndez 0001, Gema Bello Orgaz |
ICCSA (1) | 2 |
| 2026 | Reducing LLMs by Searching Heterogeneous Quantizations
Enrique Alba 0001, Yazhuo Cao, Héctor D. Menéndez 0001 |
SSBSE | 3 |
| 2026 | White-Box Execution Refactoring of Transformers for Lower Energy
Enrique Alba 0001, Héctor D. Menéndez 0001 |
SSBSE | 2 |
| 2026 | Size doesn't matter: Assessing the trustworthiness of large language models in medical contexts: A focus on epidural information retrievalabstractBACKGROUND: Since the release of ChatGPT, numerous LLMs have emerged, providing easy access to information without the need for technical expertise. However, relying on these systems can influence important life decisions, such as the choice to use epidural analgesia during childbirth. Epidural analgesia is widely regarded as the "gold standard" for pain relief during childbirth. However, limited access to anaesthesiologists and gaps in knowledge may prompt individuals to seek information from unverified sources, including AI systems. Misinformation in this area can discourage the use of effective analgesia, highlighting the need to assess the accuracy of LLM-generated content. OBJECTIVE: To evaluate the reliability of LLM-generated information regarding epidural analgesia. METHODS: We posed 10 standardized questions about epidural analgesia to 12 LLMs, each question reformulated 10 times in both Spanish and English, resulting in 2400 responses. Two anaesthesiologists were involved in the assessment process. One expert performed the initial ratings, while the second independently verified the evaluations assessing the outputs using an extended SERVQUAL framework. RESULTS: ChatGPT performed best, followed by Gemini 2. Medium-sized models, such as Phi-3 and OpenChat, outperformed several larger models like Llama-2 or Llama-3, challenging the notion that "bigger is better" and offering potential advantages in low-resource settings (e.g., Phi-3 outperformed Llama-2 with an average increase of 81% across all metrics). Specialized models did not show superior performance. Except for ChatGPT, English responses were generally more reliable than Spanish, with some Spanish outputs incoherent. ChatGPT also exhibited the least variability between responses. CONCLUSIONS: Despite promising performance, LLMs display limitations in medical contexts. Collaboration between national and international medical societies is crucial to develop evidence-based resources to guide LLM training and improve information trustworthiness. Marina del Barrio, Kazim Laos, María José Vilchez Lara, Carlos Goicoechea García, Héctor D. Menéndez 0001 |
Artif. Intell. Medicine | 5 |
| 2025 | HotCat: Green and Effective Feature Selection toward Hotfix Bug Taxonomy
Luis De La Cal, Yazhuo Cao, Ayse Irmak Ercevik, Giovanni Pinna, Lukas Twist, Karine Even-Mendoza, William B. Langdon, Héctor D. Menéndez 0001, Federica Sarro |
SSBSE | 9 |
| 2025 | GreenMalloc: Allocator Optimisation for Industrial Workloads
Aidan Dakhama, William B. Langdon, Héctor D. Menéndez 0001, Karine Even-Mendoza |
SSBSE | 3 |
| 2025 | GA4GC: Greener Agent for Greener Code via Multi-objective Configuration Optimization
Jingzhi Gong, Yixin Bian, Luis De La Cal, Giovanni Pinna, Anisha Uteem, Mar Zamorano López, Karine Even-Mendoza, William B. Langdon, Héctor D. Menéndez 0001, Federica Sarro |
SSBSE | 10 |
| 2025 | Enhancing search-based testing with LLMs for finding bugs in system simulatorsabstractAbstract Despite the wide availability of automated testing techniques such as fuzzing, little attention has been devoted to testing computer architecture simulators. We propose a fully automated approach for this task. Our approach uses large language models (LLM) to generate input programs, including information about their parameters and types, as test cases for the simulators. The LLM’s output becomes the initial seed for an existing fuzzer, , which has been enhanced with three mutation operators, targeting both the input binary program and its parameters. We implement our approach in a tool called . We use it to test the system simulator. discovered 21 new bugs in , 14 where ’s software prediction differs from the real behaviour on actual hardware, and 7 where it crashed. New defects were uncovered with each of the 6 LLMs used. Aidan Dakhama, Karine Even-Mendoza, William B. Langdon, Héctor D. Menéndez 0001, Justyna Petke |
Autom. Softw. Eng. | 4 |
| 2024 | Responsible MLOps Design Methodology for an Auditing System for AI-Based Clinical Decision Support Systems
Pepita Barnard, John Robert Bautista, Aidan Dakhama, Arya Farahi, Kazim Laos, Anqi Liu 0001, Héctor D. Menéndez 0001 |
ICTSS | 7 |
| 2024 | Automatic Summarization Evaluation: Methods and Practices
Héctor D. Menéndez 0001, Aidan Dakhama |
ICTSS | 1 |
| 2024 | Summary of ObfSec: Measuring the Security of Obfuscations from a Testing Perspective
Héctor D. Menéndez 0001, Guillermo Suarez-Tangil |
ICTSS | 1 |
| 2023 | June: A Type Testability Transformation for Improved ATG PerformanceabstractStrings are universal containers: they are flexible to use, abundant in code, and difficult to test. String-controlled programs are programs that make branching decisions based on string input. Automatically generating valid test inputs for these programs considering only character sequences rather than any underlying string-encoded structures, can be prohibitively expensive. We present June, a tool that enables Java developers to expose any present latent string structure to test generation tools. June is an annotation-driven testability transformation and an extensible library, JuneLib, of structured string definitions. The core JuneLib definitions are empirically derived and provide templates for all structured strings in our test set. June takes lightly annotated source code and injects code that permits an automated test generator (ATG) to focus on the creation of mutable substrings inside a structured string. Using June costs the developer little, with an average of 2.1 annotations per string-controlled class. June uses standard Java build tools and therefore deploys seamlessly within a Java project. By feeding string structure information to an ATG tool, June dramatically reduces wasted effort; branches are effortlessly covered that would otherwise be extremely difficult, or impossible, to cover. This waste reduction both increases and speeds coverage. EvoSuite, for example, achieves the same coverage on June-ed classes in 1 minute, on average, as it does in 9 minutes on the un-June-ed class. These gains increase over time. On our corpus, June-ing a program compresses 24 hours of execution time into ca. 2 hours. We show that many ATG tools can reuse the same June-ed code: a few June annotations, a one-off cost, benefit many different testing regimes. Dan Bruce, David A. Kelly, Héctor D. Menéndez 0001, Earl T. Barr, David Clark 0001 |
ISSTA | 3 |
| 2023 | StableYolo: Optimizing Image Generation for Large Language Models
Harel Berger, Aidan Dakhama, Zishuo Ding, Karine Even-Mendoza, David A. Kelly, Héctor D. Menéndez 0001, Rebecca Moussa, Federica Sarro |
SSBSE | 6 |
| 2023 | SearchGEM5: Towards Reliable Gem5 with Search Based Software Testing and Large Language Models
Aidan Dakhama, Karine Even-Mendoza, William B. Langdon, Héctor D. Menéndez 0001, Justyna Petke |
SSBSE | 4 |
| 2022 | Measuring Machine Learning Robustness in front of Static and Dynamic AdversariesabstractAdversarial machine learning brought a new way of understanding the reliability of different learning systems. Knowing that the learning confidence depends significantly on small changes, such as noise, created a mind change in the artificial intelligence community, who started to consider the boundaries and limitations of machine learning methods. However, if we can measure these limitations, we can improve the strength of our machine learning models and their robustness. Following this motivation, this work introduces different measures of robustness for machine learning models based on false negatives. These measures can be evaluated for either static or dynamic scenarios, where an adversary performs intelligent actions to evade the system. To evaluate the metrics I have applied 11 classifiers to different benchmark datasets and created an adversary that performs an evolutionary search process aiming to reduce the classification accuracy. The results show that the most robust models are related to K-Nearest Neighbours, Logistic regression, and neural networks, although none of the systems is robust enough when the target is to reach a single misclassification. Héctor D. Menéndez 0001 |
ICTAI | 1 |
| 2022 | ObfSec: Measuring the security of obfuscations from a testing perspectiveabstractCode obfuscation protects the intellectual property of software. However, systematically altering the control- and data-flow of a program can deteriorate the security of the resulting program. There are a wide-range of obfuscation methods available that alter the layout of the program in different ways. These modifications can introduce bugs in the program or modify the nature and the severity of an existing ones. We propose a novel strategy, called ObfSec (Obfuscation Security), to understand the implications behind obfuscating software. ObfSec starts by detecting errors on software and exposes how the obfuscation can change the nature of those errors, looking in particular at transformations that turn software bugs into a exploitable vulnerable program. Our results, on a corpus of around 70,000 programs and obfuscations, show that obfuscation can deteriorate the security of a program. Héctor D. Menéndez 0001, Guillermo Suarez-Tangil |
Expert Syst. Appl. | 1 |
| 2022 | Output Sampling for Output Diversity in Automatic Unit Test GenerationabstractDiverse test sets are able to expose bugs that test sets generated with structural coverage techniques cannot discover. Input-diverse test set generators have been shown to be effective for this, but also have limitations: e.g., they need to be complemented with semantic information derived from the Software Under Test. We demonstrate how to drive the test set generation process with semantic information in the form of output diversity. We present the first totally automatic output sampling for output diversity unit test set generation tool, called OutGen. OutGen transforms a program into an SMT formula in bit-vector arithmetic. It then applies universal hashing in order to generate an output-based diverse set of inputs. The result offers significant diversity improvements when measured as a high output uniqueness count. It achieves this by ensuring that the test set’s output probability distribution is uniform, i.e., highly diverse. The use of output sampling, as opposed to any of input sampling, CBMC, CAVM, behaviour diversity or random testing improves mutation score and bug detection by up to 4150 and 963 percent respectively on programs drawn from three different corpora: the R-project, SIR and CodeFlaws. OutGen test sets achieve an average mutation score of up to 92 percent, and 70 percent of the test sets detect the defect. Moreover, OutGen is the only automatic unit test generation tool that is able to detect bugs on the real number C functions from the R-project. Héctor D. Menéndez 0001, Michele Boreale, Daniele Gorla, David Clark 0001 |
IEEE Trans. Software Eng. | 1 |
| 2022 | Hashing Fuzzing: Introducing Input Diversity to Improve Crash DetectionabstractThe utility of a test set of program inputs is strongly influenced by its diversity and its size. Syntax coverage has become a standard proxy for diversity. Although more sophisticated measures exist, such as proximity of a sample to a uniform distribution, methods to use them tend to be type dependent. We use r-wise hash functions to create a novel, semantics preserving, testability transformation for C programs that we call HashFuzz. Use of HashFuzz improves the diversity of test sets produced by instrumentation-based fuzzers. We evaluate the effect of the HashFuzz transformation on eight programs from the Google Fuzzer Test Suite using four state-of-the-art fuzzers that have been widely used in previous research. We demonstrate pronounced improvements in the performance of the test sets for the transformed programs across all the fuzzers that we used. These include strong improvements in diversity in every case, maintenance or small improvement in branch coverage – up to 4.8 perent improvement in the best case, and significant improvement in unique crash detection numbers – between 28 to 97 perent increases compared to test sets for untransformed programs. Héctor D. Menéndez 0001, David Clark 0001 |
IEEE Trans. Software Eng. | 1 |
| 2021 | Designing large quantum key distribution networks via medoid-based algorithms
Iván García-Cobo, Héctor D. Menéndez 0001 |
Future Gener. Comput. Syst. | 2 |
| 2021 | Diversifying Focused Testing for Unit TestingabstractSoftware changes constantly, because developers add new features or modifications. This directly affects the effectiveness of the test suite associated with that software, especially when these new modifications are in a specific area that no test case covers. This article tackles the problem of generating a high-quality test suite to cover repeatedly a given point in a program, with the ultimate goal of exposing faults possibly affecting the given program point. Both search-based software testing and constraint solving offer ready, but low-quality, solutions to this: Ideally, a maximally diverse covering test set is required, whereas search and constraint solving tend to generate test sets with biased distributions. Our approach, Diversified Focused Testing (DFT), uses a search strategy inspired by GödelTest. We artificially inject parameters into the code branching conditions and use a bi-objective search algorithm to find diverse inputs by perturbing the injected parameters, while keeping the path conditions still satisfiable. Our results demonstrate that our technique, DFT, is able to cover a desired point in the code at least 90% of the time. Moreover, adding diversity improves the bug detection and the mutation killing abilities of the test suites. We show that DFT achieves better results than focused testing, symbolic execution, and random testing by achieving from 3% to 70% improvement in mutation score and up to 100% improvement in fault detection across 105 software subjects. Héctor D. Menéndez 0001, Gunel Jahangirova, Federica Sarro, Paolo Tonella, David Clark 0001 |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2019 | VARMOG: A Co-Evolutionary Algorithm to Identify Manifolds on Large DataabstractDetecting clusters defining a specific shape or manifold is an open problem and has, indeed, inspired different machine learning algorithms. These methodologies normally lack scalability, as they depend on the performance of very sophisticated processes, such as extracting the Laplacian of a similarity graph in spectral clustering. When the algorithms need not only to identify manifolds on large amounts of data or streams, but also select the number of clusters, they failed either because of the robustness of their processes or by computational limitations. This paper introduces a general methodology that works in two levels: the initial step summarizes the data into a set of relevant features using the Euclidean properties of manifolds, and the second applies a robust methodology based on a co-evolutionary multi-objective clustering algorithm that identifies both, the number of manifolds and their associated manifold. The results show that this method outperforms different state of the art clustering processes for both, benchmark and real-world datasets. Héctor D. Menéndez 0001 |
CEC | 1 |
| 2019 | Dorylus: An Ant Colony Based Tool for Automated Test Case Generation
Dan Bruce, Héctor D. Menéndez 0001, David Clark 0001 |
SSBSE | 2 |
| 2019 | The arms race: Adversarial search defeats entropy used to detect malwareabstractMalware creators have been getting their way for too long now. String-based similarity measures can leverage ground truth in a scalable way and can operate at a level of abstraction that is difficult to combat from the code level. At the string level, information theory and, specifically, entropy play an important role related to detecting patterns altered by concealment strategies, such as polymorphism or encryption. Controlling the entropy levels in different parts of a disk resident executable allows an analyst to detect malware or a black hat to evade the detection. This paper shows these two perspectives into two scalable entropy-based tools: EnTS and EEE. EnTS, the detection tool, shows the effectiveness of detecting entropy patterns, achieving 100% precision with 82% accuracy. It outperforms VirusTotal for accuracy on combined Kaggle and VirusShare malware. EEE, the evasion tool, shows the effectiveness of entropy as a concealment strategy, attacking binary-based state of the art detectors. It learns their detection patterns in up to 8 generations of its search process, and increments their false negative rate from range 0–9%, up to the range 90–98.7%. Héctor D. Menéndez 0001, Sukriti Bhattacharya, David Clark 0001, Earl T. Barr |
Expert Syst. Appl. | 1 |
| 2018 | Picking on the family: Disrupting android malware triage by forcing misclassificationabstractMachine learning classification algorithms are widely applied to different malware analysis problems because of their proven abilities to learn from examples and perform relatively well with little human input. Use cases include the labelling of malicious samples according to families during triage of suspected malware. However, automated algorithms are vulnerable to attacks. An attacker could carefully manipulate the sample to force the algorithm to produce a particular output. In this paper we discuss one such attack on Android malware classifiers. We design and implement a prototype tool, called IagoDroid, that takes as input a malware sample and a target family, and modifies the sample to cause it to be classified as belonging to this family while preserving its original semantics. Our technique relies on a search process that generates variants of the original sample without modifying their semantics. We tested IagoDroid against RevealDroid, a recent, open source, Android malware classifier based on a variety of static features. IagoDroid successfully forces misclassification for 28 of the 29 representative malware families present in the DREBIN dataset. Remarkably, it does so by modifying just a single feature of the original malware. On average, it finds the first evasive sample in the first search iteration, and converges to a 100% evasive population within 4 iterations. Finally, we introduce RevealDroid*, a more robust classifier that implements several techniques proposed in other adversarial learning domains. Our experiments suggest that RevealDroid* can correctly detect up to 99% of the variants generated by IagoDroid. Alejandro Calleja, Alejandro Martín, Héctor D. Menéndez 0001, Juan Tapiador, David Clark 0001 |
Expert Syst. Appl. | 3 |
| 2017 | Analysing temporal performance profiles of UAV operators using time series clustering
Víctor Rodríguez-Fernández, Héctor D. Menéndez 0001, David Camacho |
Expert Syst. Appl. | 2 |
| 2017 | MOCDroid: multi-objective evolutionary classifier for Android malware detection
Alejandro Martín, Héctor D. Menéndez 0001, David Camacho |
Soft Comput. | 2 |
| 2016 | Genetic boosting classification for malware detectionabstractIn the last few years virus writers have made use of new obfuscation techniques with the aim of hindering malware in order to difficult their detection by Anti-Virus engines. Strategies to reverse this trend involve executing potentially malicious programs and monitor the actions they perform in runtime, what is known as dynamic analysis. In this paper we present a method able to reach a high accuracy rate without using this kind of analysis. Instead we use a static analysis approach, which discards those samples that cannot be classified with enough certainty and need, certainly, a dynamic analysis. The K-means clustering algorithm has been used to group samples into regions according to their features. Then a boosting process, guided by a genetic algorithm, is executed in each region that are evaluated using a test dataset discarding those regions which do not reach a minimum accuracy threshold. Alejandro Martín, Héctor D. Menéndez 0001, David Camacho |
CEC | 2 |
| 2015 | GANY: A genetic spectral-based Clustering algorithm for Large Data AnalysisabstractRecently, Data analysis is one of the most growing fields. The big amounts of data are making their analysis a really challenging area. The most relevant techniques are mainly divided in two sub-domains: Classification and Clustering. Even though Classification is currently growing and evolving, one of the promising techniques to deal with the Large Data Analysis is Clustering, because Classification needs human supervision, which makes the analysis more expensive. Clustering is a blind process used to group data by similarity. Currently, the most relevant methods are those based on manifold identification. The main idea behind these techniques is to group data using the form they define in the space. In order to achieve this goal, there are several techniques based on Spectral Analysis which deal with this problem. However, these techniques are not suitable for Large Data, due to they require a lot of memory to determine the groups. Besides, there are some problems of local minima convergence in these techniques which are common in statistical methodologies. This work is focused on combining Genetic Algorithms with spectral-based methodologies to deal with the Large Data Analysis problem. Here, we will combine the Nyström method with the Spectrum to generate an approximation of the problem to an accurate summary of the search space. Also a genetic algorithm is used to reduce the local minimum convergence problem in the new search space. The performance of this methodology has been evaluated using the accuracy with both, synthetic and real-world datasets extracted from the literature. Héctor D. Menéndez 0001, David Camacho |
CEC | 1 |
| 2015 | User Profile Analysis for UAV Operators in a Simulation Environment
Víctor Rodríguez-Fernández, Héctor D. Menéndez 0001, David Camacho |
ICCCI (1) | 2 |
| 2015 | A tutorial on manifold clustering using genetic algorithmsabstractAutomatic Manifold identification is currently a challenging problem in Machine Learning. This process consists on separating a dataset blindly, according to the form defined by the data instances in the space. Data are discriminated in groups defined by their form. These approaches are usually focused on continuity-based methods where the manifold follows a continuity criterion. Currently, clustering techniques try to deal with the discrimination process, but there are a few algorithms that can generate an accurate and robust discrimination. This tutorial aims to present new different approaches, specially focused on Genetic Algorithms, which can deal with these problems. Héctor D. Menéndez 0001 |
INISTA | 1 |
| 2014 | A Co-Evolutionary Multi-Objective approach for a K-adaptive graph-based clustering algorithmabstractClustering is a field of Data Mining that deals with the problem of extract knowledge from data blindly. Basically, clustering identifies similar data in a dataset and groups them in sets named clusters. The high number of clustering practical applications has made it a fertile research topic with several approaches. One recent method that is gaining popularity in the research community is Spectral Clustering (SC). It is a clustering method that builds a similarity graph and applies spectral analysis to preserve the data continuity in the cluster. This work presents a new algorithm inspired by SC algorithm, the Co-Evolutionary Multi-Objective Genetic Graph-based Clustering (CEMOG) algorithm, which is based on the Multi-Objective Genetic Graph-based Clustering (MOGGC) algorithm and extends it by introducing an adaptative number of clusters. CEMOG takes an island-model approach where each island keeps a population of candidate solutions for kiclusters. Individuals in the islands can migrate to encourage genetic diversity and the propagation of individuals around promising search regions. This new approach shows its competitive performance, compared to several classical clustering algorithms (EM, SC and K-means), through a set of experiments involving synthetic and real datasets. Héctor D. Menéndez 0001, David F. Barrero, David Camacho |
IEEE Congress on Evolutionary Computation | 1 |
| 2014 | Combining graph connectivity and genetic clustering to improve biomedical summarizationabstractAutomatic summarization is emerging as a feasible instrument to help biomedical researchers to access online literature and face information overload. The Natural Language Processing community is actively working toward the development of effective summarization applications; however, automatic summaries are sometimes less informative than the user needs. In this work, our aim is to improve a summarization graph-based process combining genetic clustering with graph connectivity information. In this way, while genetic clustering allows us to identify the different topics that are dealt with in a document, connectivity information (in particular, degree centrality) allows us to asses and exploit the relevance of the different topics. Our automatic summaries are compared with others produced by commercial and research applications, to demonstrate the appropriateness of using this combination of techniques for automatic summarization. Héctor D. Menéndez 0001, Laura Plaza, David Camacho |
IEEE Congress on Evolutionary Computation | 1 |
| 2014 | Combining Time Series and Clustering to Extract Gamer Profile Evolution
Héctor D. Menéndez 0001, Rafael Vindel, David Camacho |
ICCCI | 1 |
| 2014 | TweetSemMiner: A Meta-Topic Identification Model for Twitter Using Semantic Analysis
Héctor D. Menéndez 0001, Carlos Delgado-Calle, David Camacho |
IDEAL | 1 |
| 2014 | A Multi-Objective Graph-based Genetic Algorithm for image segmentationabstractImage Segmentation is one of the most challenging problems in Computer Vision. This process consists in dividing an image in different parts which share a common property, for example, identify a concrete object within a photo. Different approaches have been developed over the last years. This work is focused on Unsupervised Data Mining methodologies, specially on Graph Clustering methods, and their application to previous problems. These techniques blindly divide the image into different parts according to a criterion. This work applies a Multi-Objective Genetic Algorithm in order to perform good clustering results comparing to classical and modern clustering algorithms. The algorithm is analysed and compared against different clustering methods, using a precision and recall evaluation, and the Berkeley Image Database to carry out the experimental evaluation. Héctor D. Menéndez 0001, David Camacho |
INISTA | 1 |
| 2014 | A Genetic Graph-Based Approach for Partitional ClusteringabstractClustering is one of the most versatile tools for data analysis. In the recent years, clustering that seeks the continuity of data (in opposition to classical centroid-based approaches) has attracted an increasing research interest. It is a challenging problem with a remarkable practical interest. The most popular continuity clustering method is the spectral clustering (SC) algorithm, which is based on graph cut: It initially generates a similarity graph using a distance measure and then studies its graph spectrum to find the best cut. This approach is sensitive to the parameters of the metric, and a correct parameter choice is critical to the quality of the cluster. This work proposes a new algorithm, inspired by SC, that reduces the parameter dependency while maintaining the quality of the solution. The new algorithm, named genetic graph-based clustering (GGC), takes an evolutionary approach introducing a genetic algorithm (GA) to cluster the similarity graph. The experimental validation shows that GGC increases robustness of SC and has competitive performance in comparison with classical clustering methods, at least, in the synthetic and real dataset used in the experiments. Héctor D. Menéndez 0001, David F. Barrero, David Camacho |
Int. J. Neural Syst. | 1 |
| 2013 | A Multi-Objective Genetic Graph-Based Clustering algorithm with memory optimizationabstractClustering is one of the most versatile tools for data analysis. Over the last few years, clustering that seeks the continuity of data (in opposition to classical centroid-based approaches) has attracted an increasing research interest. It is a challenging problem with a remarkable practical interest. The most popular continuity clustering method is the Spectral Clustering algorithm, which is based on graph cut: it initially generates a Similarity Graph using a distance measure and then uses its Graph Spectrum to find the best cut. Memory consuption is a serious limitation in that algorithm: The Similarity Graph representation usually requires a very large matrix with a high memory cost. This work proposes a new algorithm, based on a previous implementation named Genetic Graph-based Clustering (GGC), that improves the memory usage while maintaining the quality of the solution. The new algorithm, called Multi-Objective Genetic Graph-based Clustering (MOGGC), uses an evolutionary approach introducing a Multi-Objective Genetic Algorithm to manage a reduced version of the Similarity Graph. The experimental validation shows that MOGGC increases the memory efficiency, maintaining and improving the GGC results in the synthetic and real datasets used in the experiments. An experimental comparison with several classical clustering methods (EM, SC and K-means) has been included to show the efficiency of the proposed algorithm. Héctor D. Menéndez 0001, David F. Barrero, David Camacho |
IEEE Congress on Evolutionary Computation | 1 |
| 2013 | Extracting Collective Trends from Twitter Using Social-Based Data Mining
Gema Bello Orgaz, Héctor D. Menéndez 0001, Shintaro Okazaki, David Camacho |
ICCCI | 2 |
| 2012 | A Genetic Graph-Based Clustering Algorithm
Héctor D. Menéndez 0001, David Camacho |
IDEAL | 1 |
| 2012 | Adaptive k-Means Algorithm for Overlapped Graph ClusteringabstractThe graph clustering problem has become highly relevant due to the growing interest of several research communities in social networks and their possible applications. Overlapped graph clustering algorithms try to find subsets of nodes that can belong to different clusters. In social network-based applications it is quite usual for a node of the network to belong to different groups, or communities, in the graph. Therefore, algorithms trying to discover, or analyze, the behavior of these networks needed to handle this feature, detecting and identifying the overlapped nodes. This paper shows a soft clustering approach based on a genetic algorithm where a new encoding is designed to achieve two main goals: first, the automatic adaptation of the number of communities that can be detected and second, the definition of several fitness functions that guide the searching process using some measures extracted from graph theory. Finally, our approach has been experimentally tested using the Eurovision contest dataset, a well-known social-based data network, to show how overlapped communities can be found using our method. Gema Bello Orgaz, Héctor D. Menéndez 0001, David Camacho |
Int. J. Neural Syst. | 2 |
| 2011 | Predicting Performance in Team Games - The Automatic Coach
Guillermo Jiménez-Díaz, Héctor D. Menéndez 0001, David Camacho, Pedro A. González-Calero |
ICAART (1) | 2 |
| 2011 | Using the Clustering Coefficient to Guide a Genetic-Based Communities Finding Algorithm
Gema Bello Orgaz, Héctor D. Menéndez 0001, David Camacho |
IDEAL | 2 |