Pablo García Bringas

dblp:35/6923 · DBLP profile ↗
← Back
49ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0003-3594-9534ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 2 first-author · 12 since 2021Security and privacy · 12Databases, data management, data science and information retrieval · 10 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 since 2021Human-computer interaction and ubiquitous computing · 4Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Systems, architecture and hardware · 2Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Bayesian generation of synthetic datasets for machine-learning tasks: a performance study
abstract
Performing Machine Learning (ML) tasks on large-scale datasets, as well as simply storing them for subsequent analysis or for long-term archival, could require large computational power. The described approach builds on the technique known as “Bayesian Generation” to produce synthetic datasets in such a way that the probability distribution in the source dataset is maintained as much as possible in the new synthetic ones, even if they are much smaller than the original (large) dataset. In fact, this study investigates the impact of generating smaller synthetic datasets for training ML models in place of the original dataset, adopting a twofold perspective. Firstly, the impact on the effectiveness of ML models trained on these smaller synthetic datasets is assessed. Secondly, the amount of computational resources required to generate the synthetic datasets, train ML models on them, and perform the testing phase is measured. Specifically, both execution time and main memory usage are taken into account. Finally, this research work shows that the loss in terms of effectiveness remains consistently limited and stable, and it identifies the scenarios and ML techniques for which incorporating the generation of small synthetic datasets into the ML pipeline can be beneficial for practical deployment in environments with constrained computational resources, such as mobile or industrial devices.
Paolo Fosci, Javier Nieves, Giuseppe Psaila, Jacopo Boffelli, Pablo García Bringas
Neurocomputing5
2026 On the inherent robustness of one-stage object detection against out-of-distribution data
Aitor Martínez-Seras, Javier Del Ser, Aitzol Olivares-Rad, Alain Andres, Pablo García Bringas
Neural Networks5
2025 A generative AI model for modelling and creating an emotion-aware database: Design, and analytics
abstract
In this study, we illustrate an ongoing work regarding the building of an ad hoc textual dataset for dialogues between a user and an agent of a customer care center. The user is required to express a given emotion and to use a specific level of the Common European Framework of Reference for Languages (CEFR). Once the criteria had been defined, we asked ChatGPT 3.5 to generate dialogues. We analyzed them to ensure an acceptable level of quality for the dataset. Furthermore, we measured the language complexity used both by the user and the agent through different readability measures. Moreover, we extracted from each turn of each dialogue the attitude that ChatGPT exhibited for each dialogue turn; the interaction patterns were then analyzed with algorithm Sequential Pattern Mining (SPM) to extract typical interaction mechanisms occurring in specific emotional contexts. The generated dataset of dialogues, enriched with the readability measures and the extracted patterns, can be exploited as a system that facilitates the interaction between humans and artificial agents in specific contexts and it can constitute the basis of a reinforcement learning system for improving the effectiveness of the dialogues. • We focus on a model based on generative AI for modelling and creating a database that supports emotion recognition and handling. • We define a general architecture for our generation framework. • We provide a comprehensive experimental evaluation of the proposed model.
Alfredo Cuzzocrea, Giovanni Pilato, Pablo García Bringas
Neurocomputing3
2024 Creating, Using and Assessing a Generative-AI-Based Human-Chatbot-Dialogue Dataset with User-Interaction Learning Capabilities
abstract
The study illustrates a first step towards an ongoing work aimed at developing a dataset of dialogues potentially useful for customer service conversation management between humans and AI chatbots. The approach exploits ChatGPT 3.5 to generate dialogues. One of the requirements is that the dialogue is characterized by a specific language proficiency level of the user; the other one is that the user expresses a specific emotion during the interaction. The generated dialogues were then evaluated for overall quality. The complexity of the language used by both humans and AI agents, has been evaluated by using standard complexity measurements. Furthermore, the attitudes and interaction patterns exhibited by the chatbot at each turn have been stored for further detection of common conversation patterns in specific emotional contexts. The methodology could improve human-AI dialogue effectiveness and serve as a basis for systems that can learn from user interactions.
Alfredo Cuzzocrea, Giovanni Pilato, Pablo García Bringas
IEEE Big Data3
2024 A Systematic Literature Review of Decentralized Applications in Web3: Identifying Challenges and Opportunities for Blockchain Developers
abstract
The Internet has opened the floor to stakeholders by redefining the way of organizing, communicating, and collaborating that was initiated by the Web’s development. The advancement of the World Wide Web is an outright phenomenon and significant and we witnessed the evolution of the Web. As decentralized technologies continue to gain traction, Web3, or the decentralized internet, has emerged as a promising approach to enable a more secure, transparent, and privacy-preserved digital landscape. In this paper, we thoroughly conduct a systematic study to explore the challenges and opportunities encountered by blockchain developers in the context of decentralized applications (dApps) in Web3. We analyze a set of peer-reviewed research articles, whitepapers, and technical reports and present an in-depth understanding of the current state of Web3 development and its implications. Our finding indicates the opportunities that Web3 can facilitate, such as expanded use cases, enhanced security and privacy, decentralized infrastructure, and the potential for enabling inclusive development resources for blockchain developers. Additionally, we highlight various challenges that blockchain developers deal with including scalability, security, privacy, interoperability, and the need for standardized tools and frameworks along with various challenges in the software development lifecycle (SDLC). While there are significant challenges to overcome, the potential benefits of Web3 are substantial and could lead to a more inclusive, secure, and transparent digital ecosystem. Furthermore, we emphasize the importance of continued research, collaboration, and innovation among stakeholders to address the identified challenges and capitalize on Web3’s opportunities.
Md. Jobair Hossain Faruk, Pratusha Raya, Md Kamrul Siam, Jerry Q. Cheng, Hossain Shahriar, Alfredo Cuzzocrea, Pablo García Bringas
IEEE Big Data7
2024 On the Improvement of Generalization and Stability of Forward-Only Learning via Neural Polarization
abstract
Forward-only learning algorithms have recently gained attention as alternatives to gradient backpropagation, replacing the backward step of this latter solver with an additional contrastive forward pass. Among these approaches, the so-called Forward-Forward Algorithm (FFA) has been shown to achieve competitive levels of performance in terms of generalization and complexity. Networks trained using FFA learn to contrastively maximize a layer-wise defined goodness score when presented with real data (denoted as positive samples) and to minimize it when processing synthetic data (corr. negative samples). However, this algorithm still faces weaknesses that negatively affect the model accuracy and training stability, primarily due to a gradient imbalance between positive and negative samples. To overcome this issue, in this work we propose a novel implementation of the FFA algorithm, denoted as Polar-FFA, which extends the original formulation by introducing a neural division (polarization) between positive and negative instances. Neurons in each of these groups aim to maximize their goodness when presented with their respective data type, thereby creating a symmetric gradient behavior. To empirically gauge the improved learning capabilities of our proposed Polar-FFA, we perform several systematic experiments using different activation and goodness functions over image classification datasets. Our results demonstrate that Polar-FFA outperforms FFA in terms of accuracy and convergence speed. Furthermore, its lower reliance on hyperparameters reduces the need for hyperparameter tuning to guarantee optimal generalization capabilities, thereby allowing for a broader range of neural network configurations.
Erik Terres Escudero, Javier Del Ser, Pablo García Bringas
ECAI3
2024 Managing the unknown in machine learning: Definitions, related areas, recent advances, and prospects
abstract
In the rapidly evolving domain of machine learning, the ability to adapt to unforeseen circumstances and novel data types is of paramount importance. The deployment of Artificial Intelligence is progressively aimed at more realistic and open scenarios where data, tasks, and conditions are variable and not fully predetermined, and therefore where a closed set assumption cannot be hold. In such evolving environments, machine learning is asked to be autonomous, continuous, and adaptive, requiring effective management of uncertainty and the unknown to fulfill expectations. In response, there is a vigorous effort to develop a new generation of models, which are characterized by enhanced autonomy and a broad capacity to generalize, enabling them to perform effectively across a wide range of tasks. The field of machine learning in open set environments poses many challenges and also brings together different paradigms, some traditional but others emerging, where the overlapping and confusion between them makes it difficult to distinguish them or give them the necessary relevance. This work delves into the frontiers of methodologies that thrive in these open set environments, by identifying common practices, limitations, and connections between the paradigms Open-Ended Learning, Open-World Learning, Open Set Recognition, and other related areas such as Continual Learning, Out-of-Distribution detection, Novelty Detection, and Active Learning. We seek to easy the understanding of these fields and their common roots, uncover open problems and suggest several research directions that may motivate and articulate future efforts towards more robust and autonomous systems.
Marcos Barcina-Blanco, Jesus L. Lobo, Pablo García Bringas, Javier Del Ser
Neurocomputing3
2024 Deep learning applications on cybersecurity: A practical approach
Alberto Miranda-García, Agustin Zubillaga Rego, Iker Pastor-López, Borja Sanz 0001, Alberto Tellaeche, José Gaviria de la Puerta, Pablo García Bringas
Neurocomputing7
2023 Let's do it right the first time: Survey on security concerns in the way to quantum software engineering
abstract
Quantum computing is no longer a promise of the future but a rapidly evolving reality. Advances in quantum hardware are making it possible to make tangible a computational reality that until now was only theoretical. The proof of this is that development languages and platforms are appearing that bring physical principles closer to developers, making it feasible to begin to propose, in different areas of society, solutions to problems that until now were unsolvable. However, security vulnerabilities are also emerging that could hinder the progress of quantum computing, as well as its transition and development in industry. For this reason, this article proposes a review of some of the first artefacts that are emerging in the field of quantum computing. From this analysis, we begin to identify possible security issues that could become potential vulnerabilities in the quantum software of tomorrow. Likewise, and following the experience in classical software development, the testing technique is analysed as a possible candidate for improving security in quantum software development. Following the principles of Quantum Software Engineering, we are aware of the lack of tools, techniques and knowledge necessary to guarantee the development of quantum software in the immediate future. Therefore, this article aims to offer some first clues on what would be a roadmap to guarantee secure quantum software development.
Danel Arias, Ignacio García Rodríguez de Guzmán, Moisés Rodríguez 0001, Erik Terres Escudero, Borja Sanz 0001, José Gaviria de la Puerta, Iker Pastor-López, Agustín Zubillaga, Pablo García Bringas
Neurocomputing9
2022 Can Post-hoc Explanations Effectively Detect Out-of-Distribution Samples?
abstract
Today there is consensus around the importance of explainability as a mandatory feature in practical deployments of Artificial Intelligence (AI) models. Most research activity reported so far in the eXplainable AI (XAI) research arena has stressed on proposing new techniques for eliciting such explanations, together with different approaches for measuring their effectivity to increase the trustworthiness of the audience for which explanations are furnished. However, alternative uses of explanations beyond their original purpose have been very scarcely explored. In this work we investigate whether local explanations can be utilized for detecting Out-of-Distribution (OoD) test samples in machine learning classifiers, i.e., to identify whether query examples of an already trained classification model can be thought to belong to the distribution of the training data. To this end, we devise and assess the performance of a clustering-based OoD detection approach that exemplifies how heatmaps produced by well-established local explanation methods can be of further use than explaining individual predictions issued by the model under analysis. The overarching purpose of this work is not only to expose the benefits and the limits of the proposed XAI-based OoD approach, but also to point out the enormous potential of post-hoc explanations beyond easing the interpretability of black-box models themselves.
Aitor Martínez-Seras, Javier Del Ser, Pablo García Bringas
FUZZ-IEEE3
2021 Quality assessment methodology based on machine learning with small datasets: Industrial castings defects
Iker Pastor-López, Borja Sanz 0001, Alberto Tellaeche, Giuseppe Psaila, José Gaviria de la Puerta, Pablo García Bringas
Neurocomputing6
2021 Detecting malicious Android applications based on the network packets generated
José Gaviria de la Puerta, Iker Pastor-López, Igone Porto Gómez, Borja Sanz 0001, Pablo García Bringas
Neurocomputing5
2019 Can BlockChain Technology Provide Information Systems with Trusted Database? The Case of HyperLedger Fabric
Pablo García Bringas, Iker Pastor-López, Giuseppe Psaila
FQAS1
2019 Special issue HAIS 2015: Recent advancements in hybrid artificial intelligence systems and its application to real-world problems
Pablo García Bringas, Igor Santos, Enrique Onieva, Eneko Osaba, Héctor Quintián, Emilio Corchado
Neurocomputing1
2018 Special issue SOCO 2014: Recent advancements in soft computing and its application in industrial and environmental problems
abstract
A feedback solution for approximate optimal scheduling of switched systems with autonomous subsystems and continuous-time dynamics is presented. The proposed solution is based on policy iteration algorithm which provides the optimal switching schedule. Algorithms for offline, online, and concurrent implementation of the proposed solution are presented. For online and concurrent training, gradient descent training laws are used and the performance of the training laws is analyzed. The effectiveness of the presented algorithms is verified through numerical simulations.
Pablo García Bringas, André C. P. L. F. de Carvalho, Ajith Abraham, Álvaro Herrero 0001, Héctor Quintián, Emilio Corchado
Neurocomputing1
2016 RAMBO: Run-Time Packer Analysis with Multiple Branch Observation
Xabier Ugarte-Pedrero, Davide Balzarotti, Igor Santos, Pablo García Bringas
DIMVA4
2015 A Comprehensive Interaction Model for ASystems
abstract
In this extended poster, we present a model that aims to provide developers with an extensive and extensible set of context-aware interaction techniques, greatly facilitating the creation of meaningful AR-based user experiences. To provide a complete view of the model, we detail the different aspects that form its theoretical foundations, while also discussing several considerations for its correct implementation.
Mikel Salazar, Carlos Laorden, Pablo García Bringas
ISMAR3
2015 SoK: Deep Packer Inspection: A Longitudinal Study of the Complexity of Run-Time Packers
abstract
Run-time packers are often used by malware-writers to obfuscate their code and hinder static analysis. The packer problem has been widely studied, and several solutions have been proposed in order to generically unpack protected binaries. Nevertheless, these solutions commonly rely on a number of assumptions that may not necessarily reflect the reality of the packers used in the wild. Moreover, previous solutions fail to provide useful information about the structure of the packer or its complexity. In this paper, we describe a framework for packer analysis and we propose a taxonomy to measure the runtime complexity of packers. We evaluated our dynamic analysis system on two datasets, composed of both off-the-shelf packers and custom packed binaries. Based on the results of our experiments, we present several statistics about the packers complexity and their evolution over time.
Xabier Ugarte-Pedrero, Davide Balzarotti, Igor Santos, Pablo García Bringas
IEEE Symposium on Security and Privacy4
2015 Combined special issue SOCO 2012-2013: Recent advancements in soft computing and its application in industrial and environmental problems
Emilio Corchado, Ajith Abraham, Václav Snásel, Pablo García Bringas, Ivan Zelinka, Héctor Quintián
Neurocomputing4
2014 Procedural Playable Cave Systems Based on Voronoi Diagram and Delaunay Triangulation
abstract
The volumetric approach for terrain representation is a technique used by several video-games and other graphic applications to manage both surface and geological data from the virtual world. To enhance the exploration experience and due to the amount of data required by this approach, volumetric terrains are usually generated with procedural methods. Nevertheless, one of the main issues of those methods is the generation of cave systems with playable features and a natural appearance. In this paper we propose a new method to generate playable cave systems for 2D and 3D volumetric terrains, based on Voronoi diagrams and Delaunay triangulations. Our approach is completely customizable by the designer by a set of parameters directly related to the cave itself avoiding technical concepts. Additionally, the method runs in a completely independent way, with no interactive steps.
Aitor Santamaría-Ibirika, Xabier Cantero, Sergio Huerta, Igor Santos, Pablo García Bringas
CW5
2014 On the adoption of anomaly detection for packed executable filtering
Xabier Ugarte-Pedrero, Igor Santos, Iván García-Ferreira, Sergio Huerta, Borja Sanz 0001, Pablo García Bringas
Comput. Secur.6
2014 Study on the effectiveness of anomaly detection for spam filtering
Carlos Laorden, Xabier Ugarte-Pedrero, Igor Santos, Borja Sanz 0001, Javier Nieves, Pablo García Bringas
Inf. Sci.6
2014 Procedural approach to volumetric terrain generation
Aitor Santamaría-Ibirika, Xabier Cantero, Mikel Salazar, Jaime Devesa, Igor Santos, Sergio Huerta, Pablo García Bringas
Vis. Comput.7
2013 JURD: Joiner of Un-Readable Documents to reverse tokenization attacks to content-based spam filters
abstract
Spam has become a major issue in computer security because it is a channel for threats such as computer viruses, worms and phishing. More than 85% of received e-mails are spam. Historical approaches to combating these messages, including simple techniques like sender blacklisting or the use of e-mail signatures, are no longer completely reliable. Many current solutions feature machine-learning algorithms trained using statistical representations of the terms that most commonly appear in such e-mails. However, there are attacks that can subvert the filtering capabilities of these methods. Tokenization attacks, in particular, insert characters that create divisions within words, causing incorrect representations of e-mails. In this paper, we introduce a new method that reverses the effects of tokenization attacks. Our method processes e-mails iteratively by considering possible words, starting from the first token and compares the word candidates with a common dictionary to which spam words have been previously added. We provide an empirical study of how tokenization attacks affect the filtering capability of a Bayesian classifier and we show that our method can reverse the effects of tokenization attacks.
Igor Santos, Carlos Laorden, Borja Sanz 0001, Pablo García Bringas
CCNC4
2013 Volumetric Virtual Worlds with Layered Terrain Generation
abstract
Procedural terrain generators are widely used in computer games, simulators and other virtual environments. Traditionally, these methods have been focused on generating realistic landscapes and their elements, like mountains or rivers. In contrast, there is a movement in the industry which employs virtual worlds using volumetric terrains. This movement uses its owns methods to generate terrains, but they lack of customization features and are completely centered around their specific problem domains. In this paper, we propose an alternative approach to generate procedural volumetric terrains specially suited for shared worlds. Our approach generates completely customizable volumetric terrains with layered materials and other features (such as mineral veins, material mixtures and underground material flow). We enable the designer to specify the characteristics of the terrain and its materials with intuitive parameters. Moreover, our method uses a specific representation for the terrain based on stacked material structures, reducing the memory size and making it easier to share.
Aitor Santamaría-Ibirika, Xabier Cantero, Mikel Salazar, Jaime Devesa, Pablo García Bringas
CW5
2013 Enhancing cybersecurity learning through an augmented reality-based serious game
abstract
As social networks and always-connected mobile devices grow in popularity, the control over personal information weakens. This is especially true for teenagers between 15 and 18 years old, one of the population groups that shares more information online, but also the most unaware of the risks associated with this activity. For this reason, many institutions have developed programs to educate the students in the correct use of the new communication mediums. However, the concepts about information security require a lot of expert knowledge and are very difficult to explain appropriately. In this paper, we present a serious game designed to enhance an information security presentation aimed at high school students. This is achieved through the use of augmented reality to give shape and form to the intangible cybersecurity concepts and allow the students to interact with them using the same rule set that was explained during the presentation.
Mikel Salazar, José Gaviria de la Puerta, Carlos Laorden, Pablo García Bringas
EDUCON4
2013 MADS: Malicious Android Applications Detection through String Analysis
Borja Sanz 0001, Igor Santos, Javier Nieves, Carlos Laorden, Iñigo Alonso 0001, Pablo García Bringas
NSS6
2013 Filtering Trolling Comments through Collective Classification
Jorge de-la-Peña-Sordo, Igor Santos, Iker Pastor-López, Pablo García Bringas
NSS4
2013 Instance-based Anomaly Method for Android Malware Detection
Borja Sanz 0001, Igor Santos, Xabier Ugarte-Pedrero, Carlos Laorden, Javier Nieves, Pablo García Bringas
SECRYPT6
2013 Mama: manifest Analysis for Malware Detection in Android
abstract
The use of mobile phones has increased because they offer nearly the same functionality as a personal computer. In addition, the number of applications available for Android-based mobile devices has increased. Google offers programmers the opportunity to upload and sell applications in the Android Market, but malware writers upload their malicious code there. In light of this background, we present here manifest analysis for malware detection in Android (MAMA), a new method that extracts several features from the Android manifest of the applications to build machine learning classifiers and detect malware.
Borja Sanz 0001, Igor Santos, Carlos Laorden, Xabier Ugarte-Pedrero, Javier Nieves, Pablo García Bringas, Gonzalo Álvarez
Cybern. Syst.6
2013 Opcode sequences as representation of executables for data-mining-based unknown malware detection
Igor Santos, Felix Brezo, Xabier Ugarte-Pedrero, Pablo García Bringas
Inf. Sci.4
2012 On the study of anomaly-based spam filtering using spam as representation of normality
abstract
In previous work, we presented the first spam filtering method based on anomaly detection that reduces the necessity of labelling spam messages and only employs the representation of legitimate e-mails. This method achieved high accuracy rates detecting spam while maintaining a low false positive rate and reducing the effort produced by labelling spam. In this paper, we study the performance of our previous method when using spam messages to represent normality.
Carlos Laorden, Xabier Ugarte-Pedrero, Igor Santos, Borja Sanz 0001, Javier Nieves, Pablo García Bringas
CCNC6
2012 On the automatic categorisation of android applications
abstract
The presence of mobile devices has increased in our lives offering almost the same functionality as a personal computer. Android devices have appeared lately and, since then, the number of applications available for this operating system have increased exponentially. Google already has its Android Market where applications are offered and, as happens with every popular media, is prone to misuse. A malware writer may insert a malicious application into this market without being noticed. Indeed, there are already several cases of Android malware within the Android Market. Therefore, an approach that can automatically characterise the different types of applications can be helpful for both organising the Android Market and detecting fraudulent or malicious applications. In this paper, we propose a new method for categorising Android applications through machine-learning techniques. To represent each application, our method extracts different feature sets: (i) the frequency of occurrence of the printable strings, (ii) the different permissions of the application itself and (iii) the permissions of the application extracted from the Android Market. We evaluate this approach of automatically categorisation of Android applications and show that achieves a high performance.
Borja Sanz 0001, Igor Santos, Carlos Laorden, Xabier Ugarte-Pedrero, Pablo García Bringas
CCNC5
2012 Countering entropy measure attacks on packed software detection
abstract
Malware writers usually employ several techniques to evade detection. For the last years, the number of variants detected each day has increased significantly. Traditional approaches such as signature scanning, one of the most common techniques employed by anti-virus companies, are becoming inefficient for the high amount of samples found in the wild. In order to bypass this kind of filters, malware writers usually obfuscate and transform the code of their creations. One of the methods employed is executable packing, which consists in compressing or ciphering the real malicious code, and injecting a decryption routine into the executable that will load and decompress it at run-time. Entropy is a common heuristic for the detection of packed executables. High entropy values indicate a random distribution of the bytes that compose the executable, a property very common in compressed and ciphered data. Unfortunately, this entropy measure can be altered by different techniques that modify randomness. In this paper, we detail various attacks found on real Zeus family samples, one of the most powerful and spread malware families at this moment, which are protected by custom made packers. In addition, we describe a method for obtaining an alternative entropy measure more resilient to these techniques, and evaluate it for the classification of packed/not-packed executables, obtaining satisfactory detection and false positive rates.
Xabier Ugarte-Pedrero, Igor Santos, Borja Sanz 0001, Carlos Laorden, Pablo García Bringas
CCNC5
2012 Supervised classification of packets coming from a HTTP botnet
abstract
The posibilities that the management of a vast amount of computers and/or networks offer, is attracting an increasing number of malware writers. In this document, the authors propose a methodology thought to detect malicious botnet traffic, based on the analysis of the packets flow that circulate in the network. This objective is achieved by means of the parametrization of the static characteristics of packets, which are lately analysed using supervised machine learning techniques focused on traffic labelling so as to face proactively to the huge volume of information nowadays filters work with.
Felix Brezo, José Gaviria de la Puerta, Xabier Ugarte-Pedrero, Igor Santos, Pablo García Bringas, David Barroso
CLEI5
2012 Combination of Machine-Learning Algorithms for Fault Prediction in High-Precision Foundries
Javier Nieves, Igor Santos, Pablo García Bringas
DEXA (2)3
2012 Enhanced Topic-based Vector Space Model for semantics-aware spam filtering
Igor Santos, Carlos Laorden, Borja Sanz 0001, Pablo García Bringas
Expert Syst. Appl.4
2012 Automatic categorisation of comments in social news websites
Igor Santos, Jorge de-la-Peña-Sordo, Iker Pastor-López, Patxi Galán-García, Pablo García Bringas
Expert Syst. Appl.5
2011 Boosting Scalability in Anomaly-Based Packed Executable Filtering
Xabier Ugarte-Pedrero, Igor Santos, Pablo García Bringas
Inscrypt3
2011 Anomaly Detection for the Prediction of Ultimate Tensile Strength in Iron Casting Production
Igor Santos, Javier Nieves, Xabier Ugarte-Pedrero, Pablo García Bringas
DEXA (2)4
2011 Semi-supervised learning for packed executable detection
abstract
The term malware is coined to name any software with malicious intentions. One of the methods malware writers use for hiding their creations is executable packing. Packing consists of encrypting or hiding the real code of the executable in such a way that it is decrypted or unhidden in its execution. Widespread solutions to this issue first try to identify the packer used and next apply the corresponding unpacking routine for each packing algorithm. As it happens with malware obfuscations, this approach fails to detect new and custom packers. Generic unpacking is a technique that has been proposed to solve this issue. These methods usually execute the binary in a contained environment or sandbox to retrieve the real code of the packed executable. Because these approaches incur in a high performance overhead, a filter step is required to determine whether an executable is packed or not. Supervised machine-learning approaches have been proposed to handle this filtering step. However, the usefulness of supervised learning is far to be complete because it requires a high amount of packed and not packed executables to be identified and labelled previously. In this paper, we propose a new method for packed executable detection that adopts a well-known semi-supervised learning approach to reduce the labelling requirements of completely supervised approaches. We performed an empirical validation demonstrating that the labelling efforts are lower than when supervised learning is used while the system maintains high accuracy rates.
Xabier Ugarte-Pedrero, Igor Santos, Pablo García Bringas, Mikel Gastesi, José Miguel Esparza
NSS3
2011 Collective Classification for Unknown Malware Detection
Igor Santos, Carlos Laorden, Pablo García Bringas
SECRYPT3
2011 Anomaly-based Spam Filtering
Igor Santos, Carlos Laorden, Xabier Ugarte-Pedrero, Borja Sanz 0001, Pablo García Bringas
SECRYPT5
2011 Using opcode sequences in single-class learning to detect unknown malware
abstract
Malware is any type of malicious code that has the potential to harm a computer or network. The volume of malware is growing at a faster rate every year and poses a serious global security threat. Although signature-based detection is the most widespread method used in commercial antivirus programs, it consistently fails to detect new malware. Supervised machine-learning models have been used to address this issue. However, the use of supervised learning is limited because it needs a large amount of malicious code and benign software to be labelled first. In this study, the authors propose a new method that uses single-class learning to detect unknown malware families. This method is based on examining the frequencies of the appearance of opcode sequences to build a machine-learning classifier using only one set of labelled instances within a specific class of either malware or legitimate software. The authors performed an empirical study that shows that this method can reduce the effort of labelling software while maintaining high accuracy.
Igor Santos, Felix Brezo, Borja Sanz 0001, Carlos Laorden, Pablo García Bringas
IET Inf. Secur.5
2010 Automatic Morphological Categorisation of Carbon Black Nano-aggregates
Juan López-de-Uralde, Iraide Ruiz, Igor Santos, Agustín Zubillaga, Pablo García Bringas, Ana Okariz, Teresa Guraya
DEXA (2)5
2010 Enhanced Foundry Production Control
Javier Nieves, Igor Santos, Yoseba K. Penya, Felix Brezo, Pablo García Bringas
DEXA (1)5
2010 Fraud Detection for Voice over IP Services on Next-Generation Networks
Igor Ruiz-Agundez, Yoseba K. Penya, Pablo García Bringas
WISTP3
2009 Mechanical properties prediction in high-precision foundry production
abstract
Mechanical properties are the attributes of a metal to withstand several forces and tensions. Specifically, ultimate tensile strength is the force a material can resist until it breaks. The only way to examine this mechanical property is the employment of destructive inspections that renders the casting invalid with the subsequent cost increment. In a previous work we showed that modelling the foundry process as a probabilistic constellation of interrelated variables allows Bayesian networks to infer causal relationships. In other words, they may guess the value of a variable (for instance, the value of ultimate tensile strength). Against this background, we present here the first ultimate tensile strength prediction system that, upon the basis of a Bayesian network, is able to foresee the values of this property in order to correct it before the casting is made. Further, we have tested the accuracy and error rate of the system with data of a real foundry.
Javier Nieves, Igor Santos, Yoseba K. Penya, Sendoa Rojas-Lertxundi, Mikel Salazar, Pablo García Bringas
INDIN6
2008 Efficient failure-free foundry production
abstract
Microshrinkages are known as probably the most difficult defects to avoid in high-precission foundry. Depending on the magnitude of this defect, the piece in which it appears must be rejected with the subsequent cost increment. Modelling this environment as a probabilistic constellation of interrelated variables allows Bayesian networks to infer causal relationships. In other words, they may guess the value of a variable (for instance, the presence or not of a defect). Against this background, we present here the first microshrinkage prediction system that, upon the basis of a Bayesian network, is able to foresee the apparition of this defect and to determine whether the piece is still acceptable or not. Further, after testing this system in a real foundry, we discuss the obtained results and present a risk-level-based production methodology that increases the rate of valid manufactured pieces.
Yoseba K. Penya, Pablo García Bringas, Argoitz Zabala
ETFA2