Miguel Vargas Martin

dblp:v/MiguelVargasMartin · DBLP profile ↗
← Back
28ranked-venue papers
1as first author
11since 2021 · last 2025
0000-0001-8169-6836ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 13 · 6 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1
YearPublicationVenuePosition
2025 Context is the Key for LLM-Based Text Segmentation
Amit Maraj, Miguel Vargas Martin
NLDB (1)2
2025 A Per-Bag Suspicion-Based Bagging Strategy for Fighting Poisoning Attacks in Classification
abstract
The wide adoption of machine learning-powered systems in sensitive applications, such as banking for fraud detection, has attracted malicious actors who seek to break and subvert these systems. In this work, we focus on Data Poisoning attacks, which is a well-known type of adversarial attack carried out by an adversary whose goal is to reduce the effectiveness of the learning system. Bagging, a well-known ensemble learning technique that aims to improve performance and reduce the overall system variance, has demonstrated robustness against data poisoning attacks. Bagging has been further extended to include weighted schemes designed to detect outliers and assign lower resampling probabilities to anomalous instances, thereby enhancing the robustness of the standard bagging mechanism. Weighted bagging significantly improves system performance when the dataset is poisoned; however, it often suffers from instability due to the mechanism used to estimate the resampling probabilities. To address this challenge, we propose a novel weight estimation approach that leverages the reconstruction capabilities of autoencoders to identify and down-weight anomalous training samples. In particular, we investigate a specific type of data poisoning attack known as a label-flipping attack, using the widely studied MNIST dataset of handwritten images and conduct experiments using a Convolutional Neural Network (CNN). Our results show that the proposed weighted bagging mechanism consistently outperforms standard bagging under data poisoning levels of up to $50 \%$. To our knowledge, this is the first study to introduce a per-bag anomaly-based weighting mechanism, paving the way for future adaptive ensemble defenses in adversarial machine learning.
Aghoghomena Akasukpe, Tomi Adeyemi, Pooria Madani, Li Yang 0010, Miguel Vargas Martin
PST5
2025 Beyond Rules: How Large Language Models Are Redefining Cryptographic Misuse Detection
Zohaib Masood, Miguel Vargas Martin
SECRYPT2
2024 Coherence Graphs: Bridging the Gap in Text Segmentation with Unsupervised Learning
Amit Maraj, Miguel Vargas Martin, Masoud Makrehchi
NLDB (2)2
2024 Visualizing Differential Privacy: Assessing Infographics' Impact on Layperson Data-sharing Decisions and Comprehension
abstract
Differential privacy (DP) has emerged as a promising approach for protecting users' data in the era of big data and machine learning. Despite its deployment by governments and or-ganizations, the concept of DP remains difficult for non-technical users to comprehend. Visual aids, such as infographics, have the potential to bridge this knowledge gap and enable users to make informed data-sharing decisions. In this paper, we propose to use carefully designed infographics to explain DP and compare their effectiveness with traditional text descriptions. We conducted a vignette survey study with 367 participants on Prolific and found that our static and dynamic infographic designs improved participants' understanding of DP, including its mechanism and implication compared with text descriptions. Our infographics also enhance users' understanding of DP and educate them on whether the privacy budget ∊is exposed when sharing their highly sensitive information. This research contributes to the growing body of literature on designing effective DP descriptions to communicate DP to laypeople to facilitate their data-sharing decisions.
Mst Mahamuda Sarkar Mithila, Fangyi Yu, Miguel Vargas Martin, Shengqian Wang
PST3
2023 Honey, I Chunked the Passwords: Generating Semantic Honeywords Resistant to Targeted Attacks Using Pre-trained Language Models
Fangyi Yu, Miguel Vargas Martin
DIMVA2
2022 Emotion recognition models for companion robots
Ritvik Nimmagadda, Kritika Arora, Miguel Vargas Martin
J. Supercomput.3
2021 A Novel Short-term Post-accident Traffic Prediction Model
abstract
Traffic forecasting at appropriate times is vital for a variety of urban traffic control applications. There are a plethora of factors that can severely affect the performance of such forecasting models, especially those unpredictable events that cannot be learned through previous time steps in the model. One of the important factors that can change the traffic flow pattern abruptly, is accident occurrence. This issue affects more when the severity of the accident is high which would lead to a complete anomaly behavior in the traffic flow pattern. In this paper, we detected the approximate time that the accidents happen through anomaly detection around the accident record and then pulled out the ones that change the traffic pattern suddenly. To predict the traffic flow after a severe accident, we propose a hybrid CNN-LSTM model that takes spatio-temporal matrices as discrete time sequences for each accident. Furthermore, we leverage distances between neighboring nodes and accident location, coupled with static features that are pertained to the accidents. Moreover, we evaluate our model on PeMs dataset which is enhanced with the static features of the accidents from Moosavi’s accident data set [1] [2]. Despite recent trials that are not able to forecast the results when an anomaly event happens, our results demonstrate that our proposed model can learn traffic patterns after the accidents and predict the traffic flow more accurately compared to current models.
Farimasadat Miri, Alireza A. Namanloo, Richard Werner Nelem Pazzi, Miguel Vargas Martin
DCOSS4
2021 A More Effective Sentence-Wise Text Segmentation Approach Using BERT
Amit Maraj, Miguel Vargas Martin, Masoud Makrehchi
ICDAR (4)2
2021 Fool Me Once: A Study of Password Selection Evolution over the Past Decade
abstract
Passwords have been around for many decades and have tenaciously remained the primary means of identification and authentication. Assuming that the communication channel is not intercepted, the strength of security provided by passwords is largely dependent on two factors: password selection and password storage mechanism. While both areas have been looked into by researchers in the past, there is no consensus to suggest whether or not humanity has moved towards choosing stronger passwords, notwithstanding strong password enforcement policies. One of the key reasons behind this shortcoming is the lack of data about individual credentials in leaked datasets, which usually contain only usernames and passwords. To the best of our knowledge, we are the first researchers to enrich the attribute set of any user credential database, thus allowing deeper insights. We outline the method we devised for adding new attributes (time-stamp and source inference) to a dataset of 1.4 billion user credentials. Subsequently, we use our modified dataset to determine how passwords have evolved overtime with respect to strength and whether humankind as a whole has learned from its past mistakes.
Rahul Dubey, Miguel Vargas Martin
PST2
2021 SegmentPerturb: Effective Black-Box Hidden Voice Attack on Commercial ASR Systems via Selective Deletion
abstract
Voice control systems continue becoming more pervasive as they are deployed in mobile phones, smart home devices, automobiles, etc. Commonly, voice control systems have high privileges on the device, such as making a call or placing an order. However, they are vulnerable to voice attacks, which may lead to serious consequences. In this paper, we propose SegmentPerturb which crafts hidden voice commands via inquiring the target models. The general idea of SegmentPerturb is that we separate the original command audio into multiple equal-length segments and apply maximum perturbation on each segment by probing the target speech recognition system. We show that our method is as efficient, and in some aspects outperforms other methods from previous works. We choose four popular speech recognition APIs and one mainstream smart home device to conduct the experiments. Results suggest that this algorithm can generate voice commands which can be recognized by the machine but are hard to understand by a human.
Ganyu Wang, Miguel Vargas Martin
PST2
2019 On the null relationship between personality types and passwords
abstract
We performed a preliminary investigation of the relationship between the Big-five personality traits and the strength of selected and created online security passwords. Five-hundred and ten participants were recruited through MTurk and asked a) to complete a Big-five personality inventory, b) to select from a list of available passwords the one they felt was most secure, and c) to create unique, secure passwords for their own online protection. The security strength of participants' selected/created passwords was rated via Zxcvbn, and evaluated on a variety of additional security-based criteria (e.g., password length, inclusion of special characters). In all cases results failed to identify significant relationship between the strength of the selected/chosen password and any Big-five personality traits (when employing appropriately stringent control for multiple corrections). These null results suggest that other factors beyond an individual's personality may hold greater influence over the strength of selected/created online passwords. We present detailed observations and findings from our experiment, discuss potential considerations for contradictions, and suggest possibilities for future research into the password/personality relationship, including the potential use of enhanced password strength meters and tailored security nudges.
Amit Maraj, Miguel Vargas Martin, Matthew Shane, Mohammad Mannan
PST2
2019 Inside out - A study of users' perceptions of password memorability and recall
Ruba AlOmari, Miguel Vargas Martin, Shane MacDonald, Amit Maraj, Ramiro Liscano, Christopher Bellman
J. Inf. Secur. Appl.2
2018 Reinforcing System-Assigned Passphrases Through Implicit Learning
abstract
People tend to choose short and predictable passwords that are vulnerable to guessing attacks. Passphrases are passwords consisting of multiple words, initially introduced as more secure authentication keys that people could recall. Unfortunately, people tend to choose predictable natural language patterns in passphrases, again resulting in vulnerability to guessing attacks. One solution could be system-assigned passphrases, but people have difficulty recalling them. With the goal of improving the usability of system-assigned passphrases, we propose a new approach of reinforcing system-assigned passphrases using implicit learning techniques. We design and test a system that implements this approach using two implicit learning techniques: contextual cueing and semantic priming. In a 780-participant online study, we explored the usability of 4-word system-assigned passphrases using our system compared to a set of control conditions. Our study showed that our system significantly improves usability of system-assigned passphrases, both in terms of recall rates and login time.
Zeinab Joudaki, Julie Thorpe, Miguel Vargas Martin
CCS3
2018 Classification of EEG Signals Using Neural Networks to Predict Password Memorability
abstract
Feature extraction and classification is a subject of broad and current interest in the brain-computer interfaces (BCIs) community, and remains a challenging task when working with Electroencephalogram (EEG) data, with no agreed upon optimum features set and classifier algorithm. In this 75-participant lab study, we compare different feature extraction methods and classifiers as we investigate the relationship between users' perceptions of the memorability of a number of passwords and the users' EEG data collected using BCIs when presented with these passwords. We asked the participants to rank the comparative memorability of 15 blocks of 5 passwords each, while recording their EEG data. Features from the EEG signals are extracted in three domains, power spectrum from the frequency domain, statistics from the time domain, and wavelet coefficients from the time-frequency domain. The feature subsets are submitted for classification with two classes, most memorable and least memorable, based on the user's perceived memorability of the passwords. Classification performance of Neural Networks and Support Vector Machine algorithms are compared. Results show distinctive features of EEG signals in the two different classes, achieving a classification accuracy of 89%. Our results indicate that features extracted in the time-frequency domain using wavelet transform, and classified using neural networks, resulted in the highest classification performance.
Ruba AlOmari, Miguel Vargas Martin
ICMLA2
2017 Fusion-based hybrid many-objective optimization algorithm
abstract
In the last three decades there have been a number of efficient multi-objective optimization algorithms capable of solving real-world problems. However, due to the complexity of most real-world problems (high-dimensionality of problems, computationally expensive, and unknown function properties) researchers and decision-makers are increasingly facing the challenge of selecting an optimization algorithm capable of solving their hard problems. In this paper, we propose a simple yet efficient hybridization of multi- and many-objective optimization algorithms framework called hybrid many-objective optimization algorithm using fusion of solutions obtained by several many-objective algorithms (fusion) to gain the combined benefits of several algorithms and reducing the challenge of choosing one optimization algorithm to solve complex problems. During the optimization process, the Fusion framework (1) executes all optimization algorithms in parallel, (2) it combines solutions of these algorithms and extracts well-distributed solutions using predefined structured reference points or user-defined reference points, and (3) adaptively selects best-performing algorithm to tackle the problem at different stages of the search process. A case study of the fusion framework by considering GDE3, SMPSO, and SPEA2 as multi-objective optimization algorithms is presented. Experimental results on five unconstrained and four constrained benchmark test problems with three to ten objectives show that the Fusion framework significantly outperforms all algorithms involved in the hybridization process as well as the NSGA-III algorithm in terms of diversity and convergence of obtained solutions. Furthermore, the proposed framework is consistently able to find accurate solutions for all test problems which can be interpreted as its high robustness characteristic.
Amin Ibrahim, Miguel Vargas Martin, Shahryar Rahnamayan, Kalyanmoy Deb
CEC2
2017 Fusion of Many-Objective Non-dominated Solutions Using Reference Points
Amin Ibrahim, Shahryar Rahnamayan, Miguel Vargas Martin, Kalyanmoy Deb
EMO3
2017 Use of Machine Learning for Detection of Unaware Facial Recognition Without Individual Training
abstract
Efforts in detecting unaware facial recognitions using consumer-grade brain-computer interfaces have been able to classify these recognitions with high accuracies using intra-participant datasets. Seeking a more generalized approach to classifying facial recognition, we propose an interparticipant dataset comprised of pre-recorded data where new data can be immediately tested. An experiment was conducted where participants viewed images of faces in a two-day experiment. Participants learned a number of faces for unaware recognition during the following day’s session. Three recognition classes were recorded: no recognition, unaware recognition, and aware recognition. Using data features containing the highest levels of variance, we find that an interparticipant dataset can be used to train a classifier and achieve classification F-scores of 0.71. This suggests that detecting recognition can be achieved through pre-constructed datasets and used more readily in practical applications.
Christopher Bellman, Miguel Vargas Martin
ICMLA2
2017 What Your Brain Says About Your Password: Using Brain-Computer Interfaces to Predict Password Memorability
abstract
Recent advances in brain-computer interfaces (BCI) have enabled them as affordable consumer-grade devices for nonmedical purposes such as academic research, marketing, and entertainment. We report on the possibility of using BCIs to classify passwords into two classes—one class may be deemed as memorable and the other one as non-memorable—based on electroencephalogram (EEG) potentials collected by the BCI upon presenting the passwords to human participants. The memorable set consists of the most commonly used passwords, also known as "worst passwords lists", while the non-memorable set consists of randomly generated strings of characters, symbols, and numbers. When classifying passwords as memorable vs. nonmemorable, a classification accuracy of 76.5% was achieved. We found a positive correlation between password EEG features and password recall. We also report on users' choice of passwords, where 74% of participants were found to inadvertently choose the password with higher elicited voltage, when presented with two passwords to choose from.
Ruba AlOmari, Miguel Vargas Martin, Shane MacDonald, Christopher Bellman, Ramiro Liscano, Amit Maraj
PST2
2017 System-Assigned Passwords You Can't Write Down, But Don't Need To
abstract
We explore the feasibility of Tacit Secrets: systemassigned passwords that you can remember, but cannot write down or otherwise communicate. We design an approach to creating Tacit Secrets based on Contextual Cueing, an implicit learning method previously studied in the cognitive psychology literature. Our feasibility study involving 30 participants indicates that our approach has strong security properties: resistance to brute-force attacks, online attacks, phishing attacks, and some coercion attacks. It also offers protection against leaks from other verifiers as the secrets are system-assigned. Our approach also has a high login success rate and low false positive rates. We explore the trade-offs of different configurations of our design and provide insight into valuable directions for future work.
Zeinab Joudaki, Julie Thorpe, Miguel Vargas Martin
PST3
2016 3D-RadVis: Visualization of Pareto front in many-objective optimization
abstract
In many-objective optimization, visualization of true Pareto front or obtained non-dominated solutions is difficult. A proper visualization tool must be able to show the location, range, shape, and distribution of obtained non-dominated solutions. However, existing commonly used visualization tools in many-objective optimization (e.g., parallel coordinates) fail to show the shape of the Pareto front. In this paper, we propose a simple yet powerful visualization method, called 3-dimensional radial coordinate visualization (3D-RadVis). This method is capable of mapping M-dimensional objective space to a 3-dimensional radial coordinate plot while preserving the relative location of solutions, shape of the Pareto front, distribution of solutions, and convergence trend of an optimization process. Furthermore, 3D-RadVis can be used by decision-makers to visually navigate large many-objective solution sets, observe the evolution process, visualize the relative location of a solution, evaluate trade-off among objectives, and select preferred solutions. The visual effectiveness of the proposed method is demonstrated on widely used many-objective benchmark problems containing variety of Pareto fronts (linear, concave, convex, mixed, and disconnected). In addition, we demonstrated the capability of 3D-RadVis for visual progress tracking of the NSGA-III algorithm through generations. It is worthwhile to mention that a suitable visualization is a crucial prerequisite for an effective interactive optimization.
Amin Ibrahim, Shahryar Rahnamayan, Miguel Vargas Martin, Kalyanmoy Deb
CEC3
2016 EliteNSGA-III: An improved evolutionary many-objective optimization algorithm
abstract
Evolutionary algorithms are the most studied and successful population-based algorithms for solving single- and multi-objective optimization problems. However, many studies have shown that these algorithms fail to perform well when handling many-objective (more than three objectives) problems due to the loss of selection pressure to pull the population towards the Pareto front. As a result, there has been a number of efforts towards developing evolutionary algorithms that can successfully handle many-objective optimization problems without deteriorating the effect of evolutionary operators. A reference-point based NSGA-II (NSGA-III) is one such algorithm designed to deal with many-objective problems, where the diversity of the solution is guided by a number of well-spread reference points. However, NSGA-III still has difficulty preserving elite population as new solutions are generated. In this paper, we propose an improved NSGA-III algorithm, called EliteNSGA-III to improve the diversity and accuracy of the NSGA-III algorithm. EliteNSGA-III algorithm maintains an elite population archive to preserve previously generated elite solutions that would probably be eliminated by NSGA-III's selection procedure. The proposed EliteNSGA-III algorithm is applied to II many-objective test problems with three to I5 objectives. Experimental results show that the proposed EliteNSGA-III algorithm outperforms the NSGA-III algorithm in terms of diversity and accuracy of the obtained solutions, especially for test problems with higher objectives.
Amin Ibrahim, Shahryar Rahnamayan, Miguel Vargas Martin, Kalyanmoy Deb
CEC3
2016 Detection of Subconscious Face Recognition Using Consumer-Grade Brain-Computer Interfaces
abstract
We test the possibility of tapping the subconscious mind for face recognition using consumer-grade BCIs. To this end, we performed an experiment whereby subjects were presented with photographs of famous persons with the expectation that about 20% of them would be (consciously) recognized; and since the photos are of famous persons, we expected that subjects would have seen before some of the 80% they didn’t (consciously) recognize. Further, we expected that their subconscious would have recognized some of those in the 80% pool that they had seen before. An exit questionnaire and a set of criteria allowed us to label recognitions as conscious, false, no recognitions, or subconscious recognitions. We analyzed a number of event related potentials training and testing a support vector machine. We found that our method is capable of differentiating between no recognitions and subconscious recognitions with promising accuracy levels, suggesting that tapping the subconscious mind for face recognition is feasible.
Miguel Vargas Martin, Victor Cho, Gabriel Aversano
ACM Trans. Appl. Percept.1
2015 IEEE Services Visionary Track on Security and Privacy Engineering (SPE 2015)
abstract
Message from the IEEE Services Visionary Track on Security and Privacy Engineering (SPE 2015) Program Chairs.
Claudio A. Ardagna, Meiko Jensen, Miguel Vargas Martin
SERVICES3
2014 Crypto-assistant: Towards facilitating developer's encryption of sensitive data
abstract
The lack of encryption of data at rest or in motion is one of the top 10 database vulnerabilities [1]. We suggest that this vulnerability could be prevented by encouraging developers to perform encryption-related tasks by enhancing their integrated development environment (IDE). To this end, we created the Crypto-Assistant: a modified version of the Hibernate Tools plug-in for the popular Eclipse IDE. The purpose of the Crypto-Assistant is to mitigate the impact of developers' lack of security knowledge related to encryption by facilitating the use of encryption directives via a graphical user interface that seamlessly integrates with Hibernate Tools. Two preliminary tests helped us to identify items for improvement which have been implemented in Crypto-Assistant. We discuss Crypto-Assistant's architecture, interface, changes in the developers' workflow, and design considerations.
Ricardo Rodriguez Garcia, Julie Thorpe, Miguel Vargas Martin
PST3
2009 Detecting and Preventing the Electronic Transmission of Illicit Images and Its Network Performance
Amin Ibrahim, Miguel Vargas Martin
ICDF2C2
2009 A frame handler module for a side-channel in mobile ad hoc networks
abstract
In this paper, we establish a hidden 802.11 wireless channel, with the masking of the channel achieved by inserting intentional errors in the frame check sequence (FCS). We design a frame handler module to provide a proof-of-concept model of the side-channel using MATLAB and Simulink with communication toolbox. We justify using MATLAB over the other simulation tools because of its existing functions: physical layer IEEE 802.11 wireless local area networking (WLAN) standard, existing modular channel fading models, the MAC layer cyclic redundancy checksum (CRC) generator, the CRC Syndrome detector, and the capability of modifying fields in a frame. These existing functions allow for the creation of a frame handler which generates frames, according to our design, to be inserted as erroneous frames and recovers frames from normal 802.11 traffic. Herein we provide the design and details of the implementation of the channel. Our design offers the ability to introduce error detection and correction capabilities, and protection against passive monitoring defences. This simulation framework is a step towards the development of more sophisticated environments including multi-node simulations that maintain robust and reliable side-channel communication.
Marvin Odor, Babak Nasri, Mazda Salmanian, Peter C. Mason, Miguel Vargas Martin, Ramiro Liscano
LCN5
2000 Strategies for Hotlink Assignments
Prosenjit Bose, Evangelos Kranakis, Danny Krizanc, Miguel Vargas Martin, Jurek Czyzowicz, Andrzej Pelc, Leszek Gasieniec
ISAAC4