EDBT 2026 Demo / reviewers in the wild / expert
Luis Felipe Gutiérrez
dblp:276/0793
· DBLP profile ↗
5ranked-venue papers in the field
3as first author
2since 2021 · last 2022
0000-0001-8012-4836ORCID · reported
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 5 (3 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Solar Irradiance Prediction Using Transformer-based Machine Learning ModelsabstractThis paper presents a study of irradiance prediction using a transformer-based machine learning model for the photovoltaic (PV) renewable energy system. We explore the forecast of irradiance using ten years of data from the Texas Mesonet Data Archive at the Reese Center in Lubbock, Texas. After training our transformer model with 90% of the data, we show that it can fit the irradiance trend successfully, while the testing phase with the remaining 10% of the dataset indicates that the transformer can predict irradiance trends that align with the observed values. Additionally, we borrow the concept of rolling LSTMs to generate a rolling transformer in order to extrapolate values of irradiance even when the observed values are not available. Our extrapolation results show that the transformer can extrapolate irradiance values accurately in the short-term, but it is less precise in the long-term. To remedy this, we aim to explore more thoroughly the hyperparameter configuration of our model in order to move towards our goal of including machine learning methods in the control of PV systems. Ayda Demir, Luis Felipe Gutiérrez, Akbar Siami Namin, Stephen B. Bayne |
IEEE Big Data | 2 |
| 2022 | Generating Interpretable Features for Context-Aware Document Clustering: A Cybersecurity Case StudyabstractThis paper extends the applications of ContextMiner, a framework that we proposed recently aiming to extract interpretable contextual features to conceptualize knowledge in security texts. We show that contextual features are suitable for building a document-level representation that can be used further in a downstream machine learning and natural language processing task: document clustering. Our results show that the intrinsic readability of contextual features is usable for analyzing the document clusters obtained in our experiments. Such analysis is performed through a density-based feature selection, cluster visualization, and statistical methods in order to unveil which contextual features better characterize each document cluster. Our findings suggest that statistical techniques alongside feature analysis can be utilized to discover meaningful commonalities among documents in a particular cluster without the need of querying each document manually. Luis Felipe Gutiérrez, Akbar Siami Namin |
IEEE Big Data | 1 |
| 2020 | Predicting Emotions Perceived from SoundsabstractSonification is the science of communication of data and events to users through sounds. Auditory icons, earcons, and speech are the common auditory display schemes utilized in sonification, or more specifically in the use of audio to convey information. Once the captured data are perceived, their meanings, and more importantly, intentions can be interpreted more easily and thus can be employed as a complement to visualization techniques. Through auditory perception it is possible to convey information related to temporal, spatial, or some other context-oriented information. An important research question is whether the emotions perceived from these auditory icons or earcons are predictable in order to build an automated sonification platform. This paper conducts an experiment through which several mainstream and conventional machine learning algorithms are developed to study the prediction of emotions perceived from sounds. To do so, the key features of sounds are captured and then are modeled using machine learning algorithms using feature reduction techniques. We observe that it is possible to predict perceived emotions with high accuracy. In particular, the regression based on Random Forest demonstrated its superiority compared to other machine learning algorithms. Faranak Abri, Luis Felipe Gutiérrez, Akbar Siami Namin, David R. W. Sears, Keith S. Jones |
IEEE BigData | 2 |
| 2020 | Email Embeddings for Phishing DetectionabstractThe problem of detecting phishing emails through machine learning techniques has been discussed extensively in the literature. Conventional and state-of-the-art machine learning algorithms have demonstrated the possibility of building classifiers with high accuracy. The existing research studies treat phishing and genuine emails through general indicators and thus it is not exactly clear what phishing features are contributing to variations of the classifiers. In this paper, we crafted a set of phishing and legitimate emails with similar indicators in order to investigate whether these cues are captured or disregarded by email embeddings, i.e., vectorizations. We then fed machine learning classifiers with the carefully crafted emails to find out about the performance of email embeddings developed. Our results show that using these indicators, email embeddings techniques is effective for classifying emails as phishing or legitimate. Luis Felipe Gutiérrez, Faranak Abri, Miriam Armstrong, Akbar Siami Namin, Keith S. Jones |
IEEE BigData | 1 |
| 2020 | A Concern Analysis of Federal Reserve Statements: The Great Recession vs. The COVID-19 PandemicabstractIt is important and informative to compare and contrast major economic crises in order to confront novel and unknown cases such as the COVID-19 pandemic. The 2006 Great Recession and then the 2019 pandemic have a lot to share in terms of unemployment rate, consumption expenditures, and interest rates set by Federal Reserve. In addition to quantitative historical data, it is also interesting to compare the contents of Federal Reserve statements for the period of these two crises and find out whether Federal Reserve cares about similar concerns or there are some other issues that demand separate and unique monetary policies. This paper conducts an analysis to explore the Federal Reserve concerns as expressed in their statements for the period of 2005 to 2020. The concern analysis is performed using natural language processing (NLP) algorithms and a trend analysis of concern is also presented. We observe that there are some similarities between the Federal Reserve statements issued during the Great Recession with those issued for the 2019 COVID-19 pandemic. Luis Felipe Gutiérrez, Sima Siami-Namini, Neda Tavakoli, Akbar Siami Namin |
IEEE BigData | 1 |