EDBT 2026 Demo / reviewers in the wild / expert
Ahmed Abbasi
dblp:65/64
· DBLP profile ↗
41ranked-venue papers
12as first author
18since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 16 · 5 first-author · 6 since 2021Security and privacy · 12 · 5 first-authorArtificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Systems, architecture and hardware · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Rescorla-Wagner Steering of LLMs for Undesired Behaviors over Disproportionate Inappropriate ContextabstractRushi Wang, Jiateng Liu, Cheng Qian, Yifan Shen, Yanzhou Pan, Zhaozhuo Xu, Ahmed Abbasi, Heng Ji, Denghui Zhang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Rushi Wang, Jiateng Liu, Cheng Qian 0008, Yanzhou Pan, Zhaozhuo Xu, Ahmed Abbasi, Heng Ji 0001 |
EMNLP | 7 |
| 2025 | No Simple Answer to Data Complexity: An Examination of Instance-Level Complexity Metrics for Classification TasksabstractRyan A. Cook, John P. Lalor, Ahmed Abbasi. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Ryan A. Cook, John Lalor, Ahmed Abbasi |
NAACL (Long Papers) | 3 |
| 2025 | Hierarchical Deep Document ModelabstractTopic modeling is a commonly used text analysis tool for discovering latent topics in a text corpus. However, while topics in a text corpus often exhibit a hierarchical structure (e.g., cellphone is a sub-topic of electronics), most topic modeling methods assume a flat topic structure that ignores the hierarchical dependency among topics, or utilize a predefined topic hierarchy. In this work, we present a novel Hierarchical Deep Document Model (HDDM) to learn topic hierarchies using a variational autoencoder framework. We propose a novel objective function, sum of log likelihood, instead of the widely used evidence lower bound, to facilitate the learning of hierarchical latent topic structure. The proposed objective function can directly model and optimize the hierarchical topic-word distributions at all topic levels. We conduct experiments on four real-world text datasets to evaluate the topic modeling capability of the proposed HDDM method compared to state-of-the-art hierarchical topic modeling benchmarks. Experimental results show that HDDM achieves considerable improvement over benchmarks and is capable of learning meaningful topics and topic hierarchies. To further demonstrate the practical utility of HDDM, we apply it to a real-world medical notes dataset for clinical prediction. Experimental results show that HDDM can better summarize topics in medical notes, resulting in more accurate clinical predictions. Yi Yang 0042, John Lalor, Ahmed Abbasi, Daniel Dajun Zeng |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2025 | Empirical Guidelines for Deploying LLMs onto Resource-constrained Edge DevicesabstractThe scaling laws have become the de facto guidelines for designing large language models (LLMs), but they were studied under the assumption of unlimited computing resources for both training and inference. As LLMs are increasingly used as personalized intelligent assistants, their customization (i.e., learning through fine-tuning) and deployment onto resource-constrained edge devices will become more and more prevalent. An urgent but open question is how a resource-constrained computing environment would affect the design choices for a personalized LLM. We study this problem empirically in this work. In particular, we consider the tradeoffs among a number of key design factors and their intertwined impacts on learning efficiency and accuracy. The factors include the learning methods for LLM customization, the amount of personalized data used for learning customization, the types and sizes of LLMs, the compression methods of LLMs, the amount of time afforded to learn, and the difficulty levels of the target use cases. Through extensive experimentation and benchmarking, we draw a number of surprisingly insightful guidelines for deploying LLMs onto resource-constrained devices. For example, an optimal choice between parameter learning and RAG may vary depending on the difficulty of the downstream task, the longer fine-tuning time does not necessarily help the model, and a compressed LLM may be a better choice than an uncompressed LLM to learn from limited personalized data. Ruiyang Qin, Dancheng Liu, Chenhui Xu, Zheyu Yan, Zhaoxuan Tan, Zhenge Jia, Amir Nassereldine, Jiajie Li 0002, Meng Jiang 0001, Ahmed Abbasi, Jinjun Xiong, Yiyu Shi 0001 |
ACM Trans. Design Autom. Electr. Syst. | 10 |
| 2025 | When Automated Assessment Meets Automated Content Generation: Examining Text Quality in the Era of GPTsabstractThe use of machine learning (ML) models to assess and score textual data has become increasingly pervasive in an array of contexts including natural language processing, information retrieval, search and recommendation, and credibility assessment of online content. A significant disruption at the intersection of ML and text are text-generating large-language models (LLMs) such as generative pre-trained transformers (GPTs). We empirically assess the differences in how ML-based scoring models trained on human content assess the quality of content generated by humans versus GPTs. To do so, we propose an analysis framework that encompasses essay scoring ML models, human- and ML-generated essays, and a statistical model that parsimoniously considers the impact of type of respondent, prompt genre, and the ML model used for assessment model. A rich testbed is utilized that encompasses 18,460 human-generated and GPT-based essays. Results of our benchmark analysis reveal that LLMs and transformer pretrained language models (PLMs) more accurately score human essay quality as compared to CNN/RNN and feature-based ML methods. Interestingly, we find that LLMs and transformer PLMs tend to score GPT-generated text 10–20% higher on average, relative to human-authored documents. Conversely, traditional deep learning and feature-based ML models score human text considerably higher. Further analysis reveals that even though the LLMs and transformer PLMs are exclusively fine-tuned on human text, they more prominently attend to certain tokens appearing only in GPT-generated text, possibly (in part) due to familiarity/overlap in pre-training. Our framework and results have implications for text classification settings where automated scoring of text is likely to be disrupted by generative AI. Marialena Bevilacqua, Kezia Oketch, Ruiyang Qin, Will Stamey, Yi Gan, Kai Yang 0007, Ahmed Abbasi |
ACM Trans. Inf. Syst. | 8 |
| 2024 | FL-NAS: Towards Fairness of NAS for Resource Constrained Devices via Large Language Models : (Invited Paper)abstractNeural Architecture Search (NAS) has become the de fecto tools in the industry in automating the design of deep neural networks for various applications, especially those driven by mobile and edge devices with limited computing resources. The emerging large language models (LLMs), due to their prowess, have also been incorporated into NAS recently and show some promising results. This paper conducts further exploration in this direction by considering three important design metrics simultaneously, i.e., model accuracy, fairness, and hardware deployment efficiency. We propose a novel LLM-based NAS framework, FL-NAS, in this paper, and show experimentally that FL-NAS can indeed find high-performing DNNs, beating state-of-the-art DNN models by orders-of-magnitude across almost all design considerations. Ruiyang Qin, Zheyu Yan, Jinjun Xiong, Ahmed Abbasi, Yiyu Shi 0001 |
ASPDAC | 5 |
| 2024 | Enabling On-Device Large Language Model Personalization with Self-Supervised Data Selection and SynthesisabstractAfter a large language model (LLM) is deployed on edge devices, it is desirable for these devices to learn from user-generated conversation data to generate user-specific and personalized responses in real-time. However, user-generated data usually contains sensitive and private information, and uploading such data to the cloud for annotation is not preferred if not prohibited. While it is possible to obtain annotation locally by directly asking users to provide preferred responses, such annotations have to be sparse to not affect user experience. In addition, the storage of edge devices is usually too limited to enable large-scale fine-tuning with full user-generated data. It remains an open question how to enable on-device LLM personalization, considering sparse annotation and limited on-device storage. In this paper, we propose a novel framework to select and store the most representative data online in a self-supervised way. Such data has a small memory footprint and allows infrequent requests of user annotations for further fine-tuning. To enhance fine-tuning quality, multiple semantically similar pairs of question texts and expected responses are generated using the LLM. Our experiments show that the proposed framework achieves the best user-specific content-generating capability (accuracy) and fine-tuning speed (performance) compared with vanilla baselines. To the best of our knowledge, this is the very first on-device LLM personalization framework. Ruiyang Qin, Jun Xia 0003, Zhenge Jia, Meng Jiang 0001, Ahmed Abbasi, Peipei Zhou 0001, Jingtong Hu, Yiyu Shi 0001 |
DAC | 5 |
| 2024 | Robust Implementation of Retrieval-Augmented Generation on Edge-based Computing-in-Memory ArchitecturesabstractLarge Language Models (LLMs) deployed on edge devices learn through fine-tuning and updating a certain portion of their parameters. Although such learning methods can be optimized to reduce resource utilization, the overall required resources remain a heavy burden on edge devices. Instead, Retrieval-Augmented Generation (RAG), a resource-efficient LLM learning method, can improve the quality of the LLM-generated content without updating model parameters. However, the RAG-based LLM may involve repetitive searches on the profile data in every user-LLM interaction. This search can lead to significant latency along with the accumulation of user data. Conventional efforts to decrease latency result in restricting the size of saved user data, thus reducing the scalability of RAG as user data continuously grows. It remains an open question: how to free RAG from the constraints of latency and scalability on edge devices? In this paper, we propose a novel framework to accelerate RAG via Computing-in-Memory (CiM) architectures. It accelerates matrix multiplications by performing in-situ computation inside the memory while avoiding the expensive data transfer between the computing unit and memory. Our framework, Robust CiM-backed RAG (RoCR), utilizing a novel contrastive learning-based training method and noise-aware training, can enable RAG to efficiently search profile data with CiM. To the best of our knowledge, this is the first work utilizing CiM to accelerate RAG. Ruiyang Qin, Zheyu Yan, Dewen Zeng, Zhenge Jia, Dancheng Liu, Ahmed Abbasi, Zhi Zheng 0002, Ningyuan Cao, Kai Ni 0004, Jinjun Xiong, Yiyu Shi 0001 |
ICCAD | 7 |
| 2024 | Multimodal Mental Health Digital Biomarker Analysis From Remote Interviews Using Facial, Vocal, Linguistic, and Cardiovascular PatternsabstractOBJECTIVE: Psychiatric evaluation suffers from subjectivity and bias, and is hard to scale due to intensive professional training requirements. In this work, we investigated whether behavioral and physiological signals, extracted from tele-video interviews, differ in individuals with psychiatric disorders. METHODS: Temporal variations in facial expression, vocal expression, linguistic expression, and cardiovascular modulation were extracted from simultaneously recorded audio and video of remote interviews. Averages, standard deviations, and Markovian process-derived statistics of these features were computed from 73 subjects. Four binary classification tasks were defined: detecting 1) any clinically-diagnosed psychiatric disorder, 2) major depressive disorder, 3) self-rated depression, and 4) self-rated anxiety. Each modality was evaluated individually and in combination. RESULTS: Statistically significant feature differences were found between psychiatric and control subjects. Correlations were found between features and self-rated depression and anxiety scores. Heart rate dynamics provided the best unimodal performance with areas under the receiver-operator curve (AUROCs) of 0.68-0.75 (depending on the classification task). Combining multiple modalities provided AUROCs of 0.72-0.82. CONCLUSION: Multimodal features extracted from remote interviews revealed informative characteristics of clinically diagnosed and self-rated mental health status. SIGNIFICANCE: The proposed multimodal approach has the potential to facilitate scalable, remote, and low-cost assessment for low-burden automated mental health services. Zifan Jiang, Salman Seyedi, Emily Griner, Ahmed Abbasi, Ali Bahrami Rad, Hyeokhyen Kwon, Robert O. Cotes, Gari D. Clifford |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | Should Fairness be a Metric or a Model? A Model-based Framework for Assessing Bias in Machine Learning PipelinesabstractFairness measurement is crucial for assessing algorithmic bias in various types of machine learning (ML) models, including ones used for search relevance, recommendation, personalization, talent analytics, and natural language processing. However, the fairness measurement paradigm is currently dominated by fairness metrics that examine disparities in allocation and/or prediction error as univariate key performance indicators (KPIs) for a protected attribute or group. Although important and effective in assessing ML bias in certain contexts such as recidivism, existing metrics don’t work well in many real-world applications of ML characterized by imperfect models applied to an array of instances encompassing a multivariate mixture of protected attributes, that are part of a broader process pipeline. Consequently, the upstream representational harm quantified by existing metrics based on how the model represents protected groups doesn’t necessarily relate to allocational harm in the application of such models in downstream policy/decision contexts. We propose FAIR-Frame, a model-based framework for parsimoniously modeling fairness across multiple protected attributes in regard to the representational and allocational harm associated with the upstream design/development and downstream usage of ML models. We evaluate the efficacy of our proposed framework on two testbeds pertaining to text classification using pretrained language models. The upstream testbeds encompass over fifty thousand documents associated with twenty-eight thousand users, seven protected attributes and five different classification tasks. The downstream testbeds span three policy outcomes and over 5.41 million total observations. Results in comparison with several existing metrics show that the upstream representational harm measures produced by FAIR-Frame and other metrics are significantly different from one another, and that FAIR-Frame’s representational fairness measures have the highest percentage alignment and lowest error with allocational harm observed in downstream applications. Our findings have important implications for various ML contexts, including information retrieval, user modeling, digital platforms, and text classification, where responsible and trustworthy AI is becoming an imperative. John Lalor, Ahmed Abbasi, Kezia Oketch, Yi Yang 0042, Nicole Forsgren |
ACM Trans. Inf. Syst. | 2 |
| 2024 | Data Augmentation-based Novel Deep Learning Method for Deepfaked Images DetectionabstractRecent advances in artificial intelligence have led to deepfake images, enabling users to replace a real face with a genuine one. deepfake images have recently been used to malign public figures, politicians, and even average citizens. deepfake but realistic images have been used to stir political dissatisfaction, blackmail, propagate false news, and even carry out bogus terrorist attacks. Thus, identifying real images from fakes has got more challenging. To avoid these issues, this study employs transfer learning and data augmentation technique to classify deepfake images. For experimentation, 190,335 RGB-resolution deepfake and real images and image augmentation methods are used to prepare the dataset. The experiments use the deep learning models: convolutional neural network (CNN), Inception V3, visual geometry group (VGG19), and VGG16 with a transfer learning approach. Essential evaluation metrics (accuracy, precision, recall, F1-score, confusion matrix, and AUC-ROC curve score) are used to test the efficacy of the proposed approach. Results revealed that the proposed approach achieves an accuracy, recall, F1-score and AUC-ROC score of 90% and 91% precision, with our fine-tuned VGG16 model outperforming other DL models in recognizing real and deepfakes. Farkhund Iqbal, Ahmed Abbasi, Abdul Rehman Javed, Ahmad S. Almadhor, Zunera Jalil, Sajid Anwar 0001, Imad Rida |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Examining User Heterogeneity in Digital ExperimentsabstractDigital experiments are routinely used to test the value of a treatment relative to a status quo control setting — for instance, a new search relevance algorithm for a website or a new results layout for a mobile app. As digital experiments have become increasingly pervasive in organizations and a wide variety of research areas, their growth has prompted a new set of challenges for experimentation platforms. One challenge is that experiments often focus on the average treatment effect (ATE) without explicitly considering differences across major sub-groups — heterogeneous treatment effect (HTE). This is especially problematic because ATEs have decreased in many organizations as the more obvious benefits have already been realized. However, questions abound regarding the pervasiveness of user HTEs and how best to detect them. We propose a framework for detecting and analyzing user HTEs in digital experiments. Our framework combines an array of user characteristics with double machine learning. Analysis of 27 real-world experiments spanning 1.76 billion sessions and simulated data demonstrates the effectiveness of our detection method relative to existing techniques. We also find that transaction, demographic, engagement, satisfaction, and lifecycle characteristics exhibit statistically significant HTEs in 10% to 20% of our real-world experiments, underscoring the importance of considering user heterogeneity when analyzing experiment results, otherwise personalized features and experiences cannot happen, thus reducing effectiveness. In terms of the number of experiments and user sessions, we are not aware of any study that has examined user HTEs at this scale. Our findings have important implications for information retrieval, user modeling, platforms, and digital experience contexts, in which online experiments are often used to evaluate the effectiveness of design artifacts. Sriram Somanchi, Ahmed Abbasi, Ken Kelley, David G. Dobolyi, Ted Tao Yuan |
ACM Trans. Inf. Syst. | 2 |
| 2022 | Auto-Debias: Debiasing Masked Language Models with Automated Biased PromptsabstractHuman-like biases and undesired social stereotypes exist in large pretrained language models.Given the wide adoption of these models in real-world applications, mitigating such biases has become an emerging and important task.In this paper, we propose an automatic method to mitigate the biases in pretrained language models.Different from previous debiasing work that uses external corpora to finetune the pretrained models, we instead directly probe the biases encoded in pretrained models through prompts.Specifically, we propose a variant of the beam search method to automatically search for biased prompts such that the cloze-style completions are the most different with respect to different demographic groups.Given the identified biased prompts, we then propose a distribution alignment loss to mitigate the biases.Experiment results on standard datasets and metrics show that our proposed Auto-Debias approach can significantly reduce biases, including gender and racial bias, in pretrained language models such as BERT, RoBERTa and ALBERT.Moreover, the improvement in fairness does not decrease the language models' understanding abilities, as shown using the GLUE benchmark. Yue Guo 0009, Yi Yang 0042, Ahmed Abbasi |
ACL (1) | 3 |
| 2022 | Benchmarking Intersectional Biases in NLPabstractJohn Lalor, Yi Yang, Kendall Smith, Nicole Forsgren, Ahmed Abbasi. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. John Lalor, Yi Yang 0042, Kendall Smith, Nicole Forsgren, Ahmed Abbasi |
NAACL-HLT | 5 |
| 2022 | Deep Learning for Adverse Event Detection From Web SearchabstractAdverse event detection is critical for many real-world applications including timely identification of product defects, disasters, and major socio-political incidents. In the health context, adverse drug events account for countless hospitalizations and deaths annually. Since users often begin their information seeking and reporting with online searches, examination of search query logs has emerged as an important detection channel. However, search context - including query intent and heterogeneity in user behaviors – is extremely important for extracting information from search queries, and yet the challenge of measuring and analyzing these aspects has precluded their use in prior studies. We propose DeepSAVE, a novel deep learning framework for detecting adverse events based on user search query logs. DeepSAVE uses an enriched variational autoencoder encompassing a novel query embedding and user modeling module that work in concert to address the context challenge associated with search-based detection of adverse events. Evaluation results on three large real-world event datasets show that DeepSAVE outperforms existing detection methods as well as comparison deep learning auto encoders. Ablation analysis reveals that each component of DeepSAVE significantly contributes to its overall performance. Collectively, the results demonstrate the viability of the proposed architecture for detecting adverse events from search query logs. Ahmed Abbasi, Brent Kitchens, Donald A. Adjeroh, Daniel Dajun Zeng |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2021 | Detecting Drug-Drug Interactions using Protein Sequence-Structure Similarity NetworksabstractAdverse drug events represent a key challenge in public health, especially with respect to drug safety profiling and drug surveillance. Drug-drug interactions represent one of the most popular types of adverse drug events. Most computational approaches to this problem have used different types of data, such as drug chemical structure, information about protein targets, side effects, pathways, etc to predict potential interactions between drugs. In this work, we study the question of whether using just genetic information about the drugs can provide significant information about the potential safety profile for a given drug. We propose a novel neural network model to predict adverse drug events using only data about the protein sequence and protein structure associated with the drug targets. We compare the results with those from the state-of-the-art methods on this problem. Our results show that the proposed method is quite competitive, at times outperforming the state-of-the-art. Saminur Islam, Ahmed Abbasi, Nitin Agarwal 0001, Wanhong Zheng, Gianfranco Doretto, Donald A. Adjeroh |
BIBM | 2 |
| 2021 | Constructing a Psychometric Testbed for Fair Natural Language ProcessingabstractPsychometric measures of ability, attitudes, perceptions, and beliefs are crucial for understanding user behavior in various contexts including health, security, e-commerce, and finance.Traditionally, psychometric dimensions have been measured and collected using survey-based methods.Inferring such constructs from user-generated text could allow timely, unobtrusive collection and analysis.In this work we construct a corpus for psychometric natural language processing (NLP) related to important dimensions such as trust, anxiety, numeracy, and literacy, in the health domain.We discuss our multi-step process to align user text with their survey-based response items and provide an overview of the resulting testbed, which encompasses surveybased psychometric measures and accompanying user-generated text from 8,502 respondents.Our testbed also encompasses selfreported demographic information, including race, sex, age, income, and education, allowing for measuring bias and benchmarking fairness of text classification methods.We report preliminary results on use of the text to predict/categorize users' survey response labels and on the fairness of these models.We also discuss the important implications of our work and resulting testbed for future NLP research on psychometrics and fairness. Ahmed Abbasi, David G. Dobolyi, John Lalor, Richard G. Netemeyer, Kendall Smith, Yi Yang 0042 |
EMNLP (1) | 1 |
| 2021 | Trust calibration of automated security IT artifacts: A multi-domain study of phishing-website detection tools
Yan Chen 0016, Ahmed Abbasi, David G. Dobolyi |
Inf. Manag. | 3 |
| 2020 | An Ordinal Approach to Modeling and Visualizing Phishing SusceptibilityabstractPhishing is a significant, ongoing cybersecurity threat facing both individuals and organizations, and the consequences of falling victim to phishing can be dire, particularly in terms of financial loss. While much research has focused on understanding or predicting user susceptibility to phishing with the aim of preventing it, little of this research has focused on modeling the process of being phished holistically. Specifically, while the outcome of interacting with a phishing website may be thought of as binary (e.g., "Were you phished or not?"), the actual process typically involves a sequence of stages spanning from the choice to visit (or not visit) a link all the way to transacting with a website and giving away valuable personal or financial information. To better understand the variables that influence phishing website traversal, we conducted a controlled lab experiment with a large sample of 908 participants. In this experiment, each participant was given a task such as opening an online checking account and presented with a series of simulated search results that included both legitimate and phishing websites. By monitoring participants' interactions with these websites and collecting additional information via surveys, we evaluated which variables were the most likely to result in behavior that could lead to greater phishing exposure using a multi-model comparison approach. The results of our analyses shed light on the key variables that can lead to a greater propensity for being phished and may prove invaluable to researchers interested in designing new interventions. David G. Dobolyi, Ahmed Abbasi, Anthony Vance |
ISI | 2 |
| 2020 | Phishcasting: Deep Learning for Time Series Forecasting of Phishing AttacksabstractPhishing attacks remain pervasive and continue to be a source of significant monetary loss, identity theft, and malware. One of the challenges is that in most organizational settings, the detection paradigm is inherently about identifying and reacting to threats in real-time, as they are unfolding. As a way to complement these efforts with greater foresight, we introduce the idea of phishcasting — forecasting of phishing threat levels weeks or months into the future. Given that phishing attack volume time series data is noisy and devoid of traditional seasonal and cyclical trends, we extend the time series forecasting framework to utilize multiple time series, auxiliary information and alternate representations. We also introduce CoT-Net, a flexible, end-to-end CNN-LSTM based deep learning method for forecasting of complex phishing attack volume time series. CoT-Net uses time series embeddings to uncover correlations between organizational attack patterns within and across industry sectors. Using a publicly available test bed featuring multiple organizations’ attack volume over time, we find CoT-Net to outperform most state-of-the-art time series forecasting methods. By showing that phishcasting might be possible and practical, our work has important proactive implications for cybersecurity. Syed Hasan Amin Mahmood, Syed Mustafa Ali Abbasi, Ahmed Abbasi, Fareed Zaffar |
ISI | 3 |
| 2020 | A Deep Learning Architecture for Psychometric Natural Language ProcessingabstractPsychometric measures reflecting people’s knowledge, ability, attitudes, and personality traits are critical for many real-world applications, such as e-commerce, health care, and cybersecurity. However, traditional methods cannot collect and measure rich psychometric dimensions in a timely and unobtrusive manner. Consequently, despite their importance, psychometric dimensions have received limited attention from the natural language processing and information retrieval communities. In this article, we propose a deep learning architecture, PyNDA, to extract psychometric dimensions from user-generated texts. PyNDA contains a novel representation embedding, a demographic embedding, a structural equation model (SEM) encoder, and a multitask learning mechanism designed to work in unison to address the unique challenges associated with extracting rich, sophisticated, and user-centric psychometric dimensions. Our experiments on three real-world datasets encompassing 11 psychometric dimensions, including trust, anxiety, and literacy, show that PyNDA markedly outperforms traditional feature-based classifiers as well as the state-of-the-art deep learning architectures. Ablation analysis reveals that each component of PyNDA significantly contributes to its overall performance. Collectively, the results demonstrate the efficacy of the proposed architecture for facilitating rich psychometric analysis. Our results have important implications for user-centric information extraction and retrieval systems looking to measure and incorporate psychometric dimensions. Ahmed Abbasi, David G. Dobolyi, Richard G. Netemeyer, Gari D. Clifford, Hsinchun Chen |
ACM Trans. Inf. Syst. | 2 |
| 2019 | Using discussion logic in analyzing online group discussions: A text mining approach
Shasha Deng, Yilu Zhou, Pengzhu Zhang, Ahmed Abbasi |
Inf. Manag. | 4 |
| 2016 | Phishing susceptibility: The good, the bad, and the uglyabstractPhishing website-based attacks remain pervasive, with high user susceptibility continuing to be a major factor. In this study we use cluster analysis coupled with an elaborate controlled experiment involving hundreds of participants to identify and examine high susceptibility user segments in terms of their perceptions, demographics, and phishing website traversal behavior. The results reveal three sets of users, including two sets that exhibit highly problematic behavior. The results have important implications for training programs, usability of anti-phishing tools, and security policies. Ahmed Abbasi, Yan Chen 0016 |
ISI | 1 |
| 2016 | PhishMonger: A free and open source public archive of real-world phishing websitesabstractThe number of active, online phishing websites continues to grow unabated in recent years. This has created an ever-increasing security risk for both individual and enterprise users in terms of identity theft, malware, financial loss, etc. Although resources exist for tracking, cataloguing, and blacklisting these types of sites (e.g., PhishTank.com), the ephemeral nature of phishing websites makes in-depth analysis exceptionally difficult. In order to better understand how these phishing sites exploit user and system weaknesses, we have crafted a platform named PhishMonger for capturing live phishing websites in real-time on an ever-present, rolling basis, which we outline in this paper. Moreover, we present details regarding our growing database of verified phishing websites, which currently encompasses over 88,754 sites, spanning 10,956,415 files and folders, utilizing 108GB of compressed storage. We offer recommendations on how this corpus can be leveraged by the cybersecurity and security informatics research communities to examine several important research problems. David G. Dobolyi, Ahmed Abbasi |
ISI | 2 |
| 2014 | Benchmarking Twitter Sentiment Analysis Tools
Ahmed Abbasi, Ammar Hassan, Milan Dhar |
LREC | 1 |
| 2013 | Evaluating text visualization: An experiment in authorship analysisabstractAnalyzing authorship of online texts is an important analysis task in security-related areas such as cybercrime investigation and counter-terrorism, and in any field of endeavor in which authorship may be uncertain or obfuscated. This paper presents an automated approach for authorship analysis using machine learning methods, a robust stylometric feature set, and a series of visualizations designed to facilitate analysis at the feature, author, and message levels. A testbed consisting of 506,554 forum messages, in English and Arabic, from 14,901 authors was first constructed. A prototype portal system was then developed to support feasibility analysis of the approach. A preliminary evaluation to assess the efficacy of the text visualizations was conducted. The evaluation showed that task performance with the visualization functions was more accurate and more efficient than task performance without the visualizations. Victor A. Benjamin, Wingyan Chung, Ahmed Abbasi, Joshua Chuang, Catherine A. Larson, Hsinchun Chen |
ISI | 3 |
| 2012 | Impact of anti-phishing tool performance on attack success ratesabstractPhishing website-based attacks continue to present significant problems for individual and enterprise-level security, including identity theft, malware, and viruses. While the performance of anti-phishing tools has improved considerably, it is unclear how effective such tools are at protecting users. In this study, an experiment involving over 400 participants was used to evaluate the impact of anti-phishing tools' accuracy on users' ability to avoid phishing threats. Each of the participants was given either a high accuracy (90%) or low accuracy (60%) tool and asked to make various decisions about several legitimate and phishing websites. Experiment results revealed that participants using the high accuracy anti-phishing tool significantly outperformed those using the less accurate tool in their ability to: (1) differentiate legitimate websites from phish; (2) avoid visiting phishing websites; and (3) avoid transacting with phishing websites. However, even users of the high accuracy tool often disregarded its correct recommendations, resulting in users' phish detection rates that were approximately 15% lower than those of the anti-phishing tool used. Consequently, on average, participants visited between 74% and 83% of the phishing websites and were willing to transact with as many as 25% of the phishing websites. Anti-phishing tools were also less effective against one particular type of threat. The results suggest that while the accuracy of anti-phishing tools is a critical factor, reducing the success rates of phishing attacks requires other considerations such as improving tool interface/warning design and enhancing users' knowledge of phishing. Given the prevalence of phishing-based web fraud, the findings have important implications for individual and enterprise security. Ahmed Abbasi, Yan Chen 0016 |
ISI | 1 |
| 2012 | Exploratory experiments to identify fake websites by using features from the network stackabstractUsers on the web are unknowingly becoming more susceptible to scams from cyber deviants and malicious websites. There has been much work in the identification of malicious websites using application layer features based on content (HTML, images, links, etc.) and a plethora of classification techniques. However, there has been little work on using features from the other layers in the Open Systems Interconnection (OSI) network stack. Capturing features from the transport and internet layers of the network stack based on responses to various Hypertext Transfer Protocol (HTTP) requests may allow for increased classification accuracy. In this paper, we use learning techniques (Winnow, Logit Regression, Naïve Bayes, J48, and Bayesian) utilizing these new features to identify fake pharmacy websites. The results show that using transport and Internet layer features yields an accuracy of 80% to 95% for detecting fake websites using standard machine learning algorithms. The results suggest that many organizations may be hosting multiple websites using shared code and hosting services to enable them to produce the maximum number of fraudulent websites. Jason Koepke, Siddharth Kaza, Ahmed Abbasi |
ISI | 3 |
| 2012 | Detecting Fake Medical Web Sites Using Recursive Trust LabelingabstractFake medical Web sites have become increasingly prevalent. Consequently, much of the health-related information and advice available online is inaccurate and/or misleading. Scores of medical institution Web sites are for organizations that do not exist and more than 90% of online pharmacy Web sites are fraudulent. In addition to monetary losses exacted on unsuspecting users, these fake medical Web sites have severe public safety ramifications. According to a World Health Organization report, approximately half the drugs sold on the Web are counterfeit, resulting in thousands of deaths. In this study, we propose an adaptive learning algorithm called recursive trust labeling (RTL). RTL uses underlying content and graph-based classifiers, coupled with a recursive labeling mechanism, for enhanced detection of fake medical Web sites. The proposed method was evaluated on a test bed encompassing nearly 100 million links between 930,000 Web sites, including 1,000 known legitimate and fake medical sites. The experimental results revealed that RTL was able to significantly improve fake medical Web site detection performance over 19 comparison content and graph-based methods, various meta-learning techniques, and existing adaptive learning approaches, with an overall accuracy of over 94%. Moreover, RTL was able to attain high performance levels even when the training dataset composed of as little as 30 Web sites. With the increased popularity of eHealth and Health 2.0, the results have important implications for online trust, security, and public safety. Ahmed Abbasi, Siddharth Kaza |
ACM Trans. Inf. Syst. | 1 |
| 2012 | Sentimental Spidering: Leveraging Opinion Information in Focused CrawlersabstractDespite the increased prevalence of sentiment-related information on the Web, there has been limited work on focused crawlers capable of effectively collecting not only topic-relevant but also sentiment-relevant content. In this article, we propose a novel focused crawler that incorporates topic and sentiment information as well as a graph-based tunneling mechanism for enhanced collection of opinion-rich Web content regarding a particular topic. The graph-based sentiment (GBS) crawler uses a text classifier that employs both topic and sentiment categorization modules to assess the relevance of candidate pages. This information is also used to label nodes in web graphs that are employed by the tunneling mechanism to improve collection recall. Experimental results on two test beds revealed that GBS was able to provide better precision and recall than seven comparison crawlers. Moreover, GBS was able to collect a large proportion of the relevant content after traversing far fewer pages than comparison methods. GBS outperformed comparison methods on various categories of Web pages in the test beds, including collection of blogs, Web forums, and social networking Web site content. Further analysis revealed that both the sentiment classification module and graph-based tunneling mechanism played an integral role in the overall effectiveness of the GBS crawler. Tianjun Fu, Ahmed Abbasi, Daniel Dajun Zeng, Hsinchun Chen |
ACM Trans. Inf. Syst. | 2 |
| 2011 | Selecting Attributes for Sentiment Classification Using Feature Relation NetworksabstractA major concern when incorporating large sets of diverse n-gram features for sentiment classification is the presence of noisy, irrelevant, and redundant attributes. These concerns can often make it difficult to harness the augmented discriminatory potential of extended feature sets. We propose a rule-based multivariate text feature selection method called Feature Relation Network (FRN) that considers semantic information and also leverages the syntactic relationships between n-gram features. FRN is intended to efficiently enable the inclusion of extended sets of heterogeneous n-gram features for enhanced sentiment classification. Experiments were conducted on three online review testbeds in comparison with methods used in prior sentiment classification research. FRN outperformed the comparison univariate, multivariate, and hybrid feature selection methods; it was able to select attributes resulting in significantly better classification accuracy irrespective of the feature subset sizes. Furthermore, by incorporating syntactic information about n-gram relations, FRN is able to select features in a more computationally efficient manner than many multivariate and hybrid techniques. Ahmed Abbasi, Stephen L. France, Hsinchun Chen |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2010 | A focused crawler for Dark Web forumsabstractAbstract The unprecedented growth of the Internet has given rise to the Dark Web, the problematic facet of the Web associated with cybercrime, hate, and extremism. Despite the need for tools to collect and analyze Dark Web forums, the covert nature of this part of the Internet makes traditional Web crawling techniques insufficient for capturing such content. In this study, we propose a novel crawling system designed to collect Dark Web forum content. The system uses a human‐assisted accessibility approach to gain access to Dark Web forums. Several URL ordering features and techniques enable efficient extraction of forum postings. The system also includes an incremental crawler coupled with a recall‐improvement mechanism intended to facilitate enhanced retrieval and updating of collected content. Experiments conducted to evaluate the effectiveness of the human‐assisted accessibility approach and the recall‐improvement‐based, incremental‐update procedure yielded favorable results. The human‐assisted approach significantly improved access to Dark Web forums while the incremental crawler with recall improvement also outperformed standard periodic‐ and incremental‐update approaches. Using the system, we were able to collect over 100 Dark Web forums from three regions. A case study encompassing link and content analysis of collected forums was used to illustrate the value and importance of gathering and analyzing content from such online communities. Tianjun Fu, Ahmed Abbasi, Hsinchun Chen |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2008 | Developing ideological networks using social network analysis and writeprints: A case study of the international Falun Gong movementabstractThe convenience of the Internet has made it possible for activist groups to easily form alliances through their websites to appeal to wider audience and increase their impact. In this study, we investigate the potential of using Social Network Analysis (SNA) and Writeprints to discover the fusion of activitst ideas on the Internet, focusing on the Falun Gong movement. We find that network visualization is very useful to reveal how different types of websites or ideas are associated and, in some cases, mixed together. Furthermore, the measures of centrality in SNA help to reveal which websites most prominently link to other websites. We find that Writeprints can be used to identify the ideas which an author gradually introduces and combines through a series of messages. Yi-Da Chen, Ahmed Abbasi, Hsinchun Chen |
ISI | 2 |
| 2008 | A hybrid approach to Web forum interactional coherence analysisabstractAbstract Despite the rapid growth of text‐based computer‐mediated communication (CMC), its limitations have rendered the media highly incoherent. This poses problems for content analysis of online discourse archives. Interactional coherence analysis (ICA) attempts to accurately identify and construct CMC interaction networks. In this study, we propose the Hybrid Interactional Coherence (HIC) algorithm for identification of web forum interaction. HIC utilizes a bevy of system and linguistic features, including message header information, quotations, direct address, and lexical relations. Furthermore, several similarity‐based methods including a Lexical Match Algorithm (LMA) and a sliding window method are utilized to account for interactional idiosyncrasies. Experiments results on two web forums revealed that the proposed HIC algorithm significantly outperformed comparison techniques in terms of precision, recall, and F‐measure at both the forum and thread levels. Additionally, an example was used to illustrate how the improved ICA results can facilitate enhanced social network and role analysis capabilities. Tianjun Fu, Ahmed Abbasi, Hsinchun Chen |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2008 | Affect Analysis of Web Forums and Blogs Using Correlation EnsemblesabstractAnalysis of affective intensities in computer-mediated communication is important in order to allow a better understanding of online users' emotions and preferences. Despite considerable research on textual affect classification, it is unclear which features and techniques are most effective. In this study, we compared several feature representations for affect analysis, including learned n-grams and various automatically and manually crafted affect lexicons. We also proposed the support vector regression correlation ensemble (SVRCE) method for enhanced classification of affect intensities. SVRCE uses an ensemble of classifiers each trained using a feature subset tailored toward classifying a single affect class. The ensemble is combined with affect correlation information to enable better prediction of emotive intensities. Experiments were conducted on four test beds encompassing web forums, blogs, and online stories. The results revealed that learned n-grams were more effective than lexicon-based affect representations. The findings also indicated that SVRCE outperformed comparison techniques, including Pace regression, semantic orientation, and WordNet models. Ablation testing showed that the improved performance of SVRCE was attributable to its use of feature ensembles as well as affect correlation information. A brief case study was conducted to illustrate the utility of the features and techniques for affect analysis of large archives of online discourse. Ahmed Abbasi, Hsinchun Chen, S. Thoms, Tianjun Fu |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2008 | Writeprints: A stylometric approach to identity-level identification and similarity detection in cyberspaceabstractOne of the problems often associated with online anonymity is that it hinders social accountability, as substantiated by the high levels of cybercrime. Although identity cues are scarce in cyberspace, individuals often leave behind textual identity traces. In this study we proposed the use of stylometric analysis techniques to help identify individuals based on writing style. We incorporated a rich set of stylistic features, including lexical, syntactic, structural, content-specific, and idiosyncratic attributes. We also developed the Writeprints technique for identification and similarity detection of anonymous identities. Writeprints is a Karhunen-Loeve transforms-based technique that uses a sliding window and pattern disruption algorithm with individual author-level feature sets. The Writeprints technique and extended feature set were evaluated on a testbed encompassing four online datasets spanning different domains: email, instant messaging, feedback comments, and program code. Writeprints outperformed benchmark techniques, including SVM, Ensemble SVM, PCA, and standard Karhunen-Loeve transforms, on the identification and similarity detection tasks with accuracy as high as 94% when differentiating between 100 authors. The extended feature set also significantly outperformed a baseline set of features commonly used in previous research. Furthermore, individual-author-level feature sets generally outperformed use of a single group of attributes. Ahmed Abbasi, Hsinchun Chen |
ACM Trans. Inf. Syst. | 1 |
| 2008 | Sentiment analysis in multiple languages: Feature selection for opinion classification in Web forumsabstractThe Internet is frequently used as a medium for exchange of information and opinions, as well as propaganda dissemination. In this study the use of sentiment analysis methodologies is proposed for classification of Web forum opinions in multiple languages. The utility of stylistic and syntactic features is evaluated for sentiment classification of English and Arabic content. Specific feature extraction components are integrated to account for the linguistic characteristics of Arabic. The entropy weighted genetic algorithm (EWGA) is also developed, which is a hybridized genetic algorithm that incorporates the information-gain heuristic for feature selection. EWGA is designed to improve performance and get a better assessment of key features. The proposed features and techniques are evaluated on a benchmark movie review dataset and U.S. and Middle Eastern Web forum postings. The experimental results using EWGA with SVM indicate high performance levels, with accuracies of over 91% on the benchmark dataset as well as the U.S. and Middle Eastern forums. Stylistic features significantly enhanced performance across all testbeds while EWGA also outperformed other feature selection methods, indicating the utility of these features and techniques for document-level classification of sentiments. Ahmed Abbasi, Hsinchun Chen, Arab Salem |
ACM Trans. Inf. Syst. | 1 |
| 2007 | Affect Intensity Analysis of Dark Web ForumsabstractAffects play an important role in influencing people's perceptions and decision making. Affect analysis is useful for measuring the presence of hate, violence, and the resulting propaganda dissemination across extremist groups. In this study we performed affect analysis of U.S. and Middle Eastern extremist group forum postings. We constructed an affect lexicon using a probabilistic disambiguation technique to measure the usage of violence and hate affects. These techniques facilitate in depth analysis of multilingual content. The proposed approach was evaluated by applying it across 16 U.S. supremacist and Middle Eastern extremist group forums. Analysis across regions reveals that the Middle Eastern test bed forums have considerably greater violence intensity than the U.S. groups. There is also a strong linear relationship between the usage of hate and violence across the Middle Eastern messages. Ahmed Abbasi, Hsinchun Chen |
ISI | 1 |
| 2007 | Interaction Coherence Analysis for Dark Web ForumsabstractInteraction coherence analysis (ICA) attempts to accurately identify and construct interaction networks by using various features and techniques. It is useful to identify user roles, user's social and information value, as well as the social network structure of Dark Web communities. In this study, we applied interaction coherence analysis for Dark Web forums using the hybrid interaction coherence (HIC) algorithm. Our algorithm utilizes both system features such as header information and quotations, and linguistic features such as direct address and lexical relation. Furthermore, several similarity-based methods, for example vector space model, dice equation, and sliding window, are used to address various types of noises. Two experiments have been conducted to compare our HIC algorithm with traditional linkage-based method, similarity-based method, and a simplified HIC method that does not address noise issues. The results demonstrate the effectiveness of our HIC algorithm for identifying interactions in Dark Web forums. Tianjun Fu, Ahmed Abbasi, Hsinchun Chen |
ISI | 2 |
| 2006 | Visualizing Authorship for Identification
Ahmed Abbasi, Hsinchun Chen |
ISI | 1 |
| 2005 | Applying Authorship Analysis to Arabic Web Content
Ahmed Abbasi, Hsinchun Chen |
ISI | 1 |