VLDB 2026 Research / reviewers in the wild / expert
Hsinchun Chen
dblp:c/HsinchunChen
· DBLP profile ↗
289ranked-venue papers
45as first author
16since 2021 · last 2025
0000-0003-3251-2433ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 111 · 8 first-author · 9 since 2021Databases, data management, data science and information retrieval · 83 · 21 first-author · 6 since 2021Artificial intelligence and machine learning · 60 · 10 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 24 · 1 first-authorHuman-computer interaction and ubiquitous computing · 16 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-authorTheory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning Contextualized Action Representations in Sequential Decision Making for Adversarial Malware OptimizationabstractDeep learning (DL)-based malware detectors have shown promise in swiftly detecting unseen malware without expensive dynamic malware behavior analysis. These detectors have been shown to be susceptible to adversarial malware variants generated from meticulously modifying known malware to mislead detectors into recognizing them as benign. Being able to automatically generate optimized functional adversarial malware variants by defenders is crucial to effective cyber defense and staying ahead of the adversary. Current adversarial malware example generation methods often assume threat models with any of the following four restrictions: (1) requiring access to insider knowledge about malware detectors, (2) an unlimited size of adversarial modifications, (3) an unlimited number of queries to malware detector, and (4) relying on dynamic analysis of malware behavior in a sandbox. Drawing on Actor-Critic Reinforcement Learning (RL), we propose a novel closed-box binary manipulation method for adversarial malware optimization, named Actor-Critic with Contextualized Action Representations (AC-CAR), to generate malware variants without these restrictions. AC-CAR leverages two novel components, a contextualized policy and a neural language model-based RL-augmented top-$k$sampling method. Unlike current methods, AC-CAR can utilize tens of thousands of actions to augment malware executables for evading DL-based malware detectors. AC-CAR yields an approximately 2-fold performance increase over the current methods on average, while decreasing the payload size to 20 times smaller than leading methods. We show that using the malware variants generated by AC-CAR in an adversarial re-training procedure improves malware detector’ robustness against adversarial variants by 29.65% on average. Reza Ebrahimi 0001, Jason L. Pacheco, James Lee Hu, Hsinchun Chen |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2024 | The 4th Workshop on Artificial Intelligence-enabled Cybersecurity AnalyticsabstractCybersecurity remains a grand societal challenge. Large and constantly changing attack surfaces are non-trivial to protect against malicious actors. Entities like the United States and the European Union have recently emphasized the value of Artificial Intelligence (AI) for advancing cybersecurity. For example, the National Science Foundation has called for AI systems that can enhance cyber threat intelligence, detect new and evolving threats, and analyze massive troves of cybersecurity data. The 4th Workshop on Artificial Intelligence-enabled Cybersecurity Analytics (co-located with ACM KDD) sought to make significant and novel contributions within these relevant topics. Submissions were reviewed by highly qualified AI for cybersecurity researchers and practitioners spanning academia and private industry firms. Steven Ullman, Benjamin Ampel, Sagar Samtani, Shanchieh Jay Yang, Hsinchun Chen |
KDD | 5 |
| 2023 | Evading Deep Learning-Based Malware Detectors via Obfuscation: A Deep Reinforcement Learning ApproachabstractAdversarial Malware Generation (AMG), the generation of adversarial malware variants to strengthen Deep Learning (DL)-based malware detectors has emerged as a crucial tool in the development of proactive cyberdefense. However, the majority of extant works offer subtle perturbations or additions to executable files and do not explore full-file obfuscation. In this study, we show that an open-source encryption tool coupled with a Reinforcement Learning (RL) framework can successfully obfuscate malware to evade state-of-the-art malware detection engines and outperform techniques that use advanced modification methods. Our results show that the proposed method improves the evasion rate from 27%-49% compared to widely-used state-of-the-art reinforcement learning-based methods. Brian Etter, James Lee Hu, Reza Ebrahimi 0001, Weifeng Li 0002, Xin Li 0108, Hsinchun Chen |
ICDM | 6 |
| 2023 | Disrupting Ransomware Actors on the Bitcoin Blockchain: A Graph Embedding ApproachabstractRansomware is a growing problem and significant threat to cybersecurity in the United States. One primary vector for ransomware payments is the Bitcoin network. Network science techniques are a potential approach to analyze ransomware payment networks to discover salient ransomware actors. In this study, we propose a design framework for labeling nodes in a ransomware payment network and identifying key ransomware Bitcoin addresses that can be targeted for disruption. By leveraging semi-supervised graph embedding methodology and updating the loss function of a prevailing algorithm, GraphSAGE, to manage dataset imbalance, we identify key wallets in our ransomware network. We demonstrate the utility of our approach with a case study identifying a Bitcoin wallet that has been reported as a ransomware actor as recently as December 2021 and has transferred over $450 million in Bitcoin. Benjamin Ampel, Kaeli Otto, Sagar Samtani, Hsinchun Chen |
ISI | 4 |
| 2023 | Mapping Exploit Code on Paste Sites to the MITRE ATT&CK Framework: A Multi-label Transformer ApproachabstractCyber-criminals often use information-sharing platforms such as paste sites (e.g., Pastebin) to share vast amounts of malicious text content, such as exploit source code. Careful analysis of malicious paste site content can provide Cyber Threat Intelligence (CTI) about potential threats. In this research, we propose a Convolutional BiLSTM Transformer multi-label classification method that automatically maps paste site exploit source code to the MITRE ATT&CK framework to identify adversarial techniques in support of proactive CTI. The Convolutional BiLSTM Transformer combines a convolutional neural network layer placed before a Transformer block, a concatenated pooling from a global max pooling and global average, and a BiLSTM pair-wise function within the Transformer to capture word and sequence orders. We conducted an multi-label classification experiment where our proposed Convolutional BiLSTM Transformer model achieved state-of-the-art results in terms of accuracy, recall, F1-score, and hamming loss. The results of a case study showed the tactics and tools that are used by malicious actors on paste sites. Benjamin Ampel, Tala Vahedi, Sagar Samtani, Hsinchun Chen |
ISI | 4 |
| 2023 | The 3rd Workshop on Artificial Intelligence-enabled Cybersecurity AnalyticsabstractArtificial Intelligence (AI) has gripped modern society as a viable approach to revolutionize operational capabilities across multiple industries. One critical application area that could stand to benefit from the capabilities of AI is cybersecurity. Increasingly, federal funding agencies such as the National Science Foundation are calling for enhanced AI-enabled analytics capabilities to improve cyber threat intelligence, cyber defense generation, and more. To this end, this half-day workshop, not in its third year at ACM KDD, sought to attain significant contributions related to various aspects of AI-enabled cybersecurity analytics. This workshop received a record number of submissions. Submissions were reviewed by a highly-qualified, interdisciplinary group of AI for cybersecurity researchers and practitioners spanning academia and private industry firms. Sagar Samtani, Shanchieh Jay Yang, Hsinchun Chen |
KDD | 3 |
| 2023 | Heterogeneous Domain Adaptation With Adversarial Neural Representation Learning: Experiments on E-Commerce and CybersecurityabstractLearning predictive models in new domains with scarce training data is a growing challenge in modern supervised learning scenarios. This incentivizes developing domain adaptation methods that leverage the knowledge in known domains (source) and adapt to new domains (target) with a different probability distribution. This becomes more challenging when the source and target domains are in heterogeneous feature spaces, known as heterogeneous domain adaptation (HDA). While most HDA methods utilize mathematical optimization to map source and target data to a common space, they suffer from low transferability. Neural representations have proven to be more transferable; however, they are mainly designed for homogeneous environments. Drawing on the theory of domain adaptation, we propose a novel framework, Heterogeneous Adversarial Neural Domain Adaptation (HANDA), to effectively maximize the transferability in heterogeneous environments. HANDA conducts feature and distribution alignment in a unified neural network architecture and achieves domain invariance through adversarial kernel learning. Three experiments were conducted to evaluate the performance against the state-of-the-art HDA methods on major image and text e-commerce benchmarks. HANDA shows statistically significant improvement in predictive performance. The practical utility of HANDA was shown in real-world dark web online markets. HANDA is an important step towards successful domain adaptation in e-commerce applications. Reza Ebrahimi 0001, Yidong Chai, Hao Helen Zhang 0001, Hsinchun Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | ACM KDD AI4Cyber/MLHat: Workshop on AI-enabled Cybersecurity Analytics and Deployable DefenseabstractFederal funding agencies and industry entities are seeking innovative approaches to address the ever-growing cybersecurity crisis. Increasingly, numerous cybersecurity thought leaders are indicating that Artificial Intelligence (AI)-enabled analytics can help tackle key cybersecurity tasks and deploy defenses. This half-day workshop, co-located with ACM KDD, sought to attain significant research contributions to various aspects of AI-enabled analytics for cybersecurity applications and deployable defense solutions from academics and practitioners. This workshop was a joint workshop of the 2021 AI-enabled Cybersecurity Analytics and 2021 International Workshop on Deployable Machine Learning for Security Defense. As such, we developed an interdisciplinary Program Committee with significant experience in various aspects of AI, cybersecurity, and/or deployable defense. Sagar Samtani, Gang Wang 0011, Ali Ahmadzadeh, Arridhana Ciptadi, Shanchieh Jay Yang, Hsinchun Chen |
KDD | 6 |
| 2022 | Explainable Artificial Intelligence for Cyber Threat Intelligence (XAI-CTI)abstractThe papers in this special section focus on explainable artificial intelligence for cyber threat intelligence. Despite concerted efforts from industry, academia, and government on improving cybersecurity capabilities, cyber-threats such as ransomware, fake news, advanced malware, and others, continue to exact a substantial toll on modern infrastructure and day-to-day societal operations. To help combat the ever-growing quantity and severity of cyber-threats, many organizations are adopting Cyber Threat Intelligence (CTI). At its core, CTI is a data-driven process that aims to identify emerging threats and key threat actors to help enable effective cybersecurity decision-making. Sagar Samtani, Hsinchun Chen, Murat Kantarcioglu, Bhavani Thuraisingham |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2021 | Distilling Contextual Embeddings Into A Static Word Embedding For Improving Hacker Forum AnalyticsabstractHacker forums provide malicious actors with a large database of tutorials, goods, and assets to leverage for cyber-attacks. Careful research of these forums can provide tremendous benefit to the cybersecurity community through trend identification and exploit categorization. This study aims to provide a novel static word embedding, Hack2Vec, to improve performance on hacker forum classification tasks. Our proposed Hack2Vec model distills contextual representations from the seminal pre-trained language model BERT to a continuous bag-of-words model to create a highly targeted hacker forum static word embedding. The results of our experimental design indicate that Hack2Vec improves performance over prominent embeddings in accuracy, precision, recall, and F1-score for a benchmark hacker forum classification task. Benjamin Ampel, Hsinchun Chen |
ISI | 2 |
| 2021 | Single-Shot Black-Box Adversarial Attacks Against Malware Detectors: A Causal Language Model ApproachabstractDeep Learning (DL)-based malware detectors are increasingly adopted for early detection of malicious behavior in cybersecurity. However, their sensitivity to adversarial malware variants has raised immense security concerns. Generating such adversarial variants by the defender is crucial to improving the resistance of DL-based malware detectors against them. This necessity has given rise to an emerging stream of machine learning research, Adversarial Malware example Generation (AMG), which aims to generate evasive adversarial malware variants that preserve the malicious functionality of a given malware. Within AMG research, black-box method has gained more attention than white-box methods. However, most black-box AMG methods require numerous interactions with the malware detectors to generate adversarial malware examples. Given that most malware detectors enforce a query limit, this could result in generating non-realistic adversarial examples that are likely to be detected in practice due to lack of stealth. In this study, we show that a novel DL-based causal language model enables single-shot evasion (i.e., with only one query to malware detector) by treating the content of the malware executable as a byte sequence and training a Generative Pre-Trained Transformer (GPT). Our proposed method, MalGPT, significantly outperformed the leading benchmark methods on a real-world malware dataset obtained from VirusTotal, achieving over 24.51% evasion rate. MalGPT enables cybersecurity researchers to develop advanced defense capabilities by emulating large-scale realistic AMG. James Lee Hu, Reza Ebrahimi 0001, Hsinchun Chen |
ISI | 3 |
| 2021 | Automated PII Extraction from Social Media for Raising Privacy Awareness: A Deep Transfer Learning ApproachabstractInternet users have been exposing an increasing amount of Personally Identifiable Information (PII) on social media. Such exposed PII can be exploited by cybercriminals and cause severe losses to the users. Informing users of their PII exposure in social media is crucial to raise their privacy awareness and encourage them to take protective measures. To this end, advanced techniques are needed to extract users’ exposed PII in social media automatically, whereas most existing studies remain manual. While Information Extraction (IE) techniques can be used to extract the PII automatically, Deep Learning (DL)-based IE models alleviate the need for feature engineering and further improve the efficiency. However, DL-based IE models often require large-scale labeled data for training, but PII-labeled social media posts are difficult to obtain due to privacy concerns. Also, these models rely heavily on pre-trained word embeddings, while PII in social media often varies in forms and thus has no fixed representations in pre-trained word embeddings. In this study, we propose the Deep Transfer Learning for PII Extraction (DTL-PIIE) framework to address these two limitations. DTL-PIIE transfers knowledge learned from publicly available PII data to social media in order to address the problem of rare PII-labeled data. Moreover, our framework leverages Graph Convolutional Networks (GCNs) to incorporate syntactic patterns to guide PIIE without relying on pre-trained word embeddings. Evaluation against benchmark IE models indicates that our approach outperforms state-of-the-art DL-based IE models. An ablation analysis further confirms the efficacy of each component in our model. Our proposed framework can facilitate various applications, such as PII misuse prediction and privacy risk assessment, thereby protecting the privacy of internet users. Fang Yu Lin, Reza Ebrahimi 0001, Weifeng Li 0002, Hsinchun Chen |
ISI | 5 |
| 2021 | Exploring the Evolution of Exploit-Sharing Hackers: An Unsupervised Graph Embedding ApproachabstractCybercrime was estimated to cost the global economy $945 billion in 2020. Increasingly, law enforcement agencies are using social network analysis (SNA) to identify key hackers from Dark Web hacker forums for targeted investigations. However, past approaches have primarily focused on analyzing key hackers at a single point in time and use a hacker’s structural features only. In this study, we propose a novel Hacker Evolution Identification Framework to identify how hackers evolve within hacker forums. The proposed framework has two novelties in its design. First, the framework captures features such as user statistics, node-level metrics, lexical measures, and post style, when representing each hacker with unsupervised graph embedding methods. Second, the framework incorporates mechanisms to align embedding spaces across multiple time-spells of data to facilitate analysis of how hackers evolve over time. Two experiments were conducted to assess the performance of prevailing graph embedding algorithms and nodal feature variations in the task of graph reconstruction in five time-spells. Results of our experiments indicate that Text-Associated Deep-Walk (TADW) with all of the proposed nodal features outperforms methods without nodal features in terms of Mean Average Precision in each time-spell. We illustrate the potential practical utility of the proposed framework with a case study on an English forum with 51,612 posts. The results produced by the framework in this case study identified key hackers posting piracy assets. Kaeli Otto, Benjamin Ampel, Sagar Samtani, Hongyi Zhu 0001, Hsinchun Chen |
ISI | 5 |
| 2021 | Identifying and Categorizing Malicious Content on Paste Sites: A Neural Topic Modeling ApproachabstractMalicious cyber activities impose substantial costs on the U.S. economy and global markets. Cyber-criminals often use information-sharing social media platforms such as paste sites (e.g., Pastebin) to share vast amounts of plain text content related to Personally Identifiable Information (PII), credit card numbers, exploit code, malware, and other sensitive content. Paste sites can provide targeted Cyber Threat Intelligence (CTI) about potential threats and prior breaches. In this research, we propose a novel Bidirectional Encoder Representation from Transformers (BERT) with Latent Dirichlet Allocation (LDA) model to categorize pastes automatically. Our proposed BERT-LDA model leverages a neural network transformer architecture to capture sequential dependencies when representing each sentence in a paste. BERT-LDA replaces the Bag-of-Words (BoW) approach in the conventional LDA with a Bag-of-Labels (BoL) that encompasses class labels at the sequence level. We compared the performance of the proposed BERT-LDA against the conventional LDA and BERT-LDA variants (e.g., GPT2-LDA) on 4,254,453 pastes from three paste sites. Experiment results indicate that the proposed BERT-LDA outperformed the standard LDA and each BERT-LDA variant in terms of perplexity on each paste site. Results of our BERT-LDA case study suggest that significant content relating to hacker community activities, malicious code, network and website vulnerabilities, and PII are shared on paste sites. The insights provided by this study could be used by organizations to proactively mitigate potential damage on their infrastructure. Tala Vahedi, Benjamin Ampel, Sagar Samtani, Hsinchun Chen |
ISI | 4 |
| 2021 | ACM KDD AI4Cyber: The 1st Workshop on Artificial Intelligence-enabled Cybersecurity AnalyticsabstractDespite significant contributions to various aspects of cybersecurity, cyber-attacks remain on the unfortunate rise. Increasingly, internationally recognized entities such as the National Science Foundation and National Science & Technology Council have noted Artificial Intelligence can help analyze billions of log files, Dark Web data, malware, and other data sources to help execute fundamental cybersecurity tasks. Our objective for the 1st Workshop on Artificial Intelligence-enabled Cybersecurity Analytics (half-day; co-located with ACM KDD) was to gather academic and practitioners to contribute recent work pertaining to AI-enabled cybersecurity analytics. We composed an outstanding, inter-disciplinary Program Committee with significant expertise in various aspects of AI-enabled Cybersecurity Analytics to evaluate the submitted work. Significant contributions to the half-day workshop were made in the areas of CTI, vulnerability assessment, and malware analysis. Sagar Samtani, Shanchieh Jay Yang, Hsinchun Chen |
KDD | 3 |
| 2021 | A Multimodal Event-Driven LSTM Model for Stock Prediction Using Online NewsabstractIn finance, it is believed that market information, namely, fundamentals and news information, affects stock movements. Such media-aware stock movements essentially comprise a multimodal problem. Two unique challenges arise in processing these multimodal data. First, information from one data mode will interact with information from other data modes. A common strategy is to concatenate various data modes into one compound vector; however, this strategy ignores the interactions among different modes. The second challenge is the heterogeneity of the data in terms of sampling time. Specifically, fundamental data consist of continuous values sampled at fixed time intervals, whereas news information emerges randomly. This heterogeneity can cause valuable information to be partially missing or can distort the feature spaces. In addition, the study of media-aware stock movements in previous work has focused on the one-to-one problem, in which it is assumed that news affects only the performance of the stocks mentioned in the reports. However, news articles also impact related stocks and cause stock co-movements. In this article, we propose a tensor-based event-driven LSTM model to address these challenges. Experiments performed on the China securities market demonstrate the superiority of the proposed approach over state-of-the-art algorithms, including AZFinText, eMAQT, and TeSIA. Qing Li 0005, Jinghua Tan, Jun Wang 0089, Hsinchun Chen |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2020 | Labeling Hacker Exploits for Proactive Cyber Threat Intelligence: A Deep Transfer Learning ApproachabstractWith the rapid development of new technologies, vulnerabilities are at an all-time high. Companies are investing in developing Cyber Threat Intelligence (CTI) to counteract these new vulnerabilities. However, this CTI is generally reactive based on internal data. Hacker forums can provide proactive CTI value through automated analysis of new trends and exploits. One way to identify exploits is by analyzing the source code that is posted on these forums. These source code snippets are often noisy and unlabeled, making standard data labeling techniques ineffective. This study aims to design a novel framework for the automated collection and categorization of hacker forum exploit source code. We propose a deep transfer learning framework, the Deep Transfer Learning for Exploit Labeling (DTL-EL). DTL-EL leverages the learned representation from professional labeled exploits to better generalize to hacker forum exploits. This model classifies the collected hacker forum exploits into eight predefined categories for proactive and timely CTI. The results of this study indicate that DTL-EL outperforms other prominent models in hacker forum literature. Benjamin Ampel, Sagar Samtani, Hongyi Zhu 0001, Steven Ullman, Hsinchun Chen |
ISI | 5 |
| 2020 | Identifying Vulnerable GitHub Repositories and Users in Scientific Cyberinfrastructure: An Unsupervised Graph Embedding ApproachabstractThe scientific cyberinfrastructure community heavily relies on public internet-based systems (e.g., GitHub) to share resources and collaborate. GitHub is one of the most powerful and popular systems for open source collaboration that allows users to share and work on projects in a public space for accelerated development and deployment. Monitoring GitHub for exposed vulnerabilities can save financial cost and prevent misuse and attacks of cyberinfrastructure. Vulnerability scanners that can interface with GitHub directly can be leveraged to conduct such monitoring. This research aims to proactively identify vulnerable communities within scientific cyberinfrastructure. We use social network analysis to construct graphs representing the relationships amongst users and repositories. We leverage prevailing unsupervised graph embedding algorithms to generate graph embeddings that capture the network attributes and nodal features of our repository and user graphs. This enables the clustering of public cyberinfrastructure repositories and users that have similar network attributes and vulnerabilities. Results of this research find that major scientific cyberinfrastructures have vulnerabilities pertaining to secret leakage and insecure coding practices for high-impact genomics research. These results can help organizations address their vulnerable repositories and users in a targeted manner. Ben Lazarine, Sagar Samtani, Mark W. Patton, Hongyi Zhu 0001, Steven Ullman, Benjamin Ampel, Hsinchun Chen |
ISI | 7 |
| 2020 | Identifying, Collecting, and Monitoring Personally Identifiable Information: From the Dark Web to the Surface WebabstractPersonally identifiable information (PII) has become a major target of cyber-attacks, causing severe losses to data breach victims. To protect data breach victims, researchers focus on collecting exposed PII to assess privacy risk and identify at-risk individuals. However, existing studies mostly rely on exposed PII collected from either the dark web or the surface web. Due to the wide exposure of PII on both the dark web and surface web, collecting from only the dark web or the surface web could result in an underestimation of privacy risk. Despite its research and practical value, jointly collecting PII from both sources is a non-trivial task. In this paper, we summarize our effort to systematically identify, collect, and monitor a total of 1,212,004,819 exposed PII records across both the dark web and surface web. Our effort resulted in 5.8 million stolen SSNs, 845,000 stolen credit/debit cards, and 1.2 billion stolen account credentials. From the surface web, we identified and collected over 1.3 million PII records of the victims whose PII is exposed on the dark web. To the best of our knowledge, this is the largest academic collection of exposed PII, which, if properly anonymized, enables various privacy research inquiries, including assessing privacy risk and identifying at-risk populations. Fang Yu Lin, Zara Ahmad-Post, Reza Ebrahimi 0001, James Lee Hu, Jingyu Xin, Weifeng Li 0002, Hsinchun Chen |
ISI | 9 |
| 2020 | Smart Vulnerability Assessment for Scientific Cyberinfrastructure: An Unsupervised Graph Embedding ApproachabstractThe accelerated growth of computing technologies has provided interdisciplinary teams a platform for producing innovative research at an unprecedented speed. Advanced scientific cyberinfrastructures, in particular, provide data storage, applications, software, and other resources to facilitate the development of critical scientific discoveries. Users of these environments often rely on custom developed virtual machine (VM) images that are comprised of a diverse array of open source applications. These can include vulnerabilities undetectable by conventional vulnerability scanners. This research aims to identify the installed applications, their vulnerabilities, and how they vary across images in scientific cyberinfrastructure. We propose a novel unsupervised graph embedding framework that captures relationships between applications, as well as vulnerabilities identified on corresponding GitHub repositories. This embedding is used to cluster images with similar applications and vulnerabilities. We evaluate cluster quality using Silhouette, Calinski-Harabasz, and Davies-Bouldin indices, and application vulnerabilities through inspection of selected clusters. Results reveal that images pertaining to genomics research in our research testbed are at greater risk of high-severity shell spawning and data validation vulnerabilities. Steven Ullman, Sagar Samtani, Ben Lazarine, Hongyi Zhu 0001, Benjamin Ampel, Mark W. Patton, Hsinchun Chen |
ISI | 7 |
| 2020 | A Generative Adversarial Learning Framework for Breaking Text-Based CAPTCHA in the Dark WebabstractCyber threat intelligence (CTI) necessitates automated monitoring of dark web platforms (e.g., Dark Net Markets and carding shops) on a large scale. While there are existing methods for collecting data from the surface web, large-scale dark web data collection is commonly hindered by anti-crawling measures. Text-based CAPTCHA serves as the most prohibitive type of these measures. Text-based CAPTCHA requires the user to recognize a combination of hard-to-read characters. Dark web CAPTCHA patterns are intentionally designed to have additional background noise and variable character length to prevent automated CAPTCHA breaking. Existing CAPTCHA breaking methods cannot remedy these challenges and are therefore not applicable to the dark web. In this study, we propose a novel framework for breaking text-based CAPTCHA in the dark web. The proposed framework utilizes Generative Adversarial Network (GAN) to counteract dark web-specific background noise and leverages an enhanced character segmentation algorithm. Our proposed method was evaluated on both benchmark and dark web CAPTCHA testbeds. The proposed method significantly outperformed the state-of-the-art baseline methods on all datasets, achieving over 92.08% success rate on dark web testbeds. Our research enables the CTI community to develop advanced capabilities of large-scale dark web monitoring. Reza Ebrahimi 0001, Weifeng Li 0002, Hsinchun Chen |
ISI | 4 |
| 2020 | Proactively Identifying Emerging Hacker Threats from the Dark Web: A Diachronic Graph Embedding Framework (D-GEF)abstractCybersecurity experts have appraised the total global cost of malicious hacking activities to be $450 billion annually. Cyber Threat Intelligence (CTI) has emerged as a viable approach to combat this societal issue. However, existing processes are criticized as inherently reactive to known threats. To combat these concerns, CTI experts have suggested proactively examining emerging threats in the vast, international online hacker community. In this study, we aim to develop proactive CTI capabilities by exploring online hacker forums to identify emerging threats in terms of popularity and tool functionality. To achieve these goals, we create a novel Diachronic Graph Embedding Framework (D-GEF). D-GEF operates on a Graph-of-Words (GoW) representation of hacker forum text to generate word embeddings in an unsupervised manner. Semantic displacement measures adopted from diachronic linguistics literature identify how terminology evolves. A series of benchmark experiments illustrate D-GEF's ability to generate higher quality than state-of-the-art word embedding models (e.g., word2vec) in tasks pertaining to semantic analogy, clustering, and threat classification. D-GEF's practical utility is illustrated with in-depth case studies on web application and denial of service threats targeting PHP and Windows technologies, respectively. We also discuss the implications of the proposed framework for strategic, operational, and tactical CTI scenarios. All datasets and code are publicly released to facilitate scientific reproducibility and extensions of this work. Sagar Samtani, Hongyi Zhu 0001, Hsinchun Chen |
ACM Trans. Priv. Secur. | 3 |
| 2020 | Corrections to "NATERGM: A Model for Examining the Role of Nodal Attributes in Dynamic Social Media Networks"abstractPresents corrections to affiliation information in the above named paper. Shan Jiang 0002, Hsinchun Chen |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2020 | A Deep Learning Architecture for Psychometric Natural Language ProcessingabstractPsychometric measures reflecting people’s knowledge, ability, attitudes, and personality traits are critical for many real-world applications, such as e-commerce, health care, and cybersecurity. However, traditional methods cannot collect and measure rich psychometric dimensions in a timely and unobtrusive manner. Consequently, despite their importance, psychometric dimensions have received limited attention from the natural language processing and information retrieval communities. In this article, we propose a deep learning architecture, PyNDA, to extract psychometric dimensions from user-generated texts. PyNDA contains a novel representation embedding, a demographic embedding, a structural equation model (SEM) encoder, and a multitask learning mechanism designed to work in unison to address the unique challenges associated with extracting rich, sophisticated, and user-centric psychometric dimensions. Our experiments on three real-world datasets encompassing 11 psychometric dimensions, including trust, anxiety, and literacy, show that PyNDA markedly outperforms traditional feature-based classifiers as well as the state-of-the-art deep learning architectures. Ablation analysis reveals that each component of PyNDA significantly contributes to its overall performance. Collectively, the results demonstrate the efficacy of the proposed architecture for facilitating rich psychometric analysis. Our results have important implications for user-centric information extraction and retrieval systems looking to measure and incorporate psychometric dimensions. Ahmed Abbasi, David G. Dobolyi, Richard G. Netemeyer, Gari D. Clifford, Hsinchun Chen |
ACM Trans. Inf. Syst. | 7 |
| 2019 | Performance Modeling of Hyperledger Sawtooth BlockchainabstractWith the rapid development of blockchain platforms, it is important that different implementations are tested and analyzed for comparative purposes. One such implementation is Hyperledger Sawtooth, a new member of the Hyperledger family. Sawtooth blockchain is a permissioned implementation developed in part by Intel. While research has been done on Hyperledger Fabric, research on Sawtooth is not well documented. Using the Hyperledger Caliper benchmarking tool, we aim to test the performance of the blockchain and identify potential issues. Benjamin Ampel, Mark W. Patton, Hsinchun Chen |
ISI | 3 |
| 2019 | Dark-Net Ecosystem Cyber-Threat Intelligence (CTI) ToolabstractThe frequency and costs of cyber-attacks are increasing each year. By the end of 2019, the total cost of data breaches is expected to reach $2.1 trillion through the ever-growing online presence of enterprises and their consumers. The tools to perform these attacks and the breached data can often be purchased within the Dark-net. Many of the threat actors within this realm use its various platforms to broker, discuss, and strategize these cyber-threat assets. To combat these attacks, researchers are developing Cyber-Threat Intelligence (CTI) tools to proactively monitor the ever-growing online hacker community. This paper will detail the creation and use of a CTI tool that leverages a social network to identify cyber-threats across major Dark-net data sources. Through this network, emerging threats can be quickly identified so proactive or reactive security measures can be implemented. Nolan Arnold, Reza Ebrahimi 0001, Ben Lazarine, Mark W. Patton, Hsinchun Chen, Sagar Samtani |
ISI | 6 |
| 2019 | Identifying High-Impact Opioid Products and Key Sellers in Dark Net Marketplaces: An Interpretable Text Analytics ApproachabstractAs the Internet based applications become more and more ubiquitous, drug retailing on Dark Net Marketplaces (DNMs) has raised public health and law enforcement concerns due to its highly accessible and anonymous nature. To combat illegal drug transaction among DNMs, authorities often require agents to impersonate DNM customers in order to identify key actors within the community. This process can be costly in time and resource. Research in DNMs have been conducted to provide better understanding of DNM characteristics and drug sellers' behavior. Built upon the existing work, researchers can further leverage predictive analytics techniques to take proactive measures and reduce the associated costs. To this end, we propose a systematic analytical approach to identify key opioid sellers in DNMs. Utilizing machine learning and text analysis, this research provides prediction of high-impact opioid products in two major DNMs. Through linking the high-impact products and their sellers, we then identify the key opioid sellers among the communities. This work intends to help law enforcement authorities to formulate strategies by providing specific targets within the DNMs and reduce the time and resources required for prosecuting and eliminating the criminals from the market. Po-Yi Du, Reza Ebrahimi 0001, Hsinchun Chen, Randall A. Brown, Sagar Samtani |
ISI | 4 |
| 2018 | Identifying, Collecting, and Presenting Hacker Community Data: Forums, IRC, Carding Shops, and DNMsabstractCyber-attacks cost the global economy over $450 billion annually. To combat this issue, researchers and practitioners put enormous efforts into developing Cyber Threat Intelligence, or the process of identifying emerging threats and key hackers. However, the reliance on internal network data to has resulted in inherently reactive intelligence. CTI experts have urged the importance of proactively studying the large, ever-evolving online hacker community. Despite their CTI value, collecting data from hacker community platforms is a non-trivial task. In this paper, we summarize our efforts in systematically identifying and automatically collecting a large-scale of hacker forums, carding shops, Internet-Relay-Chat, and Dark Net Marketplaces. We also present our efforts to provide this data to the larger CTI community via the AZSecure Hacker Assets Portal (www.azsecure-hap.com). With our methodology, we collected 102 platforms for a total of 43,981,647 records. To the best of our knowledge, this compilation of hacker community data is the largest such collection in academia. Po-Yi Du, Reza Ebrahimi 0001, Sagar Samtani, Ben Lazarine, Nolan Arnold, Rachael Dunn, Sandeep Suntwal, Guadalupe Angeles, Robert Schweitzer, Hsinchun Chen |
ISI | 11 |
| 2018 | Detecting Cyber Threats in Non-English Dark Net Markets: A Cross-Lingual Transfer Learning ApproachabstractRecent advances in proactive cyber threat intelligence rely on early detection of cyber threats in hacker communities. Dark Net Markets (DNMs) are growing platforms in hacker community that provide hackers with highly- specialized tools and products which may not be found in other platforms. While text classification techniques have been used for cyber threat detection in English DNMs, the task is hindered in non-English platforms due to the language barrier and lack of ground-truth data. Current approaches use monolingual models on machine translated data to overcome these challenges. However, the translation errors can deteriorate the classification results. The abundance of data in English DNMs can be leveraged in learning non-English threats without using machine translation. In this study, we show that a deep cross-lingual model that can jointly learn the common language representation from two languages, significantly outperforms a monolingual model learned on machine translated data for identifying cyber threats in non-English DNMs. Unlike most studies, our approach does not require any external data source such as bilingual word embeddings or bilingual lexicons. Our experiments on Russian DNMs show that this approach can achieve better performance than state-of-the-art methods for non-English cyber threat detection in malicious hacker community. Reza Ebrahimi 0001, Mihai Surdeanu, Sagar Samtani, Hsinchun Chen |
ISI | 4 |
| 2018 | Vulnerability Assessment, Remediation, and Automated Reporting: Case Studies of Higher Education InstitutionsabstractScientific advances of higher education institutions make them attractive targets for malicious cyber-attacks. Modern scanners such as Nessus and Burp can pinpoint an organization's vulnerabilities for subsequent mitigation. However, the remediation reports generated from the tools often cause significant information overload while failing to provide actionable solutions. Consequently, higher education institutions lack the appropriate knowledge to improve their cybersecurity posture. In this study, we conduct a large-scale vulnerability assessment of 272 higher education institutions. From the results, we identified vulnerabilities that fail to provide comprehensive remediation strategies. Selected flaws are recreated and remediated in a virtual environment to develop enhanced, automated reporting mechanisms that provide succinct reports to enable the efficient vulnerability remediation. Our enhanced reports address 27.80% of vulnerabilities found in scanned higher education institutions. Christopher R. Harrell, Mark W. Patton, Hsinchun Chen, Sagar Samtani |
ISI | 3 |
| 2018 | Benchmarking Vulnerability Assessment Tools for Enhanced Cyber-Physical System (CPS) ResiliencyabstractCyber-Physical Systems (CPSs) are engineered systems seamlessly integrating computational algorithms and physical components. CPS advances offer numerous benefits to domains such as health, transportation, smart homes and manufacturing. Despite these advances, the overall cybersecurity posture of CPS devices remains unclear. In this paper, we provide knowledge on how to improve CPS resiliency by evaluating and comparing the accuracy, and scalability of two popular vulnerability assessment tools, Nessus and OpenVAS. Accuracy and suitability are evaluated with a diverse sample of pre-defined vulnerabilities in Industrial Control Systems (ICS), smart cars, smart home devices, and a smart water system. Scalability is evaluated using a large-scale vulnerability assessment of 1,000 Internet accessible CPS devices found on Shodan, the search engine for the Internet of Things (IoT). Assessment results indicate several CPS devices from major vendors suffer from critical vulnerabilities such as unsupported operating systems, OpenSSH vulnerabilities allowing unauthorized information disclosure, and PHP vulnerabilities susceptible to denial of service attacks. Emma McMahon, Mark W. Patton, Sagar Samtani, Hsinchun Chen |
ISI | 4 |
| 2018 | Incremental Hacker Forum Exploit Collection and Classification for Proactive Cyber Threat Intelligence: An Exploratory StudyabstractCyber threats have emerged as a key societal concern. To counter the growing threat of cyber-attacks, organizations, in recent years, have begun investing heavily in developing Cyber Threat Intelligence (CTI). Fundamentally a data driven process, many organizations have traditionally collected and analyzed data from internal log files, resulting in reactive CTI. The online hacker community can offer significant proactive CTI value by alerting organizations to threats they were not previously aware of. Amongst various platforms, forums provide the richest metadata, data permanence, and tens of thousands of freely available Tools, Techniques, and Procedures (TTP). However, forums often employ anti-crawling measures such as authentication, throttling, and obfuscation. Such limitations have restricted many researchers to batch collections. This exploratory study aims to (1) design a novel web crawler augmented with numerous anti-crawling countermeasures to collect hacker exploits on an ongoing basis, (2) employ a state-of-the-art deep learning approach, Long Short-Term Memory (LSTM) Recurrent Neural Network (RNN), to automatically classify exploits into pre-defined categories on-the-fly, and (3) develop interactive visualizations enabling CTI practitioners and researchers to explore collected exploits for proactive, timely CTI. The results of this study indicate, among other findings, that system and network exploits are shared significantly more than other exploit types. Sagar Samtani, Mark W. Patton, Hsinchun Chen |
ISI | 4 |
| 2018 | A sequence-to-sequence model-based deep learning approach for recognizing activity of daily living for senior care
Hongyi Zhu 0001, Hsinchun Chen, Randall A. Brown |
J. Biomed. Informatics | 2 |
| 2018 | Hidden Markov Model-Based Fall Detection With Motion Sensor Orientation Calibration: A Case for Real-Life Home MonitoringabstractFalls are a major threat for senior citizens' independent living. Motion sensor technologies and automatic fall detection systems have emerged as a reliable low-cost solution to this challenge. We develop a hidden Markov model (HMM) based fall detection system to detect falls automatically using a single motion sensor for real-life home monitoring scenarios. We propose a new representation for acceleration signals in HMMs to avoid feature engineering and developed a sensor orientation calibration algorithm to resolve sensor misplacement issues (misplaced sensor location and misaligned sensor orientation) in real-world scenarios. HMM classifiers are trained to detect falls based on acceleration signal data collected from motion sensors. We collect a dataset from experiments of simulated falls and normal activities and acquired a dataset from a real-world fall repository (FARSEEING) to evaluate our system. Our system achieves positive predictive value of 0.981 and sensitivity of 0.992 on the experiment dataset with 200 fall events and 385 normal activities, and positive predictive value of 0.786 and sensitivity of 1.000 on the real-world fall dataset with 22 fall events and 2618 normal activities. Our system's results significantly outperform benchmark systems, which shows the advantage of our HMM-based fall detection system with sensor orientation calibration. Our fall detection system is able to precisely detect falls in real-life home scenarios with a reasonably low false alarm ratet. Shuo Yu 0002, Hsinchun Chen, Randall A. Brown |
IEEE J. Biomed. Health Informatics | 2 |
| 2018 | Web Media and Stock Markets : A Survey and Future Directions from a Big Data PerspectiveabstractStock market volatility is influenced by information release, dissemination, and public acceptance. With the increasing volume and speed of social media, the effects of Web information on stock markets are becoming increasingly salient. However, studies of the effects of Web media on stock markets lack both depth and breadth due to the challenges in automatically acquiring and analyzing massive amounts of relevant information. In this study, we systematically reviewed 229 research articles on quantifying the interplay between Web media and stock markets from the fields of Finance, Management Information Systems, and Computer Science. In particular, we first categorized the representative works in terms of media type and then summarized the core techniques for converting textual information into machine-friendly forms. Finally, we compared the analysis models used to capture the hidden relationships between Web media and stock movements. Our goal is to clarify current cutting-edge research and its possible future directions to fully understand the mechanisms of Web information percolation and its impact on stock markets from the perspectives of investors cognitive behaviors, corporate governance, and stock market regulation. Qing Li 0005, Yan Chen 0016, Jun Wang 0089, Yuanzhu Peter Chen, Hsinchun Chen |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2018 | Supervised Topic Modeling Using Hierarchical Dirichlet Process-Based Inverse Regression: Experiments on E-Commerce ApplicationsabstractThe proliferation of e-commerce calls for mining consumer preferences and opinions from user-generated text. To this end, topic models have been widely adopted to discover the underlying semantic themes (i.e., topics). Supervised topic models have emerged to leverage discovered topics for predicting the response of interest (e.g., product quality and sales). However, supervised topic modeling remains a challenging problem because of the need to prespecify the number of topics, the lack of predictive information in topics, and limited scalability. In this paper, we propose a novel supervised topic model, Hierarchical Dirichlet Process-based Inverse Regression (HDP-IR). HDP-IR characterizes the corpus with a flexible number of topics, which prove to retain as much predictive information as the original corpus. Moreover, we develop an efficient inference algorithm capable of examining large-scale corpora (millions of documents or more). Three experiments were conducted to evaluate the predictive performance over major e-commerce benchmark testbeds of online reviews. Overall, HDP-IR outperformed existing state-of-the-art supervised topic models. Particularly, retaining sufficient predictive information improved predictive R-squared by over 17.6 percent; having topic structure flexibility contributed to predictive R-squared by at least 4.1 percent. HDP-IR provides an important step for future study on user-generated texts from a topic perspective. Weifeng Li 0002, Junming Yin, Hsinchun Chen |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2017 | Benchmarking vulnerability scanners: An experiment on SCADA devices and scientific instrumentsabstractCybersecurity is a critical concern in society today. One common avenue of attack for malicious hackers is exploiting vulnerable websites. It is estimated that there are over one million websites that are attacked daily. Two emerging targets of such attacks are Supervisory Control and Data Acquisition (SCADA) devices and scientific instruments. Vulnerability assessment tools can help provide owners of these devices with the knowledge on how to protect their infrastructure. However, owners face difficulties in identifying which tools are ideal for their assessments. This research aims to benchmark two state-of-the-art vulnerability assessment tools, Nessus and Burp Suite, in the context of SCADA devices and scientific instruments. We specifically focus on identifying the accuracy, scalability, and vulnerability results of the scans. Results of our study indicate that both tools together can provide a comprehensive assessment of the vulnerabilities in SCADA devices and scientific instruments. Malaka El, Emma McMahon, Sagar Samtani, Mark W. Patton, Hsinchun Chen |
ISI | 5 |
| 2017 | Identifying mobile malware and key threat actors in online hacker forums for proactive cyber threat intelligenceabstractCyber-attacks are constantly increasing and can prove difficult to mitigate, even with proper cybersecurity controls. Currently, cyber threat intelligence (CTI) efforts focus on internal threat feeds such as antivirus and system logs. While this approach is valuable, it is reactive in nature as it relies on activity which has already occurred. CTI experts have argued that an actionable CTI program should also provide external, open information relevant to the organization. By finding information about malicious hackers prior to an attack, organizations can provide enhanced CTI and better protect their infrastructure. Hacker forums can provide a rich data source in this regard. This research aims to proactively identify mobile malware and associated key authors. Specifically, we use a state-of-the-art neural network architecture, recurrent neural networks, to identify mobile malware attachments followed by social network analysis techniques to determine key hackers disseminating the mobile malware. Results of this study indicate that many identified attachments are zipped Android apps made by threat actors holding administrative positions in hacker forums. Our identified mobile malware attachments are consistent with some of the emerging mobile malware concerns as highlighted by industry leaders. John Grisham, Sagar Samtani, Mark W. Patton, Hsinchun Chen |
ISI | 4 |
| 2017 | Assessing medical device vulnerabilities on the Internet of ThingsabstractInternet enabled medical devices offer patients with a level of convenience. In recent years, the healthcare industry has seen a surge in the number of cyber-attacks. Given the potentially fatal impact of a compromised medical device, this study aims to identify vulnerabilities of medical devices. Our approach uses Shodan to obtain a large collection of IP addresses that will be passed through Nessus to verify if any vulnerabilities exist. We determined some devices manufactured by primary vendors such as Omron Corporation, FORA, Roche, and Bionet contain serious vulnerabilities such as Dropbear SSH Server and MS17-010. These allow remote execution of code and authentication bypassing potentially giving attackers control of their systems. Emma McMahon, Malaka El, Sagar Samtani, Mark W. Patton, Hsinchun Chen |
ISI | 6 |
| 2017 | Identifying vulnerabilities of consumer Internet of Things (IoT) devices: A scalable approachabstractThe Internet of Things becomes more defined year after year. Companies are looking for novel ways to implement various smart capabilities into their products that increase interaction between users and other network devices. While many smart devices offer greater convenience and value, they also present new security vulnerabilities that can have a detrimental effect on consumer privacy. Given the societal impact of IoT device vulnerabilities, this study aims to perform a large-scale vulnerability assessment of consumer IoT devices exposed on the Internet. Specifically, Shodan is used to collect a large testbed of consumer IoT devices which are then passed through Nessus to determine whether potential vulnerabilities exist. Results of this study indicate that a significant number of consumer IoT devices are vulnerable to exploits that can compromise user information and privacy. Emma McMahon, Sagar Samtani, Mark W. Patton, Hsinchun Chen |
ISI | 5 |
| 2016 | Deep Learning Based Topic Identification and Categorization: Mining Diabetes-Related Topics on Chinese Health Websites
Xinhuan Chen, Yong Zhang 0002, Jennifer Jie Xu 0001, Chunxiao Xing, Hsinchun Chen |
DASFAA (1) | 5 |
| 2016 | NATERGM: A model for examining the role of Nodal Attributes in dynamic social media networksabstractSocial media networks are dynamic. As such, the order in which network ties develop is an important aspect of the network dynamics.This study proposes a novel dynamic network model, the Nodal Attribute-based Temporal Exponential Random Graph Model (NATERGM) for dynamic network analysis. The proposed model focuses on how the nodal attributes of a network affect the order in which the network ties develop. Empirical results showed that the NATERGM demonstrated an enhanced pattern testing capability compared to benchmark models. The proposed NATERGM model helps explain the roles of nodal attributes in the formation process of dynamic networks. Shan Jiang 0002, Hsinchun Chen |
ICDE | 2 |
| 2016 | Identifying language groups within multilingual cybercriminal forumsabstractOnline cybercriminal communities exist in various geopolitical regions, including America, China, Russia, and more. Some multilingual forums exist where cybercriminals of differing geopolitical origin interact and exchange hacking knowledge and cybercriminal assets. Researchers can study such forums to better understand the global cybercriminal supply chain and cybercrime trends. However, little work has focused on identifying members of different language groups and geopolitical origin within such forums. One challenge is the necessity of a technique that scales across multiple languages. We are motivated to explore computational techniques that support automated and scalable categorization of cybercriminal forum participants into varying language groups. In particular, we make use of Paragraph Vectors, a state-of-the-art neural network language model to generate fixed-length vector representations (i.e., document embeddings) of messages posted by forum participants. Results indicate Paragraph Vectors outperforms traditional n-gram frequency approaches for generating document embeddings that are useful for clustering cybercriminals into language groups. Victor A. Benjamin, Hsinchun Chen |
ISI | 2 |
| 2016 | Shodan visualizedabstractThe purpose of this paper is to discuss how using Gephi to visualize the open ports at IP addresses in Shodan may provide a means of identifying SCADA devices. Visualizations were created using both IP addresses and open ports as nodes. Modularity, centralities, and layout were used to enhance the visualizations. From these visualizations we hope to gather a better understanding of what devices are on the network. Vincent J. Ercolani, Mark W. Patton, Hsinchun Chen |
ISI | 3 |
| 2016 | Exploring key hackers and cybersecurity threats in Chinese hacker communitiesabstractChinese hacker communities are of interest to cybersecurity researchers and investigators. When examining Chinese hacker communities, researchers and investigators face many challenges, including understanding the Chinese language, detecting variations in topic evolution, and identifying key hackers with their specialty areas. Therefore, we are motivated to develop a framework for analyzing key hackers and emerging threats in Chinese hacker communities. Specifically, we develop a set of topic models for extracting popular topics, tracking topic evolution, and identifying key hackers with their specialty topics. We applied our framework to 19 major Chinese hacker communities. As a result, we identified five major popular topics, including trading, fraud prevention & identification, calling for cooperation, casual chat, and monetizing. Moreover, we found several trends related to new communication channels, new stolen cards of interest, and new operating mechanism. Further, we also found the key hackers in each extracted area. Our work contributes to the cybersecurity literature by providing an advanced and scalable framework for analyzing Chinese hacker communities. Qiang Wei 0001, Yong Zhang 0002, Chunxiao Xing, Weifeng Li 0002, Hsinchun Chen |
ISI | 8 |
| 2016 | Identifying top listers in Alphabay using Latent Dirichlet AllocationabstractThis poster analyzes the Alphabay underground marketplace - an anonymous trading grounds for illicit goods and services. Listing data was collected and interpreted using Latent-Dirichlet Allocation (LDA), to determine common topics in the listings. Results found offer insight to the types of goods being sold and who is selling them. John Grisham, Calvin Barreras, Cyran Afarin, Mark W. Patton, Hsinchun Chen |
ISI | 5 |
| 2016 | Exploring the online underground marketplaces through topic-based social network and clusteringabstractCyber fraud causes significant losses to the economy and has become a lucrative form of illicit business by leveraging the Internet as a communication channel. Criminals in the cyber fraud underground economy use online underground marketplaces and other forms of social media to exchange information and trade stolen information. Analyzing these underground marketplaces is challenging due to the variability of both members and marketplaces. To understand more about the underground economy and the actors in it, we propose a topic-based social network analysis and clustering approach to identify the key members and their roles in the cyber fraud value chain. An experiment is conducted using data from several online underground marketplaces in China. Results suggest that the proposed method can aid in identifying key members in terms of roles, influence levels, and their social relationships. Shin-Ying Huang, Hsinchun Chen |
ISI | 2 |
| 2016 | SCADA honeypots: An in-depth analysis of ConpotabstractSupervisory Control and Data Acquisition (SCADA) honeypots are key tools not only for determining threats which pertain to SCADA devices in the wild, but also for early detection of potential malicious tampering within a SCADA device network. An analysis of one such SCADA honeypot, Conpot, is conducted to determine its viability as an effective SCADA emulating device. A long-term analysis is conducted and a simple scoring mechanism leveraged to evaluate the Conpot honeypot. Arthur Jicha, Mark W. Patton, Hsinchun Chen |
ISI | 3 |
| 2016 | Identifying devices across the IPv4 address spaceabstractMany of today's devices are internet-enabled with IPv4 internet addresses, exposing them to internet threats. To determine the true scale of vulnerabilities being introduced, particularly in the IPv4 internet address space, a new methodology of scanning the entire IPv4 internet space is required. To improve scanning speeds we created a framework combining fast connectionless port scanners with a thorough and accurate connection-oriented scanner to verify results. The results are stored to a database. This combined framework provides more robust results than current connectionless scanners, yet still scans the IPv4 internet fast enough to be practically usable for mass scanning. Ryan Jicha, Mark W. Patton, Hsinchun Chen |
ISI | 3 |
| 2016 | Identifying the socio-spatial dynamics of terrorist attacks in the Middle EastabstractTerrorist attacks change dynamically in social and geographic spaces. In this paper, terrorist attacks in the Middle East are analyzed using methods of network science, statistical methods, geographic information science, and artificial neural networks designed from a socio-spatial perspective. Based on the Global Terrorism Database (GTD), firstly the distribution and trends of terrorist attacks are detected. Then approaches for building diffusion network and identifying diffusion patterns of transnational and transyearly attacks are developed. Finally a Back Propagation Neural Network (BPNN) model is built for predicting future attacks. Results lead to a greater understanding of socio-spatial dependencies and diffusion regularities of terrorist attacks. The findings have significant implications for multinational security and the need to coordinate transnationally. Ze Li 0002, Duoyong Sun, Hsinchun Chen, Shin-Ying Huang |
ISI | 3 |
| 2016 | Targeting key data breach services in underground supply chainabstractOver the past decade, a growing body of cybercriminals have founded an underground supply chain to facilitate data breaches, leading to the leak of personal information for hundreds of millions of individuals. As many service providers in the supply chain are rippers, cybercriminals tend to rely on a few key services. Identifying key services is of great interest to both cybersecurity researchers and practitioners. This study presents a text-mining framework for identifying key data breach services based on analysis of service reviews. The framework includes crawlers with counter anti-crawling measures, text preprocessing, and supervised topic models. In our experiment, more than 70% of the key services were identified by our framework. Weifeng Li 0002, Junming Yin, Hsinchun Chen |
ISI | 3 |
| 2016 | Anonymous port scanning: Performing network reconnaissance through TorabstractThe anonymizing network Tor is examined as one method of anonymizing port scanning tools and avoiding identification and retaliation. Performing anonymized port scans through Tor is possible using Nmap, but parallelization of the scanning processes is required to accelerate the scan rate. Rodney Rohrmann, Mark W. Patton, Hsinchun Chen |
ISI | 3 |
| 2016 | Using social network analysis to identify key hackers for keylogging tools in hacker forumsabstractCyber-attacks are critical cybersecurity concerns across the world. Catching malicious hackers prior to a cyber-attack can save significant financial cost as well as avoid devastating cyber-attacks. Current methods of identifying and reprimanding hackers generally occurs after an attack and is reactive in nature. This research aims to proactively identify key hackers who are creating and disseminating malicious tools within hacker forums. Specifically, we utilize social network analysis techniques to systematically identify key hackers for keylogging tools within a large English hacker forum. Results of this study indicate that many key hackers are the most senior, longest tenured participants of their community. Sagar Samtani, Hsinchun Chen |
ISI | 2 |
| 2016 | AZSecure Hacker Assets Portal: Cyber threat intelligence and malware analysisabstractCyber threats pose grave national security dangers to the US. Many cyber-attacks today are executed with ever-growing collection of malicious tools. Cyber threat intelligence (CTI) and malware analysis portals aim to provide knowledge and tools to help prevent and mitigate attacks. However, current CTI and malware analysis portals and techniques have been criticized for being too reactive as they rely on data collected from past cyber-attacks. Online hacker forums provide a novel source of data that can inform a proactive CTI and malware portal. This research demonstrates the AZSecure Hacker Assets Portal. This portal collects and analyzes malicious assets directly from the largely untapped and rich data source of online hacker communities by utilizing state-of-the-art machine learning techniques. This paper explores the development and evolution of the AZSecure Hacker Assets Portal. We also present key portal functionalities such as asset searching, browsing, and downloading, source code visualizations and code comparison analytics, and an interactive CTI dashboard. Sagar Samtani, Kory Chinn, Cathy Larson, Hsinchun Chen |
ISI | 4 |
| 2016 | Identifying SCADA vulnerabilities using passive and active vulnerability assessment techniquesabstractCritical infrastructure such as power plants, oil refineries, and sewage are at the core of modern society. Supervisory Control and Data Acquisition (SCADA) systems were designed to allow human operators supervise, maintain, and control critical infrastructure. Recent years has seen an increase in connectivity of SCADA systems to the Internet. While this connectivity provides an increased level of convenience, it also increases their susceptibility to cyber-attacks. Given the potentially severe ramifications of exploiting SCADA systems, the purpose of this study is to utilize passive and active vulnerability assessment techniques to identify the vulnerabilities of Internet enabled SCADA systems. Specifically, we collect a large testbed of SCADA devices from Shodan, a search engine for the IoT, and assess their vulnerabilities with Nessus and against the National Vulnerability Database (NVD). Results of this study indicate that many SCADA systems from major vendors such as Rockwell Automation and Siemens are vulnerable to default credential, man-in-the-middle, and SSH exploit attacks. Sagar Samtani, Shuo Yu 0002, Hongyi Zhu 0001, Mark W. Patton, Hsinchun Chen |
ISI | 5 |
| 2016 | Chinese underground market jargon analysis based on unsupervised learningabstractWith the rapid growth of online population, China has become the world's largest online market. This also gives rise to the Chinese underground market, which has facilitated many of the cybercrimes in China. Consequently, there is a need for research scrutinizing Chinese underground markets. One major challenge facing cybersecurity researchers is to understand the unfamiliar cybercriminal jargons. To this end, we are motivated to analyze jargons in Chinese underground market. Particularly, we utilize the recent advancements in unsupervised machine learning methods, word embedding and Latent Dirichlet Allocation. We evaluate our work on a research testbed encompassing 29 exclusive underground market QQ groups with 23,000 members. Specifically, we test the ability of the proposed approach to learn semantically similar words of known cybersecurity-related jargons. Results suggest the state-of-the-art unsupervised learning approaches can help better understand cybercriminal language, providing promising insights for future research on Chinese underground markets. Kangzhi Zhao, Yong Zhang 0002, Chunxiao Xing, Weifeng Li 0002, Hsinchun Chen |
ISI | 5 |
| 2016 | NATERGM: A Model for Examining the Role of Nodal Attributes in Dynamic Social Media NetworksabstractSocial media networks are dynamic. As such, the order in which network ties develop is an important aspect of the network dynamics. This study proposes a novel dynamic network model, the Nodal Attribute-based Temporal Exponential Random Graph Model (NATERGM) for dynamic network analysis. The proposed model focuses on how the nodal attributes of a network affect the order in which the network ties develop. Temporal patterns in social media networks are modeled based on the nodal attributes of individuals and the time information of network ties. Using social media data collected from a knowledge sharing community, empirical tests were conducted to evaluate the performance of the NATERGM on identifying the temporal patterns and predicting the characteristics of the future networks. Results showed that the NATERGM demonstrated an enhanced pattern testing capability and an increased prediction accuracy of network characteristics compared to benchmark models. The proposed NATERGM model helps explain the roles of nodal attributes in the formation process of dynamic networks. Shan Jiang 0002, Hsinchun Chen |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2016 | A Tensor-Based Information Framework for Predicting the Stock MarketabstractTo study the influence of information on the behavior of stock markets, a common strategy in previous studies has been to concatenate the features of various information sources into one compound feature vector, a procedure that makes it more difficult to distinguish the effects of different information sources. We maintain that capturing the intrinsic relations among multiple information sources is important for predicting stock trends. The challenge lies in modeling the complex space of various sources and types of information and studying the effects of this information on stock market behavior. For this purpose, we introduce a tensor-based information framework to predict stock movements. Specifically, our framework models the complex investor information environment with tensors. A global dimensionality-reduction algorithm is used to capture the links among various information sources in a tensor, and a sequence of tensors is used to represent information gathered over time. Finally, a tensor-based predictive model to forecast stock movements, which is in essence a high-order tensor regression learning problem, is presented. Experiments performed on an entire year of data for China Securities Index stocks demonstrate that a trading system based on our framework outperforms the classic Top- N trading strategy and two state-of-the-art media-aware trading algorithms. Qing Li 0005, Yuanzhu Peter Chen, LiLing Jiang, Ping Li 0060, Hsinchun Chen |
ACM Trans. Inf. Syst. | 5 |
| 2015 | Tensor-Based Learning for Predicting Stock MovementsabstractStock movements are essentially driven by new information. Market data, financial news, and social sentiment are believed to have impacts on stock markets. To study the correlation between information and stock movements, previous works typically concatenate the features of different information sources into one super feature vector. However, such concatenated vector approaches treat each information source separately and ignore their interactions. In this article, we model the multi-faceted investors’ information and their intrinsic links with tensors. To identify the nonlinear patterns between stock movements and new information, we propose a supervised tensor regression learning approach to investigate the joint impact of different information sources on stock markets. Experiments on CSI 100 stocks in the year 2011 show that our approach outperforms the state-of-the-art trading strategies. Qing Li 0005, LiLing Jiang, Ping Li 0060, Hsinchun Chen |
AAAI | 4 |
| 2015 | Identifying Novel Adverse Drug Events from Health Social Media Using Distant Supervision
Xiao Liu 0016, Hsinchun Chen |
AMIA | 2 |
| 2015 | Developing understanding of hacker language through the use of lexical semanticsabstractThe need for more research scrutinizing online hacker communities is a common suggestion in recent years. However, researchers and practitioners face many challenges when attempting to do so. In particular, they may encounter hacking-specific terms, concepts, tools, and other items that are unfamiliar and may be challenging to understand. For these reasons, we are motivated to develop an automated method for developing understanding of hacker language. We utilize the latest advancements in recurrent neural network language models (RNNLMs) to develop an unsupervised machine learning technique for learning hacker language. The selected RNNLM produces state-of-the-art word embeddings that are useful for understanding the relations between different hacker terms and concepts. We evaluate our work by testing the RNNLMs ability to learn relevant relations between known hacker terms. Results suggest that the latest work in RNNLMs can aid in modeling hacker language, providing promising direction for future research. Victor A. Benjamin, Hsinchun Chen |
ISI | 2 |
| 2015 | Exploring threats and vulnerabilities in hacker web: Forums, IRC and carding shopsabstractCybersecurity is a problem of growing relevance that impacts all facets of society. As a result, many researchers have become interested in studying cybercriminals and online hacker communities in order to develop more effective cyber defenses. In particular, analysis of hacker community contents may reveal existing and emerging threats that pose great risk to individuals, businesses, and government. Thus, we are interested in developing an automated methodology for identifying tangible and verifiable evidence of potential threats within hacker forums, IRC channels, and carding shops. To identify threats, we couple machine learning methodology with information retrieval techniques. Our approach allows us to distill potential threats from the entirety of collected hacker contents. We present several examples of identified threats found through our analysis techniques. Results suggest that hacker communities can be analyzed to aid in cyber threat detection, thus providing promising direction for future work. Victor A. Benjamin, Weifeng Li 0002, Thomas Holt, Hsinchun Chen |
ISI | 4 |
| 2015 | Exploring hacker assets in underground forumsabstractMany large companies today face the risk of data breaches via malicious software, compromising their business. These types of attacks are usually executed using hacker assets. Researching hacker assets within underground communities can help identify the tools which may be used in a cyberattack, provide knowledge on how to implement and use such assets and assist in organizing tools in a manner conducive to ethical reuse and education. This study aims to understand the functions and characteristics of assets in hacker forums by applying classification and topic modeling techniques. This research contributes to hacker literature by gaining a deeper understanding of hacker assets in well-known forums and organizing them in a fashion conducive to educational reuse. Additionally, companies can apply our framework to forums of their choosing to extract their assets and appropriate functions. Sagar Samtani, Ryan Chinn, Hsinchun Chen |
ISI | 3 |
| 2015 | The roles of sharing, transfer, and public funding in nanotechnology knowledge-diffusion networksabstractUnderstanding the knowledge‐diffusion networks of patent inventors can help governments and businesses effectively use their investment to stimulate commercial science and technology development. Such inventor networks are usually large and complex. This study proposes a multidimensional network analysis framework that utilizes Exponential Random Graph Models (ERGMs) to simultaneously model knowledge‐sharing and knowledge‐transfer processes, examine their interactions, and evaluate the impacts of network structures and public funding on knowledge‐diffusion networks. Experiments are conducted on a longitudinal data set that covers 2 decades (1991–2010) of nanotechnology‐related US Patent and Trademark Office (USPTO) patents. The results show that knowledge sharing and knowledge transfer are closely interrelated. High degree centrality or boundary inventors play significant roles in the network, and National Science Foundation (NSF) public funding positively affects knowledge sharing despite its small fraction in overall funding and upstream research topics. Shan Jiang 0002, Hsinchun Chen, Mihail C. Roco |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2015 | A research framework for pharmacovigilance in health social media: Identification and evaluation of patient adverse drug event reports
Xiao Liu 0016, Hsinchun Chen |
J. Biomed. Informatics | 2 |
| 2014 | Analyzing market performance via social media: a case study of a banking industry crisis
Cuiqing Jiang, Hsinchun Chen, Yong Ding 0007 |
Sci. China Inf. Sci. | 3 |
| 2014 | An integrated framework for analyzing multilingual content in Web 2.0 social media
Yan Dang 0001, Gavin Yulei Zhang, Paul Jen-Hwa Hu, Susan A. Brown, Yungchang Ku, Jau-Hwang Wang, Hsinchun Chen |
Decis. Support Syst. | 7 |
| 2014 | Analyzing firm-specific social media and market: A stakeholder-based event analysis framework
Shan Jiang 0002, Hsinchun Chen, Jay F. Nunamaker Jr., David Zimbra |
Decis. Support Syst. | 2 |
| 2014 | Bridging the virtual and real: The relationship between web content, linkage, and geographical proximity of social movementsabstractAs the Internet becomes ubiquitous, it has advanced to more closely represent aspects of the real world. Due to this trend, researchers in various disciplines have become interested in studying relationships between real‐world phenomena and their virtual representations. One such area of emerging research seeks to study relationships between real‐world and virtual activism of social movement organization (SMOs). In particular, SMOs holding extreme social perspectives are often studied due to their tendency to have robust virtual presences to circumvent real‐world social barriers preventing information dissemination. However, many previous studies have been limited in scope because they utilize manual data‐collection and analysis methods. They also often have failed to consider the real‐world aspects of groups that partake in virtual activism. We utilize automated data‐collection and analysis methods to identify significant relationships between aspects of SMO virtual communities and their respective real‐world locations and ideological perspectives. Our results also demonstrate that the interconnectedness of SMO virtual communities is affected specifically by aspects of the real world. These observations provide insight into the behaviors of SMOs within virtual environments, suggesting that the virtual communities of SMOs are strongly affected by aspects of the real world. Victor A. Benjamin, Hsinchun Chen, David Zimbra |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2014 | Text mining self-disclosing health information for public health serviceabstractUnderstanding specific patterns or knowledge of self‐disclosing health information could support public health surveillance and healthcare. This study aimed to develop an analytical framework to identify self‐disclosing health information with unusual messages on web forums by leveraging advanced text‐mining techniques. To demonstrate the performance of the proposed analytical framework, we conducted an experimental study on 2 major human immunodeficiency virus (HIV)/acquired immune deficiency syndrome (AIDS) forums in Taiwan. The experimental results show that the classification accuracy increased significantly (up to 83.83%) when using features selected by the information gain technique. The results also show the importance of adopting domain‐specific features in analyzing unusual messages on web forums. This study has practical implications for the prevention and support of HIV/AIDS healthcare. For example, public health agencies can re‐allocate resources and deliver services to people who need help via social media sites. In addition, individuals can also join a social media site to get better suggestions and support from each other. Yungchang Ku, Chaochang Chiu, Gavin Yulei Zhang, Hsinchun Chen, Handsome Su |
J. Assoc. Inf. Sci. Technol. | 4 |
| 2013 | Machine learning for attack vector identification in malicious source codeabstractAs computers and information technologies become ubiquitous throughout society, the security of our networks and information technologies is a growing concern. As a result, many researchers have become interested in the security domain. Among them, there is growing interest in observing hacker communities for early detection of developing security threats and trends. Research in this area has often reported hackers openly sharing cybercriminal assets and knowledge with one another. In particular, the sharing of raw malware source code files has been documented in past work. Unfortunately, malware code documentation appears often times to be missing, incomplete, or written in a language foreign to researchers. Thus, analysis of such source files embedded within hacker communities has been limited. Here we utilize a subset of popular machine learning methodologies for the automated analysis of malware source code files. Specifically, we explore genetic algorithms to resolve questions related to feature selection within the context of malware analysis. Next, we utilize two common classification algorithms to test selected features for identification of malware attack vectors. Results suggest promising direction in utilizing such techniques to help with the automated analysis of malware source code. Victor A. Benjamin, Hsinchun Chen |
ISI | 2 |
| 2013 | Evaluating text visualization: An experiment in authorship analysisabstractAnalyzing authorship of online texts is an important analysis task in security-related areas such as cybercrime investigation and counter-terrorism, and in any field of endeavor in which authorship may be uncertain or obfuscated. This paper presents an automated approach for authorship analysis using machine learning methods, a robust stylometric feature set, and a series of visualizations designed to facilitate analysis at the feature, author, and message levels. A testbed consisting of 506,554 forum messages, in English and Arabic, from 14,901 authors was first constructed. A prototype portal system was then developed to support feasibility analysis of the approach. A preliminary evaluation to assess the efficacy of the text visualizations was conducted. The evaluation showed that task performance with the visualization functions was more accurate and more efficient than task performance without the visualizations. Victor A. Benjamin, Wingyan Chung, Ahmed Abbasi, Joshua Chuang, Catherine A. Larson, Hsinchun Chen |
ISI | 6 |
| 2013 | Recommendation as link prediction in bipartite graphs: A graph kernel-based machine learning approach
Xin Li 0004, Hsinchun Chen |
Decis. Support Syst. | 2 |
| 2013 | Research note: Examining gender emotional differences in Web forum communication
Gavin Yulei Zhang, Yan Dang 0001, Hsinchun Chen |
Decis. Support Syst. | 3 |
| 2013 | MedTime: A temporal information extraction system for clinical narratives
Yu-Kai Lin, Hsinchun Chen, Randall A. Brown |
J. Biomed. Informatics | 2 |
| 2012 | SHB 2012: international workshop on smart health and wellbeingabstractThe Smart Health and Wellbeing workshop is organized to develop a platform for authors to discuss fundamental principles, algorithms or applications of intelligent data acquisition, processing and analysis of healthcare data. We are particularly interested in information and knowledge management papers, in which the approaches are accompanied by an in-depth experimental evaluation with real world data. This paper provides an overview of the workshop and the accepted contributions. Christopher C. Yang, Hsinchun Chen, Howard D. Wactlar, Carlo Combi, Xuning Tang |
CIKM | 2 |
| 2012 | Dark Web: Exploring and Mining the Dark Side of the Web
Hsinchun Chen |
ICFCA | 1 |
| 2012 | Securing cyberspace: Identifying key actors in hacker communitiesabstractAs the computer becomes more ubiquitous throughout society, the security of networks and information technologies is a growing concern. Recent research has found hackers making use of social media platforms to form communities where sharing of knowledge and tools that enable cybercriminal activity is common. However, past studies often report only generalized community behaviors and do not scrutinize individual members; in particular, current research has yet to explore the mechanisms in which some hackers become key actors within their communities. Here we explore two major hacker communities from the United States and China in order to identify potential cues for determining key actors. The relationships between various hacker posting behaviors and reputation are observed through the use of ordinary least squares regression. Results suggest that the hackers who contribute to the cognitive advance of their community are generally considered the most reputable and trustworthy among their peers. Conversely, the tenure of hackers and their discussion quality were not significantly correlated with reputation. Results are consistent across both forums, indicating the presence of a common hacker culture that spans multiple geopolitical regions. Victor A. Benjamin, Hsinchun Chen |
ISI | 2 |
| 2012 | Evaluating an integrated forum portal for terrorist surveillance and analysisabstractWe experimentally evaluated the Dark Web Forum Portal by focusing on user task performance, usability, cognitive processing requirements, and societal benefits. Our results show that the portal performs perform well when compared with a benchmark forum. Paul Jen-Hwa Hu, Xing Wan, Yan Dang 0001, Catherine A. Larson, Hsinchun Chen |
ISI | 5 |
| 2012 | Using burst detection techniques to identify suspicious vehicular traffic at border crossingsabstractBorder safety is a critical part of national and international security. The Department of Homeland Security (DHS) searches vehicles entering the country at land borders for drugs and other contraband. However, this process is time-consuming and operational efficiency is needed for smooth operations at the border. To aid in the screening of vehicles, we propose to examine traffic patterns at checkpoints using burst detection algorithms. Our results show that the overall traffic at the border shows bursting patterns attributable to week days and the holiday seasons. In addition, using local law-enforcement data we also find that traffic with prior contacts with law-enforcement shows a bursting pattern distinct from other traffic. We also find that such bursts in suspicious traffic can be attributable to increases in vehicular traffic associated with certain kinds of criminal activity. This information can be used to specifically target vehicles searches during primary screening at ports and in the surrounding areas. Siddharth Kaza, Hsin-Min Lu, Daniel Dajun Zeng, Hsinchun Chen |
ISI | 4 |
| 2012 | An event-driven SIR model for topic diffusion in web forumsabstractSocial media is being increasingly used as a communication channel. Among social media, web forums, where people in online communities disseminate and receive information by interaction, provide a good environment to examine information diffusion. In this research, we aim to understand the mechanisms and properties of the information diffusion in the web forum. For that, we model topic-level information diffusion in web forums using the baseline epidemic model, the SIR(Susceptible, Infective, and Recovered) model, frequently used in previous research to analyze disease outbreaks and knowledge diffusion. In addition, we propose an event-driven SIR model that reflects the event effect on information diffusion in the web forum. The proposed model incorporates the effect of news postings on the web forum. We evaluate two models using a large longitudinal dataset from the web forum of a major company. The event-SIR model outperforms the SIR model in fitting on major spikey topics that have peaks of author participation. Hsinchun Chen |
ISI | 2 |
| 2012 | Partially supervised learning for radical opinion identification in hate group web forumsabstractWeb forums are frequently used as platforms for the exchange of information and opinions, as well as propaganda dissemination. But online content can be misused when the information being distributed, such as radical opinions, is unsolicited or inappropriate. However, radical opinion is highly hidden and distributed in Web forums, while non-radical content is unspecific and topically more diverse. It is costly and time consuming to label a large amount of radical content (positive examples) and non-radical content (negative examples) for training classification systems. Nevertheless, it is easy to obtain large volumes of unlabeled content in Web forums. In this paper, we propose and develop a topic-sensitive partially supervised learning approach to address the difficulties in radical opinion identification in hate group Web forums. Specifically, we design a labeling heuristic to extract high quality positive examples and negative examples from unlabeled datasets. The empirical evaluation results from two large hate group Web forums suggest that our proposed approach generally outperforms the benchmark techniques and exhibits more stable performance than its counterparts. Ming Yang 0037, Hsinchun Chen |
ISI | 2 |
| 2012 | Scalable sentiment classification across multiple Dark Web ForumsabstractThis study examines several approaches to sentiment classification in the Dark Web Forum Portal, and opportunities to transfer classifiers and text features across multiple forums to improve scalability and performance. Although sentiment classifiers typically perform poorly when transferred across domains, experimentation reveals the devised approaches offer performance equivalent to the traditional forum-specific approach in classification in an unknown domain. Furthermore, incorporating the text features identified as significant indicators of sentiment in other forums can greatly improve the classification accuracy of the traditional forum-specific approach. David Zimbra, Hsinchun Chen |
ISI | 2 |
| 2012 | Evaluating sentiment in financial news articles
Robert P. Schumaker, Gavin Yulei Zhang, Chunneng Huang, Hsinchun Chen |
Decis. Support Syst. | 4 |
| 2012 | Artificial immune system for illicit content identification in social mediaabstractAbstract Social media is frequently used as a platform for the exchange of information and opinions as well as propaganda dissemination. But online content can be misused for the distribution of illicit information, such as violent postings in web forums. Illicit content is highly distributed in social media, while non‐illicit content is unspecific and topically diverse. It is costly and time consuming to label a large amount of illicit content (positive examples) and non‐illicit content (negative examples) to train classification systems. Nevertheless, it is relatively easy to obtain large volumes of unlabeled content in social media. In this article, an artificial immune system‐based technique is presented to address the difficulties in the illicit content identification in social media. Inspired by the positive selection principle in the immune system, we designed a novel labeling heuristic based on partially supervised learning to extract high‐quality positive and negative examples from unlabeled datasets. The empirical evaluation results from two large hate group web forums suggest that our proposed approach generally outperforms the benchmark techniques and exhibits more stable performance. Ming Yang 0037, Melody Y. Kiang, Hsinchun Chen, Yijun Li 0004 |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2012 | Sentimental Spidering: Leveraging Opinion Information in Focused CrawlersabstractDespite the increased prevalence of sentiment-related information on the Web, there has been limited work on focused crawlers capable of effectively collecting not only topic-relevant but also sentiment-relevant content. In this article, we propose a novel focused crawler that incorporates topic and sentiment information as well as a graph-based tunneling mechanism for enhanced collection of opinion-rich Web content regarding a particular topic. The graph-based sentiment (GBS) crawler uses a text classifier that employs both topic and sentiment categorization modules to assess the relevance of candidate pages. This information is also used to label nodes in web graphs that are employed by the tunneling mechanism to improve collection recall. Experimental results on two test beds revealed that GBS was able to provide better precision and recall than seven comparison crawlers. Moreover, GBS was able to collect a large proportion of the relevant content after traversing far fewer pages than comparison methods. GBS outperformed comparison methods on various categories of Web pages in the test beds, including collection of blogs, Web forums, and social networking Web site content. Further analysis revealed that both the sentiment classification module and graph-based tunneling mechanism played an integral role in the overall effectiveness of the GBS crawler. Tianjun Fu, Ahmed Abbasi, Daniel Dajun Zeng, Hsinchun Chen |
ACM Trans. Inf. Syst. | 4 |
| 2011 | The Dark Web Forum Portal: From multi-lingual to videoabstractCounter-terrorism, intelligence analysts, and other investigators continue to analyze the Internet presence of terrorists, hate groups, and other extremists through the study of primary sources including terrorists' own websites, videos, chat sites, and Internet forums. Forums and videos are both particularly rich sources of information. Forums - discussion sites supporting online conversations - capture each conversation in a “thread” and the ensuing postings are usually time-stamped and attributable to a particular online poster (author). With careful analysis, they can reveal trends in topics and discussions, the sequencing of ideas, and the relationships between posters. Videos gain a global audience when posted to YouTube, but identifying and finding videos relating to a specific interest or topic can be difficult among the tens of millions of available items. The Dark Web Forum Portal was originally constructed to allow the examination, from a broad perspective, of the use of Web forums by terrorist and extremist groups. The Video Portal module has been added to facilitate the study of video as it is used by these groups. Both portals are available to researchers on a request basis. In this paper, we examine the evolution of the Dark Web Forum Portal's system design, share the results of a user evaluation, and provide an overview of the development of the new video portal. Hsinchun Chen, Dorothy E. Denning, Nancy Roberts, Catherine A. Larson, Ximing Yu, Chunneng Huang |
ISI | 1 |
| 2011 | The Geopolitical Web: Assessing societal risk in an uncertain worldabstractCountry risk - the likelihood that a state will weaken or fail - and the methods of assessing it continue to be of serious concern to the international community. Country risk has traditionally been assessed by monitoring economic and financial indicators. However, social media (such as forums, blogs, and websites) are now important transporters of citizens' daily conversations and opinions, and as such may carry discernible indicators of risk, but they have been as yet little-used for this task. The Geopolitical Web project is a research effort with the ultimate goal of developing computational approaches for monitoring public opinion in regions of conflict, assessing country risk indicators in the social media of fragile or weakening states, and correlating these risk signals with commonly accepted quantitative geopolitical risk assessments. This paper presents the initial motivation for this data-driven project, collection procedures adopted, preliminary results of an automated topical analysis of the collection's content, and expected future work. By catching and deciphering possible signals of country risk in social discourse we hope to offer the international community an additional means of assessing the need for intervention in or support for fragile or weakening states. Hsinchun Chen, Catherine A. Larson, Theodore Elhourani, David Zimbra, David Ware |
ISI | 1 |
| 2011 | An SIR model for violent topic diffusion in social mediaabstractSocial media is being increasingly used as a political communication channel. The web makes it easy to spread extreme opinions or ideologies that were once restricted to small groups. Terrorists and extremists use the web to deliver their extreme ideology to people and encourage them to get involved in fanatic behaviors. In this research, we aim to understand the mechanisms and properties of the exposure process to extreme opinions through these new publication methods, especially web forums. We propose the topic diffusion model for web forums, based on the SIR (Susceptible, Infective, and Recovered) model frequently used in previous research to analyze disease outbreaks and knowledge diffusion. The logistic growth of possible authors, the interaction between possible authors and current authors, and the influence decay of past authors are incorporated in a novel topic-based SIR model. From the proposed model we can estimate the maximum number of authors on a topic, the degree of infectiousness of a topic, and the rate describing how fast past authors lose influence over others. We apply the proposed model to a major international Jihadi forum where extreme ideology is expounded and evaluate the model on the diffusion of major violent topics. The fitting results show that it is plausible to describe the mechanism of violent topic diffusion in web forums with the SIR epidemic model. Jaebong Son, Hsinchun Chen |
ISI | 3 |
| 2011 | Dynamic user-level affect analysis in social media: Modeling violence in the Dark WebabstractAffect represents a person's emotions toward objects, issues or other persons. Recent years have witnessed a surge in studies of users' affect in social media, as marketing literature has shown that users' affect influences decision making. The current literature in this area, however, has largely focused on the message level, using text-based features and various classification approaches. Such analyses not only overlook valuable information about the user who posts the messages, but also fail to consider that users' affect may change over time. To overcome these limitations, we propose a new research design for social media affect analysis by specifically incorporating users' characteristics and the time dimension. We illustrate our research design by applying it to a major Dark Web forum of international Jihadists. Empirical results show that our research design allows us to draw on theories from other disciplines, such as social psychology, to provide useful insights on the dynamic change of users' affect in social media. Shuo Zeng, Mingfeng Lin, Hsinchun Chen |
ISI | 3 |
| 2011 | Enterprise risk and security management: Data, text and Web mining
Hsinchun Chen, Michael Chau, Shu-Hsing Li |
Decis. Support Syst. | 1 |
| 2011 | Giving context to accounting numbers: The role of news coverage
Kuo-Tay Chen, Hsin-Min Lu, Tsai-Jyh Chen, Shu-Hsing Li, Jian-Shuen Lian, Hsinchun Chen |
Decis. Support Syst. | 6 |
| 2011 | Knowledge mapping for rapidly evolving domains: A design science approach
Yan Dang 0001, Gavin Yulei Zhang, Paul Jen-Hwa Hu, Susan A. Brown, Hsinchun Chen |
Decis. Support Syst. | 5 |
| 2011 | A hierarchical Naïve Bayes model for approximate identity matching
G. Alan Wang, Homa Atabakhsh, Hsinchun Chen |
Decis. Support Syst. | 3 |
| 2011 | Determining inventor status and its effect on knowledge diffusion: A study on nanotechnology literature from China, Russia, and IndiaabstractIn an increasingly global research landscape, it is important to identify the most prolific researchers in various institutions and their influence on the diffusion of knowledge. Knowledge diffusion within institutions is influenced by not just the status of individual researchers but also the collaborative culture that determines status. There are various methods to measure individual status, but few studies have compared them or explored the possible effects of different cultures on the status measures. In this article, we examine knowledge diffusion within science and technology-oriented research organizations. Using social network analysis metrics to measure individual status in large-scale coauthorship networks, we studied an individual's impact on the recombination of knowledge to produce innovation in nanotechnology. Data from the most productive and high-impact institutions in China (Chinese Academy of Sciences), Russia (Russian Academy of Sciences), and India (Indian Institutes of Technology) were used. We found that boundary-spanning individuals influenced knowledge diffusion in all countries. However, our results also indicate that cultural and institutional differences may influence knowledge diffusion. Xuan Liu 0004, Siddharth Kaza, Pengzhu Zhang, Hsinchun Chen |
J. Assoc. Inf. Sci. Technol. | 4 |
| 2011 | Disease named entity recognition using semisupervised learning and conditional random fieldsabstractAbstract Information extraction is an important text‐mining task that aims at extracting prespecified types of information from large text collections and making them available in structured representations such as databases. In the biomedical domain, information extraction can be applied to help biologists make the most use of their digital‐literature archives. Currently, there are large amounts of biomedical literature that contain rich information about biomedical substances. Extracting such knowledge requires a good named entity recognition technique. In this article, we combine conditional random fields (CRFs), a state‐of‐the‐art sequence‐labeling algorithm, with two semisupervised learning techniques, bootstrapping and feature sampling, to recognize disease names from biomedical literature. Two data‐processing strategies for each technique also were analyzed: one sequentially processing unlabeled data partitions and another one processing unlabeled data partitions in a round‐robin fashion. The experimental results showed the advantage of semisupervised learning techniques given limited labeled training data. Specifically, CRFs with bootstrapping implemented in sequential fashion outperformed strictly supervised CRFs for disease name recognition. The project was supported by NIH/NLM Grant R33 LM07299–01, 2002–2005. Nichalin S. Summerfield, Hsinchun Chen |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2011 | Selecting Attributes for Sentiment Classification Using Feature Relation NetworksabstractA major concern when incorporating large sets of diverse n-gram features for sentiment classification is the presence of noisy, irrelevant, and redundant attributes. These concerns can often make it difficult to harness the augmented discriminatory potential of extended feature sets. We propose a rule-based multivariate text feature selection method called Feature Relation Network (FRN) that considers semantic information and also leverages the syntactic relationships between n-gram features. FRN is intended to efficiently enable the inclusion of extended sets of heterogeneous n-gram features for enhanced sentiment classification. Experiments were conducted on three online review testbeds in comparison with methods used in prior sentiment classification research. FRN outperformed the comparison univariate, multivariate, and hybrid feature selection methods; it was able to select attributes resulting in significantly better classification accuracy irrespective of the feature subset sizes. Furthermore, by incorporating syntactic information about n-gram relations, FRN is able to select features in a more computationally efficient manner than many multivariate and hybrid techniques. Ahmed Abbasi, Stephen L. France, Hsinchun Chen |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2011 | Special Issue on Social Media Analytics: Understanding the Pulse of the SocietyabstractThe four papers in this special issue focus on the advanced modeling and simulation, human organizational interactions, web spidering, digital archiving, cyber archeology, social network analysis, sentiment analysis, and data/text/web mining techniques and methodologies, which can contribute to the understanding of the pulse of the society. Hsinchun Chen, Christopher C. Yang |
IEEE Trans. Syst. Man Cybern. Part A | 1 |
| 2011 | Gender Classification for Web ForumsabstractMore and more women are participating in and exchanging opinions through community-based online social media. Questions concerning gender differences in the new media have been raised. This paper proposes a feature-based text classification framework to examine online gender differences between Web forum posters by analyzing writing styles and topics of interest. Our experiment on an Islamic women's political forum shows that feature sets containing both content-free and content-specific features perform significantly better than those consisting of only content-free features, feature selection can improve the classification results significantly, and female and male participants have significantly different topics of interest. Gavin Yulei Zhang, Yan Dang 0001, Hsinchun Chen |
IEEE Trans. Syst. Man Cybern. Part A | 3 |
| 2010 | Global disease surveillance using social media: HIV/AIDS content intervention in web forumsabstractCollecting potential data sources for use in proactively analyzing and evaluating strategies for syndromic surveillance and bio-defense have become critical issues in and challenges to infectious disease informatics [1]. Web forums link highly relevant information about patients' needs, disease pain, health conditions, and concerns for medical practice [2]. Health departments or medical service groups can use the information in the forums to identify disease sources and scope and to detect outbreaks while the possibility for intervention remains. Yungchang Ku, Chaochang Chiu, Gavin Yulei Zhang, Hsinchun Chen |
ISI | 5 |
| 2010 | Developing a Dark Web collection and infrastructure for computational and social sciencesabstractIn recent years, there have been numerous studies from a variety of perspectives analyzing the Internet presence of hate and extremist groups. Yet the websites and forums of extremist and terrorist groups have long remained an underutilized resource for terrorism researchers due to their ephemeral nature and access and analysis problems. The purpose of the Dark Web archive is to provide a research infrastructure for use by social scientists, computer and information scientists, policy and security analysts, and others studying a wide range of social and organizational phenomena and computational problems. The Dark Web Forum Portal provides web enabled access to critical international jihadist and other extremist web forums. The focus of this paper is on the significant extensions to previous work including: increasing the scope of data collection, adding an incremental spidering component for regular data updates; enhancing the searching and browsing functions; enhancing multilingual machine-translation for Arabic, French, German and Russian; and advanced Social Network Analysis. A case study on identifying active participants is shown at the end. Gavin Yulei Zhang, Shuo Zeng, Chunneng Huang, Ximing Yu, Yan Dang 0001, Catherine A. Larson, Dorothy E. Denning, Nancy Roberts, Hsinchun Chen |
ISI | 10 |
| 2010 | Comparing the virtual linkage intensity and real world proximity of social movementsabstractThe relationships between phenomena observed in the real world and their representations in virtual contexts have generated interest among researchers. In particular, the manifestations of social movements in virtual environments have been examined, with many studies dedicated to the analysis of the virtual linkages between groups. In this research, a form of link analysis was performed to examine the relationship between virtual linkage intensity and real world physical proximity among the social movement groups identified in the Southern Poverty Law Center Spring 2009 Intelligence Report. Findings indicate the existence of significant relationships between virtual linkage intensity and physical proximity, distinctive to various ideological categorizations. The results provide valuable insights into the behaviors of social movements in virtual environments. David Zimbra, Hsinchun Chen |
ISI | 2 |
| 2010 | Visualizing social network concepts
Bin Zhu 0001, Stephanie Watts, Hsinchun Chen |
Decis. Support Syst. | 3 |
| 2010 | Evaluating the use of search engine development tools in IT educationabstractAbstract It is important for education in computer science and information systems to keep up to date with the latest development in technology. With the rapid development of the Internet and the Web, many schools have included Internet‐related technologies, such as Web search engines and e‐commerce, as part of their curricula. Previous research has shown that it is effective to use search engine development tools to facilitate students' learning. However, the effectiveness of these tools in the classroom has not been evaluated. In this article, we review the design of three search engine development tools, SpidersRUs, Greenstone, and Alkaline, followed by an evaluation study that compared the three tools in the classroom. In the study, 33 students were divided into 13 groups and each group used the three tools to develop three independent search engines in a class project. Our evaluation results showed that SpidersRUs performed better than the two other tools in overall satisfaction and the level of knowledge gained in their learning experience when using the tools for a class project on Internet applications development. Michael Chau, Cho Hung Wong, Yilu Zhou, Jialun Qin, Hsinchun Chen |
J. Assoc. Inf. Sci. Technol. | 5 |
| 2010 | A focused crawler for Dark Web forumsabstractAbstract The unprecedented growth of the Internet has given rise to the Dark Web, the problematic facet of the Web associated with cybercrime, hate, and extremism. Despite the need for tools to collect and analyze Dark Web forums, the covert nature of this part of the Internet makes traditional Web crawling techniques insufficient for capturing such content. In this study, we propose a novel crawling system designed to collect Dark Web forum content. The system uses a human‐assisted accessibility approach to gain access to Dark Web forums. Several URL ordering features and techniques enable efficient extraction of forum postings. The system also includes an incremental crawler coupled with a recall‐improvement mechanism intended to facilitate enhanced retrieval and updating of collected content. Experiments conducted to evaluate the effectiveness of the human‐assisted accessibility approach and the recall‐improvement‐based, incremental‐update procedure yielded favorable results. The human‐assisted approach significantly improved access to Dark Web forums while the incremental crawler with recall improvement also outperformed standard periodic‐ and incremental‐update approaches. Using the system, we were able to collect over 100 Dark Web forums from three regions. A case study encompassing link and content analysis of collected forums was used to illustrate the value and importance of gathering and analyzing content from such online communities. Tianjun Fu, Ahmed Abbasi, Hsinchun Chen |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2010 | Agency satisfaction with electronic record management systems: A large-scale surveyabstractAbstract We investigated agency satisfaction with an electronic record management system (ERMS) that supports the electronic creation, archival, processing, transmittal, and sharing of records (documents) among autonomous government agencies. A factor model, explaining agency satisfaction with ERMS functionalities, offers hypotheses, which we tested empirically with a large‐scale survey that involved more than 1,600 government agencies in Taiwan. The data showed a good fit to our model and supported all the hypotheses. Overall, agency satisfaction with ERMS functionalities appears jointly determined by regulatory compliance, job relevance, and satisfaction with support services. Among the determinants we studied, agency satisfaction with support services seems the strongest predictor of agency satisfaction with ERMS functionalities. Regulatory compliance also has important influences on agency satisfaction with ERMS, through its influence on job relevance and satisfaction with support services. Further analyses showed that satisfaction with support services partially mediated the impact of regulatory compliance on satisfaction with ERMS functionalities, and job relevance partially mediated the influence of regulatory compliance on satisfaction with ERMS functionalities. Our findings have important implications for research and practice, which we also discuss. Paul Jen-Hwa Hu, Fang-Ming Hsu, Han-fen Hu, Hsinchun Chen |
J. Assoc. Inf. Sci. Technol. | 4 |
| 2010 | Text-based video content classification for online video-sharing sitesabstractAbstract With the emergence of Web 2.0, sharing personal content, communicating ideas, and interacting with other online users in Web 2.0 communities have become daily routines for online users. User‐generated data from Web 2.0 sites provide rich personal information (e.g., personal preferences and interests) and can be utilized to obtain insight about cyber communities and their social networks. Many studies have focused on leveraging user‐generated information to analyze blogs and forums, but few studies have applied this approach to video‐sharing Web sites. In this study, we propose a text‐based framework for video content classification of online‐video sharing Web sites. Different types of user‐generated data (e.g., titles, descriptions, and comments) were used as proxies for online videos, and three types of text features (lexical, syntactic, and content‐specific features) were extracted. Three feature‐based classification techniques (C4.5, Naïve Bayes, and Support Vector Machine) were used to classify videos. To evaluate the proposed framework, user‐generated data from candidate videos, which were identified by searching user‐given keywords on YouTube, were first collected. Then, a subset of the collected data was randomly selected and manually tagged by users as our experiment data. The experimental results showed that the proposed approach was able to classify online videos based on users' interests with accuracy rates up to 87.2%, and all three types of text features contributed to discriminating videos. Support Vector Machine outperformed C4.5 and Naïve Bayes techniques in our experiments. In addition, our case study further demonstrated that accurate video‐classification results are very useful for identifying implicit cyber communities on video‐sharing Web sites. Chunneng Huang, Tianjun Fu, Hsinchun Chen |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2010 | Gene function prediction with gene interaction networks: a context graph kernel approachabstractPredicting gene functions is a challenge for biologists in the postgenomic era. Interactions among genes and their products compose networks that can be used to infer gene functions. Most previous studies adopt a linkage assumption, i.e., they assume that gene interactions indicate functional similarities between connected genes. In this study, we propose to use a gene's context graph, i.e., the gene interaction network associated with the focal gene, to infer its functions. In a kernel-based machine-learning framework, we design a context graph kernel to capture the information in context graphs. Our experimental study on a testbed of p53-related genes demonstrates the advantage of using indirect gene interactions and shows the empirical superiority of the proposed approach over linkage-assumption-based methods, such as the algorithm to minimize inconsistent connected genes and diffusion kernels. Xin Li 0004, Hsinchun Chen, Jiexun Li |
IEEE Trans. Inf. Technol. Biomed. | 2 |
| 2010 | Prospective Infectious Disease Outbreak Detection Using Markov Switching ModelsabstractAccurate and timely detection of infectious disease outbreaks provides valuable information which can enable public health officials to respond to major public health threats in a timely fashion. However, disease outbreaks are often not directly observable. For surveillance systems used to detect outbreaks, noises caused by routine behavioral patterns and by special events can further complicate the detection task. Most existing detection methods combine a time series filtering procedure followed by a statistical surveillance method. The performance of this "two-step” detection method is hampered by the unrealistic assumption that the training data are outbreak-free. Moreover, existing approaches are sensitive to extreme values, which are common in real-world data sets. We considered the problem of identifying outbreak patterns in a syndrome count time series using Markov switching models. The disease outbreak states are modeled as hidden state variables which control the observed time series. A jump component is introduced to absorb sporadic extreme values that may otherwise weaken the ability to detect slow-moving disease outbreaks. Our approach outperformed several state-of-the-art detection methods in terms of detection sensitivity using both simulated and real-world data. Hsin-Min Lu, Daniel Dajun Zeng, Hsinchun Chen |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2010 | Interaction Analysis of the ALICE Chatterbot: A Two-Study Investigation of Dialog and Domain QuestioningabstractThis paper analyzes and compares the data gathered from two previously conducted artificial linguistic Internet chat entity (ALICE) chatterbot studies that were focused on response accuracy and user satisfaction measures for six chatterbots. These chatterbots were further loaded with varying degrees of conversational, telecommunications, and terrorism knowledge. From our prior experiments using 347 participants, we obtained 33 446 human/chatterbot interactions. It was found that asking the ALICE chatterbots ¿are¿ and ¿where¿ questions resulted in higher response satisfaction levels, as compared to other interrogative-style inputs because of their acceptability to vague, binary, or clichE¿d chatterbot responses. We also found a relationship between the length of a query and the users perceived satisfaction of the chatterbot response, where shorter queries led to more satisfying responses. Robert P. Schumaker, Hsinchun Chen |
IEEE Trans. Syst. Man Cybern. Part A | 2 |
| 2010 | Burst Detection From Multiple Data Streams: A Network-Based ApproachabstractModeling and detecting bursts in data streams is an important area of research with a wide range of applications. In this paper, we present a novel method to analyze and identify correlated burst patterns by considering multiple data streams that coevolve over time. The main technical contribution of our research is the use of a dynamic probabilistic network to model the dependency structures observed within these data streams. Such dependencies provide meaningful information concerning the overall system dynamics and should be explicitly integrated into the burst detection process. Using both synthetic scenarios and two real-world datasets, we compare our method with an existing burst-detection algorithm. Initial experimental results indicate that our approach allows for more balanced and accurate burst quantification. Aaron Sun, Daniel Dajun Zeng, Hsinchun Chen |
IEEE Trans. Syst. Man Cybern. Part C | 3 |
| 2009 | Law enforcement officers' acceptance of advanced e-government technology: a survey study of COPLNK mobileabstractTimely information access and effective knowledge support is crucial to law enforcement officers' crime fighting and investigations. An expanding array of e-government initiatives target the development of advanced information technologies and their deployment in law enforcement agencies. Abase in point is COPLINK, an integrated system that provides law enforcement officers with timely data access, effective information support, integrated knowledge sharing, and improved collaboration within or beyond the agency boundaries. In this study, we examine law enforcement officers' acceptance of COPLONK Mobile by proposing and testing a factor model premised in established theoretical foundations. According to our results, the model is capable of explaining or predicting officers' intentions to use the technology. Our survey data support the proposed model and the hypotheses it suggests. Among the acceptance determinants we investigated, perceived usefulness appears to have the most significant influence on individual officers' intention to use COPLONK Mobile. Paul Jen-Hwa Hu, Hsinchun Chen, Han-fen Hu |
ICEC | 2 |
| 2009 | IEDs in the dark web: Lexicon expansion and genre classificationabstractImprovised explosive device web pages represent a significant source of knowledge for security organizations. In this paper, we present significant improvements to our approach to the discovery and classification of IED related web pages in the Dark Web. We present a statistical feature ranking approach to the expansion of the keyword lexicon used to discover IED related web pages, which identified new relevant terms for inclusion. Additionally, we present an improved web page feature representation designed to better capture the structural and stylistic cues revealing of genres of communication, and a series of experiments comparing the classification performance of the new representation with our existing approach. Hsinchun Chen |
ISI | 1 |
| 2009 | Identification of extremist videos in online video sharing sitesabstractWeb 2.0 has become an effective grassroots communication platform for extremists to promote their ideas, share resources, and communicate among each other. As an important component of Web 2.0, online video sharing sites such as YouTube and Google video have also been utilized by extremist groups to distribute videos. This study presented a framework for identifying extremist videos in online video sharing sites by using user-generated text content such as comments, video descriptions, and titles without downloading the videos. Text features including lexical features, syntactic features and content specific features were first extracted. Then Information Gain was used for feature selection, and Support Vector Machine was deployed for classification. The exploratory experiment showed that our proposed framework is effective for identifying online extremist videos, with the F-measure as high as 82%. Tianjun Fu, Chunneng Huang, Hsinchun Chen |
ISI | 3 |
| 2009 | Gender difference analysis of political web forums: An experiment on an international islamic women's forumabstractAs an important type of social media, the political Web forum has become a major communication channel for people to discuss and debate political, cultural and social issues. Although the Internet has a male-dominated history, more and more women have started to share their concerns and express opinions through online discussion boards and Web forums. This paper presents an automated approach to gender difference analysis of political Web forums. The approach uses rich textual feature representation and machine learning techniques to examine the online gender differences between female and male participants on political Web forums by analyzing writing styles and topics of interest. The results of gender difference analysis performed on a large and long-standing international Islamic women's political forum are presented, showing that female and male participants have significantly different topics of interest. Gavin Yulei Zhang, Yan Dang 0001, Hsinchun Chen |
ISI | 3 |
| 2009 | Dark web forums portal: Searching and analyzing jihadist forumsabstractWith the advent of Web 2.0, the Web is acting as a platform which enables end-user content generation. As a major type of social media in Web 2.0, Web forums facilitate intensive interactions among participants. International Jihadist groups often use Web forums to promote violence and distribute propaganda materials. These Dark Web forums are heterogeneous and widely distributed. Therefore, how to access and analyze the forum messages and interactions among participants is becoming an issue. This paper presents a general framework for Web forum data integration. Specifically, a Web-based knowledge portal, the Dark Web Forums Portal, is built based on the framework. The portal incorporates the data collected from different international Jihadist forums and provides several important analysis functions, including forum browsing and searching (in single forum and across multiple forums), forum statistics analysis, multilingual translation, and social network visualization. Preliminary results of our user study show that the Dark Web Forums Portal helps users locate information quickly and effectively. Users found the forum statistics analysis, multilingual translation, and social network visualization functions of the portal to be particularly valuable. Gavin Yulei Zhang, Shuo Zeng, Yan Dang 0001, Catherine A. Larson, Hsinchun Chen |
ISI | 6 |
| 2009 | Automatic online news monitoring and classification for syndromic surveillance
Gavin Yulei Zhang, Yan Dang 0001, Hsinchun Chen, Mark Thurmond, Cathy Larson |
Decis. Support Syst. | 3 |
| 2009 | A quantitative stock prediction system based on financial news
Robert P. Schumaker, Hsinchun Chen |
Inf. Process. Manag. | 2 |
| 2009 | Browsing the underdeveloped Web: An experiment on the Arabic Medical Web DirectoryabstractAbstract While the Web has grown significantly in recent years, some portions of the Web remain largely underdeveloped, as shown in a lack of high‐quality content and functionality. An example is the Arabic Web, in which a lack of well‐structured Web directories limits users' ability to browse for Arabic resources. In this research, we proposed an approach to building Web directories for the underdeveloped Web and developed a proof‐of‐concept prototype called the Arabic Medical Web Directory (AMedDir) that supports browsing of over 5,000 Arabic medical Web sites and pages organized in a hierarchical structure. We conducted an experiment involving Arab participants and found that the AMedDir significantly outperformed two benchmark Arabic Web directories in terms of browsing effectiveness, efficiency, information quality, and user satisfaction. Participants expressed strong preference for the AMedDir and provided many positive comments. This research thus contributes to developing a useful Web directory for organizing the information in the Arabic medical domain and to a better understanding of how to support browsing on the underdeveloped Web. Wingyan Chung, Hsinchun Chen |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2009 | Business stakeholder analyzer: An experiment of classifying stakeholders on the WebabstractAbstract As the Web is used increasingly to share and disseminate information, business analysts and managers are challenged to understand stakeholder relationships. Traditional stakeholder theories and frameworks employ a manual approach to analysis and do not scale up to accommodate the rapid growth of the Web. Unfortunately, existing business intelligence (BI) tools lack analysis capability, and research on BI systems is sparse. This research proposes a framework for designing BI systems to identify and to classify stakeholders on the Web, incorporating human knowledge and machine‐learned information from Web pages. Based on the framework, we have developed a prototype called Business Stakeholder Analyzer (BSA) that helps managers and analysts to identify and to classify their stakeholders on the Web. Results from our experiment involving algorithm comparison, feature comparison, and a user study showed that the system achieved better within‐class accuracies in widespread stakeholder types such as partner/sponsor/supplier and media/reviewer, and was more efficient than human classification. The student and practitioner subjects in our user study strongly agreed that such a system would save analysts' time and help to identify and classify stakeholders. This research contributes to a better understanding of how to integrate information technology with stakeholder theory, and enriches the knowledge base of BI system design. Wingyan Chung, Hsinchun Chen, Edna Reid |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2009 | Arizona Literature Mapper: An integrated approach to monitor and analyze global bioterrorism research literatureabstractAbstract Biomedical research is critical to biodefense, which is drawing increasing attention from governments globally as well as from various research communities. The U.S. government has been closely monitoring and regulating biomedical research activities, particularly those studying or involving bioterrorism agents or diseases. Effective surveillance requires comprehensive understanding of extant biomedical research and timely detection of new developments or emerging trends. The rapid knowledge expansion, technical breakthroughs, and spiraling collaboration networks demand greater support for literature search and sharing, which cannot be effectively supported by conventional literature search mechanisms or systems. In this study, we propose an integrated approach that integrates advanced techniques for content analysis, network analysis, and information visualization. We design and implement Arizona Literature Mapper, a Web‐based portal that allows users to gain timely, comprehensive understanding of bioterrorism research, including leading scientists, research groups, institutions as well as insights about current mainstream interests or emerging trends. We conduct two user studies to evaluate Arizona Literature Mapper and include a well‐known system for benchmarking purposes. According to our results, Arizona Literature Mapper is significantly more effective for supporting users' search of bioterrorism publications than PubMed. Users consider Arizona Literature Mapper more useful and easier to use than PubMed. Users are also more satisfied with Arizona Literature Mapper and show stronger intentions to use it in the future. Assessments of Arizona Literature Mapper's analysis functions are also positive, as our subjects consider them useful, easy to use, and satisfactory. Our results have important implications that are also discussed in the article. Yan Dang 0001, Gavin Yulei Zhang, Hsinchun Chen, Paul Jen-Hwa Hu, Susan A. Brown, Cathy Larson |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2009 | Identifying significant facilitators of dark network evolutionabstractAbstract Social networks evolve over time with the addition and removal of nodes and links to survive and thrive in their environments. Previous studies have shown that the link‐formation process in such networks is influenced by a set of facilitators. However, there have been few empirical evaluations to determine the important facilitators. In a research partnership with law enforcement agencies, we used dynamic social‐network analysis methods to examine several plausible facilitators of co‐offending relationships in a large‐scale narcotics network consisting of individuals and vehicles. Multivariate Cox regression and a two‐proportion z‐test on cyclic and focal closures of the network showed that mutual acquaintance and vehicle affiliations were significant facilitators for the network under study. We also found that homophily with respect to age, race, and gender were not good predictors of future link formation in these networks. Moreover, we examined the social causes and policy implications for the significance and insignificance of various facilitators including common jails on future co‐offending. These findings provide important insights into the link‐formation processes and the resilience of social networks. In addition, they can be used to aid in the prediction of future links. The methods described can also help in understanding the driving forces behind the formation and evolution of social networks facilitated by mobile and Web technologies. Daning Hu, Siddharth Kaza, Hsinchun Chen |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2009 | Topological Analysis of Criminal Activity Networks: Enhancing Transportation SecurityabstractThe security of border and transportation systems is a critical component of the national strategy for homeland security. The security concerns at the border are not independent of law enforcement in border-area jurisdictions because the information known by local law enforcement agencies may provide valuable leads that are useful for securing the border and transportation infrastructure. The combined analysis of law enforcement information and data generated by vehicle license plate readers at international borders can be used to identify suspicious vehicles and people at ports of entry. This not only generates better quality leads for border protection agents but may also serve to reduce wait times for commerce, vehicles, and people as they cross the border. This paper explores the use of criminal activity networks (CANs) to analyze information from law enforcement and other sources to provide value for transportation and border security. We analyze the topological characteristics of CAN of individuals and vehicles in a multiple jurisdiction scenario. The advantages of exploring the relationships of individuals and vehicles are shown. We find that large narcotic networks are small world with short average path lengths ranging from 4.5 to 8.5 and have scale-free degree distributions with power law exponents of 0.85-1.3. In addition, we find that utilizing information from multiple jurisdictions provides higher quality leads by reducing the average shortest-path lengths. The inclusion of vehicular relationships and border-crossing information generates more investigative leads that can aid in securing the border and transportation infrastructure. Siddharth Kaza, Jennifer Jie Xu 0001, Byron Marshall, Hsinchun Chen |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2009 | Textual analysis of stock market prediction using breaking financial news: The AZFin text systemabstractOur research examines a predictive machine learning approach for financial news articles analysis using several different textual representations: bag of words, noun phrases, and named entities. Through this approach, we investigated 9,211 financial news articles and 10,259,042 stock quotes covering the S&P 500 stocks during a five week period. We applied our analysis to estimate a discrete stock price twenty minutes after a news article was released. Using a support vector machine (SVM) derivative specially tailored for discrete numeric prediction and models containing different stock-specific variables, we show that the model containing both article terms and stock price at the time of article release had the best performance in closeness to the actual future stock price (MSE 0.04261), the same direction of price movement as the future price (57.1% directional accuracy) and the highest return using a simulated trading engine (2.06% return). We further investigated the different textual representations and found that a Proper Noun scheme performs better than the de facto standard of Bag of Words in all three metrics. Robert P. Schumaker, Hsinchun Chen |
ACM Trans. Inf. Syst. | 2 |
| 2008 | Homeland security data mining using social network analysisabstractThe tragic events of September 11th have caused drastic effects on many aspects of society. Academics in the fields of computational and information science have been called upon to help enhance the government’s ability to fight terrorism and other crimes. Keeping in mind the special characteristics of crimes and security-related data, data mining techniques can contribute in six areas of research: information sharing and collaboration, security association mining, classification and clustering, intelligence text mining, spatial and temporal crime pattern mining, and criminal/terrorist network analysis. Grounded on social network analysis (SNA) research, criminal network analysis and terrorist network analysis have been shown to be most promising for public safety and homeland security. Based on the University of Arizona’s highly successful COPLINK and Dark Web projects, we will discuss relevant SNA for “dark networks” (criminal and terrorist networks). Selected techniques, examples, and case studies will be presented based on gang/narcotic networks, US extremist networks, Al Qaeda member networks, and international Jihadist web site and forum networks. Unique homeland security challenges and future directions will also be presented. Hsinchun Chen |
ISI | 1 |
| 2008 | Discovery of improvised explosive device content in the Dark WebabstractImprovised explosive device related Web content offers a wealth of knowledge to members of the security and intelligence communities. However, acquiring the desired topical information remains a challenge for analysts due to issues including site identification, accessibility, and language. This paper presents a focused crawling approach for the discovery and collection of improvised explosive device content from the dark Web. Results and examples from an exploratory collection effort are described. Site map and link analyses were also performed, offering insight into the communication dynamics and publication of improvised explosive device Web content. Hsinchun Chen |
ISI | 1 |
| 2008 | IEDs in the Dark Web: Genre classification of improvised explosive device web pagesabstractImprovised explosive device web pages represent a significant source of knowledge for security organizations. These web pages exist in distinctive genres of communication, providing different types and levels of information for the intelligence community. This paper presents a framework for the classification of improvised explosive device web pages by genre. The approach uses a complex feature extractor, extended feature representation, and support vector machine learning algorithms. Improvised explosive device web pages were collected from the Dark Web and two classification models were examined, one using feature selection. Classification accuracy exceeded 88%. Hsinchun Chen |
ISI | 1 |
| 2008 | Sentiment and affect analysis of Dark Web forums: Measuring radicalization on the internetabstractDark Web forums are heavily used by extremist and terrorist groups for communication, recruiting, ideology sharing, and radicalization. These forums often have relevance to the Iraqi insurgency or Al-Qaeda and are of interest to security and intelligence organizations. This paper presents an automated approach to sentiment and affect analysis of selected radical international Ahadist Dark Web forums. The approach incorporates a rich textual feature representation and machine learning techniques to identify and measure the sentiment polarities and affect intensities expressed in forum communications. The results of sentiment and affect analysis performed on two large-scale Dark Web forums are presented, offering insight into the communities and participants. Hsinchun Chen |
ISI | 1 |
| 2008 | Developing ideological networks using social network analysis and writeprints: A case study of the international Falun Gong movementabstractThe convenience of the Internet has made it possible for activist groups to easily form alliances through their websites to appeal to wider audience and increase their impact. In this study, we investigate the potential of using Social Network Analysis (SNA) and Writeprints to discover the fusion of activitst ideas on the Internet, focusing on the Falun Gong movement. We find that network visualization is very useful to reveal how different types of websites or ideas are associated and, in some cases, mixed together. Furthermore, the measures of centrality in SNA help to reveal which websites most prominently link to other websites. We find that Writeprints can be used to identify the ideas which an author gradually introduces and combines through a series of messages. Yi-Da Chen, Ahmed Abbasi, Hsinchun Chen |
ISI | 3 |
| 2008 | Cyber extremism in Web 2.0: An exploratory study of international Jihadist groupsabstractAs part of the NSF-funded Dark Web research project, this paper presents an exploratory study of cyber extremism on the Web 2.0 media: blogs, YouTube, and Second Life. We examine international Jihadist extremist groups that use each of these media. We observe that these new, interactive, multimedia-rich forms of communication provide effective means for extremists to promote their ideas, share resources, and communicate among each other. The development of automated collection and analysis tools for Web 2.0 can help policy makers, intelligence analysts, and researchers to better understand extremistspsila ideas and communication patterns, which may lead to strategies that can counter the threats posed by extremists in the second-generation Web. Hsinchun Chen, S. Thoms, Tianjun Fu |
ISI | 1 |
| 2008 | An integrated approach to mapping worldwide bioterrorism research capabilitiesabstractBiomedical research used for defense purposes may also be applied to biological weapons development. To mitigate risk, the U.S. Government has attempted to monitor and regulate biomedical research labs, especially those that study bioterrorism agents/diseases. However, monitoring worldwide biomedical researchers and their work is still an issue. In this study, we developed an integrated approach to mapping worldwide bioterrorism research literature. By utilizing knowledge mapping techniques, we analyzed the productivity status, collaboration status, and emerging topics in bioterrorism domain. The analysis results provide insights into the research status of bioterrorism agents/diseases and thus allow a more comprehensive view of bioterrorism researchers and ongoing work. Yan Dang 0001, Gavin Yulei Zhang, Nichalin S. Summerfield, Cathy Larson, Hsinchun Chen |
ISI | 5 |
| 2008 | Analysis of cyberactivism: A case study of online free Tibet activitiesabstractCyberactivism refers to the use of the Internet to advocate vigorous or intentional actions to bring about social or political change. Cyberactivism analysis aims to improve the understanding of cyber activists and their online communities. In this paper, we present a case study of online Free Tibet activities. For web site analysis, we use the inlink and outlink information of five selected seed URLs to construct the network of Free Tibet web sites. The network shows the close relationships between our five seed sites. Centrality measures reveal that tibet.org is probably an information hub site in the network. Further content analysis tells us that common hub site words are most popular in tibet.org whereas dalailama.com focuses mostly on religious words. For forum analysis, descriptive statistics such as the number of posts each month and the post distribution of forum users illustrate that the two large forums FreeTibetAndYou and RFAnews-Tibbs have experienced significant reduction in activities in recent years and that a small percentage of their users contribute the majority of posts. Important phrases of several long threads and active forum users are identified by using mutual information and TF-IDF scores. Such topical analyses help us understand the topics discussed in the forums and the ideas and interest of those forum users. Finally, social network analyses of the forum users are conducted to reflect their interactions and the social structure of their online communities. Tianjun Fu, Hsinchun Chen |
ISI | 2 |
| 2008 | PRM-based identity matching using social contextabstractIdentity management is critical for many intelligence and security applications. Identity information is not reliable due to the problems of unintentional errors and intentional deception by the criminals. Most of existing identity matching techniques consider personal identity features only. In this article we propose a PRM-based identity matching technique that takes both personal identity features and social contexts into account. We identify two groups of social context features, namely social activity and social relation features. Experiments show that the social activity features significantly improve the matching performance while the social relation features effectively reduce false positive and false negative. Jiexun Li, G. Alan Wang, Hsinchun Chen |
ISI | 3 |
| 2008 | Bioterrorism event detection based on the Markov switching model: A simulated anthrax outbreak studyabstractThe threat of infectious disease outbreaks and bioterrorism attacks has stimulated the development of syndromic surveillance systems, which focus on using pre-diagnostic data such as emergency department chief complaints and over-the-counter (OTC) drug sales to detect bioterrorism events in a timely manner. A key function of syndromic surveillance systems is detecting possible bioterrorism events from time series data. In this paper, we propose a novel temporal outbreak detection method based on the Markov switching model, a special case of hidden Markov models. The model is motivated to address several computational problems with existing detection schemes concerning the inconsistency in parameter estimation and the resulting undesired detection performance. Preliminary evaluation using simulated outbreaks injected on authentic time series shows that our method outperforms benchmark methods in terms of outbreak detection speed and detection sensitivity at given levels of false alarm rates. Hsin-Min Lu, Daniel Dajun Zeng, Hsinchun Chen |
ISI | 3 |
| 2008 | Botnets, and the cybercriminal undergroundabstractAn underground community of cyber criminals has grown in recent years with powerful technologies capable of inflicting serious economic and infrastructural harm in the digital age. This paper serves as an introduction to the world of botnets and to the efforts of the nonprofit group “The ShadowServer Foundation” to track them. A data mining exploration is performed on ShadowServer’s datasets to investigate possible classification mechanisms for threat assessment. Clinton J. Mielke, Hsinchun Chen |
ISI | 2 |
| 2008 | A stack-based prospective spatio-temporal data analysis approach
Wei Chang 0006, Daniel Dajun Zeng, Hsinchun Chen |
Decis. Support Syst. | 3 |
| 2008 | A machine learning approach to web page filtering using content and structure analysis
Michael Chau, Hsinchun Chen |
Decis. Support Syst. | 2 |
| 2008 | SpidersRUs: Creating specialized search engines in multiple languages
Michael Chau, Jialun Qin, Yilu Zhou, Chunju Tseng, Hsinchun Chen |
Decis. Support Syst. | 5 |
| 2008 | Evaluating ontology mapping techniques: An experiment in public safety information sharing
Siddharth Kaza, Hsinchun Chen |
Decis. Support Syst. | 2 |
| 2008 | Editors' introduction special issue on multilingual knowledge management
Christopher C. Yang, Chih-Ping Wei, Hsinchun Chen |
Decis. Support Syst. | 3 |
| 2008 | Communication-Garden System: Visualizing a computer-mediated communication process
Bin Zhu 0001, Hsinchun Chen |
Decis. Support Syst. | 2 |
| 2008 | Organizing domain-specific information on the Web: An experiment on the Spanish business Web directory
Wingyan Chung, Guanpi Lai, Alfonso Bonillas, Hsinchun Chen |
Int. J. Hum. Comput. Stud. | 5 |
| 2008 | Uncovering the dark Web: A case study of Jihad on the WebabstractAbstract While the Web has become a worldwide platform for communication, terrorists share their ideology and communicate with members on the “Dark Web”—the reverse side of the Web used by terrorists. Currently, the problems of information overload and difficulty to obtain a comprehensive picture of terrorist activities hinder effective and efficient analysis of terrorist information on the Web. To improve understanding of terrorist activities, we have developed a novel methodology for collecting and analyzing Dark Web information. The methodology incorporates information collection, analysis, and visualization techniques, and exploits various Web information sources. We applied it to collecting and analyzing information of 39 Jihad Web sites and developed visualization of their site contents, relationships, and activity levels. An expert evaluation showed that the methodology is very useful and promising, having a high potential to assist in investigation and understanding of terrorist activities by producing results that could potentially help guide both policymaking and intelligence research. Hsinchun Chen, Wingyan Chung, Jialun Qin, Edna Reid, Marc Sageman, Gabriel Weimann |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2008 | A hybrid approach to Web forum interactional coherence analysisabstractAbstract Despite the rapid growth of text‐based computer‐mediated communication (CMC), its limitations have rendered the media highly incoherent. This poses problems for content analysis of online discourse archives. Interactional coherence analysis (ICA) attempts to accurately identify and construct CMC interaction networks. In this study, we propose the Hybrid Interactional Coherence (HIC) algorithm for identification of web forum interaction. HIC utilizes a bevy of system and linguistic features, including message header information, quotations, direct address, and lexical relations. Furthermore, several similarity‐based methods including a Lexical Match Algorithm (LMA) and a sliding window method are utilized to account for interactional idiosyncrasies. Experiments results on two web forums revealed that the proposed HIC algorithm significantly outperformed comparison techniques in terms of precision, recall, and F‐measure at both the forum and thread levels. Additionally, an example was used to illustrate how the improved ICA results can facilitate enhanced social network and role analysis capabilities. Tianjun Fu, Ahmed Abbasi, Hsinchun Chen |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2008 | Kernel-based learning for biomedical relation extractionabstractAbstract Relation extraction is the process of scanning text for relationships between named entities. Recently, significant studies have focused on automatically extracting relations from biomedical corpora. Most existing biomedical relation extractors require manual creation of biomedi‐cal lexicons or parsing templates based on domain knowledge. In this study, we propose to use kernel‐based learning methods to automatically extract biomedical relations from literature text. We develop a framework of kernel‐based learning for biomedical relation extraction. In particular, we modified the standard tree kernel function by incorporating a trace kernel to capture richer contextual information. In our experiments on a biomedi‐cal corpus, we compare different kernel functions for biomedical relation detection and classification. Theexperimental results show that a tree kernel outperforms word and sequence kernels for relation detection, our trace‐tree kernel outperforms the standard tree kernel, and a composite kernel outperforms individual kernels for relation extraction. Jiexun Li, Xin Li 0004, Hsinchun Chen |
J. Assoc. Inf. Sci. Technol. | 4 |
| 2008 | Using importance flooding to identify interesting networks of criminal activityabstractAbstract Effectively harnessing available data to support homeland‐security‐related applications is a major focus in the emerging science of intelligence and security informatics (ISI). Many studies have focused on criminal‐network analysis as a major challenge within the ISI domain. Though various methodologies have been proposed, none have been tested for usefulness in creating link charts. This study compares manually created link charts to suggestions made by the proposed importance‐flooding algorithm. Mirroring manual investigational processes, our iterative computation employs association‐strength metrics, incorporates path‐based node importance heuristics, allows for case‐specific notions of importance, and adjusts based on the accuracy of previous suggestions. Interesting items are identified by leveraging both node attributes and network structure in a single computation. Our data set was systematically constructed from heterogeneous sources and omits many privacy‐sensitive data elements such as case narratives and phone numbers. The flooding algorithm improved on both manual and link‐weight‐only computations, and our results suggest that the approach is robust across different interpretations of the user‐provided heuristics. This study demonstrates an interesting methodology for including user‐provided heuristics in network‐based analysis, and can help guide the development of ISI‐related analysis tools. Byron Marshall, Hsinchun Chen, Siddharth Kaza |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2008 | Evaluating a news-aware quantitative trader: The effect of momentum and contrarian stock selection strategiesabstractAbstract We study the coupling of basic quantitative portfolio selection strategies with a financial news article prediction system, AZFinText. By varying the degrees of portfolio formation time, we found that a hybrid system using both quantitative strategy and a full set of financial news articles performed the best. With a 1‐week portfolio formation period, we achieved a 20.79% trading return using a Momentum strategy and a 4.54% return using a Contrarian strategy over a 5‐week holding period. We also found that trader overreaction to these events led AZFinText to capitalize on these short‐term surges in price. Robert P. Schumaker, Hsinchun Chen |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2008 | Ontology-enhanced automatic chief complaint classification for syndromic surveillance
Hsin-Min Lu, Daniel Dajun Zeng, Lea Trujillo, Ken Komatsu, Hsinchun Chen |
J. Biomed. Informatics | 5 |
| 2008 | Affect Analysis of Web Forums and Blogs Using Correlation EnsemblesabstractAnalysis of affective intensities in computer-mediated communication is important in order to allow a better understanding of online users' emotions and preferences. Despite considerable research on textual affect classification, it is unclear which features and techniques are most effective. In this study, we compared several feature representations for affect analysis, including learned n-grams and various automatically and manually crafted affect lexicons. We also proposed the support vector regression correlation ensemble (SVRCE) method for enhanced classification of affect intensities. SVRCE uses an ensemble of classifiers each trained using a feature subset tailored toward classifying a single affect class. The ensemble is combined with affect correlation information to enable better prediction of emotive intensities. Experiments were conducted on four test beds encompassing web forums, blogs, and online stories. The results revealed that learned n-grams were more effective than lexicon-based affect representations. The findings also indicated that SVRCE outperformed comparison techniques, including Pace regression, semantic orientation, and WordNet models. Ablation testing showed that the improved performance of SVRCE was attributable to its use of feature ensembles as well as affect correlation information. A brief case study was conducted to illustrate the utility of the features and techniques for affect analysis of large archives of online discourse. Ahmed Abbasi, Hsinchun Chen, S. Thoms, Tianjun Fu |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2008 | Guest Editors' Introduction: Special Section on Intelligence and Security InformaticsabstractThe 12 papers in this special section focus on intelligence and security informatics. They are summarized here. Daniel Dajun Zeng, Hsinchun Chen, Fei-Yue Wang 0001, Hillol Kargupta |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2008 | Writeprints: A stylometric approach to identity-level identification and similarity detection in cyberspaceabstractOne of the problems often associated with online anonymity is that it hinders social accountability, as substantiated by the high levels of cybercrime. Although identity cues are scarce in cyberspace, individuals often leave behind textual identity traces. In this study we proposed the use of stylometric analysis techniques to help identify individuals based on writing style. We incorporated a rich set of stylistic features, including lexical, syntactic, structural, content-specific, and idiosyncratic attributes. We also developed the Writeprints technique for identification and similarity detection of anonymous identities. Writeprints is a Karhunen-Loeve transforms-based technique that uses a sliding window and pattern disruption algorithm with individual author-level feature sets. The Writeprints technique and extended feature set were evaluated on a testbed encompassing four online datasets spanning different domains: email, instant messaging, feedback comments, and program code. Writeprints outperformed benchmark techniques, including SVM, Ensemble SVM, PCA, and standard Karhunen-Loeve transforms, on the identification and similarity detection tasks with accuracy as high as 94% when differentiating between 100 authors. The extended feature set also significantly outperformed a baseline set of features commonly used in previous research. Furthermore, individual-author-level feature sets generally outperformed use of a single group of attributes. Ahmed Abbasi, Hsinchun Chen |
ACM Trans. Inf. Syst. | 2 |
| 2008 | Sentiment analysis in multiple languages: Feature selection for opinion classification in Web forumsabstractThe Internet is frequently used as a medium for exchange of information and opinions, as well as propaganda dissemination. In this study the use of sentiment analysis methodologies is proposed for classification of Web forum opinions in multiple languages. The utility of stylistic and syntactic features is evaluated for sentiment classification of English and Arabic content. Specific feature extraction components are integrated to account for the linguistic characteristics of Arabic. The entropy weighted genetic algorithm (EWGA) is also developed, which is a hybridized genetic algorithm that incorporates the information-gain heuristic for feature selection. EWGA is designed to improve performance and get a better assessment of key features. The proposed features and techniques are evaluated on a benchmark movie review dataset and U.S. and Middle Eastern Web forum postings. The experimental results using EWGA with SVM indicate high performance levels, with accuracies of over 91% on the benchmark dataset as well as the U.S. and Middle Eastern forums. Stylistic features significantly enhanced performance across all testbeds while EWGA also outperformed other feature selection methods, indicating the utility of these features and techniques for document-level classification of sentiments. Ahmed Abbasi, Hsinchun Chen, Arab Salem |
ACM Trans. Inf. Syst. | 2 |
| 2007 | Graph Kernel-Based Learning for Gene Function Prediction from Gene Interaction NetworkabstractPrediction of gene functions is a major challenge to biologists in the post-genomic era. Interactions between genes and their products compose networks and can be used to infer gene functions. Most previous studies used heuristic approaches based on either local or global information of gene interaction networks to assign unknown gene functions. In this study, we propose a graph kernel-based method that can capture the structure of gene interaction networks to predict gene functions. We conducted an experimental study on a test-bed of P53-related genes. The experimental results demonstrated better performance for our proposed method as compared with baseline methods. Xin Li 0004, Hsinchun Chen, Jiexun Li |
BIBM | 3 |
| 2007 | Affect Intensity Analysis of Dark Web ForumsabstractAffects play an important role in influencing people's perceptions and decision making. Affect analysis is useful for measuring the presence of hate, violence, and the resulting propaganda dissemination across extremist groups. In this study we performed affect analysis of U.S. and Middle Eastern extremist group forum postings. We constructed an affect lexicon using a probabilistic disambiguation technique to measure the usage of violence and hate affects. These techniques facilitate in depth analysis of multilingual content. The proposed approach was evaluated by applying it across 16 U.S. supremacist and Middle Eastern extremist group forums. Analysis across regions reveals that the Middle Eastern test bed forums have considerably greater violence intensity than the U.S. groups. There is also a strong linear relationship between the usage of hate and violence across the Middle Eastern messages. Ahmed Abbasi, Hsinchun Chen |
ISI | 2 |
| 2007 | Interaction Coherence Analysis for Dark Web ForumsabstractInteraction coherence analysis (ICA) attempts to accurately identify and construct interaction networks by using various features and techniques. It is useful to identify user roles, user's social and information value, as well as the social network structure of Dark Web communities. In this study, we applied interaction coherence analysis for Dark Web forums using the hybrid interaction coherence (HIC) algorithm. Our algorithm utilizes both system features such as header information and quotations, and linguistic features such as direct address and lexical relation. Furthermore, several similarity-based methods, for example vector space model, dice equation, and sliding window, are used to address various types of noises. Two experiments have been conducted to compare our HIC algorithm with traditional linkage-based method, similarity-based method, and a simplified HIC method that does not address noise issues. The results demonstrate the effectiveness of our HIC algorithm for identifying interactions in Dark Web forums. Tianjun Fu, Ahmed Abbasi, Hsinchun Chen |
ISI | 3 |
| 2007 | Dynamic Social Network Analysis of a Dark Network: Identifying Significant Facilitatorsabstract"Dark Networks" refer to various illegal and covert social networks like criminal and terrorist networks. These networks evolve over time with the formation and dissolution of links to survive control efforts by authorities. Previous studies have shown that the link formation process in such networks is influenced by a set of facilitators. However, there have been few empirical evaluations to determine the significant facilitators. In this study, we used dynamic social network analysis methods to examine several plausible link formation facilitators in a large-scale real-world narcotics network. Multivariate Cox regression showed that mutual acquaintance and vehicle affiliations were significant facilitators in the network under study. These findings provide insights into the link formation processes and the resilience of dark networks. They also can be used to help authorities predict co-offending in future crimes. Siddharth Kaza, Daning Hu, Hsinchun Chen |
ISI | 3 |
| 2007 | Medical Ontology-Enhanced Text Processing for Infectious Disease InformaticsabstractInfectious disease informatics, as a sub-field of security informatics, is concerned with development of the science and technologies needed for collecting, sharing, reporting, analyzing, and visualizing infectious disease data; and for providing data and decision-making support for infectious disease prevention, detection, and management. Syndromic surveillance is a major study area of infectious disease informatics, focusing on identifying in a timely manner possible infectious disease outbreaks based on pre-diagnostic data. Free-text chief complaints (CCs), short phrases describing reasons for patients' emergency department visits, are a major source of data for syndromic surveillance. For surveillance purposes, CCs need to be classified into syndromic categories. However, the lack of standard vocabulary and high-quality encoding of CCs hinder effective classification. To meet this challenge, we have developed an ontology-enhanced automatic CC classification approach. Exploiting semantic relations in the UMLS, a medical ontology, this approach is motivated to address the CC word variation problem in general and to meet the specific need for a flexible classification approach capable of handling multiple sets of syndrome categories. Hsin-Min Lu, Daniel Dajun Zeng, Hsinchun Chen |
ISI | 3 |
| 2007 | The Arizona IDMatcher: A Probabilistic Identity Matching SystemabstractVarious law enforcement and intelligence tasks require managing identity information in an effective and efficient way. However, the quality issues of identity information make this task non-trivial. Various heuristic based systems have been developed to tackle the identity matching problem. However, deploying such systems may require special expertise in system configuration and customization for optimal system performance. In this paper, we propose an alternative system called the Arizona IDMatcher. The system relies on a machine learning algorithm to automatically generate a decision model for identity matching. Such a system requires minimal human configuration effort. Experiments show that the Arizona IDMatcher is very efficient in detecting matching identity records. Compared to IBM Identity Resolution (a commercial, heuristic-based system), the Arizona IDMatcher achieves better recall and overall F-measures in identifying matching identities in two large-scale real-world datasets. G. Alan Wang, Siddharth Kaza, Shailesh Joshi, Kris Chang, Homa Atabakhsh, Hsinchun Chen |
ISI | 6 |
| 2007 | Large-scale regulatory network analysis from microarray data: modified Bayesian network learning and association rule mining
Zan Huang, Jiexun Li, George S. Watts, Hsinchun Chen |
Decis. Support Syst. | 5 |
| 2007 | Enhancing border security: Mutual information analysis to identify suspect vehicles
Siddharth Kaza, Hsinchun Chen |
Decis. Support Syst. | 3 |
| 2007 | Introduction to the special issue on decision support in medicine
Gondy Leroy, Hsinchun Chen |
Decis. Support Syst. | 2 |
| 2007 | Cyberinfrastructure for homeland security: Advances in information sharing, data mining, and collaboration systems
T. S. Raghu 0001, Hsinchun Chen |
Decis. Support Syst. | 2 |
| 2007 | Leveraging Question Answer technology to address terrorism inquiry
Robert P. Schumaker, Hsinchun Chen |
Decis. Support Syst. | 2 |
| 2007 | An evaluation of the chat and knowledge delivery components of a low-level dialog system: The AZ-ALICE experiment
Robert P. Schumaker, Mark Ginsburg, Hsinchun Chen |
Decis. Support Syst. | 3 |
| 2007 | Analyzing terror campaigns on the internet: Technical sophistication, content richness, and Web interactivity
Jialun Qin, Yilu Zhou, Edna Reid, Guanpi Lai, Hsinchun Chen |
Int. J. Hum. Comput. Stud. | 5 |
| 2007 | Mapping the contemporary terrorism research domain
Edna Reid, Hsinchun Chen |
Int. J. Hum. Comput. Stud. | 2 |
| 2007 | Redips: Backlink search and analysis on the Web for business intelligence analysisabstractAbstract The World Wide Web presents significant opportunities for business intelligence analysis as it can provide information about a company's external environment and its stakeholders. Traditional business intelligence analysis on the Web has focused on simple keyword searching. Recently, it has been suggested that the incoming links, or backlinks, of a company's Web site (i.e., other Web pages that have a hyperlink pointing to the company of interest) can provide important insights about the company's “online communities.” Although analysis of these communities can provide useful signals for a company and information about its stakeholder groups, the manual analysis process can be very time‐consuming for business analysts and consultants. In this article, we present a tool called Redips that automatically integrates backlink meta‐searching and text‐mining techniques to facilitate users in performing such business intelligence analysis on the Web. The architectural design and implementation of the tool are presented in the article. To evaluate the effectiveness, efficiency, and user satisfaction of Redips, an experiment was conducted to compare the tool with two popular business intelligence analysis methods—using backlink search engines and manual browsing. The experiment results showed that Redips was statistically more effective than both benchmark methods (in terms of Recall and F‐measure) but required more time in search tasks. In terms of user satisfaction, Redips scored statistically higher than backlink search engines in all five measures used, and also statistically higher than manual browsing in three measures. Michael Chau, Boby Shiu, Ivy Chan, Hsinchun Chen |
J. Assoc. Inf. Sci. Technol. | 4 |
| 2007 | Automated criminal link analysis based on domain knowledgeabstractAbstract Link (association) analysis has been used in the criminal justice domain to search large datasets for associations between crime entities in order to facilitate crime investigations. However, link analysis still faces many challenging problems, such as information overload, high search complexity, and heavy reliance on domain knowledge. To address these challenges, this article proposes several techniques for automated, effective, and efficient link analysis. These techniques include the co‐occurrence analysis, the shortest path algorithm, and a heuristic approach to identifying associations and determining their importance. We developed a prototype system called CrimeLink Explorer based on the proposed techniques. Results of a user study with 10 crime investigators from the Tucson Police Department showed that our system could help subjects conduct link analysis more efficiently than traditional single‐level link analysis tools. Moreover, subjects believed that association paths found based on the heuristic approach were more accurate than those found based solely on the co‐occurrence analysis and that the automated link analysis system would be of great help in crime investigations. Jennifer Schroeder, Jennifer Jie Xu 0001, Hsinchun Chen, Michael Chau |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2007 | Global mapping of gene/protein interactions in PubMed abstracts: A framework and an experiment with P53 interactions
Xin Li 0004, Hsinchun Chen, Zan Huang, Jesse D. Martinez |
J. Biomed. Informatics | 2 |
| 2007 | System for Infectious Disease Information Sharing and Analysis: Design and EvaluationabstractMotivated by the importance of infectious disease informatics (IDI) and the challenges to IDI system development and data sharing, we design and implement BioPortal, a Web-based IDI system that integrates cross-jurisdictional data to support information sharing, analysis, and visualization in public health. In this paper, we discuss general challenges in IDI, describe BioPortal's architecture and functionalities, and highlight encouraging evaluation results obtained from a controlled experiment that focused on analysis accuracy, task performance efficiency, user information satisfaction, system usability, usefulness, and ease of use. Paul Jen-Hwa Hu, Daniel Dajun Zeng, Hsinchun Chen, Cathy Larson, Wei Chang 0006, Chunju Tseng |
IEEE Trans. Inf. Technol. Biomed. | 3 |
| 2007 | Optimal Search-Based Gene Subset Selection for Gene Array Cancer ClassificationabstractHigh dimensionality has been a major problem for gene array-based cancer classification. It is critical to identify marker genes for cancer diagnoses. We developed a framework of gene selection methods based on previous studies. This paper focuses on optimal search-based subset selection methods because they evaluate the group performance of genes and help to pinpoint global optimal set of marker genes. Notably, this paper is the first to introduce tabu search (TS) to gene selection from high-dimensional gene array data. Our comparative study of gene selection methods demonstrated the effectiveness of optimal search-based gene subset selection to identify cancer marker genes. TS was shown to be a promising tool for gene subset selection. Jiexun Li, Hsinchun Chen, Bernard W. Futscher |
IEEE Trans. Inf. Technol. Biomed. | 3 |
| 2007 | User-Centered Evaluation of Arizona BioPathway: An Information Extraction, Integration, and Visualization SystemabstractExplosive growth in biomedical research has made automated information extraction, knowledge integration, and visualization increasingly important and critically needed. The Arizona BioPathway (ABP) system extracts and displays biological regulatory pathway information from the abstracts of journal articles. This study uses relations extracted from more than 200 PubMed abstracts presented in a tabular and graphical user interface with built-in search and aggregation functionality. This paper presents a task-centered assessment of the usefulness and usability of the ABP system focusing on its relation aggregation and visualization functionalities. Results suggest that our graph-based visualization is more efficient in supporting pathway analysis tasks and is perceived as more useful and easier to use as compared to a text-based literature-viewing method. Relation aggregation significantly contributes to knowledge-acquisition efficiency. Together, the graphic and tabular views in the ABP Visualizer provide a flexible and effective interface for pathway relation browsing and analysis. Our study contributes to pathway-related research and biological information extraction by assessing the value of a multiview, relation-based interface that supports user-controlled exploration of pathway information across multiple granularities. Karin D. Quiñones, Byron Marshall, Shauna Eggers, Hsinchun Chen |
IEEE Trans. Inf. Technol. Biomed. | 5 |
| 2007 | Incorporating Web Analysis Into Neural Networks: An Example in Hopfield Net SearchingabstractNeural networks have been used in various applications on the World Wide Web, but most of them only rely on the available input-output examples without incorporating Web-specific knowledge, such as Web link analysis, into the network design. In this paper, we propose a new approach in which the Web is modeled as an asymmetric Hopfield Net. Each neuron in the network represents a Web page, and the connections between neurons represent the hyperlinks between Web pages. Web content analysis and Web link analysis are also incorporated into the model by adding a page content score function and a link score function into the weights of the neurons and the synapses, respectively. A simulation study was conducted to compare the proposed model with traditional Web search algorithms, namely, a breadth-first search and a best-first search using PageRank as the heuristic. The results showed that the proposed model performed more efficiently and effectively in searching for domain-specific Web pages. We believe that the model can also be useful in other Web applications such as Web page clustering and search result ranking Michael Chau, Hsinchun Chen |
IEEE Trans. Syst. Man Cybern. Part C | 2 |
| 2006 | Visualizing Authorship for Identification
Ahmed Abbasi, Hsinchun Chen |
ISI | 2 |
| 2006 | Suspect Vehicle Identification for Border Safety with Modified Mutual Information
Siddharth Kaza, Hsinchun Chen |
ISI | 3 |
| 2006 | Spatial-Temporal Cross-Correlation Analysis: A New Measure and a Case Study in Infectious Disease Informatics
Daniel Dajun Zeng, Hsinchun Chen |
ISI | 3 |
| 2006 | Using Importance Flooding to Identify Interesting Networks of Criminal Activity
Byron Marshall, Hsinchun Chen |
ISI | 2 |
| 2006 | Content Analysis of Jihadi Extremist Groups' Videos
Arab Salem, Edna Reid, Hsinchun Chen |
ISI | 3 |
| 2006 | A Multi-layer Naïve Bayes Model for Approximate Identity Matching
G. Alan Wang, Hsinchun Chen, Homa Atabakhsh |
ISI | 2 |
| 2006 | A Framework for Exploring Gray Web Forums: Analysis of Forum-Based Communities in Taiwan
Jau-Hwang Wang, Tianjun Fu, Hong-Ming Lin, Hsinchun Chen |
ISI | 4 |
| 2006 | On the Topology of the Dark Web of Terrorist Groups
Jennifer Jie Xu 0001, Hsinchun Chen, Yilu Zhou, Jialun Qin |
ISI | 2 |
| 2006 | A Review of Public Health Syndromic Surveillance Systems
Daniel Dajun Zeng, Hsinchun Chen |
ISI | 3 |
| 2006 | Exploring the Dark Side of the Web: Collection and Analysis of U.S. Extremist Online Forums
Yilu Zhou, Jialun Qin, Guanpi Lai, Edna Reid, Hsinchun Chen |
ISI | 5 |
| 2006 | Ontology-Based Automatic Chief Complaints Classification for Syndromic SurveillanceabstractThis paper presents a novel ontology-based approach to classify free-text chief complaints (CCs) into syndrome categories. This approach exploits the semantic relations in a medical ontology to address the CC word variation problem. Initial computational experiments indicate that this ontology-based approach is able to improve significantly the probability that a CC can be correctly classified as a syndrome. Hsin-Min Lu, Daniel Dajun Zeng, Hsinchun Chen |
SMC | 3 |
| 2006 | A framework of integrating gene relations from heterogeneous data sources: an experiment on Arabidopsis thalianaabstractOne of the most important goals of biological investigation is to uncover gene functional relations. In this study we propose a framework for extraction and integration of gene functional relations from diverse biological data sources, including gene expression data, biological literature and genomic sequence information. We introduce a two-layered Bayesian network approach to integrate relations from multiple sources into a genome-wide functional network. An experimental study was conducted on a test-bed of Arabidopsis thaliana. Evaluation of the integrated network demonstrated that relation integration could improve the reliability of relations by combining evidence from different data sources. Domain expert judgments on the gene functional clusters in the network confirmed the validity of our approach for relation integration and network inference. Jiexun Li, Xin Li 0004, Hsinchun Chen, David W. Galbraith |
Bioinform. | 4 |
| 2006 | Building a scientific knowledge web portal: The NanoPort experience
Michael Chau, Zan Huang, Jialun Qin, Yilu Zhou, Hsinchun Chen |
Decis. Support Syst. | 5 |
| 2006 | Intelligence and security informatics: information systems perspective
Hsinchun Chen |
Decis. Support Syst. | 1 |
| 2006 | Supporting non-English Web searching: An experiment on the Spanish business and the Arabic medical intelligence portals
Wingyan Chung, Alfonso Bonillas, Guanpi Lai, Hsinchun Chen |
Decis. Support Syst. | 5 |
| 2006 | Fighting cybercrime: a review and the Taiwan experience
Wingyan Chung, Hsinchun Chen, Weiping Chang, Shihchieh Chou |
Decis. Support Syst. | 2 |
| 2006 | Expertise visualization: An implementation and study based on cognitive fit theory
Zan Huang, Hsinchun Chen, Jennifer Jie Xu 0001, Soushan Wu, Wun-Hwa Chen |
Decis. Support Syst. | 2 |
| 2006 | Matching knowledge elements in concept maps using a similarity flooding algorithm
Byron Marshall, Hsinchun Chen, Therani Madhusudan |
Decis. Support Syst. | 2 |
| 2006 | Process-driven collaboration support for intra-agency crime analysis
J. Leon Zhao, Henry H. Bi, Hsinchun Chen, Daniel Dajun Zeng, Chienting Lin, Michael Chau |
Decis. Support Syst. | 3 |
| 2006 | CMedPort: An integrated approach to facilitating Chinese medical information seeking
Yilu Zhou, Jialun Qin, Hsinchun Chen |
Decis. Support Syst. | 3 |
| 2006 | Evaluating mass knowledge acquisition using the ALICE chatterbot: The AZ-ALICE dialog system
Robert P. Schumaker, Mark Ginsburg, Hsinchun Chen |
Int. J. Hum. Comput. Stud. | 4 |
| 2006 | Multilingual Web retrieval: An experiment in English-Chinese business intelligenceabstractAbstract As increasing numbers of non‐English resources have become available on the Web, the interesting and important issue of how Web users can retrieve documents in different languages has arisen. Cross‐language information retrieval (CLIR), the study of retrieving information in one language by queries expressed in another language, is a promising approach to the problem. Cross‐language information retrieval has attracted much attention in recent years. Most research systems have achieved satisfactory performance on standard Text REtrieval Conference (TREC) collections such as news articles, but CLIR techniques have not been widely studied and evaluated for applications such as Web portals. In this article, the authors present their research in developing and evaluating a multilingual English–Chinese Web portal that incorporates various CLIR techniques for use in the business domain. A dictionary‐based approach was adopted and combines phrasal translation, co‐occurrence analysis, and pre‐ and posttranslation query expansion. The portal was evaluated by domain experts, using a set of queries in both English and Chinese. The experimental results showed that co‐occurrence‐based phrasal translation achieved a 74.6% improvement in precision over simple word‐by‐word translation. When used together, pre‐ and posttranslation query expansion improved the performance slightly, achieving a 78.0% improvement over the baseline word‐by‐word translation approach. In general, applying CLIR techniques in Web applications shows promise. Jialun Qin, Yilu Zhou, Michael Chau, Hsinchun Chen |
J. Assoc. Inf. Sci. Technol. | 4 |
| 2006 | A framework for authorship identification of online messages: Writing-style features and classification techniquesabstractAbstract With the rapid proliferation of Internet technologies and applications, misuse of online messages for inappropriate or illegal purposes has become a major concern for society. The anonymous nature of online‐message distribution makes identity tracing a critical problem. We developed a framework for authorship identification of online messages to address the identity‐tracing problem. In this framework, four types of writing‐style features (lexical, syntactic, structural, and content‐specific features) are extracted and inductive learning algorithms are used to build feature‐based classification models to identify authorship of online messages. To examine this framework, we conducted experiments on English and Chinese online‐newsgroup messages. We compared the discriminating power of the four types of features and of three classification techniques: decision trees, backpropagation neural networks, and support vector machines. The experimental results showed that the proposed approach was able to identify authors of online messages with satisfactory accuracy of 70 to 95%. All four types of message features contributed to discriminating authors of online messages. Support vector machines outperformed the other two classification techniques in our experiments. The high performance we achieved for both the English and Chinese datasets showed the potential of this approach in a multiple‐language context. Jiexun Li, Hsinchun Chen, Zan Huang |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2006 | Moving digital libraries into the student learning space: The GetSmart experienceabstractThe GetSmart system was built to support theoretically sound learning processes in a digital library environment by integrating course management, digital library, and concept mapping components to support a constructivist, six-step, information search process. In the fall of 2002 more than 100 students created 1400 concept maps as part of selected computing classes offered at the University of Arizona and Virginia Tech. Those students conducted searches, obtained course information, created concept maps, collaborated in acquiring knowledge, and presented their knowledge representations. This article connects the design elements of the GetSmart system to targeted concept-map-based learning processes, describes our system and research testbed, and analyzes our system usage logs. Results suggest that students did in fact use the tools in an integrated fashion, combining knowledge representation and search activities. After concept mapping was included in the curriculum, we observed improvement in students' online quiz scores. Further, we observed that students in groups collaboratively constructed concept maps with multiple group members viewing and updating map details. Byron Marshall, Hsinchun Chen, Rao Shen, Edward A. Fox |
ACM J. Educ. Resour. Comput. | 2 |
| 2006 | Aggregating Automatically Extracted Regulatory Pathway RelationsabstractAutomatic tools to extract information from biomedical texts are needed to help researchers leverage the vast and increasing body of biomedical literature. While several biomedical relation extraction systems have been created and tested, little work has been done to meaningfully organize the extracted relations. Organizational processes should consolidate multiple references to the same objects over various levels of granularity, connect those references to other resources, and capture contextual information. We propose a feature decomposition approach to relation aggregation to support a five-level aggregation framework. Our BioAggregate tagger uses this approach to identify key features in extracted relation name strings. We show encouraging feature assignment accuracy and report substantial consolidation in a network of extracted relations. Byron Marshall, Daniel McDonald, Shauna Eggers, Hsinchun Chen |
IEEE Trans. Inf. Technol. Biomed. | 5 |
| 2006 | Summary in context: Searching versus browsingabstractThe use of text summaries in information-seeking research has focused on query-based summaries. Extracting content that resembles the query alone, however, ignores the greater context of the document. Such context may be central to the purpose and meaning of the document. We developed a generic, a query-based, and a hybrid summarizer, each with differing amounts of document context. The generic summarizer used a blend of discourse information and information obtained through traditional surface-level analysis. The query-based summarizer used only query-term information, and the hybrid summarizer used some discourse information along with query-term information. The validity of the generic summarizer was shown through an intrinsic evaluation using a well-established corpus of human-generated summaries. All three summarizers were then compared in an information-seeking experiment involving 297 subjects. Results from the information-seeking experiment showed that the generic summaries outperformed all others in the browse tasks, while the query-based and hybrid summaries outperformed the generic summary in the search tasks. Thus, the document context of generic summaries helped users browse, while such context was not helpful in search tasks. Such results are interesting given that generic summaries have not been studied in search tasks and the that majority of Internet search engines rely solely on query-based summaries. Daniel McDonald, Hsinchun Chen |
ACM Trans. Inf. Syst. | 2 |
| 2005 | Applying Authorship Analysis to Arabic Web Content
Ahmed Abbasi, Hsinchun Chen |
ISI | 2 |
| 2005 | Evaluating an Infectious Disease Information Sharing and Analysis System
Paul Jen-Hwa Hu, Daniel Dajun Zeng, Hsinchun Chen, Catherine A. Larson, Wei Chang 0006, Chunju Tseng |
ISI | 3 |
| 2005 | BorderSafe: Cross-Jurisdictional Information Sharing, Analysis, and Visualization
Siddharth Kaza, Byron Marshall, Jennifer Jie Xu 0001, G. Alan Wang, Hemanth Gowda, Homa Atabakhsh, Tim Petersen, Chuck Violette, Hsinchun Chen |
ISI | 9 |
| 2005 | Analyzing Terrorist Networks: A Case Study of the Global Salafi Jihad Network
Jialun Qin, Jennifer Jie Xu 0001, Daning Hu, Marc Sageman, Hsinchun Chen |
ISI | 5 |
| 2005 | The Dark Web Portal Project: Collecting and Analyzing the Presence of Terrorist Groups on the Web
Jialun Qin, Yilu Zhou, Guanpi Lai, Edna Reid, Marc Sageman, Hsinchun Chen |
ISI | 6 |
| 2005 | Mapping the Contemporary Terrorism Research Domain: Researchers, Publications, and Institutions Analysis
Edna Reid, Hsinchun Chen |
ISI | 2 |
| 2005 | Collecting and Analyzing the Presence of Terrorists on the Web: A Case Study of Jihad Websites
Edna Reid, Jialun Qin, Yilu Zhou, Guanpi Lai, Marc Sageman, Gabriel Weimann, Hsinchun Chen |
ISI | 7 |
| 2005 | Question Answer TARA: A Terrorism Activity Resource Application
Robert P. Schumaker, Hsinchun Chen |
ISI | 2 |
| 2005 | Discovering Identity Problems: A Case Study
G. Alan Wang, Homa Atabakhsh, Tim Petersen, Hsinchun Chen |
ISI | 4 |
| 2005 | BioPortal: Sharing and Analyzing Infectious Disease Information
Daniel Dajun Zeng, Hsinchun Chen, Chunju Tseng, Catherine A. Larson, Wei Chang 0006, Millicent Eidson, Ivan Gotham, Cecil Lynch, Michael Ascher |
ISI | 2 |
| 2005 | Newsmap: a knowledge map for online news
Thian-Huat Ong, Hsinchun Chen, Wai-Ki Sung, Bin Zhu 0001 |
Decis. Support Syst. | 2 |
| 2005 | Visualizing criminal relationships: comparison of a hyperbolic tree and a hierarchical list
Michael Chau, Homa Atabakhsh, Hsinchun Chen |
Decis. Support Syst. | 4 |
| 2005 | Frame-based argumentation for group decision task generation and identification
Pengzhu Zhang, Jingle Sun, Hsinchun Chen |
Decis. Support Syst. | 3 |
| 2005 | Using 3D interfaces to facilitate the spatial knowledge retrieval: a geo-referenced knowledge repository system
Bin Zhu 0001, Hsinchun Chen |
Decis. Support Syst. | 2 |
| 2005 | Evaluating event visualization: a usability study of COPLINK spatio-temporal visualizer
Wingyan Chung, Hsinchun Chen, Luis G. Chaboya, Christopher D. O'Toole, Homa Atabakhsh |
Int. J. Hum. Comput. Stud. | 2 |
| 2005 | Introduction to the special topic issue: Intelligence and security informatics
Hsinchun Chen |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2005 | User acceptance of Intelligence and Security Informatics technology: A study of COPLINKabstractAbstract The importance of Intelligence and Security Informatics (ISI) has significantly increased with the rapid and large‐scale migration of local/national security information from physical media to electronic platforms, including the Internet and information systems. Motivated by the significance of ISI in law enforcement (particularly in the digital government context) and the limited investigations of officers' technology‐acceptance decision‐making, we developed and empirically tested a factor model for explaining law‐enforcement officers' technology acceptance. Specifically, our empirical examination targeted the COPLINK technology and involved more than 280 police officers. Overall, our model shows a good fit to the data collected and exhibits satisfactory power for explaining law‐enforcement officers' technology acceptance decisions. Our findings have several implications for research and technology management practices in law enforcement, which are also discussed. Paul Jen-Hwa Hu, Chienting Lin, Hsinchun Chen |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2005 | Genescene: An ontology-enhanced integration of linguistic and co-occurrence based relations in biomedical textsabstractAbstract The increasing amount of publicly available literature and experimental data in biomedicine makes it hard for biomedical researchers to stay up‐to‐date. Genescene is a toolkit that will help alleviate this problem by providing an overview of published literature content. We combined a linguistic parser with Concept Space, a co‐occurrence based semantic net. Both techniques extract complementary biomedical relations between noun phrases from MEDLINE abstracts. The parser extracts precise and semantically rich relations from individual abstracts. Concept Space extracts relations that hold true for the collection of abstracts. The Gene Ontology, the Human Genome Nomenclature, and the Unified Medical Language System, are also integrated in Genescene. Currently, they are used to facilitate the integration of the two relation types, and to select the more interesting and high‐quality relations for presentation. A user study focusing on p53 literature is discussed. All MEDLINE abstracts discussing p53 were processed in Genescene. Two researchers evaluated the terms and relations from several abstracts of interest to them. The results show that the terms were precise (precision 93%) and relevant, as were the parser relations (precision 95%). The Concept Space relations were more precise when selected with ontological knowledge (precision 78%) than without (60%). Gondy Leroy, Hsinchun Chen |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2005 | CrimeNet explorer: a framework for criminal network knowledge discoveryabstractKnowledge about the structure and organization of criminal networks is important for both crime investigation and the development of effective strategies to prevent crimes. However, except for network visualization, criminal network analysis remains primarily a manual process. Existing tools do not provide advanced structural analysis techniques that allow extraction of network knowledge from large volumes of criminal-justice data. To help law enforcement and intelligence agencies discover criminal network knowledge efficiently and effectively, in this research we proposed a framework for automated network analysis and visualization. The framework included four stages: network creation, network partition, structural analysis, and network visualization. Based upon it, we have developed a system called CrimeNet Explorer that incorporates several advanced techniques: a concept space approach, hierarchical clustering, social network analysis methods, and multidimensional scaling. Results from controlled experiments involving student subjects demonstrated that our system could achieve higher clustering recall and precision than did untrained subjects when detecting subgroups from criminal networks. Moreover, subjects identified central members and interaction patterns between groups significantly faster with the help of structural analysis functionality than with only visualization functionality. No significant gain in effectiveness was present, however. Our domain experts also reported that they believed CrimeNet Explorer could be very useful in crime investigation. Jennifer Jie Xu 0001, Hsinchun Chen |
ACM Trans. Inf. Syst. | 2 |
| 2004 | Information Sharing and Collaboration Policies within Government Agencies
Homa Atabakhsh, Cathy Larson, Tim Petersen, Chuck Violette, Hsinchun Chen |
ISI | 5 |
| 2004 | Terrorism Knowledge Discovery Project: A Knowledge Discovery Approach to Addressing the Threats of Terrorism
Edna Reid, Jialun Qin, Wingyan Chung, Jennifer Jie Xu 0001, Yilu Zhou, Robert P. Schumaker, Marc Sageman, Hsinchun Chen |
ISI | 8 |
| 2004 | Analyzing and Visualizing Criminal Network Dynamics: A Case Study
Jennifer Jie Xu 0001, Byron Marshall, Siddharth Kaza, Hsinchun Chen |
ISI | 4 |
| 2004 | West Nile Virus and Botulism Portal: A Case Study in Infectious Disease Informatics
Daniel Dajun Zeng, Hsinchun Chen, Chunju Tseng, Catherine A. Larson, Millicent Eidson, Ivan Gotham, Cecil Lynch, Michael Ascher |
ISI | 2 |
| 2004 | Extracting gene pathway relations using a hybrid grammar: the Arizona Relation ParserabstractMOTIVATION: Text-mining research in the biomedical domain has been motivated by the rapid growth of new research findings. Improving the accessibility of findings has potential to speed hypothesis generation. RESULTS: We present the Arizona Relation Parser that differs from other parsers in its use of a broad coverage syntax-semantic hybrid grammar. While syntax grammars have generally been tested over more documents, semantic grammars have outperformed them in precision and recall. We combined access to syntax and semantic information from a single grammar. The parser was trained using 40 PubMed abstracts and then tested using 100 unseen abstracts, half for precision and half for recall. Expert evaluation showed that the parser extracted biologically relevant relations with 89% precision. Recall of expert identified relations with semantic filtering was 35 and 61% before semantic filtering. Such results approach the higher-performing semantic parsers. However, the AZ parser was tested over a greater variety of writing styles and semantic content. AVAILABILITY: Relations extracted from over 600 000 PubMed abstracts are available for retrieval and visualization at http://econport.arizona.edu:8080/NetVis/index.html. Daniel McDonald, Hsinchun Chen, Byron Marshall |
Bioinform. | 2 |
| 2004 | Credit rating analysis with support vector machines and neural networks: a market comparative study
Zan Huang, Hsinchun Chen, Chia-Jung Hsu, Wun-Hwa Chen, Soushan Wu |
Decis. Support Syst. | 2 |
| 2004 | Fighting organized crimes: using shortest-path algorithms to identify associations in criminal networks
Jennifer Jie Xu 0001, Hsinchun Chen |
Decis. Support Syst. | 2 |
| 2004 | Internet searching and browsing in a multilingual world: An experiment on the Chinese Business Intelligence Portal (CBizPort)abstractAbstract The rapid growth of the non‐English‐speaking Internet population has created a need for better searching and browsing capabilities in languages other than English. However, existing search engines may not serve the needs of many non‐English‐speaking Internet users. In this paper, we propose a generic and integrated approach to searching and browsing the Internet in a multilingual world. Based on this approach, we have developed the Chinese Business Intelligence Portal (CBizPort), a meta‐search engine that searches for business information of mainland China, Taiwan, and Hong Kong. Additional functions provided by CBizPort include encoding conversion (between Simplified Chinese and Traditional Chinese), summarization, and categorization. Experimental results of our user evaluation study show that the searching and browsing performance of CBizPort was comparable to that of regional Chinese search engines, and CBizPort could significantly augment these search engines. Subjects' verbal comments indicate that CBizPort performed best in terms of analysis functions, cross‐regional searching, and user‐friendliness, whereas regional search engines were more efficient and more popular. Subjects especially liked CBizPort's summarizer and categorizer, which helped in understanding search results. These encouraging results suggest a promising future of our approach to Internet searching and browsing in a multilingual world. Wingyan Chung, Zan Huang, Gang Wang 0011, Thian-Huat Ong, Hsinchun Chen |
J. Assoc. Inf. Sci. Technol. | 6 |
| 2004 | A graph model for E-commerce recommender systemsabstractAbstract Information overload on the Web has created enormous challenges to customers selecting products for online purchases and to online businesses attempting to identify customers' preferences efficiently. Various recommender systems employing different data representations and recommendation methods are currently used to address these challenges. In this research, we developed a graph model that provides a generic data representation and can support different recommendation methods. To demonstrate its usefulness and flexibility, we developed three recommendation methods: direct retrieval, association mining, and high‐degree association retrieval. We used a data set from an online bookstore as our research test‐bed. Evaluation results showed that combining product content information and historical customer transaction information achieved more accurate predictions and relevant recommendations than using only collaborative information. However, comparisons among different methods showed that high‐degree association retrieval did not perform significantly better than the association mining method or the direct retrieval method in our test‐bed. Zan Huang, Wingyan Chung, Hsinchun Chen |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2004 | EBizPort: Collecting and analyzing business intelligence informationabstractAbstract To make good decisions, businesses try to gather good intelligence information. Yet managing and processing a large amount of unstructured information and data stand in the way of greater business knowledge. An effective business intelligence tool must be able to access quality information from a variety of sources in a variety of forms, and it must support people as they search for and analyze that information. The EBizPort system was designed to address information needs for the business/IT community. EBizPort's collection‐building process is designed to acquire credible, timely, and relevant information. The user interface provides access to collected and metasearched resources using innovative tools for summarization, categorization, and visualization. The effectiveness, efficiency, usability, and information quality of the EBizPort system were measured. EBizPort significantly outperformed Brint, a business search portal, in search effectiveness, information quality, user satisfaction, and usability. Users particularly liked EBizPort's clean and user‐friendly interface. Results from our evaluation study suggest that the visualization function added value to the search and analysis process, that the generalizable collection‐building technique can be useful for domain‐specific information searching on the Web, and that the search interface was important for Web search and browse support. Byron Marshall, Daniel McDonald, Hsinchun Chen, Wingyan Chung |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2004 | Intelligence and security informatics for homeland security: information, communication, and transportationabstractIntelligence and security informatics (ISI) is an emerging field of study aimed at developing advanced information technologies, systems, algorithms, and databases for national- and homeland-security-related applications, through an integrated technological, organizational, and policy-based approach. This paper summarizes the broad application and policy context for this emerging field. Three detailed case studies are presented to illustrate several key ISI research areas, including cross-jurisdiction information sharing; terrorism information collection, analysis, and visualization; and "smart-border" and bioterrorism applications. A specific emphasis of this paper is to note various homeland-security-related applications that have direct relevance to transportation researchers and to advocate security informatics studies that tightly integrate transportation research and information technologies. Hsinchun Chen, Fei-Yue Wang 0001, Daniel Dajun Zeng |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2004 | Applying associative retrieval techniques to alleviate the sparsity problem in collaborative filteringabstractRecommender systems are being widely applied in many application settings to suggest products, services, and information items to potential consumers. Collaborative filtering, the most successful recommendation approach, makes recommendations based on past transactions and feedback from consumers sharing similar interests. A major problem limiting the usefulness of collaborative filtering is the sparsity problem, which refers to a situation in which transactional or feedback data is sparse and insufficient to identify similarities in consumer interests. In this article, we propose to deal with this sparsity problem by applying an associative retrieval framework and related spreading activation algorithms to explore transitive associations among consumers through their past transactions and feedback. Such transitive associations are a valuable source of information to help infer consumer interests and can be explored to deal with the sparsity problem. To evaluate the effectiveness of our approach, we have conducted an experimental study using a data set from an online bookstore. We experimented with three spreading activation algorithms including a constrained Leaky Capacitor algorithm, a branch-and-bound serial symbolic search algorithm, and a Hopfield net parallel relaxation search algorithm. These algorithms were compared with several collaborative filtering approaches that do not consider the transitive associations: a simple graph search approach, two variations of the user-based approach, and an item-based approach. Our experimental results indicate that spreading activation-based approaches significantly outperformed the other collaborative filtering methods as measured by recommendation precision, recall, the F-measure, and the rank score. We also observed the over-activation effect of the spreading activation approach, that is, incorporating transitive associations with past transactional data that is not sparse may "dilute" the data used to infer user preferences and lead to degradation in recommendation performance. Zan Huang, Hsinchun Chen, Daniel Dajun Zeng |
ACM Trans. Inf. Syst. | 2 |
| 2003 | A Spatio Temporal Visualizer for Law Enforcement
Ty Buetow, Luis G. Chaboya, Christopher D. O'Toole, Tom Cushna, Damien Daspit, Tim Petersen, Homa Atabakhsh, Hsinchun Chen |
ISI | 8 |
| 2003 | An International Perspective on Fighting Cybercrime
Weiping Chang, Wingyan Chung, Hsinchun Chen, Shihchieh Chou |
ISI | 3 |
| 2003 | Examining Technology Acceptance by Individual Law Enforcement Officers: An Exploratory Study
Paul Jen-Hwa Hu, Chienting Lin, Hsinchun Chen |
ISI | 3 |
| 2003 | CrimeLink Explorer: Using Domain Knowledge to Facilitate Automated Crime Association Analysis
Jennifer Schroeder, Jennifer Jie Xu 0001, Hsinchun Chen |
ISI | 3 |
| 2003 | Untangling Criminal Networks: A Case Study
Jennifer Jie Xu 0001, Hsinchun Chen |
ISI | 2 |
| 2003 | COPLINK Agent: An Architecture for Information Monitoring and Sharing in Law Enforcement
Daniel Dajun Zeng, Hsinchun Chen, Damien Daspit, Fu Shan, Suresh Nandiraju, Michael Chau, Chienting Lin |
ISI | 2 |
| 2003 | Collaborative Workflow Management for Interagency Crime Analysis
J. Leon Zhao, Henry H. Bi, Hsinchun Chen |
ISI | 3 |
| 2003 | Authorship Analysis in Cybercrime Investigation
Zan Huang, Hsinchun Chen |
ISI | 4 |
| 2003 | Design and evaluation of a multi-agent collaborative Web mining system
Michael Chau, Daniel Dajun Zeng, Hsinchun Chen, David Hendriawan |
Decis. Support Syst. | 3 |
| 2003 | Special issue: "Web retrieval and mining"
Hsinchun Chen |
Decis. Support Syst. | 1 |
| 2003 | Digital Government: technologies and practices
Hsinchun Chen |
Decis. Support Syst. | 1 |
| 2003 | COPLINK Connect: information and knowledge management for law enforcement
Hsinchun Chen, Jennifer Schroeder, Roslin V. Hauck, Linda Ridgeway, Homa Atabakhsh, Chris Boarman, Kevin Rasmussen, Andy W. Clements |
Decis. Support Syst. | 1 |
| 2003 | Visualization of large category map for Internet browsing
Christopher C. Yang, Hsinchun Chen, Kay Hong |
Decis. Support Syst. | 2 |
| 2003 | Testing a Cancer Meta Spider
Hsinchun Chen, Haiyan Fan, Michael Chau, Daniel Dajun Zeng |
Int. J. Hum. Comput. Stud. | 1 |
| 2003 | Introduction to the JASIST Special Topic issue on web retrieval and mining: A machine learning perspective
Hsinchun Chen |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2003 | elpfulMed: Intelligent searching for medical information over the internetabstractAbstract Medical professionals and researchers need information from reputable sources to accomplish their work. Unfortunately, the Web has a large number of documents that are irrelevant to their work, even those documents that purport to be “medically‐related.” This paper describes an architecture designed to integrate advanced searching and indexing algorithms, an automatic thesaurus, or “concept space,” and Kohonen‐based Self‐Organizing Map (SOM) technologies to provide searchers with fine‐grained results. Initial results indicate that these systems provide complementary retrieval functionalities. HelpfulMed not only allows users to search Web pages and other online databases, but also allows them to build searches through the use of an automatic thesaurus and browse a graphical display of medical‐related topics. Evaluation results for each of the different components are included. Our spidering algorithm outperformed both breadth‐first search and PageRank spiders on a test collection of 100,000 Web pages. The automatically generated thesaurus performed as well as both MeSH and UMLS—systems which require human mediation for currency. Lastly, a variant of the Kohonen SOM was comparable to MeSH terms in perceived cluster precision and significantly better at perceived cluster recall. Hsinchun Chen, Ann M. Lally, Bin Zhu 0001, Michael Chau |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2003 | A shallow parser based on closed-class words to capture relations in biomedical text
Gondy Leroy, Hsinchun Chen, Jesse D. Martinez |
J. Biomed. Informatics | 2 |
| 2003 | Teaching key topics in computer science and information systems through a web search engine projectabstractAdvances in computer and Internet technologies have made it more and more important for information technology professionals to acquire experience in a variety of aspects, including new technologies, system integration, database administration, and project management. To provide students with a chance to acquire such skills, we designed a project called "Build Your Search Engine in 90 Days," in which students were required to build a domain-specific Web search engine in a semester. In this paper we review the tools and resources available to students and report our experiences in having students to work on this project in a course at the University of Arizona. We also review two tools, called AI Spider and AI Indexer, we developed for students in this project. We highlight a few search engines that were created by the students and suggest some future directions in improving the tools and expanding the project. Michael Chau, Zan Huang, Hsinchun Chen |
ACM J. Educ. Resour. Comput. | 3 |
| 2003 | The use of dynamic context to improve casual internet searchingabstractResearch has shown that most users' online information searches are suboptimal. Query optimization based on a relevance feedback or genetic algorithm using dynamic query contexts can help casual users search the Internet. These algorithms can draw on implicit user feedback based on the surrounding links and text in a search engine result set to expand user queries with a variable number of keywords in two manners. Positive expansion adds terms to a user's keywords with a Boolean "and," negative expansion adds terms to the user's keywords with a Boolean "not." Each algorithm was examined for three user groups, high, middle, and low achievers, who were classified according to their overall performance. The interactions of users with different levels of expertise with different expansion types or algorithms were evaluated. The genetic algorithm with negative expansion tripled recall and doubled precision for low achievers, but high achievers displayed an opposed trend and seemed to be hindered in this condition. The effect of other conditions was less substantial. Gondy Leroy, Ann M. Lally, Hsinchun Chen |
ACM Trans. Inf. Syst. | 3 |
| 2002 | CI Spider: a tool for competitive intelligence on the Web
Hsinchun Chen, Michael Chau, Daniel Dajun Zeng |
Decis. Support Syst. | 1 |
| 2001 | Information navigation on the web by clustering and summarizing query results
Dmitri Roussinov, Hsinchun Chen |
Inf. Process. Manag. | 2 |
| 2001 | MetaSpider: Meta-searching and categorization on the WebabstractAbstract It has become increasingly difficult to locate relevant information on the Web, even with the help of Web search engines. Two approaches to addressing the low precision and poor presentation of search results of current search tools are studied: meta‐search and document categorization. Meta‐search engines improve precision by selecting and integrating search results from generic or domain‐specific Web search engines or other resources. Document categorization promises better organization and presentation of retrieved results. This article introduces MetaSpider, a meta‐search engine that has real‐time indexing and categorizing functions. We report in this paper the major components of MetaSpider and discuss related technical approaches. Initial results of a user evaluation study comparing MetaSpider, NorthernLight, and MetaCrawler in terms of clustering performance and of time and effort expended show that MetaSpider performed best in precision rate, but disclose no statistically significant differences in recall rate and time requirements. Our experimental study also reveals that MetaSpider exhibited a higher level of automation than the other two systems and facilitated efficient searching by providing the user with an organized, comprehensive view of the retrieved documents. Hsinchun Chen, Haiyan Fan, Michael Chau, Daniel Dajun Zeng |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2001 | Meeting medical terminology needs-the ontology-enhanced Medical Concept MapperabstractThis paper describes the development and testing of the Medical Concept Mapper, a tool designed to facilitate access to online medical information sources by providing users with appropriate medical search terms for their personal queries. Our system is valuable for patients whose knowledge of medical vocabularies is inadequate to find the desired information, and for medical experts who search for information outside their field of expertise. The Medical Concept Mapper maps synonyms and semantically related concepts to a user's query. The system is unique because it integrates our natural language processing tool, i.e., the Arizona (AZ) Noun Phraser, with human-created ontologies, the Unified Medical Language System (UMLS) and WordNet, and our computer generated Concept Space, into one system. Our unique contribution results from combining the UMLS Semantic Net with Concept Space in our deep semantic parsing (DSP) algorithm. This algorithm establishes a medical query context based on the UMLS Semantic Net, which allows Concept Space terms to be filtered so as to isolate related terms relevant to the query. We performed two user studies in which Medical Concept Mapper terms were compared against human experts' terms. We conclude that the AZ Noun Phraser is well suited to extract medical phrases from user queries, that WordNet is not well suited to provide strictly medical synonyms, that the UMLS Metathesaurus is well suited to provide medical synonyms, and that Concept Space is well suited to provide related medical terms, especially when these terms are limited by our DSP algorithm. Gondy Leroy, Hsinchun Chen |
IEEE Trans. Inf. Technol. Biomed. | 2 |
| 2000 | Exploring the use of concept spaces to improve medical information retrieval
Andrea Houston, Hsinchun Chen, Bruce R. Schatz, Susan Molloy Hubbard, Robin R. Sewell, Tobun Dorbin Ng |
Decis. Support Syst. | 2 |
| 2000 | Estimating drug/plasma concentration levels by applying neural networks to pharmacokinetic data sets
Kristin M. Tolle, Hsinchun Chen, Hsiao-Hui Chow |
Decis. Support Syst. | 2 |
| 2000 | Intelligent internet searching agent based on hybrid simulated annealing
Christopher C. Yang, Jerome Yen, Hsinchun Chen |
Decis. Support Syst. | 3 |
| 2000 | Introduction to the special topic issue: Digital Libraries - Part 1
Hsinchun Chen |
J. Am. Soc. Inf. Sci. | 1 |
| 2000 | Introduction to the special topic issue: Digital Libraries - Part 2
Hsinchun Chen |
J. Am. Soc. Inf. Sci. | 1 |
| 2000 | Comparing noun phrasing techniques for use with medical digital library toolsabstractIn an effort to assist medical researchers and professionals in accessing information necessary for their work, the A1 Lab at the University of Arizona is investigating the use of a natural language processing (NLP) technique called noun phrasing. The goal of this research is to determine whether noun phrasing could be a viable technique to include in medical information retrieval applications. Four noun phrase generation tools were evaluated as to their ability to isolate noun phrases from medical journal abstracts. Tests were conducted using the National Cancer Institute's CANCERLIT database. The NLP tools evaluated were Massachusetts Institute of Technology's (MIT's) Chopper, The University of Arizona's Automatic Indexer, Lingsoft's NPtool, and The University of Arizona's AZ Noun Phraser. In addition, the National Library of Medicine's SPECIALIST Lexicon was incorporated into two versions of the AZ Noun Phraser to be evaluated against the other tools as well as a nonaugmented version of the AZ Noun Phraser. Using the metrics relative subject recall and precision, our results show that, with the exception of Chopper, the phrasing tools were fairly comparable in recall and precision. It was also shown that augmenting the AZ Noun Phraser by including the SPECIALIST Lexicon from the National Library of Medicine resulted in improved recall and precision. Kristin M. Tolle, Hsinchun Chen |
J. Am. Soc. Inf. Sci. | 2 |
| 2000 | Validating a geographical image retrieval systemabstractThis paper summarizes a prototype geographical image retrieval system that demonstrates how to integrate image processing and information analysis techniques to support large-scale content-based image retrieval. By using an image as its interface, the prototype system addresses a troublesome aspect of traditional retrieval models, which require users to have complete knowledge of the low-level features of an image. In addition we describe an experiment to validate the performance of this image retrieval system against that of human subjects in an effort to address the scarcity of research evaluating performance of an algorithm against that of human beings. The results of the experiment indicate that the system could do as well as human subjects in accomplishing the tasks of similarity analysis and image categorization. We also found that under some circumstances texture features of an image are insufficient to represent a geographic image. We believe, however, that our image retrieval system provides a promising approach to integrating image processing techniques and information retrieval algorithms. Bin Zhu 0001, Hsinchun Chen |
J. Am. Soc. Inf. Sci. | 2 |
| 2000 | Creating a large-scale content-based airphoto image digital libraryabstractThis paper describes a content-based image retrieval digital library that supports geographical image retrieval over a testbed of 800 aerial photographs, each 25 megabytes in size. In addition, this paper also introduces a methodology to evaluate the performance of the algorithms in the prototype system. There are two major contributions: we suggest an approach that incorporates various image processing techniques including Gabor filters, image enhancement and image compression, as well as information analysis techniques such as the self-organizing map (SOM) into an effective large-scale geographical image retrieval system. We present two experiments that evaluate the performance of the Gabor-filter-extracted features along with the corresponding similarity measure against that of human perception, addressing the lack of studies in assessing the consistency between an image representation algorithm or an image categorization method and human mental model. Bin Zhu 0001, Marshall Ramsey, Hsinchun Chen |
IEEE Trans. Image Process. | 3 |
| 1999 | Interactive Internet Search through Automatic Clustering (poster abstract): an empirical studyabstractNo abstract available. Dmitri Roussinov, Kristin M. Tolle, Marshall Ramsey, Hsinchun Chen |
SIGIR | 4 |
| 1999 | Visualizing Internet Search Results with Adaptive Self-Organizing Maps (demonstration abstract)abstractNo abstract available. Dmitri Roussinov, Kristin M. Tolle, Marshall Ramsey, Michael J. McQuaid, Hsinchun Chen |
SIGIR | 5 |
| 1999 | Multidimensional scaling for group memory visualization
Michael J. McQuaid, Thian-Huat Ong, Hsinchun Chen, Jay F. Nunamaker Jr. |
Decis. Support Syst. | 3 |
| 1999 | Document clustering for electronic meetings: an experimental comparison of two techniquesabstractIn this article, we report our implementation and comparison of two text clustering techniques. One is based on Ward's clustering and the other on Kohonen's Self-organizing Maps. We have evaluated how closely clusters produced by a computer resemble those created by human experts. We have also measured the time that it takes for an expert to “clean up” the automatically produced clusters. The technique based on Ward's clustering was found to be more precise. Both techniques have worked equally well in detecting associations between text documents. We used text messages obtained from group brainstorming meetings. Dmitri Roussinov, Hsinchun Chen |
Decis. Support Syst. | 2 |
| 1999 | A Collection of Visual Thesauri for Browsing Large Collections of Geographic ImagesabstractDigital libraries of geo-spatial multimedia content are currently deficient in providing fuzzy, concept-based retrieval mechanisms to users. The main challenge is that indexing and thesaurus creation are extremely labor-intensive processes for text documents and especially for images. Recently, 800,000 declassified satellite photographs were made available by the United States Geological Survey. Additionally, millions of satellite and aerial photographs are archived in national and local map libraries. Such enormous collections make human indexing and thesaurus generation methods impossible to utilize. In this article we propose a scalable method to automatically generate visual thesauri of large collections of geo-spatial media using fuzzy, unsupervised machine-learning techniques. Marshall Ramsey, Hsinchun Chen, Bin Zhu 0001, Bruce R. Schatz |
J. Am. Soc. Inf. Sci. | 2 |
| 1998 | An intelligent personal spider (agent) for dynamic Internet/Intranet searching
Hsinchun Chen, Yi-Ming Chung, Marshall Ramsey, Christopher C. Yang |
Decis. Support Syst. | 1 |
| 1998 | Introduction to the Special Topic Issue: Artificial Intelligence Techniques for Emerging Information Systems Applications
Hsinchun Chen |
J. Am. Soc. Inf. Sci. | 1 |
| 1998 | A Smart Itsy Bitsy Spider for the WebabstractAs part of the ongoing Illinois Digital Library Initiative project, this research proposes an intelligent agent approach to Web searching. In this experiment, we developed two Web personal spiders based on best first search and genetic algorithm techniques, respectively. These personal spiders can dynamically take a user's selected starting homepages and search for the most closely related homepages in the Web, based on the links and keyword indexing. A graphical, dynamic, Java-based interface was developed and is available for Web access. A system architecture for implementing such an agent-based spider is presented, followed by detailed discussions of benchmark testing and user evaluation results. In benchmark testing, although the genetic algorithm spider did not outperform the best first search spider, we found both results to be comparable and complementary. In user evaluation, the genetic algorithm spider obtained significantly higher recall value than that of the best first search spider. However, their precision values were not statistically different. The mutation process introduced in genetic algorithm allows users to find other potential relevant homepages that cannot be explored via a conventional local search process. In addition, we found the Java-based interface to be a necessary component for design of a truly interactive and dynamic Web agent. © 1998 John Wiley & Sons, Inc. Hsinchun Chen, Yi-Ming Chung, Marshall Ramsey, Christopher C. Yang |
J. Am. Soc. Inf. Sci. | 1 |
| 1998 | Internet Browsing and Searching: User Evaluations of Category Map and Concept Space TechniquesabstractThe Internet provides an exceptional testbed for developing algorithms that can improve browsing and searching large information spaces. Browsing and searching tasks are susceptible to problems of information overload and vocabulary differences. Much of the current research is aimed at the development and refinement of algorithms to improve browsing and searching by addressing these problems. Our research was focused on discovering whether two of the algorithms our research group has developed, a Kohonen algorithm category map for browsing, and an automatically generated concept space algorithm for searching, can help improve browsing and/or searching the Internet. Our results indicate that a Kohonen self-organizing map (SOM)-based algorithm can successfully categorize a large and eclectic Internet information space (the Entertainment sub-category of Yahoo!) into manageable sub-spaces that users can successfully navigate to locate a homepage of interest to them. The SOM algorithm worked best with browsing tasks that were very broad, and in which subjects skipped around between categories. Subjects especially liked the visual and graphical aspects of the map. Subjects who tried to do a directed search, and those that wanted to use the more familiar mental models (alphabetic or hierarchical organization) for browsing, found that the map did not work well. The results from the concept space experiment were especially encouraging. There were no significant differences among the precision measures for the set of documents identified by subject-suggested terms, thesaurus-suggested terms, and the combination of subject- and thesaurus-suggested terms. The recall measures indicated that the combination of subject- and thesaurus-suggested terms exhibited significantly better recall than subject-suggested terms alone. Furthermore, analysis of the home pages indicated that there was limited overlap between the homepages retrieved by the subject-suggested and thesaurus-suggested terms. Since the retrieved homepages for the most part were different, this suggests that a user can enhance a keyword-based search by using an automatically generated concept space. Subjects especially liked the level of control that they could exert over the search, and the fact that the terms suggested by the thesaurus were “real” (i.e., originating in the homepages) and therefore guaranteed to have retrieval success. © 1998 John Wiley & Sons, Inc. Hsinchun Chen, Andrea Houston, Robin R. Sewell, Bruce R. Schatz |
J. Am. Soc. Inf. Sci. | 1 |
| 1998 | Alleviating Search Uncertainty Through Concept Associations: Automatic Indexing, Co-Occurrence Analysis, and Parallel ComputingabstractIn this article, we report research on an algorithmic approach to alleviating search uncertainty in a large information space. Grounded on object filtering, automatic indexing, and co-occurrence analysis, we performed a large-scale experiment using a parallel supercomputer (SGI Power Challenge) to analyze 400,000+ abstracts in an INSPEC computer engineering collection. Two system-generated thesauri, one based on a combined object filtering and automatic indexing method, and the other based on automatic indexing only, were compared with the human-generated INSPEC subject thesaurus. Our user evaluation revealed that the system-generated thesauri were better than the INSPEC thesaurus in concept recall, but in concept precision the 3 thesauri were comparable. Our analysis also revealed that the terms suggested by the 3 thesauri were complementary and could be used to significantly increase “variety” in search terms and thereby reduce search uncertainty. © 1998 John Wiley & Sons, Inc. Hsinchun Chen, Joanne Martinez, Amy Kirchhoff, Tobun Dorbin Ng, Bruce R. Schatz |
J. Am. Soc. Inf. Sci. | 1 |
| 1998 | A Machine Learning Approach to Inductive Query by Examples: An Experiment Using Relevance Feedback, ID3, Genetic Algorithms, and Simulated AnnealingabstractInformation retrieval using probabilistic techniques has attracted significant attention on the part of researchers in information and computer science over the past few decades. In the 1980s, knowledge-based techniques also made an impressive contribution to “intelligent” information retrieval and indexing. More recently, information science researchers have turned to other newer inductive learning techniques including symbolic learning, genetic algorithms, and simulated annealing. These newer techniques, which are grounded in diverse paradigms, have provided great opportunities for researchers to enhance the information processing and retrieval capabilities of current information systems. In this article, we first provide an overview of these newer techniques and their use in information retrieval research. In order to familiarize readers with the techniques, we present three promising methods: The symbolic ID3 algorithm, evolution-based genetic algorithms, and simulated annealing. We discuss their knowledge representations and algorithms in the unique context of information retrieval. An experiment using a 8000-record COMPEN database was performed to examine the performances of these inductive query-by-example techniques in comparison with the performance of the conventional relevance feedback method. The machine learning techniques were shown to be able to help identify new documents which are similar to documents initially suggested by users, and documents which contain similar concepts to each other. Genetic algorithms, in particular, were found to out-perform relevance feedback in both document recall and precision. We believe these inductive machine learning techniques hold promise for the ability to analyze users' preferred documents (or records), identify users' underlying information needs, and also suggest alternatives for search for database management systems and Internet applications. © 1998 John Wiley & Sons, Inc. Hsinchun Chen, G. Shankaranarayanan, Linlin She, Anand Iyer |
J. Am. Soc. Inf. Sci. | 1 |
| 1997 | Semantic Search and Semantic Categorization (Abstract)abstractNo abstract available. Hsinchun Chen, Andrea Houston, Robin R. Sewell, Bruce R. Schatz |
SIGIR | 1 |
| 1997 | Visual SOM (Abstract)abstractNo abstract available. Hsinchun Chen, Marshall Ramsey, Terence R. Smith |
SIGIR | 1 |
| 1997 | A Concept Space Approach to Addressing the Vocabulary Problem in Scientific Information Retrieval: An Experiment on the Worm Community SystemabstractThis research presents an algorithmic approach to addressing the vocabulary problem in scientific information retrieval and information sharing, using the molecular biology domain as an example. We first present a literature review of cognitive studies related to the vocabulary problem and vocabulary-based search aids (thesauri) and then discuss techniques for building robust and domain-specific thesauri to assist in cross-domain scientific information retrieval. Using a variation of the automatic thesaurus generation techniques, which we refer to as the concept space approach, we recently conducted an experiment in the molecular biology domain in which we created a C. elegans worm thesaurus of 7,657 worm-specific terms and a Drosophila fly thesaurus of 15,626 terms. About 30% of these terms overlapped, which created vocabulary paths from one subject domain to the other. Based on a cognitive study of term association involving four biologists, we found that a large percentage (59.6–85.6%) of the terms suggested by the subjects were identified in the conjoined fly-worm thesaurus. However, we found only a small percentage (8.4–18.1%) of the associations suggested by the subjects in the thesaurus. In a follow-up document retrieval study involving eight fly biologists, an actual worm database (Worm Community System), and the conjoined fly-worm thesaurus, subjects were able to find more relevant documents (an increase from about 9 documents to 20) and to improve the document recall level (from 32.41 to 65.28%) when using the thesaurus, although the precision level did not improve significantly. Implications of adopting the concept space approach for addressing the vocabulary problem in internet and digital libraries applications are also discussed. Hsinchun Chen, Tobun Dorbin Ng, Joanne Martinez, Bruce R. Schatz |
J. Am. Soc. Inf. Sci. | 1 |
| 1997 | A Graphical, Self-Organizing Approach to Classifying Electronic Meeting OutputabstractThis article describes research in the application of a Kohonen Self-Organizing Map (SOM) to the problem of classification of electronic brainstorming output and an evaluation of the results. Electronic brainstorming is one of the most productive tools in the Electronic Meeting System called GroupSystems. A major step in group problem solving involves the classification of electronic brainstorming output into a manageable list of concepts, topics, or issues that can be further evaluated by the group. This step is problematic due to information overload and the cognitive demand of processing a large quantity of textual data. This research builds upon previous work in automating the meeting classification process using a Hopfield neural network. Evaluation of the Kohonen output comparing it with Hopfield and human expert output using the same set of data found that the Kohonen SOM performed as well as a human expert in representing term association in the meeting output and outperformed the Hopfield neural network algorithm. In addition, recall of consensus meeting concepts and topics using the Kohonen algorithm was equivalent to that of the human expert. However, precision of the Kohonen results was poor. The graphical representation of textual data produced by the Kohonen SOM suggests many opportunities for improving information organization of textual information. Increasing uses of electronic mail, computer-based bulletin board systems, and world-wide web services present unique challenges and opportunities for a system-aided classification approach. This research has shown that the Kohonen SOM may be used to automatically create “a picture that can represent a thousand (or more) words.” © 1997 John Wiley & Sons, Inc. Richard E. Orwig, Hsinchun Chen, Jay F. Nunamaker Jr. |
J. Am. Soc. Inf. Sci. | 2 |
| 1996 | Internet Categorization and Search: A Self-Organizing Approach
Hsinchun Chen, Chris Schuffels, Richard E. Orwig |
J. Vis. Commun. Image Represent. | 1 |
| 1996 | A Parallel Computing Approach to Creating Engineering Concept Spaces for Semantic Retrieval: The Illinois Digital Library Initiative ProjectabstractThis research presents preliminary results generated from the semantic retrieval research component of the Illinois Digital Library Initiative (DLI) project. Using a variation of the automatic thesaurus generation techniques, to which we refer to as the concept space approach, we aimed to create graphs of domain-specific concepts (terms) and their weighted co-occurrence relationships for all major engineering domains. Merging these concept spaces and providing traversal paths across different concept spaces could potentially help alleviate the vocabulary (difference) problem evident in large-scale information retrieval. In order to address the scalability issue related to large-scale information retrieval and analysis for the current Illinois DLI project, we conducted experiments using the concept space approach on parallel supercomputers. Our test collection included computer science and electrical engineering abstracts extracted from the INSPEC database. The concept space approach called for extensive textual and statistical analysis (a form of knowledge discovery) based on automatic indexing and co-occurrence analysis algorithms, both previously tested in the biology domain. Initial testing results using a 512-node CM-5 and a 16-processor SGI Power Challenge were promising. Hsinchun Chen, Bruce R. Schatz, Tobun Dorbin Ng, Joanne Martinez, Amy Kirchhoff, Chienting Lin |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1996 | An automatic indexing and neural network approach to concept retrieval and classification of multilingual (Chinese-English) documentsabstractAn automatic indexing and concept classification approach to a multilingual (Chinese and English) bibliographic database is presented. We introduced a multi-linear term-phrasing technique to extract concept descriptors (terms or keywords) from a Chinese-English bibliographic database. A concept space of related descriptors was then generated using a co-occurrence analysis technique. Like a man-made thesaurus, the system-generated concept space can be used to generate additional semantically-relevant terms for search. For concept classification and clustering, a variant of a Hopfield neural network was developed to cluster similar concept descriptors and to generate a small number of concept groups to represent (summarize) the subject matter of the database. The concept space approach to information classification and retrieval has been adopted by the authors in other scientific databases and business applications, but multilingual information retrieval presents a unique challenge. This research reports our experiment on multilingual databases. Our system was initially developed in the MS-DOS environment, running ETEN Chinese operating system. For performance reasons, it was then tested on a UNIX-based system. Due to the unique ideographic nature of the Chinese language, a Chinese term-phrase indexing paradigm considering the ideographic characteristics of Chinese was developed as a multilingual information classification model. By applying the neural network based concept classification technique, the model presents a novel way of organizing unstructured multilingual information. Chung-Hsin Lin, Hsinchun Chen |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 1995 | Machine Learning for Information Retrieval: Neural Networks, Symbolic Learning, and Genetic AlgorithmsabstractInformation retrieval using probabilistic techniques has attracted significant attention on the part of researchers in information and computer science over the past few decades. In the 1980s, knowledge-based techniques also made an impressive contribution to “intelligent” information retrieval and indexing. More recently, information science researchers have turned to other newer artificial-intelligence-based inductive learning techniques including neural networks, symbolic learning, and genetic algorithms. These newer techniques, which are grounded on diverse paradigms, have provided great opportunities for researchers to enhance the information processing and retrieval capabilities of current information storage and retrieval systems. In this article, we first provide an overview of these newer techniques and their use in information science research. To familiarize readers with these techniques, we present three popular methods: the connectionist Hopfield network; the symbolic ID3/ID5R; and evolution-based genetic algorithms. We discuss their knowledge representations and algorithms in the context of information retrieval. Sample implementation and testing results from our own research are also provided for each technique. We believe these techniques are promising in their ability to analyze user queries, identify users' information needs, and suggest alternatives for search. With proper user-system interactions, these methods can greatly complement the prevailing full-text, keyword-based, probabilistic, and knowledge-based techniques. © 1995 John Wiley & Sons, Inc. Hsinchun Chen |
J. Am. Soc. Inf. Sci. | 1 |
| 1995 | An Algorithmic Approach to Concept Exploration in a Large Knowledge Network (Automatic Thesaurus Consultation): Symbolic Branch-and-Bound Search vs. Connectionist Hopfield Net ActivationabstractThis paper presents a framework for knowledge discovery and concept exploration. In order to enhance the concept exploration capability of knowledge-based systems and to alleviate the limitations of the manual browsing approach, we have developed two spreading activation-based algorithms for concept exploration in large, heterogeneous networks of concepts (e.g., multiple thesauri). One algorithm, which is based on the symbolic AI paradigm, performs a conventional branch-and-bound search on a semantic net representation to identify other highly relevant concepts (a serial, optimal search process). The second algorithm, which is based on the neural network approach, executes the Hopfield net parallel relaxation and convergence process to identify “convergent” concepts for some initial queries (a parallel, heuristic search process). Both algorithms can be adopted for automatic, multiple-thesauri consultation. We tested these two algorithms on a large text-based knowledge network of about 13,000 nodes (terms) and 80,000 directed links in the area of computing technologies. This knowledge network was created from two external thesauri and one automatically generated thesaurus. We conducted experiments to compare the behaviors and performances of the two algorithms with the hypertext-like browsing process. Our experiment revealed that manual browsing achieved higher-term recall but lower-term precision in comparison to the algorithmic systems. However, it was also a much more laborious and cognitively demanding process. In document retrieval, there were no statistically significant differences in document recall and precision between the algorithms and the manual browsing process. In light of the effort required by the manual browsing process, our proposed algorithmic approach presents a viable option for efficiently traversing large-scale, multiple thesauri (knowledge network). © 1995 John Wiley & Sons, Inc. Hsinchun Chen, Tobun Dorbin Ng |
J. Am. Soc. Inf. Sci. | 1 |
| 1995 | Automatic Thesaurus Generation for an Electronic Community SystemabstractThis research reports an algorithmic approach to the automatic generation of thesauri for electronic community systems. The techniques used included term filtering, automatic indexing, and cluster analysis. The testbed for our research was the Worm Community System, which contains a comprehensive library of specialized community data and literature, currently in use by molecular biologists who study the nematode worm C. elegans. The resulting worm thesaurus included 2709 researchers' names, 798 gene names, 20 experimental methods, and 4302 subject descriptors. On average, each term had about 90 weighted neighboring terms indicating relevant concepts. The thesaurus was developed as an online search aide. We tested the worm thesaurus in an experiment with six worm researchers of varying degrees of expertise and background. The experiment showed that the thesaurus was an excellent “memory-jogging” device and that it supported learning and serendipitous browsing. Despite some occurrences of obvious noise, the system was useful in suggesting relevant concepts for the researchers' queries and it helped improve concept recall. With a simple browsing interface, an automatic thesaurus can become a useful tool for online search and can assist researchers in exploring and traversing a dynamic and complex electronic community system. © 1995 John Wiley & Sons, Inc. Hsinchun Chen, Tak Yim, David Fye, Bruce R. Schatz |
J. Am. Soc. Inf. Sci. | 1 |
| 1992 | Browsing in hypertext: a cognitive studyabstractSeveral dimensions of browsing are examined to find out: what browsing is and what cognitive processes are associated with it; whether there is a browsing strategy and, if so, whether there are any differences between how subject-area experts and novices browse; and how this knowledge can be applied to improve the design of hypertext systems. Two groups of students, subject-area experts and novices, were studied while browsing a Macintosh HyperCard application. Three browsing strategies were identified: (1) search-oriented browse: scanning and reviewing information relevant to a fixed task; (2) review-browse: scanning and reviewing interesting information in the presence of transient browse goals that represent changing tasks; and (3) scan-browse: scanning for interesting information without review. Most subjects used review-browse interspersed with search-oriented browse. Within this strategy, comparisons showed that experts browsed in more depth, and viewed information differently than did novices. Based on these findings, suggestions are made to hypertext developers.> Erran Carmel, Stephen Crawford, Hsinchun Chen |
IEEE Trans. Syst. Man Cybern. | 3 |
| 1992 | Automatic construction of networks of concepts characterizing document databasesabstractTwo East-bloc computing knowledge bases, both based on a semantic network structure, were created automatically from large, operational textual databases using two statistical algorithms. The knowledge bases were evaluated in detail in a concept-association experiment based on recall and recognition tests. In the experiment, one of the knowledge bases, which exhibited the asymmetric link property, outperformed four experts in recalling relevant concepts in East-bloc computing. The knowledge base, which contained 20000 concepts (nodes) and 280000 weighted relationships (links), was incorporated as a thesaurus-like component in an intelligent retrieval system. The system allowed users to perform semantics-based information management and information retrieval via interactive, conceptual relevance feedback.> Hsinchun Chen, Kevin J. Lynch |
IEEE Trans. Syst. Man Cybern. | 1 |
| 1991 | Cognitive process as a basis for intelligent retrieval systems design
Hsinchun Chen, Vasant Dhar |
Inf. Process. Manag. | 1 |
| 1990 | A Knowledge-Based Design for Hypertext-Based Document Retrieval Systems
Hsinchun Chen |
DEXA | 1 |
| 1990 | Online Query Refinement on Information Retrieval Systems: A Process Model of Searcher/System InteractionsabstractThis article reports findings of empirical research that investigated information searchers' online query refinement process. Prior studies have recognized the information specialists' role in helping searchers articulate and refine queries. Using a semantic network and a Problem Behavior Graph to represent the online search process, our study revealed that searchers also refined their own queries in an online task environment. The information retrieval system played a passive role in assisting online query refinement, which was, however, or that confirmed Taylor's four-level query formulation model. Based on our empirical findings, we proposed using a process model to facilitate and improve query refinement in an online environment. We believe incorporating this model into retrieval systems can result in the design of more “intelligent” and useful information retrieval systems. Hsinchun Chen, Vasant Dhar |
SIGIR | 1 |
| 1990 | User Misconceptions of Information Retrieval Systems
Hsinchun Chen, Vasant Dhar |
Int. J. Man Mach. Stud. | 1 |
| 1987 | Reducing Indeterminism in Consultation: A Cognitive Model of User/Librarian Interactions
Hsinchun Chen, Vasant Dhar |
AAAI | 1 |