Shimei Pan

dblp:p/ShimeiPan · DBLP profile ↗
← Back
80ranked-venue papers
18as first author
15since 2021 · last 2026
0000-0002-5989-8543ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 46 · 14 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 24 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 7 first-author · 2 since 2021Databases, data management, data science and information retrieval · 12 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021
YearPublicationVenuePosition
2026 An LLM-Based Agentic AI System for Automated Construction of Knowledge Taxonomies in Data Science Problem Solving
abstract
The widespread integration of artificial intelligence into data science practice is fundamentally reconfiguring the roles of data science practitioners. As procedural and execution-oriented tasks such as coding, data preprocessing, and routine analysis become increasingly automated, human contribution is shifting toward complex cognitive task of problem solving including higher-order reasoning and decision making. This shift challenges prevailing models of data science education, which has been focusing on tools and techniques while offering limited support to develop those critical problem solving competencies. A prerequisite for addressing this gap is formalizing the Data Science Problem Solving (DSPS) competency as structured knowledge that can be systematically represented and evaluated, for example, by adopting principled approach of constructing knowledge taxonomies for DSPS. However, constructing a coherent and verifiable taxonomy of such knowledge remains a nontrivial knowledge engineering task. In this paper, we present a multi-agent LLM framework for the automated generation and evaluation of a DSPS knowledge taxonomy, grounded in established taxonomy construction and evaluation methods. The system comprises an Author agent that proposes and revises taxonomies and a Critique agent that evaluates them and provides feedback based on explicit assessment criteria. Additionally, an Orchestrator agent manages iterative refinement by routing feedback and enforcing stopping conditions. Our results demonstrate the potential of organizing LLMs into collaborative-adversarial architectures that support reflective, iterative knowledge construction and refinement rather than single-pass content generation. When further developed, the proposed multi-agent LLM systems can function as a scalable framework for principled knowledge engineering, enabling the systematic design of DSPS curriculum, assessments, instructional scaffolds, and AI-assisted learning environments at scale.
Md Sakib Ul Rahman Sourove, Shimei Pan, Lujie Karen Chen
L@S2
2026 Developing Problem-Solving Competency in Data Science: Exploring A Case-Based Approach
Lujie Karen Chen, Maryam Alomair, Muhammad Ali Yousuf, Shimei Pan
SIGCSE (1)4
2025 GenderAlign: An Alignment Dataset for Mitigating Gender Bias in Large Language Models
abstract
Large Language Models (LLMs) are prone to generating content that exhibits gender biases, raising significant ethical concerns. Alignment, the process of fine-tuning LLMs to better align with desired behaviors, is recognized as an effective approach to mitigate gender biases. Although proprietary LLMs have made significant strides in mitigating gender bias, their alignment datasets are not publicly available. The commonly used and publicly available alignment dataset, HH-RLHF, still exhibits gender bias to some extent. There is a lack of publicly available alignment datasets specifically designed to address gender bias. Hence, we developed a new dataset named GenderAlign, aiming at mitigating a comprehensive set of gender biases in LLMs. This dataset comprises 8k single-turn dialogues, each paired with a “chosen” and a “rejected” response. Compared to the “rejected” responses, the “chosen” responses demonstrate lower levels of gender bias and higher quality. Furthermore, we categorized the gender biases in the “rejected” responses of GenderAlign into 4 principal categories. The experimental results show the effectiveness of GenderAlign in reducing gender bias in LLMs.
Tao Zhang 0019, Ziqian Zeng, YuxiangXiao YuxiangXiao, Huiping Zhuang, Cen Chen 0002, James R. Foulds, Shimei Pan
ACL (1)7
2025 fairGNN-WOD: Fair Graph Learning Without Complete Demographics
abstract
Graph Neural Networks (GNNs) have excelled in diverse applications due to their outstanding predictive performance, yet they often overlook fairness considerations, prompting numerous recent efforts to address this societal concern. However, most fair GNNs assume complete demographics by design, which is impractical in most real-world socially sensitive applications due to privacy, legal, or regulatory restrictions. For example, the Consumer Financial Protection Bureau (CFPB) mandates that creditors ensure fairness without requesting or collecting information about an applicant’s race, religion, nationality, sex, or other demographics. To this end, this paper proposes fairGNN-WOD, a first-of-its-kind framework that considers mitigating unfairness in graph learning without using demographic information. In addition, this paper provides a theoretical perspective on analyzing bias in node representations and establishes the relationship between utility and fairness objectives. Experiments on three real-world graph datasets illustrate that fairGNN-WOD outperforms state-of-the-art baselines in achieving fairness but also maintains comparable prediction performance.
Zichong Wang, Shimei Pan, Jun Liu 0075, Fahad Saeed, Meikang Qiu, Wenbin Zhang 0002
IJCAI3
2025 Systematically Identifying, Defining and Organizing Knowledge Components for Data Science Problem Solving through Human-LLM Collaboration
abstract
As demand grows for job-ready data science professionals, there is increasing recognition that traditional training often falls short in cultivating the higher-order reasoning and real-world problem-solving skills essential to the field. A foundational step toward addressing this gap is the identification and organization of knowledge components (KCs) that underlie data science problem solving (DSPS). KCs represent conditional knowledge-knowing about appropriate actions given particular contexts or conditions-and correspond to the critical decisions data scientists must make throughout the problem-solving process. While existing taxonomies in data science education support curriculum development, they often lack the granularity and focus needed to support the assessment and development of DSPS skills. In this paper, we present a novel framework that combines the strengths of large language models (LLMs) and human expertise to identify, define, and organize KCs specific to DSPS. We treat LLMs as ''knowledge engineering assistants'' capable of generating candidate KCs by drawing on their extensive training data, which includes a vast amount of domain knowledge and diverse sets of real-world DSPS cases. Our process involves prompting multiple LLMs to generate decision points, synthesizing and refining KC definitions across models, and using sentence-embedding models to infer the underlying structure of the resulting taxonomy. Human experts then review and iteratively refine the taxonomy to ensure validity. This human-AI collaborative workflow offers a scalable and efficient proof-of-concept for LLM-assisted knowledge engineering. The resulting KC taxonomy lays the groundwork for developing fine-grained assessment tools and adaptive learning systems that support deliberate practice in DSPS. Furthermore, the framework illustrates the potential of LLMs not just as content generators but as partners in structuring domain knowledge to inform instructional design. Future work will involve extending the framework by generating a directed graph of KCs based on their input-output dependencies and validating the taxonomy through expert consensus and learner studies. This approach contributes to both the practical advancement of DSPS coaching in data science education and the broader methodological toolkit for AI-supported knowledge engineering.
Priyanka Rani, Maryam Alomair, Shimei Pan, Lujie Karen Chen
L@S3
2024 DoubleDistillation: Enhancing LLMs for Informal Text Analysis using Multistage Knowledge Distillation from Speech and Text
abstract
Traditional large language models (LLMs) leverage extensive text corpora but lack access to acoustic and para-linguistic cues present in speech. There is a growing interest in enhancing text-based models with audio information. However, current models often require an aligned audio-text dataset which is frequently much smaller than typical language model training corpora. Moreover, these models often require both text and audio streams during inference/testing. In this study, we introduce a novel two-stage knowledge distillation (KD) approach that enables language models to (a) incorporate rich acoustic and paralinguistic information from speech, (b) utilize text corpora comparable in size to typical language model training data, and (c) support text-only analysis without requiring an audio stream during inference/testing. Specifically, we employ a pre-trained speech embedding teacher model (OpenAI Whisper) to train a Teacher Assistant (TA) model on an aligned audio-text dataset in the first stage. In the second stage, the TA’s knowledge is transferred to a student language model trained on a conventional text dataset. Thus, our two-stage KD method leverages both the acoustic and paralinguistic cues in the aligned audio-text data and the nuanced linguistic knowledge in a large text-only dataset. Based on our evaluation, this DoubleDistillation system consistently outperforms traditional LLMs in 15 informal text understanding tasks.
Fatema Hasan, James R. Foulds, Shimei Pan, Bishwaranjan Bhattacharjee
ICMI4
2023 Modeling Metacognitive and Cognitive Processes in Data Science Problem Solving (Student Abstract)
abstract
Data Science (DS) is an interdisciplinary topic that is applicable to many domains. In this preliminary investigation, we use caselet, a mini-version of a case study, as a learning tool to allow students to practice data science problem solving (DSPS). Using a dataset collected from a real-world classroom, we performed correlation analysis to reveal the structure of cognition and metacognition processes. We also explored the similarity of different DS knowledge components based on students’ performance. In addition, we built a predictive model to characterize the relationship between metacognition, cognition, and learning gain.
Maryam Alomair, Shimei Pan, Lujie Karen Chen
AAAI2
2023 When Biased Humans Meet Debiased AI: A Case Study in College Major Recommendation
abstract
Currently, there is a surge of interest in fair Artificial Intelligence (AI) and Machine Learning (ML) research which aims to mitigate discriminatory bias in AI algorithms, e.g., along lines of gender, age, and race. While most research in this domain focuses on developing fair AI algorithms, in this work, we examine the challenges which arise when humans and fair AI interact. Our results show that due to an apparent conflict between human preferences and fairness, a fair AI algorithm on its own may be insufficient to achieve its intended results in the real world. Using college major recommendation as a case study, we build a fair AI recommender by employing gender debiasing machine learning techniques. Our offline evaluation showed that the debiased recommender makes fairer career recommendations without sacrificing its accuracy in prediction. Nevertheless, an online user study of more than 200 college students revealed that participants on average prefer the original biased system over the debiased system. Specifically, we found that perceived gender disparity is a determining factor for the acceptance of a recommendation. In other words, we cannot fully address the gender bias issue in AI recommendations without addressing the gender bias in humans. We conducted a follow-up survey to gain additional insights into the effectiveness of various design options that can help participants to overcome their own biases. Our results suggest that making fair AI explainable is crucial for increasing its adoption in the real world.
Clarice Wang, Kathryn Wang, Andrew Bian, Rashidul Islam, Kamrun Keya, James R. Foulds, Shimei Pan
ACM Trans. Interact. Intell. Syst.7
2022 Do Humans Prefer Debiased AI Algorithms? A Case Study in Career Recommendation
abstract
Currently, there is a surge of interest in fair Artificial Intelligence (AI) and Machine Learning (ML) research which aims to mitigate discriminatory bias in AI algorithms, e.g. along lines of gender, age, and race. While most research in this domain focuses on developing fair AI algorithms, in this work, we examine the challenges which arise when human- fair-AI interact. Our results show that due to an apparent conflict between human preferences and fairness, a fair AI algorithm on its own may be insufficient to achieve its intended results in the real world. Using college major recommendation as a case study, we build a fair AI recommender by employing gender debiasing machine learning techniques. Our offline evaluation showed that the debiased recommender makes fairer and more accurate college major recommendations. Nevertheless, an online user study of more than 200 college students revealed that participants on average prefer the original biased system over the debiased system. Specifically, we found that the perceived gender disparity associated with a college major is a determining factor for the acceptance of a recommendation. In other words, our results demonstrate we cannot fully address the gender bias issue in AI recommendations without addressing the gender bias in humans. They also highlight the urgent need to extend the current scope of fair AI research from narrowly focusing on debiasing AI algorithms to including new persuasion and bias explanation technologies in order to achieve intended societal impacts.
Clarice Wang, Kathryn Wang, Andrew Bian, Rashidul Islam, Kamrun Keya, James R. Foulds, Shimei Pan
IUI7
2022 Incorporating LIWC in Neural Networks to Improve Human Trait and Behavior Analysis in Low Resource Scenarios
abstract
Psycholinguistic knowledge resources have been widely used in constructing features for text-based human trait and behavior analysis. Recently, deep neural network (NN)-based text analysis methods have gained dominance due to their high prediction performance. However, NN-based methods may not perform well in low resource scenarios where the ground truth data is limited (e.g., only a few hundred labeled training instances are available). In this research, we investigate diverse methods to incorporate Linguistic Inquiry and Word Count (LIWC), a widely-used psycholinguistic lexicon, in NN models to improve human trait and behavior analysis in low resource scenarios. We evaluate the proposed methods in two tasks: predicting delay discounting and predicting drug use based on social media posts. The results demonstrate that our methods perform significantly better than baselines that use only LIWC or only NN-based feature learning methods. They also performed significantly better than published results on the same dataset.
Isil Doga Yakut Kiliç, Shimei Pan
LREC2
2021 Can We Obtain Fairness For Free?
abstract
There is growing awareness that AI and machine learning systems can in some cases learn to behave in unfair and discriminatory ways with harmful consequences. However, despite an enormous amount of research, techniques for ensuring AI fairness have yet to see widespread deployment in real systems. One of the main barriers is the conventional wisdom that fairness brings a cost in predictive performance metrics such as accuracy which could affect an organization's bottom-line. In this paper we take a closer look at this concern. Clearly fairness/performance trade-offs exist, but are they inevitable? In contrast to the conventional wisdom, we find that it is frequently possible, indeed straightforward, to improve on a trained model's fairness without sacrificing predictive performance. We systematically study the behavior of fair learning algorithms on a range of benchmark datasets, showing that it is possible to improve fairness to some degree with no loss (or even an improvement) in predictive performance via a sensible hyper-parameter selection strategy. Our results reveal a pathway toward increasing the deployment of fair AI methods, with potentially substantial positive real-world impacts.
Rashidul Islam, Shimei Pan, James R. Foulds
AIES2
2021 Incorporating medical knowledge in BERT for clinical relation extraction
abstract
In recent years pre-trained language models (PLM) such as BERT have proven to be very effective in diverse NLP tasks such as Information Extraction, Sentiment Analysis and Question/Answering.Trained with massive generaldomain text, these pre-trained language models capture rich syntactic, semantic and discourse information in the text.However, due to the differences between general and specific domain text (e.g., Wikipedia text versus clinic notes), these models may not be ideal for domain-specific tasks (e.g., extracting clinical relations).Furthermore, it may require additional medical knowledge to understand clinical text properly.To solve these issues, in this research, we conduct a comprehensive examination of different techniques to add medical knowledge into a pre-trained BERT model for clinical relation extraction.Our best model outperformed the state-of-the-art systems on the benchmark i2b2/VA 2010 clinical relation extraction dataset.
Arpita Roy, Shimei Pan
EMNLP (1)2
2021 Fair Representation Learning for Heterogeneous Information Networks
Ziqian Zeng, Rashidul Islam, Kamrun Keya, James R. Foulds, Yangqiu Song, Shimei Pan
ICWSM6
2021 Equitable Allocation of Healthcare Resources with Fair Survival Models
abstract
Healthcare programs such as Medicaid provide crucial services to vulnerable populations, but due to limited resources, many of the individuals who need these services the most languish on waiting lists.Survival models, e.g. the Cox proportional hazards model, can potentially improve this situation by predicting individuals' levels of need, which can then be used to prioritize the waiting lists.Providing care to those in need can prevent institutionalization for those individuals, which both improves quality of life and reduces overall costs.While the benefits of such an approach are clear, care must be taken to ensure that the prioritization process is fair, and does not reinforce harmful systemic bias.We develop multiple fairness definitions and corresponding fair learning algorithms for survival models to ensure equitable allocation of healthcare resources.We demonstrate the utility of our methods in terms of fairness and predictive accuracy on three publicly available survival datasets.
Kamrun Keya, Rashidul Islam, Shimei Pan, Ian Stockwell, James R. Foulds
SDM3
2021 Debiasing Career Recommendations with Neural Fair Collaborative Filtering
abstract
A growing proportion of human interactions are digitized on social media platforms and subjected to algorithmic decision-making, and it has become increasingly important to ensure fair treatment from these algorithms. In this work, we investigate gender bias in collaborative-filtering recommender systems trained on social media data. We develop neural fair collaborative filtering (NFCF), a practical framework for mitigating gender bias in recommending career-related sensitive items (e.g. jobs, academic concentrations, or courses of study) using a pre-training and fine-tuning approach to neural collaborative filtering, augmented with bias correction techniques. We show the utility of our methods for gender de-biased career and college major recommendations on the MovieLens dataset and a Facebook dataset, respectively, and achieve better performance and fairer behavior than several state-of-the-art models.
Rashidul Islam, Kamrun Keya, Ziqian Zeng, Shimei Pan, James R. Foulds
WWW4
2020 Machine Learning and Student Performance in Teams
Rohan Ahuja, Daniyal Khan, Sara Tahir, Magdalene Wang, Danilo Symonette, Shimei Pan, Simon Stacey, Don Engel
AIED (2)6
2020 An Intersectional Definition of Fairness
abstract
We propose differential fairness, a multi-attribute definition of fairness in machine learning which is informed by intersectionality, a critical lens arising from the humanities literature, leveraging connections between differential privacy and legal notions of fairness. We show that our criterion behaves sensibly for any subset of the set of protected attributes, and we prove economic, privacy, and generalization guarantees. We provide a learning algorithm which respects our differential fairness criterion. Experiments on the COMPAS criminal recidivism dataset and census data demonstrate the utility of our methods.
James R. Foulds, Rashidul Islam, Kamrun Keya, Shimei Pan
ICDE4
2020 Integrating Text Embedding with Traditional NLP Features for Clinical Relation Extraction
abstract
Recently, text embedding techniques such as Word2Vec and BERT have produced state-of-the-art results in a wide variety of NLP tasks. As a result, traditional NLP features frequently used in Information Extraction (IE) such as POS tags, dependency relations and semantic types have received less attention. In this paper, we investigate whether traditional NLP features can be combined with word and sentence embeddings to improve relation extraction. We have explored diverse feature sets and different neural network architectures and evaluated our models on a benchmark clinical text dataset. Our new models significantly outperformed all the baselines on the same dataset.
Fatema Hasan, Arpita Roy, Shimei Pan
ICTAI3
2020 Incorporating Extra Knowledge to Enhance Word Embedding
abstract
Word embedding, a process to automatically learn the mathematical representations of words from unlabeled text corpora, has gained a lot of attention recently. Since words are the basic units of a natural language, the more precisely we can represent the morphological, syntactic and semantic properties of words, the better we can support downstream Natural Language Processing (NLP) tasks. Since traditional word embeddings are mainly designed to capture the semantic relatedness between co-occurred words in a predefined context, it may not be effective in encoding other information that is important for different NLP applications. In this survey, we summarize the recent advances in incorporating extra knowledge to enhance word embedding. We will also identify the limitations of existing work as well as point out a few promising future directions.
Arpita Roy, Shimei Pan
IJCAI2
2020 Bayesian Modeling of Intersectional Fairness: The Variance of Bias
abstract
Intersectionality is a framework that analyzes how interlocking systems of power and oppression affect individuals along overlapping dimensions including race, gender, sexual orientation, class, and disability. Intersectionality theory therefore implies it is important that fairness in artificial intelligence systems be protected with regard to multi-dimensional protected attributes. However, the measurement of fairness becomes statistically challenging in the multi-dimensional setting due to data sparsity, which increases rapidly in the number of dimensions, and in the values per dimension. We present a Bayesian probabilistic modeling approach for the reliable, data-efficient estimation of fairness with multidimensional protected attributes, which we apply to two existing intersectional fairness metrics. Experimental results on census data and the COMPAS criminal justice recidivism dataset demonstrate the utility of our methodology, and show that Bayesian methods are valuable for the modeling and measurement of fairness in intersectional contexts.
James R. Foulds, Rashidul Islam, Kamrun Keya, Shimei Pan
SDM4
2020 Intersectional AI: A Study of How Information Science Students Think about Ethics and Their Impact
abstract
Recent literature has demonstrated the limited and, in some instances, waning role of ethical training in computing classes in the US. The capacity for artificial intelligence (AI) to be inequitable or harmful is well documented, yet it's an issue that continues to lack apparent urgency or effective mitigation. The question we raise in this paper is how to prepare future generations to recognize and grapple with the ethical concerns of a range of issues plaguing AI, particularly when they are combined with surveillance technologies in ways that have grave implications for social participation and restriction?from risk assessment and bail assignment in criminal justice, to public benefits distribution and access to housing and other critical resources that enable security and success within society. The US is a mecca of information and computer science (IS and CS) learning for Asian students whose experiences as minorities renders them familiar with, and vulnerable to, the societal bias that feeds AI bias. Our goal was to better understand how students who are being educated to design AI systems think about these issues, and in particular, their sensitivity to intersectional considerations that heighten risk for vulnerable groups. In this paper we report on findings from qualitative interviews with 20 graduate students, 11 from an AI class and 9 from a Data Mining class. We find that students are not predisposed to think deeply about the implications of AI design for the privacy and well-being of others unless explicitly encouraged to do so. When they do, their thinking is focused through the lens of personal identity and experience, but their reflections tend to center on bias, an intrinsic feature of design, rather than on fairness, an outcome that requires them to imagine the consequences of AI. While they are, in fact, equipped to think about fairness when prompted by discussion and by design exercises that explicitly invite consideration of intersectionality and structural inequalities, many need help to do this empathy 'work.' Notably, the students who more frequently reflect on intersectional problems related to bias and fairness are also more likely to consider the connection between model attributes and bias, and the interaction with context. Our findings suggest that experience with identity-based vulnerability promotes more analytically complex thinking about AI, lending further support to the argument that identity-related ethics should be integrated into IS and CS curriculums, rather than positioned as a stand-alone course.
Nora McDonald, Shimei Pan
Proc. ACM Hum. Comput. Interact.2
2020 Special Issue on Data-Driven Personality Modeling for Intelligent Human-Computer Interaction
abstract
Elsevier’s Scopus, the largest abstract and citation database of peer-reviewed literature. Search and access research from the science, technology, medicine, social sciences and arts and humanities fields.
Shimei Pan, Oliver Brdiczka, Andrea Kleinsmith, Yangqiu Song
ACM Trans. Interact. Intell. Syst.1
2019 Supervising Unsupervised Open Information Extraction Models
abstract
Arpita Roy, Youngja Park, Taesung Lee, Shimei Pan. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Arpita Roy, Youngja Park, Taesung Lee, Shimei Pan
EMNLP/IJCNLP (1)4
2019 Incorporating Domain Knowledge in Learning Word Embedding
abstract
Word embedding is a Natural Language Processing (NLP) technique that automatically maps words from a vocabulary to vectors of real numbers in an embedding space. It has been widely used in recent years to boost the performance of a variety of NLP tasks such as named entity recognition, syntactic parsing and sentiment analysis. Classic word embedding methods such as Word2Vec and GloVe work well when they are given a large text corpus. When the input texts are sparse as in many specialized domains (e.g., cybersecurity), these methods often fail to produce high-quality vectors. In this paper, we describe a novel method, called Annotation Word Embedding (AWE), to train domain-specific word embeddings from sparse texts. Our method is generic and can leverage diverse types of domain knowledge such as domain vocabulary, semantic relations and attribute specifications. Specifically, our method encodes diverse types of domain knowledge as text annotations and incorporates the annotations in word embedding. We have evaluated AWE in two cybersecurity applications: identifying malware aliases and identifying relevant Common Vulnerabilities and Exposures (CVEs). Our evaluation results have demonstrated the effectiveness of our method over state-of-the-art baselines.
Arpita Roy, Youngja Park, Shimei Pan
ICTAI3
2019 Social Media-based User Embedding: A Literature Review
abstract
Automated representation learning is behind many recent success stories in machine learning. It is often used to transfer knowledge learned from a large dataset (e.g., raw text) to tasks for which only a small number of training examples are available. In this paper, we review recent advance in learning to represent social media users in low-dimensional embeddings. The technology is critical for creating high performance social media-based human traits and behavior models since the ground truth for assessing latent human traits and behavior is often expensive to acquire at a large scale. In this survey, we review typical methods for learning a unified user embeddings from heterogeneous user data (e.g., combines social media texts with images to learn a unified user representation). Finally we point out some current issues and future directions.
Shimei Pan, Tao Ding 0007
IJCAI1
2018 Predicting Delay Discounting from Social Media Likes with Unsupervised Feature Learning
abstract
Delay discounting, a behavioral measure of impulsivity, is often used to quantify the human tendency to choose a smaller, sooner reward (e.g., $1 today) over a larger, later reward ($2 tomorrow). Delay discounting and its relation to human decision making is a hot topic in economics and behavior science since pitting the demands of long-term goals against short-term desires is among the most difficult tasks in human decision making. Previously, small-scale studies based on questionnaires were used to analyze an individual's delay discounting rate (DDR) and its relation to his/her real-world behavior such as substance abuse, pathological gambling and poor academic performance. In this research, we employ large-scale social media analytics to study DDR and its relation to people's social media behavior (e.g., their Likes on Facebook). We also build computational models to automatically infer DDR from Social Media Likes. Since the predicting feature space is very large and the size of the delay discounting ground truth dataset is relatively small, we focus on studying the impact of different unsupervised feature learning methods on predicting performance. Our results demonstrate the significant role unsupervised feature learning plays in this task.
Tao Ding 0007, Warren K. Bickel, Shimei Pan
ASONAM3
2018 Interpreting Social Media-Based Substance Use Prediction Models with Knowledge Distillation
abstract
People nowadays spend a significant amount of time on social media such as Twitter, Facebook, and Instagram. As a result, social media data capture rich human behavioral evidence that can be used to help us understand their thoughts, behavior and decision making process. Social media data, however, are mostly unstructured (e.g., text and images) and may involve a large number of raw features (e.g., millions of raw text and image features). Moreover, the ground truth data about human behavior and decision making could be difficult to obtain at a large scale. As a result, most state-of-the-art social media-based human behavior models employ sophisticated unsupervised feature learning to leverage a large amount of unsupervised data. Unfortunately, these advanced models often rely on latent features that are hard to explain. Since understanding the knowledge captured in these models is important for behavior scientists, public health providers as well as policymakers, in this research, we focus on employing a knowledge distillation framework to build machine learning models with not only state-of-the-art predictive performance but also interpretable results. We evaluate the effectiveness of the proposed framework in explaining Substance Use Disorder (SUD) prediction models. Our best models achieved 87% ROC AUC for predicting tobacco use, 84% for alcohol use and 93% for drug use, which are comparable to existing state-of-the-art SUD prediction models. Since these models are also interpretable (e.g., a logistics regression model and a gradient boosting tree model), we combine the results from these models to gain insight into the relationship between a user's social media behavior (e.g., social media likes and word usage) and substance use.
Tao Ding 0007, Fatema Hasan, Warren K. Bickel, Shimei Pan
ICTAI4
2017 Multi-View Unsupervised User Feature Embedding for Social Media-based Substance Use Prediction
abstract
In this paper, we demonstrate how the state-of-the-art machine learning and text mining techniques can be used to build effective social media-based substance use detection systems.Since a substance use ground truth is difficult to obtain on a large scale, to maximize system performance, we explore different unsupervised feature learning methods to take advantage of a large amount of unsupervised social media data.We also demonstrate the benefit of using multi-view unsupervised feature learning to combine heterogeneous user information such as Facebook "likes" and "status updates" to enhance system performance.Based on our evaluation, our best models achieved 86% AUC for predicting tobacco use, 81% for alcohol use and 84% for illicit drug use, all of which significantly outperformed existing methods.Our investigation has also uncovered interesting relations between a user's social media behavior (e.g., word usage) and substance use.
Tao Ding 0007, Warren K. Bickel, Shimei Pan
EMNLP3
2017 Automated Detection of Substance Use-Related Social Media Posts Based on Image and Text Analysis
abstract
Nowadays, teens and young adults spend a significant amount of time on social media. According to the national survey of American attitudes on substance abuse, American teens who spend time on social media sites are at increased risk of smoking, drinking and illicit drug use. Reducing teens' exposure to substance use-related social media posts may help minimize their risk of future substance use and addiction. In this paper, we present a method for automated detection of substance userelated social media posts. With this technology, substance userelated content can be automatically filtered out from social media. To detect substance use related social media posts, we employ the state-of-the-art social media analytics that combines Neural Network-based image and text processing technologies. Our evaluation results demonstrate that image features derived using Convolutional Neural Network and textual features derived using neural document embedding are effective in identifying substance use-related social media posts.
Arpita Roy, Anamika Paul, Hamed Pirsiavash, Shimei Pan
ICTAI4
2017 Machine learning on big data: Opportunities and challenges
Lina Zhou, Shimei Pan, Jianwu Wang 0001, Athanasios V. Vasilakos
Neurocomputing2
2016 What's Hot in Intelligent User Interfaces
abstract
The ACM Conference on Intelligent User Interfaces (IUI) is the annual meeting of the intelligent user interface community and serves as a premier international forum for reporting outstanding research and development on intelligent user interfaces. ACM IUI is where the Human-Computer Interaction (HCI) community meets the Artificial Intelligence (AI) community. Here we summarize the latest trends in IUI based on our experience organizing the 20th ACM IUI Conference in Atlanta in 2015.
Shimei Pan, Oliver Brdiczka, Giuseppe Carenini, Polo Chau, Per Ola Kristensson
AAAI1
2016 Analyzing and retrieving illicit drug-related posts from social media
abstract
Illicit drug use is a serious problem around the world. Social media has increasingly become an important tool for analyzing drug use patterns and monitoring emerging drug abuse trends. Accurately retrieving illicit drug-related social media posts is an important step in this research. Frequently, hashtags are used to identify and retrieve posts on a specific topic. However hashtags are highly ambiguous. Posts with the same hashtags are not always on the same topic. Moreover, hashtags are evolving, especially those related to illicit drugs. New street names are introduced constantly to avoid detection. In this paper, we employ topic modeling to disambiguate hashtags and track the changes of hashtags using semantic word embedding. Our preliminary evaluation shows the promise of these methods.
Tao Ding 0007, Arpita Roy, Zhiyuan Chen 0003, Qian Zhu 0003, Shimei Pan
BIBM5
2016 Cross-Domain Error Correction in Personality Prediction
abstract
In this paper, we analyze domain bias in automated text-based personality prediction, and proposes a novel method to correct domain bias. The proposed approach is very general since it requires neither retraining a personality prediction system using examples from a new domain, nor any knowledge of the original training data used to develop the system. We conduct several experiments to evaluate the effectiveness of the method, and the findings indicate a significant improvement of prediction accuracy.
Isil Doga Yakut Kiliç, Shimei Pan
ECAI2
2016 Personalized Emphasis Framing for Persuasive Message Generation
abstract
In this paper, we present a study on personalized emphasis framing which can be used to tailor the content of a message to enhance its appeal to different individuals. With this framework, we directly model content selection decisions based on a set of psychologically-motivated domain-independent personal traits including personality (e.g., extraversion and conscientiousness) and basic human values (e.g., self-transcendence and hedonism). We also demonstrate how the analysis results can be used in automated personalized content selection for persuasive message generation.
Tao Ding 0007, Shimei Pan
EMNLP2
2016 Analyzing and Preventing Bias in Text-Based Personal Trait Prediction Algorithms
abstract
Personality prediction based on textual data is one topic gaining attention recently for its potential in large-scale personalized applications such as social media-based marketing and political campaigning. However, when applying this technology in real-world applications, users often encounter situations in which the personality traits derived from different sources (e.g., social media posts versus emails) are inconsistent. Varying results for the same individual renders the tool ineffective and hard to trust. This paper demonstrates the impact of domain bias in automated text-based personality prediction, and proposes a novel method to correct domain bias. The proposed approach is generic since it requires neither retraining the system using examples from an application domain, nor any knowledge of the original training data used by a personal trait analysis tool. We conduct comprehensive experiments to evaluate the effectiveness of the method, and the findings indicate a significant improvement of prediction accuracy (e.g., a 20-30% relative error reduction) with the proposed method.
Isil Doga Yakut Kiliç, Shimei Pan
ICTAI2
2016 Predicting Gene Functional Interactions with Semantic Word Embedding
abstract
In this paper, we present a novel method for predicting gene functional interactions. We study the effectiveness of various raw and derived features from neural word embedding learned from biomedical literature. Our evaluation results demonstrate that the information captured in neural word embedding is very useful and our learned classification models are capable of predicting gene functional interactions with high accuracy.
Arpita Roy, Shimei Pan
ICTAI2
2016 Improving Topic Model Stability for Effective Document Exploration
Yi Yang 0042, Shimei Pan, Yangqiu Song, Jie Lu 0002, Mercan Topkara
IJCAI2
2016 An Empirical Study of the Effectiveness of using Sentiment Analysis Tools for Opinion Mining
Tao Ding 0007, Shimei Pan
WEBIST (2)2
2016 The Stability and Usability of Statistical Topic Models
abstract
Statistical topic models have become a useful and ubiquitous tool for analyzing large text corpora. One common application of statistical topic models is to support topic-centric navigation and exploration of document collections. Existing work on topic modeling focuses on the inference of model parameters so the resulting model fits the input data. Since the exact inference is intractable, statistical inference methods, such as Gibbs Sampling, are commonly used to solve the problem. However, most of the existing work ignores an important aspect that is closely related to the end user experience: topic model stability. When the model is either re-trained with the same input data or updated with new documents, the topic previously assigned to a document may change under the new model, which may result in a disruption of end users’ mental maps about the relations between documents and topics, thus undermining the usability of the applications. In this article, we propose a novel user-directed non-disruptive topic model update method that balances the tradeoff between finding the model that fits the data and maintaining the stability of the model from end users’ perspective. It employs a novel constrained LDA algorithm to incorporate pairwise document constraints, which are converted from user feedback about topics, to achieve topic model stability. Evaluation results demonstrate the advantages of our approach over previous methods.
Yi Yang 0042, Shimei Pan, Jie Lu 0002, Mercan Topkara, Yangqiu Song
ACM Trans. Interact. Intell. Syst.2
2016 An Uncertainty-Aware Approach for Exploratory Microblog Retrieval
abstract
Although there has been a great deal of interest in analyzing customer opinions and breaking news in microblogs, progress has been hampered by the lack of an effective mechanism to discover and retrieve data of interest from microblogs. To address this problem, we have developed an uncertainty-aware visual analytics approach to retrieve salient posts, users, and hashtags. We extend an existing ranking technique to compute a multifaceted retrieval result: the mutual reinforcement rank of a graph node, the uncertainty of each rank, and the propagation of uncertainty among different graph nodes. To illustrate the three facets, we have also designed a composite visualization with three visual components: a graph visualization, an uncertainty glyph, and a flow map. The graph visualization with glyphs, the flow map, and the uncertainty analysis together enable analysts to effectively find the most uncertain results and interactively refine them. We have applied our approach to several Twitter datasets. Qualitative evaluation and two real-world case studies demonstrate the promise of our approach for retrieving high-quality microblog data.
Mengchen Liu, Shixia Liu, Xizhou Zhu, Qinying Liao, Furu Wei, Shimei Pan
IEEE Trans. Vis. Comput. Graph.6
2015 Using Personal Traits For Brand Preference Prediction
abstract
In this paper, we present a comprehensive study of the relationship between an individual's personal traits and his/her brand preferences.In our analysis, we included a large number of character traits such as personality, personal values and individual needs.These trait features were obtained from both a psychometric survey and automated social media analytics.We also included an extensive set of brand names from diverse product categories.From this analysis, we want to shed some light on (1) whether it is possible to use personal traits to infer an individual's brand preferences (2) whether the trait features automatically inferred from social media are good proxies for the ground truth character traits in brand preference prediction.
Shimei Pan, Jalal Mahmud, Huahai Yang, Padmini Srinivasan
EMNLP2
2015 Signals of Expertise in Public and Enterprise Social Q&A
Shimei Pan, Elijah Mayfield, Jie Lu 0002, Jennifer C. Lai
ICWSM1
2015 User-directed Non-Disruptive Topic Model Update for Effective Exploration of Dynamic Content
abstract
Statistical topic models have become a useful and ubiquitous text analysis tool for large corpora. One common application of statistical topic models is to support topic-centric navigation and exploration of document collections at the user interface by automatically grouping documents into coherent topics. For today's constantly expanding document collections, topic models need to be updated when new documents become available. Existing work on topic model update focuses on how to best fit the model to the data, and ignores an important aspect that is closely related to the end user experience: topic model stability. When the model is updated with new documents, the topics previously assigned to old documents may change, which may result in a disruption of end users' mental maps between documents and topics, thus undermining the usability of the applications. In this paper, we describe a user-directed non-disruptive topic model update system, nTMU, that balances the tradeoff between finding the model that fits the data and maintaining the stability of the model from end users' perspective. It employs a novel constrained LDA algorithm (cLDA) to incorporate pair-wise document constraints, which are converted from user feedback about topics, to achieve topic model stability. Evaluation results demonstrate advantages of our approach over previous methods.
Yi Yang 0042, Shimei Pan, Yangqiu Song, Jie Lu 0002, Mercan Topkara
IUI2
2014 Composite Search for Distributed Multimedia Recorded Meetings
abstract
To encourage enterprise knowledge sharing especially, to facilitate the discovery and sharing of enterprise meetings, we develop an end-to-end enterprise meeting service called Agora that manages the full cycle of hosting web meetings and sharing multimedia recorded meeting artifacts. In this paper, we focus on Agora's composite search engine that allows users to seamlessly search distributed multimedia meeting artifacts. Agora has been deployed as a cloud service which allows selected IBM customers to test new collaborative technologies on IBM's Smart Cloud platform.
Shimei Pan, Mercan Topkara, Steve Wood, Jeff Boston, Jennifer C. Lai
ISM1
2014 Expediting expertise: supporting informal social learning in the enterprise
abstract
In this paper, we present Expediting Expertise, a system designed to provide structured support to the otherwise informal process of social learning in the enterprise. It employs a data-driven approach where online content is automatically analyzed and categorized into relevant topics, topic-specific user expertise is calculated by comparing the models of individual users against those of the experts, and personalized recommendation of learning activities is created accordingly to facilitate expertise development. The system's UI is designed to provide users with ongoing feedback of current expertise, progress, and comparison with others. Learning recommendation is visualized with an interactive treemap which presents estimated return on investment and distance to current expertise for each recommended learning activity. Evaluation of the system showed very positive results.
Jennifer C. Lai, Jie Lu 0002, Shimei Pan, Danny Soroker, Mercan Topkara, Justin D. Weisz, Jeff Boston, Jason Crawford
IUI3
2013 Optimizing temporal topic segmentation for intelligent text visualization
abstract
We are building a topic-based, interactive visual analytic tool that aids users in analyzing large collections of text. To help users quickly discover content evolution and significant content transitions within a topic over time, here we present a novel, constraint-based approach to temporal topic segmentation. Our solution splits a discovered topic into multiple linear, non-overlapping sub-topics along a timeline by satisfying a diverse set of semantic, temporal, and visualization constraints simultaneously. For each derived sub-topic, our solution also automatically selects a set of representative keywords to summarize the main content of the sub-topic. Our extensive evaluation, including a crowd-sourced user study, demonstrates the effectiveness of our method over an existing baseline.
Shimei Pan, Michelle X. Zhou, Yangqiu Song, Weihong Qian, Fei Wang 0001, Shixia Liu
IUI1
2013 Constrained Text Coclustering with Supervised and Unsupervised Constraints
abstract
In this paper, we propose a novel constrained coclustering method to achieve two goals. First, we combine information-theoretic coclustering and constrained clustering to improve clustering performance. Second, we adopt both supervised and unsupervised constraints to demonstrate the effectiveness of our algorithm. The unsupervised constraints are automatically derived from existing knowledge sources, thus saving the effort and cost of using manually labeled constraints. To achieve our first goal, we develop a two-sided hidden Markov random field (HMRF) model to represent both document and word constraints. We then use an alternating expectation maximization (EM) algorithm to optimize the model. We also propose two novel methods to automatically construct and incorporate document and word constraints to support unsupervised constrained clustering: 1) automatically construct document constraints based on overlapping named entities (NE) extracted by an NE extractor; 2) automatically construct word constraints based on their semantic distance inferred from WordNet. The results of our evaluation over two benchmark data sets demonstrate the superiority of our approaches against a number of existing approaches.
Yangqiu Song, Shimei Pan, Shixia Liu, Furu Wei, Michelle X. Zhou, Weihong Qian
IEEE Trans. Knowl. Data Eng.2
2012 "You've got video": increasing clickthrough when sharing enterprise video with email
abstract
In this Note we summarize our research on increasing the information scent of video recordings that are shared via email in a corporate setting. We compare two types of email messages for sharing recordings: the first containing basic information (e.g. title, speaker, abstract) with a link to the video; the second with the same information plus a set of video thumbnails (hyperlinked to the segments they represent), which are automatically created by video summarization technology. We report on the results of two user studies. The first one compares the quality of the set of thumbnails selected by the technology to sets selected by 31 humans. The second study examines the clickthrough rates for both email formats (with and without hyperlinked thumbnails) as well as gathering subjective feedback via survey. Results indicate that the email messages with the thumbnails drove significantly higher clickthrough rates than the messages without, even though people clicked on the main video link more frequently than the thumbnails. Survey responses show that users found the email with the thumbnail set significantly more appealing and novel.
Mercan Topkara, Shimei Pan, Jennifer C. Lai, Ahmet Dirik, Steve Wood, Jeff Boston
CHI2
2012 In case you missed it: benefits of attendee-shared annotations for non-attendees of remote meetings
abstract
Corporate meetings are increasingly being held remotely using web technologies. With such remote meetings being recorded and made available after the fact, there is a pressing need for tools to access and utilize these recordings efficiently. Our work explores the utility of using annotations generated by meeting attendees to meet this need. We conducted a controlled lab study to evaluate the benefits of sharing annotations. Attendee-created annotations were shared with non-attendees to assist them on typical information retrieval tasks. Results indicate that (a) non-attendees given access to shared annotations performed about as well as attendees provided with their own and shared annotations, (b) non-attendees were more confident in their responses when they used shared annotations as access cues into the recording than when they directly skimmed the video, and (c) attendees utilized shared annotations more than their own, with similar success and confidence as using their own annotations.
Mukesh Nathan, Mercan Topkara, Jennifer C. Lai, Shimei Pan, Steve Wood, Jeff Boston, Loren G. Terveen
CSCW4
2012 EPIC: a multi-tiered approach to enterprise email prioritization
abstract
We present Enterprise Priority Inbox Classifier (EPIC), an automatic personalized email prioritization system based on a topic-based user model built from the user's email data and relevant enterprise information. The user model encodes the user's topics of interest and email processing behaviors (e.g. read/reply/file) at the granularity of pair-wise interactions between the user and each of his/her email contacts. Given a new message, the user model is used in combination with the message metadata and content to determine the values of a set of contextual features. Contextual features include people-centric features representing information about the user's interaction history and relationship with the email sender, as well as message-centric features focusing on the properties of the message itself. Based on these feature values, EPIC uses a dynamic strategy to combine a global priority classifier with a user-specific classifier for determining the message's priority. An evaluation of EPIC based on 2,064 annotated email messages from 11 users, using 10-fold cross-validation, showed that the system achieves an average accuracy of 81.3%. The user-specific classifier contributed an improvement of 11.5%. Lastly we report on findings regarding the relative value of different contextual features for email prioritization.
Jie Lu 0002, Shimei Pan, Jennifer C. Lai
IUI3
2012 VISA: a visual sentiment analysis system
abstract
Sentiment plays a critical role in many information-centric business scenarios. The opinion mining methods proposed in the recent decade have formed a solid foundation to investigate the sentiment analysis tasks, but are often too complicated and scattered to serve the needs of real customers. We introduce the VISA system in this paper, which applies the visualization technology to synthesize the sentiment analysis results and present to the end user in an interactive manner. VISA builds on the generic sentiment tuple based data model and consumes the different facets of sentiment data with coordinated multiple views, hence is scalable to work with most of existing sentiment analysis engines on various application domains. We showcase the usage of VISA in a real world example and demonstrate the system's effectiveness through the user trail in finding an appropriate hotel for his family trip.
Dongxu Duan, Weihong Qian, Shimei Pan, Lei Shi 0002, Chuang Lin 0002
VINCI3
2012 TIARA: Interactive, Topic-Based Visual Text Summarization and Analysis
abstract
We are building an interactive visual text analysis tool that aids users in analyzing large collections of text. Unlike existing work in visual text analytics, which focuses either on developing sophisticated text analytic techniques or inventing novel text visualization metaphors, ours tightly integrates state-of-the-art text analytics with interactive visualization to maximize the value of both. In this article, we present our work from two aspects. We first introduce an enhanced, LDA-based topic analysis technique that automatically derives a set of topics to summarize a collection of documents and their content evolution over time. To help users understand the complex summarization results produced by our topic analysis technique, we then present the design and development of a time-based visualization of the results. Furthermore, we provide users with a set of rich interaction tools that help them further interpret the visualized results in context and examine the text collection from multiple perspectives. As a result, our work offers three unique contributions. First, we present an enhanced topic modeling technique to provide users with a time-sensitive and more meaningful text summary. Second, we develop an effective visual metaphor to transform abstract and often complex text summarization results into a comprehensible visual representation. Third, we offer users flexible visual interaction tools as alternatives to compensate for the deficiencies of current text summarization techniques. We have applied our work to a number of text corpora and our evaluation shows promise, especially in support of complex text analyses.
Shixia Liu, Michelle X. Zhou, Shimei Pan, Yangqiu Song, Weihong Qian, Weijia Cai, Xiaoxiao Lian
ACM Trans. Intell. Syst. Technol.3
2011 Enterprise blogging in a global context: comparing Chinese and American practices
abstract
We present three studies that compare adoption and appropriation between China and the United States of BlogCentral, an internal blogging tool employed by a large global enterprise. We first analyzed 23 months of usage logs for users in both countries and found that compared to the U.S., Chinese users were much less active, with less activity and sparser user interaction. We then conducted 25 interviews and surveyed 213 bloggers in both countries to understand user motivations and behaviors of corporate blogging in more detail. The results show that Chinese employees use internal blogs for organizing personal work and short term team formation, unlike U.S. users who are driven more by the goal of sharing information to a broad community. The Chinese users were seeking (and not finding) greater social interaction and felt instead a sense of alienation from the global blogging community. We further identify a gap between Chinese users' requirement for local community and the existing global aspect of enterprise blogging tools. Lastly we discuss implications from the data analysis.
Qinying Liao, Shimei Pan, Jennifer C. Lai
CSCW2
2011 Analytic Trails: Supporting Provenance, Collaboration, and Reuse for Visual Data Analysis by Business Users
Jie Lu 0002, Shimei Pan, Jennifer C. Lai
INTERACT (4)3
2011 Information at your fingertips: contextual IR in enterprise email
abstract
We present ICARUS, a contextual information retrieval system, which uses the current email message and a multi-tiered user model to retrieve relevant content and make it available in a sidebar widget embedded in the email client. The system employs a dynamic retrieval strategy to conduct automated contextual search across multiple information sources including the user's hard drive, online documents (wikis, blogs and files) and other email messages. It also presents the user with information about the sender of the current message, which varies in detail and degree based on how often the user interacts with this sender. We conducted a formative evaluation which compared three retrieval methods that used different context information: current message plus a multi-tiered user model; current message plus a single-tiered, aggregate user model; and lastly, cur-rent message only. Results indicate that the multi-tiered user modeling approach yields better retrieval performance than the other two. In addition, the study suggests that dynamically determining which sources to search, what query parameters to use, and how to filter/re-rank results can further improve the effectiveness of contextual IR.
Jie Lu 0002, Shimei Pan, Jennifer C. Lai
IUI2
2010 Constrained Coclustering for Textual Documents
abstract
In this paper, we present a constrained co-clustering approach for clustering textual documents. Our approach combines the benefits of information-theoretic co-clustering and constrained clustering. We use a two-sided hidden Markov random field (HMRF) to model both the document and word constraints. We also develop an alternating expectation maximization (EM) algorithm to optimize the constrained co-clustering model. We have conducted two sets of experiments on a benchmark data set: (1) using human-provided category labels to derive document and word constraints for semi-supervised document clustering, and (2) using automatically extracted named entities to derive document constraints for unsupervised document clustering. Compared to several representative constrained clustering and co-clustering approaches, our approach is shown to be more effective for high-dimensional, sparse text data.
Yangqiu Song, Shimei Pan, Shixia Liu, Furu Wei, Michelle X. Zhou, Weihong Qian
AAAI2
2010 Natural Language Aided Visual Query Building for Complex Data Access
abstract
Over the past decades, there have been significant efforts on developing robust and easy-to-use query interfaces to databases. So far, the typical query interfaces are GUI-based visual query interfaces. Visual query interfaces however, have limitations especially when they are used for accessing large and complex datasets. Therefore, we are developing a novel query interface where users can use natural language expressions to help author visual queries. Our work enhances the usability of a visual query interface by directly addressing the "knowledge gap" issue in visual query interfaces. We have applied our work in several real-world applications. Our preliminary evaluation demonstrates the effectiveness of our approach
Shimei Pan, Michelle X. Zhou, Keith Houck, Peter Kissa
IAAI1
2010 TIARA: a visual exploratory text analytic system
abstract
In this paper, we present a novel exploratory visual analytic system called TIARA (Text Insight via Automated Responsive Analytics), which combines text analytics and interactive visualization to help users explore and analyze large collections of text. Given a collection of documents, TIARA first uses topic analysis techniques to summarize the documents into a set of topics, each of which is represented by a set of keywords. In addition to extracting topics, TIARA derives time-sensitive keywords to depict the content evolution of each topic over time. To help users understand the topic-based summarization results, TIARA employs several interactive text visualization techniques to explain the summarization results and seamlessly link such results to the original text. We have applied TIARA to several real-world applications, including email summarization and patient record analysis. To measure the effectiveness of TIARA, we have conducted several experiments. Our experimental results and initial user feedback suggest that TIARA is effective in aiding users in their exploratory text analytic tasks.
Furu Wei, Shixia Liu, Yangqiu Song, Shimei Pan, Michelle X. Zhou, Weihong Qian, Lei Shi 0002
KDD4
2009 Interactive, topic-based visual text summarization and analysis
abstract
We are building an interactive, visual text analysis tool that aids users in analyzing a large collection of text. Unlike existing work in text analysis, which focuses either on developing sophisticated text analytic techniques or inventing novel visualization metaphors, ours is tightly integrating state-of-the-art text analytics with interactive visualization to maximize the value of both. In this paper, we focus on describing our work from two aspects. First, we present the design and development of a time-based, visual text summary that effectively conveys complex text summarization results produced by the Latent Dirichlet Allocation (LDA) model. Second, we describe a set of rich interaction tools that allow users to work with a created visual text summary to further interpret the summarization results in context and examine the text collection from multiple perspectives. As a result, our work offers two unique contributions. First, we provide an effective visual metaphor that transforms complex and even imperfect text summarization results into a comprehensible visual summary of texts. Second, we offer users a set of flexible visual interaction tools as the alternatives to compensate for the deficiencies of current text summarization techniques. We have applied our work to a number of text corpora and our evaluation shows the promise of the work, especially in support of complex text analyses.
Shixia Liu, Michelle X. Zhou, Shimei Pan, Weihong Qian, Weijia Cai, Xiaoxiao Lian
CIKM3
2009 Topic and keyword re-ranking for LDA-based topic modeling
abstract
Topic-based text summaries promise to help average users quickly understand a text collection and derive insights. Recent research has shown that the Latent Dirichlet Allocation (LDA) model is one of the most effective approaches to topic analysis. However, the LDA-based results may not be ideal for human understanding and consumption. In this paper, we present several topic and keyword re-ranking approaches that can help users better understand and consume the LDA-derived topics in their text analysis. Our methods process the LDA output based on a set of criteria that model a user's information needs. Our evaluation demonstrates the usefulness of the methods in summarizing several large-scale, real world data sets.
Yangqiu Song, Shimei Pan, Shixia Liu, Michelle X. Zhou, Weihong Qian
CIKM2
2007 Natural Language Query Recommendation in Conversation Systems
Shimei Pan, James Shaw
IJCAI1
2007 Building ubiquitous and robust speech and natural language interfaces
Gary Geunbae Lee, Shimei Pan
IUI2
2006 Responsive Information Architect: Enabling Context-Sensitive Information Seeking
Michelle X. Zhou, Keith Houck, Shimei Pan, James Shaw, Vikram Aggarwal
AAAI3
2006 Enabling context-sensitive information seeking
abstract
Information seeking is an important but often difficult task, especially when it involves large and complex data sets. We hypothesize that a context-sensitive interaction paradigm would greatly assist users in their information seeking. Such a paradigm would allow users to both express their requests and receive requested information in context. Driven by this hypothesis, we have taken rigorous steps to design, develop, and evaluate a full-fledged, context-sensitive information system. We started with a Wizard-of-OZ (WOZ) study to verify the effectiveness of our envi-sioned system. We then built a fully automated system based on the findings from our WOZ study. We targeted the development and integration of two sets of technologies: context-sensitive mul-timodal input interpretation and multimedia output generation. Finally, we formally evaluated the usability of our system in real world conditions. The results show that our system greatly improves the users' ability to perform practical information-seek-ing tasks. These results not only confirm our initial hypothesis, but they also indicate the practicality of our approaches.
Michelle X. Zhou, Keith Houck, Shimei Pan, James Shaw, Vikram Aggarwal
IUI3
2005 Instance-based Sentence Boundary Determination by Optimization for Natural Language Generation
abstract
This paper describes a novel instance-based sentence boundary determination method for natural language generation that optimizes a set of criteria based on examples in a corpus. Compared to existing sentence boundary determination approaches, our work offers three significant contributions. First, our approach provides a general domain independent framework that effectively addresses sentence boundary determination by balancing a comprehensive set of sentence complexity and quality related constraints. Second, our approach can simulate the characteristics and the style of naturally occurring sentences in an application domain since our solutions are optimized based on their similarities to examples in a corpus. Third, our approach can adapt easily to suit a natural language generation system's capability by balancing the strengths and weaknesses of its subcomponents (e.g. its aggregation and referring expression generation capability). Our final evaluation shows that the proposed method results in significantly better sentence generation outcomes than a widely adopted approach.
Shimei Pan, James Shaw
ACL1
2005 Two-way adaptation for robust input interpretation in practical multimodal conversation systems
abstract
Multimodal conversation systems allow users to interact with computers effectively using multiple modalities, such as natural language and gesture. However, these systems have not been widely used in practical applications mainly due to their limited input understanding capability. As a result, conversation systems often fail to understand user requests and leave users frustrated. To address this issue, most existing approaches focus on improving a system's interpretation capability. Nonetheless, such improvements may still be limited, since they would never cover the entire range of input expressions. Alternatively, we present a two-way adaptation framework that allows both users and systems to dynamically adapt to each other's capability and needs during the course of interaction. Compared to existing methods, our approach offers two unique contributions. First, it improves the usability and robustness of a conversation system by helping users to dynamically learn the system's capabilities in context. Second, our approach enhances the overall interpretation capability of a conversation system by learning new user expressions on the fly. Our preliminary evaluation shows the promise of this approach.
Shimei Pan, Siwei Shen, Michelle X. Zhou, Keith Houck
IUI1
2004 Responsive Information Architect: A Context-Sensitive Multimedia Conversation Framework for Information Seeking
Michelle X. Zhou, Keith Houck, Rosario Uceda-Sosa, Shimei Pan, Vikram Aggarwal, James Shaw
AAAI4
2004 SEGUE: A Hybrid Case-Based Surface Natural Language Generator
Shimei Pan, James Shaw
INLG1
2004 A multi-layer conversation management approach for information seeking applications
Shimei Pan
INTERSPEECH1
2002 Context-Based Multimodal Input Understanding in Conversational Systems
abstract
In a multimodal human-machine conversation, user inputs are often abbreviated or imprecise. Sometimes, merely fusing multimodal inputs together cannot derive a complete understanding. To address these inadequacies, we are building a semantics-based multimodal interpretation framework called MIND (Multimodal Interpretation for Natural Dialog). The unique feature of MIND is the use of a variety of contexts (e.g., domain context and conversation context) to enhance multimodal fusion. In this paper we present a semantically rich modeling scheme and a context-based approach that enable MIND to gain a full understanding of user inputs, including ambiguous and incomplete ones.
Joyce Y. Chai, Shimei Pan, Michelle X. Zhou, Keith Houck
ICMI2
2002 Designing a Speech Corpus for Instance-based Spoken Language Generation
Shimei Pan, Wubin Weng
INLG1
2002 Exploring features from natural language generation for prosody modeling
Shimei Pan, Kathy McKeown, Julia Hirschberg
Comput. Speech Lang.1
2002 Designing and Evaluating an Adaptive Spoken Dialogue System
Diane J. Litman, Shimei Pan
User Model. User Adapt. Interact.2
2001 Semantic abnormality and its realization in spoken language
abstract
In this paper we investigate the relationship between various lexical and prosodic features and semantic abnormality, the occurrence of unusual or unexpected events, in generating speech for MAGIC, which employs a Concept-to-Speech system to generate post-operative reports for patients who have undergone bypass surgery. Using the speech corpus collected for this application, we conducted empirical analysis to systematically discover significantly correlated prosodic and lexical features. The automatically learned abnormality model not only can be used in building comprehensive prosody prediction systems for Concept-to-Speech generation, but also help identify unusual information during speech analysis and understanding.
Shimei Pan, Kathy McKeown, Julia Hirschberg
INTERSPEECH1
2001 Automated authoring of coherent multimedia discourse in conversation systems
abstract
We are building a full-fledged multimedia conversation framework called Responsive Information Architect (RIA), using a combination of AI and multimedia techniques. Here we describe RIA's capability of automated authoring of a coherent multimedia discourse, which is used by RIA to express itself when conversing with a user. Specifically, we focus on explaining three unique features of our automated authoring approach: automated authoring of multimedia interaction acts, dynamic insertion of multimedia punctuation acts, and systematic design of cross-media acts.
Michelle X. Zhou, Shimei Pan
ACM Multimedia2
2000 Modeling Local Context for Pitch Accent Prediction
abstract
Pitch accent placement is a major topic in intonational phonology research and its application to speech synthesis. What factors influence whether or not a word is made intonationally prominent or not is an open question. In this paper, we investigate how one aspect of a word's local context --- its collocation with neighboring words --- influences whether it is accented or not. Results of experiments on two transcribed speech corpora in a medical domain show that such collocation information is a useful predictor of pitch accent placement.
Shimei Pan, Julia Hirschberg
ACL1
1999 Word Informativeness and Automatic Pitch Accent Modeling
Shimei Pan, Kathy McKeown
EMNLP1
1996 Spoken language generation in a multimedia system
Shimei Pan, Kathy McKeown
ICSLP1
1996 Negotiation for Automated Generation of Temporal Multimedia Presentations
abstract
Creating high-quality multimedia presentations requires much skill, time, and effort.This is particularly true when temporal media, such as speech and animation, are involved.We describe the design and implementation of a knowledge-based system that generates customized temporal multimedia presentations.We provide art overview of the system's architecture, and explain how speech, written text, and graphics are generated and coordinated.Our emphasis is on how temporal media are coordinated by the system through a multi-stage negotiation process.In negotiation, media-specific generation components interact with a novel coordination component that solves temporal constraints provided by the generators.We illustrate our work with a set of examples generated by the system in a testbed application intended to update hospital caregivers on the status of patients who have undergone a cardiac bypass operation.
Mukesh Dalal, Steven K. Feiner, Kathy McKeown, Shimei Pan, Michelle X. Zhou, Tobias Höllerer, James Shaw, Jeanne C. Fromer
ACM Multimedia4
1992 Knowledge Acquisition And Chinese Parsing Based On Corpus
Chunfa Yuan, Changning Huang, Shimei Pan
COLING3