Lisa Singh

dblp:80/3925 · DBLP profile ↗
← Back
44ranked-venue papers in the field
12as first author
16since 2021 · last 2024
0000-0002-8300-2970ORCID · corroborated

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 28 (7 first)Database Systems & Data Management · 7 (2 first)Information Retrieval & Web Search · 5 (1 first)Big Data, Cloud & Distributed Data Systems · 4 (2 first)
YearPublicationVenuePosition
2024 It is Time to Develop an Auditing Framework to Promote Value Aware Chatbots
abstract
The launch of ChatGPT in November 2022 marked the beginning of a new era in AI, the availability of generative AI tools for everyone to use. ChatGPT and other similar chatbots boast a wide range of capabilities from answering student homework questions to creating music and art. Given the large amounts of human data chatbots are built on, it is inevitable that they will inherit human errors and biases. These biases have the potential to inflict significant harm or increase inequity on different subpopulations. Because chatbots do not have an inherent understanding of societal values, they may create new content that is contrary to established norms. Examples of concerning generated content includes child pornography, inaccurate facts, and discriminatory posts. In this position paper, we argue that the speed of advancement of this technology requires us, as computer and data scientists, to mobilize and develop a values-based auditing framework containing a community established standard set of measurements to monitor the health of different chatbots and LLMs. To support our argument, we use a simple audit template to share the results of basic audits we conduct that are focused on measuring potential bias in search engine style tasks, code generation, and story generation. We identify responses from GPT 3.5 and GPT 4 that are both consistent and not consistent with values derived from existing law. While the findings come as no surprise, they do underscore the urgency of developing a robust auditing framework for openly sharing results in a consistent way so that mitigation strategies can be developed by the academic community, government agencies, and companies when our values are not being adhered to. We conclude this paper with recommendations for value-based strategies for improving the technologies.
Yanchen Wang, Lisa Singh
DATA2
2024 Let Me Generate That for You: Generative Data Augmentation for Misinformation Detection in Low-Resource Environments
abstract
Misinformation detection is a rapidly moving target, as new topics emerge and evolve in high volume on social media platforms. Annotated and fact-checked datasets are necessary for detection model training, but are laborious to curate. Thus, many misinformation detection models are trained in low-resource environments and rely on machine learning techniques to improve performance with small ground-truth datasets. Generative data augmentation methods enable topic-specific examples that increase a model's training dataset without incurring the cost and time investment associated with manual annotation. In this work, we assess the value of using generative augmentation for different classes of learning models: a classic neural model, a fine-tuned deep learning model, a reinforcement learning model, and an active learning model. We find that generated training data is not effective for all learning paradigms for the misinformation detection task, highlighting the need to use different quality measures to assess its value for low-resource machine learning tasks.
Autumn Toney, Lisa Singh
DSAA2
2024 Intermedia Agenda Setting during the 2016 and 2020 U.S. Presidential Elections
abstract
Intermedia agenda setting (IAS) theory suggests that different news sources can influence each other's agenda. While this theory has been well-established in existing literature, whether it still holds in today's high-choice media environment, which includes news producers of different credibility and ideology dispositions, is an open question. Through two case studies--the 2016 and 2020 U.S. presidential elections--we show that media are still largely aligned, especially in broad topics they choose to cover, and that the level of alignment along the credibility dimension is comparable to that along the ideology dimension. Furthermore, we find that the coverage of the Republican candidate is better aligned across different media types than that of the Democratic candidate, and that media divergence has increased along both dimensions from 2016 to 2020. Finally, we demonstrate that high-credibility media still plays a dominant role in the IAS process, yet with a cautious warning of its declining IAS power for the Democratic candidate over the course of four years.
Yaguang Liu, Lisa Singh, Ceren Budak
ICWSM3
2023 Identifying High-Quality Training Data for Misinformation Detection
Jaren Haber, Kornraphop Kawintiranon, Lisa Singh, Alexander Chen, Aidan Pizzo, Anna Pogrebivsky, Joyce Yang
DATA3
2023 All Translation Tools Are Not Equal: Investigating the Quality of Language Translation for Forced Migration
abstract
As the volume and complexity of forced movement continues to grow, there is an urgent need to use new data sources to better understand emerging crises. Organic sources, like social media and newspapers, can offer insights in near real time when administrative data are unavailable for timely and detailed analysis. However, in order to flexibly switch to different contexts, we need the ability to contextualize the drivers of movement for different locations and languages. Recent advances in natural language processing and specifically, neural machine translation, have shown impressive results on standard benchmark datasets for well-studied language pairs. However, the effectiveness of these models in a real-world scenario remains less known. To advance our understanding of real-world, contextual translation, we systematically study the performance of multiple widely used off-the-shelf machine translation tools using words associated with drivers of forced movement in both high- and low-resource languages. Our empirical results suggest significant variation between the performance of these machine translation tools in terms of accuracy and efficiency, highlighting a problem that must be faced by those conducting migration research using multilingual contexts. We conclude by suggesting strategies for obtaining reasonable translations from off-the-shelf language tools.
Ameeta Agrawal, Lisa Singh, Elizabeth Jacobs, Yaguang Liu, Gwyneth Dunlevy, Rhitabrat Pokharel, Varun Uppala
DSAA2
2023 Combining vs. Transferring Knowledge: Investigating Strategies for Improving Demographic Inference in Low Resource Settings
abstract
For some learning tasks, generating a large labeled data set is impractical. Demographic inference using social media data is one such task. While different strategies have been proposed to mitigate this challenge, including transfer learning, data augmentation, and data combination, they have not been explored for the task of user level demographic inference using social media data. This paper explores two of these strategies: data combination and transfer learning. First, we combine labeled training data from multiple data sets of similar size to understand when the combination is valuable and when it is not. Using data set distance, we quantify the relationship between our data sets to help explain the performance of the combination strategy. Then, we consider supervised transfer learning, where we pretrain a model on a larger labeled data set, fine-tune the model on smaller data sets, and incorporate regularization as part of the transfer learning process. We empirically show the strengths and limitations of the proposed techniques on multiple Twitter data sets.
Yaguang Liu, Lisa Singh
WSDM2
2023 Using topic-noise models to generate domain-specific topics across data sources
Rob Churchill, Lisa Singh
Knowl. Inf. Syst.2
2022 Students or Mechanical Turk: Who Are the More Reliable Social Media Data Labelers?
Lisa Singh, Rebecca Vanarsdall, Yanchen Wang, Carole Roan Gresenz
DATA1
2022 Inferring #MeToo Experience Tweets using Classic and Neural Models
Julianne Zech, Lisa Singh, Kornraphop Kawintiranon, Naomi Mezey, Jamillah Williams
DATA2
2022 Dynamic Topic-Noise Models for Social Media
Rob Churchill, Lisa Singh
PAKDD (2)2
2022 DeMis: Data-Efficient Misinformation Detection Using Reinforcement Learning
Kornraphop Kawintiranon, Lisa Singh
ECML/PKDD (2)2
2022 A Guided Topic-Noise Model for Short Texts
abstract
Researchers using social media data want to understand the discussions occurring in and about their respective fields. These domain experts often turn to topic models to help them see the entire landscape of the conversation, but unsupervised topic models often produce topic sets that miss topics experts expect or want to see. To solve this problem, we propose Guided Topic-Noise Model (GTM), a semi-supervised topic model designed with large domain-specific social media data sets in mind. The input to GTM is a set of topics that are of interest to the user and a small number of words or phrases that belong to those topics. These seed topics are used to guide the topic generation process, and can be augmented interactively, expanding the seed word list as the model provides new relevant words for different topics. GTM uses a novel initialization and a new sampling algorithm called Generalized Polya Urn (GPU) seed word sampling to produce a topic set that includes expanded seed topics, as well as new unsupervised topics. We demonstrate the robustness of GTM on open-ended responses from a public opinion survey and four domain-specific Twitter data sets.
Rob Churchill, Lisa Singh, Rebecca Ryan, Pamela Davis-Kean
WWW2
2021 Text Analytic Research Portals: Supporting Large-Scale Social Science Research
abstract
Large-scale organic data generated from newspapers, social media, television, and radio require an expertise in infrastructure management, data collection, and data processing in order to gain research value from them. We have developed text analytic research portals to help social science researchers who do not have the resources necessary to collect, store, and process these large-scale data sets. Our portals allow researchers to use an intuitive point and click interface to generate variables from large, dynamic data sets using state of the art text mining and learning methods. These timely variables constructed from noisy text can then be used to advance social science research in areas such as political science, economics, public health, and psychology research.
Lisa Singh, Colton Padden, Pamela Davis-Kean, Rabin David, Virinche Marwadi, Yiqing Ren, Rebecca Vanarsdall
IEEE BigData1
2021 Age Inference Using A Hierarchical Attention Neural Network
abstract
While demographic attributes, such as age, gender, and location, have been extensively studied, most previous studies usually combine different sources of data, such as the user's biography, pictures, posts, and the user's network to obtain reasonable inference accuracies. However, it is not always practical to collect all those different forms of data. Therefore, in this paper, we consider methods for inferring age that only use Twitter posts (tweet text and emojis). We propose a hierarchical attention neural model that integrates independent linguistic knowledge gained from text and emojis when making a prediction. This hierarchical model is able to capture the intra-post relationship between these different post components, as well as the inter-post relationships of a user's posts. Our empirical evaluation using a data set generated from Wikidata demonstrates that our model achieves better performance than the state-of-the-art models, and still performs well when the number of posts per user is reduced in the training data set.
Yaguang Liu, Lisa Singh
CIKM2
2021 textPrep: A Text Preprocessing Toolkit for Topic Modeling on Social Media Data
Rob Churchill, Lisa Singh
DATA2
2021 Topic-Noise Models: Modeling Topic and Noise Distributions in Social Media Post Collections
abstract
Most topic models define a document as a mixture of topics and each topic as a mixture of words. Generally, the difference in generative topic models is how these mixtures of topics are generated. We propose looking at topic models in a new way, as topic-noise models. Our topic-noise model defines a document as a mixture of topics and noise. Topic Noise Discriminator (TND) estimates both the topic and noise distributions using not only the relationships between words in documents, but also the linguistic relationships found using word embeddings. This type of model is important for short, sparse social media posts that contain both random and non-random noise. We also understand that topic quality is subjective and that researchers may have preferences. Therefore, we propose a variant of our model that combines the pre-trained noise distribution from TND in an ensemble with any generative topic model to filter noise words and produce more coherent and diverse topic sets. We present this approach using Latent Dirichlet Allocation (LDA) and show that it is effective for maintaining high quality LDA topics while removing noise within them. Finally, we show the value of using a context-specific noise list generated from TND to remove noise statically, after topics have been generated by any topic model, including non-generative ones. We demonstrate the effectiveness of all three of these approaches that explicitly model context-specific noise in document collections.
Rob Churchill, Lisa Singh
ICDM2
2020 Information Exposure From Relational Background Knowledge on Social Media
abstract
While some users share large amounts of information, others share very little. However, even with limited amounts of sharing, users may still have high levels of exposure. Previous research has shown that for certain attributes like gender, adversaries can determine a target's hidden attribute value by taking a majority vote of its community or by finding others in the site population with similar profiles. However, for some attributes, these attacks fail because of the diversity of the attribute value in the community. In this paper, we present a new privacy attack - a relational background attack (RBA), where an adversary builds inference models for a hidden attribute of the target by using the target's relational background set. Doing this allows the adversary to build a "biased" model that captures the significant local features for inferring the hidden attribute. We empirically demonstrate the effectiveness of this attack on a special case of the relational background set (a local community) using a Twitter data set. We then consider the case when an adversary only has access to different subsets of the target's local community, and show that the attack can still be conducted effectively with certain approximations of the target's local community.
Shuo Liu 0011, Lisa Singh, Kevin Tian
DSAA2
2019 Exploring the Relationship Between Conversation Using #MeToo and University Harassment Policies
abstract
While identifying those who are most vocal on social media movements can be straight-forward, finding hidden groups can be challenging. This poster presents a case study focused on the relationship between mentions of universities in the #MeToo Twitter conversation and policies universities have implemented with regards to harassment and assault. Preliminary results suggest that there is variation in terms of policies, resources and responses to sexual misconduct across campuses and that there is also variation in the number of mentions of different universities. However, there is not a clear relationship between policies and online discussion involving universities.
Julianne Zech, Fransiska Dale, Lisa Singh, Jamillah Williams, Naomi Mezey
DSAA3
2019 Blending Noisy Social Media Signals with Traditional Movement Variables to Predict Forced Migration
abstract
Worldwide displacement due to war and conflict is at all-time high. Unfortunately, determining if, when, and where people will move is a complex problem. This paper proposes integrating both publicly available organic data from social media and newspapers with more traditional indicators of forced migration to determine when and where people will move. We combine movement and organic variables with spatial and temporal variation within different Bayesian models and show the viability of our method using a case study involving displacement in Iraq. Our analysis shows that incorporating open-source generated conversation and event variables maintains or improves predictive accuracy over traditional variables alone. This work is an important step toward understanding how to leverage organic big data for societal--scale problems.
Lisa Singh, Laila Wahedi, Yanchen Wang, Yifang Wei, Christo Kirov, Susan Martin, Katharine M. Donato, Yaguang Liu, Kornraphop Kawintiranon
KDD1
2018 Detecting and Using Buzz from Newspapers to Understand Patterns of Movement
abstract
Meaningful leading indicators of mass movement are difficult to discover given the dearth of available data about involuntary movement. As a first step, we propose analyzing whether we can use the changing dynamics of newspaper content as one possible indirect indicator of such displacement. Specifically, we explore whether news media buzz correlates with patterns of migration in Iraq. We consider different methods for detecting buzz and empirically evaluate them on a corpus of 1.4 million articles.
Julia Hocket, Yaguang Liu, Yifang Wei, Lisa Singh, Nathan Schneider 0001
IEEE BigData4
2018 A Temporal Topic Model for Noisy Mediums
Rob Churchill, Lisa Singh, Christo Kirov
PAKDD (2)2
2017 EOS: A multilingual text archive of international newspaper & blog articles
abstract
The Expandable Open Source (EOS) database maintains an archive of Internet accessible newspaper and blog articles from across the world. Today, the archive contains over 700 million articles and is growing by approximately 100,000 articles each day. In this work, we describe the components of EOS, our community of data users and data providers, some real world use cases, challenges associated with maintaining this archive, and a vision for its future as a text analytic portal.
Lisa Singh, Raghu Pemmaraju
IEEE BigData1
2017 Understanding the impact of sampling and noise on detecting events using twitter
abstract
While social media sites can be used to identify events rapidly, many data streams are partial because of rate limiting while others are large, but particularly noisy. This poster investigates the impact of sample size and noise on event detection accuracy. We conduct a sensitivity analysis to understand how robust different methods for event detection on Twitter are, given a noisy, partial data stream. We find that the detection accuracy decreases as the sample fraction decreases and as the SNR decreases; however, the rate of decrease changes at different sample sizes and SNRs for different methods.
Yifang Wei, Lisa Singh
IEEE BigData2
2017 Using Network Flows to Identify Users Sharing Extremist Content on Social Media
Yifang Wei, Lisa Singh
PAKDD (1)2
2016 ASONAM 2016 panel: Social network analysis for social good
abstract
No abstract or record of the panel discussion was made available for publication as part of the conference proceedings.
V. S. Subrahmanian, Lada A. Adamic, Lise Getoor, Evimaria Terzi, Brian Uzzi, Lisa Singh
ASONAM6
2016 Identification of extremism on Twitter
abstract
Identifying extremist-associated conversations on Twitter is an open problem. Extremist groups have been leveraging Twitter (1) to spread their message and (2) to gain recruits. In this paper, we investigate the problem of determining whether a particular Twitter user engages in extremist conversation. We explore different Twitter metrics as proxies for misbehavior, including the sentiment of the user's published tweets, the polarity of the user's ego-network, and user mentions. We compare different known classifiers using these different features on manually annotated tweets involving the ISIS extremist group and find that combining all these features leads to the highest accuracy for detecting extremism on Twitter.
Yifang Wei, Lisa Singh, Susan Martin
ASONAM2
2016 Generating risk reduction recommendations to decrease vulnerability of public online profiles
abstract
Preserving online privacy is becoming increasingly challenging due in large part to the continued growth of social media. Those who choose to share their information publicly may not realize what features of their profiles make their public data more identifiable and potentially vulnerable to cross-site record linkage. This paper proposes a risk reduction recommendation method that suggests removal or modification of a small number of attributes to make a profile less unique, thereby reducing the identifiability and vulnerability of the user. Empirical results on data collected from Google+, LinkedIn, and Foursquare show that users' vulnerability in terms of identifiability and data exposure level can be significantly reduced while public profile utility can be maintained using our proposed approach.
Janet Zhu, Sicong Zhang, Lisa Singh, Grace Hui Yang, Micah Sherr
ASONAM3
2016 Overlapping Target Event and Story Line Detection of Online Newspaper Articles
abstract
Event detection from text data is an active area of research. While the emphasis has been on event identification and labeling using a single data source, this work considers event and story line detection when using a large number of data sources. In this setting, it is natural for different events in the same domain, e.g. violence, sports, politics, to occur at the same time and for different story lines about the same event to emerge. To capture events in this setting, we propose an algorithm that detects events and story lines about events for a target domain. Our algorithm leverages a multi-relational sentence level semantic graph and well known graph properties to identify overlapping events and story lines within the events. We evaluate our approach on two large data sets containing millions of news articles from a large number of sources. Our empirical analysis shows that our approach improves the detection precision and recall by 10% to 25%, while providing complete event summaries.
Yifang Wei, Lisa Singh, Brian Gallagher, David Buttler
DSAA2
2016 Anonymizing Query Logs by Differential Privacy
abstract
Query logs are valuable resources for Information Retrieval (IR) research. However, because they are also rich in private and personal information, the huge concern of leaking user privacy prevents query logs from being shared from the search companies to the broad research community. Bothered by the lack of good research data for years, the authors of this paper are motivated to explore ways to generate anonymized query logs that can still be effectively used to support the search task. We introduce a framework to anonymize query logs by differential privacy, the latest development in privacy research. The framework is empirically evaluated against multiple search algorithms on their retrieval utility, measured in standard IR evaluation metrics, using the anonymized logs. The experiments show that our framework is able to achieve a good balance between retrieval utility and privacy.
Sicong Zhang, Grace Hui Yang, Lisa Singh
SIGIR3
2015 Public Information Exposure Detection: Helping Users Understand Their Web Footprints
abstract
To help users better understand the potential risks associated with publishing data publicly, as well as the quantity and sensitivity of information that can be obtained by combining data from various online sources, we introduce a novel information exposure detection framework that generates and analyzes the web footprints users leave across the social web. Web footprints are the traces of one's online social activities represented by a set of attributes that are known or can be inferred with a high probability by an adversary who has basic information about a user from his/her public profiles. Our framework employs new probabilistic operators, novel pattern-based attribute extraction from text, and a population-based inference engine to generate web footprints. Using a web footprint, the framework then quantifies a user's level of information exposure relative to others with similar traits, as well as with regard to others in the population. Evaluation over public profiles from multiple sites (Google+, LinkeIn, FourSquare, and Twitter) shows that the proposed framework effectively detects and quantifies information exposure using a small amount of initial knowledge.
Lisa Singh, Grace Hui Yang, Micah Sherr, Andrew Hian-Cheong, Kevin Tian, Janet Zhu, Sicong Zhang
ASONAM1
2014 Membership Detection Using Cooperative Data Mining Algorithms
abstract
More and more companies are providing data mining and analytics solutions to customers using social media data. The general approach taken by these companies is to continually collect data from social media sites and then use the collected snapshot of the content for a data mining or analytics task. Unfortunately, given the exponential increase in the volume of social media data, building local database snapshots and running computationally expensive algorithms is not always plausible. As an alternative to the centralized approach, in this paper, we study the feasibility of cooperative algorithms where data never leaves the mined social media network, and instead the network users themselves work together, using only the communication primitives provided by the social media site, to solve data mining problems. While cooperative algorithms can be built for many different data mining tasks, to show the viability of this approach, we focus on a task fundamental to many different social mining applications - membership detection (an individual using the social media site wants to efficiently get a request to a member of a known group with unknown membership). Using Twitter as our specific social graph, we seek cooperative algorithms that solve this problem with high probability even when we assume only a small fraction of the Twitter network participates and we enforce a bound on the number of tweets generated. After validating the potential of cooperative solutions on Twitter, we empirically evaluate a collection of cooperative strategies on a snapshot of the Twitter network containing over 50 million users. Our best solution, which we call brokered token passing, can reliably and efficiently detect group membership while requiring only a small number of tweets be sent and a small percentage of users participate.
Calvin C. Newport, Lisa Singh, Yiqing Ren
SDM2
2013 Understanding evolving group structures in time-varying networks
abstract
This paper presents a framework for identifying persistent groups and individuals across multiple time granularities in dynamic graphs. Understanding the longevity of groups and the relevance of individuals within a group is important in many fields, including sociology, biology, economics, psychology, and political science. Different clustering algorithms have been proposed for static and dynamic graphs. However, using the clustering results to understand the changing dynamics of groups can be difficult. In order to better understand how groups evolve and the level of cohesion within these groups, we propose a holistic dynamic clustering framework that allows the user to adjust the underlying algorithms for clustering nodes in a graph that changes over time and then use the final clusters to produce a time hierarchy that highlights the groups and individuals persistent during different time periods. We test our framework and algorithm both on synthetic and real world data. Our findings indicate that our approach not only yields highly accurate results, but also detects unexpected variations in group structure.
Paul Caravelli, Yifang Wei, Daniel Subak, Lisa Singh, Janet Mann
ASONAM4
2013 Comparison Queries for Uncertain Graphs
Denis Dimitrov, Lisa Singh, Janet Mann
DEXA (2)2
2012 EWNI: Efficient Anonymization of Vulnerable Individuals in Social Networks
Frank Nagle, Lisa Singh, Aris Gkoulalas-Divanis
PAKDD (2)2
2012 SHARD: A Framework for Sequential, Hierarchical Anomaly Ranking and Detection
Jason Robinson, Margaret Lonergan, Lisa Singh, Allison Candido, Mehmet Sayal
PAKDD (2)3
2009 Can Friends Be Trusted? Exploring Privacy in Online Social Networks
abstract
In this paper, we present a case study describing the privacy and trust that exist within a small population of online social network users. We begin by formally characterizing different graphs in social network sites like Facebook. We then determine how often people are willing to divulge personal details to an unknown online user, an adversary. While most users in our sample did not share sensitive information when asked by an adversary, we found that more users were willing to divulge personal details to an adversary if there is a mutual friend connected to the adversary and the user. We then summarize the results and observations associated with this Facebook case study.
Frank Nagle, Lisa Singh
ASONAM2
2009 The Dynamics of Actor Loyalty to Groups in Affiliation Networks
abstract
In this paper, we introduce a method for analyzing the temporal dynamics of affiliation networks. We define affiliation groups which describe temporally related subsets of actors and describe an approach for exploring changing memberships in these affiliation groups over time. To model the dynamic behavior in these networks, we consider the concept of loyalty and introduce a measure that captures an actorpsilas loyalty to an affiliation group as the degree of dasiacommitmentpsila an actor shows to the group over time. We evaluate our measure using two real world affiliation networks: a senate bill co-sponsorship network and a dolphin network. The results show how the behavior of actors in different affiliation groups change dynamically over time, reinforcing the utility of our measure for understanding the loyalty of actors to time-varying affiliation groups.
Hossam Sharara, Lisa Singh, Lise Getoor, Janet Mann
ASONAM2
2009 Privately detecting bursts in streaming, distributed time series data
Lisa Singh, Mehmet Sayal
Data Knowl. Eng.1
2008 Structure-Based Hierarchical Transformations for Interactive Visual Exploration of Social Networks
Lisa Singh, Mitchell Beard, Brian Gopalan, Gregory Nelson
PAKDD1
2007 Privacy Preserving Burst Detection of Distributed Time Series Data Using Linear Transforms
abstract
In this paper, we consider burst detection within the context of privacy. In our scenario, multiple parties want to detect a burst in aggregated time series data, but none of the parties want to disclose their individual data. Our approach calculates bursts directly from linear transform coefficients using a cumulative sum calculation. In order to reduce the chance of a privacy breech, we present multiple data perturbation strategies and compare the varying degrees of privacy preserved. Our strategies do not share raw time series data and still detect significant bursts. We empirically demonstrate this using both real and synthetic distributed data sets. When evaluating both privacy guarantees and burst detection accuracy, we find that our percentage thresholding heuristic maintains a high degree of privacy while accurately identifying bursts of varying widths
Lisa Singh, Mehmet Sayal
CIDM1
2005 Pruning Social Networks Using Structural Properties and Descriptive Attributes
abstract
Scale is often an issue with understanding and making sense of large social networks. Here we investigate methods for pruning social networks by determining the most relevant relationships. We measure importance in terms of predictive accuracy on a set of target attributes of the social network. Our goal is to create a pruned network that models only the most informative affiliations and relationships. We present methods for pruning networks based on both structural properties and descriptive attributes demonstrate it on a network of NASDAQ and NYSE businesses and on a bibliographic network.
Lisa Singh, Lise Getoor, Louis Licamele
ICDM1
1999 An Algorithm for Constrained Association Rule Mining in Semi-structured Data
Lisa Singh, Rebecca Haight, Peter Scheuermann
PAKDD1
1998 A Robust System Architecture for Mining Semi-Structured Data
Lisa Singh, Rebecca Haight, Peter Scheuermann, Kiyoko Aoki
KDD1
1997 Generating Association Rules from Semi-Structured Documents Using an Extended Concept Hierarchy
abstract
Most data mining research has focused on generating rules within databases containing structured values while essentially ignoring the potentially valuable information that exists in the unstructured blocks of text. This paper suggests an approach for generating association rules that relates structured data values to concepts extracted from unstructured data. Our approach involves the use of an extended concept hierarchy (ECH) to maintain parent, child, and sibling relationships between concepts. This structure allows us to generate rules that relate a given concept in the ECH and a given structured attribute value to the neighbors of the given concept in the ECH. We also describe an efficient implementation of the ECH that keeps track of concepts and pointers to documents associated with them. Experimental results on documents from the ABI/Inform Information Retrieval System are presented. 1 Introduction With the abundant amounts of information available to businesses today, an urg...
Lisa Singh, Peter Scheuermann
CIKM1