Acar Tamersoy

dblp:79/9389 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
4since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 5 · 3 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-authorHuman-computer interaction and ubiquitous computing · 5 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2022 Trauma-Informed Computing: Towards Safer Technology Experiences for All
abstract
Trauma is the physical, emotional, or psychological harm caused by deeply distressing experiences. Research with communities that may experience high rates of trauma has shown that digital technologies can create or exacerbate traumatic experiences. Via three vignettes, we discuss how considering the possible effects of trauma and traumatic stress reactions provides an explanatory lens with new insights into people’s technology experiences. Then, we present a framework—trauma-informed computing—in which we adapt and show how to apply six key principles of trauma-informed approaches to computing: safety, trust, peer support, collaboration, enablement, and intersectionality. Through specific examples, we describe how to apply trauma-informed computing in four areas of computing research and practice: user experience research & design, security & privacy, artificial intelligence & machine learning, and organizational culture in tech companies. We conclude by discussing how adopting trauma-informed computing will lead to benefits for all users, not only those experiencing trauma.
Janet X. Chen, Allison McDonald, Yixin Zou, Emily Tseng, Kevin A. Roundy, Acar Tamersoy, Florian Schaub, Thomas Ristenpart, Nicola Dell
CHI6
2021 Towards Stalkerware Detection with Precise Warnings
abstract
Stalkerware enables individuals to conduct covert surveillance on a targeted person’s device. Android devices are a particularly fertile ground for stalkerware, most of which spy on a single communication channel, sensor, or category of private data, though 27% of stalkerware surveil multiple of private data sources. We present Dosmelt, a system that enables stalkerware warnings that precisely characterize the types of surveillance conducted by Android stalkerware so that surveiled individuals can take appropriate mitigating action. Our methodology uses active learning in a semi-supervised learning setting to tackle this task at scale, which would otherwise require expert labeling of significant number of stalkerware apps. Dosmelt leverages the observation that stalkerware differs from other categories of spyware in its open advertising of its surveillance capabilities, which we detect on the basis of the titles and self-descriptions of stalkerware apps that are posted on Android app stores. Dosmelt achieves up to 96% AUC for stalkerware detection with a 91% Macro-F1 score of surveillance capability attribution for stalkerware apps. Dosmelt has detected hundreds of new stalkerware apps that we have added to the Stalkerware Threat List.
Yufei Han 0001, Kevin A. Roundy, Acar Tamersoy
ACSAC3
2021 The Role of Computer Security Customer Support in Helping Survivors of Intimate Partner Violence
Yixin Zou, Allison McDonald, Julia Narakornpichit, Nicola Dell, Thomas Ristenpart, Kevin A. Roundy, Florian Schaub, Acar Tamersoy
USENIX Security Symposium8
2021 Secure and Utility-Aware Data Collection with Condensed Local Differential Privacy
abstract
Local Differential Privacy (LDP) is popularly used in practice for privacy-preserving data collection. Although existing LDP protocols offer high utility for large user populations (100,000 or more users), they perform poorly in scenarios with small user populations (such as those in the cybersecurity domain) and lack perturbation mechanisms that are effective for both ordinal and non-ordinal item sequences while protecting sequence length and content simultaneously. In this paper, we address the small user population problem by introducing the concept of Condensed Local Differential Privacy (CLDP) as a specialization of LDP, and develop a suite of CLDP protocols that offer desirable statistical utility while preserving privacy. Our protocols support different types of client data, ranging from ordinal data types in finite metric spaces (numeric malware infection statistics), to non-ordinal items (OS versions, transaction categories), and to sequences of ordinal and non-ordinal items. Extensive experiments are conducted on multiple datasets, including datasets that are an order of magnitude smaller than those used in existing approaches, which show that proposed CLDP protocols yield high utility. Furthermore, case studies with Symantec datasets demonstrate that our protocols accurately support key cybersecurity-focused tasks of detecting ransomware outbreaks, identifying targeted and vulnerable OSs, and inspecting suspicious activities on infected machines.
Mehmet Emre Gursoy, Acar Tamersoy, Stacey Truex, Wenqi Wei 0001, Ling Liu 0001
IEEE Trans. Dependable Secur. Comput.2
2020 Examining the Adoption and Abandonment of Security, Privacy, and Identity Theft Protection Practices
abstract
Users struggle to adhere to expert-recommended security and privacy practices. While prior work has studied initial adoption of such practices, little is known about the subsequent implementation and abandonment. We conducted an online survey (n=902) examining the adoption and abandonment of 30 commonly recommended practices. Security practices were more widely adopted than privacy and identity theft protection practices. Manual and fully automatic practices were more widely adopted than practices requiring recurring user interaction. Participants' gender, education, technical background, and prior negative experience are correlated with their levels of adoption. Furthermore, practices were abandoned when they were perceived as low-value, inconvenient, or when users overrode them with subjective judgment. We discuss how security, privacy, and identity theft protection recommendations and tools can be better aligned with user needs.
Yixin Zou, Kevin A. Roundy, Acar Tamersoy, Saurabh Shintre, Johann Roturier, Florian Schaub
CHI3
2020 The Many Kinds of Creepware Used for Interpersonal Attacks
abstract
Technology increasingly facilitates interpersonal attacks such as stalking, abuse, and other forms of harassment. While prior studies have examined the ecosystem of software designed for stalking, there exists an unstudied, larger landscape of apps-what we call creepware-used for interpersonal attacks. In this paper, we initiate a study of creepware using access to a dataset detailing the mobile apps installed on over 50 million Android devices. We develop a new algorithm, CreepRank, that uses the principle of guilt by association to help surface previously unknown examples of creepware, which we then characterize through a combination of quantitative and qualitative methods. We discovered apps used for harassment, impersonation, fraud, information theft, concealment, and even apps that purport to defend victims against such threats. As a result of our work, the Google Play Store has already removed hundreds of apps for policy violations. More broadly, our findings and techniques improve understanding of the creepware ecosystem, and will inform future efforts that aim to mitigate interpersonal attacks.
Kevin A. Roundy, Paula Barmaimon Mendelberg, Nicola Dell, Damon McCoy, Daniel Nissani, Thomas Ristenpart, Acar Tamersoy
SP7
2018 VIGOR: Interactive Visual Exploration of Graph Query Results
abstract
Finding patterns in graphs has become a vital challenge in many domains from biological systems, network security, to finance (e.g., finding money laundering rings of bankers and business owners). While there is significant interest in graph databases and querying techniques, less research has focused on helping analysts make sense of underlying patterns within a group of subgraph results. Visualizing graph query results is challenging, requiring effective summarization of a large number of subgraphs, each having potentially shared node-values, rich node features, and flexible structure across queries. We present VIGOR, a novel interactive visual analytics system, for exploring and making sense of query results. VIGOR uses multiple coordinated views, leveraging different data representations and organizations to streamline analysts sensemaking process. VIGOR contributes: (1) an exemplar-based interaction technique, where an analyst starts with a specific result and relaxes constraints to find other similar results or starts with only the structure (i.e., without node value constraints), and adds constraints to narrow in on specific results; and (2) a novel feature-aware subgraph result summarization. Through a collaboration with Symantec, we demonstrate how VIGOR helps tackle real-world problems through the discovery of security blindspots in a cybersecurity dataset with over 11,000 incidents. We also evaluate VIGOR with a within-subjects study, demonstrating VIGOR's ease of use over a leading graph database management system, and its ability to help analysts understand their results at higher speed and make fewer errors.
Robert S. Pienta, Fred Hohman, Alex Endert, Acar Tamersoy, Kevin A. Roundy, Christopher Gates 0002, Shamkant B. Navathe, Polo Chau
IEEE Trans. Vis. Comput. Graph.4
2017 Smoke Detector: Cross-Product Intrusion Detection With Weak Indicators
abstract
The central task of a Security Incident and Event Manager (SIEM) or Managed Security Service Provider (MSSP) is to detect security incidents on the basis of tens of thousands of event types coming from many kinds of security products. We present Smoke Detector, which processes trillions of security events with the Random Walk with Restart (RWR) algorithm, inferring high order relationships between known security incidents and imperfect secondary security events (smoke) to find undiscovered security incidents (fire). By finding previously undetected incidents, Smoke Detector's RWR algorithm is able to increase the MSSP's critical incident count by 19% with a 1.3% FP rate.
Kevin A. Roundy, Acar Tamersoy, Michael Spertus, Michael Hart, Daniel Kats, Matteo Dell'Amico, Robert Scott
ACSAC2
2017 Visual Graph Query Construction and Refinement
abstract
Locating and extracting subgraphs from large network datasets is a challenge in many domains, one that often requires learning new querying languages. We will present the first demonstration of VISAGE, an interactive visual graph querying approach that empowers analysts to construct expressive queries, without writing complex code (see our video: https://youtu.be/l2L7Y5mCh1s). VISAGE guides the construction of graph queries using a data-driven approach, enabling analysts to specify queries with varying levels of specificity, by sampling matches to a query during the analyst's interaction. We will demonstrate and invite the audience to try VISAGE on a popular film-actor-director graph from Rotten Tomatoes.
Robert S. Pienta, Fred Hohman, Acar Tamersoy, Alex Endert, Shamkant B. Navathe, Hanghang Tong, Polo Chau
SIGMOD Conference3
2016 VISAGE: Interactive Visual Graph Querying
abstract
Extracting useful patterns from large network datasets has become a fundamental challenge in many domains. We present Visage, an interactive visual graph querying approach that empowers users to construct expressive queries, without writing complex code (e.g., finding money laundering rings of bankers and business owners). Our contributions are as follows: (1) we introduce graph autocomplete, an interactive approach that guides users to construct and refine queries, preventing over-specification; (2) Visage guides the construction of graph queries using a data-driven approach, enabling users to specify queries with varying levels of specificity, from concrete and detailed (e.g., query by example), to abstract (e.g., with "wildcard" nodes of any types), to purely structural matching; (3) a twelve-participant, within-subject user study demonstrates Visage's ease of use and the ability to construct graph queries significantly faster than using a conventional query language; (4) Visage works on real graphs with over 468K edges, achieving sub-second response times for common queries.
Robert S. Pienta, Acar Tamersoy, Alex Endert, Shamkant B. Navathe, Hanghang Tong, Polo Chau
AVI2
2015 Understanding variations in pediatric asthma care processes in the emergency department using visual analytics
abstract
Health care delivery processes consist of complex activity sequences spanning organizational, spatial, and temporal boundaries. Care is human-directed so these processes can have wide variations in cost, quality, and outcome making systemic care process analysis, conformance testing, and improvement challenging. We designed and developed an interactive visual analytic process exploration and discovery tool and used it to explore clinical data from 5784 pediatric asthma emergency department patients.
Rahul C. Basole, Mark L. Braunstein, Hyunwoo Park 0003, Minsuk Kahng, Polo Chau, Acar Tamersoy, Daniel A. Hirsh, Nicoleta Serban, James Bost, Burton Lesnick, Beth L. Schissel
J. Am. Medical Informatics Assoc.7
2014 MAGE: Matching approximate patterns in richly-attributed graphs
abstract
Given a large graph with millions of nodes and edges, say a social network where both its nodes and edges have multiple attributes (e.g., job titles, tie strengths), how to quickly find subgraphs of interest (e.g., a ring of businessmen with strong ties)? We present MAGE, a scalable, multicore subgraph matching approach that supports expressive queries over large, richly-attributed graphs. Our major contributions include: (1) MAGE supports graphs with both node and edge attributes (most existing approaches handle either one, but not both); (2) it supports expressive queries, allowing multiple attributes on an edge, wildcards as attribute values (i.e., match any permissible values), and attributes with continuous values; and (3) it is scalable, supporting graphs with several hundred million edges. We demonstrate MAGE's effectiveness and scalability via extensive experiments on large real and synthetic graphs, such as a Google+ social network with 460 million edges.
Robert S. Pienta, Acar Tamersoy, Hanghang Tong, Polo Chau
IEEE BigData2
2014 Guilt by association: large scale malware detection by mining file-relation graphs
abstract
The increasing sophistication of malicious software calls for new defensive techniques that are harder to evade, and are capable of protecting users against novel threats. We present AESOP, a scalable algorithm that identifies malicious executable files by applying Aesop's moral that "a man is known by the company he keeps." We use a large dataset voluntarily contributed by the members of Norton Community Watch, consisting of partial lists of the files that exist on their machines, to identify close relationships between files that often appear together on machines. AESOP leverages locality-sensitive hashing to measure the strength of these inter-file relationships to construct a graph, on which it performs large scale inference by propagating information from the labeled files (as benign or malicious) to the preponderance of unlabeled files. AESOP attained early labeling of 99% of benign files and 79% of malicious files, over a week before they are labeled by the state-of-the-art techniques, with a 0.9961 true positive rate at flagging malware, at 0.0001 false positive rate.
Acar Tamersoy, Kevin A. Roundy, Polo Chau
KDD1
2013 Inside insider trading: patterns & discoveries from a large scale exploratory analysis
abstract
How do company insiders trade? Do their trading behaviors differ based on their roles (e.g., CEO vs. CFO)? Do those behaviors change over time (e.g., impacted by the 2008 market crash)? Can we identify insiders who have similar trading behaviors? And what does that tell us?
Acar Tamersoy, Bo Xie 0002, Stephen L. Lenkey, Bryan R. Routledge, Polo Chau, Shamkant B. Navathe
ASONAM1
2013 Click traffic analysis of short URL spam on Twitter
abstract
With an average of 80% length reduction, the URL shorteners have become the norm for sharing URLs on Twitter, mainly due to the 140-character limit per message. Unfortunately, spammers have also adopted the URL shorteners to camouflage and improve the user click-through of their spam URLs. In this p
De Wang, Shamkant B. Navathe, Ling Liu 0001, Danesh Irani, Acar Tamersoy, Calton Pu
CollaborateCom5
2012 Anonymization of Longitudinal Electronic Medical Records
abstract
Electronic medical record (EMR) systems have enabled healthcare providers to collect detailed patient information from the primary care domain. At the same time, longitudinal data from EMRs are increasingly combined with biorepositories to generate personalized clinical decision support protocols. Emerging policies encourage investigators to disseminate such data in a deidentified form for reuse and collaboration, but organizations are hesitant to do so because they fear such actions will jeopardize patient privacy. In particular, there are concerns that residual demographic and clinical features could be exploited for reidentification purposes. Various approaches have been developed to anonymize clinical data, but they neglect temporal information and are, thus, insufficient for emerging biomedical research paradigms. This paper proposes a novel approach to share patient-specific longitudinal data that offers robust privacy guarantees, while preserving data utility for many biomedical investigations. Our approach aggregates temporal and diagnostic information using heuristics inspired from sequence alignment and clustering methods. We demonstrate that the proposed approach can generate anonymized data that permit effective biomedical analysis using several patient cohorts derived from the EMR system of the Vanderbilt University Medical Center.
Acar Tamersoy, Grigorios Loukides, Mehmet Ercan Nergiz, Yücel Saygin, Bradley A. Malin
IEEE Trans. Inf. Technol. Biomed.1
2011 Instant anonymization
abstract
Anonymization-based privacy protection ensures that data cannot be traced back to individuals. Researchers working in this area have proposed a wide variety of anonymization algorithms, many of which require a considerable number of database accesses. This is a problem of efficiency, especially when the released data is subject to visualization or when the algorithm needs to be run many times to get an acceptable ratio of privacy/utility. In this paper, we present two instant anonymization algorithms for the privacy metricsk-anonymity and ℓ-diversity. Proposed algorithms minimize the number of data accesses by utilizing the summary structure already maintained by the database management system for query selectivity. Experiments on real data sets show that in most cases our algorithm produces an optimal anonymization, and it requires a single scan of data as opposed to hundreds of scans required by the state-of-the-art algorithms.
Mehmet Ercan Nergiz, Acar Tamersoy, Yücel Saygin
ACM Trans. Database Syst.2